A mobile underwater wireless sensor network opportunity routing candidate set screening method based on fuzzy logic and Q-learning combined decision

CN122621969BActive Publication Date: 2026-09-25HAINAN SHUIZHISHENG MARINE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611066098.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-25
Estimated Expiration
2046-07-17

AI Technical Summary

Technical Problem

然而,这类方法本质上仍属于基于当前局部状态的瞬时评价体系,存在如下明显缺陷:其一,难以反映当前选择对后续网络拓扑演化、链路持续性以及长期能量分布的影响,导致长期运行中反复选择少数局部评分较高的节点,造成节点负载集中、剩余能量快速下降,进而诱发网络空洞和路径不稳定;其二,多数方法未充分考虑候选节点集内部节点之间的互相监听关系,容易造成重复转发和额外冲突,削弱机会路由的协同转发效果

Benefits of technology

1、本发明通过模糊逻辑对邻居节点的剩余能量、链路质量及移动性进行多指标综合评价,避免了单一指标决策对复杂水下环境适应性不足的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122621969B_ABST
    Figure CN122621969B_ABST
Patent Text Reader

Abstract

The application discloses a mobile underwater wireless sensor network opportunity routing candidate set screening method based on fuzzy logic and Q-learning combined decision, and relates to the technical field of underwater wireless sensor network communication. The method comprises the following steps: acquiring multi-dimensional state indexes of neighbor nodes and inputting the fuzzy logic system to obtain node applicability values representing comprehensive forwarding applicability; constructing a reward function of Q-learning according to the node applicability values; constructing feasible candidate subsets satisfying mutual monitoring constraints and calculating reward values of each subset by using the reward function; and iteratively updating Q values of the feasible candidate subsets by using a Q-learning algorithm to determine an optimal candidate node set. The application quantifies real-time local applicability of nodes by using fuzzy logic, further introduces long-term reward optimization of candidate subsets satisfying mutual monitoring constraints by using Q-learning, overcomes short-sightedness of only relying on instantaneous local evaluation, balances energy consumption on the premise of guaranteeing cooperative forwarding, and significantly improves packet delivery success rate and routing stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of underwater wireless sensor network communication technology, and in particular to a method for selecting a candidate set of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making using fuzzy logic and Q-learning. Background Technology

[0002] In mobile underwater wireless sensor networks, nodes are highly mobile, causing network topology and link states to change dynamically over time, making data packet forwarding prone to instability. Opportunistic routing, by introducing a candidate forwarding node set mechanism, allows multiple neighboring nodes to participate in packet forwarding contention, effectively improving the packet transmission success rate in dynamic environments. The quality of the candidate node set construction directly impacts network performance.

[0003] Existing candidate node selection methods typically evaluate nodes based on single or combined metrics such as node depth, link quality, and remaining energy. Some methods introduce fuzzy logic to integrate multiple metrics, handle network state uncertainties, and obtain a comprehensive evaluation value. However, these methods are essentially still instantaneous evaluation systems based on the current local state, and have the following significant drawbacks: First, they fail to reflect the impact of the current selection on subsequent network topology evolution, link sustainability, and long-term energy distribution, leading to repeated selection of a few nodes with high local scores during long-term operation. This results in concentrated node load, rapid decline in remaining energy, and consequently, network holes and path instability. Second, most methods do not fully consider the mutual listening relationships between nodes within the candidate node set, which can easily lead to duplicate forwarding and additional conflicts, weakening the cooperative forwarding effect of opportunistic routing.

[0004] Therefore, a new candidate node set selection method is urgently needed to overcome the shortcomings of existing technologies, which rely solely on local instantaneous evaluation, lack long-term adaptive optimization, and ignore the mutual monitoring constraints of candidate nodes. Summary of the Invention

[0005] The purpose of this invention is to provide a method for selecting the candidate set of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, so as to solve the problems mentioned in the background art.

[0006] This invention is achieved through the following technical solution: a method for selecting an opportunistic route candidate set for mobile underwater wireless sensor networks based on joint decision-making using fuzzy logic and Q-learning, the method comprising the following steps: Step 1: Obtain the multi-dimensional state indicators of each neighboring node of the current node. The multi-dimensional state indicators include at least the remaining energy factor, link quality factor, and mobility factor. The remaining energy factor is calculated as follows:

[0007] in, and Each is the current node and next-hop neighbor node initial energy, and These are the nodes at the current forwarding time. and nodes The remaining energy; Step 2: Input the multi-dimensional state indicators of each neighbor node into a preset fuzzy logic system, perform fuzzification, fuzzy reasoning and defuzzification processing, and output a node applicability value that represents the comprehensive forwarding applicability of each neighbor node in the current local network state. Step 3: Select nodes that satisfy the communication distance constraint and the signal decodability constraint from the neighboring nodes to form a pre-candidate node set. Based on the mutual listening constraint that any two nodes in the feasible candidate subset satisfy the mutual reachability condition, divide the pre-candidate node set into one or more feasible candidate subsets. Use the node applicability value as input to construct the reward function of the Q-learning algorithm, and use the reward function to calculate the reward value of each feasible candidate subset. Step 4: Use the Q-learning algorithm to iteratively update the Q-values ​​of each feasible candidate subset, and determine the feasible candidate subset with the best updated Q-value as the candidate node set for opportunistic routing at the current time.

[0008] Furthermore, the Link Quality Factor (LQF) mentioned in step 1 is defined based on the successful transmission probability of data packets, and its expression is:

[0009] in, Indicates the bit length of the data packet. Indicates distance and frequency Bit error rate under given conditions.

[0010] Furthermore, the mobility factor MF mentioned in step 1 is composed of the proportion of connectivity time within the prediction time window multiplied by a penalty factor that decreases with increasing relative velocity, and the expression is:

[0011] in, It is the observation time window. , It refers to the communication range of the node. For the speed of sound at the node, It is a scale constant. A larger value indicates a higher time connectivity ratio, while a smaller value imposes more penalties on high-speed relative motion. For maximum speed, express and The relative velocity vector, Indicates its modulus length, It is the connectivity time within the predicted time window.

[0012] Furthermore, in step 2, the fuzzy logic system uses a trapezoidal membership function to fuzzify the remaining energy factor and the movement factor, and uses a triangular membership function to defuzzify the output variable, i.e., the node applicability, in order to reduce computational costs and form a clear single-peak score.

[0013] Furthermore, the fuzzy logic system described in step 2 contains 27 fuzzy IF-THEN rules. These rules are based on different combinations of linguistic values ​​of input variables and mapped to linguistic values ​​of output variables. Defuzzification is then performed using the centroid method to obtain accurate node applicability values.

[0014] Furthermore, the reward function for constructing the Q-learning algorithm described in step 3 is specifically as follows:

[0015] in, For neighboring nodes The node suitability value; To further enhance rewards, when neighboring nodes... The depth is less than the current node At a depth, For positive rewards, when neighboring nodes The depth is greater than the current node At a depth, As a penalty, when the two depths are equal, Zero, Indicates the current node to neighboring nodes The reward function.

[0016] Furthermore, step 3, which involves dividing the pre-candidate node set into one or more feasible candidate subsets, specifically includes: The selection criteria for the pre-candidate node set are to select nodes that simultaneously satisfy the communication distance constraint and the signal decodability constraint from the current node's one-hop neighbor set. The process of dividing the pre-candidate node set into one or more feasible candidate subsets is to exhaustively search all subsets of the pre-candidate node set and determine the subsets in which any two nodes satisfy the condition of mutual reachability as the feasible candidate subsets.

[0017] Furthermore, in step 3, the reward value of each feasible candidate subset is calculated using the reward function. This is achieved by aggregating the single-node rewards of all candidate nodes within the subset. The calculation formula is as follows:

[0018] in, For feasible candidate subsets The utility value, For the current node To subset Middle node The reward value.

[0019] Furthermore, in step 4, the Q-learning algorithm is used to iteratively update the Q-values ​​of each feasible candidate subset, and the update formula is as follows:

[0020] in, For learning rate, As a discount factor, To take action in the present moment Select feasible candidate subsets The reward value obtained, For the next state The maximum Q value corresponding to all actions. Indicates the state before the update. Take action below Q value, Indicates the updated status Take action below The Q value.

[0021] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: 1. This invention uses fuzzy logic to comprehensively evaluate the remaining energy, link quality, and mobility of neighboring nodes using multiple indicators, thus avoiding the problem of insufficient adaptability of single-indicator decision-making to complex underwater environments.

[0022] 2. This invention introduces the long-term optimization decision-making capability of Q-learning into the candidate set screening process. It takes feasible candidate subsets as decision objects, aggregates the overall utility of subsets through reward functions and iteratively updates Q-values, so that the selection can take into account both instantaneous link status and long-term network operation benefits, effectively overcoming the short-sightedness of existing technologies.

[0023] 3. This invention strictly follows the collaborative forwarding requirements of opportunistic routing. During the screening process, it explicitly introduces mutual listening constraints between candidate nodes, thereby avoiding duplicate forwarding and data conflicts caused by the inability of candidate nodes to listen to each other, and significantly improving the efficiency of collaborative forwarding and the success rate of packet delivery.

[0024] 4. The method of the present invention is significantly better than existing methods based on randomness, deep greediness or only using fuzzy logic in terms of cumulative data packet delivery success rate (PDR), and has higher stability and reliability in long-term operation. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only preferred embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating a method for selecting a candidate set of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making using fuzzy logic and Q-learning, as provided by the present invention.

[0027] Figure 2 This is a schematic diagram of the membership function curves of each input and output variable in the fuzzy logic system of this invention.

[0028] Figure 3 This diagram illustrates a comparison of simulation results between the method of this invention and different comparison protocols in terms of cumulative packet delivery success rate (PDR). Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the present invention more apparent, exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely a part of the embodiments of the present invention, and not all of the embodiments of the present invention. It should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described herein, all other embodiments obtained by those skilled in the art without inventive effort should fall within the protection scope of the present invention.

[0030] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without one or more of these details. In other instances, certain technical features well-known in the art have not been described in order to avoid obscuring the invention.

[0031] It should be understood that the invention can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0032] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. When used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising” and / or “including,” when used in this specification, identify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups. When used herein, the term “and / or” includes any and all combinations of the associated listed items.

[0033] To fully understand this invention, a detailed structure will be presented in the following description to illustrate the technical solution proposed by this invention. Optional embodiments of the invention are described in detail below; however, in addition to these detailed descriptions, the invention may have other embodiments.

[0034] This embodiment applies the candidate node set selection method based on joint decision-making of fuzzy logic and Q-learning proposed in this invention to a typical opportunistic routing scenario of underwater wireless sensor networks.

[0035] The simulation experiment was conducted within a two-dimensional rectangular underwater monitoring area (1000m × 1000m). Fifty sensor nodes with identical initial energy were randomly deployed in the network. The nodes used a random walk model with a movement speed of 0.8m / s to simulate the maneuverability of the underwater nodes. The communication range of all nodes was [not specified]. All distances are set to 300m. The destination node (i.e., the sink node) is fixed at coordinates (500m, 1000m) and is responsible for collecting data from the network. The network layer adopts a contention-based opportunistic routing protocol, where the selection of candidate node sets is accomplished by the method of this invention. The relevant physical layer parameters are set as follows: data rate is 350bps, data packet size... The maximum signal-to-noise ratio (SNR) is 1dB, with a control packet size of 400 bits and a target SNR of 80 bits. The maximum forwarding hop count for a single data packet is limited to 18 hops to prevent unrestricted packet propagation within the network. (See Table 1.)

[0036] Table 1 Simulation Parameter Settings

[0037] Reference Figure 1The candidate node set selection method based on joint decision-making of fuzzy logic and Q-learning proposed in this embodiment includes the following steps: Step 1: Obtain the multi-dimensional state indicators of each neighboring node of the current node. The multi-dimensional state indicators include at least the remaining energy factor, link quality factor, and mobility factor. The remaining energy factor is calculated as follows:

[0038] in, and Each is the current node and next-hop neighbor node initial energy, and These are the nodes at the current forwarding time. and nodes The remaining energy; Step 2: Input the multi-dimensional state indicators of each neighbor node into a preset fuzzy logic system, perform fuzzification, fuzzy reasoning and defuzzification processing, and output a node applicability value that represents the comprehensive forwarding applicability of each neighbor node in the current local network state. Step 3: Select nodes that satisfy the communication distance constraint and the signal decodability constraint from the neighboring nodes to form a pre-candidate node set. Based on the mutual listening constraint that any two nodes in the feasible candidate subset satisfy the mutual reachability condition, divide the pre-candidate node set into one or more feasible candidate subsets. Use the node applicability value as input to construct the reward function of the Q-learning algorithm, and use the reward function to calculate the reward value of each feasible candidate subset. Step 4: Use the Q-learning algorithm to iteratively update the Q-values ​​of each feasible candidate subset, and determine the feasible candidate subset with the best updated Q-value as the candidate node set for opportunistic routing at the current time.

[0039] For the calculation of multidimensional state indicators of neighboring nodes, at each forwarding decision time, the current node... It is necessary to evaluate its one-hop neighbor set. The overall status. This embodiment specifically collects and calculates the following three indicators: Remaining Energy Factor (REF), Link Quality Factor (LQF), and Mobility Factor (MF).

[0040] The Residual Energy Factor (REF) is used to characterize the degree of energy balance among candidate forwarding node pairs, and its calculation formula is as follows: (1) in, and Each is the current node and next-hop neighbor node The initial energy is set to the same value in this embodiment; and These are the current forwarding time nodes. and nodes The remaining energy. By using the harmonic mean, the energy at both ends is equally weighted. For example, if one end has high energy and the other end has low energy, the factor value will be significantly reduced, thus giving lower scores to links with unbalanced energy in subsequent fuzzy evaluations and avoiding excessive consumption of low-energy nodes.

[0041] For the calculation of the Link Quality Factor (LQF), LQF directly reflects the probability of a data packet being successfully received under the current communication distance and underwater acoustic channel conditions, and is defined as: (2) in, Indicates the bit length of the data packet. Indicates distance and frequency The bit error rate under the given conditions can be obtained based on the specific underwater acoustic channel model (such as Rayleigh fading or Rice fading model). In this embodiment, the distance... The farther away, The higher the value, the lower the LQF, indicating that long-distance links will be penalized in fuzzy evaluation.

[0042] For the calculation of the mobility factor (MF), considering the presence of maneuvering nodes in underwater scenarios and the time-varying characteristics of the network topology, the pre-searched packet forwarding paths are prone to failure. Therefore, the probability that a node is within the communication radius at the current moment is introduced to characterize its mobility, and the mobility factor MF is designed accordingly.

[0043] (3) in, It is the observation time window. . It refers to the communication range of the node. The velocity of sound at the node is taken as 1500 m / s here. It is a scale constant. A larger value indicates a higher time connectivity ratio, while a smaller value results in more penalties for high-speed relative motion. This is the maximum speed. express and The relative velocity vector, This indicates its modulus length. This refers to the connectivity time within the predicted time window, and its derivation is as follows. At time... When the distance between the two nodes is : (4) in, For two nodes The relative displacement vector at time t. This represents the Euclidean norm of a vector.

[0044] To keep a node within communication range, the following conditions must be met during the observation period: (5) Substituting (4) into (5) and squaring both sides, we get: (6) Expanding further: (7) make , , Here, a, b, and c are the coefficients of the quadratic term, the linear term, and the constant term, respectively. This represents the transpose of a vector. The discriminant is defined as: (8) When a = 0, it means the relative velocity between the two nodes is zero, that is, their relative positions do not change with time. At this time: (1) If c ≤ 0, then the two nodes always satisfy the communication constraint throughout the entire prediction window, therefore .

[0045] (2) If c>0, then the two nodes will never satisfy the communication constraint, therefore .

[0046] When a > 0, the connected interval is determined by the quadratic inequality: (1) If Δ < 0, then there is no moment within the prediction time window that satisfies the communication constraints. .

[0047] (2) If Δ ≥ 0, then the equation The two real roots are: Because the parabola opens upwards, therefore, within the prediction window... The actual connectivity duration within an interval is the length of the intersection of the two intervals, that is:

[0048] The node suitability comprehensive evaluation based on fuzzy logic, after obtaining the residual energy factor (REF), link quality factor (LQF), and mobility factor (MF) indicators, introduces a fuzzy logic system to flexibly handle the interaction and uncertainty between the indicators, and outputs a comprehensive "node suitability (FP)".

[0049] For fuzzification, this embodiment defines the linguistic values ​​of each input and output variable, as shown in Table 2.

[0050] Table 2 Language values ​​of each impact factor

[0051] To cover the potentially large variations in input values ​​with low computational overhead, trapezoidal membership functions are used for RET, MF, and LQF; to ensure the output evaluation results exhibit a clear unimodal shape for decision-making, a triangular membership function is used for FP. The corresponding membership function curves are shown below. Figure 2 As shown. For example, at a certain moment, a neighboring node calculates a REF of 0.75. According to the preset membership function, it may belong to the language value "many" with a membership degree of 0.8 and to "medium" with a membership degree of 0.2.

[0052] For fuzzy inference, the IF-THEN rule base is used. This rule base consists of 27 rules, which comprehensively cover the mapping relationship between different combinations of linguistic values ​​(more / medium / few, large / medium / small, good / medium / bad) of the three input variables and the output FP (good / slightly good / medium / slightly bad / bad). The rules are shown in Table 3.

[0053] Table 3 Mapping rules for each impact factor

[0054] This is a multiple-input single-output (MISO) system, where the activation of each rule is determined by the minimum value of the membership of each input under that rule.

[0055] For deblurring, to synthesize a precise FP value from multiple activated rules, this embodiment employs the Centroid Method for deblurring. The calculation formula is as follows:

[0056] in, The number of rules that are activated. No. The activation level of the rule, For the first The rule outputs the geometric center of the fuzzy set. Through this step, the system calculates the geometric center for each neighbor node. Output an accurate applicability score in the [0,1] interval. .

[0057] Based on Q-learning, the long-term utility decision of candidate node subsets is based on the fact that fuzzy logic evaluation is essentially an instantaneous scoring of the current state. To compensate for its short-sightedness and enable the current selection decision to take into account its long-term impact on the future network state (such as energy consumption distribution and link changes caused by node movement), this invention models the selection of candidate node sets as a Markov decision process and introduces Q-learning for solving it.

[0058] Regarding the definition of state and action, in this embodiment, the current node is defined. The local network view is defined as the state, which may include information such as the energy level distribution and link quality level distribution of its neighboring nodes after discretization and abstraction. The action is defined as selecting a specific subset from the set of feasible candidate subsets that satisfy the mutual listening constraint.

[0059] To construct a feasible subset of candidates that satisfies the mutual listening constraint, opportunistic routing requires that candidate nodes be able to listen to each other to achieve priority coordination and forwarding suppression. This embodiment explicitly models this as a "mutual listening constraint." The construction process is as follows: From the neighborhood set Nodes that meet two basic conditions are selected from the pool to form a pre-candidate set. : Communication distance constraints: distance

[0060] Signal decodeability constraint: Received signal-to-noise ratio ( (For successful group decoding threshold).

[0061] Solve The power set, and for each subset Perform constraint checks: if subset Any two nodes and distance Less than or equal to If a subset is found to be a valid candidate subset that satisfies the mutual listening constraint, then that subset is a feasible candidate subset. All such subsets... This constitutes the action space in the current state. .

[0062] For the reward function, in order to guide Q-learning to learn the correct long-term policy, the action... Selected subset Each member node in Its single-node reward consists of real-time forwarding quality and task-oriented deep advancement rewards:

[0063] Among them, the reward for in-depth advancement It is the key to driving the flow of data towards the Sink node:

[0064] Then, define the action. (Selecting a subset) The overall reward obtained Aggregate the rewards of all single nodes in this subset:

[0065] For iterative Q-value updates and optimal candidate set decisions, a Q-value is maintained for each state-action pair. In the initial phase, all Q-values ​​are initialized to 0. After each hop, upon making a forwarding decision and receiving feedback from the environment, the Q-values ​​are updated using a time-difference method:

[0066] in, The learning rate (set to 0.8) controls the degree to which new information overwrites old information. The discount factor (set to 0.9) represents the importance of future rewards relative to immediate rewards. That is, the action to be performed this time Rewards After multiple rounds of learning and Q-value updates, the system's decisions will gradually converge. In actual operation, the current node... Using the updated Q-table, given the current state, employ a greedy strategy (or - Greedy strategy (exploration) selects the action with the maximum Q value. The feasible candidate subset corresponding to this action This refers to the set of nodes identified as final opportunity route candidate forwarding nodes. Broadcast the data packet to ,Depend on Internally, the final coordination and forwarding is completed according to the preset priority rules (such as FP descending order).

[0067] Simulation results and performance analysis: To verify the effectiveness of the method of the present invention (denoted as FL-Q-Learning), it is compared with the following three methods in the same network environment: The fuzzy logic-only method (FL-Only) calculates FP using only steps 1-2 of this invention and directly selects the nodes with the highest FP that are listening to each other to form a candidate set, without Q-learning for long-term optimization.

[0068] Random method: Randomly select from candidate nodes that satisfy basic communication and reachability constraints.

[0069] Depth-Greedy method: Prioritizes selecting the shallowest neighbor node as the next hop.

[0070] Simulation results are as follows Figure 3 As shown, the horizontal axis represents the number of simulation rounds, and the vertical axis represents the cumulative group delivery success rate (Cumulative PDR).

[0071] from Figure 3 As can be seen, the FL-Q-Learning method of this invention exhibits the fastest convergence speed and the highest final delivery success rate. Its cumulative PDR reaches approximately 0.96 by around round 20 and remains stable at a high level of approximately 0.90 even after 120 rounds of long-term operation. In comparison, the cumulative PDRs of the FL-Only, Random, and Depth-Greedy methods at the end of the simulation are approximately 0.87, 0.78, and 0.67, respectively.

[0072] This comparative result profoundly reveals the combined benefits of the various mechanisms in this invention: FL-Only outperforms Random, proving that fuzzy logic's multi-index comprehensive evaluation is better at selecting high-quality current links than blind random selection.

[0073] FL-Q-Learning achieves a final performance improvement of approximately 0.03 compared to FL-Only, with superior stability. This precisely demonstrates the value of introducing Q-learning for long-term decision optimization. Through learning, it avoids repeatedly exploiting a single node that currently appears optimal, thereby balancing network energy consumption and preventing routing gaps caused by premature energy depletion or removal of critical nodes.

[0074] The Depth-Greedy method performed the worst (0.67), indicating that in highly mobile underwater environments, a "greedy" strategy that simply pursues depth is short-sighted and unreliable. It ignores energy and link quality, easily leading to unreachable or rapidly failing shallow nodes. In this invention, depth is only used as a reward function. As part of it, it must be weighed together with link quality, mobility, and energy state represented by FP, which is the fundamental reason for the robustness of this invention.

[0075] This embodiment illustrates in detail how the present invention, through the organic combination of fuzzy logic and Q-learning, while strictly adhering to the opportunistic routing mutual eavesdropping constraint, takes into account both local link quality and long-term network benefits, ultimately significantly improving the routing performance in mobile underwater wireless sensor networks.

[0076] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for selecting candidate sets of opportunistic routes in mobile underwater wireless sensor networks based on joint decision-making using fuzzy logic and Q-learning, characterized in that, The method includes the following steps: Step 1: Obtain the multi-dimensional state indicators of each neighboring node of the current node. The multi-dimensional state indicators include at least the remaining energy factor, link quality factor, and mobility factor. The remaining energy factor is calculated as follows: in, and Each is the current node and next-hop neighbor node initial energy, and These are the nodes at the current forwarding time. and nodes The remaining energy; Step 2: Input the multi-dimensional state indicators of each neighbor node into a preset fuzzy logic system, perform fuzzification, fuzzy reasoning and defuzzification processing, and output a node applicability value that represents the comprehensive forwarding applicability of each neighbor node in the current local network state. Step 3: Select nodes that satisfy the communication distance constraint and the signal decodability constraint from the neighboring nodes to form a pre-candidate node set. Based on the mutual listening constraint that any two nodes in the feasible candidate subset satisfy the mutual reachability condition, divide the pre-candidate node set into one or more feasible candidate subsets. Use the node applicability value as input to construct the reward function of the Q-learning algorithm, and use the reward function to calculate the reward value of each feasible candidate subset. Step 4: Use the Q-learning algorithm to iteratively update the Q-values ​​of each feasible candidate subset, and determine the feasible candidate subset with the best updated Q-value as the candidate node set for opportunistic routing at the current time.

2. The method for selecting candidate sets of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, as described in claim 1, is characterized in that... The Link Quality Factor (LQF) mentioned in step 1 is defined based on the successful transmission probability of data packets, and its expression is: in, Indicates the bit length of the data packet. Indicates distance and frequency Bit error rate under given conditions.

3. The method for selecting candidate sets of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, as described in claim 1, is characterized in that... The mobility factor MF mentioned in step 1 is composed of the proportion of connectivity time within the prediction time window multiplied by a penalty factor that decreases with increasing relative velocity, and the expression is: in, It is the observation time window. , It refers to the communication range of the node. For the speed of sound at the node, It is a scale constant. A larger value indicates a higher time connectivity ratio, while a smaller value imposes more penalties on high-speed relative motion. For maximum speed, express and The relative velocity vector, Indicates its modulus length, It is the connectivity time within the predicted time window.

4. The method for selecting candidate sets of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, as described in claim 1, is characterized in that... In step 2, the fuzzy logic system uses a trapezoidal membership function to fuzzify the remaining energy factor and the movement factor, and uses a triangular membership function to defuzzify the output variable, i.e., the node applicability, in order to reduce computational costs and form a clear unimodal score.

5. The method for selecting candidate sets of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, as described in claim 1, is characterized in that... The fuzzy logic system described in step 2 contains 27 fuzzy IF-THEN rules. These rules are based on different combinations of linguistic values ​​of input variables and mapped to linguistic values ​​of output variables. Defuzzification is then performed using the centroid method to obtain accurate node applicability values.

6. The method for selecting candidate sets of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, as described in claim 1, is characterized in that... The reward function for constructing the Q-learning algorithm described in step 3 is as follows: in, Neighboring nodes The node suitability value; To further enhance rewards, when neighboring nodes... The depth is less than the current node At a depth, For positive rewards, when neighboring nodes The depth is greater than the current node At a depth, As a penalty, when the two depths are equal, Zero, Indicates the current node to neighboring nodes The reward function.

7. The method for selecting candidate sets of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, as described in claim 1, is characterized in that... Step 3, which involves dividing the pre-candidate node set into one or more feasible candidate subsets, specifically includes: The selection criteria for the pre-candidate node set are to select nodes that simultaneously satisfy the communication distance constraint and the signal decodability constraint from the current node's one-hop neighbor set. The process of dividing the pre-candidate node set into one or more feasible candidate subsets is to exhaustively search all subsets of the pre-candidate node set and determine the subsets in which any two nodes satisfy the condition of mutual reachability as the feasible candidate subsets.

8. The method for selecting candidate sets of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, as described in claim 1, is characterized in that... Step 3 involves calculating the reward value for each feasible candidate subset using the reward function. This is achieved by aggregating the single-node rewards of all candidate nodes within the subset. The calculation formula is as follows: in, For feasible candidate subsets The utility value, For the current node To subset Middle node The reward value.

9. The method for selecting candidate sets of opportunistic routes for mobile underwater wireless sensor networks based on joint decision-making of fuzzy logic and Q-learning, as described in claim 1, is characterized in that... In step 4, the Q-learning algorithm is used to iteratively update the Q-values ​​of each feasible candidate subset. The update formula is as follows: in, For learning rate, As a discount factor, To take action in the present moment Select feasible candidate subsets The reward value obtained, For the next state The maximum Q value corresponding to all actions. Indicates the state before the update. Take action below Q value, Indicates the updated status Take action below The Q value.

Citation Information

Patent Citations

  • Air-sea cross-domain network routing protocol method based on fuzzy logic and Q learning optimization

    CN122093889A

  • Intelligent routing and adaptive communication method suitable for mobile underwater sensor network

    CN122205555A