Ocean communication method based on self-adaptive Q-learning factor setting

Through the method of adaptive Q-learning factor setting and dual-Q network asynchronous update, the environmental disconnection problem caused by static parameters in the marine communication network is solved, efficient routing decision-making and resource optimization are achieved, the delivery rate and communication stability are improved, and it adapts to the dynamic changes of the marine environment.

CN120639683APending Publication Date: 2025-09-12SHANGHAI MARITIME UNIVERSITY

Patent Information

Application Number
CN202510921786.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In existing marine communication networks, the traditional Q-Learning algorithm's static parameter design causes the node routing algorithm to be out of touch with the dynamic environment, and is unable to cope with real-time changes in topology, load, and channel quality, resulting in resource waste, uncontrolled congestion, and insufficient delivery efficiency and reliability.

Method used

An adaptive Q-learning factor setting method is adopted. Nodes periodically collect status information, dynamically update the learning rate and discount factor, and combine the dual-Q network asynchronous update to achieve real-time routing decisions and parameter adjustments, optimize the Q value table, select the forwarding node that maximizes the reward value, reduce the learning rate when the link fails, and start the detection protocol and satellite communication backup.

Benefits of technology

It improves the delivery rate of the marine communication network, can better adapt to the changing environment, reduce message transmission delay and resource consumption, ensure communication continuity and stability, adapt to the high-speed random mobility characteristics of marine nodes, and improve the overall throughput and resource utilization efficiency of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639683A_ABST
    Figure CN120639683A_ABST
Patent Text Reader

Abstract

The invention relates to an ocean communication method based on self-adaptive Q-learning factor setting. In the method, each node periodically collects own state information and broadcasts the own state information to surrounding nodes, and after each node receives a message and stores the message in a cache region, routing decision and parameter adjustment processes are circularly executed until the message is completely transmitted to a target node to complete communication; wherein the routing decision and parameter adjustment process comprises the following steps of: calling a message from a corresponding buffer area by a current node, selecting a next forwarding node of the message based on a current Q value table, updating a learning rate and a discount factor based on the current state information of the node after sending the message, and performing routing decision and parameter adjustment by combining a reward function, the updated learning rate and the discount factor. And a new Q value table is obtained through asynchronous updating of the double-Q network. Compared with the prior art, the method has the advantages that the delivery rate is higher, the method can better adapt to changeable environments, and the problem of real-time changes of topology, load and channel quality can be well solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to an ocean communication method based on adaptive Q-learning factor setting. Background Art

[0002] The present invention belongs to the field of marine communication network technology, specifically to an opportunistic routing and forwarding algorithm based on adaptive Q-learning factor settings. Marine communication networks are dynamic, heterogeneous networks composed of nodes such as ships, buoys, drones, and offshore platforms. These nodes are widely distributed, highly mobile, and significantly affected by natural environments such as ocean currents and weather. This results in frequent changes in network topology, unstable end-to-end paths, and highly limited resources (such as storage and bandwidth). Based on this, some researchers have proposed routing algorithms based on Q-Learning (a reinforcement learning method) to address these issues. However, traditional Q-Learning algorithms also suffer from numerous problems, such as overestimation bias, neglect of multiple factors, and static parameter design. In particular, the use of static parameter design can cause node routing algorithms to become disconnected from the dynamic environment and unable to cope with real-time changes in topology, load, and channel quality. It can also lead to resource waste and uncontrolled congestion, as redundant message replication and inefficient cache cleanup mechanisms can rapidly exhaust storage resources. Furthermore, it can result in insufficient delivery efficiency and reliability, lacking a dynamic trade-off between message timeliness and node reliability, and making critical data susceptible to expiration or loss.

[0003] For example, the invention patent with publication number CN111510956A discloses a hybrid routing method and marine communication system based on clustering and reinforcement learning. This hybrid routing method based on clustering and reinforcement learning improves the convergence speed and accuracy of Q learning, while limiting the impact of network topology changes to a local range, making the overall network load more balanced and the overall network performance better. However, it still uses a static learning rate, and the discount factor is simply assigned a piecewise function, which means that the node routing algorithm is still out of touch with the dynamic environment, unable to cope with real-time changes in topology, load, and channel quality, and has poor environmental adaptability. Summary of the Invention

[0004] The purpose of the present invention is to provide a marine communication method based on adaptive Q-learning factor setting in order to overcome the defects of the above-mentioned prior art.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] According to one aspect of the present invention, a method for ocean communication based on adaptive Q-learning factor setting is provided. In this method, each node periodically collects its own status information and broadcasts it to surrounding nodes; the surrounding nodes specifically refer to other nodes within a preset range centered on the node. After receiving a message, each node stores the message in a buffer area. The routing decision and parameter adjustment process are repeated until all messages are transmitted to the target node, completing the communication process.

[0007] The routing decision and parameter adjustment process includes:

[0008] S1. The current node retrieves the message from the corresponding buffer and selects the next forwarding node for the message based on the current Q value table;

[0009] S2. Check whether the next forwarding node is connected to the current node. If so, send the message to the next forwarding node and continue to execute S3. Otherwise, jump to execute S1.

[0010] S3, based on the node's current state information, update the learning rate and discount factor;

[0011] S4. Based on the reward function, the updated learning rate and the discount factor, the dual Q network is asynchronously updated to obtain the updated Q value table.

[0012] As a preferred technical solution, the next forwarding node of the message in S1 is selected as follows: based on the current Q value table, the action a that maximizes the reward value is selected. * , the specific formula is:

[0013]

[0014] Among them, a * The next forwarding node for the message; is the action space, i.e., the set of neighbor nodes; is the reward function; β is the discount factor; Q(s′,a′) is the Q value under the next state s′ action a′.

[0015] As a preferred technical solution, if it is detected in S2 that the next forwarding node is not connected to the current node, the learning rate is reduced and the detection protocol is started. If it is not restored within the preset period, the communication is switched to satellite communication mode.

[0016] As a preferred technical solution, in S3, when updating the learning rate, its own state information includes the network scale, the current node's cumulative message sending volume, and the network congestion. The specific formula for updating the learning rate is:

[0017]

[0018] Among them, α(t) is the learning rate, α0 is the learning rate baseline value; N(S) is the total number of network nodes, reflecting the network scale; N(m) is the cumulative number of messages sent by the current node; C(t) is the network congestion, which is the ratio of the number of cached messages to the maximum capacity.

[0019] As a preferred technical solution, in S3, when updating the discount factor, the state information includes the number of hops the message has been forwarded and the node movement speed. The remaining life cycle information and the number of hops the message has been forwarded are also used in the discount factor update process. The specific formula is:

[0020]

[0021] Where β(t) is the discount factor; For the message The time it takes to transmit to the destination node; T is the remaining life cycle information of the message; current is the current network average delivery delay; Hops is the number of hops that the message has been forwarded; avg is the average value of hop count; β0 is the benchmark value of discount factor.

[0022] As a preferred technical solution, the learning rate baseline value and the discount factor baseline value are adjusted as the network scale changes. The specific formula is:

[0023]

[0024] Among them, α0 is the learning rate reference value; α init is the initial learning rate benchmark value; N(S) is the total number of network nodes, which is used to reflect the network scale; β0 is the discount factor benchmark value; β init is the initial discount factor benchmark value.

[0025] As a preferred technical solution, the specific formula of the reward function in S4 is:

[0026]

[0027] Among them, R a is the reward function; η t and η h is the dynamic weight; T current is the current network average delivery delay, is the estimated delay after message forwarding. The lower the delay, the higher the reward value. current is the average number of network hops, is the estimated number of hops after the message is forwarded. The fewer the hops, the greater the reward weight. b is the historical delivery success rate of target node b; λ is the reliability weight coefficient.

[0028] As a preferred technical solution, the dynamic weight η t and η h are the average of recent delay and hop count calculated by sliding window respectively. The specific formula is:

[0029]

[0030] Among them, η t and η h is the dynamic weight; W is the window size; is the average network delay; The time it takes for a message to travel from node a to node b; is the average number of hops for message delivery in the network; is the number of hops that a message takes from node a to node b.

[0031] As a preferred technical solution, the specific process of the dual-Q network asynchronous update in S4 is:

[0032]

[0033] Among them, Q1'(s,a) and Q2'(s,a) are the updated Q networks respectively; Q1(s,a) and Q2(s,a) are the Q networks before the update respectively; α(t) is the learning rate; R is the reward function value; β(t) is the discount factor; s and s ′ are the current state and the state after the message is forwarded; a and a ′ They are the current action and the action after the message is forwarded.

[0034] As a preferred technical solution, the buffer periodically clears redundant messages, and coordinates with neighboring nodes to reduce the sending frequency when clearing redundant messages.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] 1. The present invention completes communication by cyclically executing the routing decision and parameter adjustment process until the message is fully transmitted to the target node. During the routing decision and parameter adjustment process, the next forwarding node for the message is selected based on the current Q value table. After the message is sent, the learning rate and discount factor are updated based on the node's current state information. Then, combined with the reward function, the updated learning rate and discount factor, a new Q value table is obtained by asynchronously updating the dual-Q network. This enables the method to perceive changes in the external environment in real time and change its own value as the environment changes, making learning more efficient and achieving better learning results. Compared with methods that do not use dynamic factor adjustment, this marine communication method has a higher delivery rate and can better adapt to changing environments. This marine communication method can effectively deal with real-time changes in topology, load, and channel quality.

[0037] 2. The present invention uses dynamic programming based on the Q-value table to maximize reward value routing, combining immediate rewards with long-term benefits, giving priority to forwarding nodes with "high rewards and low latency", effectively reducing the number of message transmission hops and latency, and improving the overall throughput of the marine communication network.

[0038] 3. In the present invention, the learning rate is dynamically reduced when the link fails to avoid ineffective exploration and resource consumption. At the same time, the detection protocol and satellite communication backup are started to build a self-repair and multi-modal communication mechanism to solve the problem of frequent link interruptions caused by node movement in the marine environment and ensure communication continuity in extreme scenarios.

[0039] 4. In this invention, through adaptive learning rate adjustment, the learning rate is strongly correlated with network scale, transmission volume, and congestion: the larger the network scale, the smoother the learning rate decay, balancing the exploration and utilization of large-scale networks; the learning rate is automatically reduced in congestion, reducing blind forwarding and alleviating buffer pressure; by introducing the remaining message life cycle, the number of forwarded hops, and node speed into the discount factor, the short-term reward weight is increased when the life cycle is urgent, and direct paths are preferred; high-speed mobile nodes reduce long-term profit expectations, avoid relying on unstable links, and adapt to the high-speed random movement characteristics of ocean nodes. The learning rate / discount factor baseline value is decoupled from the network scale through the logarithm / square root function, allowing the algorithm to adapt to cross-scale applications from small formations to large-scale ocean observation networks without manual parameter adjustment, significantly improving the algorithm's versatility.

[0040] 5. In this invention, the reward function integrates latency, hop count, and historical success rate, using dynamic weights to assess link quality in real time. Paths with low latency and few hops receive higher rewards, guiding nodes to prioritize efficient links. This, combined with the target node's historical reliability, avoids forwarding to nodes with frequent packet loss, improving message transmission stability. Furthermore, the dynamic weight is calculated based on the average of recent latency and hop count within a sliding window, allowing it to dynamically adjust with real-time network load. During sudden congestion, the latency weight is automatically reduced, allowing for rapid response to sudden changes in node distribution caused by tides and currents in marine environments, enhancing real-time decision-making.

[0041] 6. In the present invention, by separating action selection and value evaluation, asynchronous updates reduce computational coupling, support distributed node parallel learning, and are suitable for the decentralized architecture of marine networks; routing errors caused by model overfitting are reduced, especially in complex scenarios such as turbulence and storms where node states change frequently.

[0042] 7. In the present invention, the buffer is deduplicated regularly to avoid the accumulation of invalid messages, freeing up storage resources to process urgent data; neighbors are coordinated to reduce the sending frequency, reduce channel competition caused by redundant broadcasts, reduce overall energy consumption, and extend the life of ocean nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 Schematic diagram of the routing decision and parameter adjustment process in the present invention;

[0044] Figure 2 A schematic diagram of an ocean communication process based on adaptive Q-learning factor setting in the present invention;

[0045] Figure 3 This is an analysis of the impact of the buffer size on the delivery rate in the embodiment. DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0047] Example 1

[0048] In this embodiment, a marine communication method based on adaptive Q-learning factor setting is adopted to solve the problems of poor adaptability and low resource utilization of traditional methods in dynamic marine environments. Figure 1 As shown in the figure, in this method, each node periodically collects its own status information and broadcasts it to surrounding nodes; after receiving a message, each node stores the message in a buffer area; and repeats the routing decision and parameter adjustment process until all messages are transmitted to the target node, completing the communication process;

[0049] The routing decision and parameter adjustment process includes:

[0050] S1. The current node retrieves the message from the corresponding buffer and selects the next forwarding node for the message based on the current Q value table;

[0051] S2. Check whether the next forwarding node is connected to the current node. If so, send the message to the next forwarding node and continue to execute S3. Otherwise, jump to execute S1.

[0052] S3, based on the node's current state information, update the learning rate and discount factor;

[0053] S4. Based on the reward function, the updated learning rate and the discount factor, the dual Q network is asynchronously updated to obtain the updated Q value table.

[0054] The next forwarding node for the message in S1 is: Based on the current Q value table, select the action a that maximizes the reward value. * , the specific formula is:

[0055]

[0056] Among them, a * The next forwarding node for the message; is the action space, i.e., the set of neighbor nodes; is the reward function; β is the discount factor; Q(s′,a′) is the Q value under the next state s′ action a′.

[0057] In S2, if it is detected that the next forwarding node is not connected to the current node, the learning rate is reduced and the detection protocol is started. If it is not restored within the preset period, the communication is switched to the satellite communication mode.

[0058] In S3, when updating the learning rate, its own state information includes the network scale, the current node's cumulative message sending volume, and network congestion. The specific formula for updating the learning rate is:

[0059]

[0060] Among them, α(t) is the learning rate, α0 is the learning rate baseline value; N(S) is the total number of network nodes, reflecting the network scale; N(m) is the cumulative number of messages sent by the current node; C(t) is the network congestion, which is the ratio of the number of cached messages to the maximum capacity.

[0061] In S3, when updating the discount factor, its own state information includes the number of hops the message has been forwarded and the node's moving speed. The remaining life cycle information and the number of hops the message has been forwarded are also used in the discount factor update process. The specific formula is:

[0062]

[0063] Where β(t) is the discount factor; For the message The time it takes to transmit to the destination node; T is the remaining life cycle information of the message; current is the current network average delivery delay; Hops is the number of hops that the message has been forwarded; avg is the average value of hop count; β0 is the benchmark value of discount factor.

[0064] The learning rate baseline value and the discount factor baseline value are adjusted as the network scale changes. The specific formula is:

[0065]

[0066] Among them, α0 is the learning rate reference value; α initis the initial learning rate benchmark value; N(S) is the total number of network nodes, which is used to reflect the network scale; β0 is the discount factor benchmark value; β init is the initial discount factor benchmark value.

[0067] The specific formula of the reward function in S4 is:

[0068]

[0069] Among them, R a is the reward function; η t and η h is the dynamic weight; T current is the current network average delivery delay, is the estimated delay after message forwarding. The lower the delay, the higher the reward value. current is the average number of network hops, is the estimated number of hops after the message is forwarded. The fewer the hops, the greater the reward weight. b is the historical delivery success rate of target node b; λ is the reliability weight coefficient.

[0070] Dynamic weight η t and η h are the average of recent delay and hop count calculated by sliding window respectively. The specific formula is:

[0071]

[0072] Among them, η t and η h is the dynamic weight; W is the window size; is the average network delay; The time it takes for a message to travel from node a to node b; is the average number of hops for message delivery in the network; is the number of hops that a message takes from node a to node b.

[0073] The specific process of dual-Q network asynchronous update in S4 is as follows:

[0074]

[0075]

[0076] Among them, Q1'(s,a) and Q2'(s,a) are the updated Q networks respectively; Q1(s,a) and Q2(s,a) are the Q networks before the update respectively; α(t) is the learning rate; R is the reward function value; β(t) is the discount factor; s and s ′ are the current state and the state after the message is forwarded; a and a ′ They are the current action and the action after the message is forwarded.

[0077] The buffer periodically clears redundant messages and coordinates with neighboring nodes to reduce the sending frequency when clearing redundant messages.

[0078] In this method, a dual Q network is used to decouple the evaluation and decision-making mechanism. Two independent Q-value functions Q1(s,a) and Q2(s,a) are used to perform action selection and value evaluation respectively, eliminating the overestimation bias caused by the traditional single network. In the action selection stage, the current optimal action a′=argmax is selected based on the Q1 network. a Q1(s′,a); In the value evaluation stage, the target value Q is calculated using the Q2 network target =R+β·Q2(s′,a′), and by alternately updating the network parameters of Q1 and Q2, the decoupling of the evaluation and decision-making processes is ensured to avoid the overestimation of action values.

[0079] The learning rate α is dynamically adjusted according to the real-time state of the network, and its calculation formula is: Among them, α * is the baseline learning rate (default is 0.8), which is used to control the overall learning speed; N(S) is the total number of network nodes, reflecting the network scale. The larger the scale, the more significant the learning rate decay; N(m) is the cumulative number of messages sent by the current node, which suppresses the excessive influence of historical experience on policy updates through exponential decay; C is the network congestion, defined as the ratio of the number of cached messages to the maximum capacity.

[0080] The discount factor β is designed as a multi-dimensional dynamic parameter that integrates message timeliness, transmission efficiency, and node mobility. Its calculation formula is:

[0081]

[0082] Where β0 is the benchmark discount factor (default 0.9), which is used to control the weight of long-term returns; For the message The time it takes to transmit to the destination node; The remaining lifetime of the message. The shorter the remaining lifetime, the higher the β weight, and the critical messages that are about to expire are forwarded first. is the number of hops that the message has been forwarded. When the number of hops exceeds the network average, the β weight is reduced to suppress redundant transmission. The greater the node movement speed v, the more additional β is improved, strengthening the consideration of long-term path stability.

[0083] Reward Function The dynamic weight mechanism is used to integrate real-time network performance and node reliability, and its expression is:

[0084]

[0085] Among them, Tcurrent is the current network average delivery delay, is the estimated delay after message forwarding. The lower the delay, the higher the reward value. current is the average number of network hops, The estimated number of hops after the message is forwarded. The fewer the hops, the greater the reward weight. Dynamic weight η t and η h The average of recent delay and hop count is calculated using a sliding window (window size W = 10). The formula is:

[0086]

[0087] In addition, R b is the historical delivery success rate of target node b (R b =N success / N total ), combined with the reliability weight coefficient λ (default 0.3), the incentive algorithm gives priority to high-reliability nodes for message forwarding.

[0088] This method can perceive changes in the external environment in real time and adjust its own values ​​as the environment changes, making learning more efficient and effective. Compared with algorithms that do not use dynamic factor adjustment, the adaptive routing algorithm developed based on this method has a higher delivery rate and can better adapt to changing environments.

[0089] Example 2

[0090] In this embodiment, an opportunistic routing forwarding algorithm based on adaptive Q-learning factor setting is adopted, and the algorithm is applied to the method in Example 1 to implement the ocean communication process, such as Figure 2 As shown in the figure, it is the core flow chart of this algorithm. The specific process is as follows:

[0091] 1. Network initialization: Nodes configure positioning and communication modules, periodically broadcast their own status information (coordinates, congestion level, delivery success rate), and maintain a dynamic neighbor table;

[0092] 2. Message processing: After receiving a message, the buffer status is determined. If it is full, a cleanup strategy (regular / emergency cleanup) is triggered, and messages are sorted by priority.

[0093] 3. Dynamic parameter adjustment: The learning rate (α) and discount factor (β) are updated in real time based on network congestion, message remaining life cycle, and node movement speed;

[0094] 4. Routing decision: Based on the dual-Q network, the Q value is updated asynchronously and the forwarding action that maximizes the overall reward (delay, hop count, node reliability) is selected;

[0095] 5. Resource management: Regularly clean up redundant messages (when the number of hops exceeds the limit or is close to expiration), and coordinate with neighboring nodes to reduce the sending frequency during emergency cleanup;

[0096] 6. Abnormal recovery: After detecting a disconnection, the learning rate is reduced and the detection protocol is started. If no recovery occurs, the system switches to satellite communication mode.

[0097] It fully demonstrates how the algorithm of the present invention can achieve efficient routing and resource optimization in the marine Internet environment through dynamic perception and adaptive mechanisms.

[0098] The routing algorithm includes the following steps:

[0099] Step 1: Network status awareness and dynamic parameter adjustment

[0100] Real-time perception of network congestion, message remaining life cycle, node movement speed, and historical delivery success rate;

[0101] Dynamically adjust the learning rate α(t) and discount factor β(t) according to the rules, and calculate the adaptive reward value;

[0102] Step 2: Q value update and routing decision

[0103] Adopt dual Q network asynchronous update mechanism to update Q value according to dynamic parameters;

[0104] Select the action that maximizes the overall reward, giving priority to forwarding messages with a short remaining lifespan, a high target node delivery success rate, or optimized hop count;

[0105] Step 3: Resource Management and Exception Handling

[0106] Regularly clean up redundant messages, and trigger emergency cleanup when storage utilization exceeds the threshold;

[0107] Detect network disconnection, reduce the learning rate, and initiate a detection protocol to restore the connection;

[0108] Step 4: Network performance feedback and parameter coordination

[0109] The reward function weight is dynamically modified through the sliding window mechanism and linked with the resource management strategy to suppress the generation of redundant messages in high congestion scenarios.

[0110] The opportunistic routing forwarding algorithm based on adaptive Q-learning factor setting proposed in the present invention realizes efficient message forwarding in the ocean Internet through the following steps:

[0111] In the initial stage of the algorithm, the hardware configuration and status synchronization of the network nodes are completed. The nodes are equipped with positioning modules, communication modules and computing units, and periodically broadcast their own status information (including location, network congestion C, historical delivery success rate R b), and maintain a dynamic neighbor table to perceive the real-time network topology.

[0112] In the message forwarding decision stage, the dual Q network decouples the evaluation and decision mechanism. Based on the Q1(s,a) network, the current optimal action a′=argmax is selected. a Q1(s′,a), using Q2(s,a) network to calculate the target value Q target =R+β·Q2(s′,a′), and by alternately updating the network parameters of Q1 and Q2, the independence of the evaluation and decision-making processes is ensured, avoiding the over-estimation problem of action value caused by the traditional single network.

[0113] The dynamic learning rate α is adjusted in real time according to the network size N(S), the cumulative number of messages sent by nodes N(m) and the congestion degree C. The calculation formula is:

[0114]

[0115] When the network scale expands or the congestion increases, the learning rate adaptively decays to suppress redundant exploration behavior.

[0116] The discount factor β is dynamically generated through multi-dimensional parameters, taking into account the remaining lifetime of the message. Forwarded hops And the node moving speed v. Its calculation formula is:

[0117]

[0118] Where β(t) is the discount factor; For the message The time it takes to transmit to the destination node; T is the remaining life cycle information of the message; current is the current network average delivery delay; Hops is the number of hops that the message has been forwarded; avg is the average value of hop count; β0 is the benchmark value of discount factor.

[0119] When the remaining lifetime is shorter or the node moves faster, the β weight increases significantly, and messages with high timeliness or stable paths are forwarded first.

[0120] Reward Function Dynamically integrate network performance indicators through a sliding window mechanism. Real-time calculation of delay rewards Bonus items with hop count where η t and η h Dynamic correction based on the average of the recent 10 time units. At the same time, combined with the historical delivery success rate R of the target node b b, the weight coefficient λ (default 0.3) is used to encourage the selection of high-reliability nodes to ensure that forwarding decisions take into account both efficiency and stability.

[0121] In the Q value update phase, an asynchronous interaction mechanism is used. After each decision is made, the formula Q(s,a)←Q(s,a)+α·[R+β·Q target -Q(s,a)] to update network parameters. Through the alternating optimization of the dual-Q network, the adaptability and convergence speed of the strategy are continuously improved, ultimately achieving efficient routing and forwarding in dynamic ocean environments.

[0122] In this embodiment, some routing algorithms are used for comparative analysis:

[0123] Epidemic: "Epidemic routing algorithm" is a routing protocol based on the "flooding" concept. It draws on the epidemic propagation mechanism. Nodes continuously forward message copies to encounter nodes to improve the delivery rate. It is suitable for networks with large dynamic topology changes.

[0124] The "Spray and Wait" algorithm is a classic two-phase routing protocol for delay-tolerant networks. In the first phase (Spray), message copies are distributed to a limited number of nodes. In the second phase (Wait), nodes holding copies wait for the target node to arrive before forwarding the message. This algorithm balances resource consumption and delivery efficiency by controlling the number of copies.

[0125] FQLR: Fuzzy Q-Learning Routing is a reinforcement learning routing algorithm that combines fuzzy logic and Q-learning. It fuzzifies network states (such as node cache and link quality) and uses Q-learning to dynamically optimize routing decisions, enabling more efficient message forwarding in resource-constrained or dynamic networks.

[0126] like Figure 3 The following is a line chart showing the impact of buffer size on delivery rate, comparing the delivery rates of different routing algorithms at different buffer sizes. The horizontal axis represents buffer size (unit: storage capacity) and the vertical axis represents delivery rate (range 0.2-0.9). The following performance curves are included for the following four algorithms:

[0127] Epidemic is marked by a dotted square. The delivery rate increases slowly as the buffer size increases, but is affected by redundant copies and reaches a maximum of only 0.52.

[0128] S&W is marked by a solid triangle. The delivery rate increases rapidly in the early stages, but is limited by the fixed number of replicas and tends to be saturated after the buffer exceeds 20.

[0129] FQLR (Fuzzy Q-Learning Routing) is marked by a long dashed line. Its delivery rate, combined with dynamic parameters, significantly outperforms the previous two algorithms, reaching a maximum of 0.705.

[0130] AQORA (the algorithm of the present invention) is marked with a solid star. Through dynamic learning rate adjustment, multi-dimensional discount factor design, and hierarchical resource management, its delivery rate consistently leads, reaching a maximum of 0.807, and the curve rises smoothly without saturation.

[0131] This figure demonstrates the superiority of our algorithm (AQORA, Adaptive Q-learning Optimized Routing Algorithm) in dynamic ocean networks. Traditional protocols (such as Epidemic and S&W) suffer from resource waste and low delivery rates due to static policies. AQORA significantly improves delivery efficiency through dynamic parameter coordination and resource management strategies. Especially when the buffer capacity is large (e.g., 30), AQORA's delivery rate is 54.7% higher than Epidemic's, resolving the "low delivery efficiency" and "waste of resources" issues mentioned in the background technology.

[0132] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A marine communication method based on adaptive Q-learning factor setting, characterized in that: In this method, each node periodically collects its own status information and broadcasts it to surrounding nodes; after receiving the message, each node stores the message in the buffer area; The routing decision and parameter adjustment process is repeated until all messages are transmitted to the target node, completing the communication process. The routing decision and parameter adjustment process includes: S1. The current node retrieves the message from the corresponding buffer and selects the next forwarding node for the message based on the current Q value table; S2. Check whether the next forwarding node is connected to the current node. If so, send the message to the next forwarding node and continue to execute S3. Otherwise, jump to execute S1. S3, based on the current state information of the node, update the learning rate and discount factor; S4. Based on the reward function, the updated learning rate and the discount factor, the dual Q network is asynchronously updated to obtain the updated Q value table.

2. The method for ocean communication based on adaptive Q-learning factor setting according to claim 1, characterized in that: The next forwarding node of the message in S1 is selected as follows: based on the current Q value table, the action a that maximizes the reward value is selected. * , the specific formula is: Among them, a * The next forwarding node for the message; is the action space, i.e., the set of neighbor nodes; is the reward function; β is the discount factor; Q(s′,a′) is the Q value under the next state s′ action a′.

3. The marine communication method based on adaptive Q-learning factor setting according to claim 1, characterized in that: If it is detected in S2 that the next forwarding node is not in a connected state with the current node, the learning rate is reduced and the detection protocol is started. If it is not restored within a preset period, the communication is switched to the satellite communication mode.

4. The method for ocean communication based on adaptive Q-learning factor setting according to claim 1, characterized in that: In S3, when updating the learning rate, the self-state information includes the network scale, the current node's cumulative message sending volume, and the network congestion. The specific formula for updating the learning rate is: Among them, α(t) is the learning rate, α0 is the learning rate baseline value; N(S) is the total number of network nodes, reflecting the network scale; N(m) is the cumulative number of messages sent by the current node; C(t) is the network congestion, which is the ratio of the number of cached messages to the maximum capacity.

5. The marine communication method based on adaptive Q-learning factor setting according to claim 4, characterized in that: In S3, when updating the discount factor, the state information includes the number of hops the message has been forwarded and the node's moving speed. The remaining life cycle information and the number of hops the message has been forwarded are also used in the discount factor update process. The specific formula is: Where β(t) is the discount factor; For news The time it takes to transmit to the destination node; T is the remaining life cycle information of the message; current is the current network average delivery delay; Hops is the number of hops that the message has been forwarded; avg is the average value of hop count; β0 is the benchmark value of discount factor.

6. The method for ocean communication based on adaptive Q-learning factor setting according to claim 5, characterized in that: The learning rate baseline value and discount factor baseline value are adjusted as the network scale changes. The specific formula is: Among them, α0 is the learning rate reference value; α init is the initial learning rate benchmark value; N(S) is the total number of network nodes, which is used to reflect the network scale; β0 is the discount factor benchmark value; β init is the initial discount factor benchmark value.

7. The method for ocean communication based on adaptive Q-learning factor setting according to claim 1, characterized in that: The specific formula of the reward function in S4 is: Among them, R a is the reward function; η t and η h is the dynamic weight; T current is the current network average delivery delay, is the estimated delay after message forwarding. The lower the delay, the higher the reward value. current is the average number of network hops, is the estimated number of hops after the message is forwarded. The fewer the hops, the greater the reward weight. b is the historical delivery success rate of target node b; λ is the reliability weight coefficient.

8. The method for ocean communication based on adaptive Q-learning factor setting according to claim 7, characterized in that: The dynamic weight η t and η h are the average of recent delay and hop count calculated by sliding window respectively. The specific formula is: Among them, η t and η h is the dynamic weight; W is the window size; is the average network delay; The time it takes for a message to travel from node a to node b; is the average number of hops for message delivery in the network; is the number of hops that a message takes from node a to node b.

9. The method for ocean communication based on adaptive Q-learning factor setting according to claim 7, characterized in that: The specific process of the dual-Q network asynchronous update in S4 is as follows: Among them, Q1'(s,a) and Q2'(s,a) are the updated Q networks respectively; Q1(s,a) and Q2(s,a) are the Q networks before the update respectively; α(t) is the learning rate; R is the reward function value; β(t) is the discount factor; s and s ′ are the current state and the state after the message is forwarded; a and a ′ They are the current action and the action after the message is forwarded.

10. The method for ocean communication based on adaptive Q-learning factor setting according to claim 1, characterized in that: The buffer periodically clears redundant messages and coordinates neighboring nodes to reduce the sending frequency when clearing redundant messages.

Citation Information

Patent Citations

  • Hybrid routing method based on clustering and reinforcement learning and ocean communication system

    CN111510956A

Cited By

  • Steady-state abrupt-change adaptive deep Q learning underwater routing method

    CN121924557A

  • Steady-state mutation adaptive deep q-learning underwater routing method

    CN121924557B