Meta-Learning-Based Hyperparameter Reweighting MAC Method for Underwater Acoustic Networks

By designing a MAC framework based on double-layer optimization in the water acoustic sensor network, and using meta-learning and Q learning to optimize the hyperparameter weight function, the problem of inefficient transmission in the water acoustic network is solved, efficient data transmission and energy management are achieved, and network life cycle is extended.

CN116321431BActive Publication Date: 2025-07-25XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310298915.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2025-07-25
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

The existing water acoustic sensor network MAC protocol has problems such as high energy consumption, high computational complexity, and large-scale changes in the number of nodes in the underwater data transmission, and has not fully considered parameter optimization, resulting in low transmission efficiency.

Method used

A water acoustic network MAC framework based on double-layer optimization is designed, using meta-learning to optimize the hyperparameter weight function and Q learning to avoid conflicts, combining the information key to the transmission node and residual energy as parameters, and optimizing the reward function of Q learning through gradient descent method to realize medium access control.

Benefits of technology

It improves the efficiency of water acoustic data transmission, extends the life cycle of water acoustic sensor network, reduces energy consumption, and avoids data transmission conflicts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116321431B_ABST
    Figure CN116321431B_ABST
Patent Text Reader

Abstract

Meta-Learning-Based Hyperparameter Reweighting MAC Method for Underwater Acoustic Networks, which relates to underwater acoustic networks. In a single-hop underwater acoustic sensor network, a two-layer optimization-based framework is designed to improve the efficiency of underwater acoustic data transmission by adopting a meta-learning-based algorithm. In the core layer, the information criticality and the remaining energy of the transmission node are introduced as parameters into the underwater acoustic network, and meta-learning is used to reweight their weights, and the optimized weights are imported into the nested layer; in the nested layer, in view of the characteristics of large transmission delay of underwater acoustic data and limited energy of transmission nodes, the Q-learning algorithm is adopted in the medium access control to avoid data loss caused by data collisions during data transmission. By combining meta-learning with Q-learning, the data transmission efficiency is improved, transmission conflicts and data loss are effectively avoided, the system energy consumption of underwater acoustic data transmission is reduced, and the stability of the underwater acoustic sensor network is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an underwater acoustic network, and in particular to a hyperparameter reweighted underwater acoustic network medium access control method based on meta-learning. Background Art

[0002] The foundation that supports the rapid development of the Internet of Things is wireless communication technology represented by 5G. Its low cost and high speed have enabled communication networks to cover the globe. However, in the application scenarios of the marine Internet of Things, the inherent characteristics of underwater acoustic sensor networks, such as long transmission delay, narrow available bandwidth, high signal-to-noise ratio, and strong multipath effects, make it extremely challenging to implement underwater acoustic communication technology. Since electromagnetic waves and light waves attenuate severely underwater, acoustic waves are still mainly relied on for data transmission underwater.

[0003] In an underwater acoustic sensor network, the medium access control (MAC) protocol is located in the data link layer of the underwater node architecture. Its main task is to allocate communication channels to nodes in the underwater acoustic network in a certain efficient manner, ensuring that the communication processes between nodes do not conflict with each other, thereby improving the efficiency of underwater acoustic communication and the performance of the underwater acoustic network. Therefore, whether the underwater acoustic network MAC protocol can reasonably and efficiently utilize the underwater acoustic channel is an important factor affecting the efficiency of the underwater acoustic sensor network. Selecting a suitable underwater acoustic network MAC protocol is of great significance for avoiding collisions at the receiving end and improving the overall performance of the underwater acoustic sensor network.

[0004] Underwater acoustic sensor nodes are powered by batteries and consume a large amount of energy during underwater data transmission. To save energy, Junho Cho et al. (Junho Cho, et al. Power Control for MACA-based Underwater MAC Protocol: A Q-Learning Approach [c]. 2021 IEEE Region 10 Symposium (TENSYMP), 2021.) proposed a power control MAC protocol for underwater multiple access collision avoidance based on Q-learning. This protocol uses Q-learning to enable sensor nodes to prevent collisions without any prior knowledge of interference while maintaining high energy efficiency, eliminating the need for additional signaling, and significantly improving the energy efficiency and throughput of the underwater acoustic network. However, this protocol has problems such as hidden terminals, inability to scale the number of nodes, and low computing power and complexity.

[0005] At present, there is little research on MAC algorithms for underwater acoustic sensor networks considering parameter optimization. The present invention introduces meta-learning into underwater acoustic networks to make up for the deficiencies. Different from most deep learning, meta-learning requires retraining each time. Meta-learning hopes to enable the model to obtain an ability of "learning to learn", so that it can quickly learn new tasks based on the prior knowledge already acquired. Meta-learning is the self-update of the algorithm, which helps to find the algorithm with the best performance and the hyperparameters of the corresponding algorithm, and optimize the number of experiments, so as to make better predictions in a shorter time. Jun Shu et al. (Jun Shu, et al. Meta-Weight-Net: Learning an Explicit Mapping For Sample Weighting[J]. 33rd Conference on Neural Information Processing Systems, 2019.) proposed a method for adaptively learning an explicit weight function directly from data, and its explicit weight function is automatically learned from data by parameterizing a fully connected neural network (MLP) in a meta-learning manner. The present invention intends to design a MAC design framework for underwater acoustic networks based on double-layer optimization. Among them, the core layer is a weight function of hyperparameters that can be automatically learned from meta-data based on meta-learning, and the nested layer is a single-hop underwater acoustic network MAC protocol based on Q-learning. Summary of the Invention

[0006] The purpose of the present invention is to use the information criticality and remaining energy of the transmission node as hyperparameters, provide a method for parameter optimization using a meta-learning weight function, apply it to the MAC time slot allocation of a single-hop underwater acoustic network, avoid data loss caused by collisions during the data collection process of the source node, so as to improve the energy utilization efficiency of the underwater acoustic communication network and extend the life cycle of the underwater acoustic sensor network.

[0007] The present invention includes the following steps:

[0008] 1) Consider an underwater acoustic network composed of randomly arranged underwater acoustic sensor nodes, including 1 sink node S and M transmission nodes T i (i = 1, 2,..., M), the sink node S is within the data communication range of all transmission nodes; the transmission node T i is responsible for sensing information from the ocean, and the sink node S is responsible for collecting the information sensed by the transmission node T i ; assume that the initial energy of each transmission node T i is E0; the sink node S grades the collected data according to criticality, expressed as IK = 1, IK = 2, IK = 3, IK = 4, IK = 5, representing the five information criticality levels of "level one, level two, level three, level four, level five" respectively, and the higher the level, the higher the criticality; use IK i(i = 1, 2, ..., M) represents the information criticality of the data of the transmission node.

[0009] 2) To ensure that each transmission node T i can send data to the sink node S, the process of the sink node S collecting data is divided into M time slots; in the Q - learning algorithm of the MAC protocol applied in the underwater acoustic network, the Q - matrix is an M×M matrix, where the row number represents the serial number of the transmission node T i , and the column number represents the time - slot serial number; each transmission node T i only needs to store a row sub - matrix of the selected time slot internally; the initial Q - matrix is a zero matrix, and the maximum number of iterations is K.

[0010] 3) In Q - learning, the value of Q(s, a) is the state - action, representing the Q - value when the s - numbered transmission node occupies the a - numbered time slot to send data to the sink node S. The expected reward obtained after taking this action is Q * (s, a), which can be approximated by iteration:

[0011] Q * (s, a)=(1 - α)Q(s, a)+α[r + γmax a∈A Q(s′, a′)] (1)

[0012] Among them, the value of Q(s′, a′) represents the value of Q(s, a) next time; α is the learning rate, which determines the speed of Q - value update. It should not be too large, otherwise it may oscillate around the minimum point continuously. It should not be too small, otherwise it will take a lot of iteration times to reach the minimum point; γ is the discount factor, representing the influence of the future on the present. When γ = 0, the system is short - sighted and only considers the current result of the action. When γ approaches 1, the future rewards become more important when taking the optimal action;

[0013] r represents the reward when the s - numbered transmission node occupies the a - numbered time slot to send data to the sink node S, and its formula is defined as:

[0014] r=-g - β1e(M m )+β2k(M m ) (2)

[0015]

[0016]

[0017] β1 + β2 = 1 (5)

[0018] Among them, g represents the continuous penalty; E m represents the remaining energy of the transmission node T i , e(M m ) represents the transmission node Ti Reward for the remaining energy, where β1 represents the weight of the remaining energy; k(M m ) represents the reward for the information criticality, and β2 represents the weight of the information criticality.

[0019] 4) During a period when the underwater acoustic network starts to operate, the energy of the transmission node T i is sufficient, the weight β2 of the information criticality has a relatively large proportion, and the weight β1 of the remaining energy of the transmission node T i has a relatively small proportion; as the energy consumption increases, the energy of the transmission node T i becomes insufficient, the weight β2 of the information criticality becomes smaller, and the weight β1 of the remaining energy of the transmission node T i becomes larger.

[0020] 5) Since β1 and β2 can be determined from each other, it is only necessary to optimize the parameter of the weight β1 of the reward function r in the Q - learning algorithm according to meta - learning.

[0021] 6) Design a two - layer optimization framework and adopt the method of gradient descent in the core layer.

[0022] In meta - learning, according to prior knowledge, determine the linear function of the relationship between β1 and the remaining energy E of the transmission node m , that is, the weight function can be expressed as

[0023]

[0024] where the parameters A and B are constants, collectively referred to as hyperparameters The loss function l of meta - learning is defined as follows. Seek a series of alternative parameters through the gradient descent method to achieve self - optimization of the hyperparameters:

[0025]

[0026] where N represents the number of metadata, is the predicted value of β1 after one optimization, and θ i is the actual value of β1. Its gradient is as follows:

[0027]

[0028] Set and the learning rate ε, and the gradient descent can be expressed as:

[0029]

[0030] And so on, a series of alternative parameters λ 1 , λ 2 , λ 3 ,...., λN .

[0031] 7) In the nested layer, substitute the alternative parameters obtained in the core layer into the loss function L of Q-learning, and calculate the hyperparameters that minimize the loss function L That is

[0032]

[0033] Let the transmission node be T i The remaining energy is E m1 、E m2 、E m3 When the value of the Q-learning reward function β1 corresponds to

[0034] Substitute E m1 、E m2 、E m3 into the weight function of the alternative parameter λ 1 respectively, and obtain the value of the Q-learning reward function β1 respectively Furthermore, the value calculated by the loss function is the loss value of the alternative parameter λ 1 in the weight function

[0035] Similarly, substitute E m1 、E m2 、E m3 into the weight function of the alternative parameter λ i respectively, and obtain the loss value L of the alternative parameter λ i in the weight function i . Select the alternative parameter λ N corresponding to the minimum L i from L1, L2,......, L i , which is the hyperparameter of the best-performing weight function, denoted as

[0036] 8) When the transmission node T i sends a data transmission request to the destination node S, if the time slot number corresponding to the maximum Q value in the transmission node submatrix is unique and does not conflict with other transmission nodes, that is, select this time slot number to send a data transmission signal; if the time slot number corresponding to the maximum Q value in the transmission node submatrix is not unique, or two or more transmission nodes select time slots that conflict, after calculating the weight function to update the parameters of the Q-learning reward function r, then calculate the optimal time slot for data transmission

[0037] The present invention can effectively improve the efficiency of underwater acoustic data transmission. By introducing the information criticality and the remaining energy of transmission nodes as parameters into the underwater acoustic sensor network, and using parameter optimization of Q-learning and meta-learning to avoid conflicts in the medium access control layer of the underwater acoustic network.

[0038] The present invention has the following outstanding advantages:

[0039] 1) In an underwater acoustic single-hop network, a framework based on double-layer optimization is designed. In the core layer, the weights of information criticality and the remaining energy of transmission nodes are used as hyperparameters. Utilizing the self-update ability of the meta-learning algorithm, which helps to find the best-performing hyperparameters, a parameter optimization scheme for the medium access control protocol is proposed according to the information change of the remaining energy of transmission nodes, extending the usage cycle of the underwater acoustic communication network.

[0040] 2) In the nested layer of the present invention, the optimized hyperparameters are used to update the reward function of Q-learning to select time slots, avoiding data transmission conflicts and improving data transmission efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a random node coordinate distribution diagram of the underwater acoustic sensor network for the hyperparameter reweighting underwater acoustic network medium access control method based on meta-learning of the present invention.

[0042] Figure 2 It is a flowchart of parameter optimization using meta-learning in the underwater acoustic sensor network for the hyperparameter reweighting underwater acoustic network medium access control method based on meta-learning of the present invention.

[0043] Figure 3 It is a graph showing the change of the remaining energy of transmission nodes in the underwater acoustic sensor network for the hyperparameter reweighting underwater acoustic network medium access control method based on meta-learning of the present invention with the number of iterations.

[0044] Figure 4 It is a graph showing the change of the remaining energy weight β1 with the number of iterations in the underwater acoustic sensor network for the hyperparameter reweighting underwater acoustic network medium access control method based on meta-learning of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0045] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] 1) As Figure 1 , in an underwater acoustic sensor network, 9 nodes are randomly arranged, including 1 sink node S (node Sink) and 8 transmission nodes T i (i = 1, 2,..., 8). The sink node S is within the data communication range of all transmission nodes; the transmission node T i is responsible for sensing information from the ocean, and the sink node S is responsible for collecting the transmission node T iPerceived information; let each transmission node T i have an initial energy of E0; the sink node S grades the collected data according to its criticality, denoted as IK = 1, IK = 2, IK = 3, IK = 4, IK = 5, representing the five information criticality levels of "first level, second level, third level, fourth level, fifth level" respectively. The higher the level, the higher the criticality; use IK i (i = 1, 2,..., 8) to represent the information criticality of the data of the transmission node.

[0047] 2) To ensure that each transmission node T i can send data to the sink node S, the process of the sink node S collecting data is divided into 8 time slots; in the Q - learning algorithm of the MAC protocol applied in the underwater acoustic network, the Q - matrix is an 8×8 matrix, where the row number represents the serial number of the transmission node T i , and the column number represents the time slot serial number; each transmission node T i only needs to store a row sub - matrix of the selected time slot internally; the initial Q - matrix is a zero matrix, and the maximum number of iterations is K = 2000.

[0048] 3) In Q - learning, the value of Q(s, a) is the state - action, representing the Q - value when the s - numbered transmission node occupies the a - numbered time slot to send data to the sink node S. The expected reward obtained after taking this action is Q * (s, a), which can be approximated through iteration:

[0049] Q * (s, a)=(1 - α)Q(s, a)+α[r + γmax a∈A Q(s′, a′)] (1)

[0050] Among them, the value of Q(s′, a′) represents the value of Q(s, a) next time; α is the learning rate, which determines the speed of Q - value update. It should not be too large, otherwise it may oscillate around the minimum point continuously. It should not be too small either, otherwise it will take a lot of iteration times to reach the minimum point; γ is the discount factor, representing the influence of the future on the present. When γ = 0, the system is short - sighted and only considers the current result of the action. When γ approaches 1, the future rewards become more important when taking the optimal action;

[0051] r represents the reward when the s - numbered transmission node occupies the a - numbered time slot to send data to the sink node S, and its formula is defined as:

[0052] r=-g - β1e(M m )+β2k(M m ) (2)

[0053]

[0054]

[0055] β1 + β2 = 1 (5)

[0056] Among them, g represents the continuous penalty; E m represents the remaining energy of the transmission node T i , e(M m ) represents the reward for the remaining energy of the transmission node T i , β1 represents the weight of the remaining energy; k(M m ) represents the reward for the information criticality, and β2 represents the weight of the information criticality.

[0057] 4) During a period when the underwater acoustic network starts to operate, the transmission node T i has sufficient energy, the weight β2 of the information criticality has a relatively large proportion, and the weight β1 of the remaining energy of the transmission node T i has a relatively small proportion; as the energy consumption increases, the energy of the transmission node T i becomes insufficient, the weight β2 of the information criticality becomes smaller, and the weight β1 of the remaining energy of the transmission node T i becomes larger.

[0058] 5) Since β1 and β2 can be determined by each other, it is only necessary to optimize the parameter of the weight β1 of the reward function r in the Q - learning algorithm according to meta - learning.

[0059] 6) Design a two - layer optimization framework and adopt the method of gradient descent in the core layer.

[0060] In meta - learning, according to prior knowledge, determine the linear function of the relationship between β1 and the remaining energy E of the transmission node m , that is, the weight function can be expressed as

[0061]

[0062] Among them, the parameters A and B are constants, collectively called hyperparameters The loss function l of meta - learning is defined as follows. Seek a series of alternative parameters through the gradient descent method to achieve self - optimization of hyperparameters:

[0063]

[0064] Among them, N represents the number of metadata, is the predicted value of β1 after one optimization, and θ i is the actual value of β1. Its gradient is as follows:

[0065]

[0066] Set And the learning rate ε, gradient descent can be expressed as:

[0067]

[0068] And so on, a series of alternative parameters λ can be obtained 1 , λ 2 , λ 3 ,...., λ N .

[0069] 7) In the nested layer, substitute the alternative parameters obtained in the core layer into the loss function L of Q-learning, and calculate the hyperparameter that makes the loss function L reach the minimum That is

[0070]

[0071] Let the remaining energy of the transmission node T i be E m1 , E m2 , E m3 When the Q-learning reward function β1 value corresponds to

[0072] Substitute E m1 , E m2 , E m3 into the weight function of the alternative parameter λ 1 respectively, and obtain the Q-learning reward function β1 value in it respectively Furthermore, the value calculated by the loss function is the loss value of the alternative parameter λ 1 in the weight function

[0073] Similarly, substitute E m1 , E m2 , E m3 into the weight function of the alternative parameter λ i respectively, and obtain the loss value L of the alternative parameter λ i in the weight function i . Among L1, L2,......, L N , select the alternative parameter λ i corresponding to the minimum L i , which is the hyperparameter of the weight function with the best performance, denoted as

[0074] 8) When the transmission node T iWhen sending a data transmission request to the destination node S, if the time slot number corresponding to the maximum Q value in the transmission node sub-matrix is unique and does not conflict with other transmission nodes, then select this time slot number to send the data transmission signal; if the time slot number corresponding to the maximum Q value in the transmission node sub-matrix is not unique, or the time slots selected by two or more transmission nodes conflict, then calculate the weight function After updating the parameters of the Q-learning reward function r, calculate the optimal time slot for data transmission again.

[0075] The feasibility of the method described in the present invention is verified by computer simulation below.

[0076] The simulation platform is MATLAB R2022a.

[0077] The parameter settings are as follows: the initial energy of the node E0 = 70J; the signal-to-noise ratio SNR = 25dB; the directivity index DI = 0dB; the data transmission frequency F = 10KHZ.

[0078] The parameter settings of the Q-learning model are as follows: the exploration iteration number K = 2000; the learning rate α = 0.9; the discount factor γ = 0.8; the continuous penalty g = 0.1; the weight of the initial remaining energy β1 = 0.64; the weight reward weight of the information criticality β2 = 1 - β1 = 0.36; the communication range of the destination node is 1000m; the Q-value table is initialized as an 8×8 zero matrix.

[0079] The parameter setting of β1 of the parameter optimization model based on meta-learning is as follows: the initial learning rate ε = 0.004;

[0080] (1) As shown by Figure 2 In the core layer, the weight β1 of the reward function r in the Q-learning algorithm is updated according to meta-learning. The specific steps are as follows:

[0081] ① According to formulas (8) and (9), a series of alternative parameters λ 1 , λ 2 , λ 3 ,...., λ N are obtained according to gradient descent.

[0082] ② According to the loss value L of the alternative parameter λ i in the weight function i , select the alternative parameter λ N in L1, L2,......, L i that makes L i the smallest, which is the best-performing weight function hyperparameter, denoted as

[0083] ③ The remaining energy E of the transmission nodem Substitute the hyperparameters into the weight function to obtain the parameter β1 of the Q-learning reward function r, and realize the parameter optimization of the reward function.

[0084] (2) Reward matrix R of the transmission node 8×8 , and its reward is defined as:

[0085] r = -g - β1e(M m ) + (1 - β1)k(M m )

[0086]

[0087]

[0088] where g represents the continuous penalty; e(M m ) represents the reward for the remaining energy of the transmission node T i , β1 is the weight of the remaining energy after meta-learning update, and k(M m ) represents the reward for the information criticality.

[0089] Figure 3 and Figure 4 are respectively the graph of the remaining energy of the transmission node in the underwater acoustic sensor network changing with the number of iterations and the graph of the remaining energy weight β1 of the underwater acoustic sensor network changing with the number of iterations. It can be seen from the graph that the remaining energy of the transmission node gradually decreases with the increase of the number of iterations; at the same time, in the core layer, the remaining energy weight β1 of the transmission node is optimized through the parameter optimization model based on meta-learning, making β1 larger and larger. When the energy of the transmission node is sufficient, the transmission node with higher information criticality has a larger weight to send data to the sink node, which can ensure that important data can be transmitted in time; when the energy of the transmission node is insufficient, it has a larger weight to communicate with the sink node, which can balance the energy consumption of each node, reduce the energy hole, and effectively extend the life cycle of the underwater acoustic sensor network.

Claims

1. A hyperparameter reweighting underwater acoustic network medium access control method based on meta-learning, characterized in that It includes the following steps: 1) Consider an underwater acoustic network composed of randomly deployed underwater acoustic sensor nodes, including 1 sink node S and M transmission nodes T i , where i = 1, 2,..., M; the sink node S is within the data communication range of all transmission nodes; the transmission node T i is responsible for sensing information from the ocean, and the sink node S is responsible for collecting the information sensed by the transmission node T i ; assume that the initial energy of each transmission node T i is E0; the sink node S classifies the collected data according to the criticality level, denoted as IK = 1, IK = 2, IK = 3, IK = 4, IK = 5, representing the five information criticality levels of "level 1, level 2, level 3, level 4, level 5" respectively, and the higher the level, the higher the criticality; use IK i to represent the information criticality of the transmission node data, where i = 1, 2,..., M; 2) To ensure that each transmission node T i sends data to the sink node S, the process of the sink node S collecting data is divided into M time slots; in the Q-learning algorithm of the MAC protocol applied in the underwater acoustic network, the Q matrix is an M×M matrix, where the row number represents the serial number of the transmission node T i and the column number represents the serial number of the time slot; each transmission node T i only needs to store a row sub-matrix of the selected time slot internally; The initial Q matrix is a zero matrix, and the maximum number of iterations is K; 3) r represents the reward when the s-th transmission node occupies the a-th time slot to send data to the destination node S, and its formula is defined as: r = -g - β1e(M m ) + β2k(M m ) (2) β1+β2=1 (5) Among them, g represents continuous punishment; E m represents the remaining energy of the transmission node T i , e(M m ) represents the reward for the remaining energy of the transmission node T i , β1 represents the weight of the remaining energy; k(M m ) represents the reward for information criticality, and β2 represents the weight of information criticality; 4) Design a two-layer optimization framework and adopt the gradient descent method in the core layer; In meta - learning, a linear function that determines the relationship between β1 and the remaining energy E of the transmission node according to prior knowledge, that is, the weight function m is expressed as denoted as Among them, parameters A and B are constants, collectively referred to as hyperparameters The loss function l of meta-learning is defined as follows. A series of alternative parameters are sought through the gradient descent method to achieve self-optimization of hyperparameters: where N represents the number of metadata, is the predicted β1 value after the first optimization, and θ i is the actual β1 value, and its gradient is as follows: Set and learning rate ε, gradient descent is expressed as: And so on, a series of alternative parameters λ are obtained 1 , λ 2 , λ 3 ,...., λ N ; 5) In the nested layer, substitute the alternative parameters obtained in the core layer into the loss function L of Q-learning, and calculate the hyperparameters that minimize the loss function L That is: Set the transmission node T i The remaining energy is E m1 、E m2 、E m3 When the Q-learning reward function β1 value corresponds to Substitute E m1 , E m2 , and E m3 into the weight function of the alternative parameter λ 1 respectively, and obtain the Q-learning reward function β1 values respectively; Furthermore, the value calculated by the loss function is the loss value of the alternative parameter λ 1 in the weight function; Similarly, E m1 、E m2 、E m3 Substitute the candidate parameters λ respectively i The weight function In the weight function, we can get the alternative parameter λ i The loss value L i ; In L1, L2, ..., L N Select L i The minimum corresponding candidate parameter λ i , which is the best performing weight function hyperparameter, denoted as

Citation Information

Patent Citations

  • Q learning-based medium access control method for underwater acoustic network with variable number of nodes

    CN113691391A

  • Underwater acoustic network medium access control method based on Q learning and data importance

    CN114423083A