A Trust Management Method for Underwater Sensor Networks Based on Federated Deep Reinforcement Learning
By adopting the trust management method of federal deep reinforcement learning in underwater sensor networks, the evaluation problem of traditional trust models in heterogeneous scenarios is solved, achieving higher evaluation accuracy and lower update costs.
Patent Information
- Application Number
- CN202310994515.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-09
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-08-09
AI Technical Summary
In the trust evaluation task in heterogeneous scenarios, traditional trust model for isomorphic network design has problems such as poor evidence quality, low detection accuracy and high update cost.
The trust management method of underwater sensor network based on federated deep reinforcement learning is adopted to build a cross-domain joint trust management architecture, use deep reinforcement learning to achieve adaptive adjustment and parameter matching of trust models, and design a cross-domain trust update framework through federated learning to realize cross-domain update of global models and periodic adjustment of local models.
It improves the evaluation accuracy of trust models in space-time and space-changing scenarios, reduces update costs, and avoids additional energy consumption caused by large-scale evidence sharing.
Smart Images

Figure CN117082492B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a trust management method for underwater sensor networks based on federated deep reinforcement learning, belonging to the technical field of underwater acoustic sensor network communication support. Background Art
[0002] Underwater sensor networks are a new paradigm formed to meet the integrated application requirements of complex underwater information perception, data transmission, and information processing. By deploying sensor nodes in the monitored water area and using underwater acoustic communication as the main medium for data transmission between underwater nodes, underwater sensor networks can be widely applied in many fields such as marine disaster warning, hydrological information monitoring, marine resource exploration, underwater assisted navigation, and offshore military defense. The open and unmanned nature of underwater sensor networks makes it easy for normal nodes inside them to be invaded and compromised by malicious programs, and then turn into malicious nodes with legitimate network identities but performing network attack behaviors such as packet dropping, data content tampering, and high-frequency interactive requests. The trust management mechanism can establish a prediction model based on the historical interaction experience of both parties in the interaction, so as to predict the future behavior of the evaluation object using the obtained trust evidence and avoid the threat of malicious attacks in advance. In recent years, with the rapid development of the underwater robot industry, a new generation of intelligent sensing devices such as autonomous underwater vehicles, underwater gliders, and wave gliders have joined the underwater acoustic sensor network system, making the traditional homogeneous underwater acoustic sensor network composed of a single sensor / sensor array gradually evolve into a heterogeneous underwater acoustic sensor network composed of multiple types of underwater / surface sensing devices. There are differences in capabilities such as communication, computing, data storage, and mobility between different types of devices. Traditional trust models designed for small-scale homogeneous networks are difficult to avoid problems such as poor evidence quality, low detection accuracy, and high update costs. Therefore, it is necessary to design a cross-domain joint trust management model for cross-domain heterogeneous underwater acoustic sensor networks to achieve low-cost cross-domain joint trust modeling and update while ensuring the security of local data.
[0003] In order to obtain a high-quality trust model for underwater sensor networks, researchers at home and abroad have proposed various solutions, and the relevant literature is as follows:
[0004] 1. In 2020, aiming at the impacts of unreliable underwater acoustic channels, weak connections, long time delays and other characteristics on trust evaluation in UASNs, Su et al. proposed a highly robust trust evaluation method for underwater nodes in "A trust model for underwater acoustic sensor networks based on fast link quality assessment". First, based on fast link quality assessment, a geometric triangle model composed of signal-to-noise ratio, link quality indicator and origin is established, and the link quality is equated to the equivalent distance between parameters, so as to determine the link quality through a preset threshold. Then, trust evidence is obtained based on the link quality assessment result, and the trust value of the node is comprehensively calculated according to direct trust evidence, indirect trust evidence and recommended trust evidence. Finally, a malicious node detection method based on threshold isolates malicious nodes in the network, thus providing stable and accurate node trust evaluation.
[0005] 2. In 2021, in order to further accurately evaluate nodes in a complex underwater environment and reduce the possibility of misjudging normal nodes, Su et al. proposed a redeemable SVM-DS fusion-based trust management mechanism in "A redeemable SVM-DS fusion-based trust management mechanism for underwater acoustic sensor networks". This mechanism selects three types of trust evidence according to the characteristics of internal attacks: packet-based evidence, data-based evidence and energy-based evidence. The nodes are classified according to the trained SVM model, and then the D-S evidence theory is applied to fuse the classification results of these three types of trust evidence on the target node to obtain the overall trust evaluation. In addition, a trust redemption process is introduced. When the historical trust degree of the node is relatively high, it is decided whether to perform trust redemption processing by evaluating the degree of environmental impact, so as to reduce the possibility of misjudging normal nodes, thereby improving the security of nodes in UASNs.
[0006] 3. In 2020, Jiang et al. proposed an underwater dynamic trust evaluation and update mechanism (TEUC) based on the C4.5 decision tree in the paper "A dynamic trust evaluation and update mechanism based on C4.5 decision tree in underwater wireless sensor networks". The trust level of nodes is evaluated more accurately through the combination of different types of trust evidence. The TEUC algorithm includes three main stages: trust evidence collection, decision tree training, and trust evaluation and update. In the first stage, different types of trust evidence are collected from different sources. Among them, data-based trust evidence is obtained by analyzing the data generated by sensor nodes, link-based trust evidence is obtained by analyzing the communication link quality between nodes, and node-based trust evidence is obtained by analyzing the behavior of individual nodes. In the second stage, for the collected trust evidence, the trust evidence is used as the features of the samples, and the trust level of the nodes is used as the label to train the C4.5 decision tree. Finally, the trust level of the nodes is dynamically evaluated and updated using the trained decision tree model. When new information is available (such as new data samples or link quality measurements), the trust level of the corresponding nodes is updated in real time, and the evaluation model is also updated, which enables TEUC to adapt to changes in network conditions over time.
[0007] 4. In 2020, Du et al. proposed the trust model ITrust in the paper "ITrust: An anomaly-resilient trust model based on isolation forest for underwater acoustic sensor networks". ITrust focuses on solving the problem of trust instability caused by noise in the underwater environment. The impact of environmental noise is quantified as an index, called environmental trust; then it is combined with communication trust, data trust, and energy trust as the features of the dataset samples; finally, the isolation forest algorithm in machine learning is used to classify good and bad behaviors.
[0008] 5. In 2022, Du et al. further proposed the trust model LTrust in "LTrust: an adaptive trust model based on LSTM for underwater acoustic sensor networks". LTrust pays more attention to the problem of trust relationship fluctuations caused by network topology changes. This scheme chooses to quantify the topological relationship between adjacent nodes as node importance and embeds it into the node trust evaluation process. The trust dataset in LTrust consists of four attributes: node trust, communication trust, environmental trust, and recommendation trust. The final trust evaluation of node behavior is achieved through the LSTM algorithm. Both ITrust and LTrust use supervised learning for the final trust classification. However, the complex and changeable underwater environment leads to the problem that the trust prediction model based on a specific training set often shows insufficient generalization ability. Summary of the Invention
[0009] The technical problem to be solved by the present invention is that when the trust model designed for homogeneous networks performs trust evaluation tasks in heterogeneous scenarios, problems such as poor evidence quality, low detection accuracy, and high update cost occur. Combining deep reinforcement learning and federated learning technologies, a trust management method for underwater sensor networks based on federated deep reinforcement learning is proposed. First, a cross-domain joint trust management architecture is constructed; then, taking trust evidence as the state and trust model parameters as the action, which are used as the input and output of deep reinforcement learning respectively, and the adaptive adjustment and parameter matching of the trust model are realized by using reinforcement learning; for the trust model parameters within each local subnet, a cross-domain trust update framework is designed by using federated learning to realize the cross-domain update of the global model and the periodic adjustment of the local model, while avoiding the additional energy consumption caused by large-scale evidence sharing and further improving the evaluation accuracy of the trust model in scenarios with spatio-temporal changes.
[0010] To achieve the above object, the present invention is realized by the following technical solutions:
[0011] A trust management method for underwater sensor networks based on federated deep reinforcement learning, comprising the following steps:
[0012] Step 1: Construction of the trust management framework
[0013] Logically, the trust management framework divides the underwater sensor network into a control layer and a data layer; various devices in the network are divided into a global controller, local controllers, and data collectors according to their functions; the only global controller is responsible for the policy regulation of the entire network and the global trust update; each local controller trains and maintains a local trust model by using the data aggregated from its respective area; the data collectors transmit the interaction information between each other to the local controller to which they belong while collecting underwater data.
[0014] Step 2: Deep Reinforcement Learning Trust Modeling
[0015] Deep reinforcement learning trust modeling includes two stages: pre-training on the global controller and actual training on the local controller;
[0016] After network initialization, the global controller performs model pre-training based on the virtual interaction environment to make up for the lack of initial interaction experience; subsequently, the pre-trained model parameters are passed to the local controller as the initialization parameters of the local model, and each local controller further trains the model according to the actual interaction data among the collectors in its respective area;
[0017] Step 3: Federal Learning Trust Model Update
[0018] Each local controller periodically sends the parameters of the local model to the global controller, instead of directly transmitting the historical interaction data, to achieve the decoupling of model update and data storage, thereby reducing the update cost while protecting local data privacy; after receiving the parameters of each local model, the global controller performs global update of the parameters using the federated averaging method and republishes the updated model parameters to all local controllers; the local controller replaces the local model parameters with the received global parameters and continues to update the model based on the actual interaction data; continuously iterate the above update process so that the trust prediction model can continuously adapt to the dynamically changing underwater network environment.
[0019] In the above Step 1, the method for constructing the trust management framework is as follows:
[0020] An underwater sensor network composed of various types of devices such as wave gliders, underwater gliders, AUVs, subsea flight nodes, and underwater mooring arrays, where each type of device is divided into a global controller, a local controller, and a data collector according to its function, and the entire network is logically divided into a control layer and a data layer;
[0021] The control layer includes wave gliders deployed on the water surface and multiple AUVs underwater; the wave glider serves as the only global controller, responsible for updating the trust model and task scheduling, and at the same time receiving commands from the shore-based data center through satellite relay to adjust the tasks executed by the network; each AUV serves as a local controller, and manages the data layer devices in its respective responsible area by receiving scheduling commands from the global controller; both the global controller and the local controller have sufficient computing, communication, and storage capabilities to support the trust management system;
[0022] The data layer includes various heterogeneous underwater data collection devices such as underwater gliders, underwater anchored nodes, and seabed flying nodes, collectively referred to as collectors. Each collector collaborates to execute tasks and generate interaction information; the underwater gliders, seabed flying nodes, and underwater anchored nodes in the data layer represent underwater devices with different mobility capabilities respectively; the underwater glider represents a collector with higher mobility and can move between areas covered by different local controllers; the seabed flying node represents a collector with limited mobility and can only move within a local range and cannot move across regions; the underwater anchored node represents a fixed-position underwater collector;
[0023] The deployed trust management architecture is divided into three modules: trust evidence collection, trust modeling, and trust update; first, during the data collection process, underwater collectors record the interaction information between devices, generate trust evidence, and send it to the subordinate local controller; then, the local controller trains a trust model based on the obtained trust evidence and periodically sends the model parameters to the global controller; finally, the global controller performs joint trust update based on the model parameters from different local controllers and feeds back the updated model parameters to each local controller;
[0024] The local controller also publishes the trust prediction model to the collectors within its area for trust evaluation between collectors during the interaction process.
[0025] In the above step two, the deep reinforcement learning trust modeling method is as follows:
[0026] When a node in the network needs to select a neighbor node for data forwarding or query tasks, the current evaluator, i.e., the agent, based on its current state s, executes an action a through the policy π and transfers to a new state s ′ , and at the same time obtains the corresponding reward r for this process; the trust model aims to use the interaction experience between entities to predict the trustworthiness of the evaluation object, which is similar to giving corresponding actions according to the agent's state in reinforcement learning. Therefore, communication evidence energy evidence ε, data evidence These three types of trust evidence are defined as the state of the evaluator; each state is represented by a triple as where The action corresponds to the output of the trust model. The action is defined as the trust score given by the evaluator to the evaluation object, satisfying a ∈ [0, 1]; the reward in reinforcement learning refers to the feedback score given by the interaction environment to the action executed by the agent. The role of the reward is to guide the agent to gradually learn a behavior strategy adapted to the current interaction environment; the reward is defined as the cumulative deviation between the action and the trust evidence: where s[i] represents the i-th trust evidence in the state triple; the weights satisfy The policy in reinforcement learning represents a mapping function between states and actions, and the policy is a trust model deployed in the evaluator;
[0027] The reinforcement learning problem is not only faced with a continuous state space, that is, but also faced with a continuous action space, that is, a ∈ [0, 1]. Therefore, the deep deterministic policy gradient algorithm applicable to solving such problems is used to train the reinforcement learning model;
[0028] The model training is mainly divided into two stages: 1) model pre-training carried out in the global controller; 2) actual training carried out in each local controller; both training stages are based on the deep deterministic policy gradient algorithm framework;
[0029] After the network is initialized, the global controller performs model pre-training based on the virtual interaction environment; the working process of the virtual interaction environment includes five modules; the input a of the virtual interaction environment t represents the trust score of the evaluator for the evaluation object. At the same time, module 1 takes a t as the probability of determining whether the evaluator interacts with the evaluation object; module 2 initially randomly generates the absolute credibility of the evaluation object, and this value is equivalent to the probability that the evaluation object performs normal interaction behaviors; module 3 simulates the interaction between the evaluator and the evaluation object according to the results of the previous two modules, including updating attributes such as the number of successful / failed communications, the remaining energy of the node, and whether data is tampered with, etc.; then, the updated attributes are input into module 4 and the trust evidence ε, and the reward r are calculated; finally, module 5 combines the trust evidence into a state and outputs s t 、r t as well as s t+1 ; based on the deep deterministic policy gradient algorithm training framework and the virtual interaction environment, the global controller obtains a set of neural network parameters (θ, θ′, w, w′) after training converges, where θ represents the policy network weight, θ′ represents the target policy network weight, w represents the value network weight, and w′ represents the target value network weight; finally, the mean value of the parameters obtained after multiple training convergences is passed as the pre-training result to all local controllers;
[0030] After the local controller receives the pre-training parameters , it first distributes the trust evaluation model to the collectors in its area. Then, the collectors perform interactions based on the trust evaluation model and regularly send the interaction experiences (s t , a t , r t , s t+1)Send it to the affiliated local controller; finally, the local controller stores the interaction experiences from different collectors in the experience cache and performs local model training based on the deep deterministic policy gradient algorithm training framework using the method of small-batch random sampling.
[0031] In the above step 3, the method for updating the federated learning trust model is as follows:
[0032] After receiving the local model parameters from each local controller, the global controller uses the federated averaging method to perform global updates of the parameters and republishes the updated model parameters to all local controllers; the local controller replaces its local model parameters with the received global parameters and continues to update the model based on the actual interaction experiences; the above update process is iterated continuously, so as to ensure that the trust prediction model can continuously adapt to the dynamically changing network environment.
[0033] The network contains a unique global controller and m local controllers {LC 1 , LC 2 , …, LC m-1 , LC m}; every time after time T, all local controllers send their latest converged model parameters to the global controller; then, the global controller performs global model parameter updates: where θ T represents the global model parameters of the previous round, and η represents the soft update coefficient; finally, the global controller sends the updated model parameters θ T+1 to all local controllers; after receiving the global parameters, the local controller replaces its local model parameters with the global parameters; then, the local controller continues to update its local model parameters according to the interaction experiences from the collectors and repeats the above global model update process with a period of T, so that the trust prediction model in the network adapts to the dynamically changing underwater environment.
[0034] By adopting the above technical means, the beneficial effects of the present invention are as follows: for the problems of poor evidence quality, low detection accuracy, and high update cost that occur when the traditional trust model designed for homogeneous networks performs trust evaluation tasks in heterogeneous scenarios. By using deep reinforcement learning and federated learning, a trust management method for underwater sensor networks based on federated deep reinforcement learning is proposed. Using trust evidence as the state and trust model parameters as the action, respectively, as the input and output of deep reinforcement learning, so as to utilize its learning ability to achieve adaptive adjustment and parameter matching of the trust model; for the trust model parameters within each local sub-network, a cross-domain trust update framework is designed using federated learning to achieve cross-domain updates of the global model and periodic adjustments of the local model, while avoiding additional energy consumption caused by large-scale evidence sharing and further improving the evaluation accuracy of the trust management model in scenarios with spatio-temporal changes. Description of the Drawings
[0035] Figure 1 This is a schematic diagram of the cross - domain joint trust management framework of the present invention;
[0036] Figure 2 This is a schematic diagram of the trust model training framework based on DDPG of the present invention;
[0037] Figure 3 This is a schematic diagram of the virtual interaction environment module of the present invention;
[0038] Figure 4 This is a schematic diagram of the trust model update method based on federated learning of the present invention; Detailed Embodiment
[0039] The following further elaborates on the present invention in conjunction with the drawings and embodiments.
[0040] A trust management method for underwater sensor networks based on federated deep reinforcement learning, the steps of which include:
[0041] Step 1: Method for constructing a trust management framework
[0042] As shown in Figure 1 An underwater sensor network composed of various types of devices such as wave gliders, underwater gliders, AUVs, sub - sea flying nodes, and underwater moored buoy arrays. Each type of device is functionally divided into a global controller, a local controller, and a data collector, and the entire network is logically divided into a control layer and a data layer;
[0043] The control layer includes wave gliders deployed on the water surface and multiple AUVs underwater; the wave glider serves as the only global controller, responsible for updating the trust model and task scheduling, and at the same time receiving commands from the shore - based data center through satellite relay to adjust the tasks executed by the network; each AUV serves as a local controller, and manages the data - layer devices within its respective responsible area by receiving scheduling commands from the global controller; both the global controller and the local controller have sufficient computing, communication, and storage capabilities to support the trust management system;
[0044] The data layer includes various heterogeneous underwater data collection devices such as underwater gliders, underwater anchored nodes, and sub - sea flying nodes, collectively referred to as collectors. Each collector collaborates to execute tasks and generate interaction information; the underwater gliders, sub - sea flying nodes, and underwater anchored nodes in the data layer represent underwater devices with different mobility capabilities; the underwater glider represents a collector with relatively high mobility, capable of moving between areas covered by different local controllers; the sub - sea flying node represents a collector with limited mobility, which can only move within a local range and cannot move across regions; the underwater anchored node represents a fixed - position underwater collector;
[0045] The deployed trust management architecture is divided into three modules: trust evidence collection, trust modeling, and trust update. First, during the data collection process, the underwater collectors record the interaction information between devices, generate trust evidence, and send it to the subordinate local controllers. Then, the local controllers train trust models based on the obtained trust evidence and periodically send the model parameters to the global controller. Finally, the global controller performs joint trust update based on the model parameters from different local controllers and feeds back the updated model parameters to each local controller.
[0046] The local controllers also publish the trust prediction models to the collectors within their respective regions for trust evaluation between collectors during the interaction process.
[0047] Step 2: Deep Reinforcement Learning Trust Modeling Method
[0048] When nodes in the network need to select neighbor nodes for data forwarding and query tasks, the current evaluator, i.e., the agent, based on its current state s, executes an action a through the policy π and transfers to a new state s′, while obtaining the corresponding reward r for this process. The trust model aims to predict the trustworthiness of the evaluation object using the interaction experience between entities, which is similar to giving corresponding actions according to the agent's state in reinforcement learning. Therefore, the communication evidence energy evidence ε, and data evidence These three types of trust evidence are defined as the states of the evaluator. Each state is represented by a triple as where The action corresponds to the output of the trust model. The action is defined as the trust score given by the evaluator to the evaluation object, satisfying a ∈ [0, 1]. The reward in reinforcement learning refers to the feedback score given by the interaction environment to the action executed by the agent. The role of the reward is to guide the agent to gradually learn the behavioral strategy adapted to the current interaction environment. The reward is defined as the cumulative deviation between the action and the trust evidence: where s[i] represents the i-th trust evidence in the state triple; the weights satisfy The policy in reinforcement learning represents the mapping relationship function between the state and the action. The policy is the trust model deployed on the evaluator.
[0049] The reinforcement learning problem is not only oriented to a continuous state space, i.e., but also to a continuous action space, i.e., a ∈ [0, 1]. Therefore, the deep deterministic policy gradient algorithm suitable for solving such problems is used to train the reinforcement learning model.
[0050] The model training is mainly divided into two stages: 1) model pre-training in the global controller; 2) actual training in each local controller. Both training stages are based on as Figure 2Perform according to the shown deep deterministic policy gradient algorithm framework;
[0051] After the network is initialized, the global controller performs model pre-training based on the virtual interaction environment; as Figure 3 shown, the workflow of the virtual interaction environment includes five modules; the input a of the virtual interaction environment t represents the trust score of the evaluator for the evaluation object. At the same time, module 1 uses a t as the probability of determining whether the evaluator interacts with the evaluation object; module 2 initially randomly generates the absolute credibility of the evaluation object, which is equivalent to the probability of the evaluation object performing normal interaction behaviors; module 3 simulates the interaction between the evaluator and the evaluation object based on the results of the previous two modules, including updating attributes such as the number of successful / failed communications, the remaining energy of the node, and whether data is tampered with; then, the updated attributes are input into module 4 and the trust evidence ε, and the reward r are calculated; finally, module 5 combines the trust evidence into a state and outputs s t , r t and s t+1 ; based on the deep deterministic policy gradient algorithm training framework and the virtual interaction environment, the global controller obtains a set of neural network parameters (θ, θ′, w, w′) after training convergence, where θ represents the policy network weight, θ′ represents the target policy network weight, w represents the value network weight, and w′ represents the target value network weight; finally, the mean value of the parameters obtained after multiple training convergences is passed to all local controllers as the pre-training result;
[0052] After receiving the pre-training parameters , the local controller first distributes the trust evaluation model to the collectors in its area. Then, the collectors perform interactions based on the trust evaluation model and regularly send the interaction experiences (s t , a t , r t , s t+1 ) to the local controller they belong to; finally, the local controller stores the interaction experiences from different collectors in the experience cache and performs local model training based on the deep deterministic policy gradient algorithm training framework using the method of small-batch random sampling.
[0053] Step 3: Federal learning trust model update method
[0054] As Figure 4As shown, after receiving the local model parameters, the global controller uses the federated averaging method to globally update the parameters and republishes the updated model parameters to all local controllers; the local controllers replace their local model parameters with the received global parameters and continue to update the model based on the actual interaction experience; the above update process is iterated continuously, so as to ensure that the trust prediction model can continuously adapt to the dynamically changing network environment.
[0055] The network contains a unique global controller and m local controllers {LC 1 , LC 2 , …, LC m-1 , LC m}; every time after time T, all local controllers send their latest converged model parameters to the global controller; then, the global controller performs global model parameter update: where θ T represents the global model parameters of the previous round, and η represents the soft update coefficient; finally, the global controller sends the updated model parameters θ T+1 to all local controllers; after receiving the global parameters, the local controllers replace their local model parameters with the global parameters; then, the local controllers continue to update their local model parameters according to the interaction experience from the collector and repeat the above global model update process with a period of T, so that the trust prediction model in the network adapts to the dynamically changing underwater environment.
Claims
1. A trust management method for underwater sensor networks based on federated deep reinforcement learning, characterized in that: It includes the following steps: Step 1: Construction of the trust management framework Logically, the trust management framework divides the underwater sensor network into a control layer and a data layer; various devices in the network are divided into a global controller, local controllers, and data collectors according to their functions; the only global controller is responsible for the policy regulation of the entire network and the global trust update; each local controller trains and maintains a local trust model using the data aggregated from its respective area; while collecting underwater data, the data collectors transmit the interaction information between each other to the local controller to which they belong. Step 2: Trust modeling based on deep reinforcement learning Trust modeling based on deep reinforcement learning includes two stages: pre-training on the global controller and actual training on the local controllers; After the network is initialized, the global controller performs model pre-training based on a virtual interaction environment to make up for the lack of initial interaction experience; subsequently, the pre-trained model parameters are passed to the local controllers as the initialization parameters of the local models, and each local controller further trains the model according to the actual interaction data between the collectors within its respective area. Step 3: Update of the federated learning trust model Each local controller periodically sends the parameters of the local model to the global controller, instead of directly transmitting the historical interaction data, to achieve the decoupling of model update and data storage, thereby reducing the update cost while protecting local data privacy; after receiving the parameters of each local model, the global controller performs global update of the parameters using the federated averaging method and republishes the updated model parameters to all local controllers; the local controllers replace the local model parameters with the received global parameters and continue to update the model based on the actual interaction data; continuously iterate the above global update and model update processes so that the trust prediction model can continuously adapt to the dynamically changing underwater network environment. In the above Step 1, the method for constructing the trust management framework is as follows: The underwater sensor network consists of various types of devices such as wave gliders, underwater gliders, AUVs, seabed flying nodes, and underwater mooring arrays. Among them, various devices are divided into a global controller, local controllers, and data collectors according to their functions, and the entire network is logically divided into a control layer and a data layer; The deployed trust management architecture is divided into three modules: trust evidence collection, trust modeling, and trust update; first, during the data collection process, the underwater collectors record the interaction information between devices, generate trust evidence, and send it to the subordinate local controllers; then, the local controllers train a trust model based on the obtained trust evidence and periodically send the model parameters to the global controller; finally, the global controller performs joint trust update based on the model parameters from different local controllers and feeds back the updated model parameters to each local controller; The local controllers also publish the trust prediction model to the collectors within their respective areas for trust evaluation between collectors during the interaction process. In the above Step 2, the method for trust modeling based on deep reinforcement learning is as follows: When nodes in the network need to select neighbor nodes for data forwarding and query tasks, the current evaluator, i.e., the agent, based on its current state s, executes an action a through the policy π and transfers to a new state s ′ , and at the same time obtains the corresponding reward r; the trust model aims to use the interaction experience between entities to predict the trustworthiness of the evaluation object, which is similar to giving corresponding actions according to the agent's state in reinforcement learning. Therefore, the communication evidence energy evidence ε, data evidence These three types of trust evidence are defined as the state of the evaluator; each state is represented by a triple as where The action corresponds to the output of the trust model. The action is defined as the trust score given by the evaluator to the evaluation object, satisfying a ∈ [0, 1]; the reward in reinforcement learning refers to the feedback score given by the interaction environment to the action executed by the agent. The role of the reward is to guide the agent to gradually learn the behavioral strategy adapted to the current interaction environment; Define the reward as the cumulative deviation between the action and the trust evidence: where s[i] represents the i-th trust evidence in the state triple; the weights satisfy In reinforcement learning, the policy represents the mapping relationship function between the state and the action, and the policy is the trust model deployed in the evaluator; The reinforcement learning problem is not only oriented to a continuous state space, that is but also to a continuous action space, that is, a ∈ [0, 1]. Therefore, the deep deterministic policy gradient algorithm applicable to solving the reinforcement learning problem is used to train the reinforcement learning model; Model training is mainly divided into two stages: 1) Model pre-training conducted within the global controller; 2) Actual training conducted within each local controller; Both training stages are based on the deep deterministic policy gradient algorithm framework; After network initialization, the global controller performs model pre-training based on the virtual interaction environment; the working process of the virtual interaction environment includes five modules; the input a of the virtual interaction environment t represents the trust score of the evaluator for the evaluation object. At the same time, Module 1 uses a t as the probability of determining whether the evaluator interacts with the evaluation object; Module 2 initially randomly generates the absolute credibility of the evaluation object, and the absolute credibility is equivalent to the probability of the evaluation object performing normal interaction behaviors; Module 3 simulates the interaction between the evaluator and the evaluation object according to the results of the previous two modules, including updating attributes such as the number of successful / failed communications, the remaining energy of the node, and whether data is tampered with; then, the updated attributes are input into Module 4 and the trust evidence ε, and the reward r are calculated; finally, Module 5 combines the trust evidence into a state and outputs s t , r t and s t+1 ; based on the deep deterministic policy gradient algorithm training framework and the virtual interaction environment, the global controller obtains a set of neural network parameters (θ, θ ′ , w, w ′ ) after training convergence, where the vector θ represents the policy network weights, the vector θ ′ represents the target policy network weights, the vector w represents the value network weights, and the vector w ′ represents the target value network weights; finally, the mean values of various neural network parameters obtained after multiple training convergences are passed to all local controllers as the pre-training results; After the local controller receives the pre-trained parameters , it first distributes the trust evaluation model to the collectors in its area; then, the collectors interact based on the trust evaluation model and regularly send the interaction experiences (s t , a t , r t , s t+1 ) to the local controller to which they belong; finally, the local controller stores the interaction experiences from different collectors in the experience cache and performs local model training based on the deep deterministic policy gradient algorithm training framework using the method of mini-batch random sampling; In step three, the method for updating the federated learning trust model is as follows: After receiving the local model parameters, the global controller uses the federated averaging method to globally update the parameters and republishes the updated model parameters to all local controllers; The local controllers replace the local model parameters with the received global parameters and continue to update the model based on the actual interaction experience; The above global update and model update processes are continuously iterated to ensure that the trust prediction model can continuously adapt to the dynamically changing network environment; The network contains a unique global controller and m local controllers {LC 1 , LC 2 , …, LC m-1 , LC m}; every T time units, all local controllers send their most recently converged model parameters to the global controller; then, the global controller performs global model parameter update: where θ T represents the global model parameters of the previous round, and η represents the soft update coefficient; finally, the global controller sends the updated model parameters θ T+1 to all local controllers; after receiving the global parameters, the local controllers replace their local model parameters with the global parameters; then, the local controllers continue to update their local model parameters according to the interaction experience from the collector, and repeat the above global update and model update processes with a period of T, so that the trust prediction model in the network adapts to the dynamically changing underwater environment.
2. A trust management method for an underwater sensor network based on federated deep reinforcement learning according to claim 1, characterized in that: In step one, the control layer and the data layer are as follows: The control layer includes a wave glider deployed on the water surface and multiple AUVs underwater; The wave glider serves as the only global controller, responsible for updating the trust model and task scheduling, and at the same time receives commands from the shore-based data center through satellite relay to adjust the tasks executed by the network; Each AUV serves as a local controller, and manages the data layer devices within its respective responsible area by receiving scheduling commands from the global controller; Both the global controller and the local controllers have sufficient computing, communication, and storage capabilities to support the trust management system; The data layer includes various heterogeneous underwater data collection devices such as underwater gliders, underwater anchor nodes, and seabed flying nodes, collectively referred to as collectors, and each collector collaborates to execute tasks and generate interaction information; The underwater gliders, seabed flying nodes, and underwater anchor nodes in the data layer respectively represent underwater devices with different mobility capabilities; The underwater glider represents a collector with high mobility and can move between areas covered by different local controllers; The seabed flying node represents a collector with limited mobility and can only move within a local range and cannot move across regions; The underwater anchor node represents a fixed-position underwater collector.
Citation Information
Patent Citations
Model synchronization sharing method suitable for deep reinforcement learning agent
CN116384483A