Wireless sensor network coverage optimization method and system based on deep reinforcement learning

By reconstructing the wireless sensor network coverage optimization into the minimum vertex coverage problem, using the deep reinforcement learning method, using the Transformer encoder layer and global information fusion operation, the limitations of wireless sensor network topological characterization are solved, network coverage quality and energy efficiency are improved, and network life is extended.

CN120343566APending Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478459.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing deep reinforcement learning methods are difficult to capture long-range spatial correlations between nodes in wireless sensor network topological characterization, affecting network coverage quality and energy efficiency balance, especially in dynamic environments, and are difficult to meet real-time decision-making needs.

Method used

The wireless sensor network coverage optimization is reconstructed into the minimum vertex coverage problem, and a deep reinforcement learning method is adopted, using the Transformer encoder layer and global information fusion operation, combining the adjacency matrix mask, optimized state representation and decision-making process.

Benefits of technology

It improves the representation ability of the model under complex topology, shortens the convergence time, and improves the life and monitoring reliability of wireless sensor networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343566A_ABST
    Figure CN120343566A_ABST
Patent Text Reader

Abstract

The invention discloses a wireless sensor network coverage optimization method and system based on deep reinforcement learning, and the method comprises the steps: obtaining the state information of a wireless sensor network diagram, enabling the diagram to serve as the basic structure and state representation of a decision environment, enabling each node to represent a possible action, and enabling the node to be a possible motion; the wireless sensor network coverage is optimized and reconstructed into the minimum vertex coverage problem, state representation is simplified, the state space scale is reduced, the model convergence speed is increased, the search space is effectively reduced, the minimum vertex coverage problem is solved in combination with a deep learning model, an optimal decision is selected according to a Q value, and the network coverage is optimized and reconstructed. The limitation of a deep reinforcement learning method in the aspect of network topology representation is overcome, the problem of long-range space association between wireless sensor nodes can be more effectively captured, the service life of a wireless sensor network is prolonged, and the monitoring reliability of the wireless sensor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless sensor optimization, and relates to a method and system for optimizing the coverage of a wireless sensor network based on deep reinforcement learning. Background Art

[0002] The coverage optimization of wireless sensor networks (WSNs) is a key fundamental problem in the field of the Internet of Things. Its core goal is to minimize the network energy consumption through intelligent node scheduling while ensuring the coverage quality of the monitoring area. This problem essentially belongs to a combinatorial optimization problem in a dynamic environment and has important application values in fields such as environmental monitoring (e.g., disaster warning), smart cities (e.g., traffic perception), and industrial Internet of Things (e.g., equipment monitoring).

[0003] Traditional solutions are mainly divided into two categories: mathematical programming methods and heuristic strategies. Although mathematical programming methods can obtain theoretical optimal solutions, their computational complexity increases exponentially with the network scale and it is difficult to support real-time decision-making requirements. Heuristic strategies (such as greedy sleep scheduling and genetic algorithm optimization) can quickly generate feasible solutions, but they have defects such as incomplete elimination of coverage blind spots and severe energy consumption fluctuations, and cannot adapt to dynamic changes in network topologies. Especially in the scenario of mobile nodes, existing methods generally face bottlenecks such as policy lag and large repetitive calculation overheads. For example, in scenarios where the sudden monitoring demand surges or nodes fail abnormally, traditional methods often require a response time of several minutes and it is difficult to meet the stringent real-time requirements of critical tasks.

[0004] In recent years, deep reinforcement learning (DRL) has become a new paradigm for wireless sensor coverage optimization due to its environmental adaptability. Typical research realizes node sleep decision-making through the Q-learning framework, but there are two major limitations:

[0005] (1) Traditional graph neural networks (such as GCN) adopt a fixed neighborhood aggregation mechanism and are difficult to model long-range spatial associations between nodes (such as cross-regional collaborative coverage), resulting in insufficient representation ability of the policy network for complex topologies.

[0006] (2) Research attempts to introduce the Transformer architecture to enhance the global modeling ability, but its standard self-attention mechanism has problems with difficult training convergence.

[0007] Therefore, developing a new DRL framework that integrates dynamic graph neural networks and lightweight attention mechanisms has become an important research direction for breaking through the technical bottlenecks of wireless sensor network coverage optimization and has important application values for extending the lifespan of wireless sensor networks and improving monitoring reliability. Summary of the Invention

[0008] The object of the present invention is to solve the limitations of existing deep reinforcement learning methods in network topology representation, making it difficult to capture long-range spatial correlations between wireless sensor nodes, which affects the network coverage quality and energy efficiency balance in the dynamic environment of wireless sensors, and to provide a wireless sensor network coverage optimization method and system based on deep reinforcement learning.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] A wireless sensor network coverage optimization method based on deep reinforcement learning, comprising the following steps:

[0011] Obtain the state information of the wireless sensor network graph, where the state information of the wireless sensor network graph includes wireless sensor node information and inter-node communication link information;

[0012] Obtain a deep learning model, input the state information of the wireless sensor network graph into the deep learning model, output the expected reward Q value of each action in the wireless sensor network graph, select the action corresponding to the highest Q value as the optimal decision, and obtain the optimization result.

[0013] A further improvement of the present invention lies in:

[0014] The obtaining of the state information of the wireless sensor network graph includes:

[0015] Reconstruct the wireless sensor network coverage optimization into a minimum vertex cover problem, and establish the corresponding relationship between the network topology and the graph theory model, where the wireless sensor nodes are mapped to graph vertices, and the inter-node communication links are characterized as graph edges.

[0016] The deep learning model includes an input layer, multiple stacked Transformer encoder layers, a feature fusion layer, and an output layer.

[0017] The calculation process of the input layer includes:

[0018] The vertex feature matrix X ∈ R in the wireless sensor network graph di×N , is linearly transformed to generate an embedding vector matrix

[0019] The calculation process of the Transformer encoder layer includes:

[0020] Introduce an adjacency matrix masking operation, perform self-attention calculation based on the embedding vector matrix , obtain the output of self-attention, and obtain the local information between the nodes and their neighbors;

[0021] Introduce a global information fusion operation between multiple stacked Transformer encoder layers, concatenate the output of self-attention with a tensor of all 1s in the vertex dimension to obtain a fused feature matrix;

[0022] Input the fused feature matrix into two Transformer encoder layers to obtain an intermediate node representation matrix H temp ;

[0023] Input the intermediate node representation matrix H temp into two Transformer encoder layers without global information fusion operation to obtain a vertex representation matrix H final ;

[0024] Map the vertex representation matrix to the original feature space to obtain the Q value of the vertex action.

[0025] The introducing a global information fusion operation between multiple stacked Transformer encoder layers, concatenating the output of self-attention with a tensor of all 1s in the vertex dimension to obtain a fused feature matrix includes:

[0026] Let the output of the l-th Transformer encoder layer be matrix H (l) , and its elements are represented as where i represents the vertex index and j represents the feature dimension index;

[0027] Introduce a matrix of all 1s, 1, with the same dimension as H (l) , that is

[0028] Concatenate H (l) and 1 in the feature dimension to obtain a fused feature matrix whose dimension is

[0029] The intermediate node representation matrix H temp is obtained by the following formula:

[0030]

[0031] The vertex representation matrix H final is obtained by the following formula:

[0032] H final = TransformerEnc(H temp ).

[0033] where TransformerEnc is the functional representation of the Transformer encoder layer.

[0034] A wireless sensor network coverage optimization system based on deep reinforcement learning, comprising:

[0035] A wireless sensor network graph construction module, configured to obtain the state information of the wireless sensor network graph, where the state information of the wireless sensor network graph includes wireless sensor node information and inter-node communication link information;

[0036] An optimization module, configured to obtain a deep learning model, input the state information of the wireless sensor network graph into the deep learning model, output the expected reward Q value of each action in the wireless sensor network graph, select the action corresponding to the highest Q value as the optimal decision, and obtain the optimization result.

[0037] A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any method of the present invention when executing the computer program.

[0038] A computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any method of the present invention when executed by a processor.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The present invention discloses a wireless sensor network coverage optimization method based on deep reinforcement learning, which obtains the state information of the wireless sensor network graph. The graph serves as the basic structure and state representation of the decision-making environment, and each node represents a possible action. The wireless sensor network coverage optimization is reconstructed into a minimum vertex cover problem, which simplifies the state representation, reduces the state space scale, speeds up the model convergence speed, and effectively reduces the search space. Combining a deep learning model to solve the minimum vertex cover problem, and selecting the optimal decision according to the Q value, overcomes the limitations of the deep reinforcement learning method in network topology representation, can more effectively capture the problem of long-range spatial correlation between wireless sensor nodes, extends the lifespan of the wireless sensor network, and improves the monitoring reliability of the wireless sensor.

[0041] Furthermore, in the present invention, the deep learning model includes an input layer, multiple stacked Transformer encoder layers, a feature fusion layer, and an output layer, abandoning the explicit division of the traditional encoder-decoder, and adopting an end-to-end processing flow to more efficiently solve the wireless sensor network coverage optimization problem.

[0042] Furthermore, in the present invention, an adjacency matrix masking operation is introduced, enabling the neural network to fully extract the local information between vertices and their neighbors.

[0043] Furthermore, in the present invention, a global information fusion operation is introduced between multiple stacked Transformer encoder layers. The output of self-attention is concatenated with a tensor of all ones in the vertex dimension to obtain a fused feature matrix, which can capture the global information in graph data more comprehensively and further improve the model's representation ability for complex graph structures. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 : The overall framework of the DRL-MVC of the present invention;

[0046] Figure 2 : Schematic diagram of the deep neural network based on the improved Transformer of the present invention;

[0047] Figure 3 : Schematic diagram of the Transformer encoding layer of the present invention;

[0048] Figure 4 : Ablation experiment of the improved MDP of the present invention;

[0049] Figure 5 : Ablation experiment of the improved Transformer network of the present invention;

[0050] Figure 6 : Ablation experiment of the global information fusion operation of the present invention;

[0051] Figure 7 : Schematic diagram of solving the MDP instance of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0053] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0054] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings.

[0055] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is customarily placed during use, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0056] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0057] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set", "install", "connect", "couple" are understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0058] The present invention will be further described in detail below with reference to the accompanying drawings:

[0059] See Figures 1 to 7 , an embodiment of the present invention discloses a method for optimizing the coverage of a wireless sensor network based on deep reinforcement learning. The coverage optimization of the wireless sensor network is modeled as a minimum vertex cover problem, where network nodes correspond to graph vertices, communication links correspond to graph edges, and the coverage optimization objective is transformed into finding the minimum set of active nodes that can monitor all communication links.

[0060] First, by reconstructing the wireless sensor network coverage optimization as the Minimum Vertex Cover Problem (MVCP), the correspondence between the network topology and the graph theory model is established: where the sensor nodes are mapped to the graph vertices, the communication links between nodes are characterized as the graph edges, and the coverage optimization goal is transformed into finding the minimum set of active nodes that can monitor all communication links.

[0061] This embodiment proposes a DRL algorithm for solving MVCP, denoted as DRL-MVC. From the overall architecture, it can be divided into two stages: the offline training stage and the online inference stage. The two stages cooperate closely to jointly constitute the operation process of the entire DRL-MVC. See Figure 1 .

[0062] Specifically, it includes:

[0063] In the offline training stage:

[0064] The agent is placed in an environment composed of random graphs, which simulates the complexity and uncertainty of MVCP and provides the agent with rich interaction scenarios. Through continuous interaction with the environment, the agent collects data on vertex actions and their consequences. In each interaction, the agent will, according to the current vertex state s t and the optional action a t , use the neural network to predict the Q(s t , a t ) value (i.e., the expected reward). Subsequently, the agent selects an action based on these predictions and executes it, and then observes the reward r t fed back by the environment and the new graph state s t+1 , and stores the quadruple (s t , a t , r t , s t+1 ) into the experience pool. This process is repeated continuously to form a closed-loop learning cycle. In each learning cycle, the agent will update the parameters Θ of the neural network based on the quadruples (s t , a t , r t , s t+1 ) in the experience pool using the backpropagation algorithm. Through a large number of interactions and iterative learning, the decision-making strategy of the agent is gradually optimized until it reaches a relatively stable state. At this time, the agent has learned how to more effectively solve MVCP in the random graph environment.

[0065] In the online inference stage:

[0066] Instead of interacting with the environment to collect data or update the policy, the agent uses the trained neural network for prediction. Specifically, when the agent faces a new MVCP, it first obtains the current graph state information and inputs it into the trained neural network. The neural network calculates the Q-values of each optional action and outputs a sorted action list. The agent selects the action with the highest Q-value in this list as the current optimal decision. After executing this action, the agent continues to use the neural network for the next decision based on the change of the environmental state (i.e., the update of the graph). This process continues until the termination condition is met. Finally, the agent continuously updates and selects the best action sequence through this greedy strategy, thus efficiently solving the MVCP.

[0067] Step 1: Markov decision process modeling

[0068] When solving the MVCP modeled by the wireless sensor network optimization problem, the traditional Markov decision process (MDP) provides a general framework for solving multiple graph optimization problems, but it lacks pertinence and the training and solving efficiency is not high enough. Therefore, the present invention proposes an improved MDP to more effectively solve the MVCP. The specific modeling is as follows:

[0069] (1) State representation: the graph structure at different time steps;

[0070] (2) Action space: any vertex in the graph state;

[0071] (3) State transition: remove the action point and its isolated neighbor vertices, and retain the edges in the updated vertex set;

[0072] (4) Termination condition: there are no remaining vertices in the graph;

[0073] (5) Reward function: -1, indicating that a reward of -1 is obtained each time an action is executed;

[0074] (6) Discount factor: 1, indicating that future rewards have the same value as current rewards.

[0075] Furthermore, the dynamic graph structure state constructs the state space by maintaining a binary tuple of the real-time vertex set and the edge set, where the edge set always maintains a subset relationship with the original graph edge set.

[0076] Vertex elimination state transition mechanism: When executed, it satisfies that when a selected vertex is used as an action, the neighbor vertices with a degree of 1 are synchronously removed.

[0077] Progressive penalty reward function: By setting a single-step fixed negative reward, a linear association with the solution set size is established, so that the cumulative reward value directly corresponds to the cardinality of the vertex cover set.

[0078] Compared with the traditional MDP, the improved MDP shows the following advantages:

[0079] Simplification of state representation, directly using the current graph structure as the state, reducing redundant information;

[0080] By introducing a vertex removal mechanism, the scale of the state space is reduced;

[0081] By removing isolated neighbor vertices, not only the convergence speed of the model is accelerated, but also the search space is effectively reduced;

[0082] The termination condition based on vertex exhaustion intuitively reflects the solution process.

[0083] Improved Markov Decision Process (MDP): By strategies such as simplifying state representation, introducing a vertex removal mechanism, and removing isolated neighbor vertices, the scale of the state space is reduced, the convergence speed of the model is accelerated, and the search space is effectively reduced.

[0084] Step 2: Deep neural network based on improved Transformer

[0085] This embodiment proposes a deep neural network based on an improved Transformer, aiming to solve the problem of effectively predicting the vertex action Q-value in the graph state under the MDP framework of MVCP. The design of this neural network abandons the clear division of the traditional encoder-decoder and adopts an end-to-end processing flow, specifically including an input layer, multiple stacked Transformer encoder layers, a feature fusion layer, and an output layer. The overall framework is as Figure 2 shown.

[0086] First, considering the complexity and non-linear characteristics of graph data, this embodiment designs an input layer to map the original vertex features to a higher-dimensional embedding space. This layer provides richer input information for the subsequent processing layers and enhances the model's ability to capture complex features. The input layer achieves this goal through linear transformation. Specifically, the vertex feature matrix (d i is the dimension of vertex features, and N is the number of vertices in the graph) generates an embedding vector matrix (d e is the dimension of the embedding space) through linear transformation, as shown in the following formula:

[0087] H = WX + b,

[0088] where the weight matrix maps the features from d i dimension to d e dimension, and the bias vector is a column vector.

[0089] Secondly, this embodiment introduces a Transformer encoder layer as the core component of the network. The core of the Transformer encoding layer is the self-attention mechanism, which can effectively capture long-range dependencies in sequence data. In graph data, the complex relationships between vertices also need to be accurately modeled. Through the self-attention mechanism, the model can dynamically adjust the degree of attention to different vertices, thereby capturing the relationship information between vertices in the graph. The model is specifically as Figure 3 shown.

[0090] In the specific implementation, the model first calculates the query (Q), key (K), and value (V) matrices through the input feature matrix H. The formulas are as follows:

[0091] Q = W Q H, K = W K H, V = W V H.

[0092] By calculating the dot product of the query and the key, the attention scores are obtained and normalized by softmax, so that the sum of the attention weights of all keys corresponding to each query is 1:

[0093]

[0094] However, relying solely on the self-attention mechanism may not be able to fully utilize the local structural information in graph data. Therefore, this embodiment introduces an adjacency matrix masking operation to ensure that the model can also pay attention to the neighborhood features of vertices while capturing global features. Specifically, in this embodiment, the masking matrix M is combined with the adjacency matrix adj to adjust the attention calculation:

[0095]

[0096] Among them, the masking matrix M ensures that only adjacent vertex pairs contribute to the attention calculation, thereby enhancing the local perception ability of the model. The definition of the masking matrix M is as follows:

[0097]

[0098] In the above way, the masking matrix M ensures that when calculating the attention scores, the attention weights of non-adjacent vertex pairs will be set to a very small value, so as to ensure that the model mainly focuses on the relationships between adjacent vertices, and the following output of self-attention is obtained:

[0099] Z = A masked V.

[0100] In addition, the present invention uses a residual connection operation to avoid the problem of gradient disappearance and improve the learning ability of the model:

[0101] H' = H + Z.

[0102] To further enhance the non - linear fitting ability, in this embodiment, a feed - forward neural network (FFN) is introduced after the self - attention output. It consists of two fully - connected networks and uses the ReLU activation function for non - linear transformation:

[0103] FFN(H') = W2ReLU(W1H' + b1) + b2.

[0104] Furthermore, a residual connection is introduced again:

[0105] H” = H' + FFN(H').

[0106] After introducing the masking operation, the neural network can fully extract the local information between vertices and their neighbors, but still faces the challenge of extracting long - distance information. For this reason, a global information fusion operation is introduced between the Transformer encoder layers in this embodiment. Specifically, the output of the Transformer encoder layer is concatenated with a tensor of all - ones in the vertex dimension to introduce global information. This operation enables the model to capture the global information in the graph data more comprehensively, also integrates the attention mechanism, and thus removes the decoding process to solve the problem that the original Transformer model cannot effectively learn and solve the MVCP.

[0107] Let the output of the l - th Transformer encoder layer be the matrix H (l) , and its elements are denoted as where i represents the vertex index and j represents the feature dimension index. Then, a matrix of all - ones 1 is introduced, with the same dimension as H (l) , that is Concatenate H (l) and 1 in the feature dimension to obtain the fused feature matrix H fused , whose dimension becomes The concatenation operation can be expressed as:

[0108] H fused = [H (l) 1].[[]END]]

[0109] Input H fused to the next two Transformer encoder layers for further processing. The Transformer encoder layer can be represented as the function TransformerEnc, which encodes the input matrix and outputs the processed matrix. Therefore, after being processed by two additional Transformer encoder layers, the intermediate node representation matrix H temp is obtained:

[0110] H temp= TransformerEnc(TransformerEnc(H fused )).

[0111] Then, perform a dimensionality reduction operation on H temp to output H (4) : And input H (4) into two Transformer encoder layers without global information fusion operations to obtain the final vertex representation matrix H final :

[0112] H (4) = H temp [:,:,N:]

[0113] H final = TransformerEnc(TransformerEnc(H (4) )).

[0114] Finally, the output layer maps the final vertex representation back to the original feature space to generate the vertex action Q-value for MVCP decision-making. This step not only completes the complete mapping from input to output but also ensures that the generated Q-value can be directly used in the decision-making process of subsequent algorithms. The operation of the output layer can be expressed as:

[0115] Q = σ(W O H final + b O ),

[0116] where H O is the weight matrix of the output layer, b O is the bias term of the output layer, and σ is the activation function. In the present invention, the ReLU function is selected.

[0117] Deep neural network based on improved Transformer: Abandon the traditional encoder-decoder structure and adopt an end-to-end processing flow. Through the self-attention mechanism combined with the adjacency matrix masking operation, the model can capture both local and global information of the graph. In addition, a global information fusion operation is introduced to further improve the model's representation ability for complex graph structures.

[0118] The algorithm flow of the method disclosed in this embodiment is shown in Table 1:

[0119] Table 1 Algorithm Flow

[0120]

[0121]

[0122] In the implementation process of the training strategy and optimization method, the present invention follows the mature technical solutions under the traditional deep reinforcement learning framework, specifically including the implementation of the following six basic modules:

[0123] 1) Loss function construction module: The mean squared error is used as the loss function for neural network training, and parameter optimization is carried out by calculating the difference between the predicted Q value and the target Q value;

[0124] 2) Delayed reward estimation module: An n-step delay mechanism is introduced to improve the deep Q network, and the evaluation of immediate rewards and long-term benefits is balanced through the multi-step temporal difference method;

[0125] 3) Sample reuse module: An experience replay buffer pool is configured, and the temporal correlation between samples is broken by randomly sampling historical experience data to improve the data utilization efficiency;

[0126] 4) Network update module: An independent target network is constructed as the value estimation benchmark, and the stable control of the training process is realized through the periodic parameter synchronization mechanism;

[0127] 5) Parameter optimization module: The Adam adaptive moment estimation algorithm is used to iteratively update the weights of the neural network, and the convergence efficiency is improved by adjusting the adaptive learning rate;

[0128] 6) Exploration strategy module: The ε-greedy action selection strategy is implemented, and the balance between exploration and exploitation is dynamically adjusted during the model training process, where the ε value is annealed according to the preset decay plan.

[0129] It should be particularly noted that the above six technical modules all belong to the well-known technical means in the field of deep reinforcement learning, and their specific implementation details follow the conventional technical specifications in this field.

[0130] Furthermore, the technical effects of the present invention are verified through the following comparative experiments and ablation experiments:

[0131] Comparative experiment

[0132] To verify the generalization performance advantage of the algorithm (DRL-MVC) of the present invention when expanding the graph scale, the representative deep reinforcement learning algorithm S2V-DQN in the field of the MVCP problem is selected as the benchmark algorithm.

[0133] The experiment uses the original dataset in reference [1], and the comparability is ensured through unified experimental settings:

[0134] Training set: ER random graph (connection probability P = 0.15), divided into two scale groups:

[0135] 1) Basic group: The number of vertices N = 15 - 20;

[0136] 2) Medium-scale group: N = 50 - 100.

[0137] Test set: Large-scale graphs with 100 - 600 vertices.

[0138] Benchmark solution: The optimal solution obtained by using an exact solver after 1 hour of calculation.

[0139] The data in Table 2 shows that DRL-MVC exhibits better approximation ratio performance on various test sets. Specifically:

[0140] Training model for the basic group: On the test set with 500 - 600 vertices, the approximation ratio of DRL-MVC (1.0099) is 1.7% higher than that of S2V-DQN (1.0276);

[0141] Training model for the medium-scale group: On the test set with 200 - 300 vertices, DRL-MVC (1.0123) achieves a 4.3% performance improvement compared to the benchmark algorithm (1.0570).

[0142] Table 2 Comparison of generalization performance (approximation ratio)

[0143] 15-20 50-100 100-200 200-300 300-400 400-500 500-600 S2V-DQN (basic group) 1.0032 1.0941 1.0710 1.0484 1.0365 1.0276 1.0246 DRL-MVC (basic group) 1.0028 1.0195 1.0236 1.0258 1.0184 1.0099 1.0106 S2V-DQN (medium-scale group) - 1.0079 1.0304 1.0570 1.0532 1.0463 1.0427 DRL-MVC (medium-scale group) - 1.0065 1.0093 1.0123 1.0211 1.0103 1.0169

[0144] Design a special comparative experiment for the ability to handle long-range dependencies in the graph structure.

[0145] Long-range graph generation method:

[0146] 1) Construct a basic linear graph: containing L ordered vertices (L ≤ N);

[0147] 2) Add random edges: Supplement the edge set with a probability of P = 0.15.

[0148] Experiment configuration:

[0149] 1) Total number of vertices N = 200;

[0150] 2) Long-range parameter L ∈ {10, 25, 50, 100};

[0151] 3) Use the PyTorch framework to reproduce the benchmark algorithm.

[0152] As shown in Table 3, when the long-range parameter L ≥ 50, DRL-MVC has the following advantages compared to S2V-DQN:

[0153] 1) When L = 50, the approximation ratio is increased by 15.3% (1.0658 vs 1.2699);

[0154] 2) When L = 100, the performance advantage is expanded to 17.8% (1.1158 vs 1.3585).

[0155] This verifies the robustness of the present invention in complex topological structures.

[0156] Table 3 Comparison of long-distance graph performance (approximation ratio)

[0157] 10 25 50 100 S2V-DQN 1.0241 1.0536 1.2699 1.3585 DRL-MVC 1.0082 1.0122 1.0658 1.1158

[0158] (2) Ablation experiment

[0159] In this embodiment, ablation experiments are carried out on the proposed algorithm to evaluate the effectiveness of its innovative design. All ablation experiments are trained and tested on an ER random network with the number of vertices N = 50 and the edge generation probability P = 0.5.

[0160] First, this embodiment conducts an ablation experiment on the improved MDP, comparing the training situations of the traditional MDP model and the improved MDP model, as Figure 4 shown. The results show that the improved MDP significantly speeds up the training speed, and its convergent approximation value is slightly lower than that of the traditional MDP. The improved MDP reduces the scale of the state space through the vertex and edge deletion strategy, reducing the number of states processed and updated in the DQN model, thereby reducing the computational amount. In addition, the improved MDP makes each step of the decision-making more explicit and simple. The agent only needs to consider which vertex or edge to remove under the current graph structure, rather than considering all vertex marking combinations like the traditional marking point strategy. This simplified decision-making process reduces the decision complexity, enables the DQN algorithm to explore the strategy space more efficiently, find a better strategy combination, and thus makes the convergent approximation value slightly lower than that of the traditional MDP.

[0161] Secondly, this embodiment conducts an ablation experiment on the deep neural network design based on Transformer. The original Transformer model refers to the Transformer network with a masked adjacency matrix added to the network, including multiple layers of encoders and decoders, as Figure 5 shown. The experimental results show that directly applying the original Transformer network with a masked adjacency matrix cannot learn an effective strategy for solving MVCP, and its average solution approximation ratio is only about 1.4, much lower than about 1.01 of the improved Transformer model. This shows that although the introduction of the masked adjacency matrix can provide certain structural information, if it is not properly optimized and adjusted, it will instead limit the learning ability of the model and cause it to be unable to effectively capture the key features of the problem.

[0162] Finally, this embodiment conducts an ablation experiment on the global information fusion operation of adding a matrix of all 1s to capture global information, as Figure 6As shown. The results show that the addition of the global information fusion operation can significantly optimize the training effect, and the average solution approximation ratio is increased by about 0.13. This indicates that the global information fusion operation can effectively integrate the global feature information in the network, enhance the model's understanding and grasp of the overall structure of the problem, and thus improve the model's solution performance.

[0163] The present invention verifies its effectiveness through comparative experiments and ablation experiments. Compared with the existing representative algorithm S2V-DQN, the present invention shows significant advantages in the generalization performance of large-scale graphs, especially when dealing with long-range dependencies in graphs, and the performance improvement is more obvious. In addition, the results of the ablation experiments show that the improved MDP and the optimized Transformer network design play a key role in improving the model performance.

[0164] Furthermore, in this embodiment, a specific graph instance is constructed, and the value of the selected covering point is calculated using a deep neural network in the state to make a decision. As Figure 7 shown, in this embodiment, a connected graph containing 5 nodes is constructed as the decision-making environment. This MDP constructs a covering solution set through multiple selections, and the selected points at each decision-making stage are highlighted in yellow, thus intuitively showing the decision-making trajectory during the state transition until an empty graph state appears, indicating the end of the process.

[0165] The specific steps are as follows:

[0166] (1) Graph instance construction: First, a connected graph containing 5 nodes is constructed. This graph serves as the basic structure and state representation of the decision-making environment, and each node represents a possible action.

[0167] (2) Deep neural network value calculation: At each decision-making stage, the designed deep neural network of the present invention is used to calculate the values of selecting each point in the current state. The value represents the expected cumulative reward that can be obtained by selecting a certain covering point in the current state.

[0168] (3) Decision-making stage: At each decision-making stage, according to the calculated values, the covering point with the highest value is selected as the current decision. This selected point is Figure 7 highlighted in yellow in

[0169] (4) State transition: After selecting a covering point, the system state will be transferred according to the selected covering point. The new state will be used as the input for the next decision-making stage, and the above calculations and selection processes are repeated.

[0170] (5) Process termination: When there are no longer uncovered nodes in the graph, that is, an empty graph state appears, it indicates that the covering solution set has been constructed and the process ends.

[0171] Through the above steps, this embodiment demonstrates how to construct a coverage solution set in a connected graph using a deep neural network and an MDP process, and visually displays the decision-making trajectory through visualization means. The specific calculation formula of the deep neural network is shown in detail in the previous text.

[0172] This embodiment also discloses a wireless sensor network coverage optimization system based on deep reinforcement learning, including:

[0173] A wireless sensor network graph construction module, configured to obtain the state information of the wireless sensor network graph, where the state information of the wireless sensor network graph includes wireless sensor node information and inter-node communication link information;

[0174] An optimization module, configured to obtain a deep learning model, input the state information of the wireless sensor network graph into the deep learning model, output the expected reward Q value of each action in the wireless sensor network graph, select the action corresponding to the highest Q value as the optimal decision, and obtain the optimization result.

[0175] The schematic diagram of the terminal device provided by an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above various method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above various device embodiments are implemented.

[0176] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention.

[0177] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.

[0178] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0179] The memory can be used to store the computer program and / or module. By running or executing the computer program and / or module stored in the memory, and invoking the data stored in the memory, the processor implements various functions of the terminal device.

[0180] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0181] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A wireless sensor network coverage optimization method based on deep reinforcement learning, characterized in that It includes the following steps: Obtain the status information of the wireless sensor network graph, where the status information of the wireless sensor network graph includes wireless sensor node information and inter-node communication link information; Obtain a deep learning model, input the status information of the wireless sensor network graph into the deep learning model, output the expected reward Q value of each action in the wireless sensor network graph, select the action corresponding to the highest Q value as the optimal decision, and obtain the optimization result.

2. The method for optimizing the coverage of a wireless sensor network based on deep reinforcement learning according to claim 1, wherein The obtaining of the status information of the wireless sensor network graph includes: Reconstruct the wireless sensor network coverage optimization into a minimum vertex cover problem, and establish the corresponding relationship between the network topology and the graph theory model, where the wireless sensor nodes are mapped to graph vertices, and the inter-node communication links are characterized as graph edges.

3. A method for optimizing the coverage of a wireless sensor network based on deep reinforcement learning according to claim 1, wherein The deep learning model includes an input layer, multiple stacked Transformer encoder layers, a feature fusion layer, and an output layer.

4. A method for optimizing the coverage of a wireless sensor network based on deep reinforcement learning according to claim 3, characterized in that, The calculation process of the input layer includes: The vertex feature matrix in the wireless sensor network graph is linearly transformed to generate an embedded vector matrix 5. A method for optimizing the coverage of a wireless sensor network based on deep reinforcement learning according to claim 4, characterized in that, The calculation process of the Transformer encoder layer includes: The adjacency matrix masking operation is introduced, based on the embedding vector matrix Self-attention calculation is performed to obtain the output of self-attention and acquire the local information between nodes and their neighbors; Introduce a global information fusion operation between multiple stacked Transformer encoder layers, splice the output of the self-attention with a tensor of all 1s in the vertex dimension to obtain a fused feature matrix; Input the fused feature matrix into two Transformer encoder layers to obtain the intermediate node representation matrix H temp ; Input the intermediate node representation matrix H temp into two Transformer encoder layers without global information fusion operations to obtain the vertex representation matrix H final ; Map the vertex representation matrix to the original feature space to obtain the Q value of the vertex action.

6. The method for optimizing the coverage of a wireless sensor network based on deep reinforcement learning according to claim 5, characterized in that, The introducing of the global information fusion operation between multiple stacked Transformer encoder layers, splicing the output of the self-attention with a tensor of all 1s in the vertex dimension to obtain a fused feature matrix includes: Let the output of the $l$-th Transformer encoder layer be the matrix $\mathbf{H}$ (l) , and its elements are denoted as where $i$ represents the vertex index and $j$ represents the feature dimension index; Introduce a matrix of all 1s, with the same dimension as H (l) That is Concatenate H (l) with 1 in the feature dimension to obtain the fused feature matrix whose dimension is 7. A method for optimizing the coverage of a wireless sensor network based on deep reinforcement learning according to claim 5, characterized in that, The intermediate node represents the matrix H temp Obtained by the following formula: The vertex representation matrix H final is obtained by the following formula: H final = TransformerEnc(H temp ). where TransformerEnc is the function representation of the Transformer encoder layer.

8. A wireless sensor network coverage optimization system based on deep reinforcement learning, characterized in that, It includes: A wireless sensor network graph construction module for obtaining the status information of the wireless sensor network graph, where the status information of the wireless sensor network graph includes wireless sensor node information and inter-node communication link information; An optimization module for obtaining a deep learning model, inputting the status information of the wireless sensor network graph into the deep learning model, outputting the expected reward Q value of each action in the wireless sensor network graph, selecting the action corresponding to the highest Q value as the optimal decision, and obtaining the optimization result.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.