Underwater dynamic network resource scheduling and time slot conflict optimization method
By adopting the Q-learning reinforcement learning method in the underwater wireless sensor network, the transmission priority and time slot allocation of nodes are optimized, and the problems of long propagation delay and dynamic topological changes in the underwater network are solved, channel utilization and transmission efficiency are improved, and energy consumption is reduced.
Patent Information
- Application Number
- CN202510680906.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art is difficult to effectively solve the problems of long propagation delay, dynamic topological changes and limited node energy consumption in underwater wireless sensor networks, resulting in low channel utilization and frequent communication conflicts.
Using a reinforcement learning method based on Q-learning, through cluster head election, cluster member allocation, node state feature modeling, decision matrix initialization and multi-factor collaborative reward function, node transmission priority and time slot allocation are optimized, resource conflicts are avoided, and resource allocation is achieved.
It significantly improves the utilization rate and transmission efficiency of underwater channels, reduces node conflicts and energy consumption, and makes the network more stable in dynamic environments.
Smart Images

Figure CN120456110A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-sensor resource allocation problems, and in particular relates to an underwater dynamic network resource scheduling and time slot conflict optimization method. Background Art
[0002] With the growing demand for underwater wireless sensor networks in fields such as ocean exploration, resource development, environmental monitoring, and military applications, designing efficient underwater communication protocols has become a research priority and a challenge. Underwater communication primarily relies on acoustic propagation, which is slow underwater, far slower than the propagation speed of electromagnetic waves in air. This leads to long propagation delays in underwater communication networks. Furthermore, acoustic propagation is severely affected by multipath effects, frequency-selective attenuation, and environmental noise. Unlike terrestrial communications, underwater communication propagation delay exhibits significant spatiotemporal characteristics. Communication delays between nodes are not only related to physical distance but also influenced by node transmission times and dynamic changes in network topology. This propagation characteristic significantly increases scheduling complexity, making traditional discrete-time slot-based MAC protocols inadequate for underwater communication. In underwater wireless sensor networks, the uncertainty of packet arrival times can lead to more frequent communication collisions between nodes, especially in large-scale networks, where the cumulative effect of propagation delays is more pronounced. Furthermore, underwater nodes are typically energy-constrained. Reducing collisions while improving channel utilization and minimizing node energy consumption are key challenges in designing underwater MAC protocols. Furthermore, the dynamic nature of underwater environments presents new challenges for network scheduling. Factors such as currents and tides in the ocean can cause dynamic changes in network topology, leading to constant fluctuations in the relative positions of nodes and the quality of communication links. This makes communication protocols based on fixed topologies difficult to adapt to real-world scenarios. Furthermore, the low bandwidth and high latency of acoustic communication further limit network communication efficiency. The uneven resource allocation and low time slot utilization found in traditional MAC protocols are particularly pronounced in underwater scenarios. These characteristics reduce underwater channel utilization, necessitating the development of an efficient protocol that places higher demands on network scheduling and resource allocation.
[0003] To address communication conflicts caused by long propagation delays in underwater wireless sensor networks (UWSNs), the paper "UW-ALOHA-Q: A Reinforcement Learning-Based MAC Protocol for Underwater Acoustic Sensor Networks" proposed a distributed MAC protocol based on reinforcement learning. This protocol utilizes asynchronous operation and a random backoff mechanism to improve channel utilization in dynamic environments. However, its learning convergence time is long, making it difficult to quickly adapt to drastic changes in the network environment. To address this issue, the paper "Reinforcement Learning-Based MAC Protocol UW-ALOHA-QM for Mobile Underwater Acoustic Sensor Networks" combines reinforcement learning with a focus on the dynamic adaptability of mobile node environments. It improves channel utilization efficiency through asynchronous distributed scheduling, but its performance is limited by the long convergence time in the initial learning phase. To address data priority scheduling in dynamic environments, the paper "AQ-Learning and Data Importance Rating-Based MAC Protocol for Dynamic Clustering Underwater Acoustic Networks" combines Q-learning with data importance rating (DIR) to implement priority-based data transmission scheduling. However, this method relies too much on the data importance evaluation mechanism and is not adaptable enough to the overall network load and conflict situations.
[0004] However, the above method only considers the transmission and reception issues between nodes, and does not consider the application in the inter-cluster environment, and the underwater background environment is not comprehensive enough. Therefore, "Cluster-Based Spatial-Temporal MAC Scheduling Protocol for Underwater Sensor Networks" proposes a MAC scheduling method that reduces the spatiotemporal conflicts between nodes by constructing a spatiotemporal conflict graph (ST-CG) and an intra-cluster priority scheduling mechanism, which significantly improves channel utilization. However, this method requires accurate spatiotemporal location information, and the complex construction process of the conflict graph may increase the computational overhead in a dynamic network. To solve this problem, the literature
[0005] The paper "MR-SFAMA-Q: A MAC Protocol Based on Q-Learning for Underwater Acoustic Sensor Networks" introduces the Q-learning algorithm to update data transmission scheduling strategies in real time in dynamic environments, thereby reducing network latency and packet loss. However, it fails to fully address the issue of channel resource utilization efficiency. The paper "An Adaptive MAC Protocol for Underwater Acoustic Networks Based on Deep Reinforcement Learning" proposes a MAC protocol based on a dual deep Q network (D3QN). By utilizing a delayed reward mechanism and short-term cumulative reward optimization, it significantly improves throughput and energy efficiency in dynamic environments. However, the algorithm's high complexity may limit its use in scenarios with high real-time requirements.
[0006] In summary, although the above studies have respectively solved the problems of long propagation delay, energy consumption optimization, dynamic scheduling and conflict in underwater networks, they are all separate studies and fail to comprehensively consider the multiple problems of long propagation delay, dynamic environmental changes and limited node energy consumption. Summary of the Invention
[0007] The present invention discloses an underwater dynamic network resource scheduling and time slot conflict optimization method, which can reduce the collision rate of data packets between nodes and improve the utilization rate of underwater node channels.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] The method for scheduling underwater dynamic network resources and optimizing time slot conflicts includes the following steps:
[0010] Step 1: Cluster head election and cluster member allocation, specifically:
[0011] Cluster head candidate node n is calculated based on the residual energy E remaining The node broadcasts its status and location information. Ordinary nodes select the nearest cluster head based on the received broadcast information to join their cluster. During the election process, candidate cluster heads broadcast their remaining energy and location information, allowing other nodes to determine their suitability for the position. When selecting a cluster head, nodes prioritize those with a closer proximity and more sufficient remaining energy, achieving more efficient resource allocation.
[0012] Step 2: Broadcasting node status feature information within the cluster and modeling conflict topology. Specifically:
[0013] Each cluster head broadcasts the status information of its cluster members and calculates the status characteristics of each node, including the number of conflicting edges E, the number of conflicting vertices V, the distance between the node and the cluster head D, and the node's transmission radius R. Through these characteristics, a state matrix can be constructed for each node as the basis for subsequent optimization priority allocation.
[0014] Step 3: Initialize the intra-cluster topology decision matrix and define the action space. Specifically:
[0015] Define the decision matrix Q(s t ,a t ), where s t Indicates the state of the node, a t Represents the priority assigned to the node. Each node's decision matrix is initialized using the number of cluster heads as prior information, accelerating the convergence of the subsequent reinforcement learning process. The decision matrix records the rewards obtained by the node for taking different actions in different states, providing a basis for the subsequent learning process.
[0016] Step 4: Construct a multi-factor collaborative reward function and iteratively update the Q value. Specifically:
[0017] A reward function P is constructed based on the node allocation information within the cluster. A priority formula is used to calculate the current node state characteristics, determine whether a conflict has occurred, and determine the positive reward generated after successful data transmission. Using Q-learning, the Q value is updated to optimize the node priority allocation strategy. The reward function reflects the effectiveness of the priority actions taken by each node in a specific state, prompting the system to converge towards the optimal solution.
[0018] Step 5: Priority conflict detection and dynamic allocation mechanism, specifically:
[0019] Based on the priority assignments for nodes within the cluster, the system checks whether multiple nodes have selected the same priority. If a conflict is detected, the action for the conflicting node is reselected to ensure that the priority of each node in the cluster is independent. If there is no conflict, the current assignment result is recorded and used for reward calculation. This step ensures the independence of node priorities within the cluster and effectively avoids resource conflicts.
[0020] Step 6: Final time slot allocation strategy extraction and global scheduling implementation, specifically:
[0021] After the Q-value iterations are complete, the optimal action a is extracted from the decision matrix for each node, which serves as the node's final priority and time slot allocation. By extracting the optimal priority and time slot allocation from the decision matrix, each node in the network can be scheduled according to the optimal strategy, achieving efficient communication and resource utilization.
[0022] The remaining energy E of the underwater node in step 1remaining Calculated for the following model:
[0023] E remaining =E initial -(E transmit +E receive +E process +E idle +
[0024] E sensing )(1) Among them, E initial is the initial energy, and the rest are related energy consumption parameters such as sending energy consumption, receiving energy consumption, idle energy consumption, processing energy consumption, etc. transmit =P tx,elec ·t tx +∈·d a ·N
[0025] Among them, P tx,elec ·t tx is the circuit power consumption, ∈·d a N is the signal amplifier number, P tx,elec is the transmitting circuit power, t tx is the transmission time, N is the data volume, ∈ is the amplifier energy efficiency coefficient, d a is the transmission distance under path loss;
[0026] E receive =P rx,elec ·t rx
[0027] Among them, P rx,elec is the receiving circuit power, t rx Receiving time
[0028] E sensing =P sensor ·t sensing
[0029] Among them, P sensor Sampling power for the sensor
[0030] E idle =P idle ·t idle
[0031] Among them, P idle is the idle state power, t idle The total time the node was in idle mode
[0032] E process =E bit ·N
[0033] Among them, E bit is the processing energy consumption per bit of data.
[0034] The cluster head candidates in step 1 are calculated using the following method:
[0035]
[0036] Among them, E rrmaining is the residual energy of candidate node n, E max is the maximum residual energy of the node in the network, d, H is the distance between node i and cluster head H, w1, w2 are the weight factors of energy and distance, which are used to adjust the influence of energy and distance in cluster head election. Cluster head candidate node n broadcasts its priority P to the network CH and node ID, the ordinary node selects the candidate node n with the highest priority as the cluster head.
[0037] In step 1, cluster member assignment is achieved by selecting the nearest cluster head from the common node, and the distance to each cluster head is calculated according to the following formula:
[0038]
[0039] Among them, (x i ,y i ) is the position coordinate of node i, (x H ,y H ) is the location coordinate of cluster head H. Ordinary nodes select d i,H The smallest cluster head H joins the cluster and updates the cluster member information recorded by the cluster head.
[0040] In step 1, in order to optimize the intra-cluster communication, the cluster head divides the clusters according to the location and number of the cluster members. i,j Calculated according to the following formula:
[0041]
[0042] Among them, d i,j The cluster head divides the cluster members into multiple communication groups according to the distance between nodes and records the node information of each group.
[0043] In step 2, in order to make all nodes in the underwater network aware of the information of the corresponding neighbor nodes, the nodes broadcast each other's energy, location and status information.
[0044] In step 3, after each node receives the broadcast, a decision matrix will be established for subsequent Q-learning algorithm calculations.
[0045] In step 4, the priority score P of each node is calculated to guide the reward distribution in the time slot allocation process. The reward is calculated according to the following formula:
[0046]
[0047] Among them, E is the number of conflict edges used to characterize the connection or influence of the node with other nodes, V is the number of conflict vertices used to characterize the degree of influence of the node, D is the distance of the node, R is the transmission radius of the node, n is the number of cluster heads perceived by the node within the transmission radius, and w is the weight coefficient used to adjust the impact of the number of perceived cluster heads n on the priority. 2 (EV) represents the impact of a node conflict. The relationship between the number of edges (E) and the number of vertices (V) is used to measure the conflict priority of a node. The greater the number of edges, the higher the priority. If the number of vertices is close to the number of edges, the conflict impact is greater and the priority is lower. is the distance ratio of the node. The larger the distance D, the lower the priority, but the larger the transmission radius, the higher the priority. w·log(n+1) is the impact of cluster head perception. The greater the number of cluster heads n, the higher the priority and the more likely they are to affect important target nodes.
[0048] In step 4, after obtaining the reward R, the decision matrix is updated using the Q-learning formula, which is calculated according to the following formula:
[0049]
[0050] Among them, Q(s t ,a t ) is the Q value of the current node, α is the learning rate, and Rt is the reward value of the current action. is the maximum Q value in the next state. γ is the discount factor that determines the importance of future rewards. t Afterwards, the decision matrix is updated through the Q-learning formula so that the updated Q(s t ,a t ) is closer to the true action value. After all training rounds are completed, the values in the decision matrix represent the cumulative rewards of each node under different priority assignments. Finally, the optimal action of each node can be extracted from the decision matrix.
[0051] During the priority allocation process of nodes within the cluster in step 5, it is necessary to detect whether there is a priority conflict and adjust the allocation plan. According to the action results learned by the current node in the decision matrix, the priority of each node in the cluster is preliminarily allocated to ensure that all nodes obtain the preliminary allocation results.
[0052] Define the node set in the cluster as The node priority assignment result is If there are any two different nodes Satisfaction: a i =a j(i≠j), it is determined to be a priority conflict.
[0053] If a conflict is detected, the conflict node set Reselect action:
[0054]
[0055] in, is the available priority set, ∑ j≠i δ(a,a j ) is the conflict indicator function:
[0056]
[0057] Among them, R is the total reward value of the current round, For node i in state s i Next select action a i The Q value of β is the conflict avoidance reward weight.
[0058] By minimizing the number of conflicts, the adjusted priorities are ensured to be independent. If there is no conflict in the allocation, the current allocation plan is recorded and the reward function is updated, where β is the conflict avoidance reward weight, which is used for reinforcement learning policy optimization.
[0059] Repeat the above steps to ensure the uniqueness and rationality of the priority allocation scheme within the entire cluster, and gradually improve the allocation results through training.
[0060] In step 6, the cluster head node allocates communication time slots and priorities according to the following formula:
[0061] For each cluster head node Extracting actions from the decision matrix
[0062]
[0063] Among them, Q(s i ,a) is the node status s i The Q value of action a. Mapping to communication time slots Define the mapping function, and the function f must satisfy the time slot uniqueness constraint:
[0064]
[0065] The cluster head node will globally distribute the results Broadcast to the member nodes in the cluster to ensure that each node schedules communication according to the allocated time slot.
[0066] Beneficial effects of the present invention:
[0067] Through the above technical solution, the present invention proposes an adaptive MAC protocol optimization method based on reinforcement learning to address the problems of long propagation delay, dynamic topology changes and serious channel conflicts in underwater wireless sensor networks.
[0068] Firstly, by comprehensively considering key factors such as propagation delay, node energy consumption, and channel conflicts in underwater networks, a reinforcement learning model combining node status and network environment was constructed.
[0069] Secondly, a MAC protocol optimization problem with the goal of maximizing channel utilization and minimizing energy consumption is proposed, and dynamic data scheduling is combined to achieve efficient resource allocation among multiple nodes.
[0070] Finally, by introducing the Q-learning algorithm, combining time asynchronous operation and random backoff mechanism, the node transmission scheduling strategy is gradually optimized, and an optimized MAC protocol implementation scheme is obtained.
[0071] The present invention comprehensively considers multiple influencing factors such as propagation delay, node energy consumption and channel conflict in underwater wireless networks, significantly improves channel utilization and transmission efficiency, reduces node conflict and energy consumption, makes the network performance more stable in dynamic environments, and provides a new solution for efficient resource allocation in underwater communications. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0073] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0074] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0075] like Figure 1 As shown: The underwater dynamic network resource scheduling and time slot conflict optimization method of the present invention includes the following steps:
[0076] Step 1: Cluster head election and cluster member allocation, cluster head candidate node n is based on the residual energy E remainingThe node broadcasts its own status and location information. Ordinary nodes select the nearest cluster head to join their cluster based on the received broadcast information. During the election process, the candidate cluster head node broadcasts its own residual energy and location information to allow other nodes to determine whether it is suitable to become a cluster head. When selecting a cluster head, the node gives priority to cluster heads with closer distances and more sufficient residual energy to achieve more efficient resource allocation. Specifically:
[0077] The remaining energy E of the underwater node in step 1 remaining Calculated for the following model:
[0078] E remaining =E initial -(E transmit +...+E sensing ) (1)
[0079] Among them, E initial is the initial energy, and the rest are related energy consumption parameters such as sending energy consumption, receiving energy consumption, idle energy consumption, and processing energy consumption.
[0080] The cluster head candidates in step 1 are calculated using the following method:
[0081]
[0082] Among them, E remaining is the residual energy of candidate node n, E max is the maximum residual energy of the node in the network, d, H is the distance between node i and cluster head H, w1, w2 are the weight factors of energy and distance, which are used to adjust the influence of energy and distance in cluster head election. Cluster head candidate node n broadcasts its priority P to the network CH and node ID, the ordinary node selects the candidate node n with the highest priority as the cluster head.
[0083] In step 1, cluster member assignment is achieved by selecting the nearest cluster head from the common node, and the distance to each cluster head is calculated according to the following formula:
[0084]
[0085] Among them, (x i ,y i ) is the position coordinate of node i, (x H ,y H ) is the location coordinate of cluster head H. Ordinary nodes select cluster head H with the smallest d, H and join the cluster, while updating the cluster member information recorded by the cluster head. In step 1, in order to optimize intra-cluster communication, cluster heads are clustered according to the location and number of cluster members. Their locations are calculated according to the following formula:
[0086]
[0087] Among them, d i,j The cluster head divides the cluster members into multiple communication groups according to the distance between nodes and records the node information of each group.
[0088] Step 2: Broadcasting node status feature information within the cluster and modeling conflict topology. Each cluster head broadcasts the status information of its cluster members and calculates the status features of each node. The status features include the number of conflicting edges E, the number of conflicting vertices V, the distance D between the node and the cluster head, and the transmission radius R of the node. Based on the status features, a state matrix is constructed for each node as the basis for subsequent optimization priority allocation. Specifically:
[0089] In order to make all nodes in the underwater network aware of the information of the corresponding neighbor nodes, the nodes broadcast each other's energy, location and status information.
[0090] Step 3: Initialize the intra-cluster topology decision matrix and define the action space. After each node receives the broadcast, a decision matrix will be established for subsequent Q-learning algorithm calculations. Initialize the intra-cluster topology decision matrix and define the action space. Specifically:
[0091] Define the decision matrix Q(s t ,a t ), where s t Indicates the state of the node, a t Indicates the priority assigned to the node; the decision matrix of each node will be initialized with the number of cluster heads as prior information, thereby accelerating the convergence of the subsequent reinforcement learning process; Priority allocation: The decision matrix provides a basis for the subsequent learning process by recording the reward values obtained by the node when taking different actions in different states; After each node is broadcast, a decision matrix will be established for subsequent Q-learning algorithm calculations.
[0092] Step 4: Construct a multi-factor collaborative reward function, iteratively update the Q value, and construct the reward function Q(s) based on the node allocation information within the cluster. t ,a t ), calculate the current state characteristics of the node through the priority calculation formula, determine whether a conflict occurs, and the positive reward generated after the data is successfully sent; update the Q value based on Q-learning to optimize the node priority allocation strategy; the reward function reflects the effect of the priority action taken by each node in a specific state, prompting the system to converge to the optimal solution; specifically:
[0093] In order to calculate the priority score P of each node, and thus guide the reward distribution in the time slot allocation process, the reward is calculated according to the following formula:
[0094]
[0095] Among them, E is the number of conflict edges used to characterize the connection or influence of the node with other nodes, V is the number of conflict vertices used to characterize the degree of influence of the node, D is the distance of the node, R is the transmission radius of the node, n is the number of cluster heads perceived by the node within the transmission radius, and w is the weight coefficient used to adjust the impact of the number of perceived cluster heads n on the priority. 2 (EV) represents the impact of a node conflict. The relationship between the number of edges (E) and the number of vertices (V) is used to measure the conflict priority of a node. The greater the number of edges, the higher the priority. If the number of vertices is close to the number of edges, the conflict impact is greater and the priority is lower. is the distance ratio of the node. The larger the distance D, the lower the priority, but the larger the transmission radius, the higher the priority. w·log(n+1) is the impact of cluster head perception. The greater the number of cluster heads n, the higher the priority, and the more likely it is to affect important target nodes. In step 4, after obtaining the reward R, the decision matrix is updated using the Q-learning formula, which is calculated according to the following formula:
[0096]
[0097] Among them, Q(s t ,a t ) is the Q value of the current node, α is the learning rate, R t is the reward value of the current action. is the maximum Q value in the next state. γ is the discount factor that determines the importance of future rewards. After obtaining the reward Rt, the decision matrix is updated through the Q-learning formula so that the updated Q(s t ,a t ) is closer to the true action value. After all training rounds are completed, the values in the decision matrix represent the cumulative rewards of each node under different priority assignments. Finally, the optimal action of each node can be extracted from the decision matrix.
[0098] Step 5: Priority conflict detection and dynamic allocation mechanism. Based on the priority allocation of nodes within the cluster, it is detected whether multiple nodes have selected the same priority. If a conflict is detected, the action of the conflicting node is reselected to ensure that the priority of each node in the cluster is independent. If there is no conflict, the current allocation result is recorded and used for reward calculation. This step ensures the independence of the priorities of nodes within the cluster and effectively avoids resource conflicts.
[0099] Specifically: During the process of assigning priorities to nodes within a cluster, it is necessary to detect whether there are priority conflicts and adjust the assignment scheme. The specific implementation method is as follows:
[0100] According to the action results learned by the current node in the decision matrix, the priority of each node in the cluster is preliminarily assigned to ensure that all nodes obtain the initial allocation results. The node set in the cluster is defined as The node priority assignment result is If there are any two different nodes Satisfaction: a i =a j (i≠j), it is determined to be a priority conflict.
[0101] If a conflict is detected, the conflict node set Reselect action:
[0102]
[0103] in, is the available priority set, ∑ j≠i δ(a,a j ) is the conflict indicator function:
[0104]
[0105] Among them, R is the total reward value of the current round, For node i in state s i Next select action a i The Q value of β is the conflict avoidance reward weight.
[0106] By minimizing the number of conflicts, the adjusted priorities are ensured to be independent. If there is no conflict in the allocation, the current allocation plan is recorded and the reward function is updated, where β is the conflict avoidance reward weight, which is used for reinforcement learning policy optimization.
[0107] Repeat the above steps to ensure the uniqueness and rationality of the priority allocation scheme within the entire cluster, and gradually improve the allocation results through training.
[0108] Step 6: Final time slot allocation strategy extraction and global scheduling implementation. After the Q value iteration is completed, for each node, the optimal action a is extracted from the decision matrix. t , as the final priority and time slot allocation scheme for the node; by extracting the optimal priority allocation and time slot allocation from the decision matrix, it ensures that each node in the network can perform resource scheduling according to the optimal strategy, achieving efficient communication and resource utilization. Specifically:
[0109] For each cluster head node Extracting actions from the decision matrix
[0110]
[0111] Among them, Q(s i,a) is the node status s i The Q value of action a. Mapping to communication time slots Define the mapping function, and the function f must satisfy the time slot uniqueness constraint:
[0112]
[0113] The cluster head node will globally distribute the results Broadcast to the member nodes in the cluster to ensure that each node schedules communication according to the allocated time slot.
[0114] Through the above technical solution, the present invention addresses the resource allocation problem in underwater wireless sensor networks by proposing a method for underwater dynamic network resource scheduling and time slot conflict optimization. First, the initial node state is constructed based on the state characteristics of underwater sensor nodes. Second, reinforcement learning techniques are combined to dynamically optimize the node's transmission priority and time slot allocation to improve the network's overall communication efficiency. Then, an optimization mechanism for time slot allocation conflict detection is implemented to resolve resource competition and transmission conflicts between nodes. Finally, the intra-cluster network dynamically adjusts the node's transmission strategy by constructing a reward function, gradually converging to the optimal resource allocation solution. This method comprehensively considers inter-node communication conflicts, resource allocation efficiency, and dynamic transmission queue adjustment in multi-cluster underwater network environments. Through Q-learning and time slot optimization mechanisms, it can effectively reduce communication conflicts and improve node resource utilization efficiency in complex underwater acoustic communication environments. Compared with traditional algorithms, this method is applicable and effective in dynamically changing multi-cluster network environments, and can more rationally optimize underwater node transmission strategies.
[0115] A computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the device containing the computer-readable storage medium executes the underwater dynamic network resource scheduling and time slot conflict optimization method described above. The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory, or other memory.
[0116] An electronic device includes: a memory and a processor, wherein the memory stores a program that can be run on the processor, and when the processor executes the program, the underwater dynamic network resource scheduling and time slot conflict optimization method described above is implemented.
[0117] If the modules / units integrated in the electronic device described in this application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment methods by instructing the relevant hardware devices to complete them through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments.
[0118] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0119] Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in the electronic device to implement the underwater dynamic network resource scheduling and time slot conflict optimization method described in any of the above embodiments.
[0120] The present invention comprehensively considers the problems of long propagation delay, significant dynamic changes and serious channel conflicts of underwater network nodes in the underwater environment, and proposes a multi-cluster node transmission allocation optimization problem based on Q-Learning reinforcement learning. Through the Q-learning algorithm, the topological state of the node and the priority of the data packet are introduced into the reinforcement learning model, and the data transmission scheduling scheme is gradually optimized through multiple rounds of interactive learning between the node and the environment. The present invention comprehensively considers multiple influences, improves the adaptability of network scheduling, optimizes the allocation of channel resources, and effectively improves network throughput and system efficiency. Through dynamic learning and scheduling mechanisms, the present invention can adapt to changes in the underwater environment in real time, improve communication stability, and more reasonably optimize the multi-node resource allocation scheme of the underwater network.
[0121] Different from traditional underwater communication protocols, the present invention uses a reinforcement learning algorithm to comprehensively consider multiple factors such as propagation delay, channel conflict and node energy consumption, thereby proposing a dynamic resource allocation method based on Q-learning. It can optimize resource allocation in real time under dynamic topologies and complex environments, reduce network conflicts, and improve the security and stability of the system.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and other division methods may be used in actual implementation.
[0123] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0124] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
[0125] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products or apparatuses.
[0126] Note that the above are only preferred embodiments of the present invention and the principles of the technology used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention is described in detail through the above embodiments, the present invention is not limited to the specific embodiments described herein. Without departing from the concept of the present invention, it may also include many other effective embodiments, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. Underwater dynamic network resource scheduling and time slot conflict optimization method, characterized by: The steps include: Step 1: Cluster head election and cluster member allocation, specifically: Cluster head candidate node n is calculated based on the residual energy E remaining The node broadcasts its own status and location information, and the ordinary node selects the nearest cluster head to join its cluster according to the received broadcast information: During the election process, the candidate cluster head node broadcasts its own residual energy and location information to allow other nodes to judge whether it is suitable to become a cluster head; When the node selects a cluster head, it gives priority to the cluster head P that is closer and has more residual energy. CH , to achieve more efficient resource allocation; Step 2: Broadcasting node status feature information within the cluster and modeling conflict topology. Specifically: Each cluster head broadcasts the status information of its cluster members and calculates the status characteristics of each node, including the number of conflicting edges E, the number of conflicting vertices V, the distance D between the node and the cluster head, and the transmission radius R of the node. Based on the status characteristics, a state matrix is constructed for each node, which serves as the basis for subsequent optimization priority allocation. Step 3: After each node receives the broadcast, a decision matrix is established for subsequent Q-learning algorithm calculations, the cluster topology decision matrix is initialized, and the action space is defined. Specifically: Define the decision matrix Q(s t ,a t ), where s t Indicates the state of the node, a t Indicates the priority assigned to the node; the decision matrix of each node will be initialized with the number of cluster heads as prior information, thereby accelerating the convergence of the subsequent reinforcement learning process; Priority allocation: The decision matrix provides a basis for the subsequent learning process by recording the reward values obtained by the node taking different actions in different states; Step 4: Construct a multi-factor collaborative reward function and iteratively update the Q value. Specifically: Construct the reward function Q(s) based on the node allocation information within the cluster t ,a t ), calculate the current state characteristics of the node through the sending priority calculation formula, determine whether a conflict occurs, and the positive reward generated after the data is successfully sent; According to Q-learning, update the Q value to optimize the node priority allocation strategy; The reward function reflects the effect of the priority actions taken by each node in a specific state, prompting the system to converge towards the optimal solution; Step 5: Priority conflict detection and dynamic allocation mechanism, specifically: Based on the priority allocation of nodes within the cluster, check whether multiple nodes have selected the same priority. If a conflict is detected, reselect the action of the conflicting node to ensure that the priority of each node in the cluster is independent. If there is no conflict, record the current allocation result for reward calculation. This step ensures the independence of the priorities of nodes within the cluster and effectively avoids resource conflicts. Step 6: Final time slot allocation strategy extraction and global scheduling implementation, specifically: After the Q value iteration is completed, for each node, the optimal action a is extracted from the decision matrix. t , as the final priority and time slot allocation scheme of nodes; By extracting the optimal priority allocation and time slot allocation from the decision matrix, it ensures that each node in the network can perform resource scheduling according to the optimal strategy to achieve efficient communication and resource utilization.
2. The underwater dynamic network resource scheduling and time slot conflict optimization method according to claim 1, characterized in that: The remaining energy E of the underwater node in step 1 remaining For the following model: AND remaining =And initial -(AND transmit +E receive +E process +E idle +E sensing ) (1) Among them, E transmit =P tx,elec ·t tx +∈·d a ·N Among them, P tx,elec ·t tx is the circuit power consumption, ∈·d a N is the signal amplifier number, P tx,elec is the transmitting circuit power, t tx is the transmission time, N is the data volume, ∈ is the amplifier energy efficiency coefficient, d a is the transmission distance under path loss; E receive =P rx,elec ·t rx Among them, P rx,elec is the receiving circuit power, t rx Receiving time E sensing =P sensor ·t sensing Among them, P sensor Sampling power for the sensor E idle =P idle ·t idle Among them, P idle is the idle state power, t idle The total time the node was in idle mode AND process =And bit ·N Among them, E bit is the processing energy consumption per bit of data.
3. The underwater dynamic network resource scheduling and time slot conflict optimization method according to claim 1, characterized in that: The cluster head node selected in step 1 is denoted as P CH Calculate it as follows: Among them E remaining is the residual energy of candidate node n, E max is the maximum residual energy of the nodes in the network, d i,H is the distance from node i to cluster head H, w1 and w2 are weight factors of energy and distance, which are used to adjust the influence of energy and distance in cluster head election; cluster head candidate node n broadcasts its priority P to the network CH and node ID, the ordinary node selects the candidate node n with the highest priority as the cluster head.
4. The underwater dynamic network resource scheduling and time slot conflict optimization method according to claim 1, characterized in that: In step 1, cluster member allocation is achieved by selecting the nearest cluster head from the common node, where the distance d between each cluster head is i,H Calculated according to the following formula: Among them, (x i ,y i ) is the position coordinate of node i, (x H ,y H ) is the location coordinate of cluster head H; ordinary nodes choose d i,H The smallest cluster head H joins the cluster and updates the cluster member information recorded by the cluster head.
5. The underwater dynamic network resource scheduling and time slot conflict optimization method according to claim 1, characterized in that: In step 1, in order to optimize the intra-cluster communication, the cluster head divides the clusters according to the location and number of the cluster members. i,j Calculated according to the following formula: Among them, d i,j is the distance between nodes. The cluster head divides the cluster members into multiple communication groups according to the distance between nodes and records the node information of each group.
6. The underwater dynamic network resource scheduling and time slot conflict optimization method according to claim 1, characterized in that: In step 2, in order to ensure that all nodes in the underwater network receive information about their neighboring nodes, the nodes broadcast each other's energy, location, and status information.
7. The method for scheduling underwater dynamic network resources and optimizing time slot conflicts according to claim 1, characterized in that: In step 4, the priority score P of each node is calculated to guide the reward allocation during the time slot allocation process; the priority calculation formula is as follows: Among them, E is the number of conflict edges used to characterize the connection or influence of the node with other nodes, V is the number of conflict vertices used to characterize the degree of influence of the node, D is the distance of the node, R is the transmission radius of the node, n is the number of cluster heads perceived by the node within the transmission radius, and w is the weight coefficient used to adjust the influence of the number of perceived cluster heads n on the priority; E 2 (EV) represents the impact of node conflicts. The relationship between the number of edges (E) and the number of vertices (V) is used to measure the conflict priority of a node. The more edges there are, the higher the priority. If the number of vertices is close to the number of edges, the conflict impact is greater and the priority is lower. is the distance ratio of the node; the larger the distance D, the lower the priority, but the larger the transmission radius, the higher the priority; w·log(n+1) is the influence of cluster head perception; the more cluster heads n, the higher the priority, and the more likely it is to affect important target nodes.
8. The underwater dynamic network resource scheduling and time slot conflict optimization method according to claim 1, characterized in that: In step 4, after obtaining the reward R, the decision matrix is updated using the Q-learning formula, which is calculated according to the following formula: Among them, Q(s t ,a t ) is the Q value of the current node, α is the learning rate, R t is the reward value of the current action; is the maximum Q value in the next state; γ is the discount factor, which determines the importance of future rewards; after obtaining the reward Rt, the decision matrix is updated through the Q-learning formula so that the updated Q(s t ,a t ) is closer to the true action value. After all training rounds are completed, the values in the decision matrix represent the cumulative rewards of each node under different priority assignments. Finally, the optimal action of each node can be extracted from the decision matrix.
9. The multi-cluster node priority allocation optimization method based on Q-Learning reinforcement learning according to claim 1, characterized in that: During the priority allocation process of nodes within the cluster in step 5, it is necessary to detect whether there is a priority conflict and adjust the allocation plan. The specific implementation method is as follows: Based on the action results learned by the current node in the decision matrix, the priority of each node in the cluster is preliminarily assigned to ensure that all nodes obtain the initial allocation results; Define the node set in the cluster as The node priority assignment result is If there are any two different nodes Satisfaction: a i =a j (i≠j), it is determined to be a priority conflict; If a conflict is detected, the conflict node set Reselect action: in, is the available priority set, ∑ j≠i δ(a,a j ) is the conflict indicator function: Among them, R is the total reward value of the current round, For node i in state s i Next select action a i The Q value of , β is the conflict avoidance reward weight; By minimizing the number of conflicts, the adjusted priorities are ensured to be independent. If there is no conflict in the allocation, the current allocation plan is recorded and the reward function is updated, where β is the conflict avoidance reward weight, which is used for reinforcement learning strategy optimization. Repeat the above steps to ensure the uniqueness and rationality of the priority allocation scheme within the entire cluster, and gradually improve the allocation results through training.
10. The multi-cluster node priority allocation optimization method based on Q-Learning reinforcement learning according to claim 1, characterized in that: The specific implementation method of the cluster head node allocating communication time slots and priorities in step 6 is as follows: For each cluster head node Extracting actions from the decision matrix Among them, Q(s i ,s) is the node state a i The Q value of action a; the above priority Mapping to communication time slots Define the mapping function, and the function f must satisfy the time slot uniqueness constraint: The cluster head node will globally distribute the results Broadcast to the member nodes in the cluster to ensure that each node schedules communication according to the allocated time slot.
Citation Information
Cited By
Time slot allocation selection method for multi-terminal communication resources of wind power plant
CN120658345A
A time slot allocation selection method for wind farm multi-terminal communication resources
CN120658345B
Underwater acoustic sensor network topology clustering method, system, device, medium and product
CN122120183A