Underwater acoustic network cross-layer transmission method and system based on multi-agent reinforcement learning

By applying multi-agent reinforcement learning technology in water acoustic communication, the cross-layer transmission method of water acoustic network is realized, solving the problem of signal propagation delay and insufficient adaptability of time-varying channels, and significantly improving transmission efficiency and stability.

CN120200696APending Publication Date: 2025-06-24NANHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337467.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing water acoustic communication technology is difficult to effectively reduce signal propagation delay and adapt to time-varying water acoustic channels, resulting in insufficient transmission efficiency and stability.

Method used

The cross-layer transmission method of the acoustic network based on multi-agent reinforcement learning is adopted. By constructing the acoustic sensor network and the MARL-TS protocol, the receiving node is regarded as an agent, and action decisions are made independently, so as to realize cross-layer joint optimization between the physical layer and the MAC layer.

Benefits of technology

Under time-varying channel conditions, intelligent optimization of transmission strategies is achieved, the transmission failure rate caused by channel changes is reduced, the system's adaptability and stability is improved, channel resources are maximized, and transmission efficiency is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200696A_ABST
    Figure CN120200696A_ABST
Patent Text Reader

Abstract

According to the underwater acoustic network cross-layer transmission method and system based on multi-agent reinforcement learning, an underwater acoustic sensor network is constructed, a plurality of receiving nodes on the water surface and a plurality of sending nodes underwater are paired one to one for information transmission, and the receiving nodes perform information sharing and centralized training through a center node; each receiving node is regarded as an intelligent agent through an MARL-TS protocol, and the action of each time slot of the intelligent agent is defined as idle or transmission at specific power in a transmission link between the intelligent agent and the paired sending node; and each agent independently makes an action decision in each period according to untransmitted data of the paired sending nodes and channel parameters of corresponding transmission links, and performs power allocation and time slot resource scheduling under a time-varying channel to realize cross-layer joint optimization of a physical layer and an MAC layer. Compared with the existing MAC protocol based on reinforcement learning driving, the method can significantly improve the transmission energy efficiency and effectively reduce the information transmission delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of underwater acoustic communication and multi-agent reinforcement learning, and particularly to an underwater acoustic network cross-layer transmission method and system based on multi-agent reinforcement learning. Background Art

[0002] Reducing the signal propagation delay problem is one of the key steps in underwater communication. The medium access control (MAC) protocols and related design strategies widely used in existing research results are difficult to be directly applied to the underwater acoustic communication scenario. The propagation speed of sound waves in water is about 1500 m / s, which is significantly different from the speed of electromagnetic waves. The resulting signal propagation delay problem cannot be ignored. The high propagation delay further leads to the spatio-temporal uncertainty in underwater acoustic communication: the arrival time of the acoustic signal depends not only on the transmission time, but also is significantly affected by the distance between the transmitter and the receiver. This characteristic undoubtedly exacerbates the difficulty of achieving efficient transmission scheduling in a complex underwater interference environment. In addition, the underwater acoustic channel is highly dynamic, and its characteristics will fluctuate with changes in time and environmental conditions. For example, various factors such as water flow, temperature gradient, and surface movement will have a significant impact on the channel, resulting in the instability of signal quality and the intermittent characteristics of communication connections.

[0003] The current solutions are mostly based on the MAC layer design, assuming that the underwater acoustic propagation delay is an integer multiple of the time slot length, and do not consider the complex changes in the physical layer underwater acoustic channel. Therefore, it is difficult to adapt to the situation of arbitrary propagation delays in the real scenario, nor can it optimize the transmission energy efficiency for channel changes, thereby improving the underwater communication efficiency. Summary of the Invention

[0004] One of the purposes of the present invention is to provide an underwater acoustic network cross-layer transmission method based on multi-agent reinforcement learning to optimize the transmission energy efficiency, reduce the transmission delay, and show stronger adaptability and stability during the communication process.

[0005] To solve the above technical problems, the present invention adopts the following technical solutions: An underwater acoustic network cross-layer transmission method based on multi-agent reinforcement learning: Construct an underwater acoustic sensor network, enabling multiple receiving nodes on the water surface to be paired one-to-one with multiple transmitting nodes underwater for information transmission. Each receiving node conducts information sharing and centralized training through a central node. Regarding each receiving node as an agent through the MARL-TS protocol, and defining the action of each agent in each time slot as either the transmission link with its paired transmitting node being idle or transmitting at a specific power. Each agent independently makes action decisions in each epoch based on the untransmitted data of the paired transmitting node and the channel parameters of the corresponding transmission link, performs power allocation and time slot resource scheduling under time-varying channels to achieve cross-layer joint optimization of the physical layer and the MAC layer, and realizes collaborative learning through the interaction between multiple agents and the environment to optimize the overall system benefit.

[0006] Preferably, the receiving node is a buoy floating on the water surface, and the transmitting node is a sensor arranged underwater. Each sensor conducts sensing by collecting data from the environment and is equipped with a data buffer for storing signal packets that have not been successfully transmitted in time.

[0007] More preferably, the epoch is a plurality of non-consecutive time series into which the system operation time is divided, so as to facilitate the system being periodically awakened to transmit data to save energy. Each epoch consists of a certain number of time slots, and the length of the time slot matches the signal packet transmission time.

[0008] More preferably, the underwater acoustic sensor network further includes a signal transmission and reception model, a spatio-temporal channel model, and a signal decoding model.

[0009] More preferably, the multi-agent reinforcement learning is carried out as follows: First, the receiving node determines the transmission time slot and power allocation of the transmitting node and issues instructions to the transmitting node for execution. Subsequently, the transmitting node conducts data transmission according to the instructions. However, due to the influence of the external environment, there may be differences in the transmission effects of different nodes. The receiving node estimates the underwater acoustic channel state based on the received signal and monitors the change in the data buffer state, thereby calculating the corresponding reward value. Finally, the receiving node continuously adjusts the transmission time slot and power allocation strategy according to the rewards obtained under different states and actions to optimize the transmission performance.

[0010] More preferably, the actor-critic method is adopted to construct the MARL-TS algorithm framework, and the MAPPO algorithm is used to update the policy network, and the stability of the training process is improved by restricting the degree of policy update.

[0011] In addition, the present invention also provides an underwater acoustic network cross-layer transmission system based on multi-agent reinforcement learning, which includes:

[0012] A sending node, which is an underwater sensor equipped with a data buffer;

[0013] A receiving node, which is a buoy located on the water surface and uses underwater acoustic communication to transmit information with underwater nodes;

[0014] A central node, which is used to communicate with each receiving node and deploys the MARL-TS protocol. Thus, each receiving node is regarded as an agent, and the actions of the agents are defined as idle or transmitting at a specific power. Each agent independently makes action decisions in each period according to the data cache amount and channel parameters to achieve cross-layer joint optimization of power allocation and time slot resource scheduling at the physical layer and MAC layer under time-varying channels.

[0015] This system operates using the above-mentioned cross-layer transmission method for underwater acoustic networks based on multi-agent reinforcement learning.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] 1. By adopting multi-agent reinforcement learning technology, the present invention enables the underwater acoustic network system to automatically adapt to the dynamic changes of the channel under time-varying channel conditions and achieve intelligent optimization of transmission strategies. This adaptive ability enables the system to maintain stable communication performance in the face of complex underwater environments, effectively reducing the transmission failure rate caused by channel changes and significantly enhancing the adaptability and stability of the system.

[0018] 2. By jointly adjusting the transmission power and transmission time slots, it is possible to maximize the utilization of channel resources on the basis of ensuring successful signal transmission, reduce signal conflicts and interference, and thus significantly improve the transmission efficiency of the system. The multi-agent reinforcement learning algorithm can dynamically allocate the optimal transmission power and time slots for each sending node according to the real-time channel state and network topology structure, avoiding the problems of transmission resource waste and transmission delay caused by the failure to consider channel changes in traditional methods and achieving the maximization of transmission efficiency.

[0019] 3. The present invention innovatively applies multi-agent reinforcement learning technology to the transmission of underwater acoustic networks under time-varying channels. Thus, in view of the characteristics of high propagation delay and time-varying underwater acoustic channels, a transmission mode adapted to complex underwater acoustic environments is designed, and a joint transmission scheduling and power allocation method based on online learning is proposed, realizing the dynamic optimization of resource utilization rate and communication performance. Compared with existing reinforcement learning-driven MAC protocols, this method significantly improves the transmission energy efficiency and can effectively reduce the information transmission delay. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of an underwater acoustic sensor network in the embodiment;

[0021] Figure 2 Signal transceiver schematic diagram in the embodiment;

[0022] Figure 3 Schematic diagram of the interaction between multiple agents and the environment in the MARL-TS protocol in the embodiment;

[0023] Figure 4 Schematic diagram of the complete MARL-TS algorithm framework in the embodiment. Detailed implementation manner

[0024] For the convenience of those skilled in the art, the present invention will be further described below in conjunction with the embodiments and the accompanying drawings. The content mentioned in the implementation manner does not limit the present invention.

[0025] An underwater acoustic network cross-layer transmission method based on multi-agent reinforcement learning. The method mainly includes: constructing an underwater acoustic sensor network, enabling multiple receiving nodes on the water surface to be paired one-to-one with multiple transmitting nodes underwater for information transmission. Each receiving node performs information sharing and centralized training through a central node. Regarding each receiving node as an agent through the MARL-TS protocol, and defining the action of each agent in each time slot as idle or transmitting at a specific power. Each agent independently makes action decisions in each period according to the untransmitted data and channel parameters to achieve cross-layer joint optimization of the physical layer and the MAC layer for power allocation and time slot resource scheduling under a time-varying channel, and through the interaction between multiple agents and the environment, realizes collaborative learning to optimize the overall system benefit.

[0026] This embodiment is aimed at the underwater Internet of Things scenario, where multiple underwater sensors need to transmit information to the water surface in a timely manner. To enhance the real-time performance and reliability of information transmission, the system uses multiple buoys as receiving nodes to expand the communication coverage and provide stable services. In addition, in the unmanned system of this embodiment, the receiving nodes and sensors are paired one-to-one to optimize the utilization rate of communication resources, improve the underwater information transmission efficiency, and achieve high-reliability and low-latency data transmission. As Figure 1As shown, the receiving node is a buoy floating on the water surface, which uses underwater acoustic communication to transmit information with underwater nodes and can send the collected information to the land data center through a high-speed radio link. As a sending node, each sensor senses by collecting data from the environment and is equipped with a data buffer for storing signal packets that have not been successfully sent. In addition, signals that have not been received and decoded will remain in the data buffer and wait for subsequent retransmission. The larger the data cache volume, the higher the information transmission delay will be, causing network congestion. Each receiver uniquely corresponds to one sending node, forming a one-to-one communication relationship to ensure the efficient transmission of sensor node data. Of course, due to the existence of other unpaired sending nodes, the receiver is inevitably interfered by them, thus affecting the communication transmission efficiency. In addition, each receiving node can share information, conduct centralized training and collaborative transmission through the central node.

[0027] Figure 2 It shows the operating mode of the system in terms of time, that is, the system consists of epochs, and each epoch consists of a certain number of time slots. For the scenario where the underwater unmanned system has limited energy supply, this embodiment considers that the sending nodes are periodically awakened, and the epochs are non-continuously arranged with a fixed time T epoch . In addition, considering that the time slot is consistent with the length of the signal packet to be transmitted, each sending node can transmit a single signal packet in each time slot.

[0028] In this system, since each receiving node is connected to the central node and can perform fast information sharing, each receiving node can be precisely clock-synchronized. Therefore, the start time of each sending node's epoch can be adjusted to achieve the synchronization of signal reception at each node. Due to the different distances of each interference link, the signal propagation delays are different. Therefore, the interference signals from other sending nodes are not completely aligned with the useful signals, resulting in interference overlap between the received useful signals and interference signals. This partial overlap causes different interference levels in different time slots, depending on the relative time and propagation delay of the interference signals. After the receiving node finishes decoding all the useful signals, the ACK packet will be fed back to the sending node, containing channel information and information on whether the signal packet has been successfully decoded.

[0029] The above-mentioned underwater acoustic sensor network also includes a signal transmission and reception model, a channel model, and a signal decoding model. Specifically:

[0030] In the signal transmission and reception model, let and represent the sets of sending nodes and receiving nodes respectively, where the number of nodes in the two sets is equal This embodiment considers each sending node corresponding to a unique receiving node and assumes that this one-to-one correspondence relationship is known as prior knowledge. Let Denoted as the transmission start time of the l-th signal packet from the sending node n i , and denoted as the propagation delay between the sending node n i and its corresponding receiving node r i . The time-of-arrival (ToA) of the l-th signal packet at the receiving node r i can be expressed as:

[0031]

[0032] For a non-corresponding node m k ≠n i , the signal packet sent from m k will interfere with the receiving node r i . Let denote the propagation delay from the interfering node m k to the receiving node r i , and denote the ToA of the l m -th interfering signal packet at the receiving node r i . It can be calculated by Equation (2):

[0033]

[0034] The difference in the ToA between the l i -th signal packet from the sending node n n and the l k -th interfering signal packet from the interfering node m m at the receiving node r i can be expressed as:

[0035]

[0036] The interference pattern in the time domain is determined by the ToA difference and the duration T bl of the signal packet. Denoted as the interference of the l i -th signal packet from the sending node n n at the receiving node r i by the l k -th signal packet from the interfering node m m . Denoted as the complete separation of two signal packets.

[0037] In the channel model, this embodiment considers that the underwater acoustic channel gain is affected by both space and time, which is expressed as:

[0038]

[0039] where d is the distance between the transmitting node and the receiving node, t is the time, h0 is the normalization constant, β is the path loss exponent, α is the absorption coefficient that depends on the medium and signal frequency, and f X is the probability density function of the channel gain:

[0040]

[0041] where and depend on the distance d and time t. These parameters capture the channel fading characteristics, which affect the received signal strength as a function of time and distance. These parameters are constructed into a channel state vector:

[0042]

[0043] The variation of the channel parameters over time is modeled as a first-order Markov chain:

[0044]

[0045] where is the transition matrix that captures the dynamic changes of the channel parameters, is the noise vector representing the random fluctuations. This model allows each parameter in to evolve based on its previous state and the influence of the random noise, thus effectively characterizing the time-varying nature of the channel conditions over a certain distance.

[0046] In the signal decoding model, using the attenuation factor in Equation (4), the received power of the l-th signal packet from transmitting node n i at receiving node r i can be calculated as:

[0047]

[0048] Let denote the signal-to-interference-and-noise ratio (SINR) of the l-th signal packet from transmitting node n i at receiving node r i . For single signal packet processing and orthogonal frequency division multiplexing (OFDM) systems, the SINR can be calculated as:

[0049]

[0050] where N0 is the noise power, B is the transmission bandwidth, and N bl is the number of time slots in each period. [·] += max{·, 0}. The interference signal of the l-th signal packet is contributed by the signal packets of all non-corresponding transmitting nodes.

[0051] This embodiment considers determining whether the current link transmission is successful based on whether the received SINR of the useful signal block exceeds the decoding threshold λ. th Define as the indication function of whether the transmission of the l-th signal block of receiving node i at time t is successful:

[0052]

[0053] The total number of successfully transmitted signal packets S i (t) at time t can be obtained by the following formula:

[0054]

[0055] For transmitting node i, the change in the data buffer during the t-th period is given by:

[0056] Q i (t + 1) = Q i (t) + I i (t) - S i (t) (12)

[0057] where Q i (t + 1) represents the data buffer size of transmitting node i at the beginning of the (t + 1)-th period, Q i (t) is the data buffer size at the beginning of the t-th period, and I i (t) represents the number of signal packets flowing into the data buffer during the t-th period. When the signal transmission fails, the signal will remain in the data buffer for retransmission in subsequent periods.

[0058] The system objective constructed in this embodiment is to improve network throughput, enhance energy efficiency, and reduce transmission delay. However, achieving these goals faces many challenges, such as the high dynamics of the underwater acoustic channel, the need for comprehensive optimization among multiple objectives, and the sequential decision-making nature of the problem itself. Traditional optimization methods (such as genetic algorithms) can solve multi-agent problems, but they are usually more suitable for static environments. RL provides a powerful alternative, which can automatically extract features and make intelligent decisions in dynamic environments. Especially in systems involving multiple cooperative agents, MARL becomes a more natural choice. MARL can not only efficiently handle the dynamics and distributed characteristics of the problem, but also provide a robust solution for the joint optimization of throughput and energy management. Based on this, this embodiment models the problem as a Markov Decision Process (MDP) for solution.

[0059] Through the interaction of multiple agents in a shared environment, MARL is used to achieve collaborative learning to optimize the overall system benefit. Figure 3 The basic framework of MARL is shown, and its specific transmission and learning steps are as follows: First, the receiving node determines the transmission time slot and power allocation of the sending node and issues instructions to the sending node for execution. Subsequently, the sending node performs data transmission according to the instructions. However, due to the influence of the external environment, the transmission effects of different nodes may vary. The receiving node estimates the underwater acoustic channel state based on the received signal and monitors the change of the data cache state, thereby calculating the corresponding reward value. Finally, the receiving node continuously adjusts the transmission time slot and power allocation strategy according to the rewards obtained under different states and actions to optimize the transmission performance.

[0060] The following are the MARL elements considered in this embodiment:

[0061] Agent: In this embodiment, each receiving node is regarded as an agent, and each agent makes independent decisions in each period.

[0062] Action: The action of agent i in the t-th period is represented as:

[0063] a i (t)=[a i1 (t),a i2 (t),...,a iN (t)] (10)

[0064] Each element a ij (t) represents the specific action in the j-th time slot. Assume P max is the upper bound of the transmission power. If 0 < a ij (t) ≤ P max , it means that the sending node transmits at the power of a ij (t); if a ij (t) = 0, it means that the node is idle in this time slot and the transmission power is 0.

[0065] Observation: At the t-th period, agent i can obtain its individual observation value o i (t), as follows:

[0066] o i (t)=[Q i (t),c i,1 (t),c i,2 (t),...,c i,N (t)](11)

[0067] Among them, the evolution formula of the data cache quantity Q i (t) is shown in Equation (12), {c i1(t),c i2 (t),...,c iN (t)} represents the parameters of N channels, including the useful channels and interference channels, between agent i and all sending nodes.

[0068] State: At the t-th period, the environmental state s(t) consists of the observations of all agents:

[0069] s(t) = [o1(t), o2(t),..., o N (t)] (13)

[0070] Reward: In this embodiment, a strategy of cooperation among agents is adopted, that is, all agents share a common reward signal. In this way, the system encourages cooperation among agents and promotes their joint efforts to optimize the global goal. For the t-th period, the shared reward r(t) is defined as:

[0071]

[0072] In this formula: ΣS i (t) represents the total number of successfully transmitted signal packets of all agents; E total is the total energy consumed by agents during the t-th period; the weights w1 and w2 are used to balance the contributions of data caching and energy consumption.

[0073] Figure 4 This is the framework schematic diagram of the MARL-TS algorithm in this embodiment. This embodiment adopts an Actor-Critic method, and its general working process is: at each time step t, each agent samples an action from the policy network according to the local observation o i (t) described by Equation (14). The joint action {a1(t),..., a N (t)} of all agents acts on the environment via the sending nodes, and the environment returns the reward r(t) and the next global state s(t + 1). These transition information {s(t), {a1(t),..., a N (t)}, r(t), s(t + 1)} are stored in the experience replay buffer D for subsequent training use.

[0074] The policy network (i.e., the Actor network) π θ is deployed at the receiving node and generates transmission parameters according to the current system state. Generally, the Actor network is implemented by a neural network, with the system state as its input and the corresponding transmission time slot and transmission power as its output:

[0075] a(t) = π θ (s(t)) (17)

[0076] where θ is the learning weight parameter of the neural network.

[0077] To optimize the policy, the system randomly samples a set of state-action trajectories from the buffer D:

[0078] τ = [s(0), a(0), s(1), a(1),..., s(T), a(T)], (18)

[0079] and calculates the discounted cumulative reward based on this trajectory:

[0080]

[0081] Then, the expected reward of the system represents the expected value of the cumulative reward obtained by the agent interacting with the environment by generating the action a(t) based on the history of the Actor network starting from the initial state under the current policy. To optimize the policy of the Actor network to generate better transmission parameters, the policy can be optimized by maximizing the following objective function:

[0082]

[0083] Based on the policy gradient theorem, the parameter update rule of the Actor network is:

[0084]

[0085] where η θ is the learning rate. Through this optimization process, the Actor network gradually increases the probability of selecting the optimal policy, thereby improving the transmission benefit of the system.

[0086] Based on the sampled trajectory, the expected reward can be further expanded as:

[0087]

[0088] where, π θ (a(t)|s(t)) represents the probability of generating a(t) based on the state s(t); δ(t) is the temporal difference error, calculated as follows:

[0089]

[0090] where represents the value function of the agent in the given state, reflecting the expected benefit that can be obtained in the future regardless of the action taken starting from the state s(t). δ(t) measures the improvement in the benefit obtained by selecting the action a(t) in the state s(t) compared to the average benefit of the next state, thereby guiding the Actor network to optimize the transmission parameters.

[0091] Generally, the value function is represented by the Critic network, where φ is the learning weight parameter of the neural network. During the transmission process, the Critic network needs to continuously learn to improve its estimation accuracy of the average return. Essentially, the Critic network can be used to evaluate the quality of a state-action pair, and its optimization objective is:

[0092]

[0093] where η φ is the learning rate; is the optimization objective. This optimization process reduces the difference between the average return estimated based on the immediate reward and the theoretical return so as to improve the value estimation ability of the Critic network and assist the Actor network to optimize the policy more efficiently.

[0094] The MARL-TS proposed in this embodiment updates the policy network by using the MAPPO (Multi-Agent Proximal Policy Optimization) algorithm on the basis of Actor-Critic. Its core is to perform importance sampling on the ratio of the old and new policies and a clipping mechanism for the deviation degree of the policy update ratio during policy update. By restricting the degree of policy update, while ensuring the sampling efficiency, it significantly improves the stability of the training process, which is particularly important for collaborative learning in a multi-agent environment.

[0095] To make it easier for those of ordinary skill in the art to understand the improvements of the present invention over the prior art, some of the drawings and descriptions of the present invention have been simplified, and the above embodiments are preferred implementation schemes of the present invention. In addition, the present invention can also be implemented in other ways. Any obvious replacement without departing from the concept of the technical solution of the present invention is within the protection scope of the present invention.

Claims

1. A cross-layer transmission method for underwater acoustic networks based on multi-agent reinforcement learning, characterized by: An underwater acoustic sensor network is constructed to pair multiple receiving nodes on the water surface with multiple sending nodes underwater one-to-one for information transmission. Each receiving node shares information, conducts centralized training and collaborative transmission through a central node. Each receiving node is regarded as an intelligent agent through the MARL-TS protocol, and the action of the intelligent agent in each time slot is defined as when the transmission link between it and the paired sending node is idle or transmitted at a specific power. Each intelligent agent makes independent action decisions in each period based on the untransmitted data of the paired sending node and the channel parameters of the corresponding transmission link. Power allocation and time slot resource scheduling are performed under time-varying channels to achieve cross-layer joint optimization of the physical layer and the MAC layer. Collaborative learning is achieved through interaction between multiple agents and the transmission environment to optimize the overall benefit of the system.

2. The method for cross-layer transmission of underwater acoustic network based on multi-agent reinforcement learning according to claim 1 is characterized in that: The receiving node is a buoy floating on the water surface, and the sending node is a sensor set underwater. Each sensor senses by collecting data from the environment and is equipped with a data buffer for storing signal packets that are not successfully sent in time.

3. The method for cross-layer transmission of underwater acoustic network based on multi-agent reinforcement learning according to claim 1 is characterized in that: The period is a plurality of non-continuous time sequences into which the system operation time is divided, so that the system is periodically awakened to transmit data to save energy. Each period consists of a certain number of time slots, and the length of the time slot matches the signal packet transmission time.

4. The method for cross-layer transmission of underwater acoustic network based on multi-agent reinforcement learning according to claim 1 is characterized in that: The underwater acoustic sensor network is provided with a spatiotemporal channel model. In the channel model, the underwater acoustic channel gain is affected by both space and time, which is expressed as: Where d is the distance between the sending node and the receiving node, t is the time, h0 is the normalization constant, β is the path loss exponent, α is the absorption coefficient that depends on the medium and the signal frequency, and f X is the probability density function of the channel gain: in and Depending on the distance d and time t, these parameters capture the channel fading characteristics, which affect the received signal strength over time and distance. These parameters are constructed into a channel state vector: The change of channel parameters over time is constructed as a first-order Markov chain model: in, is the transfer matrix that captures the dynamic changes of channel parameters, is a noise vector representing random fluctuations. The model allows Each parameter in evolves based on its previous state and the influence of random noise, effectively characterizing the time-varying characteristics of channel conditions at a certain distance.

5. The method for underwater acoustic network cross-layer transmission based on multi-agent reinforcement learning according to claim 4 is characterized in that: The underwater acoustic sensor network is also provided with a signal decoding model. In the signal decoding model, the attenuation factor in equation (4) is used to send the node n i The e-th signal packet is received at the receiving node r i The received power at is calculated as: use Represents the sending node n i The e-th signal packet is received at the receiving node r i For single signal packet processing and orthogonal frequency division multiplexing system, the signal to noise ratio is calculated as: Where N0 is the noise power, B is the transmission bandwidth, and N bl is the number of time slots in each epoch, [·] + =max{·,0}, the interference signal of the lth signal packet is contributed by the signal packets of all non-corresponding sending nodes. It considers the spatiotemporal information of the underwater acoustic channel to more accurately reflect the signal reception quality of each intelligent agent under each spatial distribution in the time-varying underwater acoustic channel.

6. The method for cross-layer transmission of underwater acoustic network based on multi-agent reinforcement learning according to claim 5 is characterized in that: In the MARL-TS protocol, the action of agent i in the tth period is expressed as: a i (t)=[a i1 (t),a i2 (t),...,a iN (t)] (10); Each element a ij (t) represents the specific action of the jth time slot, assuming that P max is the upper limit of the transmission power, if 0<a ij (t)≤P max , indicating that the sending node is a ij (t) of power transfer; if a ij (t) = 0, indicating that the node is idle in this time slot and the transmission power is 0; At the tth period, agent i can obtain its individual observation value o i (t), as follows: o i (t)=[Q i (t),c i,1 (t),c i,2 (t),...,c i,N (t)] (11); Among them, {c i1 (t),c i2 (t),...,c iN (t)} represents the parameters describing the useful channels and interference channels between agent i and all sending nodes, a total of N channels; the data buffer size Q i The evolution of (t) is expressed as: Q i (t+1)=Q i (t)+I i (t)-S i (t) (12); In this formula, Q i (t+1) represents the data cache size of the sending node i at the beginning of the t+1th period, Q i (t) is the data cache size at the beginning of the tth period, I i (t) represents the number of signal packets flowing into the data cache during the tth period; when the signal transmission fails, the signal will be retained in the data cache for retransmission in the subsequent period; At the tth epoch, the environment state s(t) consists of the observations of all agents: s(t)=[o1(t),o2(t),...,o N (t)] (13); Through the strategy of collaborative cooperation among agents, all agents share a common reward signal so that the system encourages collaboration between agents and drives them to work together to achieve the optimization of the global goal. For the tth period, the shared reward r(t) is defined as: Where: ∑S i (t) represents the total number of signal packets successfully transmitted by all agents; E total is the total energy consumed by the agent during the tth period; weights w1 and w2 are used to balance the contributions of data caching and energy consumption.

7. The method for cross-layer transmission of underwater acoustic network based on multi-agent reinforcement learning according to claim 6 is characterized in that: The actor-critic method is used to construct the MARL-TS algorithm framework, and the MAPPO algorithm is used to update the policy network. By limiting the degree of policy updating, the stability of the training process is improved.

8. A cross-layer transmission system of underwater acoustic network based on multi-agent reinforcement learning, characterized in that: include: The sending node is a sensor located underwater and equipped with a data buffer; the receiving node is a buoy located on the surface of the water and uses hydroacoustic communication to transmit information with the underwater node; The central node is used to communicate with each receiving node and deploy the MARL-TS protocol, so that each receiving node is regarded as an intelligent agent, and the action of the intelligent agent in each time slot is defined as when the transmission link between it and the paired sending node is idle or transmitted at a specific power. Each intelligent agent makes independent action decisions in each period according to the untransmitted data of the paired sending node and the channel parameters of the corresponding transmission link, and performs power allocation and time slot resource scheduling under the time-varying channel to achieve cross-layer joint optimization of the physical layer and the MAC layer; The system operates using the underwater acoustic network cross-layer transmission method based on multi-agent reinforcement learning as described in claim 7.

9. An application of the underwater acoustic network cross-layer transmission method based on multi-agent reinforcement learning as described in claim 8 in a multi-node and wide-coverage underwater Internet of Things scenario.