Offshore communication cross-layer collaborative resource allocation method and system

By constructing a dynamic dual-threshold water injection model and a twin DRL service allocation model, cross-layer collaborative allocation of maritime communication resources is achieved, which solves the problem of resource utilization efficiency in traditional methods and improves the service quality and overall performance of maritime communication.

CN120456328APending Publication Date: 2025-08-08NAVAL AVIATION UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510658801.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional offshore wireless communication resource allocation methods are difficult to take into account global efficiency and individual service quality when dealing with strong dynamic scenarios and multi-type service concurrency. The existing solutions are difficult to achieve cross-layer collaboration between physical layer power control and network layer service scheduling, resulting in limited resource utilization efficiency.

Method used

A dynamic dual-threshold water injection model based on service quality perception and a twin DRL business allocation model based on domain knowledge enhancement are built. Through a two-dimensional adaptive adjustment mechanism and a two-channel feature decoupling mechanism, elastic control of power distribution and efficient coordination of business allocation are achieved.

Benefits of technology

It significantly improves the service quality of different service types in maritime communications, meets diversified QoS needs, and optimizes the overall performance and communication quality of cross-layer resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120456328A_ABST
    Figure CN120456328A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wireless communication, in particular to a cross-layer collaborative resource allocation method and system for maritime communication. The method comprises the following steps: constructing a dynamic dual-threshold water injection model based on service quality perception; constructing a twin DRL business distribution model based on domain knowledge enhancement; and allocating wireless network resources based on the dynamic dual-threshold water injection model and the twin DRL service allocation model. According to the method, the dynamic dual-threshold water injection model and the twin DRL service allocation model are constructed, so that flexible allocation of wireless network resources based on QoS perception enhancement is realized, different scene requirements of maritime communication are met, the resource allocation efficiency and quality in a maritime communication scene are improved, and the QoS requirements of various services are powerfully guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a method and system for cross-layer collaborative resource allocation for maritime communications. Background Art

[0002] Maritime wireless communications are core infrastructure for marine economic development and maritime security, supporting over 90% of global ocean trade and carrying out critical emergency rescue missions. However, unlike terrestrial communication networks, the marine environment presents unique technical challenges: complex and variable channel conditions and highly dynamic node distribution make traditional resource allocation methods difficult to apply.

[0003] At the dynamic topology level, various types of nodes (such as ships and aircraft) move continuously at different speeds, forming spatial heterogeneity in the horizontal and vertical dimensions. The wide spatial distribution differences directly affect the wireless propagation path. In terms of time-varying channel characteristics, the combined effects of sea surface wave reflection, near-sea evaporation waveguides, and atmospheric scattering and refraction constitute a unique marine wireless transmission channel. Especially in harsh sea conditions, the channel quality exhibits significant random fluctuations. The reality of limited resources further exacerbates the complexity of the problem. Sparsely deployed marine base stations face the dual constraints of energy supply and coverage range, and at the same time need to adapt to the strict power restrictions of various shipborne terminals.

[0004] In this context, intelligent resource allocation has become a core breakthrough for improving the efficiency of maritime communications: the resource constraints and competition caused by the sparse distribution of communication infrastructure may lead to a cliff-like drop in system throughput, while the different service quality requirements of various types of services such as emergency control and real-time video require methods with fine-grained differentiated processing capabilities. As a core dimension of wireless resource management, the optimization of power allocation has a dual strategic value for breaking through the above bottlenecks - it not only determines the energy efficiency of nodes, but also restricts the ability to ensure service QoS in dynamic scenarios. Especially under the dual constraints of strong dynamic topology and multi-scale channel attenuation, how to build an intelligent power allocation mechanism that adapts to the characteristics of the marine environment has become a key breakthrough for improving the robustness and transmission efficiency of maritime communication systems. It is urgent to establish a highly adaptable and low-complexity solution.

[0005] Although existing power allocation schemes have achieved certain results in improving system performance in specific scenarios, traditional schemes face performance bottlenecks when dealing with highly dynamic scenarios, especially in scenarios where multiple types of services are concurrent, making it difficult to balance global efficiency and individual service quality. The root cause is that there is a strong coupling between physical layer power control and network layer service scheduling - a single optimization dimension often leads to cross-layer performance imbalance. For example, the burst access requirements of high-priority services require the physical layer to dynamically adjust the power allocation strategy, while continuous video streaming transmission depends on a stable supply of channel resources. Existing fragmented optimization methods make it difficult to achieve cross-layer parameter coordination, resulting in limited resource utilization efficiency. At present, there are few intelligent resource allocation solutions for maritime scenarios that combine the physical layer and the network layer. Therefore, establishing a joint optimization mechanism for power and service resources has become an inevitable choice to improve the performance of communication systems in maritime scenarios, which is of great practical significance for ensuring high-quality maritime operations. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for cross-layer collaborative resource allocation for maritime communications in order to solve at least one of the above technical problems.

[0007] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0008] A cross-layer collaborative resource allocation method for maritime communications, comprising the following steps:

[0009] Construct a dynamic dual-threshold water injection model based on service quality perception;

[0010] Build a twin DRL business allocation model based on domain knowledge enhancement;

[0011] Wireless network resources are allocated based on the dynamic dual-threshold water injection model and the twin DRL service allocation model.

[0012] Furthermore, a dynamic dual-threshold water injection model based on service quality perception is constructed, including:

[0013] Constructing a wireless network node transmission model, determining a minimum quality of service requirement for each service type, and calculating, based on the wireless network node transmission model, a predicted transmission rate corresponding to a first preset score and a predicted transmission rate corresponding to a second preset score for the quality of service of the service currently being executed by the node; wherein the first preset score is higher than the second preset score;

[0014] The signal-to-noise ratio corresponding to the boundary is calculated based on the Shannon capacity formula;

[0015] Obtaining a user power allocation strategy for initializing service quality awareness;

[0016] Iterate the power allocation waterline with throughput as the waterline according to the channel gain;

[0017] Initialize the user power allocation strategy based on the real-time signal-to-noise ratio as the dynamic water level benchmark;

[0018] Subtract the power allocation strategies of the two waterlines to obtain the maximum difference and the minimum difference of the power allocation strategies;

[0019] Determining node power allocation based on the maximum difference in the power allocation strategy;

[0020] The power allocation of the end user is determined based on the remaining unallocated power of the base station and the node power allocation.

[0021] Furthermore, the predicted transmission rate corresponding to the first preset score and the predicted transmission rate corresponding to the second preset score are calculated by the following formula:

[0022]

[0023] in, Indicates the theoretical rate at which the service quality score of the current business executed by node i corresponds to the first preset score; I represents the theoretical rate at which the service quality score of the current business executed by node i corresponds to the second preset score; s Indicates the number of node connections; represents the average packet delay; J s Indicates packet jitter; L s Indicates the packet loss rate; Indicates the experience rate corresponding to the first preset score of the service quality score of the current business executed by node i; Indicates the experience rate corresponding to the second preset score of the service quality score of the current business executed by node i;

[0024] The signal-to-noise ratio corresponding to the boundary is calculated by the following formula:

[0025]

[0026] Among them, SNR i,max / min Indicates the maximum signal-to-noise ratio / minimum signal-to-noise ratio corresponding to the boundary; B i represents the bandwidth of node i;

[0027] The power allocation strategy for initializing service quality awareness is calculated using the following formula:

[0028]

[0029] in, Indicates the user power allocation strategy corresponding to the first preset score / the second preset score of the service quality score currently executed by node i; N 0,i represents the noise of node i; H(d i, f i ) represents the channel gain of node i, d i is the distance between node i and the base station, f i is the frequency used by node i;

[0030] The power allocation waterline with throughput as the waterline is calculated using the following formula:

[0031]

[0032] Among them, α i represents the node signal-to-noise ratio; μ represents the power allocation waterline with throughput as the waterline; P total represents the maximum allocatable power of the base station; i represents the i-th node of the channel;

[0033] The user power allocation strategy is calculated using the following formula:

[0034]

[0035] in, represents the user power allocation strategy of the i-th node; μ * Indicates the power allocation waterline with throughput as the waterline after multiple iterations;

[0036] The maximum and minimum power allocation strategy differences are calculated using the following formula:

[0037]

[0038] Where ΔP i,max The power allocation strategy difference corresponding to the first preset score of the service quality score of the business currently executed by node i, that is, the maximum power allocation strategy difference; Indicates the user power allocation strategy corresponding to the first preset score for the service quality score of the current execution of the business by node i; ΔP i,min The power allocation strategy difference between the service quality score of the business currently executed by node i and the second preset score, that is, the minimum power allocation strategy difference; Indicates the user power allocation strategy corresponding to the second preset score of the service quality score currently executed by node i;

[0039] The node power allocation is determined by the following formula:

[0040]

[0041] Among them, P i represents the node power allocation;

[0042] The power allocation to the end user is determined by the following formula:

[0043]

[0044] Where ΔP i,surplus Represents the remaining unallocated power ΔP surplus The power allocated to node i.

[0045] Furthermore, constructing a twin DRL business allocation model based on domain knowledge enhancement includes: constructing a twin network deep reinforcement learning model based on a dual-channel feature encoder and constructing a strategy optimization filter based on channel state perception.

[0046] Furthermore, a twin network deep reinforcement learning model based on a dual-channel feature encoder is constructed, including:

[0047] The node motion state vector, which consists of motion coordinates, instantaneous speed, instantaneous angle, and service type, is mapped to a relative coordinate system through coordinate normalization and fully connected layer feature embedding to generate a low-dimensional feature vector, thereby encoding and extracting node features.

[0048] For the service demand vector composed of delay-sensitive flags, packet delay thresholds, packet loss thresholds, data volume, and service type, logarithmic data compression and weighted aggregation using a gated attention mechanism are used to achieve service demand feature encoding and adaptive selection.

[0049] Constructing a dual-branch twin network with shared initial parameters and asynchronous updates; wherein the dual-branch twin network includes a main network and a twin branch network;

[0050] The encoded dual-channel feature vectors are respectively input into the weight-sharing main network and the twin branch network to extract features and calculate the business and node embedding similarity values, and the node allocation strategy based on business features is output in this order.

[0051] Furthermore, coordinate normalization is performed using the following formula:

[0052]

[0053] Among them, p t Indicates absolute position; p base represents the base station coordinates; ||·||2 represents the L2 norm; ε represents the numerical stability constant; Indicates p t Position in the relative coordinate system;

[0054] Generate low-dimensional feature vectors using the following formula:

[0055]

[0056] Among them, h nrepresents the low-dimensional feature vector; σ(·) represents the ReLU activation function; W n represents the customer training weight matrix; v t represents the instantaneous speed; θ t Indicates the instantaneous angle; e t Indicates the type of business the node is executing; b n represents a trainable bias vector;

[0057] Logarithmic compression is performed using the following formula:

[0058]

[0059] Wherein, m represents the amount of business data; Indicates the amount of compressed business data;

[0060] The gated attention mechanism weighted aggregation is performed using the following formula:

[0061]

[0062] Among them, h b represents fusion features; W b represents a trainable weight matrix; α represents the attention weight; τ represents the delay sensitivity flag; d represents the packet delay threshold; l represents the packet loss threshold; q represents the service type;

[0063] The business and node embedding similarity values are calculated using the following formula:

[0064]

[0065] Among them, L siamese (φ) represents the embedding similarity value; K represents the number of nodes; N represents the number of services; f φ (s i ) represents the eigenvalue of node i; f φ (s j ) represents the eigenvalue of node j.

[0066] Furthermore, a strategy optimization filter based on channel state perception is constructed, including:

[0067] By combining the node's measured transmission rate, the system's maximum rate benchmark, the service data volume requirements, and the node's channel quality coefficient, the node's adaptability to different service types is calculated, and a channel state-driven conflict resolution matrix is constructed.

[0068] Based on the conflict resolution matrix, the initial allocation probability matrix output by the two-branch twin network is reconstructed to obtain a new allocation probability matrix; the allocation probability of each type of business is normalized to ensure that the total allocation probability of each type of business is 1, and the strategy optimization filter construction is completed.

[0069] Furthermore, the adaptability of a node to different service types is calculated using the following formula:

[0070]

[0071] Among them, c nq (r) represents the adaptability of node n to service q; R n (t) represents the measured transmission rate of node n at time t; R max Indicates the current system maximum rate benchmark; D q Indicates the data volume requirement of business type q; Indicates the average data volume required for all business types; ξ n (t) represents the channel quality coefficient of node i;

[0072] The probability of allocating services to each type is normalized using the following formula:

[0073]

[0074] in, represents the probability of node i assigning service m; Represents normalized

[0075] A cross-layer collaborative resource allocation system for maritime communications, comprising:

[0076] The first model building module is used to build a dynamic dual-threshold water injection model based on service quality perception;

[0077] The second model building module is used to build a twin DRL business allocation model based on domain knowledge enhancement;

[0078] An allocation module is used to allocate wireless network resources based on the dynamic dual-threshold water injection model and the twin DRL service allocation model.

[0079] An electronic device comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the method for cross-layer collaborative resource allocation for maritime communications as described above is implemented.

[0080] The beneficial effects of the present invention are:

[0081] The present invention allocates wireless network resources based on a constructed dynamic dual-threshold water injection model and a twin DRL service allocation model.

[0082] The dynamic dual-threshold water-injection model overcomes the limitations of classic water-injection algorithms by innovatively introducing a two-dimensional adaptive adjustment mechanism. On the one hand, a dynamic water-level benchmark is established based on the real-time signal-to-noise ratio (SINR), enabling power allocation to adapt to channel changes in real time. On the other hand, a dynamic water-level benchmark is established based on quality of service (QoS). By designing a dual-water-level joint mapping function, flexible control of power allocation is achieved. This approach significantly improves service quality while sacrificing a small amount of system throughput, meeting the diverse QoS requirements of different service types in maritime communications.

[0083] The Twin DRL service allocation model addresses the dynamic optimization challenge of service-node matching by proposing a dual-channel feature decoupling mechanism and constructing a dual-channel feature encoder to accurately extract node motion features and service demand features, effectively solving the challenge of integrating heterogeneous feature spaces. Furthermore, a policy optimization filter based on channel state perception is designed, and a dynamic policy correction mechanism is introduced. This allows high-demand services to reconstruct allocation strategies using a conflict resolution matrix, achieving efficient collaborative mapping of heterogeneous feature spaces. This significantly improves the accuracy and rationality of service allocation and optimizes the overall performance of cross-layer collaborative resource allocation for maritime communications. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 This is a flow chart of a method for cross-layer collaborative resource allocation for maritime communications according to an embodiment of the present invention;

[0085] Figure 2 This is a principle block diagram of a cross-layer collaborative resource allocation method for maritime communications based on QoS awareness enhancement according to an embodiment of the present invention;

[0086] Figure 3 A comparison chart of QoS scores of different technical methods according to an embodiment of the present invention;

[0087] Figure 4 A comparison chart of energy consumption efficiency of round systems using different technical methods according to one embodiment of the present invention;

[0088] Figure 5 A comparison chart of round business completion status of different technical methods in one embodiment of the present invention;

[0089] Figure 6 A comparison chart of different types of business completion status of different technical methods in one embodiment of the present invention;

[0090] Figure 7 This is a schematic diagram of a cross-layer collaborative resource allocation system for maritime communications according to an embodiment of the present invention. DETAILED DESCRIPTION

[0091] The present invention will now be discussed with reference to exemplary embodiments. It should be understood that the embodiments discussed are only intended to enable those skilled in the art to better understand and implement the present invention, rather than to imply any limitation on the scope of the present invention.

[0092] As used herein, the term "including" and variations thereof are to be interpreted as open-ended terms meaning "including, but not limited to." The term "based on" is to be interpreted as "based, at least in part, on." The terms "one embodiment" and "an embodiment" are to be interpreted as "at least one embodiment."

[0093] The present invention proposes a cross-layer collaborative resource allocation method for maritime communications, which allocates unlimited network resources based on QoS awareness enhancement (Dual-domain Collaborative Joint Resource Allocation, DCJRA); the method mainly includes two stages: first, constructing a dynamic dual-threshold water injection method based on service quality awareness; second, constructing a twin DRL business allocation method based on domain knowledge enhancement.

[0094] In the first phase, a two-dimensional adaptive adjustment mechanism was introduced into the classic water injection algorithm framework: on the one hand, a dynamic water level benchmark based on the real-time signal-to-noise ratio (SINR) was constructed, and on the other hand, a dynamic water level benchmark for QoS was constructed. By designing a dual-water level joint mapping function, flexible control of power allocation was achieved, sacrificing a small amount of system throughput in exchange for improved service quality.

[0095] In the second phase, in response to the dynamic optimization needs of service-node matching, a dual-channel feature decoupling mechanism is proposed, and a dual-channel feature encoder is constructed to extract node motion features and service demand features respectively. A policy optimization filter based on channel state perception is designed, and a policy dynamic correction mechanism is introduced to allow high-demand category services to use conflict resolution matrices to reconstruct allocation strategies, thereby realizing collaborative mapping of heterogeneous feature spaces.

[0096] Example 1

[0097] Figure 1 FIG1 is a flow chart of a cross-layer collaborative resource allocation method for maritime communications according to an embodiment of the present invention. Figure 1 As shown, according to one embodiment of the present invention, a method for cross-layer collaborative resource allocation for maritime communications, which allocates unlimited network resources based on QoS awareness enhancement, includes the following steps:

[0098] Step S02: constructing a dynamic dual-threshold water injection model based on service quality perception;

[0099] Step S04: constructing a twin DRL business allocation model based on domain knowledge enhancement;

[0100] Step S06: Allocate wireless network resources based on the dynamic dual-threshold water injection model and the twin DRL service allocation model.

[0101] In this embodiment, a cross-layer collaborative resource allocation method for maritime communications is proposed. The core of the method is to realize wireless network resource allocation based on QoS perception enhancement. Specifically, the method first constructs a dynamic dual-threshold water injection model based on service quality perception. This model can flexibly adjust the resource allocation threshold according to the dynamic changes in service quality to adapt to the resource requirements in different scenarios; then constructs a twin deep reinforcement learning (DRL) business allocation model based on domain knowledge enhancement. This model integrates domain knowledge into the deep reinforcement learning framework and uses the twin structure to improve the accuracy and stability of the model's business allocation decisions; finally, the dynamic dual-threshold water injection model and the twin DRL business allocation model are comprehensively utilized to collaboratively complete the wireless network resource allocation task, aiming to improve the efficiency and quality of resource allocation in maritime communication scenarios and ensure the QoS requirements of various services.

[0102] By constructing a dynamic dual-threshold water injection model and a twin DRL service allocation model, the present invention realizes flexible allocation of wireless network resources based on QoS perception enhancement to adapt to the requirements of different maritime communication scenarios, improves the efficiency and quality of resource allocation in maritime communication scenarios, and effectively guarantees the QoS requirements of various services.

[0103] According to one embodiment of the present invention, step S02 includes:

[0104] Constructing a wireless network node transmission model, determining a minimum quality of service requirement for each service type, and calculating, based on the wireless network node transmission model, a predicted transmission rate corresponding to a first preset score and a predicted transmission rate corresponding to a second preset score for the quality of service of the service currently being executed by the node; wherein the first preset score is higher than the second preset score;

[0105]

[0106] in, Indicates the theoretical rate at which the service quality score of the current business executed by node i corresponds to the first preset score; I represents the theoretical rate at which the service quality score of the current business executed by node i corresponds to the second preset score; s Indicates the number of node connections; represents the average packet delay; J s Indicates packet jitter; L s Indicates the packet loss rate; Indicates the experience rate corresponding to the first preset score of the service quality score of the current business executed by node i; Indicates the experience rate corresponding to the second preset score of the service quality score of the current business executed by node i;

[0107] The signal-to-noise ratio corresponding to the boundary is calculated based on the Shannon capacity formula;

[0108]

[0109] Among them, SNR i,max / min Indicates the maximum signal-to-noise ratio / minimum signal-to-noise ratio corresponding to the boundary; B i represents the bandwidth of node i;

[0110] Get the user power allocation strategy P for initializing service quality awareness QoS ;

[0111]

[0112] in, Indicates the user power allocation strategy corresponding to the first preset score / the second preset score of the service quality score currently executed by node i; N 0,i represents the noise of node i; H(d i , f i ) represents the channel gain of node i, d i is the distance between node i and the base station, f i is the frequency used by node i;

[0113] Iterate the power allocation waterline with throughput as the waterline according to the channel gain;

[0114]

[0115] Among them, α i represents the node signal-to-noise ratio; μ represents the power allocation waterline with throughput as the waterline; P total represents the maximum allocatable power of the base station; i represents the i-th node of the channel;

[0116] Initialize the user power allocation strategy P with real-time signal-to-noise ratio as the dynamic water level benchmark SNR ;

[0117]

[0118] in, represents the user power allocation strategy of the i-th node; μ * Indicates the power allocation waterline with throughput as the waterline after multiple iterations;

[0119] Subtract the power allocation strategies of the two waterlines to obtain the maximum difference and the minimum difference of the power allocation strategies;

[0120]

[0121] Where ΔP i,max The power allocation strategy difference corresponding to the first preset score of the service quality score of the business currently executed by node i, that is, the maximum power allocation strategy difference; Indicates the user power allocation strategy corresponding to the first preset score for the service quality score of the current execution of the business by node i; ΔP i,min The power allocation strategy difference between the service quality score of the business currently executed by node i and the second preset score, that is, the minimum power allocation strategy difference; Indicates the user power allocation strategy corresponding to the second preset score of the service quality score currently executed by node i;

[0122] Determine node power allocation based on the maximum difference in power allocation strategy;

[0123]

[0124] Among them, P i represents the node power allocation;

[0125] Determine the power allocation for the end user based on the remaining unallocated power of the base station and the node power allocation;

[0126]

[0127] Where ΔP i,surplus Represents the remaining unallocated power ΔP surplus The power allocated to node i;

[0128] By dividing the remaining unallocated power of the base station by ΔP surplus Proportional to ΔP i,max Nodes with a value less than 0 determine the power allocation for the end user.

[0129] In this embodiment, the content of step S02 is further explained: first, a wireless network node transmission model is constructed and the minimum QoS requirements for each service type are determined. Based on this model, the predicted transmission rate corresponding to the node's current service execution at different preset scores (the first preset score is higher than the second preset score) is calculated, including the theoretical rate and the experienced rate; then, the Shannon capacity formula is used to calculate the signal-to-noise ratio corresponding to the boundary; then, the user power allocation strategy for initializing QoS perception is obtained; then, the power allocation waterline with throughput as the waterline is obtained based on the channel gain iteration, and the user power allocation strategy with real-time signal-to-noise ratio as the dynamic waterline benchmark is initialized; the power allocation strategies of the two waterlines are then subtracted to obtain the maximum and minimum difference between the power allocation strategies; finally, the node power allocation is determined based on the maximum difference, and the final user power allocation is determined in combination with the remaining unallocated power of the base station, and the remaining power is proportionally allocated to the nodes with a power difference less than 0. The specific proportional allocation is set according to actual needs. For example, the remaining unallocated power of the base station is evenly distributed to ΔP i,max Nodes that are less than 0.

[0130] The present invention can accurately and dynamically and reasonably allocate user power according to the service quality requirements of the node's current business, fully utilize the remaining power of the base station, and ensure efficient and stable power distribution in different business scenarios, thereby ensuring the rationality and effectiveness of wireless network resource allocation in maritime communications and improving communication quality.

[0131] According to one embodiment of the present invention, step S02 includes:

[0132] Step S022: constructing a Siamese network deep reinforcement learning model based on a dual-channel feature encoder, including:

[0133] Step 1: The node motion state vector, consisting of motion coordinates, instantaneous speed, instantaneous angle, and service type, is mapped to a relative coordinate system through coordinate normalization and fully connected layer feature embedding to generate a low-dimensional feature vector, thereby encoding and extracting node features.

[0134] Step 2: The service demand vector, which consists of the delay sensitivity flag, packet delay threshold, packet loss threshold, data volume, and service type, is subjected to logarithmic data compression and weighted aggregation using a gated attention mechanism to achieve service demand feature encoding and adaptive selection.

[0135] Step 3: Build a dual-branch twin network with shared initial parameters and asynchronous updates; the dual-branch twin network includes a main network and a twin branch network;

[0136] In step 4, the encoded dual-channel feature vectors are input into the weight-sharing main network and the twin branch network respectively to extract features and calculate the business and node embedding similarity values, and then sort and output the node allocation strategy based on business features.

[0137] Preferably, coordinate normalization is performed using the following formula:

[0138]

[0139] Among them, p t Indicates absolute position; p base represents the base station coordinates; ||·||2 represents L 2 norm; ε represents the numerical stability constant; Indicates p t Position in the relative coordinate system;

[0140] Generate low-dimensional feature vectors using the following formula:

[0141]

[0142] Among them, h n represents the low-dimensional feature vector; σ(·) represents the ReLU activation function; W n represents the customer training weight matrix; v t represents the instantaneous speed; θ t Indicates the instantaneous angle; e t Indicates the type of business the node is executing; b n represents a trainable bias vector;

[0143] Logarithmic compression is performed using the following formula:

[0144]

[0145] Wherein, m represents the amount of business data; Indicates the amount of compressed business data;

[0146] The gated attention mechanism weighted aggregation is performed using the following formula:

[0147]

[0148] Among them, h b represents fusion features; W b represents the trainable weight matrix; α represents the attention weight, Among them, w α represents the initial attention weight, (·) T represents transposition; τ represents delay sensitivity flag; d represents packet delay threshold; l represents packet loss threshold; q represents service type;

[0149] The business and node embedding similarity values are calculated using the following formula:

[0150]

[0151] Among them, L siamese (φ) represents the embedding similarity value; K represents the number of nodes; N represents the number of services; f φ (s i ) represents the eigenvalue of node i; f φ (s j ) represents the eigenvalue of node j.

[0152] Step S024, constructing a policy optimization filter based on channel state perception, including:

[0153] By combining the node's measured transmission rate, the system's maximum rate benchmark, the service data volume requirements, and the node's channel quality coefficient, the node's adaptability to different service types is calculated, and a channel state-driven conflict resolution matrix is constructed.

[0154] Based on the conflict resolution matrix, the initial allocation probability matrix output by the two-branch twin network is reconstructed to obtain a new allocation probability matrix; the allocation probability of each type of business is normalized to ensure that the total allocation probability of each type of business is 1, and the strategy optimization filter is constructed.

[0155] Preferably, the adaptability of a node to different service types is calculated using the following formula:

[0156]

[0157] Among them, c nq (t) represents the adaptability of node n to service q; R n (t) represents the measured transmission rate of node n at time t; R max Indicates the current system maximum rate benchmark; D q Indicates the data volume requirement of business type q; Indicates the average data volume required for all business types; ξ n (t) represents the channel quality coefficient of node i;

[0158] The new allocation probability matrix is obtained by the following formula:

[0159]

[0160] Among them, P0 represents the initial allocation probability matrix output by the twin network, C t represents the conflict resolution matrix driven by the channel state, that is, the conflict resolution matrix based on channel state perception, in, represents a real number matrix of dimension N0×M0; N0 represents the number of nodes, and M0 represents the number of service types; Represents the new allocation probability matrix after strategy reconstruction;

[0161] The probability of allocating services to each type is normalized using the following formula:

[0162]

[0163] in, represents the probability of node i assigning service m; Represents normalized

[0164] In this embodiment, step S02 is further explained: first, a twin network deep reinforcement learning model based on a dual-channel feature encoder is constructed, specifically including coordinate normalization of the node motion state vector, fully connected layer feature embedding to realize node feature encoding and extraction, logarithmic compression of the business demand vector and weighted aggregation of the gated attention mechanism to realize business demand feature encoding and adaptive selection, and a dual-branch twin network with shared initial parameters and asynchronous updates is constructed. The encoded dual-channel feature vector is input into the network to extract features, calculate the similarity value between the business and node embedding, and output the node allocation strategy; then, a policy optimization filter based on channel state perception is constructed, and the conflict resolution matrix is constructed by calculating the adaptability based on the measured transmission rate of the node, the maximum rate benchmark of the system, the business data volume requirement and the node channel quality coefficient. Based on the matrix, the initial allocation probability matrix is reconstructed to obtain a new matrix, and the probability of each type of business allocation is normalized to complete the construction of the policy optimization filter.

[0165] The present invention can fully extract the characteristics of nodes and business requirements, and reasonably output node allocation strategies based on these characteristics; at the same time, the policy optimization filter based on channel state perception further optimizes the allocation strategy, making resource allocation more in line with the actual channel state, effectively improving the rationality and accuracy of wireless network resource allocation in maritime communications, and ensuring the stability and efficiency of communications.

[0166] Example 2

[0167] Figure 2 This is a principle block diagram of a cross-layer collaborative resource allocation method for maritime communications based on QoS awareness enhancement according to an embodiment of the present invention. Figure 2 As shown, according to an embodiment of the present invention, a method for cross-layer collaborative resource allocation for maritime communications allocates unlimited network resources based on QoS awareness enhancement; the method includes the following steps:

[0168] Step S10, constructing a dynamic dual-threshold water injection model based on service quality perception;

[0169] Traditional waterfilling algorithms typically aim to maximize system capacity or increase transmission rates, neglecting the importance of Quality of Service (QoS) constraints to ensure that minimum user or service requirements are met. With QoS constraints incorporated, power allocation must not only optimize channel gain but also ensure that the minimum requirements of different services are met.

[0170] Step S10 specifically includes:

[0171] Step S101: construct a wireless network node transmission model to determine the minimum QoS requirements (such as minimum SNR, throughput, delay, etc.) for each service type; and calculate the predicted transmission rate corresponding to the QoS of the service currently executed by the node at 90 points and 60 points based on the wireless network node transmission model.

[0172]

[0173] in, I represents the theoretical rate corresponding to the service quality scores of 90 and 60 points for the current business executed by node i. Since various factors in the real world affect data transmission, the theoretical rate is always greater than or equal to the experienced rate; I s Indicates the number of node connections; represents the average packet delay; J s Indicates packet jitter; L s Indicates the packet loss rate; Indicates the experience rate corresponding to the service quality score of 90 points and 60 points for the current execution of the business by node i, representing the upper and lower limits of the transmission rate experience value; packet delay

[0174] Step S102: Calculate the signal-to-noise ratio corresponding to the boundary based on the Shannon capacity formula:

[0175]

[0176] Among them, SNR i,max / min Indicates the signal-to-noise ratio corresponding to the boundary; B i represents the bandwidth of node i;

[0177] Step S103: Initialize the user power allocation strategy P for quality of service awareness Qos

[0178]

[0179] in, Indicates the user power allocation strategy corresponding to the service quality scores of 90 and 60 points currently executed by node i; N 0,i represents the noise of node i; H(d i ,f i) represents the channel gain of node i, d i is the distance between node i and the base station, f i is the frequency used by node i.

[0180] Step S104: According to the channel gain H(d i , f i ) Iterate the power allocation waterline μ with throughput as the waterline;

[0181]

[0182] Among them, α i represents the node signal-to-noise ratio; μ represents the power allocation waterline with throughput as the waterline; P total represents the maximum allocatable power of the base station; i represents the i-th node of the channel.

[0183] Step S105: Initialize the user power allocation strategy P with the real-time signal-to-noise ratio as the dynamic water level benchmark. SNR

[0184]

[0185] in, represents the user power allocation strategy of the i-th node; μ * Indicates the power allocation waterline with throughput as the waterline after multiple iterations;

[0186] Step S106: Subtract the power allocation strategies of the two waterlines to obtain the maximum power allocation strategy difference ΔP i,max and the minimum difference ΔP of the power allocation strategy i,min :

[0187]

[0188] Where ΔP i,max Indicates the power allocation strategy difference corresponding to the service quality score of 90 points for the current business executed by node i; Indicates the user power allocation strategy corresponding to the service quality score of 90 points for the current execution of the business by node i; ΔP i,min Indicates the power allocation strategy difference corresponding to the service quality score of 60 points for the current business executed by node i; Indicates the user power allocation strategy corresponding to the service quality score of 60 points currently executed by node i;

[0189] Step S107: Determine the node power allocation P according to the positive or negative value of the difference. i

[0190]

[0191] Among them, Pi represents the node power allocation;

[0192] Step S108: Ensure that the total power allocation does not exceed the maximum allocable power P of the base station. total , if the base station still has remaining unallocated power ΔP surplus , then it is distributed proportionally to ΔP i,max For nodes with <0, the power allocation for the end user is:

[0193]

[0194] Where ΔP i,surplus Indicates ΔP surplus The power allocated to node i;

[0195] Step S20, constructing a dynamic decision architecture model;

[0196] A domain knowledge-enhanced Siamese DRL (DKES-DRL) business allocation method is proposed to construct a dynamic decision architecture model by structurally embedding maritime communication prior knowledge.

[0197] The dynamic decision-making architecture model consists of two parts: a twin network deep reinforcement learning model based on a dual-channel feature encoder and a policy optimization filter based on channel state perception.

[0198] Step S20 specifically includes:

[0199] Step S201, constructing a Siamese network deep reinforcement learning model based on a dual-channel feature encoder;

[0200] In the knowledge-enhanced twin DRL architecture, the dual-channel feature encoder transforms the heterogeneous information in the original observation space into a parsable feature representation through a structured data preprocessing mechanism. Specifically, the following data preprocessing process is designed to address the heterogeneity of node motion features and business demand features:

[0201] Step 1: Encode node features;

[0202] Define the node motion state vector as s n =[p t , v t ,θ t , e t ],

[0203] Among them, p t represents the motion coordinates, is a three-dimensional real space; v t Indicates instantaneous speed, Indicates instantaneous speed, is a set of positive real numbers; θ t represents the instantaneous angle, θ t ∈[-0.25π, 0.25π]; indicates the type of business the node is executing, e t ∈{1, 2, ..., 6}.

[0204] Feature extraction is achieved through the following steps:

[0205] (1) Coordinate normalization: Mapping the absolute position to a relative coordinate system centered on the base station;

[0206]

[0207] Among them, p t Indicates absolute position; p base represents the base station coordinates; ||·||2 represents L 2 norm; ε represents the numerical stability constant; Indicates p t Position in the relative coordinate system;

[0208] (2) Feature embedding: Generate low-dimensional feature vectors through fully connected layers;

[0209]

[0210] Among them, h n represents the low-dimensional feature vector; σ(·) represents the ReLU activation function; W n represents the customer training weight matrix; b n represents a trainable bias vector;

[0211] Step 2: Encode the business demand characteristics;

[0212] Define the business demand vector as s b =[τ, d, l, m, q];

[0213] Where τ represents the delay sensitivity flag, τ∈{0,1}; d represents the packet delay threshold; l represents the packet loss threshold; m represents the service data volume, is a set of positive real numbers; q represents the business type, is a set of non-negative integers.

[0214] Feature preprocessing is achieved through the following steps:

[0215] (1) If the business data volume is large, logarithmic compression is performed;

[0216]

[0217] Wherein, m represents the amount of business data; Indicates the amount of compressed business data;

[0218] (2) Feature fusion:

[0219] Adopting gated attention mechanism weighted aggregation to achieve adaptive feature selection;

[0220]

[0221] Among them, h b represents fusion features; W b represents the trainable weight matrix; α represents the attention weight, Among them, w α represents the initial attention weight, (·) T represents transpose;

[0222] Step 3: Build a dual-branch network structure, including the main network Q θ With the twin branch network Q φ , the two initially share parameters (φ←θ), and the update of the two network parameters adopts asynchronous update method.

[0223] The encoded dual-channel feature vectors are input into the weight-sharing twin network model for feature extraction, and the embedding similarity value of each service feature and the node to be connected is calculated:

[0224]

[0225] Among them, L siamese (φ) represents the embedding similarity value; K represents the number of nodes; N represents the number of services; f φ (s i ) represents the eigenvalue of node i; f φ (s j ) represents the eigenvalue of node j;

[0226] Sorting by similarity values finally outputs a node allocation strategy based on business characteristics.

[0227] Step S202: constructing a policy optimization filter based on channel state perception;

[0228] Step 1: Construct a channel state-driven conflict resolution matrix;

[0229] Define CCRM as Represents a real number matrix of dimension N0×M0, where N0 is the number of nodes and M0 is the number of service types; the matrix element c nm (t) represents the adaptability of node n to service m, which is calculated as follows:

[0230]

[0231] Among them, c nq (t) represents the adaptability of node n to service q; R n (t) represents the measured transmission rate of node n at time t; R max Indicates the current system maximum rate benchmark, R max =max{R1(t), ..., R N (t)};D q Indicates the data volume requirement of business type q; represents the average data volume demand of all business types; ξ n (t) represents the channel quality coefficient of node i, SINR i (t) represents the signal-to-noise ratio of node i, SINR th Indicates the demodulation threshold;

[0232] Step 2: Construct a strategy optimization filter;

[0233] Initial assignment probability matrix P0∈[0,1] based on the output of the twin network N×M ,First, the initial strategy is reconstructed through CCRM;

[0234]

[0235] Among them, P0 represents the initial allocation probability matrix output by the twin network, C t represents the conflict resolution matrix driven by the channel state, that is, the conflict resolution matrix based on channel state perception, in, represents a real number matrix of dimension N0×M0, where N0 represents the number of nodes and M0 represents the number of service types; Represents the new allocation probability matrix after strategy reconstruction;

[0236] Then the allocation probability of each type of business is normalized to ensure that the total allocation probability of each type of business is 1;

[0237]

[0238] in, represents the probability of node i assigning service m; Represents normalized

[0239] Repeat step S20 until the end.

[0240] The dynamic dual-threshold water injection model based on service quality perception proposed in this paper introduces a two-dimensional adaptive adjustment mechanism into the classic water injection algorithm framework: on the one hand, a dynamic water level benchmark based on real-time signal-to-noise ratio (SINR) is established, and on the other hand, a dynamic water level benchmark for QoS is constructed. By designing a dual-water level joint mapping function, flexible control of power allocation is achieved, thereby sacrificing a small part of system throughput in exchange for improved service quality.

[0241] The twin DRL service allocation method proposed in this invention, which is based on domain knowledge enhancement, proposes a dual-channel feature decoupling mechanism to meet the dynamic optimization requirements of service-node matching, and constructs a dual-channel feature encoder to extract node motion features and service demand features respectively; designs a policy optimization filter based on channel state perception, introduces a policy dynamic correction mechanism, and allows high-demand category services to use conflict resolution matrices to reconstruct allocation strategies, thereby realizing collaborative mapping of heterogeneous feature spaces.

[0242] Example 3

[0243] In this embodiment, a simulation analysis is performed using a wireless network at sea as a scenario. Figure 2 As shown, the cross-layer collaborative resource allocation method for maritime communication (i.e. DCJRA) of the present invention is adopted, and the relevant parameters are set as follows:

[0244] The number of agents is 1, the agent speed is [0, 5] m / s, the number of users is 8, the user node speed is [20, 50] m / s, the node turning angle is [-0.25π, 0.25π], the mission duration is 600s, the agent transmission power is 300W, the test scenario channel interference is 10±4dB, the node spectrum range is 2000-2100MHz, the neural network learning rate is 1e-4, the number of network layers is 3, the number of neurons is 300 / 150 / 8, the sampling time is 1s, the discount factor is 0.95, the batch sample size is 600, and the experience buffer length is 1.2*10 -5 , the update times (training / testing) are 50 / 50. Under the above simulation conditions, the present invention is simulated and verified according to specific modulation and demodulation steps.

[0245] like Figure 2 As shown, resource allocation is performed at both the physical and network layers. At the physical layer, a QoS-aware dynamic dual-threshold water injection algorithm generates a power allocation strategy, improving service quality at the expense of system throughput. At the network layer, environmental parameters are fed into a domain-knowledge-enhanced twin deep reinforcement learning model to output a network-layer service allocation strategy. These two resource allocation strategies interact with the environment to collect rewards and calculate task losses. The strategies are then iterated and updated, repeating the cycle until the task completes.

[0246] Based on the simulation results, the performance differences between the present invention and the prior art are compared and analyzed.

[0247] Among them, the cross-layer collaborative resource allocation method for maritime communication of the present invention is referred to as DDW-DKES. The existing power allocation methods for technical effect comparison with the present invention include: average allocation algorithm (abbreviated as Average), traditional water filling algorithm (abbreviated as WaterFill) [Zhang Dongmei, Xu Youyun, Cai Yueming. Linear water filling power allocation algorithm in OFDMA system [J]. Journal of Electronics and Information Technology, 2007, 29(6): 1286-1289.] and genetic algorithm (abbreviated as Genetic) [Lu Yin, Wang Chenggong, Wang Huiru, et al. A NOMA power allocation method and process based on genetic algorithm: CN201811038383[P]. 2021-02-12.], DQN algorithm for service allocation (abbreviated as DQN) [MNIH V, KAVUKCUOGLU K, SILVER D, etc. Playing atari with deep reinforcement learning [C] / / Neural Information Processing Systems. 2015: 1-9.], DDPG algorithm (abbreviated as DDPG) [LILLICRAP TP, HUNT JJ, PRITZEL A, etc. Continuous control with deep reinforcement learning [C / OL] / / International Conference on Learning Representations (ICLR). 2016. http: / / arxiv.org / abs / 1509.02971.] and Siamese neural network algorithm (abbreviated as Siamese) [BROMLEY J, GUYON I, LECUN Y, etc. Signature verification using a “Siamese” time delay neural network [C] / / Neural Information Processing Systems. 737-744.].

[0248] During the experiment, we used a pairwise combination method for comparison and verification. A total of 7 combinations were designed, namely WaterFill-DQN, Genetic-DQN, Average-DQN, DDW-DQN, DDW-DDPG, DDW-DKES, and DDW-Siamese.

[0249] The simulation scenario is set within an area of radius R, encompassing the sea surface and its surroundings. Within this area resides one low-speed base station node and n freely moving high-speed user nodes, all of which undergo continuous random motion. The current area is assumed to be within the base station's communication range. This means that nodes entering the area are connected directly and exclusively to the base station node, and are disconnected upon leaving the area. To account for distance differences between nodes due to altitude, the communication frequency bands are selected from satellite communication bands. To enhance randomness, the initial positions, velocities, and angles of the nodes within the area are randomly generated.

[0250] The technical effects were evaluated from three aspects: QoS score, energy efficiency performance, and number of completed services. The results are as follows:

[0251] QoS score: Figure 3 This is a comparison chart of QoS scores of different technical methods in one embodiment of the present invention. Figure 3 The comparison of the average node QoS scores of various methods is shown. The horizontal axis represents simulation time, with the number of test steps increasing as simulation time increases. A horizontal comparison reveals that the average QoS score of the proposed framework's nodes fluctuates continuously due to environmental randomness, but the mean square error (MSD) remains within a controllable range. Other methods experience dramatic fluctuations in average scores exceeding 5 points within a certain time period, demonstrating the robustness of the proposed method. A vertical comparison, through ablation experiments comparing WaterFill-DQN and DDW-DQN, shows that the dynamic dual-threshold water filling algorithm improves QoS scores by 9.51% compared to the traditional water filling method. This is because the proposed method stops allocating power after a node meets the highest QoS score for its service, significantly improving the channel environment of the current node compared to other nodes. This strategy tends to allocate remaining power to other nodes whose services are still underserved, but this "weak" allocation strategy comes at the cost of reduced overall system throughput. At the same time, the average service quality score of the proposed framework always maintains a high level, which is 3.25% higher than that of DDW-Siamese. This is due to the dual-channel feature encoder and policy optimization filter that allocates services that better match the node status to nodes, reducing the situation where the power allocation range is smaller than the service QoS score range.

[0252] Energy efficiency performance: Figure 4 This is a comparison chart of energy consumption efficiency of different technical methods of a round system in one embodiment of the present invention. Figure 4The energy efficiency comparison chart of each method is presented. Since both the traditional water filling algorithm and the proposed dynamic dual-threshold water filling algorithm almost fully allocate the maximum power of the base station, the energy efficiency comparison is actually a comparison of system throughput. As can be seen, the proposed power allocation algorithm sacrifices 9.3% of system throughput to improve nodes with low QoS scores, a price that is acceptable. Because the genetic algorithm has difficulty solving high-dimensional problems, its performance is inferior to the average power allocation algorithm. Furthermore, both algorithms are subject to significant fluctuations in dynamic environments and have low robustness.

[0253] Number of completed businesses: Figure 5 This is a comparison chart of the completion of round business of different technical methods in one embodiment of the present invention; wherein, Figure 5 (A) is the comparison chart of the completed quantity. Figure 5 (B) is the completion value comparison chart. Figure 6 A comparison chart of different types of business completion status of different technical methods in one embodiment of the present invention; Figure 6 (A) is a comparison chart of the completion status of business type 1. Figure 6 (B) is a comparison chart of the completion status of business type 2. Figure 6 (C) is a comparison chart of the completion status of business type 3. Figure 6 (D) is a comparison chart of the completion status of business type 4. Figure 6 (E) is a comparison chart of the completion status of business type 5. Figure 6 (F) is a comparison chart of the completion status of business type 6, where business types 1-6 are different business types. Figure 5 The comparison chart of the number of completed services of each method framework is shown. It can be seen that when the system throughput of the proposed method is lower than that of the traditional water injection algorithm, the number of completed services is still relatively advantageous, and the average number of completed services exceeds that of the traditional water injection algorithm by 1.3%. This is because under the proposed twin DRL service allocation method, services are allocated to nodes with appropriate status intervals according to their service quality requirements, that is, nodes with good channel status are allocated to services with high quality requirements and large data volumes, and vice versa. By making full use of the channel status, the loss of throughput is minimized to the greatest extent. Figure 6 The number of completed business types of different types shows that the proposed framework has completed more high-value business types 2, 3, and 5 than other methods. These high-value businesses are also businesses with high quality requirements and large data volumes. Figure 5 The business completion value of the framework proposed in (B) is higher than that of other methods, which is also related to Figure 5 The results of (A) confirm each other. The performance comparison of the methods is shown in Table 1. It can be seen that the proposed method outperforms similar methods in terms of average QoS score, number of completed services, and value, at the cost of sacrificing some energy efficiency, which is acceptable for scenarios with high service quality requirements.

[0254]

[0255]

[0256] Table 1 Performance comparison of different technical methods

[0257] Example 4

[0258] Figure 7 This is a schematic diagram of a cross-layer collaborative resource allocation system for maritime communications according to an embodiment of the present invention.

[0259] like Figure 7 As shown, a cross-layer collaborative resource allocation system for maritime communications includes:

[0260] A first model building module 10 is used to build a dynamic dual-threshold water injection model based on service quality perception;

[0261] The second model construction module 20 is used to construct a twin DRL business allocation model based on domain knowledge enhancement;

[0262] The allocation module 30 is used to allocate wireless network resources based on the dynamic dual-threshold water injection model and the twin DRL service allocation model.

[0263] According to one embodiment of the present invention, an electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, any cross-layer collaborative resource allocation method for maritime communications of the present invention is implemented.

[0264] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0265] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in this application (but not limited to). It should be understood that the size of the serial numbers of the steps in the content of the invention and the embodiments of the present invention does not absolutely mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

Claims

1. A cross-layer collaborative resource allocation method for maritime communications, characterized in that: The following steps are involved: Construct a dynamic dual-threshold water injection model based on service quality perception; Build a twin DRL business allocation model based on domain knowledge enhancement; Wireless network resources are allocated based on the dynamic dual-threshold water injection model and the twin DRL service allocation model.

2. The cross-layer collaborative resource allocation method for maritime communications according to claim 1, characterized in that: Construct a dynamic dual-threshold water injection model based on service quality perception, including: Constructing a wireless network node transmission model, determining a minimum quality of service requirement for each service type, and calculating, based on the wireless network node transmission model, a predicted transmission rate corresponding to a first preset score and a predicted transmission rate corresponding to a second preset score for the quality of service of the service currently being executed by the node; wherein the first preset score is higher than the second preset score; The signal-to-noise ratio corresponding to the boundary is calculated based on the Shannon capacity formula; Obtaining a user power allocation strategy for initializing service quality awareness; Iterate the power allocation waterline with throughput as the waterline according to the channel gain; Initialize the user power allocation strategy based on the real-time signal-to-noise ratio as the dynamic water level benchmark; Subtract the power allocation strategies of the two waterlines to obtain the maximum difference and the minimum difference of the power allocation strategies; Determining node power allocation based on the maximum difference in the power allocation strategy; The power allocation of the end user is determined based on the remaining unallocated power of the base station and the node power allocation.

3. The cross-layer collaborative resource allocation method for maritime communications according to claim 2, characterized in that: The predicted transmission rate corresponding to the first preset score and the predicted transmission rate corresponding to the second preset score are calculated by the following formula: in, Indicates the theoretical rate at which the service quality score of the current business executed by node i corresponds to the first preset score; I represents the theoretical rate at which the service quality score of the current business executed by node i corresponds to the second preset score; s Indicates the number of node connections; represents the average packet delay; J s Indicates packet jitter; L s Indicates the packet loss rate; Indicates the experience rate corresponding to the first preset score of the service quality score of the current business executed by node i; Indicates the experience rate corresponding to the second preset score of the service quality score of the current business executed by node i; The signal-to-noise ratio corresponding to the boundary is calculated by the following formula: Among them, SNR i,max / min Indicates the maximum signal-to-noise ratio / minimum signal-to-noise ratio corresponding to the boundary; B i represents the bandwidth of node i; The power allocation strategy for initializing service quality awareness is calculated using the following formula: in, Indicates the user power allocation strategy corresponding to the first preset score / the second preset score of the service quality score currently executed by node i; N 0,i represents the noise of node i; H(d i , f i ) represents the channel gain of node i, d i is the distance between node i and the base station, f i is the frequency used by node i; The power allocation waterline with throughput as the waterline is calculated using the following formula: Among them, α i represents the node signal-to-noise ratio; μ represents the power allocation waterline with throughput as the waterline; P total represents the maximum allocatable power of the base station; i represents the i-th node of the channel; The user power allocation strategy is calculated using the following formula: in, represents the user power allocation strategy of the i-th node; μ * Indicates the power allocation waterline with throughput as the waterline after multiple iterations; The maximum and minimum power allocation strategy differences are calculated using the following formula: Where ΔP i,max The power allocation strategy difference corresponding to the first preset score of the service quality score of the business currently executed by node i, that is, the maximum power allocation strategy difference; Indicates the user power allocation strategy corresponding to the first preset score of the service quality score currently executed by node i; ΔP i,min The power allocation strategy difference between the service quality score of the business currently executed by node i and the second preset score, that is, the minimum power allocation strategy difference; Indicates the user power allocation strategy corresponding to the second preset score of the service quality score currently executed by node i; The node power allocation is determined by the following formula: Among them, P i represents the node power allocation; The power allocation to the end user is determined by the following formula: Where ΔP i,surplus Represents the remaining unallocated power ΔP surplus The power allocated to node i.

4. The cross-layer collaborative resource allocation method for maritime communications according to claim 1, characterized in that: Building a twin DRL business allocation model enhanced by domain knowledge includes: building a twin network deep reinforcement learning model based on a dual-channel feature encoder and building a strategy optimization filter based on channel state perception.

5. The cross-layer collaborative resource allocation method for maritime communications according to claim 4, characterized in that: Build a Siamese network deep reinforcement learning model based on a dual-channel feature encoder, including: The node motion state vector, which consists of motion coordinates, instantaneous speed, instantaneous angle, and service type, is mapped to a relative coordinate system through coordinate normalization and fully connected layer feature embedding to generate a low-dimensional feature vector, thereby encoding and extracting node features. For the service demand vector composed of delay-sensitive flags, packet delay thresholds, packet loss thresholds, data volume, and service type, logarithmic data compression and weighted aggregation using a gated attention mechanism are used to achieve service demand feature encoding and adaptive selection. Constructing a dual-branch twin network with shared initial parameters and asynchronous updates; wherein the dual-branch twin network includes a main network and a twin branch network; The encoded dual-channel feature vectors are respectively input into the weight-sharing main network and the twin branch network to extract features and calculate the business and node embedding similarity values, and then sorted and output the node allocation strategy based on business features.

6. The cross-layer collaborative resource allocation method for maritime communications according to claim 5, characterized in that: The coordinates are normalized using the following formula: Among them, p t Indicates absolute position; p base represents the base station coordinates; ||·||2 represents L 2 norm; ε represents the numerical stability constant; Indicates p t Position in the relative coordinate system; Generate low-dimensional feature vectors using the following formula: Among them, h n represents the low-dimensional feature vector; σ(·) represents the ReLU activation function; W n represents the customer training weight matrix; v t represents the instantaneous speed; θ t Indicates the instantaneous angle; e t Indicates the type of business the node is executing; b n represents a trainable bias vector; Logarithmic compression is performed using the following formula: Wherein, m represents the amount of business data; Indicates the amount of compressed business data; The gated attention mechanism weighted aggregation is performed using the following formula: Among them, h b represents fusion features; W b represents a trainable weight matrix; α represents the attention weight; τ represents the delay sensitivity flag; d represents the packet delay threshold; l represents the packet loss threshold; q represents the service type; The business and node embedding similarity values are calculated using the following formula: Among them, L siamese (φ) represents the embedding similarity value; K represents the number of nodes; N represents the number of services; f φ (s i ) represents the eigenvalue of node i; f φ (s j ) represents the eigenvalue of node j.

7. The method for cross-layer collaborative resource allocation for maritime communications according to claim 4, characterized in that: Construct a policy optimization filter based on channel state perception, including: By combining the node's measured transmission rate, the system's maximum rate benchmark, the service data volume requirements, and the node's channel quality coefficient, the node's adaptability to different service types is calculated, and a channel state-driven conflict resolution matrix is constructed. Based on the conflict resolution matrix, the initial allocation probability matrix output by the two-branch twin network is reconstructed to obtain a new allocation probability matrix; the allocation probability of each type of business is normalized to ensure that the total allocation probability of each type of business is 1, and the strategy optimization filter construction is completed.

8. The cross-layer collaborative resource allocation method for maritime communications according to claim 7, characterized in that: The adaptability of a node to different service types is calculated using the following formula: Among them, c nq (t) represents the adaptability of node n to service q; R n (t) represents the measured transmission rate of node n at time t; R max Indicates the current system maximum rate benchmark; D q Indicates the data volume requirement of business type q; represents the average data volume demand of all business types; ξ n (t) represents the channel quality coefficient of node i; The probability of allocating services to each type is normalized using the following formula: in, represents the probability of node i assigning service m; Represents normalized 9. A cross-layer collaborative resource allocation system for maritime communications, characterized in that: include: The first model building module is used to build a dynamic dual-threshold water injection model based on service quality perception; The second model building module is used to build a twin DRL business allocation model based on domain knowledge enhancement; An allocation module is used to allocate wireless network resources based on the dynamic dual-threshold water injection model and the twin DRL service allocation model.

10. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the method for cross-layer collaborative resource allocation for maritime communications according to any one of claims 1 to 8 is implemented.