Marine wireless network joint resource allocation method and system based on knowledge embedding
Optimizing the resource allocation of offshore wireless networks through action distribution alignment and domain knowledge embedded physical boot loss function, solving the problems of high algorithm complexity and strong data dependence in offshore wireless communication, and achieving more efficient resource allocation and communication performance.
Patent Information
- Application Number
- CN202510409479.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art has problems of high algorithm complexity and strong data dependence in offshore wireless communication, which is difficult to adapt to the rapidly changing environment and business needs of offshore nodes, resulting in insufficient scalability and flexibility of resource allocation strategies.
The joint resource allocation method of maritime wireless network based on knowledge embedding is adopted, and the physical guided loss function of domain knowledge embedding is achieved through action distribution alignment and physical guidance loss function, combined with deep reinforcement learning algorithms, resource allocation strategies are optimized, and decision conflicts and data dependence are reduced.
It improves the flexibility and generalization ability of resource allocation, reduces policy conflicts, reduces data dependence of model training, and improves the communication performance of maritime wireless networks.
Smart Images

Figure CN120264428A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of maritime wireless communication, and specifically to a method and system for joint resource allocation of a maritime wireless network based on knowledge embedding. Background Art
[0002] The evolution of science and technology has driven the continuous movement of human activity areas towards the open sea and the deep sea. There is an urgent need for high-quality communication with full-time and full-domain coverage in maritime wireless communication, which has also become the development goal of the new generation of communication technologies B5G / 6G. To achieve this vision, the maritime wireless communication network needs to integrate a variety of efficient transmission means through means such as wide-area coverage and heterogeneous resource orchestration to provide ubiquitous intelligent information services integrating communication, computing, and sensing for the sky, space, sea, and complex tasks in specific areas. However, it is difficult to densely build and support high-power base stations in the marine environment, and only mobile platforms such as ships, large unmanned aerial vehicles, and airplanes can be used instead, resulting in scarce wireless network resources and limited service capabilities. Human maritime activities can use various nodes such as ships, unmanned aerial vehicles, airplanes, submersibles, buoys, and satellites. To achieve high-quality communication with wide-area coverage, it is necessary to uniformly manage and allocate heterogeneous resources. The ocean area is vast, and there are many types of nodes and services with obvious differences in demand, resulting in high real-time requirements, large span, and high difficulty in resource allocation. Existing time-frequency resource allocation technologies using fixed frequency bands and subcarriers are difficult to meet new requirements, and there is an urgent need for time-frequency resource allocation technologies with high flexibility and larger capacity.
[0003] At present, one of the core radio resource allocation technologies for terrestrial wireless transmission systems is Orthogonal Frequency Division Multiple Access (OFDMA), which has flexible spectrum allocation capabilities and high spectrum efficiency. The OFDMA system provides a structured framework for resource allocation, supports dynamic frequency selection schemes, allows resource blocks to be divided on multiple subcarriers, and can manage the modulation method and power of each subcarrier, but the set of solutions for the allocation scheme is relatively large. At the same time, the problem of power and spectrum allocation in wireless communication is a typical NP-hard non-convex problem, and it is challenging for traditional solution methods to solve such problems. Therefore, heuristic methods or deep learning methods are often selected for the radio communication network resource allocation method of the OFDMA system. However, due to the fast speed, wide distribution, large variety differences, and diverse service requirements of marine nodes, this means that the node energy consumption is limited, the communication channel changes rapidly over time, and the real-time requirement for the allocation strategy is high. Therefore, the above-mentioned terrestrial OFDMA heuristic or deep learning resource allocation technologies still face challenges when directly applied to marine scenarios, such as being difficult to adapt to dynamic environmental conditions and service requirements, resulting in low scalability. Meta-learning, especially in the context of reinforcement learning, can solve the limitations of static and heuristic methods in dynamic environments. The introduction of meta-reinforcement learning technology provides a new research perspective for solving such problems, and relevant scholars have carried out a series of studies on the combination of OFDMA-based resource allocation technologies and meta-learning methods. Although the existing dynamic resource allocation methods based on meta-learning-based deep reinforcement learning (meta-DRL) have made certain progress, the main problems they face are: (1) When the meta-DRL method allocates joint resources, if multiple algorithms are combined to output the strategy, the algorithm complexity is high; if the joint resources are allocated through the output action combination strategy, the performance is limited. (2) Existing meta-DRL methods usually require a large amount of data for pre-training to achieve the expected effect and are difficult to apply to scenarios where it is difficult to collect data. Summary of the Invention
[0004] The object of the present invention is to provide a joint resource allocation method and system for marine wireless networks based on knowledge embedding to solve at least one of the above technical problems.
[0005] In a first aspect, an embodiment of the present invention provides a method for joint resource allocation in a maritime wireless network based on knowledge embedding, which is applied to an orthogonal frequency division multiple access system; the method includes: receiving global meta-parameters and multiple resource allocation tasks from an initialization model of an outer meta-network; the multiple resource allocation tasks are a set of random tasks generated based on the outer meta-network; based on a preset action distribution alignment method, adjusting the output power and the orthogonal frequency division multiple access resource block allocation action combination of each resource allocation task to obtain a time-frequency resource allocation strategy for each resource allocation task; interacting based on the global meta-parameters and the environmental state to obtain an environmental state feedback value obtained by each resource allocation task after executing the time-frequency resource allocation strategy; based on a physical guidance loss function with domain knowledge embedding and the environmental state feedback value, updating the task parameters of each resource allocation task through a deep reinforcement learning algorithm to obtain the inner network parameters and inner task losses of each resource allocation task; based on the inner network parameters and the inner task losses, updating the global meta-parameters through backpropagation to obtain a target resource allocation model; performing maritime wireless network resource allocation based on the target resource allocation model.
[0006] Further, the preset action distribution alignment method includes: making the output action to be aligned approach the reference output action based on weight parameters; wherein, if the output distributions of the output action to be aligned and the reference output action do not completely conflict, a weighted mapping method is adopted to approach the weight parameters; if the output distributions of the output action to be aligned and the reference output action completely conflict, a migration mapping method is adopted for non-zero distributions to adjust the weight parameters.
[0007] Further, the mapping method of the weight parameters includes:
[0008]
[0009] In the formula, A is the reference output action, B is an output action different from A, β is the weight parameter, B' is the output action after alignment adjustment, Bsum is the sum of the output values of the B action, and A[i] is the i-th action output value of the reference output action.
[0010] Further, the physical guidance loss function with domain knowledge embedding includes:
[0011] Loss new = Loss CLIP + α·Loss knowledge
[0012] In the formula, Lossnew is the physical guidance loss function with domain knowledge embedding, Loss CLIPis the basic loss term, Lossknowledge is the physics-guided loss term based on domain knowledge, and α is the adjustment factor.
[0013] Furthermore, the basic loss term includes:
[0014]
[0015] In the formula, r t (θ) is the ratio of the new policy probability to the original policy probability, is the advantage estimate, denotes the expectation with respect to time t, clip denotes the clipping of r t (θ) such that the clipped r t (θ) varies in the range [1 - ε, 1 + ε], where ε is the clipping threshold.
[0016] Furthermore, the physics-guided loss term based on domain knowledge includes:
[0017]
[0018] In the formula, N is the number of nodes, vi represents the actual rate under the power and resource block allocation of the i-th node, and E{vtotal} is the expected throughput.
[0019] Furthermore, updating the global meta-parameters through backpropagation includes:
[0020]
[0021] In the formula, θ (0) is the global meta-parameter, θ′ (0) is the updated global meta-parameter, D i is the environmental state feedback value of the i-th resource allocation task, is the loss function of the i-th resource allocation task.
[0022] Second aspect, an embodiment of the present invention further provides a joint resource allocation system for a maritime wireless network based on knowledge embedding, which is applied to an orthogonal frequency division multiple access system; it includes: an initialization module, an action distribution alignment module, an interaction module, a first update module, a second update module, and an allocation module; wherein, the initialization module is used to receive global meta-parameters and multiple resource allocation tasks from the initialization model of the outer meta-network; the multiple resource allocation tasks are a set of random tasks generated based on the outer meta-network; the action distribution alignment module is used to adjust the output power and the orthogonal frequency division multiple access resource block allocation action combination of each resource allocation task based on a preset action distribution alignment method to obtain the time-frequency resource allocation strategy of each resource allocation task; the interaction module is used to interact based on the global meta-parameters and the environmental state to obtain the environmental state feedback value obtained by each resource allocation task after executing the time-frequency resource allocation strategy; the first update module is used to update the task parameters of each resource allocation task through a deep reinforcement learning algorithm based on the physical guidance loss function embedded with domain knowledge and the environmental state feedback value to obtain the inner network parameters and inner task losses of each resource allocation task; the second update module is used to update the global meta-parameters through backpropagation based on the inner network parameters and the inner task losses to obtain a target resource allocation model; the allocation module is used to perform maritime wireless network resource allocation based on the target resource allocation model.
[0023] Third aspect, an embodiment of the present invention further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the method provided by the embodiment of the present invention.
[0024] Fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, they implement the method provided by the embodiment of the present invention.
[0025] The present invention provides a joint resource allocation method and system for a maritime wireless network based on knowledge embedding, provides an action distribution alignment method for an orthogonal frequency division multiple access system, adopts different distribution mapping methods according to the degree of conflict between actions, enables the intelligent agent to output actions within the framework of real-world rules, reduces the decision-making conflict when the intelligent agent outputs action combinations, and reduces resource loss and policy hedging; the method provided by the present invention embeds a meta-learning network of reinforcement learning, uses a physical guidance loss function based on domain knowledge, constructs a physical guidance term by using known domain knowledge as a soft constraint, guides the optimization and adjustment of the model, reduces the data dependence of model training, and improves the generalization ability of the model. The present invention alleviates the technical problems of high algorithm complexity and strong data dependence existing in the prior art. Brief Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is a flowchart of a method for joint resource allocation in a maritime wireless network based on knowledge embedding provided by an embodiment of the present invention;
[0028] Figure 2 It is a flowchart of another method for joint resource allocation in a maritime wireless network based on knowledge embedding provided by an embodiment of the present invention;
[0029] Figure 3 It is a comparison chart of the total throughput of the nodes connected to the base station within the task time provided by an embodiment of the present invention;
[0030] Figure 4 It is a comparison chart of the average throughput of nodes independent of quantity provided by an embodiment of the present invention;
[0031] Figure 5 It is a comparison chart of the power allocation fairness coefficient provided by an embodiment of the present invention;
[0032] Figure 6 It is a comparison chart of the throughput under different maximum output powers of the method for joint resource allocation in a maritime wireless network based on knowledge embedding provided by an embodiment of the present invention;
[0033] Figure 7 It is a comparison chart of the total throughput under different training data of different methods provided by an embodiment of the present invention;
[0034] Figure 8 It is a schematic diagram of a system for joint resource allocation in a maritime wireless network based on knowledge embedding provided by an embodiment of the present invention. Detailed Embodiments
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0036] Embodiment 1
[0037] Figure 1It is a flowchart of a method for joint resource allocation in a maritime wireless network based on knowledge embedding according to an embodiment of the present invention. This method is applied to an orthogonal frequency division multiple access system. As Figure 1 shown, the method specifically includes the following steps:
[0038] Step S102, receive global meta-parameters and multiple resource allocation tasks from the initialization model of the outer meta-network; the multiple resource allocation tasks are a set of random tasks generated based on the outer meta-network.
[0039] Step S104, based on a preset action distribution alignment method, adjust the output power and the orthogonal frequency division multiple access resource block allocation action combination of each resource allocation task to obtain the time-frequency resource allocation strategy of each resource allocation task.
[0040] Step S106, interact based on the global meta-parameters and the environmental state to obtain the environmental state feedback value obtained by each resource allocation task after executing the time-frequency resource allocation strategy.
[0041] Step S108, based on the physical guidance loss function embedded with domain knowledge and the environmental state feedback value, update the task parameters of each resource allocation task through a deep reinforcement learning algorithm to obtain the inner network parameters and the inner task loss of each resource allocation task.
[0042] Step S110, update the global meta-parameters through backpropagation based on the inner network parameters and the inner task loss to obtain the target resource allocation model.
[0043] Step S112, perform maritime wireless network resource allocation based on the target resource allocation model.
[0044] In the embodiment of the present invention, when the agent performs joint resource allocation of action combinations, there will be a phenomenon that the action space does not match and policy conflicts occur. To reduce the resource loss caused by policy contradictions, it is necessary to align the action distribution.
[0045] Specifically, the preset action distribution alignment method includes:
[0046] Make the output action to be aligned approach the reference output action based on the weight parameter; where
[0047] If the output distributions of the output action to be aligned and the reference output action do not completely conflict, adopt a weighted mapping method to approach the weight parameter;
[0048] If the output distributions of the output action to be aligned and the reference output action completely conflict, adopt a migration mapping method to adjust the weight parameter for the non-zero distribution.
[0049] Specifically, assume that the agent outputs actions A and B simultaneously. If action A is selected as the benchmark, the distribution of action B is made to converge towards A: when the output distributions of actions A and B do not completely conflict, a weighted mapping method is used to converge the weights, and the parameter β is used to adjust the degree of retention of the information of action B; when the output distributions of actions A and B completely conflict, a migration mapping method is used to adjust the weights for non-zero distributions. The specific mapping formula is as follows:
[0050]
[0051] In the formula, A is the benchmark output action, B is the output action different from A, β is the weight parameter, B’ is the output action after alignment adjustment, B sum is the sum of the output values of action B, and A[i] is the i-th action output value of the benchmark output action.
[0052] Specifically, the physical guidance loss function for domain knowledge embedding includes:
[0053] Loss new = Loss CLIP + α·Loss knowledge
[0054] In the formula, Loss new is the physical guidance loss function for domain knowledge embedding, Loss CLIP is the basic loss term, Loss knowledge is the physical guidance loss term based on domain knowledge, and α is the adjustment factor used to control the proportion.
[0055] Specifically, the basic loss term Loss CLIP adopts a clipped objective function that optimizes the surrogate loss. This loss function measures the ratio of the advantage values of the new policy and the old policy and constrains the policy update:
[0056]
[0057] In the formula, r t (θ) is the ratio of the new policy probability to the original policy probability, is the advantage estimate, represents the expectation with respect to time t, clip represents the clipping of r t (θ) such that the clipped r t (θ) varies in the range [1 - ε, 1 + ε], where ε is the clipping threshold.
[0058] Specifically, the physical guidance loss term based on domain knowledge is composed of the difference between the expected throughput and the actual throughput, expecting the agent to continuously approach the theoretical maximum throughput:
[0059]
[0060] Where N is the number of nodes, v i represents the actual rate under the power and resource block allocation of the i-th node, and E{v total} is the expected throughput.
[0061] In the embodiment of the present invention, the Shannon capacity formula is used as domain knowledge to construct the expected throughput:
[0062]
[0063] Where B represents the node bandwidth, SNR represents the signal-to-noise ratio of the node, and P r represents the node power, and N0 represents the noise.
[0064] Given that both N0B and B are much greater than 1, and the signal-to-noise ratio part is moderately amplified, the final formula can be written as:
[0065]
[0066] Where P max represents the maximum value of the power that can be allocated to the node
[0067] The final form of the physical-guided loss function for domain knowledge embedding is obtained as follows:
[0068]
[0069] Specifically, in step S110, the global meta-parameters are updated through backpropagation, including:
[0070]
[0071] Where θ (0) is the global meta-parameter, θ′ (0) is the updated global meta-parameter, D i is the environmental state feedback value of the i-th resource allocation task, is the loss function of the i-th resource allocation task.
[0072] Embodiment 2
[0073] Figure 2 is a flowchart of another method for joint resource allocation in a maritime wireless network based on knowledge embedding provided according to the embodiment of the present invention. As Figure 2As shown, the entire process is divided into two parts: internal and external. The external part is a model network based on meta-learning, and the internal part is a policy network based on deep reinforcement learning. The external model network generates a set of random tasks and an initial policy, which are provided to the inner loop for learning and optimization. In each task, the inner loop uses an action distribution alignment module (DAM) to optimize the policy, then interacts with the environment to collect rewards and calculate the task loss, iteratively updating the initial policy until the task ends. Finally, the policy losses of multiple tasks are backpropagated to the outer loop through backpropagation. After summarizing the output of the inner loop, the outer loop adjusts the global policy parameters in the model according to the feedback gradient, making the policy more universal and generalizable.
[0074] Specifically, it includes the following steps:
[0075] 1) Receive the global parameter θ from the initialized model (0) and i tasks T i ;
[0076] 2) Adjust the output power and the OFDMA resource block allocation action combination through a preset action distribution alignment method to obtain the agent's power and spectrum resource allocation policy a t i , execute the time-frequency resource allocation policy according to task T i and store the obtained environmental state feedback value in the experience pool;
[0077] 3) Calculate the loss value through a physics-guided loss function based on domain knowledge embedding;
[0078] Repeat steps 1)-3) until the current task T i ends;
[0079] 4) Optimize the initialized global parameter θ of the model according to the losses of all tasks (0) ;
[0080] Repeat steps 1)-4) of step 2 until all tasks T i end.
[0081] For example, the relevant parameters are set as follows: the number of agents is 1, the agent speed is [10, 30] m / s, the number of users is 2 - 8, the user node speed is [15, 18] / [125, 280] m / s, the node turning angle is [-0.25π, 0.25π], the task duration is 60 s, the agent transmission power is 100 W, the channel interference in the training scenario is 10 ± 4, 15 ± 3, 20 ± 2 dB, the channel interference in the test scenario is 10, 25, 35 dB, the optional range of node spectrum is 2010 - 2030 MHz, the resource block bandwidth is 15 kHz, the neural network learning rate is 1e-4, the number of network layers is 3, the number of neurons is 300 / 150 / 8, the sampling time is 1 s, the discount factor is 0.95, the batch sample size is 600, and the length of the experience buffer is 1.2*10 -5 , the number of outer loop updates (training / testing) is 30 / 0, and the number of inner loop updates (training / testing) is 20 / 5.
[0082] Embodiment III
[0083] To facilitate the comparative analysis of the performance differences between the present invention and the prior art, a simulation analysis was carried out with a maritime wireless network as the scenario. A maritime wireless network joint resource allocation method based on knowledge embedding (KE-MAML) provided by the embodiment of the present invention was compared with the existing meta-learning algorithm MAML-PPO and RL 2 . Among them, in order to verify the functions and performance of the proposed module, an ablation comparison experiment was conducted. Among them, Loss_MAML-PPO is based on the MAML-PPO algorithm and only modifies its loss function to a physical guidance loss function that embeds domain knowledge; DAM_MAML-PPO is based on the MAML-PPO algorithm and only adds the DAM module. The specific details of the scenario are as follows:
[0084] The simulation scenario is set in the sea surface and the upper area range of R*R. There is a low-speed base station node of an agent and n high-speed user nodes moving freely in the area. The high-speed node users are divided into ship nodes and aircraft nodes, and the ratio between the two is randomly generated. All nodes perform continuous random motion. It is assumed that the current area is the communicable range of the base station, that is, when a node enters the current area, it is default to be directly and only connected to the base station node, and when it leaves the current area, it is default to be disconnected. Considering the distance difference caused by the height of the node, a part of the satellite communication frequency band is selected for the communication frequency band. To investigate the generalization of the method, different combinations are randomly selected from 3 channel environments with different interferences, n - 1 different numbers of base station-connected nodes, randomly generated initial positions of nodes, and node speeds as tasks for training and testing. In the training stage, 600 tasks are randomly generated in three channel environments, and each task lasts for 60 s (a total of 36000 samples). In the test stage, 20 different tasks are randomly generated in each of the three channel environments (a total of 60 tasks).
[0085] Figure 3 It is a comparison chart of the total throughput of the nodes connected to the base station within the task time provided by the embodiments of the present invention. As Figure 3 shown, the channel environment deteriorates in turn from left to right. First, focus on the performance comparison of the ablation experiment. Looking vertically at the performance comparison chart in the channel 1 environment, it can be found from the figure that compared with MAML-PPO, the ablation experimental group DAM_MAML has increased by 14.28%, indicating that the DAM module can effectively reduce the output of non-standard actions and reduce the idle or waste of resources caused by policy conflicts. However, there is still a 23.04% gap from the method provided by the embodiments of the present invention, because when the DAM module makes the output distribution closer, due to the information loss in the mapping process, part of the original policy becomes inefficient or ineffective, which further limits the further improvement of performance and affects the ability of the MAML-PPO algorithm itself to optimize the policy; the performance of the ablation experimental group Loss_MAML-PPO has not been improved but has decreased by 4.21%. This is because in the case of conflicts between the resource block allocation actions and power allocation actions made by the agent, the embedding of physical knowledge alone cannot play a strong guiding role for the model and it is difficult to make up for the performance decline caused by policy conflicts.
[0086] When the physical knowledge of the loss function guides the model to find a better solution, the conflicting resource block and power allocation actions increase the difficulty of the agent to explore the optimal solution to a certain extent. However, the performance shown when the two modules are used in combination has a huge increase of 40.62% compared with the benchmark MAML-PPO algorithm. This is because while the physical knowledge loss function guides the agent to explore, the DAM module reduces unnecessary exploration by avoiding action combinations that do not conform to physical reality. Correspondingly, the physical knowledge loss function makes up for the information loss in the action distribution mapping process, so that the two modules have a good effect when used in combination. RL 2 The algorithm is difficult to handle tasks with extremely high complexity, has a poor adaptability to environments with extremely strong dynamics and scarce data, and coupled with the policy conflicts between the multiple output actions of the agent, the overall performance is inferior to MAML-PPO and its variant algorithms.
[0087] Comparing the performance horizontally in different channel environments, as the channel environment gradually deteriorates, the method provided by the embodiments of the present invention can still maintain a good performance advantage. Compared with the benchmark MAML-PPO algorithm, it has increased by 41.52% and 20.42% in the channel 2 and channel 3 environments respectively, but the numerical value of the advantage difference is gradually decreasing. This is because the harsh channel environment limits the maximum channel communication capacity, and the number of optimal solutions to the problem and the differences between various allocation schemes are both decreasing.
[0088] Figure 4 It is a comparison chart of the average throughput of nodes independent of quantity provided by the embodiments of the present invention,Figure 5 is a comparison graph of power distribution fairness coefficients provided according to an embodiment of the present invention. As Figure 4 and Figure 5 shown in the vertical comparison of the average throughput of Channel 1 environment, the method provided by the embodiment of the present invention is 49.64% higher than the benchmark method MAML-PPO, and the fairness coefficient is on average 0.04 higher, indicating that the method provided by the embodiment of the present invention can take into account a certain degree of fairness while improving throughput.
[0089] Figure 6 is a comparison graph of the throughput of the joint resource allocation method for maritime wireless networks based on knowledge embedding provided according to an embodiment of the present invention under different maximum output powers. As Figure 6 shown, the curve trends of the average throughput and total throughput of the method provided by the embodiment of the present invention for nodes with no connection are not exactly the same, indicating that the strategy of the present invention is to strengthen rather than support the weak, which is in line with the understanding of the resource allocation strategy for unfamiliar resource-constrained scenarios.
[0090] Specifically, as Figure 6 shown, the method provided by the embodiment of the present invention achieves good continuous adaptation ability at each P max value, that is, when the model is updated according to the channel environment, the throughput performance on the channel generally remains stable and does not drop significantly, which further demonstrates the advantage of the method provided by the embodiment of the present invention in terms of adaptability. Table 1 shows the average algorithm performance in the same-channel condition interval tasks:
[0091] Table 1 Comparison of algorithm performance
[0092]
[0093] It can be seen from Table 1 that the method provided by the present invention has achieved great advantages at the cost of sacrificing fairness in the throughput performance of various channel environments and can meet the requirements of resource-constrained environments.
[0094] Figure 7 is a comparison graph of the total throughput of different methods under different training data provided according to an embodiment of the present invention. As Figure 7As shown in the figure, when the method provided by the embodiment of the present invention is longitudinally compared with other control algorithms (such as MAML-PPO), although the average throughput of the channel 1 environment of the present invention is relatively high, the fairness coefficient is relatively low, indicating that fairness and throughput are strategy-inverse in the current scenario. In channels 2 and 3, when horizontally comparing with other control algorithms, it will be found that the fairness coefficient of the algorithm provided by the present invention is gradually increasing. This is because under the guidance of the physical guidance loss function based on knowledge embedding, the agent discovers that the more severe the environment is, the smaller the difference in the allocation schemes of each node is, and the greater the throughput of the more fair allocation scheme will be. This shows that the proposed algorithm is more sensitive to environmental changes, can quickly find a better solution to maximize throughput, and has high generalization ability for unfamiliar environments.
[0095] The excellent performance of the meta-reinforcement learning algorithm is closely related to the pre-trained global parameter model. Figure 7 and Table 2 respectively show the curve comparison, mean value and mean square deviation comparison of the maximum throughput performance of the model under different numbers of inner and outer loop trainings, where m in m*n refers to the number of outer loop times, and n refers to the number of tasks learned in each outer loop:
[0096] Table 2 Comparison of the average total throughput of the algorithm under different amounts of training data
[0097]
[0098] As Figure 7 shown, by comparing the left and right figures, it can be seen that for the method provided by the present invention in three channel environments, the gap between the model performances corresponding to each training amount is relatively small, while the gap between the model performances of the benchmark MAML-PPO algorithm is relatively large, indicating that the benchmark algorithm has a greater demand for the amount of pre-trained data than the proposed method. This is because even if the proposed algorithm lacks a large amount of pre-trained data, it can rely on the physical guidance loss function based on knowledge embedding to guide the model optimization. At the same time, the DAM module can greatly reduce the action combinations that do not conform to the actual situation, greatly reducing the amount of data required for the model to explore and learn, and improving the utilization efficiency of the pre-trained data. Therefore, in comparison, the influence of the amount of pre-trained learning data on the benchmark algorithm is greater than that of the proposed algorithm. As shown in Table 2, by calculating the mean square deviation of the average total throughput, it is obvious that the mean square deviation of the method provided by the present invention is less than that of the benchmark algorithm, which indicates that the present invention can still maintain good performance in the scenario where it is difficult to obtain a large amount of pre-trained data, and the dependence on pre-trained data is reduced.
[0099] From the above description, it can be known that the embodiment of the present invention provides a method for joint resource allocation in a maritime wireless network based on knowledge embedding. Compared with the prior art, it has the following technical effects:
[0100] 1. A general action distribution alignment method is designed, which adopts different distribution mapping methods according to the conflict degree between actions, enabling the agent to output actions within the framework of real-world rules, reducing decision-making conflicts when the agent outputs action combinations, and reducing resource losses and policy hedging.
[0101] 2. The method provided by the present invention consists of two parts of networks inside and outside. The outer layer is a meta-learning network, and the inner layer is a deep reinforcement learning network. A knowledge embedding method is introduced, and a physical guidance loss function based on domain knowledge is designed for the inner layer network. The known domain knowledge is used as a soft constraint to construct a physical guidance term to guide the optimization and adjustment of the model, reducing the data dependence of model training and further improving the generalization ability of the model.
[0102] Embodiment 4
[0103] Figure 8 is a schematic diagram of a joint resource allocation system for a maritime wireless network based on knowledge embedding provided by an embodiment of the present invention, which is applied to an orthogonal frequency division multiple access system. As Figure 8 shown, it includes: an initialization module 10, an action distribution alignment module 20, an interaction module 30, a first update module 40, a second update module 50, and an allocation module 60.
[0104] Specifically, the initialization module 10 is used to receive global meta-parameters and multiple resource allocation tasks from the initialization model of the outer meta-network; the multiple resource allocation tasks are a set of random tasks generated based on the outer meta-network;
[0105] The action distribution alignment module 20 is used to adjust the output power of each resource allocation task and the orthogonal frequency division multiple access resource block allocation action combination based on a preset action distribution alignment method to obtain the time-frequency resource allocation strategy of each resource allocation task;
[0106] The interaction module 30 is used to interact based on the global meta-parameters and the environmental state to obtain the environmental state feedback value obtained by each resource allocation task after executing the time-frequency resource allocation strategy;
[0107] The first update module 40 is used to update the task parameters of each resource allocation task through a deep reinforcement learning algorithm based on the physical guidance loss function embedded with domain knowledge and the environmental state feedback value to obtain the inner layer network parameters and inner layer task losses of each resource allocation task;
[0108] The second update module 50 is used to update the global meta-parameters through backpropagation based on the inner layer network parameters and inner layer task losses to obtain the target resource allocation model;
[0109] The allocation module 60 is used to perform maritime wireless network resource allocation based on the target resource allocation model.
[0110] The present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method provided by the embodiment of the present invention is implemented.
[0111] The present invention also provides a computer-readable storage medium storing computer instructions, which when executed by a processor implement the method provided by the embodiment of the present invention.
[0112] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claimed invention.
[0113] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in the various embodiments can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A joint resource allocation method for maritime wireless networks based on knowledge embedding, characterized in that, Applied to an orthogonal frequency division multiple access system; including: Receiving global meta-parameters and multiple resource allocation tasks from the initialization model of the outer meta-network; the multiple resource allocation tasks are a set of random tasks generated based on the outer meta-network; Based on a preset action distribution alignment method, adjusting the output power of each resource allocation task and the orthogonal frequency division multiple access resource block allocation action combination to obtain the time-frequency resource allocation strategy of each resource allocation task; Interacting based on the global meta-parameters and the environmental state to obtain the environmental state feedback value obtained by each resource allocation task after executing the time-frequency resource allocation strategy; Based on the physical guidance loss function embedded with domain knowledge and the environmental state feedback value, updating the task parameters of each resource allocation task through a deep reinforcement learning algorithm to obtain the inner network parameters and inner task losses of each resource allocation task; Based on the inner network parameters and the inner task losses, updating the global meta-parameters through backpropagation to obtain the target resource allocation model; Performing maritime wireless network resource allocation based on the target resource allocation model.
2. The method according to claim 1, wherein: The preset action distribution alignment method includes: Based on the weight parameters, making the output action to be aligned approach the reference output action; where If the output distribution of the output action to be aligned and the reference output action does not completely conflict, then approach the weight parameters in a weighted mapping manner; If the output distribution of the output action to be aligned and the reference output action completely conflicts, then adjust the weight parameters in a migration mapping manner for the non-zero distribution.
3. The method according to claim 2, wherein: The mapping method of the weight parameters includes: Wherein, A is the reference output action, B is an output action different from A, β is the weight parameter, B' is the output action after alignment adjustment, and B sum is the sum of the B-action output values, and A[i] is the i-th action output value of the reference output action.
4. The method according to claim 1, wherein: The physical guidance loss function embedded with domain knowledge includes: Loss new = Loss CLIP + α·Loss knowledge where Loss new is the physical guidance loss function of the domain knowledge embedding, Loss CLIP is the basic loss term, Loss knowledge is the physical guidance loss term based on domain knowledge, and α is the adjustment factor.
5. The method according to claim 4, characterized in that: The basic loss term includes: where r t (θ) is the ratio of the new policy probability to the original policy probability, is the advantage estimate, denotes the expectation with respect to time t, and clip denotes the clipping of r t (θ) such that the changed range of the clipped r t (θ) is between [1 - ε, 1 + ε], where ε is the clipping threshold.
6. The method according to claim 4, characterized in that: The physical guidance loss term based on domain knowledge includes: where N is the number of nodes, and v i represents the actual rate under the power and resource block allocation of the i-th node, and E{v total} is the expected throughput.
7. The method according to claim 1, characterized in that: Updating the global meta-parameters through backpropagation includes: where θ (0) is the global meta-parameter, θ′ (0) is the updated global meta-parameter, D i is the environmental state feedback value of the i-th resource allocation task, is the loss function of the i-th resource allocation task.
8. A joint resource allocation system for maritime wireless networks based on knowledge embedding, characterized in that Applied to an orthogonal frequency division multiple access system; including: an initialization module, an action distribution alignment module, an interaction module, a first update module, a second update module, and an allocation module; where The initialization module is used to receive global meta-parameters and multiple resource allocation tasks from the initialization model of the outer meta-network; the multiple resource allocation tasks are a set of random tasks generated based on the outer meta-network; The action distribution alignment module is used to adjust the output power of each resource allocation task and the orthogonal frequency division multiple access resource block allocation action combination based on a preset action distribution alignment method to obtain the time-frequency resource allocation strategy of each resource allocation task; The interaction module is used to interact based on the global meta-parameters and the environmental state to obtain the environmental state feedback value obtained by each resource allocation task after executing the time-frequency resource allocation strategy; The first update module is used to update the task parameters of each resource allocation task through a deep reinforcement learning algorithm based on the physical guidance loss function embedded with domain knowledge and the environmental state feedback value to obtain the inner network parameters and inner task losses of each resource allocation task; The second update module is configured to update the global meta-parameters based on the inner-layer network parameters and the inner-layer task loss through backpropagation to obtain a target resource allocation model; The allocation module is configured to perform maritime wireless network resource allocation based on the target resource allocation model.
9. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method according to any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by the processor, the method according to any one of claims 1-7 is implemented.