Satellite Internet communication resource allocation method, device and readable storage medium

By obtaining the information age, channel gain and queue backlog of satellite communication networks, and using deep reinforcement learning models for dynamic power allocation, the information timeliness of satellite Internet communication is solved, and more efficient communication resource utilization and timeliness are achieved.

CN116318371BActive Publication Date: 2025-08-12HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310375156.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-10
Publication Date
2025-08-12
Estimated Expiration
2043-04-10

AI Technical Summary

Technical Problem

In the prior art, the information timeliness of satellite Internet communication is poor, especially under the influence of time-varying channels, update rates and transmission rates, and the timeliness of communication cannot be effectively improved.

Method used

By obtaining the current environmental status information of the satellite communication network, including the information age, channel gain and queue backlog of each user, input the trained deep reinforcement learning model, obtain the channel gain weight and power allocation coefficient, and dynamic allocation of communication resources is performed based on these parameters.

Benefits of technology

The information timeliness of satellite Internet communication is improved, and the utilization rate of communication resources is optimized by dynamically adjusting power allocation sorting, and adapting to complex and changeable environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116318371B_ABST
    Figure CN116318371B_ABST
Patent Text Reader

Abstract

This application discloses a satellite internet communication resource allocation method, device, and computer-readable storage medium, relating to the field of satellite communication technology. This satellite internet communication resource allocation method includes the following steps: obtaining current environmental status information of the satellite communication network, where the current environmental status information includes the information age, channel gain, and queue backlog of each user; inputting the information age, channel gain, and queue backlog into a trained deep reinforcement learning model to obtain the current system action, where the current system action includes the channel gain weight and power allocation coefficient of each user; determining the power allocation ranking of each user based on the channel gain, queue backlog, and channel gain weight; and allocating communication resources of the satellite communication network based on the power allocation coefficient and power allocation ranking. This application addresses the technical problem of poor information timeliness in satellite internet communications in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of satellite communication technology, and in particular to a communication resource allocation method, device and readable storage medium for satellite Internet. Background Art

[0002] In recent years, the rapid deployment of medium- and low-orbit satellite constellations, represented by Starlink and OneWeb, has fueled the rapid development of satellite internet. With its wide coverage and high transmission rates, satellite internet has found applications in emergency response, aviation and maritime surveillance, and remote sensing.

[0003] Due to the significant distance between ground users and satellites, communication latency between them is often non-negligible. Furthermore, satellite internet requires sharing frequency bands with terrestrial mobile networks. As the number of users increases, available bandwidth resources become increasingly limited. While the introduction of NOMA (Non-Orthogonal Multiple Access) technology can alleviate spectrum shortages, the timeliness of information delivered by existing satellite internet technologies remains poor due to multiple factors, including time-varying channels, update rates, and transmission speeds. Summary of the Invention

[0004] The main purpose of this application is to provide a method for allocating communication resources for satellite Internet, aiming to solve the technical problem of poor information timeliness in satellite Internet communication in the existing technology.

[0005] To achieve the above objectives, the present application provides a method for allocating communication resources for a satellite internet, the method comprising the following steps:

[0006] Acquiring current environmental status information of the satellite communication network, wherein the current environmental status information includes information age, channel gain, and queue backlog of each user;

[0007] Inputting the information age, the channel gain, and the queue backlog into a trained deep reinforcement learning model to obtain a current system action, wherein the current system action includes a channel gain weight and a power allocation coefficient for each user;

[0008] Determining a power allocation order for each user according to the channel gain, the queue backlog, and the channel gain weight;

[0009] Based on the power allocation coefficient and the power allocation ranking, communication resources of the satellite communication network are allocated.

[0010] Optionally, before the step of inputting the information age, the channel gain, and the queue backlog into a trained deep reinforcement learning model to obtain the current system action, the step includes:

[0011] obtaining first environmental state information of a satellite communication network and an untrained deep reinforcement learning model;

[0012] Inputting the first environmental state information into an untrained deep reinforcement learning model to obtain a first system action;

[0013] Executing the first system action to obtain second environment state information generated based on the first system action;

[0014] Generate a corresponding first system reward value according to the first environmental state information, the second environmental state information and a preset system reward function;

[0015] Based on the first system reward value and a preset loss function, parameters of the untrained deep reinforcement learning model are updated until the untrained deep reinforcement learning model converges to obtain a trained deep reinforcement learning model.

[0016] Optionally, the untrained deep reinforcement learning model includes a new policy network, an old policy network, and an evaluation network, and the step of updating parameters of the untrained deep reinforcement learning model based on the first system reward value and a preset loss function includes:

[0017] Calculating a fitted cumulative discounted reward based on a preset value function and the evaluation network;

[0018] Calculate the current cumulative discount reward based on the first system reward value;

[0019] Calculate a corresponding loss function value based on the current cumulative discount reward, the fitted cumulative discount reward, and the preset loss function;

[0020] Parameters of the new policy network, the old policy network, and the evaluation network are updated based on the loss function value.

[0021] Optionally, the preset loss function includes a new strategy loss function and an evaluation loss function, and the step of calculating a corresponding loss function value based on the current cumulative discount reward, the fitted cumulative discount reward, and the preset loss function includes:

[0022] Obtaining a new-to-old strategy probability ratio between the new strategy network and the old strategy network;

[0023] taking the difference between the current cumulative discounted reward and the fitted cumulative discounted reward as the advantage function value;

[0024] Calculating a new strategy loss function value based on the advantage function value, the probability ratio of the new strategy to the old strategy, and the new strategy loss function;

[0025] Calculating an evaluation loss function value according to the current cumulative discount reward, the fitted cumulative discount reward, and the evaluation loss function;

[0026] The new strategy loss function value and the evaluation loss function value are used as the loss function values corresponding to the untrained deep reinforcement learning model.

[0027] Optionally, the step of updating parameters of the new policy network, the old policy network, and the evaluation network based on the loss function value includes:

[0028] Acquire new policy parameters of the new policy network, and update old policy parameters of the old policy network to the new policy parameters;

[0029] According to the new strategy loss function value, the new strategy parameters of the new strategy network are updated by gradient ascent;

[0030] According to the evaluation loss function value, the evaluation parameters of the evaluation network are updated by gradient descent.

[0031] Optionally, before the step of generating a corresponding first system reward value according to the first environmental state information, the second environmental state information and a preset system reward function, the step includes:

[0032] Constructing a first optimization function based on preset constraints and long-term average information age information;

[0033] Based on Lyapunov optimization theory, the first optimization function is converted into a drift plus penalty term, and a minimization upper bound corresponding to the drift plus penalty term is determined;

[0034] According to the minimization upper bound, a corresponding preset system reward function is constructed.

[0035] Optionally, the step of converting the first optimization function into a drift plus penalty term based on Lyapunov optimization theory includes:

[0036] constructing a power consumption liability queue, a queue backlog, and a throughput liability queue according to the first optimization function, and generating a queue vector consisting of the power consumption liability queue, the queue backlog, and the throughput liability queue;

[0037] Based on Lyapunov optimization theory, a quadratic Lyapunov function is generated according to the queue vector, and a Lyapunov drift corresponding to the quadratic Lyapunov function is determined;

[0038] Constructing a single-slot penalty function according to the queue vector and the first optimization function;

[0039] According to the Lyapunov drift and single-slot penalty function, a corresponding drift plus penalty term is determined.

[0040] Optionally, the step of determining the power allocation order of each user according to the channel gain, the queue backlog, and the channel gain weight includes:

[0041] Calculating a ranking function value for each user based on the channel gain, the queue backlog, the channel gain weight, and a preset ranking function;

[0042] The users are sorted according to the sorting function value to obtain the power allocation sorting of the users.

[0043] In addition, to achieve the above-mentioned purpose, the present application also provides a communication resource allocation device for satellite Internet, and the communication resource allocation device for satellite Internet includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the computer program is executed by the processor, the steps of the communication resource allocation method for satellite Internet as described in any one of the above items are implemented.

[0044] In addition, to achieve the above-mentioned purpose, the present application also provides a computer-readable storage medium, on which a communication resource allocation program for satellite Internet is stored. When the communication resource allocation program for satellite Internet is executed by a processor, the steps of the communication resource allocation method for satellite Internet as described in any of the above items are implemented.

[0045] This application proposes a satellite internet communication resource allocation method, device, and computer-readable storage medium. The method obtains current environmental status information of the satellite communication network, including the information age, channel gain, and queue backlog of each user; inputs the information age, channel gain, and queue backlog into a trained deep reinforcement learning model to obtain the current system action, including the channel gain weight and power allocation coefficient for each user. Compared to conventional approaches that rank channels from best to worst, this method inputs information age and other environmental status information (channel gain and queue backlog) into the trained deep reinforcement learning model. The trained deep reinforcement learning model then dynamically adjusts the information age and other environmental status information to obtain the channel gain weight and power allocation coefficient for each user. This method is more adaptable to the complex and changing environment of the satellite communication network and can effectively improve the timeliness of satellite internet communications. Furthermore, the power allocation ranking of each user is determined based on the channel gain, queue backlog, and channel gain weight; and communication resources of the satellite communication network are allocated based on the power allocation coefficient and power allocation ranking. Compared to simply using the user's power allocation coefficient as the current system action, this application also determines the power allocation ranking of each user based on channel gain, the queue backlog, and the channel gain weight. Based on the appropriate power allocation coefficient and power allocation ranking, the communication resources of the satellite communication network are allocated, thereby generating gains in the information age performance of the satellite communication network, thereby effectively improving the information timeliness of satellite Internet communications. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a schematic diagram of a scenario involving a satellite communication network according to an embodiment of the present application;

[0047] Figure 2 A schematic diagram of an update model for information age involved in an embodiment of the present application;

[0048] Figure 3 This is a flow chart of a first embodiment of the satellite Internet communication resource allocation method of the present application;

[0049] Figure 4 This is a flow chart of a second embodiment of the satellite Internet communication resource allocation method of the present application;

[0050] Figure 5 A schematic diagram of a training scenario for an untrained deep reinforcement learning model according to an embodiment of the present application;

[0051] Figure 6 This is a flow chart of a third embodiment of the satellite Internet communication resource allocation method of the present application;

[0052] Figure 7 Schematic diagram of simulation results of the first simulation experiment involved in the embodiment of the present application;

[0053] Figure 8 Schematic diagram of simulation results of the second simulation experiment involved in the embodiment of the present application;

[0054] Figure 9 Schematic diagram of simulation results of the third simulation experiment involved in the embodiment of the present application;

[0055] Figure 10 Schematic diagram of simulation results of the fourth simulation experiment involved in the embodiment of the present application;

[0056] Figure 11 Schematic diagram of simulation results of the fifth simulation experiment involved in the embodiment of the present application;

[0057] Figure 12 Schematic diagram of simulation results of the sixth simulation experiment involved in the embodiment of the present application;

[0058] Figure 13 Schematic diagram of simulation results of the seventh simulation experiment involved in the embodiment of the present application;

[0059] Figure 14 Schematic diagram of simulation results of the eighth simulation experiment involved in the embodiment of the present application;

[0060] Figure 15 This is a schematic diagram of a satellite Internet communication resource allocation device involved in an embodiment of the present application.

[0061] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0062] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0063] See also Figure 1 , Figure 1 This is a schematic diagram of a scenario involving a satellite communication network according to an embodiment of the present application.

[0064] like Figure 1As shown, in existing satellite communication networks, given the limited satellite spectrum and inter-beam interference, low-orbit satellites 10 typically serve terrestrial users using hybrid multiple access. LEO satellites 10 use an orthogonal multiple access mechanism to serve different orthogonal beams 20. The frequency band on LEO satellite 10 is divided into three sub-bands to ensure that the spectrum resources allocated by adjacent orthogonal beams 20 do not overlap. Furthermore, LEO satellites 10 use a non-orthogonal multiple access mechanism for downlink data transmission with multiple terrestrial user devices within the coverage area of non-orthogonal beams 30. Specifically, at the LEO satellite 10 end, signals with different transmit powers are fully frequency-multiplexed and distinguished only by power. At the receiving end (e.g., at the terrestrial user device), a serial interference cancellation algorithm is used to sequentially demultiplex all user signals based on different channel gains, effectively utilizing the satellite's limited spectrum resources.

[0065] See also Figure 2 , Figure 2 A schematic diagram of the update model of information age involved in the embodiments of this application. In simple terms, the age of information (AoI) is a criterion for measuring the freshness of information. Figure 2 In, a i (t) is the value of the information age, that is, the AoI value. is an indicator function indicating whether SIC (Successive Interference Cancellation) decoding is successfully completed at user i. Slots is a time slot, and time slot t∈[0,T]. For example, if d i (t) = 1, indicating that the SIC decoding of user i is successful. Assume a i (0)=0, then the updated AoI at user i is expressed as follows:

[0066]

[0067] Among them, A max is the preset AoI upper limit. It can be seen that the AoI value depends on whether the SIC decoding is successful. When the SIC decoding of user i is successful in time slot t, so as to generate a corresponding update signal in time slot t, the AoI value will become 1 in the next time slot t+1. Otherwise, the AoI value increases by 1. However, the AoI value will not exceed the preset AoI upper limit A. max , when the AoI value reaches the preset AoI upper limit value A max After that, SIC decoding will be performed to update the information for the user.

[0068] Reference Figure 3 , Figure 3 This is a flow chart of the first embodiment of the satellite Internet communication resource allocation method of the present application.

[0069] like Figure 3 As shown, an embodiment of the present application provides a method for allocating communication resources of a satellite Internet, and the method for allocating communication resources of a satellite Internet includes the following steps:

[0070] Step S100, obtaining current environment status information of the satellite communication network, wherein the current environment status information includes information age, channel gain, and queue backlog of each user;

[0071] In this embodiment, it should be noted that the satellite communication network is a satellite Internet communication network that adopts a non-orthogonal multiple access mechanism.

[0072] The execution subject of this embodiment may be a low-orbit satellite. By monitoring the satellite communication network, the low-orbit satellite can obtain the current environmental status information of the satellite communication network, wherein the current environmental status information includes the information age, channel gain, and queue backlog of each user. In the communication system of the satellite Internet, the low-orbit satellite can be regarded as an intelligent agent, and the current environment is composed of the user's information age and channel status. The channel status includes channel gain and queue backlog. The user's information age, channel gain, and queue backlog together constitute the current environmental status observed in the satellite communication network, that is, the environmental status information st = {st 1 , st 2 , st 3}, assuming that there are K users in the satellite communication network, the information age st 1 = {a1(t), a2(t), a3(t), …, a K (t)}, channel gain st 2 = {g1(t), g2(t), g3(t), …, g K (t)}, queue backlog st 3 = {Q1(t), Q2(t), Q3(t), …, Q K (t)}.

[0073] Step S200: Inputting the information age, the channel gain, and the queue backlog into a trained deep reinforcement learning model to obtain a current system action, wherein the current system action includes a channel gain weight and a power allocation coefficient for each user;

[0074] In this embodiment, it should be noted that the trained deep reinforcement learning model is a model obtained by training an untrained deep reinforcement learning model using a preset training set, wherein the preset training set includes first environmental state information, a first system action corresponding to the first environmental state information, and a first system reward value. The first system action is the action obtained by inputting the first environmental state information into the untrained deep reinforcement learning model; the first system reward value is a value generated based on the second environmental state information generated by the first system action, the first environmental state information, and a preset system reward function; the preset system reward function is a function constructed based on preset constraints and long-term average information age information.

[0075] In addition, the trained deep reinforcement learning model can be a deep reinforcement learning model such as a PPO (Proximal Policy Optimization) model and a DDPG (Deep Deterministic Policy Gradient) model.

[0076] In this embodiment, by inputting the information age, the channel gain and the queue backlog into the trained deep reinforcement learning model, the current system action corresponding to the current environmental state information output by the trained deep reinforcement learning model can be obtained. The current system action includes the channel gain weight and power allocation coefficient of each user. The current system action consists of two parts: the channel gain weight v(t) of each user and the allocated power coefficient, that is, the current system action at = {at1, at2}, where at1 = v(t) represents the weight value of the channel gain (that is, the channel gain weight). When the environmental state information of the current time slot (that is, the current environmental state information) is input into the trained deep reinforcement learning model, the current system action corresponding to the current environmental state information output by the trained deep reinforcement learning model can be obtained, thereby obtaining the channel gain weight a of each user that the trained deep reinforcement learning model considers to be optimal at the moment. t 1 and the distribution power coefficient at 2 .

[0077] Distributed power coefficient at 2 = { , , ,…, };

[0078] Among them, the distribution power coefficient at 2The allocated power includes the power allocated to each user in the satellite communication network. It is understood that the allocated power satisfies the preset single-time slot peak constraint, that is, the allocated power is less than or equal to the preset single-time slot peak. In this embodiment, by inputting the information age, the channel gain, and the queue backlog into a trained deep reinforcement learning model, the trained deep reinforcement learning model can output its optimal power allocation coefficient, making the power allocation system for each user close to optimal, thereby improving the utilization of communication resources and also facilitating the improvement of information timeliness of the satellite communication network.

[0079] Step S300, determining a power allocation order for each user based on the channel gain, the queue backlog, and the channel gain weight;

[0080] An appropriate power allocation sequence can improve the long-term average information age performance of a satellite communication network. Since channel gain and queue backlog simultaneously influence the optimal power allocation ranking of users, a corresponding preset ranking function can be constructed based on the channel gain and queue backlog. After obtaining the channel gain, queue backlog, and channel gain weight, the ranking function value for each user can be calculated based on the preset ranking function. Users are then ranked according to the ranking function value to obtain a power allocation ranking for each user. Because the channel gain weight is dynamically adjusted by a trained deep reinforcement learning model based on information age and other environmental state information that affects information age, the channel gain weight obtained in this embodiment is more adaptable to the complex and changing environment of satellite communication networks than traditional ranking methods (such as ListNet). This makes the power allocation ranking in this embodiment more aligned with the current environmental state of the satellite communication network, thereby improving the information timeliness of the satellite communication network.

[0081] The step of determining the power allocation order of each user according to the channel gain, the queue backlog, and the channel gain weight includes:

[0082] Step S310, calculating a ranking function value of each user according to the channel gain, the queue backlog, the channel gain weight, and a preset ranking function;

[0083] Step S320: sort the users according to the sorting function value to obtain the power allocation sorting of each user.

[0084] As an example, the preset sorting function may be the following formula:

[0085]

[0086] Among them, v(t) is the channel gain weight, g i (t) is the channel gain, Qi (t) is the queue backlog, F i (Q i (t), g i (t)) is the ranking function value.

[0087] In this embodiment, the ranking function value of each user can be calculated based on the channel gain, the queue backlog, the channel gain weight and a preset ranking function, and then the users can be sorted in order of the ranking function values to obtain the power allocation ranking of each user.

[0088] Step S400: Allocate communication resources of the satellite communication network based on the power allocation coefficient and the power allocation ranking.

[0089] In this embodiment, it should be noted that since communication signals with different transmission powers of non-orthogonal multiple access technology are completely multiplexed on the communication frequency and are only distinguished by communication power, the communication resources include the communication power of the satellite communication network.

[0090] In this embodiment, communication resources of the satellite communication network are allocated by allocating power to each user in the satellite communication network according to the power allocation ranking and allocating communication power values to each user in the satellite communication network based on the power allocation coefficient. Compared to traditional methods that only solve for the optimal power allocation coefficient, this embodiment optimizes both the power allocation coefficient and the power allocation ranking, effectively improving the information age of the satellite communication network and thereby improving the timeliness of the information in the satellite communication network.

[0091] In one embodiment of the present application, a method for allocating communication resources for a satellite internet network is provided. The method obtains current environmental status information of a satellite communication network, including the information age, channel gain, and queue backlog of each user; inputs the information age, channel gain, and queue backlog into a trained deep reinforcement learning model to obtain a current system action, including the channel gain weight and power allocation coefficient for each user. Compared to conventional methods that rank channels from best to worst, the present method inputs information age and other environmental status information (channel gain and queue backlog) that influence information age into the trained deep reinforcement learning model. The trained deep reinforcement learning model then dynamically adjusts the information age and other environmental status information to obtain the channel gain weight and power allocation coefficient for each user. This method is more adaptable to the complex and changing environment of the satellite communication network and can effectively improve the timeliness of information in satellite internet communications. Furthermore, a power allocation ranking is determined for each user based on the channel gain, queue backlog, and channel gain weight; and communication resources of the satellite communication network are allocated based on the power allocation coefficient and the power allocation ranking. Compared to simply using the user's power allocation coefficient as the current system action, this application also determines the power allocation ranking of each user based on channel gain, the queue backlog, and the channel gain weight. Based on the appropriate power allocation coefficient and power allocation ranking, the communication resources of the satellite communication network are allocated, thereby generating gains in the information age performance of the satellite communication network, thereby effectively improving the information timeliness of satellite Internet communications.

[0092] Reference Figure 4 , Figure 4 This is a flow chart of the second embodiment of the satellite Internet communication resource allocation method of the present application.

[0093] A second embodiment of the present application provides a method for allocating communication resources for a satellite internet. Before the step of inputting the user information age, the channel gain, and the queue backlog into a trained deep reinforcement learning model to obtain a current system action in step S200, the method includes:

[0094] Step A10, obtaining first environmental state information of the satellite communication network and an untrained deep reinforcement learning model;

[0095] Step A20: inputting the first environmental state information into an untrained deep reinforcement learning model to obtain a first system action;

[0096] Step A30: executing the first system action to obtain second environment state information generated based on the first system action;

[0097] Step A40: generating a corresponding first system reward value according to the first environment state information, the second environment state information, and a preset system reward function;

[0098] Step A50: Based on the first system reward value and the preset loss function, update the parameters of the untrained deep reinforcement learning model until the untrained deep reinforcement learning model converges to obtain a trained deep reinforcement learning model.

[0099] In this embodiment, it should be noted that the first environmental state information includes a first information age, a first channel gain, and a first queue backlog. The first system action includes a first channel gain weight and a first power allocation coefficient. The preset system reward function is constructed based on preset constraints and long-term average information age information, and is used to reward the first system action that satisfies the preset constraints and minimizes the long-term average information age.

[0100] In addition, it should be noted that the untrained deep reinforcement learning model can be a deep reinforcement learning model such as a PPO (Proximal Policy Optimization) model and a DDPG (Deep Deterministic Policy Gradient) model.

[0101] This embodiment obtains first environmental state information of a satellite communication network and an untrained deep reinforcement learning model, and inputs the first environmental state information into the untrained deep reinforcement learning model, thereby obtaining a first system action output by the untrained deep reinforcement learning model under current model parameters. After executing the first system action, second environmental state information generated based on the first system action can be obtained by monitoring the satellite communication network. The second environmental state information includes a second information age, a second channel gain, and a second queue backlog. A corresponding first system reward value is then generated based on the first and second environmental state information and a preset system reward function. The first system reward value is used to indicate whether the first system action satisfies the preset constraints and minimizes the long-term average information age. The current model parameters of the untrained deep reinforcement learning model can then be updated based on the first system reward value and a preset loss function. The untrained deep reinforcement learning model is determined to have converged when the function loss value corresponding to the preset loss function is minimized. The converged untrained deep reinforcement learning model is then used as the trained deep reinforcement learning model to obtain a trained deep reinforcement learning model. It is understood that the preset loss function includes a system reward term corresponding to the first system reward value.

[0102] The untrained deep reinforcement learning model includes a new policy network, an old policy network, and an evaluation network. The step of updating the parameters of the untrained deep reinforcement learning model based on the first system reward value and the preset loss function in step A50 includes:

[0103] Step B10, calculating the fitting cumulative discount reward based on the preset value function and the evaluation network;

[0104] Step B20, calculating the current accumulated discount reward based on the first system reward value;

[0105] Step B30, calculating a corresponding loss function value based on the current cumulative discount reward, the fitted cumulative discount reward, and the preset loss function;

[0106] Step B40: updating parameters of the new policy network, the old policy network, and the evaluation network based on the loss function value.

[0107] See also Figure 5 , Figure 5 Schematic diagram of the training scenario of the untrained deep reinforcement learning model involved in the embodiment of this application. In this embodiment, the untrained deep reinforcement learning model is a PPO model. Figure 5 Where Environment is the satellite communication network of this application, and Experience Data is a database for data storage. The untrained deep reinforcement learning model includes a policy network (Actor Network) and an evaluation network (Critic Network). The policy network includes a new policy network (Actor Network π) with a new policy parameter θ. θ ) and with the old policy parameters θ old The old policy network (Actor Networkπ θold For example, in the PPO model framework, the neural network structure consists of one input layer, two hidden layers, and one output layer. The first hidden layer has 64 neurons, the second hidden layer has 32 neurons, and the Tanh function is used as the activation function. Furthermore, the Adam optimizer is used to update the neural network parameters. The parameters during training can be determined through empirical and experimental tuning.

[0108] When the first environmental status information s t After inputting the new policy network, the new policy network outputs the corresponding first system action a based on the new policy parameter θ t , the satellite communication network performs the first system action a t After that, the second environment status information s is obtained t+1, then the satellite communication network is based on the first environmental status information s t and the second environmental status information s t+1 , and the corresponding first system reward value r is obtained t+1 In addition, the new policy network also outputs the corresponding new policy action probability π θ (a t |s t ), the old policy network also outputs the corresponding old policy action probability π θold (a t |s t ). According to the first system reward value, the current cumulative discount reward (actual reward) can be calculated, and the evaluation network calculates the fitted cumulative discount reward (fitted reward) based on the preset value function and the first system reward value. In this way, the advantage of the actual reward over the fitted average reward can be determined. According to the current cumulative discount reward, the fitted cumulative discount reward, the new and old strategy probability ratio (that is, the ratio between the new strategy action probability and the new strategy action probability) and the preset loss function, the corresponding loss function value is calculated. Therefore, the parameters of the new strategy network, the old strategy network and the evaluation network can be updated based on the loss function value. See Figure 5 , the loss function value includes the new policy loss function value L(θ) and the evaluation loss function value L(φ), so that the old policy parameters of the old policy network are updated by the new policy parameters of the new policy network, the new policy parameters of the new policy network are updated by the new policy loss function value L(θ), and the evaluation parameters of the evaluation network are updated by the evaluation loss function value L(φ).

[0109] The preset loss function includes a new strategy loss function and an evaluation loss function. The step of calculating the corresponding loss function value according to the current cumulative discount reward, the fitted cumulative discount reward, and the preset loss function in step B30 includes:

[0110] Step B31, obtaining a new and old strategy probability ratio between the new strategy network and the old strategy network;

[0111] Step B32, taking the difference between the current cumulative discount reward and the fitted cumulative discount reward as the advantage function value;

[0112] Step B33, calculating a new strategy loss function value based on the advantage function value, the probability ratio of the new strategy to the old strategy, and the new strategy loss function;

[0113] Step B34, calculating an evaluation loss function value based on the current cumulative discount reward, the fitted cumulative discount reward, and the evaluation loss function;

[0114] Step B35: Using the new strategy loss function value and the evaluation loss function value as the loss function value corresponding to the untrained deep reinforcement learning model.

[0115] In this embodiment, it should be noted that after the first environmental state information is input into the new policy network and the old policy network, the new policy network outputs the corresponding new policy action probability (i.e., the probability distribution of the action output by the new policy network given the first environmental state information), and the old policy network outputs the corresponding old policy action probability (i.e., the probability distribution of the action output by the old policy network given the first environmental state information). The ratio of the new policy action probability to the old policy action probability is then used as the new-old policy probability ratio. For example, the calculation formula for the new-old policy probability ratio is as follows:

[0116] .

[0117] Among them, r t (θ) is the probability ratio of the new and old strategies, π θ (a t |s t ) is the new strategy action probability, π θold (a t |s t ) is the old policy action probability.

[0118] The present application can calculate the new and old strategy probability ratio between the new strategy network and the old strategy network by obtaining the new strategy action probability output by the new strategy network and the old strategy action probability output by the old strategy network. Then, the difference between the current cumulative discount reward and the fitted cumulative discount reward is used as the advantage function value. For example, the preset advantage function for calculating the advantage function value can be expressed as follows:

[0119]

[0120] in, is the current cumulative discount reward, which is calculated based on the preset state-action value function and the first system reward value, indicating the first environment state information Take the first system action The actual cumulative discount reward value obtained. To fit the cumulative discount reward, the evaluation network is calculated by fitting, and its main function is to fit the state information of the first environment Take the first system action The cumulative discount reward value obtained.

[0121] For example, the calculation formula for the current accumulated discount reward value is as follows:

[0122] Current cumulative discount reward value .

[0123] Among them, γ∈[0,1) is the discount factor, which indicates the importance of the reward of the future time slot and can be selected according to actual needs. t+1 is the first system reward value.

[0124] The advantage function value is used to represent the state information of the first environment Take the first system action , the advantage of the actual reward (current cumulative discounted reward) over the fitted reward (fitted cumulative discounted reward).

[0125] Furthermore, this embodiment calculates the new strategy loss function value based on the advantage function value, the probability ratio of the new and old strategies and the new strategy loss function. For example, the new strategy loss function of the new strategy network is for:

[0126]

[0127] in, is the advantage function value, π θ (a t |s t ) / π θold (a t |s t ) is the probability ratio of the new and old strategies; is the expectation, and is the empirical average obtained by sampling the system actions through the probability distribution.

[0128] Furthermore, considering that the strategy update is very sensitive to the continuous action space, a clip function is defined in the PPO model The clip function can prevent the new policy from differing too much from the old policy by limiting the probability ratio of the new and old policies to a preset probability ratio interval [1-ε, 1+ε]. Finally, the new policy loss function of the new policy network can be calculated as follows:

[0129]

[0130] Furthermore, this embodiment calculates the evaluation loss function value based on the current cumulative discount reward, the fitting cumulative discount reward and the evaluation loss function. For example, the evaluation loss function of the evaluation network is as follows:

[0131]

[0132] in, is the expectation, and is the empirical average obtained by sampling the system actions through the probability distribution. is the discount factor, rt+1 is the first system reward value, is the cumulative discounted reward for fitting.

[0133] Thus, this embodiment obtains the new strategy loss function value and the evaluation loss function value, and the new strategy loss function value and the evaluation loss function value can be used as the loss function value corresponding to the untrained deep reinforcement learning model.

[0134] The step of updating the parameters of the new policy network, the old policy network, and the evaluation network based on the loss function value in step B40 includes:

[0135] Step B41, obtaining new policy parameters of the new policy network, and updating the old policy parameters of the old policy network to the new policy parameters;

[0136] Step B42, updating the new policy parameters of the new policy network by gradient ascent according to the new policy loss function value;

[0137] Step B43: updating the evaluation parameters of the evaluation network by gradient descent according to the evaluation loss function value.

[0138] In this embodiment, the new policy parameters of the new policy network can be obtained and the old policy parameters of the old policy network can be updated to the new policy parameters. According to the new policy loss function value, the new policy parameters of the new policy network are updated by gradient ascent. The first update formula for updating the new policy parameters of the new policy network by gradient ascent is: , θ is the new strategy parameter, is the default learning rate of the new policy network, is the new strategy loss function value. And according to the evaluation loss function value, the evaluation parameters of the evaluation network are updated by gradient descent. Among them, the second update formula for updating the evaluation parameters of the evaluation network by gradient descent is: , the φ is the evaluation parameter, is the default learning rate of the new policy network, = is the evaluation loss function value. Thus, this embodiment completes the update of the parameters of the new policy network, the old policy network, and the evaluation network in the PPO model.

[0139] In the second embodiment of the present application, by obtaining first environmental state information and an untrained deep reinforcement learning model of a satellite communication network; inputting the first environmental state information into the untrained deep reinforcement learning model to obtain a first system action; executing the first system action to obtain second environmental state information generated based on the first system action; generating a corresponding first system reward value based on the first environmental state information and the second environmental state information and a preset system reward function; updating the parameters of the untrained deep reinforcement learning model based on the first system reward value and the preset loss function until the untrained deep reinforcement learning model converges to obtain a trained deep reinforcement learning model. This embodiment designs a composite space of system actions and environmental states through a deep reinforcement learning model to simultaneously solve the two optimization problems of power allocation sorting and power allocation coefficient. Compared with only taking the user's power allocation coefficient as the system action, this embodiment not only includes the user's power allocation coefficient, but also includes the channel gain weight for determining the user's power allocation sorting, which can effectively improve the information age of the satellite communication network, thereby improving the information timeliness of the satellite communication network. In addition, by using the first environmental state information to train the untrained deep reinforcement learning model to obtain the trained deep reinforcement learning model, the computational complexity can be effectively reduced and the accuracy of the power allocation strategy can be improved.

[0140] Reference Figure 6 , Figure 6 This is a flow chart of the third embodiment of the satellite Internet communication resource allocation method of the present application.

[0141] A third embodiment of the present application provides a method for allocating communication resources for a satellite internet. Prior to the step of generating a corresponding first system reward value based on the first environmental state information, the second environmental state information, and a preset system reward function in step A40, the method includes:

[0142] Step C10: constructing a first optimization function based on preset constraints and long-term average information age information;

[0143] In this embodiment, it should be noted that the long-term average information age information includes the long-term average information age function of the satellite communication network. In order to select an appropriate power allocation strategy to improve information timeliness, the long-term average information age can be used to represent the information freshness of all K users in the satellite communication network. Therefore, the long-term average information age function is as follows:

[0144] Long-term average information age

[0145] Among them, a i (t) is the value of the information age, that is, the AoI value; T is the time slot, K is the number of users, and E[] is the expected value.

[0146] While optimizing information timeliness, we must ensure that the satellite communication network meets the corresponding preset constraints to avoid problems such as damaged equipment, poor channel conditions, user information age, and data overflow. For example, the preset constraints may include the following three long-term constraints and one short-term constraint:

[0147] (1) Long-term / short-term power constraints;

[0148] : Single time slot power peak constraint;

[0149] : Long-term average power upper limit constraint;

[0150] (2) Minimum throughput constraint;

[0151] ;

[0152] (3) Network stability constraints;

[0153] ;

[0154] in, is the user power, P max is the upper limit of the power of a single time slot, K is the number of users, P mean is the long-term average power limit value, is the long-term average throughput of the user, br i (t) is the user's data leaving rate, ar i (t) is the user's data arrival rate, Q i (t) is the queue backlog of the user.

[0155] Therefore, this embodiment can avoid damage to the device through long-term / short-term power constraints, protect users with poor channel conditions through minimum throughput constraints, and prevent data overflow through network stability constraints.

[0156] Considering the above-mentioned preset constraints and the long-term average information age information, this embodiment can construct a first optimization function based on the preset constraints and the long-term average information age information. The expression of the first optimization function is as follows:

[0157]

[0158] in, is the user power, P max is the upper limit of the power of a single time slot, K is the number of users, P mean is the long-term average power limit value, is the long-term average throughput of the user, bri (t) is the user's data leaving rate, ar i (t) is the user's data arrival rate, Q i (t) is the queue backlog of the user, h i is the minimum throughput limit.

[0159] Step C20: Based on Lyapunov optimization theory, convert the first optimization function into a drift plus penalty term, and determine a minimization upper bound corresponding to the drift plus penalty term;

[0160] The deep reinforcement learning model is solvable. Lyapunov optimization theory can be used to transform the problem of minimizing the long-term average information age under the preset constraints into the problem of minimizing the drift plus penalty term. Therefore, this embodiment transforms the first optimization function into a drift plus penalty term based on Lyapunov optimization theory and determines the upper bound for minimizing the drift plus penalty term.

[0161] The step of converting the first optimization function into a drift plus penalty term based on Lyapunov optimization theory in step C20 includes:

[0162] Step C21: constructing a power consumption liability queue, a queue backlog, and a throughput liability queue according to the first optimization function, and generating a queue vector consisting of the power consumption liability queue, the queue backlog, and the throughput liability queue;

[0163] Step C22: generating a quadratic Lyapunov function according to the queue vector based on Lyapunov optimization theory, and determining a Lyapunov drift corresponding to the quadratic Lyapunov function;

[0164] Step C23: constructing a single-slot penalty function according to the queue vector and the first optimization function;

[0165] Step C24: Determine a corresponding drift plus penalty term based on the Lyapunov drift and single-slot penalty function.

[0166] Taking the preset constraints in the first optimization function as long-term / short-term power constraints, minimum throughput constraints, and network stability as an example, three corresponding virtual queues can be constructed. It is understandable that if these three virtual queues are stable on average over the long term, then the three constraints in the corresponding preset constraints are also satisfied. The virtual queues include:

[0167] (1) Power consumption debt queue P(t):

[0168]

[0169] (2) Queue backlog Qi(t):

[0170]

[0171] (3) Throughput debt queue Ui(t):

[0172]

[0173] Assumptions Representing a vector composed of P(t), Qi(t) and Ui(t), a quadratic Lyapunov function can be generated based on the queue vector based on Lyapunov optimization theory. The quadratic Lyapunov function can be expressed as follows:

[0174]

[0175] Then, the Lyapunov drift of the quadratic Lyapunov function can be determined, which is used to describe the changes of the Lyapunov function in different time slots:

[0176] .

[0177] Then, we can reduce the Lyapunov drift To maintain the stability of the virtual queue, a lower Lyapunov drift can prevent the virtual queue from entering a congested state, so that the preset constraints are met.

[0178] In order to minimize the long-term average information age, a single-slot penalty function can be constructed based on the queue vector and the first optimization function. , The smaller it is, the better the long-term average information age performance is:

[0179]

[0180] Lyapunov optimization theory is used to convert the preset constraints into drift terms, and the optimization objective (i.e., the long-term average information age) into a penalty term, resulting in the corresponding drift plus penalty term. The drift plus penalty term is the sum of the product of the preset importance weight and the single-slot penalty function and the Lyapunov drift. The drift plus penalty term DPP can be expressed as follows:

[0181]

[0182] Where V>0 is the preset importance weight, which means the long-term average information age is used as a single time slot penalty function Importance in the final drift plus penalty term.

[0183] Furthermore, according to Lyapunov optimization theory, we can get the drift plus penalty term The upper bound of is obtained by minimizing it, thereby obtaining the minimized upper bound corresponding to the drift plus penalty term.

[0184] The minimization upper bound can be expressed as follows:

[0185]

[0186] Where c is a constant term.

[0187] Since the power allocation strategy for improving information timeliness in the embodiment of the present application is a single time slot optimization problem, the symbol t in the above formula can be ignored and the minimization upper bound can be simplified to:

[0188]

[0189] Step C30: constructing a corresponding preset system reward function according to the minimization upper bound.

[0190] For example, the reward of the preset system reward function can be set as the change value of the drift plus the penalty term in the previous and next time slots. The preset system reward function can be as follows:

[0191]

[0192] Among them, r i+1 (s,a) is the system reward value under the environment state s and system action a, DPP t Add a penalty term for the drift of the current time slot, DPP t+1 Add a penalty term for the drift of the next time slot. Considering that there is a single time slot power upper limit value P in each time slot max Therefore, when the constraint is not satisfied, the system reward value of the current time slot is set to the penalty term -PEN, and PEN>0. This makes the untrained deep reinforcement learning model choose the system reward value that satisfies the single time slot power upper limit P. max When the resource allocation strategy is used, the corresponding reward can be obtained, but the power used by the selected strategy exceeds the single time slot power upper limit P max With this setting, the trained deep reinforcement learning model will avoid choosing an option that cannot satisfy P max As iterations of the untrained deep reinforcement learning model accumulate, it converges to an optimal state, where the value of the drift plus penalty term does not change significantly and remains at a small value. This allows the power allocation strategy output by the untrained deep reinforcement learning model to approach optimality by maximizing the long-term cumulative reward.

[0193] In addition, the power allocation strategy output by the untrained deep reinforcement learning model can be made close to optimal by maximizing the long-term cumulative discounted reward. The cumulative discounted reward is calculated as follows:

[0194]

[0195] in is a discount factor that represents the importance of rewards in future time slots. This allows the power allocation strategy output by the untrained deep reinforcement learning model to better align with the user's expectations for future time slots, helping to adapt to the complex and changing environment of the satellite communication network and effectively improving the timeliness of information in satellite internet communications.

[0196] In the third embodiment of the present application, a first optimization function is constructed based on preset constraints and long-term average information age information; based on Lyapunov optimization theory, the first optimization function is converted into a drift-plus-penalty term, and an upper bound corresponding to the minimization of the drift-plus-penalty term is determined; and based on the minimization upper bound, a corresponding preset system reward function is constructed. Thus, this embodiment transforms the first optimization function through Lyapunov optimization theory modeling, thereby converting the problem of minimizing information age under preset constraints into a problem of minimizing the drift-plus-penalty term in Lyapunov theory, thereby making the deep reinforcement learning model solvable. Furthermore, this embodiment can achieve a trade-off between optimization performance and modeling accuracy by adjusting the preset importance weights in the drift-plus-penalty term, depending on the requirements for optimization performance and modeling accuracy.

[0197] In addition, the present application also conducted simulation experiments using different algorithms in the embodiments of the present application, and the experimental results are as follows:

[0198] First, the algorithm complexity analysis table between different algorithms is as follows:

[0199] Table 1 Algorithm complexity analysis table

[0200]

[0201] In this simulation experiment, using an Intel i7-9700 CPU as the platform, we tested the decision-making time overheads of the NOMA-PSO (non-orthogonal multiple access particle swarm optimization) scheme, the NOMA-DDPG (non-orthogonal multiple access deep deterministic policy gradient model) scheme, and the NOMA-PPO (non-orthogonal multiple access proximal decision optimization model) scheme under different user numbers. The results are shown in the table above. Both deep reinforcement learning-based schemes (NOMA-DDPG and NOMA-PPO) have lower decision-making time overheads than NOMA-PSO, with NOMA-DDPG having slightly lower decision-making time overhead than NOMA-PPO. Furthermore, as the number of users increases, the decision-making time overhead of NOMA-PSO increases significantly. At this point, the average time overhead ratio of PPO to PSO (i.e., PPO / PSO) decreases, and the complexity advantage of NOMA-PPO becomes increasingly significant.

[0202] Second, performance comparison simulation;

[0203] See also Figure 7 , Figure 7 This is a schematic diagram of the simulation results of the first simulation experiment involved in the embodiment of the present application. Figure 7 Characterize the impact of different fading parameters on information age in the NOMA-PPO scheme in this application. Figure 7 As shown, Figure 7 The horizontal axis represents the signal-to-noise ratio (SNR), and the vertical axis represents the information age. As the SNR increases, the information age performance of the NOMA-PPO scheme gradually improves under all three fading parameters, especially under FHS fading (frequent high-frequency fading). To obtain clear simulation results, FHS fading parameters were used in subsequent experiments. Figure 7 The number of users for each scheme is K = 3, and the default importance weight V in the drift plus penalty term is 150. PPO-AS plots the information age performance as a function of the signal-to-noise ratio for average fading; PPO-ILS plots the information age performance as a function of the signal-to-noise ratio for infrequent shallow fading; and PPO-FHS plots the information age performance as a function of the signal-to-noise ratio for frequent heavy fading.

[0204] See also Figure 8 and Figure 9 , Figure 8 Schematic diagram of simulation results of the second simulation experiment involved in the embodiment of the present application; Figure 9 This is a schematic diagram of the simulation results of the third simulation experiment involved in the embodiment of the present application. Figure 8 and Figure 9 The convergence of the model during training is simulated. Figure 8The curve of the sliding average reward convergence of the NOMA-PPO scheme is shown in Figure 2. The horizontal axis is the number of episodes (model verification) and the vertical axis is the sliding average reward. The number of users corresponding to each scheme is K=6, the signal-to-noise ratio SNR=17.5dB, and FHS fading. Figure 8 It can be seen that the sliding average reward in the NOMA-PPO scheme begins to converge after about 300 episodes. Since the system reward used in the NOMA-PPO scheme is non-positive, the sliding average reward stabilizes at a negative value close to 0. Figure 9 It can be seen that the NOMA-PPO scheme achieves good information age performance in the first dozens of episodes, and the stability of the information age performance is also better than other schemes. Figure 9 The curves in the figure correspond to the NOMA-AM (Non-Orthogonal Multiple Access-Alternating Minimization) scheme, the NOMA-DDPG scheme, and the NOMA-PPO scheme from top to bottom. Figure 9 In the figure, the horizontal axis is the number of episodes and the vertical axis is the average information age. Figure 9 The number of users corresponding to each scheme is K=8, SNR=15dB, FHS fading

[0205] See also Figure 10 and Figure 11 , Figure 10 Schematic diagram of simulation results of the fourth simulation experiment involved in the embodiment of the present application; Figure 11 This is a schematic diagram of the simulation results of the fifth simulation experiment involved in the embodiment of the present application. Figure 10 and Figure 11 The average information age performance of the NOMA-PPO scheme as the signal-to-noise ratio changes under single / multi-antenna satellite conditions is demonstrated, and its performance is compared with that of NOMA-PSO, NOMA-DDPG, NOMA-G (the power allocated to users is inversely proportional to the user's channel gain), NOMA-Q (the power allocated to users is proportional to the user's queue backlog), and NOMA-PPO (Baseline) (the part that optimizes the user power allocation order is removed from the NOMA-PPO scheme). Figure 10 In the scheme, the number of users K=3, the number of satellite antennas N=1, the preset importance weight V=150, and FHS fading are used. Figure 11 Each scheme corresponds to the number of users K = 3, the number of satellite antennas N = 4, the weight V = 150, and FHS fading. Simulation results show that the NOMA-PPO scheme, which incorporates user power allocation order, outperforms the baseline scheme without considering allocation order in terms of average information age. The NOMA-PPO scheme achieves optimal average information age performance under various conditions, regardless of the number of satellite antennas and signal-to-noise ratio.

[0206] See also Figure 12and Figure 13 , Figure 12 Schematic diagram of simulation results of the sixth simulation experiment involved in the embodiment of the present application; Figure 13 This is a schematic diagram of the simulation results of the seventh simulation experiment involved in the embodiment of the present application. Figure 12 The average information age performance of different algorithms changes with the number of users. Figure 12 In the example, the signal-to-noise ratio (SNR) of each scheme is 17.5 dB, the number of satellite antennas is 1, the preset importance weight is 150, and FHS fading occurs. Figure 13 The average information age performance of the same algorithm changes with the number of users. Figure 13 In the example, the signal-to-noise ratio (SNR) of each scheme is 17.5 dB, the number of satellite antennas is N = 4, the preset importance weight is V = 150, and FHS fading occurs. Figure 12 and Figure 13 The results show how the average information age performance of different schemes changes with the number of users in the single-antenna and multi-antenna scenarios. For a single satellite antenna, when the number of users exceeds 6, the NOMA-PPO scheme exhibits a certain performance gap with other schemes. However, it still achieves the best average information age performance with an increasing number of users. When the number of users reaches 10, the average information age of the proposed NOMA-PPO scheme is 0.864 times that of the NOMA-DDPG scheme and 0.521 times that of the NOMA-PSO scheme.

[0207] Finally, see Figure 14 , Figure 14 Schematic diagram of the simulation results of the eighth simulation experiment involved in the embodiment of this application. Figure 14 The influence of the preset importance weight V on the information age performance and the average power consumption of the time slot is simulated. Figure 14 The number of users corresponding to each scheme is K=3, and FHS fading occurs. When V increases, the average information age performance improves, but the average power consumption of users per time slot increases (but will not exceed the set long-term power average upper limit P mean The simulation results confirm that by adjusting the preset importance weight V, a trade-off can be achieved between average message age performance and the length of the virtual queue (most importantly, the power consumption liability queue). Power Consumption: The average power consumed per time slot.

[0208] To sum up, the preferred option for the deep reinforcement learning model in this application is the PPO (Proximal Policy Optimization) model.

[0209] like Figure 15 As shown, Figure 15This is a structural diagram of the satellite Internet communication resource allocation device involved in the embodiment of the present application.

[0210] Exemplarily, the communication resource allocation device of the satellite Internet can be a low-orbit satellite, a PC (Personal Computer), a tablet computer, a portable computer or a server and other devices.

[0211] like Figure 15 As shown, the satellite internet communication resource allocation device may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. It is understood that, Figure 15 The device structure shown in the figure does not constitute a limitation on the satellite Internet communication resource allocation device, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0212] like Figure 15 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a communication resource allocation application for satellite Internet.

[0213] exist Figure 15 In the device shown, the processor 1001 can be used to call the satellite Internet communication resource allocation application stored in the memory 1005 and execute the operations of the satellite Internet communication resource allocation method in the above embodiments.

[0214] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0215] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for allocating communication resources for satellite Internet, characterized in that: The communication resource allocation method of the satellite Internet comprises the following steps: Acquiring current environmental status information of the satellite communication network, wherein the current environmental status information includes information age, channel gain, and queue backlog of each user; Inputting the information age, the channel gain, and the queue backlog into a trained deep reinforcement learning model to obtain a current system action, wherein the current system action includes a channel gain weight and a power allocation coefficient for each user; Determining a power allocation order for each user according to the channel gain, the queue backlog, and the channel gain weight; Based on the power allocation coefficient and the power allocation ranking, communication resources of the satellite communication network are allocated.

2. The satellite Internet communication resource allocation method according to claim 1, wherein: Before the step of inputting the information age, the channel gain, and the queue backlog into a trained deep reinforcement learning model to obtain the current system action, the method includes: obtaining first environmental state information of a satellite communication network and an untrained deep reinforcement learning model; Inputting the first environmental state information into an untrained deep reinforcement learning model to obtain a first system action; Executing the first system action to obtain second environment state information generated based on the first system action; Generate a corresponding first system reward value according to the first environmental state information, the second environmental state information and a preset system reward function; Based on the first system reward value and a preset loss function, parameters of the untrained deep reinforcement learning model are updated until the untrained deep reinforcement learning model converges to obtain a trained deep reinforcement learning model.

3. The satellite internet communication resource allocation method according to claim 2, wherein: The untrained deep reinforcement learning model includes a new policy network, an old policy network, and an evaluation network. The step of updating parameters of the untrained deep reinforcement learning model based on the first system reward value and a preset loss function includes: Calculating a fitted cumulative discounted reward based on a preset value function and the evaluation network; Calculate the current cumulative discount reward based on the first system reward value; Calculate a corresponding loss function value based on the current cumulative discount reward, the fitted cumulative discount reward, and the preset loss function; Parameters of the new policy network, the old policy network, and the evaluation network are updated based on the loss function value.

4. The satellite Internet communication resource allocation method according to claim 3, wherein: The preset loss function includes a new strategy loss function and an evaluation loss function, and the step of calculating a corresponding loss function value based on the current cumulative discount reward, the fitted cumulative discount reward, and the preset loss function includes: Obtaining a new-to-old strategy probability ratio between the new strategy network and the old strategy network; taking the difference between the current cumulative discounted reward and the fitted cumulative discounted reward as the advantage function value; Calculating a new strategy loss function value based on the advantage function value, the probability ratio of the new strategy to the old strategy, and the new strategy loss function; Calculating an evaluation loss function value according to the current cumulative discount reward, the fitted cumulative discount reward, and the evaluation loss function; The new strategy loss function value and the evaluation loss function value are used as the loss function values corresponding to the untrained deep reinforcement learning model.

5. The satellite Internet communication resource allocation method according to claim 4, wherein: The step of updating parameters of the new policy network, the old policy network, and the evaluation network based on the loss function value includes: Acquire new policy parameters of the new policy network, and update old policy parameters of the old policy network to the new policy parameters; According to the new strategy loss function value, the new strategy parameters of the new strategy network are updated by gradient ascent; According to the evaluation loss function value, the evaluation parameters of the evaluation network are updated by gradient descent.

6. The satellite Internet communication resource allocation method according to claim 2, wherein: Before the step of generating a corresponding first system reward value according to the first environmental state information, the second environmental state information and a preset system reward function, the method includes: Constructing a first optimization function based on preset constraints and long-term average information age information; Based on Lyapunov optimization theory, the first optimization function is converted into a drift plus penalty term, and a minimization upper bound corresponding to the drift plus penalty term is determined; According to the minimization upper bound, a corresponding preset system reward function is constructed.

7. The satellite Internet communication resource allocation method according to claim 6, wherein: The step of converting the first optimization function into a drift plus penalty term based on Lyapunov optimization theory includes: constructing a power consumption liability queue, a queue backlog, and a throughput liability queue according to the first optimization function, and generating a queue vector consisting of the power consumption liability queue, the queue backlog, and the throughput liability queue; Based on Lyapunov optimization theory, a quadratic Lyapunov function is generated according to the queue vector, and a Lyapunov drift corresponding to the quadratic Lyapunov function is determined; Constructing a single-slot penalty function according to the queue vector and the first optimization function; According to the Lyapunov drift and single-slot penalty function, a corresponding drift plus penalty term is determined.

8. The satellite internet communication resource allocation method according to any one of claims 1 to 7, wherein: The step of determining the power allocation order of each user according to the channel gain, the queue backlog and the channel gain weight comprises: Calculating a ranking function value for each user based on the channel gain, the queue backlog, the channel gain weight, and a preset ranking function; The users are sorted according to the sorting function value to obtain the power allocation sorting of the users.

9. A satellite Internet communication resource allocation device, characterized in that: The satellite Internet communication resource allocation device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the satellite Internet communication resource allocation method as described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a satellite Internet communication resource allocation program, which, when executed by a processor, implements the steps of the satellite Internet communication resource allocation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Satellite-borne resource allocation method and device, computer equipment and storage medium

    CN110769512A

  • Coding calculation distribution method based on matrix-vector multiplication task in satellite-ground network

    CN114614878A