Multi-beam satellite resource allocation optimization method and system based on D3QN reinforcement learning
Through the multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning, the problem of inefficient resource allocation in the prior art is solved, and efficient allocation under limited resource conditions is achieved to adapt to dynamically changing communication scenarios.
Patent Information
- Application Number
- CN202510347853.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
AI Technical Summary
Existing satellite communication technology is difficult to effectively allocate multi-beam satellite resources under limited channels, spectrum and power resources, resulting in low communication efficiency and inability to adapt to rapidly changing user needs.
The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning is adopted, and optimization indicators and goals are constructed by reconstructing dynamic scenarios, and the optimal resource allocation scheme is solved using the D3QN algorithm to improve spectrum and energy efficiency.
It realizes efficient allocation of multi-beam satellite resources under limited resources, maximizes the combined spectrum efficiency and energy efficiency, and can adapt to communication scenarios with dynamic time changes, and meet user needs.
Smart Images

Figure CN120185690A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite communication resource allocation, and particularly to a multi-beam satellite resource allocation optimization method and system based on D3QN reinforcement learning. Background Art
[0002] Since the emergence of satellite communication, with its advantages such as long communication distance, wide frequency band, and stable communication line, it has played an important role in many fields where ground communication is difficult to achieve, such as ocean communication and long-distance communication. As an extension of satellite technology, multi-beam satellites can further efficiently utilize spectrum resources with their narrow-band beams focused on different geographical locations, and enhance capacity and coverage.
[0003] The amount of signals to be transmitted by a multi-beam satellite system is relatively large at all times, and all information transmissions through the satellite need to be carried out among various channels. Channel resources have become an important content restricting the efficiency of satellite communication. In addition, especially in recent years, with the rapid development of mobile communication technology, many institutions around the world have continuously launched satellites into the air, and the available frequency bands specified internationally have been continuously compressed and utilized. Therefore, existing satellite communication technologies need to meet the communication needs of users under the condition of limited resources. In traditional allocation algorithms, the allocation scheme has poor flexibility. Due to the influence of geographical location factors and other factors, the communication resources of users in each beam are different, and the required service of users generally shows a non-trending change in time. Therefore, traditional allocation methods cannot well adapt to rapidly changing communication conditions in time. Finally, in addition to channel and spectrum resources, the signal transmitter can only provide transmission with limited power, and the amplification function of the satellite for signals is also limited. In order to ensure that the signals transmitted to the downlink user receiver are not submerged in noise and signal interference, achieving effective transmission under limited power resources is also one of the issues that should be considered.
[0004] In view of this, the present application is specifically proposed. Summary of the Invention
[0005] In order to solve the problem of multi-beam satellite resource allocation under the conditions of limited channels, spectrum, and power resources, the purpose of the present invention is to provide a multi-beam satellite resource allocation optimization method and system based on D3QN reinforcement learning, to achieve the goal of efficient allocation of multi-beam GEO satellite resources and maximizing the joint spectrum efficiency and energy efficiency within a cycle time, so as to help the system meet the user's needs as much as possible in terms of spectrum and energy under the condition of limited resources.
[0006] The present invention is realized through the following technical solutions:
[0007] In the first aspect, the present invention provides a multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning, and the method includes:
[0008] Based on the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals, reconstruct the dynamic scenario and obtain the communication link change situation of the user terminals in each time slot;
[0009] According to the communication link change situation of the user terminals in each time slot and the total throughput of the downlink user terminal channels of the GEO multi-beam satellite system in the periodic time slot, construct an optimization index; the optimization index includes energy efficiency and spectral efficiency;
[0010] Perform weighted summation on the energy efficiency and spectral efficiency, establish the optimization objective for the GEO multi-beam satellite system, and set limiting conditions for the optimization objective;
[0011] Use the D3QN reinforcement learning algorithm to solve the optimization objective and obtain the optimal resource allocation scheme.
[0012] Furthermore, the geographical location information of the ground main axis projection of the GEO multi-beam satellite refers to the longitude and latitude information of the projection of the main axis of the GEO multi-beam satellite on the ground;
[0013] The geographical location information of the beams released by the GEO multi-beam satellite refers to the longitude and latitude information of the coverage areas of the multiple circular narrowband beams released by the GEO multi-beam satellite;
[0014] The geographical location information of the uplink and downlink user terminals includes the geographical location information of the uplink user terminal and the geographical location information of the downlink user terminal; the geographical location information of the uplink user terminal refers to the longitude and latitude position information of the uplink user transmitter; the geographical location information of the downlink user terminal refers to the longitude and latitude position information of the downlink user receiver.
[0015] Furthermore, based on the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals, reconstruct the dynamic scenario and obtain the communication link change situation of the user terminals in each time slot, including:
[0016] Within the longitude and latitude area of the preset two-dimensional scenario, set the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals;
[0017] Based on the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals, establish a stochastic process of the uplink and downlink user terminal communication links;
[0018] According to the stochastic process, simulate the dynamic scenario of the arrival of new link user terminals, the departure of new link user terminals, and the geographical location change process with Poisson distribution, and obtain the communication link change situation of the user terminals in each time slot;
[0019] Among them, the change situation of the communication link of a user terminal in a time slot refers to that within the time period t, there are x pairs of newly added user terminals, y pairs of user terminals that leave due to the completion of communication tasks, and z pairs of user terminals with changed geographical locations.
[0020] Furthermore, the change situation of the communication links of each time slot user includes the time slot, the number of communication links of the user terminal, the number of links of the leaving user terminal, the number of links of the newly added user terminal, and the number of links of the user terminal with changed geographical location.
[0021] Furthermore, the steps for obtaining the total throughput of the downlink user terminal channel of the GEO multi-beam satellite system in a periodic time slot are as follows:
[0022] Calculate the signal gain according to the main axis deviation angle of the user terminal relative to the main axis of the GEO multi-beam satellite;
[0023] Calculate the signal attenuation according to the operating frequency of the GEO multi-beam satellite and the distance between the GEO multi-beam satellite and the ground;
[0024] Calculate the useful signal and the interference signal according to the coverage situation of the user terminal in each beam;
[0025] Based on the useful signal and the interference signal, and considering the signal gain and signal loss existing in the propagation process, calculate the signal-to-interference-plus-noise ratio;
[0026] According to the signal-to-interference-plus-noise ratio, use Shannon's theorem to calculate the total throughput of the downlink user terminal channel of the GEO multi-beam satellite system in this time slot in a dynamic scenario.
[0027] Furthermore, the total throughput S(t) total The expression is:
[0028]
[0029] In the formula, m is any link number, represents the calculated signal-to-interference-plus-noise ratio of the m-th link at the n-th channel resource at time t; represents the 0-1 variable that the channel resource numbered n at time t is allocated to the link numbered m; represents the received power information of the receiver of the link numbered m at time t; represents other gains from the transmitter of the link numbered m to the receiver of the link numbered m under the channel resource n at time t; is the communication status information in the case of narrowband beam coverage when the signal is transmitted from the transmitter belonging to the link numbered m to the receiver belonging to the link numbered m; σ 2 represents the thermal noise interference received during the transmission of information except for the interference between beams; is the signal belonging to the link number During the process of transmitting from the transmitter to the receiver belonging to link number m, the communication status information under the narrowband beam coverage; For other links different from the link number m, Indicates that the channel resource numbered n at time t is allocated to the link number 0-1 variable; Indicates the received power information of the receiver with link number at time t; Indicates the link number at time t Other gains from the transmitter of link number to the receiver of link number m under channel resource n; M t Indicates the total number of user links in communication at time t; N represents the total number of channel resources; B k Indicates the bandwidth allocated to each beam, related to the number of user transmitters N rk and the number of receivers N tk in the beam coverage area.
[0030] Furthermore, the expression of the optimization objective is:
[0031]
[0032]
[0033] In the formula, η(t) is the weighted optimization objective obtained by weighted optimization of the energy efficiency and spectral efficiency; f1 and f2 are normalization coefficients; ξ is the weighting coefficient, M max is the maximum number of user communication pairs that can be borne in the GEO multi-beam satellite system, and m is any link number; is the energy efficiency, that is, the ratio of the overall downlink receiver throughput to the power consumption; is the spectral efficiency, that is, the ratio of the overall downlink receiver throughput to the total system bandwidth; S(t) total is the total throughput; Indicates the power loss of all circuit transmissions of the satellite platform; M t Indicates the total number of user links in communication at time t; Indicates the received power information of the receiver of link number m at time t; B k Indicates the bandwidth allocated to each beam, related to the number of user transmitters N rk and the number of receivers N tk in the beam coverage area.
[0034] Furthermore, the constraint conditions of the optimization objective include:
[0035] The first constraint condition: It is represented that each link occupies exactly one channel allocation resource at time t;
[0036] The second constraint condition: where represents the power information allocated to link number m at time t, that is, the power of the uplink user transmitter; this constraint condition represents the value range of the transmission power of the transmitter of link m at time t;
[0037] The third constraint condition: represents the 0-1 channel allocation variable allocated to link m at time t.
[0038] Furthermore, using the D3QN reinforcement learning algorithm, the optimization objective is solved to obtain the optimal resource allocation scheme, including:
[0039] Construct a network model of the D3QN reinforcement learning algorithm, including a Q network and a target network, and create an experience pool, etc.;
[0040] Take the GEO multi-beam satellite system as the agent, define the 0-1 variable of the channel allocation of each link and the power magnitude of the transmitter of each link as the action; define the link number, the channel state of its own link, and the interference states of its own and other links as the state; define the reward as the weighted value of the energy efficiency and spectral efficiency of the satellite communication system within a certain time slot t;
[0041] Input the initial state information, execute the action, obtain the new state and reward information, until the termination iteration condition is reached, and obtain the optimal resource allocation scheme.
[0042] In a second aspect, the present invention also provides a multi-beam satellite resource allocation optimization system based on D3QN reinforcement learning, and this system includes:
[0043] A dynamic scenario reconstruction unit, which is used to reconstruct the dynamic scenario based on the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals, and obtain the communication link change situation of the user terminals in each time slot;
[0044] An optimization index construction unit, which is used to construct optimization indexes according to the communication link change situation of the user terminals in each time slot and the total throughput of the downlink user terminal channels of the GEO multi-beam satellite system within the periodic time slot; the optimization indexes include energy efficiency and spectral efficiency;
[0045] An optimization objective establishment unit, which is used to perform weighted summation on the energy efficiency and spectral efficiency, establish the optimization objective for the GEO multi-beam satellite system, and set constraint conditions for the optimization objective;
[0046] The D3QN reinforcement learning unit is used to solve the optimization objective by using the D3QN reinforcement learning algorithm to obtain the optimal resource allocation scheme.
[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0048] 1. The multi-beam satellite resource allocation optimization method and system based on D3QN reinforcement learning of the present invention. By establishing a GEO multi-beam satellite uplink and downlink user communication system model, the present invention establishes a total rate expression function for the dynamically changing scenario within the next cycle of this system, and sets two optimization indexes of spectral efficiency and energy efficiency. Based on the real-time number of user links, the energy and spectral efficiency are dynamically weighted, comprehensively considering the beam position, user position, and satellite position, so that the invention effect can not only be applicable to static conditions, but also to time-dynamically changing scenarios.
[0049] 2. The multi-beam satellite resource allocation optimization method and system based on D3QN reinforcement learning of the present invention. The optimization objective constructed by the present invention is that under the conditions of limited resources such as channels, power, and spectrum, the total user reception rate of the system needs to simultaneously consider limited channel resources, link spectrum limitations, and transmitter power limitations, so as to help the system meet user needs as much as possible in terms of spectrum and energy under limited resources. Description of the Drawings
[0050] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0051] Figure 1 It is a flowchart of the multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning of the present invention;
[0052] Figure 2 It is a schematic diagram of the longitude and latitude positions of the beams, user transmitters, and receivers in the two-dimensional plane of the communication scenario of the present invention;
[0053] Figure 3 It is a schematic diagram of calculating the deviation angle of the user from the satellite main axis of the present invention;
[0054] Figure 4 It is the antenna gain radiation pattern diagram of the present invention;
[0055] Figure 5 It is the D3QN network structure model (i.e., the D3QN reinforcement learning algorithm) of the present invention;
[0056] Figure 6 It is a diagram of the change in the number and position information of user terminals at time t1 - t2 of the present invention;
[0057] Figure 7This is the iterative process of the algorithm of the present invention and two comparison algorithms;
[0058] Figure 8 This is the throughput of different links within the cycle time of the present invention;
[0059] Figure 9 This is the comparison curve of the weekly weighting factor of the present invention;
[0060] Figure 10 This is the structural block diagram of the multi-beam satellite resource allocation optimization system based on D3QN reinforcement learning of the present invention. Detailed implementation manners
[0061] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments and drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention.
[0062] Embodiment 1
[0063] As Figure 1 shown, the multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning of the present invention includes:
[0064] S1. Based on the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals, reconstruct the dynamic scenario and obtain the communication link change situation of each time slot user terminal;
[0065] Specifically, the geographical location information of the ground main axis projection of the GEO multi-beam satellite refers to the longitude and latitude information of the projection of the main axis of the GEO multi-beam satellite on the ground; the geographical location information of the beams released by the GEO multi-beam satellite refers to the longitude and latitude information of the coverage range of multiple circular narrowband beams released by the GEO multi-beam satellite; the geographical location information of the uplink and downlink user terminals includes the geographical location information of the uplink user terminal and the geographical location information of the downlink user terminal; the geographical location information of the uplink user terminal refers to the longitude and latitude position information of the uplink user transmitting end; the geographical location information of the downlink user terminal refers to the longitude and latitude position information of the downlink user receiving end.
[0066] Specifically, step S1 includes:
[0067] Within the longitude and latitude area of the preset two-dimensional scenario, set the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals;
[0068] Based on the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals, establish a stochastic process of the communication links of the uplink and downlink user terminals;
[0069] According to the stochastic process, a dynamic scenario of the arrival of new link user terminals, the departure of new link user terminals, and the process of geographical location change is simulated by Poisson distribution, and the communication link change of user terminals in each time slot is obtained;
[0070] Among them, the communication link change of user terminals in one time slot refers to that within the time period t, there are x pairs of newly added user terminals, y pairs of user terminals that leave due to the completion of communication tasks, and z pairs of user terminals with geographical location changes.
[0071] In this embodiment, a geostationary orbit (GEO) multi-beam satellite is stationary relative to the ground at a height of H from the ground, and the longitude and latitude of the projection of its main axis on the ground are respectively
[0072] The GEO multi-beam satellite releases K circular narrowband beams. Assuming that the ground coverage radius of each narrowband beam is R, and the longitude and latitude of the projection on the ground are respectively There are m×m uplink and downlink communication users in the two-dimensional plane of the ground. Each pair of links contains an uplink user transmitter and a downlink user receiver, and the signals and noises received by the downlink user from other transmitters are all interference signals.
[0073] The uplink users and downlink users are numbered as (t1, t2,... t m ) and (r1, r2,... r m ), respectively. Then, the longitude and latitude position information of the uplink users can be obtained as The longitude and latitude position information of the downlink users is
[0074]
[0075] As Figure 2 shown, the circular range represents the beam coverage range, the "*" legend represents the uplink user transmitter, and the "." legend represents the downlink user receiver.
[0076] The uplink users in the same beam can transmit signals to the receivers of all downlink users within the range of this beam. In addition, due to the co-channel interference between beams in the multi-beam satellite system, part of the transmission information of the uplink users within the coverage range of the beams sharing the same frequency band will also be transmitted as interference information to the receivers of the downlink users within the coverage range of other beams of the same frequency band. The entire mission cycle of the multi-beam satellite communication system is discretized into multiple separate time intervals, and the duration of each time interval is t.
[0077] In the static communication scenario, the number of uplink and downlink users remains unchanged, so the number of communication links is also a fixed value. In the dynamic scenario, within each time slot, there are user terminals that exit due to the completion of communication tasks or join due to new communication requirements. The number of communication links changes continuously over time. For the convenience of analysis, the arrival and departure processes of uplink and downlink users in the network are modeled as two independent random processes. Each time slot, the communication network arrives at the coverage area of the satellite communication system according to a Poisson distribution process with parameter λ t Therefore, the probability that x pairs of new link users arrive at the system during the time period t is:
[0078]
[0079] Similarly, the departure process also follows a Poisson distribution with parameter μ t . Therefore, the probability that y pairs of new link users leave the system during the time period t is:
[0080]
[0081] In addition, the popularization of mobile communication technology makes the positions of uplink and downlink user terminals not fixed. Assume that during the time period t, there are z pairs of uplink and downlink user terminals whose geographical locations change. The change process still follows a Poisson distribution with parameter τ t :
[0082]
[0083] Therefore, during this time period t, there are x pairs of newly added user terminals, y pairs of user terminals that leave due to the completion of communication tasks, and z pairs of user terminals whose geographical locations change. And in each subsequent time period, the link state changes according to a random process.
[0084] S2. According to the communication link change situation of user terminals in each time slot and the total throughput of the downlink user terminal channels of the GEO multi-beam satellite system in the periodic time slot, an optimization index is constructed; the optimization index includes energy efficiency and spectral efficiency;
[0085] Specifically, the steps for obtaining the total throughput of the downlink user terminal channels of the GEO multi-beam satellite system in the periodic time slot in step S2 are as follows:
[0086] Calculate the signal gain according to the main axis deviation angle of the user terminal relative to the main axis of the GEO multi-beam satellite;
[0087] Calculate the signal attenuation according to the operating frequency of the GEO multi-beam satellite and the distance of the GEO multi-beam satellite from the ground;
[0088] Calculate the useful signal and interference signal according to the coverage situation of the user terminal in each beam;
[0089] Based on the useful signal and the interference signal, and considering the signal gain and signal loss existing in the propagation process, calculate the signal-to-interference-plus-noise ratio (SINR);
[0090] According to the SINR, use Shannon's theorem to calculate the total throughput of the downlink user terminal channel of the GEO multi-beam satellite system in the dynamic scenario within this time slot.
[0091] Specifically, the total throughput S(t) total has the following expression:
[0092]
[0093] where m is any link number, represents the calculated SINR of the m-th link on the n-th channel resource at time t; represents the 0-1 variable that the n-th channel resource numbered at time t is allocated to the link numbered m; represents the received power information of the receiver of link number m at time t; represents other gains from the transmitter of link number m to the receiver of link number m under the channel resource n at time t; is the communication status information in the case of narrowband beam coverage during the process of the signal being transmitted from the transmitter belonging to link number m to the receiver belonging to link number m; σ 2 represents the thermal noise interference other than the inter-beam transmission suffered during information transmission; is the communication status information in the case of narrowband beam coverage during the process of the signal being transmitted from the transmitter belonging to to the receiver belonging to link number m; is another link different from the link numbered m, represents the 0-1 variable that the n-th channel resource numbered at time t is allocated to the link numbered ; represents the received power information of the receiver of link number at time t; represents the other gains from the transmitter of link number to the receiver of link number m under the channel resource n; M t represents the total number of user links in communication at time t; N represents the total number of channel resources; B k represents the bandwidth allocated to each beam, which is related to the number of user transmitters N rk and the number of receivers N tk within the beam coverage area, B total is the total bandwidth resource.
[0094] Next, the quantities in formulas (4) and (5) are specifically explained as follows:
[0095] It represents the power information allocated to link number m at time t. This value mainly refers to the power level of the uplink user's transmitter. Since when the uplink user's signal is transmitted to the downlink user, it first goes through on-board system processing and then is sent back to the ground base station, there are a series of losses and gains in this process. Therefore, certain transformations need to be made to this power information.
[0096] For the loss process, without considering the effects of rain and cloud attenuation and the uplink channel, the signal is transmitted in the form of electromagnetic waves through the free space between the satellite and the user terminal and is received by the downlink user's receiver. Taking the satellite antenna as a point signal source, the signal is spherical in the propagation process. If the distance between the ground terminal and the satellite antenna is d (km), then the spherical area is S s = 4πd 2 , and the effective receiving area of the user terminal is S t = λ / 4π. If the energy propagation is uniform, the propagation loss of power is expressed as the ratio of the output power to the received power, which can be approximated as the ratio of the areas, i.e.:
[0097]
[0098] The free space loss L from the satellite antenna to the downlink user f is related to the satellite operating frequency and the satellite's distance from the earth. The mathematical expression is as follows:
[0099]
[0100] In the formula, c represents the speed of light, λ is the wavelength, and f is the satellite operating frequency (GHz).
[0101] In addition to losses during propagation, when the signal is transmitted to the satellite, the satellite antenna will also generate gains to the signal power received and transmitted by the satellite respectively. The gain effect is related to the deviation angle between the main axis of the user terminal and the satellite's main axis. The greater the deviation from the projection position of the antenna's main axis, the smaller the generated gain. As Figure 3 shown, according to the projection position of the main axis on the ground and the user's position, the formula for calculating the deviation angle of the main axis is as follows:
[0102]
[0103] In the formula, R earth is the radius of the earth.
[0104] As Figure 4 shown, the antenna gain can be obtained by substituting the calculated deviation angle of the main axis into the following expression:
[0105]
[0106] Refer to the ITU-R S.672-4 protocol, Φ b = 0.28°, a = 2.88, b = 6.32, L s = -25, G m = 41.6 dBi.
[0107] Assume that the power released by the uplink user at time t is Then, after the gain and loss in the transmission process, the power transmitted to the receiving end of the downlink user is
[0108] In addition, use to represent other gains from the transmitting end of link number m to the receiving end of link number m under channel resource n at time t.
[0109] It is the communication status information in the case of narrowband beam coverage during the process of the signal being transmitted from the transmitter belonging to link number m to the receiver belonging to link number m. Assume that the geographical locations of the transmitter and the receiver do not change within each time period, then the value of this status information is fixed within this time period. However, due to the real-time change of the user terminal status information such as joining and disconnecting in different time periods, the definition is as follows:
[0110] The values are divided into four types, namely:
[0111]
[0112] Among them When and only when the geographical locations of the transmitter of the uplink user and the receiver of the downlink user in the link are within the coverage range of the same narrowband beam, that is:
[0113]
[0114] For other values and corresponding conditions, as follows:
[0115]
[0116] The variable explanations in the formula are as follows:
[0117]
[0118] This variable represents the coverage of the uplink user transmitter t m within the narrowband beam k. The value of 1 indicates that the transmitter is within the coverage range of beam k, otherwise it is not within the coverage range;
[0119]
[0120] This variable represents the downlink user receiver r mCoverage within the narrowband beam k. A value of 1 indicates that the receiver is within the coverage of beam k; otherwise, it is not within the coverage.
[0121] Set σ 2 Represents other interference received during information transmission except for inter-beam transmission. Here, it is set as thermal noise, and the noise floor limit is -174 dBm / Hz.
[0122] Then, based on the above variables, the calculated signal-to-interference-plus-noise ratio of the m-th link on the n-th channel resource at time t can be obtained
[0123]
[0124] Specifically, the optimization metrics constructed in step S2 include energy efficiency and spectral efficiency, as follows:
[0125] The energy efficiency of the satellite communication system is the ratio of the overall downlink receiver throughput to the power consumption, and can be expressed at time t as:
[0126]
[0127] The spectral efficiency of the satellite communication system is the ratio of the overall downlink receiver throughput to the total system bandwidth, and can be expressed at time t as:
[0128]
[0129] Among them, is the energy efficiency, which is the ratio of the overall downlink receiver throughput to the power consumption; is the spectral efficiency, which is the ratio of the overall downlink receiver throughput to the total system bandwidth; S(t) total is the total throughput; represents the total transmission loss power of all circuits on the satellite platform; M t represents the total number of user links in communication at time t; represents the received power information of the receiver of link number m at time t; B k represents the bandwidth allocated to each beam, which is related to the number of user transmitters N rk within the coverage of the beam and the number of receivers N tk is related.
[0130] S3. Perform a weighted sum of the energy efficiency and spectral efficiency to establish an optimization objective for the GEO multi-beam satellite system, and set constraints for the optimization objective;
[0131] Since the number of uplink and downlink channels in this system is not fixed, and there are processes of arrival and departure of user transmit-receive terminals in each time slot, it is necessary to perform weighted addition of energy efficiency and spectral efficiency to give dynamic weights. The more users there are, the higher the demand for spectrum resources, so the weighting coefficient of spectral efficiency is larger. The fewer users there are, the higher the demand for energy resources, so the weighting coefficient of energy efficiency is larger. Assume that the maximum number of user communication pairs that can be supported in this system is M max , then the weighting coefficient is M max is the maximum number of user communication pairs that can be supported in the GEO multi-beam satellite system, and m is any link number;
[0132] Then the weighted optimization objective expression is:
[0133]
[0134] In the formula, η(t) is the weighted optimization objective obtained by weighted optimization of energy efficiency and spectral efficiency; f1 and f2 are normalization coefficients.
[0135] Since the on-board resources are very limited, the following assumptions are made: (1) Each link communicates through only one channel resource at each moment without occupying other channel resources. (2) The power released by the uplink user transmitter of each link is within a certain threshold range. Then the constraint conditions of this optimization problem are:
[0136] The first constraint condition: It means that each link occupies only one channel allocation resource at time t;
[0137] The second constraint condition: where represents the power information allocated to link number m at time t, that is, the power of the uplink user transmitter; this constraint condition represents the value range of the transmission power of the transmitter of link m at time t;
[0138] The third constraint condition: represents the 0-1 channel allocation variable allocated to link m at time t.
[0139] That is, this optimization objective can be generally expressed as the following expression:
[0140]
[0141] S4. Use the D3QN reinforcement learning algorithm to solve the optimization objective and obtain the optimal resource allocation scheme.
[0142] Specifically, step S4 specifically includes:
[0143] Construct the network model of the D3QN reinforcement learning algorithm, including the Q-network and the target network, and create an experience pool, etc.;
[0144] Regard the GEO multi-beam satellite system as an agent, and define the 0-1 variable of the channel allocation for each link and the power magnitude at the transmitting end of each link as actions; define the link number, the channel state of its own link, and the interference states of its own and other links as states; define the reward as the weighted value of the energy efficiency and spectral efficiency of the satellite communication system within a certain time slot t;
[0145] Input the initial state information, execute actions, obtain new state and reward information until the termination iteration condition is reached, and obtain the optimal resource allocation scheme.
[0146] In this embodiment, the D3QN (Dueling Double Deep Q-Network) reinforcement learning algorithm regards the GEO multi-beam satellite system as an agent, and the agent executes actions and obtains rewards by interacting with the environment; it is set that the GEO multi-beam satellite system has a total of N c channel resources, and the power range is evenly discretized into N P values. In the established model, the number of user communication links and geographical locations change dynamically at each moment, which affects the weighted result of the finally calculated spectral efficiency and energy efficiency. The deep reinforcement learning network can realize the dynamic adjustment of the network structure through interaction with the environment to optimize the objective function at each moment, and has strong real-time performance. In resource allocation, the channel and power resources allocated to each spot beam at the previous moment cannot completely and accurately determine the transfer of the next state, and the state space composed of the allocation situations of all beams is also very large, so it is suitable to use a parameterized action value function to select the optimal strategy. The D3QN (Dueling Double Deep Q-Network) algorithm combines the traditional DQN (Deep Q-Network) and the double deep Q-network (DDQN, Double Deep Q-Network) technologies, that is, it solves the defect that the DQN calculates the Q value through the greedy algorithm resulting in overestimation, and also improves the defect that the DDQN algorithm is difficult to converge in the reward adjustment problem. The DDQN and the adversarial deep Q-network are introduced on the basis of the DQN, and two neural networks calculate the action and state value functions respectively, so as to reduce the estimation error of the Q value. Therefore, the D3QN reinforcement learning algorithm is used to optimize the joint problem of spectral efficiency and energy efficiency.
[0147] The model adopted by the algorithm is as Figure 5As shown in the figure. The D3QN reinforcement learning algorithm includes two Q-networks and a target network with exactly the same structure. The parameters are copied from the Q-network to the target network within a fixed time through the learning of the Q-network. Specifically, it includes an input convolutional layer, a hidden layer, a state value function output layer, and an action advantage value function output layer. The input layer is used to receive the environmental state, adopting a two-layer (FC-ReLU) structure, receiving state information with a dimension of [(2 + N c + M) × M], which is used to capture the characteristics of the input information. The hidden layer adopts a fully connected structure with 256 nodes and is used to process the state information. On the one hand, the processed result is output through the advantage value function network to obtain the advantage value function A(s, a), and on the other hand, it is output through the state value function network to obtain the state value function V(s). The calculation formula for the advantage function is:
[0148] A(s, a; ω A ) = Q(s, a; ω) - V(s; ω V ) (23)
[0149] where ω V represents the parameters of the state value function network, and ω A represents the parameters of the advantage value function network. Let ω = (ω V , ω A ). This formula represents the advantage of action a relative to V(s; ω V ), which is related to the selection of actions. V(s; ω V ) is used to evaluate the quality of the current state and represents the expected return reward for future probabilistic actions:
[0150]
[0151] The advantage function represents the advantage of a specific action compared to other actions. To show the difference and eliminate the bias of the advantage value of each action, the average action advantage value generally needs to be deducted. Then, the formula for calculating the action Q value is:
[0152] Q(s, a; ω) = V(s; ω V ) + A(s, a; ω A ) - mean a A(s, a; ω A ) (25)
[0153] where mean represents the average value of all action advantage value functions. The greedy algorithm is used for the action selection strategy, that is, the following strategy is adopted when selecting actions:
[0154]
[0155] After that, execute action a to obtain the reward value and the new environmental state.
[0156] Define the experience replay buffer and the random experience replay pool: To be able to refer to the training results at past times when training the D3QN reinforcement learning network, an experience replay buffer is created to store the state information of each link, the action information taken, the reward value, and the set of link state information at the next moment. The mathematical expression of the experience replay buffer is as follows:
[0157] Δ={s t ,a t ,s t+1 ,r t} (27)
[0158] In the formula, s t represents the state information of all links at time t, s t+1 represents the state information of all links at time t + 1, a t represents the set of actions in the state s t , and r t represents the reward information obtained at the current time t.
[0159] After that, a random experience replay pool is created. When the storage capacity of the replay pool has not been reached, the agent first fills the replay pool by pre-storing randomly selected actions in a random state. Then, during training, new experience information replaces the old experience information over time. When the channel state changes slowly, using the experience replay mechanism can break the strong correlation between consecutive experiences, making each training more independent and the results more reliable.
[0160] Update network parameters: The D3QN network initially defines two networks, an evaluation network and a target network. During the learning process, only the network weights of the evaluation network are updated. After setting a certain number of updates, the parameters of the evaluation network are assigned to the target network, and then the next batch of updates is carried out.
[0161] The D3QN algorithm uses the evaluation network to obtain the action corresponding to the optimal action value in the s t+1 state, and then uses the target network to calculate the action value of this action, which can effectively avoid the overestimation problem. The update of the parameter ω is completed by calculating the gradient of the loss function through the gradient descent method. The calculation of the loss function L(ω) uses the root mean square error formula:
[0162] y t =r t+1 +γQ(s t+1 ,argmax Q(s t+1 ,a;ω e );ω) (28)
[0163] L(ω)=E[(y t -Q(s t ,at ; ω)) 2 (29)
[0164] where γ is the discount factor, and ω e and ω represent the parameters of the evaluation network and the target network respectively. According to the value of the loss function, the influence of each parameter on the loss function is calculated by the backpropagation algorithm to obtain the parameter gradient, and the gradient descent method is used to update each parameter to minimize the loss function as much as possible until convergence.
[0165] Regarding the GEO multi-beam satellite system as an agent, the agent interacts with the environment to perform actions and obtain rewards. The action is defined as the 0-1 variable of the channel allocation for each link and the power level at the transmitting end of each link. Then the size of the action space is N c ×N P . The state is defined as the link number + the channel state of its own link + the interference state of its own and other links. The link number is the link number m used by the uplink and downlink user terminals, m ∈ {1, 2,..., M t}}, the channel state of its own link refers to where n ∈ {1, 2,..., N c}}, and the interference state of its own and other links is the variable mentioned in step (2). The size of the state space is [(2 + N + M) × M]. The reward is defined as the weighted value of the energy efficiency and spectral efficiency of the GEO multi-beam satellite system within a certain time slot t: c When the D3QN algorithm is specifically executed, the initial state information is input, actions are performed, and new state and reward information are obtained until the termination iteration condition is reached to obtain the optimal resource allocation strategy.
[0166]
[0167] Specifically in implementation, the joint spectrum and energy efficiency optimization algorithm based on D3QN reinforcement learning proposed in the present invention is compared with the DQN algorithm and the DDQN algorithm. Table 1 gives the main parameters of the simulation:
[0168] Specifically in implementation, the joint spectrum and energy efficiency optimization algorithm based on D3QN reinforcement learning proposed in the present invention is compared with the DQN algorithm and the DDQN algorithm. Table 1 gives the main parameters of the simulation:
[0169] Table 1 Algorithm Parameter Settings
[0170]
[0171] Among the set of ten time intervals, the initial number of user communication links is 30 pairs. Starting from time t1, there are always user communication links that are about to leave due to the completion of communication tasks, user communication links that join this scenario due to new communication tasks, and user communication links whose positions change due to the use of mobile terminals. Table 2 shows the changes in the number of user links within the ten time slots. The changes in the communication links of each time slot for users include the time slot, the number of user terminal communication links, the number of departing user terminal links, the number of newly added user terminal links, and the number of user terminal links with geographical location changes. To facilitate observing the specific changes, a two-dimensional scenario diagram for the (t1 - t2) period is taken, as Figure 6 shown, and the dynamic changes in the scenario can better simulate the actual communication scenario.
[0172] Table 2 Changes in user communication links for each time slot
[0173]
[0174] The comparative simulation results of the present invention and other algorithms are as Figure 7 shown. All three methods perform smoothing processing on the original data to facilitate observing the overall convergence trend. The maximum number of iterations is limited to 25,000 times. The optimized algorithm based on D3QN interacts with the environment at different time slots and adjusts the allocated spectrum and power resources according to the current environmental state. Therefore, the joint spectrum and energy efficiency within the cycle time generally show an upward trend as the number of iterations increases. The DQN and DDQN algorithms have a similar overall trend because of the similar strategies they adopt. For the convergence results, the number of times the D3QN algorithm reaches convergence is approximately 10,000 times, and the number of times the DDQN algorithm and the DQN algorithm reach convergence are approximately 12,000 and 15,500 times respectively. Additionally, in terms of the total sum of spectrum efficiency and energy efficiency within the cycle time, the reward value of the D3QN algorithm within 10 time slots is approximately 9.73, and the reward values of the DDQN and DQN algorithms after reaching convergence are approximately 7.09 and 6.40 respectively. The proposed algorithm is 16.67% ahead of the DDQN algorithm and 35.5% ahead of the DQN algorithm in terms of the number of convergence times. In terms of the convergence reward value, it is 37.2% higher than the DDQN algorithm and 51.8% higher than the DQN algorithm. This proves the superiority of the present invention in joint spectrum efficiency optimization.
[0175] Figure 8 Shows the capacity results of 6 pairs of communication links after reaching convergence. Selecting [Link 1, Link 5, Link 9, Link 13, Link 17, Link 21] where the link numbers always exist as the analysis objects, the throughput optimization results of the D3QN algorithm for the links within the cycle time are higher than those of the DDQN and DQN algorithms, and it can also meet the user link communication requirements in terms of capacity demand compared to other algorithms, proving the efficient utilization of resources and thus improving user satisfaction.
[0176] To verify the beneficial effect on the results after dynamically weighting the spectral efficiency and energy efficiency, the optimization results with different fixed weighting coefficients and dynamic weighting coefficients are compared. As Figure 9 shown, the reward values of the system in each time slot are given under the condition that the weight factor values of η SE and η EE are (1 / 2, 1 / 2), (1 / 4, 3 / 4), and (3 / 4, 1 / 4) respectively, and under the weight of the dynamic weighting factor of the present invention. It can be analyzed that due to the change of the number of user pairs and the link state in each time slot, the reward value is constantly changing. The three methods using fixed weight factors achieve better results only in some time slots and special link number scenarios, and perform poorly in other time slots, lacking flexibility. The algorithm using the dynamic weighting factor can well adapt to different link number scenarios, and the weighted spectral efficiency in the cycle time is significantly better than the other three algorithms with fixed weight factors, which proves the excellent adaptability of the present invention in a dynamic environment.
[0177] Embodiment 2
[0178] As Figure 10 shown, the difference between this embodiment and Embodiment 1 is that this embodiment provides a multi-beam satellite resource allocation optimization system based on D3QN reinforcement learning, which corresponds one-to-one with the multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning in Embodiment 1; the system includes:
[0179] A dynamic scenario reconstruction unit, which is used to reconstruct the dynamic scenario based on the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals, and obtain the communication link change situation of the user terminals in each time slot;
[0180] An optimization index construction unit, which is used to construct optimization indexes according to the communication link change situation of the user terminals in each time slot and the total throughput of the downlink user terminal channels of the GEO multi-beam satellite system in the periodic time slots; the optimization indexes include energy efficiency and spectral efficiency;
[0181] An optimization goal establishment unit, which is used to perform weighted summation of the energy efficiency and spectral efficiency, establish the optimization goal for the GEO multi-beam satellite system, and set limiting conditions for the optimization goal;
[0182] A D3QN reinforcement learning unit, which is used to use the D3QN reinforcement learning algorithm to solve the optimization goal and obtain the optimal resource allocation scheme.
[0183] Among them, the execution process of each unit can be carried out according to the process steps of the multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning in Embodiment 1, and will not be elaborated one by one in this embodiment.
[0184] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0185] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0186] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0188] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning, characterized in that: The method includes: Based on the ground principal axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite and the geographical location information of the uplink and downlink user terminals, the dynamic scene is reconstructed and the communication link changes of the user terminals in each time slot are obtained; According to the change of the communication link of the user terminal in each time slot and the total throughput of the downlink user terminal channel of the GEO multi-beam satellite system in the periodic time slot, an optimization index is constructed; the optimization index includes energy efficiency and spectrum efficiency; Taking a weighted sum of the energy efficiency and the spectrum efficiency, establishing an optimization target for the GEO multi-beam satellite system, and setting a constraint condition for the optimization target; The D3QN reinforcement learning algorithm is used to solve the optimization objective and obtain the optimal resource allocation solution.
2. The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning according to claim 1 is characterized in that: The geographical location information of the ground main axis projection of the GEO multi-beam satellite refers to the latitude and longitude information of the projection of the main axis of the GEO multi-beam satellite on the ground; The geographical location information of the beam released by the GEO multi-beam satellite refers to the longitude and latitude information of the coverage area of the multiple circular narrowband beams released by the GEO multi-beam satellite; The geographical location information of the uplink and downlink user terminals includes geographical location information of the uplink user terminal and geographical location information of the downlink user terminal; the geographical location information of the uplink user terminal refers to the longitude and latitude location information of the uplink user originating end; The geographical location information of the downlink user terminal refers to the longitude and latitude location information of the downlink user terminal.
3. The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning according to claim 1 is characterized in that: Based on the ground principal axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite and the geographical location information of the uplink and downlink user terminals, the dynamic scene is reconstructed and the communication link changes of the user terminals in each time slot are obtained, including: In a preset two-dimensional scene, the ground principal axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite, and the geographical location information of the uplink and downlink user terminals are set; A random process for establishing communication links of uplink and downlink user terminals based on the projection of the ground principal axis of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite and the geographical location information of the uplink and downlink user terminals; According to the random process, a dynamic scenario of a new link user terminal arriving, a new link user terminal leaving, and a geographical location change process is simulated using Poisson distribution to obtain a communication link change of the user terminal in each time slot; The change of the communication link of the user terminal in a time slot refers to the existence of x pairs of newly added user terminals, y pairs of user terminals that leave due to completion of communication tasks, and z pairs of user terminals whose geographical locations change within the time period t.
4. The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning according to claim 3 is characterized in that: The communication link changes of the users in each time slot include the time slot, the number of user terminal communication links, the number of links of leaving user terminals, the number of links of newly added user terminals and the number of links of user terminals with changed geographical locations.
5. The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning according to claim 1 is characterized in that: The steps for obtaining the total throughput of the downlink user terminal channel of the GEO multi-beam satellite system in a periodic time slot are: Calculate the signal gain according to the main axis deviation angle of the user terminal relative to the main axis of the GEO multi-beam satellite; Calculate signal attenuation based on the operating frequency of the GEO multi-beam satellite and the distance between the GEO multi-beam satellite and the ground; Calculate the useful signal and interference signal according to the coverage of the user terminal in each beam; Calculate the signal-to-interference-to-noise ratio based on the useful signal and the interference signal, and taking into account the signal gain and signal loss in the propagation process; According to the signal to interference and noise ratio, the Shannon theorem is used to calculate the total throughput of the downlink user terminal channel of the GEO multi-beam satellite system in the time slot in the dynamic scenario.
6. The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning according to claim 5 is characterized in that: The total throughput S(t) total The expression is: Where m is the number of any link. represents the calculated signal-to-interference-and-noise ratio of the mth link on the nth channel resource at time t; Indicates that the channel resource numbered n at time t is allocated to the 0-1 variable with link number m; Indicates the power information received by the receiver of link number m at time t; It represents the other gain from the transmitting end of link number m to the receiving end of link number m under channel resource n at time t; is the communication status information of the signal transmitted from the transmitter of link number m to the receiver of link number m under the condition of narrowband beam coverage; 2 It represents the thermal noise interference except the one transmitted between beams when transmitting information; The signal is composed of link number Communication status information in the process of transmitting from the transmitter of link number m to the receiver of link number m under the condition of narrowband beam coverage; For other links with different numbers from the link numbered m, Indicates that the channel resource numbered n at time t is allocated to the link numbered 0-1 variables; Indicates the link number at time t The power information received by the receiving end; Indicates the link number at time t Other gains from the transmitting end to the receiving end of link number m under channel resource n; M t represents the total number of user links communicating at time t; N represents the total number of channel resources; B k It represents the bandwidth allocated to each beam and the number of user terminals N within the coverage area of the beam. rk and the number of receiving terminals N tk related.
7. The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning according to claim 1 is characterized in that: The expression of the optimization objective is: Wherein, η(t) is the weighted optimization target obtained by weighted optimization of the energy efficiency and spectrum efficiency; f1 and f2 are normalization coefficients; ξ is the weighting coefficient, M max is the maximum number of user communication pairs that can be supported in the GEO multi-beam satellite system, and m is the number of any link; is the energy efficiency, i.e., the ratio of the entire downlink receiving end throughput to the power consumption; is the spectrum efficiency, which is the ratio of the entire downlink receiving end throughput to the total system bandwidth; S(t) total is the total throughput; Represents the transmission loss power of all circuits on the satellite platform; M t represents the total number of user links communicating at time t; Indicates the power information received by the receiver of link number m at time t; B k It represents the bandwidth allocated to each beam and the number of user terminals N within the coverage area of the beam. rk and the number of receiving terminals N tk related.
8. The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning according to claim 1 is characterized in that: The constraints of the optimization objective include: First restriction: It means that at time t, each link occupies only one channel allocation resource; Second restriction: in Indicates the power information allocated to link number m at time t, that is, the power of the uplink user sending end; the constraint condition is expressed as the transmission power value range of the sending end of link m at time t; The third restriction: represents the 0-1 channel allocation variable assigned to link m at time t.
9. The multi-beam satellite resource allocation optimization method based on D3QN reinforcement learning according to claim 1 is characterized in that: The D3QN reinforcement learning algorithm is used to solve the optimization objective and obtain the optimal resource allocation solution, including: Build the network model of the D3QN reinforcement learning algorithm, including the Q network and the target network, and create an experience pool; The GEO multi-beam satellite system is taken as an intelligent agent, and the channel allocation 0-1 variable of each link and the power size of each link are defined as actions; the link number, the channel state of the link itself, and the interference state of itself and other links are defined as states; the reward is defined as the weighted value of the energy efficiency and spectrum efficiency of the satellite communication system in a certain time slot t; Input the initial state information, execute the action, obtain the new state and reward information, until the termination condition is reached and the optimal resource allocation plan is obtained.
10. A multi-beam satellite resource allocation optimization system based on D3QN reinforcement learning, characterized in that: The system includes: A dynamic scene reconstruction unit is used to reconstruct the dynamic scene and obtain the communication link change of the user terminal in each time slot based on the ground main axis projection of the GEO multi-beam satellite, the beams released by the GEO multi-beam satellite and the geographical location information of the uplink and downlink user terminals; An optimization index construction unit is used to construct an optimization index according to the communication link change of the user terminal in each time slot and the total throughput of the downlink user terminal channel of the GEO multi-beam satellite system in the periodic time slot; the optimization index includes energy efficiency and spectrum efficiency; An optimization target establishing unit, configured to perform a weighted summation of the energy efficiency and the spectrum efficiency to establish an optimization target for the GEO multi-beam satellite system, and set a constraint condition for the optimization target; The D3QN reinforcement learning unit is used to solve the optimization target by using the D3QN reinforcement learning algorithm to obtain an optimal resource allocation solution.