PDT-based integrated ground and satellite network dynamic resource allocation method
By applying a dynamic resource allocation method based on PDT in integrated ground and satellite networks, the problems of user access control and power distribution in dynamic environments are solved, and the network energy efficiency and response capabilities are improved.
Patent Information
- Application Number
- CN202510488149.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively control user access and power distribution in dynamic environments, resulting in difficulty in taking into account network throughput and energy efficiency, high computational complexity and insufficient adaptability.
The integrated ground and satellite network dynamic resource allocation method based on Prompt DecisionTransformer (PDT) is adopted to achieve real-time resource allocation optimization by establishing system models, constructing Markov decision-making processes, collecting historical interactive data, designing trajectory prompt generation algorithms, and combining offline pre-training and online fine-tuning strategies.
It improves the overall energy efficiency, real-time response capability and resource scheduling accuracy, and significantly improves the adaptability and stability of the network in a dynamic environment.
Smart Images

Figure CN120018147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of communication networks, and in particular to a dynamic resource allocation method for integrated ground and satellite networks based on Prompt Decision Transformer (PDT). Background Art
[0002] With the growing demand for high-quality network access and energy efficiency, integrated ground and satellite networks have gradually become an emerging communication architecture. By combining the wide-area coverage advantages of satellites with the high-capacity characteristics of ground systems, this architecture provides strong support for multiple scenarios such as emergency communications, remote area networking, and the global Internet of Things. However, there are still many challenges in actual deployment, such as how to effectively control access and allocate power to users in a dynamic environment to balance key indicators such as network throughput and energy efficiency.
[0003] There are some existing technical solutions that attempt to optimize resource allocation from different angles. For example, some solutions use enhanced caching technology and approximate methods to handle sub-problems such as user association, bandwidth allocation, and power control separately; there are also solutions that use water-filling algorithms and geometric programming to achieve joint scheduling of cache, computing, and communication resources; there are also studies that explore the collaborative design of beamforming, carrier selection, and power allocation for cognitive scenarios; in addition, some work focuses on intelligent scheduling and satellite-ground collaborative optimization. However, in practice, the above methods usually have problems such as high computational complexity, insufficient adaptability, and insufficient generalization ability when the network status changes drastically, making it difficult to meet the diverse and dynamic communication needs of the future.
[0004] To this end, people have gradually introduced intelligent algorithms such as deep reinforcement learning into the resource allocation problem, hoping to reduce dependence on prior knowledge through self-learning and achieve better scheduling strategies in complex environments. However, conventional reinforcement learning often has shortcomings such as long training time, slow convergence speed, and difficulty in quickly adapting to environmental changes when facing a high-dimensional environment with multiple users, multiple satellites, and multiple base stations. Summary of the invention
[0005] The purpose of the present invention is to provide a PDT-based integrated ground and satellite network dynamic resource allocation method, by establishing a system model, constructing a resource allocation problem of joint user association and power control, and converting it into a Markov decision process, collecting multi-scenario historical interaction data to construct a trajectory data set, designing a trajectory prompt generation algorithm to achieve effective decoupling of state and action representation, and performing offline pre-training based on PDT and combining it with an online fine-tuning strategy, and finally using the trained model to achieve real-time resource allocation optimization, thereby solving the problems of high computational complexity, insufficient adaptability and unstable resource allocation in a dynamic environment in the prior art, and improving the overall network energy efficiency, real-time response capability and accuracy of resource scheduling.
[0006] To achieve the above object, the present invention provides a PDT-based integrated terrestrial and satellite network dynamic resource allocation method, the steps comprising:
[0007] S1. Establish a system model integrating ground and satellite networks;
[0008] S2, construct the resource allocation problem of joint user association and power control, and transform it into a Markov decision process model;
[0009] S3, collect and preprocess multi-scenario historical interaction data to construct trajectory dataset;
[0010] S4, training the PDT model based on the trajectory dataset;
[0011] S5. Use the trained PDT model to generate user association and power control decisions in real time to achieve real-time resource allocation optimization.
[0012] Preferably: the system model established in step S1 includes a mutually exclusive connection relationship between users, base stations and satellites, that is, each user can only establish a connection with a single access point in any time slot, thereby ensuring the uniqueness and stability of the connection during resource allocation, and the single access point is a base station or a satellite.
[0013] Preferably: the Markov decision process model in step S2 includes:
[0014] Define the system state space of the integrated ground and satellite network. The state space consists of the set of energy efficiency states at each moment. The formula is:
[0015] ;
[0016] ;
[0017] In the formula, Indicates that user u is at time energy efficiency state, N represents the user set, u represents the user index, Indicates time The energy efficiency state set, Indicates that user u is at time Energy efficiency, Indicates that user u is at time energy efficiency;
[0018] The action space of the system integrating terrestrial and satellite networks is defined. The action space includes user association and power control actions at each moment. The formula is:
[0019] ;
[0020] = ;
[0021] = ;
[0022] In the formula, Indicates time User association and power control actions, Indicates user-associated actions. Indicates power control action;
[0023] ;
[0024] ;
[0025] Where L represents the number of low-orbit satellites in the system integrating ground and satellite networks, M represents the number of base stations in the system integrating ground and satellite networks, Indicates time Satellite-user association indicator variable, Indicates time Base station-user association indicator variable, Indicates time The downlink transmission power allocated by satellite s to user u is: Indicates time The downlink transmission power allocated by base station b to user u is: Indicates that user u is at time The associated action, Indicates that user u is at time Power control action;
[0026] Defining the reward function For user u at time The energy efficiency of the integrated ground and satellite network system is maximized through user association and power control decisions. The formula is:
[0027] ;
[0028] In the formula, Indicates at time The total reward received by the system.
[0029] Preferably, in step S3, collecting and preprocessing multi-scenario historical interaction data to construct a trajectory data set includes:
[0030] S31. Collect historical interaction data between satellites, base stations and users in multiple network scenarios. The collected historical interaction data should include state information, associated actions, power allocation actions and corresponding reward values at each moment.
[0031] S32, cleaning the collected historical interaction data, removing abnormal and redundant data, and arranging the historical interaction data into a trajectory sequence including states, actions, and rewards in chronological order;
[0032] S33, normalizing or standardizing the rewards in the trajectory sequence, and using a bootstrap method to statistically estimate the reward distribution of each trajectory sequence, and extracting the reward distribution characteristics of each trajectory sequence;
[0033] S34, integrating the trajectory sequences and reward distribution characteristics obtained in steps S31 to S33 into a trajectory dataset as input data for PDT.
[0034] Preferably: step S3 also includes:
[0035] Design a trajectory prompt generation algorithm to achieve effective decoupling of state and action representation, and construct an enhanced trajectory dataset containing trajectory prompts.
[0036] Preferably, the steps of constructing an enhanced trajectory data set including trajectory prompts include:
[0037] S31, performing cumulative reward statistics for each trajectory sequence, and analyzing the difference in reward changes at different times;
[0038] S32, using a bootstrap method to estimate the mean and standard deviation of the cumulative rewards of each trajectory sequence, calculating the deviation value of the cumulative rewards of each trajectory sequence relative to the reward distribution characteristics, and obtaining a trajectory prompt value between 0 and 1 through normalization processing;
[0039] S33, separately encoding the obtained trajectory prompt value and the state and action information through independent encoders, so as to achieve effective decoupling of the trajectory prompt and the state-action representation, and construct an enhanced trajectory dataset containing the trajectory prompt.
[0040] Preferably, the method further comprises:
[0041] S6. Based on real-time generation of user association and power control decisions, the trained PDT model is fine-tuned online according to real-time network feedback to achieve continuous optimization of resource allocation strategy.
[0042] Preferably: The calculation formula is:
[0043] .
[0044] In the formula, Indicates at time , the signal-to-noise ratio when satellite s sends a signal to user u, Indicates at time , the signal-to-noise ratio when base station b sends a signal to user u.
[0045] Preferably: The calculation formula is:
[0046] .
[0047] In the formula, Indicates time The complex channel coefficients of the downlink from satellite s to user u are: represents the background noise power, Indicates at time The interference introduced by all base stations in the downlink to user u is Indicates at time Interference caused by satellite s sending signals to users other than user u, or by satellites other than satellite s sending signals to users other than user u.
[0048] Preferably: The calculation formula is:
[0049] .
[0050] In the formula, ) indicates time The complex channel coefficients of the downlink from base station b to user u are: represents the background noise power, Indicates at time The interference introduced by all satellites downlink to user u, Indicates at time Interference caused by base station b sending signals to users other than user u, or by base stations other than base station b sending signals to users other than user u.
[0051] Therefore, the present invention adopts the above-mentioned PDT-based integrated terrestrial and satellite network dynamic resource allocation method, which has the following beneficial effects:
[0052] (1) By jointly optimizing user association and power control, the overall network energy efficiency can be greatly improved, system energy consumption can be reduced, and communication quality and user experience can be improved;
[0053] (2) A hybrid learning strategy combining offline pre-training and online fine-tuning is adopted to enable the model to quickly respond to changes in network status and adjust resource allocation strategies in real time to meet changing communication needs;
[0054] (3) By designing a trajectory prompt generation algorithm, the state and action are effectively decoupled, the decision-making process is simplified, the online computing burden is reduced, and the stability and scalability of the system in large-scale complex networks are improved.
[0055] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flow chart of a method according to an embodiment of the present invention;
[0057] Figure 2 A comparison chart of energy efficiency optimization experimental results of the present invention and other algorithms in different scenarios in an embodiment of the present invention;
[0058] Figure 3 This is a comparison chart of the optimized convergence speed of the present invention and the baseline algorithm in different scenarios in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0060] Example
[0061] Refer to Figure 1 The present invention provides a PDT-based integrated ground and satellite network dynamic resource allocation method, the steps comprising:
[0062] S1. Establish a system model integrating ground and satellite networks, and clarify the connection relationship between users, base stations and satellites. The connection relationship includes the connection between users and satellites, the connection between users and base stations, each user can only establish a connection with one access point (satellite or base station) in the same time slot, channel gain and allocated power and other parameters.
[0063] S2. Construct the resource allocation problem of joint user association and power control and transform it into a Markov decision process model.
[0064] Specifically, the Markov decision process model of the integrated ground and satellite network system is defined as follows:
[0065] Define the state space of the integrated ground and satellite network system. The state space consists of the energy efficiency state of each time slot t, and its formula is:
[0066] ;
[0067] ;
[0068] In the formula, N represents the user set, u represents the user index, Indicates at time The state set, Indicates that user u is at time energy efficiency;
[0069] The action space of the integrated ground and satellite network system is defined. The action space includes the user association and power control decision at each time slot t, and its formula is:
[0070] ;
[0071] = ;
[0072] = ;
[0073] In the formula, Indicates user-associated actions. Indicates power control action;
[0074] ;
[0075] ;
[0076] Where L represents the number of low-orbit satellites in the integrated ground-satellite network, M represents the number of base stations in the integrated ground-satellite network, represents the satellite-user association indicator variable, which is used to identify whether user u has established a connection with satellite s. represents the base station-user association indicator variable, which is used to identify whether user u has established a connection with base station b. represents the downlink transmission power allocated by satellite s to user u, represents the downlink transmission power allocated by base station b to user u, Indicates that user u is The association decision at the time is used to constrain user u at the time Only a single connection is allowed (value 0 or 1), Indicates that user u is The power control decision at each moment is used to aggregate the transmit power received by the user from the satellite or base station to which it belongs;
[0077] Define rewards, reward function Define user u in The energy efficiency of the integrated ground and satellite network system is maximized through user association and power control decisions. The formula is:
[0078] ;
[0079] In the formula, Indicated in The total reward obtained by the system at that moment is the sum of the energy efficiency of all users, which is used to measure the energy efficiency level of the entire system at that moment.
[0080] S3. Collect and preprocess multi-scenario historical interaction data to construct a trajectory dataset. The obtained preprocessed data and its reward distribution characteristics can be integrated to construct a trajectory dataset.
[0081] S4. Design a trajectory prompt generation algorithm to achieve effective decoupling of state and action representation. This includes cumulative reward statistics for preprocessed historical trajectory data, analyzing reward changes, generating corresponding trajectory prompts based on reward deviations, and encoding the prompts separately from state and action information.
[0082] S5. Use the offline constructed trajectory dataset to perform offline pre-training based on the PDT framework, and combine the online fine-tuning strategy based on real-time network feedback to quickly adapt to the dynamic network environment.
[0083] S6. Use the trained model to achieve real-time resource allocation optimization and improve the overall network energy efficiency.
[0084] The specific steps include:
[0085] S61, generating user association and power control decisions in real time by the trained PDT model according to the current network status information;
[0086] S62, using the decision result to calculate the energy efficiency of each user, the calculation formula is:
[0087] ;
[0088] In the formula, Indicates at time , the signal-to-noise ratio when satellite s sends a signal to user u, Indicates at time , the signal-to-noise ratio when base station b sends a signal to user u.
[0089] S63: signal-to-noise ratio when satellite s sends a signal to user u in step S62 The calculation formula is:
[0090] ;
[0091] In the formula, Indicates time The complex channel coefficients of the downlink from satellite s to user u are: represents the background noise power, Indicates at time Interference introduced by other base stations sending downlink signals, Indicates at time Interference caused by the same satellite sending signals to other users, or other satellites sending signals to any user.
[0092] S64: signal-to-noise ratio when base station b sends a signal to user u in step S62 The calculation formula is:
[0093] ;
[0094] In the formula, ) indicates time The complex channel coefficients of the downlink from base station b to user u are: Indicates at time Interference introduced by other satellites transmitting downlink signals, Indicates at time Interference caused by the same satellite sending signals to other users, or other satellites sending signals to any user.
[0095] S65. The energy efficiency of each user is accumulated to obtain the overall system reward, and the model parameters are further fine-tuned online based on real-time feedback to achieve continuous optimization of the resource allocation strategy.
[0096] In order to verify the resource allocation effect of the method of the present invention in the integrated ground and satellite network, the method of the present invention was compared with other models in an experiment, and the specific contents are as follows.
[0097] Model setup
[0098] The Decision Transformer (DT) method, Proximal Policy Optimization (PPO) algorithm, and Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm are selected as baseline models. The DT method transforms the reinforcement learning problem into a sequence modeling task, and uses the Transformer architecture to perform autoregressive modeling on the historical interaction trajectory, thereby generating future decision sequences and realizing offline learning and prediction of resource allocation strategies. The PPO algorithm is a reinforcement learning method based on policy gradients. It ensures the stability and efficient convergence of training by limiting the amplitude of each policy update (using the clipping objective function), and is suitable for online policy optimization in dynamic environments. The TD3 algorithm is an improvement on DDPG. It reduces the overestimation of action value by introducing a dual Q network, and uses a delayed update mechanism to enhance training stability. It is suitable for dealing with resource scheduling problems in high-dimensional continuous action spaces.
[0099] Evaluation Metrics
[0100] (1) Energy efficiency: measured in bits / joule, it reflects the ratio between the amount of data transmitted by the system and its energy consumption, and directly evaluates the energy-saving effect of the resource allocation strategy.
[0101] (2) Convergence speed: This is the number of training rounds or time required for the algorithm to achieve stable performance, measured by the changes in energy efficiency or rewards during training.
[0102] Dataset Description
[0103] Several typical scenario datasets of simulated integrated ground and satellite networks are used. These datasets include user status information, action records, and reward data. The specific parameters are: the network consists of 4 low-orbit satellites, 4 base stations, and 16 users, and each satellite and each base station serves a maximum of 2 users. The satellite link operates in the S band with a transmit power upper limit of 30 dBm, while the base station link follows the Rayleigh fading model with a maximum transmit power of 23 dBm. To train the PDT-based resource allocation model, the Adam optimizer and ReLU activation function are used, and the training parameters are set to batch size 10, learning rate 10^-4, and discount factor 0.95. The entire training process contains 2000 rounds, and each round is managed by the central controller for 100 time slots. The offline dataset is collected by using the PPO algorithm, and some online interaction samples are selected to fine-tune the PDT model online so that the model can quickly adapt to the dynamic network environment in new scenarios.
[0104] Results Analysis
[0105] Figure 2 and Figure 3 The experimental results in different network scenarios are presented, and the test is conducted for the cases where the number of users is 40 (scenario 1) and 32 (scenario 2). The experimental results show that in terms of energy efficiency, the PDT method achieves 1330 and 999 bits / joule in the two scenarios, respectively, which is significantly better than the DT method (1210 and 940 bits / joule), the PPO method (1050 and 850 bits / joule), and the TD3 method (850 and 640 bits / joule). In addition, in terms of convergence speed, the PDT method is also better than DT, and significantly better than PPO and TD3. In summary, the use of a PDT-based resource allocation strategy can not only significantly improve the network energy efficiency, but also converge faster, thereby improving the real-time response capability and overall performance of the system in a dynamic environment.
[0106] Therefore, the present invention adopts the above-mentioned PDT-based integrated ground and satellite network dynamic resource allocation method, and uses a hybrid learning strategy that combines offline pre-training and online fine-tuning to generate optimal resource allocation decisions in real time using the trained PDT model, thereby being able to respond quickly in a dynamic network environment, significantly improving network energy efficiency and communication performance, while reducing computational complexity and enhancing system robustness.
[0107] Based on the same technical solution, the present invention also provides an electronic device, including:
[0108] Memory for storing computer programs;
[0109] The processor is used to implement the steps of the above-mentioned PDT-based integrated terrestrial and satellite network dynamic resource allocation method when executing the computer program.
[0110] Based on the same technical solution, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned PDT-based integrated ground and satellite network dynamic resource allocation method are implemented. The computer-readable storage medium may include: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0111] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A PDT-based integrated terrestrial and satellite network dynamic resource allocation method, characterized in that the steps include: S1. Establish a system model integrating ground and satellite networks; S2, construct the resource allocation problem of joint user association and power control, and transform it into a Markov decision process model; S3, collect and preprocess multi-scenario historical interaction data to construct trajectory dataset; S4, training the PDT model based on the trajectory dataset; S5. Use the trained PDT model to generate user association and power control decisions in real time to achieve real-time resource allocation optimization.
2. The method for dynamic resource allocation of integrated terrestrial and satellite networks based on PDT according to claim 1, characterized in that: The system model established in step S1 includes the mutually exclusive connection relationship between users, base stations and satellites, that is, each user can only establish a connection with a single access point in any time slot, thereby ensuring the uniqueness and stability of the connection during the resource allocation process. The single access point is a base station or a satellite.
3. The PDT-based integrated terrestrial and satellite network dynamic resource allocation method according to claim 2, characterized in that: The Markov decision process model in step S2 includes: Define the system state space of the integrated ground and satellite network. The state space consists of the set of energy efficiency states at each moment. The formula is: ; ; In the formula, Indicates that user u is at time energy efficiency state, N represents the user set, u represents the user index, Indicates time The energy efficiency state set, Indicates that user u is at time Energy efficiency, Indicates that user u is at time energy efficiency; The action space of the system integrating terrestrial and satellite networks is defined. The action space includes user association and power control actions at each moment. The formula is: ; = ; = ; In the formula, Indicates time User association and power control actions, Indicates user-associated actions. Indicates power control action; ; ; Where L represents the number of low-orbit satellites in the system integrating ground and satellite networks, M represents the number of base stations in the system integrating ground and satellite networks, Indicates time Satellite-user association indicator variable, Indicates time Base station-user association indicator variable, Indicates time The downlink transmission power allocated by satellite s to user u is: Indicates time The downlink transmission power allocated by base station b to user u is: Indicates that user u is at time The associated action, Indicates that user u is at time Power control action; Defining the reward function For user u at time The energy efficiency of the integrated ground and satellite network system is maximized through user association and power control decisions. The formula is: ; In the formula, Indicates at time The total reward received by the system.
4. The PDT-based integrated terrestrial and satellite network dynamic resource allocation method according to claim 3, characterized in that: In step S3, the multi-scenario historical interaction data is collected and preprocessed to construct a trajectory dataset, including: S31. Collect historical interaction data between satellites, base stations and users in multiple network scenarios. The collected historical interaction data should include state information, associated actions, power allocation actions and corresponding reward values at each moment. S32, cleaning the collected historical interaction data, removing abnormal and redundant data, and arranging the historical interaction data into a trajectory sequence including states, actions, and rewards in chronological order; S33, normalizing or standardizing the rewards in the trajectory sequence, and using a bootstrap method to statistically estimate the reward distribution of each trajectory sequence, and extracting the reward distribution characteristics of each trajectory sequence; S34, integrating the trajectory sequences and reward distribution characteristics obtained in steps S31 to S33 into a trajectory dataset as input data for PDT.
5. The method for dynamic resource allocation of integrated terrestrial and satellite networks based on PDT according to claim 4, characterized in that: Step S3 also includes: Design a trajectory prompt generation algorithm to achieve effective decoupling of state and action representation, and construct an enhanced trajectory dataset containing trajectory prompts.
6. The method for dynamic resource allocation of integrated terrestrial and satellite networks based on PDT according to claim 5, characterized in that: The steps to construct an enhanced trajectory dataset containing trajectory hints include: S31, performing cumulative reward statistics for each trajectory sequence, and analyzing the difference in reward changes at different times; S32, using a bootstrap method to estimate the mean and standard deviation of the cumulative rewards of each trajectory sequence, calculating the deviation value of the cumulative rewards of each trajectory sequence relative to the reward distribution characteristics, and obtaining a trajectory prompt value between 0 and 1 through normalization processing; S33, separately encoding the obtained trajectory prompt value and the state and action information through independent encoders, so as to achieve effective decoupling of the trajectory prompt and the state-action representation, and construct an enhanced trajectory dataset containing the trajectory prompt.
7. The PDT-based integrated terrestrial and satellite network dynamic resource allocation method according to claim 1 or 5, characterized in that: The method further comprises: S6. Based on real-time generation of user association and power control decisions, the trained PDT model is fine-tuned online according to real-time network feedback to achieve continuous optimization of resource allocation strategy.
8. The PDT-based integrated terrestrial and satellite network dynamic resource allocation method according to claim 3, characterized in that: The calculation formula is: , In the formula, Indicates at time , the signal-to-noise ratio when satellite s sends a signal to user u, Indicates at time , the signal-to-noise ratio when base station b sends a signal to user u.
9. The PDT-based integrated terrestrial and satellite network dynamic resource allocation method according to claim 8, characterized in that: The calculation formula is: , In the formula, Indicates time The complex channel coefficients of the downlink from satellite s to user u are: represents the background noise power, Indicates at time The interference introduced by all base stations in the downlink to user u is Indicates at time Interference caused by satellite s sending signals to users other than user u, or by satellites other than satellite s sending signals to users other than user u.
10. The PDT-based integrated terrestrial and satellite network dynamic resource allocation method according to claim 8, characterized in that: The calculation formula is: , In the formula, ) indicates time The complex channel coefficients of the downlink from base station b to user u are: represents the background noise power, Indicates at time The interference introduced by all satellites downlink to user u, Indicates at time Interference caused by base station b sending signals to users other than user u, or by base stations other than base station b sending signals to users other than user u.
Citation Information
Patent Citations
Energy efficiency optimization method and device of integrated ground satellite network
CN112543049A
Distributed task unloading method and device based on reinforcement learning and storage medium
CN116467005A
Base station energy consumption optimization method based on improved Decision Transform model
CN118945684A