A resource allocation method for near-zero power beam-hopping satellite communication system based on NOMA
Through the intelligent decision-making mechanism of multi-dimensional resource joint optimization and deep reinforcement learning, the time-sharing multiplexing mechanism and deep Q network processing discrete actions, the beam interference and energy consumption problems in beam hopping satellite systems are solved, efficient resource allocation and dynamic scheduling are achieved, and system adaptability and computing efficiency are improved.
Patent Information
- Application Number
- CN202510819871.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing NOMA-based beam hopping satellite system does not fully consider the interference between beams and the large user energy consumption and overhead, which makes it difficult for resource allocation to meet real-time requirements and multi-objective optimization, especially when large-scale user access is too large.
By building a multi-dimensional resource joint optimization model and an intelligent decision-making mechanism for deep reinforcement learning, the time-sharing multiplexing mechanism divides the beam jump time slot into data transmission and energy collection stages, uses the radio frequency energy generated by delay-sensitive users to power the delay-tolerance user, and combines the deep Q network and the deep deterministic strategy gradient algorithm for beam scheduling and power allocation to realize coordinated deployment of satellites and ground.
It realizes efficient and dynamic resource allocation of high-throughput beam hopping satellite systems, improves resource utilization and energy utilization efficiency, solves the problem of insufficient adaptability of traditional optimization algorithms in dynamic environments, and reduces the on-satellite computing burden.
Smart Images

Figure CN120320835B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of beam-hopping satellite communication systems, and in particular to a resource allocation method for a near-zero power consumption beam-hopping satellite communication system based on NOMA. Background Art
[0002] Next-generation high-throughput beam-hopping satellite systems utilize multiple spot beams and frequency reuse to build a space-based backbone network for global mobile communications services. However, their fixed power and spectrum allocation methods struggle to adapt to the growing number of users and heterogeneous service demands, resulting in low resource utilization. Furthermore, inter-beam interference caused by large-scale user access exacerbates energy consumption, further hindering efficient resource allocation.
[0003] Beam-hopping technology uses time-division multiplexing to dynamically hop between multiple beam positions, leveraging onboard phased arrays to allocate resources in four dimensions: time, space, power, and bandwidth. Non-Orthogonal Multiple Access (NOMA) simultaneously serves multiple users on the same time-frequency resources and allocates them non-orthogonally in the power domain, improving spectrum efficiency while more flexibly meeting differentiated communication needs. Existing solutions use heuristic algorithms to transform non-convex planning into a convex optimization problem, optimizing power and beam illumination patterns on a beam-by-beam-slot basis, improving the flexibility of traditional fixed allocation.
[0004] Existing NOMA-based beam-hopping satellite systems do not fully consider the interference between beams and the high energy consumption of users caused by interference. In the NOMA beam-hopping system, resource allocation needs to consider multiple dimensions such as the time scheduling of beam hopping, NOMA user grouping and power allocation, and frequency resource management, resulting in an exponential growth in the solution space of the optimization problem. Especially when the system scale expands, such as when a low-orbit beam-hopping satellite constellation needs to serve hundreds of beams and thousands of users simultaneously, traditional optimization algorithms often cannot meet real-time requirements due to excessive computational complexity. Actual beam-hopping satellite systems need to simultaneously optimize multiple objectives such as spectrum efficiency, energy efficiency, user fairness, and service latency. Traditional methods use weighted summation to transform multiple objectives into a single objective, but the selection of weights lacks theoretical guidance, and fixed weights are difficult to adapt to dynamic environments.
[0005] Therefore, a resource allocation method for near-zero power beam-hopping satellite communication system based on NOMA is needed. Summary of the Invention
[0006] In view of this, the present invention provides a resource allocation method for a near-zero-power beam-hopping satellite communication system based on NOMA. By constructing a multi-dimensional resource joint optimization model and an intelligent decision-making mechanism based on deep reinforcement learning, efficient matching of onboard resources and user request traffic of a high-throughput beam-hopping satellite system is achieved.
[0007] To this end, the present invention provides the following technical solutions:
[0008] A resource allocation method for a near-zero power consumption beam-hopping satellite communication system based on NOMA, comprising:
[0009] Users within the same beam are divided into delay-sensitive users and delay-tolerant users. The beam-hopping time slot is divided into a data transmission phase and an energy collection phase through a time-division multiplexing mechanism. The RF energy generated by delay-sensitive users is used to power delay-tolerant users, thus building a near-zero-power beam-hopping satellite communication system.
[0010] A joint optimization model for data transmission rate and power allocation is established with the goal of minimizing the difference between the total transmit power of beam-hopping satellites and user request traffic and minimizing the probability of service interruption for delay-sensitive users.
[0011] A deep reinforcement learning framework is constructed based on a joint optimization model of data transmission rate and power allocation combined with a multi-dimensional optimization space;
[0012] The deep reinforcement learning network framework is deployed collaboratively between satellite and ground, and historical status and business demand data are used at the gateway to perform offline training and verification of a joint optimization model for data transmission rate and power allocation. The trained model parameters are uploaded to the beam-hopping satellite. The beam-hopping agent receives the model parameters and dynamically adjusts the beam hopping pattern, power allocation coefficient, and time slot partitioning parameters according to real-time business needs, thereby achieving adaptive matching and scheduling of on-board resources and beam traffic.
[0013] Furthermore, a digital model of a beam-hopping satellite receiving signal is constructed to characterize the total transmit power and user request traffic of the beam-hopping satellite;
[0014] The digital model:
[0015]
[0016] in, , represents the beam hopping time slot, Represents beam hopping time slot Inner beam b down-beam satellite receiving signal, Indicates the The access status of each DTU, Indicates the The access status of each DSU; M represents the number of delay-sensitive users under beam b, N represents the number of delay-tolerant users under beam b, and S represents the number of illuminated beams in the snapshot; express The transmission power, express The transmission power; for The sending signal, for Sending signal; For service Beam The channel gain of For service Beam communication frequency band.
[0017] Furthermore, the data transmission rate and power allocation joint optimization model further includes:
[0018] The constraint condition is that the beam satellite should give priority to the service transmission needs of delay-sensitive users while meeting the power requirements of delay-tolerant users.
[0019] Furthermore, the deep reinforcement learning framework includes:
[0020] Modeling beam-hopping satellites as intelligent agents responsible for learning and decision-making;
[0021] The data packets to be sent to each beam cell within the coverage area received by the beam-hopping satellite are modeled as the environment;
[0022] The problem of which beam cell each working beam is scheduled to in different beam hopping time slots is modeled as an interaction process between the agent and the environment;
[0023] The number of packets to be transmitted with different residual delay tolerances stored in each beam queue in the beam hopping time slot is used as the state;
[0024] The action is to select the beam cell covered by the working beam in the current beam hopping time slot;
[0025] The reward is the difference between the number of packets processed in the current beam hopping time slot and the number of packets that time out.
[0026] Furthermore, the Q network structure of the deep reinforcement learning framework includes: a 5-layer convolutional neural network;
[0027] Each layer of the convolutional neural network includes: 2 convolutional layers, 1 Flatten layer, and 2 fully connected layers; and uses the ReLU activation function.
[0028] Furthermore, the satellite-ground collaborative deployment of the deep reinforcement learning network framework also includes:
[0029] Through the measurement and control link, the differential update mechanism is used to upload the Q network parameters after training convergence to the onboard processing unit.
[0030] Furthermore, the dynamic adjustment of the beam hopping pattern, power allocation coefficient, and time slot division parameters according to real-time service requirements includes:
[0031] Receive a request from a beam cell in each beam hopping time slot and update the status;
[0032] Normalize the updated state;
[0033] The agent determines the action through the policy network and calculates the Q value corresponding to the action through the Q network;
[0034] use Strategy The probability of selecting the action with the largest Q value as the action of the current beam hopping time slot;
[0035] The analytical action obtains the beam hopping pattern, power allocation coefficient and time slot division parameters.
[0036] Furthermore, the Q network of the deep reinforcement learning framework is trained by a deep deterministic policy gradient algorithm, including:
[0037] Adopt Adam method to minimize the loss function;
[0038] The loss function:
[0039]
[0040] in, Indicates the current network output, Represents the parameters of the main network, Represents the label value, Indicates status, Indicates action.
[0041] Furthermore, the data transmission rate and power allocation joint optimization model is:
[0042]
[0043]
[0044] in, is the normalization constant, is the weighted value, Beam hopping time slot The power allocated to the DSU in inner beam b; The power provided to the energy storage device, is the actual power requirement of the DTU in the beam hopping time slot, represents the total transmit power of the beam-hopping satellite in the current beam-hopping time slot T; represents the requested traffic of all DTUs in the snapshot of beam-hopping time slot T; represents the requested traffic of all DSUs in the snapshot of beam-hopping time slot T; represents the normalization constant, which is set according to the system scale; Represents a user Beam hopping time slot in beam b The business interruption cost under Represents a user The interruption indication function of the beam-hopping satellite is ; The energy used by DTU for data transmission; Beam hopping time slot the remaining energy in the energy storage device at the start; , Represents beam hopping time slot The inner beam b is activated, Represents beam hopping time slot Inner beam b is not activated.
[0045] Advantages and positive effects of the present invention:
[0046] This method achieves efficient dynamic allocation of resources in a high-throughput beam-hopping satellite system through multi-dimensional resource joint optimization and an intelligent decision-making mechanism based on deep reinforcement learning.
[0047] A joint optimization model based on NOMA is constructed: Using a time-division multiplexing mechanism, beam-hopping timeslots are divided into two phases: data transmission and energy harvesting. The RF energy generated by delay-sensitive users is used to power delay-tolerant users, ensuring service continuity while improving energy efficiency. NOMA technology's power domain reuse enables multiple users to be served within the same beam, effectively improving resource utilization under uneven traffic distribution. Traditional beam-hopping systems struggle to balance delay-sensitive services with energy-constrained users.
[0048] Intelligent scheduling based on PDQN: Combining the advantages of DQN for processing discrete actions and DDPG for processing continuous actions, a hierarchical decision network is used to achieve joint optimization of beam scheduling and power allocation, improving the system's adaptability in dynamic environments and solving the problem of insufficient adaptability of traditional optimization algorithms in complex beam-hopping satellite communication environments.
[0049] This method uses lightweight deployment of satellite-ground collaboration: offline model training is completed through ground gateway stations, and lightweight inference is performed on the satellite platform. This effectively reduces the onboard computing burden while ensuring real-time decision-making. It resolves the contradiction between the limited onboard computing power and the real-time requirements of the business. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0051] Figure 1 Flowchart of a resource allocation method for a near-zero power consumption beam-hopping satellite communication system based on NOMA in an embodiment of the present invention;
[0052] Figure 2 This is a structural diagram of a near-zero power consumption beam-hopping satellite communication system according to an embodiment of the present invention;
[0053] Figure 3 Schematic diagram of zero-power network energy harvesting for sideline communication in an embodiment of the present invention;
[0054] Figure 4 This is a schematic diagram of the deep reinforcement learning framework structure in an embodiment of the present invention;
[0055] Figure 5 This is a diagram of the PDQN network structure in an embodiment of the present invention;
[0056] Figure 6 A flowchart for online generation of beam scheduling and traffic allocation strategies in an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0058] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0059] This paper provides a resource allocation method for a near-zero-power beam-hopping satellite communication system based on NOMA. This method uses NOMA technology to achieve power domain reuse, serving multiple users within the same beam. Using a time-division multiplexing mechanism, beam-hopping time slots are divided into two phases: data transmission and energy harvesting. The radio frequency energy generated by delay-sensitive users is used to power delay-tolerant users, thereby building a self-sustaining communication network. A multidimensional optimization space is established, encompassing beam-hopping patterns, power allocation coefficients, and beam-hopping time slot partitioning parameters. Combining the advantages of the Deep Q-Network (DQN) for handling discrete actions and the Deep Deterministic Policy Gradient (DDPG) algorithm for handling continuous actions, an intelligent scheduling algorithm based on PDQN is designed. This algorithm dynamically matches onboard resources with beam traffic through a hierarchical decision network. Furthermore, a lightweight deployment scheme, combining satellite and ground collaboration, and an offline training-online inference architecture effectively reduces onboard computational load while ensuring real-time decision-making.
[0060] Example
[0061] Combine Figure 1 As shown, the method steps include:
[0062] In this embodiment, the beam-hopping satellite and multiple ground users form a near-zero power consumption beam-hopping satellite communication system, the structure of which is as follows: Figure 2 shown.
[0063] The beam-hopping satellite coverage area is geographically partitioned to create multiple beam cells. Each beam can illuminate only one beam cell during each beam-hopping time slot. Several beam cells are selected during each beam-hopping time slot as a beam snapshot for illumination. After a beam-hopping cycle, all beam cells are guaranteed to be covered at least once. Each beam cell is equipped with an energy storage device to store RF energy generated by delay-sensitive users collected from the environment by delay-tolerant users within the beam cell.
[0064] In order to improve data transmission performance, beam-hopping satellites can generate multiple beams and transmit data to different areas on the ground at the same time. The number of beams of a beam-hopping satellite is defined as , the total transmit power of the beam-hopping satellite is , the total system bandwidth is .
[0065] Each beam cell simultaneously serves N delay-sensitive users (DSUs) and M delay-tolerant users (DTUs). Since DSUs have a higher priority, the satellite prioritizes DSUs and provides services to DTUs while ensuring DSU services.
[0066] Step 1: Using a time-division multiplexing mechanism, the beam-hopping timeslot is divided into two phases: data transmission and energy harvesting. The RF energy generated by delay-sensitive users is used to power delay-tolerant users, thereby building a self-sustaining communication network as a near-zero-power communication network. Based on this near-zero-power communication network, a power model for the beam cell's own communications is constructed to quantify the energy harvesting and allocation process in the beam-hopping satellite communication system.
[0067] The delay-tolerant user in each beam collects the RF energy generated by delay-sensitive users in the environment and stores it in an energy storage device within each beam cell. It then uses a zero-power system based on sideline communication for its own communication, including:
[0068] Step 1-1: Set each beam hopping slot Divided into and There are two stages. Indicates the time-sharing parameter. During the time, DTU uses the remaining energy of the energy storage device in the previous beam hopping time slot to transmit data; In the RF energy storage device, the DTU collects RF energy from the DSU to charge the energy storage device. In a beam hopping time slot, it first provides energy for the DTU's data transmission, and then charges the storage device to ensure that the DTU's service needs can be met in the next beam hopping time slot.
[0069] Figure 3 This is a block diagram of the zero-power network energy harvesting solution based on side communication in this embodiment: and The transmitted signal is respectively subjected to the channel gain of and The channel reaches After that, the power divider allocates energy collection and data transmission phases according to the signal instructions; the collected energy Used for communication with beam-hopping satellites during the data transmission phase. and represent the noise in the signal transmission and energy harvesting links, respectively.
[0070] Step 1-2: Beam hopping time slot The signal received by the satellite in the inner beam b down-beam is modeled as follows:
[0071]
[0072] in, , Indicates the beam b down-beam satellite receiving signal, Indicates the The access status of each DTU, Indicates the The access status of each DSU; express The transmission power, express The transmission power; for The sending signal, for Sending signal; For service Beam The channel gain of For service Beam communication frequency band.
[0073] Step 1-3: Based on beam hopping time slots, DTU collects RF energy from DSU without causing additional energy consumption of DSU. The model of the beam-hopping satellite receiving signal under the inner beam b is used to determine the In beam hopping time slot The energy collected in is expressed as:
[0074]
[0075] in, For beam b In beam hopping time slot The energy collected in is the energy harvesting efficiency; for and The channel gain between .
[0076] Step 1-4: Determine beam hopping time slot At the end, the remaining energy of the energy storage device in beam b is calculated as:
[0077]
[0078] in, Beam hopping time slot the remaining energy in the energy storage device at the start; Beam hopping time slot The remaining energy in the energy storage device at the end; The energy used by DTU for data transmission; Indicates the energy storage upper limit of the energy storage device.
[0079] Step 1-5: Use the remaining energy of the energy storage device to transmit data during the beam hopping time slot In inner beam b, the energy in the energy storage device is used to provide energy for DTU data transmission. The power provided by the energy storage device to the DTU is:
[0080]
[0081] in, It indicates that the energy storage device in beam b provides power for DTU data transmission.
[0082] Step 2: With the goal of minimizing the difference between the total transmit power of the beam-hopping satellite and the user's requested traffic, a joint optimization model for data transmission rate and power allocation is constructed;
[0083] After the beam-hopping satellite collects the needs of users in the service area, a joint optimization model of data transmission rate and power allocation is constructed with the goal of minimizing the difference between the total transmission power of the beam-hopping satellite and the user request traffic, taking into account the service interruption probability of delay-sensitive users and the minimum power guarantee of delay-tolerant users.
[0084] Step 2-1: Determine the beam hopping time slot Inner beam Next service user The signal-to-noise ratio is calculated as:
[0085]
[0086] in, is the beam in the beam hopping time slot T Service users signal-to-noise ratio; To serve users Beam With users The channel gain between For users The power distribution coefficient, Assigning beams to beam-hopping satellites The power, To serve users Intra-beam interference with other users in the same beam, To serve users Inter-beam interference with other active beams, is Gaussian white noise; the calculation formula for intra-beam interference and inter-beam interference is:
[0087]
[0088]
[0089] in, represents the intra-beam interference, represents inter-beam interference; Indicates the beam hopping time slot adjacent beams in the snapshot Is it activated? Indicates that it is activated. Indicates not activated.
[0090] Step 2-2: Determine the beam hopping time slot In the snapshot, the beam The data transmission rate of the served users is calculated as follows:
[0091]
[0092] in, Indicates service users Data transmission rate; Beam bandwidth; , Represents beam hopping time slot The inner beam b is activated, Represents beam hopping time slot Inner beam b is not activated and there is , represents the set of beams to be activated in the snapshot within the beam-hopping time slot T, and S is the number of irradiated beams in the snapshot.
[0093] Step 2-3: Determine beam hopping time slot The flow rate provided by the inner-beam hopping satellite to each service user under beam b is calculated as follows:
[0094]
[0095] in, Represents beam hopping time slot Inner beam hopping satellite provides beam Service users in The total flow rate.
[0096] Step 2-4: Determine beam hopping time slot The total traffic provided by the inner beam-hopping satellite to all service users under beam b and the total transmission power of the beam-hopping satellite are:
[0097]
[0098]
[0099] in, represents the total traffic provided by the beam-hopping satellite to beam b; Indicates the total transmit power of the beam-hopping satellite.
[0100] Step 2-5: Determine beam hopping time slot The request flow of internal DTU and DSU is calculated as follows:
[0101]
[0102]
[0103] Among them, the total traffic demand of all DTUs under beam b in beam hopping time slot T is , the total traffic demand of all DSUs under beam b in beam hopping time slot T is ; represents the requested traffic of all DTUs in the snapshot of beam-hopping time slot T; represents the requested traffic of all DSUs in the snapshot of beam-hopping time slot T.
[0104] Step 2-6: The optimization objectives are to match the traffic provided by beam b in beam hopping time slot T with the user requested traffic in the cell and to minimize the probability of DSU user service interruption. The formula is:
[0105]
[0106]
[0107] in, represents the normalization constant, which is set according to the system scale; Represents a user Beam hopping time slot in beam b The business interruption cost under Represents a user The interrupt indication function.
[0108]
[0109] Where, represents the interrupt threshold of the signal-to-noise ratio, hour Indicates the beam-hopping satellite and user Normal communication, otherwise it indicates communication interruption.
[0110] Step 2-7: Determine constraints and beam hopping time slots In the constrained satellite, the power requirement of DTU transmission is met while giving priority to the DSU service transmission requirement. The constraint formula is expressed as:
[0111]
[0112] in, Beam hopping time slot The power allocated to the DSU in inner beam b; The power provided to the energy storage device, is the actual power requirement of the DTU in the beam hopping time slot, It represents the total transmit power of the beam-hopping satellite in the current beam-hopping time slot T.
[0113] The power allocated to the DSU under beam b in the beam hopping time slot T is calculated as follows:
[0114]
[0115] in, Beam hopping time slot Inner beam hopping satellites are assigned to beams power.
[0116] Step 2-8: With the goal of minimizing the difference between the total transmit power of the beam-hopping satellite and the user's requested traffic, combined with the user's power demand constraint, a joint optimization model for data transmission rate and power allocation is constructed:
[0117]
[0118] Translates to:
[0119]
[0120]
[0121] in, is the normalization constant, is the weighted value.
[0122] The beam-hopping satellite system is modeled as a discrete-time event system driven by new service arrival events, and the hopping problem in the beam-hopping satellite system is transformed into a sequential decision problem.
[0123] Combining the advantages of deep Q-networks in processing discrete action spaces and deep deterministic policy gradient algorithms in processing continuous action spaces, a beam scheduling and power allocation framework based on a parameterized deep Q-network (PDQN) is constructed. The beam-hopping satellite is modeled as an intelligent agent responsible for learning and decision-making. The data packets received by the beam-hopping satellite and sent to each beam cell within its coverage area are modeled as the environment. Since the essence of beam scheduling in a beam-hopping satellite system is to allocate the hopping beam hopping time slot resources, the question of which beam cell each working beam is scheduled to in different beam hopping time slots is modeled as an interaction process between the intelligent agent and the environment. Through continuous training and learning of the PDQN, long-term maximum benefits are achieved.
[0124] Taking the joint optimization model of data transmission rate and power allocation as the objective function, a deep reinforcement learning framework is constructed to realize offline training of Q network and online generation of beam scheduling and power allocation strategies.
[0125] The relevant parameters for model training are designed. The data transmission rate and power allocation joint optimization model is then trained and verified offline through the gateway's powerful computing and processing capabilities. Finally, the trained model parameters are uploaded to the beam-hopping satellite via the measurement and control link. Online inference and decision-making operations are performed onboard, enabling real-time scheduling of onboard resources and beam traffic based on different service QoS requirements.
[0126] Step 3: Take the data transmission rate and power allocation joint optimization model as the objective function and build a deep reinforcement learning framework. The structure is as follows: Figure 4 As shown;
[0127] Step 3-1: Determine the status;
[0128] Different types of service data packets have different residual delay tolerances. is the maximum delay tolerance among all service types. After each hopping beam slot, the remaining delay tolerance decreases by 1. Therefore, The state of a beam-hopping slot is defined as:
[0129]
[0130] Among them, the matrix express The number of data packets to be transmitted with different residual delay tolerances stored in each beam queue in the beam hopping time slot.
[0131] Step 3-2: Determine the action;
[0132] Set the current beam hopping slot The beam cell covered by the illumination beam acts as an action:
[0133]
[0134] in, Indicates output action; Indicates whether beam b is activated within the beam hopping time slot T.
[0135] Step 3-3: Determine the reward;
[0136] The optimization goal of the objective function is to minimize the difference between the total transmit power of the beam-hopping satellite and the user's requested traffic. Therefore, the difference between the number of packets processed in the current beam-hopping time slot and the number of packets that timed out is used as the reward:
[0137]
[0138] in, represents the reward function; express The total number of packets processed by the system after the beam hopping timeslot selection action, The total number of packets that have timed out in the current beam hopping slot.
[0139] Step 3-4: Determine the network structure;
[0140] like Figure 5 As shown, in this embodiment, the Q network serves as the core decision-making module and adopts a 5-layer convolutional neural network structure, including 2 convolutional layers, 1 Flatten layer and 2 fully connected layers, and uses the ReLU activation function. This design efficiently extracts the deep features of the state matrix through the convolutional layer, and then achieves accurate mapping of the state to the action value through the fully connected layer. The policy network serves as an auxiliary module, and its output is only used to guide the exploration direction of the Q network or provide initial action suggestions, and does not directly participate in the final decision. When the two work together, the gradient update of the policy network is based on the optimization target of the Q network to ensure the consistency of its auxiliary function with the main network.
[0141] Step 3-5: Introduce the experience pool to store the quadruple generated during the interaction between the agent and the environment , in order to break the correlation between data during the neural network training process, so that the neural network can learn independent and identically distributed data.
[0142] Step 3-6: Introduce the target network to update the main network. The target network has the same structure as the main network, but different parameters. The parameters of the target network are a slow copy of the main network parameters, that is, the target network parameters are copied and updated from the main network every G steps to reduce the correlation between the target Q value and the current Q value.
[0143] Step 3-7: Determine the loss function:
[0144]
[0145] in, Indicates the current network output, represents the target network output, Represents the parameters of the main network, Represents the parameters of the target network; For the tag value:
[0146]
[0147] in, γ is the discount factor for future rewards.
[0148] Step 4: Train PDQN using a deep deterministic policy gradient algorithm.
[0149] Step 4-1: Initialize the beam-hopping satellite system scenario parameters.
[0150] Step 4-2: Randomly initialize the online Q network and target network , clear the experience pool.
[0151] Step 4-3: Set the number of training cycles and the number of beam hopping time slots per cycle, and start the training process.
[0152] Step 4-4: Initialize the state of the Q network in this cycle ,action and rewards .
[0153] In each beam hopping slot In the process, the agent determines the action through the policy network and calculates the Q value corresponding to the action through the Q network. Strategy( -greedy strategy) with The probability of selecting the action with the largest Q value.
[0154]
[0155] in, represents the greed factor, is the power allocation coefficient; selecting pattern coefficients for beam hopping; It is a time-sharing parameter used to divide a beam-hopping time slot and determine the ratio of data transmission time to energy collection time.
[0156] Step 4-5: Agent performs actions , the status changes from becomes , calculate the reward value of the action :
[0157]
[0158] in, , Satisfaction with resource allocation, ; is the traffic gap for each beam, = - ; is a constant used for standardization .
[0159] Step 4-6: Quadruple Store it in the experience pool. If the experience pool is full, discard the earliest quadruple.
[0160] Steps 4-7: Select a batch size of experience samples from the experience pool and calculate the loss function and label value through steps 3-7.
[0161] Step 4-8: Use Adam method to update network parameters, target network parameter Every G steps an update is made from the Q network.
[0162] Step 5: Satellite-ground collaborative deployment of deep reinforcement learning network framework, the flowchart is as follows Figure 6 As shown; execute the training process in step 4 until the Q network converges; then execute the online decision phase to obtain the beam scheduling and power allocation strategy. The specific steps are as follows:
[0163] Step 5-1: The measurement and control link will train the converged Q network parameters and target network parameters Uploaded from the gateway to the onboard intelligent agent. Parameter upload uses a differential update mechanism, transmitting only the differences from the previous version to reduce uplink resource usage.
[0164] Step 5-2: The onboard intelligent agent makes an online decision based on the current state of each beam queue through the trained network and generates the current beam hopping time slot action;
[0165] 1) The onboard agent loads the received model parameters and initializes the environment state, agent actions, and experience pool.
[0166] 2) Establish initial state synchronization with each beam cell and obtain the current state of each beam queue.
[0167] 3) Initialize the state matrix and record the number of data packets with different residual delay tolerances for each beam.
[0168] 4) Start to make online decisions in each beam hopping time slot In the process, accept the request of the beam cell and update the state matrix , normalize the state data.
[0169] 5) Input the normalized state into the Q network to obtain the Q value estimate of the possible action .use The strategy selects the action for the current beam hopping slot .
[0170] Step 5-3: Parsing Actions from The snapshot of the decoded T-hop beam slot is Beam cells to be activated:
[0171]
[0172] in, represents the set of beams to be activated, Represents beam Is it activated, and . According to the power allocation coefficient Calculate the transmit power of beam b:
[0173]
[0174] in, Indicates the power distribution coefficient of beam b. set up The time slot structure of time slot beam b.
[0175] Step 5-4: Go to step 4) to start the next iterative decision.
[0176] By outputting resource allocation for a near-zero-power beam-hopping satellite communication system through a deep learning framework, the traffic provided by the beam-hopping satellite is matched with the traffic requested by the user, and the probability of DSU user service interruption is minimized.
[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A resource allocation method for a near-zero power beam-hopping satellite communication system based on NOMA, characterized in that: include: Users within the same beam are divided into delay-sensitive users and delay-tolerant users. The beam-hopping time slot is divided into a data transmission phase and an energy collection phase through a time-division multiplexing mechanism. The RF energy generated by delay-sensitive users is used to power delay-tolerant users, thus building a near-zero-power beam-hopping satellite communication system. With the goal of minimizing the difference between the total transmit power of beam-hopping satellites and user request traffic and minimizing the probability of service interruption for delay-sensitive users, a joint optimization model for data transmission rate and power allocation is established, including: in, is the normalization constant, is the weighted value, Beam hopping time slot The power allocated to the DSU in inner beam b; The power provided to the energy storage device, is the actual power requirement of the DTU in the beam hopping time slot, represents the total transmit power of the beam-hopping satellite in the current beam-hopping time slot T; represents the requested traffic of all DTUs in the snapshot of beam-hopping time slot T; represents the requested traffic of all DSUs in the snapshot of beam-hopping time slot T; represents the normalization constant, which is set according to the system scale; Represents a user Beam hopping time slot in beam b The business interruption cost under Represents a user The interruption indication function of the beam-hopping satellite is ; The energy used by DTU for data transmission; Beam hopping time slot the remaining energy in the energy storage device at the start; , Represents beam hopping time slot The inner beam b is activated, Represents beam hopping time slot Inner beam b is not activated; Based on the joint optimization model of data transmission rate and power allocation combined with multi-dimensional optimization space, a deep reinforcement learning framework is constructed, including: Modeling beam-hopping satellites as intelligent agents responsible for learning and decision-making; The data packets to be sent to each beam cell within the coverage area received by the beam-hopping satellite are modeled as the environment; The problem of which beam cell each working beam is scheduled to in different beam hopping time slots is modeled as an interaction process between the agent and the environment; The number of packets to be transmitted with different residual delay tolerances stored in each beam queue in the beam hopping time slot is used as the state; The action is to select the beam cell covered by the working beam in the current beam hopping time slot; The reward is the difference between the number of packets processed in the current beam hopping time slot and the number of packets that have timed out. The deep reinforcement learning network framework is deployed collaboratively between satellite and ground, and historical status and business demand data are used at the gateway to perform offline training and verification of a joint optimization model for data transmission rate and power allocation. The trained model parameters are uploaded to the beam-hopping satellite. The beam-hopping agent receives the model parameters and dynamically adjusts the beam hopping pattern, power allocation coefficient, and time slot partitioning parameters according to real-time business needs, thereby achieving adaptive matching and scheduling of on-board resources and beam traffic.
2. The method according to claim 1, characterized in that Constructing a digital model of a beam-hopping satellite receiving signal to characterize the total transmit power of the beam-hopping satellite and the user request flow; The digital model: in, , represents the beam hopping time slot, Represents beam hopping time slot Inner beam b down-beam satellite receiving signal, Indicates the The access status of each DTU, Indicates the The access status of each DSU; M represents the number of delay-sensitive users under beam b, N represents the number of delay-tolerant users under beam b, and S represents the number of illuminated beams in the snapshot; express The transmission power, express The transmission power; for The sending signal, for Sending signal; For service Beam The channel gain of For service Beam communication frequency band.
3. The method according to claim 1, characterized in that The data transmission rate and power allocation joint optimization model further includes: The constraint condition is that the beam satellite should give priority to the service transmission needs of delay-sensitive users while meeting the power requirements of delay-tolerant users.
4. The method according to claim 1, wherein The Q network structure of the deep reinforcement learning framework includes: a 5-layer convolutional neural network; Each layer of the convolutional neural network includes: 2 convolutional layers, 1 Flatten layer, and 2 fully connected layers; and uses the ReLU activation function.
5. The method according to claim 1, characterized in that The satellite-ground collaborative deployment of the deep reinforcement learning network framework also includes: Through the measurement and control link, the differential update mechanism is used to upload the Q network parameters after training convergence to the onboard processing unit.
6. The method according to claim 1, characterized in that The dynamic adjustment of the beam hopping pattern, power allocation coefficient, and time slot division parameters according to real-time service requirements includes: Receive a request from a beam cell in each beam hopping time slot and update the status; Normalize the updated state; The agent determines the action through the policy network and calculates the Q value corresponding to the action through the Q network; use Strategy The probability of selecting the action with the largest Q value as the action of the current beam hopping time slot; The analytical action obtains the beam hopping pattern, power allocation coefficient and time slot division parameters.
7. The method according to claim 1, characterized in that The Q network of the deep reinforcement learning framework is trained by a deep deterministic policy gradient algorithm, comprising: Adopt Adam method to minimize the loss function; The loss function: in, Indicates the current network output, Represents the parameters of the main network, Represents the label value, Indicates status, Indicates action.
Citation Information
Patent Citations
Beam hopping resource allocation method and system based on deep reinforcement learning, storage medium and equipment
CN113572517A
NOMA-based beam hopping satellite communication system design method
CN115882924A