Low-power-consumption Internet of Things direct-connected satellite uplink scheduling method based on deep reinforcement learning

Through the TDMA channel allocation and SAPPO algorithm based on deep reinforcement learning, the problems of low transmission efficiency and poor fairness in uplink scheduling of low-power IoT direct-connected satellites are solved, and efficient and reliable IoT device communication is achieved.

CN120263259APending Publication Date: 2025-07-04NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510198963.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing low-power IoT direct-connected satellite uplink scheduling technology cannot reasonably allocate channel resources, resulting in low transmission efficiency, poor fairness, weak priority processing capabilities, and unable to meet the needs of efficient and reliable communications of IoT devices.

Method used

The low-power IoT direct-connected satellite uplink scheduling method based on deep reinforcement learning is adopted. By constructing a time division multiple access (TDMA) channel, Markov decision-making process model and self-attention mechanism, combined with the SAPPO algorithm, transmission time is dynamically allocated to optimize transmission efficiency and fairness, and the priority processing of important tasks is considered.

Benefits of technology

It improves transmission efficiency and fairness among equipment, optimizes the processing capabilities of important tasks, and is better than traditional scheduling algorithms, and is adapted to scenarios with different node counts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263259A_ABST
    Figure CN120263259A_ABST
Patent Text Reader

Abstract

The invention discloses a low-power-consumption Internet of Things direct connection satellite uplink scheduling method based on deep reinforcement learning, and aims to solve the transmission scheduling problem of distributed low-power-consumption Internet of Things equipment in satellite communication. The method is realized through the following steps that uplink transmission is divided into different time slots in a time division multiple access mode, each time slot corresponds to a specific transmission window, and therefore transmission conflicts among different devices are avoided. And calculating the satellite visible period of the Internet of Things equipment in the target area by using the double-line orbit element. A near-end strategy optimization algorithm based on a self-attention mechanism is adopted, and an optimal transmission strategy is learned through interaction with the environment according to parameters such as the visible period, the transmission fairness and the task priority of equipment. And according to the learned strategy, transmission time is dynamically allocated to each device, so that the uplink transmission efficiency, the device fairness and the priority processing capability are improved. According to the invention, an effective solution is provided for transmission scheduling of the low-power-consumption Internet of Things equipment in satellite communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication for low-orbit satellite Internet of Things, and relates to a low-power Internet of Things device direct connection to a low-earth orbit (LEO) satellite network model. Specifically, it is a method for uplink scheduling of low-power Internet of Things direct connection to satellites based on deep reinforcement learning. Background Art

[0002] Currently, Internet of Things devices are mainly concentrated in urban areas, and ground networks such as cellular and Wi-Fi can provide stable services. However, in remote areas such as mountains and oceans, it is difficult to cover the ground Internet of Things, and the demand for the Internet of Things in these areas is equally urgent, such as agricultural monitoring and environmental detection.

[0003] Satellite communication, with its characteristics of wide coverage and low latency, has become a powerful supplement to the communication of Internet of Things devices in remote areas. Low-earth orbit (LEO) satellites can quickly respond to the communication needs of ground devices. The low-power wide-area network (LPWAN) technology, with its advantages of low power consumption and long-distance transmission, is very suitable for realizing direct-to-satellite Internet of Things (DtS-IoT). However, the visible time of the satellite gateway is short, and devices need to complete data transmission within a short time, resulting in a tense satellite channel resource. Based on the Aloha access mechanism, the collision probability is already relatively high in the ground network. When communicating with satellites, the collision problem is more serious. A large number of devices simultaneously attempt uplink transmission within the limited visible time of the gateway, further exacerbating the collision probability and affecting the data transmission efficiency and reliability.

[0004] Existing scheduling algorithms such as first-come-first-served (FCFS) and fair scheduling do not perform well in the scenario of low-power Internet of Things direct connection to satellites. FCFS may cause some devices to have no transmission opportunity for a long time, and fair scheduling reduces the overall transmission efficiency when there are many devices. At the same time, traditional algorithms have poor processing ability for tasks with higher priorities. Summary of the Invention

[0005] The technical problem to be solved by the present invention is that considering the existing uplink scheduling technology for low-power Internet of Things direct connection to satellites cannot reasonably allocate channel resources, resulting in low transmission efficiency, poor fairness, and weak priority processing ability, and cannot meet the requirements of efficient and reliable communication of Internet of Things devices. A method for uplink scheduling of low-power Internet of Things direct connection to satellites based on deep reinforcement learning is provided.

[0006] A method for uplink scheduling of low-power Internet of Things direct connection to satellites based on deep reinforcement learning includes the following steps:

[0007] Build a low-power Internet of Things (IoT) direct connection low-Earth orbit satellite network model. Divide time slots according to the time-on-air (TOA) and guard interval of the transmitted data packets, and divide channels using the time division multiple access (TDMA) method. Based on the Markov decision process problem model and the proximal policy optimization (PPO) algorithm, combine the self-attention mechanism. Construct a low-power IoT direct satellite uplink scheduling strategy based on the self-attention proximal policy optimization (SAPPO) algorithm.

[0008] S1. Build a network system model, mainly including a low-power IoT direct connection low-Earth orbit satellite network model.

[0009] In the present invention, take the case where the device communicates using the LoRa method as an example.

[0010] The network consists of N IoT terminal devices, a central network server, a ground satellite gateway, and a low-Earth orbit (LEO) satellite equipped with a LoRa gateway. The IoT devices deployed on the ground communicate with the LoRaWAN gateway on the low-Earth orbit satellite via the LoRa method.

[0011] Adopt a network architecture with a single satellite and a single gateway. The satellite gateway only serves as a forwarding device, encapsulating the original data in an IP packet and sending it from the terminal device to the network server. The network server is responsible for setting and executing the LoRaWAN protocol in the network, and at the same time collecting the received data packets and decoding them.

[0012] The network server knows the location of each device and the visible time of the satellite. All devices use a spreading factor SF = 12 for transmission, and only schedule one transmission at a time without conflicts. A single uplink transmission makes the channel busy during the TOA period, and other devices cannot use the channel for communication at this time.

[0013] S2. Divide time slots according to the time-on-air (TOA) and guard interval of the transmitted data packets, and divide channels using the time division multiple access (TDMA) method.

[0014] ToA is the time-on-air when the node sends a data packet using SF, and is defined as follows:

[0015]

[0016] n preamble is the preamble of the LoRa frame, is the number of symbols in the payload of the LoRa frame. The symbol time can be calculated by the following formula:

[0017]

[0018] BW is the signal bandwidth. The higher the SF value, the more suitable for long-distance communication. We assume that the SF selected by the terminal device for LoRa communication is 12.

[0019] A guard time is added before and after each TOA. Whenever the network server schedules a transmission for a device, a reserved time is allocated to it, which is defined as follows:

[0020] T r = TOA + 2×T g

[0021] Each device completes its uplink transmission within the T allocated to it by the network server. Devices not allocated T cannot send information at this time, avoiding conflicts. r Devices not allocated T r cannot send information at this time, avoiding conflicts.

[0022] S3. Establish a problem model in the Markov decision process, combined with the self-attention mechanism.

[0023] Define the state space model in the Markov decision process:

[0024] The state space consists of four parts: the satellite visibility time T of the device, the time interval τ since the end of the last transmission, the number of uplink transmissions N of the device in the current cycle, and whether the current visibility time contains important tasks I.

[0025] The state is further defined as:

[0026] s t = {T, τ, N, I}

[0027] s t represents the state observed from the environment and satisfies is the state space

[0028] Define the action space model in the Markov decision process:

[0029] In the satellite Internet of Things network, the scheduling of time resources needs to comprehensively consider the remaining time slots from the end of the last transmission to the end of the current visibility time, the number of uplink transmissions that the device has completed in the current cycle, and whether the device has important tasks to transmit during the current visibility time.

[0030] The action space is defined as a discrete space. The action vector includes two parts: allocating transmission time for the device at the current moment and not allocating transmission time for the device. Among them, '1' means allocating a T r , and '0' means not allocating.

[0031] Define the reward function model in the Markov decision process:

[0032] The scheduling in the present invention needs to consider multiple objectives for optimization. In addition to the total number of device uplink transmissions, the scheduling rate of the device and the number of important tasks obtaining uplink transmissions also need to be ensured.

[0033] Therefore, the reward function r t is defined as:

[0034]

[0035] where R is the reward value obtained for successfully scheduling the uplink transmission of the device, α is the penalty coefficient in case of scheduling conflict, φ represents whether there is a conflict in the uplink scheduling, "1" represents a conflict, "0" represents no conflict, and the conflict here means that after allocating the uplink transmission time of T r to the device at the current moment, it exceeds the satellite visibility time of the device, k i is the importance coefficient, and a higher importance coefficient is obtained for transmitting important tasks, β i is the reward decay for multiple transmissions of the same device, a t = 1 indicates that the channel resources are allocated to the device at the current moment, indicating successful scheduling, otherwise represents scheduling failure due to conflicts or other reasons, p is the penalty value for not allocating channel resources, and η is the compensation coefficient in extreme cases. When the channel resources are tight and the agent chooses not to schedule the transmission for the device, a certain compensatory reward is given.

[0036] The self-attention mechanism is a special attention mechanism that is widely used in sequence processing tasks. Assume the input sequence is X, with dimensions (d model , N), d model is the dimension of the model, and N is the length of the sequence. Define three weight matrices W Q 、W K 、W V to generate the query Q, key K, and value V, where Q = XW Q , K = XW K , V = XW V The calculation process of the self-attention mechanism can be expressed as:

[0037]

[0038] O is the output of the self-attention, d k is the dimension of the key, is a scaling factor. The self-attention mechanism can also be extended to the multi-head attention mechanism. The input data will be divided into multiple "heads" and the self-attention calculation will be performed independently. Then the output vectors will be concatenated together to form the final output as follows:

[0039] MultiHead(Q, K, V) = Concat(O1,..., O h )W O

[0040] where WO is the output weight matrix, and Concat is the concatenation operation.

[0041] The self-attention mechanism is introduced into the network. Each network contains two sub-layers. The first sub-layer is the MHSA sub-layer, and the second sub-layer is the fully connected layer. Layer normalization is performed after the first sub-layer. Therefore, the output of the MHSA layer is:

[0042] MHSAOutput = LayerNorm(x + MHSA(x))

[0043] Among them, MHSA(x) is the function implemented by the MHSA sub-layer. Each hidden layer of the fully connected layer is a ReLU activation function after a linear transformation. The output of the fully connected layer can be expressed as:

[0044] FCN(x) = ReLU(xW train + b train )

[0045] S4. Low-power IoT direct satellite uplink scheduling strategy based on the SAPPO algorithm.

[0046] In the present invention, the construction of the SAPPO algorithm neural network is based on the self-attention mechanism and the fully connected network. A multi-head self-attention mechanism layer is added to the actor-critic network of the PPO algorithm, and the policy gradient method of the PPO algorithm is retained.

[0047] The goal of the policy gradient method is to find the policy parameter π that can maximize the expected discounted reward θ , and such optimization problems are solved using the gradient ascent algorithm with gradient estimation. The gradient estimator formula is expressed as:

[0048]

[0049] is the estimator of the advantage function A(s t , a t ), is the expectation on the finite batch of trajectories. The PPO algorithm uses the generalized advantage estimation (GAE) method to estimate the advantage function:

[0050]

[0051] Among them, λ is the coefficient used in GAE, is the state value function given by the critic network, and the parameter θ is updated using the gradient ascent method:

[0052]

[0053] where α ∈ [0, 1] represents the learning rate, and L(θ) is the loss function of the actor network. During the parameter update process, the new parameter θ and the old parameter θ old should be kept close or similar. Therefore, the following clipping function is introduced in PPO:

[0054]

[0055] The loss function after passing through the clipping function is obtained as:

[0056]

[0057] The parameters in the critic network are updated using the gradient descent method as follows:

[0058]

[0059] where represents the loss function of the critic network, which is defined as:

[0060]

[0061] The technical effects of the technical solution of the present invention are as follows:

[0062] The present invention uses time division multiple access technology for time slot division to avoid data packet collisions, and proposes an uplink transmission scheduling algorithm based on deep reinforcement learning - the SAPPO algorithm. This algorithm uses the self - attention mechanism to improve the algorithm performance, not only optimizing the number of uplink transmissions, but also considering the fairness among users.

[0063] In addition, the present invention also takes the priority processing problem of important tasks not considered by traditional algorithms as an optimization goal, which is not available in existing algorithms.

[0064] The SAPPO algorithm proposed by the present invention has better transmission efficiency than traditional scheduling algorithms in scenarios with different numbers of nodes, and better stability and convergence effect than existing DRL algorithms (such as DQN and A2C), which proves the effectiveness of the algorithm in the uplink scheduling of low - power Internet of Things direct - connected satellites. Brief Description of the Drawings

[0065] Figure 1 Flowchart for constructing a network system model.

[0066] Figure 2 Time slot division for different devices.

[0067] Figure 3 Flowchart for establishing a problem model in the Markov decision process.

[0068] Figure 4Flowchart of uplink transmission of low-power Internet of Things directly connected to satellites based on SAPPO.

[0069] Figure 5 Schematic diagram of the SAPPO network architecture. Detailed implementation manners

[0070] The present invention proposes a low-power Internet of Things direct satellite uplink scheduling method based on deep reinforcement learning. This method dynamically allocates the time for uplink transmission according to the visible period information of Internet of Things devices to improve the transmission efficiency, including the following steps:

[0071] Construct a network system model. The network consists of Internet of Things terminal devices, a ground satellite gateway, a central network server, and a LEO satellite carrying the gateway. Slots are divided according to the time-of-arrival (TOA) in space and the guard interval of the transmitted data packets, and each slot corresponds to a specific transmission window to avoid transmission conflicts between different devices;

[0072] Based on the Markov decision process problem model, combined with the low-power Internet of Things direct satellite uplink transmission, define the state space model and action space model in the Markov decision process, and transform the scheduling problem into a sequential decision problem;

[0073] Define the reward function model in the Markov decision process, incorporate the number of uplink transmissions into the reward design of the deep reinforcement learning algorithm, and add fairness constraints. At the same time, set higher priorities for important tasks, and give certain compensation for giving up transmission in extreme channel environments;

[0074] Construct a SAPPO network for low-power Internet of Things direct satellite uplink transmission scheduling, incorporate the self-attention mechanism into the network design of the PPO algorithm, and construct an uplink scheduling algorithm based on SAPPO.

[0075] The following combines the accompanying drawings to describe in detail the specific implementation manners of each step in the method of the present invention. However, it should be understood that the protection scope of the present invention is not limited by the specific implementation manners.

[0076] Figure 1 Construct a low-power Internet of Things direct low-earth orbit satellite network model, specifically including:

[0077] The network model mainly consists of the following parts: N Internet of Things terminal devices, a central network server, a ground satellite gateway, and a LEO satellite carrying the gateway.

[0078] N IoT terminal devices deployed on the ground communicate with the LoRaWAN gateway on the low-earth orbit satellite via LoRa. These IoT terminal devices are all equipped with LoRa modules and can transmit data with a spreading factor SF = 12. The satellite footprint is the coverage area of the satellite, and only the terminal devices within the satellite footprint can reach the gateway. The area covered by the satellite during its x-th flyover is given by the following formula:

[0079] A[x] = 2πE 2 (1 - cosψ[x])

[0080] where E is the radius of the Earth, in kilometers, and ψ[x] is the angular radius of the coverage circle during the x-th flyover. The calculation formula for ψ[x] is as follows:

[0081]

[0082] where h[x] is the altitude of the satellite during its x-th pass, in kilometers.

[0083] According to their own business requirements, the IoT terminal devices send the collected data in the form of LoRa signals. Due to the use of the spreading factor SF = 12, the data transmission has strong anti-interference ability and long transmission distance, and can effectively cover the visible range of the satellite.

[0084] After receiving the LoRa signal sent by the IoT terminal device, the LoRaWAN gateway on the LEO satellite performs encapsulation processing on the original data. The encapsulated IP packet is forwarded from the LoRaWAN gateway on the LEO satellite to the ground satellite gateway through the satellite communication link. The satellite gateway only acts as a data forwarding device and does not parse or process the received data.

[0085] After receiving the IP packet forwarded by the LEO satellite, the ground satellite gateway performs unpacking processing on it, extracts the original LoRa data packet, and sends the unpacked LoRa data packet to the central LoRaWAN network server.

[0086] The network server has the function of recording and managing the location information of each IoT terminal device and the visible time of the satellite, and reasonably schedules the transmission of the IoT terminal devices. During a single uplink transmission process, the channel will be busy during the transmission time occupancy (TOA), and other devices cannot use this channel for communication during this period, thus ensuring the orderliness and reliability of data transmission.

[0087] Figure 2 The time slot division of different devices includes:

[0088] Based on the characteristics of the data packets sent by the Internet of Things (IoT) terminal devices, the network server calculates the transmission time occupancy in the air (TOA) for each data packet. The calculation of TOA can be based on factors such as the size of the data packet, the transmission rate, and the propagation characteristics of the channel.

[0089] ToA is the time in the air for a node to send a data packet using SF and is defined as follows:

[0090]

[0091] n preamble is the preamble of the LoRa frame. is the number of symbols in the payload of the LoRa frame. The symbol time can be calculated by the following formula:

[0092]

[0093] BW is the signal bandwidth. The higher the SF value, the more suitable for long-distance communication. We assume that the SF selected by the terminal device for LoRa communication is 12.

[0094] Next, the network server adds a guard time before and after each TOA. The setting of the guard time is to avoid the possible edge overlap phenomenon when different devices transmit data packets and ensure the integrity of data transmission.

[0095] Whenever the network server arranges a transmission for a device, it will allocate a reserved time for it, which is defined as follows:

[0096] T r = TOA + 2 × T g

[0097] Each device completes the uplink transmission within the T r allocated to it by the network server. Devices not allocated T r cannot send information at this time.

[0098] Based on the TOA and the guard interval, the network server divides the entire available uplink transmission time into multiple time slots, and the length of each time slot is equal to TOA plus the two guard times before and after.

[0099] During the transmission time slot of each device, other devices not allocated this time slot will remain silent and not perform data transmission operations. In this way, even if multiple devices are simultaneously within the visible range of the satellite, channel conflicts will not occur due to simultaneous data transmission.

[0100] For example, when device 1 transmits in the time slot from 0 - 11 ms, device B remains silent; when device 2 transmits in the time slot from 11 ms - 22 ms, device A remains silent.

[0101] In this way, the time-division multiplexing of multiple devices and satellite uplink transmission is effectively realized, improving the utilization rate of the channel and the overall performance of the system.

[0102] Figure 3 , a problem model in the Markov decision process is proposed, including:

[0103] Determine the state space in the Markov decision process, which consists of four parts: the satellite visibility time T of the device, the time interval τ since the end of the last transmission, the number of uplink transmissions N of the device in the current cycle, and whether the current visibility time contains important tasks I.

[0104] The state is further defined as s t ={T, τ, N, I}, where T represents the discrete value corresponding to the satellite visibility time, τ represents the discrete value corresponding to the time interval, N represents the number of uplink transmissions, I represents whether there are important tasks (0 or 1), and s t represents the state observed from the environment, satisfying s t ∈S.

[0105] Define the action space model in the Markov decision process:

[0106] In the satellite Internet of Things network, the scheduling of time resources needs to comprehensively consider the remaining time slots from the end of the last transmission to the end of the current visibility time, the number of uplink transmissions that the device has completed in the current cycle, and whether the device has important tasks to transmit during the current visibility time.

[0107] Define the action space S as a discrete space, and the action vector consists of two parts: allocating transmission time for the device at the current moment and not allocating transmission time for the device. The action is represented in binary, where "1" represents allocating a transmission time slot and "0" represents not allocating.

[0108] Define the reward function model in the Markov decision process:

[0109] The scheduling in the present invention needs to consider multiple objectives for optimization, including the total number of device uplink transmissions, the scheduling rate of the device, and the number of times important tasks obtain uplink transmissions, etc. Therefore, a reward function is defined to measure the advantages and disadvantages of different scheduling decisions, expressed as:

[0110]

[0111] where R is the reward value obtained by successfully scheduling uplink transmission for the device, α is the penalty coefficient when there is a scheduling conflict, φ represents whether there is a scheduling conflict, "1" represents a conflict occurs, "0" represents no conflict, and the conflict here means that after allocating the uplink transmission time of T r for the device at the current moment, it exceeds the satellite visibility time of the device, ki is the importance coefficient. Higher importance coefficients are obtained for transmitting important tasks, β i is the reward decay for multiple transmissions by the same device, a t a = 1 indicates that channel resources are allocated to the device at the current moment, indicates successful scheduling, and conversely represents failed scheduling due to conflicts or other reasons. p is the penalty value for not allocating channel resources, and η is the compensation coefficient in extreme cases. When channel resources are scarce, a certain compensatory reward is given when the agent chooses not to schedule transmissions for the device.

[0112] The self-attention mechanism is a special attention mechanism that is widely used in sequence processing tasks. Assume the input sequence is X, with dimension (d model , N), where d model is the dimension of the model and N is the length of the sequence. Define three weight matrices W Q , W K , W V for generating the query Q, key K, and value V, where Q = XW Q , K = XW K , V = XW V . The calculation process of the self-attention mechanism can be expressed as:

[0113]

[0114] O is the output of self-attention, d k is the dimension of the key, is a scaling factor. The self-attention mechanism can also be extended to the multi-head attention mechanism. The input data is divided into multiple "heads" and self-attention calculations are performed independently. Then the output vectors are concatenated together to form the final output as follows:

[0115] MultiHead(Q, K, V) = Concat(O1,..., O h )W O

[0116] where W O is the output weight matrix and Concat is the concatenation operation.

[0117] In the embodiments of the present invention, according to the actual application scenario, each parameter in the reward function is assigned values and adjusted. The specific values of these parameters need to be comprehensively considered and optimized according to factors such as the performance requirements of the network, the transmission characteristics of the device, and the importance of the service.

[0118] Figure 4 Proposed is a low-power IoT direct satellite uplink transmission scheduling based on SAPPO, including:

[0119] For the selection of the algorithm, considering the complexity of the states in the satellite Internet of Things scheduling scenario, the PPO algorithm that can optimize both the policy and the value function simultaneously is selected as the basic algorithm. The PPO algorithm restricts the amplitude of policy updates through the truncated probability ratio, which can effectively avoid large fluctuations during the policy update process, thereby improving the stability and convergence speed of the algorithm.

[0120] However, in the satellite Internet of Things scheduling scenario, due to the complexity of the state space and action space, the PPO algorithm may face the problem of low sample efficiency when dealing with high-dimensional data. In the present invention, the PPO algorithm incorporated with the self-attention mechanism, namely SAPPO, is selected. The self-attention mechanism enables the algorithm to better capture the global dependencies between state features, enhancing the model's perception ability of complex states, and thus making more accurate action decisions.

[0121] In the present invention, the construction of the SAPPO algorithm neural network is based on the self-attention mechanism and the fully connected network. A multi-head self-attention mechanism layer is added to the actor-critic network of the PPO algorithm, while retaining the policy gradient method of the PPO algorithm.

[0122] The goal of the policy gradient method is to find the policy parameter π that can maximize the expected discounted reward θ , and such optimization problems are solved using the gradient ascent algorithm with gradient estimation. The gradient estimator formula is expressed as:

[0123]

[0124] is the estimator of the advantage function A(s t ,a t ), is the expectation on the finite batch of trajectories. The PPO algorithm uses the Generalized Advantage Estimation (GAE) method to estimate the advantage function:

[0125]

[0126] where λ is the coefficient used in GAE, is the state value function given by the critic network, and the parameter θ is updated using the gradient ascent method:

[0127]

[0128] where α ∈ [0, 1] represents the learning rate, and L(θ) is the loss function of the actor network. During the parameter update process, the new parameter θ and the old parameter θ old should be kept close or similar. Therefore, the following clipping function is introduced in PPO:

[0129]

[0130] The loss function after passing through the clipping function is as follows:

[0131]

[0132] In the critic network, the parameters are updated using gradient descent as follows:

[0133]

[0134] where represents the loss function of the critic network, which is defined as:

[0135]

[0136] The SAPPO network structure is as Figure 5 shown. There are three types of network structures in the SAPPO algorithm framework, namely the critic network with φ as a parameter, the new actor network with θ as a parameter, and the old actor network with as a parameter.

[0137] The of the old actor network is used to limit the change of the new policy θ .

[0138] The critic network takes the state provided by the environment and the actions generated by the actor network as inputs together. Its output can be understood as the expected cumulative discounted reward in the current state, which is used to evaluate the performance of the new actor network θ .

[0139] The self-attention mechanism is introduced into the network. Each network contains two sub-layers. The first sub-layer is the MHSA sub-layer, and the second sub-layer is the fully connected layer. Layer normalization is performed after the first sub-layer. The output of the MHSA layer is

[0140] MHSAOutput = LayerNorm(x + MHSA(x))

[0141] where MHSA(x) is the function implemented by the MHSA sub-layer. Each hidden layer of the fully connected layer is a ReLU activation function after a linear transformation. The output of the fully connected layer can be expressed as

[0142] FCN(x) = ReLU(xW train + b train )

[0143] The state-action function Q π (s,a) (i.e., Q-function) and the value function V π (s) can be expressed as

[0144]

[0145] At the start of training, initialize the parameters θ of the new actor network, the parameters θ of the old actor network old and the parameters of the critic network as well as the data buffer pool D. In each training step t, the agent first captures the current state s t and stores it in the buffer pool. The old actor network outputs a probability distribution of an action, and then samples from the action distribution to obtain the action a t . After executing the action a t , the agent interacts with the environment to obtain the corresponding reward and the new state s t+1 . After one round of iteration, the parameters θ old of the old actor network will be replaced by the parameters θ of the new actor network. After the parameter update is completed, SAPPO will clear the accumulated data from the training process to prepare for a new iteration.

[0146] The complete low-power IoT direct satellite uplink transmission scheduling algorithm based on SAPPO of the present invention is as follows:

[0147]

[0148]

[0149] For those skilled in the art, it is obvious that this application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of this application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of this application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within this application. Any reference signs in the claims should not be regarded as limiting the claimed rights.

Claims

1. A low-power Internet of Things direct satellite uplink scheduling method based on deep reinforcement learning, characterized in that Dynamically allocate the uplink transmission time according to the visibility period information of Internet of Things (IoT) devices to improve the transmission efficiency, including the following steps: Construct a network system model. The network consists of IoT terminal devices, a ground satellite gateway, a central network server, and a LEO satellite carrying the gateway. Divide time slots according to the time-of-arrival (TOA) in the air and the guard interval of the transmitted data packet. Each time slot corresponds to a specific transmission window to avoid transmission conflicts between different devices; Based on the Markov decision process problem model, combined with the low-power IoT direct satellite uplink transmission, define the state space model and action space model in the Markov decision process, and transform the scheduling problem into a sequential decision problem; Define the reward function model in the Markov decision process, incorporate the number of uplink transmissions into the reward design of the deep reinforcement learning algorithm, and add fairness constraints. At the same time, set higher priorities for important tasks and give certain compensation for giving up transmission in extreme channel environments; Construct a SAPPO network for low-power IoT direct satellite uplink transmission scheduling, integrate the self-attention mechanism into the network design of the PPO algorithm, and construct an uplink scheduling algorithm based on SAPPO; 2. The uplink scheduling method for low-power Internet of Things direct connection to satellites based on deep reinforcement learning according to claim 1, characterized in that: Construct a network system model. The network consists of N IoT terminal devices (EDs), a ground satellite gateway, a central network server (LNS), and a LEO satellite carrying the gateway. The IoT devices deployed on the ground communicate with the gateway on the LEO satellite through a wireless link. The satellite footprint is the coverage area of the satellite. Only the terminal devices within the satellite footprint can reach the gateway. The area covered by the satellite during the x-th flyover is given by the following formula: A[x] = 2πE 2 (1 - cosψ[x]) where E is the radius of the Earth in kilometers, and ψ[x] is the angular radius of the coverage circle during the x-th flyover. The calculation formula of ψ[x] is as follows: Where h[x] is the altitude of the satellite at the x-th pass, in kilometers, θ is the angle between the direction in which the beam is directly pointed at the satellite and the local horizontal plane, and whenever the LNS schedules a transmission for a device, a reserved time T is allocated to it. r , T r = TOA + 2×T g , T g is the set guard interval. Each device completes its uplink transmission within the T r allocated to it by the LNS. Devices not allocated T r cannot send information at this time, thus avoiding collisions.

3. The uplink scheduling method for low-power Internet of Things direct connection to satellites based on deep reinforcement learning according to claim 1, wherein: Based on the Markov decision process problem model and combined with the uplink transmission of low-power Internet of Things directly connected to satellites, the LNS can collect the key status information of devices in real time, including the satellite visibility time T of the device, the time interval τ since the end of the last transmission, the number of uplink transmissions N of the device in the current cycle, whether the current visible time contains important tasks I, and the state is further defined as s t ={T, τ, N, I}, s t represents the state observed from the environment and satisfies is the state space. The action space is defined as a discrete space. The action vector includes two parts: allocating transmission time for the device at the current moment and not allocating transmission time for the device. Among them, '1' means allocating a T r , and '0' means not allocating.

4. The uplink scheduling method for low-power Internet of Things direct connection to satellites based on deep reinforcement learning according to claim 1, characterized in that: Define the reward function model in the Markov decision process, and the reward function r t is expressed as Among them, R is the reward value obtained for successfully scheduling uplink transmission for the device, α is the penalty coefficient in case of scheduling conflict, φ represents whether uplink scheduling conflicts occur, "1" represents a conflict, "0" represents no conflict, and the conflict here means that after allocating the uplink transmission time of T r to the device at the current moment, it exceeds the satellite visibility time of the device, and k i is the importance coefficient. Higher importance coefficients are obtained for transmitting important tasks, and β i is the reward attenuation for multiple transmissions of the same device, and a t = 1 indicates that channel resources are allocated to the device at the current moment, indicating successful scheduling, represents scheduling failure due to conflict reasons. p is the penalty value for not allocating channel resources, and η is the compensation coefficient in extreme cases. When channel resources are tight and the agent chooses not to schedule transmission for the device, a certain compensation reward is given.

5. The uplink scheduling method for low-power Internet of Things direct connection to satellites based on deep reinforcement learning according to claim 1, wherein: Construct a SAPPO network for low-power IoT direct satellite uplink transmission scheduling. The SAPPO algorithm introduces the self-attention mechanism into the actor network and critic network of PPO. Each network contains two sub-layers. The first sub-layer is the MHSA sub-layer, and the second sub-layer is the fully connected layer. Layer normalization is performed after the first sub-layer. The output of the MHSA layer is MHSAOutput = LayerNorm(x + MHSA(x)) where MHSA(x) is the function implemented by the MHSA sub-layer. Each hidden layer of the fully connected layer is a ReLU activation function after a linear transformation. The output of the fully connected layer is expressed as FCN(x) = ReLU(xW train + b train ) State-action function Q π (s,a) is the Q-function and the value function V π (s) is expressed as At each training step t, the agent first captures the current device state s t and stores it in the buffer pool. Then, it obtains the scheduling action a through the actor network t . After executing the action a t , the agent interacts with the environment to obtain the corresponding reward and the new state s t+1 . After one round of iteration, the parameters θ of the old actor network old will be replaced by the parameters θ of the new actor network. After the parameter update is completed, SAPPO will clear the accumulated data from the training process to prepare for the new iteration.