A method for spreading factor allocation for lora wan terrestrial connected satellite systems

CN117221919BActive Publication Date: 2026-09-25TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311377712.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-23
Publication Date
2026-09-25
Estimated Expiration
2043-10-23

AI Technical Summary

Technical Problem

[0005]目前在地上无线传感器网络领域已经提出了各种扩频因子分配策略,但这些策略未考虑到地下土壤特性、同扩频因子干扰影响以及网络能效指标,难以适用基于LoRaWAN的大规模地下直连卫星场景

Benefits of technology

[0035]1)本发明在考虑准确的信道模型和同扩频因子干扰的影响下,采用多智能体强化学习模型为基于LoRaWAN的地下直连卫星系统推导出最优的扩频因子分配策略,相对于其他现有的扩频因子分配策略,该策略实现了更高的网络能效,为部署可持续的地下直连卫星系统奠定了基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117221919B_ABST
    Figure CN117221919B_ABST
Patent Text Reader

Abstract

The application relates to a spread factor allocation method for a LoRaWAN underground direct connection satellite system, which comprises the following steps: S1, constructing a LoRaWAN underground direct connection satellite system, initializing the spread factor configuration of all underground sensor nodes, and uploading the current state information of the underground sensor nodes to a LoRaWAN gateway arranged on a low earth orbit satellite by using the initialized spread factor configuration; S2, after the LoRaWAN gateway receives the state information of all the underground sensor nodes, a pre-trained multi-intelligent reinforcement learning model is used to calculate a spread factor allocation result maximizing network energy efficiency; and S3, the LoRaWAN gateway broadcasts the spread factor configuration result to all the underground sensor nodes, and the underground sensor nodes use the updated spread factor configuration for subsequent sensing data uploading. Compared with the prior art, the application can effectively reduce the power consumption of the underground sensor nodes, further improve the operation cycle of the system, and promote the field deployment of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of underground direct satellite connections and low-power wide area networks, and in particular to a spreading factor allocation method for LoRaWAN underground direct satellite connection systems. Background Technology

[0002] The LoRaWAN underground direct satellite system combines a LoRaWAN-based underground wireless sensor network with low Earth orbit satellites to provide a large-scale, low-cost underground environmental monitoring solution in remote or disaster-stricken areas. It can be widely used in fields such as automated agriculture, underground pipeline monitoring in remote areas, and disaster relief and rescue in post-disaster areas.

[0003] LoRa, as a modulation scheme in LoRaWAN, employs chirped spread spectrum modulation and introduces quasi-orthogonal characteristics, giving it excellent anti-interference capabilities. Simultaneously, LoRa modulation technology achieves different trade-offs between communication distance and energy consumption by adjusting the spreading factor. However, in large-scale underground direct satellite connection scenarios, the use of a medium access protocol similar to Aloha in LoRa modulation technology severely limits its network capacity and collision robustness. For example, when a large number of underground sensor nodes are configured with the same spreading factor, the probability of data packets carrying the same spreading factor experiencing interference on the same channel increases, resulting in a lower packet success rate.

[0004] Therefore, how to utilize the quasi-orthogonal characteristics of LoRa modulation technology to design the optimal spreading factor strategy for large-scale underground sensor nodes in order to improve the overall network energy efficiency is of great practical significance for realizing the actual deployment of large-scale LoRaWAN underground direct-connect satellite systems.

[0005] Various spreading factor allocation strategies have been proposed in the field of terrestrial wireless sensor networks, but these strategies do not take into account the characteristics of underground soil, the interference of the same spreading factor, and the network energy efficiency index, making them difficult to apply to large-scale underground direct satellite connection scenarios based on LoRaWAN.

[0006] Therefore, for LoRaWAN underground direct satellite systems, there is an urgent need to design a spreading factor allocation method applicable to large-scale LoRaWAN underground direct satellite scenarios, in order to effectively reduce the power consumption of underground sensor nodes, further improve the system's operating cycle, and promote the field deployment of the system. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a spreading factor allocation method for LoRaWAN underground direct satellite systems, which can effectively reduce the power consumption of underground sensor nodes, further improve the system's operating cycle, and promote the field deployment of the system.

[0008] The objective of this invention can be achieved through the following technical solutions:

[0009] This invention provides a spreading factor allocation method for LoRaWAN underground direct satellite systems, the method comprising:

[0010] Step S1: Construct a LoRaWAN underground direct satellite system and initialize the spreading factor configuration of all underground sensor nodes. The underground sensor nodes upload their current status information to the LoRaWAN gateway deployed on the low Earth orbit satellite using the initialized spreading factor configuration.

[0011] Step S2: After receiving the status information of all underground sensor nodes, the LoRaWAN gateway uses a pre-trained multi-intelligence reinforcement learning model to calculate the spreading factor allocation result that maximizes network energy efficiency.

[0012] Step S3: The LoRaWAN gateway broadcasts the spreading factor configuration result to all underground sensor nodes, and the underground sensor nodes use the updated spreading factor configuration to upload subsequent sensor data.

[0013] Preferably, the status information of each underground sensor node includes the current spreading factor configuration and the packet success rate P under interference-free conditions. SNR The packet success rate P under the capture effect SIR The overall probability of successful packet reception, P S And the average transmitted data packet energy (EPP).

[0014] Preferably, the construction process of the multi-agent reinforcement learning model includes:

[0015] 1) Intelligent Agent: Each underground sensor node is considered as a separate intelligent agent, and the intelligent agent elements... Meaning: The nth underground sensor node observes the environmental state in discrete time slot t according to the current strategy. Selected Action When the agent performs an action You will then receive a reward. And switch to the next state in the next time slot t+1.

[0016] 2) Action Space: Each agent can choose a spreading factor configuration from the candidate spreading factor configurations, so its action space is represented as A = {SF}. k}, k corresponds to the number of the candidate spreading factor configuration; then the set of actions selected by all agents in each time slot t is represented as N is the number of agents;

[0017] 3) State Space: The environmental state observed by each agent includes the selected spreading factor configuration and the packet success rate P under interference-free conditions. SNR The packet success rate P under the capture effect SIR The overall probability of successful packet reception, P S And the average transmitted data packet energy (EPP), the state space of each agent is represented as The set of environmental states observed by all agents in each time slot t can be represented as:

[0018] 4) Reward: The goal of the multi-agent reinforcement learning model is to minimize the average energy EPP of each underground sensor node. The reward function for each agent in each time slot t is expressed as follows: The sum of rewards for all agents can be represented as

[0019] Preferably, the candidate spreading factor configuration includes SF7, SF8, SF9, SF10, SF11, and SF12, which respectively represent the number of symbols that need to be transmitted for one information bit as 2^7, 2^8, 2^9, 2^10, 2^11, and 2^12.

[0020] Preferably, the spreading factor for initializing all underground sensor nodes is configured as SF12.

[0021] Preferably, the packet success rate P under the capture effect SIR The calculation process is as follows:

[0022]

[0023] Where δ is a given threshold, Φ is the set of interference signals that arrive simultaneously with the target signal in the same channel and with the same spreading factor, and N k It is the spreading factor configuration SF k The number of nodes, ToA k When configured as SF k The time it takes for data packets to travel through the air, T p N represents the data packet upload period. c This represents the number of uplink channels, 2F1() represents the Gaussian hypergeometry function, and d max η represents the maximum distance between the underground sensor node and the gateway, and η represents the air path loss index.

[0024] Preferably, the interference-free packet reception success rate P SNR The expression is:

[0025]

[0026] Among them, Gt and G r Represent the antenna gains of the underground sensor node and the gateway, respectively, |h| 2 This indicates that the channel gain from the underground sensor node to the gateway follows a Rayleigh distribution. represents the variance of additive white Gaussian noise, q is the set signal-to-noise ratio demodulation threshold; g(d) represents the path loss at a distance d from the underground sensor node to the gateway.

[0027] Preferably, the multi-agent reinforcement learning model is MAD3QN, which is a value-function-based multi-agent combat dual-deep Q-network; each agent is trained through a deep Q-network. Estimate the expected value of the Q-value distribution corresponding to the action, where ω n This represents the weights of the nth deep Q-network.

[0028] Preferably, the MAD3QN obtains the optimal depth Q-network weights by minimizing a loss function, the expression of which is:

[0029]

[0030] Where γ is a discount factor, used to characterize the importance of future rewards to the current moment. The weights of the nth target depth Q-network are generated by cloning the current depth Q-network and are updated based on the weights of the current depth Q-network after a fixed number of iterations.

[0031] Preferably, the MAD3QN divides the output layer of each agent into two predictors, namely the dominance function. and value function The output network layer of each agent is represented as follows:

[0032]

[0033] in, κ n ,ν n These are the network weight parameters for the common part, value function, and advantage function of the nth agent network, respectively.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] 1) Considering the accurate channel model and the impact of co-spreading factor interference, this invention uses a multi-agent reinforcement learning model to derive the optimal spreading factor allocation strategy for LoRaWAN-based underground direct-connect satellite systems. Compared with other existing spreading factor allocation strategies, this strategy achieves higher network energy efficiency, laying the foundation for deploying sustainable underground direct-connect satellite systems.

[0036] 2) The multi-agent reinforcement learning model used in this invention is a value function-based multi-agent dueling double-deep Q-network (MAD3QN). Compared with traditional multi-agent reinforcement learning models, it effectively solves the problem of Q-value overestimation and improves the learning convergence speed. Attached Figure Description

[0037] Figure 1 This is a flowchart of the method of the present invention;

[0038] Figure 2 A schematic diagram of the LoRaWAN underground direct-connect satellite system deployment;

[0039] Figure 3 This is a flowchart of the method in the embodiment;

[0040] Figure 4 The flowchart of the MAD3QN algorithm in the embodiment is shown below;

[0041] Figure 5 The average EPP of four spread factor benchmark schemes and the proposed MAD3QN scheme is compared with those of the four spread factor benchmark schemes under different numbers of nodes.

[0042] Figure 6 This shows the spread factor distribution under different spread factor allocation schemes when the number of nodes is 5000. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0044] Example

[0045] For large-scale underground direct-connect satellite systems based on LoRaWAN, this embodiment considers co-spreading factor interference and underground soil characteristics, and presents a spreading factor allocation method for LoRaWAN underground direct-connect satellite systems. An optimal spreading factor configuration is calculated using a multi-agent reinforcement learning model to minimize the energy consumption of underground sensor nodes. Figure 1 and Figure 3 As shown, the method includes the following steps:

[0046] Step S1: Construct a LoRaWAN underground direct satellite system, initialize the spreading factor configuration of all underground sensor nodes, and upload the current status information of the underground sensor nodes to the LoRaWAN gateway deployed on the low Earth orbit satellite using the initialized spreading factor configuration.

[0047] Step S2: After receiving the status information of all underground sensor nodes, the LoRaWAN gateway uses a pre-trained multi-agent reinforcement learning model to calculate the spread spectrum factor allocation result that maximizes energy efficiency.

[0048] Step S3: The LoRaWAN gateway broadcasts the spreading factor configuration result to all underground sensor nodes, and the underground sensor nodes use the updated spreading factor configuration to upload subsequent sensing data.

[0049] Next, each step will be explained in detail.

[0050] 1. Construct a LoRaWAN underground direct satellite connection system, initialize the spreading factor configuration of all underground sensor nodes, and the underground sensor nodes upload the current status information to the LoRaWAN gateway deployed on the low Earth orbit satellite using the initialized spreading factor configuration.

[0051] like Figure 2 The LoRaWAN-based underground direct satellite connection system shown depicts a LoRaWAN gateway deployed on a low Earth orbit satellite, using a single beam to cover N locations buried at a depth of d within a circular area of ​​radius R. u The LoRa modulation uses each bit to represent the payload information, and the spreading factor represents the number of symbols transmitted for each bit. There are six spreading factor configurations in LoRa modulation: SF7, SF8, SF9, SF10, SF11, and SF12. For example, SF7 means that 2^7 symbols need to be transmitted for one bit, and SF12 means that 2^12 symbols need to be transmitted for one bit. Therefore, from SF7 to SF12, it means stronger anti-interference capability and longer communication range, but also longer air propagation time and higher transmission power consumption. During initialization of the spreading factor configuration, each underground sensor node is configured with the strongest spreading factor by default, SF12, to ensure that the gateway deployed on the satellite can collect the status information of all underground sensor nodes. The status information of each underground sensor node includes the current spreading factor configuration and the packet success rate P under interference-free conditions. SNR The packet success rate P under the capture effect SIR The overall probability of successful packet reception, P SAnd the average transmitted data packet energy (EPP).

[0052] First, in the absence of interference, the probability of a data packet being successfully demodulated can be expressed as:

[0053]

[0054] Among them, G t and G r Represent the antenna gains of the underground sensor node and the gateway, respectively, |h| 2 This indicates that the power attenuation in Ruili follows an exponential distribution with a unit mean. Let q = {-6, -9, -12, -15, -17.5, -20} dB represent the variance of additive white Gaussian noise, respectively. k The signal-to-noise ratio demodulation threshold ({SF) for |k=1,…,6} k |k=1,…,6} corresponds to SF7~SF12). Additionally, g(d) represents the path loss at a distance d from the underground sensor node to the gateway, which can be expressed as…

[0055]

[0056] Where f represents the carrier frequency, c represents the speed of light in free space, and η represents the air path loss exponent. The distance of propagation in the underground soil is represented by α, which is the attenuation constant and can be expressed as:

[0057]

[0058] β is the phase shift constant, which can be expressed as

[0059]

[0060] Where, μ r ε0 is the relative permeability of the soil, μ0 is the permeability of free space, ε0 is the permittivity of free space, and ε′ and ε″ are the real and imaginary parts of the soil permittivity, respectively, which can be determined by the MBSDM model, i.e., (5)-(20). Only the soil clay percentage C, the electromagnetic wave frequency f propagating in the soil, and the soil volumetric water content m0 are given. v Then we can obtain:

[0061]

[0062]

[0063]

[0064] Where, n m ,n d ,nb ,n f and κ m ,κ d ,κ b ,κ f These represent the values ​​of refractive index and normalized attenuation coefficient, respectively. The subscripts m, d, b, and f represent wet soil, dry soil, soil bound water, and soil free water, respectively. vt This represents the maximum volume of bound water in a given soil layer. The refractive index and normalized attenuation coefficient for dry soil, bound water, and free water are expressed as:

[0065]

[0066]

[0067] The real and imaginary parts of the dielectric constant and loss factor for bound water and free water in the soil are derived from the Debye relaxation equation:

[0068]

[0069]

[0070] Where σ b,f , τ b,f and ε 0b,0f These are the electrical conductivity, relaxation time, and dielectric constant at the low-frequency limit, which are related to soil bound water and soil free water, respectively. ∞ The dielectric constant at the high-frequency limit is 4.9 in both bound water and free water in the soil. The spectral parameters in (5)-(11) above can be obtained from a large number of soil samples using the following empirical formulas (12)-(20):

[0071] n d =1.634 - 0.539 × 10 -2 C + 0.2748 × 10 -4 C 2 (12)

[0072] k d =0.03952 - 0.04038 × 10 -2 C, (13)

[0073] m vt =0.02863 + 0.30673 × 10 -2 C, (14)

[0074] ε 0b =79.8 - 85.4 × 10 -2 C + 32.7 × 10 -4 C 2 (15)

[0075] ε 0f =100, (16)

[0076] τ b =1.062×10 -11 +3.450×10 -12 ×10 -2 C, (17)

[0077] τ f =8.5×10 -12 (18)

[0078] σ b =0.3112 + 0.467 × 10 -2 C, (19)

[0079] σ f =0.3631 + 1.217 × 10 -2 C. (20)

[0080] Studies have shown that in LoRa modulation, utilizing the capture effect, when the signal-to-interference ratio (SIR) is higher than a threshold δ, the target signal can be successfully demodulated from the interfering signals. Therefore, given a threshold δ and a set Φ of interfering signals arriving simultaneously with the target signal in the same channel and with the same spreading factor, the probability that the target signal can be demodulated under the capture effect can be expressed as:

[0081]

[0082] Where i represents the interference signal, N k Indicates to configure SF k The number of nodes, ToA k When configured as SF k The time it takes for data packets to travel through the air, T p N represents the data packet upload period. c This represents the number of uplink channels, 2F1() represents the Gaussian hypergeometry function, and d max This indicates the maximum distance between the underground sensor node and the gateway.

[0083] Therefore, the overall probability of successful packet transmission can be expressed as:

[0084] P S (d)=P SNR (d)P SIR (d). (22)

[0085] The optimized energy efficiency index in this embodiment is characterized by the average energy consumed by each underground sensor node in successfully uploading a data packet to the gateway, and can therefore be expressed as:

[0086]

[0087] Among them, V supply I represents the supply voltage of the underground sensor node. t It is determined by the transmission power P t The determined emission current.

[0088] 2. After receiving the status information of all underground sensor nodes, the LoRaWAN gateway uses a pre-trained multi-agent reinforcement learning model to calculate the spread spectrum factor allocation result that maximizes energy efficiency.

[0089] After the gateway receives the state information of all nodes, it immediately adopts the proposed multi-agent reinforcement learning method to derive the optimal spreading factor allocation strategy by minimizing the EPP of the underground sensor nodes. The proposed multi-agent reinforcement learning method mainly consists of the following four parts:

[0090] 1) Intelligent Agent: Each underground sensor node can be considered as a separate intelligent agent. Each intelligent agent includes four elements, namely... This means that the nth underground sensor node, according to the current strategy, observes the environmental state in discrete time slot t. Selected action When the agent performs an action You will then receive a reward. And switch to the next state in the next time slot t+1.

[0091] 2) Action Space: Each agent can choose one spreading factor configuration from six options, therefore its action space can be represented as A = {SF} k Therefore, the set of actions selected by all agents in each time slot t can be represented as:

[0092] 3) State Space: The environmental state observed by each agent includes the selected spreading factor configuration and the packet success rate P under interference-free conditions. SNR The packet success rate P under the capture effect SIR The overall probability of successful packet reception, P S And the average transmitted data packet energy (EPP), therefore the state space of each agent can be represented as Therefore, the set of environmental states observed by all agents in each time slot t can be represented as:

[0093] 4) Reward: The proposed multi-intelligence reinforcement learning aims to minimize the average energy EPP of each underground sensor node. Therefore, the reward function for each agent in each time slot t can be expressed as: Therefore, the sum of rewards for all agents can be represented as

[0094] As another preferred embodiment, the multi-agent reinforcement learning model is a value function-based multi-agent dueling double-deep Q-network (MAD3QN). MAD3QN can be used to obtain an optimal policy that maps states to action distributions. Each agent's policy is derived through the deep Q-network. Estimate the expected value of the Q-value distribution corresponding to the action, where ω n This represents the weights of the nth deep Q-network. MAD3QN needs to minimize the policy's loss function to obtain the optimal deep Q-network weights; therefore, the loss function can be defined as...

[0095]

[0096] Where γ is a discount factor, used to characterize the importance of future rewards to the current moment. The weights representing the nth target depth Q-network are generated by cloning the current depth Q-network and updated based on the current depth Q-network weights after a fixed number of iterations. Furthermore, MAD3QN divides the output layer of each agent into two predictors: an advantage function and a state value function, to improve the convergence speed of the algorithm. Therefore, the output network layer of each agent can be represented as...

[0097]

[0098] in, κ n ,ν n These are the network weight parameters for the common part, value function, and advantage function of the nth agent network, respectively. The MAD3QN algorithm flow is as follows: Figure 4 As shown.

[0099] 3. The LoRaWAN gateway broadcasts the spreading factor configuration result to all underground sensor nodes, which then use the updated spreading factor configuration to upload subsequent sensor data.

[0100] Once the gateway derives the optimal spreading factor allocation scheme using the MAD3QN algorithm, it broadcasts the result to all underground sensor nodes. After all underground sensor nodes receive the allocated spreading factor configuration, they will use that configuration for their next uplink data transmission, thereby maximizing the energy efficiency of the LoRaWAN-based underground direct-connect satellite network.

[0101] Below, this embodiment uses numerical simulation to quantitatively analyze the actual performance of the spread spectrum factor allocation scheme based on multi-agent reinforcement learning in the LoRaWAN-based underground direct satellite connection scenario, and compares it with four benchmark schemes to verify the effectiveness of the present invention.

[0102] Simulation Setup: In this specific implementation, the LoRaWAN gateway is considered to be deployed on a low Earth orbit satellite at an altitude of 200 km, providing single-beam coverage of the entire central irrigation farm. To ensure that the simulation results reflect real-world system performance, on-site soil parameters, such as volumetric water content and clay content, are used. Simultaneously, underground sensor nodes buried at a depth of 0.4 m are configured to upload a 23-byte physical layer data packet every 600 seconds. Specific simulation parameters are shown in Table 1 below.

[0103] Table 1. Description of Initialization Condition Parameters

[0104]

[0105]

[0106] Four baseline spreading factor strategies: (1) Same SF: The spreading factor configuration of all underground sensor nodes is consistent; (2) EIB: The monitored circular area is divided into six equal-width rings (i.e., R / 6), and SF7 to SF12 are assigned sequentially from the innermost to the outermost ring; (3) EAB: The monitored circular area is divided into six equal-area rings (i.e., πR / 6). 2 / 6), and SF7 to SF12 are assigned sequentially from the inside to the outside of the ring; (4) PLB: All underground sensor nodes are assigned their respective spreading factor configurations according to the path loss model (i.e. formula (2)) and the signal-to-noise ratio threshold of each spreading factor d.

[0107] Simulation results: such as Figure 5 As shown, the obtained optimal spreading factor allocation scheme is compared with four benchmark schemes under different numbers of nodes to perform average EPP (i.e., A comparison of ). From Figure 5It is evident that, compared to the four benchmark schemes, the method proposed in this invention exhibits superior average EPP across different numbers of nodes, and this performance improvement becomes more significant as the number of nodes increases. Specifically, when the number of nodes is 5000, the average EPP of the proposed spread spectrum allocation strategy is only 2.46J, while the lowest average EPP among the four benchmark schemes is as high as 7.52J. Figure 6 The distribution of spreading factors in the monitoring area under different spreading factor allocation strategies is listed when there are 5000 nodes. Figure 6 It can be observed that the proposed spreading factor allocation scheme places SF8-10 near the inner ring, and only allocates a small number of SF11 or SF12 in the outer ring. This aims to minimize transmit power and the probability of interference with the same spreading factor while ensuring reliable connectivity. Therefore, compared to other benchmark schemes, the proposed scheme exhibits a significant improvement in network energy efficiency.

[0108] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for spreading factor allocation in a LoRaWAN underground direct-connect satellite system, characterized in that, The method includes: Step S1: Construct a LoRaWAN underground direct satellite system and initialize the spreading factor configuration of all underground sensor nodes. The underground sensor nodes upload their current status information to the LoRaWAN gateway deployed on the low Earth orbit satellite using the initialized spreading factor configuration. Step S2: After receiving the status information of all underground sensor nodes, the LoRaWAN gateway uses a pre-trained multi-agent reinforcement learning model to calculate the spread spectrum factor allocation result that maximizes energy efficiency. Step S3: The LoRaWAN gateway broadcasts the spreading factor configuration result to all underground sensor nodes, and the underground sensor nodes use the updated spreading factor configuration to upload subsequent sensor data. The status information for each underground sensor node includes the current spreading factor configuration and the packet success rate under interference-free conditions. P SNR Packet success rate under capture effect P SIR The overall probability of successful packet reception P S and average data packet energy EPP ; The capture effect packet success rate P SIR The calculation process is as follows: in, δ Given a threshold, Φ is the set of interference signals that arrive simultaneously with the target signal in the same channel and with the same spreading factor. N k It is the configuration of the spreading factor. SF k The number of nodes, ToA k When configured as SF k The time it takes for data packets to travel through the air. T p Represents the data packet upload cycle. N c It represents the number of uplink channels, 2. F 1() represents the Gaussian hypergeometric function. d max This indicates the maximum distance between the underground sensor node and the gateway. η This represents the air path loss index. g ( d () represents the distance from the underground sensor node to the gateway as... d Path loss; The interference-free packet reception rate P SNR The expression is: in, G t and G r Representing the antenna gains of the underground sensor node and the gateway, respectively. h | 2 This indicates that the channel gain from the underground sensor node to the gateway follows a Rayleigh distribution. This represents the variance of additive white Gaussian noise. q The set signal-to-noise ratio demodulation threshold; The multi-agent reinforcement learning model is MAD3QN, which is a value-function-based multi-agent combat dual-depth model. Q Network; each agent through deep learning Q network To estimate the expected distribution of Q-values ​​corresponding to the actions, where ω n Representing the n Each depth Q Network weights For the first n An underground sensor node in discrete time slots t Based on the current strategy, observe the environmental state.

2. The spreading factor allocation method for a LoRaWAN underground direct-connect satellite system according to claim 1, characterized in that, The construction process of the multi-agent reinforcement learning model includes: 1) Intelligent Agent: Each underground sensor node is considered as a separate intelligent agent, and the intelligent agent elements... Meaning: The first n An underground sensor node in discrete time slots t Based on the current strategy, observe the environmental state Selected Action When the intelligent agent performs an action You will then receive a reward. And in the next time slot t +1 Switch to the next state ; 2) Action Space: Each agent can choose a spreading factor configuration from the candidate spreading factor configurations, therefore its action space is represented as follows: A ={ SF k }, The corresponding candidate spreading factor configuration number; then all agents in each time slot t The selected set of actions is represented as , The number of intelligent agents; 3) State Space: The environmental state observed by each agent includes the selected spreading factor configuration and the packet success rate under interference-free conditions. P SNR Packet success rate under capture effect P SIR The overall probability of successful packet reception P S and average data packet energy EPP The state space of each agent is represented as: Then all intelligent agents in each time slot t The set of observed environmental states can be represented as ; 4) Reward: The goal of the multi-intelligence reinforcement learning model is to minimize the average energy of each underground sensor node. EPP Each agent in each time slot t The reward function is expressed as follows: The sum of rewards for all agents can be represented as .

3. The spreading factor allocation method for a LoRaWAN underground direct satellite system according to claim 2, characterized in that, The candidate spreading factor configurations include SF7, SF8, SF9, SF10, SF11, and SF12, which respectively represent the number of symbols that need to be sent for one information bit as 2^7, 2^8, 2^9, 2^10, 2^11, and 2^12.

4. The spreading factor allocation method for a LoRaWAN underground direct satellite system according to claim 3, characterized in that, The initialization spread factor configuration for all underground sensor nodes is set to SF12.

5. A spreading factor allocation method for a LoRaWAN underground direct-connect satellite system according to claim 1, characterized in that, The MAD3QN obtains the optimal weights for the deep Q-network by minimizing a loss function, the expression of which is: in, γ It is a discount factor used to characterize the importance of future rewards to the current moment. Representing the n Target depth Q The network weights are determined by cloning the current depth. Q Network generation, and after a fixed number of iterations based on the current depth Q The purpose of updating network weights is to update network weights.

6. The spreading factor allocation method for a LoRaWAN underground direct satellite system according to claim 5, characterized in that, The MAD3QN divides the output layer of each agent into two predictors, namely the dominance function. and value function The output network layer of each agent is represented as: in, The first n The network weight parameters of the common part, value function, and advantage function of each agent network.