Diffusion model enhanced low-altitude wireless network path selection method

By enhancing the state and action space of the low-altitude wireless network path selection algorithm through a diffusion model, the adaptability and reliability issues of the low-altitude wireless network in time-varying environments are resolved, and the performance of the path selection algorithm in emergency and tactical communications is improved.

CN121486926BActive Publication Date: 2026-04-07NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing low-altitude wireless network path selection algorithms suffer from insufficient adaptability and reliable decision-making capabilities in environments with changing network conditions. In particular, they struggle to effectively cope with dynamic changes in topology and link status, especially in emergency and tactical communication scenarios.

Method used

We design a state-space enhancement mechanism based on noise injection and an action-space enhancement mechanism based on progressive denoising. By using a diffusion model, we enhance the network state input and action exploration capabilities of the path selection decision agent, thereby improving the adaptability and reliable decision-making capabilities of the path selection algorithm.

Benefits of technology

It significantly improves the adaptability of the path selection algorithm under scenarios of changes in topology and link status, enhances the service quality assurance capability of differentiated services, breaks through the bottleneck of local optima, and constructs a better strategy exploration space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486926B_ABST
    Figure CN121486926B_ABST
Patent Text Reader

Abstract

This invention discloses a path selection method for low-altitude wireless networks enhanced by a diffusion model, comprising: constructing a network state representation for the network environment; constructing a virtual network state representation using a diffusion model; combining the network state representation and the virtual network state representation with weights to construct an enhanced perception state representation, inputting it into a path selection decision agent, and outputting a corresponding action; constructing a virtual action using a diffusion model; combining the output action of the path selection decision agent and the virtual action with weights to construct an enhanced action; performing path selection, calculating a reward and feeding it back to the path selection decision agent; and applying the trained path selection decision agent to perform path selection in the network environment. This invention improves the adaptability of path selection algorithms to time-varying network states, providing key support for the effective application of low-altitude wireless network path selection methods in complex time-varying scenarios such as emergency communication and tactical communication.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information engineering, and in particular to a low-altitude wireless network path selection method enhanced by a diffusion model. BACKGROUND

[0002] Low-altitude wireless networks guarantee the on-demand delivery of differentiated services through air-ground collaborative networking and dynamic path selection. However, in complex application scenarios such as emergency communication and tactical communication, low-altitude wireless networks have the characteristic of time-varying network state, specifically, time-varying topology structure and time-varying link state. These typical characteristics pose new challenges to the adjustment and adaptation capabilities and global reliable decision-making capabilities of path selection algorithms.

[0003] Existing path selection algorithms can be summarized into three categories: traditional path selection algorithms, heuristic path selection algorithms, and intelligent path selection algorithms. Traditional path selection algorithms, represented by Open Shortest Path First (OSPF) and Routing Information Protocol (RIP), rely on static topology modeling and fixed rule decision-making as their core mechanisms. Although they have the advantage of low computational complexity, they are difficult to effectively adapt to the dynamic characteristics of time-varying topology structure of low-altitude wireless networks. Heuristic path selection algorithms, including Particle Swarm Optimization (PSO), Genetic Algorithm (GA), and Ant Colony Optimization (ACO), partially alleviate the above problems by introducing dynamic modeling and experience-driven parameter tuning mechanisms. However, their convergence efficiency and solution quality are significantly dependent on parameter configuration, and in the scenario of time-varying link state of low-altitude wireless networks, decision actions are prone to local optimum or oscillation state. In contrast, Deep Reinforcement Learning (DRL) provides an innovative solution for adaptive path optimization in dynamic network environments. Through an interactive learning mechanism between agents and the environment, this technology can dynamically adjust path selection strategies based on real-time reward feedback, effectively dealing with the time-varying characteristics of network topology and link state.

[0004] However, in the face of low-altitude wireless networks with time-varying network state, path selection algorithms based on deep reinforcement learning still have two defects. On the one hand, the dynamic changes in network topology structure and link state can lead to distorted state information, and such path selection algorithms lack a rapid response mechanism and are difficult to adapt to highly dynamic network environments. On the other hand, such path selection algorithms rely heavily on the current strategy sampling action. Due to the spike distribution characteristics of the strategy network output, the action exploration space may be limited, which can cause the algorithm to fall into a local optimum dilemma and affect its decision reliability. SUMMARY

[0005] The purpose of this invention is to provide a diffusion model-enhanced path selection method for low-altitude wireless networks. This method aims to overcome the dual deficiencies of existing low-altitude wireless network path selection algorithms, which generally suffer from insufficient adaptability and reliable decision-making capabilities in time-varying network conditions. The method employs a noise-injection-based state-space enhancement mechanism to provide the path selection decision-making agent with richer network state inputs, thereby improving the algorithm's adaptability to time-varying network states. Furthermore, it designs a progressive denoising-based action-space enhancement mechanism to encourage the agent to continuously explore new actions, thus breaking through local optima bottlenecks and enhancing the algorithm's reliable decision-making capabilities. This provides crucial support for the effective application of low-altitude wireless network path selection methods in complex time-varying scenarios such as emergency communication and tactical communication.

[0006] To achieve the above functions, this invention designs a path selection method for low-altitude wireless networks enhanced with a diffusion model. For a network environment containing a path selection decision agent, the following steps S1-S6 are executed to complete the path selection:

[0007] Step S1: Collect network state information in real time, extract network state information at the current moment and several previous historical moments, and construct a time-series network state representation;

[0008] Step S2: Use a diffusion model to progressively denoise the network state representation to generate a virtual network state representation, and then weightedly fuse it with the network state representation to form an enhanced perception state representation.

[0009] Step S3: Input the enhanced perception state representation into the path selection decision intelligent agent model. The path selection decision intelligent agent model outputs corresponding actions based on the enhanced perception state representation and the DPPO model.

[0010] Step S4: Using the diffusion model, based on the enhanced perception state representation, virtual actions are obtained through a progressive denoising process. The output actions of the path selection decision intelligent agent model and the virtual actions are weighted and combined to construct the enhanced actions.

[0011] Step S5: Input the enhanced action into the network environment. Based on the enhanced action execution path selection, the network environment constructs a reward function, calculates the reward, and feeds the reward back to the path selection decision-making intelligent agent model.

[0012] Step S6: Iteratively train the intelligent agent model for path selection decision until the DPPO model on which it is based converges, obtain the trained intelligent agent model for path selection decision, and apply it to complete path selection in the network environment.

[0013] Beneficial effects: Compared with the prior art, the advantages of the present invention include:

[0014] This invention designs a path selection method for low-altitude wireless networks enhanced by a diffusion model. This method employs a noise-injection-based state space enhancement mechanism to provide the path selection decision agent with richer network state inputs, thereby improving the algorithm's adaptability to time-varying network states. Furthermore, it designs an action space enhancement mechanism based on progressive denoising to encourage the path selection decision agent to continuously explore new actions, thus overcoming local optima bottlenecks and enhancing the algorithm's reliable decision-making capability. Verification results show that, compared to other baseline algorithms, the DPPO algorithm exhibits stronger adaptability in two typical scenarios: topology changes and link state changes. While constructing a policy exploration space with better coverage, it significantly improves the quality of service (QoS) assurance capability for differentiated services. Attached Figure Description

[0015] Figure 1 This is a structural diagram of the DPPO model provided according to an embodiment of the present invention;

[0016] Figure 2 This is an end-to-end average delay diagram of five algorithms under the topology change scenario provided by the embodiments of the present invention;

[0017] Figure 3 This is a network throughput diagram of five algorithms under a topology change scenario provided by an embodiment of the present invention;

[0018] Figure 4 This is a packet loss rate diagram of five algorithms under the topology change scenario provided by the embodiments of the present invention;

[0019] Figure 5 This is an end-to-end average latency diagram of five algorithms in a resource-constrained scenario provided by an embodiment of the present invention;

[0020] Figure 6 This is a network throughput diagram of five algorithms in a resource-constrained scenario provided by an embodiment of the present invention;

[0021] Figure 7 This is a packet loss rate diagram of five algorithms in a resource-constrained scenario provided by an embodiment of the present invention;

[0022] Figure 8 This is an action selection density and frequency diagram of the conventional PPO algorithm provided according to an embodiment of the present invention;

[0023] Figure 9 This is an action selection density and frequency diagram of the DPPO model provided according to an embodiment of the present invention;

[0024] Figure 10 These are end-to-end average delay diagrams for five algorithms under different flow intensities provided in embodiments of the present invention;

[0025] Figure 11 These are network throughput diagrams for five algorithms under different traffic intensities, provided by embodiments of the present invention.

[0026] Figure 12 This is a packet loss rate diagram for five algorithms under different traffic intensities provided by an embodiment of the present invention. Detailed Implementation

[0027] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0028] This invention provides a diffusion model-enhanced path selection method for low-altitude wireless networks. For a network environment containing a path selection decision agent, the following steps S1-S6 are performed to complete path selection:

[0029] Step S1: Collect network state information in real time, extract network state information at the current moment and several previous historical moments, and construct a time-series network state representation;

[0030] The specific method for step S1 is as follows:

[0031] Collect any Network status at any time Including global network service requests Remaining bandwidth of the path Path delay and path packet loss rate The information is as follows:

[0032] ;

[0033] in, Dimensions representing network state ; Indicates the number of paths;

[0034] Based on the network status at each historical moment, construct The network state at time t is represented by the following formula:

[0035] ;

[0036] in, express Network state representation at time t, express Time and Time before The network state at a time step.

[0037] Step S2: Use a diffusion model to progressively denoise the network state representation to generate a virtual network state representation, and then weightedly fuse it with the network state representation to form an enhanced perception state representation.

[0038] The specific steps of step S2 are as follows:

[0039] Step S2.1: In Network state representation at time step Gaussian noise is gradually added to the top. ,in It follows a normal distribution, and 0 indicates that the mean vector is zero. This represents the identity matrix. A state diffusion transformation is then performed using this matrix, gradually converting it into pure noise. The expression for forward diffusion is:

[0040] ;

[0041] Among them, when At that time, initial noise , Preset noise scheduling parameters;

[0042] Step S2.2: Predict the noise using a diffusion model to obtain the predicted noise. For prediction noise The output result is obtained by state reconstruction. ,make , For representing the virtual network state, the expression for state reconstruction is:

[0043] ;

[0044] in, This is a variance control term. , For the diffusion model, steps 1 to 2 are... Step noise scheduling parameters The product of consecutive terms. Predicting noise. training expression for:

[0045] ;

[0046] Step S2.3: Represent the network state and virtual network state representation By performing weighted combination, an enhanced perception state representation is constructed. Its expression is:

[0047] ;

[0048] ;

[0049] in, This is a parameterized mapping function, specifically the sigmoid function, which restricts the weights to the range [0,1]. This represents the rate of change of network state characteristics. and Preset learnable parameters; These are the weighting coefficients.

[0050] Step S3: Input the enhanced perception state representation into the path selection decision intelligent agent model. The path selection decision intelligent agent model outputs corresponding actions based on the enhanced perception state representation and the DPPO model.

[0051] The working mechanism of the intelligent agent model for path selection decision-making is as follows: it enhances the perception of state representation. As input to the intelligent agent model for path selection decision, the intelligent agent model for path selection decision outputs path selection decision results based on the DPPO algorithm framework. The optional output of the intelligent agent model for path selection decision corresponds to all feasible data transmission paths in the network. Each path is identified as an independent decision option. The output result of the intelligent agent model for path selection decision corresponds to selecting a specific path from the set of feasible paths to carry the data transmission task of the current network service.

[0052] Step S4: Using the diffusion model, based on the enhanced perception state representation, virtual actions are obtained through a progressive denoising process. The output actions of the path selection decision intelligent agent model and the virtual actions are weighted and combined to construct the enhanced actions.

[0053] The specific steps of step S4 are as follows:

[0054] Step S4.1: Initialize a pure noise using a diffusion model. ,in It follows a normal distribution, and 0 indicates that the mean vector is zero. Represent the identity matrix; and represent it according to the enhanced sensing state at the current time. right Stepwise noise reduction, the expression is:

[0055] ;

[0056] in, To preset noise scheduling parameters, , For the diffusion model, steps 1 to 2 are... Step noise scheduling parameters The product of consecutive products; This is a variance control term. ;

[0057] Step S4.2: Use the diffusion model to predict noise and obtain the action space prediction noise. and to The output result is obtained by reverse denoising. ,make , Expression for inverse denoising of virtual actions for:

[0058] ;

[0059] Add noise directly to an action that has already been performed to obtain the network state:

[0060] ;

[0061] in, The network state at time t;

[0062] in, To enhance the action, the specific formula is as follows:

[0063] ;

[0064] ;

[0065] in, These are the weighting coefficients. and Preset parameters for learning; Representing virtual actions, the path selection decision-making agent uses probability. Select virtual actions, to Probabilistic action selection strategy π θ ,in To control the exploration decay rate, its expression is:

[0066] ;

[0067] in, For controllable parameters, This is the decay scaling factor. If it is the first action, the DPPO model is used to select the action first, and the diffusion model is not trained.

[0068] Step S5: Input the enhanced action into the network environment. Based on the enhanced action execution path selection, the network environment constructs a reward function, calculates the reward, and feeds the reward back to the path selection decision agent.

[0069] The specific method for step S5 is as follows:

[0070] Will enhance action Input the network environment and obtain the corresponding augmented sensing state representation of the network environment. In enhancing the representation of perceived states Execute enhanced actions The composite reward function value is then generated. Its expression is:

[0071] ;

[0072] in, Basic rewards, For controllable fusion coefficient, Virtual rewards, virtual rewards The expression is:

[0073] ;

[0074] in, The scaling factor is controllable. This represents the network state at time t+1. This represents the virtual network state at time t+1.

[0075] Step S6: Iteratively train the intelligent agent model for path selection decision until the DPPO model on which it is based converges, obtain the trained intelligent agent model for path selection decision, and apply it to complete path selection in the network environment.

[0076] The training process for the intelligent agent model for path selection decision-making in step S6 is as follows:

[0077] Step S6.1: Initialize the reward discount factor γ, the total number of training rounds T, and the interaction step size. The learning rate λ1 of the policy module and the learning rate λ2 of the value module in the DPPO model are: the batch size of the empirical samples D, the parameters of the policy module θ, and the parameters of the value module v.

[0078] Step S6.2: For each training round, the path selection decision-making intelligent agent model obtains the initial network state representation s0;

[0079] Step S6.3: For each interaction step, obtain... Network state representation at time step Enhance the representation of perceived states , will the action Enhance movement And execute;

[0080] Step S6.4: The intelligent agent model for path selection decision-making calculates the composite reward function value. The enhanced perception state representation of the next moment Storing experience samples ;

[0081] Step S6.5: Based on the enhanced perception state representation of the next time step Update the network state representation;

[0082] Step S6.6: Repeat steps S6.3-S6.5 to complete. The interaction between the intelligent agent for secondary path selection decision-making and the model network architecture;

[0083] Step S6.7: For any empirical sample in a batch, collect an augmented perception state representation. Enhancement actions corresponding to experience sample i The composite reward function value corresponding to experience sample i The enhanced perception state representation of the next moment Update the parameters of the intelligent agent model for path selection decision-making;

[0084] DPPO model reference Figure 1 The method for updating the parameters of the intelligent agent model for path selection decision is as follows:

[0085] Use the advantage function To measure the quality of an action, the following formula is used:

[0086] ;

[0087] in, It is the composite reward function value that the intelligent agent model for path selection decision-making actually receives and uses for learning at each step. This represents the reward discount factor. and The value module is in the parameters The output values ​​below represent the current augmented perception state representation. and the enhanced perception state representation in the next moment The assessed value;

[0088] Gradient clipping method is used to define the objective function. The expression is:

[0089] ;

[0090] in, Represents the mathematical expectation. To limit the input value to a lower limit of The upper limit is Clipping functions within the range, Indicates the current path selection strategy Network state representation Take action below Probability and old path selection strategy in network state representation Take action below The ratio of the probabilities;

[0091] The expression is as follows:

[0092] ;

[0093] in, Let be the probability of the current policy, representing the probability when the parameter is . Under the current policy module, when the environment is in an enhanced perception state, the representation... At that time, the intelligent agent model for path selection decision-making selects the original action. The probability of; The probability of the old policy is represented by the parameter being... Under the old policy network, the same augmented perception state representation Select the same original action below The probability of;

[0094] The expression for updating the policy network parameters in the DPPO model is:

[0095] ;

[0096] in, This indicates the updated strategy module parameters; The pruning factor, usually represented by a small positive number, limits the magnitude of the strategy module parameter updates to a certain range. Within the range;

[0097] In one embodiment, the parameter settings during the DPPO model training process are shown in Table 1 below:

[0098] Table 1. DPPO Model Parameter Settings

[0099]

[0100] Step S6.8: Calculate and update the diffusion model parameters φ and w;

[0101] Step S6.9: Repeat steps S6.7-S6.8 to update the parameters of the intelligent agent model for path selection decision and the diffusion model using D experience samples from the same batch, respectively.

[0102] Step S6.10: Repeat steps S6.1-S6.9 to complete T training rounds and obtain the trained intelligent agent model for path selection decision.

[0103] Figure 2 ,Figure 3 , Figure 4 The graphs show the end-to-end average latency, network throughput, and packet loss rate for five algorithms under different topology scenarios. Figure 2 The results show that as the number of broken links increases from 0 to 3, the latency of all algorithms increases, but the DPPO model shows a latency increase of 25.2%, demonstrating its adaptability to changing topology scenarios. Similarly, Figure 3 and Figure 4 This indicates that the DPPO model also demonstrates significant advantages in network throughput and packet loss rate performance.

[0104] Figure 5 , Figure 6 , Figure 7 The graphs show the end-to-end average latency, network throughput, and packet loss rate for five algorithms in resource-constrained scenarios. Figure 5 The results show that as the effective bandwidth ratio decreases from 100% to 20%, the latency of all algorithms increases. After a 75% bandwidth reduction, the DPPO model's latency only increases by 25.8%, demonstrating its adaptability to changing link states. Similarly, Figure 6 and Figure 7 This indicates that the DPPO model also demonstrates significant advantages in network throughput and packet loss rate performance.

[0105] Figure 8 , Figure 9 The figures show the action selection density and frequency plots for the traditional PPO algorithm and the DPPO model, respectively. The results indicate that the traditional PPO algorithm, constrained by centralized updates, is prone to policy convergence in specific regions of the action space. In contrast, the DPPO model exhibits significant exploration advantages, with its action frequency distribution displaying multiple extrema and its corresponding action density curve exhibiting multimodal distribution characteristics. This suggests that the DPPO model can construct a policy search space with better coverage.

[0106] Figure 10 , Figure 11 , Figure 12 The graphs show the end-to-end average latency, network throughput, and packet loss rate for five algorithms under different traffic intensities. The results show that the DPPO model consistently outperforms other comparative algorithms in terms of average end-to-end latency, network throughput, and packet loss rate under different traffic intensities. At 100 kbps, the DPPO model reduces average end-to-end latency by at least 11.2%, increases network throughput by at least 3.8%, and reduces packet loss rate by at least 4.1%, indicating that the DPPO model possesses stronger reliable decision-making capabilities under different traffic intensities.

[0107] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A path selection method for low-altitude wireless networks enhanced by a diffusion model, characterized in that, A two-stage augmentation decision framework driven by a diffusion model is constructed for the network environment, and the following steps S1-S6 are executed to complete path selection: Step S1: Collect network state information in real time, extract network state information at the current moment and several previous historical moments, and construct a time-series network state representation; Step S2: Use a diffusion model to progressively denoise the network state representation to generate a virtual network state representation, and then weightedly fuse it with the network state representation to form an enhanced perception state representation. Step S3: Input the enhanced perception state representation into the path selection decision intelligent agent model. The path selection decision intelligent agent model outputs corresponding actions based on the enhanced perception state representation and the DPPO model. In step S3, the working mechanism of the intelligent agent model for path selection decision-making is as follows: The enhanced perception state representation is used... As input to the intelligent agent model for path selection decision, the intelligent agent model for path selection decision outputs path selection decision results based on the DPPO algorithm framework. The optional output of the intelligent agent model for path selection decision corresponds to all feasible data transmission paths in the network. Each path is identified as an independent decision option. The output result of the intelligent agent model for path selection decision corresponds to selecting a specific path from the set of feasible paths to carry the data transmission task of the current network service. Step S4: Using the diffusion model, based on the enhanced perception state representation, virtual actions are obtained through a progressive denoising process. The output actions of the path selection decision intelligent agent model and the virtual actions are weighted and combined to construct the enhanced actions. Step S5: Input the enhanced action into the network environment. Based on the enhanced action execution path selection, the network environment constructs a reward function, calculates the reward, and feeds the reward back to the path selection decision-making intelligent agent model. Step S6: Iteratively train the intelligent agent model for path selection decision until the DPPO model on which it is based converges, obtain the trained intelligent agent model for path selection decision, and apply it to complete path selection in the network environment.

2. The diffusion model-enhanced low-altitude wireless network path selection method according to claim 1, characterized in that, The specific method for step S1 is as follows: Collect any Network status at any time Including global business request intensity Available bandwidth resources for each path Path transmission delay and path data packet loss rate The information is as follows: ; in, Dimensions representing network state ; Indicates the number of paths; Based on the network status at each historical moment, construct The network state at time t is represented by the following formula: ; in, express Network state representation at time t, express Time and Time before The network state at a time step.

3. The diffusion model-enhanced low-altitude wireless network path selection method according to claim 2, characterized in that, The specific steps of step S2 are as follows: Step S2.1: In Network state representation at time step Gaussian noise is gradually added to the top. ,in It follows a normal distribution, and 0 indicates that the mean vector is zero. Representing the identity matrix, Gradually becomes pure noise The expression for forward diffusion is: ; Among them, when hour, , Preset noise scheduling parameters; Step S2.2: Predict the noise using a diffusion model to obtain the predicted noise. For prediction noise The output result is obtained by state reconstruction. ,make , For representing the virtual network state, the expression for state reconstruction is: ; in, This is a variance control term. , For the diffusion model, steps 1 to 2 are... Step noise scheduling parameters The product of consecutive products; predicting noise training expression for: ; Step S2.3: Represent the state of the network and virtual network state representation By performing weighted combination, an enhanced perception state representation is constructed. Its expression is: ; ; in, For parameterized mapping functions, This represents the rate of change of network state characteristics. and Preset learnable parameters; These are the weighting coefficients.

4. The diffusion model-enhanced low-altitude wireless network path selection method according to claim 3, characterized in that, The specific steps of step S4 are as follows: Step S4.1: Initialize a pure noise using a diffusion model. ,in It follows a normal distribution, and 0 indicates that the mean vector is zero. Represents the identity matrix; And based on the enhanced perception state representation at the current moment right Stepwise noise reduction, the expression is: ; in, To preset noise scheduling parameters, , For the diffusion model, steps 1 to 2 are... Step noise scheduling parameters The product of consecutive products; This is a variance control term. ; Step S4.2: Use the diffusion model to predict noise and obtain the action space prediction noise. and to The output result is obtained by reverse denoising. ,make , Expression for inverse denoising of virtual actions for: ; Add noise directly to an action that has already been performed to obtain the network state: ; in, The network state at time t; To enhance the action, the specific formula is as follows: ; ; in, These are the weighting coefficients. and Preset parameters for learning; Representing virtual actions, the intelligent agent makes path selection decisions based on probability. Select virtual actions, to Probabilistic action selection strategy π θ ,in To control the exploration decay rate, its expression is: ; in, For controllable parameters, This is the attenuation scaling factor.

5. The diffusion model-enhanced low-altitude wireless network path selection method according to claim 4, characterized in that, The specific method for step S5 is as follows: Will enhance action Input the network architecture and obtain the corresponding augmented sensing state representation of the network architecture. In enhancing the representation of perceived states Execute enhanced actions The composite reward function value is then generated. Its expression is: ; in, Basic rewards, For controllable fusion coefficient, Virtual rewards, virtual rewards The expression is: ; in, The scaling factor is controllable. This represents the network state at time t+1. This represents the virtual network state at time t+1.

6. The diffusion model-enhanced low-altitude wireless network path selection method according to claim 5, characterized in that, The training process for the intelligent agent model for path selection decision-making in step S6 is as follows: Step S6.1: Initialize the reward discount factor γ, the total number of training rounds T, and the interaction step size. The learning rate λ1 of the policy module and the learning rate λ2 of the value module in the DPPO model are: the batch size of the empirical samples D, the parameters of the policy module θ, and the parameters of the value module v. Step S6.2: For each training round, the path selection decision-making intelligent agent model obtains the initial network state representation s0; Step S6.3: For each interaction step, obtain... Network state representation at time step Enhance the representation of perceived states , will the action Enhance movement And execute; Step S6.4: The intelligent agent model for path selection decision-making calculates the composite reward function value. The enhanced perception state representation of the next moment Storing experience samples ; Step S6.5: Based on the enhanced perception state representation of the next time step Update the network state representation; Step S6.6: Repeat steps S6.3-S6.5 to complete. The interaction between the intelligent agent model for secondary path selection decision-making and the network architecture; Step S6.7: For any empirical sample in a batch, collect an augmented perception state representation. Enhancement actions corresponding to experience sample i The composite reward function value corresponding to experience sample i The enhanced perception state representation of the next moment Update the parameters of the intelligent agent for path selection decision-making; Step S6.8: Calculate and update the diffusion model parameters φ and w; Step S6.9: Repeat steps S6.7-S6.8 to update the parameters of the intelligent agent model for path selection decision and the diffusion model using D experience samples from the same batch, respectively. Step S6.10: Repeat steps S6.1-S6.9 to complete T training rounds and obtain the trained intelligent agent model for path selection decision.

Citation Information

Patent Citations

  • Intelligent path optimization method and system based on link state perception enhancement

    CN119011463A

  • Unmanned aerial vehicle cluster dynamic strategy optimization method based on super network and diffusion model

    CN121143469A