Radio map construction method and apparatus, electronic device, and storage medium
By optimizing UAV flight paths through grid partitioning and modeling, and combining conditional diffusion and Markov decision models, the problem of radio map accuracy caused by UAV trajectory constraints was solved, and high-precision radio environment reconstruction was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-01-30
- Publication Date
- 2026-06-02
Smart Images

Figure CN122130054A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of radio mapping, and more particularly to a radio mapping method, apparatus, electronic device, and storage medium. Background Technology
[0002] It should be noted that the above description of the technical background is only for the purpose of providing a clear and complete explanation of the technical solutions of the present invention and facilitating understanding by those skilled in the art. It should not be assumed that the above technical solutions are known to those skilled in the art simply because they have been described in the background section of this invention.
[0003] With the rapid development of information technology, emerging wireless services such as the Internet of Things and autonomous driving are constantly appearing, placing higher demands on the management and optimization of spectrum resources. Radio environment maps, by visualizing the spatial distribution of radio frequency signals, reveal key information in the invisible electromagnetic spectrum, providing crucial support for applications such as dynamic spectrum access, spectrum sharing, and indoor / outdoor positioning. Therefore, the rapid and accurate construction of radio maps is particularly important.
[0004] However, drone trajectories are often restricted by physical obstacles and regulations, resulting in irregular distribution of observation points and blank areas, making it difficult to capture fine textures. How to optimize drone trajectories, achieve real-time and accurate global inference based on a minimum number of data collection points, and reconstruct high-precision radio maps has become an urgent problem to be solved. Summary of the Invention
[0005] In view of the above, the purpose of one or more embodiments of this disclosure is to provide a radio map construction method, apparatus, electronic device and storage medium to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the first aspect of this disclosure provides a method for constructing a radio map, comprising:
[0007] Based on the monitoring granularity requirements, the target area is divided into multiple grid units; The first observation data of the target area is collected based on the first flight path of the UAV; Based on the first observation data and the geographic data, monitoring data corresponding to the multiple grid cells is generated through a conditional diffusion model, and the monitoring data is used to generate a first radio map; Based on the monitoring data, the flight path of the UAV is optimized using a Markov decision model to obtain the second flight path of the UAV. Collect second observation data of the target area according to the second flight path; A second radio map is generated based on the second observation data and the geographic data.
[0008] Optionally, the geographic data includes building distribution data, vegetation distribution data, and signal source data; Based on the first observation data and the geographic data, monitoring data corresponding to the multiple grid cells are generated through a conditional diffusion model, including: Initial noise is obtained by random sampling from a standard normal distribution; The first observation data, the initial noise, and the building distribution data are concatenated and then input into UNet to obtain upsampled features and predicted noise. The upsampling features, the vegetation distribution features obtained based on the vegetation distribution data, and the signal source features obtained based on the signal source data are fused to obtain the fused features. The fused features and the predicted noise are input into the denoising diffusion implicit model to obtain the monitoring data.
[0009] Optionally, the upsampling features, the vegetation distribution features obtained based on the vegetation distribution data, and the signal source features obtained based on the signal source data are fused to obtain fused features, including: The vegetation distribution features and the signal source features are convolved, and a weight matrix is generated by a normalized exponential function. After performing a Hadaman product on the vegetation distribution features, the signal source features, and the weight matrix, the fused features are concatenated with the upsampled features and then convolved to obtain the fused features.
[0010] Optionally, the loss function of the denoising diffusion implicit model includes a loss term for the predicted noise and a physical constraint loss, wherein the physical constraint loss is determined based on the Helmholtz equation.
[0011] Optionally, the state space of the Markov decision model is determined based on the geographic data and the energy consumption of the UAV; the geographic data includes building distribution data, vegetation distribution data, and signal source data.
[0012] Optionally, the reward function of the Markov decision model is determined based on the reconstruction error of the first radio map and the travel distance of the UAV.
[0013] Optionally, the step size of each iteration of the Markov decision model is an integer multiple of the side length of the grid cell.
[0014] A second aspect of this disclosure provides a radio map building apparatus, comprising: The partitioning module is configured to divide the target area into multiple grid cells based on the monitoring granularity requirements; The first acquisition module is configured to acquire first observation data of the target area according to the first flight path of the UAV; The first generation module is configured to generate monitoring data corresponding to the plurality of grid cells based on the first observation data and the geographic data through a conditional diffusion model, and the monitoring data is used to generate a first radio map. The optimization module is configured to optimize the flight path of the UAV based on the monitoring data using a Markov decision model to obtain a second flight path for the UAV. The second acquisition module is configured to acquire second observation data of the target area according to the second flight path; The second generation module is configured to generate a second radio map based on the second observation data and the geographic data.
[0015] A third aspect of this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in the first aspect.
[0016] In a fourth aspect, this disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method as described in the first aspect.
[0017] As can be seen from the above, the radio map construction method, apparatus, electronic device and storage medium provided in this disclosure, through the technical process of first dividing the grid into grid units according to the monitoring granularity and geographic data, then collecting observation data in stages by UAVs, combining the conditional diffusion model to generate full grid monitoring data, and optimizing the flight path by Markov decision model, improves the utilization of UAV resources by gradually and accurately perceiving and efficiently reconstructing the radio environment of the target area, thereby reducing scheduling costs and improving the reconstruction accuracy of the radio map. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in one or more embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only one or more embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating one or more embodiments of a radio map construction method disclosed herein; Figure 2 This is a schematic diagram of the feature extraction process of one or more embodiments of this disclosure; Figure 3 This is a schematic diagram illustrating the process of a radio map construction method according to an embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of a radio map building apparatus according to one or more embodiments of the present disclosure; Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to one or more embodiments of this disclosure. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0021] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar words used in one or more embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0022] refer to Figure 1 The present disclosure discloses a radio map construction method according to one or more embodiments, including the following steps: Step S101: Divide the target area into multiple grid cells based on the monitoring granularity requirements; Step S102: Collect first observation data of the target area according to the first flight path of the UAV; Step S103: Based on the first observation data and the geographic data, generate monitoring data corresponding to the multiple grid cells through a conditional diffusion model, and use the monitoring data to generate a first radio map; Step S104: Based on the monitoring data, optimize the flight path of the UAV using a Markov decision model to obtain the second flight path of the UAV; Step S105: Collect second observation data of the target area according to the second flight path; Step S106: Generate a second radio map based on the second observation data and the geographic data.
[0023] In this embodiment of the disclosure, a grid quantization method can be used to divide a geographic region into a series of discrete grid cells of fixed size, providing a unified spatial reference for subsequent data acquisition, map reconstruction, and route optimization. The set of grid cells can be represented as... ,in Indicates the target region number row and number The grid cell of the column, which corresponds to the radio map. row and number Sub-regions of a column.
[0024] In this embodiment of the disclosure, the observation data collected by the UAV is measured data related to radio signals within the target area. For example, radio parameters such as Received Signal Strength Indicator (RSSI), phase, and polarization direction, as well as associated data such as the spatial location of sampling points and sampling time. The observation data is used to reconstruct a radio map, and all observation data capable of achieving this purpose are within the scope of protection of this disclosure, which does not specifically limit them.
[0025] In developing this disclosure, the inventors discovered that in complex urban environments, the combined effects of tall buildings and dense vegetation create complex non-line-of-sight (NLOS) propagation effects, resulting in fine and highly non-uniform texture features on radio maps. Specifically, buildings completely block radio signals, creating significant shadow fading areas behind them, severely impacting radio coverage and making it difficult to obtain effective radio information from buildings. While vegetation doesn't completely block signals, it causes gradual signal attenuation, resulting in relatively weaker signal strength in areas behind vegetation. Therefore, reducing these texture features is crucial for spectrum management, network planning, and interference prediction.
[0026] However, most related technologies focus on the obstruction effect of buildings, paying insufficient attention to the signal attenuation and scattering caused by vegetation. Considering that in scenarios such as urban-rural transition zones, urban parks, and roadside green belts, large areas of vegetation can significantly alter the spatial distribution of signals, ignoring the obstruction effect of vegetation leads to poor accuracy in reconstructed radio maps.
[0027] Therefore, embodiments of this disclosure propose to reconstruct a first radio map based on first observation data and geographic data through conditional diffusion probability, thereby more accurately simulating the influence of buildings and vegetation on signals and reducing the occlusion effect of areas behind buildings and the occlusion and attenuation effect of areas behind vegetation.
[0028] In some embodiments, the aforementioned geographic data may include building distribution data, vegetation distribution data, and signal source data. Building distribution data represents the distribution information of buildings within the target area, vegetation distribution data represents the distribution information of vegetation within the target area, and signal source data represents the source location matrix obtained by maxima estimation based on sampling point information. In this paper, environmental data can be denoted as... ,in, Represents building distribution data, Representing vegetation distribution data and This represents the signal source data.
[0029] In some embodiments, the process of generating monitoring data corresponding to the plurality of grid cells using a conditional diffusion model may include forward noise addition and reverse noise reduction, thereby reconstructing sparse first observation data into dense monitoring data, where each monitoring data has a corresponding grid cell. In some embodiments, each grid cell has corresponding monitoring data.
[0030] In some embodiments, generating monitoring data corresponding to the plurality of grid cells based on the first observation data and the geographic data through a conditional diffusion model may include: randomly sampling initial noise from a standard normal distribution; concatenating the first observation data, the initial noise, and the building distribution data and inputting them into a UNet to obtain upsampling features and predicted noise; fusing the upsampling features, the vegetation distribution features obtained based on the vegetation distribution data, and the signal source features obtained based on the signal source data to obtain fused features; and inputting the fused features and the predicted noise into a denoising diffusion implicit model to obtain the monitoring data.
[0031] In the process of realizing this disclosure, the inventors discovered that the data in a sample map Based on the corresponding sampling information The environmental information E can follow a certain conditional probability distribution. This can be achieved through a diffusion probability model. Convert to standard normal distribution By constructing a reverse process, a normal distribution can be established to reconstruct the data. The mapping.
[0032] In some embodiments, Gaussian noise can be added to the original data using the forward process of the Denoising Diffusion Probabilistic Models (DDPM). Its expression is: ; in, express Noisy variables at time intervals, express Noisy variables at time intervals, The mean is variance is Gaussian distribution, Represents the identity matrix. Gaussian white noise is added to control the time step t.
[0033] Using the chain rule and Markov properties, exist The joint distribution under the condition can be expressed as: .
[0034] The noise here is the standard Gaussian noise from the sampling. When the total number of time steps is large enough, It approximates a standard normal distribution.
[0035] In other words, after a sufficient number of noise-adding steps, the data will be transformed into noise distributed according to a standard normal distribution. Thus, in the application phase, initial noise can be directly sampled randomly from a normal distribution.
[0036] After generating initial noise, the UNet (U-shaped network) in the conditional diffusion model is used to generate predicted noise and upsampled features.
[0037] In some embodiments, the initial noise, first observation data, and geographic data described above can be concatenated and input into UNet to obtain upsampled features and predicted noise. This data concatenation method can be called an early fusion strategy. The early fusion strategy is simple in structure and effective, enabling the model to learn radio propagation characteristics from observation data and environmental factors (such as buildings, vegetation, etc.) simultaneously during training.
[0038] In developing this disclosure, the inventors discovered that the limitations of this early fusion strategy become apparent when the first observation data is missing or sparse in certain areas. In observation-limited areas, due to the lack of observation signals corresponding to environmental information, the model struggles to accurately learn features from the joint input, potentially leading to unstable predictions. Furthermore, real-world radio propagation is often influenced by multiple environmental factors, and the interaction mechanisms between these factors are complex, making it difficult to model clearly using only a simple early fusion strategy. This further exacerbates the prediction uncertainty of the model in complex scenarios.
[0039] Therefore, in some embodiments, the initial noise, the first observation data, and the building distribution data can be spliced together and used as the input of the U-Net. The network can then gradually learn the influence pattern of buildings on signal propagation. Considering that vegetation is usually distributed in patches and has a continuous attenuation effect on the signal, and that its attenuation characteristics are relatively stable and easy to model, vegetation distribution data and signal source data are introduced in the upsampling stage of the U-Net. The features extracted by the decoder are then fused with the vegetation distribution data and signal source data to more accurately simulate the influence of vegetation on the signal, especially the shading and attenuation effect on the area behind the vegetation.
[0040] Specifically, in some embodiments, it is assumed that This represents the features obtained during the U-Net upsampling stage. This represents the vegetation distribution features and signal source features to be fused. First, the vegetation distribution features and signal source features are processed by two-dimensional convolution, and then the Softmax function is used to obtain the weight matrix W. This weight matrix can be used to capture and enhance key region features, achieving effective focusing on important information. Subsequently, Perform a Hadaman product with the weight matrix W and then... The features are concatenated and then convolved again to obtain the final fused feature H. This process can be represented as: ; ; in The symbol represents the Hadaman product, and Soft represents the softmax function. This represents convolution.
[0041] Fusion feature H after 1 The final network output is obtained by convolution with 1.
[0042] By inputting the fused features and the predicted noise into the denoising diffusion implicit model, the monitoring data can be obtained.
[0043] Figure 2 A schematic diagram of the feature extraction process according to one or more embodiments of the present disclosure is shown.
[0044] like Figure 2 As shown, the rich edge texture features generated by multipath propagation and reflection effects during radio wave transmission are illustrated. In realizing this disclosure, the inventors discovered that preserving these details during fine-grained reconstruction can improve the accuracy of reconstructed radio maps. In other words, enabling neural networks to more effectively learn the spatial characteristics of electromagnetic wave propagation and analyze the physical behavior of electromagnetic fields in the environment is crucial.
[0045] In some embodiments, the Helmholtz equation can be introduced as the physical basis for capturing the propagation of dynamic signals. The Helmholtz equation can be expressed as: .
[0046] in, Wave number represents the spatial propagation characteristics of waves in a medium. Scalar field variables representing the propagation characteristics of dynamic signals.
[0047] Specifically, the two-dimensional Helmholtz equations involve applying a Labras operator to the electric field strength, as stated below: ; in, , This represents two directions in two-dimensional space that are orthogonal.
[0048] In practical applications, since the monitoring data is sampled on a uniform grid, the finite difference method can be used to approximate the differential operator. For a spacing of... Uniform grid, and The second derivative in the direction is approximated using the central difference scheme. The specific details of the sub-region (grid cell) are as follows: ; .
[0049] The above expressions yield discrete gradient operators with similar bounded errors, capable of accurately assessing changes in the orientation field in radio charts. Although these finite-difference approximations can compute the Laplace operator and gradient terms, practical obstacles remain, making… If the actual electromagnetic map data is used Substitute In this context, non-zero terms are typically obtained at the edges of bounded regions such as buildings, e.g., This non-zero term Often referred to as the gradient term of the Helmholtz equation, it corresponds to areas in a radio channel where signals change rapidly due to obstruction by buildings or vegetation. These areas are crucial for the construction of radio maps.
[0050] However, since the gradient term of the real map cannot be directly obtained, this paper extends the noise prediction network of the diffusion model to jointly predict the gradient term corresponding to the Helmholtz operator while predicting the diffusion noise. The gradient prediction branch participates in the update process as a physical guide term during the sampling phase, thereby effectively constraining the spatial structure of the generated results. This enables the model to more accurately reconstruct the detailed changes caused by occlusions such as buildings and vegetation, ultimately generating an electromagnetic map with higher fidelity and physical consistency.
[0051] Therefore, in some embodiments, the model’s loss function may include two key components: a loss term for prediction noise and a loss term for physical constraints.
[0052] In some embodiments, the predicted noise loss can be constructed as the mean square error (MSE) between the predicted noise component and the actual injected noise: = ; The physical constraint loss can be expressed as follows, and the physical constraint loss is determined based on the Helmholtz equation: =
[0053] in, This represents the residual term obtained from calculating the Helmholtz equation using a real electromagnetic map. This represents the residual term predicted by the model. The above expression provides a physical constraint on the learned denoising trajectory. This auxiliary term enhances training stability and facilitates more accurate recovery of the latent representation, especially under a constrained denoising budget. Combining the above two components, the overall loss function can be defined as follows: ; in, and This represents a scalar weight used to balance the influence of the various components. Scalar weights can be adjusted based on empirical performance and specific task constraints.
[0054] In order to implement a data sampling method based on physical information conditions and direct gradient descent during the sampling process, gradient descent can be introduced into the conditional sampling process of DDIM. Therefore, the expression for x can be modified as follows: ; in, , , The noise estimate obtained by the denoising network. Standard Gaussian white noise is used to introduce randomness. Noise scale parameters are used to control the randomness of sampling.
[0055] The sampling process introduced in the above formula is essentially a linear combination of the diffusion model score function and the Helmholtz gradient residual.
[0056] In some embodiments, a conditional diffusion model can be used to predict and generate monitoring data for the entire target area based on geographic data of the target area and first observation data of sparse grid sub-regions (e.g., sparse measurements with a coverage of less than 5%), and to generate a Radio Environment Map (REM). This involves intelligently inferring and reconstructing the monitoring data (e.g., a high-resolution RSSI heatmap) of all grid sub-regions within the target area based on limited UAV flight path measurements and the denoising inverse process of the diffusion model. In this way, only a limited number of UAV scheduling tasks need to be deployed to intelligently infer the UAV monitoring data for the entire target area, reducing deployment costs and system complexity.
[0057] In some embodiments, the above-described conditional diffusion model can be trained based on simulated monitoring data and geographic data.
[0058] In some embodiments, Altair WinProp simulation software can be used to perform high-fidelity simulations of urban radio map distribution to generate simulated monitoring data. Specifically, the simulation is based on an Open Street Map (OSM) data source, which covers building and vegetation information in cities and urban-rural fringe areas such as Beijing, the Yangtze River Delta region, Washington, D.C., and Chicago. Each simulation sample is a 256×256 pixel resolution raster map, including a corresponding top-down view map and an obstacle height distribution map. The transmit power is fixed at 60 dBm, the carrier frequency is 1.80 GHz, and the transmitter is configured as an omnidirectional antenna, supporting multipath propagation and reflection / diffraction effect simulation, thereby generating radio field distribution data covering the map area. The monitoring data is normalized to obtain normalized monitoring data information. This includes data features such as signal strength, phase, and polarization direction. Here, f represents the monitoring data frequency. The Z-score method is used for normalization, converting data of different magnitudes (e.g., RSSI and location coordinates) to the same magnitude to ensure comparability between data. The formula for the Z-score method is expressed as: ; in, It is the indicator item corresponding to the current time slice in the monitoring data. It is the average value of the monitored data indicators. It is the variance of this indicator item in the monitoring data. It is a normalized indicator. (This applies to training and monitoring data.) For each sample, m sample values are randomly selected to obtain the observed sparse data. .
[0059] In realizing this disclosure, the inventors discovered that, considering the varying values of different sampling points for the radio map reconstruction model, this value depends not only on the geographical data and information gain of the sampling points, but also on the impact of UAV flight energy consumption when performing path planning.
[0060] To address the problem of rationally planning UAV sampling paths within a target area, some embodiments use Markov decision processes and multi-agent reinforcement learning methods to determine UAV trajectory optimization strategies within the target area.
[0061] It should be noted that this disclosure allows for the optimization of unfinished trajectories of the drone during flight, as well as the optimization of the trajectory for the next flight after the drone has completed one flight.
[0062] For example, assuming the UAV can sample 100 points, it can either plan the specific locations of subsequent sampling points based on the acquired observation data after sampling the first 20 points, and iteratively optimize the location of each subsequent sampling point until sampling is complete; or it can use data from the previous flight of 100 sampling points to optimize the specific locations of the current 100 sampling points to generate an optimized flight trajectory; or it can combine the above methods, performing route optimization not only before each flight but also during the flight. This disclosure does not limit this approach.
[0063] In some embodiments, the objective of the trajectory optimization strategy is to plan the optimal path for the UAV to move from its initial position to a high information gain sampling point under energy-constrained conditions.
[0064] The goal of a Markov decision process is to minimize the map reconstruction error through iterative sampling and movement, while considering constraints such as obstacles, energy consumption, and velocity limitations. A Markov game can be defined as a tuple. . It is a state space. It is the action space. It is the expected reward for drones. From state to state The transition probability.
[0065] In some embodiments, the state space of the Markov decision model can be determined based on the geographic data and the energy consumption of the UAV; the geographic data includes building distribution data, vegetation distribution data, and signal source data. That is, the state space can be represented as: ; in, It is a multi-channel matrix, specifically expressed as follows: .in, Indicates the number of channels. Indicates the current map size. This indicates the location information of buildings on the map (e.g., whether each location is obscured by a building). For the signal strength distribution information of the current sampling point, This indicates the current location information of the UAV. This indicates the location estimation information of the signal source. The vector features represent the energy consumption information and remaining energy of each drone in its current state.
[0066] In some embodiments, action space This represents the set of all drone actions, including all available locations the drone might choose, excluding obstacle locations. For the... A drone, action This represents the drone's decision at time step t, i.e., the drone's choice of the next sampling location or flight path in its current state. Action It consists of two parts: speed and direction. ,in This means that the drone can only move a certain number of grids in the four directions of north, south, east, and west in each step.
[0067] To adapt to the grid granularity, in some embodiments, the drone's speed can be quantized into several intervals, i.e. Where K is a constant. It is an integer multiple of the grid size. In the time slot. At that time, state drones According to strategy From the action space Choose one action In other words, the step size of each iteration of the Markov decision model is guaranteed to be an integer multiple of the side length of the grid cell.
[0068] S×A→R represents the agent in a given state. The reward expected immediately after taking action.
[0069] In some embodiments, the reward function of the Markov decision model is determined based on the reconstruction error of the first radio map and the distance traveled by the UAV. Therefore, the immediate reward... It consists of two components: energy consumption and current inference error. The calculation is as follows: ; The total energy consumed by a drone is divided into two parts: distance energy consumption and transmission energy consumption, defined as follows: Energy consumed per unit distance The energy consumed to collect a unit of data, The total energy consumption collected by the drone is .in For the first A drone The distance moved at any moment For the first A UAV in The amount of data collected at any given time.
[0070] In some embodiments, calculations are performed based on the UAV's kinematic constraints and obstacle distribution. Considering the UAV's strong spatial characteristics, obstacles, sampling points, and their relationships, a neural network framework can be used to... Features are extracted, and the output is then flattened into a one-dimensional vector. (For UAV information...) The feature extraction layer is converted into a vector using a single-layer linear fully connected layer. These two processed vectors are stacked into a single, unified vector, which serves as the input to subsequent network layers. The subsequent policy and value networks share this feature extraction layer.
[0071] Proximal Policy Optimization (PPO) is used to train agents in environments with multidimensional discrete action spaces due to its stable performance and robust results. The AC method updates its policy, but this can lead to unstable and drastic policy updates due to its reliance on noisy gradient estimations in the environment. To address these issues, Trust Region Policy Optimization (TRPO) was developed, introducing a constraint that limits the policy update range to a small, predefined region to maintain stability. PPO builds upon the TRPO framework but simplifies it by using a sheared alternative objective that limits the policy update rate to prevent significant deviations from the current policy. In PPO, the policy update rule is adjusted to include a shearing mechanism in the objective function, as shown in the following formula: ; ; in, Indicates the state Take action The ratio of the probability of the new strategy to the probability of the old strategy. This is known as the generalized advantage estimate (GAE), and its formula is: ; in, It is a discount factor that measures the importance of future rewards. Here, K is the GAE smoothing parameter, and K is the GAE cutoff step size, used to limit the time range of advantage estimation. Utilizing GAE helps PPO smoothly estimate the advantage of the state and reward sequences, thus providing a more stable and consistent gradient for policy updates.
[0072] Figure 3 The diagram illustrates the process of an embodiment of this disclosure. It can be seen that, with the increase in the number of iterations, a radio environment heatmap covering the entire grid can be accurately reconstructed based on the conditional diffusion model using only a small number of discrete sparse observation points from UAVs within the target area (corresponding to the discrete marker points in the diagram). The different color gradients in the diagram clearly show the spatial distribution differences in signal strength, accurately restoring not only the signal attenuation areas caused by environmental factors such as building obstruction, but also filling in the signal details in unsampled areas, achieving the technical effect of obtaining a high-fidelity radio environment map with low sampling cost.
[0073] The one or more technical solutions provided in this disclosure first divide the target area into grid cells based on the monitoring granularity requirements and geographic data. Then, the UAV collects sparse observation data according to the initial flight path. Combined with the geographic data, the data is used to generate full-grid monitoring data through a conditional diffusion model to construct a first radio map. Subsequently, based on this map, the flight path is optimized using a Markov decision model (the state space includes geographic and UAV energy consumption data, the reward function associates reconstruction error and travel distance, and the iteration step size matches the grid side length). This guides the UAV to collect supplementary observation data and generate a more accurate second radio map. The entire process is a closed loop of "sparse sampling - model reconstruction - path optimization - precise supplementary sampling". While reducing the energy consumption and cost of UAV sampling, it effectively solves the accuracy bottleneck of traditional methods in complex scenarios such as building obstruction and vegetation attenuation, and achieves low-cost and high-fidelity radio environmental perception.
[0074] It is understandable that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.
[0075] It should be noted that the methods of one or more embodiments of this disclosure can be executed by a single device, such as a computer or server. The methods of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to complete the process. In such a distributed scenario, one of these devices may execute only one or more steps of the methods of one or more embodiments of this disclosure, and the multiple devices will interact with each other to complete the method described.
[0076] It should be noted that the above description describes specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims may be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0077] Based on the same inventive concept, corresponding to any of the methods in the above embodiments, this disclosure also provides a radio map construction apparatus. For example... Figure 4 As shown, the device includes: The partitioning module 11 is configured to divide the target area into multiple grid cells based on the monitoring granularity requirements; The first acquisition module 12 is configured to acquire first observation data of the target area according to the first flight path of the UAV; The first generation module 13 is configured to generate monitoring data corresponding to the plurality of grid cells based on the first observation data and the geographic data through a conditional diffusion model, and the monitoring data is used to generate a first radio map. The optimization module 14 is configured to optimize the flight path of the UAV based on the monitoring data using a Markov decision model to obtain a second flight path for the UAV. The second acquisition module 15 is configured to acquire second observation data of the target area according to the second flight path; The second generation module 16 is configured to generate a second radio map based on the second observation data and the geographic data.
[0078] Optionally, the geographic data includes building distribution data, vegetation distribution data, and signal source data; The first generation module 13 is specifically configured as follows: Initial noise is obtained by random sampling from a standard normal distribution; The first observation data, the initial noise, and the building distribution data are concatenated and then input into UNet to obtain upsampled features and predicted noise. The upsampling features, the vegetation distribution features obtained based on the vegetation distribution data, and the signal source features obtained based on the signal source data are fused to obtain the fused features. The fused features and the predicted noise are input into the denoising diffusion implicit model to obtain the monitoring data.
[0079] Optionally, the first generation module 13 is specifically configured as follows: The vegetation distribution features and the signal source features are convolved, and a weight matrix is generated by a normalized exponential function. After performing a Hadaman product on the vegetation distribution features, the signal source features, and the weight matrix, the fused features are concatenated with the upsampled features and then convolved to obtain the fused features.
[0080] Optionally, the loss function of the denoising diffusion implicit model includes a loss term for the predicted noise and a physical constraint loss, wherein the physical constraint loss is determined based on the Helmholtz equation.
[0081] Optionally, the state space of the Markov decision model is determined based on the geographic data and the energy consumption of the UAV; the geographic data includes building distribution data, vegetation distribution data, and signal source data.
[0082] Optionally, the reward function of the Markov decision model is determined based on the reconstruction error of the first radio map and the travel distance of the UAV.
[0083] Optionally, the step size of each iteration of the Markov decision model is an integer multiple of the side length of the grid cell.
[0084] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, when implementing one or more embodiments of this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0085] The apparatus described above is used to implement the corresponding methods in the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0086] Figure 5 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0087] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure.
[0088] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this disclosure are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0089] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0090] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0091] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0092] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this disclosure, and not necessarily all the components shown in the figures.
[0093] The electronic devices described above are used to implement the corresponding methods in the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0094] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0095] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0096] Additionally, to simplify the description and discussion, and to avoid obscuring one or more embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring one or more embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which one or more embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) are set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that one or more embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0097] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0098] This disclosure includes one or more embodiments intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for constructing radio maps, characterized in that, include: Based on the monitoring granularity requirements, the target area is divided into multiple grid units; The first observation data of the target area is collected based on the first flight path of the UAV; Based on the first observation data and geographic data, monitoring data corresponding to the multiple grid cells is generated through a conditional diffusion model, and the monitoring data is used to generate a first radio map. Based on the monitoring data, the flight path of the UAV is optimized using a Markov decision model to obtain the second flight path of the UAV. Collect second observation data of the target area according to the second flight path; A second radio map is generated based on the second observation data and the geographic data.
2. The method according to claim 1, characterized in that, The geographic data includes building distribution data, vegetation distribution data, and signal source data; Based on the first observation data and the geographic data, monitoring data corresponding to the multiple grid cells are generated through a conditional diffusion model, including: Initial noise is obtained by random sampling from a standard normal distribution; The first observation data, the initial noise, and the building distribution data are concatenated and then input into UNet to obtain upsampled features and predicted noise. The upsampling features, the vegetation distribution features obtained based on the vegetation distribution data, and the signal source features obtained based on the signal source data are fused to obtain the fused features. The fused features and the predicted noise are input into the denoising diffusion implicit model to obtain the monitoring data.
3. The method according to claim 2, characterized in that, The upsampling features, the vegetation distribution features obtained based on the vegetation distribution data, and the signal source features obtained based on the signal source data are fused to obtain fused features, including: The vegetation distribution features and the signal source features are convolved, and a weight matrix is generated by a normalized exponential function. After performing a Hadaman product on the vegetation distribution features, the signal source features, and the weight matrix, the fused features are concatenated with the upsampled features and then convolved to obtain the fused features.
4. The method according to claim 2, characterized in that, The loss function of the denoising diffusion implicit model includes a loss term for the predicted noise and a physical constraint loss, the physical constraint loss being determined based on the Helmholtz equation.
5. The method according to claim 1, characterized in that, The state space of the Markov decision model is determined based on the geographic data and the energy consumption of the UAV; the geographic data includes building distribution data, vegetation distribution data, and signal source data.
6. The method according to claim 1, characterized in that, The reward function of the Markov decision model is determined based on the reconstruction error of the first radio map and the travel distance of the UAV.
7. The method according to claim 1, characterized in that, The step size of each iteration of the Markov decision model is an integer multiple of the side length of the grid cell.
8. A radio map building device, characterized in that, include: The partitioning module is configured to divide the target area into multiple grid cells based on the monitoring granularity requirements; The first acquisition module is configured to acquire first observation data of the target area according to the first flight path of the UAV; The first generation module is configured to generate monitoring data corresponding to the plurality of grid cells based on the first observation data and geographic data through a conditional diffusion model, and the monitoring data is used to generate a first radio map. The optimization module is configured to optimize the flight path of the UAV based on the monitoring data using a Markov decision model to obtain a second flight path for the UAV. The second acquisition module is configured to acquire second observation data of the target area according to the second flight path; The second generation module is configured to generate a second radio map based on the second observation data and the geographic data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executed by the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.