Method for determining optimal spatiotemporal sampling granularity of mobile crowd sensing based on online learning

CN115988646BActive Publication Date: 2026-09-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211468081.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2026-09-11
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

本发明解决了在无先验知识情况下的区域最佳采样时空粒度确认问题,便于以最能反映区域分布的采样策略高效地执行感知任务

Benefits of technology

[0045] This invention solves the problem of determining the optimal spatiotemporal granularity of regional sampling in the absence of prior knowledge, facilitating efficient execution of perception tasks with a sampling strategy that best reflects the regional distribution. Faced with an unknown sampling environment, a multi-armed slot machine-like mechanism is employed to select the current optimal sampling strategy by balancing exploration and utilization, thus refining the estimation of actual rewards. Through multiple iterations, the sampling strategy with the highest reward is found, achieving the goal of optimal efficiency—that is, the sampling strategy that best restores the data distribution at the lowest cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115988646B_ABST
    Figure CN115988646B_ABST
Patent Text Reader

Abstract

This invention discloses a method for determining the optimal spatiotemporal sampling granularity for mobile crowdsourcing sensing based on online learning. By constructing a reward function that includes the distribution of sampled data and sampling costs, the sampling results are transformed into rewards. These rewards are modeled as a generalized linear model. Utilizing a method of sensing the data distribution in unknown regions using an optimal arm, the model is iteratively refined through multiple rounds to find the optimal spatiotemporal sampling granularity while minimizing sensing costs. This invention solves the problem of determining the optimal spatiotemporal sampling granularity for a region in the absence of prior knowledge, facilitating efficient execution of sensing tasks with a sampling strategy that best reflects the region's distribution. Faced with an unknown sampling environment, a multi-armed machine-like mechanism is employed to balance exploration and utilization, selecting the current optimal sampling strategy and refining the estimation of actual rewards. Through multiple iterations, the sampling strategy with the highest reward is found, achieving the goal of optimal efficiency—that is, the sampling strategy that best reflects the data distribution at the lowest cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) technology, specifically relating to a method for determining the optimal spatiotemporal sampling granularity for mobile swarm intelligence perception. Background Technology

[0002] In recent years, with the widespread adoption of mobile devices, continuous crowdsourcing sensing has been widely applied to data detection in real-world areas. With the development of 5G technology, mobile devices can rapidly collect and share data in the Internet of Things (IoT) environment. The rise of crowdsourcing sensing has played a significant role in areas such as urban temperature measurement, noise detection, urban traffic congestion detection, and smart agriculture. A common approach to continuous crowdsourcing sensing of a region is to divide the region into spatiotemporal granularities and then sample according to these granularities. The spatiotemporal granularity division is typically driven by historical data, meaning that the spatiotemporal sampling granularity is set based on prior knowledge and the spatiotemporal correlations of the data. In compressed sensing, the sampling granularity is also usually set based on historical or prior data, utilizing the spatiotemporal correlations of the data to set an appropriate sampling granularity and reduce sensing costs. Therefore, historical data plays a crucial role in the quality of sensing data.

[0003] However, in many real-world scenarios, emergencies occur in unknown areas, such as disasters. In these situations, the environment changes rapidly, rendering previously held historical experience in that environment ineffective. Furthermore, in cases like flood disaster monitoring, it's crucial to quickly collect data on specific physical observations to determine the distribution of these observations within the area, facilitating rescue operations based on the findings. In such cases, due to the rapid changes in data, the sampling granularity determined by historical data is clearly unsuitable for the current environment. Therefore, it's necessary to find an online learning method that adjusts the sampling scheme based on real-time feedback.

[0004] Multi-armed sensors arise from the trade-off between utilizing known information and exploring unknown information, seeking to maximize the benefits. In recent years, multi-armed sensors have been increasingly applied in the field of mobile crowdsourcing sensing, becoming a feasible solution for finding the most efficient sensing strategy in the sensing field.

[0005] Current data collection methods often rely on prior knowledge to obtain the spatiotemporal granularity of data sampling that best reflects the distribution of regional data. However, in the face of completely unknown environments or rapidly changing environments, prior knowledge is either absent or ineffective. Summary of the Invention

[0006] To overcome the shortcomings of existing technologies, this invention provides a method for determining the optimal spatiotemporal sampling granularity for mobile crowdsourcing sensing based on online learning. By constructing a reward function that includes the distribution of sampled data and sampling costs, the sampling results are transformed into rewards, which are then modeled as a generalized linear model. Utilizing a method of sensing the data distribution in unknown regions with optimal arms, the model is iteratively refined through multiple rounds to find the optimal spatiotemporal sampling granularity while minimizing sensing costs. This invention solves the problem of determining the optimal spatiotemporal sampling granularity for a region in the absence of prior knowledge, facilitating efficient execution of sensing tasks with a sampling strategy that best reflects the regional distribution. Faced with an unknown sampling environment, a multi-armed machine-like mechanism is employed to balance exploration and utilization, selecting the current optimal sampling strategy and refining the estimation of actual rewards. Through multiple iterations, the sampling strategy with the highest reward is found, achieving the goal of optimal efficiency—that is, the sampling strategy that best reproduces the data distribution at the lowest cost.

[0007] The technical solution adopted by this invention to solve its technical problem includes the following steps:

[0008] Step 1: Assume an unknown region A m×n The sampled area is divided into m×n finest spatial granularity regions, which represent the sampling area of ​​the UAV. Let the candidate strategy set for sampling be π = {(Area1, Time1), (Area2, Time2), ..., (Area...}. K Time K There are K sampling strategies in the strategy set, and the sampling strategy π i (i∈[K]) represents the spatial policy Area i ——Unknown region A m×n The division into several sub-regions, and the time strategy. i —The sampling frequency corresponding to each sub-region; During sampling, the unknown region is divided into several sampling sub-regions, and each sub-region is treated as a whole for sampling. The sub-regions are irregular and consist of several continuous regions with the finest spatial granularity; In the unknown region, there are base stations and multiple drones. Multiple drones sample data in the unknown region, and the data distribution of the unknown region is obtained based on the data sampling results of the drones.

[0009] The UAV sampling strategy is to randomly select the finest spatial granularity region in each spatial sub-region for sampling, and use the sampling result of this finest spatial granularity region as the sampling result of each finest spatial granularity region in the spatial sub-region;

[0010] Suppose that the feature vector x corresponding to the UAV sampling strategy, the unknown reward parameter vector θ, and the reward are bounded, satisfying: ||x||² ≤ L, ||θ||² ≤ S, r ≤ R; μ is the reward function modeled as a generalized linear model, and x tLet represent the feature vector corresponding to the sampling strategy in the t-th round; θ represents the estimated reward value for the sampling strategy selected in round t. t Let θ be the estimated value of the t-th round, and let d be the spatiotemporal feature dimension.

[0011] Step 2: Let the spacetime dimension be d;

[0012] Step 2-1: Using the multi-armed slot machine algorithm, during the first E iterations, in the candidate sampling strategy set π = {(Area1,Time1),(Area2,Time2),......(Area... K Time K Random selection strategy in )} Sampling is performed, M E The matrix formed by the eigenvectors of the sampling strategy. This represents the feature vector of the sampling strategy selected in the vth round of the first E rounds; and the selection strategy satisfies the information matrix M. E The smallest eigenvalue λ0 > 0;

[0013] Step 2-2: Starting from the E+1th iteration, estimate the current optimal sampling strategy based on all the information obtained so far. Most likely optimal sampling strategy And sampling strategies that can obtain the most unknown information And adopt strategies and strategy For A respectively m×n Distribute drones to conduct actual sampling and obtain Sampling results S in the finest spatial granularity region t ={p t,1 ,p t,2 ,......,p t,mn} and estimated area data volume N t ,as well as The sampling result p t ;

[0014] In round t≥E+1, the expected reward of each sampling strategy is obtained based on the currently known information, and the strategy with the highest expected reward is found.

[0015]

[0016] In the formula, This represents the optimal strategy selected in round t. The corresponding feature vector;

[0017] The strategy most likely to yield the highest reward

[0018]

[0019] △(j t i t )for and Estimated sampled reward difference:

[0020]

[0021]

[0022] in, C t This is a time-varying parameter used to scale the sampling strategy in the model. The estimated error threshold width; Estimate the feature vector corresponding to the most likely optimal sampling strategy for round t; α is an adjustable parameter. definition k μ and c μ Here are the parameters related to the reward function; c and c' are the parameters in [c μ ,k μ A constant that takes any value within the range.

[0023] And strategies that can obtain the most information about unknown areas

[0024]

[0025] in, M t-1 Let a represent the information matrix consisting of the eigenvector of the optimal sampling strategy, the eigenvector of the most likely optimal sampling strategy, and the eigenvectors used in the first t-1 iterations, respectively; t The index corresponding to the sampling strategy containing the most unknown information;

[0026] Step 3: Model the reward function related to regional data distribution and sampling cost as a generalized linear model, based on the currently obtained strategy. The true sampling data distribution updates the unknown parameter θ in the generalized linear model that determines the estimated reward value, and updates the reward estimate for all sampling strategies;

[0027] Assuming the reward function follows a Poisson distribution, the estimated reward function, obtained from the generalized linear model, is μ(z) = 1 / (1+e^(-1 / 2)). -z The maximum likelihood estimation is used to determine the parameter vector θ that determines the reward in round t. t The estimate;

[0028] Step 4: Calculate the estimated optimal sampling strategy And estimating the asymptotic optimal strategy If the reduction in the reward gap threshold is less than the set error threshold, then the best sampling strategy estimated from the currently obtained information is determined to be the best sampling strategy, and the iteration stops; otherwise, the next round of drone sampling continues.

[0029] If for step 2-2 and The estimated reward value is less than the set threshold The optimal sampling strategy can be determined by the probability of δ. That is, P[μ(θ)] T x * )-μ(θ T x)≤ε]≥δ, where δ takes the value of:

[0030]

[0031] In the formula, ζ is a user-defined parameter; μ(θ) T x i ) is the strategy π i The reward function value, x i For strategy π i eigenvectors;

[0032] When the decrease in β(i,j) between iterations is less than a certain value, it is determined that the error threshold no longer changes and the stopping condition is met, i.e., β(i t-1 ,j t-1 )-β(i t ,j t If the stopping condition is not met, it is determined that the data distribution collected by the UAV does not reflect the true data distribution of the region, and the process continues for the unknown region A. m×n Distribute drones for sampling, i.e., use a sampling strategy. Sampling was conducted to continue obtaining environmental information;

[0033] Step 5: Estimate the optimal sampling strategy Sampling results Calculate the true regional reward value r, which takes into account both data quality and sampling cost. t ;

[0034] For step 1 Sampling results S t Calculate the actual reward value:

[0035]

[0036]

[0037] In the formula, Representation Strategy Sampling cost, A user-defined function representing the cost. The impact on the reward function value is as follows: P represents the regional distribution obtained by sampling the data during the iteration process as a summary of historical experience based on the current information; KL(.) represents the relative entropy, and JS(.) represents the JS divergence.

[0038] Step 6: Sampling strategy that contains the most unknown information in this iteration Sampling results And the estimated area data volume N t Based on historical experience, update the current historical weighted average sample data distribution P = {p1, p2, ..., p...} mn},in

[0039] Furthermore, the drone sampling process is as follows:

[0040] Observe the unknown region A m×n The distribution of the physical quantities, and the unknown region A. m×n It is divided into m×n finest spatial granularity regions. The spatial strategy is a combination of the finest spatial granularity regions, while the temporal strategy is to select different sampling frequencies on the time axis.

[0041] The spatiotemporal partitioning strategy set is defined as: π = {(Area1, Time1), (Area2, Time2), ..., (Area...} K Time K Area i and Time i Representing sampling strategy π i The spatial and temporal partitioning methods for sampling; in Each region is composed of continuously varying finest granular regions; time granularity This indicates the sampling frequency for each spatial division;

[0042] Let X = {x1, x2, ..., x} K}(x i =(z i ,t i (), i = 1, 2, ..., K) represent the spatiotemporal granularity feature set of the policy set π, where z i Represents spatial feature labels, t i For time feature labels; the most suitable region A needs to be selected from π. m×nThe sampling method and spatial partitioning method are continuous and irregular, and the given spatial granularity is a combination of the finest granular regions;

[0043] Set a strategy The sampling results are In the sub-regions respectively Within the region, the finest spatial granularity region is randomly selected as a representative for sampling, and the sampling result is used as the region. The sampling is therefore represented as the probability distribution of the regional data.

[0044] The beneficial effects of this invention are as follows:

[0045] This invention solves the problem of determining the optimal spatiotemporal granularity of regional sampling in the absence of prior knowledge, facilitating efficient execution of perception tasks with a sampling strategy that best reflects the regional distribution. Faced with an unknown sampling environment, a multi-armed slot machine-like mechanism is employed to select the current optimal sampling strategy by balancing exploration and utilization, thus refining the estimation of actual rewards. Through multiple iterations, the sampling strategy with the highest reward is found, achieving the goal of optimal efficiency—that is, the sampling strategy that best restores the data distribution at the lowest cost. Attached Figure Description

[0046] Figure 1 This is a flowchart of the method of the present invention.

[0047] Figure 2 This is the spatial division strategy for the sampling area in Embodiment 4×4 of the present invention.

[0048] Figure 3 This is the time division strategy for the sampling region in Embodiment 4×4 of the present invention. Detailed Implementation

[0049] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0050] In real-world scenarios, when facing unknown areas or rapidly changing environments during emergencies, historical data for these areas is often lacking. However, due to cost and task considerations, it is necessary to quickly collect data on certain physical observations to determine the distribution of observations within the area. For example, in the event of a fire or flood, it is crucial to quickly identify the most severely affected location. In such cases, because the distribution of the observed data changes rapidly, data distribution determined based on historical data is clearly unsuitable for the current environment. To address this practical problem, in the absence of prior data, a suitable spatiotemporal sampling granularity division method is sought. This method can best reflect the regional distribution of data while also considering sampling costs, finding a comprehensive optimal sampling scheme from the perspectives of real-time performance, economy, and accuracy. The goal is to establish an understanding of unknown areas through multiple iterative real-time feedback. Due to limited perception budgets and the time sensitivity of perception targets, the number of iterations should be minimized. This invention provides a spatiotemporal sampling granularity confirmation mechanism for unknown areas, finding a relatively ideal spatiotemporal sampling granularity while minimizing perception budget.

[0051] The objective of this invention can be achieved by adopting the following technical solutions:

[0052] A method for determining the granularity of data sampling based on online learning and mobile crowd sensing includes the following steps:

[0053] A method for determining the granularity of data sampling based on online learning and mobile crowd sensing includes the following steps:

[0054] Step 1: Assume an unknown environment A m×n The area is divided into m×n finest spatial granularity regions, where each region represents the area sampled by the UAV. The sampling strategy is represented as a spatial strategy, i.e., region A. m×n The system is divided into several spatial sub-regions, and a temporal strategy is defined, i.e., the sampling frequency corresponding to each sub-region. The spatial sub-regions are irregular and consist of several consecutive finest-grained regions. The UAV sampling strategy is represented by randomly selecting a finest-grained spatial sampling region within each spatial sub-region for sampling, and using the sampling result of this region as the sampling result of each finest-grained region within the corresponding spatial sub-region.

[0055] Step 2: Let the spacetime dimension be d;

[0056] Step 2-1: Given parameter E, during the first E iterations, in the candidate sampling strategy set π = {(Area1,Time1),(Area2,Time2),......(Area...} K Time KA random sampling strategy is selected in the process to assign drones to perform preliminary sampling in the unknown environment, thereby establishing a preliminary understanding of the unknown environment.

[0057] Step 2-2: Starting from the E+1th iteration, estimate the current optimal sampling strategy based on all the information obtained so far. Most likely optimal sampling strategy And sampling strategies that can obtain the most unknown information And adopt strategies and strategy For A respectively m×n Distribute drones to conduct actual sampling and obtain Sampling results S in the finest granular sampling region t ={p t,1 ,p t,2 ,......,p t,mn} and estimated area data volume N t ,as well as The sampling result p t .

[0058] Step 3: Model the reward function related to regional data distribution and sampling cost as a generalized linear model, based on the current strategy. The true sampling data distribution updates the unknown parameter θ in the generalized linear model that determines the estimated reward value, and updates the reward estimate for all sampling strategies.

[0059] Step 4: Calculate the estimated optimal sampling strategy And estimating the asymptotic optimal strategy If the reduction in the reward gap threshold is less than the set error threshold, we have high confidence that the optimal sampling strategy we estimated based on the currently obtained information is the optimal sampling strategy, and the iteration stops; otherwise, we continue with the next round of drone sampling.

[0060] Step 5: Estimate the optimal sampling strategy Sampling results S t Calculate the true regional reward value r, which takes into account both data quality and sampling cost. t .

[0061] Step 6: Sampling strategy that contains the most unknown information in this iteration Sampling results Based on historical experience, update the current historical weighted average sample data distribution P = {p1, p2, ..., p...} mn},in

[0062] Step 4 introduces the Chernov bound to measure whether the asymptotically optimal sampling strategy reliably reflects the data distribution of the unknown environment. Specifically, when the difference in reward between the optimal sampling strategy and the near-optimal strategy is less than a certain value ε, there is a confidence level of ζ that the near-optimal strategy is the optimal one. The Chernov bound is used to regulate the values ​​of ε and δ. When the optimal strategy... and near-optimal strategy Reward function gap threshold Where ζ is a given parameter, and δ takes the following values:

[0063]

[0064] Where ε takes the value of For strategy The reward function value, For strategy The feature vector. However, in practical applications, considering the high sampling cost and rapidly changing environment, we should set a fast stopping condition, that is, when β(i t ,j t When the decrease between iterations is less than a certain value, the error threshold is considered to no longer change and the stopping condition is met.

[0065] Step 5 introduces JS divergence to measure data quality and participates in the design of the reward function, as follows: Assuming the reward function follows an exponential family distribution, a generalized linear model (GLM) is applied to fit the vector θ that determines the reward function, given a policy set π = {π1, π2, ..., π...}. K In the t-th iteration, the sampling strategy is selected. The sampling results include the sampling target distribution S t ={p t,1 ,p t,2 ,......,p t,mn} and estimated area data volume N t This paper constructs an understanding of the reward function under the current circumstances. The reward function considers two factors: the quality of the sampled data and the cost. A JavaScript divergence metric is introduced to measure data quality, using the JavaScript divergence between the sampled distribution and the historical weighted average distribution as a representation of the sampled data quality. The cost is the sampling strategy. Pre-set fixed costs Rewards are positively correlated with data quality and negatively correlated with cost. The reward function is designed as follows:

[0066]

[0067] After the current iteration ends, the sampling results that provide the most information are incorporated into the historical weighted average distribution P = {p1, p2, ..., p...} mn} Specific implementation examples:

[0069] This invention designs a method for determining the optimal spatiotemporal sampling granularity of mobile crowd sensing based on online learning, see reference. Figure 1 As shown, the specific steps of the present invention are as follows:

[0070] Assume an unknown environment A m×n The system is divided into m×n finest spatial granularity regions, each representing the area sampled by the drones. In an unknown region, there are base stations and multiple drones. Multiple drones collect data within this region, and the data distribution of the region is determined based on the collected drone data. The sampling strategy is for region A. m×n The system is divided into several spatial sub-regions and their corresponding sampling frequencies. The UAV sampling strategy is represented by randomly selecting the finest spatial sampling granularity region within each spatial sub-region for sampling, and using the sampling result of this region as the sampling result of each finest granularity region within the corresponding spatial sub-region.

[0071] 1. For the relevant parameter settings, assume that the feature vector x corresponding to the sampling strategy, the unknown reward parameter vector θ, and the reward are bounded, satisfying: ||x||²≤L, ||θ||²≤S, r≤R. μ is the reward function modeled as a generalized linear model. Let d represent the estimated reward value of the sampling strategy selected in round t, with a spatiotemporal feature dimension of d.

[0072] 1-1. First, in the first E rounds, a random strategy is selected. Sampling is performed, and the results constitute a preliminary understanding of the distribution of data in the unknown environment. M E The matrix formed by the feature vectors of the sampling strategy used to initially understand the region information. And the chosen strategy satisfies the information matrix M E The smallest eigenvalue λ0 > 0.

[0073] 1-2. In round t≥E+1, based on the estimated currently known information, obtain the expected reward for each sampling strategy, and find the strategy with the highest expected reward.

[0074]

[0075] The strategy most likely to yield the highest reward

[0076]

[0077] △(j t i t )for and Estimated sampled reward difference:

[0078]

[0079]

[0080] in, C t It is a time-varying parameter that can be used to scale the model's application to the sampling strategy. The estimated error threshold width. α is an adjustable parameter. definition k μ and c μ For the parameters related to the reward function, c and c' take values ​​in the range [c μ ,k μ The parameters between ].

[0081] And strategies that can obtain the most information about unknown areas :

[0082]

[0083] 2. Assuming the reward function follows a Poisson distribution, according to the generalized linear model, the estimated reward function is μ(z) = 1 / (1+e^(z-1)). -z The maximum likelihood estimation is used to determine the reward vector θ in round t. t The estimate.

[0084] 3. If for steps 1-2 and The estimated reward value is less than the set threshold There is a probability of δ that the optimal sampling strategy has been found. That is, P[|△(j)] t i t )+β(i t ,j t )|≤ε]≥δ, where δ takes the value of:

[0085]

[0086] In practical applications, considering the high sampling cost and rapidly changing environment, a fast stopping condition is used to reduce the number of samples, when β(i t ,j t When the decrease between iterations is less than a certain value, the error threshold is considered to no longer change and the stopping condition is reached, i.e., β(i t-1 ,j t-1 )-β(i t ,j tIf the rapid stopping condition is not met, it is considered that the data distribution collected by the UAV is insufficient to reflect the true data distribution of the area, and further investigation of the unknown area A is required. m×n Deploying drones for sampling enhances our understanding of unknown areas; this is a sampling strategy. Sampling was conducted to continue obtaining environmental information.

[0087] 4. Regarding step 1 Calculate the true reward value based on the sampling results:

[0088]

[0089]

[0090] In the formula, KL(.) represents KL divergence, i.e., relative entropy, and JS(.) represents JS divergence.

[0091] 5. The sampling strategy that contains the most unknown information in this iteration Sampling results As historical experience, update the current historical weighted average sampling data distribution P = {p1, p2, ..., p...}, which is accumulated from historical experience through multiple samplings. mn},in

[0092] The drone sampling process is as follows:

[0093] Observe the unknown region A m×n Regarding the distribution of a certain physical quantity in region A, we have no knowledge of its distribution within that region. m×n ={A1,A2,......,A mn The spatiotemporal partitioning strategy set is composed of the finest-grained division of m×n regions. The spatial strategy is a combination of these fine-grained regions, while the temporal strategy involves selecting different sampling frequencies along the time axis. The spatiotemporal partitioning strategy set is defined as: π = {(Area1, Time1), (Area2, Time2), ..., (Area...} K Time K Area i and Time i Representing sampling strategy π i The spatial and temporal partitioning methods for sampling. in Each is composed of continuously varying regions of the finest granularity. (Time granularity) This represents the sampling frequency for each spatial partition. Let X = {x1, x2, ..., x...} K}(x i =(z i,t i (), i = 1, 2, ..., K) represent the spatiotemporal granularity feature set of the policy set π, where z i Represents spatial feature labels, t i This is a time feature label. We need to select the most suitable region A from π. m×n The sampling method here allows for spatial partitioning that can be continuous or irregular, with the given spatial granularity being a combination of the finest-grained regions. Let the strategy be... The sampling results are Among them we are Within the region, the finest-grained region is randomly selected as a representative for sampling, and the sampling result is used as the region. The sampling is therefore represented as the probability distribution of the regional data.

Claims

1.A method for determining optimal spatiotemporal sampling granularity for mobile crowd sensing based on online learning, characterized in that, Includes the following steps: Step 1: Assume an unknown region Divided into m × n Let there be a minimum spatial granularity region, which represents the area sampled by the UAV; let the candidate strategy set for sampling be... The strategy is concentrated. K Each sampling strategy, sampling strategy ( ) represents spatial strategy ——Unknown area The division into several sub-regions and the time strategy —The sampling frequency corresponding to each sub-region; During sampling, the unknown region is divided into several sampling sub-regions, and each sub-region is treated as a whole for sampling. The sub-regions are irregular and consist of several continuous regions with the finest spatial granularity; In the unknown region, there are base stations and multiple drones. Multiple drones sample data in the unknown region, and the data distribution of the unknown region is obtained based on the data sampling results of the drones. The UAV sampling strategy is to randomly select the finest spatial granularity region in each spatial sub-region for sampling, and use the sampling result of this finest spatial granularity region as the sampling result of each finest spatial granularity region in the spatial sub-region; Let the feature vector corresponding to the UAV sampling strategy be... Unknown reward parameter vector And the rewards are bounded, satisfying: , , ; It is a reward function modeled as a generalized linear model. Let represent the feature vector corresponding to the sampling strategy in the t-th round; Representing the t The estimated reward value of the round-selection sampling strategy. For round t The estimated value, with a spatiotemporal feature dimension of d ; Step 2: Let the spacetime dimension be... d ; Step 2-1: Using the multi-armed slot machine algorithm, during the first E iterations, in the candidate sampling strategy set... Random selection strategy Perform sampling. The matrix formed by the eigenvectors of the sampling strategy. , Indicates the preceding E Wheel of Life v The sampling strategy feature vector selected in each round; and the selection strategy satisfies the information matrix. The smallest eigenvalue ; Step 2-2: Starting from the E+1th iteration, estimate the current optimal sampling strategy based on all the information obtained so far. The most likely optimal sampling strategy And sampling strategies that can obtain the most unknown information And adopt strategies and strategy To each Distribute drones to conduct actual sampling and obtain Sampling results in the finest spatial granularity region and estimated regional data volume ,as well as Sampling results ; In the In each round, the expected reward of each sampling strategy is estimated based on the currently known information, and the strategy with the highest expected reward is found. : In the formula, Indicates the first t The best strategy selected in rounds The corresponding feature vector; The strategy most likely to yield the highest reward : for and Estimated sampled reward difference: in, , Let be a time-varying parameter, where t represents the iteration round and d represents the dimension of the policy vector. The confidence level representing the error threshold is used to scale the model to the sampling strategy. The estimated error threshold width; Estimate the feature vector corresponding to the sampling strategy that is most likely to be optimal for round t; It is an adjustable parameter. ,definition , and These are parameters related to the reward function; and In order to be in A constant that takes any value within the range. ; This represents the decision matrix obtained through cold start rounds. The smallest eigenvalue; And strategies that can obtain the most information about unknown areas : in, , , They represent the first time. t -1 rounds of iteration, the feature vector of the best sampling strategy, the estimated feature vector of the most likely best sampling strategy, and the previous... t -1 rounds use an information matrix composed of eigenvectors; The index corresponding to the sampling strategy containing the most unknown information; Step 3: Model the reward function related to regional data distribution and sampling cost as a generalized linear model, based on the currently obtained strategy. The true sampling data distribution updates the unknown parameter θ in the generalized linear model that determines the estimated reward value, and updates the reward estimate for all sampling strategies; Assuming the reward function follows a Poisson distribution, the estimated reward function, obtained from the generalized linear model, is: The maximum likelihood estimation was used to complete the first step. t The parameter vector that determines the reward for each wheel pair The estimate; Step 4: Calculate the estimated optimal sampling strategy And estimating the asymptotic optimal strategy If the reduction in the reward gap threshold is less than the set error threshold, then the best sampling strategy estimated from the currently obtained information is determined to be the best sampling strategy, and the iteration stops; otherwise, the next round of drone sampling continues. If for step 2-2 and The estimated reward value is less than the set threshold There will be The optimal sampling strategy is obtained by determining the probability. Right now , The value can be: In the formula, These are custom parameters; For strategy The reward function value, For strategy eigenvectors; when When the decrease between iterations is less than a certain value, the error threshold is determined to no longer change and the stopping condition is met. If the stopping condition is not met, it is determined that the data distribution collected by the drone does not reflect the true data distribution of the area, and the process continues for the unknown area. Distribute drones for sampling, i.e., use a sampling strategy. Sampling was conducted to continue obtaining environmental information; Step 5: Estimate the optimal sampling strategy Sampling results Calculate the true regional reward value that takes into account both data quality and sampling cost. ; For step 1 Sampling results Calculate the actual reward value: In the formula, Representation Strategy Sampling cost, A user-defined function representing the cost. The impact on the value of the reward function. This represents the regional distribution obtained by sampling data during the iteration process, which is a summary of historical experience and based on current information. Represents relative entropy. Indicates JS divergence; Step 6: Sampling strategy that contains the most unknown information in this iteration Sampling results and estimated regional data volume Based on historical experience, update the current historical weighted average sample data distribution. ,in . The specific drone sampling process is as follows: Observing unknown areas Distribution of physical quantities, unknown region Divided into m × n The finest spatial granularity region is the combination of the finest spatial granularity regions, while the temporal strategy is to select different sampling frequencies on the time axis. The set of spatiotemporal partitioning strategies is defined as follows: , and They represent sampling strategies respectively. The spatial and temporal partitioning methods for sampling; ,in Each is composed of continuously varying regions of the finest granularity; time granularity This indicates the sampling frequency for each spatial division; Let's assume Representation strategy set The spatiotemporal granularity feature set, in which Indicates spatial feature labels, For time feature labels; need to be derived from Choose the most suitable area The sampling method and spatial partitioning method are continuous and irregular, and the given spatial granularity is a combination of the finest granular regions; Set a strategy The sampling results are In the sub-regions, respectively Within the region, the finest spatial granularity region is randomly selected as a representative for sampling, and the sampling result is used as the region. The sampling is therefore represented as the probability distribution of the regional data. .