Sample dynamic augmentation method, system, device and medium based on environmental covariates
By using a dynamic supplementation method based on environmental covariates, similarity and uncertainty values are calculated, and supplementary samples are selected hierarchically. This solves the problem of insufficient sample set reusability and improves the evaluation accuracy and sample utilization efficiency of map products.
Patent Information
- Application Number
- CN202610037819.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2046-01-13
AI Technical Summary
Existing sampling methods are difficult to reuse sample sets when the research scope expands or the region changes, which leads to increased time and economic costs. Furthermore, they lack full utilization of environmental covariates, resulting in insufficient adaptability of supplementary samples in complex geographical environments.
By acquiring legacy sample point sets and environmental variable data, standardizing them, calculating similarity and uncertainty values, dividing unknown points into layers, and selecting supplementary samples based on the mean uncertainty value to form an updated sample set.
It improves the representativeness and accuracy of the sample set without re-deploying the sample, reduces costs, is applicable to areas with strong environmental heterogeneity, and enhances the evaluation accuracy of map products.
Smart Images

Figure CN121505392B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of geographic information science and spatial sampling technology, in particular relates to a sample dynamic supplement method and system based on environmental covariates. BACKGROUND
[0002] In the field of geographic information science and cartography, sampling methods as a key means to obtain spatial distribution information of geographic variables have been widely used in the process of cartographic precision evaluation for a long time. Traditional sampling methods mainly include simple random sampling, stratified random sampling, systematic sampling and other design-based sampling methods. These methods can effectively obtain representative samples in a specific study area, but they all face the common limitation that when the study scope expands or the area changes, the original sample set is often difficult to reuse directly, and it is necessary to reposition the sampling points, which not only increases the time and economic cost, but also causes waste of existing sampling resources.
[0003] With the diversification of geographic information data sources, non-probabilistic sampling data such as volunteer geographic information has gradually become popular. Although such data has the advantages of low cost and wide coverage, due to the lack of a unified design framework in the collection process, the data quality is uneven, and it is difficult to effectively integrate with sample sets based on probabilistic sampling, which limits its direct application in map product precision evaluation. At the same time, existing sample supplement methods such as conditional Latin hypercube sampling and spatial simulated annealing algorithm can supplement sampling based on the remaining samples, but most of them cannot provide the priority collection order of new samples, which lacks practical guiding value in the case of limited field sampling resources.
[0004] The deeper problem is that existing sampling methods often regard mapping sampling and evaluation sampling as two independent links, and lack a unified framework to organically combine the two. Traditional supplement algorithms focus more on the spatial prediction effect of geographic variables, and fail to fully consider the special needs of map product precision evaluation. In addition, these methods do not fully utilize environmental covariates, and fail to systematically consider the spatial dependence of sample points and the representativeness of environmental characteristics, resulting in insufficient adaptability of supplement samples in complex geographical environments.
[0005] The root of these problems lies in the lack of systematic consideration of sample reusability and spatial dependence in existing technologies, the lack of effective data fusion mechanism between different sampling frameworks, and the neglect of the internal relationship between evaluation objects and environmental covariates in traditional algorithms which overemphasize spatial uniform distribution. SUMMARY
[0006] Therefore, it is necessary to provide a sample dynamic supplement method and system based on environmental covariates to solve the above technical problems.
[0007] In a first aspect, the application provides a sample dynamic supplement method based on environmental covariates, comprising:
[0008] S1, obtaining a remaining sample point set and an environmental variable data set in a study area, and performing standardization processing on the environmental variable data set to generate a standardized environmental variable vector; wherein the environmental variable data set comprises at least one of a terrain factor, a climate factor, a vegetation index, and a soil property;
[0009] S2, based on the standardized environmental variable vector, calculating the similarity value of each unknown point and each remaining sample point in the remaining sample point set by a similarity algorithm; wherein the unknown point represents a to-be-sampled position in the study area that is not covered by the remaining sample points in the remaining sample point set;
[0010] S3, calculating the uncertainty value of each unknown point corresponding to the remaining sample point set according to the similarity value;
[0011] S4, calculating the uncertainty mean value of all unknown points corresponding to the remaining sample point set according to the uncertainty value, forming an uncertainty mean value set; based on the uncertainty mean value set, dividing the unknown points into multiple layers to generate a layer division result; wherein each layer in the layer division result corresponds to an uncertainty mean value range;
[0012] S5, determining the number of supplement samples for each layer according to the given total number of supplement samples and the proportion of the number of unknown points in each layer in the layer division result; selecting supplement points from each layer according to the number of supplement samples to form a supplement sample set;
[0013] S6, merging the supplement sample set and the remaining sample point set to obtain an updated sample set.
[0014] In a second aspect, the application further provides a sample dynamic supplement system based on environmental covariates, which is used to implement the method in the first aspect, and the system comprises:
[0015] An environmental feature preprocessing module is configured to obtain a remaining sample point set and an environmental variable data set in a study area, and perform standardization processing on the environmental variable data set to generate a standardized environmental variable vector; wherein the environmental variable data set comprises at least one of a terrain factor, a climate factor, a vegetation index, and a soil property;
[0016] A similarity calculation module is configured to calculate the similarity value of each unknown point and each remaining sample point in the remaining sample point set by a similarity algorithm based on the standardized environmental variable vector; wherein the unknown point represents a to-be-sampled position in the study area that is not covered by the remaining sample points in the remaining sample point set;
[0017] An uncertainty quantification module is configured to calculate the uncertainty value of each unknown point corresponding to the remaining sample point set according to the similarity value;
[0018] The hierarchical optimization division module is configured to calculate, according to the uncertainty values, uncertainty mean values corresponding to the remaining sample point set for all unknown points, to form an uncertainty mean value set; and to divide the unknown points into multiple layers based on the uncertainty mean value set to generate a layer division result, wherein each layer in the layer division result corresponds to an uncertainty mean value range.
[0019] The supplement strategy decision module is configured to determine the number of samples to be supplemented in each layer according to the given total number of supplement samples and the proportion of the number of unknown points in each layer in the layer division result; and to select supplement points from each layer to form a supplement sample set according to the number of samples to be supplemented.
[0020] The sample set dynamic fusion module is configured to merge the supplement sample set and the remaining sample point set to obtain an updated sample set.
[0021] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the sample dynamic supplement method based on the environmental covariate as in the first aspect when executing the computer program.
[0022] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the sample dynamic supplement method based on the environmental covariate as in the first aspect.
[0023] The sample dynamic supplement method, system, device and medium based on environmental covariates, by obtaining the remaining sample point set and the corresponding environmental variable data set, generating a standardized environmental variable vector after standardization, calculating the environmental similarity of each unknown point in the study area and the remaining sample point set based on this, and quantifying the uncertainty of each unknown point; further, by simulating each unknown point as a candidate point to join the sample set in turn, calculating the uncertainty mean of the newly formed sample set to the remaining unknown points, thereby measuring the strength of its representation to the overall region; according to the uncertainty mean, all unknown points are sorted and equally layered to construct a structured sampling layer; according to the preset total number of supplement samples, the sample quota is allocated according to the number proportion of unknown points in each layer, and the supplement sample set is selected from each layer; finally, the supplement sample set is combined with the original remaining sample point set to obtain an updated comprehensive evaluation sample set. The method accurately guides the supplement process through environmental covariants, and scientifically stratifies based on the uncertainty mean, realizing effective reuse and optimization of existing sample resources without the need for comprehensive layout. The core technical effect is that the obtained updated sample set, when performing map product precision evaluation, the evaluation result is obviously improved compared with the traditional stratified random sampling method in overall precision, and can more stably and more closely approach the theoretical true precision level of the region, thereby ensuring the scientificity of evaluation while significantly improving the sample utilization efficiency and cost benefit. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the embodiment or related art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0025] Figure 1 A flowchart of a sample dynamic supplement method based on environmental covariants provided by the present application;
[0026] Figure 2 A structural diagram of a sample dynamic supplement system based on environmental covariants provided by the present application. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of the present application more clear, the following will further describe the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0028] REFERENCE Figure 1It shows a flowchart of a sample dynamic supplement method based on environmental covariates provided in the application, which includes the following steps:
[0029] S1, obtain the remaining sample point set and the environmental variable data set in the study area, and standardize the environmental variable data set to generate a standardized environmental variable vector; wherein the environmental variable data set includes at least one of topographic factor, climatic factor, vegetation index and soil property.
[0030] Specifically, the remaining sample point set is a collection of sampling points in the study area that have geographical variable observation records, and its sources include historical sampling data, field survey samples, etc., and when obtaining, it is necessary to ensure that each sample point contains complete spatial coordinate information. The environmental variable data set needs to select covariates related to the geographical variable distribution of the study area, at least covering one of topographic factor, climatic factor, vegetation index and soil property, and these covariates can quantitatively represent the environmental characteristics of each spatial position. The core purpose of standardizing the environmental variable data set is to eliminate the calculation bias caused by the dimensional difference of different covariates, unify the data scale, and make each covariate have equal weight in subsequent similarity calculation. After standardization, the environmental variable data of each spatial position (including the remaining sample points and unknown points) will be integrated into a standardized environmental variable vector, providing uniform format basic data for subsequent environmental vector construction and similarity calculation.
[0031] S2, based on the standardized environmental variable vector, the similarity value of each unknown point and each remaining sample point in the remaining sample point set is calculated by a similarity algorithm; wherein the unknown point represents a to-be-sampled position in the study area that is not covered by the remaining sample points in the remaining sample point set.
[0032] Among them, based on the standardized environmental variable vector, the similarity value of each unknown point and each remaining sample point in the remaining sample point set is calculated by a similarity algorithm, including:
[0033] S21, based on the standardized environmental variable vector, the environmental vector of each remaining sample point j is constructed and the environmental vector of each unknown point i , wherein z is the number of environmental covariates; wherein, represents the kth environmental covariate in the environmental vector , represents the kth environmental covariate in the environmental vector , .
[0034] S22, for unknown point i and remaining sample point j, the monovariant similarity on the kth environmental covariate is calculated according to the formula , wherein This represents the maximum value of the k-th environmental covariate in the environmental vector of all unknown points. This represents the minimum value of the k-th environmental covariate in the environmental vector of all remaining sample points.
[0035] S23. Among all the univariate similarities calculated from all z environmental covariates, the minimum value is taken as the comprehensive similarity between the unknown point i and the remaining sample point j; whereby the expression for the comprehensive similarity between the unknown point i and the remaining sample point j is: ;in, This indicates taking the minimum value among all similarities.
[0036] S24. Traverse all remaining sample points in the remaining sample point set to obtain the similarity set between the unknown point i and the remaining sample point set. ,in This represents the number of remaining sample points.
[0037] Specifically, firstly, an environmental vector is constructed based on a standardized environmental variable vector. The dimension of the environmental vector for each legacy sample point and unknown point is consistent with the number of environmental covariates z, ensuring that each covariate occupies an independent dimension in the vector, accurately quantifying the combination of environmental features at each spatial location. Then, univariate similarity is calculated. The numerical difference between unknown points and legacy sample points on a single covariate dimension is normalized to the [0,1] interval using a formula. The closer the value is to 1, the more similar the environmental features are on that dimension. The denominator is the difference between the maximum value of that covariate for all unknown points and the minimum value for that covariate for all legacy sample points, ensuring the rationality and consistency of the difference quantification. The comprehensive similarity adopts a "minimum value" integration method. The core logic is that the overall similarity of the geographical environment is determined by the covariate dimension with the largest difference. Only when the similarity of all covariate dimensions is high does the overall environmental feature have a strong correlation, avoiding misjudgment of overall similarity due to similarity in some dimensions. Finally, all sample points in the legacy sample point set are traversed, generating a set for each unknown point containing the comprehensive similarity with all legacy sample points, fully presenting the distribution of the environmental correlation between the unknown point and the existing sample set.
[0038] S3. Based on the similarity values, calculate the uncertainty value of each unknown point corresponding to the set of remaining sample points. This calculation includes: based on the similarity set between each unknown point i and the set of remaining sample points, using the formula... Calculate the uncertainty value between the unknown point i and the remaining sample point set. in, Let represent the uncertainty value between the unknown point i and the remaining sample point set. Let i represent the set of similarities between the unknown point i and the set of remaining sample points.
[0039] Specifically, the calculation of the uncertainty value is based on a maximum value in the similarity set , which represents the comprehensive similarity of the unknown point and the sample point in the environment feature most similar to all the remaining sample points. Formula The reverse correlation between similarity and uncertainty is constructed, where is the uncertainty value of unknown point i, and the value range is [0, 1]. When approaches 1, it indicates that the unknown point is highly similar to the environment feature of a sample point in the existing sample set, and the existing sample represents the unknown point to a high degree, and the uncertainty value approaches 0; when approaches 0, it indicates that the unknown point is significantly different from the environment feature of all sample points in the existing sample set, and the existing sample cannot effectively represent the unknown point, and the uncertainty value approaches 1. Through the formula, each unknown point is calculated one by one, and finally a set containing the uncertainty values of all unknown points is generated, providing a core quantitative basis for subsequent layering operations.
[0040] S4, according to the uncertainty value, calculate the uncertainty mean value of all unknown points corresponding to the remaining sample point set, form the uncertainty mean value set; based on the uncertainty mean value set, divide the unknown points into multiple layers to generate the layer division result; wherein each layer in the layer division result corresponds to an uncertainty mean value range.
[0041] Wherein, according to the uncertainty value, calculate the uncertainty mean value of all unknown points corresponding to the remaining sample point set, form the uncertainty mean value set; based on the uncertainty mean value set, divide the unknown points into multiple layers to generate the layer division result, including:
[0042] S41, after adding unknown point i as a candidate point to the remaining sample point set, calculate the uncertainty value of the new sample set and each remaining unknown point according to the formula , and obtain the mean value of these uncertainty values, and sequentially traverse all unknown points and calculate the uncertainty mean value obtained after adding them as candidate points to the remaining sample point set , form the uncertainty mean value set; wherein the uncertainty mean value is used as an index to measure the sample representativeness of the new sample set, and the smaller the uncertainty mean value, the better the representativeness of the new sample set.
[0043] S42, based on the uncertainty mean value set, sort the uncertainty mean values of all unknown points in ascending order, and divide the numerical range of the sorted uncertainty mean values into r equidistant intervals; wherein r is the preset number of layers.
[0044] S43, according to the uncertainty mean value of each unknown point For each interval, the unknown points are assigned to the corresponding layers, and the number of unknown points in each layer is counted.
[0045] S44. Based on the given total number of supplementary samples and the number of unknown points in each layer, allocate the samples according to the proportional formula. Calculate the theoretical sample allocation number for the i-th layer; where, The number of unknown points in the i-th layer is represented by L, the total number of supplementary samples is L, and the preset number of layers is r.
[0046] S45. By rounding the theoretical sample allocation number and ensuring that at least one sample is allocated to each layer, determine the final number of supplementary samples required for each layer; based on the final number of supplementary samples required, generate the layer partitioning result containing the hierarchical structure and the number of sample allocations.
[0047] Specifically, first, all unknown points are iterated sequentially, and the mean uncertainty of each unknown point as a candidate point is calculated. Then, these means are sorted in ascending order and divided into r intervals, where the number of layers equals the number of intervals, r. Each interval contains the number of unknown points. This method constructs stratified sampling levels based on the magnitude of the uncertainty mean, and allocates supplementary samples to each level according to the proportion of unknown points in each interval. Assuming the number of supplementary samples each time is L, the number of samples is determined based on the number of unknown points in the i-th level. According to the formula The theoretical sample allocation number for the i-th layer is calculated according to the proportional allocation principle. If the number of samples allocated to the i-th layer is less than 1, then at least 1 sample is allocated to that layer. The sample number of all layers is allocated by rounding down, and finally, a layer partitioning result containing the hierarchical structure and the sample allocation number is generated.
[0048] S5. Based on the given total number of supplementary samples and the proportion of unknown points in each layer in the layer division results, determine the number of supplementary samples needed for each layer; based on the number of supplementary samples needed, select supplementary points from each layer to form a supplementary sample set.
[0049] Specifically, the number of supplementary samples required for each layer is determined directly based on the final number of supplementary samples obtained from S4. This number has been adjusted through proportional allocation and rounding to balance the scale of unknown points and the value of supplementation for each layer. When selecting supplementary points, a corresponding number of unknown points are selected within each layer according to preset rules. The selection rules can be randomized to ensure that the selected supplementary points can represent the uncertainty level and environmental characteristics of that layer. By selecting supplementary points for each layer, the supplementary sample set can cover regions with different supplementary value gradients. This prioritizes supplementing high-value regions (layers with low mean uncertainty and significant improvement in sample representativeness) while also considering medium- and low-value regions, avoiding representativeness imbalance caused by samples being concentrated in a single region. Ultimately, this results in a well-structured and comprehensively covered supplementary sample set.
[0050] S6, merging the augmented sample set and the remaining sample point set to obtain an updated sample set.
[0051] Specifically, the merging operation needs to ensure the uniformity of the data structure of the two types of sample points, integrate the spatial coordinates, environmental vectors, and uncertainty-related information of each point in the augmented sample set with the corresponding fields of the remaining sample point set, and form a complete updated sample set. The updated sample set not only retains the effective data resources of the original remaining sample points, avoiding sample waste, but also fills in the coverage blank areas of the original sample set through the augmented samples, improving the representativeness of the environmental characteristics of the study area. This updated sample set can be directly used for subsequent applications such as map product precision evaluation, solving the problem of needing to re-deploy all samples when the study area changes or the remaining samples are missing in traditional methods, and realizing the efficient reuse and optimal allocation of sample resources.
[0052] In addition, the method is significantly superior to traditional methods such as stratified sampling in terms of sample utilization efficiency and prediction accuracy through dynamic uncertainty assessment and precise sample augmentation. Specifically, the comparison is carried out through three core indicators: user accuracy, producer accuracy, and overall accuracy.
[0053] 1) Overall accuracy: Overall accuracy reflects the overall prediction accuracy of the sample set for the geographical variables of the study area. Traditional stratified sampling only allocates samples according to the size of the region or a preset proportion, without considering the uncertainty differences of local environmental characteristics, which can easily lead to insufficient coverage of samples in high-uncertainty areas and redundancy of samples in low-uncertainty areas, resulting in high overall prediction error. This method quantifies the value of each unknown point by the mean uncertainty and allocates samples precisely by level, ensuring that high-value areas (low mean uncertainty, high augmentation contribution) are fully covered, and that samples are reasonably allocated in low-value areas to avoid resource waste. Ultimately, the overall accuracy is significantly improved compared to traditional stratified sampling (the improvement is positively correlated with the heterogeneity of the study area).
[0054] 2) User accuracy: User accuracy focuses on the error rate of predicting a certain category but actually being of other categories, directly affecting the reliability of sample application. Traditional stratified sampling lacks targeted sample allocation, which may result in insufficient sample representativeness in areas with complex environmental characteristics (such as terrain transition zones and climate boundary areas), leading to confusion in category prediction and low user accuracy. This method adjusts extreme uncertainty values by spatial weighting, selects candidate points based on the principle of minimum uncertainty, and verifies spatial distribution to ensure the spatial relevance and environmental representativeness of augmented samples and remaining samples, reducing category prediction confusion and significantly improving user accuracy compared to traditional methods, especially in areas with strong environmental heterogeneity.
[0055] 3) Producer's precision: Producer's precision reflects the "miss rate of actual but not correctly predicted" of a certain category, which is directly related to the completeness of sample coverage on the true geographical variable. Traditional stratified sampling lacks dynamic adjustment mechanism, and when there are local environmental feature mutations in the study area, it is easy to have sample blank area, resulting in the decrease of producer's precision; this method dynamically identifies the weak sample coverage area through temporary new sample set simulation and uncertainty mean sorting, and preferentially supplements high contribution samples, ensuring that the key features of each category of geographical variable can be captured by the sample, effectively improving the producer's precision of traditional stratified sampling, and significantly reducing the risk of category omission.
[0056] In summary, the core limitation of the traditional method is that the sample allocation lacks dynamic response to environmental uncertainty, while the method through the "uncertainty quantification-hierarchical division-accurate supplement" logic closed loop makes the sample set more consistent with the environmental feature distribution of the study area, and the three precision indicators are significantly improved, especially suitable for scenes with strong environmental heterogeneity and limited sample resources.
[0057] This method makes full use of the correlation between existing residual sample points and environmental covariants, combined with dynamic uncertainty assessment, and is suitable for the following specific situations:
[0058] A) Some residual sample points in the study area are missing or cannot be updated: In long-term monitoring, regional investigation and other scenarios, some residual sample points may be missing due to natural destruction, human disturbance or observation equipment failure, or cannot be updated due to high cost, traffic restrictions and other factors. Traditional methods often need to re-deploy a large number of samples to ensure coverage completeness, resulting in high resource consumption and low efficiency; this method does not need to rely on complete residual sample point set, only through the correlation model between effective sample points and environmental covariants, the uncertainty value of unknown points is calculated, and then the sample in the missing area is accurately supplemented, on the basis of preserving the value of effective samples, only a small number of key samples are needed to restore or even improve the representativeness of the sample set, greatly reducing the cost of sample updating.
[0059] B) With the change of the study area, the original sample set of the region cannot be used: When the study area changes due to boundary adjustment, range expansion or core study area transfer, the spatial coverage of the original sample set may not match the environmental features of the new area, and the traditional sampling method needs to completely re-deploy sampling points, which not only consumes time and effort, but also may cause research interruption; this method can generate a sample set that adapts to the new region based on the environmental covariant data of the new region, combined with the effective sample points in the original sample set (not limited to the original region, only the environmental covariant dimension needs to be consistent), identify the sample blank area and high value supplement area of the new region through uncertainty assessment, dynamically generate a sample set that adapts to the new region, without completely re-sampling, significantly shortening the sample deployment period, improving research efficiency, especially suitable for application scenarios with frequent regional dynamic adjustment.
[0060] The above-mentioned sample dynamic supplement method based on environmental covariates obtains a set of remaining sample points and a corresponding set of environmental variable data, generates a standardized environmental variable vector after standardization processing, calculates the environmental similarity of each unknown point in the study area with the set of remaining sample points based on this, and quantifies the uncertainty of each unknown point accordingly; further, each unknown point is sequentially simulated as a candidate point to be added to the sample set, the uncertainty mean of the newly formed sample set to the remaining unknown points is calculated to measure the strength of its representation of the overall region; according to the uncertainty mean, all unknown points are sorted and equally layered to construct a structured sampling layer; according to the total number of supplement samples, the sample quota is allocated according to the number proportion of unknown points in each layer, and the supplement sample set is selected from each layer; finally, the supplement sample set is combined with the original set of remaining sample points to obtain an updated comprehensive evaluation sample set. This method accurately guides the supplement process through environmental covariants and scientifically stratifies based on the uncertainty mean, effectively reuses and optimizes existing sample resources without the need for comprehensive layout. The core technical effect is that the updated sample set obtained by the method can significantly improve the overall accuracy of the evaluation results compared to the traditional stratified random sampling method, and can more stably and more closely approximate the true accuracy level of the region, thereby ensuring the scientificity of the evaluation while significantly improving the sample utilization efficiency and cost-effectiveness.
[0061] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately executed with at least part of other steps or steps or stages in other steps.
[0062] Based on the same inventive concept, the embodiments of the present application also provide a system for implementing the above-mentioned sample dynamic supplement method based on environmental covariants. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme described in the above method, so the specific limitations in one or more sample dynamic supplement systems based on environmental covariants provided below can refer to the limitations of the sample dynamic supplement method based on environmental covariants described above, which will not be repeated here.
[0063] In one exemplary embodiment, as shown in Figure 2 An environment covariate-based sample dynamic supplement system 20 is provided for implementing the method in the above method embodiments, and the system comprises:
[0064] An environment feature preprocessing module 21 is configured to obtain a set of legacy sample points and a set of environment variable data in a study area, and perform standardization processing on the set of environment variable data to generate a standardized environment variable vector; wherein the set of environment variable data comprises at least one of a terrain factor, a climate factor, a vegetation index, and a soil property.
[0065] A similarity calculation module 22 is configured to calculate, based on the standardized environment variable vector, a similarity value of each unknown point with each legacy sample point in the set of legacy sample points by a similarity algorithm; wherein the unknown point represents a to-be-sampled position in the study area that is not covered by the legacy sample points in the set of legacy sample points.
[0066] An uncertainty quantification module 23 is configured to calculate, according to the similarity value, an uncertainty value of each unknown point corresponding to the set of legacy sample points.
[0067] A hierarchical optimization division module 24 is configured to calculate, according to the uncertainty value, a mean uncertainty value of all unknown points corresponding to the set of legacy sample points to form a set of mean uncertainty values; and divide the unknown points into a plurality of layers based on the set of mean uncertainty values to generate a layer division result; wherein each layer in the layer division result corresponds to a range of mean uncertainty values.
[0068] A supplement strategy decision module 25 is configured to determine, according to a given total number of supplement samples and a proportion of the number of unknown points in each layer in the layer division result, a number of supplement samples for each layer; and select, according to the number of supplement samples, supplement points from each layer to form a set of supplement samples.
[0069] A sample set dynamic fusion module 26 is configured to merge the set of supplement samples with the set of legacy sample points to obtain an updated sample set.
[0070] Embodiments of the present application also provide a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0071] Embodiments of the present application also provide a computer-readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the steps in the above method embodiments.
[0072] For the device embodiment, since it basically corresponds to the method embodiment, the relevant part can be seen from the part of the method embodiment. The device embodiment described above is only schematic, wherein the components shown as separate components can or can not be physically separate, and the components shown as a unit can or can not be a physical unit, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure. Those skilled in the art can understand and implement it without creative labor.
[0073] The above-described embodiments only express several implementation manners of the present application, which are described in detail, but cannot be understood as a limitation on the patent scope of the application. It should be pointed out that, for those skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application.
Claims
1. A method for dynamic sample supplementation based on environmental covariates, characterized in that, The method includes: S1. Obtain the legacy sample point set and the environmental variable dataset within the study area, and standardize the environmental variable dataset to generate a standardized environmental variable vector; wherein, the environmental variable dataset includes at least one of topographic factors, climate factors, vegetation indices, and soil properties; S2. Based on the standardized environmental variable vector, calculate the similarity value between each unknown point and each remaining sample point in the remaining sample point set using a similarity algorithm; wherein, the unknown point represents the sampling location in the study area that is not covered by the remaining sample points in the remaining sample point set; S3. Based on the similarity value, calculate the uncertainty value of each unknown point corresponding to the set of remaining sample points; S4. Based on the uncertainty value, calculate the mean uncertainty of all the unknown points corresponding to the set of remaining sample points to form a set of mean uncertainty values; based on the set of mean uncertainty values, divide the unknown points into multiple layers to generate layer partitioning results; wherein, each layer in the layer partitioning results corresponds to a range of mean uncertainty values; S5. Based on the given total number of supplementary samples and the proportion of unknown points in each layer in the layer division result, determine the number of supplementary samples needed for each layer; based on the number of supplementary samples needed, select supplementary points from each layer to form a supplementary sample set; S6. Merge the supplementary sample set with the legacy sample point set to obtain an updated sample set; Wherein, S2 includes: S21. Based on the standardized environmental variable vector, construct the environmental vector for each legacy sample point j and the environmental vector for each unknown point i; S22. For the unknown point i and the remaining sample point j, calculate the univariate similarity on the k-th environmental covariate; S23. Among all the univariate similarities calculated from all z environmental covariates, the minimum value is taken as the comprehensive similarity between the unknown point i and the legacy sample point j. S24. Traverse all the remaining sample points in the set of remaining sample points to obtain the similarity set between the unknown point i and the set of remaining sample points; Wherein, S3 includes: based on the similarity set between each unknown point i and the legacy sample point set, using the formula Calculate the uncertainty value between the unknown point i and the remaining sample point set; wherein, The uncertainty value between the unknown point i and the set of remaining sample points represents the uncertainty value. This represents the set of similarities between the unknown point i and the set of remaining sample points; In step S4, calculating the mean uncertainty of all unknown points corresponding to the set of remaining sample points based on the uncertainty value, forming a set of mean uncertainty values, includes: adding unknown point i as a candidate point to the set of remaining sample points, and then... Calculate the uncertainty values of the new sample set and each remaining unknown point, and obtain the mean of these uncertainty values. Iterate through all unknown points and calculate the mean uncertainty value obtained after adding them as candidate points to the remaining sample point set. This forms the set of uncertainty mean values; wherein the uncertainty mean value is used as an indicator to measure the representativeness of the new sample set, and the smaller the uncertainty mean value, the better the representativeness of the new sample set.
2. The method according to claim 1, characterized in that, In S21, the environmental vector of the remaining sample point j is The environment vector of the unknown point i is Where z is the number of the environmental covariates; where, Represents the environment vector The kth environmental covariate in Represents the environment vector The kth environmental covariate in ; In step S22, the univariate similarity between the unknown point i and the legacy sample point j on the k-th environmental covariate is... The calculation formula is ,in This represents the maximum value of the k-th environmental covariate in the environmental vector of all unknown points. This represents the minimum value of the k-th environmental covariate in the environmental vector of all remaining sample points; In step S23, the expression for the comprehensive similarity between the unknown point i and the remaining sample point j is: ;in, This indicates taking the minimum value among all similarities; In step S24, the similarity set between the unknown point i and the set of remaining sample points is: ,in This represents the number of remaining sample points.
3. The method according to claim 2, characterized in that, The process of dividing the unknown points into multiple layers based on the set of uncertain mean values and generating layer partitioning results includes: S41. Based on the set of uncertainty mean values, calculate the uncertainty mean value corresponding to all unknown points. Sort in ascending order, and then calculate the mean of the uncertainties after sorting. The numerical range is divided into r equally spaced intervals; where r is the preset number of layers; S42. Based on the mean uncertainty of each unknown point For the interval in which the unknown points are located, the unknown points are assigned to the corresponding layers, and the number of unknown points contained in each layer is counted. S43. Based on the given total number of supplementary samples and the number of unknown points in each layer, allocate the samples according to the proportional formula. Calculate the theoretical sample allocation number for the i-th layer; where, The number of unknown points in the i-th layer is represented by L, where L is the total number of supplementary samples and r is the preset number of layers. S44. By rounding the theoretical sample allocation number and ensuring that at least one sample is allocated to each layer, the final number of samples to be supplemented in each layer is determined; based on the final number of samples to be supplemented, the layer partitioning result containing the hierarchical structure and the number of sample allocations is generated.
4. A sample dynamic supplementation system based on environmental covariates, used to implement the method according to any one of claims 1 to 3, characterized in that, The system includes: An environmental feature preprocessing module is used to acquire a set of legacy sample points and a dataset of environmental variables within the study area, and to standardize the dataset to generate a standardized environmental variable vector; wherein, the dataset of environmental variables includes at least one of topographic factors, climate factors, vegetation indices, and soil properties; The similarity calculation module is used to calculate the similarity value between each unknown point and each legacy sample point in the legacy sample point set based on the standardized environmental variable vector using a similarity algorithm; wherein, the unknown point represents the sampling location in the study area that is not covered by the legacy sample points in the legacy sample point set; An uncertainty quantification module is used to calculate the uncertainty value of each unknown point corresponding to the set of legacy sample points based on the similarity value; The hierarchical optimization partitioning module is used to calculate the mean uncertainty of all unknown points corresponding to the legacy sample point set based on the uncertainty value, forming a set of mean uncertainty values; based on the set of mean uncertainty values, the unknown points are divided into multiple layers to generate a layer partitioning result; wherein, each layer in the layer partitioning result corresponds to a range of mean uncertainty values; The supplementary strategy decision module is used to determine the number of supplementary samples needed for each layer based on the given total number of supplementary samples and the proportion of unknown points in each layer in the layer division result; and to select supplementary points from each layer to form a supplementary sample set based on the number of supplementary samples needed. The sample set dynamic fusion module is used to merge the supplementary sample set with the legacy sample point set to obtain an updated sample set; The operations performed by the similarity calculation module include: S21. Based on the standardized environmental variable vector, construct the environmental vector for each legacy sample point j and the environmental vector for each unknown point i; S22. For the unknown point i and the remaining sample point j, calculate the univariate similarity on the k-th environmental covariate; S23. Among all the univariate similarities calculated from all z environmental covariates, the minimum value is taken as the comprehensive similarity between the unknown point i and the legacy sample point j. S24. Traverse all the remaining sample points in the set of remaining sample points to obtain the similarity set between the unknown point i and the set of remaining sample points; The uncertainty quantification module performs the following operations: based on the similarity set between each unknown point i and the legacy sample point set, it uses the formula... Calculate the uncertainty value between the unknown point i and the remaining sample point set; wherein, The uncertainty value between the unknown point i and the set of remaining sample points represents the uncertainty value. This represents the set of similarities between the unknown point i and the set of remaining sample points; In the operation performed by the hierarchical optimization partitioning module, the step of calculating the mean uncertainty of all unknown points corresponding to the legacy sample point set based on the uncertainty value, forming a set of mean uncertainty values, includes: adding unknown point i as a candidate point to the legacy sample point set, and then, according to the formula... Calculate the uncertainty values of the new sample set and each remaining unknown point, and obtain the mean of these uncertainty values. Iterate through all unknown points and calculate the mean uncertainty value obtained after adding them as candidate points to the remaining sample point set. This forms the set of uncertainty mean values; wherein the uncertainty mean value is used as an indicator to measure the representativeness of the new sample set, and the smaller the uncertainty mean value, the better the representativeness of the new sample set.
5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 3.
Citation Information
Patent Citations
Forest canopy height mapping method and system based on individual representativeness of sample points
CN117541679A
Multi-level track disease identification system based on vehicle body vibration data
CN119669869A