Random sampling method for precision inspection of remote sensing image information products

By using the method of area stratification and random sampling within each layer, the problem of low efficiency in traditional manual sampling is solved, achieving balanced distribution and rapid feedback in the accuracy inspection of remote sensing image information products, and improving the quality control capability of remote sensing products.

CN121074453BActive Publication Date: 2026-07-03SIWEI SHIJING TECH (BEIJING) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SIWEI SHIJING TECH (BEIJING) CO LTD
Filing Date
2025-09-11
Publication Date
2026-07-03

Smart Images

  • Figure CN121074453B_ABST
    Figure CN121074453B_ABST
Patent Text Reader

Abstract

The application discloses a kind of remote sensing image information product precision inspection random sampling method, it is related to image processing technical field.The method includes: step 1: the area of each product graph spot is calculated;Step 2: determine a modified total sample capacity;Step 3: according to the area numerical distribution of the product graph spot, the set of the product graph spot is divided into multiple graph spot layers;Step 4: based on the area variance of product graph spot in each graph spot layer, determine the sample capacity of each graph spot layer;Step 5: in each graph spot layer, according to its corresponding sample capacity, randomly extract product graph spot;Step 6: all product graph spots extracted from the multiple graph spot layers are combined to constitute the final sample set for precision inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a random sampling method for accuracy verification of remote sensing image information products. Background Technology

[0002] With the rapid development of satellite remote sensing technology and geographic information systems, remote sensing image information products have been widely applied in various fields such as land resource monitoring, ecological environment assessment, agricultural production management, and disaster monitoring and early warning. These products are typically based on remote sensing imagery, and after image interpretation, classification, modeling, and data processing, they generate data results containing spatially structured information, such as land cover classification maps, vegetation cover index products, burned area monitoring products, and digital elevation models. These products directly serve government management, scientific research, and commercial applications, and their quality and accuracy directly affect the scientific validity and reliability of decision-making. Therefore, how to efficiently and objectively verify the accuracy of remote sensing image information products has become an important research topic in this field.

[0003] Traditional accuracy verification methods largely rely on manual sampling and comparison. Taking land cover classification products as an example, a common practice is to manually select several verification sample points from remote sensing imagery and then compare them with high-precision reference data or field survey results to calculate indicators such as overall accuracy, user accuracy, producer accuracy, and Kappa coefficient. However, with the continuous operation of high-resolution satellites, the amount of remote sensing data generated daily has reached the petabyte level, and the number of automated remote sensing information products is enormous. The efficiency of manual sampling verification is severely lagging behind and can no longer meet the needs of operational production for rapid quality feedback. In existing technologies, statistical methods such as random sampling, systematic sampling, stratified sampling, and multi-stage sampling are widely used in the accuracy verification of remote sensing products. The advantage of simple random sampling is that it is theoretically clear, and each sample unit has an equal probability of being selected, which can avoid obvious selection bias. However, in practice, simple random sampling has two problems: first, the spatial distribution of samples may not be uniform, resulting in insufficient coverage of highly heterogeneous areas; second, when the study area is large, the sample locations are scattered, increasing the cost of field surveys or manual comparisons. Summary of the Invention

[0004] In view of this, the present invention provides a random sampling method for accuracy verification of remote sensing image information products. By calculating overall characteristics to determine a reasonable sample size, and combining area quantile stratification with intra-stratum random sampling, a balanced spatial distribution of samples is achieved. Furthermore, a difference-driven allocation mechanism improves coverage of complex and heterogeneous regions. This method not only overcomes the problems of low efficiency, large subjective bias, and easy neglect of sparse land features in traditional manual sampling, but also possesses repeatability and traceability. While ensuring representativeness and scientific rigor, it significantly shortens the verification cycle and improves the efficiency and reliability of accuracy verification, thus providing robust technical support for the quality control and large-scale application of remote sensing products.

[0005] The technical solution adopted in this invention is as follows:

[0006] A random sampling method for accuracy verification of remote sensing image information products, the method comprising:

[0007] Step 1: Calculate the area of ​​each of the multiple remote sensing image information product patches within the target area;

[0008] Step 2: Based on the area variance and area mean of all product patches, and combined with the preset confidence level and allowable error, determine a corrected total sample size;

[0009] Step 3: Based on the area distribution of the product patches, divide the set of product patches into multiple patch layers;

[0010] Step 4: Based on the area variance of product patches within each patch layer, the corrected total sample size is allocated among multiple patch layers to determine the sample size for each patch layer.

[0011] Step 5: Within each image patch layer, randomly select product image patches according to their corresponding sample size; and

[0012] Step 6: Merge all product patches extracted from multiple patch layers to form the final sample set for accuracy verification.

[0013] Furthermore, step 2 includes: First, based on the area variance S 2 Average area Calculate an initial sample size n0 based on the confidence level t and the allowable error r; then, adjust the initial sample size n0 according to the total number N of product patches to obtain the adjusted total sample size n.

[0014] Furthermore, the initial sample size n0 is determined by the following formula:

[0015] Furthermore, the corrected total sample size (n) is determined by the following formula:

[0016] Furthermore, in step 3, the set of product patches is divided into multiple patch layers based on the distribution of the area values ​​of the product patches using the quantile method.

[0017] Furthermore, in step 4, the corrected total sample size (n) is allocated among multiple patch layers in such a way that the sample size allocated to the patch layers with larger area variance is greater than that allocated to the patch layers with smaller area variance.

[0018] Furthermore, the number of layers in step 3 is 4 to 5.

[0019] Furthermore, in step 5, the method of randomly selecting product patches within each patch layer is simple random sampling.

[0020] By adopting the above technical solutions, this invention achieves the following beneficial effects: Firstly, this invention uses the area of ​​map features for statistical analysis and incorporates stratification, resulting in a more balanced spatial distribution of samples. This avoids the clustering and omission problems common in traditional simple random sampling, thus ensuring representativeness. Secondly, this method introduces a data-driven correction mechanism in determining the sample size, dynamically adjusting the sample size based on overall characteristics. This avoids resource waste due to oversampling and prevents distorted conclusions due to insufficient samples, making the testing process more scientific and reasonable. Thirdly, in the stratified sample allocation stage, by considering the differences within different strata, more samples are allocated to strata with larger area differences, thereby improving the coverage of complex heterogeneous regions and compensating for the shortcomings of existing methods in accurately reflecting highly heterogeneous regions. Furthermore, this method employs a random strategy without replacement in the sampling operation, ensuring the objectivity of the sampling results. Simultaneously, it enables process repeatability through preset rules, enhancing the stability and traceability of the test results. In terms of efficiency, this method can reliably infer overall accuracy with a limited sample size, significantly shortening the verification cycle and meeting the operational needs of rapid production of remote sensing image information products. Furthermore, the combined design of stratification and sample merging preserves the overall structural characteristics while improving the coverage of different area patches by the sampling results, providing strong support for product quality control and optimization. Overall, this invention achieves comprehensive improvements in efficiency, accuracy, representativeness, and reproducibility, effectively solving problems such as low sampling efficiency, severe systematic bias, and insufficient coverage of rare classes in existing technologies, providing a solid technical guarantee for the large-scale application of remote sensing image information products. Attached Figure Description

[0021] Figure 1This is a schematic diagram of the random sampling method for accuracy verification of remote sensing image information products in an embodiment of the present invention. Detailed Implementation

[0022] All features disclosed in this specification, or steps in all methods or processes disclosed herein, may be combined in any way, except for mutually exclusive features and / or steps.

[0023] Any feature disclosed in this specification, unless otherwise stated, may be replaced by other equivalent or similar features. That is, unless otherwise stated, each feature is merely one example of a series of equivalent or similar features.

[0024] refer to Figure 1 Example 1: Repeatable random sampling based on stratification and temperature soft allocation is implemented on a set of remote sensing image information product patches within a region to be inspected. The patches are stored in a vector file as area features, and each patch has a unique identifier and area attribute.

[0025] The set of image features is denoted as Where N is the total number of map features, g k Let A be the k-th polygon. The area of ​​the k-th polygon is denoted as A. k (Unit can be m) 2 or hm 2 (The following are not limited). Each map patch has a unique identifier ID. k .

[0026] Define the following parameters or symbols: t: confidence level parameter, corresponding to the quantile of the normal distribution. r: upper limit of the allowable relative error (dimensionless). Total area mean S 2 Total area variance n0: Initial sample size, determined by the mean, variance, confidence level, and relative error. n: Corrected total sample size, considering a finite population and sampling without replacement. m: Number of strata; in this embodiment, m = 4. Q(p): Empirical quantile function for area, i.e., the p-th quantile of the area. B j : The threshold value of the j-th layer boundary, B j =Q(p j ). The subset of patterns in the i-th layer. N i : Number of patches in the i-th layer Area variance within the i-th layer. The mean of the area variance of each layer, λ: Temperature parameter used to control the "steepness" of interlayer distribution. i : The variance scaling index for the i-th layer. pi : The proportion of target samples in the i-th layer. i : Target sample size of layer i (integer). s: Global random seed (integer). H(·): 64-bit hash function (e.g., based on SipHash or xxHash, implementation is not limited below). u k : The intralayer random key value of the k-th patch (a uniformly distributed value between 0 and 1).

[0027] Step 1: Area Calculation. Read the vector file of the polygons, unify the projected coordinates, and calculate the area A of each polygon. k .get

[0028] Step 2: Determine sample size (including finite population correction) and calculate the population area mean. With the variance of the total area S 2 Given the confidence parameter t and the upper limit of relative error r, the initial sample size is determined by the following formula.

[0029] When sampling without replacement and the population is finite (total number of patches is N), the total sample size n is obtained by applying a finite population correction to n0: In actual execution, round(n) is taken as the final target total sample size, denoted as...

[0030] Step 3: Determine the quantile threshold set {p1, p2, p3} based on the area quantiles (in this embodiment, quartiles are selected: {0.25, 0.50, 0.75}), and calculate the boundary values ​​B1 = Q(0.25), B2 = Q(0.50), and B3 = Q(0.75). Based on this, the population is divided into m = 4 layers. Calculate the intra-layer variance for each layer. With layer size N i .

[0031] Step 4: Temperature-based soft assignment (unconventional assignment algorithm) based on intra-layer variance. To achieve adaptive tilting of highly heterogeneous layers under the constraint of "relying only on intra-layer area variance", a temperature-based soft assignment strategy is adopted. First, a scaling index is constructed based on the intra-layer variance. in To avoid extremely small positive numbers with a denominator of 0 (e.g., ∈ = 10) -12 Then, the amplification degree of interlayer differences is adjusted by using the temperature parameter λ>0, and the target proportion is defined. i = 1, ..., m.

[0032] Based on this, a real-valued solution for the target sample capacity within the layer is given. Will Round to the nearest integer. Afterwards, it may occur In this case, the maximum residual method is used for integer consistency correction: the decimal residual is denoted as... like Then, decrease the sum by 1 in each of the layers with the smallest residuals until the sum is 1. like Then add 1 to each of the layers with the largest residuals until the sum is 1. The corrected integer sample size is obtained.

[0033] When n appears i >N i When n is set first i :=N i And press p to adjust the excess portion. j In the remaining allocatable layer {j|n j <N j Redistribute between them, and repeat the above integer consistency steps until it is feasible.

[0034] Note: This temperature soft distribution degenerates into a uniform distribution as λ→0 (p i ≈1 / m), as λ increases, the height... The layer gives higher p i This is an unconventional but statistically intuitive implementation of "sample allocation based on intra-layer variance," and it satisfies the requirement of using only... The limitation.

[0035] Step 5: Intra-layer random sampling (simple random sampling without replacement based on hash sorting) To simultaneously satisfy "random sampling", "no replacement", "repeatability" and "high efficiency", hash-based sorting sampling is adopted: Set a global random seed s. For each patch in the i-th layer... calculate Where id k As a unique identifier for the polygon, || represents concatenation, and H(·) maps the string to a 64-bit unsigned integer. Therefore, u k ∈(0,1) can be considered as independent and identically distributed uniform random numbers. 3) Arrange the plots within the layer according to u k Sort by size in ascending order and take the first n. i Each individual is used as a sample in this layer. This process is equivalent to assigning independent uniform random bonds to all individuals within the layer and selecting the smallest n. i 1, belonging to simple random sampling without replacement; by id k The unique determination of u by s k This enables complete repeatability.

[0036] Step 6: Sample Merging and Metadata Recording. Merge the sample sets from each layer to obtain the final sample set used for accuracy verification. The smallest n in this layeri indivual}.

[0037] Simultaneously generate a sampling record table, with fields including at least: t, r, n0, n, m,{B j}, {n i},s and the IDs of all samples k Floor number and area A k This record form is archived with the product for verification and traceability.

[0038] The confidence level and error t can be selected according to the application; for example, t = 1.96 corresponds to a confidence level of approximately 95%. r can be set within the range of 0.05 to 0.10, balancing timeliness and cost. The temperature parameter λ is preferably selected using a grid search within the range of 0.5 to 2.0: the larger λ is, the stronger the skewness towards high-variance layers; when λ→0, it approaches uniform distribution. λ can be selected based on historical projects, using minimizing the estimated variance of the validation set or the total test cost as indicators. The hash function H(·) needs to have good uniformity and speed. A high-speed unencrypted hash function for 64-bit systems (e.g., xxHash64) can be used. For cross-platform consistency, a fixed implementation version is required.

[0039] Empty layers and small layers If a certain layer N i =0, then set its target sample size n i Press P for all j Reassign to other layers. If 0 <N i <n i Let n i =N i The difference is then redistributed in the remaining layers. The quantile boundaries of juxtaposed areas occur when multiple A values ​​correspond to the quantiles. k When identical data leads to boundary instability, a "left-closed, right-open" layer interval definition is used to avoid duplicate assignments; alternatively, the order of patches with the same area is broken up at the boundary without changing the overall statistic. Repeatability and deduplication are related to u k By id k The condition 's' determines that any repeated execution under the same input will necessarily yield the same sample. If version iteration causes 'id'... k When changes occur, a mapping from the old identifier to the new identifier should be provided for traceability.

[0040] The computational complexity for area, statistics, and stratified calculations is O(N); intra-stratum sampling is based on key-value sorting, and if the "top n" selection method is used... i Linear time selection algorithms (such as Quickselect) can achieve an expected time complexity of O(N). The overall expected complexity is approximately O(N), and it can be linearly scaled to millions of patches. Parallelizing key generation and selection within each layer allows for parallel execution; quantile calculations can employ parallel selection algorithms or approximate quantile data structures (such as t-digest) to save memory.

[0041] The statistically consistent temperature soft assignment only changes the weights of the sample sizes between strata; within each stratum, simple random sampling without replacement remains, maintaining equal probability within the stratum. Since the assigned weights are... The monotonically increasing function can improve coverage of highly heterogeneous layers without introducing external priors, thereby reducing the upper bound of the variance of the overall estimate. Compared with the equal distribution after uniform stratification by area, the soft distribution of temperature can improve the sensitivity of error detection in highly heterogeneous scenarios (such as urban-rural transition zones and fractured mountainous areas) with the same total sample size; the intra-layer sampling based on hash sorting avoids the expensive random access and repetition problems, and is convenient for stable reproduction in batch processing pipelines.

[0042] In a typical region, t = 1.96, r = 0.05, m = 4, λ = 1.0, and s = 1234567 are set. The values ​​are calculated according to the steps of Example 1. S 2 Given n0 and n, the population is divided into 4 layers by the quartiles {Q(0.25), Q(0.50), Q(0.75)}. Calculate the values ​​for each layer. And obtain {p i}, obtained by correcting using the maximum residue method, {n} i Finally, press u. k =H(id) k ||s) / 2 64 Select the top n in each layer i The sample patches are merged to form the final sample set, and the sampling record is output. This instance can be run repeatedly on different platforms to obtain completely consistent sample results, meeting the business requirements of quality verification and version tracking.

[0043] Preferably, when stratifying by area quantiles, let the candidate set of stratification numbers be . To ensure lower intra-layer variance and avoid excessive management costs associated with increasing the number of layers, an objective function is introduced. Where m is the number of candidate layers. Let be the area variance of the i-th layer, N be the total number of patches, and γ>0 be a penalty coefficient (appearing for the first time; γ is used to balance the trade-off between "decreasing intra-layer dispersion" and "increasing layer number," preferably γ∈[0.1,1.0]). Calculate... The final number of layers is determined. This strategy, without changing the "layering based on area quantiles", preferably uses 4 or 5 layers to stabilize the dispersion within each layer.

[0044] Preferably, to mitigate the impact of extreme area values ​​on quantile boundaries, a Windsorization process is introduced. Let A k Define the Windsorized area as the area of ​​the k-th patch (its first appearance). Q(1-τ)), where Q(·) is the area quantile function (appearing for the first time), and τ∈[0.005,0.02] is the truncation ratio at both ends (preferably τ=0.01). Calculate Q(0.25), Q(0.50), and Q(0.75) as the stratification boundaries to maintain the principle of "stratification by area" while improving boundary stability.

[0045] Preferably, under the premise of "allocation based on intra-layer area variance", a family of optional monotonically increasing mappings is provided to adapt to the heterogeneity intensity of different regions. Let The area variance of the i-th layer (first appearance), δ>0 is the smoothing term (first appearance, preferably δ=10). -12 ), β≥1 is the power exponent (first appearance), give Where p i This represents the proportion of samples in the i-th layer (appearing for the first time). This represents the total target sample size after rounding (first appearance). When β = 1, the weights are linear; as β ∈ [1.0, 2.5] increases, the tilt towards high-variance layers is enhanced. This family is interchangeable with the "temperature soft assignment" family; either can be chosen in engineering.

[0046] Preferably, for Rounding to the nearest whole number yields the result. Afterwards, if Using the maximum residual method and introducing deterministic decision-making: calculating residuals Generate the adjudication key using the layer identifier i and the global random seed s (first occurrence, integer). Where H(·) is a 64-bit hash function (appearing for the first time), and || represents concatenation. When it is necessary to reduce the number of samples, it is done by (δ) i ,c i Decrease in ascending order; when supplementation is needed, adjust according to (δ). i ,c i The order increases sequentially in descending order until the condition is met.

[0047] The rule provides a definite and reproducible sequence of decisions in each run.

[0048] Preferably, to support ultra-large-scale patches or streaming input, a modified version of the reservoir sampling algorithm R is used in each layer. Let the reservoir capacity be n. i Sequentially scan samples within the layer: the first n i The first map patch is directly added to the database; for the t-th map patch that arrives (t>n)... i ), with probability n iThe ` / t` method replaces an element at a random location in the reservoir; this process maintains equal probability for each individual within the layer and operates without replacement, making it an equivalent alternative to batch processing based on hash sorting to select the smallest key. Randomness is determined by a pseudo-random number generator driven by a seed `s`, ensuring complete repeatability.

[0049] Preferably, while ensuring "allocation based on intra-layer variance", to avoid a few layers being assigned 0 samples and thus losing coverage, the following settings are configured: Where N i Let i be the number of the i-th layer patch (first appearance). This is an indicator function (first appearance). Let... like Then, according to the deterministic decision of option 4, in high... Recycle in lower layers until satisfied This processing falls under the integer and feasibility correction stage and does not change the principle of "variance-based skewed allocation".

[0050] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.

Claims

1. A random sampling method for accuracy verification of remote sensing image information products, characterized in that, The method includes: Step 1: Calculate the area of ​​each of the multiple remote sensing image information product patches within the target area; Step 2: Based on the area variance and area mean of all product patches, and combined with the preset confidence level and allowable error, determine a corrected total sample size; Step 2 includes: First, based on the area variance... Average area Confidence level and allowable error Calculate an initial sample size Then, based on the total amount of the product patch... For the initial sample size Make corrections to obtain the corrected total sample size. The initial sample size Determined by the following formula: The corrected total sample size Determined by the following formula: ; Step 3: Based on the area distribution of the product patches, divide the set of product patches into multiple patch layers; Step 4: Based on the area variance of product patches within each patch layer, a temperature-based soft allocation strategy is adopted to distribute the corrected total sample size among the multiple patch layers, thereby determining the sample size for each patch layer; wherein, the temperature-based soft allocation strategy includes: constructing a scaling index based on the intra-layer area variance of each patch layer. ,in For the first Intra-layer area variance of the layer This represents the mean of the area variances of each layer. To avoid extremely small positive numbers with a denominator of zero; using temperature parameters Define the proportion of target samples for each patch layer. ,in The number of image patch layers is used to calculate the target sample size for each image patch layer. Step 5: Within each image patch layer, a simple random sampling method without replacement based on hash sort is used to randomly select product image patches according to their corresponding sample size; wherein, the simple random sampling without replacement based on hash sort includes: setting a global random seed. For each patch within each patch layer, based on the patch's unique identifier... With global random seed Through hash function Calculate random key values ,in This indicates splicing; all patches within the layer of this image patch are arranged according to random key values. Sort by size from smallest to largest, and take the first few. One sample is used as the layering sample for this patch, among which The sample size for stratifying the patch; and Step 6: Merge all product patches extracted from the multiple patch layers to form the final sample set for accuracy verification.

2. The method according to claim 1, characterized in that, Step 3 uses the quantile method to divide the set of product patches into multiple patch layers based on the area distribution of the product patches.

3. The method according to claim 1, characterized in that, The number of layers for dividing the map patches in step 3 is 4 to 5.

4. The method according to claim 1, characterized in that, Step 4 also includes: after rounding the target sample capacity of each patch layer, if the sum of the sample capacities of each patch layer is not equal to the total sample capacity, the maximum residue method is used for integer consistency correction.