Ground observation point site selection method based on regional representativeness and pixel representativeness

By constructing a grid-based surface parameter matrix and principal component analysis, combined with K-means clustering and multi-parameter temporal characteristic error calculation, representative sample areas were selected, solving the problem of low accuracy in ground observation point location, realizing efficient and accurate multi-parameter collaborative observation, and improving the application effect of quantitative remote sensing.

CN121996734APending Publication Date: 2026-05-08SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST JIAOTONG UNIV
Filing Date
2026-01-30
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies lack regional and pixel representativeness in the selection of ground observation points, resulting in low accuracy in the selection of ground observation points. This makes it impossible to accurately reflect the spatial characteristics of remote sensing data, affecting the accuracy and reliability of quantitative remote sensing applications.

Method used

A ground observation point selection method based on regional and pixel representativeness was adopted. A grid surface parameter matrix was constructed using high-resolution remote sensing image time series data. Principal component analysis and K-means clustering algorithm were used to classify grid categories. Representative sample point areas that meet the spatial distance constraint threshold were selected. The representativeness error of multi-parameter time series features was calculated. The high-resolution pixel with the smallest total representativeness error was selected as the observation point.

Benefits of technology

It improves the accuracy of ground observation point site selection, achieves precise matching between observation data and remote sensing pixel spatial characteristics, reduces the impact of spatial heterogeneity of surface parameters on site selection, and enhances the application effectiveness of quantitative remote sensing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996734A_ABST
    Figure CN121996734A_ABST
Patent Text Reader

Abstract

The invention discloses a ground observation point site selection method based on regional representativeness and pixel representativeness, and relates to the technical field of remote sensing. The method comprises the following steps: constructing a ground surface parameter matrix for each grid in a target area, and dividing the categories of the grids; for each grid category, selecting a grid which meets a spatial distance constraint threshold and is closest to the clustering center as a representative sampling point area; for each representative sample point area, acquiring multi-parameter time sequence characteristics of all high-resolution pixels in the representative sample point area, calculating representative degree errors of all parameters in the multi-parameter time sequence characteristics, and screening candidate pixel combinations meeting a preset threshold condition; and traversing random combinations in the candidate pixel combinations, selecting a plurality of high-resolution pixels with the minimum total representative error as ground observation points of the representative sample point area, and integrating the ground observation points of all the representative sample point areas into ground observation points of the target area. According to the method, the site selection precision of the ground observation points is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing technology, and in particular to a method for selecting ground observation points based on regional representativeness and pixel representativeness. Background Technology

[0002] In the field of quantitative remote sensing technology, with the rapid improvement of sensor accuracy, data processing algorithms, and remote sensing platform performance, quantitative remote sensing has been widely used in many important scenarios such as agricultural monitoring, ecological assessment, and resource exploration. Its core value lies in retrieving key surface parameters (such as vegetation index, leaf area index, surface temperature, soil moisture, and vegetation cover) from remote sensing data. These parameters can provide accurate spatial data support for various scientific research, engineering applications, and decision-making, and are the core foundation for promoting the leap of remote sensing technology from qualitative description to quantitative analysis.

[0003] In the practical application of quantitative remote sensing technology, ground station observation data is an indispensable key supporting data. Existing technologies are limited by factors such as terrain, accessibility, and power supply when selecting ground observation points. They are mostly set in areas with convenient conditions and uniform surface, which lacks regional typicality and leads to low accuracy in the selection of ground observation points. Summary of the Invention

[0004] Therefore, it is necessary to provide a ground observation point location method based on regional representativeness and pixel representativeness to address the aforementioned technical problems. This method improves the accuracy of ground observation point location.

[0005] The present invention adopts the following technical solution: This invention provides a method for selecting ground observation points based on regional representativeness and pixel representativeness, comprising: The target area is divided into multiple grids based on the latitude and longitude information of low-resolution remote sensing data in the target area; A surface parameter matrix is ​​constructed for each grid cell, and the grid cells are classified according to the surface parameter matrix. The surface parameter matrix includes normalized vegetation index, normalized water index, normalized soil moisture index, and near-infrared reflectance. For each grid category, the grid that meets the spatial distance constraint threshold and is closest to the cluster center is selected as the representative sample area; the spatial distance constraint threshold is half of the maximum range value calculated based on the semivariogram analysis results of the principal components. For each representative sample area, the multi-parameter temporal features of all high-resolution pixels within the representative sample area are obtained, the representativeness error of all parameters in the multi-parameter temporal features is calculated, and candidate pixel combinations that meet the preset threshold conditions are selected. The random combinations among the candidate pixel combinations are traversed, and the multiple high-resolution pixels with the smallest total representative error are selected as ground observation points of the representative sample area. The ground observation points of all representative sample areas are then integrated into the ground observation points of the target area.

[0006] Preferably, the grid is classified according to the surface parameter matrix, specifically including: Principal component analysis was used to reduce the dimensionality of the surface parameter matrix. Based on the data reduced by principal component analysis, the K-means clustering algorithm was used to classify the grid into categories.

[0007] Preferably, the process of determining the spatial distance constraint threshold specifically includes: A semi-variogram function is constructed based on the data after dimensionality reduction via principal component analysis; the semi-variogram function is: ; in, The interval between sample points is h The semivariogram value at time, For all intervals, the distance is h The set of sample point pairs, In pixels The eigenvalues ​​of the principal components, For in pixels The eigenvalues ​​of the principal components at that location, = + h The principal components are the data after PCA dimensionality reduction; The calculated semivariogram values ​​are fitted based on the theoretical model, and the range parameters are extracted from the fitting results. The theoretical model includes a spherical model. The range parameters are used to characterize the maximum distance at which spatial variables are correlated in space. If there are multiple principal components in the data after dimensionality reduction by principal component analysis, then half of the maximum value of the range parameter of all principal components is determined as the spatial distance constraint threshold. If the data after dimensionality reduction by principal component analysis contains only one principal component, then half of the maximum value of the range parameter of that single principal component is determined as the spatial distance constraint threshold.

[0008] Preferably, the formula for calculating the representativeness error of the parameters is: ; in, RE This represents the representativeness error of the parameters. The time series vector of the candidate sampling points. This is a time series vector representing the mean value of all pixels within the sampling area. n The length of the time series. The overall average of the region mean vector. i For time indexing; The formula for calculating the total representativeness error is: ; in, The total representativeness error, As the weight of the normalized vegetation index, As the weight of the normalized water index, To normalize the soil moisture index weights, As the weight of near-infrared reflectivity, The representativeness error of the normalized vegetation index is represented by [missing information]. The representativeness error of the normalized water index is represented by... This represents the representativeness error of the normalized soil moisture index. This represents the representativeness error of near-infrared reflectance. The formula corresponding to the preset threshold condition is: ; in, This indicates the pixel positions where the representative error across all parameters is less than the set threshold. Indicates location, The representative error of the normalized vegetation index is... The representative error threshold for the normalized vegetation index is denoted as . The representative error of the normalized water index is... The representative error threshold for the normalized water index is... The representative error of the normalized soil moisture index is... The representative error threshold for the normalized soil moisture index is... This represents the typical error in near-infrared reflectance. This represents the representative error threshold for near-infrared reflectance.

[0009] Preferably, a surface parameter matrix is ​​constructed for each grid cell, specifically including: For each grid cell, extract the temporal data of the high-resolution remote sensing image within the grid cell; Band information of the grid is obtained by spatial aggregation averaging based on time series data; Based on the band information, construct the surface parameter matrix of the grid.

[0010] Preferably, the method further includes: After acquiring medium- and low-resolution remote sensing data of the target area, the medium- and low-resolution remote sensing data is subjected to UTM projection to transform the target area from a sphere to a plane.

[0011] This invention provides a ground observation point location device based on regional representativeness and pixel representativeness, comprising: The segmentation module is used to divide the target area into multiple grids based on the latitude and longitude information of the low-resolution remote sensing data in the target area. The classification module is used to construct a surface parameter matrix for each grid and classify the grid according to the surface parameter matrix; the surface parameter matrix includes normalized vegetation index, normalized water index, normalized soil moisture index and near-infrared reflectance. The first filtering module is used to select, for each grid category, the grid that meets the spatial distance constraint threshold and is closest to the cluster center as the representative sample area; the spatial distance constraint threshold is half of the maximum range value calculated based on the semivariogram analysis results of the principal components. The second filtering module is used to obtain the multi-parameter temporal features of all high-resolution pixels in each representative sample area, calculate the representativeness error of all parameters in the multi-parameter temporal features, and filter candidate pixel combinations that meet the preset threshold conditions. The determination module is used to traverse the random combinations in the candidate pixel combinations, select the multiple high-resolution pixels with the smallest total representative error as ground observation points of the representative sample area, and integrate the ground observation points of all representative sample areas into ground observation points of the target area.

[0012] The present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for selecting ground observation points based on regional representativeness and pixel representativeness.

[0013] The present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the above-described method for selecting ground observation points based on regional representativeness and pixel representativeness.

[0014] The above-mentioned at least one technical solution adopted in this invention can achieve the following beneficial effects: This invention first constructs a gridded surface parameter matrix by spatial aggregation and averaging of high-resolution remote sensing image temporal data. Based on the dimensionality reduction results of principal component analysis, it achieves scientific classification of the grids, ensuring that grids of the same type possess consistent spatiotemporal characteristics of surface parameters. For each type of grid, the grid closest to the cluster center and satisfying the spatial distance constraint threshold is selected as a representative sample area. This invention first selects images at the regional scale, choosing representative sample areas based on the spatiotemporal characteristics of surface parameters. Within the representative sample area, by extracting multi-parameter temporal features of high-resolution pixels, calculating the representativeness error, and screening candidate pixel combinations, the pixel with the smallest total representativeness error is finally selected as the observation point. Then, at the pixel scale, the specific location of the ground observation point is determined based on the total representativeness error. The pixel scale achieves accurate matching between the observation data and the spatial average features of remote sensing pixels, significantly reducing the impact of spatial heterogeneity of surface parameters on site selection. This invention optimizes sampling at two levels while simultaneously considering dual scales and multiple parameters, improving the site selection accuracy of ground observation points. Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0016] Figure 1 A schematic diagram of a ground observation point selection method based on regional representativeness and pixel representativeness provided by the present invention; Figure 2 The correlation diagram of four indicators for PCA dimensionality reduction and the k-means clustering method based on PC1 and PC2 provided by this invention; Figure 3 The graph shows the variation of pixel scale representativeness error of each parameter with the number of points under multi-parameter collaborative sampling provided by the present invention. Figure 4 A diagram showing the experimental area, the selected 1km representative sample area, and the location of multi-parameter collaborative sampling points provided by this invention. Figure 5 This is a schematic diagram of the multi-level sample point layout method provided by the present invention; Figure 6 This is a comparison chart of the minimum representativeness error of a single indicator and the collaborative representativeness error of multiple parameters for regional samples proposed in this invention. Figure 7 A schematic diagram of a ground observation point location device based on regional representativeness and pixel representativeness provided by the present invention; Figure 8 This is a schematic diagram of a computer device for implementing a ground observation point location method based on regional representativeness and pixel representativeness, as provided by the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0018] Devices such as desktop computers, servers, and laptops are capable of executing the present invention. For ease of explanation, the following description will focus on servers as the executing entity.

[0019] The combined application of ground station observation data and remote sensing pixel data is a key step in improving the accuracy of quantitative remote sensing inversion. This step needs to achieve two core objectives simultaneously: First, to ensure optimal pixel-scale representativeness, meaning that ground observation data must accurately characterize the spatial average features of the corresponding individual remote sensing pixels to ensure spatial matching between observation data and remote sensing pixels. Lack of typicality in site selection can easily lead to observation data failing to accurately correspond to pixel features. Second, to achieve optimal interpolation results at the regional scale, meaning that by interpolating and expanding discrete ground observation data, full coverage of the study area can be achieved to support large-scale quantitative remote sensing analysis. Uneven site distribution will reduce the reliability of interpolation results, ultimately limiting the application effectiveness of quantitative remote sensing technology in practical scenarios.

[0020] Currently, ground observation points are often limited by factors such as terrain, accessibility, and power supply, and are mostly set up in areas with convenient conditions and uniform surface. Their observation values ​​may have systematic biases and lack regional typicality, which leads to a decrease in accuracy when the trained models or validated products are applied in the region.

[0021] There are three fundamental scientific challenges in jointly applying ground station observation data with remote sensing image data: Insufficient regional representativeness: Both authenticity testing and model training require that the sample data used not only represent individual pixels but also the entire study area (e.g., an ecosystem, a range of data). Scale mismatch and representativeness error: Ground observations are typically conducted at the meter scale, while medium- and low-resolution remote sensing pixels (e.g., 250m-1km in MODIS, 30m in Landsat) cover areas of hundreds to tens of thousands of square meters. Surface parameters (e.g., soil moisture, vegetation cover) exhibit significant spatial heterogeneity, making it difficult for single-point or small-area ground measurements to represent the spatial average of the entire pixel. This introduces a large representativeness error, severely reducing the credibility of authenticity testing and the accuracy of training samples. Conducting representative observations at the pixel scale is crucial to reducing this uncertainty.

[0022] Fragmented observations and isolated parameters: In traditional field observation practices, observation equipment for different surface parameters, such as flux, spectrum, soil parameters, and meteorological parameters, is often independently deployed at different locations. This single-parameter, single-point model leads to several drawbacks: (a) high cost, as each parameter requires independent site infrastructure and maintenance investment; (b) asynchronous data, with different parameters acquired at different locations, lacking strict temporal synchronization and spatial consistency, making it difficult to use for studying the coupling mechanisms between multiple parameters; and (c) lack of correlation between multiple parameters, making it impossible to obtain comprehensive surface state information at the same location. Multivariate collaborative observation can achieve resource sharing, optimized configuration, and improved efficiency, and provide a scientific basis for enhancing the intrinsic consistency among multi-source remote sensing products.

[0023] Currently, optimized sampling methods for ground-based observations can be mainly divided into two categories, but neither can simultaneously address the three challenges mentioned above: 1. Sampling methods based on geostatistical models, such as kriging and spatially stratified random sampling.

[0024] Technical principle: This method assumes that the spatial distribution of surface parameters follows certain statistical laws (such as normal distribution and second-order stationarity). It analyzes spatial autocorrelation through a prior semivariance function to guide the spatial layout of sampling points. The goal is to minimize the spatial interpolation error (such as Kriging variance) of the entire region.

[0025] The drawbacks of statistical model-based sampling methods include: Scale Applicability Conflict: This method aims to achieve optimal spatial interpolation at the regional scale. However, optimal spatial distribution at the regional scale (such as hierarchical randomization) does not guarantee representativeness for every pixel scale of the ground station. Uncaptured heterogeneity may exist within pixels, leading to representativeness errors in station data when used for pixel-scale modeling and validation.

[0026] Limitations of assumptions: Its performance heavily depends on assumptions such as spatial stability, while real surface parameters (especially in highly heterogeneous regions) often do not meet these strict assumptions, leading to model failure or suboptimal deployment schemes.

[0027] Single-parameter-oriented: Traditional statistical sampling typically optimizes for a single parameter. When multiple parameters need to be observed simultaneously, simply adopting a scheme based on a dominant parameter will result in other parameters with different spatial variability characteristics being observed in a suboptimal state. If a scheme aiming for overall optimization of multiple parameters is adopted, it cannot be guaranteed that each parameter will be in an optimal state.

[0028] 2. Sampling methods based on spatiotemporal representativeness models, such as upsampling or downsampling representativeness models.

[0029] Technical principle: This method primarily addresses the issue of pixel-scale representativeness. It pre-analyzes the spatial variation structure within pixels using high-resolution spatial data (such as UAV imagery and Sentinel-2 imagery) to identify the "representative point" or "representative region" that best represents the pixel's average value. The goal is to minimize the error between the observed point value and the pixel mean.

[0030] The shortcomings of sampling methods based on spatiotemporal representativeness models include: Scale limitations: This method focuses on optimizing representativeness at the pixel scale but does not consider the typicality of that point at the regional scale. A point that is most representative within a pixel may be located on an uncommon or anomalous land cover in the region, resulting in poor regional representativeness and making it unsuitable for training intelligent models at the regional scale.

[0031] Limitations of Single Parameters: Existing methods calculate representativeness for a single land surface variable (such as vegetation index or land surface temperature). Because different parameters exhibit different spatial variation patterns, a point optimal for one parameter may be suboptimal or even unrepresentative for others. Currently, there is a lack of a collaborative site selection method capable of simultaneously optimizing pixel representativeness for multiple parameters.

[0032] The shortcomings of existing technologies can be summarized as follows: Scale fragmentation: Existing methods are mostly single-scale oriented, either focusing only on the overall representativeness of the region or pursuing the best within a single pixel. They lack a unified optimization framework across scales and cannot simultaneously take into account the two different and often conflicting goals of optimal regional scale interpolation and the highest pixel scale representativeness.

[0033] Parameter isolation: Existing methods are designed for single parameters and cannot achieve overall optimization of multi-parameter collaborative observation under the strict constraint of a limited number of stations, making it difficult to meet the needs of multi-parameter remote sensing product validation and model training.

[0034] The approaches are rigid: geostatistical methods make overly strong assumptions and have poor adaptability; representativeness-based methods ignore regional context. Both lack a flexible and universal framework to balance the complex needs of multiple scales and objectives. Furthermore, they rely heavily on prior assumptions and are difficult to adapt to the complex and variable spatial distribution characteristics of surface parameters.

[0035] Therefore, there is an urgent need in this field for an innovative site optimization and selection method that can overcome the limitations of existing technologies and provide core technical support for the construction of an efficient, accurate, and economical multi-parameter collaborative ground observation network.

[0036] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0037] Figure 1This is a schematic diagram of a ground observation point selection method based on regional representativeness and pixel representativeness in this invention, which specifically includes the following steps: S101: Divide the target area into multiple grids based on the latitude and longitude information of the low-resolution remote sensing data in the target area.

[0038] In an exemplary embodiment, the method further includes: after acquiring medium- and low-resolution remote sensing data of the target area, performing UTM projection on the medium- and low-resolution remote sensing data to transform the target area from a sphere into a plane.

[0039] Specifically, determine the target area, download medium- and low-resolution satellite products (such as MODISMOD13A2) covering the area, unify their projection coordinate system (such as UTM projection), and use their spatial extent as the initial 1km×1km scale grid.

[0040] Specifically, this experiment uses a 3km × 20km area in the middle reaches of a river basin as an example. The remote sensing data used include: low-to-medium resolution data: MODIS MOD13A2 surface albedo product, used to define the initial grid at a 1km scale; and high-resolution data: Landsat 8 OLI sensor data, used to extract multi-temporal feature indicators for 30m resolution pixels, including NDVI, NDWI, NSMI, and NIR. The MOD13A2 data was uniformly converted to UTM projection, and a 1km × 1km grid was defined as the basic unit.

[0041] S102: Construct a surface parameter matrix for each grid cell and classify the grid cells according to the surface parameter matrix; the surface parameter matrix includes normalized vegetation index, normalized water index, normalized soil moisture index and near-infrared reflectance.

[0042] In an exemplary embodiment, a surface parameter matrix is ​​constructed for each grid cell, specifically including: for each grid cell, extracting time-series data of high-resolution remote sensing images within the grid cell; obtaining the band information of the grid cell by spatial aggregation averaging based on the time-series data; and constructing the surface parameter matrix of the grid cell based on the band information.

[0043] The surface parameter matrix is ​​given by formula (1): (1); in, The surface parameter matrix of the grid. This is a column vector composed of the NDVI values ​​of all pixels within the study area arranged in spatial order. This is a column vector composed of the NDWI values ​​of all pixels within the study area arranged in spatial order. This is a column vector composed of the NSMI values ​​of all pixels within the study area arranged in spatial order. This is a column vector composed of the NIR values ​​of all pixels within the study area arranged in spatial order. , The NDVI value of the first pixel within the grid. The NDVI value of the second pixel within the grid. For the first in the grid m The NDVI value of each pixel, , The NDWI value of the first pixel within the grid. The NDWI value of the second pixel within the grid. For the first in the grid m The NDWI value of each pixel. , The NSMI value of the first pixel within the grid. The NSMI value of the second pixel within the grid. For the first in the grid m The NSMI value of each pixel. , The NIR value of the first pixel within the grid. The NIR value of the second pixel within the grid. For the first in the grid m NIR values ​​of each pixel m This represents the total number of pixels within the grid. T This represents the transpose of a vector.

[0044] Specifically, based on the segmented grid, high-resolution satellite imagery (such as Landsat 8) is acquired within the grid. The land surface parameter matrix for each high-resolution pixel in a specified time series (e.g., June-August 2023-2024) is extracted, including but not limited to Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), Normalized Difference Soil Moisture Index (NSMI), and Near-Infrared Band Reflectance (NIR). The temporal mean values ​​of NDVI, NDWI, NSMI, and NIR for all Landsat 8 pixels within each grid during the 2023-2024 growing season (June-August) are calculated to form the land surface parameter matrix for each grid.

[0045] In an exemplary embodiment, classifying the grid according to the surface parameter matrix specifically includes: performing principal component analysis to reduce the dimensionality of the surface parameter matrix; and using the K-means clustering algorithm to classify the grid according to the data after dimensionality reduction by principal component analysis.

[0046] Specifically, Principal Component Analysis (PCA) was performed on all multi-parameter features extracted from the 1km×1km grid to reduce dimensionality, eliminate redundancy, and highlight core variation features. Subsequently, K-means clustering was used based on the principal components to divide the grid into several categories. PCA was performed on four feature indicators, such as... Figure 2 As shown in Figure (a), the correlation between NDVI, NDWI, and NIR is greater than 0.8, indicating a strong correlation among the three indicators. Dimensionality reduction yields PC1 (primarily carrying NDVI, NDWI, and NIR information) and PC2 (primarily carrying NSMI information). Based on PC1 and PC2, the K-means clustering algorithm is used to divide all grids into 5 categories.

[0047] S103: For each grid category, select the grid that meets the spatial distance constraint threshold and is closest to the cluster center as the representative sample area; the spatial distance constraint threshold is half of the maximum range value calculated based on the semivariogram analysis results of the principal components.

[0048] In an exemplary embodiment, the process of determining the spatial distance constraint threshold specifically includes: constructing a semivariogram based on the data after dimensionality reduction by principal component analysis; fitting the calculated semivariogram value based on a theoretical model and extracting range parameters from the fitting results; the theoretical model includes a spherical model; the range parameters are used to characterize the maximum distance at which spatial variables are correlated in space; if there are multiple principal components in the data after dimensionality reduction by principal component analysis, half of the maximum value of the range parameters of all principal components is determined as the spatial distance constraint threshold; if there is only one principal component in the data after dimensionality reduction by principal component analysis, half of the maximum value of the range parameters of the only principal component is determined as the spatial distance constraint threshold.

[0049] The semivariogram is given by formula (2): (2); in, The interval between sample points is h The semivariogram value at time, For all intervals, the distance is h The set of sample point pairs, In pixels The eigenvalues ​​of the principal components, For in pixels The eigenvalues ​​of the principal components at that location, = + h The principal components are the data after PCA dimensionality reduction.

[0050] Specifically, while taking into account the principle of uniform spatial distribution, spatial distance constraints are introduced to avoid excessive clustering of sample points. The grid closest to the cluster center (i.e., the one with the largest feature variation and the strongest representativeness) in each category is selected as the representative sample point area at the regional scale.

[0051] Within each category, select the grid cell that is closest to the cluster center and satisfies the spatial distance constraint, such as... Figure 2 As shown in Figure (b), five representative sample areas were selected, and the covariance determinant was used as the regional representativeness evaluation index. This index can measure the volume of the sample distribution in the principal component space, thus reflecting the degree of variation covered by the sample. The principal component covariance matrices of the entire sample and the representative sample were calculated separately. The covariance determinant of the entire sample was 3.1086, and the covariance determinant of the representative sample was 3.637. The information retention rate (representative sample / all) was 117%. This indicates that the selected samples not only have strong spatial independence and typicality, but also have a wider coverage in the principal component feature space, and have high representativeness and discriminative power.

[0052] S104: For each representative sample area, obtain the multi-parameter temporal features of all high-resolution pixels within the representative sample area, calculate the representativeness error of all parameters in the multi-parameter temporal features, and select candidate pixel combinations that meet the preset threshold conditions.

[0053] For each selected 1km×1km representative sample area, multi-parameter temporal features of NDVI, NDWI, NSMI, and NIR for all 30m Landsat 8 pixels within it are extracted. For example... Figure 3 As shown, experimental results in different regions indicate that the representativeness errors of NDVI, NDWI, NSMI, and NIR all decrease with the increase of the number of sampling points. When the number of sampling points is 1, the error level is relatively high, especially for NDVI and NDWI, indicating that single-point sampling cannot effectively characterize the spectral features of the entire region. When the number of sampling points increases to 2, the error decreases significantly, and the representativeness is significantly improved. However, when the number of sampling points further increases to 3, although the errors of each index still decrease, the rate of decrease becomes gradual, showing obvious convergence characteristics. Combining the performance of the four indices in different regions, it can be seen that a maximum of 3 sampling points can effectively balance accuracy and efficiency, ensuring that the sampling results maintain a low representativeness error without causing excessive redundancy. Figure 3This invention believes that, under the conditions of satisfying the representativeness error of all parameters being less than 3% and saving costs, it is more reasonable to finally determine the number of sampling points in sample area 1, sample area 2, and sample area 3 as 3, and the number of sampling points in sample area 4 and sample area 5 as 2.

[0054] S105: Traverse the random combinations in the candidate pixel combinations, select the multiple high-resolution pixels with the smallest total representative error as ground observation points of the representative sample area, and integrate the ground observation points of all representative sample areas into ground observation points of the target area.

[0055] For each selected representative sample area, multi-parameter temporal features of all high-resolution pixels (30m) within it are extracted. To comprehensively consider spatial and temporal representativeness, this invention introduces a vector constructed from time-series data to calculate the representativeness error. The formula for the representativeness error of the parameters is formula (3):

[0056] (3); in, RE This represents the representativeness error of the parameters. The time series vector of the candidate sampling points. This is a time series vector representing the mean value of all pixels within the sampling area. n The length of the time series. The overall average of the region mean vector. i For time indexing.

[0057] To ensure that the sampling points are representative of each parameter, a candidate combination is defined: that is, among the four parameters, the representativeness error of each parameter is less than a specific threshold. The determination of the threshold involves two considerations: firstly, ensuring that the representativeness error of each parameter is not too large, which requires… Q The value cannot be too large; on the other hand, the representativeness error requirement for a single parameter cannot be too stringent, i.e., it cannot be too small, so that there are a certain number of candidate combinations to find the combination with the smallest total representativeness error. Taking both aspects into consideration, Define a parameter in The 30th percentile of the representativeness error calculated for each combination, sorted from smallest to largest. The formula corresponding to the preset threshold condition is formula (4):

[0058] (4); in, This indicates that the representative error across all four parameters is less than the set threshold position. Indicates location, This represents the representative error of parameter 1. This represents the representative error threshold for parameter 1. This represents the representative error of parameter 2. This represents the representative error threshold for parameter 2. This represents the representative error of parameter 3. This represents the representative error threshold for parameter 3. This represents the representative error of parameter 4. This represents the representative error threshold for parameter 4.

[0059] To comprehensively evaluate the representativeness of candidate sample point combinations in the multidimensional feature space and avoid the one-sidedness of single-parameter evaluation, a total representativeness error function is constructed, the value of which is the weighted sum of the representativeness errors of each parameter. The total representativeness error function is given by formula (5):

[0060] (5); in, The total representativeness error, As the weight of the normalized vegetation index, As the weight of the normalized water index, To normalize the soil moisture index weights, As the weight of near-infrared reflectivity, Representativeness error of the normalized vegetation index. The representativeness error of the normalized water index is represented by [missing information]. The representativeness error of the normalized soil moisture index is represented by [missing information]. Representativeness error of near-infrared reflectance.

[0061] Tests show that when the number of samples is less than or equal to 2, the method of assigning weights in the objective function of multivariate collaborative optimization sampling has no significant impact on the results; that is, the difference in the spatial heterogeneity of different parameters does not have a substantial impact on the weight allocation. Considering the operability and simplicity of field sampling, this study adopts an equal-weight allocation method (i.e., letting...). = = = =1), directly summing the representativeness errors of multiple parameters to construct the objective function. This strategy not only simplifies the optimization process but also meets the practical need for equal emphasis on each parameter in multi-parameter collaborative inversion.

[0062] The number of ground sampling points to be deployed within the sampling area is determined. Then, the mean of the multi-parameter features of all possible sampling point combinations (i.e., high-resolution pixel combinations) is calculated, and the representativeness error (RE) is calculated by comparing it with the true value of the overall multi-parameter features of the sampling area (i.e., the mean of all high-resolution pixel features). This study adopts a complete traversal method based on random combinations. Assuming there are N pixels in a 1km area, when the number of sampling points is... k At that time, from N Randomly selected from each pixelk There are 1 pixel when The case is represented by formula (6):

[0063] (6); in, k For the number of samples, N For the number of pixels, .

[0064] As the number of quadrats increases, the computational cost of traversing all combinations increases dramatically. At a 1km scale, typically three small quadrats (pixels) are sufficient to obtain representative observations; therefore, in this invention… k The maximum value is set to 3.

[0065] Specifically, after determining the number of sampling points in different regions, the optimal pixel combination is calculated using the full traversal method, and candidate combinations that meet the constraints are selected using formula (6). After the candidate combinations are determined, the sum of the absolute values ​​of the representativeness errors of the four indicators is calculated for each candidate combination according to formulas (3) and (5), and the combination with the smallest total representativeness error is searched among all candidate combinations. The pixel position in this combination is the optimal observation point position corresponding to the number of samples. This method can avoid the bias caused by a single indicator and achieve a comprehensive representation of vegetation conditions, water conditions, soil information, and near-infrared reflectance characteristics. The final optimal multi-point combination satisfies the balanced representativeness of the overall region in terms of spectral characteristics, and finally outputs the position of the optimal pixel combination for each region. Figure 4 The experimental area, the selected 1km representative sample area, and the location of multi-parameter coordinating points provided for this invention are shown in the diagram.

[0066] In one exemplary embodiment, the present invention provides as follows Figure 5 The diagram illustrates a multi-level sample point layout method. Figure 5 As shown, the typicality (representativeness) and multi-parameter synergy at the regional scale: Abandoning the traditional single-pixel, single-parameter optimization approach, this method starts from the overall regional perspective and uses the multi-parameter feature space as the optimization criterion. Through analysis, candidate sample areas that can comprehensively represent the multi-dimensional variation characteristics of the region are selected, laying a solid foundation for solving the problem of "where to sample". Representativeness and multi-parameter synergy at the pixel scale: Within each representative sample area, with the goal of minimizing the comprehensive representativeness error of multiple parameters, the optimal combination of sampling points that best represents the average condition of the region is selected through traversal calculations, simultaneously achieving pixel-scale point placement and multi-parameter synergistic optimization.

[0067] The basic idea of ​​this invention is as follows: First, at the regional scale, representative sample areas are selected based on the comprehensive characteristics of multiple parameters; then, at the pixel scale, specific locations within the sample area are determined based on the principle of minimizing the overall error of multiple parameters. These two levels of optimization are interconnected, ultimately using a sparse sampling scheme to simultaneously overcome the two major challenges of dual-scale and multi-parameter problems.

[0068] The multi-level sampling point deployment method proposed in this invention integrates representative sampling area delineation at the regional scale with combined optimization strategies at the pixel scale, which significantly reduces the cost of ground observation while meeting the requirements of multi-scale authenticity verification of remote sensing. The method has the following advantages: (1) Multi-index fusion: It integrates multiple remote sensing indices such as NDVI, NDWI, NSMI, and NIR to ensure multi-dimensional representativeness; (2) Hierarchical optimization design: It realizes cross-scale point deployment scheme from top to bottom, from region to pixel; (3) Combined optimization evaluation: The pixel selection adopts the principle of minimizing error to improve the representativeness of ground points to the target area; (4) Strong adaptability: It is applicable to various remote sensing product (high resolution / medium resolution / low resolution) authenticity verification scenarios and has good versatility and promotion value.

[0069] Figure 6 This is a comparison chart of the minimum representativeness error of a single indicator and the collaborative representativeness error of multiple parameters for regional samples proposed in this invention, as shown in the figure. Figure 6 As shown, the representative errors of the multi-point combination selected by the invention in terms of NDVI, NDWI, NSMI, and NIR are very close to the optimal values ​​achievable by independent optimization of each parameter. This indicates that, with the same number of stations, the present invention can simultaneously achieve high-precision observation of four key surface parameters, theoretically increasing the observation efficiency to more than four times that of traditional single-parameter optimization methods, thereby significantly reducing the manpower, material resources, and time costs of field sampling. The experimental results fully verify the overall representativeness of the selected station combination for the multi-parameter characteristics of the region, providing an efficient and reliable ground observation scheme for multi-parameter collaborative remote sensing verification and inversion.

[0070] Through practical verification in the aforementioned experimental areas, the method of the present invention exhibits the following specific and quantifiable superior effects: 1. Excellent dual-scale collaborative optimization capability, achieving a unity of global and local representativeness.

[0071] This invention, by constructing a two-level linkage optimization framework of "region-pixel," mathematically achieves a synergistic trade-off between dual-scale representativeness for the first time. At the region scale, the principal component spatial covariance determinant is innovatively introduced as a quantitative indicator. The covariance determinant value of the selected representative sample area (3.6370) is significantly higher than the global benchmark value (3.1086), with an information retention rate of 117%, proving that it can comprehensively cover the multi-dimensional variation characteristics of the region. At the pixel scale, based on the regional optimization results, the most representative specific points within each sample area are further precisely calculated through a combined optimization algorithm. This top-down optimization strategy not only ensures the spatial variation coverage capability of the sample point layout across the global scope but also ensures its high representativeness accuracy within local pixels, fundamentally overcoming the limitations of traditional single-scale optimization methods and achieving the organic unity and synergistic optimization of dual-scale representativeness.

[0072] 2. Significant advantages of multi-parameter collaborative sampling, achieving a dual improvement in efficiency and consistency.

[0073] This invention analyzes the variation patterns of representative errors for multiple indicators (NDVI, NDWI, NSMI, NIR) under different numbers of sampling points, and clarifies the optimal balance between accuracy and cost based on the multi-parameter coupling relationship. Considering both computational load and cost, this experiment selects three sampling points as the maximum sampling quantity, using a representative error of less than 3% for all parameters to measure the number of sampling points in different sample areas. After determining the number of points, a full-traversal combinatorial optimization algorithm is used to search for the globally optimal sample point combination, aiming to minimize the sum of the absolute values ​​of the errors of multiple indicators. The ultimately selected sample point combination has a comprehensive representative error that approximates the theoretical minimum value of independent optimization for each single indicator. This demonstrates that this invention successfully achieves the design goal of "one-time station deployment, multi-parameter sharing," fundamentally solving the deviation problem of traditional single-parameter-guided methods, significantly improving sampling efficiency, and ensuring the multidimensional consistency of observation data.

[0074] 3. Strong adaptive optimization and reliability ensure the robustness of the solution in complex environments.

[0075] The entire optimization process of this method completely abandons any prior assumptions about the spatial distribution of surface parameters (such as normality and stationarity). Whether it's cluster analysis at the regional scale or combinatorial optimization at the pixel scale, it is directly driven by the real spatial heterogeneity of the surface reflected by high-resolution remote sensing imagery. Results show that the sampling scheme obtained based on this adaptive optimization mechanism can achieve high-precision comprehensive characterization of multidimensional features such as vegetation, water, soil, and biochemical properties with a very small number of sampling points (maximum 3 per region). This effectively avoids the representativeness problem caused by assumptions deviating from reality in traditional geostatistical methods, and significantly improves the adaptability, reliability, and robustness of the sampling scheme in complex real-world surface environments.

[0076] In summary, this invention addresses the core needs of verifying the authenticity of remote sensing data and constructing highly representative samples, innovatively proposing a ground observation point selection method that integrates regional and pixel representativeness. This method overcomes the limitations of traditional single-parameter, single-scale site selection. Through a multi-parameter collaborative optimization mechanism, it can simultaneously meet the inversion and verification requirements of multiple remote sensing parameters, effectively avoiding redundant site selection work. Its refined sampling framework, linking "region" and "pixel," significantly improves the comprehensive representativeness and verification reliability of ground observation data. While significantly reducing costs (requiring only 3 or fewer sample points), it greatly improves work efficiency and scientific value, possessing broad prospects for widespread application.

[0077] When applying the ground observation point selection method based on regional representativeness and pixel representativeness provided by this invention, it is not necessary to consider... Figure 1 The steps shown are executed in sequence. The specific execution order of each step can be determined as needed, and this invention does not impose any restrictions on it.

[0078] The above describes a ground observation point location method based on regional representativeness and pixel representativeness, provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding ground observation point location device based on regional representativeness and pixel representativeness, such as... Figure 7 As shown.

[0079] Figure 7 A schematic diagram of a ground observation point location device based on regional representativeness and pixel representativeness provided by the present invention includes: The partitioning module 701 is used to divide the target area into multiple grids based on the latitude and longitude information of the low-resolution remote sensing data in the target area.

[0080] The classification module 702 is used to construct a surface parameter matrix for each grid and classify the grid according to the surface parameter matrix; the surface parameter matrix includes normalized vegetation index, normalized water index, normalized soil moisture index and near-infrared reflectance.

[0081] The first screening module 703 is used to select, for each grid category, the grid that meets the spatial distance constraint threshold and is closest to the cluster center as the representative sample area; the spatial distance constraint threshold is half of the maximum range value calculated based on the semivariogram analysis results of the principal components.

[0082] The second screening module 704 is used to obtain the multi-parameter temporal features of all high-resolution pixels in each representative sample area, calculate the representativeness error of all parameters in the multi-parameter temporal features, and screen candidate pixel combinations that meet the preset threshold conditions.

[0083] The determination module 705 is used to traverse the random combinations in the candidate pixel combinations, select multiple high-resolution pixels with the smallest total representative error as ground observation points of the representative sample area, and integrate the ground observation points of all representative sample areas into ground observation points of the target area.

[0084] Specific limitations regarding the ground observation point location device based on regional and pixel representativeness can be found in the limitations of the ground observation point location method based on regional and pixel representativeness mentioned above, and will not be repeated here. Each module in the aforementioned ground observation point location device based on regional and pixel representativeness can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0085] The present invention also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The proposed method for selecting ground observation points is based on regional representativeness and pixel representativeness.

[0086] The present invention also provides Figure 8 The schematic diagram of the computer device shown is as follows: Figure 8 As shown, at the hardware level, this computer device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then executes it to achieve the above. Figure 1 The proposed method for selecting ground observation points is based on regional representativeness and pixel representativeness.

[0087] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0088] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this invention.

Claims

1. A ground observation point site selection method based on regional representativeness and pixel representativeness, characterized in that, include: The target area is divided into multiple grids based on the latitude and longitude information of low-resolution remote sensing data in the target area; Construct a surface parameter matrix for each grid cell and classify the grid cells according to the surface parameter matrix; The surface parameter matrix includes normalized vegetation index, normalized water index, normalized soil moisture index, and near-infrared reflectance. For each grid category, the grid that meets the spatial distance constraint threshold and is closest to the cluster center is selected as the representative sample area; the spatial distance constraint threshold is half of the maximum range value calculated based on the semivariogram analysis results of the principal components. For each representative sample area, the multi-parameter temporal features of all high-resolution pixels within the representative sample area are obtained, the representativeness error of all parameters in the multi-parameter temporal features is calculated, and candidate pixel combinations that meet the preset threshold conditions are selected. The random combinations among the candidate pixel combinations are traversed, and the multiple high-resolution pixels with the smallest total representative error are selected as ground observation points of the representative sample area. The ground observation points of all representative sample areas are then integrated into the ground observation points of the target area.

2. The method as described in claim 1, characterized in that, The classification of grids based on the surface parameter matrix specifically includes: Principal component analysis was used to reduce the dimensionality of the surface parameter matrix. Based on the data reduced by principal component analysis, the K-means clustering algorithm was used to classify the grid into categories.

3. The method as described in claim 2, characterized in that, The process of determining the spatial distance constraint threshold specifically includes: A semi-variogram function is constructed based on the data after dimensionality reduction by principal component analysis; the semi-variogram function is: ; in, The interval between sample points is h The semivariogram value at time, For all interval distances h The set of sample point pairs, In pixels The eigenvalues ​​of the principal components, For in pixels The eigenvalues ​​of the principal components at that location, = + h The principal components are the data after PCA dimensionality reduction; The calculated semivariogram values ​​are fitted based on a theoretical model, and the range parameters are extracted from the fitting results; the theoretical model includes a spherical model; the range parameters are used to characterize the maximum distance at which spatial variables are correlated in space. If there are multiple principal components in the data after dimensionality reduction by principal component analysis, then half of the maximum value of the range parameter of all principal components is determined as the spatial distance constraint threshold. If the data after dimensionality reduction by principal component analysis contains only one principal component, then half of the maximum value of the range parameter of that single principal component is determined as the spatial distance constraint threshold.

4. The method as described in claim 1, characterized in that, The formula for calculating the representativeness error of the parameter is: ; in, RE This represents the representativeness error of the parameters. The time series vector of the candidate sampling points. This is a time series vector representing the mean value of all pixels within the sampling area. n The length of the time series. The overall average of the region mean vector. i For time indexing; The formula for calculating the total representativeness error is: ; in, The total representativeness error, As the weight of the normalized vegetation index, As the weight of the normalized water index, To normalize the soil moisture index weights, As the weight of near-infrared reflectivity, The representativeness error of the normalized vegetation index is represented by [missing information]. The representativeness error of the normalized water index is represented by... This represents the representativeness error of the normalized soil moisture index. This represents the representativeness error of near-infrared reflectance. The formula corresponding to the preset threshold condition is: ; in, This indicates the pixel positions where the representative error across all parameters is less than the set threshold. Indicates location, The representative error of the normalized vegetation index is... The representative error threshold for the normalized vegetation index is denoted as . The representative error of the normalized water index is... The representative error threshold for the normalized water index is... The representative error of the normalized soil moisture index is... The representative error threshold for the normalized soil moisture index is... This represents the typical error in near-infrared reflectance. This represents the representative error threshold for near-infrared reflectance.

5. The method as described in claim 1, characterized in that, The construction of the surface parameter matrix for each grid cell specifically includes: For each grid cell, extract the temporal data of the high-resolution remote sensing image within the grid cell; Based on the aforementioned time-series data, the band information of the grid is obtained through spatial aggregation and averaging; Based on the band information, a grid surface parameter matrix is ​​constructed.

6. The method as described in claim 1, characterized in that, The method further includes: After acquiring medium- and low-resolution remote sensing data of the target area, the target area is transformed from a sphere to a plane by performing UTM projection on the medium- and low-resolution remote sensing data.

7. A ground observation point location device based on regional representativeness and pixel representativeness, characterized in that, include: The segmentation module is used to divide the target area into multiple grids based on the latitude and longitude information of the low-resolution remote sensing data in the target area. The classification module is used to construct a surface parameter matrix for each grid and classify the grids according to the surface parameter matrix; The surface parameter matrix includes normalized vegetation index, normalized water index, normalized soil moisture index, and near-infrared reflectance. The first filtering module is used to select, for each grid category, the grid that meets the spatial distance constraint threshold and is closest to the cluster center as the representative sample area; the spatial distance constraint threshold is half of the maximum range value calculated based on the semi-variogram analysis results of the principal components. The second filtering module is used to obtain the multi-parameter temporal features of all high-resolution pixels in each representative sample area, calculate the representativeness error of all parameters in the multi-parameter temporal features, and filter candidate pixel combinations that meet the preset threshold conditions. The determination module is used to traverse the random combinations in the candidate pixel combinations, select multiple high-resolution pixels with the smallest total representative error as ground observation points of the representative sample area, and integrate the ground observation points of all representative sample areas into ground observation points of the target area.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method as described in any one of claims 1 to 6.

9. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 6.