Sampling Point Attribute Feature Representativeness Measurement Method and Its Device

Through the spatial overlay analysis of the auxiliary attribute data set and Tyson polygon, the representative measurement value of the attribute characteristics of the sampling points is calculated, which solves the problem of insufficient and bias of the representative measurement of the sampling points, and improves the accuracy and reliability of data analysis.

CN120011756BActive Publication Date: 2025-07-01BEIJING RES CENT FOR INFORMATION TECH & AGRI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510459223.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-01
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

In the process of representative measurement of the attribute characteristics of sampling points, the prior art has problems of insufficient and deviation in representative measurements, which affects the accuracy and reliability of data analysis.

Method used

By determining the auxiliary attributes related to the target attribute, a auxiliary attribute data set is constructed, and spatially overlayed with the Tyson polygon as a measurement unit, the representative measure value of the attribute characteristics of each sampling point is calculated.

Benefits of technology

It improves the accuracy of the evaluation of the quality of the sampling point data, reduces the uncertainty in the data analysis process, and ensures the application accuracy and reliability of the sampling point data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011756B_ABST
    Figure CN120011756B_ABST
Patent Text Reader

Abstract

The sampling point attribute feature representativeness measurement method and device provided by the present invention belong to the technical field of data processing, and include: determining target sampling points based on target attributes, and obtaining the auxiliary attribute data set of the target sampling points and its measurement unit; performing spatial overlay analysis on the auxiliary attribute data set and the measurement unit, and comprehensively measuring the proportion of the number of attribute categories of each auxiliary attribute in the measurement unit, the proportion of the category patch area of the target sampling point in the measurement unit, and the attribute weight of the auxiliary attribute to determine the attribute feature representativeness measurement value of the target sampling point. By using the Thiessen polygon of the sampling point as the measurement unit, the present invention quantifies the contribution degree of each auxiliary attribute related to the target attribute data to the target attribute, obtains the attribute feature representativeness of each sampling point, and can evaluate the data quality and usability of the sampling point through the attribute feature representativeness of the sampling point, reduce the uncertainty of the sampling point data, and ensure the accuracy and reliability of the specific analysis and practical application of the sampling point data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a sampling point attribute feature representativeness measurement method and device. Background Art

[0002] Spatial sampling is the basis for investigating the spatial variation of soil properties and spatial mapping. The spatial uniformity of sampling points and the representativeness of attribute characteristics are key parameters for evaluating the quality of sampling point data, and are also important optimization targets for the geographic space and feature space of sampling points, respectively. The representativeness of sampling point attribute characteristics means that the attribute characteristics (such as environmental parameters, physical and chemical properties, etc.) of the selected sampling points can reflect the overall characteristics or distribution patterns of the target area or target population. The key lies in the consistency of attribute characteristics between the sampling points and the population. Before analyzing and applying the sampling point data, the attribute characteristic representativeness of the sampling point data must be measured to detect and evaluate the data quality of the sampling points. According to the sampling point data quality evaluation results, the corresponding data refinement processing of the low-representative sampling points can not only reduce the uncertainty in the subsequent analysis of the sampling point data, but also ensure the accuracy and reliability of the actual application of the sampling point data.

[0003] The higher the representativeness of the attribute characteristics of the sampling points, the more the spatial distribution law of the soil attributes reflected by the sampling points can reflect the overall distribution law of the soil attributes in the study area. In the numerical space, the soil attribute values ​​of the sampling points with high representativeness of attribute characteristics should contain the typical values ​​of the soil attributes in the study area as much as possible, and the value range of the sampling point attributes should be highly consistent with the value range of the soil attributes in the study area.

[0004] At present, there are the following problems in the process of measuring the representativeness of sampling point attribute features:

[0005] First, the data analysis is directly applied without measuring the representativeness of the attribute characteristics of the sampling points. This may result in poor representativeness of the attribute characteristics of some sampling points, increase the uncertainty in the data analysis process, and affect the precision and accuracy of the sampling point data analysis results.

[0006] The second is to measure the representativeness of individual sample points based on the similarity of environmental conditions. This indicator indirectly reflects the representativeness of the sampling points from the perspective of environmental similarity, and cannot directly reflect the representativeness of the attribute characteristics of the sampling points, which makes the measurement results prone to deviations, thereby affecting the accuracy and reliability of data analysis. Summary of the invention

[0007] The present invention provides a method and device for measuring the representativeness of sampling point attribute features, which are used to measure the representativeness of sampling point attribute features and evaluate the quality of sampling point data, thereby reducing the uncertainty of sampling points in the data analysis process and ensuring the accuracy and reliability of sampling point data mining and analysis.

[0008] The present invention provides a method for measuring the representativeness of sampling point attribute features, including the following steps:

[0009] Based on the target attribute, determine the target sampling points in the sampling area that need to measure the representativeness of attribute features, where the target attribute is determined according to the research objective;

[0010] Determine the auxiliary attributes associated with the target attribute to obtain the auxiliary attribute data of each target sampling point and construct an auxiliary attribute dataset;

[0011] Determine the Thiessen polygon where each target sampling point is located as the measurement unit for measuring the representativeness of its attribute features;

[0012] Perform a spatial overlay analysis of the auxiliary attribute dataset and all the measurement units to determine the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit and the proportion of the area of the category patches of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs;

[0013] The proportion of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the proportion of the area of the category patches is determined based on the area of the category patches corresponding to the specific category of the target sampling point in the measurement unit to which it belongs and the total area of the category patches of the specific category in the measurement unit to which it belongs;

[0014] Based on the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the area of the category patches of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weights of each auxiliary attribute, determine the measurement value of the representativeness of the attribute features of each target sampling point.

[0015] According to the method for measuring the representativeness of sampling point attribute features provided by the present invention, performing a spatial overlay analysis of the auxiliary attribute dataset and all the measurement units to determine the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit and the proportion of the area of the category patches of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs includes:

[0016] Take each auxiliary attribute in the auxiliary attribute dataset as an independent layer to generate a corresponding auxiliary attribute data layer;

[0017] Perform a spatial overlay analysis of each auxiliary attribute data layer and the measurement unit layer to obtain the number of categories of the i th auxiliary attribute in each measurement unit and the total area of the category patches of each category in the i th auxiliary attribute; the measurement unit layer is composed of the measurement units of all the target sampling points;

[0018] According to the first ratio between the number of categories of the i types of auxiliary attributes in each measurement unit and the total number of categories of the i types of auxiliary attributes in the sampling area, obtain the proportion of the number of attribute categories of the i types of auxiliary attributes in each measurement unit;

[0019] According to the specific category of the auxiliary attribute at the spatial position of the target sampling point in the measurement unit to which it belongs, determine the area of the category patch where the specific category of the target sampling point in the measurement unit is located, and the total area of the category patches corresponding to the specific category in the measurement unit;

[0020] According to the second ratio between the category patch area and the total area of the category patches, determine the category patch area ratio;

[0021] Among them, i is a positive integer.

[0022] According to a method for measuring the representativeness of sampling point attribute characteristics provided by the present invention, before generating corresponding auxiliary attribute data layers by taking each auxiliary attribute in the auxiliary attribute dataset as an independent layer, it further includes:

[0023] According to the spatial distribution characteristics of the continuous attribute data, reclassify each continuous attribute data in the auxiliary attribute dataset into categorical attribute data.

[0024] According to a method for measuring the representativeness of sampling point attribute characteristics provided by the present invention, the mathematical calculation model for determining the representativeness measurement value of the attribute characteristics of each target sampling point by comprehensively considering the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the category patch area of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weights of each auxiliary attribute is:

[0025] ;

[0026] Among them, is the representativeness measurement value of the attribute characteristics of the target sampling point y, is the number of categories of the i types of auxiliary attributes in the measurement unit where the target sampling point y is located; is the total number of categories of the i types of auxiliary attributes in the sampling area; n is the number of auxiliary attributes in the auxiliary attribute dataset; is the proportion of the category patch area of the i types of auxiliary attributes of the target sampling point y in the measurement unit to which it belongs; is the iThe attribute weights corresponding to the auxiliary attributes.

[0027] According to a method for measuring the representativeness of sampling point attribute features provided by the present invention, the determination of auxiliary attributes associated with the target attribute includes:

[0028] Determine a set of candidate auxiliary attributes, which includes a plurality of continuous attributes and a plurality of categorical attributes;

[0029] Calculate the correlation coefficient between each of the continuous attributes and the target attribute respectively, and select a preset number of continuous attributes as continuous auxiliary attributes according to the absolute value of the correlation coefficient from large to small;

[0030] Set each categorical attribute as an independent variable, set the target attribute as a dependent variable, perform an analysis of variance to obtain the significance probability value of each categorical attribute, and use the categorical attributes with the significance probability value less than the preset critical threshold as categorical auxiliary attributes.

[0031] According to a method for measuring the representativeness of sampling point attribute features provided by the present invention, the attribute weights of each auxiliary attribute are calculated based on the classification principal component analysis method, specifically including:

[0032] Perform classification principal component analysis on the auxiliary attribute data set to generate a principal component load matrix and eigenvalues;

[0033] Determine the principal components with eigenvalues greater than 1 among all the principal components as effective principal components, so as to screen out the principal component loads corresponding to the effective principal components from the principal component load matrix;

[0034] Calculate the common factor variance of each auxiliary attribute according to the principal component loads of all effective principal components, and the common factor variance is determined according to the square of the load of the auxiliary attribute on each principal component;

[0035] Based on the common factor variance of each auxiliary attribute, determine the attribute weight of each auxiliary attribute.

[0036] According to a method for measuring the representativeness of sampling point attribute features provided by the present invention, the mathematical calculation model for calculating the common factor variance is:

[0037] ;

[0038] Where, is the common factor variance of the t th auxiliary attribute, m is the number of retained effective principal components, is the t th auxiliary attribute on the load corresponding to the k th effective principal component.

[0039] A method for measuring the representativeness of sampling point attribute features provided by the present invention, determining the attribute weights of each auxiliary attribute based on the common factor variance of each auxiliary attribute, includes:

[0040] Standardize the common factor variances of all auxiliary attributes;

[0041] Determine the standardized value of each obtained auxiliary attribute as the attribute weight of each auxiliary attribute;

[0042] The standardization process is one of linear normalization, Z-score standardization, and range standardization.

[0043] The present invention also provides a device for measuring the representativeness of sampling point attribute features, including the following modules:

[0044] A first processing unit, configured to determine target sampling points within a sampling area that need to measure the representativeness of attribute features based on a target attribute, where the target attribute is determined according to a research objective;

[0045] A second processing unit, configured to determine auxiliary attributes associated with the target attribute, so as to obtain auxiliary attribute data of each target sampling point and construct an auxiliary attribute data set;

[0046] A third processing unit, configured to determine the Thiessen polygon where each target sampling point is located as a measurement unit for measuring the representativeness of its attribute features;

[0047] A fourth processing unit, configured to perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units, determine the proportion of the number of attribute categories of each auxiliary attribute within each measurement unit and the proportion of the area of the category map of each auxiliary attribute of the target sampling point within the measurement unit to which it belongs; the proportion of the number of attribute categories is determined based on the number of categories of each auxiliary attribute within the measurement unit and the total number of categories of the same auxiliary attribute within the sampling area; the proportion of the area of the category map is determined based on the area of the category map corresponding to the specific category of the target sampling point within the measurement unit to which it belongs and the total area of the category map of the specific category within the measurement unit;

[0048] A fifth processing unit, configured to comprehensively determine the measurement value of the representativeness of the attribute features of each target sampling point based on the proportion of the number of attribute categories of each auxiliary attribute within each measurement unit, the proportion of the area of the category map of each auxiliary attribute of each target sampling point within the measurement unit to which it belongs, and the attribute weights of each auxiliary attribute.

[0049] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for measuring the representativeness of the sampling point attribute features as described in any one of the above is implemented.

[0050] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for measuring the representativeness of the sampling point attribute features as described in any one of the above is implemented.

[0051] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for measuring the representativeness of the sampling point attribute features as described in any one of the above is implemented.

[0052] The method and device for measuring the representativeness of the sampling point attribute features provided by the present invention quantify the contribution degree of each auxiliary attribute related to the target attribute data to the target attribute by taking the Thiessen polygon of the sampling point as the measurement unit, obtain the representativeness of the attribute features of each sampling point, and can evaluate the data quality and usability of the sampling point data through the representativeness of the sampling point attribute features, reduce the uncertainty of the sampling point data, so as to ensure the accuracy and reliability of the specific analysis and practical application of the sampling point data. Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 is one of the flow schematic diagrams of the method for measuring the representativeness of the sampling point attribute features provided by the present invention.

[0055] Figure 2 is the second flow schematic diagram of the method for measuring the representativeness of the sampling point attribute features provided by the present invention.

[0056] Figure 3 is the scatter diagram of the representativeness of the sampling point attribute features provided by the present invention.

[0057] Figure 4 is the structural schematic diagram of the device for measuring the representativeness of the sampling point attribute features provided by the present invention.

[0058] Figure 5 is the structural schematic diagram of the electronic device provided by the present invention. Detailed Embodiments

[0059] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention may be understood according to specific circumstances.

[0061] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0062] In the process of evaluating and analyzing the target attribute data of sampling points and actual application, due to the limitation of sampling point data, there may be situations where the number of sampling points in the research area is insufficient or the representativeness of sampling points in a local research area is poor, which will affect the accuracy and reliability of the specific analysis and actual application of the target attribute data of sampling points. Therefore, it is necessary to use data of multiple other attributes (referred to as auxiliary attributes in the present invention) associated with the target attribute data to assist in carrying out mining analysis to make up for the deficiencies in the analysis of the target attribute data. In this data analysis process, how to select the auxiliary attribute data associated with the target attribute and measure the representativeness of the attribute characteristics of sampling points is very crucial. In view of the specific business requirements of such application scenarios, the present invention proposes a method for measuring the representativeness of attribute characteristics of sampling points, which is used to measure the representativeness of the attribute characteristics of the auxiliary attribute data associated with the target attribute data.

[0063] The following combination Figures 1 - 5Describe the representative measurement method and device for the attribute characteristics of sampling points provided by the present invention, in order to evaluate the data quality of sampling points, reduce the uncertainty of sampling points in the data analysis process, and ensure the accuracy and reliability of data mining and analysis of sampling points.

[0064] Figure 1 It is one of the schematic flowcharts of the representative measurement method for the attribute characteristics of sampling points provided by the present invention, as Figure 1 shown, including but not limited to the following steps:

[0065] Step 101, determine the target sampling points in the sampling area that need to measure the representative attribute characteristics based on the target attribute.

[0066] Among them, the target attribute is determined according to the research objective, and it can be soil moisture, soil pH value or other soil attributes related to the research objective. For the convenience of description, in this embodiment, without further explanation, the attribute of soil fertility is used as the target attribute to elaborate on the entire scheme, which is not regarded as a specific limitation on the protection scope of the present invention.

[0067] In this step, mainly consider determining the target sampling points in the sampling area that need to measure the representative attribute characteristics based on the target attribute. These target sampling points are pre-selected according to the research objective. For example, several sampling points can be selected in the sampling area as target sampling points by methods such as random sampling, systematic sampling or stratified sampling.

[0068] Step 102, determine the auxiliary attributes associated with the target attribute to obtain the auxiliary attribute data of each target sampling point and construct an auxiliary attribute data set.

[0069] In this embodiment, the other attributes that have a certain association with the target attribute except the target attribute are uniformly called auxiliary attributes by the present invention. For example, the attributes such as soil texture, vegetation coverage, and terrain slope associated with soil fertility are called auxiliary attributes, and these auxiliary attributes can provide additional information support for the measurement of soil fertility.

[0070] Based on the determined auxiliary attributes, obtain the auxiliary attribute data of each target sampling point and construct an auxiliary attribute data set. For example, through means such as on-site measurement, remote sensing image interpretation or Geographic Information System (GIS) data query, obtain the auxiliary attribute data such as soil texture, vegetation coverage, and terrain slope of each target sampling point, and organize these auxiliary attribute data into an auxiliary attribute data set for subsequent analysis and use.

[0071] Step 103: determine the Voronoi Diagram where each target sampling point is located as a measurement unit for measuring the representativeness of its attribute features.

[0072] Thiessen polygons are a distance-based space segmentation method. The core idea is that on a plane, the distance from any point in each Thiessen polygon to its corresponding sampling point is strictly less than the distance to any other sampling point. Through this segmentation, each sampling point is assigned an independent polygon area (i.e., measurement unit) to characterize the representativeness of the sampling point to the surrounding spatial attribute characteristics. The spatial size and shape of each measurement unit reflects the distribution law of the sampling point density and the surrounding environment characteristics.

[0073] The boundary of the Thiessen polygon is composed of the perpendicular bisectors of adjacent sampling points. All Thiessen polygons are seamlessly connected to cover the entire study area to avoid omissions or overlaps.

[0074] Step 104, performing spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of ​​each auxiliary attribute of the target sampling point in the measurement unit to which it belongs.

[0075] Specifically, this embodiment quantifies the distribution characteristics of each auxiliary attribute in each Thiessen polygon (measurement unit) through spatial overlay analysis, and provides basic data for the subsequent calculation of representative measurement values ​​of attribute characteristics.

[0076] The ratio of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area, and is used to reflect the diversity of the auxiliary attributes in the measurement unit.

[0077] The category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of ​​the specific category in the measurement unit to which it belongs. The sampling area is generally composed of multiple internally continuous patches, and the patches show different degrees of fragmented distribution. When measuring the representativeness of the attribute characteristics of the sampling points, the present invention takes into account the fragmented characteristics of the sampling area, and introduces the category patch area ratio to facilitate a more comprehensive evaluation of the representativeness of the attribute characteristics of the target sampling points. It integrates the category information of spatial distribution and auxiliary attributes, and can reflect the spatial coverage capability of the target sampling points.

[0078] Assuming that the study area contains several soil sampling points and the target attribute is soil fertility (with organic matter content as the core indicator), it is necessary to evaluate the distribution characteristics of the auxiliary attributes in each measurement unit through auxiliary attributes (such as pH value, soil texture, parent material, etc.), and then calculate the representativeness of the attribute characteristics of soil fertility.

[0079] If it is assumed that the auxiliary attribute of soil texture in the study area has four categories, including sandy loam, light loam, medium loam, and heavy loam, and a certain measurement unit contains two of these categories, such as sandy loam and light loam, then the proportion of attribute categories corresponding to the auxiliary attribute of soil texture in the measurement unit is 50%.

[0080] Furthermore, when calculating the area ratio of the category patch corresponding to the target sampling point, first determine the specific spatial position of the target sampling point in the measurement unit to which it belongs, and then determine the specific category of the auxiliary attribute at this spatial position. Assuming that the specific category of the auxiliary attribute of soil texture is sandy loam, it is necessary to count the total area of ​​the category patch of sandy loam in the measurement unit to which the target sampling point belongs, assuming that the total area of ​​this category patch is 50 km². Then, based on the ratio between the area of ​​the category patch where the target sampling point is located (assuming it is 10 km²) and the total area of ​​the category patch, the area ratio of the category patch corresponding to the target sampling point can be calculated to be 20%.

[0081] Based on the above examples, the ratio of the number of attribute categories of each auxiliary attribute in the measurement unit to which each target sampling point belongs and the ratio of the category patch area of ​​each auxiliary attribute of each target sampling point can be determined.

[0082] Step 105, comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of ​​each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weight of each auxiliary attribute, to determine the representative measurement value of the attribute feature of each target sampling point.

[0083] Among them, the attribute weight can be determined according to the degree of correlation between the auxiliary attribute and the target attribute. For example, if the correlation between soil texture and soil fertility is high, the soil texture has a larger weight; if the correlation between vegetation coverage and soil fertility is low, the vegetation coverage has a smaller weight. Of course, other methods can also be used to more accurately calculate the attribute weight of each auxiliary attribute, which will be described in detail in the subsequent embodiments.

[0084] In this embodiment, a mathematical method such as weighted summation can be used to combine the ratio of the number of attribute categories of each auxiliary attribute, the ratio of the area of ​​the category patches of each target sampling point in the measurement unit to which the auxiliary attribute belongs, and the attribute weight of the auxiliary attribute to calculate the attribute feature representativeness measurement value of each target sampling point. The attribute feature representativeness measurement value can reflect the representativeness of the target sampling point in the measurement unit to which it belongs relative to the target attribute. The larger the attribute feature representativeness measurement value, the stronger the representativeness of the target sampling point.

[0085] The sampling point attribute feature representativeness measurement method provided by the present invention quantifies the contribution degree of each auxiliary attribute related to the target attribute data by using the Thiessen polygon of the sampling point as the measurement unit, obtains the attribute feature representativeness of each sampling point, and can evaluate the data quality and usability of the sampling point data through the sampling point attribute feature representativeness, reduce the uncertainty of the sampling point data, and ensure the accuracy and reliability of the specific analysis and actual application of the sampling point data.

[0086] Based on the content of the above embodiments, as an alternative embodiment, a spatial overlay analysis is performed on the auxiliary attribute dataset and all measurement units to determine the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit and the proportion of the area of the category patches of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs. Specifically, it includes:

[0087] Step 1: Generate a corresponding auxiliary attribute data layer for each auxiliary attribute in the auxiliary attribute dataset as an independent layer.

[0088] Step 2: Perform a spatial overlay analysis on each auxiliary attribute data layer and the measurement unit layer to obtain the number of categories of the i th auxiliary attribute and the total area of the category patches of each category in the i th auxiliary attribute in each measurement unit. Wherein, the measurement unit layer is composed of the measurement units of all the target sampling points.

[0089] Step 3: According to the first ratio between the number of categories of the i th auxiliary attribute in each measurement unit and the total number of categories of the i th auxiliary attribute in the sampling area, obtain the proportion of the number of attribute categories of the i th auxiliary attribute in each measurement unit.

[0090] Step 4: According to the specific category of the auxiliary attribute at the spatial position of the target sampling point in the measurement unit to which it belongs, determine the area of the category patch where the specific category of the target sampling point in the measurement unit is located, and the total area of the category patches corresponding to the specific category in the measurement unit.

[0091] Step 5: Determine the proportion of the category patch area according to the second ratio between the category patch area and the total category patch area.

[0092] Wherein, i is a positive integer.

[0093] In this embodiment, the specific implementation steps for realizing the spatial overlay analysis will be introduced in detail. Its basic principle is to calculate the proportion of the number of attribute categories of the auxiliary attribute and the proportion of the category patch area in each measurement unit through the overlay of the auxiliary attribute data layer and the measurement unit layer.

[0094] In order to perform spatial overlay analysis, it is first necessary to generate an independent auxiliary attribute data layer for each auxiliary attribute. Assuming that the auxiliary attributes include soil texture, vegetation coverage, and terrain slope, an auxiliary attribute data layer can be generated for each auxiliary attribute. These auxiliary attribute data layers can be generated through geographic information system (GIS) software or other related tools.

[0095] Furthermore, each auxiliary attribute data layer is spatially overlaid with the measurement unit layer. The measurement unit layer is composed of the Thiessen polygons (i.e., measurement units) of all target sampling points, which are used to characterize the spatial range represented by each target sampling point. Through spatial overlay analysis, the spatial distribution of each measurement unit can be obtained. i The number of categories and the i The total area of ​​each category in the auxiliary attribute.

[0096] For example, soil texture, an auxiliary attribute, may contain multiple categories, such as sandy loam, light loam, medium loam, and heavy loam. Through spatial overlay analysis, the number of soil texture categories within each measurement unit and the total area of ​​each category can be determined.

[0097] As an optional embodiment, assuming that the target attribute is soil fertility, the auxiliary attributes include soil texture (categories include sandy loam, light loam, and medium loam) and vegetation coverage (categories include high, medium, and low).

[0098] For the measurement unit of a target sampling point, through spatial overlay analysis, if the following data is obtained: the soil texture in the measurement unit includes two categories: sandy loam and light loam, and the number of categories is 2; the total area of ​​the category patches of sandy loam in the measurement unit is 20 m², and the total area of ​​the category patches of light loam is 30 m². The vegetation coverage in the measurement unit includes high vegetation coverage and low vegetation coverage, and the number of categories is 2; the total area of ​​the category patches of high vegetation coverage in the measurement unit is 10 m², and the total area of ​​the category patches of low vegetation coverage is 20 m².

[0099] Based on the data obtained from the above spatial overlay analysis, the ratio of the number of attribute categories of each auxiliary attribute in the measurement unit and the ratio of the category patch area of ​​each auxiliary attribute of each target sampling point in the measurement unit can be calculated.

[0100] Among them, the proportion of attribute categories of the auxiliary attribute of soil texture is 2 / 3, and the proportion of attribute categories of the auxiliary attribute of vegetation coverage is 2 / 3.

[0101] Suppose that the categories of the positions of the target sampling points in the measurement unit are sandy loam and high vegetation coverage, and the area of the sandy loam category patch corresponding to the position of the target sampling point in the measurement unit is 4 m², and the area of the high vegetation coverage category patch is 5 m². Then, the proportion of the patch area of the category under the auxiliary attribute of soil texture can be calculated as 20% (obtained by 4 / 20), and the proportion of the patch area of the category under the auxiliary attribute of vegetation coverage is 50% (obtained by 5 / 10).

[0102] Finally, mathematical methods such as weighted summation can be used to calculate the representative measurement values of the attribute characteristics of each target sampling point based on the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the patch area of each category of each auxiliary attribute of each target sampling point in its affiliated measurement unit, and the attribute weights of each auxiliary attribute.

[0103] The method for measuring the representativeness of the attribute characteristics of the sampling points provided by the present invention can accurately calculate the proportion of the number of categories of each auxiliary attribute in each measurement unit and the proportion of the patch area of the category of each target sampling point in its affiliated measurement unit by generating independent data layers from the auxiliary attribute data and performing spatial overlay analysis with the Thiessen polygon (measurement unit). This refined spatial analysis method makes the representative measurement of the sampling points more accurate and can fully reflect the attribute distribution characteristics of the sampling points within their spatial ranges.

[0104] As an alternative embodiment, before generating the corresponding auxiliary attribute data layer for each auxiliary attribute in the auxiliary attribute dataset, it further includes:

[0105] According to the spatial distribution characteristics of the continuous attribute data, each continuous attribute data in the auxiliary attribute dataset is reclassified into categorical attribute data.

[0106] Since in the auxiliary attribute dataset, some of the auxiliary attribute data exist in continuous form, such as soil moisture (expressed as a percentage), soil organic matter content (expressed in grams per kilogram), or terrain slope (expressed in degrees). Although these continuous data have obvious gradient changes in spatial distribution, it is difficult to directly reflect their hierarchical distribution characteristics.

[0107] For better spatial overlay analysis, these continuous attribute data need to be reclassified into categorical data according to their spatial distribution characteristics. The specific method of reclassification can be determined according to the actual application scenario and research objectives. The reclassification results should reflect the spatial differences of the attribute data. In addition to customizing the classification intervals, one or a combination of the following reclassification methods can also be adopted: natural breaks classification method, equal interval classification method, standard deviation classification method, and quantile classification method, etc. Which reclassification method to choose specifically can be determined by combining factors such as the spatial distribution characteristics and value range of the attribute data.

[0108] Among them, the natural breaks classification method refers to dividing the continuous attribute data into several categories according to its natural distribution characteristics, so that the data differences within each category are minimized, while the differences between categories are maximized.

[0109] The equal interval classification method refers to dividing the value range of the continuous attribute data into several equally spaced intervals, and the length of each interval is equal. It is applicable to scenarios with a clear value range and relatively uniform distribution, such as the classification of geographical data (such as terrain elevation, temperature distribution, etc.).

[0110] The standard deviation classification method refers to dividing the data categories by calculating the mean (Mean) and standard deviation (Standard Deviation) of the continuous attribute data. It is applicable to data with a normal distribution or an approximately normal distribution. The continuous attribute data is divided into several intervals, each interval centered on the mean and spaced by the standard deviation. It is applicable to scenarios where the data distribution is relatively uniform and conforms to the normal distribution, and can reflect the degree of data dispersion.

[0111] The quantile classification method is to divide the continuous attribute data into several equal intervals after arranging them in ascending order, and each interval contains the same number of data points. It is applicable to scenarios where the data distribution is uneven and the relative position needs to be highlighted, such as the classification of plant density data.

[0112] After completing the reclassification of each continuous attribute data in the auxiliary attribute dataset, each auxiliary attribute is used as an independent layer to generate the corresponding auxiliary attribute data layer, which is more convenient for subsequent spatial overlay analysis.

[0113] For example, soil organic matter can be reclassified into three categories: "low (0 - 5 g / kg)", "medium (5 - 15 g / kg)", and "high (>15 g / kg)" according to the content. After reclassification, a soil organic matter data layer is generated, and this soil organic matter data layer contains three categories: low, medium, and high.

[0114] The sampling point attribute feature representativeness measurement method provided by the present invention converts continuous attribute data into categorical attribute data through reclassification processing, which not only simplifies the complexity of spatial analysis but also can more intuitively reflect the spatial distribution characteristics of sampling points. This processing step provides a more reliable data basis for subsequent spatial overlay analysis.

[0115] As an alternative embodiment, the mathematical calculation model for determining the attribute feature representativeness measurement value of each target sampling point by integrating the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the area of the category patches of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weights of each auxiliary attribute can be:

[0116] ;

[0117] where, is the attribute feature representativeness measurement value of the target sampling point y, is the number of categories of the i th auxiliary attribute in the measurement unit where the target sampling point y is located; is the total number of categories of the i th auxiliary attribute in the sampling area; n is the number of auxiliary attributes in the auxiliary attribute dataset; is the proportion of the area of the category patches of the i th auxiliary attribute of the target sampling point y in the measurement unit to which it belongs; is the i th auxiliary attribute corresponding attribute weight.

[0118] In the specific calculation process of the attribute feature representativeness measurement value of any target sampling point in this embodiment, it is mainly implemented based on the following steps:

[0119] Step 1, determine the proportion of the number of attribute categories of each auxiliary attribute in the target sampling point and the proportion of the area of the category patches of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs.

[0120] Step 2, determine the attribute weights of each auxiliary attribute, which reflect the contribution degree of different auxiliary attributes to the target attribute. The attribute weights can be determined by correlation analysis, expert experience or statistical methods (such as principal component analysis). In this embodiment, it is assumed that the weights of each auxiliary attribute have been calculated by a certain method and normalized so that their sum is 1.

[0121] Step 3: According to the mathematical calculation model of the representative measure value of the above attribute characteristics, successively accumulate the product of the proportion of the number of attribute categories of each auxiliary attribute in the measurement unit where the target sampling point is located and the proportion of the class patch area of each auxiliary attribute of the target sampling point in its measurement unit. During the accumulation process, considering the attribute weights of different auxiliary attributes at the same time, the representative measure value of the attribute characteristics of the target sampling point can be obtained.

[0122] The following is a detailed description in combination with a specific embodiment.

[0123] Suppose the target attribute is soil fertility, and the auxiliary attributes include soil texture (including 3 categories: sandy loam, light loam, medium loam) and vegetation coverage (including 3 categories: high, medium, low). The total number of categories of each auxiliary attribute in the sampling area is shown in Table 1:

[0124] Table 1 List of auxiliary attribute data in the sampling area

[0125]

[0126] At the same time, assume that the target sampling point y The calculation data of the proportion of the number of attribute categories of each auxiliary attribute in its measurement unit is shown in Table 2:

[0127] Table 2 List of relevant data on the proportion of the number of attribute categories of the target sampling point y

[0128]

[0129] Furthermore, assume that the relevant data for calculating the proportion of the class patch area corresponding to the specific category of the target sampling point y in its measurement unit is shown in Table 3:

[0130] Table 3 List of relevant data on the proportion of the class patch area of the target sampling point y

[0131]

[0132] Assume further that the calculated attribute weights of each auxiliary attribute are: the attribute weight of soil texture is P 1 = 0.6, and the attribute weight of vegetation coverage is P 2 = 0.4. Then, the representative contribution value of each auxiliary attribute can be calculated. For example, the representative contribution value of the auxiliary attribute of soil texture is 2 / 3 20% 0.6 = 0.08. The representative contribution value of the auxiliary attribute of vegetation coverage is 2 / 3 50% 0.4 ≈ 0.13.

[0133] Finally, by accumulating the representative contribution values of all the above-mentioned auxiliary attribute categories, the representative measurement value of the attribute characteristics of the target sampling point can be obtained. y :

[0134] D y = 0.08 + 0.13 = 0.21.

[0135] Through the above calculation, the representative measurement value of the attribute characteristics of the target sampling point can be obtained. y is 0.21.

[0136] The method for measuring the representativeness of the attribute characteristics of the sampling point provided by the present invention can comprehensively and objectively evaluate the representativeness of each target sampling point by comprehensively considering the proportion of the number of categories of each auxiliary attribute and the proportion of the area of the category patches of each auxiliary attribute, and taking into account the attribute weights of the auxiliary attributes, providing a scientific basis for the evaluation of the data quality of the sample points.

[0137] Figure 2 is the second flow schematic diagram of the method for measuring the representativeness of the attribute characteristics of the sampling point provided by the present invention. The following provides a specific implementation process of the method for measuring the representativeness of the attribute characteristics of the sampling point in conjunction with Figure 2 shown:

[0138] Obtain each target sampling point and its attribute data in the research area. The attribute data is divided into target attribute data and other candidate auxiliary attribute data. To solve the problems of redundant variable interference, high computational complexity, and poor model interpretability that may exist in the analysis of the representativeness of the attribute characteristics of the sampling point, the present invention screens the auxiliary attributes strongly related to the target attribute from the coarse-grained candidate auxiliary attribute set through statistical methods, including performing correlation analysis or variance analysis on the target attribute and the candidate auxiliary attributes to screen out multiple auxiliary attributes related to the target attribute.

[0139] As an optional embodiment, the steps of determining the auxiliary attributes associated with the target attribute mainly include:

[0140] Determine a candidate auxiliary attribute set, which includes multiple continuous attributes and multiple categorical attributes.

[0141] Calculate the correlation coefficient between each continuous attribute and the target attribute respectively, and select a preset number of continuous attributes as continuous auxiliary attributes according to the absolute value of the correlation coefficient from large to small.

[0142] Set each categorical attribute as the independent variable and the target attribute as the dependent variable, and perform an Analysis of Variance (ANOVA) to obtain the significance probability value (abbreviated as P-value) of each categorical attribute. Then, take the categorical attributes with P-values less than the preset critical threshold as categorical auxiliary attributes.

[0143] As an optional embodiment, nine continuous attributes including soil thickness, soil bulk density, elevation, slope, aspect, Normalized Difference Vegetation Index (NDVI), annual average temperature, annual average precipitation, and sunshine hours, and three categorical attributes including soil parent material, soil type, and soil texture are combined to form a candidate auxiliary attribute set. Taking soil fertility as the target attribute, perform a correlation analysis to screen out the auxiliary attributes that are closely related to the target attribute.

[0144] Specifically, for the screening steps of continuous attributes, the screening is mainly achieved by calculating the correlation coefficient between the target attribute and each continuous attribute. Its mathematical calculation model can be expressed as:

[0145] ;

[0146] where is the continuous attribute x and the target attribute y between the correlation coefficients, is the first i continuous attribute value of the sampling point, is the first i target attribute value of the sampling point, is the mean of the continuous attributes of all sampling points, is the mean of the target attributes of all sampling points, n is the total number of sampling points.

[0147] After calculating the correlation coefficient corresponding to each continuous attribute, the continuous attributes can be sorted in descending order of the absolute value of the correlation coefficient, and the first preset number of continuous attributes that are significantly correlated with the target attribute are selected as auxiliary attributes.

[0148] For the screening of categorical auxiliary attributes, it is mainly through ANOVA analysis to evaluate whether the influence of different categorical attributes on the target attribute is significant. If the mean difference of different categories of a certain categorical attribute on the target attribute is significant, then this categorical attribute can be considered to be related to the target attribute and can be used as an auxiliary attribute.

[0149] Perform appropriate preprocessing on the sampling data of the sample points, such as missing value processing, outlier detection, etc. Then, use statistical software (such as SPSS, R, Python, etc.) to perform ANOVA analysis. Taking SPSS software as an example, the main operation steps include:

[0150] In the SPSS software, select the option of comparing means under the analysis menu. Finally, select one-way ANOVA and run it. Then, the target attribute can be set as the dependent variable, and by sequentially setting different categorical attributes as the independent variables, the P-value corresponding to each categorical attribute can be calculated.

[0151] According to the results of ANOVA analysis, select those categorical attributes whose P-values are less than the preset significance level (evaluated by the preset critical threshold). These attributes are considered auxiliary attributes related to the target attribute and can be used for subsequent measurement of the representativeness of attribute characteristics.

[0152] The method for measuring the representativeness of sampling point attribute characteristics provided by the present invention can systematically screen out continuous and categorical auxiliary attributes significantly related to the target attribute through correlation coefficients and ANOVA analysis, thereby ensuring that the selected auxiliary attributes can be effectively used to evaluate the representativeness of sampling point attribute characteristics and guaranteeing the accuracy of data analysis results.

[0153] Reference Figure 2 As shown, next, a brief introduction will be given on how to calculate the weights of each auxiliary attribute related to the target attribute in the method for measuring the representativeness of sampling point attribute characteristics provided by the present invention.

[0154] Considering that principal component analysis (PCA) is only applicable to data analysis of continuous attributes; while categorical principal components analysis (CATPCA) uses optimal scaling transformation to convert categorical labels into numerical values and ensures that the variance of the quantified data is maximized, and further uses the quantified data for dimensionality reduction of the overall data, which is applicable to data analysis containing categorical attributes. Therefore, the present invention calculates the attribute weights of each auxiliary attribute by using the categorical principal components analysis method, mainly including but not limited to the following steps:

[0155] Perform categorical principal components analysis on the auxiliary attribute dataset to generate a principal component loading matrix and eigenvalues. Categorical principal components analysis can identify the main variation directions in the data and project the original high-dimensional data into a lower-dimensional space while retaining as much information of the original data as possible.

[0156] Determine the principal components with eigenvalues greater than 1 among all principal components as effective principal components, and select the principal component loadings of the effective principal components. This is because the principal components with eigenvalues greater than 1 are considered to retain more information than a single original variable and are therefore considered effective.

[0157] Further, calculate the common factor variance of each auxiliary attribute according to the filtered principal component load matrix. The common factor variance is determined based on the square of the load of the auxiliary attribute on each principal component, reflecting the contribution degree of each auxiliary attribute in the principal component analysis. Its calculation formula can be expressed as:

[0158] ;

[0159] where, is the common factor variance of the t th auxiliary attribute, m is the number of effective principal components retained, is the t th auxiliary attribute, and k is the load corresponding to the k th effective principal component.

[0160] Finally, based on the common factor variance of each auxiliary attribute, determine the attribute weight corresponding to each auxiliary attribute.

[0161] As an optional embodiment, the above-mentioned determining the attribute weight of each auxiliary attribute based on the common factor variance of each auxiliary attribute may include the following steps:

[0162] Step 1, perform standardization processing on the common factor variances of all auxiliary attributes.

[0163] Step 2, determine the standardized value of each obtained auxiliary attribute as the attribute weight of each auxiliary attribute.

[0164] Among them, the standardization processing adopted may be one of linear normalization processing, Z-score standardization processing, and range standardization processing.

[0165] Taking the linear normalization processing as an example, ensure that the sum of all attribute weights is 1, so that the attribute weights can directly reflect the relative contribution degree of each auxiliary attribute.

[0166] The normalization formula adopted by the present invention may be:

[0167] ;

[0168] where, is the attribute weight of the i th auxiliary attribute, n is the total number of auxiliary attributes, is the i th weight value of the auxiliary attribute.

[0169] The following is an example to illustrate how to specifically calculate the attribute weight of each auxiliary attribute.

[0170] Suppose the auxiliary attribute dataset includes three auxiliary attributes: soil moisture, soil organic matter, and soil type. Through classified principal component analysis, the principal component load matrix and eigenvalues are obtained. The eigenvalues of principal component 1, principal component 2, and principal component 3 are 1.3, 1.1, and 0.4 respectively. The eigenvalues of the first two principal components are greater than 1, that is, the first 2 principal components are effective principal components, and the obtained principal component load matrix is shown in Table 4.

[0171] Table 4 Principal Component Load Matrix

[0172]

[0173] Furthermore, according to the principal component load matrix, the common factor variance of each auxiliary attribute is calculated. The common factor variance is the sum of the squares of the loads of the auxiliary attribute on each principal component.

[0174] For the common factor variance of soil moisture f The calculation formula for 1 is:

[0175] f 1 = 0.6 2 + 0.4 2 = 0.52.

[0176] For the common factor variance of soil organic matter f The calculation formula for 2 is:

[0177] f 2 = 0.7 2 + 0.3 2 = 0.58.

[0178] For the common factor variance of soil type f The calculation formula for 3 is:

[0179] f 3 = 0.5 2 + 0.5 2 = 0.50.

[0180] Then, based on the common factor variance of each auxiliary attribute, the weight attribute of each auxiliary attribute is calculated. To simplify the calculation, in this embodiment, a linear proportional relationship is directly set between the common factor variance and the attribute weight.

[0181] Specifically, after calculating the total common factor variance = 0.52 + 0.58 + 0.50 = 1.60, the attribute weights of each auxiliary attribute can be calculated respectively:

[0182] The auxiliary weight of soil moisture is 0.52 / 1.60, and the result is 0.3250;

[0183] The auxiliary weight of soil organic matter is 0.58 / 1.60, and the result is 0.3625;

[0184] The auxiliary weight of the soil type is 0.50 / 1.60, and the result is 0.3125.

[0185] The method for measuring the representativeness of the sampling point attribute features provided by the present invention calculates the auxiliary attribute weights through the classification principal component analysis method, effectively quantifying the contribution degree of each auxiliary attribute to the target attribute.

[0186] The following combines Figure 2 As shown, a complete embodiment is used to illustrate the specific implementation process of a method for measuring the representativeness of sampling point attribute features provided by the present invention.

[0187] Step 1: Obtain the data of 433 target sampling points in a certain sampling area, which altogether include 13 soil attributes. Taking the organic matter content as the target attribute and the remaining 12 as auxiliary attributes, a candidate auxiliary attribute set is formed. Among them, there are 9 continuous attributes: soil thickness, soil bulk density, elevation, slope, aspect, normalized difference vegetation index, annual average temperature, annual average precipitation, sunshine hours, and 3 categorical attributes: soil parent material, soil type, soil texture.

[0188] Further, based on the correlation analysis, continuous variables are screened to obtain 5 continuous auxiliary attributes associated with the target attribute of organic matter content: soil thickness, soil bulk density, elevation, normalized difference vegetation index, annual average precipitation. Based on the variance analysis, categorical variables are screened to obtain 3 categorical auxiliary attributes associated with the target attribute data: soil parent material, soil type, soil texture, that is, finally 8 auxiliary attributes related to the target attribute are screened out.

[0189] Step 2: Based on the above 433 target sampling points, generate corresponding Thiessen polygons, and perform spatial clipping to obtain the Thiessen polygon where each target sampling point is located, that is, a measurement unit for measuring the representativeness of the attribute features of each target sampling point is obtained. Among them, the operations of generating Thiessen polygons and spatial clipping can be implemented based on ArcGIS software.

[0190] Step 3: Use the method provided in the above embodiment to calculate the attribute weights of 8 auxiliary attributes related to the target attribute of the target sampling point. Since the auxiliary attributes include continuous attributes and categorical attributes, the classification principal component analysis method is selected to calculate the attribute weights. After selecting the principal components with eigenvalues greater than 1, the common factor variance is calculated according to the principal component load matrix; then, by normalizing the common factor variance, the attribute weights of soil parent material, soil type, soil texture, soil thickness, soil bulk density, elevation, normalized difference vegetation index, and annual average precipitation are calculated as: 0.1020, 0.1311, 0.1374, 0.1285, 0.1358, 0.1461, 0.0682, 0.1509 respectively.

[0191] Step 4: According to the data space distribution characteristics and value range, reclassify the 5 continuous attributes into categorical attributes, including: classifying soil thickness into 7 categories, soil bulk density into 6 categories, normalized difference vegetation index into 4 categories, elevation into 7 categories, and annual average precipitation into 5 categories.

[0192] Then, based on the 433 Thiessen polygons generated by the target sampling points, perform spatial overlay analysis on the 8 auxiliary attribute data layers of parent material, soil type, soil texture, soil thickness, soil bulk density, elevation, normalized difference vegetation index, and annual average precipitation with the Thiessen polygons (measurement unit layers) respectively, and calculate the proportion of the number of attribute categories of each auxiliary attribute within each Thiessen polygon, as well as the proportion of the area of the category patches of each auxiliary attribute of each target sampling point within its affiliated measurement unit.

[0193] Finally, the representative measurement values of the attribute characteristics of 433 target sampling points are calculated and obtained.

[0194] Figure 3 It is a schematic scatter diagram of the representativeness of the attribute characteristics of the sampling points provided by the present invention. As Figure 3 shown, the maximum value of the representative measurement values of the attribute characteristics of 433 target sampling points is 0.6701, and the minimum value is 0.1706. It can be seen that the representativeness of the attribute characteristics of 2 target sampling points is significantly low. In the next step, specific analysis or data refinement processing can be carried out on these 2 target sampling points.

[0195] Through the above embodiments, it is fully illustrated that the method for measuring the representativeness of the attribute characteristics of the sampling points provided by the present invention generates Thiessen polygons through sampling point data, spatially cuts to obtain the measurement units representative of the attribute characteristics of each sampling point, calculates the attribute weights of the auxiliary attributes related to the target attribute, and then calculates the representativeness of the attribute characteristics of each target sampling point in the research area, draws a scatter diagram of the representativeness of the attribute characteristics of the sampling points to measure the data quality of the sampling points, evaluates the data quality and usability of the sampling points through the representativeness of the attribute characteristics of the sampling points, reduces the uncertainty of the sampling point data, and can effectively ensure the accuracy and reliability of the specific analysis and practical application of the sampling point data.

[0196] Figure 4 It is a schematic structural diagram of the device for measuring the representativeness of the attribute characteristics of the sampling points provided by the present invention. As Figure 4 shown, it mainly includes the following components:

[0197] The first processing unit 41 is mainly used to determine the target sampling points that need to measure the representativeness of the attribute characteristics within the sampling area based on the target attribute, and the target attribute is determined according to the research objective.

[0198] The second processing unit 42 is mainly used to determine the auxiliary attributes associated with the target attribute, so as to obtain the auxiliary attribute data of each target sampling point and construct an auxiliary attribute data set.

[0199] The third processing unit 43 is mainly used to determine the Thiessen polygon where each target sampling point is located as the measurement unit for measuring the representativeness of its attribute characteristics.

[0200] The fourth processing unit 44 is mainly used to perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units, and determine the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit and the proportion of the area of the category map of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; the proportion of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the proportion of the area of the category map is determined based on the area of the category map corresponding to the specific category of the target sampling point in the measurement unit to which it belongs and the total area of the category map of the specific category in the measurement unit to which it belongs.

[0201] The fifth processing unit 45 is mainly used to comprehensively determine the representativeness measurement value of the attribute characteristics of each target sampling point based on the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the area of the category map of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weights of each auxiliary attribute.

[0202] It should be noted that the sampling point attribute characteristic representativeness measurement device provided by the present invention can execute the sampling point attribute characteristic representativeness measurement method described in any of the above embodiments during specific operation, and this embodiment will not be elaborated herein.

[0203] The sampling point attribute characteristic representativeness measurement device provided by the present invention uses the Thiessen polygon of the sampling point as the measurement unit, quantifies the contribution degree of each auxiliary attribute related to the target attribute data to the target attribute, obtains the representativeness of the attribute characteristics of each sampling point, and can evaluate the data quality and usability of the sampling point through the representativeness of the sampling point attribute characteristics, reduce the uncertainty of the sampling point data, and ensure the accuracy and reliability of the specific analysis and practical application of the sampling point data.

[0204] Figure 5 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute the representative measurement method for the attribute characteristics of sampling points. The method includes: determining target sampling points in the sampling area that need to be measured for the representative attribute characteristics based on the target attribute, where the target attribute is determined according to the research objective; determining auxiliary attributes associated with the target attribute to obtain auxiliary attribute data of each target sampling point and constructing an auxiliary attribute data set; determining the Thiessen polygon where each target sampling point is located as the measurement unit for measuring the representative attribute characteristics of the target sampling point; performing a spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit and the proportion of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; the proportion of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the proportion of the category patch area is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs and the total category patch area of the specific category in the measurement unit to which it belongs; comprehensively determining the representative measurement value of the attribute characteristics of each target sampling point based on the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the category patch area of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weight of each auxiliary attribute.

[0205] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0206] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the sampling point attribute feature representativeness measurement method provided in the above embodiments. The method includes: determining target sampling points in a sampling area that need to be measured for attribute feature representativeness based on a target attribute, where the target attribute is determined according to a research objective; determining an auxiliary attribute associated with the target attribute to obtain auxiliary attribute data of each target sampling point and constructing an auxiliary attribute data set; determining the Thiessen polygon where each target sampling point is located as a measurement unit for measuring the attribute feature representativeness of the target sampling point; performing a spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit and the proportion of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which the target sampling point belongs; the proportion of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the proportion of the category patch area is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which the target sampling point belongs and the total category patch area of the specific category in the measurement unit to which the target sampling point belongs; comprehensively determining the attribute feature representativeness measurement value of each target sampling point based on the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the category patch area of each auxiliary attribute of each target sampling point in the measurement unit to which the target sampling point belongs, and the attribute weight of each auxiliary attribute.

[0207] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the sampling point attribute feature representative measurement method provided in the above embodiments. The method includes: determining target sampling points in the sampling area that need to be measured for attribute feature representativeness based on a target attribute, where the target attribute is determined according to the research objective; determining auxiliary attributes associated with the target attribute to obtain auxiliary attribute data of each target sampling point and constructing an auxiliary attribute data set; determining the Thiessen polygon where each target sampling point is located as a measurement unit for measuring the attribute feature representativeness of the target sampling point; performing a spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit and the proportion of the area of the category patch of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; the proportion of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the proportion of the area of the category patch is determined based on the area of the category patch corresponding to the specific category of the target sampling point in the measurement unit to which it belongs and the total area of the category patches of the specific category in the measurement unit to which it belongs; comprehensively considering the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the area of the category patch of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weights of each auxiliary attribute, determining the representative measurement value of the attribute feature of each target sampling point.

[0208] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0209] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0210] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for measuring the representativeness of sampling point attribute features, characterized in that: include: Determining target sampling points within the soil spatial sampling area that require representative measurement of attribute characteristics based on target attributes, wherein the target attributes are determined according to the research objectives; The target sampling point is a soil sampling point, and the target attribute refers to one of the soil attributes; Determine the auxiliary attribute associated with the target attribute to obtain auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; Determine the Thiessen polygon where each target sampling point is located as a measurement unit for representative measurement of its attribute characteristics; Perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of ​​each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; The attribute category number ratio is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of ​​the specific category in the measurement unit to which it belongs; The representative measurement value of the attribute feature of each target sampling point is determined by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of ​​each auxiliary attribute of each target sampling point in the corresponding measurement unit, and the attribute weight of each auxiliary attribute.

2. The representativeness measurement method of sampling point attribute features according to claim 1 is characterized in that: The auxiliary attribute data set is spatially overlaid with all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of ​​each auxiliary attribute of the target sampling point in the measurement unit to which it belongs, including: Generate a corresponding auxiliary attribute data layer by taking each auxiliary attribute in the auxiliary attribute data set as an independent layer; Perform spatial overlay analysis on each auxiliary attribute data layer and the measurement unit layer to obtain the first i The number of categories and the i The total area of ​​the category spots of each category in the auxiliary attribute; the measurement unit layer is composed of the measurement units of all the target sampling points; According to the measurement unit i The number of categories of the auxiliary attribute is related to the number of i The first ratio between the total number of categories of the auxiliary attributes is obtained in each of the measurement units. i The ratio of the number of attribute categories of the auxiliary attributes; Determine, according to the specific category of the auxiliary attribute at the spatial position of the target sampling point in the measurement unit to which it belongs, the category patch area where the specific category of the target sampling point in the measurement unit is located, and the total area of ​​the category patches corresponding to the specific category in the measurement unit; Determining the area ratio of the category spots according to a second ratio between the area of ​​the category spots and the total area of ​​the category spots; in, i Is a positive integer.

3. The representativeness measurement method of sampling point attribute features according to claim 2 is characterized in that: Before generating a corresponding auxiliary attribute data layer by taking each auxiliary attribute in the auxiliary attribute data set as an independent layer, the method further includes: According to the spatial distribution characteristics of the continuous attribute data, each of the continuous attribute data in the auxiliary attribute data set is reclassified into categorical attribute data.

4. The representativeness measurement method of sampling point attribute features according to claim 1 is characterized in that: The mathematical calculation model for determining the representative measurement value of the attribute feature of each target sampling point by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of ​​each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weight of each auxiliary attribute is: ; in, is the representative metric value of the attribute feature of the target sampling point y, is the number of units in which the target sampling point y is located. i The number of categories of auxiliary attributes; In the sampling area i The total number of categories of auxiliary attributes; n is the number of auxiliary attributes in the auxiliary attribute dataset; The target sampling point y is the first i The area ratio of the category patches of the auxiliary attributes; For the i The attribute weight corresponding to the auxiliary attribute.

5. The representativeness measurement method of sampling point attribute features according to claim 1 is characterized in that: The determining of the auxiliary attribute associated with the target attribute comprises: Determine a candidate auxiliary attribute set, wherein the candidate auxiliary attribute set includes a plurality of continuous attributes and a plurality of categorical attributes; Calculating the correlation coefficient between each of the continuous attributes and the target attribute respectively, and selecting a preset number of continuous attributes from large to small according to the absolute values ​​of the correlation coefficients as continuous auxiliary attributes; Each categorical attribute is set as an independent variable, the target attribute is set as a dependent variable, variance analysis is performed to obtain a significance probability value of each categorical attribute, and the categorical attributes whose significance probability values ​​are less than a preset critical threshold are used as categorical auxiliary attributes.

6. The representativeness measurement method of sampling point attribute features according to any one of claims 1 to 5, characterized in that: The attribute weights of each auxiliary attribute are calculated based on the classification principal component analysis method, including: Performing classified principal component analysis on the auxiliary attribute data set to generate a principal component loading matrix and eigenvalues; Determine the principal component whose eigenvalue is greater than 1 among all principal components as the effective principal component, so as to screen out the principal component load corresponding to the effective principal component from the principal component load matrix; Calculate the common factor variance of each auxiliary attribute based on the principal component loadings of all valid principal components, where the common factor variance is determined based on the square of the auxiliary attribute loading on each principal component; Based on the common factor variance of each auxiliary attribute, the attribute weight of each auxiliary attribute is determined.

7. The representativeness measurement method of sampling point attribute features according to claim 6 is characterized in that: The mathematical calculation model used to calculate the common factor variance is: ; in, For the t The common factor variance of auxiliary attributes, m is the number of valid principal components retained, For the t The auxiliary attribute is k The loadings corresponding to the effective principal components.

8. The representativeness measurement method of sampling point attribute features according to claim 6, characterized in that: The determining the attribute weight of each auxiliary attribute based on the common factor variance of each auxiliary attribute includes: Standardize the common factor variances of all auxiliary attributes; Determine the obtained standardized value of each auxiliary attribute as the attribute weight of each auxiliary attribute; The standardization process is one of linear normalization process, Z-score standardization process, and range standardization process.

9. A sampling point attribute feature representativeness measurement device, characterized in that: include: A first processing unit is used to determine a target sampling point in a soil spatial sampling area that needs to be representatively measured for attribute characteristics based on a target attribute, wherein the target attribute is determined according to a research objective; the target sampling point is a soil sampling point, and the target attribute refers to one of the soil attributes; A second processing unit is used to determine an auxiliary attribute associated with the target attribute to obtain auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; A third processing unit is used to determine the Thiessen polygon where each target sampling point is located as a measurement unit for measuring the representativeness of its attribute characteristics; A fourth processing unit is used to perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of ​​each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; The attribute category number ratio is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of ​​the specific category in the measurement unit to which it belongs; The fifth processing unit is used to comprehensively determine the representative measurement value of the attribute feature of each target sampling point by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of ​​each auxiliary attribute of each target sampling point in the corresponding measurement unit, and the attribute weight of each auxiliary attribute.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the representativeness measurement method of sampling point attribute features according to any one of claims 1 to 8 is implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the representativeness measurement method of sampling point attribute features according to any one of claims 1 to 8 is implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the representativeness measurement method of sampling point attribute features according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Mapping soil properties with satellite data using machine learning approaches

    CN113196294A

  • Sample point layout method and device for machine learning space prediction model, and medium

    CN118656634A