Sampling point attribute characteristic representativeness measurement method and device
By using auxiliary attributes and spatial overlay analysis methods of Tyson polygons in the representative measurement of the attribute characteristics of the sampling point, the problems of insufficient representative measurement and result deviation are solved, and the accuracy and reliability of data analysis are improved.
Patent Information
- Application Number
- CN202510459223.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The prior art has problems of insufficient representative measurement and deviation in the representative measurement of the attribute characteristics of sampling point, which affects the accuracy and reliability of data analysis.
By determining the auxiliary attributes related to the target attribute, a auxiliary attribute data set is constructed, and spatially overlayed with the Tyson polygon as a measurement unit, the representative measure value of the attribute characteristics of each sampling point is calculated.
It improves the accuracy of the evaluation of the quality of the sampling point data, reduces the uncertainty in the data analysis process, and ensures the application accuracy and reliability of the sampling point data.
Smart Images

Figure CN120011756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a sampling point attribute feature representativeness measurement method and device. Background Art
[0002] Spatial sampling is the basis for investigating the spatial variation of soil properties and spatial mapping. The spatial uniformity of sampling points and the representativeness of attribute characteristics are key parameters for evaluating the quality of sampling point data, and are also important optimization targets for the geographic space and feature space of sampling points, respectively. The representativeness of sampling point attribute characteristics means that the attribute characteristics (such as environmental parameters, physical and chemical properties, etc.) of the selected sampling points can reflect the overall characteristics or distribution patterns of the target area or target population. The key lies in the consistency of attribute characteristics between the sampling points and the population. Before analyzing and applying the sampling point data, the attribute characteristic representativeness of the sampling point data must be measured to detect and evaluate the data quality of the sampling points. According to the sampling point data quality evaluation results, the corresponding data refinement processing of the low-representative sampling points can not only reduce the uncertainty in the subsequent analysis of the sampling point data, but also ensure the accuracy and reliability of the actual application of the sampling point data.
[0003] The higher the representativeness of the attribute characteristics of the sampling points, the more the spatial distribution law of the soil attributes reflected by the sampling points can reflect the overall distribution law of the soil attributes in the study area. In the numerical space, the soil attribute values of the sampling points with high representativeness of attribute characteristics should contain the typical values of the soil attributes in the study area as much as possible, and the value range of the sampling point attributes should be highly consistent with the value range of the soil attributes in the study area.
[0004] At present, there are the following problems in the process of measuring the representativeness of sampling point attribute features: First, the data analysis is directly applied without measuring the representativeness of the attribute characteristics of the sampling points. This may result in poor representativeness of the attribute characteristics of some sampling points, increase the uncertainty in the data analysis process, and affect the precision and accuracy of the sampling point data analysis results.
[0005] The second is to measure the representativeness of individual sample points based on the similarity of environmental conditions. This indicator indirectly reflects the representativeness of the sampling points from the perspective of environmental similarity, and cannot directly reflect the representativeness of the attribute characteristics of the sampling points, which makes the measurement results prone to deviations, thereby affecting the accuracy and reliability of data analysis. Summary of the invention
[0006] The present invention provides a method and device for measuring the representativeness of sampling point attribute features, which are used to measure the representativeness of sampling point attribute features and evaluate the quality of sampling point data, thereby reducing the uncertainty of sampling points in the data analysis process and ensuring the accuracy and reliability of sampling point data mining and analysis.
[0007] The present invention provides a method for measuring the representativeness of sampling point attribute features, comprising the following steps: Determine the target sampling points in the sampling area for which attribute characteristic representativeness measurement is required based on the target attributes, wherein the target attributes are determined according to the research objectives; Determine the auxiliary attribute associated with the target attribute to obtain auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; Determine the Thiessen polygon where each target sampling point is located as a measurement unit for representative measurement of its attribute characteristics; Perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; The attribute category number ratio is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of the specific category in the measurement unit to which it belongs; The representative measurement value of the attribute feature of each target sampling point is determined by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the corresponding measurement unit, and the attribute weight of each auxiliary attribute.
[0008] According to a representative measurement method for attribute features of sampling points provided by the present invention, the auxiliary attribute data set and all the measurement units are spatially overlaid for analysis, and the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs are determined, including: Generate a corresponding auxiliary attribute data layer by taking each auxiliary attribute in the auxiliary attribute data set as an independent layer; Perform spatial overlay analysis on each auxiliary attribute data layer and the measurement unit layer to obtain the first i The number of categories and the i The total area of the category spots of each category in the auxiliary attribute; the measurement unit layer is composed of the measurement units of all the target sampling points; According to the measurement unit i The number of categories of the auxiliary attribute is related to the number of i The first ratio between the total number of categories of the auxiliary attributes is obtained in each of the measurement units. i The ratio of the number of attribute categories of the auxiliary attributes; Determine, according to the specific category of the auxiliary attribute at the spatial position of the target sampling point in the measurement unit to which it belongs, the category patch area where the specific category of the target sampling point in the measurement unit is located, and the total area of the category patches corresponding to the specific category in the measurement unit; Determining the area ratio of the category spots according to a second ratio between the area of the category spots and the total area of the category spots; in, i Is a positive integer.
[0009] According to a sampling point attribute feature representativeness measurement method provided by the present invention, before taking each auxiliary attribute in the auxiliary attribute data set as an independent layer to generate a corresponding auxiliary attribute data layer, the method further comprises: According to the spatial distribution characteristics of the continuous attribute data, each of the continuous attribute data in the auxiliary attribute data set is reclassified into categorical attribute data.
[0010] According to a sampling point attribute feature representativeness measurement method provided by the present invention, the mathematical calculation model for determining the attribute feature representativeness measurement value of each target sampling point is as follows, which comprehensively considers the attribute category number ratio of each auxiliary attribute in each measurement unit, the category patch area ratio of each auxiliary attribute in the corresponding measurement unit of each target sampling point, and the attribute weight of each auxiliary attribute: ; in, is the representative metric value of the attribute feature of the target sampling point y, is the number of units in the measurement unit where the target sampling point y is located. i The number of categories of auxiliary attributes; In the sampling area i The total number of categories of auxiliary attributes; n is the number of auxiliary attributes in the auxiliary attribute dataset; The target sampling point y is the first i The area ratio of the category patches of the auxiliary attributes; For the i The attribute weight corresponding to the auxiliary attribute.
[0011] According to a sampling point attribute feature representativeness measurement method provided by the present invention, the determining of the auxiliary attribute associated with the target attribute includes: Determine a candidate auxiliary attribute set, wherein the candidate auxiliary attribute set includes a plurality of continuous attributes and a plurality of categorical attributes; Calculating the correlation coefficient between each of the continuous attributes and the target attribute respectively, and selecting a preset number of continuous attributes from large to small according to the absolute values of the correlation coefficients as continuous auxiliary attributes; Each categorical attribute is set as an independent variable, the target attribute is set as a dependent variable, variance analysis is performed to obtain a significance probability value of each categorical attribute, and the categorical attributes whose significance probability values are less than a preset critical threshold are used as categorical auxiliary attributes.
[0012] According to a sampling point attribute feature representativeness measurement method provided by the present invention, the attribute weights of each auxiliary attribute are calculated based on a classification principal component analysis method, specifically including: Performing classified principal component analysis on the auxiliary attribute data set to generate a principal component loading matrix and eigenvalues; Determine the principal component whose eigenvalue is greater than 1 among all principal components as the effective principal component, so as to screen out the principal component load corresponding to the effective principal component from the principal component load matrix; Calculate the common factor variance of each auxiliary attribute based on the principal component loadings of all valid principal components, where the common factor variance is determined based on the square of the auxiliary attribute loading on each principal component; Based on the common factor variance of each auxiliary attribute, the attribute weight of each auxiliary attribute is determined.
[0013] According to a representative measurement method of sampling point attribute features provided by the present invention, the mathematical calculation model used to calculate the common factor variance is: ; in, For the t The common factor variance of auxiliary attributes, m is the number of valid principal components retained, For the t The auxiliary attribute is k The loadings corresponding to the effective principal components.
[0014] According to a sampling point attribute feature representativeness measurement method provided by the present invention, determining the attribute weight of each auxiliary attribute based on the common factor variance of each auxiliary attribute includes: Standardize the common factor variances of all auxiliary attributes; Determine the obtained standardized value of each auxiliary attribute as the attribute weight of each auxiliary attribute; The standardization process is one of linear normalization process, Z-score standardization process, and range standardization process.
[0015] The present invention also provides a sampling point attribute feature representativeness measurement device, comprising the following modules: A first processing unit is used to determine target sampling points in a sampling area that need to be measured for attribute characteristic representativeness based on target attributes, wherein the target attributes are determined according to a research objective; A second processing unit is used to determine an auxiliary attribute associated with the target attribute to obtain auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; A third processing unit is used to determine the Thiessen polygon where each target sampling point is located as a measurement unit for measuring the representativeness of its attribute characteristics; The fourth processing unit is used to perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; the ratio of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the ratio of the category patch area is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs and the total area of the category patch of the specific category in the measurement unit to which it belongs; The fifth processing unit is used to comprehensively determine the representative measurement value of the attribute feature of each target sampling point by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the corresponding measurement unit, and the attribute weight of each auxiliary attribute.
[0016] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the representativeness measurement method of attribute features of sampling points as described above is implemented.
[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the representativeness measurement method of attribute features of sampling points as described in any one of the above is implemented.
[0018] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned representativeness measurement methods for attribute features of sampling points.
[0019] The representativeness measurement method and device of sampling point attribute features provided by the present invention quantify the contribution of each auxiliary attribute related to the target attribute data to the target attribute by taking the Thiessen polygon of the sampling point as the measurement unit, and obtain the representativeness of the attribute features of each sampling point. The quality and availability of the sampling point data can be evaluated by the representativeness of the sampling point attribute features, and the uncertainty of the sampling point data can be reduced to ensure the accuracy and reliability of the specific analysis and practical application of the sampling point data. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 This is one of the flow charts of the representativeness measurement method of sampling point attribute features provided by the present invention.
[0022] Figure 2 This is the second flow chart of the representativeness measurement method of sampling point attribute features provided by the present invention.
[0023] Figure 3 It is a representative scatter diagram of the attribute characteristics of the sampling points provided by the present invention.
[0024] Figure 4 It is a structural schematic diagram of a representative measurement device for sampling point attribute characteristics provided by the present invention.
[0025] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] It should be noted that, in the description of the present invention, the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "includes a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0028] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0029] In the process of evaluating, analyzing and actually applying the target attribute data of the sampling points, due to the limitation of the sampling point data, there may be a situation where the number of sampling points in the study area is insufficient or the representativeness of the sampling points in the local study area is poor, which in turn affects the accuracy and reliability of the specific analysis and actual application of the target attribute data of the sampling points. It is necessary to use multiple other attribute data (referred to as auxiliary attributes in this invention) associated with the target attribute data to assist in the mining analysis to make up for the deficiencies in the target attribute data analysis. In this data analysis process, how to select the auxiliary attribute data associated with the target attribute and measure the representativeness of the attribute features of the sampling points is very critical. In view of the specific business needs of such application scenarios, the present invention proposes a sampling point attribute feature representativeness measurement method for measuring the attribute feature representativeness of the auxiliary attribute data associated with the target attribute data.
[0030] Combine the following Figure 1-Figure 5 The present invention describes a representative measurement method and device for sampling point attribute characteristics, in order to reduce the uncertainty of sampling points in the data analysis process by evaluating the quality of sampling point data, and ensure the accuracy and reliability of sampling point data mining analysis.
[0031] Figure 1 This is one of the flow charts of the representativeness measurement method of sampling point attribute features provided by the present invention, such as Figure 1 As shown, including but not limited to the following steps: Step 101 : determining target sampling points in a sampling area for which attribute feature representativeness measurement is required based on target attributes.
[0032] The target attribute is determined according to the research objective, and may be soil moisture, soil pH, or other soil attributes related to the research objective. For ease of description, this embodiment uses the attribute of soil fertility as the target attribute to describe the entire solution without further explanation, which is not considered as a specific limitation on the scope of protection of the present invention.
[0033] In this step, the target sampling points in the sampling area that need to be representatively measured for attribute characteristics are mainly determined based on the target attributes. These target sampling points are pre-selected according to the research objectives. For example, several sampling points can be selected as target sampling points in the sampling area through random sampling, systematic sampling or stratified sampling.
[0034] Step 102: determine the auxiliary attribute associated with the target attribute to obtain auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set.
[0035] In this embodiment, the attributes other than the target attributes that are related to the target attributes are collectively referred to as auxiliary attributes in the present invention. For example, attributes such as soil texture, vegetation coverage, terrain slope, etc. that are related to soil fertility are referred to as auxiliary attributes. These auxiliary attributes can provide additional information support for the measurement of soil fertility.
[0036] Based on the determined auxiliary attributes, the auxiliary attribute data of each target sampling point is obtained, and an auxiliary attribute data set is constructed. For example, auxiliary attribute data such as soil texture, vegetation coverage and terrain slope of each target sampling point are obtained through field measurement, remote sensing image interpretation or Geographic Information System (GIS) data query, and these auxiliary attribute data are organized into an auxiliary attribute data set for subsequent analysis.
[0037] Step 103: determine the Voronoi Diagram where each target sampling point is located as a measurement unit for measuring the representativeness of its attribute features.
[0038] Thiessen polygons are a distance-based space segmentation method. The core idea is that on a plane, the distance from any point in each Thiessen polygon to its corresponding sampling point is strictly less than the distance to any other sampling point. Through this segmentation, each sampling point is assigned an independent polygon area (i.e., measurement unit) to characterize the representativeness of the sampling point to the surrounding spatial attribute characteristics. The spatial size and shape of each measurement unit reflects the distribution law of the sampling point density and the surrounding environment characteristics.
[0039] The boundary of the Thiessen polygon is composed of the perpendicular bisectors of adjacent sampling points. All Thiessen polygons are seamlessly connected to cover the entire study area to avoid omissions or overlaps.
[0040] Step 104, performing spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs.
[0041] Specifically, this embodiment quantifies the distribution characteristics of each auxiliary attribute in each Thiessen polygon (measurement unit) through spatial overlay analysis, and provides basic data for the subsequent calculation of representative measurement values of attribute characteristics.
[0042] The ratio of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area, and is used to reflect the diversity of the auxiliary attributes in the measurement unit.
[0043] The category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of the specific category in the measurement unit to which it belongs. The sampling area is generally composed of multiple internally continuous patches, and the patches show different degrees of fragmented distribution. When measuring the representativeness of the attribute characteristics of the sampling points, the present invention takes into account the fragmented characteristics of the sampling area, and introduces the category patch area ratio to facilitate a more comprehensive evaluation of the representativeness of the attribute characteristics of the target sampling points. It integrates the category information of spatial distribution and auxiliary attributes, and can reflect the spatial coverage capability of the target sampling points.
[0044] Assuming that the study area contains several soil sampling points and the target attribute is soil fertility (with organic matter content as the core indicator), it is necessary to evaluate the distribution characteristics of the auxiliary attributes in each measurement unit through auxiliary attributes (such as pH value, soil texture, parent material, etc.), and then calculate the representativeness of the attribute characteristics of soil fertility.
[0045] If it is assumed that the auxiliary attribute of soil texture in the study area has four categories, including sandy loam, light loam, medium loam, and heavy loam, and a certain measurement unit contains two of these categories, such as sandy loam and light loam, then the proportion of attribute categories corresponding to the auxiliary attribute of soil texture in the measurement unit is 50%.
[0046] Furthermore, when calculating the area ratio of the category patch corresponding to the target sampling point, first determine the specific spatial position of the target sampling point in the measurement unit to which it belongs, and then determine the specific category of the auxiliary attribute at this spatial position. Assuming that the specific category of the auxiliary attribute of soil texture is sandy loam, it is necessary to count the total area of the category patch of sandy loam in the measurement unit to which the target sampling point belongs, assuming that the total area of this category patch is 50 km². Then, based on the ratio between the area of the category patch where the target sampling point is located (assuming it is 10 km²) and the total area of the category patch, the area ratio of the category patch corresponding to the target sampling point can be calculated to be 20%.
[0047] Based on the above examples, the ratio of the number of attribute categories of each auxiliary attribute in the measurement unit to which each target sampling point belongs and the ratio of the category patch area of each auxiliary attribute of each target sampling point can be determined.
[0048] Step 105, comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weight of each auxiliary attribute, to determine the representative measurement value of the attribute feature of each target sampling point.
[0049] Among them, the attribute weight can be determined according to the degree of correlation between the auxiliary attribute and the target attribute. For example, if the correlation between soil texture and soil fertility is high, the soil texture has a larger weight; if the correlation between vegetation coverage and soil fertility is low, the vegetation coverage has a smaller weight. Of course, other methods can also be used to more accurately calculate the attribute weight of each auxiliary attribute, which will be described in detail in the subsequent embodiments.
[0050] In this embodiment, a mathematical method such as weighted summation can be used to combine the ratio of the number of attribute categories of each auxiliary attribute, the ratio of the area of the category patches of each target sampling point in the measurement unit to which the auxiliary attribute belongs, and the attribute weight of the auxiliary attribute to calculate the attribute feature representativeness measurement value of each target sampling point. The attribute feature representativeness measurement value can reflect the representativeness of the target sampling point in the measurement unit to which it belongs relative to the target attribute. The larger the attribute feature representativeness measurement value, the stronger the representativeness of the target sampling point.
[0051] The representativeness measurement method of sampling point attribute features provided by the present invention quantifies the contribution of each auxiliary attribute related to the target attribute data to the target attribute by taking the Thiessen polygon of the sampling point as the measurement unit, and obtains the representativeness of the attribute features of each sampling point. The quality and availability of the sampling point data can be evaluated by the representativeness of the sampling point attribute features, and the uncertainty of the sampling point data can be reduced to ensure the accuracy and reliability of the specific analysis and practical application of the sampling point data.
[0052] Based on the content of the above embodiment, as an optional embodiment, the auxiliary attribute data set is spatially overlaid with all measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs, specifically including: Step 1: Generate a corresponding auxiliary attribute data layer by treating each auxiliary attribute in the auxiliary attribute dataset as an independent layer.
[0053] Step 2: Perform spatial overlay analysis on each auxiliary attribute data layer and the measurement unit layer to obtain the first i The number of categories and the i The total area of the category spots of each category in the auxiliary attributes. Wherein, the measurement unit layer is composed of the measurement units of all the target sampling points.
[0054] Step 3: According to the first i The number of categories of the auxiliary attribute is related to the number of i The first ratio between the total number of categories of the auxiliary attributes is obtained in each of the measurement units. i The ratio of attribute categories of auxiliary attributes.
[0055] Step 4: Determine the category patch area of the specific category of the target sampling point in the measurement unit and the total category patch area corresponding to the specific category in the measurement unit according to the specific category of the auxiliary attribute at the spatial position of the target sampling point in the measurement unit.
[0056] Step 5, determining the category patch area ratio according to the second ratio between the category patch area and the total category patch area.
[0057] in, i Is a positive integer.
[0058] In this embodiment, the specific execution steps for implementing spatial overlay analysis will be introduced in detail. The basic principle is to calculate the proportion of the attribute categories and the proportion of the category patch area of the auxiliary attributes in each measurement unit by superimposing the auxiliary attribute data layer and the measurement unit layer.
[0059] In order to perform spatial overlay analysis, it is first necessary to generate an independent auxiliary attribute data layer for each auxiliary attribute. Assuming that the auxiliary attributes include soil texture, vegetation coverage, and terrain slope, an auxiliary attribute data layer can be generated for each auxiliary attribute. These auxiliary attribute data layers can be generated through geographic information system (GIS) software or other related tools.
[0060] Furthermore, each auxiliary attribute data layer is spatially overlaid with the measurement unit layer. The measurement unit layer is composed of the Thiessen polygons (i.e., measurement units) of all target sampling points, which are used to characterize the spatial range represented by each target sampling point. Through spatial overlay analysis, the spatial distribution of each measurement unit can be obtained. i The number of categories and the i The total area of each category in the auxiliary attribute.
[0061] For example, soil texture, an auxiliary attribute, may contain multiple categories, such as sandy loam, light loam, medium loam, and heavy loam. Through spatial overlay analysis, the number of soil texture categories within each measurement unit and the total area of each category can be determined.
[0062] As an optional embodiment, assuming that the target attribute is soil fertility, the auxiliary attributes include soil texture (categories include sandy loam, light loam, and medium loam) and vegetation coverage (categories include high, medium, and low).
[0063] For the measurement unit of a target sampling point, through spatial overlay analysis, if the following data is obtained: the soil texture in the measurement unit includes two categories: sandy loam and light loam, and the number of categories is 2; the total area of the category patches of sandy loam in the measurement unit is 20 m², and the total area of the category patches of light loam is 30 m². The vegetation coverage in the measurement unit includes high vegetation coverage and low vegetation coverage, and the number of categories is 2; the total area of the category patches of high vegetation coverage in the measurement unit is 10 m², and the total area of the category patches of low vegetation coverage is 20 m².
[0064] Based on the data obtained from the above spatial overlay analysis, the ratio of the number of attribute categories of each auxiliary attribute in the measurement unit and the ratio of the category patch area of each auxiliary attribute of each target sampling point in the measurement unit can be calculated.
[0065] Among them, the proportion of attribute categories of the auxiliary attribute of soil texture is 2 / 3, and the proportion of attribute categories of the auxiliary attribute of vegetation coverage is 2 / 3.
[0066] Assuming that the category of the target sampling point in the measurement unit is determined to be sandy loam and high vegetation coverage, the sandy loam category patch area corresponding to the target sampling point in the measurement unit is 4 m², and the high vegetation coverage category patch area is 5 m². It can be calculated that the category patch area ratio under the auxiliary attribute of soil texture is 20% (calculated by 4 / 20), and the category patch area ratio under the auxiliary attribute of vegetation coverage is 50% (calculated by 5 / 10).
[0067] Finally, we can use mathematical methods such as weighted summation to calculate the representative measurement value of the attribute characteristics of each target sampling point by combining the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit, the proportion of the category patch area of each auxiliary attribute of each target sampling point in its corresponding measurement unit, and the attribute weight of each auxiliary attribute.
[0068] The representativeness measurement method of sampling point attribute features provided by the present invention can accurately calculate the proportion of the number of categories of each auxiliary attribute in each measurement unit and the proportion of the category patch area of the target sampling point in the measurement unit to which it belongs, by generating an independent data layer from the auxiliary attribute data and performing spatial overlay analysis with the Thiessen polygon (measurement unit). This refined spatial analysis method makes the representativeness measurement of the sampling points more accurate and can fully reflect the attribute distribution characteristics of the sampling points within their spatial range.
[0069] As an optional embodiment, before generating a corresponding auxiliary attribute data layer by taking each auxiliary attribute in the auxiliary attribute data set as an independent layer, the method further includes: According to the spatial distribution characteristics of the continuous attribute data, each continuous attribute data in the auxiliary attribute dataset is reclassified into categorical attribute data.
[0070] Since some auxiliary attribute data in the auxiliary attribute dataset are in continuous form, such as soil moisture (expressed as a percentage), soil organic matter content (expressed in grams / kilograms) or terrain slope (expressed in degrees), these continuous data have obvious gradient changes in spatial distribution, but it is difficult to directly reflect their hierarchical distribution characteristics.
[0071] In order to better conduct spatial overlay analysis, these continuous attribute data need to be reclassified into categorical data according to their spatial distribution characteristics. The specific method of reclassification can be determined according to the actual application scenario and research objectives. The reclassification results should reflect the spatial differences of the attribute data. In addition to custom classification intervals, one or a combination of the following reclassification methods can also be used: natural break point classification method, equal interval classification method, standard deviation classification method, and quantile classification method. The specific choice of reclassification method can be determined by combining factors such as the spatial distribution characteristics and value range of the attribute data.
[0072] Among them, the natural break point classification method refers to dividing continuous attribute data into several categories according to their natural distribution characteristics, so that the data differences within each category are minimized and the differences between categories are maximized.
[0073] The equal interval classification method refers to dividing the value range of continuous attribute data into several equally spaced intervals, with each interval of equal length. It can be applied to scenarios with a clear value range and a relatively uniform distribution, such as the classification of geographic data (such as terrain elevation, temperature distribution, etc.).
[0074] The standard deviation classification method is to classify data by calculating the mean and standard deviation of continuous attribute data. It can be applied to data with normal distribution or approximate normal distribution. Continuous attribute data is divided into several intervals, each of which is centered on the mean and separated by the standard deviation. It is suitable for scenarios where the data distribution is relatively uniform and conforms to the normal distribution, and can reflect the degree of data dispersion.
[0075] The quantile classification rule is to arrange the continuous attribute data in order of size and divide it into several equal intervals, each of which contains the same number of data points. It is suitable for scenarios where data distribution is uneven and relative positions need to be highlighted, such as the classification of plant density data.
[0076] After completing the reclassification of each continuous attribute data in the auxiliary attribute dataset, each auxiliary attribute is used as an independent layer to generate a corresponding auxiliary attribute data layer, which is more convenient for subsequent spatial overlay analysis.
[0077] For example, soil organic matter can be reclassified into three categories according to its content: "low (0-5 g / kg)", "medium (5-15 g / kg)" and "high (>15 g / kg)". After reclassification, a soil organic matter data layer is generated, which contains three categories: low, medium and high.
[0078] The representativeness measurement method of sampling point attribute features provided by the present invention converts continuous attribute data into categorical attribute data through reclassification processing, which not only simplifies the complexity of spatial analysis, but also can more intuitively reflect the spatial distribution characteristics of sampling points. This processing step provides a more reliable data basis for subsequent spatial overlay analysis.
[0079] As an optional embodiment, the mathematical calculation model for determining the representative measurement value of the attribute feature of each target sampling point by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the corresponding measurement unit, and the attribute weight of each auxiliary attribute can be: ; in, is the representative metric value of the attribute feature of the target sampling point y, is the number of units in the measurement unit where the target sampling point y is located. iThe number of categories of auxiliary attributes; In the sampling area i The total number of categories of auxiliary attributes; n is the number of auxiliary attributes in the auxiliary attribute dataset; The target sampling point y is the first i The area ratio of the category patches of the auxiliary attributes; For the i The attribute weight corresponding to the auxiliary attribute.
[0080] In the specific process of calculating the representative metric value of the attribute feature of any target sampling point in this embodiment, it is mainly implemented based on the following steps: Step 1: determine the ratio of the number of attribute categories of each auxiliary attribute in the target sampling point and the ratio of the category patch area of each auxiliary attribute in the measurement unit to which the target sampling point belongs.
[0081] Step 2: Determine the attribute weight of each auxiliary attribute, which reflects the contribution of different auxiliary attributes to the target attribute. The attribute weight can be determined by correlation analysis, expert experience or statistical methods (such as principal component analysis). In this embodiment, it is assumed that the weight of each auxiliary attribute has been calculated by a certain method and normalized so that its sum is 1.
[0082] Step 3, according to the mathematical calculation model of the representative measurement value of the attribute feature, the product of the ratio of the number of attribute categories of each auxiliary attribute in the measurement unit where the target sampling point is located and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs are successively accumulated. In the process of accumulation, the attribute weights of different auxiliary attributes are taken into account at the same time, so as to obtain the representative measurement value of the attribute feature of the target sampling point.
[0083] The following is a detailed description with reference to a specific embodiment.
[0084] Assuming the target attribute is soil fertility, the auxiliary attributes include soil texture (including 3 categories: sandy loam, light loam, medium loam) and vegetation cover (including 3 categories: high, medium, and low). The total number of categories of each auxiliary attribute in the sampling area is shown in Table 1: Table 1 List of auxiliary attribute data in the sampling area
[0085] At the same time, assuming that the target sampling point y The calculation data of the proportion of the number of attribute categories of each auxiliary attribute in the measurement unit is shown in Table 2: Table 2 Target sampling points y List of data related to the proportion of attribute categories
[0086] Furthermore, assuming that the target sampling point y The relevant data for calculating the area ratio of the category patches corresponding to the specific categories in the measurement unit are shown in Table 3: Table 3 Target sampling points y List of data related to the area ratio of category patches
[0087] Assume that the calculated weights of the auxiliary attributes are: The attribute weight of soil texture is P 1=0.6, the attribute weight of vegetation coverage is P 2=0.4, then the representative contribution value of each auxiliary attribute can be calculated. For example, the representative contribution value of the auxiliary attribute of soil texture is 2 / 3 20% 0.6=0.08. The representative contribution value of the auxiliary attribute of vegetation coverage is 2 / 3 50% 0.4≈0.13.
[0088] Finally, the representative contribution values of all the above auxiliary attribute categories are accumulated to obtain the target sampling point y Representative measure of attribute characteristics: D y =0.08+0.13=0.21.
[0089] Through the above calculation, the target sampling point can be obtained y The attribute feature representativeness measure value is 0.21.
[0090] The representativeness measurement method of sampling point attribute features provided by the present invention can comprehensively and objectively evaluate the representativeness of each target sampling point by comprehensively considering the proportion of the number of categories of each auxiliary attribute and the proportion of the category patch area of each auxiliary attribute, while taking into account the attribute weights of the auxiliary attributes, thereby providing a scientific basis for the evaluation of the quality of sample point data.
[0091] Figure 2 This is the second flow chart of the representativeness measurement method of sampling point attribute features provided by the present invention. Figure 2 As shown in the figure, a specific implementation process of a sampling point attribute feature representativeness measurement method is provided: Each target sampling point and its attribute data in the study area are obtained, and the attribute data are divided into target attribute data and other candidate auxiliary attribute data. In order to solve the problems of redundant variable interference, high computational complexity and poor model interpretability that may exist in the representative analysis of sampling point attribute characteristics, the present invention uses a statistical method to screen auxiliary attributes that are strongly correlated with the target attribute from a coarse-grained set of candidate auxiliary attributes, including correlation analysis or variance analysis of the target attribute and the candidate auxiliary attributes to screen out multiple auxiliary attributes related to the target attribute.
[0092] As an optional embodiment, the step of determining the auxiliary attribute associated with the target attribute mainly includes: A candidate auxiliary attribute set is determined, wherein the candidate auxiliary attribute set includes a plurality of continuous attributes and a plurality of categorical attributes.
[0093] The correlation coefficient between each continuous attribute and the target attribute is calculated respectively, and a preset number of continuous attributes are selected from large to small according to the absolute values of the correlation coefficients as continuous auxiliary attributes.
[0094] Each categorical attribute is set as an independent variable, and the target attribute is set as a dependent variable. Analysis of Variance (ANOVA) is performed to obtain the significance probability value (P value for short) of each categorical attribute, and the categorical attributes with P values less than the preset critical threshold are used as categorical auxiliary attributes.
[0095] As an optional embodiment, nine continuous attributes, namely soil thickness, soil bulk density, elevation, slope, aspect, normalized difference vegetation index (NDVI), average annual temperature, average annual precipitation, and sunshine hours, and three categorical attributes, namely parent material, soil type, and soil texture, are combined to form a set of candidate auxiliary attributes. Soil fertility is used as the target attribute for correlation analysis, and auxiliary attributes closely related to the target attribute are screened out.
[0096] Specifically, for the screening step of the continuous attributes, the screening is mainly achieved by calculating the correlation coefficient between the target attribute and each continuous attribute. Its mathematical calculation model can be expressed as: ; in, Continuous attribute x With target attributes y The correlation coefficient between For the i The continuous attribute value of the sampling points, For the i The target attribute value of each sampling point, is the mean of the continuous attributes of all sampling points, is the mean of the target attributes of all sampling points, n is the total number of sampling points.
[0097] After calculating the correlation coefficient corresponding to each continuous attribute, the correlation coefficients can be sorted from large to small according to their absolute values, and a preset number of continuous attributes that are significantly correlated with the target attribute can be selected as auxiliary attributes.
[0098] The screening of categorical auxiliary attributes is mainly done through ANOVA analysis to evaluate whether the influence of different categorical attributes on the target attribute is significant. If the mean values of different categories of a categorical attribute on the target attribute are significantly different, then this categorical attribute can be considered to be related to the target attribute and can be used as an auxiliary attribute.
[0099] Perform appropriate preprocessing on the sample point sampling data, such as missing value processing, outlier detection, etc., and then use statistical software (such as SPSS, R, Python, etc.) to perform ANOVA analysis. Taking SPSS software as an example, the main operation steps include: In SPSS software, select the Compare Means option under the Analysis menu, and finally select One-Way ANOVA and run it. Then, you can set the target attribute as the dependent variable, and by setting different categorical attributes as independent variables in turn, you can calculate the P value corresponding to each categorical attribute.
[0100] According to the results of ANOVA analysis, categorical attributes with P values less than the preset significance level are selected (evaluated by the preset critical threshold). These attributes are considered to be auxiliary attributes related to the target attribute and can be used for subsequent attribute feature representativeness measurement.
[0101] The representativeness measurement method of attribute features of sampling points provided by the present invention can systematically screen out continuous and categorical auxiliary attributes that are significantly correlated with the target attributes through correlation coefficient and ANOVA analysis, thereby ensuring that the selected auxiliary attributes can be effectively used to evaluate the representativeness of attribute features of sampling points and ensuring the accuracy of data analysis results.
[0102] refer to Figure 2 As shown, the following briefly introduces how to calculate the weights of each auxiliary attribute related to the target attribute in the sampling point attribute feature representativeness measurement method provided by the present invention.
[0103] Considering that principal component analysis (PCA) is only applicable to data analysis of continuous attributes, and categorical principal component analysis (CATPCA) uses optimal scale change to convert classification labels into numerical values, and ensures that the variance of the variables converted into quantization is maximized, and further uses the quantitative data to reduce the overall dimension of the data, it is applicable to data analysis containing categorical attributes. Therefore, the present invention adopts the method based on categorical principal component analysis to calculate the attribute weights of each auxiliary attribute, which mainly includes but is not limited to the following steps: The auxiliary attribute data set is subjected to a categorical principal component analysis to generate a principal component loading matrix and eigenvalues. The categorical principal component analysis can identify the main variation directions in the data and project the original high-dimensional data into a lower-dimensional space while retaining the information of the original data as much as possible.
[0104] The principal components with eigenvalues greater than 1 among all the principal components are determined as valid principal components to screen out the principal component loadings of the valid principal components. This is because the principal components with eigenvalues greater than 1 are considered to retain more information than a single original variable and are therefore considered valid.
[0105] Furthermore, the common factor variance of each auxiliary attribute is calculated according to the screened principal component loading matrix. The common factor variance is determined according to the square of the auxiliary attribute loading on each principal component, reflecting the contribution of each auxiliary attribute in the principal component analysis. The calculation formula can be expressed as: ; in, For the t The common factor variance of auxiliary attributes, m is the number of valid principal components retained, For the t The auxiliary attribute is k The loadings corresponding to the effective principal components.
[0106] Finally, based on the common factor variance of each auxiliary attribute, the attribute weight corresponding to each auxiliary attribute is determined.
[0107] As an optional embodiment, the determining the attribute weight of each auxiliary attribute based on the common factor variance of each auxiliary attribute may include the following steps: Step 1: Standardize the common factor variance of all auxiliary attributes.
[0108] Step 2: determining the obtained standardized value of each auxiliary attribute as the attribute weight of each auxiliary attribute.
[0109] The standardization process used may be one of linear normalization, Z-score standardization, and range standardization.
[0110] Taking linear normalization as an example, it is ensured that the sum of all attribute weights is 1, so that the attribute weight can directly reflect the relative contribution of each auxiliary attribute.
[0111] The normalization formula used in the present invention may be: ; in, For the i The attribute weight of auxiliary attributes, n is the total number of auxiliary attributes, For the i The weight of the auxiliary attribute.
[0112] The following example illustrates how to calculate the attribute weight of each auxiliary attribute.
[0113] Assume that the auxiliary attribute data set includes three auxiliary attributes: soil moisture, soil organic matter, and soil type. Through classification principal component analysis, the principal component load matrix and eigenvalues are obtained. The eigenvalues of principal component 1, principal component 2, and principal component 3 are 1.3, 1.1, and 0.4, respectively. The eigenvalues of the first two principal components are greater than 1, that is, the first two principal components are effective principal components. The principal component load matrix obtained can be shown in Table 4.
[0114] Table 4 Principal component loading matrix
[0115] Furthermore, the common factor variance of each auxiliary attribute is calculated according to the principal component loading matrix. The common factor variance is the sum of the squares of the auxiliary attribute loadings on each principal component.
[0116] For the common factor variance of soil moisture f The calculation formula for 1 is: f 1=0.6 2 +0.4 2 =0.52.
[0117] Common factor variance for soil organic matter f The calculation formula for 2 is: f 2=0.7 2 +0.3 2 =0.58.
[0118] Common factor variance for soil type f The calculation formula for 3 is: f 3=0.5 2 +0.5 2 =0.50.
[0119] Then, based on the common factor variance of each auxiliary attribute, the weight attribute of each auxiliary attribute is calculated. To simplify the calculation, in this embodiment, the common factor variance and the attribute weight are directly set to a linear positive proportional relationship.
[0120] Specifically, we can first calculate the total common factor variance = 0.52 + 0.58 + 0.50 = 1.60, and then calculate the attribute weight of each auxiliary attribute separately: The auxiliary weight of soil moisture is 0.52 / 1.60, and the result is 0.3250; The auxiliary weight of soil organic matter is 0.58 / 1.60, and the result is 0.3625; The auxiliary weights for soil type are 0.50 / 1.60, which results in 0.3125.
[0121] The representativeness measurement method of sampling point attribute features provided by the present invention calculates auxiliary attribute weights by a classification principal component analysis method, and effectively quantifies the contribution of each auxiliary attribute to the target attribute.
[0122] Combine the following Figure 2 As shown, a complete embodiment is used to illustrate the specific implementation process of a sampling point attribute feature representativeness measurement method provided by the present invention.
[0123] Step 1: Obtain data from 433 target sampling points in a sampling area, which contains a total of 13 soil attributes. Take organic matter content as the target attribute and the remaining 12 as auxiliary attributes to form a candidate auxiliary attribute set. Among them, there are 9 continuous attributes: soil thickness, soil bulk density, elevation, slope, slope aspect, normalized vegetation index, average annual temperature, average annual precipitation, sunshine hours, and 3 categorical attributes: soil parent material, soil type, and soil texture.
[0124] Furthermore, based on correlation analysis, continuous variables were screened to obtain five continuous auxiliary attributes associated with the target attribute of organic matter content: soil thickness, soil bulk density, elevation, normalized difference vegetation index, and average annual precipitation. Based on variance analysis, categorical variables were screened to obtain three categorical auxiliary attributes associated with the target attribute data: soil parent material, soil type, and soil texture, that is, 8 auxiliary attributes related to the target attribute were finally screened out.
[0125] Step 2: Generate corresponding Thiessen polygons based on the above 433 target sampling points, and spatially clip to obtain the Thiessen polygons where each target sampling point is located, that is, obtain a measurement unit for measuring the representativeness of the attribute characteristics of each target sampling point. The generation of Thiessen polygons and spatial clipping operations can be implemented based on ArcGIS software.
[0126] Step 3, using the method provided in the above embodiment, calculate the attribute weights of 8 auxiliary attributes related to the target attribute of the target sampling point. Auxiliary attributes include continuous attributes and categorical attributes, so the classification principal component analysis method is used to calculate the attribute weights. After selecting the principal component with an eigenvalue greater than 1, the common factor variance is calculated according to the principal component loading matrix; then the common factor variance is standardized to calculate the attribute weights of soil parent material, soil type, soil texture, soil thickness, soil bulk density, elevation, normalized vegetation index, and average annual precipitation, respectively: 0.1020, 0.1311, 0.1374, 0.1285, 0.1358, 0.1461, 0.0682, 0.1509.
[0127] Step 4: According to the spatial distribution characteristics and value range of the data, the five continuous attributes are reclassified into categorical attributes, including: soil thickness is divided into 7 categories, soil bulk density is divided into 6 categories, normalized vegetation index is divided into 4 categories, elevation is divided into 7 categories, and average annual precipitation is divided into 5 categories.
[0128] Then, based on the 433 Thiessen polygons generated by the target sampling points, eight auxiliary attribute data layers, including parent material, soil type, soil texture, soil thickness, soil bulk density, elevation, normalized difference vegetation index, and average annual precipitation, were spatially overlaid with the Thiessen polygons (metric unit layer), and the proportion of the number of attribute categories of each auxiliary attribute in each Thiessen polygon and the proportion of the category patch area of each auxiliary attribute of each target sampling point in the corresponding metric unit were calculated.
[0129] Finally, the representative measurement values of the attribute features of 433 target sampling points were calculated.
[0130] Figure 3 is a representative scatter diagram of the attribute characteristics of the sampling points provided by the present invention, such as Figure 3 As shown in the figure, the maximum value of the attribute feature representativeness measurement value of the 433 target sampling points is 0.6701, and the minimum value is 0.1706. It can be seen that the attribute feature representativeness of two target sampling points is obviously low. The next step is to conduct specific analysis or data refinement on these two target sampling points.
[0131] Through the above embodiments, the representativeness measurement method of sampling point attribute features provided by the present invention is fully explained. Thiessen polygons are generated through sampling point data, spatial cutting is performed to obtain the representativeness measurement unit of the attribute features of each sampling point, the attribute weights of the auxiliary attributes related to the target attributes are calculated, and then the representativeness of the attribute features of each target sampling point in the study area is calculated, and a representative scatter plot of the attribute features of the sampling points is drawn to measure the quality of the sampling point data. The quality and availability of the sampling point data are evaluated through the representativeness of the attribute features of the sampling points, the uncertainty of the sampling point data is reduced, and the accuracy and reliability of the specific analysis and practical application of the sampling point data can be effectively guaranteed.
[0132] Figure 4 is a schematic diagram of the structure of the representative measurement device for sampling point attribute characteristics provided by the present invention, such as Figure 4 As shown, it mainly includes the following components: The first processing unit 41 is mainly used to determine the target sampling points in the sampling area that need to be measured for attribute characteristic representativeness based on the target attributes, and the target attributes are determined according to the research objectives.
[0133] The second processing unit 42 is mainly used to determine the auxiliary attribute associated with the target attribute to obtain the auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set.
[0134] The third processing unit 43 is mainly used to determine the Thiessen polygon where each target sampling point is located as a measurement unit for measuring the representativeness of its attribute characteristics.
[0135] The fourth processing unit 44 is mainly used to perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; the attribute category ratio is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of the specific category in the measurement unit to which it belongs.
[0136] The fifth processing unit 45 is mainly used to comprehensively determine the representative measurement value of the attribute feature of each target sampling point by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weight of each auxiliary attribute.
[0137] It should be noted that the sampling point attribute feature representativeness measurement device provided by the present invention can execute the sampling point attribute feature representativeness measurement method described in any of the above embodiments during specific operation, which will not be described in detail in this embodiment.
[0138] The sampling point attribute feature representativeness measurement device provided by the present invention quantifies the contribution of each auxiliary attribute related to the target attribute data to the target attribute by taking the Thiessen polygon of the sampling point as the measurement unit, and obtains the attribute feature representativeness of each sampling point. The sampling point data quality and availability can be evaluated through the sampling point attribute feature representativeness, and the uncertainty of the sampling point data can be reduced to ensure the accuracy and reliability of the specific analysis and practical application of the sampling point data.
[0139] Figure 5 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 5 As shown, the electronic device may include: a processor (Processor) 510, a communication interface (Communications Interface) 520, a memory (Memory) 530 and a communication bus (Communication Bus) 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logic instructions in the memory 530 to execute the sampling point attribute feature representativeness measurement method, the method comprising: determining the target sampling points in the sampling area that need to be measured for attribute feature representativeness based on the target attribute, the target attribute is determined according to the research target; determining the auxiliary attributes associated with the target attribute to obtain the auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; determining the Thiessen polygon where each of the target sampling points is located as the measurement unit for measuring the attribute feature representativeness; performing spatial overlay analysis on the auxiliary attribute data set and all the measurement units, determining the proportion of the number of attribute categories of each auxiliary attribute in each measurement unit and the The ratio of the category patch areas of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; the attribute category number ratio is determined based on the category number of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of the specific category in the measurement unit to which it belongs; the representative measurement value of the attribute feature of each target sampling point is determined by comprehensively considering the attribute category number ratio of each auxiliary attribute in each measurement unit, the category patch area ratio of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weight of each auxiliary attribute.
[0140] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0141] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the sampling point attribute feature representativeness measurement method provided in the above embodiments, and the method includes: determining the target sampling points in the sampling area that need to be measured for attribute feature representativeness based on the target attribute, and the target attribute is determined according to the research objective; determining the auxiliary attributes associated with the target attribute to obtain the auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; determining the Thiessen polygon where each of the target sampling points is located as a measurement unit for measuring its attribute feature representativeness; comparing the auxiliary attribute data set with all the measurement units A spatial overlay analysis is performed to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; the ratio of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the ratio of the category patch area is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of the specific category in the measurement unit to which it belongs; the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weight of each auxiliary attribute are comprehensively considered to determine the representative measurement value of the attribute feature of each target sampling point.
[0142] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the processor executes the sampling point attribute feature representativeness measurement method provided in the above embodiments, the method comprising: determining the target sampling points in the sampling area that need to be measured for attribute feature representativeness based on the target attribute, wherein the target attribute is determined according to the research objective; determining the auxiliary attribute associated with the target attribute to obtain the auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; determining the Thiessen polygon where each of the target sampling points is located as the measurement unit for measuring the attribute feature representativeness of the target sampling point; performing spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the auxiliary attribute data set of each measurement unit. The ratio of the number of attribute categories of each auxiliary attribute in the unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the unit of measurement to which it belongs; the ratio of the number of attribute categories is determined based on the number of categories of each auxiliary attribute in the unit of measurement and the total number of categories of the same auxiliary attribute in the sampling area; the ratio of the category patch area is determined based on the category patch area corresponding to the specific category of the target sampling point in the unit of measurement to which it belongs, and the total category patch area of the specific category in the unit of measurement to which it belongs; the representative measurement value of the attribute feature of each target sampling point is determined by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each unit of measurement, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the unit of measurement to which it belongs, and the attribute weight of each auxiliary attribute.
[0143] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0144] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for measuring the representativeness of sampling point attribute features, characterized in that: include: Determine the target sampling points in the sampling area for which attribute characteristic representativeness measurement is required based on the target attributes, wherein the target attributes are determined according to the research objectives; Determine the auxiliary attribute associated with the target attribute to obtain auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; Determine the Thiessen polygon where each target sampling point is located as a measurement unit for representative measurement of its attribute characteristics; Perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; The attribute category number ratio is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of the specific category in the measurement unit to which it belongs; The representative measurement value of the attribute feature of each target sampling point is determined by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the corresponding measurement unit, and the attribute weight of each auxiliary attribute.
2. The representativeness measurement method of sampling point attribute features according to claim 1 is characterized in that: The auxiliary attribute data set is spatially overlaid with all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs, including: Generate a corresponding auxiliary attribute data layer by taking each auxiliary attribute in the auxiliary attribute data set as an independent layer; Perform spatial overlay analysis on each auxiliary attribute data layer and the measurement unit layer to obtain the first i The number of categories and the i The total area of the category spots of each category in the auxiliary attribute; the measurement unit layer is composed of the measurement units of all the target sampling points; According to the measurement unit i The number of categories of the auxiliary attribute is related to the number of i The first ratio between the total number of categories of the auxiliary attributes is obtained in each of the measurement units. i The ratio of the attribute categories of the auxiliary attributes; Determine, according to the specific category of the auxiliary attribute at the spatial position of the target sampling point in the measurement unit to which it belongs, the category patch area where the specific category of the target sampling point in the measurement unit is located, and the total area of the category patches corresponding to the specific category in the measurement unit; Determining the area ratio of the category spots according to a second ratio between the area of the category spots and the total area of the category spots; in, i Is a positive integer.
3. The representativeness measurement method of sampling point attribute features according to claim 2 is characterized in that: Before generating a corresponding auxiliary attribute data layer by taking each auxiliary attribute in the auxiliary attribute data set as an independent layer, the method further includes: According to the spatial distribution characteristics of the continuous attribute data, each of the continuous attribute data in the auxiliary attribute data set is reclassified into categorical attribute data.
4. The representativeness measurement method of sampling point attribute features according to claim 1 is characterized in that: The mathematical calculation model for determining the representative measurement value of the attribute feature of each target sampling point by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the measurement unit to which it belongs, and the attribute weight of each auxiliary attribute is: ; in, is the representative measurement value of the attribute feature of the target sampling point y, is the number of units in the measurement unit where the target sampling point y is located. i The number of categories of auxiliary attributes; In the sampling area i The total number of categories of auxiliary attributes; n is the number of auxiliary attributes in the auxiliary attribute dataset; The target sampling point y is the first i The area ratio of the category patches of the auxiliary attributes; For the i The attribute weight corresponding to the auxiliary attribute.
5. The representativeness measurement method of sampling point attribute features according to claim 1 is characterized in that: The determining of the auxiliary attribute associated with the target attribute comprises: Determine a candidate auxiliary attribute set, wherein the candidate auxiliary attribute set includes a plurality of continuous attributes and a plurality of categorical attributes; Calculating the correlation coefficient between each of the continuous attributes and the target attribute respectively, and selecting a preset number of continuous attributes from large to small according to the absolute values of the correlation coefficients as continuous auxiliary attributes; Each categorical attribute is set as an independent variable, the target attribute is set as a dependent variable, variance analysis is performed to obtain a significance probability value of each categorical attribute, and the categorical attributes whose significance probability values are less than a preset critical threshold are used as categorical auxiliary attributes.
6. The representativeness measurement method of sampling point attribute features according to any one of claims 1 to 5, characterized in that: The attribute weights of each auxiliary attribute are calculated based on the classification principal component analysis method, including: Performing classified principal component analysis on the auxiliary attribute data set to generate a principal component loading matrix and eigenvalues; Determine the principal component whose eigenvalue is greater than 1 among all principal components as the effective principal component, so as to screen out the principal component load corresponding to the effective principal component from the principal component load matrix; Calculate the common factor variance of each auxiliary attribute based on the principal component loadings of all valid principal components, where the common factor variance is determined based on the square of the auxiliary attribute loading on each principal component; Based on the common factor variance of each auxiliary attribute, the attribute weight of each auxiliary attribute is determined.
7. The representativeness measurement method of sampling point attribute features according to claim 6 is characterized in that: The mathematical calculation model used to calculate the common factor variance is: ; in, For the t The common factor variance of auxiliary attributes, m is the number of valid principal components retained, For the t The auxiliary attribute is k The loadings corresponding to the effective principal components.
8. The representativeness measurement method of sampling point attribute features according to claim 6, characterized in that: The determining the attribute weight of each auxiliary attribute based on the common factor variance of each auxiliary attribute includes: Standardize the common factor variances of all auxiliary attributes; Determine the obtained standardized value of each auxiliary attribute as the attribute weight of each auxiliary attribute; The standardization process is one of linear normalization process, Z-score standardization process, and range standardization process.
9. A device for measuring the representativeness of sampling point attribute features, characterized in that: include: A first processing unit is used to determine target sampling points in a sampling area that need to be measured for attribute characteristic representativeness based on target attributes, wherein the target attributes are determined according to a research objective; A second processing unit is used to determine an auxiliary attribute associated with the target attribute to obtain auxiliary attribute data of each target sampling point to construct an auxiliary attribute data set; A third processing unit is used to determine the Thiessen polygon where each target sampling point is located as a measurement unit for measuring the representativeness of its attribute characteristics; A fourth processing unit is used to perform spatial overlay analysis on the auxiliary attribute data set and all the measurement units to determine the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit and the ratio of the category patch area of each auxiliary attribute of the target sampling point in the measurement unit to which it belongs; The attribute category number ratio is determined based on the number of categories of each auxiliary attribute in the measurement unit and the total number of categories of the same auxiliary attribute in the sampling area; the category patch area ratio is determined based on the category patch area corresponding to the specific category of the target sampling point in the measurement unit to which it belongs, and the total category patch area of the specific category in the measurement unit to which it belongs; The fifth processing unit is used to comprehensively determine the representative measurement value of the attribute feature of each target sampling point by comprehensively considering the ratio of the number of attribute categories of each auxiliary attribute in each measurement unit, the ratio of the category patch area of each auxiliary attribute of each target sampling point in the corresponding measurement unit, and the attribute weight of each auxiliary attribute.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the representativeness measurement method of sampling point attribute features according to any one of claims 1 to 8 is implemented.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the representativeness measurement method of sampling point attribute features according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the representativeness measurement method of sampling point attribute features according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Mapping soil properties with satellite data using machine learning approaches
CN113196294A
Sample point layout method and device for machine learning space prediction model, and medium
CN118656634A
Medical reachability analysis method and system for urban inland inundation emergencies
CN118737405A
Fault-tolerant method for improving underwater robot networking robustness
CN118741573A
Cultivation assistance device, cultivation assistance method, and recording medium for storing program
US20160179779A1