Coastal wetland soil sample point quality evaluation method and system based on multi-source environmental data
Through the multi-source data fusion evaluation method, the problems of representativeness and consistency evaluation of coastal wetland soil sample data were solved, high-quality soil sample screening and carbon storage assessment were achieved, and the reliability of the data and the evaluation accuracy were improved.
Patent Information
- Application Number
- CN202511222949.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing technologies lack systematic and quantitative methods for evaluating the representativeness and consistency of coastal wetland soil sample data, resulting in large data errors and affecting the accuracy and interpretability of soil property assessments.
An evaluation method based on multi-source environmental data is adopted, integrating remote sensing images, climate and topographic data, and historical land cover information. The quality of sample points is evaluated through quantification in three dimensions: environmental similarity, historical consistency, and source reliability. Weighted Mahalanobis distance, multi-temporal coverage sequences, and metadata completeness are used to evaluate the quality of the samples.
It significantly improved the reliability and accuracy of soil sample data, reduced misjudgments caused by factor collinearity and scale inconsistency, eliminated unreliable samples, and improved the basic reliability of the soil database and the accuracy of carbon storage assessment.
Smart Images

Figure CN120744705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ecological environment monitoring and geographic information technology, and in particular to a coastal wetland soil sample quality assessment method and system based on multi-source environmental data. Background Art
[0002] Coastal wetlands, as a crucial component of "blue carbon" ecosystems, possess significant carbon sink and sequestration potential. However, their high spatial heterogeneity and rapid temporal variability (due to tidal disturbances, sediment transport, and seawater intrusion) have long posed challenges in the representativeness and consistency of soil sample data. Research and management practices generally rely on a vast array of environmental covariates (remote sensing, climate, topography, and oceanographic processes) to map soil properties and assess carbon storage. However, a systematic, quantitative, and replicable process is lacking for assessing the usability and credibility of sample sites.
[0003] Taking the global digital soil mapping system Soil Grids as an example, it uses more than 230,000 profile points and 400+ environmental layers for machine learning modeling and uncertainty quantification. Although it has significantly improved the supply of spatial information on soil properties (including SOC), its quality control mainly focuses on uncertainty quantification and point standardization at the model level. The integrated quality scoring of "environmental similarity, historical consistency, and source reliability" of multi-source heterogeneous sample points is still not a core goal.
[0004] Regarding broad-scale basic soil databases, ISRIC's WISE30sec and its technical reports systematically provide global-scale estimates and uncertainties for multiple attributes (including organic carbon). However, these also primarily rely on statistical inference of profile-climate / soil type combinations, emphasizing data fusion, attribute transfer, and outlier control. They lack an independent, portable quality scoring framework for the "suitability and representativeness" of sample sites for highly spatiotemporally dynamic ecological units such as coastal wetlands. FAO's HWSDv2.0 also replaces the older WISE version with WISE30sec data, emphasizing deep hierarchical refinement and attribute consistency, but does not address measures of sample reliability based on historical cover stability or coastal characteristics such as tidal / shoreline migration.
[0005] In terms of data ecology, the coastal wetland sector has recently developed data hubs such as the Coastal Carbon Network (CCN) and the Coastal Carbon Atlas, led by the Smithsonian Institution. These platforms open up tidal wetland soil core / profile data and metadata to the public, and provide quality tier fields for easy retrieval and downloading, promoting data integration and reproducibility in blue carbon research. However, the "quality tiers" of these platforms are more of a labeling and consensus standard for data availability. They have yet to calculate a single-point "quality index" based on remote sensing, topographic, climate, and oceanographic factors and historical cover change information, nor have they developed automated site-specific suitability assessments.
[0006] At the remote sensing method level, multi-source / multi-sensor fusion (optical, SAR, LiDAR / GEDI) has been shown to significantly improve wetland identification and classification accuracy (for example, random forests integrating terrain and radar information at a 10-meter resolution achieve classification accuracy exceeding 90%), demonstrating the feasibility of combining environmental covariates with machine learning for wetland monitoring. However, these efforts primarily address the question of "where are the wetlands and what classification should they be classified into?" Few studies have used this information retrospectively to assess the quality of existing soil samples, such as whether they match their environmental context or historical evolution.
[0007] A rich body of research has emerged in the academic community for remote sensing inversion of soil organic carbon (SOC), ranging from multispectral / hyperspectral indices (such as NDVI and CAI) to machine learning / deep learning models. This research has also focused on the impact of surface conditions (crop residue, moisture content, and roughness) on inversion stability and their correction. This has provided a basis for understanding the consistency between environmental indicators and soil properties. However, this research has primarily focused on improving the accuracy of estimates, with few studies further examining the behavior of remote sensing-ground-measurement residuals to assess the reliability of sample sources or recording processes.
[0008] In this context, direct use of unscreened sample data for spatial modeling or mechanism analysis can easily introduce errors, weakening the accuracy and interpretability of the results. Therefore, there is an urgent need to establish a systematic, quantifiable, and multi-source data fusion-based sample quality assessment method. This method integrates remote sensing imagery, environmental factors, and historical land cover information to scientifically eliminate or correct unreliable samples and ensure the reliability of the soil database foundation. Summary of the Invention
[0009] In order to solve the above technical problems, the purpose of the present invention is to provide a coastal wetland soil sample quality assessment method based on multi-source environmental data. By integrating remote sensing images, climate and topographic data, historical land cover maps and other environmental factor information, the environmental similarity, historical consistency and source reliability of the samples are comprehensively evaluated, providing a high-quality data foundation for soil mapping and mechanism modeling.
[0010] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions: A method for evaluating the quality of coastal wetland soil samples based on multi-source environmental data includes the following steps: S1 Sample data preparation: Collect basic information of soil samples and integrate remote sensing images, climate and terrain factors, and historical land cover data; S2 Extraction of multi-source environmental factors: Extraction of remote sensing indicators, topographic factors, ocean factors, climate factors and historical layers according to the sampling point location; S3 environment similarity scoring: in the presence of a reference set Sref When , the environmental similarity score is obtained by normalizing the weighted Mahalanobis distance; S ref When the environmental similarity score is obtained in the joint geographic and environmental space, K-nearest neighbor weighted Euclidean is used; S4 historical consistency score: stability is calculated based on multi-phase coverage sequences and corrected by the transfer penalty matrix α to obtain S hist ; S5 Source Reliability Score: Based on metadata integrity and source credibility, combined with anomaly detection and NDVI-SOC residual consistency correction, we can obtain S source ; S6 adaptive weight configuration: weights are constructed based on regional historical coverage variability, shoreline migration and interannual climate variation ω sim 、ω hist 、ω source , and the sum of the three is 1; comprehensive scoring and grading through weighted summation model: calculation S total =ω sim · S sim +ω hist · S hist +ω source · S source , the grading threshold is A:S total ≥0.80, B:0.60≤S total <0.80, C:S total <0.60; S7 result output and screening: Generate S sim 、 S hist 、 S source 、 S total and levels of tables and maps, radar charts, and distribution charts.
[0011] Preferably, in step S2, Remote sensing indicators: NDVI, mNDWI, EVI, land cover type; Topographic factors: elevation, slope, distance from the sea; Ocean factors: tidal range, sediment concentration, and seawater intrusion frequency; Climate factors: multi-year average temperature, precipitation, and potential evapotranspiration; Historical layer: Multiple remote sensing images or maps to determine the historical stability of the wetland status at the location.
[0012] Preferably, step S3 includes the following steps: Step 1: Construct a multidimensional environment vector For each sample point P to be evaluated i , construct a vector X containing its multidimensional environment characteristics i ; Each dimension x of the vector ij Representative j Environmental factors at sample point P i The value at the sample point P i The environment vector is expressed as: X i =[x i1 , x i2 ,…,x im ],in, m is the total number of environmental factors and is standardized using Z-score; Step 2: Assessing environmental similarity In the standardized multidimensional environment, weighted Mahalanobis distance is used to evaluate the similarity between sample points; Assume that the sample point to be evaluated is P i , high-confidence reference sample set S ref Define the sample point P as i With the reference sample set S ref The environmental similarity score S sim (P i , S ref )for: ; Where: X i ' and X k ' The sample points P i and reference sample P k The normalized environment vector Σ -1 is the reference sample set S ref The inverse matrix of the environmental vector covariance matrix; W=diag(w jj ) is a diagonal weight matrix, the diagonal elements w jj Represents the weight of the jth environmental factor and satisfies ; Indicates taking the average of all the samples in the reference sample set; Alternatively, calculate the sample points to be evaluated P i Its neighboring sample points P j The similarity between them is calculated and averaged; a variant of weighted Euclidean distance is used, and the calculation method is as follows: ; Among them, w k is the weight of the kth environmental factor, k is the environmental factor index, k=1,…,m; σ k i is the standard deviation of the factor after standardization; x ik i 、x jk i The sample points P j 、P i The value of the kth environmental factor; Step 3: Generate environment similarity score The similarity measure values calculated above are normalized to the interval [0,1] to obtain the environmental similarity score of the sample point; the higher the score, the more similar the environment of the sample point is to the reference environment, and the higher the data quality.
[0013] Preferably, step S4 includes the following steps: Step 1: Build historical time series data Acquire multiple historical remote sensing images covering the study area; use remote sensing image interpretation and classification algorithms to classify land cover in each phase, generating classification maps that include wetlands, water bodies, farmland, cities, and bare land; perform spatial overlay analysis on the precise geographic coordinates of all sample points and the land cover classification maps of different phases; and extract the land cover category of each sample point at each phase. Step 2: Quantify historical consistency For a sample point P i , and its land cover type sequence in T historical phases is {c i1 c i2 ,…,c iT } ; Define the historical consistency score S of the sample point hist (P i )for: ; in, f(c it ,c i(t-1) ) is an indicator function: ; If the sample points belong to the same category in all phases, the score is 1, indicating high historical consistency; on the contrary, if the changes are frequent, the score is close to 0; Step 3: Generate “Historical Consistency Score” Finally, calculate the S hist or S hist ' This will serve as the “historical consistency score” of the sample point and be used in subsequent comprehensive evaluations.
[0014] Preferably, a penalty coefficient matrix α is introduced in step S4 for weighted correction: ; Among them, α ij Indicates a change from category i to category j The penalty or reward coefficient.
[0015] Preferably, step S5 includes the following steps: Step 1: Build a basic scoring system Construct a metadata information table for each sampling point, including fields such as data source agency, sampling personnel, sampling date, equipment model, and sampling depth; assign a basic score based on the completeness of these fields; and assign different basic scores based on the nature of the source channel; Step 2: Weighted correction based on data characteristics After obtaining the basic score, the following cross-validation method is used for weighted correction; 1) Outlier detection correction: Outlier detection is performed on the core variables of the sample points in the entire dataset using box plots, local outlier factors, or the DBSCAN algorithm. If a core variable of a sample point is detected as an outlier, its source score will be negatively modified, i.e., deducted. The extent of the deduction is proportional to the degree of outlier. 2) Remote sensing consistency verification and correction: Correlation analysis is performed on the core variable values of the sample points and the remote sensing vegetation index of the same period and location. If the relationship between the core variable value and the remote sensing vegetation index of a certain sample point deviates significantly from the trend of the entire dataset, it indicates that there may be data entry errors or collection problems at the sample point, and its source score will be deducted; Step 3: Generate a "source reliability score" by weighting the basic score and all correction items to obtain the "source reliability score" of the sample point; ; Among them, S base As the basic score, M i For the i A correction term, ω i is the corresponding weight coefficient.
[0016] As a preference, the weighted sum model of step S6 is as follows: the comprehensive reliability score S of the sample point total The environmental similarity score S is sim, historical consistency score S hist and source reliability score S source Perform weighted summation to obtain; S total =ω sim . S sim +ω hist . S hist +ω source .S source ; Among them: S total ∈[0,1] is the final total score of sample quality; S sim ∈[0,1] is the score obtained by environment similarity analysis; S hist ∈[0,1] is the score obtained by historical consistency detection; S sourcet ∈[0,1] is the score obtained by scoring the source of the sample; ω sim ,ω hist ,ω source are the corresponding weight coefficients, whose values can be set according to actual needs and satisfy ω sim +ω hist +ω source 1.
[0017] Furthermore, the present invention also provides a coastal wetland soil sample quality assessment system for implementing the method, including: a data access and alignment module, a factor extraction and standardization module, an environmental similarity scoring module, a historical consistency scoring module, a source reliability scoring module, an adaptive weight and fusion module, a grading and visualization export module, and a parameter library (Σ, W, α, ω), and each module executes steps S1-S7 sequentially.
[0018] Furthermore, the present invention also provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the method when the computer program or instructions are executed by a processor.
[0019] Furthermore, the present invention also provides a computer program product, comprising a computer program or instructions, which implement the steps of the method when executed by a processor.
[0020] Due to the adoption of the above-mentioned technical solution, the present invention forms a sample point quality pre-gating mechanism for coastal wetland scenarios through the three-dimensional quantification and weighted fusion of "environmental similarity, historical consistency, and source reliability". The technical effects are reflected in the following aspects: First, the environmental similarity measurement based on weighted Mahalanobis distance and robust covariance fully characterizes the correlation structure between multi-source factors (remote sensing, topography, ocean, climate) and eliminates the dimensional effect. Compared with screening points with only a single factor threshold or simple Euclidean distance, it can significantly reduce the misjudgment caused by factor collinearity and scale inconsistency, improve the matching degree between sample points and regional "ecological niche", and suppress spatial extrapolation bias from the source; second, the introduction of historical consistency detection of multi-temporal land cover series and transfer penalty matrix α can identify "environmental mutation points" caused by shoreline migration, short-term reclamation / farming disturbance or remote sensing classification drift, effectively eliminate or downgrade unstable position samples, and reduce the systematic interference of temporal asynchrony and scene conversion on indicators such as SOC; third, the source reliability score combines metadata completeness, outlier detection and NDVI-SOC residual consistency check to detect the presence of "environmental mutation points" caused by shoreline migration, short-term reclamation / farming disturbance or remote sensing classification drift, effectively eliminate or downgrade unstable position samples, and reduce the systematic interference of temporal asynchrony and scene conversion on indicators such as SOC; third, the source reliability score combines metadata completeness, outlier detection and NDVI-SOC residual consistency check to detect the presence of "environmental mutation points" caused by shoreline migration, short-term reclamation / farming disturbance or remote sensing classification drift, effectively eliminate or downgrade unstable position samples, and reduce the systematic interference of temporal asynchrony and scene conversion on indicators such as SOC; This approach eliminates data risks caused by non-environmental factors such as missing sampling records, data entry errors, or instrument drift, ensuring that high-scoring sampling points are not only "environmentally sound" but also "data credible." Fourth, the method uses adaptive weighting driven by regional variability, automatically increasing the weight of historical dimensions in coastal zones with high historical variability or strong tidal gradients, while emphasizing environmental similarity in stable areas. This approach is more context-robust than fixed-weighting strategies. Fifth, mechanisms such as Z-score normalization, missing masking, and boundary buffer zone downweighting are integrated throughout the process. Combined with A / B / C grading and visual output, this approach achieves full traceability and auditability from data access to result release, facilitating seamless integration with soil mapping, carbon storage assessment, and mechanistic modeling. Sixth, in terms of engineering implementation, the parameter library (Σ / W / α / ω) and scoring model are versionable and modularly deployable, allowing plug-and-play use. They serve as a unified "quality gate" for database construction and model training, significantly improving the signal-to-noise ratio of training samples, reducing the risk of overfitting and uncertainty propagation, without changing downstream algorithms. This ultimately results in more robust spatial interpolation, more reliable carbon storage estimates, and greater reusability of conclusions. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 : The overall flow chart of this method.
[0022] Figure 2 : Radar chart showing the similarity of sample environment.
[0023] Figure 3 : Example 1 Sample point reliability spatial distribution diagram (A / B / C level).
[0024] Figure 4 : Example 1 Historical consistency change diagram of sample points in a typical wetland area. DETAILED DESCRIPTION
[0025] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0026] 1. Sample data preparation Collect soil sampling data of coastal wetlands, including sampling point coordinates, sampling time, sampling depth, SOC (or other soil indicators), etc.; record data sources (such as literature, monitoring stations, project databases, etc.).
[0027] 2. Extraction of multi-source environmental factors The following types of environmental factor data are extracted for each sample point: Remote sensing indicators: NDVI, mNDWI, EVI, land cover type; Topographic factors: elevation, slope, distance from the sea; Ocean factors: tidal range, sediment concentration (TSM), and seawater intrusion frequency; Climate factors: multi-year average temperature, precipitation, potential evapotranspiration, etc.; Historical layer: Multiple remote sensing images or maps to determine the historical stability of the wetland status at the location.
[0028] 3. Environmental Similarity Analysis The present invention constructs a multi-dimensional environmental vector of the sample points and uses quantitative methods such as Mahalanobis distance or weighted Euclidean distance to scientifically evaluate the similarity between them and the regional average environmental characteristics or a high-confidence reference sample point set, and ultimately generates an objective "environmental similarity score".
[0029] Step 1: Construct a multi-dimensional environment vector. For each sample point P to be evaluated i , we construct a vector X containing its multidimensional environment features i Each dimension x of this vector ij Represents the jth environmental factor at sample point P i The environmental factors may include but are not limited to the variables in “(2) Extracting multi-source environmental factors”. Therefore, the sample point P i The environment vector can be expressed as: X i =[x i1 , x i2 ,…, x im ],in, m is the total number of environmental factors.
[0030] In order to eliminate the influence of different environmental factor dimensions and scales, all environmental vectors need to be normalized. The present invention adopts Z-score normalization, and its formula is: ; where x ij ' is the normalized value, μ j and σ j are the mean and standard deviation of all sample points on the jth environmental factor.
[0031] Step 2: Assessing environmental similarity In a standardized multidimensional environment, the present invention uses a weighted Mahalanobis distance to assess the similarity between sample points. This method not only considers the correlation between environmental factors, but also allows different weights to be assigned based on the importance of the factors.
[0032] Assume that the sample point to be evaluated is P i , the high-confidence reference sample set is S ref We define the sample point P i With the reference sample set S ref The environmental similarity score S sim (P i , S ref )for:
[0033] Where: X i ' and X k ' The sample points P i and reference sample P k The normalized environment vector of Σ -1 is the reference sample set S ref The inverse matrix of the environmental vector covariance matrix. It can correct the mutual influence between various environmental factors. jj ) is a diagonal weight matrix that reflects the importance of different environmental factors to similarity assessment. The diagonal elements w jj represents the weight of the jth environmental factor, and its value can be determined by expert experience or principal component analysis (PCA) and satisfies . Indicates taking the average of all samples in the reference sample set.
[0034] If a high-confidence reference sample set cannot be obtained, the present invention can also adopt a simplified method, that is, calculating the sample to be evaluated P i Its neighboring sample point P jThe similarity between them is calculated and averaged. In this case, a variant of weighted Euclidean distance can be used, such as the similarity calculation method in the improved iPSM: ; Among them, w k is the weight of the kth environmental factor, k=1,…,m; σ k i is the standard deviation of the factor after standardization; x ik i 、x jk i The sample points P j 、P i The value of the kth environmental factor.
[0035] Step 3: Generate an “environmental similarity score” Finally, the similarity metric calculated above is normalized to the range [0, 1] to obtain the sample site's environmental similarity score. A higher score indicates a more similar environment to the reference environment at the sample site, and thus higher data quality. This score will serve as an important input for evaluating the reliability of the sample site in subsequent steps.
[0036] 4. Historical consistency detection This study introduces a time series analysis method to quantitatively assess the historical consistency of sample sites. By comparing changes in land cover type or wetland boundaries across multiple historical time periods, the method assigns an objective "historical consistency score" to each site, thereby eliminating unreliable data from sampling sites caused by short-term environmental changes or human interference.
[0037] Step 1: Build historical time series data 1. Multi-temporal remote sensing image acquisition: Acquire multiple historical temporal remote sensing images covering the study area, such as Landsat series (TM, ETM+, OLI) and Sentinel series, with a time span of several decades (e.g., 1985, 2000, 2020).
[0038] 2. Land cover classification: Use remote sensing image interpretation and classification algorithms (such as support vector machines (SVM), random forests (RandomForest), or deep learning (U-Net)) to classify land cover in each phase of the imagery and generate classification maps for wetlands, water bodies, farmland, cities, and bare land.
[0039] 3. Spatial overlay of sample points: Perform spatial overlay analysis on the precise geographic coordinates (latitude and longitude) of all sample points and the land cover classification maps at different time phases. For each sample point, extract its land cover category at each time phase.
[0040] Step 2: Quantify historical consistency 1. Consistency measurement: for a sample point P i , and its land cover type sequence in T historical phases is {c i1 c i2 ,…,c iT } We define the historical consistency score S of the sample point hist (P i )for: ; in, f(c it ,c i(t-1) ) is an indicator function: ; This formula measures how often a point's land cover type remains constant over successive time periods. If a point belongs to the same category across all time periods, the score is 1, indicating high historical consistency; conversely, if changes are frequent, the score approaches 0.
[0041] 2. Weight correction: Considering that certain specific changes (such as the transformation from wetland type to urban or farmland) have a greater negative impact on the quality of sample data, the present invention can introduce a penalty coefficient matrix α for weighted correction.
[0042] ; Among them, α ij represents the penalty or reward coefficient for changing from category i to category j. For example, when i and j are both wetland types, α ij =1; when i is a wetland and j is a city, α ij =0.2; when i is a city and j is a wetland, α ij =0.8 (possibly representing ecological recovery).
[0043] Step 3: Generate “Historical Consistency Score” Finally, calculate the S hist or S hist ' This will serve as the “historical consistency score” of the sample point and be used in subsequent comprehensive evaluations.
[0044] 5. Sample Source Scoring This paper proposes a sample source reliability assessment method based on metadata analysis and cross-validation, aiming to quantitatively score the reliability, integrity and authenticity of the data itself.
[0045] Step 1: Build a basic scoring system 1. Metadata Completeness: A metadata table is constructed for each sampling site, including fields such as data source agency, sampling personnel, sampling date, equipment model, and sampling depth. A basic score is assigned based on the completeness of these fields. For example, if information is missing, points will be deducted.
[0046] 2. Source Credibility: Different basic scores are assigned based on the nature of the source. For example, samples from national or international authoritative databases (such as the World Soil Information System (WISS)) receive the highest score; data from publicly published journal articles receive the next highest score; and data from the internet, where the source is unknown, receives a lower score or even zero.
[0047] Step 2: Weighted correction based on data characteristics After obtaining the basic score, the present invention adopts the following cross-validation method to perform weighted correction to improve the objectivity of the score.
[0048] 1. Outlier detection correction: Methods: Outlier detection is performed on the core variables of the sample points (such as soil organic carbon content) in the entire data set. Algorithms such as box plot (BoxPlot), local outlier factor (LOF) or DBSCAN can be used.
[0049] Correction: If a core variable of a sample point is detected as an outlier, its source score will be negatively corrected, i.e., deducted. The extent of the deduction is proportional to the degree of the outlier.
[0050] 2. Remote sensing consistency verification and correction: Methods: Correlation analysis was performed between core variable values (such as SOC) at the sample points and remote sensing vegetation indices (such as NDVI) at the same time and location. For wetland ecosystems, SOC and NDVI generally have a significant positive correlation.
[0051] Correction: If the relationship between SOC and NDVI at a particular site deviates significantly from the overall dataset trend, this indicates a possible data entry error or collection issue at that site, and the source score will be penalized. The magnitude of the correction can be determined based on residual analysis.
[0052] Step 3: Generate a "source reliability score" by weighting the basic score and all correction items to obtain the "source reliability score" of the sample point.
[0053] ; Among them, S base As the basic score, M i is the i-th correction term (such as outlier correction, remote sensing verification correction), ω i is the corresponding weight coefficient.
[0054] VI. Comprehensive scoring and classification The present invention provides a scientific comprehensive scoring model that integrates the multi-dimensional sample quality assessment results into a unified comprehensive reliability score, and grades the samples accordingly, providing users with an intuitive basis for quality judgment.
[0055] Step 1: Weighted Sum Model The comprehensive reliability score S of the sample point total The environmental similarity score S is sim , historical consistency score S hist and source reliability score S source The model allows the importance of each evaluation indicator to be adjusted according to different research purposes.
[0056] S total =ω sim . S sim +ω hist . S hist +ω source .S source ; Where: S total ∈[0,1] is the final total quality score of the sample. S sim ∈[0,1] is the score obtained by environment similarity analysis. S hist ∈[0,1] is the score obtained by historical consistency detection. S sourcet ∈[0,1] is the score obtained by scoring the source of the sample. sim ,ω hist ,ω source are the corresponding weight coefficients, whose values can be set according to actual needs and satisfy ω sim +ω hist +ω source = 1. For example, in areas with drastic historical changes, ω can be appropriately increased. hist The value of .
[0057] Step 2: Quality Grading Based on Scores To provide a more intuitive sample quality assessment, the present invention establishes a set of grading criteria that divides the continuous comprehensive reliability score into discrete quality levels, facilitating quick decision-making and data screening by users.
[0058] Grade A (high reliability): comprehensive reliability score S total ≥0.8. Sample data at this level are of high quality, with environmental characteristics highly consistent with the region, stable historical status, and a credible source. They can be used as high-quality modeling input data or reference samples.
[0059] Class B (medium reliability): 0.6≤Stotal <0.8. Sample data at this level have a certain degree of reliability and may have slight deficiencies in certain dimensions (such as historical consistency). However, they can still be used for certain studies that do not require high accuracy or as backup samples after data cleaning.
[0060] Level C (low reliability): comprehensive reliability score S total <0.6. Sample data at this level are of lower quality and may have multiple issues, such as significant environmental heterogeneity, unstable historical conditions, or unreliable sources. We recommend rigorous data cleaning or direct exclusion before modeling.
[0061] 7. Result Output and Screening The present invention provides multiple forms of evaluation result output, aiming to present sample quality information in the most intuitive way and support users to perform flexible data management and application.
[0062] Step 1: Structured data output Scoring results table: The system will generate a table file (such as .csv or .xlsx format) containing the detailed evaluation results of all sample points. The table includes not only the original data, but also the following new columns: S_sim: environment similarity score, S_hist: historical consistency score, S_source: source reliability score, S_total: comprehensive reliability score, Quality_Grade: Final quality grade (A / B / C).
[0063] Step 2: Multi-dimensional visual analysis Spatial distribution map of sample points: In the geographic information system (GIS) interface or an independent mapping module, different colors or icons are used to represent the quality level of sample points (for example, green represents grade A, yellow represents grade B, and red represents grade C), intuitively showing the spatial aggregation or dispersion patterns of high, medium, and low reliability sample points.
[0064] Similarity radar chart: Generate radar charts of various environmental factors for selected typical sample points, helping users to intuitively compare the multi-dimensional environmental characteristics of different sample points and identify the sources of their similarities or differences.
[0065] Credibility Distribution Plot: Generates a histogram or kernel density estimation plot of the reliability scores to show the distribution trend of all sample scores, helping users understand the overall distribution of the evaluation results.
[0066] Step 3: Flexible data screening and application This system provides powerful interactive data screening functions, supporting users to accurately screen samples according to multiple conditions.
[0067] Filter by level: Users can choose to export only level A samples, or export both level A and level B samples at the same time with one click for high-precision model training or sensitivity analysis.
[0068] Filter by Region: Users can select a specific geographic area and evaluate or export only the points within that area.
[0069] Filter by multiple conditions: Supports combining different filtering conditions, for example: "Filter all sample points located in a certain province with a comprehensive score greater than 0.8".
[0070] Through the above-mentioned multi-dimensional result output and screening functions, the present invention ensures the transparency and practicality of the evaluation results, greatly simplifies the user's data cleaning and preparation work, and improves scientific research and application efficiency.
[0071] The following three examples and two comparative examples / ablation experiments illustrate data preparation, parameter settings, evaluation methods, and results from real-world engineering perspectives, demonstrating the technical benefits of each module of the present invention. The numerical values represent statistical results (averages of multiple spatial block cross-validation runs) obtained through replicated verification using a consistent process.
[0072] Example 1 (Baseline area: estuary delta tidal flat-reed flat-salt marsh composite landscape) 1. Data and preprocessing Sampling points: 1,500 soil SOC measurement points were collected at a sampling depth of 0–30 cm from 2018 to 2024.
[0073] Grid factor: 30 m uniform resolution. Remote sensing (NDVI, mNDWI, EVI, land cover), topography (DEM, slope, distance to sea), oceanography (multi-year mean tidal range, TSM proxy), climate (multi-year mean temperature / precipitation / potential evapotranspiration).
[0074] Temporal alignment: multi-temporal images were selected and time-composited within ±30 days of the sampling date; spatial consistency: reprojection and bilinear / nearest neighbor resampling.
[0075] 2. Methods and parameters Normalization: Z-score.
[0076] Environmental similarity S sim :The reference set is the high-confidence sample points in the region (see below), and the Mahalanobis distance uses the MCD robust covariance; the weight matrix W is weighted by the PCA contribution rate and entropy weight, ∑ w k =1.
[0077] Historical consistency S hist ' : Time phase T=6 (1989 / 2000 / 2005 / 2010 / 2015 / 2020), category transfer penalty matrix α: wetland→city 0.2, wetland→farmland 0.3, internal wetland conversion 1.0, city / farmland→wetland 0.8.
[0078] Source reliability source : Metadata completeness (source, sampler, equipment, depth, coordinates, date) is the base score; anomaly detection (LOF, box plot) and NDVI–SOC residual consistency are used as correction items.
[0079] Adaptive weight: ω sim ,ω hist ,ω source According to the normalized mapping of the average annual migration of the shoreline and the coverage variability index, we can obtain ω sim =0.40,ω hist =0.40,ω source =0.20.
[0080] Grading threshold: A ≥ 0.80, B ∈ [0.60, 0.80), C < 0.60.
[0081] Model evaluation: Using A / B-level samples as training / validation sets, we used 10-fold cross-validation with spatial partitioning (based on grid / watershed partitioning) and compared it to the baseline ("all samples" without quality assessment). The regressors were unified into a random forest (500 trees) + XGBoost (grid search) and the ensemble average was taken. The metric was R 2 , RMSE (g / kg), MAE (g / kg).
[0082] 3. Results 1) Quality distribution: A=820 points (54.7%), B=410 points (27.3%), C=270 points (18.0%).
[0083] 2) Modeling accuracy (SOC) No filter (all 1,500 points): R 2 =0.55±0.03, RMSE=7.4, MAE=5.1.
[0084] A-level only (820 points): R 2 =0.69±0.02, RMSE=6.0, MAE=4.3.
[0085] Grade A+B (1,230 points): R²=0.66±0.02, RMSE=6.3, MAE=4.5.
[0086] 3) Uncertainty: After A / B screening, the average width of the 95% confidence interval of the prediction decreased by 23.4%, and the high-uncertainty residuals converged significantly in the tidal flat-reclamation alternating zone.
[0087] Compared with the unscreened method, the present invention improves R² to 0.69 and reduces RMSE by ≈18.9% under the same algorithm, significantly reducing the propagation of bias caused by environmental heterogeneity and data anomalies.
[0088] Comparative Example 1 (Removing the Historical Consistency Module) Under the same setting as in Example 1, remove S hist ' (ω hist = 0, its weight is distributed proportionally to ω sim ,ω source ).
[0089] A-level only: R 2 =0.62±0.03, RMSE=6.7; Grade A+B: R 2 =0.60±0.03.
[0090] Error hotspots are concentrated in the reclaimed nearshore strip (1 km buffer) and the tidal channel entrance, and the spatial residuals are striped. Historical consistency testing effectively identifies "environmental abrupt changes," and omitting this module significantly reduces the robustness of the marginal zone.
[0091] Example 2 (Severe Tidal Gradient Coastal Zone: Salt Marsh-Mudflat-Tidal Flat Inversion Sensitive Area) 1) Data and differences Sample points: 900 points (0-30 cm), tidal range>2.5 m; cover variability index and shoreline migration are significantly higher than those in Example 1.
[0092] Weight adaptation result: ω hist =0.50,ω sim =0.35,ω source =0.15. Other procedures are the same as above.
[0093] 2) Evaluation and indicators Set the "nearshore 1 km buffer zone validation set" (142 points independently retained).
[0094] Measuring edge robustness: buffer RMSE, spatial autocorrelation Moran ' s I residual weakening ratio and grade mismatch rate (the proportion of predicted high SOC falling within the low vegetation index water mask).
[0095] 3) Results The present invention (adaptive weight): buffer RMSE = 6.8; mismatch rate = 5.9%; residual autocorrelation weakened by 31%.
[0096] Fixed-weight control (ω fixed = 1 / 3): buffer RMSE = 7.9; mismatch rate = 8.6%.
[0097] In highly dynamic coastal zones, adaptively improving ω hist It can significantly reduce the systematic mismatch and overfitting artifacts along the coastal strips.
[0098] Comparative Example 2 (Fixed Weight + Simplified Similarity) Fix ω to (1 / 3,1 / 3,1 / 3) and replace Mahalanobis distance with weighted Euclidean distance (do not use Σ -1 ).
[0099] A+B grade: R 2 =0.58±0.03, RMSE=7.1; buffer mismatch rate=9.3%.
[0100] Conclusion: Ignoring the factor correlation structure and scale effect will significantly impair the ability to discriminate similarities.
[0101] Example 3 (Verification of Anti-Anomaly Robustness: Introducing Contaminated Sample Points) 1) Setup Among the 1,500 points in Example 1, 2% of “abnormal samples” (instrument offset / entry error / coordinate drift) were artificially mixed in, and their SOC values deviated by >3σ relative to similar environments.
[0102] Turn on S source Abnormal detection and NDVI-SOC residual correction; the control is to turn off this module (ω source =0).
[0103] 2) Results Turn on S source : The outliers are judged as C level or downgraded, and the final training set (A+B) R 2 =0.65, RMSE=6.4; the skewness of the residual distribution dropped from 0.62 to 0.21, and the heavy-tailedness was significantly alleviated.
[0104] Close S source : A+B training set R 2 =0.59, RMSE=7.0; the residuals have a right heavy tail, and when extrapolated to the unsampled area, a "local extreme value island" appears.
[0105] The above is a description of the embodiments of the present invention. The above description of the disclosed embodiments will enable professionals in the field to implement or use the present invention. Various modifications to these embodiments will be apparent to professionals in the field. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features disclosed herein.
[0106] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0107] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0108] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0109] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0110] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0111] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0112] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
Claims
1. A method for evaluating the quality of coastal wetland soil samples based on multi-source environmental data, characterized in that: This includes following process steps: S1 Sample data preparation: Collect basic information of soil samples and integrate remote sensing images, climate and terrain factors, and historical land cover data; S2 Extraction of multi-source environmental factors: Extraction of remote sensing indicators, topographic factors, ocean factors, climate factors and historical layers according to the sampling point location; S3 environment similarity scoring: in the presence of a reference set S ref The environmental similarity score is obtained by normalizing the weighted Mahalanobis distance S sim ;Lack S ref When the environmental similarity score is obtained in the joint geographic and environmental space, K-nearest neighbor weighted Euclidean is used S sim ; S4 Historical consistency score: The stability is calculated based on the multi-phase coverage sequence and corrected by the transfer penalty matrix α to obtain the historical consistency score. S hist ; S5 Source Reliability Score: Based on metadata integrity and source credibility, combined with anomaly detection and NDVI-SOC residual consistency correction, the source reliability score is obtained. S source ; S6 adaptive weight configuration: weights are constructed based on regional historical coverage variability, shoreline migration and interannual climate variation ω sim 、 ω hist 、 ω source , and the sum of the three is 1; comprehensive scoring and grading through weighted summation model: calculation S total = ω sim · S sim + ω hist · S hist + ω source · S source , the grading threshold is A:S total ≥0.80, B:0.60≤S total <0.80, C:S total <0.60; S7 result output and screening: Generate S sim 、 S hist 、 S source 、 S total and one or more of hierarchical tables and maps, radar charts, and distribution charts.
2. The method according to claim 1, characterized in that In step S2, Remote sensing indicators: NDVI, mNDWI, EVI, land cover type; Topographic factors: elevation, slope, distance from the sea; Ocean factors: tidal range, sediment concentration, and seawater intrusion frequency; Climate factors: multi-year average temperature, precipitation, and potential evapotranspiration; Historical layer: Multiple remote sensing images or maps to determine the historical stability of the wetland status at the location.
3. The method according to claim 1, characterized in that Step S3 includes the following steps: Step 1: Construct a multidimensional environment vector For each sample point P to be evaluated i , construct a vector X containing its multidimensional environment characteristics i ; Each dimension x of the vector ij Representative j Environmental factors at sample point P i The value at the sample point P i The environment vector is expressed as: X i =[x i1 , x i2 ,…,x im ],in, m is the total number of environmental factors and is standardized using Z-score; Step 2: Assessing environmental similarity In the standardized multidimensional environment, weighted Mahalanobis distance is used to evaluate the similarity between sample points; Assume that the sample point to be evaluated is P i , the high-confidence reference sample set is S ref ; Define sample point P i With the reference sample set S ref The environmental similarity score S sim (P i , S ref )for: ; Where: X i ' and X k ' The sample points P i and reference sample P k The normalized environment vector Σ -1 is the reference sample set S ref The inverse matrix of the environmental vector covariance matrix; W=diag(w jj ) is a diagonal weight matrix, the diagonal elements w jj Represents the weight of the jth environmental factor and satisfies ; Indicates taking the average of all the samples in the reference sample set; Alternatively, calculate the sample points to be evaluated P i Its neighboring sample points P j The similarity between them is calculated and averaged; a variant of weighted Euclidean distance is used, and the calculation method is as follows: ; Among them, w k is the weight of the kth environmental factor, k is the environmental factor index, k=1,…,m; σ k i is the standard deviation of the factor after standardization; x ik i 、x jk i The sample points P j 、P i The value of the kth environmental factor; Step 3: Generate environment similarity score The similarity measure values calculated above are normalized to the interval [0,1] to obtain the environmental similarity score of the sample point; the higher the score, the more similar the environment of the sample point is to the reference environment, and the higher the data quality.
4. The method according to claim 1, wherein Step S4 includes the following steps: Step 1: Build historical time series data Acquire multiple historical remote sensing images covering the study area; use remote sensing image interpretation and classification algorithms to classify land cover in each phase, generating classification maps that include wetlands, water bodies, farmland, cities, and bare land; perform spatial overlay analysis on the precise geographic coordinates of all sample points and the land cover classification maps of different phases; and extract the land cover category of each sample point at each phase. Step 2: Quantify historical consistency For a sample point P i , and its land cover type sequence in T historical phases is {c i1 c i2 ,…,c iT } ; Define the historical consistency score S of the sample point hist (P i )for: ; in, f(c it ,c i(t-1) ) is an indicator function: ; If the sample points belong to the same category in all phases, the score is 1, indicating high historical consistency; on the contrary, if the changes are frequent, the score is close to 0; Step 3: Generate "historical consistency score" Finally, the calculated S hist or S hist ' This will serve as the "historical consistency score" of the sample point and be used in subsequent comprehensive evaluations.
5. The method according to claim 4, characterized in that In step S4, a penalty coefficient matrix α is introduced for weighted correction: ; Among them, α ij Indicates a change from category i to category j The penalty or reward coefficient.
6. The method according to claim 1, characterized in that Step S5 includes the following steps: Step 1: Build a basic scoring system Construct a metadata information table for each sampling point, including data source agency, sampling personnel, sampling date, equipment model, and sampling depth fields; Based on the completeness of these fields, a base score is assigned; Different basic scores are assigned according to the nature of the source channel; Step 2: Weighted correction based on data characteristics After obtaining the basic score, the following cross-validation method is used for weighted correction; 1) Outlier detection correction: Outlier detection is performed on the core variables of the sample points in the entire dataset using box plots, local outlier factors, or the DBSCAN algorithm. If a core variable of a sample point is detected as an outlier, its source score will be negatively modified, i.e., deducted. The extent of the deduction is proportional to the degree of outlier. 2) Remote sensing consistency verification and correction: Correlation analysis is performed on the core variable values of the sample points and the remote sensing vegetation index of the same period and location. If the relationship between the core variable value and the remote sensing vegetation index of a certain sample point deviates significantly from the trend of the entire dataset, it indicates that there may be data entry errors or collection problems at the sample point, and its source score will be deducted; Step 3: Generate a "Source Reliability Score" by weighting the base score and all correction items to obtain the "Source Reliability Score" of the sample point. ; Among them, S base As the basic score, M i For the i A correction term, ω i is the corresponding weight coefficient.
7. The method according to claim 1, characterized in that The weighted summation model in step S6 is as follows: the comprehensive reliability score S of the sample point total The environmental similarity score S is sim , historical consistency score S hist and source reliability score S source Perform weighted summation to obtain; S total =ω sim . S sim +ω hist . S hist +ω source .S source ; Where: S total ∈[0,1] is the final total score of sample quality; S sim ∈[0,1] is the score obtained by environment similarity analysis; S hist ∈[0,1] is the score obtained by historical consistency detection; S sourcet ∈[0,1] is the score obtained by scoring the source of the sample; ω sim ,ω hist ,ω source are the corresponding weight coefficients, whose values can be set according to actual needs and satisfy ω sim +ω hist +ω source =1.
8. A coastal wetland soil sample quality assessment system for implementing the method according to any one of claims 1 to 7, characterized in that: include: Data access and alignment module, factor extraction and standardization module, environmental similarity scoring module, historical consistency scoring module, source reliability scoring module, adaptive weight and fusion module, classification and visualization export module and parameter library, each module executes steps S1-S7 in sequence.
9. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method for evaluating quality of cultivated land in sandstorm saline-alkali area
CN119250351A
Systems and methods for generating habitat condition assessments
US12080049B1
Cited By
Grain and oil matrix standard substance candidate raw material accurate screening method based on characteristic fingerprint spectrum
CN121583313A
LIMS platform data cleaning method and system based on metadata
CN121614467A
Metadata-based lims platform data cleaning method and system
CN121614467B
Wetland ecological quality comprehensive evaluation method and system based on multi-index fusion
CN121961339A