A method and system for quality assessment of coastal wetland soil sampling points based on multi-source environmental data

By using a multi-source environmental data fusion assessment method, the problem of representativeness and consistency assessment of soil sample data in coastal wetlands was solved, and high-quality soil property assessment and accurate carbon storage calculation were achieved.

CN120744705BActive Publication Date: 2025-11-14GUANGDONG LABORATORY OF SOUTHERN OCEAN SCIENCE AND ENGINEERING (GUANGZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511222949.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-14
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing technologies lack systematic and quantitative methods for assessing the representativeness and consistency of soil sampling data in coastal wetlands, resulting in large data errors and affecting the accuracy and interpretability of soil property assessments.

Method used

An assessment method based on multi-source environmental data is adopted, which integrates remote sensing imagery, climate and topographic data and historical land cover information. It is quantified through three dimensions: environmental similarity, historical consistency and source reliability. The quality of sample points is evaluated using weighted Mahalanobis distance, multi-temporal cover sequence and metadata integrity. An adaptive weight model is constructed for comprehensive scoring.

Benefits of technology

It significantly improved the reliability and accuracy of soil sampling data, reduced errors caused by environmental heterogeneity and data anomalies, and enhanced the basic reliability of the soil database and the accuracy of carbon storage assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744705B_ABST
    Figure CN120744705B_ABST
Patent Text Reader

Abstract

This invention relates to the fields of ecological environment monitoring and geographic information technology, and particularly to a method and system for quality assessment of coastal wetland soil sampling points based on multi-source environmental data. The method integrates remote sensing, climate, topography, marine, and historical land cover data, extracts and standardizes indicators, and constructs environmental similarity scores (weighted Mahalanobis / nearest neighbor weighted Euclidean), historical consistency scores (multi-temporal cover + migration penalty), and source reliability scores (metadata integrity + anomaly detection + NDVI-SOC residual). An adaptive weighted fusion is used to obtain a comprehensive score and classify the data into A / B / C levels, outputting a report and a high-quality sampling point set. The system includes data access, a scoring engine, and a visualization module, which can significantly improve the accuracy and robustness of soil mapping and carbon storage assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological environment monitoring and geographic information technology, and in particular to a method and system for quality assessment of coastal wetland soil samples based on multi-source environmental data. Background Technology

[0002] Coastal wetlands, as an important component of "blue carbon" ecosystems, possess strong carbon sink and sequestration potential. However, their high spatial heterogeneity and rapid temporal changes (tidal disturbances, sediment transport, seawater intrusion, etc.) have long posed challenges to the representativeness and consistency of soil sampling data. Research and management practices generally rely on massive environmental covariates (remote sensing, climate, topography, marine processes) to conduct soil property mapping and carbon storage assessment. However, a systematic, quantitative, and replicable process is still lacking in the quality assessment stage of "whether the sampling points themselves are usable / reliable."

[0003] Taking the global digital soil mapping system Soil Grids as an example, it uses more than 230,000 profile points and 400+ environmental layers to perform machine learning modeling and quantify uncertainty. Although it has significantly improved the supply of spatial information on soil properties (including SOC), its quality control mainly focuses on the quantification of uncertainty and the standardization of points at the model level. The integrated quality score of "environmental similarity, historical consistency and source reliability" of multi-source heterogeneous sample points is still not the core objective.

[0004] Regarding broad-scale basic soil databases, ISRIC's WISE30sec and its technical report systematically provide global-scale multi-attribute (including organic carbon) estimates and uncertainties. However, it also relies primarily on statistical inferences from profile-climate / soil combination, emphasizing data fusion, attribute transfer, and outlier control. It has not developed an independent and transferable quality scoring framework for the "applicability and representativeness of sampling points" in highly spatiotemporally dynamic ecological units such as coastal wetlands. FAO's HWSDv2.0 also replaces the old WISE with WISE30sec data, emphasizing in-depth hierarchical refinement and attribute consistency, but it does not address the measurement of sampling point reliability based on coastal zone characteristics such as historical cover stability or tidal / shoreline migration.

[0005] In terms of data ecosystem, the coastal wetland sector has recently seen the formation of data hubs such as the Coastal Carbon Network (CCN) and Coastal Carbon Atlas, led by the Smithsonian Institution. These hubs open up tidal wetland soil core / profile data and metadata to the community, providing quality tier fields for retrieval and download, thus promoting data integration and reproducibility in blue carbon research. However, the "quality tiers" of these platforms are more of a labeling and consensus specification for data availability, and have not yet calculated remote sensing-topography-climate-ocean factors and historical cover change information into single-point "quality indices," nor have they developed automated adaptation assessments for sample points.

[0006] At the remote sensing methodology level, multi-source / multi-sensor fusion (optical, SAR, LiDAR / GEDI) has been proven to significantly improve wetland identification and classification accuracy (e.g., random forests, integrating topographic and radar information at 10m resolution, achieve a classification accuracy of over 90%), demonstrating the feasibility of combining environmental covariates with machine learning for wetland monitoring. However, these works primarily address the questions of "where are the wetlands and what classifications are they," rarely using this information retrospectively for quality metrics such as "whether existing soil samples match their environmental background / historical evolution."

[0007] For remote sensing inversion of SOC (soil organic carbon), the academic community has developed a relatively rich methodological system, ranging from multispectral / hyperspectral indices (such as NDVI and CAI) to machine learning / deep learning models. There is also a growing focus on the impact and correction of surface conditions (crop residues, moisture content, roughness) on inversion stability, which provides a basis for understanding the consistency between "environmental indicators and soil properties." However, these studies mostly focus on "improving the accuracy of estimates," rarely using "remote sensing-ground measurement residual behavior" to evaluate the reliability of sample source or recording processes.

[0008] In this context, using unfiltered sample data directly for spatial modeling or mechanism analysis is highly susceptible to introducing errors, weakening the accuracy and interpretability of the results. Therefore, there is an urgent need to establish a systematic, quantifiable, multi-source data fusion-based sample quality assessment method that integrates remote sensing imagery, environmental factors, and historical land cover information to scientifically eliminate or correct unreliable samples, ensuring the reliability of the soil database foundation. Summary of the Invention

[0009] To address the aforementioned technical problems, the present invention aims to provide a method for assessing the quality of coastal wetland soil samples based on multi-source environmental data. By integrating environmental factors such as remote sensing images, climate and topographic data, and historical land cover maps, the method comprehensively evaluates the environmental similarity, historical consistency, and source reliability of the samples, providing a high-quality data foundation for soil mapping and mechanism modeling.

[0010] To achieve the above objectives, the present invention adopts the following technical solution:

[0011] A method for quality assessment of coastal wetland soil samples based on multi-source environmental data includes the following steps:

[0012] S1 sample point data preparation: Collect basic information on soil samples and integrate remote sensing images, climate and topographic factors, and historical land cover data;

[0013] S2 extracts multi-source environmental factors: extracts remote sensing indicators, topographic factors, marine factors, climate factors, and historical layers according to the sampling point location;

[0014] S3 Environment Similarity Score: In the presence of a reference set S ref Environmental similarity scores were obtained by weighted Mahalanobis distance normalization; lacking S ref At that time, the K-nearest neighbor weighted Euclidean algorithm was used to obtain the environmental similarity score in the combined geographic and environmental spatial data.

[0015] S4 Historical Consistency Score: Stability is calculated based on multi-temporal covered sequences and corrected according to the transition penalty matrix α, resulting in... S hist ;

[0016] S5 Source Reliability Score: Based on metadata integrity and source credibility, combined with anomaly detection and NDVI-SOC residual consistency correction, the score is obtained. S source ;

[0017] S6 Adaptive Weight Configuration: Weights are constructed based on historical regional cover variability, shoreline migration, and interannual climate variability. ω sim ω hist ω source And the sum of the three is 1; the comprehensive score and classification are calculated using a weighted summation model: S total =ω sim · S sim +ω hist · S hist +ω source · S source The grading threshold is A:S total ≥0.80, B: 0.60≤S total <0.80, C:S total <0.60;

[0018] S7 Result Output and Filtering: Generate results containing... S sim , S hist , S source , S total And one or more of the following: tables and maps, radar charts and distribution maps.

[0019] Preferably, in step S2...

[0020] Remote sensing indicators: NDVI, mNDWI, EVI, land cover type;

[0021] Topographic factors: elevation, slope, distance from the sea;

[0022] Marine factors: tidal range, sediment concentration, frequency of seawater intrusion;

[0023] Climate factors: multi-year average temperature, precipitation, and potential evapotranspiration;

[0024] Historical layers: Multiple remote sensing images or maps to determine the stability of wetland conditions at this location throughout history.

[0025] Preferably, step S3 includes the following steps:

[0026] Step 1: Construct a multidimensional environment vector

[0027] For each sample point P to be evaluated i Construct a vector X that includes its multidimensional environmental characteristics. i Each dimension x of the vector ij Representing the j Environmental factors at sample point P i The value at point P; sample point P i The environment vector is represented as: X i =[x i1 , x i2 ,…,x im ],in, m This represents the total number of environmental factors; and is standardized using Z-score.

[0028] Step 2: Assess environmental similarity

[0029] In the standardized multidimensional environment, weighted Mahalanobis distance is used to evaluate the similarity between samples;

[0030] Assume the sample point to be evaluated is P. i High-confidence reference sample set S ref Define sample point P. i With reference sample set S ref Environmental similarity score S sim (P i , S ref )for:

[0031] ;

[0032] Where: X i ' and X k ' Sample point P i and reference sample point P k Standardized environment vector; Σ -1 It is the reference sample set S refThe inverse matrix of the environmental vector covariance matrix; W=diag(w jj ) is a diagonal weight matrix, with diagonal elements w jj Let represent the weight of the j-th environmental factor, and satisfy . ; This indicates that the average value is taken over all samples in the reference sample set;

[0033] Alternatively, calculate the sample points to be evaluated. P i and its neighboring samples P j The similarity between them is calculated and averaged; a variant of weighted Euclidean distance is used, and the calculation method is as follows:

[0034] ;

[0035] Among them, w k σ represents the weight of the k-th environmental factor, where k is the index of the environmental factor, k=1,…,m; k i x represents the standard deviation of this factor after standardization. ik i x jk i Sample point P j P i The value of the k-th environmental factor;

[0036] Step 3: Generate environment similarity scores

[0037] The calculated similarity metric values ​​are normalized to the [0,1] interval to obtain the environmental similarity score of the sample point; the higher the score, the more similar the environment of the sample point is to the reference environment, and the higher the data quality.

[0038] Preferably, step S4 includes the following steps:

[0039] Step 1: Constructing historical time-series data

[0040] Acquire multiple historical remote sensing images covering the study area; use remote sensing image interpretation and classification algorithms to classify land cover in each time phase image, generating classification maps including wetlands, water bodies, farmland, cities, and bare land; perform spatial overlay analysis of the precise geographic coordinates of all sample points with land cover classification maps of different time phases; for each sample point, extract its land cover category in each time phase.

[0041] Step 2: Quantifying Historical Consistency

[0042] For a sample point P i Its land cover type sequence over T historical time phases is as follows{c i1 c i2 ,…,c iT } Define the historical consistency score S for this sample point. hist (P i )for:

[0043]

[0044] in, f(c it ,c i(t-1) ) It is an indicator function:

[0045] ;

[0046] If the sample points belong to the same category in all time phases, the score is 1, indicating high historical consistency; conversely, if the changes are frequent, the score approaches 0.

[0047] Step 3: Generate the "Historical Consistency Score" Finally, the calculated S hist or S hist ' The "historical consistency score" for this sample point is used for subsequent comprehensive evaluation.

[0048] Preferably, a penalty coefficient matrix α is introduced in step S4 for weighted correction:

[0049] ;

[0050] Where, α ij This indicates a change from category i to category i. j The penalty or reward coefficient.

[0051] Preferably, step S5 includes the following steps:

[0052] Step 1: Constructing a basic scoring system

[0053] A metadata information table is constructed for each sample point, including fields for data source organization, sampling personnel, sampling date, equipment model, and sampling depth; a base score is assigned based on the completeness of these fields; and different base scores are assigned based on the nature of the source channel.

[0054] Step 2: Weighted Adjustment Based on Data Characteristics After obtaining the basic scores, the following cross-validation method is used for weighted adjustment;

[0055] 1) Outlier detection and correction:

[0056] Outlier detection is performed on the core variables of the sample points across the entire dataset using box plots, local outlier factor analysis, or the DBSCAN algorithm. If a core variable of a sample point is detected as an outlier, its source score will be negatively corrected, i.e., a deduction will be made. The magnitude of the deduction is proportional to the degree of outlierness of the outlier.

[0057] 2) Remote sensing consistency verification correction:

[0058] Correlation analysis was performed between the core variable values ​​of the sample points and the remote sensing vegetation indices of the same period and location. If the relationship between the core variable values ​​of a sample point and the remote sensing vegetation index deviates significantly from the trend of the overall dataset, it indicates that there may be an input error or collection problem in the data of that sample point, and its source score will be deducted.

[0059] Step 3: Generate "Source Reliability Score" by weighting and summing the base score and all correction items to obtain the "Source Reliability Score" of the sample.

[0060] ;

[0061] Among them, S base Based on the score, M i For the first i One correction term, ω i These are the corresponding weighting coefficients.

[0062] As a preferred option, the weighted summation model in step S6 is as follows: the comprehensive reliability score S of the sample points total It is achieved by scoring environmental similarity S sim Historical consistency score S hist and source reliability score S source We obtain S by weighted summation. total =ω sim S sim +ω hist S hist +ω source .S source ; where: S total ∈[0,1] is the final total sample quality score; S sim ∈[0,1] is the score obtained through environmental similarity analysis; S hist ∈[0,1] is the score obtained through historical consistency detection; S sourcet ∈[0,1] is the score obtained through the sampling source scoring; ω sim ,ω hist ,ω source These are the corresponding weighting coefficients, whose values ​​can be set according to actual needs, and satisfy ω. sim +ω hist +ω source1.

[0063] Furthermore, the present invention also provides a coastal wetland soil sampling point quality assessment system for implementing the method, comprising: a data access and alignment module, a factor extraction and standardization module, an environmental similarity scoring module, a historical consistency scoring module, a source reliability scoring module, an adaptive weighting and fusion module, a grading and visualization export module, and a parameter library (Σ, W, α, ω), wherein each module executes steps S1-S7 sequentially.

[0064] Furthermore, the present invention also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the method.

[0065] Furthermore, the present invention also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the method described herein.

[0066] This invention, employing the aforementioned technical solution, quantifies and weights the three dimensions of "environmental similarity, historical consistency, and source reliability" to form a pre-gating mechanism for sample point quality in coastal wetland scenarios. The technical effects are as follows: First, based on weighted Mahalanobis distance and robust covariance, the environmental similarity measure fully characterizes the correlation structure between multiple source factors (remote sensing, topography, ocean, and climate) and eliminates the influence of dimensions. Compared to screening points using only single-factor thresholds or simple Euclidean distance, it significantly reduces misjudgments caused by factor collinearity and scale inconsistency, improves the matching degree between sample points and regional "niche," and suppresses spatial extrapolation bias from the source. Second, the introduction of historical consistency detection of multi-temporal land cover sequences and the transfer penalty matrix α can identify "environmental mutation points" caused by shoreline migration, short-term reclamation / aquaculture disturbances, or remote sensing classification drift, effectively eliminating or downweighting unstable location sample points and reducing systematic interference from time asynchrony and scene transitions on indicators such as SOC. Third, the source reliability score combines metadata integrity, outlier detection, and NDVI-SOC residual consistency verification, enabling... The data risks caused by non-environmental factors such as missing sampling records, entry errors, or instrument drift ensure that high-scoring sampling points are not only "environmentally reasonable" but also "data reliable." Fourth, the method adopts regional variability-driven adaptive weights, which automatically increases the weight of historical dimensions in coastal zones with large historical variability or strong tidal gradients, and emphasizes environmental similarity in stable areas, making it more context-robust than fixed-weight strategies. Fifth, mechanisms such as Z-score standardization, missing data masking, and boundary buffer zone weight reduction are implemented throughout the process. Combined with A / B / C grading and visualization output, it achieves full-link traceability and auditability from data access to result publication, facilitating seamless integration with soil mapping, carbon storage assessment, and mechanism models. Sixth, in engineering implementation, the parameter library (Σ / W / α / ω) and scoring model can be versioned and modularly deployed, plug-and-play, and can serve as a unified "quality gate" for database construction and model training. Without changing the downstream algorithm, it significantly improves the signal-to-noise ratio of training samples, reduces the risk of overfitting and uncertainty propagation, and ultimately results in more robust spatial interpolation, more reliable carbon storage estimation, and higher reusability of conclusions. Attached Figure Description

[0067] Figure 1 : Overall flowchart of this method.

[0068] Figure 2 : Radar diagram illustrating the environmental similarity of sample points.

[0069] Figure 3 Example 1: Spatial distribution map of sample point reliability (A / B / C level).

[0070] Figure 4 Example 1: Historical consistency change of a sample point in a typical wetland area. Detailed Implementation

[0071] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0072] I. Sample Data Preparation

[0073] Collect soil sampling data from coastal wetlands, including sampling point coordinates, sampling time, sampling depth, SOC (or other soil indicators), etc.; record the data source (such as literature, monitoring stations, project databases, etc.).

[0074] II. Extraction of Multi-Source Environmental Factors

[0075] Extract the following types of environmental factor data for each sampling point:

[0076] Remote sensing indicators: NDVI, mNDWI, EVI, land cover type;

[0077] Topographic factors: elevation, slope, distance from the sea;

[0078] Marine factors: tidal range, sediment concentration (TSM), frequency of seawater intrusion;

[0079] Climate factors: multi-year average temperature, precipitation, potential evapotranspiration, etc.

[0080] Historical layers: Multiple remote sensing images or maps to determine the stability of wetland conditions at this location throughout history.

[0081] III. Environmental Similarity Analysis

[0082] This invention constructs a multidimensional environmental vector of sample points and uses quantitative methods such as Mahalanobis distance or weighted Euclidean distance to scientifically evaluate the similarity between the sample points and the regional average environmental characteristics or a set of highly reliable reference sample points, ultimately generating an objective "environmental similarity score".

[0083] Step 1: Construct a multi-dimensional environment vector. For each sample point P to be evaluated... i We construct a vector X that contains its multidimensional environmental features. i Each dimension x of this vector ij This represents the j-th environmental factor at sample point P. i The value at point P. Environmental factors may include, but are not limited to, the variables in “(2) Extracting multi-source environmental factors”. Therefore, the sample point P i The environment vector can be represented as: X i =[x i1 , x i2 ,…, xim ],in, m This represents the total number of environmental factors.

[0084] To eliminate the influence of different environmental factors' dimensions and scales, all environmental vectors need to be standardized. This invention uses Z-score standardization, the formula of which is: ; where x ij ' μ is the standardized value. j and σ j and are the mean and standard deviation of all samples on the j-th environmental factor, respectively.

[0085] Step 2: Assess environmental similarity

[0086] In a standardized multidimensional environment, this invention employs a weighted Mahalanobis distance to assess the similarity between sample points. This method not only considers the correlation between various environmental factors but also allows for assigning different weights based on the importance of the factors.

[0087] Assume the sample point to be evaluated is P. i The high-confidence reference sample set is S. ref We define sample point P. i With reference sample set S ref Environmental similarity score S sim (P i , S ref )for:

[0088]

[0089] Where: X i ' and X k ' Sample point P i and reference sample point P k The standardized environment vector. Σ -1 It is the reference sample set S ref The inverse matrix of the environmental vector covariance matrix. It can correct for the mutual influence between various environmental factors. W=diag(w jj () is a diagonal weight matrix used to reflect the importance of different environmental factors in similarity assessment. Diagonal elements w jj The weight representing the j-th environmental factor can be determined by expert experience or methods such as principal component analysis (PCA), and satisfies the following conditions: . This indicates that the average value is taken from all samples in the reference sample set.

[0090] If a high-confidence reference sample set is unavailable, this invention can also employ a simplified method, namely, calculating the sample P to be evaluated. i Its neighboring sample point P j The similarity between them is calculated and averaged. In this case, a variant of weighted Euclidean distance can be used, such as the similarity calculation method in the improved iPSM:

[0091] ;

[0092] Among them, w k σ represents the weight of the k-th environmental factor, k=1,…,m; k i x represents the standard deviation of this factor after standardization. ik i x jk i Sample point P j P i The value of the k-th environmental factor.

[0093] Step 3: Generate "Environmental Similarity Score"

[0094] Finally, the calculated similarity metrics are normalized to the [0,1] interval to obtain the environmental similarity score of the sample point. The higher the score, the more similar the environment of the sample point is to the reference environment, and the higher the data quality. This score will serve as an important input for evaluating the reliability of the sample points in subsequent steps.

[0095] IV. Historical Consistency Detection

[0096] This invention introduces a time series analysis method for quantitatively evaluating the historical consistency of sample points. This method assigns an objective "historical consistency score" to sample points by comparing changes in land cover type or wetland boundaries over multiple historical periods, thereby eliminating the unreliability of sampling data caused by short-term environmental changes or human interference.

[0097] Step 1: Constructing historical time-series data

[0098] 1. Acquisition of multi-temporal remote sensing images: Acquire multiple historical temporal remote sensing images covering the study area, such as the Landsat series (TM, ETM+, OLI), Sentinel series, etc., with a time span covering the past decades (e.g., 1985, 2000, 2020).

[0099] 2. Land cover classification: Using remote sensing image interpretation and classification algorithms (such as Support Vector Machine (SVM), Random Forest (RandomForest), or Deep Learning U-Net, etc.), land cover is classified for each temporal image, generating classification maps for wetlands, water bodies, farmland, cities, bare land, etc.

[0100] 3. Spatial Overlay of Sample Points: Spatial overlay analysis is performed on the precise geographic coordinates (latitude and longitude) of all sample points and land cover classification maps of different time phases. For each sample point, its land cover category in each time phase is extracted.

[0101] Step 2: Quantifying Historical Consistency

[0102] 1. Consistency measure: For a sample point P i Its land cover type sequence over T historical time phases is as follows {c i1 c i2 ,…,c iT } We define the historical consistency score S for this sample point. hist (P i )for:

[0103] ;

[0104] in, f(c it ,c i(t-1) ) It is an indicator function:

[0105] ;

[0106] This formula measures how frequently a sample point's land cover type remains unchanged across consecutive time phases. If a sample point belongs to the same category in all time phases, it scores 1, indicating high historical consistency; conversely, if changes are frequent, the score approaches 0.

[0107] 2. Weighting Correction: Considering that certain specific changes (such as changing from wetland type to urban or farmland) have a greater negative impact on the quality of sample data, this invention can introduce a penalty coefficient matrix α for weighting correction.

[0108] ;

[0109] Where, α ij This represents the penalty or reward coefficient for changing from category i to category j. For example, when both i and j are wetland types, α ij =1; when i is a wetland and j is a city, α ij =0.2; when i is a city and j is a wetland, α ij =0.8 (which may represent ecological restoration).

[0110] Step 3: Generate the "Historical Consistency Score" Finally, the calculated Shist or S hist ' The "historical consistency score" for this sample point is used for subsequent comprehensive evaluation.

[0111] V. Sampling Source Scoring

[0112] This invention proposes a sample source reliability assessment method based on metadata analysis and cross-validation, which aims to quantify and score the reliability, integrity and authenticity of the data itself.

[0113] Step 1: Constructing a basic scoring system

[0114] 1. Metadata Completeness: A metadata information table is constructed for each sample point, including fields such as data source organization, sampling personnel, sampling date, equipment model, and sampling depth. A base score is assigned based on the completeness of these fields. For example, points are deducted if information is missing.

[0115] 2. Credibility of Source: Different base scores are assigned based on the nature of the source. For example, samples from national or international authoritative databases (such as the World Soil Information System (WISS)) receive the highest scores; data from publicly published journal articles receive the next highest scores; while online data from unknown sources receive low scores, or even zero scores.

[0116] Step 2: Weighted Correction Based on Data Characteristics After obtaining the basic score, this invention uses the following cross-validation method to perform weighted correction in order to improve the objectivity of the scoring.

[0117] 1. Outlier detection and correction:

[0118] Methods: Outlier detection is performed on the core variables of the sample points (such as soil organic carbon content) in the entire dataset. Algorithms such as BoxPlot, Local Outlier Factor (LOF), or DBSCAN can be used.

[0119] Correction: If a core variable of a sample point is detected as an outlier, its source score will be negatively corrected, i.e., a deduction will be made. The magnitude of the deduction is proportional to the degree of outlierness of the outlier.

[0120] 2. Remote sensing consistency verification correction:

[0121] Methods: Correlation analysis was performed on the core variable values ​​(such as SOC) of the sampling points with remote sensing vegetation indices (such as NDVI) at the same time and location. For wetland ecosystems, SOC usually shows a significant positive correlation with NDVI.

[0122] Correction: If the relationship between the SOC value and NDVI of a sample point deviates significantly from the trend of the overall dataset, it indicates that there may be an entry error or collection problem in the data of that sample point, and its source score will be deducted. The correction magnitude can be determined based on residual analysis.

[0123] Step 3: Generate the "Source Reliability Score" by weighting and summing the base score and all correction items to obtain the "Source Reliability Score" for the sample.

[0124] ;

[0125] Among them, S base Based on the score, M i For the i-th correction term (such as outlier correction, remote sensing verification correction), ω i These are the corresponding weighting coefficients.

[0126] VI. Overall Scoring and Grade Classification

[0127] This invention provides a scientific comprehensive scoring model that integrates multi-dimensional sample quality assessment results into a unified comprehensive reliability score, and classifies the samples accordingly, providing users with an intuitive basis for quality judgment.

[0128] Step 1: Weighted Summation Model

[0129] The overall reliability score S of the sample points total It is achieved by scoring environmental similarity S sim Historical consistency score S hist and source reliability score S source The result is obtained by weighted summation. This model allows for adjustments to the importance of each evaluation indicator based on different research objectives.

[0130] S total =ω sim S sim +ω hist S hist +ω source .S source ;

[0131] Wherein: S total ∈[0,1] is the final total quality score of the sample points. S sim ∈[0,1] is the score obtained through environmental similarity analysis. hist ∈[0,1] is the score obtained through historical consistency detection. sourcet ∈[0,1] represents the score obtained through sampling source scoring. ω sim ,ω hist ,ω sourceThese are the corresponding weighting coefficients, whose values ​​can be set according to actual needs, and satisfy ω. sim +ω hist +ω source =1. For example, in regions with drastic historical changes, ω can be appropriately increased. hist The value of .

[0132] Step 2: Quality Classification Based on Scores

[0133] To provide a more intuitive assessment of sample quality, this invention establishes a grading standard. This standard divides continuous comprehensive reliability scores into discrete quality levels, facilitating rapid decision-making and data filtering for users.

[0134] Grade A (High Reliability): Overall Reliability Score S total ≥0.8. Sample data at this level are of high quality, with environmental characteristics highly consistent with the region, stable historical status, and reliable sources, making them suitable as high-quality modeling input data or reference samples.

[0135] Class B (Medium Reliability): 0.6≤S total <0.8. Sample data at this level have a certain degree of reliability, but may have slight deficiencies in a certain dimension (such as historical consistency). However, they can still be used for some studies that do not require high precision or as backup samples after data cleaning.

[0136] Grade C (Low Reliability): Overall Reliability Score total <0.6. Sample data at this level are of low quality and may have multiple problems, such as high environmental heterogeneity, unstable historical status, or unreliable sources. It is recommended to perform rigorous data cleaning or exclude samples directly before modeling.

[0137] VII. Results Output and Filtering

[0138] This invention provides multiple forms of evaluation result output, aiming to present sample quality information in the most intuitive way and support users in flexible data management and application.

[0139] Step 1: Structured Data Output

[0140] Scoring Results Table: This system will generate a table file (e.g., .csv or .xlsx format) containing detailed evaluation results for all sample points. This table includes not only the original data but also the following:

[0141] S_sim: Environment similarity score

[0142] S_hist: Historical consistency score

[0143] S_source: Source reliability score

[0144] S_total: Overall reliability score

[0145] Quality_Grade: Final quality grade (A / B / C).

[0146] Step 2: Multi-dimensional Visual Analysis

[0147] Spatial distribution map of sample points: In the geographic information system (GIS) interface or a separate drawing module, different colors or icons are used to represent the quality level of sample points (for example, green represents level A, yellow represents level B, and red represents level C), which intuitively shows the spatial clustering or dispersion patterns of high, medium, and low reliability sample points.

[0148] Similarity Radar Chart: For selected typical samples, a radar chart is generated for each environmental factor, helping users to intuitively compare the multidimensional environmental characteristics of different samples and identify the sources of their similarities or differences.

[0149] Credibility distribution map: Generates a histogram or kernel density estimation map of the reliability score, showing the distribution trend of the scores of all samples, helping users understand the overall distribution of the evaluation results.

[0150] Step 3: Flexible Data Filtering and Application

[0151] This system offers powerful interactive data filtering capabilities, allowing users to accurately filter sample points based on various criteria.

[0152] Filter by level: Users can choose to export only level A sample points or export both level A and level B sample points at the same time with one click, for high-precision model training or sensitivity analysis.

[0153] Filter by region: Users can select a specific geographic region and evaluate or export only the samples within that region.

[0154] Filter by multiple conditions: Supports combining different filtering conditions, for example: "Filter all sample points located in a certain province and with a comprehensive score greater than 0.8".

[0155] Through the aforementioned multi-dimensional result output and filtering functions, this invention ensures the transparency and practicality of the evaluation results, greatly simplifies the user's data cleaning and preparation work, and improves the efficiency of scientific research and application.

[0156] The following three "Example Examples" and two "Comparative Examples / Ablation Experiments" describe the data preparation, parameter settings, evaluation methods, and results using real engineering standards, thereby demonstrating the technical effects brought about by each module of the present invention. The numerical values ​​are the statistical results obtained by replicating the verification according to a consistent process (average of multiple spatial block cross-validation).

[0157] Example 1 (Baseline area: Estuary delta tidal flat-reed flat-salt marsh composite landscape)

[0158] 1. Data and Preprocessing

[0159] Sampling points: A total of 1,500 soil SOC measurement points were collected at a depth of 0–30 cm from 2018 to 2024.

[0160] Raster factor: 30 m uniform resolution. Remote sensing (NDVI, mNDWI, EVI, land cover), topography (DEM, slope, distance from sea), ocean (multi-year mean tidal range, TSM proxy), climate (multi-year mean temperature / precipitation / potential evapotranspiration).

[0161] Temporal alignment: Select multi-temporal images with sample date ±30 days and perform temporal composite; Spatial consistency: Reprojection and bilinear / nearest neighbor resampling.

[0162] 2. Methods and Parameters

[0163] Standardization: Z-score.

[0164] Environmental similarity S sim The reference set consists of high-confidence samples within the region (see below), and the Mahalanobis distance uses MCD robust covariance; the weight matrix W is weighted by a mixture of PCA contribution rate and entropy weight, ∑ w k =1.

[0165] Historical consistency S hist ' Time phase T=6 (1989 / 2000 / 2005 / 2010 / 2015 / 2020), category transition penalty matrix α: wetland→city 0.2, wetland→farmland 0.3, wetland internal inter-transition 1.0, city / farmland→wetland 0.8.

[0166] Source reliability S source Metadata completeness (source, sampler, device, depth, coordinates, date) is the baseline score; anomaly detection (LOF, box plot) and NDVI–SOC residual consistency are used as correction items.

[0167] Adaptive weights: ω sim ,ω hist ,ω source Based on the normalized mapping of the annual average shoreline migration rate and the coverage variability index, ω is obtained. sim =0.40,ω hist =0.40,ω source =0.20.

[0168] Grading thresholds: A≥0.80, B∈[0.60,0.80), C<0.60.

[0169] Modeling and Evaluation: A / B level samples were used as the training / validation set, employing spatially partitioned 10-fold cross-validation (based on grid / watershed partitioning) and compared to a baseline ("all samples" without quality assessment). The regressor was uniformly set to Random Forest (500 trees) + XGBoost (grid search) with ensemble averaging. The metric was R0. 2 , RMSE (g / kg), MAE (g / kg).

[0170] 3. Results

[0171] 1) Quality distribution: A=820 points (54.7%), B=410 points (27.3%), C=270 points (18.0%).

[0172] 2) Modeling accuracy (SOC)

[0173] No filter (all 1,500 points): R 2 =0.55±0.03, RMSE=7.4, MAE=5.1.

[0174] Grade A only (820 points): R 2 =0.69±0.02, RMSE=6.0, MAE=4.3.

[0175] Grade A+B (1,230 points): R²=0.66±0.02, RMSE=6.3, MAE=4.5.

[0176] 3) Uncertainty: After A / B screening, the average width of the 95% confidence interval of the prediction decreased by 23.4%, and the high uncertainty residuals converged significantly in the alternating zone of tidal flats and reclamation.

[0177] Compared to the unscreened version, this invention increases R² to 0.69 and decreases RMSE by approximately 18.9% under the same algorithm, significantly reducing the propagation of bias caused by environmental heterogeneity and data anomalies.

[0178] Comparative Example 1 (removing the historical consistency module)

[0179] Under the exact same settings as in Example 1, remove S hist ' (ω) hist =0, its weight is distributed proportionally to ω. sim ,ω source ).

[0180] Grade A only: R 2 =0.62±0.03, RMSE=6.7; Grade A+B: R 2 =0.60±0.03.

[0181] Error hotspots re-concentrate on the nearshore reclamation strip (1 km buffer zone) and tidal channel entrances, with spatial residuals exhibiting a striped pattern. Historical consistency detection effectively identifies "environmental abrupt change points," and the absence of this module significantly reduces the robustness of the edge zone.

[0182] Example 2 (Strong tidal gradient coastal zone: Salt marsh-mudflat-tidal flat inversion sensitive area)

[0183] 1) Data and Differences

[0184] Sampling points: 900 points (0-30 cm), tidal range >2.5 m; the cover variability index and shoreline migration were significantly higher than those in Example 1.

[0185] Weight adaptive result: ω hist =0.50,ω sim =0.35,ω source =0.15. The rest of the process is the same as above.

[0186] 2) Evaluation and Indicators

[0187] Set up a "1 km nearshore buffer zone validation set" (extracting 142 independently reserved points).

[0188] Measuring edge robustness: Buffer RMSE, Spatial autocorrelation Moran ' S I residual weakening ratio, grade mismatch rate (predicting the proportion of high SOC falling into the water mask of low vegetation index).

[0189] 3) Results

[0190] This invention (adaptive weights): buffer RMSE = 6.8; mismatch rate = 5.9%; residual autocorrelation weakened by 31%.

[0191] Fixed weight comparison (ω fixed = 1 / 3): Buffer RMSE = 7.9; Mismatch rate = 8.6%.

[0192] In highly dynamic coastal zones, adaptively improve ω hist It can significantly reduce systematic mismatch and overfit artifacts in coastal strips.

[0193] Comparative Example 2 (Fixed Weights + Simplified Similarity)

[0194] Fix ω to (1 / 3, 1 / 3, 1 / 3), and replace Mahalanobis distance with weighted Euclidean distance (without using Σ). -1 ).

[0195] Grade A+B: R 2 =0.58±0.03, RMSE=7.1; Buffer mismatch rate=9.3%.

[0196] Conclusion: Ignoring factor-related structure and scaling effects significantly impairs similarity discrimination ability.

[0197] Example 3 (Verification of robustness against anomalies: introduction of contaminated sample points)

[0198] 1) Settings

[0199] In Example 1, 2% of “abnormal samples” (instrument offset / entry error / coordinate drift) were artificially mixed into the 1,500 samples, whose SOC values ​​deviated by more than 3σ relative to similar environments.

[0200] Enable S source Anomaly detection and NDVI-SOC residual correction; control group is the module disabled (ω) source =0).

[0201] 2) Results

[0202] Enable S source Outliers are classified as C or downweighted, resulting in a final training set (A+B) R. 2 =0.65, RMSE=6.4; the skewness of the residual distribution decreased from 0.62 to 0.21, and the heavy-tailedness was significantly alleviated.

[0203] Close S source Training set R of A+B 2 =0.59, RMSE=7.0; the residuals show right heavy tails, and "local extreme islands" appear when extrapolated to the unsampled area.

[0204] The foregoing description of embodiments of the present invention, through which those skilled in the art are able to implement or use the present invention, will be readily apparent to those skilled in the art. Various modifications to these embodiments will be readily apparent to those skilled in the art. The general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novelty disclosed herein.

[0205] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0206] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0207] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0208] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0209] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0210] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0211] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

Claims

1. A method for assessing the quality of coastal wetland soil samples based on multi-source environmental data, characterized in that, This includes following the steps below: S1 sample point data preparation: Collect basic information on soil samples and integrate remote sensing images, climate and topographic factors, and historical land cover data; S2 extracts multi-source environmental factors: extracts remote sensing indicators, topographic factors, marine factors, climate factors, and historical layers according to the sampling point location; S3 Environment Similarity Score: In the presence of a reference set S ref Environmental similarity scores are obtained by weighted Mahalanobis distance normalization. S sim ;Lack S ref Environmental similarity scores are obtained by using K-nearest neighbor weighted Euclidean algorithm in conjunction with geographic and environmental spatial data. S sim ; S4 Historical Consistency Score: The historical consistency score is obtained by calculating the stability based on multi-temporal covered sequences and correcting it according to the transition penalty matrix α. S hist ; S5 Source Reliability Score: Based on metadata integrity and source credibility, the source reliability score is obtained by combining anomaly detection and NDVI-SOC residual consistency correction. S source ; S6 Adaptive Weight Configuration: Weights are constructed based on historical regional cover variability, shoreline migration, and interannual climate variability. ω sim , ω hist , ω source And the sum of the three is 1; the comprehensive score and classification are calculated using a weighted summation model: S total = ω sim · S sim + ω hist · S hist + ω source · S source The grading threshold is A:S total ≥0.80, B: 0.60≤S total <0.80, C:S total <0.60; S7 Result Output and Filtering: Generate results containing... S sim , S hist , S source , S total And one or more of the following: hierarchical tables and maps, radar charts and distribution maps.

2. The method according to claim 1, characterized in that, In step S2, Remote sensing indicators: NDVI, mNDWI, EVI, land cover type; Topographic factors: elevation, slope, distance from the sea; Marine factors: tidal range, sediment concentration, frequency of seawater intrusion; Climate factors: multi-year average temperature, precipitation, and potential evapotranspiration; Historical layers: multi-period remote sensing images or maps to determine the stability of wetland conditions at sample point locations throughout history.

3. The method according to claim 1, characterized in that, Step S3 includes the following steps: Step 1: Construct a multidimensional environment vector For each sample point P to be evaluated i Construct a vector X that contains its multidimensional environmental features. i Each dimension x of the vector ij Representing the j Environmental factors at sample point P i The value at point P; sample point P i The environment vector is represented as: X i =[x i1 , x i2 ,…,x im ],in, m This represents the total number of environmental factors; and is standardized using Z-score. Step 2: Assess environmental similarity In the standardized multidimensional environment, weighted Mahalanobis distance is used to evaluate the similarity between samples; Assume the sample point to be evaluated is P. i The high-confidence reference sample set is S. ref Define sample point P i With reference sample set S ref Environmental similarity score S sim (P i , S ref )for: ; Where: X i ' and X k ' Sample point P i and reference sample point P k Standardized environment vector; Σ -1 It is the reference sample set S ref The inverse matrix of the environmental vector covariance matrix; W=diag(w jj ) is a diagonal weight matrix, with diagonal elements w jj The weight representing the j-th environmental factor. ; This indicates that the average value is taken over all samples in the reference sample set; Alternatively, calculate the sample points to be evaluated. P i and its neighboring samples P j The similarity between them is calculated and averaged; a variant of weighted Euclidean distance is used, and the calculation method is as follows: ; Among them, w k σ represents the weight of the k-th environmental factor, where k is the index of the environmental factor, k=1,…,m; k i x represents the standard deviation of the standardized environmental factors; ik i x jk i Sample point P i P j The value of the k-th environmental factor; Step 3: Generate environment similarity score The calculated similarity metric is normalized to the [0,1] interval to obtain the environmental similarity score of the sample point; the higher the score, the more similar the environment of the sample point is to the reference environment, and the higher the data quality.

4. The method according to claim 1, characterized in that, Step S4 includes the following steps: Step 1: Constructing historical time-series data Acquire multiple historical remote sensing images covering the study area; use remote sensing image interpretation and classification algorithms to classify land cover in each time phase image, generating classification maps including wetlands, water bodies, farmland, cities, and bare land; perform spatial overlay analysis of the precise geographic coordinates of all sample points with land cover classification maps of different time phases; for each sample point, extract its land cover category in each time phase. Step 2: Quantifying Historical Consistency For a sample point P i Its land cover type sequence over T historical time phases is as follows {c i1, c i2 ,…,c iT } Define the historical consistency score S for this sample point. hist (P i )for: ; in, f(c it ,c i(t-1) ) It is an indicator function: ; If the sample points belong to the same category in all time phases, the score is 1, indicating high historical consistency; conversely, if the changes are frequent, the score approaches 0. Step 3: Generate the final "historical consistency score" and use the calculated S hist or S hist Weighted adjustment of S hist ' The "historical consistency score" for this sample point is used for subsequent comprehensive evaluation.

5. The method according to claim 4, characterized in that, In step S4, a penalty coefficient matrix α is introduced for weighted correction: ; Where, α i,j This indicates a change from category i to category i. j The penalty or reward coefficient.

6. The method according to claim 1, characterized in that, Step S5 includes the following steps: Step 1: Constructing a basic scoring system A metadata information table is constructed for each sample point, including fields for data source organization, sampling personnel, sampling date, equipment model, and sampling depth; Assign a base score based on the completeness of these fields; Different base scores are assigned based on the nature of the source channel; Step 2: Weighted Adjustment Based on Data Characteristics After obtaining the basic scores, the following cross-validation method is used for weighted adjustment; 1) Outlier detection and correction: Outlier detection is performed on the core variables of the sample points across the entire dataset using box plots, local outlier factor analysis, or the DBSCAN algorithm. If a core variable of a sample point is detected as an outlier, its source score will be negatively corrected, i.e., a deduction will be made. The magnitude of the deduction is proportional to the degree of outlierness of the outlier. 2) Remote sensing consistency verification correction: Correlation analysis was performed between the core variable values ​​of the sample points and the remote sensing vegetation indices of the same period and location. If the relationship between the core variable values ​​of a sample point and the remote sensing vegetation index deviates significantly from the trend of the overall dataset, it indicates that there may be an input error or collection problem in the data of that sample point, and its source score will be deducted. Step 3: Generate "Source Reliability Score" by weighting and summing the base score and all correction items to obtain the "Source Reliability Score" of the sample. ; Among them, S base Based on the score, M i For the first i One correction term, ω i These are the corresponding weighting coefficients.

7. The method according to claim 1, characterized in that, Step S6, the weighted summation model, is as follows: the comprehensive reliability score S of the sample points. total It is achieved by scoring environmental similarity S sim Historical consistency score S hist and source reliability score S source We obtain S by weighted summation. total =ω sim . S sim +ω hist . S hist +ω source .S source ; Wherein: S total ∈[0,1] is the final total sample quality score; S sim ∈[0,1] is the score obtained through environmental similarity analysis; S hist ∈[0,1] is the score obtained through historical consistency detection; S sourcet ∈[0,1] is the score obtained through the sampling source scoring; ω sim ,ω hist ,ω source These are the corresponding weighting coefficients, whose values ​​can be set according to actual needs, and satisfy ω. sim +ω hist +ω source =1.

8. A quality assessment system for coastal wetland soil sampling points used to implement the method of any one of claims 1-7, characterized in that, include: The data access and alignment module, factor extraction and standardization module, environmental similarity scoring module, historical consistency scoring module, source reliability scoring module, adaptive weighting and fusion module, hierarchical and visualization export module, and parameter library are executed sequentially through steps S1-S7.

9. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method described in any one of claims 1-7.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method for evaluating quality of cultivated land in sandstorm saline-alkali area

    CN119250351A

  • Systems and methods for generating habitat condition assessments

    US12080049B1