A Method and System for Soil Property Spectral Inversion Based on Surrogate Structural Variables
Patent Information
- Application Number
- CN202611110311.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-24
- Publication Date
- 2026-09-29
AI Technical Summary
[0022]为了解决现有技术土壤属性光谱反演方法在复杂土壤环境下预测稳定性不足,以及高精度联合建模方法依赖实验室化学变量、难以满足田间原位快速检测需求的问题,提出了一种基于代理结构变量的土壤属性光谱反演方法及系统,能够利用可见—近红外光谱数据、田间环境变量及相关辅助变量对土壤属性进行快速预测
[0034]1. 提供了一种关键结构关联变量的识别方法,可以增强土壤光谱反演的结构约束能力。
Smart Images

Figure CN122835977A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of soil spectral detection and quantitative inversion of soil properties, and in particular to a method and system for spectral inversion of soil properties based on surrogate structural variables. Background Technology
[0002] Soil organic matter, soil nutrients, pH, and related physicochemical properties are important indicators for evaluating soil fertility, soil quality, and farmland productivity. Traditional soil property determination usually relies on laboratory chemical analysis. Although the results are relatively accurate, the process of sample collection, transportation, pretreatment, and chemical detection is complex and the detection cycle is long, making it difficult to meet the needs of large-scale, rapid, and in-situ soil monitoring.
[0003] Visible-near-infrared spectroscopy can indirectly characterize soil organic matter, moisture, clay, iron oxides, and other mineral components by observing the reflection response of soil samples to electromagnetic radiation in specific wavelengths. It features rapid detection, minimal sample damage, and the ability to simultaneously predict multiple indicators. Therefore, soil property inversion methods based on visible-near-infrared spectroscopy have become an important technical approach for rapid soil detection and digital soil mapping.
[0004] Currently, the existing technologies for soil property inversion using visible-near-infrared spectroscopy mainly include the following categories.
[0005] The first category is data-driven inversion methods based on pure spectral features. These methods typically preprocess the raw spectra, performing smoothing, denoising, scattering correction, derivative transformation, standard normal variable transformation, or band selection. The processed spectral features are then input into partial least squares regression, support vector machines, random forests, neural networks, or other machine learning models to establish a statistical mapping relationship between spectral features and target soil properties. The advantages of this type of method are its relatively simple process, the fact that the input data mainly comes from spectral acquisition, and its ability to maintain the characteristics of rapid, low-cost, and non-destructive testing.
[0006] The second category is inversion methods based on joint modeling of spectral data and field environmental variables. To improve the applicability of pure spectral models in complex soil environments, existing technologies also include methods that jointly model readily available external variables such as soil moisture content, electrical conductivity, pH, temperature, topographic factors, or climatic factors with spectral features. These methods typically concatenate the spectral matrix with the aforementioned auxiliary variables before inputting it into the prediction model to enhance the model's ability to represent differences in sample background. Compared to pure spectral models, these methods can, to some extent, supplement the external environmental information of soil samples, which is beneficial for improving prediction accuracy in certain scenarios.
[0007] The third category is inversion methods based on joint modeling of spectral data and laboratory-measured chemical variables. For soil samples with complex chemical compositions or significant environmental gradients, some existing technologies further incorporate total nitrogen, available iron, total iron, clay content, cation exchange capacity, or other laboratory-measured soil physicochemical indicators, inputting them along with spectral characteristics into the prediction model. These variables can reflect the internal component structure or chemical state of the soil, playing a role in improving the prediction accuracy of target attributes, especially when the spectral response of the target attribute is affected by other soil components; laboratory chemical variables can provide additional constraint information for the model.
[0008] While the aforementioned existing technologies can predict soil properties under certain conditions, their technical approaches primarily rely on two approaches: one is to directly utilize spectral characteristics or establish statistical mapping relationships between spectra and low-cost field variables; the other is to directly introduce laboratory-measured chemical variables to enhance the model's input information. The former, while maintaining the advantage of rapid detection, often struggles to fully characterize the impact of internal chemical structure changes on the spectral response of target properties in soil systems with significant differences in chemical environment, obvious pH gradients, or substantial changes in iron oxide activity. The latter, while improving the model's ability to resolve complex soil backgrounds, still requires additional laboratory chemical analysis in the unknown sample detection stage, making it difficult to meet the application requirements of in-situ, low-cost, and rapid field detection. Therefore, existing technologies still lack a technical solution that can introduce key internal structural information and improve the stability of soil property spectral inversion without relying on laboratory chemical determination of unknown samples.
[0009] Specifically, when dealing with soil samples that are complex in type, have obvious pH gradients, have large variations in iron oxide activity, or have significant differences in the state of organic matter, the following technical shortcomings still exist.
[0010] 1. Pure spectral inversion models are difficult to characterize changes in internal chemical structure and lack predictive stability under complex environments.
[0011] Existing pure spectral inversion methods mainly rely on the statistical mapping relationship between spectral features and target soil properties. These methods typically assume that there is a relatively consistent spectral response pattern between training samples and unknown samples. However, in actual soil systems, the spectral expression of target properties is not only affected by their own content, but also by factors such as soil pH, water content, iron and aluminum oxides, mineral composition, and organic-mineral binding state.
[0012] When soil samples are exposed to different acidic or alkaline environments or different degrees of weathering, the activity of iron oxides, the form of organic matter, and the state of related nutrients may change, leading to a shift in the response characteristics of the same target attribute in the spectral space. Pure spectral models, which only use reflectance spectra as input, cannot explicitly characterize the aforementioned changes in internal chemical structure, and are prone to increased prediction errors and decreased stability under cross-regional, cross-soil-type, or strong environmental gradient conditions.
[0013] 2. Simple splicing of conventional field variables is insufficient to fully supplement internal structural information.
[0014] To improve the applicability of pure spectral models, some existing techniques incorporate readily available field environmental variables such as pH, water content, electrical conductivity, temperature, and topographic factors into the model along with spectral data. While this approach can supplement some external environmental information, it typically employs a simple feature concatenation method, directly adding field variables as additional inputs to the model.
[0015] This approach primarily enhances the model's ability to identify differences in the external background of the sample, but it struggles to further characterize the correlation between the target attribute and internal structural variables such as total nitrogen, available iron, iron oxides, and organic-mineral binding state. For complex systems such as highly weathered or acidic soils, relying solely on spectral data and conventional field variables may still fail to fully explain the sources of variation in the spectral response of the target attribute, resulting in limited improvement in model accuracy.
[0016] 3. Directly introducing laboratory chemical variables will increase the detection cost and make it difficult to meet the needs of in-situ rapid detection.
[0017] To further improve prediction accuracy, some existing technologies input laboratory-measured variables such as total nitrogen, available iron, total iron, clay content, and cation exchange capacity into the model along with spectral data. These variables can reflect the internal chemical composition or structural state of the soil and do indeed help improve the model's ability to characterize complex soil backgrounds.
[0018] However, the aforementioned laboratory chemical variables typically require processes such as soil sampling, air drying, grinding, sieving, digestion, extraction, and instrumental measurement, resulting in a lengthy testing cycle and incurring costs for reagents, equipment, and labor. For in-situ field testing of unknown samples, if laboratory measurements of TN, AFe, etc., are still required in advance, the advantages of rapid, low-cost, and non-destructive application of spectroscopic detection are diminished.
[0019] 4. Existing methods lack a closed-loop process of "key structural variable identification - proxy deduction - target inversion".
[0020] In existing technologies, auxiliary variables are typically introduced in two ways: one is to directly combine low-cost variables such as pH, water content, and conductivity based on experience; the other is to directly use laboratory-measured chemical variables to enhance the model input. While the former is convenient for field acquisition, it is difficult to fully characterize internal structural information; the latter, while supplementing internal chemical constraints, relies on laboratory testing. Therefore, the core problem with existing technologies is that pure spectroscopic models cannot fully characterize changes in internal chemical structure under environmental gradients such as pH, while joint models that directly introduce laboratory-measured variables such as TN and AFe cannot meet the needs of rapid field application.
[0021] Therefore, current technologies lack a complete workflow that can first identify key structural correlation variables from calibration samples, then use spectral data and low-cost field variables to deduce surrogate values for these key structural variables, and finally use the surrogate variables for soil property inversion during the unknown sample detection stage. In other words, current technologies struggle to simultaneously meet the following three requirements: first, to introduce internal structural constraint information under complex environments; second, to avoid laboratory chemical determinations during the unknown sample detection stage; and third, to maintain the speed, low cost, and predictive stability of soil property inversion. Summary of the Invention
[0022] To address the shortcomings of existing soil property spectral inversion methods in predicting in complex soil environments, and the reliance of high-precision co-modeling methods on laboratory chemical variables, which makes them unsuitable for rapid in-situ field detection, a soil property spectral inversion method and system based on surrogate structural variables is proposed. This method and system can rapidly predict soil properties using visible-near-infrared spectral data, field environmental variables, and relevant auxiliary variables. The technical solution of this invention is as follows.
[0023] A method for spectral inversion of soil properties based on surrogate structural variables, comprising the following steps: A. Obtain multi-source data from multiple calibrated soil samples, including spectral data; B. Preprocess the spectral data of multiple calibrated soil samples to obtain standardized spectral characteristics; C. Based on multi-source data from multiple calibrated soil samples, determine key structural correlation variables; D. Train the proxy structural variable inference model using key structural correlation variables as output; E. For unknown soil samples to be tested, the surrogate structural variables of the unknown soil samples are obtained by using a surrogate structural variable inference model. F. Calculate the predicted values of target soil properties from the fusion and inversion of unknown soil samples.
[0024] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, obtaining multi-source data of calibrated soil samples includes: Multiple calibrated soil samples were acquired, and for each calibrated soil sample, spectral data, low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties were collected or measured. Among them, the spectral data of the calibrated soil samples are used to characterize the reflectance spectral response of the soil samples, including visible-near infrared spectral data, near infrared spectral data, mid-infrared spectral data, hyperspectral image data, or other spectral data that can be used to characterize changes in soil properties; Low-cost field environmental variables for calibrating soil samples include variables obtained through portable sensors, in-situ field testing equipment, rapid testing devices, or external data sources, including soil pH, water content, electrical conductivity, temperature, salinity, sampling depth, geographical location, topographic factors, or climatic factors. Laboratory chemical structure variables for soil samples include variables that reflect the internal chemical composition, mineral composition, nutrient status, or structural status of the soil and need to be obtained through laboratory analysis, including soil nutrient variables, metal oxide-related variables, mineral composition variables, particle composition variables, cation exchange-related variables, or other soil physicochemical indicators. The target soil properties for calibrating soil samples include soil properties that need to be predicted by spectral inversion methods, including one or more of soil organic matter, soil organic carbon, soil nutrients, pH, salinity, cation exchange capacity, or other soil physicochemical properties.
[0025] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, the spectral data of multiple calibrated soil samples are preprocessed to obtain standardized spectral features, including: The preprocessing operations include one or more of the following: abnormal band removal, smoothing, band aggregation, standard normal variable transformation, multivariate scattering correction, derivative transformation, normalization, and feature band selection. After preprocessing, standardized spectral characteristics of the calibrated soil samples are obtained for subsequent structural analysis, surrogate structural variable deduction, and target soil property inversion.
[0026] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, the key structural correlation variables are determined based on multi-source data from multiple calibrated soil samples, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, a structural learning method is used to construct structural relationships among variables in the dataset of the calibrated soil samples. The structural learning method can determine one or more key structurally related variables based on the conditional dependencies, stable associations, node connections, edge frequency, or the degree of association with the target soil properties.
[0027] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, the key structural correlation variables are determined based on multi-source data from multiple calibrated soil samples, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, Bayesian networks were used to construct the variable structure relationships, and a resampling strategy was combined to screen key structurally related variables that appeared stably.
[0028] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, training the surrogate structural variable inference model with key structural correlation variables as output includes: Using standardized spectral features and low-cost field environmental variables as inputs, and measured values of key structural correlation variables as outputs, a proxy structural variable inference model is trained. The surrogate structural variables for the training samples are generated using out-of-sample prediction, including: The training samples are divided into K subsets; each time, K-1 subsets are used to train the surrogate structure variable inference model, and surrogate prediction values are generated for the remaining 1 subset; after K iterations, the out-of-sample surrogate structure variable matrix corresponding to all training samples is obtained.
[0029] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, the predicted values of the target soil properties fused from the inversion of unknown soil samples include: The standardized spectral characteristics of unknown soil samples, low-cost field environmental variables, and surrogate structural variables are fused to form joint input features, which are then input into the target soil property inversion model to output the predicted value of the target soil property. The fusion inversion process of the target soil property inversion model organizes input features in the order of spectral basic information layer, field environment constraint layer and surrogate structure constraint layer. The spectral basic information layer is used to characterize the reflectance spectral response of the soil sample, the field environment constraint layer is used to supplement the external environmental state of the soil sample, and the surrogate structure constraint layer is used to supplement the internal structural information represented by key structural correlation variables.
[0030] This invention also includes a soil property spectral inversion system based on surrogate structural variables. This system comprises a multi-source data acquisition module, a spectral preprocessing module, a key structural correlation variable identification module, a surrogate structural variable inference module, a fusion inversion module, and a result output module. The multi-source data acquisition module is used to acquire multi-source data from multiple calibrated soil samples, including spectral data. The spectral preprocessing module is used to preprocess the spectral data of multiple calibrated soil samples to obtain standardized spectral characteristics; The key structural correlation variable identification module is used to determine key structural correlation variables based on multi-source data from multiple calibrated soil samples, and to train a proxy structural variable inference model with key structural correlation variables as output. The surrogate structural variable inference module is used to obtain surrogate structural variables for unknown soil samples to be detected using the surrogate structural variable inference model. The fusion and inversion module is used to fuse and invert the predicted values of target soil properties for unknown soil samples; The results output module is used to display, store, or upload the predicted soil property values.
[0031] Furthermore, in the soil property spectral inversion system based on surrogate structural variables of the present invention, the multi-source data acquisition module includes a spectral data unit, a low-cost field environmental variable unit, a laboratory chemical structural variable unit, and a target soil property measured value unit, wherein, The spectral data unit is used to collect spectral data for each calibrated soil sample. The spectral data is used to characterize the reflectance spectral response of the soil sample, including visible-near infrared spectral data, near-infrared spectral data, mid-infrared spectral data, hyperspectral image data, or other spectral data that can be used to characterize changes in soil properties. The low-cost field environmental variable unit is used to collect low-cost field environmental variable data for each calibrated soil sample; low-cost field environmental variables include variables obtained through portable sensors, in-situ field detection equipment, rapid detection devices or external data sources, including soil pH, water content, electrical conductivity, temperature, salinity, sampling depth, geographical location, topographic factors or climate factors. The Laboratory Chemical Structure Variable Unit is used to collect laboratory chemical structure variable data for each calibrated soil sample. Laboratory chemical structure variables include variables that reflect the internal chemical composition, mineral composition, nutrient status or structural status of the soil and need to be obtained through laboratory analysis, including soil nutrient variables, metal oxide-related variables, mineral composition variables, particle composition variables, cation exchange-related variables or other soil physicochemical indicators. The target soil property measurement unit is used to measure target soil property data for each calibrated soil sample. Target soil properties include soil properties that need to be predicted by spectral inversion methods, including one or more of soil organic matter, soil organic carbon, soil nutrients, pH, salinity, cation exchange capacity, or other soil physicochemical properties.
[0032] Furthermore, in the soil property spectral inversion system based on surrogate structural variables of the present invention, the key structural correlation variable identification module is used to determine key structural correlation variables based on multi-source data from multiple calibrated soil samples, and to train a surrogate structural variable inference model with the key structural correlation variables as output, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, a structural learning method is used to construct structural relationships between variables in the dataset of the calibrated soil samples. The structural learning method can determine one or more key structurally related variables based on the conditional dependencies, stable associations, node connections, edge frequency, or the degree of association with the target soil properties between variables. Alternatively, a Bayesian network can be used to construct the variable structure relationship, and a resampling strategy can be used to screen key structurally related variables that appear stably.
[0033] Through the above technical solutions, the present invention can achieve the following technical effects.
[0034] 1. A method for identifying key structural correlation variables is provided, which can enhance the structural constraint capability of soil spectral inversion.
[0035] Existing pure spectral inversion methods mainly rely on the statistical mapping relationship between spectral features and target soil properties, which makes it difficult to fully characterize the influence of factors such as pH, iron oxides, nitrogen state, and organic-mineral binding state on the spectral response of target properties. In contrast, this invention uses a structure learning method to identify key structural correlation variables related to the inversion of target soil properties from a calibration sample set containing spectral data, field environmental variables, laboratory chemical variables, and measured values of target properties.
[0036] In this way, the present invention can provide auxiliary information with clear structural constraints for subsequent model construction, avoiding information redundancy and modeling instability caused by selecting auxiliary variables based solely on experience or simply splicing variables.
[0037] 2. A proxy structural variable derivation method is provided, which can reduce the reliance on laboratory chemical determination in the unknown sample detection stage.
[0038] Existing high-precision joint inversion methods typically require direct input of laboratory-measured variables such as total nitrogen, available iron, total iron, and clay content. While these variables can improve the model's ability to characterize complex soil backgrounds, they increase sampling, pretreatment, and chemical analysis steps in the unknown sample detection stage, making it difficult to meet the needs for rapid, low-cost, and in-situ detection.
[0039] In the technical solution of this invention, during the offline training phase, spectral data and low-cost field environmental variables are used as inputs, and the measured values of key structural correlation variables are used as outputs to train a surrogate structural variable inference model. During the unknown sample detection phase, only the spectral data and field environmental variables of the soil sample to be tested need to be obtained to generate the corresponding surrogate structural variables. Therefore, this invention can introduce the internal structural constraint information represented by high-cost laboratory chemical variables without directly measuring them.
[0040] 3. A progressive fusion inversion method is provided to improve the accuracy and stability of soil property prediction under complex environments.
[0041] Existing auxiliary variable joint modeling methods typically employ simple feature stitching, making it difficult to distinguish the different roles of spectral base information, external environmental information, and internal structural constraint information in the model. To address this issue, this invention constructs a progressive fusion and inversion framework for spectral features, field environmental variables, and surrogate structural variables.
[0042] In the inversion framework of this invention, spectral features are used to characterize the basic reflectance response of soil samples, field environmental variables are used to supplement external environmental constraints, and surrogate structural variables are used to supplement internal structural constraints related to the inversion of target attributes. Through multi-level information fusion, this invention can reduce the adverse effects of complex chemical environmental changes on the spectral inversion of soil attributes and improve the predictive stability of the model in samples with different pH gradients, different soil backgrounds, or different regions.
[0043] 4. A modularly implementable framework for soil property spectral inversion system is provided.
[0044] This invention also provides a soil property spectral inversion system capable of performing the above-described methods. The system includes functional modules such as multi-source data acquisition, spectral preprocessing, identification of key structural correlation variables, surrogate structural variable deduction, progressive fusion prediction, and result output.
[0045] Through this system framework, only spectral data and low-cost field environmental variables of the soil sample to be tested can be collected during the unknown sample detection stage, and the prediction results of the target soil properties can be obtained through proxy structural variable inference and fusion inversion.
[0046] In summary, the advantages of this invention include establishing a soil property spectral inversion technology that takes into account both internal structural constraint information and the need for rapid field detection, thereby improving the accuracy, stability, and applicability of soil property inversion under complex environments without increasing the cost of laboratory chemical testing of unknown samples. Attached Figure Description
[0047] Figure 1 This is a flowchart of the soil property spectral inversion method based on surrogate structural variables according to a specific embodiment of the present invention.
[0048] Figure 2 This is a structural diagram of a soil property spectral inversion system module based on surrogate structural variables according to a specific embodiment of the present invention.
[0049] Figure 3 This is a schematic diagram of the variable association topology obtained by Bayesian network learning in the soil property spectral inversion method based on surrogate structural variables according to a specific embodiment of the present invention.
[0050] Figure 4 This is a comparison chart showing the effectiveness of the soil property spectral inversion method based on surrogate structural variables in a specific embodiment of the present invention in predicting soil organic matter content for different models.
[0051] Figure 5 This is a comparison chart showing the soil organic matter inversion effect under different pH conditions using the soil property spectral inversion method based on surrogate structural variables in a specific embodiment of the present invention. Detailed Implementation
[0052] The present invention will now be described in detail with reference to the accompanying drawings.
[0053] The following detailed exemplary embodiments are disclosed. However, the specific structural and functional details disclosed herein are merely for the purpose of describing exemplary embodiments.
[0054] However, it should be understood that the present invention is not limited to the specific exemplary embodiments disclosed, but covers all modifications, equivalents, and substitutions falling within the scope of this disclosure. Throughout the description of the drawings, the same reference numerals denote the same elements.
[0055] Referring to the accompanying drawings, the structures, proportions, sizes, etc., depicted in the drawings are merely for illustrative purposes to aid those skilled in the art in understanding and reading the content disclosed herein. They are not intended to limit the conditions under which the invention can be implemented and therefore have no substantial technical significance. Any modifications to the structure, changes in proportions, or adjustments to the size, without affecting the effects and objectives achieved by the invention, should still fall within the scope of the technical content disclosed herein. Furthermore, the positional limitations used in this specification are merely for clarity of description and are not intended to limit the scope of the invention. Changes or adjustments to their relative relationships, without substantially altering the technical content, should also be considered within the scope of the invention's implementation.
[0056] It should also be understood that the term “and / or” as used herein includes any and all combinations of one or more of the related listed items. Furthermore, it should be understood that when a component or unit is referred to as “connected” or “coupled” to another component or unit, it may be directly connected or coupled to the other component or unit, or there may be intermediate components or units. In addition, other words used to describe the relationship between components or units should be understood in the same manner (e.g., “between” versus “directly between,” “adjacent” versus “directly adjacent,” etc.).
[0057] Figure 1 This is a flowchart of a soil property spectral inversion method based on surrogate structural variables according to a specific embodiment of the present invention. The flowchart illustrates the overall process of acquiring multi-source data of calibrated soil samples, preprocessing spectral data, determining key structural correlation variables, training a surrogate structural variable inference model, generating surrogate structural variables for unknown samples, and fusing and inverting target soil properties. As shown in the figure, the specific embodiment of the present invention includes a soil property spectral inversion method based on surrogate structural variables, which includes the following steps: A. Obtain multi-source data from multiple calibrated soil samples, including spectral data; B. Preprocess the spectral data of multiple calibrated soil samples to obtain standardized spectral characteristics; C. Based on multi-source data from multiple calibrated soil samples, determine key structural correlation variables; D. Train the proxy structural variable inference model using key structural correlation variables as output; E. For unknown soil samples to be tested, the surrogate structural variables of the unknown soil samples are obtained by using a surrogate structural variable inference model. F. Calculate the predicted values of target soil properties from the fusion and inversion of unknown soil samples.
[0058] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, obtaining multi-source data of calibrated soil samples includes: Multiple calibrated soil samples were acquired, and for each calibrated soil sample, spectral data, low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties were collected or measured. Among them, the spectral data of the calibrated soil samples are used to characterize the reflectance spectral response of the soil samples, including visible-near infrared spectral data, near infrared spectral data, mid-infrared spectral data, hyperspectral image data, or other spectral data that can be used to characterize changes in soil properties; Low-cost field environmental variables for calibrating soil samples include variables obtained through portable sensors, in-situ field testing equipment, rapid testing devices, or external data sources, including soil pH, water content, electrical conductivity, temperature, salinity, sampling depth, geographical location, topographic factors, or climatic factors. Laboratory chemical structure variables for soil samples include variables that reflect the internal chemical composition, mineral composition, nutrient status, or structural status of the soil and need to be obtained through laboratory analysis, including soil nutrient variables, metal oxide-related variables, mineral composition variables, particle composition variables, cation exchange-related variables, or other soil physicochemical indicators. The target soil properties for calibrating soil samples include soil properties that need to be predicted by spectral inversion methods, including one or more of soil organic matter, soil organic carbon, soil nutrients, pH, salinity, cation exchange capacity, or other soil physicochemical properties.
[0059] The above steps form a calibration sample dataset. This dataset includes a spectral data matrix, a low-cost field environmental variable matrix, a laboratory chemical structure variable matrix, and measured values of target soil properties. It is used for subsequent identification of key structural correlation variables, training of proxy structural variable inference models, and training of target soil property inversion models.
[0060] In addition, as an optional implementation, the spectral data in the specific embodiments of the present invention are not limited to visible-near infrared spectra, but may also be near-infrared spectra, mid-infrared spectra, visible spectra, hyperspectral images or other data that can characterize the spectral response of soil samples.
[0061] The low-cost field environmental variables are not limited to soil pH, water content, and electrical conductivity, but may also include soil temperature, salinity, bulk density, topographic factors, climatic factors, sampling depth, geographical location, or other variables that can be obtained through rapid detection or external data sources.
[0062] The laboratory chemical structure variables and key structure-related variables are not limited to a specific soil physicochemical index, but may also include soil nutrient variables, metal oxide-related variables, mineral composition variables, particle composition variables, cation exchange-related variables, or other variables that can reflect the internal composition and structural state of the soil.
[0063] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, the spectral data of multiple calibrated soil samples are preprocessed to obtain standardized spectral features, including: The spectral data obtained in step A is preprocessed to obtain standardized spectral features. This preprocessing is used to reduce spectral noise, mitigate scattering effects, reduce redundant bands, and improve the stability of subsequent modeling.
[0064] The preprocessing operations include one or more of the following: abnormal band removal, smoothing, band aggregation, standard normal variable transformation, multivariate scattering correction, derivative transformation, normalization, and feature band selection. After preprocessing, standardized spectral characteristics of the calibrated soil samples are obtained for subsequent structural analysis, surrogate structural variable deduction, and target soil property inversion.
[0065] Specifically, in the soil property spectral inversion method based on surrogate structural variables in the specific embodiments of the present invention, the key structural correlation variables are determined based on multi-source data from multiple calibrated soil samples, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, a structural learning method is used to construct structural relationships among variables in the dataset of the calibrated soil samples. The structural learning method can determine one or more key structurally related variables based on the conditional dependencies, stable associations, node connections, edge frequency, or the degree of association with the target soil properties.
[0066] The key structural correlation variables refer to soil physicochemical variables that are related to the inversion of target soil properties, reflect the source of changes in the spectral response of target soil properties, and are not suitable for direct acquisition through rapid in-situ methods during the unknown sample detection stage. These variables can provide internal structural constraint information for the inversion of target soil properties.
[0067] Through this step, the specific implementation of the present invention can screen out variables that have a constraining effect on the inversion of target soil properties from the calibration sample dataset, rather than selecting auxiliary variables based solely on experience or directly stacking all laboratory chemical variables.
[0068] Specifically, in the soil property spectral inversion method based on surrogate structural variables in the specific embodiments of the present invention, the key structural correlation variables are determined based on multi-source data from multiple calibrated soil samples, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, Bayesian networks were used to construct the variable structure relationships, and a resampling strategy was combined to screen key structurally related variables that appeared stably.
[0069] Optionally, the key structural correlation variable identification method described in the specific embodiments of the present invention is not limited to using Bayesian networks, but may also employ structural equation modeling, graphical modeling, conditional independence testing, stability selection, correlation network analysis, or other methods capable of identifying variable structural relationships.
[0070] Specifically, in the soil property spectral inversion method based on surrogate structural variables of the present invention, training the surrogate structural variable inference model with key structural correlation variables as output includes: Using standardized spectral features and low-cost field environmental variables as inputs, and measured values of key structural correlation variables as outputs, a proxy structural variable inference model is trained. The surrogate structural variables for the training samples are generated using out-of-sample prediction, including: The training samples are divided into K subsets; each time, K-1 subsets are used to train the surrogate structure variable inference model, and surrogate prediction values are generated for the remaining 1 subset; after K iterations, the out-of-sample surrogate structure variable matrix corresponding to all training samples is obtained.
[0071] The surrogate structural variable inference model is used to establish a mapping relationship from "standardized spectral characteristics and low-cost field environmental variables" to "key structural correlation variables". Using this model, key structural correlation variables can be obtained by inferring corresponding surrogate values based on collected spectral data and low-cost field environmental variables, instead of directly measuring them during the unknown soil sample testing stage.
[0072] Specifically, the proxy structural variable inference model can be a partial least squares regression model, or it can be a support vector regression, random forest, gradient boosting tree, extreme gradient boosting, neural network, ensemble learning model or other regression prediction model.
[0073] In one implementation, a surrogate structural variable inference model is used to generate surrogate values for one or more key structural correlation variables. These surrogate values correspond to the respective key structural correlation variables and are used to replace the laboratory-measured values of the key structural correlation variables in subsequent inversion during the unknown sample detection phase.
[0074] During training, to avoid the target soil property inversion model indirectly accessing the true values of laboratory chemical structure variables from the same sample, out-of-sample prediction can be used to generate surrogate structure variables for the training samples. Specifically, the training samples are divided into K subsets; each time, K-1 subsets are used to train the surrogate structure variable inference model, and surrogate prediction values are generated for the remaining subset; after K iterations, the out-of-sample surrogate structure variable matrix corresponding to all training samples is obtained. This out-of-sample surrogate structure variable matrix is used for subsequent training of the target soil property inversion model.
[0075] It should be noted that the surrogate structural variable inference model is not limited to a specific regression algorithm. As long as it can generate surrogate values for key structural correlation variables based on spectral data and low-cost field environmental variables, it can be used as the surrogate structural variable inference model of this invention.
[0076] For unknown soil samples to be tested, their spectral data and low-cost field environmental variables are collected, and the spectral data are processed in the same or corresponding preprocessing method as in step B to obtain the standardized spectral characteristics of the unknown samples.
[0077] The standardized spectral characteristics and low-cost field environmental variables of the unknown soil sample are input into the surrogate structural variable inference model trained in step D to generate surrogate values for key structural correlation variables of the unknown sample.
[0078] In this step, surrogate structural variables for subsequent target soil property inversion can be obtained without determining the laboratory measured values of the corresponding key structural correlation variables for unknown samples. Therefore, this invention can reduce reliance on laboratory chemical analysis in the unknown sample detection stage while retaining the constraining effect of key structural correlation variables on the target soil property inversion.
[0079] Specifically, in the soil property spectral inversion method based on surrogate structural variables according to a specific embodiment of the present invention, the predicted values of the target soil properties fused from the inversion of unknown soil samples include: The standardized spectral characteristics of unknown soil samples, low-cost field environmental variables, and surrogate structural variables are fused to form joint input features, which are then input into the target soil property inversion model to output the predicted value of the target soil property. The fusion inversion process of the target soil property inversion model organizes input features in the order of spectral basic information layer, field environment constraint layer and surrogate structure constraint layer. The spectral basic information layer is used to characterize the reflectance spectral response of the soil sample, the field environment constraint layer is used to supplement the external environmental state of the soil sample, and the surrogate structure constraint layer is used to supplement the internal structural information represented by key structural correlation variables.
[0080] Specifically, the target soil property inversion model can be a partial least squares regression model, or it can be a support vector regression, random forest, gradient boosting tree, extreme gradient boosting, neural network, ensemble learning model or other models that can perform quantitative prediction of soil properties.
[0081] The target soil properties are not limited to soil organic matter, but may also include soil organic carbon, soil nutrients, pH, salinity, cation exchange capacity, water content, or other soil physicochemical properties.
[0082] The standardized spectral features of the unknown samples, low-cost field environmental variables, and surrogate structural variables generated in step E are fused to form joint input features, which are then input into the target soil property inversion model to output the predicted values of the target soil properties.
[0083] In one embodiment of the present invention, the joint input features include standardized spectral features, low-cost field environmental variables, and one or more surrogate structural variables. The target soil property inversion model outputs predicted values of the target soil properties based on the aforementioned joint input features.
[0084] In another embodiment of the present invention, the fusion inversion process organizes the input features in the order of spectral basic information layer, field environment constraint layer, and surrogate structural constraint layer. The spectral basic information layer characterizes the reflectance spectral response of the soil sample, the field environment constraint layer supplements the external environmental state of the soil sample, and the surrogate structural constraint layer supplements the internal structural information represented by key structural correlation variables.
[0085] Through the above steps, this invention can supplement internal structural constraint information by using proxy structural variables without measuring the laboratory measured values of key structural correlation variables in the unknown sample detection stage, thereby improving the accuracy and stability of soil property spectral inversion.
[0086] Figure 2 This is a structural diagram of a soil property spectral inversion system based on surrogate structural variables according to a specific embodiment of the present invention. The diagram illustrates the connection and data flow relationships between the multi-source data acquisition module, spectral preprocessing module, key structural correlation variable identification module, surrogate structural variable inference module, fusion inversion module, and result output module in the system of the present invention. As shown in the figure, the specific embodiment of the present invention includes a soil property spectral inversion system based on surrogate structural variables. This system includes a multi-source data acquisition module, a spectral preprocessing module, a key structural correlation variable identification module, a surrogate structural variable inference module, a fusion inversion module, and a result output module, wherein; A multi-source data acquisition module is used to acquire multi-source data from multiple calibrated soil samples, including spectral data. In one embodiment, the module includes a spectral acquisition device, a field environmental variable acquisition device, and a data transmission interface. The spectral acquisition device is used to obtain the spectral data of the soil samples, the field environmental variable acquisition device is used to obtain low-cost field environmental variables of the soil samples, and the data transmission interface is used to transmit the acquired data to subsequent modules.
[0087] The spectral preprocessing module is used to preprocess the spectral data of multiple calibrated soil samples to obtain standardized spectral features. The preprocessing includes one or more of the following: outlier band removal, smoothing, band aggregation, standard normal variable transformation, multivariate scattering correction, derivative transformation, normalization, or feature band screening.
[0088] The key structural association variable identification module is used to determine key structural association variables based on multi-source data from multiple calibrated soil samples, and to train a surrogate structural variable inference model using these key structural association variables as outputs. This includes constructing variable structure relationships and identifying key structural association variables during the offline calibration phase, based on low-cost field environmental variables, laboratory chemical structural variables, and measured values of target soil properties in the calibration samples. Therefore, the key structural association variable identification module is primarily used in the system initialization, model training, or model update phases.
[0089] The surrogate structural variable inference module is used to obtain surrogate structural variables for unknown soil samples to be tested using a surrogate structural variable inference model; the surrogate values are used to replace the laboratory measured values of the corresponding key structural correlation variables in the unknown sample detection stage and are input into the fusion inversion module.
[0090] The fusion and inversion module is used to fuse and invert the predicted values of target soil properties for unknown soil samples; The results output module is used to display, store, or upload soil property prediction results. It can also record sample number, collection time, collection location, spectral data, low-cost field environmental variables, surrogate structural variables, and predicted values.
[0091] Through the above system structure, the specific implementation of the present invention can support the generation of proxy structural variables and the inversion of target soil properties by relying only on spectral data and low-cost field environmental variables in the unknown sample detection stage.
[0092] The system deployment method is not limited to portable computing terminals, but can also be deployed on host computers, servers, edge computing devices, mobile terminals, field inspection vehicle platforms, cloud platforms, or other software and hardware systems with data acquisition, model calculation, and result output capabilities.
[0093] The above alternative implementation methods can be used individually or in combination according to different soil types, detection targets, spectrometer types and application scenarios. None of these methods affect the basic technical concept of this invention to achieve spectral inversion of soil properties through proxy structural variables.
[0094] Furthermore, in the soil property spectral inversion system based on surrogate structural variables of the present invention, the multi-source data acquisition module includes a spectral data unit, a low-cost field environmental variable unit, a laboratory chemical structural variable unit, and a target soil property measured value unit, wherein, The spectral data unit is used to collect spectral data for each calibrated soil sample. The spectral data is used to characterize the reflectance spectral response of the soil sample, including visible-near infrared spectral data, near-infrared spectral data, mid-infrared spectral data, hyperspectral image data, or other spectral data that can be used to characterize changes in soil properties. The low-cost field environmental variable unit is used to collect low-cost field environmental variable data for each calibrated soil sample; low-cost field environmental variables include variables obtained through portable sensors, in-situ field detection equipment, rapid detection devices or external data sources, including soil pH, water content, electrical conductivity, temperature, salinity, sampling depth, geographical location, topographic factors or climate factors. The Laboratory Chemical Structure Variable Unit is used to collect laboratory chemical structure variable data for each calibrated soil sample. Laboratory chemical structure variables include variables that reflect the internal chemical composition, mineral composition, nutrient status or structural status of the soil and need to be obtained through laboratory analysis, including soil nutrient variables, metal oxide-related variables, mineral composition variables, particle composition variables, cation exchange-related variables or other soil physicochemical indicators. The target soil property measurement unit is used to measure target soil property data for each calibrated soil sample. Target soil properties include soil properties that need to be predicted by spectral inversion methods, including one or more of soil organic matter, soil organic carbon, soil nutrients, pH, salinity, cation exchange capacity, or other soil physicochemical properties.
[0095] Furthermore, in the soil property spectral inversion system based on surrogate structural variables of the present invention, the key structural correlation variable identification module is used to determine key structural correlation variables based on multi-source data from multiple calibrated soil samples, and to train a surrogate structural variable inference model with the key structural correlation variables as output, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, a structural learning method is used to construct structural relationships between variables in the dataset of the calibrated soil samples. The structural learning method can determine one or more key structurally related variables based on the conditional dependencies, stable associations, node connections, edge frequency, or the degree of association with the target soil properties between variables. Alternatively, a Bayesian network can be used to construct the variable structure relationship, and a resampling strategy can be used to screen key structurally related variables that appear stably.
[0096] The following two application examples, using soil organic matter inversion in tropical acidic soils as examples, illustrate the data acquisition, model construction, surrogate structural variable generation, and prediction effects of the specific embodiments of the present invention in specific soil property prediction tasks. It should be noted that the following application examples are only used to illustrate specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention.
[0097]
Application Example 1
[0098] 1. Experimental conditions and data acquisition
[0099] Sixty-eight tropical acidic soil samples were selected as the experimental dataset. Reflectance spectral data of the soil samples in the range of 350–2500 nm were acquired under laboratory conditions using an ASD FieldSpec 4 spectrometer. Low-cost field environmental variables, including soil pH, water content (WC), and electrical conductivity (EC), were simultaneously measured. Laboratory chemical structure variables, including total nitrogen (TN) and available iron (AFe), were also measured. The target soil property, namely soil organic matter (SOM) content, was also determined.
[0100] In this application example, the target soil property is soil organic matter, the key structural correlation variables are total nitrogen (TN) and available iron (AFe), and the corresponding generated surrogate structural variables are the total nitrogen surrogate value TN_proxy and the available iron surrogate value AFE_proxy.
[0101] 2. Parameter Setting and Implementation Process
[0102] First, the raw spectral data is preprocessed. Specifically, the raw spectra in the range of 350–2500 nm are aggregated in 5 nm bands and subjected to Standard Normal Transform (SNV), resulting in 430 effective bands as standardized spectral features.
[0103] Secondly, key structural association variables were identified. Low-cost field environmental variables, laboratory chemical structural variables, and measured values of target soil properties were input into a Bayesian network for structure learning. The K2Score scoring function was used, the HillClimbSearch structure search algorithm was employed, the bootstrap resampling times were set to 200, and the edge retention threshold was set to 40%. Based on the variable structure learning results, total nitrogen (TN) and available iron (AFe) were identified as key structural association variables in this embodiment, and their variable association topology is as follows: Figure 3 As shown.
[0104] Next, a surrogate structural variable inference model was trained. Using 430 preprocessed spectral bands, along with pH, WC, and EC, as inputs, and measured TN and AFe as outputs, a partial least squares regression (PLSR) model was established to infer surrogate structural variables. To avoid information leakage during training, out-of-sample prediction with 10-fold cross-validation was used to generate surrogate structural variables for the training samples, namely TN_proxy and AFe_proxy. When generating surrogate structural variables for the training samples, an out-of-sample prediction method was used, ensuring that the surrogate structural variable for each sample was generated by a model that did not include the measured key structural variables for that sample, thus reducing the risk of information leakage during the generation of surrogate structural variables.
[0105] Finally, a soil organic matter fusion inversion model was constructed. Standardized spectral features, pH, WC, EC, TN_proxy, and AFe_proxy were concatenated into a joint input feature matrix. Measured SOM values were used as training labels, and the final soil organic matter inversion model was established using PLSR. In this application example, the number of potential components nc in PLSR was set to 15.
[0106] 3. Implementation Results
[0107] In this embodiment, a model using only standardized spectral features is used as the pure spectral baseline model, while a model incorporating standardized spectral features, low-cost field environmental variables, and surrogate structural variables is used as the method of this invention. Test results are as follows: Figure 4 As shown.
[0108] Depend on Figure 4The results show that, compared with the pure spectral baseline model, the method of this invention increases the predicted R² of soil organic matter from 0.484 to 0.618, a relative improvement of approximately 27.7%; and decreases the RMSE from 3.90 g / kg to 3.35 g / kg, a relative reduction of approximately 14.1%. These results demonstrate that, even when TN and AFe are not directly measured in unknown samples, this invention can still supplement internal structural constraint information through proxy structural variables, thereby improving the accuracy of soil organic matter spectral inversion.
[0109]
Application Example 2
[0110] 1. Test conditions
[0111] To verify the predictive stability of the method of the present invention under different pH conditions, the 608 soil samples in Application Example 1 were divided into two sample groups according to the median pH of 5.37: a low pH group and a high pH group. The low pH group consisted of samples with pH ≤ 5.37, and the high pH group consisted of samples with pH > 5.37.
[0112] 2. Parameter Setting and Implementation Process
[0113] First, pure spectral models were established in both the low and high pH groups, and their prediction results were tested. Then, following the method in Application Example 1, low-cost field environmental variables and TN_proxy and AFe_proxy generated from the surrogate structural variable model were introduced into different pH groups to re-establish the soil organic matter fusion inversion model. In this application example, the number of potential components nc in the final PLSR inversion model was set to 12.
[0114] It should be noted that TN_proxy and AFe_proxy in this application example are not laboratory measured values of unknown samples, but surrogate structural variables derived from standardized spectral characteristics and low-cost field environmental variables.
[0115] 3. Implementation Results
[0116] Test results are as follows Figure 5 As shown. By Figure 5The results show that the R² of the pure spectral model in the low pH group is 0.247, indicating that the model's predictive stability is low when relying solely on spectral features. After applying the method of this invention, the R² of the low pH group increased to 0.584, and the R² of the high pH group reached 0.686. This result demonstrates that this invention, by introducing internal structural constraint information through surrogate structural variables, can mitigate the adverse effects of changes in soil spectral response under different pH conditions on organic matter inversion results, thereby improving the model's predictive stability in complex soil environments.
[0117] As can be seen from the detailed descriptions of the above specific implementation methods and application examples, the key points of the soil property spectral inversion method and system based on surrogate structural variables in the specific implementation methods of the present invention are mainly reflected in the following four aspects:
[0118] 1. Determine key structural correlation variables based on the structural relationships of the calibrated samples.
[0119] In existing soil spectral inversion methods, the selection of auxiliary variables usually relies on empirical judgment, or directly splices together available field environmental variables, laboratory chemical variables and spectral features, which can easily lead to a lack of specificity in the selection of auxiliary variables, redundancy of input information or insufficient model stability.
[0120] One of the key aspects of this invention is that it utilizes low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties from calibrated samples to construct variable structure relationships, and identifies key structural correlation variables that constrain the inversion of target soil properties. These key structural correlation variables are not arbitrarily selected auxiliary inputs, but rather internal structural variables obtained through variable structure relationship analysis that reflect the source of changes in the spectral response of target soil properties.
[0121] In this way, the present invention can perform structured screening of auxiliary variables before model construction, providing a more targeted input basis for subsequent deduction of proxy structure variables and inversion of target soil properties.
[0122] 2. Transform key structural correlation variables into surrogate structural variables usable during the unknown sample phase.
[0123] While existing high-precision joint inversion methods can directly incorporate laboratory chemical structure variables to improve prediction results, the corresponding variables still need to be measured in the laboratory during the unknown sample detection stage, which increases detection time, detection cost and operation process, making it difficult to meet the needs of rapid in-situ detection in the field.
[0124] The second key point of this invention is that, in the offline training phase, spectral data and low-cost field environmental variables are used as inputs, and the measured values of key structural correlation variables are used as outputs to train the surrogate structural variable inference model; in the unknown sample detection phase, the laboratory measured values of key structural correlation variables are no longer directly measured, but the corresponding surrogate values are generated through the trained surrogate structural variable inference model.
[0125] In this way, the present invention transforms the laboratory chemical structure information that can be obtained in the training phase into surrogate structural variables that can be derived from low-cost data in the application phase, so that unknown samples can still introduce internal structural constraint information represented by key structural correlation variables without adding laboratory chemical determination steps.
[0126] 3. The data usage methods are separated between the training phase and the unknown sample application phase.
[0127] The third key point of this invention is: clearly distinguishing the data usage methods between the offline training phase and the unknown sample application phase.
[0128] During the offline training phase, spectral data from calibration samples, low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties can be used for key structural correlation variable identification, proxy structural variable inference model training, and target soil property inversion model training.
[0129] In the unknown sample application phase, only spectral data and low-cost field environmental variables of the sample to be tested are collected, and laboratory measured values of key structural correlation variables are no longer obtained. The system generates surrogate structural variables through a surrogate structural variable inference model, and inputs them together with spectral features and low-cost field environmental variables into the target soil property inversion model.
[0130] By separating the training and application phases as described above, this invention can both utilize laboratory chemical structure information from calibrated samples for model construction and avoid dependence on laboratory chemical determinations in the unknown sample detection phase, thus balancing the integrity of model information and the convenience of actual detection.
[0131] 4. A fusion inversion framework based on spectral features, field environmental variables, and surrogate structural variables
[0132] Existing auxiliary variable joint modeling methods typically employ simple feature splicing, making it difficult to distinguish the roles of spectral basic information, external environmental information, and internal structural constraint information in the inversion of target soil properties.
[0133] The fourth key point of this invention is that it uses standardized spectral features, low-cost field environmental variables, and surrogate structural variables as inputs to the target soil property inversion model, forming a fusion inversion framework. Specifically, spectral features characterize the basic reflectance response of the soil sample, low-cost field environmental variables supplement the external environmental state of the sample, and surrogate structural variables supplement the internal structural constraint information represented by key structural correlation variables.
[0134] Through this fusion inversion framework, the present invention can introduce internal structural constraint information into the soil property inversion process without directly measuring laboratory chemical structure variables during the unknown sample detection stage, thereby improving the prediction accuracy and stability of soil property spectral inversion under complex environments.
[0135] Through the above technical solutions, the present invention can achieve the following technical effects.
[0136] 1. Reduce reliance on laboratory chemical assays during the detection of unknown samples.
[0137] Existing high-precision joint inversion methods typically require the direct acquisition of measured values of laboratory chemical structure variables during the unknown sample detection stage. This results in the detection process still including steps such as sampling, pretreatment, chemical analysis, and instrument measurement, which makes it difficult to meet the needs of rapid field detection.
[0138] This invention utilizes a surrogate structural variable inference model. During the unknown sample detection stage, only the spectral data of the soil sample to be tested and low-cost field environmental variables are needed to generate corresponding surrogate structural variables, which can then be used for target soil property inversion. Therefore, this invention reduces the dependence of unknown samples on measured values of laboratory chemical structural variables, simplifies the detection process, reduces detection costs, and improves the applicability of rapid soil property detection.
[0139] 2. Introduce internal structural constraint information while maintaining low-cost input.
[0140] Pure spectral inversion methods rely mainly on spectral features, which are difficult to fully characterize the impact of changes in internal chemical structure in complex soil environments on the spectral response of target properties. While directly introducing laboratory chemical structure variables can supplement internal structural information, it will increase the cost of detecting unknown samples.
[0141] This invention utilizes a "key structural correlation variable identification—surrogate structural variable derivation" approach to transform laboratory chemical structural information obtainable during the training phase into surrogate structural variables that can be deduced from low-cost data during the unknown sample phase. These surrogate structural variables can serve as internal structural constraint information in the inversion of target soil properties, thereby enhancing the model's ability to characterize complex soil backgrounds without adding the laboratory chemical measurement step for unknown samples.
[0142] 3. Improve the prediction accuracy of spectral inversion of target soil properties.
[0143] This invention inputs standardized spectral features, low-cost field environmental variables, and surrogate structural variables into the target soil property inversion model, enabling the model to utilize not only the spectral response information of soil samples, but also external environmental conditions and internal structural constraints.
[0144] In one application example, soil organic matter was used as the target soil property for inversion. Compared with a baseline model that only uses spectral features, the model determination coefficient R² increased from 0.484 to 0.618 using the method of this invention, a relative improvement of approximately 27.7%; the root mean square error RMSE decreased from 3.90 g / kg to 3.35 g / kg, a relative reduction of approximately 14.1%. These results demonstrate that this invention can improve the accuracy of spectral inversion of target soil properties without directly measuring the corresponding laboratory chemical structure variables in unknown samples.
[0145] 4. Improve the stability of predictions under complex environmental conditions.
[0146] For samples with significant pH gradients, large variations in iron oxide activity, or significant differences in soil component structure, the spectral response of the target soil properties is easily affected by changes in environmental background and internal structure. Under such conditions, pure spectral models may experience increased prediction errors or decreased stability.
[0147] In one application example, samples were grouped and tested according to pH conditions. In the low-pH sample group, the R² of the pure spectral model was 0.247; after applying the method of this invention, the R² of the low-pH sample group increased to 0.584. This result demonstrates that this invention, by introducing internal structural constraint information through surrogate structural variables, can mitigate the adverse effects of spectral response changes under complex acidic environments on the inversion of target soil properties, and improve the predictive stability of the model under complex environmental conditions.
[0148] 5. Facilitates the development of a deployable rapid soil property detection system.
[0149] This invention not only provides an inversion method, but also a system architecture capable of executing that method. This approach facilitates the subsequent deployment of the model to portable computing terminals, edge computing devices, field testing equipment, or soil data management platforms.
[0150] In the unknown sample detection phase, the system only needs to collect spectral data of the soil sample to be tested and low-cost field environmental variables to complete the generation of proxy structural variables and the inversion of target soil properties. This approach facilitates the deployment of the model to portable computing terminals, edge computing devices, field testing equipment, or soil data management platforms, improving the automation level and engineering application adaptability of the soil property detection process.
[0151] In summary, the specific embodiments of the present invention can reduce the reliance on laboratory chemical structure variable determination in the unknown sample detection stage, and at the same time supplement internal structural constraint information by proxy structure variables, thereby improving the prediction accuracy, stability and rapid detection applicability of soil property spectral inversion.
[0152] The foregoing description illustrates and describes several preferred embodiments of the present invention. However, as mentioned above, it should be understood that the present invention is not limited to the forms disclosed in this specification and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the inventive concept described in this specification through the foregoing teachings or techniques or knowledge in related fields. Any modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A method for spectral inversion of soil properties based on surrogate structural variables, characterized in that, The method includes the following steps: A. Obtain multi-source data from multiple calibrated soil samples, including spectral data; B. Preprocess the spectral data of multiple calibrated soil samples to obtain standardized spectral characteristics; C. Based on multi-source data from multiple calibrated soil samples, determine key structural correlation variables; D. Train the proxy structural variable inference model using key structural correlation variables as output; E. For unknown soil samples to be tested, the surrogate structural variables of the unknown soil samples are obtained by using a surrogate structural variable inference model. F. Calculate the predicted values of target soil properties from the fusion and inversion of unknown soil samples.
2. The method for spectral inversion of soil properties based on surrogate structural variables as described in claim 1, characterized in that, Obtaining multi-source data from calibrated soil samples includes: Multiple calibrated soil samples were acquired, and for each calibrated soil sample, spectral data, low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties were collected or measured. Among them, the spectral data of the calibrated soil samples are used to characterize the reflectance spectral response of the soil samples, including visible-near infrared spectral data, near infrared spectral data, mid-infrared spectral data, hyperspectral image data, or other spectral data that can be used to characterize changes in soil properties; Low-cost field environmental variables for calibrating soil samples include variables obtained through portable sensors, in-situ field testing equipment, rapid testing devices, or external data sources, including soil pH, water content, electrical conductivity, temperature, salinity, sampling depth, geographical location, topographic factors, or climatic factors. Laboratory chemical structure variables for soil samples include variables that reflect the internal chemical composition, mineral composition, nutrient status, or structural status of the soil and need to be obtained through laboratory analysis, including soil nutrient variables, metal oxide-related variables, mineral composition variables, particle composition variables, cation exchange-related variables, or other soil physicochemical indicators. The target soil properties for calibrating soil samples include soil properties that need to be predicted by spectral inversion methods, including one or more of soil organic matter, soil organic carbon, soil nutrients, pH, salinity, cation exchange capacity, or other soil physicochemical properties.
3. The method for spectral inversion of soil properties based on surrogate structural variables as described in claim 1, characterized in that, Preprocessing of spectral data from multiple calibrated soil samples yielded standardized spectral characteristics, including: The preprocessing operations include one or more of the following: abnormal band removal, smoothing, band aggregation, standard normal variable transformation, multivariate scattering correction, derivative transformation, normalization, and feature band selection. After preprocessing, standardized spectral characteristics of the calibrated soil samples are obtained for subsequent structural analysis, surrogate structural variable deduction, and target soil property inversion.
4. The method for spectral inversion of soil properties based on surrogate structural variables as described in claim 2, characterized in that, Based on multi-source data from multiple calibrated soil samples, key structural correlation variables were identified, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, a structural learning method is used to construct structural relationships among variables in the dataset of the calibrated soil samples. The structural learning method can determine one or more key structurally related variables based on the conditional dependencies, stable associations, node connections, edge frequency, or the degree of association with the target soil properties.
5. The method for spectral inversion of soil properties based on surrogate structural variables according to claim 2, characterized in that, Based on multi-source data from multiple calibrated soil samples, key structural correlation variables were identified, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, Bayesian networks were used to construct the variable structure relationships, and a resampling strategy was combined to screen key structurally related variables that appeared stably.
6. The method for spectral inversion of soil properties based on surrogate structural variables according to claim 2, characterized in that, Using key structural correlation variables as output, training a proxy structural variable inference model includes: Using standardized spectral features and low-cost field environmental variables as inputs, and measured values of key structural correlation variables as outputs, a proxy structural variable inference model is trained. The surrogate structural variables for the training samples are generated using out-of-sample prediction, including: The training samples are divided into K subsets; each time, K-1 subsets are used to train the surrogate structure variable inference model, and surrogate prediction values are generated for the remaining 1 subset; after K iterations, the out-of-sample surrogate structure variable matrix corresponding to all training samples is obtained.
7. The method for spectral inversion of soil properties based on surrogate structural variables as described in claim 2, characterized in that, The predicted values of target soil properties obtained through fusion inversion of unknown soil samples include: The standardized spectral characteristics of unknown soil samples, low-cost field environmental variables, and surrogate structural variables are fused to form joint input features, which are then input into the target soil property inversion model to output the predicted value of the target soil property. The fusion inversion process of the target soil property inversion model organizes input features in the order of spectral basic information layer, field environment constraint layer and surrogate structure constraint layer. The spectral basic information layer is used to characterize the reflectance spectral response of the soil sample, the field environment constraint layer is used to supplement the external environmental state of the soil sample, and the surrogate structure constraint layer is used to supplement the internal structural information represented by key structural correlation variables.
8. A soil property spectral inversion system based on surrogate structural variables, characterized in that, The system includes a multi-source data acquisition module, a spectral preprocessing module, a key structural correlation variable identification module, a surrogate structural variable inference module, a fusion inversion module, and a result output module. The multi-source data acquisition module is used to acquire multi-source data from multiple calibrated soil samples, including spectral data. The spectral preprocessing module is used to preprocess the spectral data of multiple calibrated soil samples to obtain standardized spectral characteristics; The key structural correlation variable identification module is used to determine key structural correlation variables based on multi-source data from multiple calibrated soil samples, and to train a proxy structural variable inference model with key structural correlation variables as output. The surrogate structural variable inference module is used to obtain surrogate structural variables for unknown soil samples to be detected using the surrogate structural variable inference model. The fusion and inversion module is used to fuse and invert the predicted values of target soil properties for unknown soil samples; The results output module is used to display, store, or upload the predicted soil property values.
9. The soil property spectral inversion system based on surrogate structural variables as described in claim 8, characterized in that, The multi-source data acquisition module includes spectral data units, low-cost field environmental variable units, laboratory chemical structure variable units, and measured values of target soil properties units. The spectral data unit is used to collect spectral data for each calibrated soil sample. The spectral data is used to characterize the reflectance spectral response of the soil sample, including visible-near infrared spectral data, near-infrared spectral data, mid-infrared spectral data, hyperspectral image data, or other spectral data that can be used to characterize changes in soil properties. The low-cost field environmental variable unit is used to collect low-cost field environmental variable data for each calibrated soil sample; low-cost field environmental variables include variables obtained through portable sensors, in-situ field detection equipment, rapid detection devices or external data sources, including soil pH, water content, electrical conductivity, temperature, salinity, sampling depth, geographical location, topographic factors or climate factors. The Laboratory Chemical Structure Variable Unit is used to collect laboratory chemical structure variable data for each calibrated soil sample. Laboratory chemical structure variables include variables that reflect the internal chemical composition, mineral composition, nutrient status or structural status of the soil and need to be obtained through laboratory analysis, including soil nutrient variables, metal oxide-related variables, mineral composition variables, particle composition variables, cation exchange-related variables or other soil physicochemical indicators. The target soil property measurement unit is used to measure target soil property data for each calibrated soil sample. Target soil properties include soil properties that need to be predicted by spectral inversion methods, including one or more of soil organic matter, soil organic carbon, soil nutrients, pH, salinity, cation exchange capacity, or other soil physicochemical properties.
10. The soil property spectral inversion system based on surrogate structural variables according to claim 9, characterized in that, The key structural association variable identification module is used to determine key structural association variables based on multi-source data from multiple calibrated soil samples, and to train a surrogate structural variable inference model using key structural association variables as output, including: Based on low-cost field environmental variables, laboratory chemical structure variables, and measured values of target soil properties in calibrated soil samples, variable structure relationships are constructed, and key structurally related variables are identified from them. Specifically, a structural learning method is used to construct structural relationships between variables in the dataset of the calibrated soil samples. The structural learning method can determine one or more key structurally related variables based on the conditional dependencies, stable associations, node connections, edge frequency, or the degree of association with the target soil properties between variables. Alternatively, a Bayesian network can be used to construct the variable structure relationship, and a resampling strategy can be used to screen key structurally related variables that appear stably.