Soil Classification Prediction Method and System Based on Multi-Source Environmental and Remote Sensing Data
By constructing an adaptive weight configuration and closed-loop optimization mechanism, the problem of multi-source data fusion in soil classification is solved, and the accuracy, efficiency and stability are improved, adapting to complex environments and reducing computing resource consumption are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUNNAN DAOSHI GEOGRAPHIC INFORMATION SYSTEM ENGINEERING CO LTD
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional soil classification methods rely on manual sampling, which is inefficient. The fusion of multi-source data from remote sensing technology is difficult, resulting in insufficient classification accuracy. Furthermore, the fixed weight allocation method cannot adapt to dynamic environments, leading to the weakening of key features or the amplification of noise, which reduces model efficiency and accuracy, and makes classification results unstable in complex environments.
By constructing a soil feature dataset containing spectral, physicochemical, and environmental characteristics, the contribution and correlation of each feature are calculated, an adaptive weight configuration is generated, a subset of features with high contribution is selected for weighted fusion, a soil classification decision tree is constructed, and the weights are updated when the confidence threshold is not met, thus forming a closed-loop optimization mechanism.
It achieves a synergistic improvement in the accuracy, efficiency, and stability of soil classification, enhances the model's adaptability to changing scenarios, significantly improves classification accuracy in complex geographical environments, and reduces computational resource consumption.
Smart Images

Figure CN121765653B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer data processing technology, and more specifically, to a method and system for soil classification prediction based on multi-source environmental and remote sensing data. Background Technology
[0002] Traditional soil classification methods, which rely on manual sampling, are inefficient and cannot support large-scale monitoring needs. While existing remote sensing technologies have expanded the observation range, their classification accuracy is insufficient due to the limitations of relying on a single data source.
[0003] Current technical bottlenecks are concentrated in the difficulty of integrating multi-source heterogeneous data: the spectral characteristics of satellite remote sensing, the physicochemical characteristics of ground sensors, and the topographic and climatic characteristics (environmental characteristics) of environmental monitoring are difficult to integrate effectively due to differences in spatiotemporal scale and information density. In particular, the contribution of different features to soil classification changes dynamically with the environment, while traditional fixed-weight allocation methods have the following drawbacks: static weights cannot adapt to dynamic scenarios, resulting in the weakening of key features or the amplification of noise; data redundancy and information loss coexist, reducing model efficiency and accuracy; and the classification results are unstable in complex environments, restricting the reliability of practical applications.
[0004] Therefore, it is urgent to break through the core technical bottlenecks of adaptive weight allocation and dynamic feature optimization of multi-source data, and achieve soil classification prediction with synergistic optimization of accuracy and efficiency. Summary of the Invention
[0005] The purpose of this application is to provide a soil classification prediction method and system based on multi-source environmental and remote sensing data, which can solve at least one of the technical problems mentioned above. The specific solution is as follows:
[0006] According to a specific embodiment of this application, in a first aspect, this application provides a soil classification prediction method based on multi-source environmental and remote sensing data, comprising:
[0007] Construct a soil feature dataset that includes spectral features, physicochemical features, and environmental features;
[0008] The contribution of each feature to soil classification and their correlation are calculated to generate an initial weight configuration. Based on the spectral feature contribution meeting a first threshold condition and / or the correlation between the physicochemical features and the environmental features meeting a second threshold condition, the weight coefficients are updated to generate an adaptive weight configuration. Based on the adaptive weight configuration, a subset of features whose contribution exceeds the contribution threshold is weighted and fused to generate a soil classification feature vector. Based on the soil classification feature vector, a soil classification decision tree is constructed and the optimal classification path is determined. Based on the optimal classification path, the classification result of the soil sample to be classified is predicted. It is then determined whether the confidence level of the classification result meets the confidence threshold condition. If not, the weight coefficients are updated, and the process returns to the step of generating the adaptive weight configuration until the confidence threshold condition is met, resulting in the final classification result.
[0009] According to a specific embodiment of this application, in a second aspect, this application provides a soil classification prediction system based on environmental and remote sensing multi-source data, comprising:
[0010] The system comprises: a construction unit for constructing a soil feature dataset containing spectral features, physicochemical features, and environmental features; a processing unit for calculating the contribution of each feature to soil classification and the correlation between them, generating an initial weight configuration, updating the weight coefficients based on the spectral feature contribution meeting a first threshold condition, and / or the correlation between the physicochemical features and the environmental features meeting a second threshold condition, and generating an adaptive weight configuration; a weighted fusion of a subset of features whose contribution exceeds the contribution threshold based on the adaptive weight configuration, generating a soil classification feature vector; and a soil classification decision tree constructed based on the soil classification feature vector and determining the optimal classification path; and a prediction unit for predicting the classification result of the soil sample to be classified based on the optimal classification path, determining whether the confidence level of the classification result meets the confidence level threshold condition, updating the weight coefficients if not, and returning to the step of generating the adaptive weight configuration, until the confidence level threshold condition is met, and obtaining the final classification result.
[0011] Compared with the prior art, the above-described solution of this application has at least the following beneficial effects: This application provides a soil classification prediction method based on environmental and remote sensing multi-source data. By constructing a multi-source data adaptive fusion and closed-loop optimization mechanism, it achieves a synergistic improvement in accuracy, efficiency and stability in soil classification prediction.
[0012] Specifically, a unified dataset of spectral, physicochemical, and environmental features is constructed to achieve multi-dimensional information collaboration between satellite remote sensing, ground sensors, and environmental monitoring equipment, providing technical support for large-scale precision soil monitoring. By selecting feature subsets with contributions exceeding a threshold for fusion, noise interference is effectively suppressed, reducing data redundancy while retaining key classification information, thus addressing the problem of low information utilization efficiency. Combining the optimal path selection of the soil classification decision tree with a confidence threshold verification mechanism, a weight re-optimization loop is triggered for classification results that do not meet the confidence threshold conditions, significantly improving the classification reliability of marginal samples and complex scenarios. Based on the weight configuration of feature contribution and correlation, the weight configuration of spectral, physicochemical, and environmental features can be dynamically adjusted according to soil type and environmental conditions, enhancing the model's adaptability to changing scenarios and significantly improving generalization performance. Breaking through the limitations of traditional fixed-weight models, adaptive weight configuration and confidence feedback iteration significantly improve classification accuracy in complex geographical environments while greatly reducing computational resource consumption, effectively solving the problem of balancing accuracy and efficiency. Attached Figure Description
[0013] Figure 1 A flowchart of a soil classification prediction method based on multi-source environmental and remote sensing data is shown;
[0014] Figure 2 A flowchart illustrating a method for generating optimal weight configurations is shown.
[0015] Figure 3 A flowchart illustrating a method for building a dynamically scalable adaptive weight configuration library is shown.
[0016] Figure 4 A flowchart of a closed-loop feedback mechanism with optimal weight configuration is shown.
[0017] Figure 5 A flowchart illustrating a method for generating adaptive weight configurations is shown.
[0018] Figure 6 A flowchart illustrating a method for constructing a soil classification decision tree and determining the optimal classification path is shown.
[0019] Figure 7 A flowchart illustrating a method for obtaining the final classification result is shown.
[0020] Figure 8 A flowchart illustrating a method for generating initial weight configurations is shown.
[0021] Figure 9 A flowchart illustrating a method for constructing a soil feature dataset is shown.
[0022] Figure 10A block diagram of a soil classification prediction system based on environmental and remote sensing multi-source data according to an embodiment of this application is shown. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms "a," "the," and "the" as used in the embodiments of this application are also intended to include the plural forms, and "multiple" generally includes at least two unless the context clearly indicates otherwise.
[0025] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0026] It should be understood that although the terms first, second, third, etc., may be used in the embodiments of this application, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, first may also be referred to as second without departing from the scope of the embodiments of this application, and similarly, second may also be referred to as first.
[0027] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”
[0028] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system that includes said element.
[0029] It should be noted that any symbols and / or numbers present in the specification that are not marked in the accompanying drawings are not reference numerals.
[0030] The optional embodiments of this application are described in detail below with reference to the accompanying drawings.
[0031] The embodiments provided in this application are embodiments of a soil classification prediction method based on environmental and remote sensing multi-source data.
[0032] The following is combined Figure 1 The embodiments of this application will be described in detail.
[0033] Figure 1 A flowchart of a soil classification prediction method based on multi-source environmental and remote sensing data is shown, such as... Figure 1 As shown, it includes the following steps.
[0034] Step S101: Construct a soil feature dataset that includes spectral features, physicochemical features, and environmental features.
[0035] Step S102: Calculate the contribution of each feature to soil classification and the correlation between them, generate an initial weight configuration, update the weight coefficients based on the first threshold condition of the contribution of spectral features and / or the second threshold condition of the correlation between physicochemical features and environmental features, and generate an adaptive weight configuration.
[0036] Step S103: Based on adaptive weight configuration, the selected feature subsets with contribution exceeding the contribution threshold are weighted and fused to generate soil classification feature vectors.
[0037] Step S104: Based on the soil classification feature vector, construct a soil classification decision tree and determine the optimal classification path.
[0038] Step S105: Based on the optimal classification path, predict the classification result of the soil sample to be classified, determine whether the confidence of the classification result meets the confidence threshold condition, if not, update the weight coefficients, and return to the step of generating adaptive weight configuration until the confidence threshold condition is met, and obtain the final classification result.
[0039] Among these, spectral characteristics refer to soil electromagnetic wave reflectance data acquired through satellite remote sensing, including reflectance information from the visible to infrared bands, such as red, green, blue, and near-infrared values, used to characterize the physical state of the soil surface. Physicochemical characteristics refer to soil physicochemical property data collected through ground sensors, including but not limited to soil moisture, temperature, nutrient content, and pH value, used to quantify the internal composition and properties of the soil. Environmental characteristics refer to climate and topographic data acquired through environmental monitoring equipment, including rainfall, wind speed, elevation, and slope, used to describe the external environmental conditions of the soil.
[0040] Among them, contribution refers to the degree of influence of a single feature on the soil classification result, which is quantified using the information entropy method. The higher the value, the more significant the distinguishing effect of the feature on the classification decision. Correlation refers to the strength of the statistical correlation between different features, calculated using the Pearson correlation coefficient, and is used to measure the synergistic or conflicting relationships between features. The first threshold condition is the baseline value of the contribution of spectral features, and the second threshold condition is the baseline value of the correlation between physicochemical features and environmental features, used to characterize the triggering of weight coefficient updates. Adaptive weight configuration refers to the real-time adjustment of feature weight combinations through iterative optimization based on the dynamic changes in contribution and correlation, making the weight allocation adaptable to the current environmental scenario. Iterative optimization methods include simulated annealing algorithms, etc.
[0041] The feature subset refers to the set of features selected from the original feature data. The selection criterion is that the contribution exceeds a preset contribution threshold, used to remove redundant information. The soil classification feature vector is a mathematical vector generated by adaptively weighting and fusing the selected feature subsets, serving as the input data structure for the classification model. The optimal classification path refers to the branch path selected first in the classification decision tree based on the proportion of feature contribution, determined through recursive segmentation and accuracy verification to establish the final classification logic chain. The confidence threshold condition is a quantitative evaluation standard for the reliability of the classification results. When the predicted classification result does not meet the confidence threshold condition, an iterative optimization loop of the weight coefficients is triggered.
[0042] In some embodiments, spectral features, including red, green, blue and near-infrared band values, are acquired through satellite remote sensing; physical and chemical features, including soil moisture, temperature and nutrient content, are collected through ground sensors; and environmental features, including rainfall, wind speed and topographic elevation, are acquired through environmental monitoring equipment.
[0043] As a feasible implementation, a spatiotemporal registration algorithm is used to process heterogeneous data. For time synchronization, if the timestamp deviation between multispectral band data, physicochemical features, and environmental features exceeds a preset threshold, the timestamp is adjusted using an interpolation algorithm. For spatial alignment, projection transformation is used to match geographic coordinates between multispectral band data and topographic elevation data; if the coordinate deviation exceeds a preset threshold, the spatial position is adjusted through raster resampling. For data fusion, a weighted average method is used to generate a multidimensional soil feature dataset from the spatially aligned spectral features and soil moisture data.
[0044] In some embodiments, the information entropy method is used to calculate the contribution of each feature to soil classification, and the Pearson correlation coefficient method is used to calculate the correlation between features to generate an initial weight configuration.
[0045] As a feasible implementation, the weights are dynamically updated based on threshold conditions. If the contribution of spectral features meets the first threshold condition, its weight coefficient is increased through simulated annealing algorithm; if the correlation between physicochemical features and environmental features meets the second threshold condition, its weight coefficient is decreased through linear interpolation.
[0046] As a feasible implementation, a subset of features whose contribution exceeds a contribution threshold is selected based on adaptive weight configuration. For example, the importance of features is evaluated using a random forest algorithm, and the feature subsets are weighted and fused to generate a soil classification feature vector. For instance, spectral reflectance features and physicochemical features are standardized, and a weighted sum is calculated according to the weights to generate a feature vector.
[0047] In some embodiments, an adaptive weighting matrix is used to perform weighted fusion processing on the multidimensional soil feature dataset, and the subset of spectral features and physicochemical parameters that contribute the most to soil classification is selected through feature importance evaluation. Specifically, spectral reflectance data and physicochemical parameter data are obtained from the multidimensional soil dataset; the feature importance of the spectral reflectance data and physicochemical parameter data is calculated using the information gain method to generate a preliminary importance ranking of each feature; when the feature importance of the spectral reflectance data is higher than a preset threshold, its weight coefficients are iteratively updated using a weighted adjustment method to generate an adjusted weight coefficient matrix; based on the weight coefficient matrix, a dimensionality reduction processing method is used to perform feature selection on the spectral reflectance data and physicochemical parameter data to obtain a simplified feature dataset; the simplified feature dataset is converted into a classification feature vector using a vector mapping method, and the optimized soil classification feature vector is output.
[0048] As a feasible implementation, a random forest algorithm is used to calculate Gini importance to assign feature importance to each feature. For example, the input features include: spectral reflectance feature A1 (450-550nm), spectral reflectance feature A2 (650-750nm), and physicochemical parameters B1 (organic matter content) and B2 (pH value). Using the random forest model with 100 trees and a maximum depth of 10, the feature importance results are: A1 = 0.42, A2 = 0.35, B1 = 0.18, and B2 = 0.05. Assuming an importance threshold of 0.20, a high-contribution feature subset {A1, A2, B1} is selected, and B2 is removed.
[0049] As a specific implementation, the weighted fusion includes: assuming the initial weights are all equal at 0.333, adjusting the weights based on feature importance, as shown in the following formula:
[0050] (Formula 1)
[0051] in, Indicates the initial weights. This represents the sum of the initial weights. express The importance score of the i-th feature This represents the weights adjusted based on feature importance.
[0052] The above formula adjusts the weights: A1 weight is adjusted to 0.45, A2 weight is adjusted to 0.38, and B1 weight is adjusted to 0.17.
[0053] Based on Formula 1 above, the weights are adjusted as follows: A1 weight is adjusted to 0.45, A2 weight to 0.38, and B1 weight to 0.17. The standardized eigenvalues are then weighted and summed, resulting in a mean of 0 and a variance of 1: 1.2 × 0.45 + 0.8 × 0.38 + 0.5 × 0.17 = 0.929.
[0054] The classification performance of the fused vectors was evaluated using K-fold cross-validation (K=5). Assuming the initial classification performance score of the fused vectors was 0.85, gradient descent was used with a learning rate of 0.01 and 100 iterations to fine-tune the weights. After adjustment, the weights of A1 were 0.47, A2 were 0.36, and B1 were 0.16, improving the classification performance score of the fused vectors to 0.91.
[0055] In some embodiments, a multi-layer classification decision tree is constructed based on soil classification feature vectors. If the contribution of spectral features is higher than that of physicochemical features, a spectral classification branch is constructed first; if the contribution of physicochemical features is higher than that of spectral features, a physicochemical classification branch is constructed first.
[0056] As a specific implementation, decision nodes are generated through recursive segmentation, and the optimal classification path is determined by cross-validation.
[0057] In some embodiments, the classification results of soil samples are predicted based on the optimal classification path, and it is determined whether the confidence level meets the confidence threshold condition. If not, the weight coefficients are updated using a simulated annealing algorithm, and an adaptive weight configuration is regenerated. If the confidence level meets the threshold condition, the final classification result is output. For example, when the classification confidence level is lower than the threshold, the weight coefficients of spectral features and physicochemical features are adjusted until the confidence level of the classification result meets the confidence threshold condition.
[0058] The aforementioned soil classification and prediction method based on multi-source environmental and remote sensing data constructs a unified dataset of spectral, physicochemical, and environmental features, enabling multi-dimensional information collaboration among satellite remote sensing, ground sensors, and environmental monitoring equipment, providing technical support for large-scale precise soil monitoring. By selecting and fusing feature subsets with contributions exceeding a contribution threshold, noise interference is effectively suppressed, reducing data redundancy while retaining key classification information, thus addressing the problem of low information utilization efficiency. Combining the optimal path selection of the soil classification decision tree with a confidence threshold verification mechanism, a weight re-optimization loop is triggered for classification results that do not meet the confidence threshold condition, significantly improving the classification reliability of marginal samples and complex scenarios. Based on the weight configuration of feature contribution and correlation, the weight configuration of spectral, physicochemical, and environmental features can be dynamically adjusted according to soil type and environmental conditions, enhancing the model's adaptability to changing scenarios and significantly improving generalization performance. Breaking through the limitations of traditional fixed-weight models, through adaptive weight configuration and confidence feedback iteration, the classification accuracy in complex geographical environments is significantly improved, while greatly reducing computational resource consumption, effectively solving the problem of balancing accuracy and efficiency. By constructing a multi-source data adaptive fusion and closed-loop optimization mechanism, a synergistic improvement in accuracy, efficiency, and stability was achieved in soil classification prediction.
[0059] Figure 2 A flowchart illustrating the method for generating the optimal weight configuration is shown, as follows: Figure 2 As shown, it includes the following steps.
[0060] Step S201: Extract misclassified samples from the final classification results whose predicted labels do not match the true labels.
[0061] Step S202: In response to the feature distribution deviation value being greater than the deviation threshold, the dynamic range boundary of the weight coefficients is adjusted using a gradient optimization algorithm.
[0062] Step S203: Iteratively optimize the adaptive weight configuration based on the adjusted range boundary, generate the optimal weight configuration, and store it.
[0063] As a feasible embodiment, misclassified samples whose predicted labels do not match the true labels are extracted from the final classification results, and their original data distribution of spectral and physicochemical characteristics is obtained. For these misclassified samples, a clustering analysis algorithm is used to hierarchically classify the feature distribution, calculate the mean and variance of each feature subset, and generate a quantified feature distribution bias value. The clustering analysis algorithm includes the K-means clustering algorithm.
[0064] If the feature distribution deviation value exceeds a preset deviation threshold, a gradient optimization algorithm is used to dynamically adjust the range boundary of the weight coefficients. This gradient optimization algorithm includes gradient descent. Specifically, the dominant error feature dimension is identified based on the deviation analysis results. For example, if the variance of spectral features is significantly higher than that of physicochemical features, the spectral dimension is determined to be the key error source. The upper limit of the dynamic change of the key feature weights is expanded based on the importance of the error source, while the lower limit of the change of the secondary feature weights is compressed. The weight adjustment magnitude is calculated through iterative optimization to generate an updated weight distribution combination.
[0065] As a specific implementation, the updated weight distribution combination is input into a classification model for validation, where the classification model includes a support vector machine. Specifically, if the classification accuracy on the validation set improves, it is confirmed as the optimal weight configuration; soil environmental parameters, such as humidity range, are associated to construct an environmental adaptation strategy; and the optimized weight configuration is stored in a database, forming a closed-loop update mechanism.
[0066] For example, when spectral features cause classification bias due to changes in environmental humidity, their dynamic weight range can be expanded, for instance, from 0.5-0.8 to 0.6-0.85, and associated with a humidity threshold, such as 20%, to generate a weight configuration pattern for a specific environment. For example, the optimized configuration can be associated with soil environmental parameters to form a business adaptation strategy.
[0067] The above method extracts the feature distribution deviation value of misclassified samples, uses a gradient optimization algorithm to dynamically adjust the range boundary of the weight coefficients when the deviation exceeds the threshold, and iteratively optimizes to generate the optimal weight configuration, forming a closed-loop feedback mechanism. This significantly improves the environmental adaptability and classification accuracy of adaptive weight configuration. At the same time, it achieves rapid adaptation to multiple scenarios through storage optimization configuration, effectively solving the problem of insufficient generalization ability of traditional fixed weight models in complex environments.
[0068] Figure 3 A flowchart illustrating the method for building a dynamically scalable adaptive weight configuration library is shown, such as... Figure 3 As shown, it includes the following steps.
[0069] Step S301: Based on the optimal weight configuration, construct an adaptive weight configuration library for different environmental conditions.
[0070] Step S302: For the new soil sample, extract the feature distribution of the new soil sample, calculate its cosine similarity with the feature distribution of each weight configuration in the adaptive weight configuration library, and determine the maximum cosine similarity.
[0071] Step S303: Determine whether the maximum cosine similarity is less than the matching threshold.
[0072] In step S304a, in response to the maximum cosine similarity being less than the matching threshold, a new weight configuration is generated based on the environmental conditions of the new soil sample, the classification is predicted based on the new weight configuration, and the new weight configuration is synchronized to the adaptive weight configuration library.
[0073] Step S304b: In response to the maximum cosine similarity being no less than the matching threshold, predict the classification based on the weight configuration corresponding to the maximum cosine similarity.
[0074] In some embodiments, an adaptive weight configuration library for different environmental conditions is constructed, storing the optimal weight configurations and their associated spectral reflectance features and physicochemical parameter features under different geographical locations and seasons. For soil samples to be classified, their spectral reflectance features and physicochemical parameter features are extracted to generate feature distribution vectors; the cosine similarity between this vector and the feature distributions of each configuration in the adaptive weight configuration library is calculated; it is determined whether the maximum cosine similarity is less than the matching threshold. If not, classification is predicted based on the weight configuration corresponding to the maximum cosine similarity; if so, a new weight configuration is dynamically generated according to the sample environmental conditions and synchronously updated to the configuration library.
[0075] As a specific implementation, a configuration library is constructed. For spring samples from a certain region, spectral reflectance parameters are stored: mean 0.45 in the 0.8-1.2μm band; physicochemical parameters: organic matter content 2.5-3.5%, pH value 6.0-7.0; and corresponding optimal weight configurations are established, with spectral feature weight at 0.55 and physicochemical parameter weight at 0.45. For similarity matching, features of the new sample are extracted: mean 0.48 in the 0.8-1.2μm band, organic matter 3.0%, pH value 6.5. The cosine similarity with each configuration in the library is calculated, for example, results of 0.92, 0.85, and 0.78, for similarity matching. For classification prediction, the configuration with the highest similarity (0.92 in the above example) is matched, and its weight is adopted: spectral feature weight at 0.55 and physicochemical parameter weight at 0.45. The weighted features are input into a random forest classification model with 100 trees and a maximum depth of 10, outputting soil type T2.
[0076] For example, the prediction results can be associated with the sample's geographical location and seasonal data to achieve closed-loop updates and business linkage; the feature range of the configuration library can be dynamically updated, such as expanding the average spectral reflectance to 0.45-0.48; and planting suggestions can be generated in conjunction with soil management business to achieve full-process automation.
[0077] By constructing a dynamically scalable adaptive weight configuration library and combining cosine similarity matching and threshold determination mechanisms, rapid and accurate prediction of soil samples to be classified can be achieved. At the same time, a closed-loop update strategy is used to automatically expand the coverage of the configuration library, significantly improving the classification system's adaptability to changes in environmental conditions and providing automated decision support for soil management operations.
[0078] Figure 4 The flowchart of the closed-loop feedback mechanism with optimal weight configuration is shown, as follows: Figure 4 As shown, it includes the following steps.
[0079] Step S401: Calculate the adjustment range of the weight coefficients based on the adjusted range boundaries.
[0080] Step S402: The gradient descent method is used to iteratively optimize the adjustment range to obtain the updated weight configuration.
[0081] Step S403: Map and match the updated weight configuration with the feature distribution of the misclassified samples.
[0082] Step S404: Determine whether the mapping match meets the goal of continuous optimization.
[0083] Step S405a: In response to the mapping matching conforming to the goal of continuous optimization, the optimal weight configuration is generated and stored.
[0084] In step S405b, in response to the mapping match not meeting the goal of continuous optimization, the range boundary is readjusted, and the step of calculating the adjustment magnitude of the weight coefficients based on the adjusted range boundary is returned until the mapping match meets the goal of continuous optimization.
[0085] In some embodiments, an adaptive weight configuration is optimized using a feedback mechanism based on the final soil classification results and the feature distribution of misclassified samples. This includes the following steps: extracting spectral reflectance data and physicochemical parameter data of misclassified samples from the soil classification results; performing hierarchical classification of the feature distribution of misclassified samples to generate a distribution deviation quantification value; adjusting the range boundary of the weight coefficients in response to the distribution deviation quantification value exceeding a preset threshold; calculating the adjustment magnitude of the weight coefficients based on the adjusted range boundary; iteratively optimizing the adjustment magnitude using gradient descent to generate an updated weight configuration; mapping and matching the updated weight configuration with the spectral reflectance features and physicochemical parameter feature distributions; determining whether the mapping and matching meet the continuous optimization objective: if it does, generating and storing the optimal weight configuration; if it does not, readjusting the range boundary and returning the calculated adjustment magnitude of the weight coefficients based on the adjusted range boundary.
[0086] As a specific implementation, the characteristic distribution of misclassified samples is analyzed. Assume that the misclassified samples of soil type T1 have a mean spectral reflectance of 0.35 in the 1.5-2.0 μm band, a pH range of 5.5-6.5, and a soil particle size range of 0.02-0.05 mm. A K-means clustering algorithm with 3 cluster centers is used to divide the feature distribution into subsets. The variance of the spectral features is calculated to be 0.12, and the variance of the physicochemical features is 0.08, thus determining spectral reflectance as the main error feature dimension. An adaptive weight configuration is initialized with a weight of 0.6 for spectral features and 0.4 for physicochemical features. Gradient descent with a learning rate of 0.01 and 500 iterations is used to adjust the weights, increasing the weight of the spectral features to 0.68 and decreasing the weight of the physicochemical features to 0.32.
[0087] For example, when verifying the optimization effect, a support vector machine was used to test the validation set. The kernel function of the support vector machine is a radial basis function, and the penalty parameter C is 1.0. The classification accuracy improved from 0.82 to 0.87. The soil environmental adaptability analysis business was linked, and the weight configuration was calibrated in combination with soil moisture data. The soil moisture range is 10-30%. The database was updated through automated scripts to form a closed-loop optimization mechanism.
[0088] By dynamically adjusting the range of weight coefficients and combining gradient descent iterative optimization with mapping matching verification, a closed-loop feedback mechanism is formed, which significantly improves the accuracy and adaptability of the soil classification model. At the same time, the continuous optimization of weight configuration is achieved through automated processes, which effectively enhances the model's robustness to different environmental conditions and reduces the cost of manual intervention.
[0089] Figure 5 A flowchart illustrating the method for generating adaptive weight configurations is shown, such as... Figure 5 As shown, it includes the following steps.
[0090] Step S501: Compare the contribution of spectral features with a first threshold, and compare the correlation between physicochemical features and environmental features with a second threshold.
[0091] In step S502a, in response to the spectral feature contribution exceeding the first threshold, it is confirmed that the spectral feature contribution meets the first threshold condition, and the weight coefficient of the spectral feature is increased.
[0092] In step S502b, in response to the fact that the correlation between the physicochemical features and the environmental features is lower than the second threshold, it is confirmed that the correlation between the physicochemical features and the environmental features meets the second threshold condition, and the weight coefficient of the physicochemical features is reduced.
[0093] Step S503: Generate an adaptive weight configuration by increasing the weight coefficient of spectral features and / or decreasing the weight coefficient of physicochemical features.
[0094] In one specific embodiment, a first threshold is set to 0.45, and a second threshold is set to 0.35. When the spectral feature contribution is detected to be 0.5, which is greater than the first threshold, a simulated annealing algorithm is initiated for iterative optimization. Specifically, the initial temperature is set to 1000, the cooling rate is 0.95, and a total of 500 iterations are performed. In each iteration, the weight value is randomly perturbed, for example, the weight is adjusted from 0.52 to 0.54, and the objective function value, i.e., the feature discrimination, is calculated. If the objective function value increases from 0.78 to 0.82, the adjustment is accepted. After optimization loops, the spectral feature weight is finally increased to 0.58. When the correlation coefficient between physicochemical features and environmental features is detected to be 0.28, which is lower than the second threshold, a linear decay mechanism is used to reduce the weight. Specifically, the new weight is calculated using the following formula:
[0095] (Formula 2)
[0096] in, To calculate the new weights, The original weights, For the attenuation factor, specifically, It can be set to 0.1.
[0097] For example, the weight of physicochemical characteristics is reduced from 0.36 to 0.36×(1-0.1)=0.324, and the weight of environmental characteristics is simultaneously reduced from 0.29 to 0.29×(1-0.1)=0.261 to maintain the balance between characteristics.
[0098] Optionally, after adjusting the weights of each feature, the deviation of the total weights is checked. If the deviation exceeds the deviation threshold, a proportional correction is performed, calculated using the following formula:
[0099] (Formula 3)
[0100] in, The values after the new weights are calibrated. This is the sum of the new weights.
[0101] For example, if the deviation threshold is set to 0.01, and the total deviation of the weights is 0.58 + 0.324 + 0.261 = 1.165 > 1.01, then each weight is multiplied by the correction factor 1 / 1.165 ≈ 0.858 to obtain the final calibration matrix.
[0102] Optionally, to ensure the effectiveness of the configuration, soil moisture is introduced as an auxiliary validation factor. For example, if the auxiliary validation factor is set to 0.75, and the model discrimination of the adjusted weight matrix is lower than 0.75, the system automatically backtracks to the previous effective state to recalculate the optimization path. Simultaneously, the system fully records the weight changes and evaluation metrics for each iteration, forming a traceable optimization chain. This process is fully automated through a loop algorithm.
[0103] The above-mentioned dual-threshold judgment mechanism dynamically adjusts the feature weights, increasing the weights when the contribution of spectral features is high and decreasing the weights when the correlation between physicochemical and environmental features is weak. Combined with simulated annealing optimization and linear decay control, the weight configuration is significantly improved to adapt to the distribution of soil features. The closed-loop verification mechanism ensures the stability of the classification model accuracy, while the automated process reduces the cost of manual intervention.
[0104] Figure 6 A flowchart illustrating the method for constructing a soil classification decision tree and determining the optimal classification path is shown, as follows: Figure 6 As shown, it includes the following steps.
[0105] Step S601: Based on the soil classification feature vector, calculate the contribution ratio of spectral features and the contribution ratio of physicochemical features, and determine the group with the relatively higher contribution ratio between the two.
[0106] In step S602a, in response to the fact that the contribution ratio of spectral features is higher than that of physicochemical features, a spectral classification branch is preferentially constructed as the first layer of the decision tree, and child nodes are recursively generated under the spectral classification branch to form a soil classification decision tree.
[0107] In step S602b, in response to the fact that the contribution ratio of physicochemical features is higher than that of spectral features, a physicochemical classification branch is constructed first as the first layer of the decision tree, and child nodes are recursively generated under the physicochemical classification branch to form a soil classification decision tree.
[0108] Step S603: For all branches of the soil classification decision tree, cross-validation is used to evaluate the classification accuracy, and the branch with the highest classification accuracy is taken as the optimal classification path.
[0109] For example, the contribution quantification analysis is as follows: assuming the total contribution of spectral feature group C1 is 0.65 and the total contribution of physicochemical feature group C2 is 0.35. Since C1 has a higher proportion than C2, the spectral classification branch is preferentially selected as the first layer.
[0110] As a specific implementation, the decision tree construction process is as follows: Under the spectral classification branch, a support vector machine is used with a radial basis function kernel and a penalty parameter C of 1.0 to divide spectral features. A hyperspectral response group and a low-spectral response group are divided with a feature value of 0.9 as the split point. Feature values greater than 0.9 correspond to the hyperspectral response group, and feature values not greater than 0.9 correspond to the low-spectral response group. Logistic regression is applied to the hyperspectral response group for 50 iterations. Physicochemical features are analyzed. A soil moisture threshold of 0.3 is used; soils with a classification probability greater than 0.6 are classified as soil type A, otherwise as type B. Under the physicochemical classification branch, a decision tree algorithm is used with a maximum depth of 8, splitting by pH value, for example, using pH 6.5 as the split point to divide into acidity / alkalinity groups. Path accuracy is evaluated through 10-fold cross-validation. The accuracy of the spectral branch is 0.88, and the accuracy of the physicochemical branch is 0.82. Therefore, the spectral branch is selected as the optimal classification path, and a classification decision logic chain is generated and stored in the database.
[0111] The above method significantly improves classification efficiency and accuracy by dynamically selecting feature types with higher contribution as the first-level branches of the decision tree. Prioritizing key features can reduce invalid computation. After constructing a complete decision tree by recursively generating child nodes, the optimal classification path is then selected through cross-validation, ensuring that the soil classification results have both high accuracy and optimized computational resources.
[0112] Figure 7 The flowchart illustrating the method for obtaining the final classification result is shown, as follows: Figure 7 As shown, it includes the following steps.
[0113] Step S701: Calculate the classification confidence of the classification result and compare it with the confidence threshold.
[0114] In step S702, in response to the classification confidence being lower than the confidence threshold, the simulated annealing algorithm is used to iteratively optimize the weight coefficients of the spectral features and physicochemical features to update the adaptive weight configuration.
[0115] Step S703: Based on the updated adaptive weight configuration, regenerate the soil classification results and return to the step of calculating the classification confidence of the classification results until the classification confidence is not lower than the confidence threshold, and obtain the final classification result.
[0116] As a feasible implementation, spectral and physicochemical data are obtained from a soil sample dataset. Principal component analysis is used to reduce the dimensionality of the data, resulting in spectral and physicochemical feature vectors. Based on these vectors, the information gain method is used to calculate the feature contribution, yielding quantified values for spectral and physicochemical feature contributions. If the classification confidence level is lower than a confidence threshold, simulated annealing is used to iteratively optimize the weight coefficients of the quantified spectral and physicochemical feature contributions, resulting in updated weight coefficients. Using the updated weight coefficients and a preset classification path, the soil samples are predicted and classified to obtain the final classification result and classification confidence level.
[0117] In one specific embodiment, the soil sample data to be classified includes spectral feature groups and physicochemical feature groups. Features are automatically extracted and their contribution values are calculated. For example, the spectral feature group includes near-infrared reflectance (range 0.8-1.2 μm), and the physicochemical feature group includes organic matter content (range 0-5%) and cation exchange capacity (range 5-30 cmol / kg). Principal component analysis shows that the contribution value of spectral features is 0.55, and the contribution value of physicochemical features is 0.45. The system uses a random forest classifier for initial classification of the samples, with 100 trees and a maximum depth of 10, predicting the soil type as S1, S2, or S3, and outputting the classification confidence score.
[0118] For example, setting the confidence threshold for classification results to 0.75, if sample A is predicted as soil type S1 with a classification confidence of 0.72, which is lower than the confidence threshold of 0.75, a simulated annealing algorithm is initiated to optimize the feature weight coefficients. The initial weight coefficients are 0.5 for spectral features and 0.5 for physicochemical features, with an initial temperature of 100°C, a cooling rate of 0.95, and 1000 iterations. In each iteration, the weight coefficients are adjusted, and a new classification result is calculated, with the objective function being classification accuracy. After optimization, the weight coefficients are adjusted to 0.62 for spectral features and 0.38 for physicochemical features. The random forest classifier is then rerun, resulting in a classification confidence of sample A that is now 0.78, higher than the confidence threshold, confirming it as soil type S1. The classification results are stored in a database, recording the feature values, feature contributions, optimized weight coefficients, and the final classification result. If the initial classification confidence of a sample is higher than the confidence threshold, such as 0.80, the result is stored directly. The system logs the changes in weight coefficients and accuracy for each iteration to ensure that the classification logic is traceable.
[0119] The above approach significantly improves the accuracy and reliability of soil classification by dynamically optimizing weight coefficients and iteratively verifying confidence levels, ensuring that classification results meet preset confidence standards, while reducing human intervention and enhancing the model's adaptability and the credibility of the results.
[0120] Figure 8 A flowchart illustrating the method for generating initial weight configurations is shown, as follows: Figure 8 As shown, it includes the following steps.
[0121] Step S801: Use the information entropy method to calculate the contribution of spectral features, physicochemical features and environmental features to soil classification.
[0122] Step S802: Use Pearson correlation coefficient to analyze the correlation between spectral characteristics, physicochemical characteristics and environmental characteristics.
[0123] Step S803: Based on the contribution of each feature and the correlation between each feature, an initial weight configuration is generated using a weighted average method.
[0124] In some embodiments, generating the initial weight configuration includes the following steps:
[0125] The features are standardized to obtain spectral features, physicochemical features and environmental features. Specifically, the spectral features are the reflectance values in the 400-2500nm band; the physicochemical features are the pH value and organic matter content; and the environmental features are the average annual rainfall and slope.
[0126] The eigenvalues are normalized using the following formula:
[0127] (Formula 4)
[0128] in, These are the original eigenvalues. The minimum value in the dataset. The maximum value in the dataset. These are the normalized eigenvalues.
[0129] For example, for a pH value of 5.5-8.5, the result obtained by normalizing using Formula 4 above is 0.2-0.8.
[0130] The information entropy of each feature class is calculated using the following formula:
[0131] (Formula 5)
[0132] in, Indicates the first Class feature entropy value, For the eigenvalue at the th The probability within the interval, The total number of samples.
[0133] The contribution of each feature to soil classification is calculated using the following formula:
[0134] (Formula 6)
[0135] in, This represents the contribution of each feature to soil classification, satisfying the following conditions: .
[0136] The Pearson correlation coefficient r between features is calculated using the following formula:
[0137] (Formula 7)
[0138] in, Indicates the first i Features of a sample x The value, Indicates the first i Features of a sample y The value, Indicating all samples x The mean, Indicating all samples y The mean, n This represents the total number of samples.
[0139] The correlation between features is analyzed using the Pearson correlation coefficient. The formula for calculating the Pearson correlation coefficient is existing technology and will not be elaborated here. If the Pearson correlation coefficient is greater than a preset threshold, for example, r>0.5, a weighted average is performed based on the contribution weights.
[0140] Optionally, inter-data calibration can be performed, and the environmental feature data can be spatially registered using a raster resampling method: the terrain elevation data and spectral reflectance data can be unified to the same coordinate system to ensure the consistency of the feature space.
[0141] As a specific embodiment, the following implementation details are included, with data example: 100 soil samples, including:
[0142] Spectral characteristics: 10 band values in the 400-2500nm wavelength range (e.g., 0.35, 0.42); Physicochemical characteristics: 5 indicators including pH value (5.5-8.5) and organic matter content (1%-5%); Environmental characteristics: 4 indicators including average annual rainfall (500-1500mm) and slope (0-30°).
[0143] Contribution calculation: For spectral feature entropy values ranging from 2.0 to 2.8, the contribution is calculated to be 0.4 based on Formula 6 above; for physicochemical feature entropy values ranging from 1.8 to 2.5, the contribution is calculated to be 0.35 based on Formula 6 above; and for environmental feature entropy values ranging from 2.2 to 2.9, the contribution is calculated to be 0.25 based on Formula 6 above.
[0144] Correlation analysis: Based on Formula 7, the Pearson correlation coefficient between features was calculated, and this coefficient was used as the correlation degree between features. Specifically, the correlation degree between spectral features and physicochemical features was 0.75, the correlation degree between spectral features and environmental features was 0.45, and the correlation degree between physicochemical features and environmental features was 0.3. A 3×3 initial weighting matrix was generated by weighting the contribution and correlation degrees. This matrix comprehensively reflects the contribution and interrelation degree of each feature, providing a weighting basis for subsequent soil classification.
[0145] The above method objectively quantifies the contribution of features through information entropy and accurately characterizes the correlation between features by combining the Pearson correlation coefficient, thereby generating a scientific and reasonable initial weight configuration. This configuration matrix effectively integrates the independent importance of features and the strength of cross-feature correlation, providing a multi-dimensional weighting basis for the soil classification model and significantly improving classification accuracy and feature utilization efficiency.
[0146] Figure 9 A flowchart illustrating the method for constructing a soil feature dataset is shown, such as... Figure 9 As shown, it includes the following steps.
[0147] Step S901: Obtain multi-source data characterizing soil properties.
[0148] Step S902: The multi-source data is synchronized in time and aligned in space by a spatiotemporal registration algorithm to generate a soil feature dataset containing spectral features, physicochemical features and environmental features.
[0149] The multi-source data includes satellite remote sensing spectral data, physicochemical parameter data collected by ground sensors, and climate and topographic data collected by environmental monitoring equipment.
[0150] As a feasible implementation, spectral features are acquired through satellite remote sensing, including red, green, blue, and near-infrared band values; physicochemical features are collected through ground sensors, including soil moisture, temperature, and nutrient content; and environmental features are acquired through environmental monitoring equipment, including rainfall, wind speed, and topographic elevation.
[0151] As a feasible implementation, a spatiotemporal registration algorithm is used to process multi-source data. For time synchronization, if the timestamp deviations of spectral features, physicochemical features, and environmental features exceed a preset threshold, an interpolation algorithm is used to align the time axis; for example, time linear interpolation can be used to eliminate time errors between satellite data and meteorological data. For spatial alignment, for spectral features, such as multispectral band data, and environmental features, such as terrain elevation data, a projection transformation is used to match geographic coordinates; if the coordinate deviation exceeds a preset threshold, raster resampling is used to adjust the spatial position; for example, Kriging interpolation is used to spatially interpolate ground sensor data.
[0152] As a feasible implementation, spatially aligned data is fused. A weighted average method is used to integrate spectral and physicochemical features, such as soil moisture data, to generate a multidimensional soil feature dataset with a unified spatiotemporal benchmark; for example, a soil health index is generated by combining soil moisture, vegetation index, and temperature data.
[0153] The above method integrates satellite remote sensing spectral features, ground sensor physicochemical features, and environmental monitoring equipment features, and uses a spatiotemporal registration algorithm to achieve temporal synchronization and spatial alignment of multi-source data. This constructs a multidimensional soil feature dataset with unified spatiotemporal benchmarks, effectively solving the problem of representation distortion caused by data fragmentation and spatiotemporal bias in traditional methods. It provides a highly consistent and multidimensional fused data foundation for subsequent soil classification prediction, significantly improving the accuracy and reliability of the classification model.
[0154] This application also provides system embodiments that follow the above embodiments, for implementing the method steps of the above embodiments. The interpretation of the same names is the same as that of the above embodiments, and they have the same technical effects as those of the above embodiments, so they will not be repeated here.
[0155] like Figure 10 As shown, this application provides a soil classification prediction system based on environmental and remote sensing multi-source data, including:
[0156] Building unit 1001 is used to construct a soil feature dataset that includes spectral features, physicochemical features, and environmental features.
[0157] The processing unit 1002 is used to calculate the contribution of each feature to soil classification and the correlation between them, generate an initial weight configuration, update the weight coefficients and generate an adaptive weight configuration based on the spectral feature contribution meeting the first threshold condition and / or the correlation between physicochemical features and environmental features meeting the second threshold condition; based on the adaptive weight configuration, the selected feature subset with contribution exceeding the contribution threshold is weighted and fused to generate a soil classification feature vector; based on the soil classification feature vector, a soil classification decision tree is constructed and the optimal classification path is determined.
[0158] The prediction unit 1003 is used to predict the classification result of the soil sample to be classified based on the optimal classification path, determine whether the confidence of the classification result meets the confidence threshold condition, and if not, update the weight coefficients and return to the step of generating adaptive weight configuration until the confidence threshold condition is met and the final classification result is obtained.
[0159] Regarding the system in the above embodiments, the specific manner in which each module performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0160] Although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or serial order shown, or requiring all of the operations shown to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.
[0161] The methods, systems, devices, and storage media of this application can be implemented using standard programming techniques, and various method steps can be implemented using rule-based logic or other logic. It should also be noted that the terms "system" and "module" as used herein are intended to include implementations using one or more lines of software code and / or hardware implementations and / or devices for receiving input.
[0162] Any step, operation, or procedure described herein may be performed or implemented using one or more hardware or software modules, either alone or in combination with other devices. In one embodiment, the software module is implemented using a computer program product comprising a computer-readable medium containing computer program code, which is executable by a computer processor to perform any or all of the described steps, operations, or procedures.
[0163] The foregoing description of implementations of this application has been provided for illustrative and descriptive purposes. The foregoing description is not exhaustive and is not intended to limit this application to the exact forms disclosed. Various modifications and variations may exist in accordance with the foregoing teachings, or may arise from practice of this application. These embodiments were chosen and described to illustrate the principles of this application and its practical application, enabling those skilled in the art to utilize this application in various implementations and modifications to suit the specific purpose of the concept.
[0164] Regarding the system in the above embodiments, the specific manner in which each module performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0165] It can be further understood that, unless otherwise specified, "connection" includes both direct connections where no other components exist between the two parties and indirect connections where other components exist between them.
[0166] It is further understood that although the operations are described in a specific order in the accompanying drawings in the embodiments of this application, this should not be construed as requiring these operations to be performed in the specific order or serial order shown, or requiring all the operations shown to be performed to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.
[0167] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the field of this application that are not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0168] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
[0169] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A soil classification prediction method based on multi-source environmental and remote sensing data, characterized in that, include: Construct a soil feature dataset that includes spectral features, physicochemical features, and environmental features; Calculate the contribution of each feature to soil classification and the correlation between them to generate an initial weight configuration. Update the weight coefficients and generate an adaptive weight configuration based on the contribution of the spectral features meeting a first threshold condition and / or the correlation between the physicochemical features and the environmental features meeting a second threshold condition. Based on the adaptive weight configuration, the selected feature subsets with contribution exceeding the contribution threshold are weighted and fused to generate a soil classification feature vector. Based on the soil classification feature vector, a soil classification decision tree is constructed and the optimal classification path is determined; Based on the optimal classification path, the classification result of the soil sample to be classified is predicted. It is then determined whether the confidence of the classification result meets the confidence threshold condition. If not, the weight coefficient is updated, and the process returns to the step of generating adaptive weight configuration until the confidence threshold condition is met, and the final classification result is obtained. The method further includes: Extract the misclassified samples from the final classification results whose predicted labels do not match the true labels; Calculate the feature distribution deviation values between the spectral features and physicochemical features of the misclassified samples; In response to the feature distribution deviation value being greater than the deviation threshold, a gradient optimization algorithm is used to adjust the dynamic range boundary of the weight coefficients; The adaptive weight configuration is iteratively optimized based on the adjusted range boundary to generate and store the optimal weight configuration. The method further includes: Based on the optimal weight configuration, an adaptive weight configuration library is constructed for different environmental conditions; For a new soil sample, the feature distribution of the new soil sample is extracted, and the cosine similarity between the feature distribution of the new soil sample and the feature distribution of each weight configuration in the adaptive weight configuration library is calculated, and the maximum cosine similarity is determined. Determine whether the maximum cosine similarity is less than the matching threshold; In response to the maximum cosine similarity being less than the matching threshold, a new weight configuration is generated based on the environmental conditions of the new soil sample, classification is predicted based on the new weight configuration, and the new weight configuration is synchronized to the adaptive weight configuration library. In response to the maximum cosine similarity being no less than the matching threshold, a classification prediction is made based on the weight configuration corresponding to the maximum cosine similarity. The step of updating the weight coefficients and generating an adaptive weight configuration based on the spectral feature contribution satisfying a first threshold condition and / or the correlation between the physicochemical features and the environmental features satisfying a second threshold condition includes: The contribution of the spectral features is compared with a first threshold, and the correlation between the physicochemical features and the environmental features is compared with a second threshold; In response to the spectral feature contribution exceeding the first threshold, it is confirmed that the spectral feature contribution meets the first threshold condition, and the weighting coefficient of the spectral feature is increased; If the correlation between the physicochemical feature and the environmental feature is lower than the second threshold, it is confirmed that the correlation between the physicochemical feature and the environmental feature meets the second threshold condition, and the weight coefficient of the physicochemical feature is reduced. An adaptive weight configuration is generated by increasing the weight coefficients of the spectral features and / or decreasing the weight coefficients of the physicochemical features.
2. The method according to claim 1, characterized in that, The step of iteratively optimizing the adaptive weight configuration based on the adjusted range boundary, generating the optimal weight configuration, and storing it includes: The adjustment range of the weight coefficients is calculated based on the adjusted range boundaries; The adjustment range is iteratively optimized using the gradient descent method to obtain the updated weight configuration. The updated weight configuration is mapped and matched with the feature distribution of the misclassified samples; Determine whether the mapping match meets the goal of continuous optimization; In response to the mapping matching conforming to the continuous optimization objective, an optimal weight configuration is generated and stored; If the mapping match does not meet the goal of continuous optimization, the range boundary is readjusted, and the step of calculating the adjustment magnitude of the weight coefficient based on the adjusted range boundary is returned until the mapping match meets the goal of continuous optimization.
3. The method according to claim 1, characterized in that, The step of constructing a soil classification decision tree and determining the optimal classification path based on the soil classification feature vector includes: Based on the soil classification feature vector, the contribution ratio of spectral features and the contribution ratio of physicochemical features are calculated, and the group with the higher contribution ratio is determined. In response to the fact that the contribution ratio of the spectral features is higher than that of the physicochemical features, a spectral classification branch is preferentially constructed as the first layer of the decision tree, and child nodes are recursively generated under the spectral classification branch to form a soil classification decision tree. In response to the fact that the contribution ratio of the physicochemical features is higher than that of the spectral features, a physicochemical classification branch is preferentially constructed as the first layer of the decision tree, and child nodes are recursively generated under the physicochemical classification branch to form a soil classification decision tree; For all branches of the soil classification decision tree, cross-validation is used to evaluate the classification accuracy, and the branch with the highest classification accuracy is selected as the optimal classification path.
4. The method according to claim 1, characterized in that, The step of determining whether the confidence level of the classification result meets the confidence threshold condition, and if not, updating the weight coefficients and returning to the step of generating adaptive weight configuration, continues until the confidence threshold condition is met to obtain the final classification result, including: Calculate the classification confidence score of the classification result and compare it with a confidence score threshold; In response to the classification confidence score being lower than the confidence score threshold, a simulated annealing algorithm is used to iteratively optimize the weight coefficients of the spectral features and the physicochemical features in order to update the adaptive weight configuration; Based on the updated adaptive weight configuration, the soil classification results are regenerated, and the step of calculating the classification confidence of the classification results is returned until the classification confidence is not lower than the confidence threshold, thus obtaining the final classification result.
5. The method according to claim 1, characterized in that, The calculation of the contribution of each feature to soil classification and the correlation between them, to generate an initial weight configuration, includes: The contribution of the spectral features, physicochemical features, and environmental features to soil classification was calculated using the information entropy method. The correlation between the spectral features, the physicochemical features, and the environmental features was analyzed using the Pearson correlation coefficient. Based on the contribution of each feature and the correlation between the features, an initial weight configuration is generated using a weighted average method.
6. The method according to claim 1, characterized in that, The construction of the soil feature dataset, which includes spectral features, physicochemical features, and environmental features, includes: Acquire multi-source data characterizing soil properties; The multi-source data is synchronized in time and aligned in space by a spatiotemporal registration algorithm to generate a soil feature dataset containing spectral features, physicochemical features and environmental features. The multi-source data includes satellite remote sensing spectral data, physicochemical parameter data collected by ground sensors, and climate and topographic data collected by environmental monitoring equipment.
7. A soil classification and prediction system based on multi-source environmental and remote sensing data, characterized in that, For implementing the method as described in any one of claims 1-6, comprising: The building blocks are used to construct soil feature datasets that include spectral features, physicochemical features, and environmental features; The processing unit is used to calculate the contribution of each feature to soil classification and the correlation between them, generate an initial weight configuration, update the weight coefficients and generate an adaptive weight configuration based on the contribution of the spectral features meeting a first threshold condition and / or the correlation between the physicochemical features and the environmental features meeting a second threshold condition; based on the adaptive weight configuration, weighted fusion is performed on the selected feature subsets whose contribution exceeds the contribution threshold to generate a soil classification feature vector; and based on the soil classification feature vector, a soil classification decision tree is constructed and the optimal classification path is determined. The prediction unit is used to predict the classification result of the soil sample to be classified based on the optimal classification path, determine whether the confidence of the classification result meets the confidence threshold condition, and if not, update the weight coefficient and return to the step of generating adaptive weight configuration until the confidence threshold condition is met to obtain the final classification result.