Method for evaluating improvement effect of saline-alkali soil

By using the SMOTE algorithm based on quantum superposition states and the autoencoder of the feature importance attenuation function in saline-alkali soil data processing, the problems of sample imbalance and high-dimensional feature space are solved, and higher classification accuracy and adaptability are achieved.

CN119939227AActive Publication Date: 2025-05-06WATER RESOURCES RES INST OF SHANDONG PROVINCE

Patent Information

Application Number
CN202510436115.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

There are dimensional disasters caused by sample imbalance, noise interference and high-dimensional feature space in saline-alkali soil data processing, which affects the classification accuracy and generalization ability of machine learning models.

Method used

The data set is expanded by using SMOTE algorithm based on quantum superposition states, nonlinear perturbations are introduced through quantum tunneling effect, and perturbation optimization is performed in combination with local density to generate more complex and diverse samples. At the same time, the feature dimensionality reduction is performed using an autoencoder based on the feature importance attenuation function, removing redundant features and retaining key discriminant information.

Benefits of technology

The sample diversity and robustness of saline-alkali soil data are improved, and the generated samples are more in line with the real distribution, which improves the classification accuracy and adaptability of the model, especially when samples are scarce or categories are uneven.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939227A_ABST
    Figure CN119939227A_ABST
Patent Text Reader

Abstract

The invention relates to a saline-alkali soil improvement effect evaluation method, and belongs to the field of saline-alkali soil improvement evaluation. The method comprises the following steps: acquiring saline-alkali soil data, and manually marking the saline-alkali soil data to obtain an original saline-alkali soil data set; the original saline-alkali soil data set is expanded through the trained SMOTE algorithm based on the quantum superposition state, and a final expanded saline-alkali soil data set is obtained; preprocessing the expanded final saline-alkali soil data set to finally obtain preprocessed saline-alkali soil data; processing the preprocessed saline-alkali soil data through a machine learning model to obtain a classification result; and evaluating the improvement effect of the saline-alkali soil according to the classification result, and adjusting and optimizing the model according to the evaluation result. According to the method, the quantum state mapping and tunneling effect are introduced into the traditional SMOTE algorithm, so that the sample diversity and robustness of saline alkali soil data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of saline-alkali land improvement assessment, and in particular relates to a saline-alkali soil improvement effect assessment method. Background Art

[0002] With the increasing demand for saline-alkali soil governance and land use management, people have begun to explore how to use data science and machine learning methods to effectively analyze and process saline-alkali soil data. Soil data usually involves complex multidimensional features, including chemical composition, physical properties, humidity and other information. The special properties of saline-alkali soil make it an important but difficult soil type to handle. Traditional soil classification and management methods often face problems such as data imbalance, noise interference and dimensionality disaster caused by high-dimensional feature space. Especially in the process of saline-alkali soil data collection and analysis, we often face the dilemma of scarce minority class samples. The sample imbalance problem causes the machine learning model to be unable to fully learn the characteristics of the minority class, which in turn affects its classification accuracy and generalization ability. In addition, although the traditional data sampling method can increase the number of minority class samples, the samples it generates may be too simple and lack sufficient diversity, making it difficult to reflect the complex characteristics of the data. Summary of the invention

[0003] In order to solve the above problems, the present invention provides a method for evaluating the improvement effect of saline-alkali soil.

[0004] In order to achieve the above object, the present invention is implemented through the following technical solutions: The present invention provides a method for evaluating the improvement effect of saline-alkali soil, comprising the following steps: S1. Obtain saline-alkali soil data and manually annotate the saline-alkali soil data to obtain the original saline-alkali soil data set; S2. The original saline-alkali soil data set is expanded by the trained quantum superposition-based SMOTE algorithm to obtain the expanded final saline-alkali soil data set; nonlinear perturbations are introduced in the quantum superposition-based SMOTE algorithm by using the quantum tunneling effect; S3. Preprocess the expanded final saline-alkali soil data set to make it suitable for the training and analysis of the machine learning model, and finally obtain the preprocessed saline-alkali soil data; S4. The pre-processed saline-alkali soil data is processed by a machine learning model to obtain a classification result; the machine learning model includes a trained autoencoder and classifier model based on a feature importance decay function; S5. Evaluate the improvement effect of saline-alkali soil based on the classification results, and optimize the model based on the evaluation results.

[0005] Furthermore, step S1 specifically includes: The saline-alkali soil data include soil pH value, electrical conductivity, sodium adsorption ratio, total soluble salt, clay content percentage, organic matter content, exchangeable calcium-magnesium ratio, boron concentration, annual average precipitation anomaly, and temperature seasonal variation coefficient; The labeled categories include: extremely severe salinization, severe salinization, moderate salinization, mild salinization, and non-salinization.

[0006] Furthermore, in step S2, a SMOTE algorithm based on quantum superposition is constructed, and the training process of the SMOTE algorithm based on quantum superposition includes: S21. Perform quantum representation processing on each sample in the saline-alkali soil dataset to obtain a quantum superposition state representation ; S22. The strategy of generating synthetic samples based on the SMOTE algorithm generates new synthetic samples by linear interpolation. At the same time, multiple quantum state superpositions are used in the neighborhood of the original sample, combined with the quantum tunneling effect to generate new candidate samples. The formula is as follows: , in, Represents a new candidate sample; Indicates The disturbance adjustment factor of the sample is used to control the The perturbation amplitude of each sample; Indicates The local density of saline-alkali soil data sample points; Indicates The local density of saline-alkali soil data sample points; Indicates Feature representation of samples; Indicates Feature representation of samples; is a positive integer; is a positive integer; It is a quantum superposition state representation, which represents the superposition samples in the quantum state space; Represents the interpolation coefficient, which controls the distance between the generated sample and the original sample, satisfying ; represents the nonlinear perturbation introduced by quantum tunneling effect, represents the imaginary unit, Represents the phase angle of the quantum state; based on the quantum tunneling effect, nonlinear perturbations are introduced in the process of calculating the phase angle of the quantum state; S23. Adjust the direction and amplitude of the disturbance according to the local density and position of each sample. The calculation formula of the disturbance adjustment factor of the sample is expressed as follows: , in, It represents the decay rate of the control disturbance amplitude as the sample distance from the center point; It is the L2 norm, which is calculated in the same way as the Euclidean distance; The sample point representing the global center of the saline-alkali soil dataset; Indicates The local density of the sample reflects the The density of samples around a sample; represents the maximum local density in the saline-alkali soil dataset; S24. Identifying the boundaries of high-dimensional saline-alkali soil data space by local density measurement; S25. Adopt the adaptive perturbation optimization strategy to update the new sample. The formula is as follows: , in, Represents the new sample generated after updating; represents the learning rate of the perturbation; Indicates that the generated new sample is The distance measure between samples is Indicates that the generated new sample is A distance metric between samples, wherein the distance metric is calculated based on Euclidean distance; It represents the sample disturbance update amount, reflecting the disturbance direction and amplitude. The disturbance update amount is calculated based on the relative distance between samples and the superposition state of the synthetic samples after category balance. The formula is as follows: , in, represents the disturbance expansion factor; represents the optimized quantum state representation, characterizing the superposition state of the synthetic sample after category balancing; S26. The quantum state optimization mechanism is used to achieve sample category balance. By calculating the distribution of samples in each category and adjusting the generated sample ratio, the number of samples in each category in the saline-alkali soil dataset is dynamically balanced. The calculation formula of the optimized quantum state representation is as follows: , in, Indicates The optimization coefficient of each category, is a positive integer; Indicates The number of samples in each category; Represents the total number of samples in the saline-alkali soil dataset; Indicates the total number of samples in the current saline-alkali soil dataset; S27. The quality score of the new sample generated after the update is obtained by calculating the feature distance between the new sample generated after the update and the original sample. The formula is as follows: , in, represents the quality score of the new sample generated after updating, Indicates the number of samples involved in the evaluation; Indicates The weight of each sample in the quality assessment is adjusted according to its importance in the original saline-alkali soil dataset. The weight is adjusted according to the sample feature vector and is expressed as: , is a constant, set to 0.001; if the quality score of the new sample generated after the update is greater than the preset threshold, the sample is retained, otherwise the sample is discarded; S28. Combining the samples retained in step S27 with the original saline-alkali soil data to form an expanded final saline-alkali soil dataset.

[0007] Furthermore, step S3 specifically includes: The preprocessing operations include data cleaning, missing data filling, outlier detection and correction, and standardization, so that the preprocessed saline-alkali soil data obtained is suitable for the training and analysis of the machine learning model.

[0008] Furthermore, the saline-alkali soil data preprocessed in step S4 is subjected to a trained autoencoder based on a feature importance decay function to obtain low-dimensional saline-alkali soil data after feature dimension reduction, specifically including: An autoencoder based on a feature importance decay function is constructed, wherein the training process of the autoencoder based on a feature importance decay function includes: S41. An autoencoder based on a multi-layer neural network is constructed by implicit state space modeling. In the implicit state space modeling, the preprocessed saline-alkali soil data is mapped from a high-dimensional feature space to a low-dimensional implicit state space using a nonlinear activation function and an initial weight matrix, and then the weight matrix is ​​optimized by a manifold alignment regularization term. S42. In the dimensionality reduction process, the multi-scale saline-alkali soil data collaborative regularization term is used to optimize the dimensionality reduction process by integrating the saline-alkali soil data features at different scales; S43. The contribution of each feature is evaluated through the feature importance decay function. According to the contribution of the feature in the output result, the weight of the unimportant feature is decayed. The decay function dynamically adjusts the influence of the feature according to the relationship between the feature and the output, thereby optimizing the feature selection process; the feature decay process is refined through the gradient feedback mechanism; S44. In dynamic evolution and adaptive training, a feedback-based update strategy is used to update the autoencoder weights; S45. Combine global error and local error to optimize the objective function and force similar samples to maintain proximity in latent space; S46. The iteration termination condition is determined by the loss change. Training is stopped when the following conditions are met , in, Indicates The loss function of the autoencoder for the iteration is, Indicates The loss function of the autoencoder for the iteration is, is the convergence threshold.

[0009] Furthermore, in step S4, the low-dimensional saline-alkali soil data after feature dimension reduction is input into a classifier model for classification to obtain a classification result; the classifier model includes a random forest, a support vector machine, a decision tree, and a logistic regression model.

[0010] The advantages of the present invention are: The present invention adopts the SMOTE algorithm based on quantum superposition state. Through quantum state mapping and tunneling effect, the sample diversity and robustness of saline-alkali soil data can be improved. Compared with the traditional SMOTE algorithm, quantum state SMOTE can capture more complex nonlinear relationships in the feature space, and the generated synthetic samples are more consistent with the real distribution. Especially in the case of scarce samples or unbalanced categories, it can effectively expand the feature coverage of minority class samples and improve the classification accuracy of the model; the quantum tunneling effect is used to introduce nonlinear perturbations, and the perturbation optimization is performed in combination with the local density of the samples. The perturbation amplitude and direction can be adaptively adjusted according to the sample density in different regions, thereby effectively avoiding the generation of invalid noise samples and optimizing the generation quality of saline-alkali soil data sets. It can not only improve the quality of generated samples, but also effectively deal with boundary effect problems and enhance the adaptability of the model to sparse areas; the present invention adopts a feature importance attenuation function based on the feature importance attenuation function. The autoencoder dimensionality reduction method automatically evaluates and attenuates unimportant features through the feature importance decay function, retains key discriminant information in high-dimensional saline-alkali soil data, effectively removes redundant features, and optimizes the mapping process of the autoencoder in combination with the manifold alignment regularization term, thereby improving the separability and interpretability of saline-alkali soil data, thereby optimizing the feature dimensionality reduction process and avoiding the defect of possible loss of local structure in traditional dimensionality reduction methods; by adopting the collaborative regularization term of multi-scale saline-alkali soil data features, the representation ability of saline-alkali soil data at different scales is optimized, and the interaction between features of different scales is enhanced, so that the reduced-dimensional data can more truly reflect the complex nature of the soil; an adaptive perturbation optimization strategy is adopted to adaptively adjust the perturbation direction and amplitude, especially in areas with sparse data or large noise, and local density-driven perturbation optimization is used to avoid the problem of generating noise samples by traditional methods and improve the robustness and stability of generated samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0012] Figure 1 is a flow chart of the steps of the method of the present invention; Figure 2 This is a comparison chart of the classification accuracy between the method of the present invention and the existing method; Figure 3 A sample quality assessment diagram for the method of the present invention and the prior art; Figure 4 A comparison chart of the classification accuracy of the method of the present invention and the existing method at different noise levels; Figure 5 The figure is a comparison chart of algorithm efficiency between the method of the present invention and the existing method. DETAILED DESCRIPTION

[0013] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0014] Example 1 In this embodiment, Figure 1 As shown, the present invention provides a method for evaluating the effect of saline-alkali soil improvement, and the specific steps include: S1. Obtain saline-alkali soil data and manually annotate the saline-alkali soil data to obtain the original saline-alkali soil data set; Specifically, raw data related to saline-alkali soil improvement are collected from a variety of sources and preliminarily sorted and stored. The data collection sources include: 1) field sampling data: stratified soil sampling is carried out in salinized areas (such as agricultural land, industrial wasteland, and coastal saline areas), and real-time ion concentration data are obtained through electrochemical sensor arrays; 2) remote sensing vector data: combined with multi-spectral satellite remote sensing vector bands, the soil surface reflectivity characteristics are extracted; 3) meteorological and hydrological data: historical precipitation, evaporation and groundwater level dynamic monitoring data are obtained from meteorological stations.

[0015] The properties of the data include: is the soil pH value (logarithm of hydrogen ion activity), is the electrical conductivity (EC, dS / m), is the sodium adsorption ratio (SAR), is the total amount of soluble salt (g / kg), is the clay content percentage (%), is the organic matter content (g / kg), is the exchangeable calcium-magnesium ratio (CMR), is the boron concentration (mg / kg), is the annual average precipitation anomaly (mm), is the seasonal variation coefficient of temperature; it should be noted that this embodiment is only used to illustrate a data format and type of the present invention. In practical applications, the attributes of the data are usually more than 10 attributes, and the number of data attributes may reach dozens or even hundreds.

[0016] The collected data is labeled. The labeling method of the present invention is manual labeling. The labeling categories include: : Extremely severe salinization (improvement needs>90%); : Severe salinization (improvement needs 70% to 90%); : Moderate salinization (improvement requirement 50% to 70%); : Mild salinization (improvement requirement 30% to 50%); : Non-salinized soil (improvement requirement <30%).

[0017] S2. The original saline-alkali soil data set is expanded by the trained quantum superposition-based SMOTE algorithm to obtain the expanded final saline-alkali soil data set; nonlinear perturbations are introduced in the quantum superposition-based SMOTE algorithm by using the quantum tunneling effect; In order to solve the problems of sample imbalance, noise interference and sample scarcity in saline-alkali soil datasets, the quantum superposition-based SMOTE (Synthetic Minority Over-sampling Technique) algorithm is used. By mapping the original samples to the multidimensional quantum state space to generate multiple potential variants, the quantum tunneling effect is used to introduce nonlinear perturbations, dynamically adjust the perturbation amplitude and direction, and achieve category balance and sample quality assessment through quantum state optimization, thereby expanding the saline-alkali soil dataset, improving sample diversity and enhancing the generalization ability of the model. Specifically, the training process of the SMOTE algorithm based on quantum superposition is as follows: S21. Each sample in the original saline-alkali soil data set is quantum represented. Each original sample is represented by a quantum superposition state, and the sample is mapped to multiple possible states in the multidimensional feature space to retain the original feature information and expand the coverage of the feature space, so as to effectively deal with the problem of complex sample distribution. For the real-time ion concentration data of saline-alkali soil, the soil surface reflectivity characteristics, and the dynamic monitoring data of meteorological stations, these data themselves are usually uneven in spatial distribution, especially in areas with severe salinization. Some features may be relatively scarce or more complex. When collecting these data, due to the high noise interference of these data, it is easy to cause instability and overfitting of model training. Therefore, the original sample is mapped to multiple potential states in the multidimensional quantum space through quantum state mapping, which not only retains the original feature information, but also alleviates the problem of uneven data distribution by expanding the coverage of the feature space. The calculation method of the quantum superposition state is expressed as: , In the formula, It is a quantum superposition state representation, which represents the superposition samples in the quantum state space; Indicates The feature representation of each sample in the feature space, characterizing the feature vector of the sample; Indicates The quantum coefficient of a sample determines the weight of the sample in the superposition state; Indicates the total number of samples in the current saline-alkali soil dataset; is a positive integer; Furthermore, the quantum coefficient is determined based on the similarity between samples, and the calculation method is expressed as: , In the formula, represents the adjustment factor, controlling the change of quantum coefficients, , preferably, Set to 0.4; Indicates Feature representation of samples; is a positive integer; It is the L2 norm, which is calculated in the same way as the Euclidean distance; represents the normalization constant, which ensures that the sum of all quantum coefficients is 1, and is calculated as .

[0018] S22. The strategy of generating synthetic samples based on the SMOTE algorithm generates new synthetic samples by linear interpolation. At the same time, multiple quantum state superpositions are used in the neighborhood of the original sample. The nonlinear relationship between samples is explored in combination with the quantum tunneling effect to generate highly representative candidate samples. The samples are evaluated based on the sample similarity and neighborhood relationship. For example, the soil surface reflectivity may have a complex relationship in different bands. Traditional linear interpolation may not be enough, while nonlinear perturbations can generate more realistic samples. Quantum tunneling effect processing can generate multiple variants for each sample in the multidimensional quantum state space, thereby improving sample diversity, which can be expressed as: , In the formula, Represents the generated new samples, characterizing the synthetic samples in the feature space; It is The disturbance adjustment factor of samples controls the The perturbation amplitude of each sample; It is The local density of the saline-alkali soil data sample points indicates the The sample density around the saline-alkali soil data sample points; It is The local density of saline-alkali soil data sample points; Represents the interpolation coefficient, which controls the distance between the generated sample and the original sample, satisfying , preferably, Set to 0.5; represents the nonlinear perturbation introduced by quantum tunneling effect, represents the imaginary unit, represents the quantum state phase angle.

[0019] Furthermore, the phase angle is calculated based on the quantum tunneling effect to introduce nonlinear perturbations. The traditional SMOTE method may generate invalid samples when there is a lot of noise data, but the quantum tunneling effect can effectively adjust the amplitude and direction of the perturbation, especially in the boundary area of ​​saline-alkali soil data. The perturbation amplitude can be adaptively adjusted according to the local density. For example, in saline-alkali soil areas, the distribution of samples is usually uneven, and samples close to the boundary will obtain more generated samples through larger perturbations, thereby expanding the coverage of the training set and enhancing the stability of the model. The calculation method is expressed as: , In the formula, represents the adjustment factor that controls the phase angle change, , preferably, Set to 0.3; represents a small constant that prevents the denominator from being zero, , preferably, Set to 0.0001.

[0020] S23. Adjust the direction and amplitude of the disturbance according to the local density and position of each sample. The disturbance intensity of the sample depends not only on the distance between the samples, but also on the local density of the sample. For saline-alkali soil data, the disturbance amplitude of the samples located in the low-density area will be relatively increased, thereby enhancing the expansion of samples in the boundary area. For example, there may be fewer soil samples at the junction of agricultural land and industrial wasteland. The boundary is identified by local density, more samples are generated, and the poor performance of the model in the boundary area is avoided. The calculation method of the disturbance adjustment factor of the sample is expressed as: , In the formula, It controls the decay rate of the disturbance amplitude as the sample distance from the center point; It is the L2 norm, which is calculated in the same way as the Euclidean distance; It is the sample point of the global center of the saline-alkali soil dataset; It is The local density of the sample reflects the The density of samples around a sample; is the maximum local density in the saline-alkali soil dataset.

[0021] S24. Under high-dimensional saline-alkali soil data and complex distribution, traditional quantum superposition state and perturbation optimization sometimes cannot fully capture the boundary effect in the saline-alkali soil data space. The boundary effect refers to the fact that saline-alkali soil data points are often unevenly distributed in high-dimensional space, and noise data easily affects the effect of generating data. Therefore, in the process of saline-alkali soil data expansion, the present invention increases the recognition of boundary areas in the data space, and specifically identifies the boundary of the data space through local density measurement. The local density calculation method of the saline-alkali soil data sample point is expressed as: , In the formula, Indicates The local density of saline-alkali soil data sample points; It is the adjustment factor in density calculation, which controls the decay rate of sample similarity. Preferably, Set to 0.95.

[0022] S25. Adopt adaptive perturbation optimization strategy to update new samples, dynamically adjust the perturbation amplitude and direction according to the distance and distribution between samples, with smaller perturbation amplitude in dense areas and larger perturbation amplitude in sparse areas, and optimize the perturbation direction according to the local feature relationship to avoid generating noise samples. The calculation method is expressed as: , In the formula, Represents the new sample generated after updating; represents the learning rate of the perturbation, controlling the perturbation amplitude, preferably, Set to 0.01; Indicates that the generated new sample is The distance measure between samples is Indicates that the generated new sample is A distance metric between samples, wherein the distance metric is calculated based on Euclidean distance; It represents the sample disturbance update amount, reflecting the disturbance direction and amplitude.

[0023] Furthermore, the disturbance update amount is calculated based on the relative distance between samples and the superposition state of the synthetic samples after category balance, which is expressed as: , In the formula, represents the disturbance expansion factor, which controls the disturbance intensity. , preferably, Set to 0.1; It represents the optimized quantum state representation, characterizing the superposition state of the synthetic samples after category balancing.

[0024] S26. The quantum state optimization mechanism is used to achieve sample category balance. By calculating the distribution of samples in each category and adjusting the proportion of generated samples, the number of samples in each category in the saline-alkali soil dataset is dynamically balanced to avoid training bias caused by category imbalance. The boundary area samples of saline-alkali soil data usually contain more critical samples and have a greater impact on noise. Through the quantum state optimization mechanism, the direction and amplitude of the disturbance are automatically adjusted during the sample generation process to optimize the generation quality of samples in the boundary area and improve the diversity of training data and model accuracy. The calculation method is expressed as: , In the formula, Indicates The optimization coefficient of each category controls the weight of the generated samples among the categories. is a positive integer; Indicates The number of samples in each category; Represents the total number of samples in the saline-alkali soil dataset; Indicates the total number of samples in the current saline-alkali soil dataset; Indicates The feature representation of a sample.

[0025] Furthermore, the optimization coefficient is calculated according to the importance of the category samples, which is expressed as: , In the formula, Indicates The importance coefficient of the category represents The relative importance of samples of each category in the expansion process; Indicates The importance coefficient of each category, represents the sum of all category importance coefficients, Is a positive integer.

[0026] Furthermore, the importance coefficient is calculated based on the number of category samples, expressed as: , In the formula, Indicates The number of samples in each category, Is a positive integer.

[0027] S27. The new samples generated after the update are evaluated for quality. The evaluation criteria are based on the characteristic distance between the new samples generated after the update and the original samples, so as to screen out effective samples and improve the quality of the saline-alkali soil dataset. For saline-alkali soil data, combined with other multidimensional data such as historical precipitation, evaporation and groundwater level dynamic monitoring data, the new samples generated can effectively represent the saline-alkali soil characteristics in the real world. The calculation method is expressed as: , In the formula, Represents the quality score of the new sample generated after updating; Indicates The weight of each sample in the quality assessment is adjusted according to its importance in the original saline-alkali soil dataset; Indicates the number of samples involved in the evaluation.

[0028] Furthermore, the weight is adjusted according to the sample feature vector, and the calculation method is expressed as: , In the formula, Indicates The weight of the samples, is a constant. Preferably, Set to 0.001.

[0029] Furthermore, if the quality score of the new sample generated after the update is greater than a preset threshold, the sample is retained, otherwise the sample is discarded; preferably, the threshold is set to 0.8.

[0030] S28. All generated synthetic samples are combined with the original saline-alkali soil data to form the expanded final dataset, expressed as: In the formula, represents the final dataset after expansion; Respectively represent the first, second, and samples, each sample corresponds to a feature representation; Indicates the first, second, New samples generated after update, Indicates the number of synthetic samples.

[0031] S3. Preprocess the expanded final saline-alkali soil data set to make it suitable for the training and analysis of the machine learning model, and finally obtain the preprocessed saline-alkali soil data; The preprocessing operations include data cleaning, missing data filling, outlier detection and correction, and standardization, so that the preprocessed saline-alkali soil data obtained is suitable for the training and analysis of the machine learning model.

[0032] S4. The pre-processed saline-alkali soil data is processed by a machine learning model to obtain a classification result; the machine learning model includes a trained autoencoder and a classifier model based on a feature importance decay function; since saline-alkali soil data may have high-dimensional nonlinear characteristics, in order to solve the problems of redundant feature interference and local structure loss in the process of dimensionality reduction of high-dimensional saline-alkali soil data features, an autoencoder based on a feature importance decay function is used for dimensionality reduction, which effectively eliminates irrelevant features while retaining key discriminant information, thereby improving the analyzability and interpretability of saline-alkali soil data after dimensionality reduction. The training process of the autoencoder based on the feature importance decay function is as follows: S41. For the input saline-alkali soil data, an autoencoder based on a multi-layer neural network is constructed through the implicit state space modeling method. In the implicit state space modeling, the saline-alkali soil data is mapped from the high-dimensional feature space to the low-dimensional implicit state space, and the nonlinear activation function and the initial weight matrix are used for mapping to achieve preliminary saline-alkali soil data compression and mapping. The weight matrix is ​​optimized through the manifold alignment regularization term to solve the collinearity problem of high-dimensional saline-alkali soil data. The calculation method is expressed as: In the formula, represents the output features of the autoencoder, is the saline-alkali soil data input to the autoencoder; is the weight matrix of the autoencoder, which controls the linear transformation from input to latent space; is the bias term of the autoencoder, adjusting the output offset; It is the ReLU activation function to realize nonlinear mapping; is the manifold alignment regularization term, is the co-regularization term.

[0033] Furthermore, the manifold alignment regularization term optimizes the mapping accuracy by weighted adjustment of the relationship between features. The regularization maintains the manifold structure of saline-alkali soil data by weighting the distance between samples. The calculation method is expressed as: , In the formula, is the i-th row and j-th column element of the autoencoder weight matrix, i is a positive integer, and j is a positive integer; is the L2 norm; is the regularization coefficient, preferably, set to 0.1; is a scale parameter, preferably, set to 0.5; is the feature vector of the i-th sample; is the feature vector of the jth sample.

[0034] S42. In the process of dimensionality reduction, traditional dimensionality reduction methods may ignore the multi-scale characteristics of saline-alkali soil data, that is, saline-alkali soil data can be observed and processed from multiple levels and scales. The present invention adopts a multi-scale saline-alkali soil data collaborative regularization term, optimizes the dimensionality reduction process by fusing saline-alkali soil data features at different scales, decomposes the input saline-alkali soil data into representations at multiple scales, and uses collaborative regularization terms to enhance the interaction between different scales. The calculation method is expressed as: , In the formula, is the co-regularization term, The scale is The saline-alkali soil data show that The scale is The saline-alkali soil data show that; is the first saline-alkali soil data scale, indicating the front of the saline-alkali soil data Features is the second saline-alkali soil data scale, indicating the front of the saline-alkali soil data Features is a parameter that controls the similarity between scales. Preferably, Set to 0.2.

[0035] S43. The contribution of each feature is evaluated by the feature importance decay function. According to the contribution of the feature in the output result, the weight of the unimportant feature is decayed. The decay function dynamically adjusts the influence of the feature according to the relationship between the feature and the output, thereby optimizing the feature selection process and improving the dimensionality reduction effect. The calculation method of the feature decay coefficient is expressed as: , In the formula, For the The attenuation coefficient of the feature, is the decay rate, For the autoencoder Features and The weights between neurons, is the attenuation coefficient update value of the i-th feature. Preferably, Set to 2.

[0036] Furthermore, the feature attenuation process is refined through the gradient feedback mechanism. This mechanism adjusts the attenuation strength through the weight amplitude and refines the feature influence assessment in combination with the gradient direction, so that unimportant features are more effectively reduced. Local gradient feedback is used to adjust the feature attenuation coefficient to enhance the accuracy of attenuation. The calculation method of the feature attenuation coefficient update value is expressed as: , In the formula, Update the value of the attenuation coefficient of the i-th feature; is the gradient of the loss function of the autoencoder with respect to the input saline-alkali soil data; is the decay coefficient learning rate of the feature; is the connection weight between the kth neuron of the autoencoder and the i-th feature; is the decay rate. Preferably, Set to 0.01, Set to 0.95.

[0037] S44. In the dynamic evolution and adaptive training, the feedback update strategy is used to update the autoencoder weights to optimize the evolution of the autoencoder network topology, adjust the learning intensity according to the state difference between layers, and improve the adaptability of complex saline-alkali soil data. The update calculation method is expressed as: In the formula, is the parameter update operation, is the weight matrix of the autoencoder, is the learning rate for autoencoder weight updates, For the The attenuation coefficient of the feature, is the loss function of the autoencoder; is the control coefficient. Preferably, Set to 0.2, Set to 0.05.

[0038] S45. Combining global error and local error, optimizing the objective function makes the optimization process more comprehensive and accurate. This design forces similar samples to maintain proximity in the latent space while maintaining global reconstruction accuracy. The calculation method is expressed as: In the formula, is the loss function balance coefficient, is the global error of the autoencoder, is the local error of the autoencoder. Preferably, Set to 0.2.

[0039] Furthermore, the global error of the autoencoder The calculation method is expressed as: , In the formula, is the number of samples input to the autoencoder for the current batch, is the feature vector of the i-th sample, is the i-th reconstructed feature vector output by the decoder in the autoencoder.

[0040] Furthermore, the local error of the autoencoder The calculation method is expressed as: , In the formula, is the local weight coefficient, is the feature vector of the jth sample, is the jth reconstructed feature vector output by the decoder in the autoencoder. Preferably, Set to 0.2.

[0041] S46. The iteration termination condition is determined by the loss change, ensuring that the model is fully trained while avoiding overfitting. Training is stopped when the following conditions are met: , In the formula, Indicates The loss function of the autoencoder for the iteration is, Indicates The loss function of the autoencoder for the iteration is, is the convergence threshold. Preferably, Set to 1e-5.

[0042] Furthermore, the preprocessed saline-alkali soil data is passed through a trained autoencoder based on the feature importance decay function to obtain low-dimensional saline-alkali soil data after feature dimensionality reduction.

[0043] Furthermore, the low-dimensional saline-alkali soil data after feature dimension reduction is input into a classifier model for classification to obtain a classification result; the classifier model can be any one of a random forest, a support vector machine, a decision tree, and a logistic regression model, and the classification result includes: : Extremely severe salinization (improvement needs> 90%), : Severe salinization (improvement needs 70% to 90%), : Moderate salinization (improvement requirement 50% to 70%), : Mild salinization (improvement requirement 30% to 50%) and : Non-salinized soil (improvement requirement <30%).

[0044] S5. Evaluate the improvement effect of saline-alkali soil based on the classification results, and optimize the model based on the evaluation results.

[0045] Example 2 In this embodiment, Figure 2As shown in the figure, by comparing the classification accuracy of different oversampling methods, it is verified that in extreme class imbalance scenarios, this technology improves the classification performance of the model compared to the traditional oversampling method. The experimental results show that as the class imbalance ratio increases, the accuracy of traditional methods (such as random oversampling and Borderline-SMOTE) decreases significantly, especially when minority class samples are extremely scarce. The homogeneity problem of generated samples makes it difficult for the model to capture boundary features. This technology explores the nonlinear correlation between samples in the feature space through quantum superposition state mapping and tunneling perturbations, and generates diversified samples that are closer to the real distribution. Therefore, this technology can still maintain stable classification accuracy under high imbalance ratios, indicating that the samples it generates effectively expand the feature coverage of minority classes and alleviate the problem of narrow feature space caused by linear interpolation in traditional methods.

[0046] Example 3 In this embodiment, Figure 3 As shown in the figure, by evaluating the quality of generated samples and comparing the distribution consistency of samples generated by different methods with real samples, the generation quality of this technology is verified. The experimental results show that the distribution of samples generated by traditional methods shows obvious deviation, especially in areas with low data density, which is prone to outliers, reflecting the problem of noise accumulation in the oversampling process. The kernel density curve of this technology is highly consistent with the real sample, indicating that the generated samples not only retain the statistical characteristics of the original data, but also suppress noise interference through quantum state optimization. Therefore, the quantum state optimization mechanism ensures that the generated samples follow the intrinsic distribution law of the original data by dynamically evaluating the local density and phase angle adjustment of the samples, solving the distribution distortion problem caused by the fixed interpolation strategy of the traditional method.

[0047] Example 4 In this embodiment, Figure 4 As shown in the figure, by conducting noise robustness analysis, the stability of the algorithm in a noise-polluted environment is tested, and the effectiveness of the adaptive perturbation optimization strategy is verified. The experimental results show that as the noise ratio increases, the classification accuracy of the traditional method drops sharply, indicating that its generated samples are susceptible to noise propagation, resulting in blurred decision boundaries. This technology introduces nonlinear perturbations through the quantum tunneling effect, and dynamically adjusts the perturbation direction in combination with local density, so that the generated samples can filter out noise interference while retaining key features. Therefore, this technology can still maintain high accuracy in strong noise scenarios, proving that its perturbation optimization strategy can adaptively distinguish between effective features and noise, and enhance the robustness of generated samples.

[0048] Example 5 In this embodiment, Figure 5As shown in the figure, by comparing the algorithm efficiency, the computational efficiency of different oversampling algorithms is evaluated, and the optimization design of this technology is verified to save computing resources. The experimental results show that the time consumption of the traditional method increases approximately linearly with the increase of sample size, reflecting the computational redundancy of its global interpolation strategy. However, this technology significantly reduces the number of iterations of neighborhood search and interpolation optimization through local density-driven dynamic sampling and quantum state parallel calculation. Therefore, this technology can still maintain a low time overhead on large-scale data sets, indicating that it achieves a balance between computational efficiency and generation quality through the partitioning of feature space and the parallelism of quantum state superposition.

[0049] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for evaluating the effect of saline-alkali soil improvement, characterized in that: The following steps are involved: S1. Obtain saline-alkali soil data and manually annotate the saline-alkali soil data to obtain the original saline-alkali soil data set; S2. The original saline-alkali soil data set is expanded by the trained quantum superposition-based SMOTE algorithm to obtain the expanded final saline-alkali soil data set; nonlinear perturbations are introduced in the quantum superposition-based SMOTE algorithm by using the quantum tunneling effect; S3. Preprocess the expanded final saline-alkali soil data set to make it suitable for the training and analysis of the machine learning model, and finally obtain the preprocessed saline-alkali soil data; S4. The pre-processed saline-alkali soil data is processed by a machine learning model to obtain classification results; The machine learning model includes a trained autoencoder and classifier model based on a feature importance decay function; S5. Evaluate the improvement effect of saline-alkali soil based on the classification results, and optimize the model based on the evaluation results.

2. A saline-alkali soil improvement effect evaluation method according to claim 1, characterized in that: Step S1 specifically includes: The saline-alkali soil data include soil pH value, electrical conductivity, sodium adsorption ratio, total soluble salt, clay content percentage, organic matter content, exchangeable calcium-magnesium ratio, boron concentration, annual average precipitation anomaly, and temperature seasonal variation coefficient; The labeled categories include: extremely severe salinization, severe salinization, moderate salinization, mild salinization, and non-salinization.

3. A saline-alkali soil improvement effect evaluation method according to claim 2, characterized in that: In step S2, a SMOTE algorithm based on quantum superposition is constructed. The training process of the SMOTE algorithm based on quantum superposition includes: S21. Perform quantum representation processing on each sample in the saline-alkali soil dataset to obtain a quantum superposition state representation ; S22. The strategy of generating synthetic samples based on the SMOTE algorithm generates new synthetic samples by linear interpolation. At the same time, multiple quantum state superpositions are used in the neighborhood of the original sample, combined with the quantum tunneling effect to generate new candidate samples. The formula is as follows: , in, Represents a new candidate sample; Indicates The disturbance adjustment factor of the sample is used to control the The perturbation amplitude of each sample; Indicates The local density of saline-alkali soil data sample points; Indicates The local density of saline-alkali soil data sample points; Indicates Feature representation of samples; Indicates Feature representation of samples; is a positive integer; is a positive integer; It is a quantum superposition state representation, which represents the superposition samples in the quantum state space; Represents the interpolation coefficient, which controls the distance between the generated sample and the original sample, satisfying ; represents the nonlinear perturbation introduced by quantum tunneling effect, represents the imaginary unit, Represents the phase angle of the quantum state; based on the quantum tunneling effect, nonlinear perturbations are introduced in the process of calculating the phase angle of the quantum state; S23. Adjust the direction and amplitude of the disturbance according to the local density and position of each sample. The calculation formula of the disturbance adjustment factor of the sample is expressed as follows: , in, It represents the decay rate of the control disturbance amplitude as the sample distance from the center point; It is the L2 norm, which is calculated in the same way as the Euclidean distance; The sample point representing the global center of the saline-alkali soil dataset; Indicates The local density of the samples reflects the The density of samples around a sample; represents the maximum local density in the saline-alkali soil dataset; S24. Identifying the boundaries of high-dimensional saline-alkali soil data space by local density measurement; S25. Adopt the adaptive perturbation optimization strategy to update the new sample. The formula is as follows: , in, Represents the new sample generated after updating; represents the learning rate of the perturbation; Indicates that the generated new sample is The distance measure between samples is Indicates that the generated new sample is A distance metric between samples, wherein the distance metric is calculated based on Euclidean distance; It represents the sample disturbance update amount, reflecting the disturbance direction and amplitude. The disturbance update amount is calculated based on the relative distance between samples and the superposition state of the synthetic samples after category balance. The formula is as follows: , in, represents the disturbance expansion factor; represents the optimized quantum state representation, characterizing the superposition state of the synthetic sample after category balancing; S26. The quantum state optimization mechanism is used to achieve sample category balance. By calculating the distribution of samples in each category and adjusting the generated sample ratio, the number of samples in each category in the saline-alkali soil dataset is dynamically balanced. The calculation formula of the optimized quantum state representation is as follows: , in, Indicates The optimization coefficient of each category, is a positive integer; Indicates The number of samples in each category; Represents the total number of samples in the saline-alkali soil dataset; Indicates the total number of samples in the current saline-alkali soil dataset; S27. The quality score of the new sample generated after the update is obtained by calculating the feature distance between the new sample generated after the update and the original sample. The formula is as follows: , in, represents the quality score of the new sample generated after updating, Indicates the number of samples involved in the evaluation; Indicates The weight of each sample in the quality assessment is adjusted according to its importance in the original saline-alkali soil dataset. The weight is adjusted according to the sample feature vector and the formula is expressed as , is a constant, set to 0.001; if the quality score of the new sample generated after the update is greater than the preset threshold, the sample is retained, otherwise the sample is discarded; S28. Combining the samples retained in step S27 with the original saline-alkali soil data to form an expanded final saline-alkali soil dataset.

4. A saline-alkali soil improvement effect evaluation method according to claim 3, characterized in that: Step S3 specifically includes: The preprocessing operations include data cleaning, missing data filling, outlier detection and correction, and standardization, so that the preprocessed saline-alkali soil data obtained is suitable for the training and analysis of the machine learning model.

5. A saline-alkali soil improvement effect evaluation method according to claim 4, characterized in that: The saline-alkali soil data preprocessed in step S4 is subjected to a trained autoencoder based on a feature importance attenuation function to obtain low-dimensional saline-alkali soil data after feature dimension reduction, specifically including: An autoencoder based on a feature importance decay function is constructed, wherein the training process of the autoencoder based on a feature importance decay function includes: S41. An autoencoder based on a multi-layer neural network is constructed by implicit state space modeling. In the implicit state space modeling, the preprocessed saline-alkali soil data is mapped from a high-dimensional feature space to a low-dimensional implicit state space using a nonlinear activation function and an initial weight matrix, and then the weight matrix is ​​optimized by a manifold alignment regularization term. S42. In the dimensionality reduction process, the multi-scale saline-alkali soil data collaborative regularization term is used to optimize the dimensionality reduction process by integrating the saline-alkali soil data features at different scales; S43. The contribution of each feature is evaluated through the feature importance decay function. According to the contribution of the feature in the output result, the weight of the unimportant feature is decayed. The decay function dynamically adjusts the influence of the feature according to the relationship between the feature and the output, thereby optimizing the feature selection process; the feature decay process is refined through the gradient feedback mechanism; S44. In dynamic evolution and adaptive training, a feedback-based update strategy is used to update the autoencoder weights; S45. Combine global error and local error to optimize the objective function and force similar samples to maintain proximity in latent space; S46. The iteration termination condition is determined by the loss change. Training is stopped when the following conditions are met: , in, Indicates The loss function of the autoencoder for the iteration is, Indicates The loss function of the autoencoder for the iteration is, is the convergence threshold.

6. A saline-alkali soil improvement effect evaluation method according to claim 5, characterized in that: In step S4, the low-dimensional saline-alkali soil data after feature dimension reduction is input into a classifier model for classification to obtain a classification result; the classifier model includes a random forest, a support vector machine, a decision tree, and a logistic regression model.

Citation Information

Patent Citations

  • A parameter selection optimization method, system and equipment in random forest model training

    CN113591944A

  • Environmental backscatter system communication error rate assessment method based on artificial intelligence

    CN118740343A

  • Effect evaluation method for improving saline-alkali soil by desulfurized gypsum

    CN118983025A

  • Power grid toughness evaluation method and system

    CN119293488A

  • Slope soil fertility prediction method

    CN119359475A

Cited By

  • Biochar matching decision-making method driven by saline-alkali soil improvement target

    CN120765066A

  • Ploughing layer salt and alkali conditioning optimization method based on soybean response

    CN121168767A