A method for evaluating the improvement effect of saline-alkali soil

Through the SMOTE algorithm based on quantum superposition state and the autoencoder of the feature importance attenuation function, the problem of sample imbalance and noise interference in saline-alkali soil data is solved, and diverse and robust samples are generated, which improves the classification accuracy and feature selection effect of saline-alkali soil data.

CN119939227BActive Publication Date: 2025-07-25WATER RESOURCES RES INST OF SHANDONG PROVINCE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510436115.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

During the collection and analysis of saline-alkali soil data, it faces dimensional disasters caused by sample imbalance, noise interference and high-dimensional feature space. Traditional methods are difficult to deal with effectively, especially when a few samples are scarce, affecting classification accuracy and generalization ability.

Method used

The SMOTE algorithm based on quantum superposition states is adopted, and data expansion is expanded by combining the quantum tunneling effect, and dimensionality reduction is performed through the autoencoder of the feature importance attenuation function. The quantum state mapping and tunneling effect are used to generate diverse and robust samples, and the adaptive perturbation optimization strategy and multi-scale feature synergistic regularization term is combined to optimize the generation and feature selection of saline-alkali soil data.

Benefits of technology

The sample diversity and robustness of saline-alkali soil data are improved, the classification accuracy and adaptability of the model are enhanced, the boundary effect and noise interference problems are effectively solved, the characteristic dimensionality reduction process is optimized, and the separability and interpretability of saline-alkali soil data are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939227B_ABST
    Figure CN119939227B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for evaluating the improvement effect of saline-alkali soil, belonging to the field of saline-alkali land improvement evaluation. It includes the following steps: obtaining saline-alkali soil data, and manually annotating the saline-alkali soil data to obtain the original saline-alkali soil data set; expanding the original saline-alkali soil data set through the trained SMOTE algorithm based on quantum superposition states to obtain the final expanded saline-alkali soil data set; preprocessing the final expanded saline-alkali soil data set to finally obtain the preprocessed saline-alkali soil data; the preprocessed saline-alkali soil data is processed through a machine learning model to obtain a classification result; evaluating the improvement effect of the saline-alkali soil according to the classification result, and optimizing the model according to the evaluation result. By introducing quantum state mapping and tunneling effect into the traditional SMOTE algorithm, the present invention can improve the sample diversity and robustness of saline-alkali soil data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of saline-alkali soil improvement evaluation, and particularly relates to a method for evaluating the improvement effect of saline-alkali soil. Background Art

[0002] With the increasing demands for saline-alkali soil treatment and land use management, people have started to explore how to effectively analyze and process saline-alkali soil data using data science and machine learning methods. Soil data usually involves complex multi-dimensional features, including chemical composition, physical properties, humidity and other aspects of information. The special properties of saline-alkali soil make it an important but difficult-to-process soil type. Traditional soil classification and management methods often face problems such as data imbalance, noise interference, and the curse of dimensionality caused by high-dimensional feature spaces. Especially in the process of saline-alkali soil data collection and analysis, there is often a dilemma of scarce minority class samples. The sample imbalance problem causes machine learning models to be unable to fully learn the features of the minority class, thereby affecting their classification accuracy and generalization ability. In addition, although traditional data sampling methods can increase the number of minority class samples, the generated samples may be too simple and lack sufficient diversity, making it difficult to reflect the complex features of the data. Summary of the Invention

[0003] In order to solve the above problems, the present invention provides a method for evaluating the improvement effect of saline-alkali soil.

[0004] To achieve the above object, the present invention is realized through the following technical solutions:

[0005] The present invention provides a method for evaluating the improvement effect of saline-alkali soil, including the following steps:

[0006] S1. Obtain saline-alkali soil data and perform manual annotation on the saline-alkali soil data to obtain an original saline-alkali soil data set;

[0007] S2. Expand the original saline-alkali soil data set through the trained SMOTE algorithm based on quantum superposition states to obtain an expanded final saline-alkali soil data set; in the SMOTE algorithm based on quantum superposition states, introduce non-linear perturbations using the quantum tunneling effect;

[0008] S3. Preprocess the expanded final saline-alkali soil data set to make it suitable for the training and analysis of machine learning models, and finally obtain preprocessed saline-alkali soil data;

[0009] S4. The preprocessed saline-alkali soil data is processed through a machine learning model to obtain a classification result; the machine learning model includes a trained autoencoder and a classifier model based on a feature importance decay function;

[0010] S5. Evaluate the improvement effect of saline-alkali soil according to the classification results, and optimize the model according to the evaluation results.

[0011] Furthermore, step S1 specifically includes:

[0012] The saline-alkali soil data includes soil pH value, electrical conductivity, sodium adsorption ratio, total soluble salt content, percentage of clay content, organic matter content, exchangeable calcium-magnesium ratio, boron element concentration, annual precipitation anomaly value, and temperature seasonal variation coefficient;

[0013] The labeled categories include: extremely severe salinization, severe salinization, moderate salinization, mild salinization, and non-salinization.

[0014] Furthermore, in step S2, construct the SMOTE algorithm based on quantum superposition states. The training process of the SMOTE algorithm based on quantum superposition states includes:

[0015] S21. Perform quantum representation processing on each sample in the saline-alkali soil dataset to obtain a quantum superposition state representation ;

[0016] S22. Based on the strategy of the SMOTE algorithm to generate synthetic samples, generate new synthetic samples through linear interpolation. At the same time, adopt multiple quantum state superpositions within the neighborhood of the original samples, and combine the quantum tunneling effect to generate new candidate samples. The formula is expressed as follows:

[0017] ,

[0018] where, represents the new candidate sample; represents the perturbation adjustment factor of the th sample, which is used to control the perturbation amplitude of the th sample; represents the local density of the th saline-alkali soil data sample point; represents the local density of the th saline-alkali soil data sample point; represents the feature representation of the th sample; represents the feature representation of the th sample; is a positive integer; is a positive integer; is the quantum superposition state representation, which characterizes the samples superimposed in the quantum state space; represents the interpolation coefficient, which controls the distance between the generated sample and the original sample, satisfying ; represents the non-linear perturbation introduced by using the quantum tunneling effect, represents the imaginary unit, represents the quantum state phase angle; according to the quantum tunneling effect, a non-linear perturbation is introduced in the process of calculating the quantum state phase angle;

[0019] S23. Adjust the direction and amplitude of the perturbation according to the local density and position of each sample. The calculation formula of the perturbation adjustment factor of the sample is expressed as follows:

[0020] ,

[0021] where, represents the decay rate of the control perturbation amplitude with respect to the distance of the sample from the center point; is the L2 norm, the same as the Euclidean distance calculation method; represents the sample point of the global center of the saline-alkali soil dataset; represents the th local density of the sample, reflecting the sample density around the th sample; represents the maximum local density in the saline-alkali soil dataset;

[0022] S24. Identify the boundary of the high-dimensional saline-alkali soil data space through local density metrics;

[0023] S25. Update the new sample using an adaptive perturbation optimization strategy. The formula is expressed as follows:

[0024] ,

[0025] where, represents the newly generated sample after update; represents the learning rate of the perturbation; represents the distance metric between the newly generated sample and the th sample, represents the distance metric between the newly generated sample and the th sample. The distance metric is calculated based on the Euclidean distance; represents the sample perturbation update amount, reflecting the perturbation direction and amplitude; the perturbation update amount is calculated according to the relative distance between samples and the superposition state of the synthesized samples after class balance. The formula is expressed as follows:

[0026] ,

[0027] where, represents the perturbation expansion factor; represents the optimized quantum state representation, characterizing the superposition state of the synthesized samples after class balance;

[0028] S26. Implement sample class balance using a quantum state optimization mechanism. By calculating the sample distribution of each class and adjusting the proportion of generated samples, dynamically balance the number of samples of each class in the saline-alkali soil dataset. The calculation formula of the optimized quantum state representation is as follows:

[0029] ,

[0030] where, represents the optimization coefficient of the th class, is a positive integer; represents the number of samples of the th class; represents the total number of samples in the saline-alkali soil dataset; represents the current total number of samples in the saline-alkali soil dataset;

[0031] S27. By calculating the feature distance between the newly generated samples after update and the original samples, obtain the quality score of the newly generated samples after update. The formula is as follows:

[0032] ,

[0033] where, represents the quality score of the newly generated samples after update, represents the number of samples participating in the evaluation; represents the weight of the th sample in the quality evaluation, which is adjusted according to its importance in the original saline-alkali soil dataset. This weight is adjusted based on the sample feature vector, and the formula is , is a constant, set to 0.001; if the quality score of the newly generated samples after update is greater than the preset threshold, then retain the samples, otherwise discard the samples;

[0034] S28. Combine the samples retained in step S27 with the original saline-alkali soil data to form the expanded final saline-alkali soil dataset.

[0035] Furthermore, step S3 specifically includes:

[0036] The preprocessing operation includes data cleaning, missing data filling, outlier detection and correction, and standardization, so that the preprocessed saline-alkali soil data is suitable for the training and analysis of machine learning models.

[0037] Furthermore, the preprocessed saline-alkali soil data in step S4 passes through a trained autoencoder based on the feature importance decay function to obtain the low-dimensional saline-alkali soil data after feature reduction, specifically including:

[0038] Construct an autoencoder based on a feature importance decay function. The training process of the autoencoder based on the feature importance decay function includes:

[0039] S41. By means of implicit state space modeling, construct an autoencoder based on a multi-layer neural network. In implicit state space modeling, the preprocessed saline-alkali soil data is mapped from a high-dimensional feature space to a low-dimensional implicit state space, and the mapping is carried out using a non-linear activation function and an initial weight matrix, and then the weight matrix is optimized through a manifold alignment regularization term;

[0040] S42. In the dimensionality reduction process, adopt a multi-scale saline-alkali soil data collaborative regularization term to optimize the dimensionality reduction process by fusing the characteristics of saline-alkali soil data at different scales;

[0041] S43. Evaluate the contribution of each feature through a feature importance decay function, and according to the contribution degree of the feature in the output result, decay the weight of the unimportant feature. The decay function dynamically adjusts the influence of the feature according to the relationship between the feature and the output, so as to optimize the feature selection process; refine the feature decay process through a gradient feedback mechanism;

[0042] S44. In dynamic evolution and adaptive training, adopt a strategy based on feedback update to update the weights of the autoencoder;

[0043] S45. Combine the global error and the local error to optimize the objective function and force similar samples to maintain a neighboring relationship in the latent space; S46. The iteration termination condition is judged by the change in loss. When the following conditions are met, stop training

[0044] ,

[0045] wherein, represents the loss function of the autoencoder in the -th iteration, represents the loss function of the autoencoder in the -th iteration, is the convergence threshold.

[0046] Furthermore, in step S4, the low-dimensional saline-alkali soil data obtained after feature dimensionality reduction is input into a classifier model for classification to obtain a classification result; the classifier model includes a random forest, a support vector machine, a decision tree, and a logistic regression model.

[0047] The advantages of the present invention are:

[0048] The present invention adopts the SMOTE algorithm based on quantum superposition states. Through quantum state mapping and tunneling effect, it can improve the sample diversity and robustness of saline-alkali soil data. Compared with the traditional SMOTE algorithm, quantum state SMOTE can capture more complex non-linear relationships in the feature space, and the generated synthetic samples are more in line with the real distribution. Especially in the case of scarce samples or class imbalance, it can effectively expand the feature coverage of minority class samples and improve the classification accuracy of the model. By introducing non-linear perturbations using the quantum tunneling effect and optimizing the perturbations in combination with the local density of samples, it can adaptively adjust the perturbation amplitude and direction according to the sample density in different regions, thus effectively avoiding the generation of invalid noise samples and optimizing the generation quality of the saline-alkali soil data set. It can not only improve the quality of the generated samples, but also effectively address the boundary effect problem and enhance the adaptability of the model to sparse regions. The present invention adopts an autoencoder dimensionality reduction method based on a feature importance decay function. Through the feature importance decay function, it automatically evaluates and attenuates unimportant features, retains key discriminant information in high-dimensional saline-alkali soil data, effectively removes redundant features, and optimizes the mapping process of the autoencoder in combination with the manifold alignment regularization term, enhancing the separability and interpretability of saline-alkali soil data, thereby optimizing the feature dimensionality reduction process and avoiding the defect of possible loss of local structure in traditional dimensionality reduction methods. By adopting a collaborative regularization term for multi-scale saline-alkali soil data features, it optimizes the representation ability of saline-alkali soil data at different scales, enhances the interaction between features at different scales, and enables the dimensionality-reduced data to more truly reflect the complex properties of the soil. An adaptive perturbation optimization strategy is adopted. By adaptively adjusting the perturbation direction and amplitude, especially in regions with sparse data or high noise, through local density-driven perturbation optimization, the problem of generating noise samples by traditional methods is avoided, and the robustness and stability of the generated samples are improved. Description of the Drawings

[0049] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.

[0050] Figure 1 It is a flowchart of the steps of the method of the present invention;

[0051] Figure 2 It is a comparison chart of the classification accuracy between the method of the present invention and the existing method;

[0052] Figure 3 It is an evaluation chart of the quality of the generated samples between the method of the present invention and the existing method;

[0053] Figure 4 It is a comparison chart of the classification accuracy between the method of the present invention and the existing method under different noise levels;

[0054] Figure 5This is a comparison chart of the algorithm efficiency between the method of the present invention and the existing methods. Specific embodiments

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] Embodiment 1

[0057] In this embodiment, as Figure 1 shown, the present invention provides a method for evaluating the improvement effect of saline-alkali soil, and the specific steps include:

[0058] S1. Obtain saline-alkali soil data and perform manual annotation on the saline-alkali soil data to obtain the original saline-alkali soil dataset;

[0059] Specifically, collect the original data related to the improvement of saline-alkali soil from multiple sources, and perform preliminary sorting and storage. The data collection sources include: 1) Field sampling data: Conduct layered soil sampling in saline-alkali areas (such as agricultural land, industrial waste land, coastal saline areas), and obtain real-time ion concentration data through an electrochemical sensor array; 2) Remote sensing vector data: Combine multi-spectral satellite remote sensing vector bands to extract the reflectance characteristics of the soil surface layer; 3) Meteorological and hydrological data: Obtain historical precipitation, evaporation, and dynamic monitoring data of the groundwater level from meteorological stations.

[0060] The attributes of the data include: is the soil pH value (negative logarithm of hydrogen ion activity), is the electrical conductivity (EC, dS / m), is the sodium adsorption ratio (SAR), is the total soluble salt content (g / kg), is the percentage of clay content (%); is the organic matter content (g / kg), is the exchangeable calcium-magnesium ratio (CMR), is the boron element concentration (mg / kg), is the annual precipitation anomaly value (mm), is the temperature seasonal variation coefficient; it should be noted that this embodiment is only to illustrate a data format and type of the present invention. In actual applications, the attributes of the data are usually more than 10 attributes, and the number of data attributes may reach dozens or even hundreds.

[0061] Label the collected data. The labeling method of the present invention is manual labeling, and the labeled categories include:

[0062] : Extremely severe salinization (improvement demand > 90%);

[0063] : Severe salinization (improvement demand 70% - 90%);

[0064] : Moderate salinization (improvement demand 50% - 70%);

[0065] : Mild salinization (improvement demand 30% - 50%);

[0066] : Non - salinized soil (improvement demand < 30%).

[0067] S2. Augment the original saline - alkali soil dataset through the trained SMOTE algorithm based on quantum superposition state to obtain the final augmented saline - alkali soil dataset; introduce non - linear perturbations using the quantum tunneling effect in the SMOTE algorithm based on quantum superposition state;

[0068] To solve the problems of unbalanced samples, noise interference, and sample scarcity in the saline - alkali soil dataset, the SMOTE (Synthetic Minority Over - sampling Technique) algorithm based on quantum superposition state is adopted. By mapping the original samples to a multi - dimensional quantum state space to generate multiple potential variants, then introducing non - linear perturbations using the quantum tunneling effect, dynamically adjusting the perturbation amplitude and direction, and achieving class balance and sample quality evaluation through quantum state optimization, thus augmenting the saline - alkali soil dataset to improve sample diversity and enhance the generalization ability of the model. Specifically, the training process of the SMOTE algorithm based on quantum superposition state is as follows:

[0069] S21. Perform quantum representation processing on each sample in the original saline-alkali soil dataset. Each original sample is processed through quantum superposition state representation, mapping the sample to multiple possible states in a multi-dimensional feature space to retain the original feature information and broaden the feature space coverage, thereby effectively addressing the problem of complex sample distribution. For real-time ion concentration data of saline-alkali soil, soil surface reflectance characteristics, and dynamic monitoring data of meteorological stations, etc., these data themselves usually have uneven spatial distribution. Especially in severely saline-alkali areas, some features may be relatively scarce or complex. During data collection, due to the high noise interference in these data, it is easy to cause instability and overfitting phenomena in model training. Therefore, mapping the original samples to multiple potential states in a multi-dimensional quantum space through quantum state mapping not only retains the original feature information but also alleviates the problem of uneven data distribution by expanding the coverage of the feature space. The calculation method of the quantum superposition state is expressed as:

[0070] ,

[0071] In the formula, is the quantum superposition state representation, characterizing the samples superimposed in the quantum state space; represents the feature representation of the -th sample in the feature space, characterizing the feature vector of the sample; represents the -th sample's quantum coefficient, determining the weight of this sample in the superposition state; represents the total number of samples in the current saline-alkali soil dataset; is a positive integer;

[0072] Furthermore, determine the quantum coefficient based on the similarity between samples, and the calculation method is expressed as:

[0073] ,

[0074] In the formula, represents the adjustment factor, controlling the change of the quantum coefficient, , preferably, is set to 0.4; represents the feature representation of the -th sample; is a positive integer; is the L2 norm, with the same calculation method as the Euclidean distance; represents the normalization constant, ensuring that the sum of all quantum coefficients is 1, and its calculation method is .

[0075] S22. The strategy of generating synthetic samples based on the SMOTE algorithm generates new synthetic samples through linear interpolation. At the same time, multiple quantum state superpositions are adopted within the neighborhood of the original samples, and the quantum tunneling effect is combined to explore the non-linear relationships between samples, generating candidate samples with high representativeness, and evaluating them based on sample similarity and neighborhood relationships. For example, the reflectance of the soil surface layer may have complex relationships in different bands, and traditional linear interpolation may not be sufficient, while non-linear perturbations can generate more realistic samples. The quantum tunneling effect processing can generate multiple variants for each sample in the multi-dimensional quantum state space, enhancing sample diversity, expressed as:

[0076] ,

[0077] In the formula, represents the newly generated sample, characterizing the synthetic sample in the feature space; is the perturbation adjustment factor of the -th sample, controlling the perturbation amplitude of the -th sample; is the local density of the -th saline-alkali soil data sample point, representing the sample density around the -th saline-alkali soil data sample point; is the local density of the -th saline-alkali soil data sample point; represents the interpolation coefficient, controlling the distance between the generated sample and the original sample, satisfying , preferably, is set to 0.5; represents the non-linear perturbation introduced by using the quantum tunneling effect, represents the imaginary unit, represents the quantum state phase angle.

[0078] Furthermore, the phase angle is calculated based on the quantum tunneling effect to introduce non-linear perturbations. The traditional SMOTE method may generate invalid samples in the case of more noise data, but the quantum tunneling effect can effectively adjust the amplitude and direction of the perturbations. Especially in the boundary region of the saline-alkali soil data, the perturbation amplitude can be adaptively adjusted according to the local density. For example, in the saline-alkali soil area, the sample distribution is usually uneven, and the samples near the boundary will get more generated samples through larger amplitude perturbations, thus expanding the coverage of the training set and enhancing the stability of the model. The calculation method is expressed as:

[0079] ,

[0080] In the formula, represents the adjustment factor controlling the change of the phase angle, , preferably, Set to 0.3; Represents a small constant to prevent the denominator from being zero, , preferably, Set to 0.0001.

[0081] S23. Adjust the direction and amplitude of the perturbation according to the local density and position of each sample. The perturbation intensity of the sample not only depends on the distance between samples but also is adaptively adjusted according to the local density of the sample. For saline-alkali soil data, for samples located in areas with lower density, the perturbation amplitude will increase relatively, thereby enhancing the expansion of samples in the boundary area. For example, there may be fewer soil samples at the junction of agricultural land and industrial wasteland. Identify the boundary through local density and generate more samples to avoid poor performance of the model in the boundary area. The calculation method of the perturbation adjustment factor of the sample is expressed as:

[0082] ,

[0083] In the formula, is the attenuation rate that controls the perturbation amplitude with respect to the distance of the sample from the center point; is the L2 norm, the same as the calculation method of the Euclidean distance; is the sample point of the global center of the saline-alkali soil data set; is the th local density of the sample, reflecting the sample density around the th sample; is the maximum local density in the saline-alkali soil data set.

[0084] S24. Under high-dimensional saline-alkali soil data and complex distributions, traditional quantum superposition states and perturbation optimization sometimes cannot fully capture the boundary effect in the saline-alkali soil data space. The boundary effect refers to the fact that saline-alkali soil data points are often unevenly distributed in high-dimensional space, and noise data easily affects the effect of the generated data. Therefore, in the process of expanding saline-alkali soil data, the present invention increases the identification of the boundary area in the data space, specifically by identifying the boundary of the data space through local density measurement. The calculation method of the local density of the saline-alkali soil data sample point is expressed as:

[0085] ,

[0086] In the formula, represents the local density of the th saline-alkali soil data sample point; is the adjustment factor in density calculation, controlling the attenuation rate of sample similarity. Preferably, is set to 0.95.

[0087] S25. Update the new samples using an adaptive perturbation optimization strategy, dynamically adjust the perturbation amplitude and direction according to the distance and distribution between samples, with a smaller perturbation amplitude in dense regions and a larger perturbation amplitude in sparse regions, and optimize the perturbation direction based on local feature relationships to avoid generating noisy samples. The calculation method is expressed as:

[0088] ,

[0089] In the formula, represents the newly generated sample after update; represents the learning rate of the perturbation, which controls the perturbation amplitude. Preferably, is set to 0.01; represents the distance metric between the newly generated sample and the -th sample, represents the distance metric between the newly generated sample and the -th sample. The distance metric is calculated based on the Euclidean distance; represents the sample perturbation update amount, which reflects the perturbation direction and amplitude.

[0090] Furthermore, the perturbation update amount is calculated based on the relative distance between samples and the superposition state of the synthesized samples after class balance, which is expressed as:

[0091] ,

[0092] In the formula, represents the perturbation expansion factor, which controls the perturbation intensity, , preferably, is set to 0.1; represents the optimized quantum state representation, which characterizes the superposition state of the synthesized samples after class balance.

[0093] S26. Implement sample class balance using a quantum state optimization mechanism. By calculating the distribution of samples in each class and adjusting the proportion of generated samples, dynamically balance the number of samples in each class in the saline-alkali soil dataset to avoid training bias caused by class imbalance. The samples in the boundary region of the saline-alkali soil data usually contain relatively key samples and are greatly affected by noise. Through the quantum state optimization mechanism, automatically adjust the direction and amplitude of the perturbation during sample generation to optimize the generation quality of the boundary region samples and improve the diversity and model accuracy of the training data. The calculation method is expressed as:

[0094] ,

[0095] In the formula, represents the optimization coefficient of the -th class, which controls the weight of the generated samples among different classes, is a positive integer; Indicates the number of samples in the th category; Indicates the total number of samples in the saline-alkali soil dataset; Indicates the total number of samples in the current saline-alkali soil dataset; Indicates the th sample's feature representation.

[0096] Furthermore, the optimization coefficient is calculated based on the importance of category samples, expressed as:

[0097] ,

[0098] In the formula, Indicates the importance coefficient of the th category, characterizing the relative importance of the samples in the th category during the expansion process; Indicates the importance coefficient of the th category, Indicates the sum of the importance coefficients of all categories, is a positive integer.

[0099] Furthermore, the importance coefficient is calculated based on the number of category samples, expressed as:

[0100] ,

[0101] In the formula, Indicates the number of samples in the th category, is a positive integer.

[0102] S27. The newly generated samples after updating are subject to quality assessment. The assessment criteria are based on the feature distance between the newly generated samples after updating and the original samples, in order to screen out effective samples and improve the quality of the saline-alkali soil dataset. For saline-alkali soil data, combined with other multi-dimensional data such as historical precipitation, evaporation, and dynamic monitoring data of the groundwater level, the newly generated samples can effectively represent the characteristics of real-world saline-alkali soil. The calculation method is expressed as:

[0103] ,

[0104] In the formula, Indicates the quality score of the newly generated samples after updating; Indicates the weight of the th sample in the quality assessment, adjusted according to its importance in the original saline-alkali soil dataset; Indicates the number of samples participating in the assessment.

[0105] Furthermore, the weight is adjusted according to the sample feature vector, and the calculation method is expressed as:

[0106] ,

[0107] wherein, represents the weight of the th sample, is a constant. Preferably, is set to 0.001.

[0108] Further, if the quality score of the newly generated sample after updating is greater than the preset threshold, the sample is retained; otherwise, the sample is discarded; preferably, the threshold is set to 0.8.

[0109] S28. Combine all the generated synthetic samples with the original saline-alkali soil data to form an expanded final dataset, denoted as:

[0110]

[0111] wherein, represents the expanded final dataset; respectively represent the 1st, 2nd, and th samples in the original saline-alkali soil dataset, and each sample corresponds to a feature representation; represent the 1st, 2nd, and th newly generated samples after updating, represents the number of synthetic samples.

[0112] S3. Preprocess the expanded final saline-alkali soil dataset to make it suitable for the training and analysis of machine learning models, and finally obtain the preprocessed saline-alkali soil data;

[0113] The preprocessing operations include data cleaning, missing data filling, outlier detection and correction, and standardization, so that the obtained preprocessed saline-alkali soil data is suitable for the training and analysis of machine learning models.

[0114] The preprocessed saline-alkali soil data is processed by a machine learning model to obtain a classification result; the machine learning model includes a trained autoencoder and a classifier model based on a feature importance decay function; due to the characteristics of high-dimensional non-linearity that saline-alkali soil data may have, for the problems of redundant feature interference and local structure loss during the feature dimensionality reduction process of high-dimensional saline-alkali soil data features, an autoencoder based on a feature importance decay function is used for dimensionality reduction, which effectively eliminates irrelevant features while retaining key discriminant information, and improves the analyzability and interpretability of the dimensionality-reduced saline-alkali soil data. The training process of the autoencoder based on the feature importance decay function is as follows:

[0115] S41. For the input saline-alkali soil data, an autoencoder based on a multi-layer neural network is constructed through the method of implicit state space modeling. In the implicit state space modeling, the saline-alkali soil data is mapped from the high-dimensional feature space to the low-dimensional implicit state space, and the mapping is realized by using the non-linear activation function and the initial weight matrix, achieving the preliminary compression and mapping of the saline-alkali soil data. The weight matrix is optimized by the manifold alignment regularization term to solve the collinearity problem of the high-dimensional saline-alkali soil data. The calculation method is expressed as:

[0116]

[0117] In the formula, represents the output feature of the autoencoder, is the saline-alkali soil data input to the autoencoder; is the weight matrix of the autoencoder, controlling the linear transformation input to the hidden space; is the bias term of the autoencoder, adjusting the output offset; is the ReLU activation function, realizing the non-linear mapping; is the manifold alignment regularization term, is the co-regularization term.

[0118] Furthermore, the manifold alignment regularization term optimizes the mapping accuracy by weighted adjustment of the relationship between features. This regularization constrains the weights by the distance between samples to maintain the manifold structure of the saline-alkali soil data stream. The calculation method is expressed as:

[0119] ,

[0120] In the formula, is the element in the i-th row and j-th column of the autoencoder weight matrix, where i is a positive integer and j is a positive integer; is the L2 norm; is the regularization coefficient, preferably set to 0.1; is the scale parameter, preferably set to 0.5; is the feature vector of the i-th sample; is the feature vector of the j-th sample.

[0121] S42. In the dimensionality reduction process, traditional dimensionality reduction methods may ignore the multi-scale characteristics of the saline-alkali soil data, that is, the saline-alkali soil data can be observed and processed from multiple levels and scales. The present invention adopts the multi-scale saline-alkali soil data co-regularization term to optimize the dimensionality reduction process by fusing the features of the saline-alkali soil data at different scales. The input saline-alkali soil data is decomposed into representations at multiple scales, and the co-regularization term is used to enhance the interaction between different scales. The calculation method is expressed as:

[0122] ,

[0123] In the formula, is the collaborative regularization term, represents the saline-alkali soil data representation with a scale of ; represents the saline-alkali soil data representation with a scale of ; is the first saline-alkali soil data scale, representing the first features of the saline-alkali soil data; is the second saline-alkali soil data scale, representing the first features of the saline-alkali soil data; is a parameter for controlling the similarity between scales. Preferably, is set to 0.2.

[0124] S43. Evaluate the contribution of each feature through the feature importance attenuation function. According to the contribution degree of the feature in the output result, attenuate the weight of the unimportant feature. The attenuation function dynamically adjusts the influence of the feature based on the relationship between the feature and the output, thereby optimizing the feature selection process and improving the dimensionality reduction effect. The calculation method of the attenuation coefficient of the feature is expressed as:

[0125] ,

[0126] In the formula, is the attenuation coefficient of the th feature, is the attenuation rate, is the weight between the th feature of the autoencoder and the th neuron, is the updated value of the attenuation coefficient of the i-th feature. Preferably, is set to 2.

[0127] Furthermore, refine the feature attenuation process through the gradient feedback mechanism. This mechanism adjusts the attenuation intensity through the weight amplitude and refines the feature influence evaluation in combination with the gradient direction, so that the unimportant features are more effectively cut. Use local gradient feedback to adjust the feature attenuation coefficient and enhance the accuracy of attenuation. The calculation method of the updated value of the feature attenuation coefficient is expressed as:

[0128] ,

[0129] In the formula, is the updated value of the attenuation coefficient of the i-th feature; is the gradient of the loss function of the autoencoder with respect to the input saline-alkali soil data; is the learning rate of the feature attenuation coefficient; is the connection weight between the k-th neuron of the autoencoder and the i-th feature; is the attenuation rate. Preferably, is set to 0.01, is set to 0.95.

[0130] S44. In dynamic evolution and adaptive training, a feedback-update-based strategy is adopted to update the weights of the autoencoder to optimize the evolution of the autoencoder network topology, adjust the learning intensity according to the inter-layer state difference, and improve the adaptability to complex saline-alkali soil data. The update calculation method is expressed as:

[0131]

[0132] In the formula, is the parameter update operation, is the weight matrix of the autoencoder, is the learning rate for the autoencoder weight update, is the th attenuation coefficient of the feature, is the loss function of the autoencoder; is the control coefficient. Preferably, is set to 0.2, is set to 0.05.

[0133] S45. Combine the global error and the local error to optimize the objective function, making the optimization process more comprehensive and accurate. This design, while maintaining the global reconstruction accuracy, forces similar samples to maintain a neighboring relationship in the latent space. The calculation method is expressed as:

[0134]

[0135] In the formula, is the loss function balance coefficient, is the global error of the autoencoder, is the local error of the autoencoder. Preferably, is set to 0.2.

[0136] Furthermore, the global error of the autoencoder is calculated as:

[0137] ,

[0138] In the formula, is the number of samples input to the autoencoder in the current batch, is the feature vector of the i-th sample, is the i-th reconstructed feature vector output by the decoder in the autoencoder.

[0139] Furthermore, the local error of the autoencoder is calculated as:

[0140] ,

[0141] In the formula, is the local weight coefficient, is the feature vector of the j-th sample, is the j-th reconstructed feature vector output by the decoder in the autoencoder. Preferably, is set to 0.2.

[0142] S46. The iteration termination condition is judged by the change in loss, avoiding overfitting while ensuring sufficient training of the model. Stop training when the following conditions are met:

[0143] ,

[0144] In the formula, represents the loss function of the autoencoder in the -th iteration, represents the loss function of the autoencoder in the -th iteration, is the convergence threshold. Preferably, is set to 1e-5.

[0145] Furthermore, the preprocessed saline-alkali soil data passes through the trained autoencoder based on the feature importance decay function to obtain the low-dimensional saline-alkali soil data after feature dimensionality reduction.

[0146] Furthermore, the low-dimensional saline-alkali soil data after feature dimensionality reduction is input into the classifier model for classification to obtain the classification result; the classifier model can be any one of random forest, support vector machine, decision tree, and logistic regression, and the classification result includes: : Extremely severe salinization (improvement requirement > 90%), : Severe salinization (improvement requirement 70% - 90%), : Moderate salinization (improvement requirement 50% - 70%), : Mild salinization (improvement requirement 30% - 50%) and : Non-saline-alkali soil (improvement requirement < 30%).

[0147] S5. Evaluate the improvement effect of the saline-alkali soil according to the classification result, and optimize the model according to the evaluation result.

[0148] Example 2

[0149] In this example, as Figure 2As shown, by comparing the classification accuracies of different oversampling methods, the improvement effect of this technology on the model classification performance compared with traditional oversampling methods is verified in extremely imbalanced class scenarios. The experimental results show that as the class imbalance ratio increases, the accuracies of traditional methods (such as random oversampling, Borderline-SMOTE) decrease significantly. Especially when the minority class samples are extremely scarce, the problem of homogenization of the generated samples makes it difficult for the model to capture boundary features. However, through quantum superposition state mapping and tunneling perturbation, this technology explores the non-linear correlations between samples in the feature space and generates more diverse samples that are closer to the real distribution. Therefore, this technology can still maintain stable classification accuracy at high imbalance ratios, indicating that the samples it generates effectively expand the feature coverage of the minority class and alleviate the problem of narrow feature space caused by linear interpolation in traditional methods.

[0150] Example 3

[0151] In this example, as Figure 3 shown, by evaluating the quality of the generated samples and comparing the distribution consistency between the samples generated by different methods and the real samples, the generation quality of this technology is verified. The experimental results show that the distributions of the samples generated by traditional methods show obvious deviations, and outliers are easily generated in areas with low data density, reflecting the problem of noise accumulation during the oversampling process. However, the kernel density curve of this technology highly coincides with the real samples, indicating that the generated samples not only retain the statistical characteristics of the original data but also suppress noise interference through quantum state optimization. Therefore, the quantum state optimization mechanism ensures that the generated samples follow the internal distribution law of the original data through dynamic evaluation of the local density and phase angle adjustment of the samples, solving the problem of distribution distortion caused by fixed interpolation strategies in traditional methods.

[0152] Example 4

[0153] In this example, as Figure 4 shown, by conducting noise robustness analysis and testing the stability of the algorithm in a noise-polluted environment, the effectiveness of the adaptive perturbation optimization strategy is verified. The experimental results show that as the noise ratio increases, the classification accuracy of traditional methods drops sharply, indicating that the generated samples are vulnerable to noise propagation, resulting in blurred decision boundaries. However, this technology introduces non-linear perturbations through the quantum tunneling effect and dynamically adjusts the perturbation direction in combination with the local density, enabling the generated samples to filter out noise interference while retaining key features. Therefore, this technology can still maintain a high accuracy in strong noise scenarios, proving that its perturbation optimization strategy can adaptively distinguish effective features from noise and enhance the robustness of the generated samples.

[0154] Example 5

[0155] In this example, as Figure 5As shown, by comparing the algorithm efficiencies, the computational efficiencies of different oversampling algorithms are evaluated to verify the effect of the optimized design of this technology in saving computing resources. The experimental results show that the time consumption of the traditional method increases approximately linearly with the growth of the sample size, reflecting the computational redundancy of its global interpolation strategy. While this technology significantly reduces the number of iterations for neighborhood search and interpolation optimization through locally density-driven dynamic sampling and quantum state parallel computing. Therefore, this technology can still maintain a low time overhead on large-scale datasets, indicating that it achieves a balance between computational efficiency and generation quality through the partition processing of the feature space and the parallelism of quantum state superposition.

[0156] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for evaluating the improvement effect of saline-alkali soil, characterized in that, It includes the following steps: S1. Obtain saline-alkali soil data and perform manual annotation on the saline-alkali soil data to obtain the original saline-alkali soil dataset; S2. Expand the original saline-alkali soil dataset through the trained SMOTE algorithm based on quantum superposition states to obtain the final expanded saline-alkali soil dataset; introduce nonlinear perturbations using the quantum tunneling effect in the SMOTE algorithm based on quantum superposition states; Specifically, the training process of the SMOTE algorithm based on quantum superposition states includes: S21. Perform quantum representation processing on each sample in the saline-alkali soil dataset to obtain a quantum superposition state representation ; S22. Based on the strategy of generating synthetic samples by the SMOTE algorithm, generate new synthetic samples through linear interpolation. At the same time, adopt multiple quantum state superpositions within the neighborhood of the original samples and combine the quantum tunneling effect to generate new candidate samples. The formula is expressed as follows: , Among them, represents a new candidate sample; represents the perturbation adjustment factor of the -th sample, which is used to control the perturbation amplitude of the -th sample; represents the local density of the -th saline-alkali soil data sample point; represents the local density of the -th saline-alkali soil data sample point; represents the feature representation of the -th sample; represents the feature representation of the -th sample; is a positive integer; is a positive integer; is a quantum superposition state representation, which characterizes the samples superimposed in the quantum state space; represents the interpolation coefficient, which controls the distance between the generated sample and the original sample, satisfying ; represents the non-linear perturbation introduced by using the quantum tunneling effect, represents the imaginary unit, represents the quantum state phase angle; according to the quantum tunneling effect, non-linear perturbation is introduced in the process of calculating the quantum state phase angle; S23. Adjust the direction and amplitude of the perturbation according to the local density and position of each sample. The calculation formula of the perturbation adjustment factor of the sample is expressed as follows: , Among them, represents the attenuation rate of the control disturbance amplitude with respect to the sample distance from the center point; is the L2 norm, the same as the Euclidean distance calculation method; represents the sample point of the global center of the saline-alkali soil dataset; represents the th local density of the sample, reflecting the sample density around the th sample; represents the maximum local density in the saline-alkali soil dataset; S24. Identify the boundary of the high-dimensional saline-alkali soil data space through local density measurement; S25. Update the new samples using the adaptive perturbation optimization strategy. The formula is expressed as follows: , Among them, represents the new sample generated after the update; represents the learning rate of the perturbation; represents the distance metric between the generated new sample and the th sample, represents the distance metric between the generated new sample and the th sample, and the distance metric is calculated based on the Euclidean distance; represents the sample perturbation update amount, reflecting the perturbation direction and amplitude; the perturbation update amount is calculated based on the relative distance between samples and the superposition state of the synthesized samples after class balancing, and the formula is as follows: , Among them, represents the perturbation expansion factor; represents the optimized quantum state representation, characterizing the superposition state of the synthesized samples after class balance; S26. Adopt a quantum state optimization mechanism to achieve sample class balance. By calculating the distribution of samples in each class and adjusting the proportion of generated samples, dynamically balance the number of samples in each class in the saline-alkali soil dataset. The calculation formula of the optimized quantum state representation is expressed as follows: , Among them, represents the optimization coefficient of the th category, which is a positive integer; represents the number of samples in the th category; represents the feature representation of the th sample; represents the total number of samples in the saline-alkali soil dataset; represents the current total number of samples in the saline-alkali soil dataset; S27. Calculate the feature distance between the newly generated samples after update and the original samples to obtain the quality score of the newly generated samples after update. The formula is expressed as follows: , Among them, represents the quality score of the new sample generated after the update, represents the number of samples participating in the evaluation; represents the weight of the th sample in the quality evaluation, which is adjusted according to its importance in the original saline-alkali soil dataset, and this weight is adjusted based on the sample feature vector. The formula is expressed as , is a constant, set to 0.001; if the quality score of the new sample generated after the update is greater than the preset threshold, then the sample is retained, otherwise the sample is discarded; S28. Combine the samples retained in step S27 with the original saline-alkali soil data to form the final expanded saline-alkali soil dataset; S3. Preprocess the final expanded saline-alkali soil dataset to make it suitable for the training and analysis of machine learning models, and finally obtain the preprocessed saline-alkali soil data; S4. The preprocessed saline-alkali soil data is processed through a machine learning model to obtain a classification result; the machine learning model includes a trained autoencoder and a classifier model based on the feature importance attenuation function; S5. Evaluate the improvement effect of saline-alkali soil according to the classification result and optimize the model according to the evaluation result.

2. The method for evaluating the improvement effect of saline-alkali soil according to claim 1, wherein, Step S1 specifically includes: The saline-alkali soil data includes soil pH value, conductivity, sodium adsorption ratio, total soluble salt content, percentage of clay content, organic matter content, exchangeable calcium-magnesium ratio, boron element concentration, annual precipitation anomaly value, temperature seasonal variation coefficient; The labeled categories include: extremely severe salinization, severe salinization, moderate salinization, mild salinization, and non-salinization.

3. The method for evaluating the improvement effect of saline-alkali soil according to claim 2, characterized in that, Step S3 specifically includes: The preprocessing operations include data cleaning, missing data filling, outlier detection and correction, and standardization to make the obtained preprocessed saline-alkali soil data suitable for the training and analysis of machine learning models.

4. The method for evaluating the improvement effect of saline-alkali soil according to claim 3, characterized in that, In step S4, the preprocessed saline-alkali soil data passes through a trained autoencoder based on the feature importance attenuation function to obtain low-dimensional saline-alkali soil data after feature dimensionality reduction, specifically including: Construct an autoencoder based on a feature importance attenuation function. The training process of the autoencoder based on the feature importance attenuation function includes: S41. Construct an autoencoder based on a multi-layer neural network through the method of implicit state space modeling. In implicit state space modeling, the preprocessed saline-alkali soil data is mapped from a high-dimensional feature space to a low-dimensional implicit state space, and the mapping is carried out using a non-linear activation function and an initial weight matrix, and then the weight matrix is optimized through a manifold alignment regularization term; S42. In the process of dimensionality reduction, adopt a multi-scale saline-alkali soil data collaborative regularization term to optimize the dimensionality reduction process by fusing the characteristics of saline-alkali soil data at different scales; S43. Evaluate the contribution of each feature through a feature importance attenuation function. According to the contribution degree of the feature in the output result, attenuate the weights of unimportant features. The attenuation function dynamically adjusts the influence of the feature according to the relationship between the feature and the output, so as to optimize the feature selection process; refine the feature attenuation process through a gradient feedback mechanism; S44. In dynamic evolution and adaptive training, adopt a strategy based on feedback update to update the weights of the autoencoder; S45. Combine the global error and the local error to optimize the objective function and force similar samples to maintain a neighboring relationship in the latent space; S46. The iteration termination condition is judged by the change in loss. Stop training when the following conditions are met: , Among them, represents the loss function of the autoencoder for the th iteration, represents the loss function of the autoencoder for the th iteration, is the convergence threshold.

5. The method for evaluating the improvement effect of saline-alkali soil according to claim 4, characterized in that, In step S4, the low-dimensional saline-alkali soil data obtained after feature dimensionality reduction is input into a classifier model for classification to obtain a classification result; the classifier model includes a random forest, a support vector machine, a decision tree, and a logistic regression model.

Citation Information

Patent Citations

  • Slope soil fertility prediction method

    CN119359475A