Hyperspectral modeling method for soil organic matter content based on SAE-GS-TabNet

Through the improved SAE-GS-TabNet model, combined with sparse autoencoder, particle swarm optimization algorithm and attention mechanism optimization TabNet model, the non-destructive detection, time-consuming and cost-effective problems of traditional soil organic matter measurement methods are solved, and efficient and accurate soil organic matter prediction is achieved, suitable for soil in different regions.

CN120234586APending Publication Date: 2025-07-01GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271760.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-09
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Traditional soil organic matter measurement methods have problems such as non-destructive testing, time-consuming, high cost and difficulty in real-time monitoring in the field. The high dimensionality and complexity of hyperspectral remote sensing technology pose great challenges to data processing and modeling.

Method used

The improved SAE-GS-TabNet model is adopted, combining sparse autoencoder (SAE), particle swarm optimization algorithm (GSPSO) and attention mechanism optimization TabNet model to establish an efficient soil organic matter prediction model, and achieve rapid prediction of soil organic matter content through hyperspectral data preprocessing and feature selection.

Benefits of technology

It significantly improves modeling accuracy and solves the multicollinear problem between various bands. It has the advantages of small workload, low cost, high accuracy and good reliability. It is suitable for different soils in different regions. It provides a simple soil organic matter prediction method with small data demand, high computing efficiency and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234586A_ABST
    Figure CN120234586A_ABST
Patent Text Reader

Abstract

The invention provides a spectral data quantitative analysis model based on SAE-GS-TabNet. The SAE-GS-TabNet model is formed by improving a sparse auto-encoder (SAE), a parameter optimization algorithm (GSPSO) and a TabNet network model. In the construction of the improved SAE-GS-TabNet model, the TabNet model obtains the highest prediction precision (R2 is equal to 0.942, RMSE is equal to 2.973 g.kg <-1 >, and RPD is equal to 4.159) in SOM content prediction models of four kinds of standard deep learning under Adam optimization. GSPSO is used for carrying out hyper-parameter optimization, a sparse auto-encoder (SAE) is used for extracting an SAE-GS-TabNet model with effective features, and higher prediction precision (R2 is equal to 0.954, RMSE is equal to 2.632 g.kg <-1 >, and RPD is equal to 4.699) is obtained. And the forest land soil SOM content can be quickly and effectively predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil spectral acquisition and analysis. Using this method, the soil organic matter content can be quickly determined, and the measurement accuracy can be improved. Background Art

[0002] Soil, as an important part of the earth's surface system, is not only the basis of agricultural production but also a key element for the healthy operation of the ecosystem. Soil quality directly determines the growth status of crops, the sustainability of agricultural production, and the stability of the ecological environment. Among the many properties of soil, soil organic matter (SOM) is one of the most core indicators. Soil organic matter is not only an important source of soil fertility but also indirectly regulates the soil's water-holding capacity, nutrient supply capacity, and microbial activity by affecting the physical, chemical, and biological properties of the soil. In addition, soil organic matter plays an important role in the global carbon cycle, and its dynamic changes have a profound impact on climate change. Therefore, accurately monitoring and predicting the soil organic matter content is of great significance for realizing precision agriculture, optimizing soil management, and coping with climate change.

[0003] Traditional methods for determining soil organic matter mainly rely on laboratory chemical analysis, such as wet chemical analysis methods (such as potassium dichromate oxidation method) and dry combustion methods (such as high-temperature combustion method). Although these methods have high precision, they have obvious limitations: First, they require the destruction of soil samples and cannot achieve non-destructive detection; second, the experimental process is time-consuming and costly, making it difficult to meet the needs of large-scale and high-frequency soil monitoring; finally, laboratory analysis usually requires professional equipment and technical personnel, which limits its application in on-site real-time monitoring. With the increasing global demand for soil data in agriculture and the ecological environment, developing a fast, efficient, and non-destructive method for detecting soil organic matter has become an urgent need in current research.

[0004] In recent years, the rapid development of remote sensing technology has provided new solutions for the rapid monitoring of soil properties. Among them, hyperspectral remote sensing technology has gradually become an important tool for soil science research because it can provide continuous and dense spectral information. Hyperspectral remote sensing obtains the surface reflection spectrum through sensors in multiple wavelength ranges such as visible light, near-infrared, and short-wave infrared, and can capture the fine spectral characteristics of the soil. These spectral characteristics are closely related to the physical, chemical, and biological properties of the soil, such as organic matter content, water content, and mineral composition. By analyzing the hyperspectral data of the soil, a quantitative relationship between soil properties and spectral characteristics can be established, thus realizing the rapid prediction of soil properties.

[0005] In addition, the preprocessing and feature selection of hyperspectral data are also key factors affecting the performance of the model. Spectral data is usually disturbed by noise, scattering effects, and baseline drift, and needs to be denoised and corrected through preprocessing methods (such as SG filtering, wavelet transform, multiplicative scatter correction, etc.). At the same time, the high dimensionality of hyperspectral data may lead to the "curse of dimensionality" problem, so key features need to be extracted and the data dimensionality reduced through feature selection or dimensionality reduction methods (such as principal component analysis, PCA). The selection and optimization of these preprocessing and feature selection methods have an important impact on the performance of the model.

[0006] In summary, the rapid and accurate prediction of soil organic matter is of great significance for precision agriculture, soil management, and environmental protection. Hyperspectral remote sensing technology provides a new solution for the rapid monitoring of soil properties, but its high dimensionality and complexity pose great challenges to data processing and modeling. The purpose of this study is to develop an efficient soil organic matter prediction model by combining hyperspectral remote sensing technology and deep learning methods, providing technical support for soil science research and practical applications. Summary of the Invention

[0007] In view of this, the present invention proposes an improved SAE-GS-TabNet model that adds a sparse autoencoder (SAE) and a particle swarm optimization algorithm (GSPSO) to alleviate the problems of gradient disappearance and low training efficiency that occur in traditional deep learning models during training.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] The technical solution for achieving the object of the present invention is: A hyperspectral modeling method for soil organic matter content based on SAE-GS-TabNet, comprising the following steps:

[0010] 1. Collection and processing of soil samples: Collect a number of samples, air-dry and grind the samples naturally, and divide the samples evenly into two parts;

[0011] 2. Determination of the organic matter content and spectral reflectance data of soil samples: Sieve the samples through a 0.2 nm soil sieve, then use potassium dichromate oxidation heating to measure the SOM content. Sieve the samples through a 0.149 nm soil sieve and use an asd fieldspec1 4 hi-res ground object spectrometer to obtain hyperspectral data, with the spectral range including the visible and near-infrared regions, i.e., the wavelength range of 350 - 2500 nm.

[0012] 3. Preprocess the soil spectral reflectance data in Step 2: that is, remove the edge bands with large noise in the spectral reflectance data and smooth the spectral reflectance data. Among them, the edge bands with large noise are the 350 - 399 nm and 2401 - 2500 nm bands, and the smoothing process is to perform Savitzky - Golay smoothing on the spectral reflectance data;

[0013] 4. Spectral reflectance data transformation: Perform transformation processing on the spectral reflectance including first - order differential, second - order differential, moving average filtering, standard normal variate, and multiplicative scatter correction;

[0014] 5. Establish a soil organic matter prediction model using machine learning algorithms and deep learning algorithms. Among them, the machine learning algorithms are partial least squares regression and support vector machine, and the deep learning algorithm is long short - term memory network. Four - fifths of the total number of samples are used as the training set, and one - fifth is the validation set. The plsr and dbo - svr models are implemented by calling the corresponding machine learning modules in the sklearn interface. The lstm model is established using the keras library in the pycharm software with the Python 3.8 language. Establish an inversion model between five types of spectral reflectance data, namely r, 1dr, 2dr, maf, and msc, and soil organic matter content. The initial model test is to perform data model fitting using origin 2021;

[0015] 6. Establish a hyperspectral prediction model: Select the optimal model in Step 5;

[0016] 7. Model accuracy evaluation: Use the coefficient of determination and root mean square error to evaluate the soil organic matter prediction model established in Step 5, and determine the accuracy, stability, and prediction performance of the soil organic matter prediction model.

[0017] Compared with the prior art, the advantages and beneficial effects of this technical solution are as follows: Compared with modeling using the attention mechanism to optimize the traditional residual network model, this technical solution can significantly improve the modeling accuracy, solve problems such as multicollinearity between bands, and has the advantages of small workload, low cost, high accuracy, and high reliability. The generalization of the model is high and it is applicable to different soils in different regions; This technical solution provides a soil organic matter prediction method with a simple model, small data volume requirements, high operation efficiency, high accuracy, and good prediction performance.

[0018] Through soil sample collection and processing, soil organic matter determination, spectral reflectance data determination, spectral reflectance data transformation, establishment of machine learning algorithms and deep learning algorithms models, establishment of hyperspectral prediction models, and evaluation indicators, this technical solution can predict in real - time, quickly, and accurately compared with machine learning methods and traditional deep learning methods, and has good practical application value.

[0019] This method can be widely applied to engineering practice, providing a basis for the subsequent management and utilization of land resources. Brief Description of the Drawings

[0020] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably below in conjunction with the drawings, where:

[0021] Figure 1 It is the structure diagram of SAE-GS-TabNet

[0022] Figure 2 It is the architecture diagram of the sparse autoencoder model.

[0023] Figure 3 It is the overall structure diagram of TabNet.

[0024] Figure 4 It is the structure diagram of the feature transformer.

[0025] Figure 5 It is the structure diagram of the attention transformer.

[0026] Figure 6 It is the convergence diagram of the prediction accuracy curve of the modeled SOM content.

[0027] Figure 7 It is the scatter plot of the prediction accuracy of the modeled SOM content. Detailed Embodiments

[0028] The following illustrates the embodiments of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention schematically, and the following embodiments and the features in the embodiments can be combined with each other without conflict.

[0029] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention: in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged, or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted. The same or similar reference numerals in the drawings of the embodiments of the present invention correspond to the same or similar components.

[0030] The present invention provides a deep learning method for predicting soil organic matter based on near-infrared spectral data, aiming to improve the prediction accuracy. By improving the TabNet model, it achieves maintaining the optimal energy consumption and total cost, including the following steps:

[0031] Step 1: Improved model of sparse autoencoder

[0032] The sparse autoencoder is an unsupervised learning method that can extract high-level latent features from the original data. These features are more representative and discriminative than the original data. Through the non-linear activation functions ReLU or Sigmoid, the SAE can capture the complex non-linear relationships in the input data while enhancing the robustness of the model to noisy data. In soil data analysis, the original spectral data usually has high-dimensional and non-linear characteristics. The SAE can reveal the relationship between soil components and their spectral responses by learning the latent patterns of these features. Since the original spectral data may contain a large amount of redundant information and noise, the SAE compresses the high-dimensional data into a low-dimensional latent space through the encoder part to achieve feature dimensionality reduction, and reconstructs the data through the decoder part, weakening the influence of noise while retaining the main information.

[0033] The structure of the sparse autoencoder usually includes an input layer, an encoding layer, a hidden layer, a decoding layer, and an output layer. The encoder consists of the input layer and the hidden layer, which is used to compress the original input data, such as the long vector of an image or a signal, into the feature representation space; while the decoder reconstructs the compressed data into a long vector similar to the input through the hidden layer and the output layer. Specifically, the original input data x is transformed into a new representative feature h, which is further processed by the output layer to generate a similar reconstructed data. In addition, the weights W and the biases b are determined during the training process, and backpropagation is also used to reduce the loss of the SAE. The size of the hidden layer determines the feature dimensions that the model can learn and the feature compression rate. By minimizing the error between the original input data and the reconstructed data, the hidden layer can extract effective features, thus providing a more discriminative representation for subsequent analysis.

[0034] The sparse autoencoder plays an important role in feature extraction and feature weighting in the model. It preprocesses the features and guides the sparsity of the dataset divided by data processing, providing a cleaner and more valuable input feature for TabNet, enabling the subsequent TabNet to focus on important features, reducing the computational complexity, and thus improving the overall modeling effect.

[0035] During the training process of the autoencoder (AE), the choice of the training optimizer directly affects the fitting ability and training efficiency of the model. To verify the impact of the optimizer on the model performance, we selected four common optimizers: Adma, Admax, RAdma, and NAdma, and ran 1500 training iterations for each optimizer to observe their performance differences. By comparing the performance of each optimizer on the loss function, the optimal optimizer was finally selected for model training. The results show that the Adma optimizer has the fastest convergence speed, the Admax and NAdma optimizers have similar convergence speeds, and the RAdma optimizer has the slowest convergence speed. As the number of training rounds increases, the convergence degrees of these four optimizers are relatively close. After comparison, the convergence degree of the NAdma optimizer is slightly lower, and the convergence curves of the Admax and RAdma optimizers basically coincide in the later stage. The Adam optimizer performs the best among the four optimizers.

[0036] Step 2: Improving the model with the GSPSO optimization algorithm

[0037] To avoid the influence of random variables on this comparative experiment, the common hyperparameters of these models were subjected to GSPSO (Guided Search Particle Swarm Optimization), which is an algorithm improved on the basis of the standard Particle Swarm Optimization (PSO). It mainly enhances the global search ability and local optimization ability by introducing a guided search mechanism or combining gradient information, making it more suitable for complex or high-dimensional optimization problems.

[0038] The ordinary PSO algorithm is simple and easy to use, but it has limitations such as particles being easily trapped in local optima and the search efficiency of particles may be insufficient in high-dimensional spaces or for complex functions. To overcome this limitation, the GSPSO algorithm introduces a guiding direction during each particle update to help the particles approach the global optimal solution faster. It introduces gradient-assisted optimization, where the gradient term accelerates convergence when the particles are close to the global optimal region, but weakens it during global exploration to maintain randomness, and uses dynamic adjustment of weights to achieve a balance between exploration and exploitation. The main parameter optimization process of the GSPSO algorithm is subdivided into the following five steps:

[0039] (1) Initialize the relevant parameters of the algorithm, including the definition of the search space, the initialization of the particle population, and the initialization of the individual optimal and global optimal.

[0040] (2) Calculate the fitness of each particle's current position to update the individual optimal position and the global optimal position.

[0041] (3) Through the guided search mechanism, generate the guiding direction and dynamically adjust the guiding factor.

[0042] (4) Update the velocity and position of the particles. For particles that go out of bounds, boundary handling such as reflection or re-initialization is adopted.

[0043] (5) Convergence determination. Check the termination condition and record the global optimal position.

[0044] The specific parameter settings in step (1) include the upper and lower bounds of each dimension , clarify the objective function , which is used to evaluate the fitness of the particles, and initialize the initial positions of particles and the velocity of each particle , and the individual optimal position of each particle .

[0045] For the above step (3), the guiding direction is generated by additional gradient information or problem prior knowledge. Common practices include calculating the gradient of the objective function at the current position , and using it as the direction guidance, as well as designing heuristic rules according to specific problems to directly determine the high-quality search area. To avoid the excessive influence of the guided search, the guiding factor is dynamically adjusted, and the calculation formula is as follows:

[0046]

[0047] In Equation 5-1, is the maximum number of iterations, is the current iteration number

[0048] In step (4), the update of the particle velocity and position is the key step of the particle swarm optimization algorithm. By combining the inertia weight and adding the guiding direction, the particles can move in the search space and approach the global optimal solution. The particle velocity update formula is:

[0049]

[0050] In Equation 5.x, is the inertia term, which is used to maintain the moving inertia of the particles and prevent them from stagnating quickly, is the inertia weight, which controls the influence of the current velocity of the particle on the next velocity, is the particle at the th iteration. is the individual cognitive term, is the individual optimal position of the particle , is the particle 's current position, is the individual acceleration factor, which is used to weigh the degree of dependence of the particle on its own experience, is a random number in [0, 1]. is a group collaboration term that makes the particle move towards the global optimal position of the group close to is the global optimal position is the global acceleration factor, which weighs the degree of dependence of the particle on group collaboration. It is a guiding term uses external information to accelerate the convergence of the particle. According to the updated velocity the position of the particle is updated according to the following formula:

[0051]

[0052] Step Three: TabNet Network Structure

[0053] TabNet is a deep learning model for tabular data. It performs feature selection and prediction on the input data in a multi-step additive manner. The structure design takes into account both efficiency and interpretability. Its core idea is to select important features through an attention mechanism in each decision step and perform deep transformation on them to optimize the prediction performance. The overall network structure of TabNet is shown in the attached figure. The overall architecture of TabNet consists of multiple consecutive decision steps. Each step filters and transforms the input features, generates intermediate results, and accumulates them into the final output. The operations of each step are interrelated, gradually improving the model's understanding and representation ability of the data. The input of TabNet is a B×D matrix, where B is the number of samples and D is the feature dimension. This design is suitable for high-dimensional tabular data and allows the model to process all samples and features in parallel. To eliminate the inconsistent distribution of input features, all input data will first pass through a batch normalization layer (BN). This process can improve the model's convergence speed and reduce numerical instability during training, especially when dealing with features with a large numerical range span. The multi-step additive design of TabNet not only improves the model's learning ability but also makes the feature selection and transformation process of each decision step highly modular and interpretable.

[0054] Step Four: Feature Transformer

[0055] The feature transformer is a key component of TabNet, responsible for mapping the input features into new representations, and at the same time supporting the separation of general and specific features to balance the stability and flexibility of the model. The feature transformer consists of a shared feature transformer and a specific feature transformer, and the design of residual connection is also considered in the connection between the two parts.

[0056] The shared feature transformer is used to extract general features related to the overall task and shared across all decision-making steps. It consists of two sub-modules, each containing a fully connected layer (FC), batch normalization (BN), and gated linear unit (GLU). The gated linear unit can dynamically adjust the importance of features by controlling the passing and suppression of input information. By sharing parameters, the model size is reduced, the stability of feature extraction is improved, and the repeated training of similar representations in decision-making steps is avoided.

[0057] The specific feature transformer is independently designed for each decision-making step and is responsible for extracting specific features highly relevant to the current step. Similar to the shared part, it contains two FC+BN+GLU sub-modules and uses residual connection to improve the stability of training. The independent feature transformer provides flexibility for each step, enabling the model to capture more refined feature representations and thus enhancing the differentiation ability between decision-making steps.

[0058] Step Five: Attention Transformer

[0059] The attention transformer is another key component of TabNet, which is used to dynamically screen features and determine the features to be focused on in each decision-making step. Its design introduces the concepts of sparsity and prior knowledge to achieve efficient and interpretable feature selection. The attention transformer consists of a fully connected layer (FC), batch normalization (BN), prior scales, and SparseMax. The fully connected layer and batch normalization perform linear combination and normalization on the input features to generate the basic representation for feature selection. The prior scales provide information about the importance of features in each decision-making step, taking the feature selection result of the previous decision-making step as input to enhance the continuity and rationality of feature selection. SparseMax normalizes the weights of each feature through a sparsifying activation function to ensure that the model only focuses on a few key features, thereby improving the computational efficiency and interpretability of the model. During model training, the attention transformer first calculates the importance weights of each feature based on the current input, and then combines the prior information and specific sparsifying operations to select the few features most valuable for the current decision. The design of the attention transformer realizes automation and dynamization in feature selection, enabling TabNet to efficiently process spectral data and having a clear decision-making logic.

Claims

1. A modeling method for predicting soil organic matter content based on hyperspectral SAE-GS-TabNet, characterized in that: The method comprises the following steps: (1) Improved TabNet model based on sparse autoencoder: Sparse autoencoder is an unsupervised learning method that can extract high-level latent features from raw data. These features are more representative and discriminative than the raw data. (2) GSPSO optimization algorithm optimizes model hyperparameters: GSPSO (Guided Search Particle Swarm Optimization) is an algorithm improved on the basis of the standard particle swarm optimization algorithm (PSO). It mainly enhances the global search capability and local optimization capability by introducing a guided search mechanism or combining gradient information, making it more suitable for complex or high-dimensional optimization problems. (3) TabNet-based neural network modeling method: TabNet is a deep learning model for tabular data. It selects and predicts features of input data in a multi-step additive manner. Its structural design takes into account both efficiency and interpretability.

2. The method according to claim 1, characterized in that The channel attention mechanism described in step (1) is an unsupervised learning method that can extract high-level potential features from raw data. These features are more representative and discriminative than the raw data. Through the nonlinear activation function ReLU or Sigmoid, SAE can capture the complex nonlinear relationship in the input data and enhance the robustness of the model to noisy data. In soil data analysis, the original spectral data usually has high-dimensional and nonlinear characteristics. SAE can reveal the relationship between soil components and their spectral responses by learning the potential patterns of these characteristics. Since the original spectral data may contain a lot of redundant information and noise, SAE compresses the high-dimensional data into a low-dimensional latent space through the encoder part to achieve feature dimensionality reduction, and reconstructs the data through the decoder part to retain the main information while weakening the impact of noise. The structure of a sparse autoencoder usually includes an input layer, an encoding layer, a hidden layer, a decoding layer, and an output layer. The encoder consists of an input layer and a hidden layer, and is used to compress the original input data, such as a long vector of an image or signal, into a feature representation space. The decoder reconstructs the compressed data into a long vector similar to the input through the hidden layer and the output layer. Specifically, the input raw data x is converted into a new representative feature h, which is further processed by the output layer to generate similar reconstructed data. In addition, the weight W and bias b are determined during the training process, and back propagation is also used to reduce the loss of SAE. The size of the hidden layer determines the feature dimension and feature compression rate that the model can learn. By minimizing the error between the original input data and the reconstructed data, the hidden layer can extract effective features, thereby providing a more discriminative representation for subsequent analysis. The sparse autoencoder plays an important role in feature extraction and feature weighting in the model. It performs feature preprocessing and sparsity guidance on the data set that has been processed and divided, providing TabNet with purer and more valuable input features, allowing the subsequent TabNet to focus on important features, reducing computational complexity, and thus improving the overall modeling effect.

3. The method according to claim 1, characterized in that The channel attention mechanism described in step (2) is an algorithm improved on the basis of the standard particle swarm optimization algorithm (PSO). It mainly enhances the global search capability and local optimization capability by introducing a guided search mechanism or combining gradient information, making it more suitable for complex or high-dimensional optimization problems. The ordinary PSO algorithm is simple and easy to use, but it faces the limitations that particles are prone to fall into local optimality and the efficiency of example search may be insufficient in high-order spatial or complex functions. In order to overcome this limitation, the GSPSO algorithm introduces a guiding direction each time an example is updated to help particles approach the global optimal solution faster, introduces gradient-assisted optimization, and the gradient term accelerates convergence when the particle approaches the global optimal area, but is weakened during global exploration to maintain randomness, and uses dynamic adjustment weights to achieve a balance between exploration and development. The main parameter optimization process of the GSPSO algorithm is subdivided into the following five steps:

1. Initialize the relevant parameters of the algorithm, including the definition of the search space, particle population initialization, individual optimal and global optimal initialization; (ii) For each particle’s current position, calculate its fitness to update the individual optimal position and the global optimal position; (iii) Guide direction generation and dynamically adjust guiding factors through guided search mechanisms; (iv) Update the velocity and position of particles. For particles that cross the boundary, use boundary processing such as reflection or reinitialization. (V) Convergence determination, checking termination conditions, and recording the global optimal position; The specific parameter settings in step (1) include the upper and lower bounds of each dimension , clarify the objective function , used to evaluate the fitness of particles and initialize The initial position of the particle and the velocity of each particle , the individual optimal position of each particle ; For the above step (iii), guide direction Generated by additional gradient information or prior knowledge of the problem. Common practices include calculating the objective function gradient at the current position. , and use it as a direction guide, and design heuristic rules according to specific problems to directly determine the high-quality search area; to avoid excessive influence of guided search, dynamically adjust the guiding factor , the calculation formula is as follows: ; In the formula, is the maximum number of iterations, is the current iteration number; In step (iv), particle velocity and position update is the key step of the particle swarm optimization algorithm. By combining the inertia weight and the guiding direction, the particles can move in the search space and approach the global optimal solution. The particle velocity update formula is: ; In the formula, is the inertia term, which is used to maintain the moving inertia of particles and prevent them from stagnating quickly. is the inertia weight, which controls the influence of the current particle speed on the next speed. It is a particle In the The speed in iterations; is an individual cognitive item, It is a particle The individual's highest priority position, It is a particle Current location, is the individual acceleration factor, which is used to weigh the particle's dependence on its own experience. is a random number in [0,1]; is the group collaboration term, which allows particles to move to the global optimal position of the group near, is the global optimal position, is the group acceleration factor, which weighs the particle's dependence on group cooperation; is the guidance term, Use external information to accelerate the convergence of particles; according to the updated speed , the position of the particle Update according to the following formula: 。 4. The method according to claim 1, characterized in that: The channel attention mechanism described in step (3) is a deep learning model for tabular data, which performs feature selection and prediction on input data in a multi-step addition manner, and its structural design takes into account both efficiency and interpretability; The core idea is to select important features through the attention mechanism in each decision step and perform deep transformation on them to optimize the prediction performance. The overall architecture of TabNet consists of multiple consecutive decision steps. Each step screens and transforms the input features, generates intermediate results, and accumulates them into the final output. The operations of each step are interrelated, gradually improving the model's understanding and representation capabilities of the data. The input of TabNet is a B×D matrix, where B is the number of samples and D is the feature dimension. This design is suitable for high-dimensional tabular data and allows the model to process all samples and features in parallel. In order to eliminate the distribution inconsistency of input features, all input data will first pass through the batch normalization layer (BN). This process can improve the convergence speed of the model and reduce numerical instability during training, especially when processing features with a large numerical range. TabNet's multi-step additive design not only improves the learning ability of the model, but also makes the feature selection and transformation process of each decision step highly modular and interpretable. The step (3) includes the following steps: 4.1 Feature Converter: The feature converter is a key component of TabNet, responsible for mapping input features into new representations while supporting the separation of general and specific features; 4.2 Shared Feature Transformer: The shared feature transformer is used to extract common features related to the task as a whole; 4.3 Attention Converter: The attention converter is another key component of TabNet, which is used to dynamically screen features and determine the features that need to be paid attention to in each decision step.

5. The method according to claim 4, characterized in that The feature converter in step 4.1 is a key component of TabNet, which is responsible for mapping the input features into new representations and supporting the separation of general and specific features in order to balance the stability and flexibility of the model. The feature converter consists of a shared feature converter and a specific feature converter. The design of the residual connection is also taken into account in the connection between the two parts.

6. The method according to claim 4, characterized in that The shared feature converter in step 4.2 is used to extract common features related to the task as a whole, which are shared in all decision steps. It consists of two submodules, each of which contains a fully connected layer (FC), batch normalization (BN) and a gated linear unit (GLU). The gated linear unit can dynamically adjust the importance of features by controlling the passage and suppression of input information. It reduces the model size by sharing parameters, improves the stability of feature extraction, and avoids repeated training of similar representations in decision steps. The specific feature converter is designed independently for each decision step and is responsible for extracting specific features that are highly relevant to the current step. Like the shared part, it contains two FC+BN+GLU submodules and uses residual connections to improve training stability. Independent feature transformers provide flexibility for each step, allowing the model to capture more refined feature representations, thereby improving the differentiation ability between decision steps.

7. The method according to claim 4, characterized in that The attention converter in step 4.3 is another key component of TabNet, which is used to dynamically screen features and determine the features to be focused on in each decision step; Its design introduces the concepts of sparsity and prior knowledge to achieve efficient and interpretable feature selection; the attention converter consists of four parts: fully connected layer (FC), batch normalization (BN), prior scales and SparseMax; the fully connected layer and batch normalization linearly combine and normalize the input features to generate basic representations for feature selection; the prior scale provides information about the importance of features in each decision step, combined with the feature selection results of the previous decision step as input, to enhance the continuity and rationality of feature selection; SparseMax normalizes the weight of each feature through a sparse activation function to ensure that the model only focuses on a few key features, thereby improving the computational efficiency and interpretability of the model; during model training, the attention converter first calculates the importance weight of each feature based on the current input, and then structures prior information and specific sparsification operations to select a few features that are most valuable to the current decision; The design of the attention transformer realizes automation and dynamics in feature selection, enabling TabNet to process spectral data efficiently with clear decision logic.