Rock hyperspectral intelligent classification method based on Transform self-attention mechanism

By adopting a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism, the problem of CNN models struggling to model long-distance feature correlations and environmental interference in rock classification is solved, achieving fast, accurate, and non-destructive rock identification and improving the robustness and automation level of the model.

CN121962702APending Publication Date: 2026-05-01XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
Filing Date
2025-12-17
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing CNN-based rock hyperspectral classification methods are difficult to efficiently and directly model the correlation of ultra-long distance, non-local features, and the models lack robustness, are easily affected by environmental factors, and have poor generalization ability.

Method used

A rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism is adopted. By constructing a data augmentation strategy pool, including basic physical simulation, sample mixing and physical mechanism perturbation simulation transformation strategies, the spectral Transformer model is trained to capture long-range feature dependencies in spectral data. Different water-bearing states are simulated during the model training stage to improve the robustness of the model.

Benefits of technology

It achieves end-to-end intelligent classification of spectral data, reduces reliance on professional knowledge, improves the level of automation and robustness of classification, and can quickly and accurately identify rock types, making it suitable for large-scale, fast, and non-destructive rock identification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962702A_ABST
    Figure CN121962702A_ABST
Patent Text Reader

Abstract

The invention provides a rock hyperspectral intelligent classification method based on a Transform self-attention mechanism. The rock hyperspectral intelligent classification method based on the Transform self-attention mechanism is used for solving the technical problem that an existing rock hyperspectral classification method based on a CNN is difficult to efficiently and directly carry out modeling on ultra-long-distance and non-local feature correlation. According to the rock hyperspectral intelligent classification method based on the Transform self-attention mechanism provided by the invention, on-line enhancement is carried out on each piece of preprocessed spectral data by randomly selecting a transformation strategy in the data enhancement strategy pool, a large number of high-quality virtual training samples can be generated on line on the premise of not introducing artifacts, and the classification efficiency is improved. The over-fitting problem in small sample training is effectively avoided, so that the optimization of the spectrum Transform model is realized; and meanwhile, deep mining is performed through the optimized spectrum Transformer model, and a characteristic dependency relationship crossing a long-distance wave band in spectrum data is established, so that the limitation of traditional machine learning and a CNN model on characteristic extraction is effectively overcome, and end-to-end intelligent classification from spectrum input to category output is realized.
Need to check novelty before this filing date? Find Prior Art

Description

A Rock Hyperspectral Intelligent Classification Method Based on Transformer Self-Attention Mechanism Technical Field

[0001] This invention relates to rock identification and classification methods, and more particularly to a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism. Background Technology

[0002] Accurate identification and classification of rock types is fundamental to geological surveys, mineral resource exploration, engineering geological evaluation, and disaster early warning. Traditional geological methods, such as hand specimen identification, thin section microscopy, X-ray diffraction (XRD), and X-ray fluorescence (XRF), are the earliest methods for rock identification and classification. While they can provide accurate information on lithology and mineral composition, they generally suffer from problems such as complex procedures, long cycles, high costs, and destructive effects on samples. Furthermore, they are highly dependent on the experience of professionals and cannot meet the needs of large-scale, rapid, in-situ exploration.

[0003] As a fast and non-destructive alternative, hyperspectral analysis technology identifies and classifies rocks by acquiring their unique reflectance spectral "fingerprints." Early spectral analysis often employed classic machine learning algorithms such as Support Vector Machines (SVM) and Random Forests (RF). However, these algorithms face inherent bottlenecks when processing high-dimensional, nonlinear spectral data: 1) Limited feature extraction capabilities, relying heavily on shallow learning or manually designed features, making it difficult to capture the deep, nonlinear relationships in spectral curves determined by mineral composition, crystal structure, and physical state. Furthermore, they are poor at handling common problems in spectral data such as high-dimensional sparsity and feature collinearity. 2) Insufficient model robustness, as spectral signals are highly susceptible to environmental factors such as sample water content, surface roughness, and measurement illumination. Especially in regions with strong moisture absorption, the signal undergoes severe distortion, leading to a sharp decline in the performance of traditional models. 3) Strong data dependence; with limited field geological samples, models are prone to overfitting and exhibit poor generalization ability.

[0004] In recent years, deep learning methods, represented by Convolutional Neural Networks (CNNs), have been introduced into spectral analysis, leading to CNN-based rock hyperspectral classification methods. CNNs, through their fixed local receptive fields, excel at extracting morphological features such as local absorption and reflection in the spectrum, thus improving classification accuracy to some extent. However, this characteristic of CNNs also limits their ability to capture long-range spectral dependencies. Petrological spectral identification practices show that the diagnostic characteristics of rocks are often composed of a combination of features from multiple widely separated bands. For example, the naming of a specific granite may simultaneously depend on its composition of iron ions (…) in the visible light band (e.g., 400-700 nm). The broad and gentle absorption characteristics resulting from electronic transitions, and the absorption of clay minerals (such as kaolinite) in the short-wave infrared band (such as around 2200 nm). The sharp absorption valleys generated by bond vibrations. These two key features are far apart in the spectral sequence, and the CNN architecture is inherently difficult to efficiently and directly model such long-distance, non-local feature correlations. Summary of the Invention

[0005] The purpose of this invention is to solve the technical problem that existing CNN-based rock hyperspectral classification methods are difficult to efficiently and directly model the correlation of ultra-long distance, non-local features, and to provide a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism.

[0006] To achieve the above objectives, the technical solution provided by this invention is as follows: A rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism, characterized by the following steps: Step 1, constructing a training dataset and a validation dataset; acquiring the original spectral data of different rock samples, and sequentially performing spectral denoising and data standardization on multiple original spectral data to obtain multiple preprocessed spectral data; constructing a training dataset and a validation dataset based on the multiple preprocessed spectral data; Step 2, augmenting the training dataset; constructing a data augmentation strategy pool, which includes transformation strategies based on basic physical simulation, regularization transformation strategies based on sample mixing, and perturbation simulation transformation strategies based on physical mechanisms; randomly selecting transformation strategies from the data augmentation strategy pool to perform online augmentation on each preprocessed spectral data in the training dataset to generate augmented spectral data; using the augmented spectral data as augmented data for the training dataset; Step 3, ... Step 4: Training the Spectral Transformer Model. The reshaped spectral sequences from the training and validation datasets are fed into the Spectral Transformer Model for training. During training, a gradient descent-based optimizer and cross-entropy loss function are used, and the model parameters are iteratively updated using the backpropagation algorithm. The validation dataset is used to verify the model training results, and the model with the best validation performance is selected as the trained Spectral Transformer Model. Step 5: Spectral denoising and data standardization are performed on the spectral data of the rocks to be classified. The processed spectral data are then reshaped into spectral sequences suitable for the Spectral Transformer Model's input according to the spectral curves. The reshaped spectral sequences are input into the trained Spectral Transformer Model, and through one forward propagation, the rock category prediction result is directly output, completing the rock type identification and classification.

[0007] Furthermore, in step 2, the transformation strategy of the basic physical simulation includes Gaussian noise, random intensity scaling, random baseline shift, and random pruning and scaling; the regularization transformation strategy based on sample mixing includes Mixup and Intra-class Mixup; and the perturbation simulation transformation strategy based on physical mechanism is the moisture effect simulation transformation.

[0008] Furthermore, in step 4, the gradient descent-based optimizer is AdamW, SGD, or RMSprop.

[0009] Further, in step 4, the spectral Transformer model includes an input processing module, an encoder module, and a classification decision module; the input processing module includes a slicing layer, a linear projection layer, and a position encoding layer connected in sequence from input to output; the input of the slicing layer is used to receive the spectral sequence and segment the spectral sequence into multiple spectral slices; the linear projection layer is used to map the features of each spectral slice to a high-dimensional space to form a high-dimensional embedding vector; the position encoding layer is used to add position information to the embedding vector sequence; the encoder module includes N stacked encoding blocks, where N≥1; the encoding block includes a multi-head self-attention layer, a first normalization layer, a feedforward neural network layer, and a second normalization layer connected in sequence from input to output; the multi-head self-attention layer in the first encoding block... The input of the attention layer is connected to the output of the position encoding layer; the multi-head self-attention layer is used to compute the correlation weights between multiple spectral slices in parallel to capture global spectral features; the feedforward neural network layer is used to perform nonlinear transformations on the input data; the first normalization layer and the second normalization layer are used to normalize the corresponding input data; the second normalization layer in the Nth encoding block outputs the sequence features corresponding to the extracted multiple spectral slices; the classification decision module includes a pooling layer and a classification head connected in sequence according to input and output; the input of the pooling layer is connected to the output of the second normalization layer in the Nth encoding block, and is used to receive the sequence features output by the encoder module and aggregate them into a global feature vector of fixed dimensions; the classification head is used to output the rock category prediction result based on the global feature vector.

[0010] Furthermore, in step 4, N=4; the number of heads in the multi-head self-attention layer is 8.

[0011] Furthermore, in step 4, the position encoding layer adds position information using a sine / cosine position encoding method.

[0012] Further, step 1 specifically comprises: Step 1.1, using a hyperspectral analyzer to collect spectral reflectance data of different rock samples in the visible-near-infrared electromagnetic spectrum range, forming multiple original spectral data containing different wavelength points; Step 1.2, using a smoothing filter to smooth the multiple original spectral data respectively, obtaining multiple denoised spectral data; Step 1.3, performing standardization processing on the multiple denoised spectral data respectively, obtaining multiple preprocessed spectral data; Step 1.4, constructing training datasets and validation datasets based on the multiple preprocessed spectral data respectively.

[0013] Furthermore, in step 5, the spectral data of the rocks to be classified are subjected to spectral denoising and data standardization in sequence as follows: First, a smoothing filter is used to smooth the spectral data of the rocks to be classified to obtain the denoised spectral data of the rocks to be classified; then, the denoised spectral data of the rocks to be classified is standardized to obtain the preprocessed spectral data of the rocks to be classified.

[0014] Further, in steps 1.2 and 5, the smoothing filter is a Savitzky-Golay filter, a moving average filter, or a Gaussian filter; in steps 1.3 and 5, the normalization process uses Z-Score normalization.

[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: 1. This invention provides a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism. By constructing a data augmentation strategy pool, the transformation strategy in the data augmentation strategy pool is randomly selected to augment each preprocessed spectral data in the training dataset online. This can generate a large number of high-quality virtual training samples with diverse shapes and consistent features online without introducing artifacts, effectively avoiding the overfitting problem in small sample training, thereby optimizing the spectral Transformer model. At the same time, the optimized spectral Transformer model deeply mines and establishes feature dependencies across long-distance bands in the spectral data, effectively overcoming the limitations of traditional machine learning and CNN models in feature extraction. It realizes end-to-end intelligent classification from spectral input to category output. The entire process does not require manual design and extraction of complex features, greatly reducing the dependence on the professional geological knowledge of operators and improving the automation level and objectivity of rock classification work.

[0016] 2. This invention provides a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism, which directly addresses the long-standing technical pain point of decreased accuracy in the field of spectral analysis due to moisture interference. It innovatively introduces a perturbation simulation transformation strategy based on physical mechanisms, namely moisture effect simulation transformation, into the data augmentation strategy pool. This allows the spectral Transformer model to learn a large number of virtual spectral samples from different water content states, from dry to saturated, during the model training phase. This enables the spectral Transformer model to capture the severe signal distortion caused by moisture in specific wavelengths (such as around 1400nm and 1900nm), thereby accurately decoupling and identifying the masked intrinsic spectral features of rocks. Its classification performance is not significantly reduced compared to dry conditions, exhibiting extremely strong and designable environmental robustness.

[0017] 3. This invention presents a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism. It uses a spectral Transformer model for classification prediction. Once the spectral Transformer model is trained, for new unknown samples, only one forward propagation is needed to quickly obtain the classification result. The single-sample prediction time is in the millisecond range, which can meet the high timeliness requirements of scenarios such as rapid on-site exploration. At the same time, this invention inherits the advantage of being completely non-destructive of spectral analysis technology, has strong universality, and can be easily transferred to hyperspectral intelligent classification tasks of soil, vegetation, or other materials.

[0018] 4. The present invention provides a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism. By setting a cutting layer in the spectral Transformer model, the one-dimensional spectral sequence is cleverly transformed into discrete serialized words, i.e., multiple spectral slices, thus successfully adapting the spectral Transformer model to spectral analysis.

[0019] 5. This invention provides a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism. Through the visualization analysis of the self-attention matrix of the multi-head self-attention layer in the spectral Transformer model, it can reveal the key bands or band combinations that the model focuses on when making classification decisions. This not only enhances the credibility of the model's decisions, but also has the potential to help geologists discover previously overlooked new spectral features that distinguish different rock types. Compared with the local receptive field of CNN, the global attention mechanism of this method can achieve parallel and integrated modeling of the entire spectral range, which is more suitable for capturing the overall morphology and long-distance features of the spectrum. Attached Figure Description

[0020] Figure 1 is a flowchart of an embodiment of the intelligent classification method for rock hyperspectral data based on the Transformer self-attention mechanism of the present invention; Figure 2 is a logical diagram of the online enhancement of each preprocessed spectral data S' in the training dataset by randomly selecting a transformation strategy from the data augmentation strategy pool in step 2.2 of the embodiment of the present invention; Figure 3 is a structural block diagram of the spectral Transformer model in step 4 of the embodiment of the present invention. Detailed Implementation

[0021] To make the advantages and features of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] This invention aims to design a unified intelligent computing framework that can collaboratively address three key challenges in rock hyperspectral analysis: 1) how to effectively model the long-range band dependencies that are common in the spectrum; 2) how to overcome the severe spectral distortion caused by physical environmental factors such as water content in real exploration scenarios; and 3) how to ensure the generalization ability of the model under the condition of scarce training samples commonly faced in the geological field.

[0023] Based on the above concept, the present invention provides a rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism, as shown in Figure 1, which specifically includes the following steps: Step 1, constructing a training dataset and a validation dataset.

[0024] Step 1.1: Use an ASD FieldSpec series hyperspectral analyzer to collect spectral reflectance data of different rock samples in a specific electromagnetic spectrum range, such as the visible-near infrared (350nm-2500nm) electromagnetic spectrum range, to form multiple raw spectral data S containing different wavelength points.

[0025] Step 1.2: A smoothing filter is used to smooth multiple original spectral data to suppress noise and high-frequency interference, resulting in multiple denoised spectral data. The smoothing filter can be a Savitzky-Golay filter, a moving average filter, or a Gaussian filter, etc. In this embodiment, a Savitzky-Golay filter is used, with a window length of 11 and a polynomial order of 2, to effectively denoise while preserving key spectral morphology features.

[0026] Step 1.3 involves standardizing multiple denoised spectral data sets to eliminate dimensional differences and accelerate model convergence, resulting in multiple preprocessed spectral data sets. In this embodiment, Z-Score standardization is used to adjust the data distribution of each wavelength point across all samples to a preset statistical distribution (e.g., mean 0, standard deviation 1), yielding preprocessed spectral data S'. The preprocessing steps 1.2 and 1.3 effectively improve data quality.

[0027] Step 1.4: Construct training and validation datasets based on multiple preprocessed spectral data S'.

[0028] Step 2, augmentation of the training dataset.

[0029] To optimize the spectral Transformer model and improve classification accuracy, this invention first performs feature enhancement on the preprocessed spectral data S', and then feeds the feature-enhanced spectral data and the unenhanced preprocessed spectral data S' into the spectral Transformer model for training.

[0030] Step 2.1: Construct a data augmentation strategy pool, which includes transformation strategies for basic physics simulation, regularization transformation strategies based on sample mixing, and perturbation simulation transformation strategies based on physical mechanisms.

[0031] The transformation strategies for fundamental physics simulations include Gaussian noise, random intensity scaling, random baseline shift, and random clipping and scaling, used to simulate common measurement errors.

[0032] Regularization transformation strategies based on sample mixing include Mixup (inter-class mixing) and Intra-class Mixup (intra-class mixing). When there are classes with highly similar spectral features in the training dataset, Mixup or Intra-class Mixup is used for transformation. By performing linear interpolation only between samples of the same class, intra-class diversity is enriched while avoiding blurring of class boundaries.

[0033] The perturbation simulation transformation strategy based on physical mechanisms is the moisture effect simulation transformation, which addresses the industry pain point of moisture state interference with spectral identification. The moisture effect simulation transformation obtains standard moisture absorption spectrum curves. and combine it with a random strength factor (For example, After multiplication, the results are superimposed onto the preprocessed spectral data using either multiplication or addition. This generates virtual samples with different water content states that have real physical meaning. For example, it can be done through The simulation method is targeted and based on physical mechanisms. This data augmentation actively teaches the model how to decouple the intrinsic characteristics of rocks from the spectrum heavily contaminated by moisture, which is one of the key steps in achieving high robustness in this invention. After simulating the moisture effect transformation, extremely high classification accuracy was achieved under both dry and moist conditions in a specific rock dataset, significantly outperforming existing technologies and demonstrating excellent classification performance and strong generalization ability on new data.

[0034] Step 2.2, as shown in Figure 2, randomly selects a transformation strategy from the data augmentation strategy pool to perform online augmentation on each preprocessed spectral data S' in the training dataset, generating augmented spectral data S''. The augmented spectral data S'' is then used as augmented data to expand the training dataset.

[0035] Step 3: Since the spectral Transformer model can only process sequential data, and the preprocessed spectral data S' and enhanced spectral data S'' in the training and validation datasets are both one-dimensional data, it is necessary to first reshape each spectral data in the training and validation datasets into a spectral sequence according to the spectral curve to adapt to the spectral Transformer model.

[0036] Step 4, training the spectral Transformer model.

[0037] As shown in Figure 3, the spectral Transformer model in this embodiment includes an input processing module, an encoder module, and a classification decision module. The input processing module includes a slicing layer, a linear projection layer, and a position encoding layer connected sequentially in order of input and output. The input end of the slicing layer is used to receive the spectral sequence and process it into a spectral array of length [missing information]. The spectral sequence is divided into There are P spectral slices, which can be non-overlapping or partially overlapping; in this embodiment, the length P of the spectral slice is set to 16 wavelength points, i.e., the spectral slice is P-dimensional. A linear projection layer is used to map the features of each spectral slice to a high-dimensional space, forming a high-dimensional embedding vector. In this embodiment, the mapped embedding vector is 128-dimensional. A positional encoding layer is used to add positional information to the embedding vector sequence to preserve the order relationship of the bands. In this embodiment, a sine / cosine positional encoding method is preferably used to add positional information. The encoder module in this embodiment includes four stacked encoding blocks; each encoding block includes a multi-head self-attention layer, a first normalization layer, a feedforward neural network layer, and a second normalization layer connected sequentially in input and output. The input of the multi-head self-attention layer in the first encoding block is connected to the output of the positional encoding layer, used to calculate in parallel the correlation weights between multiple spectral slices after adding positional information, in order to capture global spectral features; the input of the multi-head self-attention layer in the second encoding block is connected to the output of the second normalization layer in the first encoding block, used to calculate in parallel the correlation weights between the features of multiple spectral slice sequences, and so on. In this embodiment, the multi-head self-attention mechanism in the multi-head self-attention layer has 8 heads. The feedforward neural network layer performs nonlinear transformations on the input data to enhance the model's ability to fit complex features; the first and second normalization layers normalize the corresponding input data, and the second normalization layer in the fourth encoding block outputs the sequence features corresponding to the extracted multiple spectral slices. The classification decision module includes a pooling layer and a classification head connected sequentially in input and output order; the input of the pooling layer is connected to the output of the second normalization layer in the fourth encoding block, used to receive the sequence features output by the encoder module and aggregate them into a fixed-dimensional global feature vector. In this embodiment, the aggregation method of the pooling layer is as follows: extract the representation vector corresponding to the first spectral slice in the spectral sequence and directly use it as the global feature vector. The classification head outputs the probability that a sample belongs to each rock category through one or more fully connected layers and a Softmax function, i.e., outputs the rock category prediction result.

[0038] The reconstructed spectral sequences from the training dataset are fed into the aforementioned spectral Transformer model for training. During training, a gradient descent-based optimizer and a cross-entropy loss function are used, with model parameters iteratively updated via backpropagation. The gradient descent-based optimizer can be AdamW, SGD, or RMSprop, etc.; in this embodiment, AdamW is used.

[0039] In addition, during the training process, a validation dataset is used to verify the model training results. In multiple training rounds, the performance of the spectral Transformer model on the validation dataset is continuously monitored, and the model with the best validation performance is used as the trained spectral Transformer model.

[0040] Step 5: Using the methods in Steps 1.2 and 1.3, the spectral data of the rocks to be classified are sequentially subjected to spectral denoising and data standardization. The processed spectral data are then reshaped into spectral sequences that are adapted to the input of the spectral Transformer model according to the spectral curves. The reshaped spectral sequences are then input into the trained spectral Transformer model. Through one forward propagation, the rock category prediction result is directly output, thus completing the identification and classification of rock types.

[0041] The above description is only used to illustrate the technical solutions of the present invention, and is not intended to limit them. For those skilled in the art, modifications can be made to the specific technical solutions described in the above embodiments, or equivalent substitutions can be made to some of the technical features. However, these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions protected by the present invention.

Claims

1. A rock hyperspectral intelligent classification method based on Transformer self-attention mechanism, characterized in that, Includes the following steps: Step 1: Construct training and validation datasets; obtain raw spectral data of different rock samples, and perform spectral denoising and data standardization on multiple raw spectral data in sequence to obtain multiple preprocessed spectral data; The training dataset and validation dataset are constructed based on multiple preprocessed spectral data; Step 2: Augmentation of the training dataset; Constructing a data augmentation strategy pool, which includes transformation strategies for basic physical simulation, regularization transformation strategies based on sample mixing, and perturbation simulation transformation strategies based on physical mechanisms. The transformation strategy in the data augmentation strategy pool is randomly selected to perform online augmentation on each preprocessed spectral data in the training dataset, generating augmented spectral data. Use augmented spectral data as augmentation data for the training dataset; Step 3: Reshape each spectral data point in the training and validation datasets into a spectral sequence adapted to the input of the spectral Transformer model according to the spectral curve. Step 4: Train the spectral Transformer model. Feed the reshaped spectral sequences from the training dataset into the spectral Transformer model for training. During training, an optimizer based on gradient descent and a cross-entropy loss function are used, and the model parameters are iteratively updated using the backpropagation algorithm. The model training results are validated using the validation dataset, and the model with the best validation performance is selected as the trained spectral Transformer model. Step 5: Perform spectral denoising and data standardization on the spectral data of the rocks to be classified, and reshape the processed spectral data into a spectral sequence adapted to the input of the spectral Transformer model according to the spectral curve. Input the reshaped spectral sequences into the trained spectral Transformer model, and through one forward propagation, directly output the rock category prediction result, completing the rock type identification and classification.

2. The rock hyperspectral intelligent classification method based on Transformer self-attention mechanism according to claim 1, characterized in that: In step 2, the transformation strategies of the basic physical simulation include Gaussian noise, random intensity scaling, random baseline shift, and random pruning and scaling; the regularization transformation strategies based on sample mixing include Mixup and Intra-class Mixup; and the perturbation simulation transformation strategy based on physical mechanisms is the moisture effect simulation transformation.

3. The rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism according to claim 2, characterized in that: In step 4, the gradient descent-based optimizer is AdamW, SGD, or RMSprop.

4. A rock hyperspectral intelligent classification method based on Transformer self-attention mechanism according to any one of claims 1-3, characterized in that: In step 4, the spectral Transformer model includes an input processing module, an encoder module, and a classification decision module; the input processing module includes a slicing layer, a linear projection layer, and a position encoding layer connected in sequence according to input and output; the input end of the slicing layer is used to receive the spectral sequence and divide the spectral sequence into multiple spectral slices; the linear projection layer is used to map the features of each spectral slice to a high-dimensional space to form a high-dimensional embedding vector. The positional encoding layer is used to add positional information to the embedded vector sequence; The encoder module includes N stacked coding blocks, where N≥1; each coding block includes a multi-head self-attention layer, a first normalization layer, a feedforward neural network layer, and a second normalization layer connected in sequence according to input and output. The input of the multi-head self-attention layer in the first coding block is connected to the output of the position coding layer; Multi-head self-attention layers are used to compute association weights between multiple spectral slices in parallel to capture global spectral features; Feedforward neural network layers are used to perform nonlinear transformations on the input data; The first and second normalization layers are used to normalize the corresponding input data; the second normalization layer in the Nth coding block outputs the sequence features corresponding to the extracted multiple spectral slices; The classification decision module includes a pooling layer and a classification head connected sequentially in terms of input and output. The input of the pooling layer is connected to the output of the second normalization layer in the Nth encoding block, and is used to receive the sequence features output by the encoder module and aggregate them into a global feature vector of fixed dimension. The classification head is used to output the rock category prediction result based on the global feature vector.

5. The rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism according to claim 4, characterized in that: In step 4, N=4; the number of heads in the multi-head self-attention layer is 8.

6. The rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism according to claim 5, characterized in that: In step 4, the position encoding layer adds position information using a sine / cosine position encoding method.

7. The intelligent rock hyperspectral classification method based on Transformer self-attention mechanism according to claim 6, characterized in that, Step 1 specifically comprises: Step 1.1, using a hyperspectral analyzer to collect spectral reflectance data of different rock samples in the visible-near-infrared electromagnetic spectrum range, forming multiple original spectral data containing different wavelength points; Step 1.2, using a smoothing filter to smooth the multiple original spectral data respectively, obtaining multiple denoised spectral data; Step 1.3, performing standardization processing on the multiple denoised spectral data respectively, obtaining multiple preprocessed spectral data; Step 1.4, constructing training datasets and validation datasets based on the multiple preprocessed spectral data respectively.

8. The rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism according to claim 7, characterized in that: In step 5, the spectral data of the rocks to be classified are subjected to spectral denoising and data standardization in sequence. Specifically, a smoothing filter is first used to smooth the spectral data of the rocks to be classified to obtain the denoised spectral data of the rocks to be classified; then the denoised spectral data of the rocks to be classified is standardized to obtain the preprocessed spectral data of the rocks to be classified.

9. The rock hyperspectral intelligent classification method based on the Transformer self-attention mechanism according to claim 8, characterized in that: In steps 1.2 and 5, the smoothing filter is a Savitzky-Golay filter, a moving average filter, or a Gaussian filter; in steps 1.3 and 5, the normalization process uses Z-Score normalization.