A method and system for extracting features from spectral data
Patent Information
- Application Number
- CN202611113625.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]然而,上述现有Transformer方法应用于高维小样本光谱分析时存在严重缺陷,其根本原因在于其固有的分块标记化策略与单一优化目标设计未能适配数据特性,首先,分块操作是将具有连续物理意义的吸收峰硬性截断,破坏了光谱特征的物理拓扑连续性,导致关键判别信息碎片化;其次,分块策略会产生较长的标记序列,显著增加了自注意力矩阵的计算复杂度,在小样本约束下,庞大的参数空间极易因“数据饥饿”而拟合高频噪声,加剧过拟合
本发明提出的光谱数据特征提取方法,构建的网络模型,通过特征提取模块将高维小样本光谱数据集中的全波段光谱无切分地完整投影为单一的核心特征向量,避免了将连续光谱硬性切割为多个独立片段,不仅避免了直接截断光谱中较宽的物理吸收峰,破坏局部特征的连续性和理化意义,而且在网络输入端实质上构建了一个“信息漏斗”,在保留全局物理轮廓的同时,有效过滤了高频冗余背景噪声。并行的对比约束投影分支和分类决策分支,避免了因强制归一化压缩而丢失光谱物理尺度特征的问题。
Smart Images

Figure CN122821148A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of spectral analysis and deep learning, and more specifically to a method and system for extracting spectral data features. Background Technology
[0002] Spectroscopic analysis technology, due to its advantages such as being non-destructive, rapid, and enabling simultaneous detection of multiple components, has become a core tool for physicochemical property analysis in agriculture, food, and pharmaceutical fields. However, data generated by modern spectrometers typically possesses the dual characteristics of being "high-dimensional" (thousands of bands) and "small-sample" (high-quality labeled samples are scarce). The high dimensionality stems from band overlap and nonlinear coupling caused by dense sampling and complex molecular vibrational responses, easily leading to the "curse of dimensionality." The small-sample nature is limited by the cost of expensive destructive physicochemical sampling and the need for expert annotation. In this high-dimensional small-sample (HDLSS) data scenario, traditional analytical models are prone to information loss or overfitting, making it difficult to stably extract discriminative and generalizable feature representations from massive redundant backgrounds, thus becoming a key bottleneck restricting the in-depth application of this technology.
[0003] In recent years, the Transformer architecture has become the mainstream method for processing high-dimensional sequence data (including spectral signals) due to its powerful global dependency modeling capabilities. Among them, existing spectral Transformer methods, represented by SpectralFormer, first employ a block-based labeling strategy to divide the continuous spectral signal sequence into multiple fixed-length sub-segments, and linearly project each sub-segment as a label. Then, these labeled sequences are input into a Transformer encoder composed of a multi-layer multi-head self-attention mechanism. Through self-attention calculation, the long-range dependencies between labels in different bands are captured, thereby achieving the extraction and dimensionality reduction of global spectral features. Finally, the extracted features are used for downstream classification or regression tasks.
[0004] However, the existing Transformer methods described above suffer from serious drawbacks when applied to high-dimensional, small-sample spectral analysis. The root cause lies in the fact that their inherent block-based labeling strategy and single optimization objective design fail to adapt to the data characteristics. First, the block operation rigidly truncates absorption peaks with continuous physical meaning, disrupting the physical topological continuity of spectral features and leading to fragmentation of key discriminative information. Second, the block strategy generates long label sequences, significantly increasing the computational complexity of the self-attention matrix. Under small-sample constraints, the large parameter space is prone to fitting high-frequency noise due to "data hunger," exacerbating overfitting. Finally, the optimization space of a single loss function is unstable. Existing methods typically only use cross-entropy loss or standard contrastive loss for optimization. In small-sample, small-batch scenarios, pure cross-entropy loss lacks underlying geometric constraints, easily leading to feature discretization. Standard contrastive loss, when faced with scarce positive samples within a batch and amplified high-dimensional noise, is prone to numerical instability and excessive feature clustering or dimensional collapse, ultimately resulting in the failure of the representation system and a decline in generalization ability. Summary of the Invention
[0005] To address the problems existing in the above-mentioned fields, this invention proposes a method and system for extracting spectral data features. The constructed network model, while preserving the continuity of physical topology, alleviates the technical bottleneck of parameter overfitting under small sample conditions, and the constructed loss function can achieve high-precision feature transfer.
[0006] To address the aforementioned technical problems, this invention discloses a method for extracting spectral data features, comprising the following steps: Obtain the original high-dimensional small sample spectral dataset and construct a training set based on the true labels of the spectral dataset; A network model is constructed, wherein the network model is based on a Transformer encoder, a feature extraction module is added before the Transformer encoder, and parallel contrastive constraint projection branches and classification decision branches are connected at the output of the Transformer encoder; The feature extraction module projects the full-band spectrum of the spectral dataset into a core feature vector, and sequentially adds a preset sequence dimension and position encoding to the core feature vector to form an initial feature sequence; the Transformer encoder extracts the depth feature vector of the initial feature sequence; the contrast constraint projection branch projects the depth feature vector and normalizes it to a unit hypersphere to obtain the projected features; the classification decision branch performs a linear mapping on the depth feature vector to obtain the classification prediction probability. The network model is trained based on the training set and the loss function to obtain the trained network model. The loss function is determined based on the difference between the classification prediction probability and the true label, as well as the degree of clustering of the projected features with similar samples and the degree of separation with dissimilar samples on the unit hypersphere. The high-dimensional small-sample spectral dataset to be tested is input into the trained network model, and the projection features output by the contrast constraint projection branch are used as the extracted spectral features.
[0007] Preferably, the feature extraction module projects the full-band spectrum of the spectral dataset into a core feature vector, and sequentially adds a preset sequence dimension and positional encoding to the core feature vector to form an initial feature sequence, specifically including: The whole-band spectrum of the input original high-dimensional small sample spectral dataset is linearly mapped without segmentation using a global projection matrix of preset dimensions to obtain a core feature vector with fixed low dimensions. Add a preset sequence dimension to the core feature vector and reshape it into an initial single token sequence structure containing only a single element; The preset position encoding matrix is added to and fused with the initial single token sequence structure to generate the initial feature sequence.
[0008] Preferably, the Transformer encoder extracts a deep feature vector from the initial feature sequence, specifically including: The Transformer encoder consists of a Pre-LayerNorm layer, a multi-head attention layer, a normalization layer, and a feedforward network in sequence. The Pre-LayerNorm layer performs a first-layer normalization process on the feature distribution of the initial feature sequence to obtain the features after the first normalization process. The multi-head attention layer performs channel-dimensional feature recalibration on the features after the first normalization process through multiple parallel linear projection subspaces; and adds the output of the multi-head attention layer to the features before the first normalization process through the first residual connection and normalization, and inputs it into the normalization layer for the second normalization process. The feedforward network performs nonlinear feature purification and expression enhancement on the features after the second normalization process; and adds the output of the feedforward network to the features before the second normalization process through the second residual connection to output a deep feature vector.
[0009] Preferably, the contrast-constrained projection branch projects and normalizes the depth feature vector to a unit hypersphere to obtain projected features, specifically including: The contrast-constrained projection branch includes a nonlinear projection layer and an L2 normalization layer; The nonlinear projection layer uses a projection head composed of a multilayer sensing mechanism to perform nonlinear dimensionality reduction mapping on the depth feature vector, thereby obtaining the feature vector after the projection head mapping. The L2 normalization layer maps the feature vectors mapped by the projection head onto the unit hypersphere that has been normalized by L2, generating the dimension-reduced projection features.
[0010] Preferably, the classification decision branch performs a linear mapping on the deep feature vector to obtain the classification prediction probability, specifically including: The classification decision branch uses a single-layer linear classification head to perform linear mapping on the deep feature vector and generate prediction logical values corresponding to each preset category. The generated prediction logic values corresponding to each preset category are processed by the softmax activation function to output the classification prediction probability corresponding to each category.
[0011] Preferably, the loss function is determined based on the difference between the classification prediction probability and the true label, as well as the degree of clustering of projected features with similar samples and the degree of separation with dissimilar samples on a unit hypersphere, specifically including: The cross-entropy loss is determined based on the difference between the predicted classification probability and the true label. The contrast loss is determined based on the degree of clustering with similar samples and the degree of separation from dissimilar samples on a unit hypersphere, according to the projection features. The weighted sum of the cross-entropy loss and the contrastive loss yields the loss function.
[0012] Preferably, the method for determining the cross-entropy loss based on the difference between the classification prediction probability and the true label is as follows: ; in, C The total number of categories, N This represents the total number of samples participating in training within the current training batch. for The first under the category The true label of each sample; This refers to the classification prediction probability output by the classification decision branch in the trained network model.
[0013] Preferably, the contrast loss is determined based on the degree of clustering with similar samples and the degree of separation with dissimilar samples on a unit hypersphere according to the projection features: ; ; in, To compare the losses, Represents a set The base number, This represents the set of valid anchor points within the current training batch. I Indicates the sample index within the current training batch. The set, Within the current training batch and anchor point i A set of positive samples of the same category; Indicates anchor point There must be at least one positive sample of the same category in the current training batch; Indicates anchor point Positive sample set of the same category The projected feature vector of a positive sample in the dataset; j This indicates the current training batch excluding anchor points. Index of all other samples, Indicates the currently extracted anchor point ; This indicates the current training batch excluding anchor points. Anchor points in all other samples j ; This refers to temperature hyperparameters. This indicates an internal numerical clamping operation; A quantity term to prevent logarithmic calculation crashes caused by extreme free samples.
[0014] Preferably, the step of obtaining the original high-dimensional small-sample spectral dataset and constructing a training set based on the true labels of the spectral dataset specifically includes: The acquired original high-dimensional small sample spectral dataset is divided into training set and test set according to a preset ratio; The spectral data in the training set includes the raw spectral data and its corresponding true class labels.
[0015] Preferably, it further includes a spectral data feature extraction system, comprising: The data acquisition module is used to acquire the original high-dimensional small sample spectral dataset and construct a training set based on the true labels of the spectral dataset; A network model construction module is used to construct a network model, wherein the network model is based on a Transformer encoder, with a feature extraction module added before the Transformer encoder, and parallel contrastive constraint projection branches and classification decision branches connected to the output of the Transformer encoder; the feature extraction module projects the full-band spectrum of the spectral dataset into a core feature vector, and sequentially adds a preset sequence dimension and position encoding to the core feature vector to form an initial feature sequence; the Transformer encoder extracts the depth feature vector of the initial feature sequence; the contrastive constraint projection branch projects and normalizes the depth feature vector to a unit hypersphere to obtain projected features; the classification decision branch linearly maps the depth feature vector to obtain the classification prediction probability; the network model is trained based on the training set and a loss function to obtain the trained network model, wherein the loss function is determined based on the difference between the classification prediction probability and the true label, and the degree of clustering of the projected features with similar samples and the degree of separation with dissimilar samples on the unit hypersphere; The feature extraction module is used to input the high-dimensional small sample spectral dataset to be tested into the trained network model, and the projection features output by the contrast constraint projection branch are used as the extracted spectral features.
[0016] Compared with the prior art, the present invention has the following beneficial effects: The spectral data feature extraction method proposed in this invention constructs a network model that, through a feature extraction module, projects the entire spectrum of a high-dimensional, small-sample spectral dataset into a single core feature vector without any segmentation. This avoids the rigid cutting of continuous spectra into multiple independent segments, preventing the direct truncation of broad physical absorption peaks in the spectrum and thus avoiding the destruction of the continuity and physicochemical meaning of local features. Furthermore, it essentially constructs an "information funnel" at the network input, effectively filtering high-frequency redundant background noise while preserving the global physical contour. The parallel contrast-constrained projection branch and classification decision branch avoid the loss of spectral physical scale features due to forced normalization compression.
[0017] During training, a constructed loss function is used. This function uses the degree of clustering of projected features with similar samples and the degree of separation from dissimilar samples on a unit hypersphere as the model's underlying geometric constraints. This explicitly brings similar samples closer together and pushes away dissimilar samples, addressing the feature discretization phenomenon caused by local noise interference in pure cross-entropy loss with very small sample sizes. The difference between the predicted classification probability and the true label is used as the model's top-level classification guide, assigning mutually exclusive decision boundaries to the clustered projected feature clusters. This addresses the unconstrained over-clustering phenomenon that can easily occur in pure contrastive learning when class coordinates are lacking. The weighted combination of these two loss functions constructs a spectral feature space that combines discriminative power and robustness. The projected features output by the network model after the above loss function co-optimization naturally exhibit a discriminative topology in the latent space, characterized by compact intra-class features and separation between classes. This allows the extracted features to perfectly adapt to various independent downstream machine learning analyzers even under the stringent constraint of requiring only a very small number of training samples, achieving high-precision feature transfer and providing reliable algorithmic support for reducing the expensive destructive physical and chemical sampling and labeling costs in non-destructive testing. Attached Figure Description
[0018] Figure 1 This is a flowchart of the spectral data feature extraction method proposed in this invention; Figure 2 The network architecture of the network model constructed for this invention; Figure 3 The original feature distribution curves of three representative spectral datasets provided in the embodiments of the present invention; Figure 4 Box plots comparing the classification performance of various methods on three datasets provided in this embodiment of the invention; Figure 5 Eleven sets of ablation experimental results for the mixed loss weighting coefficients provided in this embodiment of the invention; Figure 6 The small sample performance evolution trajectory of the CLS-Former framework and various baseline methods under different training set proportions provided in the embodiments of the present invention; Figure 7 The evolution of t-SNE feature manifolds for the tablet (Raman) dataset provided in this embodiment of the invention under different loss constraints and training ratios; Figure 8 The evolution of t-SNE feature manifolds for the tablet (NIR) dataset provided in this embodiment of the invention under different loss constraints and training ratios. Detailed Implementation
[0019] The following will refer to the appendices in the embodiments of the present invention. Figures 1-8The technical solutions in the embodiments of the present invention will be clearly and completely described. It should be understood that the terminology used in the present invention is only for describing particular implementation methods and is not intended to limit the present invention.
[0020] Example like Figure 1 As shown, this invention proposes a method for extracting spectral data features, which includes the following steps: S1: Obtain the original high-dimensional small sample spectral dataset and construct a training set based on the true labels of the spectral dataset; S2: Construct a network model, in which the network model is based on the Transformer encoder, a feature extraction module is added before the Transformer encoder, and parallel contrastive constraint projection branch and classification decision branch are connected at the output of the Transformer encoder. The feature extraction module projects the full-band spectrum of the spectral dataset into a core feature vector, and sequentially adds preset sequence dimensions and positional encodings to the core feature vector to form an initial feature sequence; the Transformer encoder extracts the depth feature vector of the initial feature sequence; the contrast constraint projection branch projects the depth feature vector and normalizes it to a unit hypersphere to obtain the projected features; the classification decision branch performs a linear mapping on the depth feature vector to obtain the classification prediction probability. The network model is trained based on the training set and the loss function to obtain the trained network model. The loss function is determined based on the difference between the classification prediction probability and the true label, as well as the degree of clustering of the projected features with similar samples and the degree of separation with dissimilar samples on the unit hypersphere. S3: Input the high-dimensional small sample spectral dataset to be tested into the trained network model, and use the projection features output by the contrast constraint projection branch as the extracted spectral features.
[0021] Specifically, in step S1, the acquired original high-dimensional small sample spectral dataset is divided into a training set and a test set according to a preset ratio; the spectral data in the training set includes the original spectral data and its corresponding true category labels.
[0022] To eliminate the inherent scale differences in absorbance or scattering intensity in cross-spectral analysis techniques and to adapt to the scale requirements of the feature distribution in the subsequent contrastive learning metric space, this embodiment employs Z-score normalization as a dimension-wise preprocessing scheme for the original high-dimensional spectrum. By mapping each feature dimension (i.e., the spectral response at a specific wavenumber) to a distribution with a mean of 0 and a standard deviation of 1, this operation reduces the scale bias caused by different band dimensions. This helps improve the numerical stability and gradient convergence speed of model training and provides a physically consistent input representation for the subsequent global serialization mapping mechanism. Its mathematical transformation is defined as: ; In the formula, The original high-dimensional spectral vector, and These represent the sample mean and standard deviation vectors for the corresponding feature dimensions, respectively.
[0023] To avoid the risk of data leakage that may occur in small-sample evaluations, this embodiment follows the principle of independent evaluation during data partitioning and processing. The above-mentioned standardized distribution parameters... The statistical calculations are based solely on the training set data after each independent split. Subsequently, the data transformation between the test and validation sets directly applies the fixed parameter matrix extracted from the training set. This isolation mechanism ensures the independence of the test data from the model training process, thereby guaranteeing the objectivity of the model in generalization performance evaluation.
[0024] In step S2, the network model (Contrastive-Learning and Supervised-classification Former, CLS-Former framework) provided in this embodiment is designed following the guiding principles of "preserving global physical continuity" and "multi-task decoupling and collaboration." For example... Figure 2 As shown, its forward propagation architecture consists of two core modules connected in series: a feature extraction module and a multi-task decoupled dual-branch output architecture, wherein the dual-branch output architecture includes a parallel contrast constraint projection branch and a classification decision branch.
[0025] The feature extraction module aims to circumvent the locality limitations of traditional sequence segmentation strategies, transforming the original high-dimensional small-sample spectral data ( The signal is mapped to a global single-token sequence and relies on a lightweight Transformer encoder to construct an "information funnel." While preserving the physical contours across the entire band, this mechanism achieves dimensionality reduction and feature extraction, transforming complex nonlinear signals into a single core feature vector. Subsequently, the hybrid loss optimization module receives this feature and employs a dual-branch output architecture with multi-task decoupling. Robust supervised contrastive learning is introduced. With cross-entropy Under the dual constraints of deep space, this module simultaneously optimizes the clustering of low-level feature distributions and the delineation of top-level classification decision boundaries. Through end-to-end connection of these two modules, a spectral feature extraction process under high-dimensional, small-sample conditions is jointly constructed.
[0026] To avoid physical truncation that may result from hard sequence segmentation and the risk of overfitting under small sample conditions, this study adopts a precoding method that "preserves global topological integrity" and proposes a global serialization mapping mechanism.
[0027] Using a global projection matrix of a preset dimension For a specific dimension of the input D Raw high-dimensional small sample spectral data A non-segmented linear mapping is performed on the full-band spectrum to obtain a fixed low-dimensional core feature vector. In this embodiment, the hidden dimension is... Set to 64; Add a sequence dimension to the core feature vector and reshape it into an initial single-token sequence structure containing only a single element; To adapt to the input specifications of the Transformer encoder, the dimensionality reduction vector is assigned a sequence dimension, and a pre-defined, learnable position encoding matrix is used. This is combined with the initial single-token sequence structure to generate the initial feature sequence. : ; The significance of introducing a global single-token pre-architecture in this embodiment lies in the fact that it essentially constructs an "information funnel" at the network input end. This is achieved by establishing an information funnel from a higher-dimensional space. To the low-dimensional representation subspace The projection mapping mechanism filters out high-frequency redundant noise and alleviates the physical truncation problem caused by the block strategy, enabling the model to extract global spectral features while maintaining the continuity of band topology.
[0028] In this context, feature purification relies primarily on the Pre-LayerNorm within the Transformer encoder and the feedforward network FFN with residual connections.
[0029] To further reduce the risk of overfitting of deep models in small sample scenarios, the core computational unit for feature extraction adopts a lightweight single-layer Transformer encoder with four parallel attention heads.
[0030] The Transformer encoder consists of a Pre-LayerNorm layer, a multi-head attention layer, a normalization layer, and a feedforward network. Pre-LayerNorm layer for initial feature sequence The feature distribution is subjected to layer normalization to obtain the normalized features; Since the input sequence length is limited to 1, the traditional self-attention mechanism (MHA) in the multi-head attention layer no longer performs cross-band cross-computation. Instead, it uses its multi-head linear projection subspace to perform feature recalibration of the channel dimension, extracting the discriminative information of different feature channels in the normalized features. Then, through the first residual connection and normalization, the output of the multi-head attention layer is added to the features before the first normalization, and the result is input into the normalization layer for the second normalization. The feedforward network performs nonlinear feature purification and expression enhancement on the features after the second normalization process; and through the second residual connection, the output of the feedforward network is added to the features before the second normalization process to output a deep feature vector.
[0031] The forward propagation computation process is defined as follows: ; ; in, This represents the network depth. FFN consists of two linear mapping layers with a hidden dimension of 256 and using the ReLU activation function, used to enhance the non-linear expressive power of deep features. To address the sparse nature of labeled samples and constrain the model's hypothesis space, this embodiment sets the Dropout rate to 0.3 and restricts the Transformer encoder depth to a single layer. =1. This simplified architecture retains the gradient stability of the Transformer encoder while adapting to the upper limit of information capacity for small samples, thereby reducing high-dimensional noise fitting caused by parameter redundancy, and ultimately outputting a deep feature vector. .
[0032] In small-sample scenarios, directly inputting high-dimensional features into a linear classifier can easily lead to overfitting by fitting noise from specific training samples to optimize labels. Therefore, this embodiment designs a multi-task decoupled output architecture, establishing a decoupled structure between "bottom-level distribution optimization" and "top-level boundary partitioning" at the network's end. The deep feature vector output by the Transformer encoder... The evaluation is split into two independent evaluation branches in parallel: the contrastive constraint projection branch and the classification decision branch, where:
[0033] The contrast-constrained projection branch includes a nonlinear projection layer and an L2 normalization layer. The nonlinear projection layer uses a projection head composed of a multilayer perceptron (MLP) to process the depth feature vector. A nonlinear dimensionality reduction mapping is performed to obtain the feature vector after the projection head is mapped; the L2 normalization layer maps the feature vector after the projection head mapping onto the L2 normalized unit hypersphere to generate the dimensionality-reduced projected feature vector. .
[0034] This path strips away absolute magnitude information, providing a standardized geometric metric space for subsequent feature topology aggregation.
[0035] In the classification decision branch, a single-layer linear classification head is used to process the original deep feature vector. Perform a linear mapping to generate prediction logits corresponding to each preset category; The generated predicted logical values for each preset category are processed using a softmax activation function to output the classification prediction probability for each category. .
[0036] Model training In the analysis of high-dimensional, small-sample spectral data, a single optimization objective often struggles to balance the compactness of deep feature distributions with the accurate delineation of classification decision boundaries. Therefore, this invention designs a hybrid optimization strategy that integrates robust supervised comparison and cross-entropy constraints.
[0037] Standard supervised contrastive loss uses labels to explicitly construct positive and negative sample pairs, encouraging similar features to cluster in deeper space and pushing away dissimilar features. However, when faced with high-dimensional spectra and small sample sizes, the limited batch size may lead to isolated anchor points lacking similar positive samples within a batch, triggering the division-by-zero anomaly. Simultaneously, high-dimensional background noise may increase the scalar value of feature dot products, causing numerical instability after exponential operations (such as gradient explosion or NaN anomalies). To mitigate the numerical instability of small-batch training, this embodiment employs a hybrid optimization strategy to robustly improve the contrastive loss.
[0038] Specifically, a contrast loss with numerical clamp protection is constructed based on the degree of clustering with similar samples and the degree of separation from dissimilar samples on a unit hypersphere using projection features. for:
[0039]
[0040] in, To compare the losses, Represents a set The base number, This represents the set of valid anchor points within the current training batch. I Indicates the sample index within the current training batch. The set, Within the current training batch and anchor point i A set of positive samples of the same category; Indicates anchor point There must be at least one positive sample of the same category in the current training batch; Indicates anchor point Positive sample set of the same category The projected feature vector of a positive sample in the dataset; j This indicates the current training batch excluding anchor points. Index of all other samples, Indicates the currently extracted anchor point ; This indicates the current training batch excluding anchor points. Anchor points in all other samples j ; For the temperature hyperparameter, the value is set to 0.1 in this embodiment; This indicates the internal numerical clamping operation; A quantity term to prevent logarithmic calculation crashes caused by extreme free samples.
[0041] Compared to traditional calculation methods, this invention introduces a low-level numerical protection mechanism: First, the algorithm extracts a set of valid anchor points with at least one positive sample before summing the numerators. And based on the cardinality of the set. Perform outermost mean normalization to avoid... To prevent division by zero errors caused by empty arrays, the gradient is scale-invariant with respect to batch size. Secondly, an internal numerical clamping operation is introduced. The feature cosine similarity is truncated to The interval is used to suppress the amplification effect of high-frequency noise in exponential operations; finally, a minimal constant is added to the logarithmic denominator. This prevents logarithmic calculations from crashing due to extreme free samples.
[0042] To give the feature clustering process a classification task orientation and prevent the feature distribution from falling into over-clustering or collapse, the top-level classification decision path is simultaneously driven by the standard cross-entropy loss. Cross-entropy is responsible for assigning mutually exclusive decision boundaries to the clustered feature clusters. Based on the difference between the predicted classification probabilities and the true labels, the determined cross-entropy loss is: ; The cross-entropy loss and contrastive loss are weighted and summed to construct the loss function (joint optimization objective function) of the network model: ; in, C The total number of categories, N This represents the total number of samples participating in training within the current training batch. for The first under the category The true labels of each sample are obtained through one-hot encoding; The classification prediction probability output for the classification decision branch; For weights.
[0043] Preliminary experiments show that classifying directly on the normalized hyperspherical features may compress the physical scale features of the spectrum. Therefore, the unnormalized original depth feature vector is used in the classification decision branch. Calculate the cross-entropy to preserve the discriminative power of the spectrum to the greatest extent possible.
[0044] In step S3, the trained network model is validated using a test set, and the projection features output by the contrast constraint projection branch are used as the extracted spectral features.
[0045] This invention also proposes a spectral data feature extraction system, comprising: The data acquisition module is used to acquire the original high-dimensional small sample spectral dataset and construct a training set based on the true labels of the spectral dataset; The network model construction module is used to build the network model. The network model uses a Transformer encoder as the base network, with a feature extraction module added before the Transformer encoder. Parallel contrastive constraint projection branches and classification decision branches are connected to the output of the Transformer encoder. The feature extraction module projects the full-band spectrum of the spectral dataset into core feature vectors, and sequentially adds preset sequence dimensions and positional encodings to these core feature vectors to form an initial feature sequence. The Transformer encoder extracts the depth feature vectors from the initial feature sequence. The contrastive constraint projection branch projects and normalizes the depth feature vectors to a unit hypersphere to obtain projected features. The classification decision branch linearly maps the depth feature vectors to obtain classification prediction probabilities. The network model is trained based on the training set and a loss function to obtain the trained network model. The loss function is determined based on the difference between the classification prediction probability and the true label, as well as the degree of clustering of projected features with similar samples and the degree of separation with dissimilar samples on the unit hypersphere. The feature extraction module is used to input the high-dimensional small sample spectral dataset to be tested into the trained network model, and the projection features output by the contrast constraint projection branch are used as the extracted spectral features.
[0046] Experimental Analysis To verify the generalization ability of the proposed CLS-Former framework under conditions of high-dimensional nonlinear decoupling and scarce labeled samples, this embodiment constructs a spectral evaluation benchmark covering multiple physical mechanisms. This benchmark includes three typical high-dimensional small-sample spectral datasets: tablets (Raman), tablets (NIR), and wine (FT-IR). All data in these high-dimensional small-sample spectral datasets were provided and publicly released by the Department of Food Science at the University of Copenhagen, and their acquisition was legal and legitimate. These three types of spectral datasets correspond to specific data characteristics in actual nondestructive testing, aiming to systematically evaluate the model's characterization ability under different data conditions. Their key statistical characteristics are summarized in Table 1.
[0047] Table 1. Summary of key statistical features of high-dimensional small-sample spectral datasets using cross-spectral analysis techniques.
[0048] The Raman tablet dataset, with its 3402-dimensional high-resolution features, constitutes a typical high-dimensional nonlinear evaluation benchmark. Because its feature dimension far exceeds the total number of samples (…),… Furthermore, the complex molecular vibrational responses lead to significant nonlinear physical correlations across a wide frequency band, and key discriminative information is easily masked by redundant background noise. This dataset is primarily used to evaluate the global dependency modeling and complex nonlinear feature decoupling capabilities of the lightweight Transformer encoder in the CLS-Former framework.
[0049] The tablet (NIR) dataset consists of near-infrared spectra with a moderate feature dimension of 407. Its main characteristics are significant multicollinearity and baseline drift between bands. This dataset was introduced for parallel evaluation to examine the effectiveness of the framework in removing physical redundancy and extracting perturbation-resistant discriminative features under collinearity interference.
[0050] The wine (FT-IR) dataset represents an application scenario with highly limited sample sizes. This dataset possesses 842-dimensional Fourier transform infrared features, but the total number of samples is only 44, constituting a stringent few-sample learning condition. Under this data scale, deep models are prone to overfitting and representation system failure. Testing on this benchmark aims to verify the robustness of CLS-Former's hybrid loss synergistic strategy in correcting class centroid bias and reshaping the high-dimensional sparse spatial structure under few-sample constraints.
[0051] To evaluate the performance of the CLS-Former framework in feature decoupling, generalization ability, and few-shot scenarios, this experiment designed a two-stage progressive validation scheme. Stratified sampling was employed in all random sampling stages to ensure that the prior class distributions of the training and test sets remained consistent with the original dataset.
[0052] Phase 1: Baseline Comparison and Mixed Loss Ablation Experiment This phase aims to evaluate the performance of the CLS-Former framework under a standard sample size and optimize the weight configuration of the hybrid loss co-operation mechanism. The experiment divides the entire dataset into training and independent test sets in a 7:3 ratio, and performs 10 independent Monte Carlo cross-validations. For model evaluation, this phase employs two methods: "feature transfer" and "end-to-end direct classification." The former inputs extracted features into four independent downstream classifiers—SVM, Random Forest (RF), Logistic Regression (LR), and K-Nearest Neighbors (KNN)—to verify the generality of the features; the latter utilizes the model's built-in classification head to evaluate its discriminative ability. Classification accuracy (...) is used as the evaluation metric. The mean and standard deviation of () are used as core quantitative indicators.
[0053] The baseline comparison includes the four mainstream machine learning classifiers mentioned above, classic dimensionality reduction algorithms such as principal component analysis (PCA) and linear discriminant analysis (LDA), as well as deep metric learning models for small sample scenarios, including Siamese Network and Prototypical Network.
[0054] Phase Two: Small Sample Evolutionary Analysis This phase focuses on examining the model's feature representation and generalization performance under data-constrained conditions. Experiments concentrate on the Raman (tablet) dataset, which exhibits high-dimensional nonlinearity, and the NIR (tablet) dataset, which suffers from strong collinearity. The independent test set is set to 30%, and the training set allocation starts at 10%, increasing in 10% increments to 70%. To ensure the effectiveness of the evaluation under sparse sampling (e.g., a 10% training allocation), a class integrity verification mechanism is introduced. This mechanism avoids missing class issues through hierarchical class coverage and minimum sampling, ensuring that each class retains valid samples at each evolutionary gradient.
[0055] This stage visually reveals the feature distribution patterns of the model under different data scales. It also introduces a high-dimensional feature visualization technique based on t-Distributed Stochastic Neighbor Embedding (t-SNE). Combining the quantified "training ratio-performance" evolution curve with the t-SNE clustering pattern, this stage aims to verify whether the framework can maintain a discriminative topology of "compact within classes and separation between classes" with small samples, thereby exploring the effective sample size requirements for spectral analysis.
[0056] The network model in this embodiment is trained based on the PyTorch framework, and both model training and evaluation are completed on a computing platform equipped with an NVIDIA RTX 4070 GPU.
[0057] In terms of network optimization and parameter configuration, the full-scale experiment used the AdamW optimizer, with the initial learning rate and weight decay coefficient both set to [value missing]. To address the training oscillation problem that is prone to occur in small sample scenarios, the maximum number of training epochs is limited to 100 epochs, and the batch size is set to 8. A cosine annealing learning rate scheduling strategy is integrated throughout the training cycle, and a gradient clipping mechanism with a maximum L2 norm of 1.0 is applied to suppress the risk of gradient explosion in high-dimensional space. In the hybrid loss calculation, the feature extraction base is a single-layer, four-head parallel Transformer encoder, and the temperature hyperparameter of the supervised contrastive loss is used. Fixed at 0.1.
[0058] To prevent potential data leakage issues in hyperparameter tuning and early stopping strategies, this embodiment introduces an inner-validation mechanism. In each round of independent sampling, the algorithm stratifies and extracts 20% of the data from the current training set as an inner-validation set. The classification accuracy of the logistic regression probe (LR probe) is used as a monitoring metric for the quality of the underlying feature clustering, and an early stopping tolerance of 15 rounds is set to capture the best training epoch. After locking the target epoch, the model discards its current weights and retrains from scratch to that specified epoch on the full training set, including the original inner-validation set. This process makes full use of limited data while preventing the leakage of test set information, ensuring the objectivity of the final evaluation results.
[0059] like Figure 3 The image shows a visualization of the original curves from the three representative spectral datasets mentioned above. The horizontal axis represents the wavenumber (cm²). -¹), the vertical axis represents the spectral response intensity (Intensity, au) or absorbance (Absorbance, au) respectively, based on the physical properties of each spectral detection technology.
[0060] like Figure 3 As shown, although the spectral signals acquired by different techniques were not affected by high-frequency noise, they all exhibited different physical response characteristics and distribution features. Specifically, in Figure 3 In the Raman tablet dataset (shown as 'a'), with a feature dimension of 3402, the spectral curves exhibit dense feature peaks, and the peaks of different categories overlap nonlinearly across a wide frequency band. Figure 3 As shown in b of the image, in the tablet (NIR) dataset, the spectrum exhibits a broad absorption band accompanied by significant baseline shift and contour overlap, reflecting multicollinearity between adjacent bands; Figure 3 As shown in c, in the wine (FT-IR) dataset, the spectral trajectories of different samples highly overlap, with relatively small macroscopic inter-class variance, mainly concentrated in local wavenumber ranges (e.g., 1000-1500). With 2800-4000 There are local response differences in the vicinity.
[0061] In summary, these three datasets commonly exhibit band overlap, baseline fluctuations, and nonlinear coupling. These intuitive data characteristics of "feature redundancy and category overlap" indicate the limitations of relying solely on apparent absolute amplitudes for feature extraction. This objective challenge at the data level constitutes the main motivation for introducing the CLS-Former framework in this invention to achieve deep nonlinear decoupling and feature space optimization.
[0062] For the Raman tablet dataset (n=120) with 3402 high-resolution features, traditional feature engineering methods are prone to dimensionality reduction loss when dealing with high-dimensional small samples. As shown in Table 2, Principal Component Analysis (PCA), which relies on the assumption of linear mapping, is difficult to capture deep nonlinear correlations between bands, and the accuracy of its extracted features in the four downstream classifiers is only between 57.22% and 61.39%.
[0063] Table 2 Comparison of feature transfer and end-to-end classification accuracy on the tablet (Raman) dataset (accuracy ± standard deviation, %)
[0064] In contrast, the CLS-Former framework achieves decoupling of complex nonlinear features through global serialization mapping. Its extracted features achieve an accuracy of 91.94% on both SVM and logistic regression (LR) classifiers (a 30.55% improvement over PCA). In end-to-end evaluation, the CLS-Former framework achieves 90.83%, outperforming the Prototypical Network (61.39%), which is constrained by high-dimensional sparse space, and also surpassing the single-objective Transformer baseline (Transformer-CE, 86.11%), which does not incorporate contrastive learning constraints. These results demonstrate that hybrid loss strategies have significant architectural advantages in mitigating the "curse of dimensionality" in high-dimensional spaces and extracting nonlinear features.
[0065] As shown in Table 3, the ability of the model to extract robust features was examined in a tablet (NIR) dataset with a moderate feature dimension of 407 dimensions but accompanied by significant baseline drift and multicollinearity.
[0066] Table 3. Comparison of feature transfer and end-to-end classification accuracy on the tablet (NIR) dataset (accuracy ± standard deviation, %)
[0067] Table 3 shows that the global self-attention mechanism of the CLS-Former framework can effectively remove collinear redundancy information. In the feature transfer task, features based on the CLS-Former framework consistently achieve over 95% accuracy across all four downstream classifiers, with the highest accuracy of 96.13% achieved when combined with Random Forest (RF). In end-to-end direct classification, the CLS-Former framework maintains a classification accuracy of 94.84%. In contrast, conventional few-shot metric learning models exhibit certain performance limitations when dealing with strong collinear backgrounds; for example, the end-to-end accuracy of the Siamese Network and the prototype network drops to 87.63% and 81.83%, respectively. Experimental data demonstrate that the CLS-Former framework can more stably extract chemical composition representations resistant to physical interference.
[0068] As shown in Table 4, the wine (FT-IR) dataset suffers from a very small sample size ( The matrix response is highly similar to the matrix response, which can easily lead to model overfitting and class centroid estimation bias.
[0069] Table 4. Comparison of feature transfer and end-to-end classification accuracy on the wine (FT-IR) dataset (accuracy ± standard deviation, %)
[0070] The results in Table 4 show that at this microscale of data, mainstream metric learning baselines face representation degradation; for example, the prototype network's end-to-end accuracy is only 47.86%. However, the CLS-Former framework maintains high representation robustness under these constraints: its end-to-end accuracy reaches 71.43% (a 23.57% improvement over the prototype network). Furthermore, its extracted features achieve a performance of 72.86% in both SVM and LR downstream classifiers, and the performance standard deviation of all test metrics is kept at a low level. This result confirms that the synergistic effect of supervised comparison mechanism and cross-entropy can effectively stabilize the topological structure of high-dimensional feature space and correct class distribution bias under extremely small sample conditions.
[0071] To evaluate the training stability of each method under multiple independent sampling from a statistical distribution perspective, such as Figure 4 As shown, box plots illustrate the distribution of classification accuracy for different feature extraction methods across three datasets. Here, a represents the tablet (Raman) dataset (n=120); b represents the tablet (NIR) dataset (n=310); and c represents the wine (FT-IR) dataset (n=44). Each subplot shows the statistical distribution of feature transfer accuracy for each method across four downstream classifiers (SVM, RF, LR, KNN) and the end-to-end direct classification accuracy of the corresponding deep models in 10 independent random sampling experiments.
[0072] In 10 independent repeated samplings of the three datasets, the classification accuracy of each method exhibited different statistical distribution characteristics. Overall, the median performance (the center line within the box) of the CLS-Former framework was at a high level in all evaluation scenarios, and the interquartile range (IQR, i.e., box height), which reflects the range of performance fluctuations, was relatively small.
[0073] Specifically, Figure 4 Figure a shows that on the Raman tablet dataset with high feature dimensions, conventional feature engineering methods (such as PCA) generally have low accuracy, while the Prototypical Network exhibits significant distribution variance. In contrast, the CLS-Former framework demonstrates a more compact distribution of high scores. Figure 4 Figure b shows that in the NIR (Non-Integrated Matrix) dataset with strong collinearity, although most methods have high medians, some deep metric learning baselines (such as Siamese network's Direct end-to-end classification) exhibit significant vertical stretching and downward outlier distributions, indicating training instability under specific sampling conditions. The CLS-Former framework, on the other hand, maintains a flattened box distribution in this scenario.
[0074] Figure 4Figure c shows that in the extremely scarce wine (FT-IR) dataset, the overall IQR of various methods is amplified to varying degrees due to the very small data size. However, conventional methods (such as PCA and prototype networks) generally exhibit significant box stretching and low score distribution, reflecting their susceptibility to overfitting or getting trapped in local optima with extremely small samples. Under this constraint, the CLS-Former framework still maintains a high median and, to some extent, controls the lower bound of performance.
[0075] Based on the distribution pattern of the box plots, the smaller IQR and higher median of the CLS-Former framework indicate a hybrid loss mechanism ( This not only helps improve the average accuracy of classification, but also effectively reduces the performance fluctuation of the model under different data partitions, providing a stable representation basis for the generalization of spectral features and end-to-end classification.
[0076] To assess and monitor the loss With cross-entropy loss In this embodiment, the synergistic effect within the CLS-Former framework is investigated using hybrid weighting coefficients. The ablation experiment. The experiment will... The loss was increased from 0.0 (pure cross-entropy constraint) in increments of 0.1 to 1.0 (pure contrastive learning constraint), and 10 independent repeated samplings were performed on each of the three datasets to quantify the impact of different loss proportions on classification performance and training stability.
[0077] like Figure 5 As shown, the weighting coefficients of the mixed loss The results of 11 ablation experiments were analyzed, based on the mean evolution trend of each classification index. Different datasets showed varying sensitivity to weights, but overall classification performance remained consistent. The surrounding area has reached a relatively good level. Figure 5 In the figure, 'a' represents the performance curve exhibiting certain fluctuations in the tablet (Raman) dataset. High accuracy rates were achieved across all assessment pathways, while... When the value is 0.7 or 1.0, the performance drops significantly. Figure 5 In this context, 'b' represents the classification performance in the NIR (Non-Impact Tablet) dataset. It is relatively stable within the range, and... When it reaches its peak; when it relies entirely on contrast constraints ( When the value was 0, all indicators showed a significant decrease. Figure 5 In this context, 'c' represents the performance of the model on a very small wine (FT-IR) dataset. The fluctuations are more pronounced; a weight deviation from 0.5 (such as 0.7 or 0.9) leads to a decrease in classification accuracy. Furthermore, Figure 5 In the graph, d, e, and f are the independent error bar decomposition plots of each evaluation curve in subplots a, b, and c, respectively. From top to bottom, they represent the mean values of SVM, LR, RF, KNN, and Direct. Standard deviation is used to microscopically represent training stability under different configurations. Figure 5 The error bar distribution of df in the data further indicates that, Nearby, the performance standard deviations of each downstream evaluation task are relatively small, reflecting that the training process is relatively stable at this time.
[0078] Analyzing the mechanism of feature space optimization, a single objective function easily exposes specific structural limitations when dealing with high-dimensional small samples. When the weights are completely biased towards pure cross-entropy loss (i.e., ... When the sample size is very small, model optimization relies entirely on error feedback from the top-level labels. Under extremely small sample conditions, the feature space, lacking constraints from the underlying geometric structure, is prone to "feature discretization." The network tends to fit local noise from a small subset of samples, resulting in loosely distributed similar features in the latent space, thus weakening the model's generalization ability.
[0079] Conversely, when the network is fully driven by supervised contrastive loss (i.e. When performing small batches, the performance degradation and variance increase are mainly due to the lack of mini-batch sampling constraints and absolute class coordinates. Under limited batch processing scale, small samples easily trigger computational bottlenecks caused by insufficient positive sample anchors. More importantly, without the top-level class guide provided by cross-entropy, feature clusters relying solely on relative distance are prone to unconstrained over-clustering, ultimately leading to biases in the classification decision plane.
[0080] Experimental results and mechanistic analysis show that a single optimization strategy is insufficient for extracting complex spectral features. Setting the threshold to 0.5 effectively activates the synergistic effect of both methods: the supervised contrastive loss acts as a structural constraint, encouraging the clustering of similar features, while the cross-entropy loss defines the decision boundary. This joint optimization strategy effectively alleviates representation instability under small sample conditions and constructs a spectral feature space that combines discriminative power and robustness.
[0081] To explore the representational capabilities of the CLS-Former framework under limited data scales, this embodiment conducts gradient evolution evaluation on the tablet (Raman) (total samples n=120) and tablet (NIR) (total samples n=310) datasets. The experiment maintains an independent test set ratio of 30%, and increases the training set allocation ratio from 0.1 to 0.7. In this stage, the combined average accuracy of feature transfer precision from the four downstream classifiers is used as the quantitative metric, and 80% is set as the evaluation baseline. Figure 6 The study demonstrates the macroscopic evolution trends of various methods under different proportions of training data.
[0082] like Figure 6 As shown, the performance evolution trajectories of the CLS-Former framework and various baseline methods on small samples under different training set proportions are illustrated in the line graph. The graph shows the variation of the overall average accuracy of the features extracted by each method in the downstream classifier as the proportion of training samples increases, under the premise of a fixed 30% independent test set. Figure 6 In the figure, 'a' represents the overall improvement in classification performance of various methods with increasing training samples on the Raman tablet dataset, which has high-dimensional nonlinear characteristics. With smaller data scales (scales of 0.1 and 0.2), limited by the dependence of deep networks on basic information, the accuracy of the CLS-Former framework is similar to that of Linear Discriminant Analysis (LDA), ranging from 56.46% to 62.57%. When the training scale increases to 0.3, the accuracy of the CLS-Former framework improves to 73.12%; and at a scale of 0.4 (approximately 12 samples per class), it reaches 81.46%, surpassing the 80% baseline. Within the range of 0.5 to 0.7, its performance stabilizes between 85.49% and 86.60%. In contrast, while LDA reaches approximately 80% at a scale of 0.4, its performance tends to stagnate with subsequent increases in data volume. The combined accuracy of mainstream metric learning baselines (such as prototype networks and Siamese networks) within this test range fails to break through 80%. The above trends indicate that the CLS-Former framework can extract high-dimensional nonlinear features more effectively after reaching a certain sample threshold.
[0083] Figure 6In the figure, b represents the high overall baseline performance of the classification task in the Non-Integrated Tablet (NIR) dataset, which exhibits significant collinearity. At a low training ratio of 0.1 (approximately 8 samples per class), the CLS-Former framework achieves a comprehensive average accuracy of 86.40%, exceeding the performance of the deep metric learning baseline at the same ratio and outperforming conventional PCA dimensionality reduction methods at a high ratio of 0.7 (below 78%). As the number of training samples increases, the evolution trajectory of the CLS-Former framework remains stable, reaching 91.24% at a ratio of 0.2 and a global peak accuracy of 96.69% at a ratio of 0.7. The Siamese Network demonstrates good adaptability in this task and closely follows, but the CLS-Former framework maintains high accuracy across the entire range. These evaluation results indicate that the hybrid loss strategy has stable generalization ability when dealing with spectral multicollinearity interference.
[0084] Based on the above gradient evolution evaluation results, the CLS-Former framework exhibits a relatively stable and perturbation-resistant performance growth curve under limited data conditions. For high-dimensional and complex spectral signals, the framework maintains effective classification performance even when the proportion of training samples is reduced to 10% to 40% (approximately 8 to 12 samples per class). This objective law provides a feasible algorithmic reference for reducing the high costs of destructive physicochemical sampling and data annotation in practical applications.
[0085] To explore the intrinsic mechanism of the CLS-Former framework at the feature representation level, this embodiment employs t-SNE dimensionality reduction technology to map the high-dimensional feature distribution onto a two-dimensional visualization plane. For example... Figure 7 As shown, this embodiment selects the tablet (Raman) dataset with high feature dimensions and, as shown, the tablet (Raman) dataset with high feature dimensions. Figure 8 The NIR (Non-Inductively Coupled Pixel) dataset, exhibiting collinearity, is shown. The system compares the evolution of feature topology as the training data size is gradually reduced under three loss constraints (pure cross-entropy loss, pure supervised contrastive loss, and the hybrid loss proposed in this paper). The analysis focuses on the constraint of reducing the training ratio to 0.2 (i.e.,...). Figure 7 mo in and Figure 8 (mo in the model), the influence of different optimization mechanisms on the geometric distribution of the feature space.
[0086] in, Figure 7 The evolution of the t-SNE feature manifold for the tablet (Raman) dataset under different loss constraints and training ratios is given. Here, a–c represents the feature distribution of the original test set; d–f represents the feature distribution of the full training set (Stage 1); and g–o represents the feature evolution trajectory under the typical small sample condition (Stage 2) with the training ratio decreasing sequentially (0.6, 0.4, 0.2). Figure 7 The left column represents CLS-Former ( The middle column represents the Transformer baseline model (TFM-CE) driven by pure cross-entropy. The right column represents the Transformer baseline model driven by purely supervised contrastive learning (TFM-SupCon). ).
[0087] and Figure 7 resemblance, Figure 8 The evolution of t-SNE feature manifolds for the tablet (NIR) dataset under different loss constraints and training ratios.
[0088] When the feature extraction network is driven only by purely supervised contrastive loss ( The model tends to bring similar samples closer together by using relative distance. However, due to the lack of class boundary guidance provided by cross-entropy, feature clusters are prone to unconstrained over-clustering in the dimensionality reduction space. Figure 7 In the Raman dataset with a training ratio of 0.2, the feature distributions of different categories interweave and overlap, resulting in a significant decrease in inter-class separation. Figure 8 In the near-infrared dataset, the phenomenon manifests as a significant spatial overlap of feature clusters for specific categories (such as the green and orange samples in the lower right corner of the image). This confusion in spatial topological distribution provides a geometric explanation for the degradation of classification performance in pure contrastive learning with limited samples.
[0089] When the network is driven solely by pure cross-entropy ( ), Figure 7 and Figure 8 In the context of sufficient training data, 'e' represents the model's ability to define basic decision boundaries. However, as the training ratio decreases to 0.2, feature distribution exhibits feature discretization due to the lack of explicit constraints on the underlying spatial structure. Figure 7 In the Raman dataset, samples of the same class exhibit a relatively dispersed scatter pattern with blurred inter-class boundaries. Figure 8 In the near-infrared dataset, although the features maintain a certain striped structure, significant class confusion occurs in the central intersection region (e.g., blue, orange, and green samples intertwine locally near the origin). This distribution pattern indicates that in a sparse feature space, relying solely on cross-entropy loss is easily affected by local data noise and makes it difficult to maintain the compactness of similar features.
[0090] In contrast, CLS-Former, which employs hybrid loss co-optimization, It exhibits relatively stable feature decoupling performance under different data gradients. Under a limited training ratio of 0.2, it is effective for high-dimensional Raman spectra (…). Figure 7 In the near-infrared dataset (m), this method maps features into four relatively independent and mutually exclusive clustering regions. Figure 8 In the model (m), the feature clusters extend radially outward from the center, maintaining good intra-class compactness, and there is no obvious interweaving or adhesion between the branches of different classes. This discriminative topology, characterized by intra-class compactness and inter-class separation, visually verifies the synergistic mechanism of the hybrid loss: supervised contrastive loss acts as a structural constraint to promote the clustering of similar features, while cross-entropy loss acts as a classification guide to drive the feature centroids to partition towards mutually exclusive decision domains. The combination of the two constructs a robust spectral feature space under small sample conditions.
[0091] Experimental results demonstrate that the CLS-Former framework possesses superior feature representation capabilities when processing high-dimensional complex spectral data. Traditional dimensionality reduction methods (such as PCA and LDA) rely on linear mapping assumptions for decoupling high-dimensional nonlinear features. When processing data like Raman spectra, which have 3402 dimensions and overlapping nonlinear bands, linear models are prone to performance degradation due to the loss of discriminative features. Empirical data shows that PCA achieves an accuracy of 61.39% on this dataset. In contrast, the CLS-Former framework avoids a sequence-blocking strategy that could disrupt physical continuity. Instead, it constructs an "information funnel" while preserving physical integrity through global serialization mapping and a lightweight Transformer encoder. This mechanism suppresses the propagation of high-frequency background noise, achieving low-dimensional decoupling of nonlinear features. Its feature transfer accuracy reaches 91.94% (a 30.55% improvement over PCA), effectively mitigating information loss and feature overlap issues in high-dimensional space.
[0092] To address the model degradation problem caused by scarce samples, conventional deep metric learning methods are prone to overfitting or class centroid estimation bias. Taking the wine (FT-IR) dataset with a total of 44 samples as an example, the Prototypical Network achieves an end-to-end accuracy of 47.86% in this sparse feature space. In contrast, the CLS-Former framework combines multi-task decoupling output with a hybrid loss mechanism... This method aggregates the low-level feature distribution through supervised contrastive loss and uses cross-entropy loss to delineate the top-level decision boundary, achieving an accuracy of 71.43% on this dataset. During feature distribution optimization, this mechanism mitigates dimensionality collapse and training instability that are prone to occur under small batch and small sample conditions through low-level protective designs such as numerical clamping and effective anchor point extraction. Experimental results show that this framework maintains high classification accuracy while exhibiting low performance variance in multiple repeated samplings, providing a reliable feature extraction scheme for high-dimensional small-sample spectral analysis.
[0093] This embodiment utilizes three spectral datasets with different physical mechanisms to verify the generalization ability of the CLS-Former framework across detection technologies. From the Raman tablet dataset, which has high feature dimensionality and severe nonlinear overlap, to the NIR tablet dataset, which exhibits baseline drift and multicollinearity, and finally to the FT-IR wine dataset, characterized by small macroscopic inter-class variance and scarce samples, the CLS-Former framework consistently demonstrates stable feature transfer accuracy and end-to-end classification performance. This performance across physical mechanisms indicates that the global labeling and hybrid co-optimization mechanism employed by the framework can overcome the apparent differences in underlying optical responses, extract universal core physicochemical features, and thus exhibit generalization robustness independent of specific spectral detection hardware.
[0094] Furthermore, this embodiment quantifies the small-sample representation boundary of the CLS-Former framework through gradient evolution evaluation. Experimental results show that, with an 80% reference threshold for basic representation effectiveness, the framework requires only 40% training ratio (approximately 12 samples per class) when dealing with Raman spectroscopy, and reduces data dependency to 10% (approximately 8 samples per class) when dealing with near-infrared spectroscopy. Although the empirical boundary for a specific dataset cannot be completely equated to the theoretical lower limit for all spectral techniques, this quantitative indicator demonstrates the framework's potential in reducing data annotation requirements. Compared to the data starvation dilemma faced by traditional deep networks, the reduced minimum effective sample size requirement of the CLS-Former framework provides an algorithmic foundation for building a low-cost, high-precision non-destructive testing system.
[0095] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0096] Furthermore, unless otherwise stated, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. All references to this specification are incorporated by way of citation to disclose and describe methods relating to those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.
Claims
1. A method for extracting features from spectral data, characterized in that, Includes the following steps: Obtain the original high-dimensional small sample spectral dataset and construct a training set based on the true labels of the spectral dataset; A network model is constructed, wherein the network model is based on a Transformer encoder, a feature extraction module is added before the Transformer encoder, and parallel contrastive constraint projection branches and classification decision branches are connected at the output of the Transformer encoder; The feature extraction module projects the full-band spectrum of the spectral dataset into a core feature vector, and sequentially adds a preset sequence dimension and position encoding to the core feature vector to form an initial feature sequence; the Transformer encoder extracts the depth feature vector of the initial feature sequence; the contrast constraint projection branch projects the depth feature vector and normalizes it to a unit hypersphere to obtain the projected features; the classification decision branch performs a linear mapping on the depth feature vector to obtain the classification prediction probability. The network model is trained based on the training set and the loss function to obtain the trained network model. The loss function is determined based on the difference between the classification prediction probability and the true label, as well as the degree of clustering of the projected features with similar samples and the degree of separation with dissimilar samples on the unit hypersphere. The high-dimensional small-sample spectral dataset to be tested is input into the trained network model, and the projection features output by the contrast constraint projection branch are used as the extracted spectral features.
2. The spectral data feature extraction method according to claim 1, characterized in that, The feature extraction module projects the full-band spectrum of the spectral dataset into a core feature vector, and sequentially adds a preset sequence dimension and positional encoding to the core feature vector to form an initial feature sequence, specifically including: The whole-band spectrum of the input original high-dimensional small sample spectral dataset is linearly mapped without segmentation using a global projection matrix of preset dimensions to obtain a core feature vector with fixed low dimensions. Add a preset sequence dimension to the core feature vector and reshape it into an initial single token sequence structure containing only a single element; The preset position encoding matrix is added to and fused with the initial single token sequence structure to generate the initial feature sequence.
3. The spectral data feature extraction method according to claim 1, characterized in that, The Transformer encoder extracts a deep feature vector from the initial feature sequence, specifically including: The Transformer encoder consists of a Pre-LayerNorm layer, a multi-head attention layer, a normalization layer, and a feedforward network in sequence. The Pre-LayerNorm layer performs a first-layer normalization process on the feature distribution of the initial feature sequence to obtain the features after the first normalization process; The multi-head attention layer performs channel-dimensional feature recalibration on the features after the first normalization process through multiple parallel linear projection subspaces; and adds the output of the multi-head attention layer to the features before the first normalization process through the first residual connection and normalization, and inputs it into the normalization layer for the second normalization process. The feedforward network performs nonlinear feature purification and expression enhancement on the features after the second normalization process; and adds the output of the feedforward network to the features before the second normalization process through the second residual connection to output a deep feature vector.
4. The spectral data feature extraction method according to claim 1, characterized in that, The contrast-constrained projection branch projects and normalizes the depth feature vector to a unit hypersphere to obtain projected features, specifically including: The contrast-constrained projection branch includes a nonlinear projection layer and an L2 normalization layer; The nonlinear projection layer uses a projection head composed of a multilayer sensing mechanism to perform nonlinear dimensionality reduction mapping on the depth feature vector, thereby obtaining the feature vector after the projection head mapping. The L2 normalization layer maps the feature vectors mapped by the projection head onto the unit hypersphere that has been normalized by L2, generating the dimension-reduced projection features.
5. The spectral data feature extraction method according to claim 1, characterized in that, The classification decision branch performs a linear mapping on the deep feature vector to obtain the classification prediction probability, specifically including: The classification decision branch uses a single-layer linear classification head to perform linear mapping on the deep feature vector and generate prediction logical values corresponding to each preset category. The generated prediction logic values corresponding to each preset category are processed by the softmax activation function to output the classification prediction probability corresponding to each category.
6. The spectral data feature extraction method according to claim 1, characterized in that, The loss function is determined based on the difference between the classification prediction probability and the true label, as well as the degree of clustering of projected features with similar samples and the degree of separation with dissimilar samples on a unit hypersphere. Specifically, it includes: The cross-entropy loss is determined based on the difference between the predicted classification probability and the true label. The contrast loss is determined based on the degree of clustering with similar samples and the degree of separation from dissimilar samples on a unit hypersphere, according to the projection features. The weighted sum of the cross-entropy loss and the contrastive loss yields the loss function.
7. The spectral data feature extraction method according to claim 6, characterized in that, The cross-entropy loss, determined based on the difference between the predicted classification probability and the true label, is as follows: ; in, C The total number of categories, N This represents the total number of samples participating in training within the current training batch. for The first under the category The true label of each sample; This refers to the classification prediction probability output by the classification decision branch in the trained network model.
8. The spectral data feature extraction method according to claim 6, characterized in that, The contrast loss is determined based on the degree of clustering with similar samples and the degree of separation with dissimilar samples on a unit hypersphere, according to the projection features: ; ; in, To compare the losses, Represents a set The base number, This represents the set of valid anchor points within the current training batch. I Indicates the sample index within the current training batch. The set, Within the current training batch and anchor point i A set of positive samples of the same category; Indicates anchor point There must be at least one positive sample of the same category in the current training batch; Indicates anchor point Positive sample set of the same category The projected feature vector of a positive sample in the dataset; j This indicates the current training batch excluding anchor points. Index of all other samples, Indicates the currently extracted anchor point ; This indicates the current training batch excluding anchor points. Anchor points in all other samples j ; This refers to temperature hyperparameters. This indicates an internal numerical clamping operation; A quantity term to prevent logarithmic calculation crashes caused by extreme free samples.
9. The spectral data feature extraction method according to claim 1, characterized in that, The process of obtaining the original high-dimensional small-sample spectral dataset and constructing a training set based on the true labels of the spectral dataset specifically includes: The acquired original high-dimensional small sample spectral dataset is divided into training set and test set according to a preset ratio; The spectral data in the training set includes the raw spectral data and its corresponding true class labels.
10. A spectral data feature extraction system, characterized in that, include: The data acquisition module is used to acquire the original high-dimensional small sample spectral dataset and construct a training set based on the true labels of the spectral dataset; A network model construction module is used to construct a network model, wherein the network model is based on a Transformer encoder, with a feature extraction module added before the Transformer encoder, and parallel contrastive constraint projection branches and classification decision branches connected to the output of the Transformer encoder; the feature extraction module projects the full-band spectrum of the spectral dataset into a core feature vector, and sequentially adds preset sequence dimensions and positional encodings to the core feature vector; the Transformer encoder extracts the depth feature vector of the initial feature sequence; the contrastive constraint projection branch projects and normalizes the depth feature vector to a unit hypersphere to obtain projected features; the classification decision branch linearly maps the depth feature vector to obtain the classification prediction probability; the network model is trained based on the training set and a loss function to obtain the trained network model, wherein the loss function is determined based on the difference between the classification prediction probability and the true label, and the degree of clustering of the projected features with similar samples and the degree of separation with dissimilar samples on the unit hypersphere; The feature extraction module is used to input the high-dimensional small sample spectral dataset to be tested into the trained network model, and the projection features output by the contrast constraint projection branch are used as the extracted spectral features.