Contrastive learning feature amplification biomass adaptive channel splicing rapid detection system and method
By combining multi-channel differential spectroscopy and the BYOL framework with a feature amplification and attention mechanism, the problems of overlapping absorption peaks and weak feature signals in biomass property detection are solved, achieving high-precision, interpretable non-destructive detection of biomass properties and improving the robustness and applicability of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2026-04-10
- Publication Date
- 2026-07-14
AI Technical Summary
Existing NIRS-based methods for detecting biomass characteristics suffer from severe overlap of absorption peaks, weak feature signals, and complex background interference. Traditional preprocessing methods lack adaptive fusion capabilities, and deep learning models have insufficient generalization ability in small sample scenarios and fail to effectively incorporate physical mechanisms, resulting in insufficient prediction reliability.
The model employs a multi-channel differential spectroscopy construction module, a channel attention weighting module, a sequence encoding module, and a contrastive learning feature amplification module. Combined with the BYOL framework and physical constraints, it achieves feature enhancement and long-range dependency modeling through adaptive optimization of preprocessed channel weights and feature fusion, thereby improving the robustness and generalization ability of the model in small sample scenarios.
It significantly improves the high-precision non-destructive testing capability of biomass characteristics. The model reduces RMSE to 45.6 and MAPE to 5.01% on the test set. It is suitable for various biomass categories and complex application scenarios, has efficient feature extraction capabilities and strong generalization performance, and supports migration applications for different spectrometer platforms and biomass types.
Smart Images

Figure CN122392712A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rapid detection technology for spectral data, specifically to a rapid detection method driven by intelligent algorithms for predicting biomass characteristics, which combines contrastive learning feature amplification and adaptive channel stitching. Background Technology
[0002] Near-infrared spectroscopy (NIRS) is a highly efficient and non-destructive testing method based on the vibrational frequency harmonics and combination frequency absorption of molecules. By measuring the spectral response of biomass samples in the near-infrared band, its elemental composition, chemical composition, and physical properties can be rapidly determined. Due to its advantages such as fast analysis speed, no sample pretreatment required, and simultaneous detection of multiple components, NIRS technology has demonstrated significant application value in fields such as biomass resource utilization, agricultural quality monitoring, and industrial process control.
[0003] However, existing NIRS-based methods for detecting biomass properties still face several technical bottlenecks. First, near-infrared spectroscopy itself suffers from severe absorption peak overlap, weak characteristic signals, and complex background interference, making it difficult to extract effective information. Second, traditional preprocessing methods (such as scattering correction and derivative transformation) rely on manual experience for selection and lack the ability to adaptively fuse multi-source spectral features, making it difficult to fully exploit the effective information in complementary preprocessing channels such as RAW, first-order derivative (1D), and second-order derivative (2D). Third, at the modeling level, traditional linear methods such as partial least squares (PLS) are unable to capture the complex nonlinear relationship between the spectrum and the target properties, while deep learning models heavily rely on a large amount of labeled data and have insufficient generalization ability in small biomass sample scenarios.
[0004] Furthermore, existing methods for feature extraction and modeling strategies often employ single models or simple feature splicing, which have significant limitations. For example, directly splicing multi-channel differential spectra leads to the curse of dimensionality and information dilution, making it less effective than a single-channel approach. Traditional machine learning models (such as PLS and SVR) lack the ability to model long-range spectral dependencies, making it difficult to capture complex nonlinear coupling effects between wavelengths. More importantly, existing methods fail to effectively incorporate prior physical knowledge, neglecting the mechanistic constraints of biomass conversion processes, resulting in insufficient predictive reliability in complex real-world scenarios.
[0005] In summary, existing NIRS detection methods still have significant shortcomings in feature extraction, multi-source information fusion, small-sample modeling, and physical mechanism embedding. Therefore, there is an urgent need to construct a novel intelligent analysis framework that integrates feature amplification, adaptive channel weighting, and long-range dependency modeling. By introducing contrastive learning, attention mechanisms, and physical constraint optimization, this framework can effectively enhance, fuse, and interpretably model spectral features, thereby promoting the development of non-destructive testing technology for biomass properties towards higher precision, stronger generalization, and mechanism-driven approaches. Summary of the Invention
[0006] This invention provides a non-destructive detection method for biomass characteristics based on BYOL spectral feature amplification channel splicing and attention mechanism. This method combines multi-channel differential spectral features, attention weight allocation mechanism, and long-range dependency modeling capability to match reasonable feature enhancement strategies and modeling schemes for complex biomass samples. At the same time, it addresses the problems of spectral drift, noise interference, and feature weakening in spectral data by adaptively optimizing the preprocessing channel weights and feature fusion methods according to data characteristics. This improves the model's feature extraction capability and cross-scenario generalization performance while ensuring prediction accuracy.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: A rapid detection system for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification includes a multi-channel differential spectrum construction module, a channel attention weighting module, a sequence encoding module, a contrastive learning feature amplification module, and an end-to-end optimization module; The multi-channel differential spectrum construction module extracts the gradient and curvature features of the original spectrum through first-order and second-order differential preprocessing, respectively, to form complementary original spectrum, first-order differential spectrum and second-order differential spectrum feature representations, and splices them in the channel dimension. The channel attention weighting module employs a dual-path pooling strategy and a multilayer perceptron to adaptively generate weight coefficients for each preprocessed channel of the multi-channel differential spectrum, thereby enhancing key feature channels and suppressing redundant channels. The sequence encoding module is a pure encoder Transformer structure, which uses its multi-head self-attention mechanism to capture the long-range nonlinear correlation of weighted spectral features between wavelength sequences and extract spectral patterns that are deeply related to the physicochemical properties of biomass. The contrastive learning feature amplification module constructs a dual-branch structure of online network and target network based on the BYOL framework. It optimizes the spectral feature representation through a negative sample-free contrastive learning strategy, thereby improving the robustness and generalization ability of the model in small sample scenarios. The end-to-end optimization module combines physical constraints and task loss to jointly train the sequence encoding module and the contrastive learning feature amplification module, and finally outputs quantitative prediction results of the physicochemical properties of biomass targets.
[0008] This invention also provides a rapid detection method for biomass adaptive channel splicing based on BYOL contrast learning feature amplification, and a rapid detection system for biomass adaptive channel splicing based on BYOL contrast learning feature amplification, comprising the following steps: Step S1: Construct a multi-channel differential spectrum input, and extract gradient features and curvature features from the original spectrum through first-order and second-order differential preprocessing to form a complementary feature representation space; Step S2: Adaptively weight the multi-channel features using a channel attention mechanism, and calculate the weight coefficients of each preprocessed channel through a compression-excitation network to enhance key feature channels and suppress redundant channels; Step S3: Input the weighted multi-channel features into the pure encoder Transformer structure, capture the long-range nonlinear correlation between wavelengths through the multi-head self-attention mechanism, and extract the spectral patterns that are deeply related to the physicochemical properties of biomass; Step S4: Perform negative sample-free comparative learning based on the BYOL framework. Optimize feature representation through interaction between the online network and the target network to improve the robustness and generalization ability of the model in small sample scenarios. Step S5: Combine physical constraints and task loss to perform end-to-end training, output quantitative prediction results of biomass characteristics, and achieve high-precision, interpretable non-destructive testing.
[0009] Preferably, the method further includes a data augmentation step, which generates multiple augmented samples from the original samples by simulating the baseline drift, multiplicative noise, and additive noise error sources of the near-infrared spectroscopy measurements using a quadratic polynomial and combining it with random noise injection.
[0010] Preferably, in step S2, the channel attention mechanism employs a dual-path pooling strategy, fusing the output features of global average pooling and global max pooling, and its weight generation function is expressed as: in, and Let $\mathbf$ and $\mathbf$ represent the global average and maximum values of the $c$-th channel, respectively. and Here are the parameters for the fully connected layer, and σ is the Sigmoid function.
[0011] Preferably, in step S3, the pure encoder Transformer structure employs a scaled dot product attention mechanism, and its attention output is represented as: Where Q, K, and V are the query, key, and value vectors, respectively. This is a dimensional adjustment factor.
[0012] Preferably, in step S3, the pure encoder Transformer structure contains 4 encoder layers, each containing 4 attention heads, with an embedding dimension of 48.
[0013] Preferably, in step S4, the BYOL framework adopts a negative sample-free contrastive learning strategy, updates the target network parameters through exponential moving average, avoids interference from negative sample quality on feature learning, and improves the model's adaptability in scenarios with a small amount of labeled data.
[0014] Preferably, the performance evaluation of end-to-end training in step S5 is based on root mean square error, mean relative error, and correlation coefficient.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the above-described solution, the present invention can achieve high-precision non-destructive testing of biomass characteristics, and has the following beneficial effects: 1. Significantly enhanced feature extraction capability: Through multi-channel differential splicing and attention weighting, overlapping spectral peaks are effectively decoupled. The model's RMSE is reduced to 45.6 and MAPE reaches 5.01% on the test set, with an error reduction of more than 10% compared to traditional methods. 2. Improved model generalization performance: By combining the BYOL framework and data augmentation strategies, the field experiment error is less than 5% in a database of only 300 samples, making it suitable for a variety of biomass categories and complex application scenarios; 3. Resource allocation optimization: The channel attention mechanism automatically selects high-value feature channels, reducing redundant calculations and improving model running efficiency; 4. Enhanced interpretability of the mechanism: The attention weight distribution is consistent with the absorption characteristics of biomass functional groups, providing a physical basis for the analysis of spectral features; 5. Expanded System Applicability: The modular network structure supports migration and application across different spectrometer platforms and biomass types, promoting the standardization and industrialization of non-destructive testing technology.
[0016] This method and system have the following advantages: The system described in this invention has excellent multidimensional spectral feature extraction and decoupling capabilities, which significantly improves prediction accuracy. At the same time, it automatically optimizes feature channel configuration, reduces computational redundancy, and thus improves operating efficiency.
[0017] The system described in this invention maintains strong generalization performance even under limited sample conditions, can be adapted to various biomass and complex working conditions, and can be easily migrated to different types of spectral platforms and biomass types. Attached Figure Description
[0018] To more clearly illustrate the technical solution of the present invention, the present invention will be further described below with reference to the accompanying drawings. It should be understood that the accompanying drawings are only used to illustrate embodiments of the present invention and do not constitute a limitation on the scope of protection of the present invention.
[0019] Figure 1 This is a schematic flowchart of a rapid biomass detection method based on BYOL contrastive learning feature amplification, provided as an embodiment of the present invention.
[0020] Figure 2 This is a detailed schematic diagram illustrating the data preparation stage of a rapid biomass detection method based on BYOL contrastive learning feature amplification, provided as an embodiment of the present invention.
[0021] Figure 3 A schematic diagram of the adaptive channel splicing attention mechanism module provided in an embodiment of the present invention. Detailed Implementation
[0022] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0023] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0024] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0025] Example 1 like Figure 1 and Figure 2 As shown in this embodiment, a rapid detection method for adaptive channel stitching based on BYOL contrast learning feature amplification is provided for biomass spectral datasets. This framework implements a system that uses near-infrared spectroscopy as input to simultaneously predict the content of five key physicochemical indicators in biomass samples: cellulose, hemicellulose, lignin, ash, and volatile matter. Its core workflow encompasses data loading, preprocessing, feature enhancement, model building, training optimization, and result visualization, forming a complete, end-to-end analysis and prediction system. The rapid detection system for biomass based on BYOL contrast learning feature amplification adaptive channel stitching includes the following steps: Step S1: Construct a multi-channel differential spectral input. Gradient and curvature features from the original spectrum are extracted through first-order and second-order differential preprocessing, forming a complementary feature representation space. The data augmentation step addresses baseline drift, multiplicative noise, and additive noise error sources in near-infrared spectroscopy measurements. Baseline drift is simulated using a quadratic polynomial, combined with random noise injection, to generate multiple augmented samples from the original samples.
[0026] This module is responsible for reading the raw spectra and target values from the structured file (Excel), and completing the data partitioning and standardization to provide standardized input for subsequent modeling.
[0027] This module defines the `load_biomass_excel` function, which reads data from an Excel file at a specified path, treating all columns except the last five as spectral data, while the last five columns serve as multi-objective prediction values. This design clearly defines the task of performing simultaneous regression predictions on five components of biomass: cellulose, hemicellulose, lignin, ash, and volatile matter. The data is converted to float32 type to optimize computational efficiency and memory usage.
[0028] In the main process, the first step is to standardize and split the data. StandardScaler is used to perform Z-score standardization on both the spectral data (spectra) and the target values (targets), which involves subtracting the mean and dividing by the standard deviation to bring all feature dimensions to the same order of magnitude, thereby accelerating model convergence and improving training stability. Then, train_test_split is used to randomly divide the standardized dataset into training and test sets in an 8:2 ratio to ensure the objectivity of model evaluation.
[0029] The custom dataset and data loader module defines the PyTorch standard data interface, enabling batch loading and real-time enhancement of data, serving as the data pipeline for model training.
[0030] The BiomassDataset class, which inherits from torch.utils.data.Dataset, converts the input NumPy array into a PyTorch tensor during initialization. The key operation unsqueeze(1) adds a channel dimension to the spectral data, transforming the shape from (N, L) to (N, 1, L) to accommodate the input shape (batch_size, channels, length) requirements of a one-dimensional convolutional layer. This class also supports an optional transform parameter for applying data augmentation strategies on the fly during data readout.
[0031] Step S2: Adaptively weight the multi-channel features using a channel attention mechanism, and calculate the weight coefficients of each preprocessed channel through a compression-excitation network to enhance key feature channels and suppress redundant channels; The channel attention mechanism employs a dual-path pooling strategy, fusing the output features of global average pooling and global max pooling. Its weight generation function is expressed as: in, and Let $\mathbf$ and $\mathbf$ represent the global average and maximum values of the $c$-th channel, respectively. and Here are the parameters for the fully connected layer, and σ is the Sigmoid function.
[0032] The one-dimensional convolutional neural network regression module constructs the core deep learning model for spectral feature extraction and multi-objective regression. First, there's the Regressor1D class, inherited from nn.Module, a convolutional neural network specifically designed for one-dimensional sequence data, particularly spectral data. Its structure is shown in the deployment steps below.
[0033] Feature extraction layer: Contains two one-dimensional convolutional blocks. Each block consists of a convolutional layer (Conv1d), a batch normalization layer (BatchNorm1d), and a ReLU activation function in sequence. The kernel size is 11 for both convolutions, and padding=5 is used to keep the sequence length constant, aiming to capture the spatial correlation of the spectrum within a local wavelength range.
[0034] Global feature aggregation layer: The entire wavelength sequence of each feature channel is globally averaged using AdaptiveAvgPool1d(1), which compresses the variable-length spectral features of each sample into a fixed-length feature vector (dimension 64).
[0035] Step S3: Input the weighted multi-channel features into the pure encoder Transformer structure, capture the long-range nonlinear correlation between wavelengths through the multi-head self-attention mechanism, and extract the spectral patterns that are deeply related to the physicochemical properties of biomass; Furthermore, the pure encoder Transformer structure employs a scaled dot product attention mechanism, and its attention output is represented as: Where Q, K, and V are the query, key, and value vectors, respectively. This is a dimensional adjustment factor.
[0036] Furthermore, the pure encoder Transformer structure contains 4 encoder layers, each containing 4 attention heads, with an embedding dimension of 48.
[0037] Finally, the regression output layer flattens the pooled feature vectors and passes them through two fully connected layers. The first fully connected layer (Linear(64, feat_dim)) further performs feature transformation and dimensionality reduction; the second fully connected layer (Linear(feat_dim, n_out)) directly outputs the predicted values of the five target components.
[0038] This model leverages the local connectivity and weight sharing properties of convolution operations to efficiently learn discriminative hierarchical feature representations from high-dimensional spectral data, and finally maps them to the target space through fully connected layers.
[0039] Step S4: Perform negative-sample-free contrastive learning based on the BYOL framework. This involves optimizing feature representations through interaction between the online network and the target network, thereby improving the model's robustness and generalization ability in scenarios with limited sample size. The BYOL framework employs a negative-sample-free contrastive learning strategy, updating the target network parameters through exponential moving averages to avoid interference from negative sample quality on feature learning and improve the model's adaptability in scenarios with limited labeled data.
[0040] This module defines the core spectral differential data augmentation strategy, the `DiffAugment` class. In the `__call__` method, the following operations are performed on a single input spectral sample (shape (1, L)): Channel dimension removal (`squeeze`); calculation of the first derivative of the spectrum (`torch.diff`), using the `prepend` parameter to ensure the sequence length remains unchanged after differentiation; standardization of the differentiated sequence by subtracting the mean and dividing by the standard deviation to stabilize its numerical distribution; and stacking the original spectrum and the standardized first-derived spectrum along the channel dimension (`torch.stack`) to generate a two-channel feature map of shape (2, L). This strategy, without increasing the number of samples, enriches the model's input features by introducing differential information reflecting the rate of spectral change and inflection points. This helps the network simultaneously capture both the absolute intensity and local morphological changes of the spectrum, enhancing the model's ability to perceive subtle spectral differences and its generalization capabilities.
[0041] Finally, this module constructs a data loader, using DataLoader to encapsulate BiomassDataset, enabling batch data loading. The training set loader sets shuffle=True to shuffle the data order at the beginning of each epoch, preventing the model from learning spurious patterns caused by data sorting. The test set loader, however, does not require shuffling.
[0042] Step S5: Perform end-to-end training by combining physical constraints and task loss to output quantitative prediction results of biomass properties, achieving high-precision and interpretable non-destructive testing. Performance evaluation of end-to-end training is based on root mean square error, mean relative error, and correlation coefficient.
[0043] The model training and evaluation process encapsulates the standard workflow for model training iterations and performance evaluation. A `train_epoch` function is defined, responsible for forward propagation, loss calculation, and backpropagation for a single training epoch. It iterates through the training data loader, moves data to the specified computing device, obtains predicted values through the model, and calculates the mean squared error between the predicted and true values as the loss function. It then clears the gradients, performs backpropagation to calculate the gradients, and calls the optimizer to update the model parameters. This function returns the average loss for all samples in that epoch.
[0044] The module also includes an `evaluate` function, enabling the code to run the model in evaluation mode without calculating gradients. It iterates through the test data loader, collecting predicted and true values for all samples. A key step involves using the `inverse_transform` method of the passed-in `scaler_y` target value normalizer to inversely transform the standardized predicted and true values back to their original physical dimensions. This ensures that the evaluation metrics have practical physical meaning and are comparable across different targets, calculating the model performance evaluation metric for each target component.
[0045] Model performance is evaluated based on three complementary metrics: root mean square error (RMSE), mean absolute percentage error (MAPE), and correlation coefficient (r). The calculation methods are shown below: In the formula: For the test set label values, Its mean, For the predictions corresponding to the model, Its mean, The number of samples in the test set is denoted as . Smaller RMSE, smaller MAPE, and a value closer to 1 (r) indicate better model performance. This paper uses a five-fold cross-validation method to evaluate model performance, and averages the RMSE, MAPE, and r values calculated for each fold as the final performance metric.
[0046] The visualization and results analysis module generates a series of charts to provide multi-faceted visualization analysis of the data processing flow, model training dynamics, and final prediction performance. This mainly involves the construction and invocation of the following functions.
[0047] `plot_initial_spectra` draws the original near-infrared spectrum curves, visually displaying the overall spectral morphology of the dataset and the differences between samples; `plot_augmented_spectra` displays the original spectra and their corresponding first derivatives side-by-side, clearly showing the dual-channel input generated for each sample by the data augmentation strategy; `plot_loss_curve` plots the curve of the training loss decreasing with the iteration cycle, used to monitor model convergence and detect overfitting; `plot_fit_results` generates scatter plots for each prediction component, with the x-axis representing the true value and the y-axis representing the model's predicted value, and adds a reference line for y=x. This plot is an intuitive tool for evaluating the model's predictive performance and bias. The R-values are marked on the plot. 2 The value quantifies the degree of linear correlation between the predicted value and the actual value.
[0048] Finally, the main process and system integration. The main function `if __name__ == '__main__'` integrates all the above modules, forming a complete execution pipeline.
[0049] The environment configuration includes detecting and setting up the computing device (GPU / CPU) and creating the output directory. Data loading involves loading the raw data and calling visualization functions to generate the initial and enhanced spectra. Data preprocessing standardizes the spectra and target values, and proportionally divides them into training and test sets. A data pipeline is built to instantiate the dataset and data loader, applying DiffAugment data augmentation to both the training and test sets. Finally, the model is initialized, a one-dimensional CNN regression model is instantiated, and it is moved to the computing device while defining the Adam optimizer.
[0050] The model training method mainly involves iterative training with fixed periods, recording and saving the training loss for each period, evaluating the final model performance on the test set to achieve model evaluation and visualization, and plotting loss curves and scatter plots of fitting results for each component. Finally, the destandardized true values, predicted values, and model evaluation metrics for each component are summarized and saved to an Excel file, and a concise evaluation report is printed in the console.
[0051] In summary, this example implements a rigorous, deep learning-based spectral analysis workflow. Its innovation lies in using the first-order derivative of the spectrum as an enhancement feature, inputting it along with the original spectrum into the network. Leveraging the powerful local feature extraction capabilities of one-dimensional CNNs, it achieves high-precision, simultaneous quantitative prediction of complex multi-component biomass. The entire system features automated processing, visualized analysis, and standardized output, providing a reliable computational framework for the application of near-infrared spectroscopy in rapid biomass detection.
[0052] Example 2 like Figure 3 As shown, this example provides an end-to-end biomass component prediction system based on channel attention and Transformer deep neural network architecture, specifically designed for target regression prediction tasks of biomass physicochemical properties (especially DP values) based on near-infrared spectroscopy. This framework integrates advanced data augmentation, multi-scale feature fusion, attention mechanisms, and sequence modeling techniques, constructing a complete end-to-end analysis workflow from raw spectral data to prediction results.
[0053] The core architecture of this example primarily adopts a four-stage design concept of "feature enhancement - attention weighting - sequence modeling - regression prediction," aiming to extract and integrate discriminative information from the spectrum from multiple dimensions. Compared with conventional methods, the advantages of this framework are reflected in: 1. Multichannel spectral representation: Using spectral information at three different scales—the original spectrum, the first derivative, and the second derivative—as parallel input channels. 2. Physically Heuristic Data Augmentation: Design augmentation strategies based on actual error sources in spectral measurements, such as baseline drift and noise.
[0054] 3. Two-layer attention mechanism: Channel attention focuses on the importance of information at different scales, and Transformer self-attention captures long-range dependence between wavelengths. 4. Interpretable Design: Achieving a physical explanation of model decisions through visualization of attention weights. Based on the above significant advantages, we will first elaborate on the modular architecture steps in this example.
[0055] Step S1: Construct a multi-channel differential spectrum input, and extract gradient features and curvature features from the original spectrum through first-order and second-order differential preprocessing to form a complementary feature representation space.
[0056] The `load_spectral_data` method loads spectral data from a structured Excel file. A flexible column indexing strategy is employed: the first column is assumed to be the sample identifier, the last column to be the target DP value, and all intermediate columns to be full-band spectral data. This design ensures the code's adaptability to different data formats.
[0057] The spectral derivatives method is used to perform spectral differential calculations, realizing a core preprocessing technique in chemometrics. The first derivative eliminates baseline drift and highlights inflection points in the spectrum; the second derivative further separates overlapping peaks and identifies shoulder peaks. Mathematically, this is expressed as calculating the numerical gradient of the spectral curve for each sample along the wavelength axis (axis=1).
[0058] The channel concatenation method is a key step in constructing multi-scale spectral feature fusion. It stacks three signals with different physical meanings—the original spectrum, the first derivative, and the second derivative—along the channel dimension (axis=1) to form a three-dimensional tensor (number of samples × 3 × number of wavelengths). This fusion strategy allows the model to simultaneously utilize the absolute intensity, slope variation, and curvature information of the spectrum.
[0059] The `data_augmentation` method implements data augmentation based on physical priors. Targeting the main error sources in near-infrared spectroscopy measurements—baseline drift (simulated using a quadratic polynomial), multiplicative noise, and additive noise—it generates multiple augmented samples from a single original sample by randomly combining these perturbation factors, significantly expanding the training dataset and improving model robustness.
[0060] During the training phase, to support subsequent BYOL contrastive learning, this embodiment calls the data_augmentation method twice independently for the same original sample to generate two different random augmented views.
[0061] The `prepare_data` method integrates a pipeline of all preprocessing steps. The execution order is: data loading → standardization → differential calculation → channel concatenation → dataset partitioning → training set augmentation. It's important to note that the algorithm in this module only performs data augmentation on the training set; the test set maintains its original distribution to evaluate true generalization ability.
[0062] By using a custom dataset class (SpectralDataset), a standard PyTorch data interface is constructed, enabling batch data loading and efficient memory management. This dataset class inherits from torch.utils.data.Dataset and encapsulates spectral data and corresponding labels. The core method __getitem__ returns a tensor representation of a single sample based on its index, supporting the application of real-time augmentation transformations during data loading.
[0063] Step S2: Adaptively weight the multi-channel features using a channel attention mechanism. Calculate the weight coefficients of each preprocessed channel using a compression-excitation network to enhance key feature channels and suppress redundant channels.
[0064] The multi-channel features obtained in step S1 are input into the ChannelAttention module to dynamically learn and weight the importance of channels at different spectral scales. The core idea is that the original spectrum, first-order derivative, and second-order derivative contain complementary but not equivalent information, and their relative importance should be dynamically adjusted with the samples and the learning task.
[0065] A dual-pooling strategy is employed. First, Global Average Pooling (GAP) is used to capture the global statistical features of the channels. Second, Global Max Pooling (GMP) is used to focus on the most salient responses of the channels. The two are concatenated to provide a more comprehensive channel description, forming a dual descriptor vector.
[0066] The dual-pooling features are processed through a two-layer bottleneck structure, a multilayer perceptron (MLP). First, they are compressed through a fully connected layer, then activated by ReLU and expanded back to the original number of channels. Finally, the channel weights between 0 and 1 are output through the Sigmoid function.
[0067] The learned channel weights are multiplied channel by channel with the original input features to achieve feature selection. Features in high-weight channels are enhanced, while features in low-weight channels are suppressed. This module not only improves model performance, but its output channel weights also provide an interpretable perspective for model decision-making.
[0068] Step S3: Input the weighted multi-channel features into the pure encoder Transformer structure, capture the long-range nonlinear correlation between wavelengths through the multi-head self-attention mechanism, and extract the spectral patterns that are deeply related to the physicochemical properties of biomass.
[0069] The weighted features are input into the Transformer encoder module. In spectral analysis, complex physicochemical relationships exist between absorption peaks at different wavelengths, and the local receptive field of traditional convolutional neural networks struggles to capture these global interactions. The success of the Transformer architecture in natural language processing has inspired its application in sequence data modeling, primarily to model the long-range dependencies between any two wavelength points in a spectral sequence and capture complex nonlinear spectral-property correlations.
[0070] The Transformer itself is equivariant and lacks sequence order information. Sine-cosine position encoding injects absolute and relative position information into each wavelength position. The encoding formula is: in, It is the index of the position in the sequence starting from 0. It is the dimension of the model, i.e., the length of the embedding vector. The range of values is Dimension index. Even-numbered dimensions. Encoding using the sine function, odd-dimensional Cosine function encoding is used. This encoding allows the model to easily learn relative positional relationships.
[0071] Features are processed through multiple Transformer encoder layers. Each layer contains a standard multi-head self-attention mechanism and a feedforward network, and training stability is ensured through residual connections and layer normalization. The multi-head self-attention mechanism allows the model to extract features from different subspaces and capture global dependencies.
[0072] After Transformer encoding, each position in the sequence contains global context information. The variable-length sequence is compressed into a fixed-length global feature vector using adaptive average pooling, and then mapped to the target space through a three-layer fully connected network to extract high-level spectral modes.
[0073] Step S4: Perform negative sample-free comparative learning based on the BYOL framework. Optimize feature representation through interaction between the online network and the target network to improve the robustness and generalization ability of the model in small sample scenarios.
[0074] This system is built based on the BYOL framework. Specifically, two feature encoding networks with identical structures are constructed, referred to as the online network and the target network, respectively. Both networks include the aforementioned channel attention module (step S2) and Transformer encoder (step S3). An additional predictor, typically a 2-3 layer multilayer perceptron, is added after the online network encoder. The parameters of the target network are not updated directly through gradient descent, but rather through the exponential moving average (EMA) of the online network parameters, with the update formula: θ_target ← τ * θ_target + (1-τ) * θ_online, where τ is the momentum coefficient (e.g., 0.99). During training, the two augmented views generated in step S1 are input into the online network and the target network, respectively. The online network outputs the predicted feature z_online, and the target network outputs the feature z_target. By minimizing the mean squared error between the online network predictor output and the target network output after gradient descent is stopped, a feature consistency contrastive loss L_con = || predictor(z_online) - sg(z_target) || is constructed. 2 This negative-sample-free contrastive learning aims to enable the network to learn data augmentation-invariant feature representations, thereby improving feature quality and model generalization ability.
[0075] Step S5: Combine physical constraints and task loss to perform end-to-end training, output quantitative prediction results of biomass characteristics, and achieve high-precision, interpretable non-destructive testing.
[0076] The model's forward propagation process is clearly layered as follows: a channel attention layer, which takes input shape (batch, 3, L) and assigns weighted features and channel weights; an embedding projection layer, which projects the spectral length dimension onto the d_model dimension latent space; a Transformer encoding layer, which performs self-attention computation in the wavelength dimension to model global dependencies; and a pooling and regression layer, which aggregates sequence information and outputs the final prediction. This design achieves hierarchical feature learning from local to global and from multi-scale to ensemble.
[0077] An end-to-end training framework (ModelTrainer) is adopted. A joint loss function is used for optimization within the standard mini-batch stochastic gradient descent training loop.
[0078] The model is optimized using a standard mini-batch stochastic gradient descent training loop, employing mean squared error loss, the Adam optimizer, and adaptive learning rate scheduling based on validation loss (ReduceLROnPlateau). The learning rate is halved when the validation loss fails to decrease for 10 consecutive epochs. The optimal validation loss is tracked, and model checkpoints are saved only when performance improves to prevent overfitting. Finally, a comprehensive evaluation using multiple metrics ensures the overall performance of the model.
[0079] This embodiment demonstrates that, in the application scenario of rapid detection of biomass physicochemical properties based on intelligent algorithms, compared with traditional chemical detection and measurement methods, the machine learning-driven adaptive channel splicing method described in this invention can achieve more effective prediction of physicochemical properties under complex conditions and weak feature peaks, thereby bringing about a significant improvement in rapid detection results.
[0080] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A rapid detection system for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification, characterized in that, It includes a multi-channel differential spectroscopy construction module, a channel attention weighting module, a sequence encoding module, a contrastive learning feature amplification module, and an end-to-end optimization module; The multi-channel differential spectrum construction module extracts the gradient and curvature features of the original spectrum through first-order differential and second-order differential preprocessing, respectively, to form complementary original spectrum, first-order differential spectrum and second-order differential spectrum feature representations, and splices them in the channel dimension; The channel attention weighting module employs a dual-path pooling strategy and a multilayer perceptron to adaptively generate weight coefficients for each preprocessed channel of the multi-channel differential spectrum, thereby enhancing key feature channels and suppressing redundant channels. The sequence encoding module is a pure encoder Transformer structure, which uses its multi-head self-attention mechanism to capture the long-range nonlinear correlation of weighted spectral features between wavelength sequences and extract spectral patterns that are deeply related to the physicochemical properties of biomass. The contrastive learning feature amplification module constructs a dual-branch structure of online network and target network based on the BYOL framework. It optimizes the spectral feature representation through a negative sample-free contrastive learning strategy, thereby improving the robustness and generalization ability of the model in small sample scenarios. The end-to-end optimization module combines physical constraints and task loss to jointly train the sequence encoding module and the contrastive learning feature amplification module, and finally outputs quantitative prediction results of the physicochemical properties of biomass targets.
2. A rapid detection method for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification, characterized in that, A rapid detection system for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification includes the following steps: Step S1: Construct a multi-channel differential spectrum input, and extract gradient features and curvature features from the original spectrum through first-order and second-order differential preprocessing to form a complementary feature representation space; Step S2: Adaptively weight the multi-channel features using a channel attention mechanism, and calculate the weight coefficients of each preprocessed channel through a compression-excitation network to enhance key feature channels and suppress redundant channels; Step S3: Input the weighted multi-channel features into the pure encoder Transformer structure, capture the long-range nonlinear correlation between wavelengths through the multi-head self-attention mechanism, and extract the spectral patterns that are deeply related to the physicochemical properties of biomass; Step S4: Perform negative sample-free comparative learning based on the BYOL framework. Optimize feature representation through interaction between the online network and the target network to improve the robustness and generalization ability of the model in small sample scenarios. Step S5: Combine physical constraints and task loss to perform end-to-end training, output quantitative prediction results of biomass characteristics, and achieve high-precision, interpretable non-destructive testing.
3. The rapid detection method for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification according to claim 2, characterized in that, It also includes a data augmentation step, which targets the baseline drift, multiplicative noise and additive noise error sources of near-infrared spectroscopy measurements. The data augmentation step simulates the baseline drift using a quadratic polynomial and combines it with random noise injection to generate multiple enhanced samples from the original samples.
4. The rapid detection method for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification according to claim 2, characterized in that, In step S2, the channel attention mechanism employs a dual-path pooling strategy, fusing the output features of global average pooling and global max pooling. Its weight generation function is expressed as: in, and Let $\mathbf$ and $\mathbf$ represent the global average and maximum values of the $c$-th channel, respectively. and Here are the parameters for the fully connected layer, and σ is the Sigmoid function.
5. The rapid detection method for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification according to claim 2, characterized in that, In step S3, the pure encoder Transformer structure employs a scaled dot product attention mechanism, and its attention output is represented as follows: Where Q, K, and V are the query, key, and value vectors, respectively. This is a dimensional adjustment factor.
6. The rapid detection method for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification according to claim 5, characterized in that, In step S3, the pure encoder Transformer structure contains 4 encoder layers, each containing 4 attention heads, with an embedding dimension of 48.
7. The rapid detection method for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification according to claim 2, characterized in that, In step S4, the BYOL framework adopts a negative sample-free contrastive learning strategy, which updates the target network parameters through exponential moving average to avoid interference from negative sample quality on feature learning and improve the model's adaptability in scenarios with a small amount of labeled data.
8. A rapid detection method for biomass adaptive channel splicing based on BYOL contrastive learning feature amplification according to claim 2, characterized in that, In step S5, the performance evaluation of end-to-end training is based on the root mean square error, the average relative error, and the correlation coefficient.