Raman spectrum algorithm based on multi-channel one-dimensional convolutional neural network and spatial pyramid pooling technology

By combining multi-channel one-dimensional convolutional neural network with spatial pyramid pooling technology in Raman spectroscopy analysis, the problem that traditional methods are difficult to deal with high-dimensional spectral data is solved, and higher classification and prediction accuracy and robustness are achieved.

CN120124683APending Publication Date: 2025-06-10CHINA JILIANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510194279.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional Raman spectroscopy analysis methods are difficult to effectively process high-dimensional spectral data, cannot fully mine potential information in the data, and are not stable and accurate for noise interference and data loss.

Method used

The Raman spectroscopy algorithm SPP-1D based on multi-channel one-dimensional convolutional neural network (1D-CNN) and spatial pyramid pooling (SPP) technology is adopted. Local features are extracted through the convolution layer, combined with the SPP layer to pool on different scales, and multi-scale feature representations are generated to improve the learning ability and robustness of the model.

Benefits of technology

It significantly improves the classification and prediction accuracy of Raman spectral data, enhances the learning ability of complex spectral signals, improves the generalization ability and robustness of the model, especially in the face of noise interference and data loss, showing strong stability and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005280992880000021
    Figure BDA0005280992880000021
  • Figure BDA0005280992880000022
    Figure BDA0005280992880000022
  • Figure BDA0005280992880000023
    Figure BDA0005280992880000023
Patent Text Reader

Abstract

The invention discloses a Raman spectrum training and prediction algorithm SPP-1D of a multichannel one-dimensional convolutional neural network based on spatial pyramid pooling. According to the method, local features of spectral data are extracted through 1D-CNN, and features are pooled on multiple scales by using spatial pyramid pooling (SPP), so that global feature representation is generated. Compared with the traditional principal component analysis (PCA) and linear discriminant analysis (LDA) methods, the method can effectively capture the complex nonlinear relationship of the spectral data and enhance the anti-noise performance. The SPP-1D supports multi-channel input, a plurality of characteristic channels of the Raman spectrum can be processed at the same time, correlation and complementarity between the channels are fused, and more comprehensive characteristic description is achieved. The algorithm is suitable for application scenes such as single substance identification, complex mixture component analysis, spectral imaging and the like, and the precision, robustness and generalization ability of Raman spectrum data processing are remarkably improved. According to the invention, deep learning and a signal processing technology are creatively combined, and an efficient spectral analysis tool is provided for the fields of material science, chemical analysis and biomedicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of machine learning and computer vision, and particularly to a Raman spectroscopy algorithm SPP-1D based on multi-channel one-dimensional convolutional neural network (1D-CNN) and spatial pyramid pooling (SPP) technology, which is particularly suitable for the processing and analysis of high-dimensional spectral data. In particular, it relates to a deep learning method for feature extraction and classification using Raman spectroscopy data. SPP-1D is innovatively applied to fields such as materials science, chemical analysis, and biomedicine. Background Art

[0002] Raman spectroscopy is a spectral technology based on the Raman scattering phenomenon, used to study the molecular vibration, rotation and other characteristics of materials. Raman spectroscopy can provide detailed molecular structure information of materials, but due to the high-dimensionality and complexity of its data, traditional analysis methods are difficult to effectively process these data. Traditional data analysis methods such as principal component analysis (PCA) and linear discriminant analysis (LDA) often fail to fully exploit the potential information in the data. Therefore, combining deep learning technology to analyze and process Raman spectroscopy data has become an important research direction.

[0003] Traditional convolutional neural networks (CNNs) usually require input images to have a fixed size. However, when processing one-dimensional sequence data of different sizes, this requirement limits the flexibility of the network. Spatial pyramid pooling (SPP) technology solves the fixed-size problem by performing pooling at different scales and can effectively extract multi-scale features. In the prior art, SPP technology is mainly applied to two-dimensional image data and less applied to one-dimensional sequence data. Therefore, it is necessary to provide an improved SPP technology suitable for processing and analyzing one-dimensional sequence data. Summary of the Invention

[0004] The object of the present invention is to provide a Raman spectroscopy algorithm SPP-1D based on multi-channel one-dimensional convolutional neural network (1D-CNN) and spatial pyramid pooling technology, which can effectively process complex spectral data and greatly improve the accuracy of classification and prediction.

[0005] To achieve the above object of the invention, the main technical solutions adopted are as follows:

[0006] A Raman spectroscopy algorithm based on multi-channel one-dimensional convolutional neural network and spatial pyramid pooling technology, characterized in that it is carried out according to the following steps:

[0007] Step 1: Construct a multi-channel one-dimensional convolutional neural network (1D-CNN) model. Use multiple convolutional kernels to perform convolutional operations on the input data and extract local features. Utilize the Spatial Pyramid Pooling (SPP) layer to perform pooling operations at different scales, and concatenate the pooling results at each scale to obtain the final multi-scale feature representation.

[0008] Step 2: Normalize the input data to have the same scale, segment and augment the data to improve the generalization ability of the model. Select appropriate loss functions (cross-entropy loss and mean squared error), as well as an optimization algorithm (Adam), and evaluate the model performance on the validation set.

[0009] Step 3: The output layer is used to generate prediction results. Select appropriate activation functions (Mish and GeLU).

[0010] Step 4: Initialize the model parameters, monitor the accuracy and loss values during the training process, adjust the model parameters in real-time, and find the optimal hyperparameters.

[0011] Step 5: After the training is completed, use matplotlib to plot the curves of the loss value and accuracy during the training process.

[0012] Step 6: Normalize and smooth the input spectral data, use the trained model to predict the preprocessed spectral data, and perform visual display.

[0013] Specifically, in Step 1, construct a multi-channel one-dimensional convolutional neural network (1D-CNN) model. Use multiple convolutional kernels to perform convolutional operations on the input data to extract local features, and then utilize the Spatial Pyramid Pooling (SPP) layer to perform pooling operations at different scales, extract multi-scale features and concatenate the pooling results to obtain the final multi-scale feature representation. The specific steps are as follows:

[0014] Step (1.1): Calculate and determine the start and end positions of the pooling window:

[0015]

[0016] where ix and iy are the indices of the current pooling window, i is the level of the pooling layer, num cols and num rows are the number of columns and rows of the input feature map, respectively.

[0017] Step (1.2): Perform a pooling operation on the cropped pooling region. Use the following formula to calculate the maximum value after pooling:

[0018] PooledVal = max(x crop , axis=(1, 2))………(5)

[0019] where x crop is the cropped pooling area, and PooledVal is the maximum value after pooling.

[0020] Step (1.3): Concatenate the pooling results at different scales to form the final multi-scale feature representation. The output shape is determined by the following formula:

[0021] output shape = (input shape [0], num outputs × channels)…………(6)

[0022] where input shape is the shape of the input feature map, num outputs is the number of outputs per channel, and channels is the number of channels.

[0023] Among them, the multi-channel one-dimensional convolutional neural network model includes the following structures:

[0024] Convolutional layer: Extract local features of the input signal through the combination of the convolutional layer, batch normalization layer, activation layer, and pooling layer. The convolution operation formula is:

[0025]

[0026] where x is the input signal, ω is the convolution kernel, K is the size of the convolution kernel, and y is the convolution result.

[0027] SPP layer: Perform multi-scale pooling on the feature map output by the convolution and flatten the result as the input to the fully connected layer.

[0028] Fully connected layer: Further process the extracted features through the fully connected layer and generate the final feature representation for classification or prediction.

[0029] On this basis, in Step 2, the input data is normalized and reshaped into the shape required by the model. The following formula is used for normalization:

[0030]

[0031] where X is the input data, and X normal is the normalized data.

[0032] Furthermore, in Step 3, select an appropriate activation function:

[0033] Mish activation function:

[0034] X Mish = x · tanh(ln(1 + e x))……………(11)

[0035] In the formula, x is the input value, and X Mish is the output value of the Mish activation function.

[0036] Gaussian Error Linear Unit (GeLU) activation function:

[0037]

[0038] In the formula, x is the value input to the activation function, and X GeLU is the output value of the GeLU activation function.

[0039] Furthermore, in step four, the model training and optimization steps include:

[0040] Step (4.1), training the deep learning model using the training data;

[0041] Step (4.2), monitoring the accuracy and loss value during the training process in real time, and using the Adam optimization algorithm to adjust the model parameters. The Adam optimization formula is:

[0042] m t = β 1 m t-1 +(1 - β 1 )g t …………(14)

[0043]

[0044]

[0045] In the formula, m t and v t are the momentum and root mean square gradient, β 1 and β 2 are the momentum decay rates, g t is the gradient, η is the learning rate, ∈ is the smoothing term, θ t is the parameter, and are the momentum and root mean square gradient after bias correction.

[0046] Step (4.3), adjusting the hyperparameters, such as the learning rate, batch size, and number of network layers, through grid search or random search methods.

[0047] After that, step five displays the training progress and model performance through a visualization interface, and uses matplotlib to plot the change curves of the loss value and accuracy during the training process.

[0048] Finally, in step six, the prediction of spectral data includes smoothing the spectral data and using the trained model for prediction. The Whittaker smoothing and airPLS methods are used:

[0049] Whitaaker smoothing algorithm:

[0050] z = (W + λD T D) -1 (Wx)…………(19)

[0051] In the formula, W is the weight matrix, λ is the smoothing parameter, D is the difference matrix, x is the input signal, and z is the smoothed signal.

[0052] airPLS algorithm:

[0053]

[0054] In the formula, w is the weight, d is the residual, and i is the number of iterations.

[0055] The prediction result display includes generating the final prediction result according to the probability distribution or regression result output by the model, and using a pie chart for visual display.

[0056] Compared with the prior art, the present invention has the following advantages:

[0057] The present invention discloses a Raman spectroscopy algorithm SPP-1D based on a multi-channel one-dimensional convolutional neural network and spatial pyramid pooling technology. By innovatively combining the multi-channel one-dimensional convolutional neural network (1D-CNN) and spatial pyramid pooling (SPP) technology, a new solution is provided, significantly improving the accuracy and robustness of Raman spectroscopy analysis.

[0058] Traditional Raman spectroscopy analysis methods, such as Raman spectroscopy analysis based on principal component analysis (PCA) or linear discriminant analysis (LDA), are difficult to capture the complex non-linear relationships in spectral data and are easily affected by noise interference, resulting in information loss. However, the present invention can extract local and global features of spectral data at different scales, thereby enhancing the model's learning ability for complex spectral signals.

[0059] Introducing a multi-channel one-dimensional convolutional neural network can simultaneously process multiple feature channels of Raman spectroscopy. For example, when processing Raman spectroscopy data, spectral data of different bands or different preprocessing results can usually be used as independent input channels. For example, for a set of Raman spectroscopy data, assuming we extract three different bands (such as 400-600cm -1 , 600-800cm -1 , 800-1000cm -1)As three input channels, the network can automatically identify the correlation between each channel, thereby fusing the feature information of multiple bands, improving the classification and recognition accuracy of the spectrum.

[0060] The spatial pyramid pooling technique is introduced into Raman spectroscopy analysis for the first time. When traditional methods such as using a simple deep neural network (DNN) or support vector machine (SVM) are used for Raman spectroscopy analysis, usually only features of a single scale are extracted, ignoring the multi-level and global features of the spectral signal. In contrast, the present invention performs feature pooling at different scales through spatial pyramid pooling, retains more detailed information, and effectively improves the generalization ability and robustness of the model. Especially in the face of noise interference and data loss, it has strong stability and accuracy.

[0061] This algorithm is designed to have better versatility and flexibility, and can adapt to different types of Raman spectroscopy data and analysis tasks. It can not only be used for the Raman spectroscopy identification and analysis of a single substance, but also be extended to various application scenarios such as the component analysis of complex mixtures and Raman spectroscopy imaging, and can show good performance in different data sets and practical applications. When using this algorithm for analysis, it can identify spectra with high similarity. In the component analysis of complex mixtures, this algorithm also shows excellent performance.

[0062] The present invention realizes the deep integration of deep learning technology and Raman spectroscopy analysis technology, introduces advanced methods in the fields of computer vision and signal processing into the field of chemical analysis, provides new ideas and methods for the development and application of Raman spectroscopy technology, and promotes the technological innovation and development in interdisciplinary fields. Description of the Drawings

[0063] Figure 1 is the overall flowchart of the Raman spectroscopy algorithm SPP-1D based on a multi-channel one-dimensional convolutional neural network and spatial pyramid pooling technology. Specific Embodiment

[0064] The present invention will be further described below with reference to the accompanying drawings.

[0065] In this embodiment, as Figure 1 shown, the overall flowchart of the Raman spectroscopy algorithm SPP-1D based on a multi-channel one-dimensional convolutional neural network and spatial pyramid pooling technology, the specific implementation mainly includes the following steps:

[0066] Step 1, construct a multi-channel one-dimensional convolutional neural network (1D-CNN) model, use multiple convolutional kernels to perform convolutional operations on the input data, and extract local features. Use the spatial pyramid pooling (SPP) layer to perform pooling operations at different scales, and splice the pooling results of each scale to obtain the final multi-scale feature representation.

[0067] Step 2: Standardize the input data to have the same scale, segment and augment the data to improve the generalization ability of the model. Select appropriate loss functions (cross-entropy loss and mean squared error), as well as optimization algorithms (Adam), and evaluate the model performance on the validation set.

[0068] Step 3: The output layer is used to generate prediction results, and appropriate activation functions (Mish and GeLU) are selected.

[0069] Step 4: Initialize the model parameters, monitor the accuracy and loss values during the training process, adjust the model parameters in real time, and find the optimal hyperparameters.

[0070] Step 5: After the training is completed, use matplotlib to plot the change curves of the loss value and accuracy during the training process.

[0071] Step 6: Standardize and smooth the input spectral data, use the trained model to predict the preprocessed spectral data, and perform visual display.

[0072] Specifically, in Step 1, a multi-channel one-dimensional convolutional neural network (1D-CNN) model is constructed. Multiple convolutional kernels are used to perform convolutional operations on the input data to extract local features. Then, the spatial pyramid pooling (SPP) layer is used to perform pooling operations at different scales, extract multi-scale features, and splice the pooling results together to obtain the final multi-scale feature representation. The specific steps are as follows:

[0073] Step (1.1): Calculate and determine the start and end positions of the pooling window:

[0074]

[0075] where ix and iy are the indices of the current pooling window, i is the level of the pooling layer, num cols and num rows are the number of columns and rows of the input feature map, respectively.

[0076] Step (1.2): Perform a pooling operation on the cropped pooling area, and use the following formula to calculate the maximum value after pooling:

[0077] PooledVal = max(x crop , axis=(1, 2)) …………(5)

[0078] where x crop is the cropped pooling area, and PooledVal is the maximum value after pooling.

[0079] Step (1.3), concatenate the pooling results of different scales to form the final multi-scale feature representation, and the output shape is determined by the following formula:

[0080] output shape =(input shape [0], num outputs ×channels)…………(6)

[0081] In the formula, input shape is the shape of the input feature map, num outputs is the number of outputs per channel, and channels is the number of channels.

[0082] Among them, the multi-channel one-dimensional convolutional neural network model includes the following structures:

[0083] Convolutional layer: Through the combination of the convolutional layer, batch normalization layer, activation layer, and pooling layer, extract the local features of the input signal. The convolution operation formula is:

[0084]

[0085] In the formula, x is the input signal, ω is the convolution kernel, K is the size of the convolution kernel, and y is the convolution result.

[0086] SPP layer: Perform multi-scale pooling on the feature map output by the convolution and flatten the result as the input of the fully connected layer.

[0087] Fully connected layer: Further process the extracted features through the fully connected layer and generate the final feature representation for classification or prediction.

[0088] On this basis, in Step 2, standardize and reshape the input data into the shape required by the model, and use the following formula for standardization:

[0089]

[0090] In the formula, X is the input data, and X normal is the data after standardization.

[0091] Furthermore, in Step 3, select an appropriate activation function:

[0092] Mish activation function:

[0093] X Mish =x·tanh(ln(1 + e x ))…………(11)

[0094] In the formula, x is the input value, and X Mish is the output value of the Mish activation function.

[0095] GeLU activation function:

[0096]

[0097] where \(x\) is the value input to the activation function, \(x\) GeLU is the output value of the GeLU activation function.

[0098] Furthermore, in step four, the model training and optimization steps include:

[0099] Step (4.1), training the deep learning model using the training data;

[0100] Step (4.2), monitoring the accuracy and loss value during the training process in real time, and using the Adam optimization algorithm to adjust the model parameters. The Adam optimization formula is:

[0101] \(m\) t =\(\beta\) 1 \(m\) t-1 +(1 - \(\beta\) 1 )\(g\) t …………(14)

[0102]

[0103]

[0104] where \(m\) t and \(v\) t are the momentum and root mean square gradient, \(\beta\) 1 and \(\beta\) 2 are the momentum decay rates, \(g\) t is the gradient, \(\eta\) is the learning rate, \(\epsilon\) is the smoothing term, \(\theta\) t is the parameter, and are the momentum and root mean square gradient after bias correction.

[0105] Step (4.3), adjusting the hyperparameters, such as the learning rate, batch size, and number of network layers, through grid search or random search methods.

[0106] After that, step five displays the training progress and model performance through a visualization interface, and uses matplotlib to plot the change curves of the loss value and accuracy during the training process.

[0107] Finally, in step six, the prediction of spectral data includes smoothing the spectral data and using the trained model for prediction. Using the Whittaker smoothing and airPLS methods:

[0108] Whitaaker smoothing algorithm:

[0109] z = (W + λD T D) -1 (Wx)……………(19)

[0110] Wherein, W is the weight matrix, λ is the smoothing parameter, D is the difference matrix, x is the input signal, and z is the smoothed signal.

[0111] airPLS algorithm:

[0112]

[0113] Wherein, w is the weight, d is the residual, and i is the number of iterations.

[0114] The prediction result display includes generating the final prediction result according to the probability distribution or regression result output by the model, and visualizing it using a pie chart.

Claims

1. A Raman spectroscopy algorithm SPP-1D based on multi-channel one-dimensional convolutional neural network and spatial pyramid pooling technology, characterized in that: The steps include: Step 1: Build a multi-channel one-dimensional convolutional neural network (1D-CNN) model, use multiple convolution kernels to perform convolution operations on the input data, and extract local features. Use the spatial pyramid pooling (SPP) layer to perform pooling operations at different scales, and concatenate the pooling results of each scale to obtain the final multi-scale feature representation. Step 2: Standardize the input data to make them of the same scale, segment and enhance the data to improve the generalization ability of the model. Select the appropriate loss function (cross entropy loss and mean square error) and optimization algorithm (Adam), and evaluate the model performance on the validation set. Step 3: The output layer is used to generate prediction results and select the appropriate activation function (Mish and GeLU). Step 4: Initialize the model parameters, monitor the accuracy and loss value during training, adjust the model parameters in real time, and find the optimal hyperparameters. Step 5: After the training is completed, use matplotlib to draw the change curve of loss value and accuracy during the training process. Step six: standardize and smooth the input spectral data, use the trained model to predict the preprocessed spectral data, and perform visual display.

2. The Raman spectroscopy algorithm SPP-1D according to claim 1, characterized in that: In step 1, a multi-channel one-dimensional convolutional neural network (1D-CNN) model is constructed, and multiple convolution kernels are used to perform convolution operations on the input data to extract local features. Then, the spatial pyramid pooling (SPP) layer is used to perform pooling operations at different scales, extract multi-scale features and concatenate the pooling results to obtain the final multi-scale feature representation. The specific steps are: Step (1), calculate and determine the starting and ending positions of the pooling window: Where ix and iy are the indices of the current pooling window, i is the number of pooling layers, and num cols and num rows are the number of columns and rows of the input feature map, respectively. Step (2) performs a pooling operation on the cropped pooling area and calculates the maximum value after pooling using the following formula: PooledVal=max(x crop ,axis=(1,2))………(5) In the formula, x crop is the cropped pooled area, and PooledVal is the maximum value after pooling. Step (3) concatenates the pooling results of different scales to form the final multi-scale feature representation. The output shape is determined by the following formula: output shape =(input shape [0],num outputs ×channels)…………(t) In the formula, input shape is the shape of the input feature map, num outputs is the number of outputs per channel, and channels is the number of channels.

3. The Raman spectroscopy algorithm SPP-1D according to claim 1, characterized in that: In step 1, the multi-channel one-dimensional convolutional neural network model includes the following structure: Step (1), convolution layer: extract the local features of the input signal through the combination of convolution layer, batch normalization layer, activation layer and pooling layer. The convolution operation formula is: Where x is the input signal, ω is the convolution kernel, K is the size of the convolution kernel, and y is the convolution result. Step (2), SPP layer: multi-scale pooling is performed on the feature map output by the convolution, and the result is flattened as the input of the fully connected layer. Step (3), fully connected layer: The extracted features are further processed through the fully connected layer to generate the final feature representation for classification or prediction.

4. The Raman spectroscopy algorithm SPP-1D according to claim 1, characterized in that: In step 2, the input data is normalized and reshaped into the shape required by the model. The normalization is performed using the following formula: In the formula, X is the input data, X normal is the standardized data.

5. The Raman spectroscopy algorithm SPP-1D according to claim 1, characterized in that: In step 3, the classification and prediction steps include using the Mish activation function for multi-class classification and using the GeLU activation function for two-class classification or regression tasks: Mish activation function: X Mish =x·tanh(ln(1+e x )) In the formula, x is the input value, X Mish Output value for the Mish activation function. GeLU activation function: In the formula, x is the value input to the activation function, X GeLU Output value for the GeLU activation function.

6. The Raman spectroscopy algorithm SPP-1D according to claim 1, characterized in that: In step 4, the model training and optimization steps include: Step (1), training the deep learning model using training data; Step (2) monitors the accuracy and loss value during the training process in real time, and uses the optimization algorithm Adam to adjust the model parameters. The Adam optimization formula is: m t =β1m t-1 +(1-β1)g t ……………(13) In the formula, m t and v t are the momentum and the RMS gradient, β1 and β2 are the momentum decay rates, and g t is the gradient, η is the learning rate, ∈ is the smoothing term, θ t is a parameter, and are the bias-corrected momentum and rms gradient. In step (3), hyperparameters such as learning rate, batch size, and number of network layers are adjusted by grid search or random search methods.

7. The Raman spectroscopy algorithm SPP-1D according to claim 1, characterized in that: In step 5, the training progress and model performance are visualized, and matplotlib is used to plot the change curves of loss value and accuracy during training.

8. The Raman spectroscopy algorithm SPP-1D according to claim 8, characterized in that: In step 6, the prediction of spectral data includes smoothing the spectral data and using the trained model for prediction. Using Whittaker smoothing and airPLS method: Whitaaker smoothing algorithm: z=(W+λD T D) -1 (Wx)…………(1t) Where W is the weight matrix, λ is the smoothing parameter, D is the difference matrix, x is the input signal, and z is the smoothed signal. airPLS algorithm: Where w is the weight, d is the residual, and i is the number of iterations.

9. The Raman spectroscopy algorithm SPP-1D according to claim 1, characterized in that: In step six, the prediction result display includes generating the final prediction result based on the probability distribution or regression result output by the model, and visually displaying it using a pie chart.

Citation Information

Cited By

  • Comparison matching system of spectrum and molecular structure based on machine learning

    CN120508836A

  • A machine learning-based spectrum and molecular structure comparison and matching system

    CN120508836B