A substance analysis method based on the combination of spectral analysis and video classification algorithms
By integrating spectroscopy analysis with video classification and transfer learning, the method enhances substance identification accuracy and robustness, addressing the limitations of traditional spectroscopy and deep learning models with limited data.
Patent Information
- Application Number
- CN202211236799.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-10-10
AI Technical Summary
The prior art cannot accurately identify substance analysis when mixtures have similar spectral spectral or low substance concentrations. The lack of data support in the spectral field of deep learning methods leads to overfitting of models.
Spectral analysis is used in combination with video classification algorithm, and the sample spectral data is obtained and spectral video is converted into spectral video, the C3D network model is used for transfer learning, and the UCF101 data set pre-trained model is used for matter recognition.
More accurate substance detection is achieved, especially in similar mixtures and low concentration substance detection, improving the accuracy and robustness of identification.
Smart Images

Figure CN115512164B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biochemistry, and more specifically to a method for analyzing substances by combining spectral analysis and video classification algorithms. Background Art
[0002] Spectral analysis methods refer to a class of methods for analyzing substances, mainly including atomic emission spectrometry, atomic absorption spectrometry, ultraviolet-visible absorption spectrometry, infrared spectrometry, etc. According to the nature of electromagnetic radiation, spectral analysis can be further divided into molecular spectra and atomic spectra. It is mainly caused by the transition of valence electrons in molecules, so this absorption spectrum is determined by the distribution and binding of valence electrons in molecules. Spectral analysis methods have the advantages of fast analysis speed, simple operation, no need for pure samples, simultaneous determination of multiple elements or compounds, good selectivity, high sensitivity, and less damage to samples. However, the current methods for analyzing substances cannot accurately identify the spectra of mixtures. When the spectra of mixtures are similar or the substance concentration is low, even the most advanced deep learning methods are difficult to effectively identify each mixture.
[0003] The video classification task is to generate labels related to a video given video frames. A good video classification model can best describe the entire video based on the information of each frame in the video. It is similar to image classification, using a feature extractor (such as a convolutional neural network) to extract features from image sequences and time series, and then classifying them. Application tasks include assigning one or more labels to a video, and assigning one or more labels to each frame within a video. There has been a large amount of research in the field of video classification, such as human activity recognition, gesture recognition, anomaly detection, and monitoring, but it has never been used in the spectral field.
[0004] Deep learning has been widely applied in the field of spectral analysis. Powerful models require a large amount of training data for support. In practical applications, transfer learning is usually used to solve the problem of data volume. However, there are few public datasets in the current spectral field, so there will be situations of insufficient data volume and model overfitting. While there are rich public datasets and the most advanced models in the video field.
[0005] The goal of transfer learning is to apply the knowledge learned in one field to a different but related field. It uses a pre-trained network applied in other fields as the starting point of the model parameters for training its own dataset, and can obtain excellent results with fewer computing resources and less datasets. This technology has been widely applied in the field of deep learning. For example, after transfer learning, the cat and dog image classification model can obtain the ability to recognize wolves and tigers; the simulation images of video games are used for pre-training autonomous vehicles.
[0006] Therefore, how to provide an accurate method for material analysis by combining spectral analysis and video classification algorithms is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a method for material analysis by combining spectral analysis and video classification algorithms to achieve more accurate material detection.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A method for material analysis by combining spectral analysis and video classification algorithms, comprising the following steps:
[0010] S1. Obtain samples of pure substances and mixtures;
[0011] S2. Obtain sample spectral data at different wavelengths and different laser powers;
[0012] S3. Convert the sample spectral data into spectral images through continuous wavelet transform;
[0013] S4. Synthesize the obtained spectral images to obtain a spectral video; wherein, n spectral images of each type of sample obtained by the same spectrometer are synthesized, and the n spectral images are respectively from adjacent laser powers;
[0014] S5. Collect a public video dataset to pre-train a video deep learning classification model;
[0015] S6. Perform transfer learning on the spectral video synthesized in S4 on the pre-trained video deep learning classification model;
[0016] S7. Identify substances through the video deep learning classification model after transfer learning.
[0017] Preferably, the sample spectral data obtained in S2 is Raman spectral data.
[0018] Preferably, the specific content of converting the spectral data into spectral images in S3 includes:
[0019] 2) Use the Morse wavelet as the mother wavelet to transform the Raman spectral signal into a two-dimensional time-frequency domain signal, where the Morse wavelet is:
[0020]
[0021] where ω is the digital domain frequency, representing the rate of sequence change, and its expression is ω = 2πf*Ts, where Ts is the sampling period, It is a Morse wavelet, U(ω) is the unit step, a is the normalization constant, β is the time-bandwidth product, and γ is the symmetry parameter characterizing the symmetry of the Morse wavelet;
[0022] 2) According to the spectral resolution, set the time-bandwidth product β and the symmetry parameter γ, calculate ω, and use the Raman spectral data as the input to perform feature transformation through the Morse wavelet:
[0023]
[0024] Among them, b is a scaling factor with a fixed length;
[0025] Represent the CWT time-frequency domain two-dimensional data scale diagram as
[0026] Preferably, when the deep learning classification model is used for the first time, preprocess the data, divide each of the spectral videos into non-overlapping n pictures, and use them as the input of the deep learning classification model.
[0027] Preferably, the deep learning classification model in S5 is a C3D network model, which includes 8 convolutional layers, 5 pooling layers, 2 fully connected layers, 2 Dropout layers and 1 classification layer, and the classification layer is a fully connected layer activated by softmax;
[0028] The 8 convolutional layers are connected in sequence, and pooling layers are connected between the first convolutional layer and the second convolutional layer, between the second convolutional layer and the third convolutional layer, between the fourth convolutional layer and the fifth convolutional layer, between the sixth convolutional layer and the seventh convolutional layer, and between the eighth convolutional layer and the first fully connected layer;
[0029] The 2 fully connected layers are connected in sequence, and the Dropout layers are respectively connected between the first fully connected layer and the second fully connected layer, and between the second fully connected layer and the classification layer.
[0030] Preferably, the convolution kernels of the 8 convolutional layers of the C3D network model are all 3*3*3 in size and 1*1*1 in stride; the pooling kernel size and stride of the first pooling layer are both 1*2*2, and the size and stride of the remaining pooling layers are both 2*2*2; each fully connected layer includes 4096 output units, and the classification layer is a fully connected layer with 14 output units activated by softmax.
[0031] Preferably, the convolution of the C3D network model is three-dimensional convolution, and the feature maps in the convolutional layer are respectively connected to different adjacent frames in the previous layer to obtain motion information. The value formula at the (x, y, z) position of the jth feature map in the ith layer is as follows:
[0032]
[0033] where tanh() is the hyperbolic tangent function, and b ij is the offset of this feature map, m is the index of the set of feature maps in the (i - 1)th layer connected to the current feature map, Pi and Qi are the height and width of the kernel respectively, and R i is the size of the 3D kernel along the time dimension, is the (p, q, r)th value of the kernel connected to the mth feature map of the previous layer.
[0034] Preferably, the common video dataset for pre-training in S5 is the UCF101 dataset. Change the output of the classification layer of the C3D network model to 101, and after training for 200 epochs on the UCF101 dataset, select the model with the highest accuracy on the validation set as the pre-trained model.
[0035] Preferably, the specific content of the transfer learning in S6 includes: using the model parameters pre-trained with UCF101 as the initial parameters for the spectral video dataset, and changing the output of the classification layer of the C3D network model from 101 to 14 for training and fine-tuning.
[0036] Preferably, when performing pre-training and transfer learning training, the learning rate of the classification layer is 1e - 3, and the learning rate of the other layers is 1e - 4; the optimizer is stochastic gradient descent; the loss function is cross-entropy loss, and its calculation formula is as follows:
[0037]
[0038] where N is the number of samples, K is the number of labels, and y i,k represents that the true label of the i-th sample is k, and p i,k represents the probability that the i-th sample is predicted as the k-th label value.
[0039] Through the above technical solutions, compared with the prior art, the present invention discloses a method for substance analysis based on the combination of spectral analysis and video classification algorithm, and has the following technical effects:
[0040] 1. Compared with using traditional chemometrics analysis and machine learning for substance identification, the deep learning model has the advantages of higher accuracy, greater convenience, and stronger robustness;
[0041] 2. Compared with directly using one-dimensional spectral data for deep learning substance identification, the three-dimensional model can learn spatio-temporal features, and the three-dimensional video data has richer feature information.
[0042] 3. Compared with directly using one-dimensional spectral data for deep learning-based substance identification, converting to three-dimensional video data allows the use of more powerful and advanced models, and applying transfer learning after pre-training with a large public dataset, achieving better results than directly analyzing one-dimensional spectral data. It can be applied to tasks such as detecting similar mixtures, detecting low-concentration substances, and detecting substances in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the provided drawings.
[0044] Figure 1 It is a schematic flowchart of a substance analysis method based on the combination of spectral analysis and video classification algorithm provided by an embodiment of the present invention;
[0045] Figure 2 It is a schematic structural diagram of the C3D network model provided by an embodiment of the present invention;
[0046] Figure 3 It is a confusion matrix diagram of the validation set provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0048] An embodiment of the present invention discloses a substance analysis method based on the combination of spectral analysis and video classification algorithm, as Figure 1 shown, including the following steps:
[0049] S1. Obtain samples of pure substances and mixtures;
[0050] S2. Obtain spectral data of the samples at different wavelengths and different laser powers;
[0051] S3. Convert the sample spectral data into spectral images through continuous wavelet transform;
[0052] S4. Synthesize the obtained spectral images to obtain a spectral video; among them, n spectral images of each type of sample obtained by the same spectrometer are synthesized, and the n spectral images are respectively from adjacent laser powers;
[0053] S5. Collect a public video dataset to pre-train the video deep learning classification model;
[0054] S6. Perform transfer learning on the spectral videos synthesized in S4 using the pre-trained video deep learning classification model;
[0055] S7. Perform substance identification using the video deep learning classification model after transfer learning.
[0056] In this embodiment, the pure substances are methanol, ethanol, propanol, butanol, pentanol, hexanol, heptanol, and octanol, and the mixtures are eight alcohol mixtures of methanol and pentanol, ethanol and hexanol, propanol and heptanol, butanol and pentanol, methanol and ethanol and butanol and octanol.
[0057] To further implement the above technical solution, the sample spectral data obtained in S2 is Raman spectral data.
[0058] In this embodiment, Raman spectral data is obtained by 4 Raman spectrometers, two of which have a wavelength of 785 nm, one has a wavelength of 532 nm, and the other has a wavelength of 1064 nm. Each Raman spectrometer uses 10 different laser powers, and 10 groups of data are measured at each laser power. Each substance has a total of 400 groups of one-dimensional spectral data:
[0059] The Raman spectrometer with a laser wavelength of 532 nm uses an integration time of 500 ms, an average number of times of 3 times, a laser power of 10 mW - 100 mW (one for every 10 mW), and 10 groups of Raman spectral data are measured for each laser power, for a total of 100 groups.
[0060] One Raman spectrometer with a laser wavelength of 785 nm uses an integration time of 500 ms, an average number of times of 3 times, a laser power of 210 mW - 300 mW (one for every 10 mW), and 10 groups of Raman spectral data are measured for each laser power, for a total of 100 groups.
[0061] Another Raman spectrometer with a laser wavelength of 785 nm uses an integration time of 2000 ms, an average number of times of 1 time, a laser power of 200 mW - 380 mW (one for every 20 mW), and 10 groups of Raman spectral data are measured for each laser power, for a total of 100 groups.
[0062] The Raman spectrometer with a laser wavelength of 1064 nm uses an integration time of 5000 ms, an average number of times of 1 time, a laser power of 410 mW - 500 mW (one for every 10 mW), and 10 groups of Raman spectral data are measured for each laser power, for a total of 100 groups.
[0063] There are 400 groups of data for one type of sample, and a total of 5600 groups of data for 14 types of samples.
[0064] To further implement the above technical solution, the spectral data in S3 is converted into a spectral image through continuous wavelet transform. The specific content includes:
[0065] 3) Using the Morse wavelet as the mother wavelet to transform the Raman spectral signal into a two-dimensional time-frequency domain signal, where the Morse wavelet is:
[0066]
[0067] where ω is the digital domain frequency, representing the rate of sequence change, and its expression is ω = 2πf*Ts, where Ts is the sampling period. is the Morse wavelet, U(ω) is the unit step, a is the normalization constant, β is the time-bandwidth product, and γ is the symmetry parameter characterizing the symmetry of the Morse wavelet;
[0068] 2) According to the spectral resolution, set the time-bandwidth product β and the symmetry parameter γ, calculate ω, and use the Raman spectral data as the input to perform feature transformation through the Morse wavelet:
[0069]
[0070] where b is a scaling factor with a fixed length;
[0071] Represent the CWT time-frequency domain two-dimensional data scale diagram as
[0072] To further implement the above technical solution, when the deep learning classification model is used for the first time, the data is preprocessed, and each spectral video is segmented into non-overlapping n pictures, which are used as the input of the deep learning classification model.
[0073] All video frames are adjusted to 128*171, and the video is segmented into non-overlapping 16-frame clips, which are used as the input of the network. The input size is 3*16*128*171, and the video segment is represented as c*l*h*w, where c is the number of channels, l is the length of the number of frames, and h and w are the height and width of the frame respectively.
[0074] In this embodiment, n is 20, the fps of each Raman video data is 10, the duration is 2s, there are 14 categories of pure substances and mixture samples, each category contains 20 video data, and there are a total of 280 video data; 16 videos of each category of substances are used as the training set of the deep learning classification model, and 4 videos are used as the test set. The test set is the intermediate power video data at each wavelength; there are 224 videos in the training set and 56 videos in the test set.
[0075] To further implement the above technical solution, the deep learning classification model is a C3D network model.
[0076] In this embodiment, the network model and training code are built based on the PyTorch framework. In practical applications, a two-stream model, a model based on video transformers, etc. can also be adopted.
[0077] To further implement the above technical solution, as Figure 2 shown, the C3D network model includes 8 convolutional layers, 5 pooling layers, 2 fully connected layers, 2 Dropout layers, and 1 classification layer. The classification layer is a fully connected layer activated by softmax;
[0078] The 8 convolutional layers are connected in sequence. Pooling layers are connected between the first and second convolutional layers, the second and third convolutional layers, the fourth and fifth convolutional layers, the sixth and seventh convolutional layers, and the eighth convolutional layer and the first fully connected layer;
[0079] The 2 fully connected layers are connected in sequence, and Dropout layers are respectively connected between the first fully connected layer and the second fully connected layer, and the second fully connected layer and the classification layer.
[0080] To further implement the above technical solution, the convolutional kernels of the 8 convolutional layers are all 3*3*3 in size and 1*1*1 in stride; the pooling kernel size and stride of the first pooling layer are both 1*2*2, and the size and stride of the remaining pooling layers are both 2*2*2; each fully connected layer includes 4096 output units, and the classification layer is a fully connected layer with 14 output units activated by softmax.
[0081] The convolution of the C3D network model is a three-dimensional convolution, which is achieved by convolving a 3D kernel into a cube formed by stacking multiple adjacent frames. Through this structure, the feature maps in the convolutional layer are connected to multiple adjacent frames in the previous layer, thereby obtaining motion information. The value at the (x, y, z) position of the j-th feature map in the i-th layer is given by the following formula:
[0082]
[0083] where tanh() is the hyperbolic tangent function, b ij is the offset of this feature map, m is the index of the set of (i - 1)-layer feature maps connected to the current feature map, Pi and Qi are the height and width of the kernel respectively, R i is the size of the 3D kernel along the time dimension, is the (p, q, r)-th value of the kernel connected to the m-th feature map of the previous layer.
[0084] To further implement the above technical solution, the public large video dataset used for pre-training in S5 is the UCF101 dataset. The output of the classification layer of the C3D network model is changed to 101, and after training for 200 epochs on the UCF101 dataset, the model with the highest accuracy on the validation set is taken as the pre-training model.
[0085] To further implement the above technical solution, the specific content of transfer learning in S6 includes: using the model parameters pre-trained on UCF101 as the initial parameters of the spectral video dataset, and changing the output of the classification layer of the C3D network model from 101 to 14 for training and fine-tuning.
[0086] It should be noted that:
[0087] The transfer learning in S6 refers to transfer learning based on shared parameters, specifically referring to that after the model is pre-trained on a large-scale dataset and then fine-tuned on the target dataset.
[0088] In practical applications of the present invention, other models based on 3D convolutional neural networks, two-stream models based on temporal convolution and spatial convolution, models based on video transformers, etc. can also be adopted.
[0089] To further implement the above technical solution, when performing pre-training and transfer learning training, the learning rate of the classification layer is 1e-3, and the learning rate of the remaining layers is 1e-4; the optimizer is stochastic gradient descent; the loss function is cross-entropy loss, and its calculation formula is as follows:
[0090]
[0091] where N is the number of samples, K is the number of labels, y i,k represents that the true label of the i-th sample is k, and p i,k represents the probability that the i-th sample is predicted as the k-th label value.
[0092] In this embodiment, the number of epochs during transfer learning training is 30, and the learning rate is reduced by 10 times every 10 epochs; the optimizer is stochastic gradient descent, the momentum factor is 0.9, and the weight decay is 5e-4; after verification by the validation set, the accuracy rate is 94.64%, Figure 3 which is the confusion matrix result of 56 validation set videos.
[0093] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0094] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for material analysis based on the combined use of spectral analysis and video classification algorithms, characterized in that, It includes the following steps: S1. Obtain samples of pure substances and mixtures; S2. Obtain spectral data of the samples at different wavelengths and different laser powers; S3. Convert the sample spectral data into spectral images through continuous wavelet transform; S4. Synthesize the obtained spectral images to obtain a spectral video; wherein, n spectral images of each type of sample obtained by the same spectrometer are synthesized, and the n spectral images are respectively from adjacent laser powers; S5. Collect a public video dataset to pre-train the video deep learning classification model; S6. Perform transfer learning on the spectral video synthesized in S4 on the pre-trained video deep learning classification model; S7. Perform substance identification through the video deep learning classification model after transfer learning; The specific content of converting the spectral data into spectral images in S3 includes: Use the Morse wavelet as the mother wavelet to transform the Raman spectral signal into a two-dimensional time-frequency domain signal, where the Morse wavelet is: where ω is the digital domain frequency, representing the rate of sequence change, and its expression is ω = 2πf*Ts, where Ts is the sampling period. is the Morse wavelet, U(ω) is the unit step, a is the normalization constant, β is the time-bandwidth product, and γ is the symmetry parameter characterizing the symmetry of the Morse wavelet. 2) According to the spectral resolution, set the time-bandwidth product β and the symmetry parameter γ, calculate ω, and use the Raman spectral data as the input to perform feature transformation through the Morse wavelet: where b is a scaling factor with a fixed length; Express the scale map of the two-dimensional data in the CWT time-frequency domain as The deep learning classification model in S5 is a C3D network model, which includes 8 convolutional layers, 5 pooling layers, 2 fully connected layers, 2 Dropout layers and 1 classification layer, and the classification layer is a fully connected layer activated by softmax; The 8 convolutional layers are connected in sequence, and pooling layers are connected between the first convolutional layer and the second convolutional layer, between the second convolutional layer and the third convolutional layer, between the fourth convolutional layer and the fifth convolutional layer, between the sixth convolutional layer and the seventh convolutional layer, and between the eighth convolutional layer and the first fully connected layer; The 2 fully connected layers are connected in sequence, and Dropout layers are respectively connected between the first fully connected layer and the second fully connected layer, and between the second fully connected layer and the classification layer; The convolution kernels of the 8 convolutional layers of the C3D network model are all 3*3*3 in size, and the stride is 1*1*1; the pooling kernel size and stride of the first pooling layer are both 1*2*2, and the sizes and strides of the remaining pooling layers are both 2*2*2; each fully connected layer includes 4096 output units, and the classification layer is a fully connected layer with 14 output units activated by softmax.
2. The substance analysis method based on the combined use of spectral analysis and video classification algorithm according to claim 1, wherein The sample spectral data obtained in S2 is Raman spectral data.
3. A method for material analysis based on the combination of spectral analysis and video classification algorithm according to claim 1, characterized in that, When the deep learning classification model is used for the first time, preprocess the data, and divide each spectral video into non-overlapping n pictures for use as the input of the deep learning classification model.
4. A method for analyzing substances by combining spectral analysis and video classification algorithms according to claim 1, characterized in that, The convolution of the C3D network model is three-dimensional convolution, and the feature maps in the convolutional layer are respectively connected to different adjacent frames in the previous layer to obtain motion information. The value formula at the (x, y, z) position of the jth feature map in the ith layer is as follows: where tanh() is the hyperbolic tangent function, bij is the offset of the feature map, m is the index of the set of (i - 1)-layer feature maps connected to the current feature map, Pi and Qi are the height and width of the kernel respectively, and Ri is the size of the 3D kernel along the time dimension, is the (p, q, r)th value of the kernel connected to the mth feature map of the previous layer.
5. A method for substance analysis based on the combined use of spectral analysis and video classification algorithms according to claim 1, characterized in that, The public video dataset used for pre-training in S5 is the UCF101 dataset. The output of the classification layer of the C3D network model is changed to 101, and after training for 200 epochs on the UCF101 dataset, the model with the highest accuracy on the validation set is taken as the pre-trained model.
6. The substance analysis method based on the combination of spectral analysis and video classification algorithm according to claim 5, wherein, The specific content of the transfer learning in S6 includes: using the model parameters pre-trained on the UCF101 as the initial parameters of the spectral video dataset, and changing the output of the classification layer of the C3D network model from 101 to 14 for training and fine-tuning.
7. A method for analyzing substances by combining spectral analysis and video classification algorithms according to claim 1, characterized in that When performing pre-training and transfer learning training, the learning rate of the classification layer is 1e-3, and the learning rate of the remaining layers is 1e-4; the optimizer is stochastic gradient descent; the loss function is cross-entropy loss, and its calculation formula is as follows: where N is the number of samples, K is the number of labels, and y i,k indicates that the true label of the i-th sample is k, and p i,k represents the probability that the i-th sample is predicted as the k-th label value.
Citation Information
Patent Citations
Hyperspectral image classification method based on separable three-dimensional residual network and transfer learning
CN109754017A
Multi-industry detection-oriented laser raman spectrum intelligent identification method and system
WO2015165394A1