Raman spectrum feature extraction method and modeling method based on convolution and attention mechanism

By employing a Raman spectroscopy feature extraction method based on convolution and attention mechanisms, and integrating Raman spectroscopy and physicochemical data, the problem of real-time monitoring and accurate prediction during fermentation was solved, enabling intelligent control and optimization of the fermentation process.

CN120046099BActive Publication Date: 2025-12-09南宁桂电电子科技研究院有限公司 +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510076527.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-12-09
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Existing technologies cannot achieve real-time and accurate monitoring of components and physical quantities during fermentation. Sensor technology is limited to the measurement of specific parameters, and models based on experience and statistics have poor adaptability and are difficult to describe complex nonlinear dynamic changes.

Method used

A Raman spectroscopy feature extraction method based on convolution and attention mechanisms is adopted. By fusing Raman spectroscopy and physicochemical time series data, data preprocessing and enhancement are performed. A deep learning model is used to automatically extract key features in the fermentation process, build the model, and perform backpropagation training.

Benefits of technology

It enables real-time and accurate monitoring of the fermentation process, reduces interference with the fermentation process, improves the accuracy of predictions and the adaptability of the model, reduces data acquisition and processing costs, and supports intelligent control and optimization of the fermentation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046099B_ABST
    Figure CN120046099B_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure provides a Raman spectrum feature extraction method and modeling method based on convolution and attention mechanism, and the Raman spectrum feature extraction method comprises the following steps: acquiring Raman spectrum time series data and physical and chemical time series data generated by monitoring a target biological process in fermentation; fusing the Raman spectrum time series data and the physical and chemical time series data to generate fermentation feature sequence data; preprocessing and enhancing the fermentation feature sequence data to obtain enhanced feature sequence data; extracting Raman spectrum features from the enhanced feature sequence data based on the Raman spectrum feature extraction model to predict the Raman spectrum features and determine the components produced by the biological process in fermentation and the physical quantities thereof.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, and particularly relates to a Raman spectrum feature extraction method and modeling method based on convolution and attention mechanism. BACKGROUND

[0002] In the field of biological fermentation, accurately predicting the components and their physical quantities produced during fermentation is crucial for optimizing fermentation processes, improving product quality and production efficiency. With the development of biotechnology and the increasing demand for industrial production, precise monitoring and prediction of fermentation processes have become a research hotspot. Traditionally, predicting the components and their physical quantities produced by organisms during fermentation mainly relies on chemical analysis methods, sensor technology, and models based on experience and statistics.

[0008] Chemical analysis methods are early means for detecting fermentation components, such as high-performance liquid chromatography (HPLC), gas chromatography (GC), etc. These methods can accurately determine the types and contents of fermentation products by separating and detecting the fermentation broth. For example, in antibiotic fermentation production, HPLC can be used to quantitatively analyze the content of antibiotics in the fermentation broth. However, these methods usually require complex sample pretreatment, tedious operation, and long time, and cannot realize real-time online monitoring. Moreover, frequent sampling may interfere with the normal progress of the fermentation process, affecting the accuracy of experimental results. Sensor technology is also a commonly used monitoring means. For example, pH sensors, dissolved oxygen sensors, etc. can monitor the physicochemical parameters in real time during the fermentation process. These sensors can quickly feed back data, providing certain basis for the control of the fermentation process. However, the types of sensors are limited, and only specific parameters can be measured, which cannot fully reflect the complex changes of components in the fermentation process. In addition, the precision and stability of the sensor are easily affected by the fermentation environment, such as changes in temperature, pH, etc. Factors such as changes in the sensor measurement error increase, affecting the accuracy of the prediction.

[0009] Models based on experience and statistics are established according to a large amount of historical fermentation data. Through analysis of historical data, key factors affecting fermentation components and physical quantities are found out, and corresponding mathematical models are established. For example, linear regression models, multivariate statistical analysis models, etc. can predict some indicators in the fermentation process. These models can predict the law of historical data to some extent, but they are often based on simple linear assumptions or statistical relationships, and are difficult to accurately describe the complex nonlinear dynamic changes in the fermentation process. Moreover, when the fermentation conditions change, the adaptability of these models is poor, and a large amount of data collection and model training need to be performed again. SUMMARY

[0007] The present disclosure provides a Raman spectrum feature extraction method and modeling method based on convolution and attention mechanism.

[0008] According to a first aspect of the present disclosure, a Raman spectrum feature extraction method based on convolution and attention mechanism is provided, which comprises:

[0009] Obtaining Raman spectrum time series data and physicochemical time series data generated by monitoring a target bioprocess in fermentation;

[0010] Fusing the Raman spectrum time series data and the physicochemical time series data to generate fermentation feature sequence data;

[0011] Preprocessing and enhancing the fermentation feature sequence data to obtain enhanced feature sequence data;

[0012] Based on the Raman spectrum feature extraction model, Raman spectrum features are extracted from the enhanced feature sequence data to predict Raman spectrum features and determine the components and their physical quantities produced by the bioprocess in fermentation.

[0013] According to a second aspect of the present disclosure, a Raman spectrum feature modeling method based on convolution and attention mechanism is provided, which comprises:

[0014] Obtaining Raman spectrum time series data samples and physicochemical time series data samples generated by monitoring sample bioprocesses in fermentation;

[0015] Fusing the Raman spectrum time series data samples and the physicochemical time series data samples to generate fermentation feature sequence data samples;

[0016] Obtaining component labels and their physical quantity labels corresponding to the fermentation feature sequence data samples;

[0017] Preprocessing and enhancing the fermentation feature sequence data samples to obtain enhanced feature sequence data;

[0018] Based on the to-be-trained model, acute forward propagation is performed on the enhanced feature sequence data samples to extract Raman spectrum feature samples therefrom, so as to predict component samples and their physical quantity samples produced by the sample bioprocess in fermentation;

[0019] According to the predicted component samples and their physical quantity samples and the component labels and their physical quantity labels, a loss value of model training is calculated, and based on the loss value, back propagation is performed to adjust the model parameters of the to-be-trained model until a Raman spectrum feature extraction model is obtained.

[0020] The technical effects of the present disclosure are as follows:

[0021] Traditional chemical analysis methods (such as HPLC, GC) are cumbersome to operate and time-consuming, and cannot achieve real-time online monitoring. However, by obtaining Raman spectrum time series data and physical and chemical time series data, the method can monitor the fermentation process in real time. Raman spectroscopy can quickly detect the fermentation system without damaging the sample, and combined with real-time collection of physical and chemical data, it can timely capture the changes of components and physical quantities in the fermentation process, providing the possibility of real-time adjustment of the fermentation process and avoiding the problem of out-of-control fermentation process caused by delayed detection.

[0022] Chemical analysis methods require frequent sampling of samples, which may interfere with the normal progress of the fermentation process and affect the accuracy of experimental results. The method based on real-time monitoring of Raman spectrum and physical and chemical data does not require a large number of samples of the fermentation system, reducing the interference with the fermentation process and ensuring the naturalness of the fermentation process and the reliability of the experimental results.

[0023] Traditional sensor technology can only measure specific parameters and cannot fully reflect the complex changes of components in the fermentation process. The method fuses Raman spectrum time series data and physical and chemical time series data, and Raman spectrum can provide rich molecular structure information and reflect the types and changes of biological components; physical and chemical data can describe the state of the fermentation environment. The fermentation feature sequence data generated by the fusion of the two contains more comprehensive information, which helps to better understand the fermentation process and improve the accuracy of the prediction of the components and physical quantities of the fermentation product.

[0024] Empirical and statistical models are often based on simple linear assumptions or statistical relationships, which are difficult to accurately describe the complex nonlinear dynamic changes in the fermentation process and have poor adaptability to changes in fermentation conditions. The Raman spectrum feature extraction model based on convolution and attention mechanism can automatically learn the complex features and patterns in the data. The convolution layer can effectively extract local features in the data and capture subtle changes in Raman spectrum and physical and chemical data; the attention mechanism can dynamically focus on key features related to fermentation product components and physical quantities and ignore irrelevant information. This model structure can better adapt to the nonlinear dynamic changes in the fermentation process, and even if the fermentation conditions change, it can also accurately predict through learning new data features, improving the adaptability and generalization ability of the model.

[0025] Traditional methods have certain limitations in data processing and model establishment. Chemical analysis methods require complex sample pretreatment processes, consuming a large amount of manpower and time; empirical and statistical models need to be retrained with a large amount of data collection when the fermentation conditions change. This method can fully utilize existing data information through data fusion, preprocessing and enhancement, etc., reducing unnecessary data collection and processing costs. At the same time, the model based on convolution and attention mechanism has strong learning ability and can quickly extract effective features from data, improving the prediction efficiency.

[0026] This method uses a deep learning model to automatically extract Raman spectrum features and predict fermentation product components and their physical quantities, reducing human interference and improving the objectivity and accuracy of the prediction. The model can continuously learn and optimize, and the prediction performance will continuously improve with the accumulation of data and the training of the model, providing strong support for the intelligent control and optimization of the fermentation process.

[0027] It should be understood that the content described in the summary section is not intended to limit the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0028] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail the following embodiments with reference to the attached drawings. The attached drawings are intended to better understand the present disclosure and do not limit the present disclosure. In the drawings, the same or similar reference numerals refer to the same or similar elements, and:

[0029] Figure 1 A structure block diagram of a Raman spectrum source data collection system according to an embodiment of the present disclosure is shown;

[0030] Figure 2 A structure block diagram of the Raman spectrum feature extraction model according to an embodiment of the present disclosure is shown;

[0031] Figure 3 A flowchart of a Raman spectrum feature extraction method based on convolution and attention mechanism according to an embodiment of the present disclosure is shown;

[0032] Figure 4(a) shows the waveform diagram of two batches of Raman spectrum source data collected in an application scenario;

[0033] Figure 4(b) shows the waveform diagram of the Raman spectrum source data in Figure 4(a) after preprocessing in an application scenario;

[0034] Figure 5 A comparison diagram of the prediction effect of the feature extraction method of the present disclosure and other related methods is shown;

[0035] Figure 6 The figure shows a comparison of the root mean square error (RMSE) of the prediction performance of the feature extraction method of the present invention with other related methods. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0037] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0038] Figure 1 A block diagram of a Raman spectroscopy source data acquisition system according to an embodiment of the present disclosure is shown; as follows: Figure 1 As shown, the fermenter contains fermenting organisms. Physicochemical sensors are fixed in the fermenter to detect the fermenting organisms and generate physicochemical source data. Raman spectral source data is generated by irradiating the fermenting organisms with a Raman light source and detecting the irradiated fermenting organisms with Raman light.

[0039] Figure 2 A structural block diagram of the Raman spectroscopy feature extraction model according to an embodiment of the present disclosure is shown; the Raman spectroscopy feature extraction model includes: a convolutional layer, an attention mechanism layer, and a multilayer perceptron, and a detailed description of these structures is given in the preferred embodiment.

[0040] Figure 3 A schematic flowchart of a Raman spectral feature extraction method based on convolution and attention mechanisms according to embodiments of the present disclosure is shown. Figure 3 As shown, it includes:

[0041] Acquire Raman spectral time series data and physicochemical time series data generated by monitoring the target biological process of fermentation;

[0042] The Raman spectroscopy time series data and the physicochemical time series data are fused to generate fermentation characteristic sequence data;

[0043] The fermentation feature sequence data is preprocessed and enhanced to obtain enhanced feature sequence data;

[0044] Based on the Raman spectrum feature extraction model, Raman spectrum features are extracted from the enhanced feature sequence data to predict Raman spectrum features and determine the components produced by the organism in the fermentation process and their physical quantities accordingly.

[0045] Figure 4(a) shows the waveform diagram of two batches of Raman spectrum source data collected in an application scenario; Figure 4(b) shows the waveform diagram of the Raman spectrum source data in Figure 4(a) after preprocessing in an application scenario. From the waveform, the fermentation feature sequence data after preprocessing removes the influence of fluorescence background and highlights the Raman feature information.

[0046] Optionally, the method further comprises:

[0047] Obtaining Raman spectrum source data generated by irradiating the organism in fermentation with a Raman light source;

[0048] The Raman spectrum data is sequenced according to the time stamp of real-time collection to generate Raman spectrum time sequence data.

[0049] Optionally, the method further comprises:

[0050] Obtaining physical and chemical source data generated by detecting the organism in fermentation with a physical and chemical sensor;

[0051] The physical and chemical source data is aligned with the Raman spectrum source data to generate physical and chemical time sequence data.

[0052] Optionally, the fusion of the Raman spectrum time sequence data and the physical and chemical time sequence data to generate fermentation feature sequence data comprises merging the Raman spectrum time sequence data and the physical and chemical time sequence data to generate fermentation feature sequence data.

[0053] Optionally, the preprocessing and enhancement of the fermentation feature sequence data to obtain enhanced feature sequence data comprises:

[0054] Based on a set of standardized sliding windows, the fermentation feature sequence data is processed by sliding window, and based on a set of standard deviation multiples, the mean and standard deviation of data points in each standardized sliding window are calculated;

[0055] Based on the mean and standard deviation, the maximum range value and the minimum range value of the data points in each standardized sliding window are calculated;

[0056] normalizing the data points in each of the standardized sliding windows based on the maximum range value and the minimum range value, so that the data point distribution of each of the standardized sliding windows reaches a set distribution standard to generate standardized feature sequence data;

[0057] enhancing the standardized feature sequence data to obtain enhanced feature sequence data.

[0058] Optionally, the size of the standardized sliding window is 3, and the standard deviation multiple is 1.5. To this end, when calculating the mean value and the standard deviation of the data points in each standardized sliding window based on the set standard deviation multiple, the mean value and the standard deviation of each data point and its adjacent three data points are calculated, and the standard deviation multiple is 1.5, which means that when calculating the mean value and the standard deviation of the data points in each standardized sliding window, the calculated maximum range value and the minimum range value are located between 1.5 and 0.

[0059] Specifically, when normalizing the data points in each of the standardized sliding windows, the formula is used for the data points The normalized fermentation feature data is calculated, x is the original fermentation feature data, μ is the mean value, σ is the standard deviation, and x' is the normalized fermentation feature data. If the normalized fermentation feature data exceeds the maximum range value, it is adjusted to the maximum range value. If it is lower than the minimum range value, it is adjusted to the minimum range value.

[0060] Optionally, the enhancing the standardized feature sequence data to obtain enhanced feature sequence data comprises:

[0061] based on a set filter, the standardized feature sequence data is denoised to generate denoised feature sequence data;

[0062] deriving the denoised feature sequence data to obtain the enhanced feature sequence data.

[0063] Optionally, the filter is a Savitzky-Golay filter, and the window length of the Savitzky-Golay filter is 13. When the standardized feature sequence data is denoised based on the set filter to generate denoised feature sequence data, based on each Savitzky-Golay filter, 13 data points around each data point are used to perform least squares fitting polynomial to perform data point interpolation to generate denoised feature sequence data.

[0064] Optionally, the polynomial order of the least squares fitting polynomial is 2.

[0065] Specifically, the window length is 13, and the polynomial order is 2, meaning that in the neighborhood of each data point, 13 consecutive data points will be used to fit a second-order polynomial. The core idea is to fit a polynomial in the neighborhood of each data point using the least squares method, and then use the value of this polynomial at the data point to replace the original value, thereby achieving smoothing. For a given filter window size m (m = 13) and polynomial order n (n = 2), the Savitzky-Golay filter can be represented as:

[0066]

[0067] where a k is the polynomial coefficient determined by least squares fitting, x(t) is the normalized feature value of the center data point in the window, and x(t-i) is the normalized feature value of other points in the window. In this way, each point can be expressed as a weighted linear combination of other points in the window to achieve data point interpolation to achieve smoothing effect, and to generate denoised feature sequence data.

[0068] Optionally, when the derivative of the denoised feature sequence data is processed, the derivative order is 2, so as to perform second-order derivative processing on the denoised feature sequence data and extract the change rate of the denoised feature sequence data accordingly, to obtain the enhanced feature sequence data.

[0069] Specifically, the second-order derivative processing can be represented as:

[0070]

[0071] In discrete data, the second-order derivative can be approximated by the following formula:

[0072]

[0073] where h is the step size, x i-1 is the denoised feature value corresponding to the i-1th data point, f(x i-1 ) is the intensity corresponding to the denoised feature value of the i-1th data point; f(x i ) is the intensity corresponding to the denoised feature value of the i-1th data point; f(x i+1 ) is the intensity corresponding to the denoised feature value of the i+1th data point, x i is the denoised feature value corresponding to the i-1th data point, x i+1 is the denoised feature value corresponding to the i-1th data point, f”(x i ) is the second-order derivative, i.e., the change rate of the denoised feature sequence data, thereby identifying the extreme points therein to generate enhanced feature sequence data, which usually corresponds to the frequency of molecular vibration, to provide information about the molecular structure.

[0074] Optionally, the Raman spectrum feature extraction model comprises a convolution layer, an attention mechanism layer, and a multi-layer perceptron.

[0075] The Raman spectrum feature is extracted from the enhanced feature sequence data based on the Raman spectrum feature extraction model to predict the Raman spectrum feature and determine the ingredients and physical quantities generated by the organism in the fermentation process, comprising:

[0076] The convolution layer is used to slide on the enhanced feature sequence data to extract local feature patterns and spatial correlation features, wherein the local feature patterns are associated with the Raman spectrum intensity of the product of the organism in the fermentation process, and the spatial correlation features are associated with the spatial correlation between the physicochemical and Raman spectrum intensity of the organism in the fermentation process;

[0077] The attention mechanism layer is used to extract attention features from the local feature patterns and spatial correlation features to generate a context feature vector, wherein the context feature vector is associated with the Raman spectrum change and physicochemical change trend between different fermentation states of the organism in the fermentation process;

[0078] The multi-layer perceptron is used to map the context feature vector to generate a hyperplane feature vector, and compress the hyperplane feature vector to generate a metabolic pathway related feature vector to predict the Raman spectrum feature and determine the ingredients and physical quantities generated by the organism in the fermentation process.

[0079] Optionally, the convolution layer is a sequential model comprising a first convolution layer, a first batch normalization layer, a second convolution layer, a second batch normalization layer, a third convolution layer, and a third batch normalization layer, wherein the number of convolution kernels in the first convolution layer is greater than that in the second and third convolution layers, and the size of a single convolution kernel in the first convolution layer is greater than that in the second and third convolution layers.

[0080] The convolution layer is used to slide on the enhanced feature sequence data to extract local feature patterns and spatial correlation features, comprising:

[0081] The first convolution layer is used to slide on the enhanced feature sequence data to extract first local feature patterns and first spatial correlation features.

[0082] The first batch normalization layer is used to normalize the first local feature patterns and first spatial correlation features to generate first normalized local feature patterns and first normalized spatial correlation features.

[0083] based on the second convolutional layer, sliding on the first normalized local feature pattern and first normalized spatial correlation feature to extract a second local feature pattern and a second spatial correlation feature;

[0084] based on the second batch normalization layer, normalizing the second local feature pattern and the second spatial correlation feature to generate a second normalized local feature pattern and a second normalized spatial correlation feature;

[0085] based on the third convolutional layer, sliding on the second normalized local feature pattern and the second normalized spatial correlation feature to extract a third local feature pattern and a third spatial correlation feature;

[0086] based on the third batch normalization layer, normalizing the third local feature pattern and the third spatial correlation feature to generate a third normalized local feature pattern and a third normalized spatial correlation feature, as the local feature pattern and the spatial correlation feature extracted from the enhanced feature sequence data, respectively.

[0087] In the Raman spectrum feature extraction method based on convolution and attention mechanism, the convolutional layer is an important component, which processes the enhanced feature sequence data through a sequential model structure to extract local feature patterns and spatial correlation features. The sequential model is composed of multiple convolutional layers and batch normalization layers alternately. Each convolutional layer is responsible for feature extraction, and the batch normalization layer normalizes the extracted features to speed up model training and improve generalization ability.

[0088] Specifically, the specific technical implementation details of the convolutional layer sliding on the enhanced feature sequence data to extract local feature patterns and spatial correlation features are as follows:

[0089] The first convolutional layer is set with n1 convolutional kernels, each with a size of k1xk1. In the bio-fermentation monitoring scene, a larger number of convolutional kernels and a larger convolutional kernel size can capture more global features, such as detecting feature information related to multiple substances in a larger wave number range in Raman spectrum data, and overall trends of sensor data in a longer time interval. The specific value of n1 can be adjusted according to the complexity of the data and the performance requirements of the model, generally between 1664.

[0090] The enhanced feature sequence data X has a dimension of HxWxC, where H represents the height of the data (in Raman spectrum data, it can be understood as the length of the time series), W represents the width of the data (such as the wave number range of the spectrum), and C represents the number of channels of the data (including Raman spectrum data and physical and chemical data).

[0091] For each convolution kernel K i (i = 1, 2, …, n1), which is k1xk1xC in size, performs a sliding convolution operation on the input data X. The specific calculation process is as follows:

[0092]

[0093] where Y i is the feature map after convolution of the i-th convolution kernel, (h, w) is the position on the feature map, b i is the bias term corresponding to the i-th convolution kernel. In this way, the first convolution layer slides on the enhanced feature sequence data, extracts the first local feature pattern and the first spatial correlation feature, and obtains an output feature map with a dimension of H1xW1xn1, where s is the convolution step, usually set to 1.

[0094] The feature map Y output by the first convolution layer (including the first local feature pattern and the first spatial correlation feature) has a dimension of H1xW1xn1.

[0095] Batch normalization processing is performed on each channel c. For each small batch data B containing N samples, the calculation formula is as follows:

[0096]

[0097] where μ B,c is the mean of the small batch data in channel c, is the variance of the small batch data in channel c, is the normalized feature value, ò is a very small constant (usually set to 10 -5 ) to prevent the denominator from being zero, γ c and β c are learnable parameters for scaling and shifting the normalized feature value, Y i,:,:,c is the final normalized feature value. After processing by the first batch normalization layer, the first normalized local feature pattern and the first normalized spatial correlation feature are obtained, with a dimension of H1xW1xn1.

[0098] Second convolution layer: set n2 convolution kernels, each with a size of k2xk2, and n2

[0099] The first normalized local feature pattern and the first normalized spatial correlation feature output by the first batch normalization layer have a dimension of H1xW1xn1.

[0100] Similar to the first convolution layer, each convolution kernel K j (j = 1, 2,..., n2) of the second convolution layer slides and convolves on the input data. The calculation formula is:

[0101]

[0102] wherein Z j is the feature map after convolution of the jth convolution kernel, (h', w') is the position on the feature map, b j is the bias term corresponding to the jth convolution kernel. The output feature map has a dimension of H2xW2xn2, wherein

[0103] The feature map Z output by the second convolution layer (including the second local feature pattern and the second spatial correlation feature) has a dimension of H2xW2xn2.

[0104] Similar to the first batch normalization layer, batch normalization is performed on each channel c. For small batch data B containing N samples, the calculation formula is as follows:

[0105]

[0106] After the second batch normalization layer, the second normalized local feature pattern and the second normalized spatial correlation feature are obtained, and the dimension is still H2xW2xn2.

[0107] Third convolution layer: set n3 convolution kernels, each convolution kernel has a size of k3xk3, and n3≤n2, k3≤k2. The third convolution layer continues to refine and abstract the features of the previous layer, and extracts the most critical feature information.

[0108] The second normalized local feature pattern and the second normalized spatial correlation feature output by the second batch normalization layer have a dimension of H2xW2xn2.

[0109] Each convolution kernel K l (l = 1, 2,..., n3) of the third convolution layer has a size of k3xk3xn2 and slides and convolves on the input data. The calculation formula is:

[0110]

[0111] wherein U lis the feature map after the convolution of the lth convolution kernel (including the third local feature pattern and the third spatial correlation feature), (h", w") is the position on the feature map, b l is the bias term corresponding to the lth convolution kernel. The output feature map has a dimension of H3xW3xn3, where

[0112] The feature map U output by the third convolution layer has a dimension of H3xW3xn3.

[0113] Batch normalization is also performed on each channel c. For a small batch of data B containing N samples, the calculation formula is as follows:

[0114]

[0115]

[0116] After the third batch normalization layer, the third normalized local feature pattern and the third normalized spatial correlation feature are obtained, which still have a dimension of H3xW3xn3, serving as the final local feature pattern and spatial correlation feature extracted from the enhanced feature sequence data.

[0117] Therefore, in contrast, in traditional convolutional neural network applications, the number and size of convolution kernels are often fixed, making it difficult to adapt to feature extraction requirements at different levels and scales. In the present technology, the first convolution layer is provided with n1 convolution kernels with a size of k1xk1, the second convolution layer is provided with n2 convolution kernels with a size of k2xk2 (n2 The larger convolution kernel size k1 and the larger number of convolution kernels n1 enable the first convolution layer to capture data features in a larger spatial range, for example, in Raman spectrum data, it can cover a wider wave number range and a longer time series, detect macroscopic feature information related to multiple substances, and detect the overall trend of sensor data in a longer time interval.

[0118] As the convolution layers go deeper, the second convolution layer and the third convolution layer focus on more local and more detailed features by reducing the size and number of convolution kernels. Taking the second convolution layer as an example, The smaller k2 enables the convolution kernel to capture microscopic features in the data in more detail, such as the small changes in the molecular structure of microbial cells in the spectrum, which are manifested as features in a smaller wave number range. This multi-scale feature extraction method can comprehensively mine information at different levels in the biological fermentation process, from overall trends to microscopic details, which is more in line with the complexity and diversity of data in the biological fermentation process than the traditional fixed convolution kernel setting method.

[0119] The traditional method cannot effectively distinguish the importance of different levels of features, resulting in inaccurate feature extraction. In the present technology, with the progression of convolution layers, the reduction in the number of convolution kernels prompts the model to filter and integrate the extracted features, achieving step-by-step refinement and abstraction of features. In biological fermentation monitoring, this means that the model can gradually focus on key features that have a significant impact on the fermentation process from a large number of original Raman spectrum and sensor data features. For example, after extracting rich original features in the first convolution layer, the second convolution layer filters and recombines these features through a smaller number of convolution kernels, extracting more representative features, and the third convolution layer further refines the features so that the final extracted features can more accurately reflect the essential characteristics of the biological fermentation process, providing more valuable information for subsequent analysis and prediction.

[0120] The traditional model often shows poor generalization ability when facing different biological fermentation data sets due to differences in data distribution. The batch normalization layer normalizes the features, making the model more adaptable to different data distributions. Normalization operations on each small batch of data can unify the distribution of data to a relatively stable range, reducing data noise and interference. For example, in different batches of biological fermentation experiments, Raman spectrum data may fluctuate slightly due to small differences in experimental conditions, and the batch normalization layer can effectively handle these fluctuations, enabling the model to perform well on different batches of data and improving the model's generalization ability. Meanwhile, the learnable parameters γ c and β c can adjust the normalized features according to the characteristics of the data, further enhancing the model's adaptability to different data.

[0121] Optionally, the attention mechanism layer includes a first attention mechanism module and a second attention mechanism module.

[0122] The attention mechanism layer is used to extract attention features from the local feature patterns and spatial correlation features to generate a context feature vector. The context feature vector is associated with the Raman spectrum changes and physical and chemical change trends between different fermentation states of the biological fermentation process, including:

[0123] The first attention mechanism module is used to extract attention features from the local feature patterns to generate a first context feature vector.

[0124] The second attention mechanism module is used to extract attention features from the spatial correlation features to generate a second context feature vector.

[0125] The first context feature vector and the second context feature vector are fused to generate a fused context feature vector as the context feature vector.

[0126] Optionally, the attention feature extraction on the local feature pattern based on the first attention mechanism module includes:

[0127] The local feature pattern is multiplied by a set first linearization weight matrix to obtain a first query vector;

[0128] The query vector is multiplied by a set first attention weight matrix to obtain a first attention score;

[0129] The first attention score is normalized to obtain a first attention weight;

[0130] The local feature pattern and the first attention weight are multiplied and summed and flattened to generate a first context vector.

[0131] Optionally, the attention feature extraction on the spatial correlation feature based on the second attention mechanism module includes:

[0132] The spatial correlation feature is multiplied by a set second linearization weight matrix to obtain a second query vector;

[0133] The query vector is multiplied by a set second attention weight matrix to obtain a second attention score;

[0134] The second attention score is normalized to obtain a second attention weight;

[0135] The local feature pattern and the second attention weight are multiplied and summed and flattened to generate a second context vector.

[0136] In the scene of Raman spectrum feature extraction for biological fermentation monitoring, the design of the attention mechanism layer aims to fully mine the local feature pattern and the spatial correlation feature to generate a context feature vector that can accurately reflect the trend of Raman spectrum and physical and chemical changes in the biological fermentation process. This layer is composed of a first attention mechanism module and a second attention mechanism module, which respectively process different types of features, and finally fuse to obtain the final context feature vector.

[0137] Specifically, for the first attention mechanism module, the local feature pattern is denoted as X lis a tensor with dimensions I x J x K. In the context of biological fermentation monitoring, I can represent the length of the time series of Raman spectral data, reflecting the changes in the fermentation process over time; J represents the spectral dimension, such as the wave number range, covering the Raman characteristic information of different substances; K represents the number of channels related to local feature patterns, which may include Raman spectral intensity information of different fermentation products or other attributes related to local features.

[0138] The first linearization weight matrix is denoted as W l1 with dimensions K x D1. Here D1 is a hyperparameter that determines the dimension of the first query vector. The role of this matrix is to map the input local feature pattern X l to a new feature space for subsequent attention score calculation.

[0139] The first attention weight matrix is denoted as A l1 with dimensions D1 x D1. It is used to calculate the attention scores between different positions in the local feature pattern, reflecting the correlation and importance between positions.

[0140] The first query vector is calculated as: where i = 1,..., I, j = 1,..., J, and d = 1,..., D1. By multiplying each element of the local feature pattern X l with the corresponding element of the first linearization weight matrix W l1 and summing them up, the first query vector Q l is obtained, with dimensions I x J x D1. This step converts the original local feature pattern into a form more suitable for attention score calculation, highlighting important features related to subsequent calculations through linear transformation of the weight matrix.

[0141] The first attention score is calculated as: where i = 1,..., I, j = 1,..., J, and d2 = 1,..., D1. Matrix multiplication is performed between the first query vector Q l and the first attention weight matrix A l1 to obtain the first attention score S l with dimensions I x J x D1. This score represents the importance of each position in the local feature pattern relative to other positions and is the basis for subsequent attention weight calculation.

[0142] The first attention weight is calculated as: where i = 1,..., I, j = 1,..., J, and d2 = 1,..., D1.

[0143] The first attention score S lSoftmax normalization is performed. The Softmax function calculates the exponential of each position's attention score and divides it by the sum of the exponentials of all position attention scores, so that each position's attention weight is between 0 and 1, and the sum of all position attention weights is 1. The first attention weight P l , with dimensions I x J x D1, accurately reflects the relative importance probability distribution of each position in the local feature pattern.

[0144] The first context vector is generated: First, the local feature pattern X l is element-wise multiplied with the first attention weight P l , and then summed in the I, J, and D1 dimensions to obtain an intermediate result. Finally, this intermediate result is flattened into a one-dimensional vector through the Flatten operation to obtain the first context vector C l . This step combines the local feature pattern with the attention weight, highlighting important feature information and converting it into a vector form for subsequent processing.

[0145] For the second attention mechanism module, the spatial correlation feature is denoted as X s , which is a tensor with dimensions I x J x L. In biological fermentation monitoring, I and J have the same meaning as in the local feature pattern, and L represents the number of channels related to the spatial correlation feature, which contains spatial correlation information of physical and chemical and Raman spectrum intensity during the fermentation process of the organism, such as the relationship between the distribution of different substances in space and the Raman spectrum characteristics.

[0146] The second linearization weight matrix is denoted as W l2 , with dimensions L x D2. Where D2 is a hyperparameter that determines the dimension of the second query vector. This matrix is used to map the spatial correlation feature X s to a new feature space for subsequent calculation of attention scores.

[0147] The second attention weight matrix is denoted as A l2 , with dimensions D2 x D2. It is used to calculate the attention scores between different positions in the spatial correlation feature.

[0148] The second query vector is calculated: where i = 1,..., I, j = 1,..., J, and d = 1,..., D2. Similar to the method of calculating the query vector in the first attention mechanism module, by multiplying each element of the spatial correlation feature X s with the corresponding element of the second linearization weight matrix W l2 and summing them up, the second query vector Q s is obtained, with dimensions I x J x D2.

[0149] The second attention score is calculated as follows: where i = 1,..., I, j = 1,..., J, and d2 = 1,..., D2. The second query vector Q s is multiplied by the second attention weight matrix A l2 to obtain the second attention score S s with dimensions I x J x D2.

[0150] The second attention weight is calculated as follows: where i = 1,..., I, j = 1,..., J, and d2 = 1,..., D2.

[0151] The second attention score S s is normalized by Softmax to obtain the second attention weight P s with dimensions I x J x D2, which reflects the relative importance probability distribution of each position in the spatial correlation feature.

[0152] The second context vector is generated as follows:

[0153] The spatial correlation feature X s is multiplied element-wise by the second attention weight P s , then summed over the I, J, and D2 dimensions, and finally flattened into a one-dimensional vector by the Flatten operation to obtain the second context vector C s .

[0154] For context vector fusion, the first context vector C l and the second context vector C s have lengths I x J and I x J, respectively. An innovative fusion method is adopted, combining weighted summation, element multiplication, and nonlinear transformation.

[0155] First, the weighted sum vector C w is calculated as follows: w C l = a · C s + (1 - a) · C p , where a is a hyperparameter between 0 and 1, controlling the relative importance of the first and second context vectors in the fusion process.

[0156] Then, the element product vector C w is calculated as follows: where | denotes the element product operation, i.e., multiplying the elements at corresponding positions.

[0157] Next, C w and C p are subjected to nonlinear transformation: C nf1= tanh(C w ), C nf2 = sigmoid(C p ). Here, the tanh function and the sigmoid function are used to introduce nonlinearity, enhancing the model's expressive power.

[0158] Finally, the fused context feature vector C is calculated: C = Flatten(C nf1 + C nf2 ). The two vectors after nonlinear transformation are added together, and then the result is flattened into a one-dimensional vector through the Flatten operation to obtain the final fused context feature vector C.

[0159] To this end, unlike the traditional attention mechanism that mixes all features for processing, the present technology separates local feature patterns and spatial correlation features for processing, and deeply mines them through two independent attention mechanism modules. This way can more accurately capture key information of different types of features, more accurately extract local features related to fermentation products and spatial correlation features of physical and chemical and Raman spectrum intensity in biological fermentation monitoring, and improve the accuracy and efficiency of feature extraction.

[0160] The traditional context vector fusion method is usually simple, such as direct splicing or simple weighted summation. However, the present technology adopts an innovative fusion strategy, combining weighted summation, element multiplication, and nonlinear transformation. This complex fusion method can fully utilize the information of the two context vectors, not only considering their weight distribution, but also mining the potential relationship between the two vectors through element multiplication, and then enhancing the model's expressive power through nonlinear transformation, so that the fused context feature vector can more comprehensively and accurately reflect the Raman spectrum changes and physical and chemical change trends between different fermentation states in the biological fermentation process.

[0161] In addition, in each attention mechanism module, by setting different linearization weight matrices and attention weight matrices, and setting adjustable hyperparameters a in the fusion process, the model can be flexibly adjusted according to the specific characteristics of the biological fermentation data. This flexibility enables the model to better adapt to different experimental conditions and data characteristics, and has stronger generalization ability and adaptability compared to traditional fixed parameter models, and can achieve more accurate analysis and prediction in various biological fermentation monitoring scenarios.

[0162] Optionally, the multi-layer perceptron comprises: a first fully connected layer, a first nonlinear activation layer, a second fully connected layer, a second nonlinear activation layer, and an output layer, wherein the first fully connected layer is connected to the second fully connected layer;

[0163] based on the multi-layer perceptron, performing feature mapping on the context feature vector to generate a hyperplane feature vector, and performing compression on the hyperplane feature vector to generate a metabolic pathway related feature vector to predict Raman spectrum features and determine the components and their physical quantities produced by the organism in the fermentation process based on the Raman spectrum features, including:

[0164] based on the first fully connected layer, performing feature mapping on the context feature vector to generate a hyperplane feature vector and performing nonlinear activation on the hyperplane feature vector through the first nonlinear activation layer;

[0165] based on the second fully connected layer, performing compression on the nonlinearly activated hyperplane feature vector to generate a metabolic pathway related feature vector and performing nonlinear activation on the metabolic pathway related feature vector through the second nonlinear activation layer;

[0166] based on the output layer, performing linear weighted summation and bias adjustment on the nonlinearly activated metabolic pathway related feature vector to predict Raman spectrum features and determine the components and their physical quantities produced by the organism in the fermentation process based on the Raman spectrum features.

[0167] In the technical system of Raman spectrum feature extraction based on convolution and attention mechanism for biological fermentation monitoring, the multi-layer perceptron (MLP) plays a core role in gradually transforming the context feature vector into Raman spectrum feature prediction value, and then realizing the prediction of the components and their physical quantities of biological fermentation products. It is composed of multiple functional layers, which deeply process the input information through a series of feature mapping, compression and activation operations, providing key support for accurate prediction of the biological fermentation process.

[0168] Specifically, the technical steps of the multi-layer perceptron (MLP) are as follows:

[0169] For the first fully connected layer, the context feature vector is denoted as X, which is a one-dimensional vector with length N. In the context of biological fermentation monitoring, this vector is the result obtained after processing by the previous attention mechanism layer, containing rich information such as Raman spectrum changes and physicochemical change trends under different fermentation states in the biological fermentation process. These information is the basis for the model to understand the fermentation process and make subsequent predictions.

[0170] The weight matrix of the first fully connected layer is denoted as W1, with dimensions N x M1. M1 is a hyperparameter representing the number of features output by the first fully connected layer. This matrix determines the linear transformation of the context feature vector in this layer, and through the setting of weights, it adjusts the importance and influence of different features in subsequent calculations.

[0171] The bias vector of the first fully connected layer is denoted as b1, with dimension M1. It introduces an offset to the result of the linear transformation, increasing the flexibility of the model and enabling it to better fit complex data distributions and capture subtle features in the data.

[0172] During feature mapping: Z1 = X·W1 + b1, the context feature vector X is multiplied by the weight matrix W1 of the first fully connected layer, and then the bias vector b1 is added to obtain a vector Z1 of dimension M1. This step achieves linear feature mapping of the context feature vector, transforming the input feature vector from the original space to a new feature space. In the field of bio-fermentation monitoring, this process can be seen as a preliminary sorting and integration of the relationships between different features, highlighting the potential information related to Raman spectral features during bio-fermentation, and preparing for subsequent nonlinear activation operations.

[0173] For the first nonlinear activation layer, the output Z1 of the previous layer has a dimension of M1. The activation function chosen is the Swish function, defined as f(x) = x·σ(x), where... It is the Sigmoid function.

[0174] When performing nonlinear activation, use Applying the Swish activation function to each element of the input vector Z1 yields a vector A1 after nonlinear activation, still with dimension M1. Compared to traditional simple activation functions (such as ReLU), the Swish function possesses characteristics such as continuous differentiability and non-monotonicity, enabling it to more accurately capture the complex nonlinear relationships between data in the biofermentation process. Biofermentation is a highly complex nonlinear system involving numerous interrelated physical, chemical, and biochemical reactions. The Swish function can better simulate the complex interactions between features in these reactions, significantly enhancing the model's expressive power and allowing it to learn more complex Raman spectral feature patterns and product relationships during biofermentation.

[0175] For the second fully connected layer, the input vector is the output A1 of the first nonlinear activation layer, with dimension M1.

[0176] The weight matrix of the second fully connected layer is denoted as W2, with dimensions M1×M2. M2 is a hyperparameter, and M2 < M1, meaning that the number of features output by the second fully connected layer is less than the number of input features. This setting allows the second fully connected layer to compress the input features and extract more crucial information.

[0177] The bias vector of the second fully connected layer is denoted as b2, and its dimension is M2.

[0178] In the process of feature compression and mapping, Z2 = A1·W2 + b2 is used to perform matrix multiplication between the vector A1 processed by the first nonlinear activation layer and the second fully connected layer weight matrix W2, and then add the bias vector b2 to obtain a vector Z2 with a dimension of M2. This step further linearly transforms the input features while compressing the features by reducing the number of output features. In biological fermentation monitoring, this operation helps to remove redundant information and focus on key Raman spectral features more directly related to biological metabolic pathways, laying the foundation for subsequent generation of more targeted metabolic pathway-related feature vectors.

[0179] For the second nonlinear activation layer, the input vector is the output Z2 of the previous layer, with a dimension of M2.

[0180] The activation function uses the ELU (Exponential Linear Unit) function, which is defined as where α is an adjustable parameter, usually set to 1.

[0181] In the process of nonlinear activation, A2 = ELU(Z2) is used. Apply the ELU activation function to each element of the input vector Z2 to obtain the vector A2 after nonlinear activation, still with a dimension of M2. The ELU function cleverly combines the linear characteristics of the ReLU function in the positive part and the smooth transition characteristics in the negative part, which is of great significance in the training of complex models in biological fermentation monitoring.

[0182] For the output layer, the input vector is the output A2 of the second nonlinear activation layer, with a dimension of M2.

[0183] The output layer weight vector is denoted as W3, with a dimension of M2×1. Its role is to map the metabolic pathway-related feature vector obtained after the previous multi-layer processing to the final Raman spectrum feature prediction space.

[0184] The output layer bias scalar is denoted as b3, which is a scalar value.

[0185] In the process of linear weighted summation and bias adjustment, based on the formula Y = A2·W3 + b3, the metabolic pathway-related feature vector A2 processed by the second nonlinear activation layer is multiplied by the output layer weight vector W3, and then the bias scalar b3 is added to obtain a scalar value Y. This Y value is the predicted value of the Raman spectrum feature in the biological fermentation process.

[0186] After obtaining the Raman spectrum characteristic prediction value Y, it needs to be matched with the pre-specified relationship spectrum of the composition and physical and chemical quantity. This relationship spectrum is constructed based on a large amount of experimental data, theoretical research and long-term biological fermentation practical experience. It records in detail the corresponding relationship between different Raman spectrum characteristic values and the composition and physical and chemical quantity produced in the biological fermentation process.

[0187] By matching the Y value with the relationship spectrum, the actual composition and physical quantity produced by the organism in the fermentation process can be further determined according to the known corresponding relationship. For example, the relationship spectrum may explicitly record the concentration range, yield and other physical quantity information of a specific component in the fermentation product within a specific Raman spectrum characteristic prediction value range. This matching process provides a conversion bridge from abstract Raman spectrum characteristic prediction value to specific biological fermentation product information, so that the model prediction result can be closely linked with the actual biological fermentation situation, thereby more accurately understanding the state and product situation of the biological fermentation process.

[0188] Therefore, unlike the simple activation function (such as ReLU) commonly used in traditional multi-layer perceptron, the present technology adopts Swish function in the first nonlinear activation layer and ELU function in the second nonlinear activation layer. These complex activation functions can more accurately capture the complex nonlinear relationship in the biological fermentation process due to their excellent nonlinear expression ability and unique gradient characteristics, greatly improving the prediction accuracy of the Raman spectrum characteristics, and thus more accurately reflecting the changes of the composition and physical quantity in the biological fermentation process.

[0189] In addition, by setting the output feature quantity of the second fully connected layer to be less than that of the first fully connected layer, effective compression of the features is realized. This hierarchical processing method can gradually focus on the key Raman spectrum characteristics closely related to the biological metabolic pathway, remove redundant information, and make the model more focused on the prediction of the composition and physical quantity of the biological fermentation product. At the same time, the nonlinear activation operation of each layer further enhances the prediction ability of the model for complex Raman spectrum characteristic patterns, and compared with the traditional single-layer or simple structure model, it can more deeply mine the potential information in the data and improve the prediction accuracy of the model.

[0190] Furthermore, each layer of the multi-layer perceptron is equipped with adjustable parameters (such as weight matrix and bias vector), and adjustable parameters (such as α in the ELU function) are also set in the activation function. This flexible parameter setting gives the model strong adaptability, and the technical personnel can make fine adjustments according to the characteristics of different biological fermentation data. By optimizing these parameters, the model can better adapt to various biological fermentation scenarios and monitoring targets, improve the generalization ability of the model, and ensure accurate prediction under different experimental conditions and production environments.

[0191] The embodiment of the application also provides a Raman spectrum feature modeling method based on convolution and attention mechanism, which comprises the following steps:

[0192] Obtaining Raman spectrum time series data samples and physicochemical time series data samples generated by monitoring a fermented sample biological process;

[0193] Fusing the Raman spectrum time series data samples and the physicochemical time series data samples to generate fermentation feature sequence data samples;

[0194] Obtaining ingredient labels and physical quantity labels corresponding to the fermentation feature sequence data samples;

[0195] Preprocessing and enhancing the fermentation feature sequence data samples to obtain enhanced feature sequence data samples;

[0196] Based on a to-be-trained model, acute forward propagation is performed on the enhanced feature sequence data samples to extract Raman spectrum feature samples therefrom, so as to predict ingredient samples and physical quantity samples generated by the sample biological process in the fermentation process;

[0197] According to the predicted ingredient samples and physical quantity samples and the ingredient labels and physical quantity labels, a loss value of model training is calculated, and based on the loss value, back propagation is performed to adjust model parameters of the to-be-trained model until a Raman spectrum feature extraction model is obtained.

[0198] Specifically, in the above model training, a professional Raman spectrometer is used to monitor the fermented sample biological process in real time, and Raman spectrum data at different time points are obtained to form a time sequence. Raman spectrum can reflect the structure and composition information of biological molecules. Different biological ingredients have specific peak values and characteristics on the Raman spectrum. These data provide important information about the change of ingredients in the biological fermentation process for the model.

[0199] In addition, various sensors are used to synchronously collect physicochemical parameters related to the fermentation process, such as temperature, pH value, dissolved oxygen concentration, pressure, etc., to form a time sequence. These physicochemical parameters have an important influence on the biological fermentation process, and they are closely related to the generation and transformation of biological ingredients. Combined with the Raman spectrum data, the fermentation process can be more comprehensively described.

[0200] The component labels and their physical quantity labels correspond to the fermentation feature sequence data samples. These labels are determined through experimental measurements, chemical analysis, etc. They represent the actual components produced by the sample organism during fermentation and the physical quantities (such as concentration, yield, etc.) of these components. The component labels are used to indicate which specific biological components are included in the fermentation product, and the physical quantity labels give quantitative information about these components. The accuracy of the labels is crucial for the training of the model, as they are the target of the model's learning. The model adjusts its parameters continuously in an attempt to make the predicted results as close as possible to the labels.

[0201] Based on the to-be-trained model, the enhanced feature sequence data samples are forward propagated to extract Raman spectrum feature samples therefrom, and to predict the component samples and their physical quantity samples produced by the sample organism during fermentation. The to-be-trained model is usually a neural network model based on convolution and attention mechanism, and the structure of the model is composed of convolution layers, attention mechanism layers, and multi-layer perceptrons, and the functions of the layers are as follows:

[0202] Convolution layer: The convolution layer performs sliding convolution operations on the input fermentation feature sequence data through convolution kernels. The convolution kernel is a small weight matrix that performs weighted summation on the local region of the data when it slides, thereby extracting local features. For example, for two-dimensional Raman spectrum data, the convolution kernel can capture local patterns in the spectrum, such as specific peak combinations; for time series data, the convolution kernel can extract local trends and changes in the time series.

[0203] The convolution layer can automatically learn local features in the data, reducing the dimensionality of the data while preserving important information. Through the stacking of multiple convolution layers, features of different levels and complexities can be extracted, from simple local patterns to more complex feature combinations.

[0204] Attention mechanism layer: aims to calculate the attention weights of the input features to dynamically adjust the importance of different features. In biological fermentation monitoring, not all features have equal importance for predicting components and physical quantities. The attention mechanism calculates the weight of each feature, allowing the model to focus on features more relevant to the target task (predicting components and physical quantities). Specifically, it maps the input features to an attention weight vector through a calculation process, and each element of the vector represents the importance of the corresponding feature.

[0205] The attention mechanism can help the model better handle complex data, improve the accuracy and relevance of feature extraction. It can ignore some irrelevant or interfering features and enhance the influence of features closely related to the components and physical quantities of the fermentation product, thereby improving the prediction performance of the model.

[0206] Multi-Layer Perceptron (MLP): A type of artificial neural network composed of multiple fully connected layers that further map and classify / regress the feature vectors processed by the convolutional layers and attention mechanism layers. In this model, the MLP receives the output features from the attention mechanism layer and maps them to the final prediction space through a series of linear transformations and nonlinear activation functions.

[0207] The MLP can learn complex nonlinear relationships between features by adjusting weights to establish an accurate mapping between input features and prediction results. When predicting component samples, the MLP can output probability distributions for each component; when predicting physical quantity samples, the MLP can output specific numerical predictions.

[0208] The predicted component samples and their physical quantity samples are compared with the corresponding component labels and their physical quantity labels to calculate the loss value. For component prediction, if it is a classification problem, the cross-entropy loss function can be used; for physical quantity prediction, the mean squared error (MSE) loss function is usually used. The loss value reflects the difference between the model's prediction results and the true labels, and the training goal of the model is to minimize this loss value.

[0209] Backpropagation: Based on the calculated loss value, backpropagation is performed to adjust the model's parameters. Backpropagation is a gradient descent-based optimization algorithm that calculates the gradient of the loss function with respect to the model's parameters (such as the convolution kernel weights of the convolutional layers, the parameters of the attention mechanism layer, and the weights and biases of the MLP) to determine the direction of parameter updates. Specifically, starting from the output layer, the gradient of each layer is calculated in reverse based on the gradient of the loss function with respect to the output, and then the parameters of each layer are updated based on the gradient. Through continuous iteration of this process, the model's parameters are gradually adjusted so that the loss value continuously decreases until a pre-set convergence condition is met (such as the loss value being less than a certain threshold or reaching the maximum number of training rounds), resulting in a Raman spectrum feature extraction model.

[0210] Based on the predicted component samples and their physical quantity samples and the component labels and their physical quantity labels, the loss value of the model training is calculated. Common loss functions include mean squared error (MSE) for physical quantity prediction, cross-entropy loss function for component classification prediction (if the component prediction is a classification problem), etc. The loss value reflects the difference between the model's prediction results and the true labels, and the goal of the model is to minimize this loss value.

[0211] Based on the loss value, the gradient of the model parameters (such as the convolution kernel weight of the convolution layer, the weight matrix of the fully connected layer, etc.) is calculated by calculating the loss function, and the gradient descent algorithm (such as stochastic gradient descent SGD, Adagrad, Adadelta, etc.) is used to adjust the model parameters. The process of back propagation is from the output layer to the input layer, and the gradient is propagated step by step, so that the model parameters are updated in the direction of reducing the loss value. In each training iteration, the model calculates the prediction result according to the current parameters, calculates the loss value and the gradient, and then updates the parameters, and repeatedly repeats this process until the loss value converges to a smaller level, or reaches the preset training number of times, at this time the Raman spectrum feature extraction model is obtained.

[0212] Figure 5 The prediction effect comparison chart of the feature extraction method of the application and other related methods is shown; Figure 6 The root mean square error (RMSE) comparison chart of the prediction effect of the feature extraction method of the application and other related methods is shown. In a specific application scenario, the predicted biological component is glucose, and the predicted physicochemical parameter is concentration. See Figure 5 、 Figure 6 In the process of implementing the application, different architecture models (such as PLS, ANN, XGBOOST, RFR, CNN, CNN+Attention: the application) are used for verification, from Figure 5 It can be seen that the ordinate represents the concentration (g / L) of the predicted biological component. Among them, the change curve of the predicted biological component concentration of the method of the application CNN+Attention is most consistent with the HPLC standard value, indicating that the prediction effect of the application is optimal; from Figure 6 It can be seen that the ordinate represents the RMSE. It can be seen from the comparison that the RMSE of the method of the application CNN+Attention is the smallest, indicating that the prediction effect of the application is optimal.

[0213] The brief descriptions of the above different models are as follows:

[0214] PLS: Partial Least Squares, partial least squares method.

[0215] ANN: Artificial Neural Network, artificial neural network.

[0216] XGBOOST: eXtreme Gradient Boosting, extreme gradient boosting.

[0217] RFR: Random Forest Regressor, random forest regressor.

[0218] CNN: Convolutional Neural Network, convolutional neural network.

[0219] CNN+Attention: Convolutional Neural Network+Attention Mechanism.

[0220] Here, it needs to be explained that in the above scheme, the biological component category to be predicted and the category of physical quantity can be determined according to the application scenario, and for this purpose, the model training is targeted for training. The biological components that can be predicted may also include but are not limited to: amino acids, enzymes, etc., and the physical quantities may also include but are not limited to purity, molecular weight, activity, etc.

[0221] The embodiment of the application further provides an electronic device, comprising:

[0222] at least one processor; and

[0223] a memory in communication with the at least one processor; wherein

[0224] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps described in the embodiments of the application.

[0225] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described modules can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0226] In the technical solution of the present disclosure, the acquisition, storage and application of user personal information involved comply with relevant laws and regulations and do not violate public order and good customs.

[0227] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device and a readable storage medium.

[0228] The electronic device of the embodiments of the present application is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are merely examples and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0229] The electronic device includes a computing unit that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM or loaded into a RAM from a storage unit. Various programs and data required for operation of the electronic device can also be stored in the RAM. The computing unit, the ROM, and the RAM are connected to each other through a bus. An I / O interface is also connected to the bus.

[0230] Various components in the electronic device are connected to the I / O interface, including: an input unit such as a keyboard, a mouse, and the like; an output unit such as various types of displays, a speaker, and the like; a storage unit such as a magnetic disk, an optical disk, and the like; and a communication unit such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0231] The computing unit can be various general-purpose and / or special-purpose processing engines having processing and computing capabilities. Some examples of the computing unit include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit performs various methods and processes described above, such as the intervention task generation method. For example, in some embodiments, the intervention task generation method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the computing unit, one or more steps of the intervention task generation method described above can be performed. Alternatively, in other embodiments, the computing unit can be configured to perform the intervention task generation method by any other appropriate means, such as by means of firmware.

[0232] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0233] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.

[0234] In the context of the present disclosure, a readable storage medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The readable storage medium can be a machine-readable signal medium or a machine-readable storage medium. The readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a readable storage medium would include one or more lines of electrical wire, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0235] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0236] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0237] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server is generally established by computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers combined with a blockchain.

[0238] It should be understood that the various forms of flow described above can be reordered, additional or fewer steps added, or steps deleted, using the above-described forms. For example, the steps described in the present disclosure can be executed in parallel, in series, or in a different order, without limitation herein, as long as the desired results of the technical solutions of the present disclosure can be achieved.

[0239] The specific embodiments described above are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present disclosure. Any further modifications, equivalent replacements, improvements, and the like made within the spirit and principles of the present disclosure should be included within the scope of the present disclosure.

Claims

1. A method for Raman spectral feature extraction based on convolution and attention mechanism, characterized in that, The method comprises: obtaining Raman spectrum time series data generated by monitoring a target biological process of fermentation and physicochemical time series data; fusing the Raman spectrum time series data and the physicochemical time series data to generate fermentation feature sequence data; preprocessing and enhancing the fermentation feature sequence data to obtain enhanced feature sequence data; extracting Raman spectrum features from the enhanced feature sequence data based on a Raman spectrum feature extraction model to predict the Raman spectrum features and determine the components and physical quantities generated by the organism in the fermentation process according to the Raman spectrum features; wherein the Raman spectrum feature extraction model comprises a convolution layer, an attention mechanism layer, and a multilayer perceptron; the extraction of the Raman spectrum features from the enhanced feature sequence data based on the Raman spectrum feature extraction model to predict the Raman spectrum features and determine the components and physical quantities generated by the organism in the fermentation process according to the Raman spectrum features comprises: sliding the convolution layer on the enhanced feature sequence data to extract local feature patterns and spatial correlation features, the local feature patterns being associated with the Raman spectrum intensity of the products of the organism in the fermentation process, and the spatial correlation features being associated with the spatial correlation between the physicochemical and the Raman spectrum intensity in the fermentation process of the organism; extracting attention features from the local feature patterns and the spatial correlation features based on the attention mechanism layer to generate a context feature vector, the context feature vector being associated with the Raman spectrum changes and physicochemical change trends between different fermentation states of the organism in the fermentation process; mapping the context feature vector to generate a hyperplane feature vector based on the multilayer perceptron, and compressing the hyperplane feature vector to generate a metabolic pathway related feature vector to predict the Raman spectrum features and determine the components and physical quantities generated by the organism in the fermentation process according to the Raman spectrum features.

2. The method of claim 1, wherein, The method further comprises: obtaining Raman spectrum source data generated by irradiating the organism in fermentation with a Raman light source; serializing the Raman spectrum data according to the timestamps of real-time collection to generate Raman spectrum time series data.

3. The method of claim 1, wherein, The method further comprises: obtaining physicochemical source data generated by detecting the organism in fermentation with a physicochemical sensor; aligning the physicochemical source data and the Raman spectrum source data to generate physicochemical time series data.

4. The method of claim 1, wherein, The fusion of the Raman spectrum time series data and the physicochemical time series data to generate fermentation feature sequence data comprises merging the Raman spectrum time series data and the physicochemical time series data to generate fermentation feature sequence data.

5. The method of claim 1, wherein, The preprocessing and enhancement of the fermentation feature sequence data to obtain enhanced feature sequence data comprises: performing sliding window processing on the fermentation feature sequence data based on a set standardized sliding window, and calculating the mean and standard deviation of the data points in each standardized sliding window based on a set standard deviation multiple; calculating a maximum range value and a minimum range value of the data points in each of the standardized sliding windows based on the mean value and the standard deviation; normalizing the data points in each of the standardized sliding windows based on the maximum range value and the minimum range value, so that the data point distribution of each of the standardized sliding windows reaches a set distribution standard to generate standardized feature sequence data; enhancing the standardized feature sequence data to obtain enhanced feature sequence data.

6. The method of claim 5, wherein, The enhancement of the standardized feature sequence data to obtain enhanced feature sequence data comprises: performing denoising processing on the standardized feature sequence data based on a set filter to generate denoised feature sequence data; performing derivation processing on the denoised feature sequence data to obtain the enhanced feature sequence data.

7. The method of claim 1, wherein, The convolution layer is a sequential model, and the sequential model comprises a first convolution layer, a first batch normalization layer, a second convolution layer, a second batch normalization layer, a third convolution layer, and a third batch normalization layer. The number of convolution kernels in the first convolution layer is greater than the number of convolution kernels in the second convolution layer and the third convolution layer, and the size of a single convolution kernel in the first convolution layer is greater than the size of a single convolution kernel in the second convolution layer and the third convolution layer. The sliding of the convolution layer on the enhanced feature sequence data to extract local feature patterns and spatial correlation features comprises: sliding the first convolution layer on the enhanced feature sequence data to extract first local feature patterns and first spatial correlation features; performing normalization processing on the first local feature patterns and the first spatial correlation features based on the first batch normalization layer to generate first normalized local feature patterns and first normalized spatial correlation features; sliding the second convolution layer on the first normalized local feature patterns and the first normalized spatial correlation features to extract second local feature patterns and second spatial correlation features; performing normalization processing on the second local feature patterns and the second spatial correlation features based on the second batch normalization layer to generate second normalized local feature patterns and second normalized spatial correlation features; sliding the third convolution layer on the second normalized local feature patterns and the second normalized spatial correlation features to extract third local feature patterns and third spatial correlation features; performing normalization processing on the third local feature patterns and the third spatial correlation features based on the third batch normalization layer to generate third normalized local feature patterns and third normalized spatial correlation features, which are respectively taken as the local feature patterns and the spatial correlation features extracted from the enhanced feature sequence data.

8. The method of claim 1, wherein, The attention mechanism layer comprises a first attention mechanism module and a second attention mechanism module. The attention mechanism layer is used for attention feature extraction on the local feature pattern and the spatial correlation feature to generate a context feature vector, which is associated with the Raman spectrum change and the physical and chemical change trend of the biological sample in different fermentation states, including: The first attention mechanism module is used for attention feature extraction on the local feature pattern to generate a first context feature vector; The second attention mechanism module is used for attention feature extraction on the spatial correlation feature to generate a second context feature vector; The first context feature vector and the second context feature vector are fused to generate a fused context feature vector as the context feature vector. 9.A method for Raman spectrum feature extraction modeling based on convolution and attention mechanism, characterized in that, The method comprises the following steps: Obtaining the Raman spectrum time series data sample and the physical and chemical time series data sample generated by monitoring the fermentation sample biological process; Fusing the Raman spectrum time series data sample and the physical and chemical time series data sample to generate a fermentation feature sequence data sample; Obtaining the ingredient label and the physical quantity label corresponding to the fermentation feature sequence data sample; Preprocessing and enhancing the fermentation feature sequence data sample to obtain an enhanced feature sequence data sample; Based on the to-be-trained model, acute forward propagation is performed on the enhanced feature sequence data sample to extract a Raman spectrum feature sample therefrom, so as to predict the ingredient sample and the physical quantity sample generated by the sample biological process in the fermentation process; According to the predicted ingredient sample and the physical quantity sample and the ingredient label and the physical quantity label, the loss value of model training is calculated, and the loss value is used for back propagation to adjust the model parameters of the to-be-trained model until a Raman spectrum feature extraction model is obtained, which is applied to realize the Raman spectrum feature extraction method based on convolution and attention mechanism in any one of claims 1-8.

Citation Information

Patent Citations

  • Method for monitoring microbial community structure in fermentation process

    CN116682494A

  • Attention-based interpretable Raman spectrum identification method, device and equipment

    CN118468141A