Raman spectrum feature extraction method and modeling method based on convolution and attention mechanism
Through the Raman spectral feature extraction method based on convolution and attention mechanism, the problem of real-time monitoring and prediction of traditional methods during the fermentation process is solved, and high accuracy and adaptability prediction of fermentation product components and physical quantity is achieved.
Patent Information
- Application Number
- CN202510076527.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-17
AI Technical Summary
In traditional fermentation, there is difficulty in real-time monitoring and prediction of biologically generated components and physical quantities, and existing models are difficult to accurately describe the complex nonlinear dynamic changes in the fermentation process.
The Raman spectral feature extraction method based on convolution and attention mechanism is adopted, and the Raman spectral time series data and physical and chemical time series data are obtained, and the fermentation feature sequence data is fused to generate, preprocess and enhance, and the Raman spectral features are extracted to predict the components of the fermentation product and its physical quantity.
Real-time monitoring of the fermentation process is achieved, which reduces interference to the fermentation process, improves the accuracy and adaptability of prediction, can better understand the fermentation process, and optimize the fermentation process.
Smart Images

Figure CN120046099A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a Raman spectrum feature extraction method and a modeling method based on convolution and attention mechanisms. Background Art
[0002] In the field of biomanufacturing, accurately predicting the components and their physical quantities generated during the fermentation process is crucial for optimizing the fermentation process, improving product quality, and production efficiency. With the development of biotechnology and the increasing demand for industrial production, the precise monitoring and prediction of the fermentation process have become a research hotspot.
[0003] Traditionally, predicting the components and their physical quantities generated by organisms during the fermentation process mainly relies on chemical analysis methods, sensor technologies, and models based on experience and statistics.
[0004] Chemical analysis methods are the means earlier used for detecting fermentation components, such as high-performance liquid chromatography (HPLC), gas chromatography (GC), etc. These methods can accurately determine the types and contents of fermentation products by separating and detecting the fermentation broth. For example, in the fermentation production of antibiotics, HPLC can be used to quantitatively analyze the content of antibiotics in the fermentation broth. However, these methods usually require complex pretreatment of samples, with cumbersome operations and long time consumption, and cannot achieve real-time online monitoring. Moreover, frequent sampling may interfere with the normal progress of the fermentation process and affect the accuracy of experimental results.
[0005] Sensor technology is also a commonly used monitoring means. For example, pH sensors, dissolved oxygen sensors, etc. can monitor the physicochemical parameters during the fermentation process in real time. These sensors can quickly feedback data and provide a certain basis for the control of the fermentation process. However, the types of sensors are limited, and they can only measure specific parameters and cannot comprehensively reflect the complex component changes during the fermentation process. In addition, the accuracy and stability of sensors are easily affected by the fermentation environment. For example, changes in factors such as temperature and acidity may cause an increase in the measurement error of sensors, affecting the accuracy of prediction.
[0006] Models based on experience and statistics are established according to a large amount of historical fermentation data. By analyzing the historical data, the key factors affecting the fermentation components and physical quantities are found, and corresponding mathematical models are established. For example, linear regression models, multivariate statistical analysis models, etc. can predict certain indicators during the fermentation process. These models can, to a certain extent, use the laws of historical data for prediction, but they often rely on simple linear assumptions or statistical relationships and are difficult to accurately describe the complex non-linear dynamic changes during the fermentation process. Moreover, when the fermentation conditions change, the adaptability of these models is poor, and a large amount of data collection and model training need to be carried out again. Summary of the Invention
[0007] The present disclosure provides a Raman spectroscopy feature extraction method and a modeling method based on convolution and attention mechanisms.
[0008] According to the first aspect of the present disclosure, there is provided a Raman spectroscopy feature extraction method based on convolution and attention mechanisms, which includes:
[0009] Obtaining Raman spectroscopy time series data and physicochemical time series data generated by monitoring a target biological process of fermentation;
[0010] Fusing the Raman spectroscopy time series data and the physicochemical time series data to generate fermentation feature sequence data;
[0011] Preprocessing and enhancing the fermentation feature sequence data to obtain enhanced feature sequence data;
[0012] Based on the Raman spectroscopy feature extraction model, extracting Raman spectroscopy features from the enhanced feature sequence data to predict the Raman spectroscopy features and thereby determining the components and their physical quantities produced by the organism during fermentation.
[0013] According to the second aspect of the present disclosure, there is provided a Raman spectroscopy feature modeling method based on convolution and attention mechanisms, which includes:
[0014] Obtaining a Raman spectroscopy time series data sample and a physicochemical time series data sample generated by monitoring a sample biological process of fermentation;
[0015] Fusing the Raman spectroscopy time series data sample and the physicochemical time series data sample to generate a fermentation feature sequence data sample;
[0016] Obtaining the component label and its physical quantity label corresponding to the fermentation feature sequence data sample;
[0017] Preprocessing and enhancing the fermentation feature sequence data sample to obtain an enhanced feature sequence data sample;
[0018] Based on the model to be trained, performing forward propagation on the enhanced feature sequence data sample to extract Raman spectroscopy feature samples therefrom, so as to predict the component samples and their physical quantity samples produced by the sample organism during fermentation;
[0019] According to the predicted component samples and their physical quantity samples and the component labels and their physical quantity labels, calculating the loss value of model training, and performing backpropagation based on the loss value to adjust the model parameters of the model to be trained until a Raman spectroscopy feature extraction model is obtained.
[0020] The technical effects of the solution of the present disclosure are as follows:
[0021] Traditional chemical analysis methods (such as HPLC, GC) are cumbersome and time-consuming, and cannot achieve real-time on-line monitoring. However, this method can monitor the fermentation process in real time by obtaining Raman spectroscopy time series data and physicochemical time series data. Raman spectroscopy technology can quickly detect the fermentation system without destroying the sample. Combining with the physicochemical data collected in real time, it can timely capture the changes in components and physical quantities during the fermentation process, providing the possibility for real-time adjustment of the fermentation process and avoiding the problem of out-of-control fermentation process caused by delayed detection.
[0022] Chemical analysis methods require frequent sampling of samples, which may interfere with the normal progress of the fermentation process and affect the accuracy of experimental results. This method is based on the real-time monitoring of Raman spectroscopy and physicochemical data, without the need for a large number of samples of the fermentation system, reducing the interference with the fermentation process and ensuring the naturalness of the fermentation process and the reliability of experimental results.
[0023] Traditional sensor technologies can only measure specific parameters and cannot comprehensively reflect the complex component changes during the fermentation process. This method fuses Raman spectroscopy time series data with physicochemical time series data. Raman spectroscopy can provide rich molecular structure information to reflect the types and changes of biological components; physicochemical data can describe the state of the fermentation environment. The fermentation characteristic sequence data generated by the fusion of the two contains more comprehensive information, which helps to understand the fermentation process more accurately, thereby improving the accuracy of predicting the components and physical quantities of fermentation products.
[0024] Models based on experience and statistics often rely on simple linear assumptions or statistical relationships, making it difficult to accurately describe the complex non-linear dynamic changes during the fermentation process and having poor adaptability to changes in fermentation conditions. However, the Raman spectroscopy feature extraction model based on convolution and attention mechanisms can automatically learn the complex features and patterns in the data. The convolutional layer can effectively extract the local features in the data and capture the subtle changes in Raman spectroscopy and physicochemical data; the attention mechanism can dynamically focus on the key features related to the components and physical quantities of fermentation products and ignore irrelevant information. This model structure can better adapt to the non-linear dynamic changes during the fermentation process. Even if the fermentation conditions change, it can accurately predict by learning new data features, improving the adaptability and generalization ability of the model.
[0025] Traditional methods have certain limitations in data processing and model establishment. Chemical analysis methods require complex sample pretreatment processes, consuming a large amount of manpower and time; models based on experience and statistics need to re - collect a large amount of data and retrain the model when fermentation conditions change. This method can make full use of existing data information through steps such as data fusion, pretreatment, and enhancement, reducing unnecessary data collection and processing costs. At the same time, the model based on convolutional and attention mechanisms has strong learning ability, can quickly extract effective features from data, and improves the prediction efficiency.
[0026] This method uses a deep - learning model to realize the automatic extraction of Raman spectral features and the prediction of fermentation product components and their physical quantities, reducing the interference of human factors and improving the objectivity and accuracy of prediction. The model can continuously learn and optimize. With the accumulation of data and the training of the model, the prediction performance will continuously improve, providing strong support for the intelligent control and optimization of the fermentation process.
[0027] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings
[0028] In combination with the accompanying drawings and with reference to the following detailed description, the above - mentioned and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. In the drawings, the same or similar reference numerals represent the same or similar elements, where:
[0029] Figure 1 shows a block diagram of a Raman spectral source data acquisition system according to an embodiment of the present disclosure;
[0030] Figure 2 shows a block diagram of the Raman spectral feature extraction model according to an embodiment of the present disclosure;
[0031] Figure 3 shows a flowchart of a Raman spectral feature extraction method based on convolutional and attention mechanisms according to an embodiment of the present disclosure;
[0032] Figure 4(a) shows a waveform diagram of two batches of Raman spectral source data collected in an application scenario;
[0033] Figure 4(b) shows a waveform diagram of pre - processing the Raman spectral source data in Figure 4(a) in an application scenario;
[0034] Figure 5 shows a comparison diagram of the prediction effects of the feature extraction method of the present invention and other related methods;
[0035] Figure 6 The figure shows the comparison chart of the root mean square error (RMSE) of the prediction effects between the feature extraction method of the present invention and other related methods. Detailed implementation manners
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0037] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0038] Figure 1 The figure shows the structural block diagram of a Raman spectroscopy source data acquisition system according to an embodiment of the present disclosure; as Figure 1 shown, the fermenter contains fermented organisms, and physicochemical sensors are fixed in the fermenter to detect the fermented organisms to generate physicochemical source data. The Raman spectroscopy source data irradiates the fermented organisms through a Raman light source, and a Raman detector performs Raman light detection on the irradiated fermented organisms to generate Raman spectroscopy source data.
[0039] Figure 2 The figure shows the structural block diagram of the Raman spectroscopy feature extraction model according to an embodiment of the present disclosure; the Raman spectroscopy feature extraction model includes: a convolutional layer, an attention mechanism layer, and a multi-layer perceptron. For the descriptions of these structures, please refer to the description of the preferred embodiment.
[0040] Figure 3 The figure shows the schematic flowchart of a Raman spectroscopy feature extraction method based on convolution and attention mechanism according to an embodiment of the present disclosure. As Figure 3 shown, it includes:
[0041] Obtain Raman spectroscopy time series data and physicochemical time series data generated by monitoring the target biological process of fermentation;
[0042] Fuse the Raman spectroscopy time series data and the physicochemical time series data to generate fermentation feature sequence data;
[0043] Preprocess and enhance the fermentation feature sequence data to obtain enhanced feature sequence data;
[0044] Based on the Raman spectroscopy feature extraction model, extract Raman spectroscopy features from the enhanced feature sequence data to predict the Raman spectroscopy features and thereby determine the components and their physical quantities produced by the organism during fermentation.
[0045] Figure 4(a) shows the waveform diagrams of two batches of Raman spectroscopy source data collected in an application scenario; Figure 4(b) shows the waveform diagram of preprocessing the Raman spectroscopy source data in Figure 4(a) in an application scenario. From the waveforms, the fluorescence background influence is removed from the preprocessed fermentation feature sequence data, and the Raman feature information is more prominent.
[0046] Optionally, the method further includes:
[0047] Obtain Raman spectroscopy source data, which is generated by irradiating the fermenting organism with a Raman light source;
[0048] Serialize the Raman spectroscopy data according to the time stamps of real-time acquisition to generate Raman spectroscopy time series data.
[0049] Optionally, the method further includes:
[0050] Obtain physicochemical source data, which is generated by detecting the fermenting organism with physicochemical sensors;
[0051] Align the physicochemical source data with the Raman spectroscopy source data to generate physicochemical time series data.
[0052] Optionally, the fusion of the Raman spectroscopy time series data and the physicochemical time series data to generate fermentation feature sequence data includes: merging the Raman spectroscopy time series data and the physicochemical time series data to generate fermentation feature sequence data.
[0053] Optionally, the preprocessing and enhancement of the fermentation feature sequence data to obtain enhanced feature sequence data includes:
[0054] Perform a sliding window process on the fermentation feature sequence data based on a set standardized sliding window, and calculate the mean and standard deviation of the data points within each standardized sliding window based on a set multiple of the standard deviation;
[0055] Based on the mean and the standard deviation, calculate the maximum range value and the minimum range value of the data points within each standardized sliding window;
[0056] Based on the maximum range value and the minimum range value, standardize the data points within each standardized sliding window so that the data point distribution within each standardized sliding window reaches a set distribution standard to generate standardized feature sequence data;
[0057] Enhance the standardized feature sequence data to obtain enhanced feature sequence data.
[0058] Optionally, the size of the standardized sliding window is 3, and the standard deviation multiple is 1.5. For this reason, when calculating the mean and standard deviation of the data points within each standardized sliding window based on the set standard deviation multiple, calculate the mean and standard deviation of each data point and its adjacent 3 data points. The standard deviation multiple of 1.5 means that when calculating the mean and standard deviation of the data points within each standardized sliding window, it is necessary to consider making the calculated maximum range value and minimum range value fall between 1.5 and 0.
[0059] Specifically, when standardizing the data points within each standardized sliding window, use the formula for the data points to calculate the standardized fermentation feature data. x is the original fermentation feature data, μ is the mean, σ is the standard deviation, and x' is the standardized fermentation feature data. If the standardized fermentation feature data exceeds the maximum range value, adjust it to the maximum range value. If it is lower than the minimum range value, adjust it to the minimum range value.
[0060] Optionally, the enhancement of the standardized feature sequence data to obtain enhanced feature sequence data includes:
[0061] Based on a set filter, perform denoising processing on the standardized feature sequence data to generate denoised feature sequence data;
[0062] Perform derivative processing on the denoised feature sequence data to obtain the enhanced feature sequence data.
[0063] Optionally, the filter is a Savitzky-Golay filter, and the window length of the Savitzky-Golay filter is 13. When performing denoising processing on the standardized feature sequence data based on the set filter to generate denoised feature sequence data, when performing denoising based on each Savitzky-Golay filter, perform least squares fitting of polynomials based on 13 data points around each data point obtained by the sliding window for data point interpolation to generate denoised feature sequence data.
[0064] Optionally, the polynomial order of the least squares fitting polynomial is 2.
[0065] Specifically, the window length is 13 and the polynomial order is 2, which means that within the neighborhood of each data point, 13 consecutive data points will be used to fit a second-order polynomial. The core idea is to fit a polynomial using the least squares method within the neighborhood of each data point, and then use the value of this polynomial at that data point to replace the original value, thereby achieving smoothing. For a given window size m (m = 13) and polynomial order n (n = 2) of the filter, the Savitzky-Golay filter can be expressed as:
[0066]
[0067] where a k are the polynomial coefficients determined by least squares fitting, x(t) is the normalized eigenvalue of the central data point in the window, and x(t - i) are the normalized eigenvalues of other points in the window. In this way, each point can be expressed as a weighted linear combination of other points in the window to achieve data point interpolation for smoothing effect and generate denoised feature sequence data.
[0068] Optionally, when performing the derivative processing on the denoised feature sequence data, the derivative order is 2 to perform second-order derivative processing on the denoised feature sequence data and extract the change rate of the denoised feature sequence data accordingly to obtain the enhanced feature sequence data.
[0069] Specifically, the second-order derivative processing can be expressed as:
[0070]
[0071] In discrete data, the second derivative can be approximately calculated by the following formula:
[0072]
[0073] where h is the step size, x i-1 is the denoised eigenvalue corresponding to the (i - 1)-th data point, f(x i-1 ) is the intensity corresponding to the denoised eigenvalue of the (i - 1)-th data point; f(x i ) is the intensity corresponding to the denoised eigenvalue of the i-th data point; f(x i+1 ) is the intensity corresponding to the denoised eigenvalue of the (i + 1)-th data point, x i is the denoised eigenvalue corresponding to the i-th data point, x i+1 is the denoised eigenvalue corresponding to the (i + 1)-th data point, f”(x i ) is the second derivative, that is, the change rate of the denoised feature sequence data is extracted to identify the extreme points therein to generate the enhanced feature sequence data. The enhanced feature sequence data usually corresponds to the frequency of molecular vibration to provide information about the molecular structure.
[0074] Optionally, the Raman spectrum feature extraction model includes: a convolutional layer, an attention mechanism layer, and a multi-layer perceptron;
[0075] Based on the Raman spectrum feature extraction model, Raman spectrum features are extracted from the enhanced feature sequence data to predict Raman spectrum features and thereby determine the components and their physical quantities produced by the organism during the fermentation process, including:
[0076] Based on the convolutional layer, sliding is performed on the enhanced feature sequence data to extract local feature patterns and spatial correlation features. The local feature patterns are associated with the Raman spectrum intensities corresponding to the products of the organism during the fermentation process, and the spatial correlation features are associated with the spatial correlation between the physical chemistry and Raman spectrum intensities of the organism during the fermentation process;
[0077] Based on the attention mechanism layer, attention feature extraction is performed on the local feature patterns and spatial correlation features to generate context feature vectors. The context feature vectors are associated with the Raman spectrum changes and physical chemistry change trends between different fermentation states of the organism during the fermentation process;
[0078] Based on the multi-layer perceptron, feature mapping is performed on the context feature vectors to generate hyperplane feature vectors, and the hyperplane feature vectors are compressed to generate metabolic pathway-related feature vectors to predict Raman spectrum features and thereby determine the components and their physical quantities produced by the organism during the fermentation process.
[0079] Optionally, the convolutional layer is a sequential model, and the sequential model includes a first convolutional layer, a first batch normalization layer, a second convolutional layer, a second batch normalization layer, a third convolutional layer, and a third batch normalization layer. The number of convolutional kernels in the first convolutional layer is respectively more than the number of convolutional kernels in the second convolutional layer and the third convolutional layer, and the size of a single convolutional kernel in the first convolutional layer is respectively larger than the size of a single convolutional kernel in the second convolutional layer and the third convolutional layer;
[0080] The sliding of the convolutional layer on the enhanced feature sequence data to extract local feature patterns and spatial correlation features includes:
[0081] Based on the first convolutional layer, sliding is performed on the enhanced feature sequence data to extract the first local feature pattern and the first spatial correlation feature;
[0082] Based on the first batch normalization layer, normalization processing is performed on the first local feature pattern and the first spatial correlation feature to generate the first normalized local feature pattern and the first normalized spatial correlation feature;
[0083] The second convolutional layer slides on the first normalized local feature pattern and the first normalized spatial correlation feature to extract a second local feature pattern and a second spatial correlation feature;
[0084] Based on the second batch normalization layer, the second local feature pattern and the second spatial correlation feature are normalized to generate a second normalized local feature pattern and a second normalized spatial correlation feature;
[0085] The third convolutional layer slides on the second normalized local feature pattern and the second normalized spatial correlation feature to extract a third local feature pattern and a third spatial correlation feature;
[0086] Based on the third batch normalization layer, the third local feature pattern and the third spatial correlation feature are normalized to generate a third normalized local feature pattern and a third normalized spatial correlation feature, which are respectively used as the local feature pattern and the spatial correlation feature extracted from the enhanced feature sequence data.
[0087] In the Raman spectrum feature extraction method based on convolution and attention mechanism, the convolutional layer is an important part. The enhanced feature sequence data is processed through a sequential model structure to extract the local feature pattern and the spatial correlation feature therein. The sequential model is alternately composed of multiple convolutional layers and batch normalization layers. Each convolutional layer is responsible for feature extraction, and the batch normalization layer normalizes the extracted features to accelerate model training and improve generalization ability.
[0088] Specifically, the specific technical implementation details of sliding the convolutional layer on the enhanced feature sequence data to extract the local feature pattern and the spatial correlation feature are as follows:
[0089] The first convolutional layer: Set n 1 convolution kernels, and the size of each convolution kernel is k 1 ×k 1 . In the scenario of bioreactor monitoring, a larger number of convolution kernels and a larger convolution kernel size can capture richer global features. For example, in Raman spectrum data, it can detect feature information related to multiple substances in a larger wavenumber range, as well as the overall change trend of sensor data in a longer time interval. Here, the specific value of n 1 can be adjusted according to the complexity of the data and the performance requirements of the model, and generally takes values between 16 and 64.
[0090] Enhanced feature sequence data X, with dimensions H×W×C, where H represents the height of the data (which can be understood as the length of the time series in Raman spectroscopy data), W represents the width of the data (such as the wavenumber range of the spectrum), and C represents the number of channels of the data (including Raman spectroscopy data and physical and chemical data).
[0091] For each convolutional kernel K i (i = 1, 2,..., n 1 ), with size k 1 ×k 1 ×C, perform a sliding convolutional operation on the input data X. The specific calculation process is as follows:
[0092]
[0093] where Y i is the feature map after convolution by the i-th convolutional kernel, (h, w) is the position on the feature map, and b i is the bias term corresponding to the i-th convolutional kernel. In this way, the first convolutional layer slides on the enhanced feature sequence data, extracts the first local feature pattern and the first spatial correlation feature, and the output feature map has dimensions H 1 ×W 1 ×n 1 , where s is the convolution stride, usually set to 1.
[0094] The feature map Y output by the first convolutional layer (including the first local feature pattern and the first spatial correlation feature) has dimensions H 1 ×W 1 ×n 1 .
[0095] Perform batch normalization on each channel c. For each mini-batch data B, containing N samples, the calculation formula is as follows:
[0096]
[0097] where μ B,c is the mean of the mini-batch data on channel c, is the variance of the mini-batch data on channel c, is the normalized eigenvalue, ò is a very small constant (usually set to 10 -5 ), used to prevent the denominator from being zero, γ c and β c are learnable parameters, used to scale and shift the normalized eigenvalue respectively, and Y i,:,:,c is the finally normalized eigenvalue. After being processed by the first batch normalization layer, the first normalized local feature pattern and the first normalized spatial correlation feature are obtained, and the dimensions remain H1 ×W 1 ×n 1 。
[0098] Second convolutional layer: Set n 2 convolution kernels, each with a size of k 2 ×k 2 , and n 2 <n 1 , k 2 <k 1 . On the basis of the first convolutional layer, the second convolutional layer further focuses on more local features and mines deeper detailed information in the data. For example, during the fermentation process, the minute changes in the molecular structure within microbial cells are manifested as features within a smaller wavenumber range in the spectrum, and smaller convolutional kernels can better capture these microscopic features.
[0099] The first normalized local feature pattern and the first normalized spatial correlation feature output by the first batch normalization layer, with dimensions of H 1 ×W 1 ×n 1 。
[0100] Similar to the first convolutional layer, each convolution kernel K of the second convolutional layer j (j = 1, 2,, n 2 ), with a size of k 2 ×k 2 ×n 1 , performs a sliding convolution operation on the input data. The calculation formula is:
[0101]
[0102] where Z j is the feature map after convolution by the j-th convolution kernel, (h′, w′) is the position on the feature map, and b j is the bias term corresponding to the j-th convolution kernel. The output feature map has dimensions of H 2 ×W 2 ×n 2 , where
[0103] The feature map Z output by the second convolutional layer (including the second local feature pattern and the second spatial correlation feature) has dimensions of H 2 ×W 2 ×n 2 。
[0104] Similar to the first batch normalization layer, batch normalization is performed on each channel c. For a small batch of data B containing N samples, the calculation formula is as follows:
[0105]
[0106] After being processed by the second batch normalization layer, the second normalized local feature pattern and the second normalized spatial correlation feature are obtained, and the dimension is still H 2 ×W 2 ×n 2 。
[0107] Third convolutional layer: Set n 3 convolution kernels, each with a size of k 3 ×k 3 and n 3 ≤n 2 ,k 3 ≤k 2 。The third convolutional layer continues to refine and abstract the features of the previous layer to extract the most critical feature information.
[0108] The second normalized local feature pattern and the second normalized spatial correlation feature output by the second batch normalization layer, with a dimension of H 2 ×W 2 ×n 2 。
[0109] Each convolution kernel K of the third convolutional layer l (l = 1, 2,, n 3 ), with a size of k 3 ×k 3 ×n 2 performs a sliding convolution operation on the input data. The calculation formula is:
[0110]
[0111] where U l is the feature map after convolution by the l-th convolution kernel (including the third local feature pattern and the third spatial correlation feature), (h″, w″) is the position on the feature map, and b l is the bias term corresponding to the l-th convolution kernel. The output feature map dimension is H 3 ×W 3 ×n 3 ,where
[0112] The feature map U output by the third convolutional layer, with a dimension of H 3 ×W 3 ×n 3 。
[0113] Batch normalization is also performed on each channel c. For a small batch of data B containing N samples, the calculation formula is as follows:
[0114]
[0115]
[0116] After being processed by the third batch normalization layer, the third normalized local feature pattern and the third normalized spatial correlation feature are obtained, and the dimension is still H 3 ×W 3 ×n 3 , as the final local feature pattern and spatial correlation feature extracted from the enhanced feature sequence data.
[0117] Therefore, in contrast, in the application of traditional convolutional neural networks, the number and size of convolutional kernels are often fixed and difficult to adapt to the feature extraction requirements of different levels and scales. In this technology, the first convolutional layer sets n 1 convolutional kernels of size k 1 ×k 1 , the second convolutional layer sets n 2 convolutional kernels of size k 2 ×k 2 (n 2 <n 1 , k 2 <k 1 ), and the third convolutional layer sets n 3 convolutional kernels of size k 3 ×k 3 (n 3 ≤n 2 , k 3 ≤k 2 ). From the perspective of convolution operation, for example, the relatively large convolutional kernel size k 1 and the relatively large number of convolutional kernels n 1 enable the first convolutional layer to capture data features in a relatively large spatial range. For example, in Raman spectroscopy data, it can cover a wider wavenumber range and a longer time series, detecting macroscopic feature information related to multiple substances and the overall change trend of sensor data over a long time interval.
[0118] As the convolutional layer deepens, the second and third convolutional layers focus on more local and finer features by reducing the convolutional kernel size and number. Taking the of the second convolutional layer as an example, the smaller k 2 enables the convolutional kernel to capture microscopic features in the data more carefully. For example, the minute changes in the molecular structure within microbial cells are manifested as features in a relatively small wavenumber range in the spectrum. This multi-scale feature extraction method can comprehensively mine information at different levels in the biological fermentation process, from the overall trend to the microscopic details, which is more in line with the complexity and diversity characteristics of the data in the biological fermentation process compared with the method of setting traditional fixed convolutional kernels.
[0119] Traditional methods cannot effectively distinguish the importance of different levels of features, resulting in inaccurate feature extraction. In this technology, as the convolutional layer progresses, the reduction in the number of convolutional kernels prompts the model to screen and integrate the extracted features, achieving gradual refinement and abstraction of the features. In biological fermentation monitoring, this means that the model can gradually focus on the key features that have an important impact on the fermentation process from a large number of original Raman spectroscopy and sensor data features. For example, after the first convolutional layer extracts rich original features, the second convolutional layer screens and recombines these features through a smaller number of convolutional kernels to extract more representative features, and the third convolutional layer further refines them, so that the finally extracted features can more accurately reflect the essential features of the biological fermentation process, providing more valuable information for subsequent analysis and prediction.
[0120] When traditional models face different biological fermentation datasets, due to the differences in data distribution, they often show poor generalization ability. The batch normalization layer normalizes the features, making the model more adaptable to different data distributions. Performing the normalization operation on each small batch of data can unify the data distribution into a relatively stable range, reducing data noise and interference. For example, in different batches of biological fermentation experiments, Raman spectroscopy data may have certain fluctuations due to slight differences in experimental conditions. The batch normalization layer can effectively handle these fluctuations, enabling the model to perform well on data from different batches and improving the model's generalization ability. At the same time, the learnable parameters γ c and β c can adjust the normalized features according to the characteristics of the data, further enhancing the model's adaptability to different data.
[0121] Optionally, the attention mechanism layer includes a first attention mechanism module and a second attention mechanism module;
[0122] Performing attention feature extraction on the local feature pattern and the spatial correlation feature based on the attention mechanism layer to generate a context feature vector, the context feature vector is associated with the Raman spectroscopy changes and the physicochemical change trends between different fermentation states during the fermentation process of the organism, including:
[0123] Performing attention feature extraction on the local feature pattern based on the first attention mechanism module to generate a first context feature vector;
[0124] Performing attention feature extraction on the spatial correlation feature based on the second attention mechanism module to generate a second context feature vector;
[0125] Fuse the first context feature vector and the second context feature vector to generate a fused context feature vector as the context feature vector.
[0126] Optionally, the attention feature extraction of the local feature pattern based on the first attention mechanism module to generate a first context feature vector includes:
[0127] Multiply the local feature pattern by a set first linearization weight matrix to obtain a first query vector;
[0128] Multiply the query vector by a set first attention weight matrix to obtain a first attention score;
[0129] Normalize the first attention score to obtain a first attention weight;
[0130] Multiply and sum the local feature pattern and the first attention weight and flatten them to generate a first context vector.
[0131] Optionally, the attention feature extraction of the spatial correlation feature based on the second attention mechanism module to generate a second context feature vector includes:
[0132] Multiply the spatial correlation feature by a set second linearization weight matrix to obtain a second query vector;
[0133] Multiply the query vector by a set second attention weight matrix to obtain a second attention score;
[0134] Normalize the second attention score to obtain a second attention weight;
[0135] Multiply and sum the local feature pattern and the second attention weight and flatten them to generate a second context vector.
[0136] In the scenario of Raman spectral feature extraction for bioreactor monitoring, the design of the attention mechanism layer aims to fully explore local feature patterns and spatial correlation features to generate context feature vectors that can accurately reflect the trends of Raman spectra and physicochemical changes during the bioreactor process. This layer consists of a first attention mechanism module and a second attention mechanism module, which process different types of features respectively, and finally fuse to obtain the final context feature vector.
[0137] Specifically, for the first attention mechanism module, the local feature pattern is denoted as X l, is a tensor of dimension I×J×K. In the context of biological fermentation monitoring, I can represent the length of Raman spectral data in a time series, reflecting the change of the fermentation process over time; J represents the spectral dimension, such as the wavenumber range, which covers the Raman characteristic information of different substances; K represents the number of channels related to local feature patterns, and these channels may contain the Raman spectral intensity information of different fermentation products or other attributes related to local features.
[0138] The first linearization weight matrix is denoted as W l1 , with dimension K×D 1 . Here, D 1 is a hyperparameter that determines the dimension of the first query vector. The role of this matrix is to map the input local feature pattern X l to a new feature space for subsequent calculation of attention scores.
[0139] The first attention weight matrix is denoted as A l1 , with dimension D 1 ×D 1 . It is used to calculate the attention scores between different positions in the local feature pattern, reflecting the correlation and importance degree between each position.
[0140] Calculate the first query vector: where i = 1,..., I, j = 1,..., J, d = 1,..., D 1 . By multiplying each element of the local feature pattern X l with the corresponding element of the first linearization weight matrix W l1 and summing them up, the first query vector Q l is obtained, with dimension I×J×D 1 . This step transforms the original local feature pattern into a form more suitable for calculating attention scores. Through the linear transformation of the weight matrix, the important features related to subsequent calculations are highlighted.
[0141] Calculate the first attention score: where i = 1,..., I, j = 1,..., J, d 2 = 1,..., D 1 . By performing matrix multiplication on the first query vector Q l and the first attention weight matrix A l1 , the first attention score S l is obtained, with dimension I×J×D 1 . This score represents the importance degree of each position in the local feature pattern relative to other positions and is the basis for subsequent calculation of attention weights.
[0142] Calculate the first attention weight: where i = 1,..., I, j = 1,..., J, d2 = 1,,D 1 。
[0143] For the first attention score S l Perform Softmax normalization. The Softmax function exponentiates the attention scores at each position and divides them by the sum of the exponents of all position attention scores, so that the attention weights at each position are between 0 and 1, and the sum of the attention weights of all positions is 1. The first attention weight P l , with dimensions I×J×D 1 is obtained, which accurately reflects the probability distribution of the relative importance of each position in the local feature pattern.
[0144] Generate the first context vector: First, multiply the local feature pattern X l element-wise with the first attention weight P l , and then perform a summation operation in the three dimensions of I, J, and D 1 to obtain an intermediate result. Finally, flatten this intermediate result into a one-dimensional vector through a Flatten operation to obtain the first context vector C l . This step combines the local feature pattern with the attention weight, highlights the important feature information, and converts it into a vector form convenient for subsequent processing.
[0145] For the second attention mechanism module, the spatial correlation feature is denoted as X s , which is a tensor with dimensions I×J×L. In bioreactor monitoring, the meanings of I and J are the same as those in the local feature pattern, and L represents the number of channels related to the spatial correlation feature, which contain the spatial correlation information of the physical chemistry and Raman spectral intensity during the fermentation process, such as the relationship between the distribution of different substances in space and the Raman spectral characteristics.
[0146] The second linearization weight matrix is denoted as W l2 , with dimensions L×D 2 . Among them, D 2 is a hyperparameter that determines the dimension of the second query vector. This matrix is used to map the spatial correlation feature X s to a new feature space for subsequent calculation of attention scores.
[0147] The second attention weight matrix is denoted as A l2 , with dimensions D 2 ×D 2 . It is used to calculate the attention scores between different positions in the spatial correlation feature.
[0148] Calculate the second query vector: where \(i = 1,\cdots,I\), \(j = 1,\cdots,J\), \(d = 1,\cdots,D\) 2 Similar to the method of calculating the query vector in the first attention mechanism module, by multiplying each element of the spatial correlation feature \(X\) s with the corresponding element of the second linearized weight matrix \(W\) l2 and summing them up, the second query vector \(Q\) s is obtained, whose dimension is \(I\times J\times D\) 2 .
[0149] Calculate the second attention score: where \(i = 1,\cdots,I\), \(j = 1,\cdots,J\), \(d\) 2 = 1,\cdots,D 2 . Multiply the second query vector \(Q\) s with the second attention weight matrix \(A\) l2 to perform a matrix multiplication operation, and the second attention score \(S\) s is obtained, with a dimension of \(I\times J\times D\) 2 .
[0150] Calculate the second attention weight: where \(i = 1,\cdots,I\), \(j = 1,\cdots,J\), \(d\) 2 = 1,\cdots,D 2 .
[0151] Perform Softmax normalization on the second attention score \(S\) s to obtain the second attention weight \(P\) s , with a dimension of \(I\times J\times D\) 2 , which reflects the relative importance probability distribution of each position in the spatial correlation feature.
[0152] Generate the second context vector:
[0153] Multiply the spatial correlation feature \(X\) s element-wise with the second attention weight \(P\) s , then sum over the three dimensions of \(I\), \(J\), and \(D\) 2 , and finally flatten it into a one-dimensional vector through a Flatten operation to obtain the second context vector \(C\) s .
[0154] For context vector fusion, the first context vector \(C\) l and the second context vector \(C\) s , their lengths are \(I\times J\) and \(I\times J\) respectively. An innovative fusion method is adopted, which combines weighted summation, element-wise product, and non-linear transformation.
[0155] First, calculate the weighted sum vector \(C\) w : \(C\) w= α·C l +(1 - α)·C s where α is a hyperparameter between 0 and 1, used to control the relative importance of the first and second context vectors in the fusion process.
[0156] Then, calculate the element-wise product vector C p : where | represents the element-wise product operation, i.e., multiplying the elements at corresponding positions.
[0157] Next, perform a non-linear transformation on C w and C p : C nf1 = tanh(C w ), C nf2 = sigmoid(C p ). Here, the tanh function and the sigmoid function are used to introduce non-linearity and enhance the expressive power of the model.
[0158] Finally, calculate the fused context feature vector C: C = Flatten(C nf1 + C nf2 ). Add the two vectors after non-linear transformation, and then flatten the result into a one-dimensional vector through the Flatten operation to obtain the final fused context feature vector C.
[0159] To this end, different from the traditional attention mechanism that mixes all features, this technology separately processes local feature patterns and spatial correlation features, and deeply mines them through two independent attention mechanism modules. This way can more specifically capture the key information of different types of features. In biological fermentation monitoring, it can more accurately extract local features related to fermentation products and spatial correlation features between physical chemistry and Raman spectral intensity, improving the accuracy and efficiency of feature extraction.
[0160] The traditional context vector fusion methods are usually relatively simple, such as direct concatenation or simple weighted summation. While this technology adopts an innovative fusion strategy, combining weighted summation, element-wise product, and non-linear transformation. This complex fusion method can make full use of the information of the two context vectors, not only considering their weight distribution, but also mining the potential relationship between the two vectors through element-wise product, and then enhancing the expressive power of the model through non-linear transformation, so that the fused context feature vector can more comprehensively and accurately reflect the Raman spectral changes and physical chemistry change trends between different fermentation states in the biological fermentation process.
[0161] In addition, in each attention mechanism module, by setting different linearization weight matrices and attention weight matrices, and setting an adjustable hyperparameter α during the fusion process, the model can be flexibly adjusted according to the characteristics of specific bioprocess data. This flexibility enables the model to better adapt to different experimental conditions and data characteristics. Compared with traditional fixed-parameter models, it has stronger generalization ability and adaptability, and can achieve more accurate analysis and prediction in various bioprocess monitoring scenarios.
[0162] Optionally, the multi-layer perceptron includes: a first fully connected layer, a first non-linear activation layer, a second fully connected layer, a second non-linear activation layer, and an output layer, where the first fully connected layer is connected to the second fully connected layer;
[0163] Based on the multi-layer perceptron, feature mapping is performed on the context feature vector to generate a hyperplane feature vector, and the hyperplane feature vector is compressed to generate a metabolic pathway-related feature vector to predict Raman spectral features and thereby determine the components and their physical quantities produced by the organism during fermentation, including:
[0164] Based on the first fully connected layer, feature mapping is performed on the context feature vector to generate a hyperplane feature vector, and the hyperplane feature vector is non-linearly activated through the first non-linear activation layer;
[0165] Based on the second fully connected layer, the non-linearly activated hyperplane feature vector is compressed to generate a metabolic pathway-related feature vector, and the metabolic pathway-related feature vector is non-linearly activated through the second non-linear activation layer;
[0166] Based on the output layer, linear weighted summation and bias adjustment are performed on the non-linearly activated metabolic pathway-related feature vector to predict Raman spectral features and thereby determine the components and their physical quantities produced by the organism during fermentation.
[0167] In the technical system of Raman spectral feature extraction based on convolution and attention mechanism for bioprocess monitoring, the multi-layer perceptron (MLP) plays a core role in gradually transforming the context feature vector into the predicted value of Raman spectral features, and then realizing the prediction of the components and their physical quantities of bioprocess products. It consists of multiple functional layers, and through a series of feature mapping, compression, and activation operations, it deeply processes the input information and provides key support for accurately predicting the bioprocess.
[0168] Specifically, the technical step details of the multi-layer perceptron (MLP) are as follows:
[0169] For the first fully connected layer, the context feature vector is denoted as X, which is a one-dimensional vector of length N. In the context of bioreactor monitoring, this vector is the result of being processed by the previous attention mechanism layer, containing rich information such as Raman spectral changes and physicochemical change trends under different fermentation states during the bioreactor process. This information is the basis for the model to understand the fermentation process and make subsequent predictions.
[0170] The weight matrix of the first fully connected layer is denoted as W 1 , with dimensions N×M 1 . Where M 1 is a hyperparameter representing the number of features output by the first fully connected layer. This matrix determines the linear transformation method of the context feature vector, and by setting the weights, adjusts the importance and influence of different features in subsequent calculations.
[0171] The bias vector of the first fully connected layer is denoted as b 1 , with dimensions M 1 . It introduces an offset to the result of the linear transformation, increasing the flexibility of the model, enabling the model to better fit complex data distributions and capture subtle features in the data.
[0172] During feature mapping: Z 1 = X·W 1 + b 1 , perform a matrix multiplication operation on the context feature vector X and the weight matrix W of the first fully connected layer 1 , and then add the bias vector b 1 , to obtain a vector Z 1 with dimensions M 1 . This step realizes the linear feature mapping of the context feature vector, transforming the input feature vector from the original space to a new feature space. In the field of bioreactor monitoring, this process can be regarded as a preliminary sorting and integration of the relationships between different features, highlighting the potential information related to Raman spectral features during the bioreactor process, and preparing for subsequent non-linear activation operations.
[0173] For the first non-linear activation layer, the output Z 1 from the previous layer, with dimensions M 1 . The activation function uses the Swish function, which is defined as f(x) = x·σ(x), where is the Sigmoid function.
[0174] During non-linear activation, use to apply the Swish activation function to each element of the input vector Z 1 , to obtain the vector A 1 after non-linear activation, with dimensions still M 1Compared with traditional simple activation functions (such as ReLU), the Swish function has characteristics such as being continuously differentiable and non-monotonic, and can more accurately capture the complex non-linear relationships between data in the biological fermentation process. The biological fermentation process is a highly complex non-linear system involving numerous interrelated physical, chemical, and biochemical reactions. The Swish function can better simulate the complex interactions between features in these reaction processes, significantly enhancing the model's expressive ability, enabling the model to deeply learn more complex Raman spectral feature patterns and product relationships in the biological fermentation process.
[0175] For the second fully connected layer, the input vector is the output A of the first non-linear activation layer 1 , with dimension M 1 .
[0176] The weight matrix of the second fully connected layer is denoted as W 2 , with dimension M 1 ×M 2 . Where M 2 is a hyperparameter, and M 2 <M 1 , which means that the number of features output by the second fully connected layer is less than the number of input features. This setting enables the second fully connected layer to compress the input features and extract more crucial information.
[0177] Bias vector of the second fully connected layer: denoted as b 2 , with dimension M 2 .
[0178] When performing feature compression and mapping, use Z 2 =A 1 ·W 2 +b 2 , multiply the vector A 1 processed by the first non-linear activation layer with the weight matrix W 2 of the second fully connected layer, and then add the bias vector b 2 to obtain a vector Z 2 with dimension M 2 . This step not only performs a further linear transformation on the input features but also compresses the features by reducing the number of output features. In biological fermentation monitoring, this operation helps to remove redundant information and focus on the key Raman spectral features more directly related to the biological metabolic pathway, laying a foundation for generating more targeted metabolic pathway-related feature vectors in the subsequent stage.
[0179] For the second non-linear activation layer, the input vector is: the output Z of the previous layer 2 , with dimension M 2 .
[0180] The activation function uses the ELU (Exponential Linear Unit) function, which is defined as where α is an adjustable parameter, usually set to 1.
[0181] When performing non-linear activation, use to apply the ELU activation function to each element of the input vector Z 2 to obtain the vector A after non-linear activation, with the dimension still being M 2 . The ELU function cleverly combines the linear characteristics of the ReLU function in the positive part and the smooth transition characteristics in the negative part, and this characteristic is of great significance in the training of complex models for biological fermentation monitoring. 2 .
[0182] For the output layer, the input vector is the output A of the second non-linear activation layer 2 , with the dimension being M 2 .
[0183] The weight vector of the output layer is denoted as W 3 , with the dimension being M 2 ×1. Its role is to map the metabolic pathway-related feature vector obtained after being processed through multiple previous layers to the final Raman spectrum feature prediction space.
[0184] Bias scalar of the output layer: denoted as b 3 , which is a scalar value.
[0185] When performing linear weighted summation and bias adjustment, based on the formula Y = A 2 ·W 3 +b 3 , to perform matrix multiplication on the metabolic pathway-related feature vector A 2 after being processed by the second non-linear activation layer and the weight vector W of the output layer 3 , and then add the bias scalar b 3 to obtain a scalar value Y. This Y value is the predicted value of the model for the Raman spectrum features in the biological fermentation process.
[0186] After obtaining the predicted value Y of the Raman spectrum features, it is necessary to match it with the relationship spectrum of the pre-specified components and physical and chemical quantities. This relationship spectrum is constructed based on a large amount of experimental data, theoretical research, and long-term biological fermentation practical experience. It details the corresponding relationships between different Raman spectrum feature values and the components and their physical and chemical quantities generated during the biological fermentation process.
[0187] By matching the Y value with the relationship spectrum, the components and physical quantities actually produced by the organism during the fermentation process can be further determined based on the known corresponding relationship. For example, the relationship spectrum may clearly record the concentration range, yield and other physical quantity information of a certain component in the fermentation product within a specific Raman spectral feature prediction value range. This matching process provides a conversion bridge from the abstract Raman spectral feature prediction value to the specific biological fermentation product information, making it possible to closely link the model prediction results with the actual biological fermentation situation, so as to more accurately understand the state of the biological fermentation process and the product situation.
[0188] To this end, different from the simple activation functions (such as ReLU) commonly used in traditional multi-layer perceptrons, this technology uses the Swish function in the first nonlinear activation layer and the ELU function in the second nonlinear activation layer. These complex activation functions, with their excellent nonlinear expression ability and unique gradient characteristics, can more accurately capture the complex nonlinear relationships in the biological fermentation process, greatly improving the prediction accuracy of Raman spectral features, thereby more accurately reflecting the changes in components and physical quantities in the biological fermentation process.
[0189] In addition, by setting the number of output features of the second fully connected layer to be smaller than that of the first fully connected layer, effective compression of features is achieved. This hierarchical processing method can gradually focus on key Raman spectral features closely related to biological metabolic pathways, remove redundant information, and make the model more focused on predicting the composition and physical quantities of biological fermentation products. At the same time, the nonlinear activation operation of each layer further enhances the model's ability to predict complex Raman spectral feature patterns. Compared with traditional single-layer or simple structure models, it can more deeply explore the potential information in the data and improve the prediction accuracy of the model.
[0190] Furthermore, each layer of the multilayer perceptron is equipped with adjustable parameters (such as weight matrix and bias vector), and adjustable parameters are also set in the activation function (such as α in the ELU function). This flexible parameter setting gives the model strong adaptability, and technicians can make fine adjustments based on the characteristics of different bio-fermentation data. By optimizing these parameters, the model can better adapt to various bio-fermentation scenarios and monitoring targets, improve the generalization ability of the model, and ensure accurate predictions under different experimental conditions and production environments.
[0191] The embodiment of the present application also provides a Raman spectral feature modeling method based on convolution and attention mechanism, which includes:
[0192] Obtaining Raman spectral time series data samples and physical and chemical time series data samples generated by monitoring the fermentation sample biological process;
[0193] Fuse the Raman spectroscopy time series data samples and the physicochemical time series data samples to generate fermentation feature sequence data samples;
[0194] Obtain the component labels and their physical quantity labels corresponding to the fermentation feature sequence data samples;
[0195] Preprocess and enhance the fermentation feature sequence data samples to obtain enhanced feature sequence data samples;
[0196] Based on the model to be trained, perform forward propagation on the enhanced feature sequence data samples to extract Raman spectroscopy feature samples therefrom, so as to predict the component samples and their physical quantity samples generated by the sample organism during the fermentation process;
[0197] According to the predicted component samples and their physical quantity samples and the component labels and their physical quantity labels, calculate the loss value of model training, and perform backpropagation based on the loss value to adjust the model parameters of the model to be trained until a Raman spectroscopy feature extraction model is obtained.
[0198] Specifically, during the above model training, a professional Raman spectrometer is used to monitor the fermentation process of the sample organism in real time, and Raman spectroscopy data at different time points are obtained to form a time series. Raman spectroscopy can reflect the structural and compositional information of biomolecules. Different biological components will have specific peaks and characteristics in the Raman spectrum, and these data provide important information about the component changes during the biological fermentation process for the model.
[0199] In addition, various sensors are used to synchronously collect physicochemical parameters related to the fermentation process, such as temperature, pH value, dissolved oxygen concentration, pressure, etc., which also form a time series. These physicochemical parameters have an important impact on the biological fermentation process. They are closely related to the generation and transformation of biological components, and combined with Raman spectroscopy data, they can more comprehensively describe the fermentation process.
[0200] Regarding the component labels and their physical quantity labels corresponding to the fermentation feature sequence data samples. These labels are determined through experimental measurements, chemical analysis, etc. They represent the components actually produced by the sample organism during the fermentation process and the physical quantities of these components (such as concentration, yield, etc.). The component labels are used to indicate which specific biological components are contained in the fermentation products, and the physical quantity labels give the quantitative information of these components. The accuracy of the labels is crucial for model training. They are the targets for the model to learn. The model tries to make the prediction results as close as possible to the labels by continuously adjusting the parameters.
[0201] Based on the model to be trained, forward propagation is performed on the enhanced feature sequence data samples to extract Raman spectroscopy feature samples, and the component samples and their physical quantity samples produced by the sample organisms during the fermentation process are predicted. The model to be trained is usually a neural network model based on convolutional and attention mechanisms. The structure of the model consists of a convolutional layer, an attention mechanism layer, and a multi-layer perceptron. The functions of each layer are as follows:
[0202] Convolutional layer: The convolutional layer performs sliding convolutional operations on the input fermentation feature sequence data through a convolutional kernel. The convolutional kernel is a small weight matrix. When it slides on the data, it performs weighted summation on the data in the local area to extract local features. For example, for two-dimensional Raman spectroscopy data, the convolutional kernel can capture local patterns in the spectrum, such as specific peak combinations; for time series data, the convolutional kernel can extract local trends and changes in the time series.
[0203] The convolutional layer can automatically learn local features in the data, reduce the dimension of the data, and at the same time retain important information. By stacking multiple convolutional layers, features of different levels and complexities can be extracted, from simple local patterns to more complex feature combinations.
[0204] Attention mechanism layer: It aims to calculate the attention weights of the input features to dynamically adjust the importance of different features. In biological fermentation monitoring, not all features are equally important for predicting components and physical quantities. The attention mechanism enables the model to focus on features more relevant to the target tasks (predicting components and physical quantities) by calculating the weights of each feature. Specifically, through a calculation process, it maps the input features to an attention weight vector, and each element of this vector represents the importance degree of the corresponding feature.
[0205] The attention mechanism can help the model better process complex data, improve the accuracy and pertinence of feature extraction. It can ignore some irrelevant or interfering features, enhance the influence of features closely related to the components and physical quantities of fermentation products, thereby improving the prediction performance of the model.
[0206] Multi-layer perceptron: It consists of multiple fully connected layers, and further performs feature mapping and classification / regression on the feature vectors processed by the convolutional layer and the attention mechanism layer. In this model, the multi-layer perceptron receives the output features from the attention mechanism layer, and through a series of linear transformations and non-linear activation functions, maps them to the final prediction space.
[0207] The multi-layer perceptron can learn complex non-linear relationships between features. By continuously adjusting the weights, an accurate mapping is established between the input features and the prediction results. When predicting component samples, the multi-layer perceptron can output the probability distribution of each component; when predicting physical quantity samples, the multi-layer perceptron can output specific numerical predictions.
[0208] Compare the component samples and their physical quantity samples predicted by the model with the corresponding component labels and their physical quantity labels, and calculate the loss value. For component prediction, if it is a classification problem, the cross-entropy loss function can be used; for physical quantity prediction, the mean squared error (MSE) loss function is usually used. The loss value reflects the degree of difference between the model's prediction result and the true label, and the training objective of the model is to minimize this loss value.
[0209] Backpropagation: Based on the calculated loss value, perform backpropagation to adjust the model's parameters. Backpropagation is an optimization algorithm based on gradient descent. It determines the update direction of the parameters by calculating the gradients of the loss function with respect to the model's parameters (such as the convolutional kernel weights of the convolutional layer, the parameters of the attention mechanism layer, the weights and biases of the multi-layer perceptron). Specifically, starting from the output layer, according to the gradient of the loss function with respect to the output, gradually calculate the gradients of each layer in reverse, and then update the parameters of each layer according to the gradients. By continuously iterating this process, the model's parameters are gradually adjusted, making the loss value continuously decrease until the preset convergence condition (such as the loss value is less than a certain threshold or the maximum number of training epochs is reached), and the Raman spectrum feature extraction model is obtained.
[0210] Calculate the loss value of the model training based on the predicted component samples and their physical quantity samples and the component labels and their physical quantity labels. Commonly used loss functions include mean squared error (MSE) for physical quantity prediction, cross-entropy loss function for component classification prediction (if component prediction is a classification problem), etc. The loss value reflects the difference between the model's prediction result and the true label, and the goal of the model is to minimize this loss value.
[0211] Perform backpropagation based on the loss value. By calculating the gradients of the loss function with respect to the model's parameters (such as the convolutional kernel weights of the convolutional layer, the weight matrix of the fully connected layer, etc.), use gradient descent algorithms (such as stochastic gradient descent SGD, Adagrad, Adadelta, etc.) to adjust the model's parameters. The process of backpropagation starts from the output layer and gradually propagates the gradients towards the input layer, so that the model's parameters are updated in the direction of reducing the loss value. In each training iteration, the model calculates the prediction result according to the current parameters, calculates the loss value and gradients, and then updates the parameters. This process is continuously repeated until the loss value converges to a small level or reaches the preset number of training epochs. At this time, the Raman spectrum feature extraction model is obtained.
[0212] Figure 5 The figure shows a comparison chart of the prediction effects of the feature extraction method of the present invention and other related methods; Figure 6 The figure shows a comparison chart of the root mean square error (RMSE) of the prediction effects of the feature extraction method of the present invention and other related methods. In a specific application scenario, taking glucose as the predicted biological component and concentration as the predicted physicochemical parameter for example, see Figure 5 , Figure 6 , during the implementation of the present invention, models with different architectures (such as PLS, ANN, XGBOOST, RFR, CNN, CNN+Attention: the present invention) were used for verification. As can be seen from Figure 5 , the vertical axis represents the concentration (g / L) of the predicted biological component. Among them, for the method of CNN+Attention of the present invention, the change curve of the concentration of the predicted biological component is most consistent with the HPLC standard value, indicating that the prediction effect of the present invention is the best; as can be seen from Figure 6 , the vertical axis represents RMSE. By comparison, for the method of CNN+Attention of the present invention, its RMSE is the smallest, indicating that the prediction effect of the present invention is the best.
[0213] A brief description of the above different models is as follows:
[0214] PLS: Partial Least Squares, partial least squares method.
[0215] ANN: Artificial Neural Network, artificial neural network.
[0216] XGBOOST: eXtreme Gradient Boosting, extreme gradient boosting.
[0217] RFR: Random Forest Regressor, random forest regressor.
[0218] CNN: Convolutional Neural Network, convolutional neural network.
[0219] CNN+Attention: Convolutional Neural Network+Attention Mechanism.
[0220] Here, it should be noted that in the above solution, according to the application scenario, the category of the biological component to be predicted and the category of the physical quantity can be determined. Therefore, targeted training for model training can be carried out. The predictable biological components can also include, but are not limited to, amino acids, enzymes, etc., and the physical quantities can also include, but are not limited to, purity, molecular weight, activity, etc.
[0221] An embodiment of the present application further provides an electronic device, which includes:
[0222] At least one processor; and
[0223] A memory communicatively connected to the at least one processor; wherein,
[0224] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps described in the embodiments of the present application.
[0225] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0226] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0227] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0228] The electronic device in the embodiments of the present application is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0229] The electronic device includes a computing unit, which can execute various appropriate actions and processes according to the computer program stored in the ROM or the computer program loaded from the storage unit into the RAM. In the RAM, various programs and data required for the operation of the electronic device can also be stored. The computing unit, ROM, and RAM are connected to each other through a bus. The I / O interface is also connected to the bus.
[0230] Multiple components in the electronic device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a magnetic disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0231] The computing unit can be various general-purpose and / or special-purpose processing engines with processing and computing capabilities. Some examples of the computing unit include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit executes the various methods and processes described above, such as the intervention task generation method. For example, in some embodiments, the intervention task generation method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the computing unit, one or more steps of the intervention task generation method described above can be executed. Alternatively, in other embodiments, the computing unit can be configured to execute the intervention task generation method by any other suitable means (e.g., by means of firmware).
[0232] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0233] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0234] In the context of the present disclosure, a readable storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The readable storage medium may be a machine-readable signal medium or a machine-readable storage medium. The readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0235] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0236] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0237] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0238] It should be understood that the various forms of processes described above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present disclosure can be achieved, and no limitation is imposed herein.
[0239] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A Raman spectroscopy feature extraction method based on convolution and attention mechanism, characterized in that: include: Acquire Raman spectral time series data and physical and chemical time series data generated by monitoring the target biological process of fermentation; fusing the Raman spectrum time series data with the physical and chemical time series data to generate fermentation characteristic sequence data; Preprocessing and enhancing the fermentation characteristic sequence data to obtain enhanced characteristic sequence data; Based on the Raman spectral feature extraction model, Raman spectral features are extracted from the enhanced feature sequence data to predict Raman spectral features and thereby determine components and physical quantities produced by the organism during the fermentation process.
2. The method according to claim 1, characterized in that The method further comprises: Acquiring Raman spectral source data, wherein the Raman spectral source data is generated by irradiating a fermented organism with a Raman light source; The Raman spectrum data is serialized according to the timestamp of real-time acquisition to generate Raman spectrum time series data.
3. The method according to claim 1, characterized in that: The method further comprises: Acquiring physicochemical source data, wherein the physicochemical source data is generated by detecting fermented organisms through physicochemical sensors; The physicochemical source data is aligned with the Raman spectroscopy source data to generate physicochemical time series data.
4. The method according to claim 1, characterized in that: The fusing the Raman spectral time series data with the physical and chemical time series data to generate fermentation characteristic sequence data includes: merging the Raman spectral time series data with the physical and chemical time series data to generate fermentation characteristic sequence data.
5. The method according to claim 1, characterized in that The preprocessing and enhancing of the fermentation characteristic sequence data to obtain enhanced characteristic sequence data comprises: Performing sliding window processing on the fermentation characteristic sequence data based on a set standardized sliding window, and calculating the mean and standard deviation of the data points in each standardized sliding window based on a set standard deviation multiple; Based on the mean and standard deviation, calculate the maximum range value and the minimum range value of the data points in each standardized sliding window; Based on the maximum range value and the minimum range value, the data points in each standardized sliding window are standardized so that the distribution of the data points in each standardized sliding window reaches a set distribution standard to generate standardized feature sequence data; The standardized feature sequence data is enhanced to obtain enhanced feature sequence data.
6. The method according to claim 5, characterized in that The step of enhancing the standardized feature sequence data to obtain enhanced feature sequence data comprises: Based on the set filter, the standardized feature sequence data is subjected to denoising processing to generate denoised feature sequence data; The denoised feature sequence data is subjected to a derivative process to obtain the enhanced feature sequence data.
7. The method according to claim 1, characterized in that The Raman spectrum feature extraction model includes: a convolution layer, an attention mechanism layer, and a multi-layer perceptron; The extracting Raman spectral features from the enhanced feature sequence data based on the Raman spectral feature extraction model to predict the Raman spectral features and thereby determine the components and physical quantities produced by the organism during the fermentation process, comprises: Based on the convolution layer sliding on the enhanced feature sequence data to extract local feature patterns and spatial correlation features, the local feature patterns are associated with the Raman spectrum intensity corresponding to the product of the organism during the fermentation process, and the spatial correlation features are associated with the spatial correlation between the physical chemistry and the Raman spectrum intensity of the organism during the fermentation process; Performing attention feature extraction on the local feature pattern and the spatial correlation feature based on the attention mechanism layer to generate a context feature vector, wherein the upper and lower feature vectors are associated with Raman spectral changes and physicochemical change trends of the organism between different fermentation states during the fermentation process; Based on the multi-layer perceptron, the context feature vector is feature mapped to generate a hyperplane feature vector, and the hyperplane feature vector is compressed to generate a metabolic pathway related feature vector to predict Raman spectral features and thereby determine the components and physical quantities produced by the organism during the fermentation process.
8. The method according to claim 7, characterized in that The convolution layer is a sequential model, which includes a first convolution layer, a first batch normalization layer, a second convolution layer, a second batch normalization layer, a third convolution layer, and a third batch normalization layer. The number of convolution kernels in the first convolution layer is respectively greater than the number of convolution kernels in the second convolution layer and the third convolution layer. The size of a single convolution kernel in the first convolution layer is respectively greater than the size of a single convolution kernel in the second convolution layer and the third convolution layer. The step of sliding the convolutional layer on the enhanced feature sequence data to extract local feature patterns and spatial correlation features includes: Sliding extracting a first local feature pattern and a first spatial correlation feature on the enhanced feature sequence data based on the first convolutional layer; Normalizing the first local feature pattern and the first spatial correlation feature based on the first batch normalization layer to generate a first normalized local feature pattern and a first normalized spatial correlation feature; Sliding on the first normalized local feature pattern and the first normalized spatial correlation feature based on the second convolutional layer to extract a second local feature pattern and a second spatial correlation feature; Normalizing the second local feature pattern and the second spatial correlation feature based on the second batch normalization layer to generate a second normalized local feature pattern and a second normalized spatial correlation feature; Sliding on the second normalized local feature pattern and the second normalized spatial correlation feature based on the third convolutional layer to extract a third local feature pattern and a third spatial correlation feature; The third local feature pattern and the third spatial correlation feature are normalized based on the third batch normalization layer to generate a third normalized local feature pattern and a third normalized spatial correlation feature, which are used as the local feature pattern and the spatial correlation feature extracted from the enhanced feature sequence data respectively.
9. The method according to claim 7, characterized in that: The attention mechanism layer includes a first attention mechanism module and a second attention mechanism module; The attention feature extraction is performed on the local feature pattern and the spatial correlation feature based on the attention mechanism layer to generate a context feature vector, wherein the upper and lower feature vectors are associated with the Raman spectrum changes and the physicochemical change trends of the organism between different fermentation states during the fermentation process, including: Performing attention feature extraction on the local feature pattern based on the first attention mechanism module to generate a first context feature vector; Performing attention feature extraction on the spatial correlation feature based on the second attention mechanism module to generate a second context feature vector; The first context feature vector and the second context feature vector are fused to generate a fused context feature vector as the context feature vector.
10. A Raman spectral feature modeling method based on convolution and attention mechanism, characterized in that: include: Obtaining Raman spectral time series data samples and physical and chemical time series data samples generated by monitoring the fermentation sample biological process; Fusion of the Raman spectrum time series data sample and the physical chemistry time series data sample to generate a fermentation characteristic sequence data sample; Obtaining component labels and physical quantity labels corresponding to the fermentation characteristic sequence data sample; Preprocessing and enhancing the fermentation characteristic sequence data sample to obtain an enhanced characteristic sequence data sample; Based on the model to be trained, the enhanced feature sequence data sample is rapidly forward propagated to extract Raman spectrum feature samples therefrom, so as to predict the component samples and physical quantity samples produced by the sample organism in the fermentation process; According to the predicted component samples and their physical quantity samples and component labels and their physical quantity labels, the loss value of the model training is calculated, and back propagation is performed based on the loss value to adjust the model parameters of the model to be trained until a Raman spectroscopy feature extraction model is obtained.
Citation Information
Patent Citations
Method for monitoring microbial community structure in fermentation process
CN116682494A
Raman spectrum qualitative analysis method, system and equipment for mixture
CN117935963A
Attention-based interpretable Raman spectrum identification method, device and equipment
CN118468141A
Automated control and prediction for a fermentation system
US20220290090A1