Method for analyzing quality of microbial fermentation product based on Raman spectrum and neural network algorithm
By combining Raman spectroscopy with neural network algorithms, the problems of complexity and insufficient accuracy in the quality detection of microbial fermentation products have been solved, achieving rapid, low-cost, and accurate quality detection that is suitable for industrial applications of various fermentation products.
Patent Information
- Application Number
- CN202511020462.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-21
AI Technical Summary
Existing methods for detecting the quality of microbial fermentation products suffer from problems such as complex detection procedures, high costs, difficult operation, long detection cycles, and insufficient detection accuracy and stability, making it difficult to meet the needs of modern production for real-time, online, and automated detection.
By combining Raman spectroscopy with neural network algorithms, and leveraging the rapid non-destructive testing characteristics and automated analysis process, along with spectral preprocessing and multi-scale adaptive feature extraction techniques, signal quality is improved. Furthermore, a deep learning architecture is used to automatically extract deep features from spectral data, achieving efficient and accurate quality detection.
It has achieved improved detection efficiency, shortened the single detection time to 10 seconds, reduced detection costs by 80%, improved detection accuracy to RMSE≤0.25g/L, and achieved a classification accuracy of ≥92.7%, adapting to the detection of various fermentation products and meeting the needs of industrial environments.
Smart Images

Figure CN120992576A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and specifically to a method for analyzing the quality of microbial fermentation products based on Raman spectroscopy and neural network algorithms. Background Technology
[0002] Microbial fermentation products are widely used in food, feed, pharmaceuticals, and the fermentation industry. Existing microbial fermentation products mainly originate from various microbial strains such as yeast, lactic acid bacteria, Bacillus, Aspergillus, and Streptomyces. Their extracts or metabolites are rich in amino acids, nucleotides, peptides, vitamins, and various functional active substances. The quality of these fermentation products directly affects the application efficacy and safety of the final products. Therefore, effectively controlling and evaluating the quality of fermentation products has become a crucial link in the microbial fermentation industry chain. In particular, different microbial strains and their fermentation process conditions (such as the degree of hydrolysis, fermentation time, and culture medium composition) have a significant impact on the composition and content of the final product. Traditional experience-based quality control methods can no longer meet the demands of modern production for precise and automated quality testing. Therefore, developing an efficient, accurate, and non-destructive quality testing and analysis method is of great significance for ensuring product consistency and safety and improving production efficiency.
[0003] Currently, the quality detection of microbial fermentation products largely relies on traditional analytical instruments such as high-performance liquid chromatography (HPLC), amino acid analyzers (AAA), and gas chromatography (GC). While these methods offer high accuracy, they suffer from drawbacks including complex procedures, long detection cycles, high instrument costs, and demanding operator skills, making them particularly unsuitable for real-time, online detection on production lines. Furthermore, these traditional detection techniques typically require complex sample pretreatment, increasing operational difficulty and human error, hindering large-scale, high-throughput quality monitoring of microbial fermentation products. Therefore, existing detection methods can no longer meet the development needs of the modern fermentation industry for rapid, non-destructive, and automated detection.
[0004] Raman spectroscopy, a spectroscopic analysis technique based on molecular vibrations, offers advantages such as speed, non-destructive nature, and the elimination of complex sample pretreatment, leading to its widespread application in food testing, pharmaceutical analysis, and biomedical detection. Particularly in the field of microbial fermentation product quality detection, Raman spectroscopy is an ideal detection tool due to its ability to obtain molecular structure information without damaging the sample. However, traditional Raman spectroscopy methods often suffer from insufficient sensitivity, poor stability, and inaccurate results when applied to the detection of microbial fermentation products, limited by factors such as the complex matrix of fermentation products, low Raman signal intensity, severe fluorescence background interference, signal overlap, and significant batch-to-batch variability. Therefore, improving Raman spectral signal quality and reducing data interference are key factors affecting its large-scale application.
[0005] Traditional data analysis methods typically employ machine learning techniques such as Partial Least Squares Regression (PLS), Principal Component Analysis (PCA), and Support Vector Machines (SVM) for feature extraction and classification prediction of Raman spectroscopy data. However, these methods have significant limitations when processing Raman spectroscopy data in complex matrices: firstly, traditional algorithms have limited ability to express nonlinear features, making it difficult to effectively reveal deep-seated features in complex spectral data; secondly, the models have poor generalization ability and are easily affected by batch differences in samples and changes in experimental conditions, resulting in insufficient model stability and applicability. Therefore, existing methods face significant bottlenecks in the accuracy and reliability of microbial fermentation product quality detection, failing to meet the demands of the modern fermentation industry for efficient and precise detection. Summary of the Invention
[0006] To address the aforementioned technical problems, this application provides a method for detecting and analyzing the quality of microbial fermentation products based on Raman spectroscopy and neural network algorithms, thereby at least solving or mitigating the problems existing in the prior art.
[0007] To achieve the above objectives, according to one aspect of this application, a method for detecting the quality of microbial fermentation products based on Raman spectroscopy and neural networks is provided, comprising:
[0008] Step 1: Collect Raman spectral data of microbial fermentation products;
[0009] Step 2: Vectorize the Raman spectral data to obtain the Raman spectral feature vectors;
[0010] Step 3: Input the Raman spectral augmentation vector into the fermentation product quality prediction model for forward propagation to predict the quality of microbial fermentation products.
[0011] This technical solution achieves a comprehensive improvement in detection efficiency, accuracy, and industrial adaptability through the deep integration of Raman spectroscopy and neural network algorithms. In terms of efficiency, it integrates rapid detection and automated analysis to meet the real-time monitoring needs of production lines. Regarding accuracy, the predicted RMSE for key components is ≤0.25g / L, and the classification accuracy is ≥92.7%. In terms of adaptability, through low-temperature noise suppression and a multi-strain universal model design, it can adapt to industrial environments ranging from -20℃ to 60℃, supporting the detection of various fermentation products such as yeast and lactic acid bacteria, providing the microbial fermentation industry with an efficient, accurate, and non-destructive intelligent quality control solution. Attached Figure Description
[0012] Figure 1 This is a schematic flowchart of a method for detecting the quality of microbial fermentation products based on Raman spectroscopy and neural networks, according to an embodiment of this application.
[0013] Figure 2 This is a plot of raw, unprocessed Raman spectra.
[0014] Figure 3 The data is denoised and standardized.
[0015] Figure 4 This is a schematic diagram of the architecture of the fermentation product quality prediction model of this application.
[0016] Figure 5 Confusion matrix diagram of the classification results of yeast extract compliance.
[0017] Figure 6 The graph shows the predicted amino acid content of yeast extract.
[0018] Figure 7 This section presents a performance comparison of different models in predicting the content of hydrolyzed and free amino acids. Detailed Implementation
[0019] Figure 1 This is a schematic flowchart illustrating a method for detecting the quality of microbial fermentation products based on Raman spectroscopy and a neural network algorithm, according to an embodiment of this application. Figure 1 As shown, it includes:
[0020] Step 1: Collect Raman spectral data of microbial fermentation products;
[0021] Step 2: Vectorize the Raman spectral data to obtain the Raman spectral feature vectors;
[0022] Step 3: Input the Raman spectral feature vector into the fermentation product quality prediction model for forward propagation to predict the quality of microbial fermentation products.
[0023] This technical solution effectively addresses the problems of complex and costly detection processes associated with traditional HPLC and GC methods by leveraging the rapid and non-destructive detection characteristics and automated analysis workflow of Raman spectroscopy. Raman spectroscopy eliminates the need for sample centrifugation and extraction pretreatment steps, with a single detection time of ≤10 seconds, simplifying the operation process by over 70% compared to traditional methods and significantly reducing human error. Simultaneously, combined with automated feature extraction and prediction using neural network models, it can be integrated into production lines for 24-hour online monitoring, reducing the cost of single-sample detection by 80% and shortening the detection cycle from 4 hours to 2 minutes, fully meeting the demands of modern production for automation and high-throughput quality control. To address issues such as weak Raman signals and fluorescence interference caused by the complex matrix of fermentation products, this solution employs a 1064nm Raman analyzer to collect Raman signals from microbial fermentation products to weaken background fluorescence interference, and improves signal quality through spectral preprocessing and multi-scale adaptive feature extraction technology. On the one hand, vectorization processing such as baseline correction and noise suppression improves the Raman fluorescence signal-to-noise ratio by 3-5 times, effectively removing more than 90% of the fluorescence background. On the other hand, dynamic wavenumber-sensing convolutional layers enhance chemical bond vibration features (such as amino acid characteristic peaks), and adaptive weight generation suppresses matrix interference, resulting in a characteristic peak retention rate of over 95%, significantly improving signal resolution capabilities under complex matrices and overcoming the bottleneck of insufficient sensitivity in traditional Raman analysis. Compared to traditional algorithms such as PLS and PCA, the neural network algorithm used in this scheme significantly improves model performance through nonlinear feature learning and adaptive generalization mechanisms. The deep learning architecture can automatically extract deep features such as high-order correlations of chemical bond vibrations and metabolic pathway semantics from spectral data, improving the ability to express nonlinear features by more than 40%. At the same time, through historical spectral statistical information and multimodal fusion strategies, the feature sensitivity of different batches of samples is dynamically adjusted, ensuring that the prediction accuracy fluctuation of the model in batch-difference scenarios is ≤5%, which is more than 20% higher than traditional methods, effectively solving the detection stability problem caused by insufficient generalization ability.
[0024] See Figure 2 This is a Raman spectral data image before vectorization. Figure 2 In the middle, the horizontal axis represents the wavenumber, with the unit being cm. -1The ordinate reflects the positions of different vibrational frequencies in the Raman spectrum, which can be used to identify characteristic vibrational peaks of different chemical bonds or functional groups in microbial fermentation products. The ordinate represents normalized intensity (au), indicating the intensity of the Raman scattered light. Normalization facilitates comparison between different spectra, and its value reflects the relative intensity of the characteristic peak at the corresponding wavenumber. Different colored curves represent Raman spectral data from different samples. For example, different labels such as FP102, FM502, and FM760 correspond to different microbial fermentation product samples. By comparing these curves, the differences in Raman signals of different samples at various wavenumbers can be analyzed, thereby studying their compositional differences.
[0025] Optionally, step 2, vectorizing the Raman spectral data to obtain Raman spectral feature vectors, includes the following steps:
[0026] Step 21: Perform baseline correction and band clipping on the Raman spectral data to obtain effective spectral data free from fluorescence interference;
[0027] Step 22: Perform median filtering and normalization on the effective spectral data to remove fluorescence interference, so as to obtain denoised and standardized spectral data;
[0028] Step 23: Vectorize the denoised and standardized spectral data to obtain Raman spectral feature vectors.
[0029] Optionally, step 21, performing baseline correction and band clipping on the Raman spectral data to obtain effective spectral data free from fluorescence interference, includes the following steps:
[0030] Step 211: Based on the distribution characteristics of Raman spectral characteristic peaks, pre-screen the characteristic intervals of the original Raman spectral data to generate the initial region of interest spectrum;
[0031] Step 212: Based on the multi-scale dynamic weight iterative fitting model, perform dynamic baseline estimation and subtraction on the initial region of interest spectrum to generate a baseline-de-drift spectrum;
[0032] Step 213: Based on the database of key component characteristic peaks of fermentation products, the baseline-drift spectrum is finely trimmed for characteristic bands to obtain effective spectral data free of fluorescence interference.
[0033] Preferably, in step 211, the raw Raman spectral data (500–3500 cm⁻¹) are used. -1 The spectral curvature (including fluorescence background, noise, and characteristic peaks) was calculated using the second derivative method across the entire wavelength range, and a curvature threshold (e.g., >0.01 AU / cm) was set. -1 Identify the positions of potential feature peaks and obtain a candidate peak list {P1, P2, ..., P}.n Furthermore, for each candidate peak P... i Calculate its full width at half maximum (FWHM) and compare it with the characteristic peak FWHM range of known components in the standard Raman database (e.g., the protein amide I band is 15–25 cm⁻¹). -1 Compare and remove abnormal peaks (such as FWHM > 50 cm⁻¹). -1 (Noise peaks). Finally, the selected characteristic peaks are sorted by wavenumber and merged if the spacing is less than 50cm. -1 Adjacent peaks (e.g., phenylalanine at 1003 cm⁻¹) -1 With 1031cm -1 (peak), and expands 20cm to both sides. -1 Form an initial region of interest (step 211) to generate the initial spectrum of step 211 (e.g., covering 600–1800 cm⁻¹). -1 2800–3100cm -1 The region contains characteristic peaks of key components such as proteins, amino acids, and carbohydrates, while removing most of the background regions that lack information.
[0034] Preferably, in step 213, the baseline-de-drift spectrum (including the complete step 211 but still containing redundant regions) is compared with the characteristic peak database of fermentation products (e.g., yeast fermentation containing phenylalanine 1003 cm⁻¹). -1 Leucine 1390cm -1 Compare 20 key peaks and calculate a similarity score (e.g., cosine similarity > 0.9). For successfully matched feature peaks, a similarity score of ±20cm is calculated centered on the peak position. -1 Determine the precise band (e.g., 1003 ± 20 cm) -1 1390±20cm -1 Remove unmatched redundant areas (such as 2000–2800 cm). -1 (Regions without characteristic peaks). Calculate the signal-to-noise ratio (SNR) of the cropped spectrum, requiring an SNR > 3 for each characteristic band. Otherwise, readjust the cropping range or return to step 212 to optimize baseline subtraction to generate effective spectral data free from fluorescence interference (e.g., retaining only 10 key characteristic bands, with a total width of approximately 400 cm). -1 This provides a clean feature input for subsequent vectorization.
[0035] Optionally, step 212, based on a multi-scale dynamic weight iterative fitting model, performs dynamic baseline estimation and subtraction on the initial region of interest spectrum to generate a baseline-de-drift spectrum, specifically including:
[0036] Step 2121: Divide the spectrum of the initial region of interest into subbands according to wavenumber, and perform sliding window decomposition on the divided subbands to obtain a spectral sequence based on window decomposition.
[0037] Step 2122: Based on the defined spectral point weighting function, perform weighted polynomial fitting on the window-decomposition-based spectral sequence to generate a preliminary estimated baseline curve.
[0038] Step 2123: Based on the statistical temperature factor, compensate for the thermal expansion and contraction effect of the initially estimated baseline curve to calculate the smoothness index of the spectrum after deducting the baseline.
[0039] Step 2124: Based on the smoothness index of the spectrum after baseline subtraction, perform median filtering on the initially estimated baseline curve to generate a baseline-de-drift spectrum.
[0040] Preferably, in step 2121, the initial step 211 spectrum (including residual fluorescence background, baseline drift, and characteristic peak signals, with a wavenumber range such as 600-2400 cm⁻¹) is used. -1 Based on the frequency characteristics of the Raman spectrum, the initial step 211 spectrum is divided into three sub-bands: the low-frequency sub-band (600-1200 cm⁻¹). -1 ): Mainly contains slowly changing fluorescence background and some low-frequency chemical bond vibration signals; mid-frequency subband (1200-1800 cm⁻¹) -1 ): Characteristic peaks of key components such as proteins and amino acids are concentrated; high-frequency subband (2800-3100cm). -1 The spectrum contains high-frequency vibrational signals such as CH bonds and some noise. Further, for each sub-band, a sliding window of different sizes is used for decomposition: a 200-point sliding window with a step size of 40 points (i.e., 80% window overlap) is used for the low-frequency sub-band to fully capture slowly changing background signals; a 100-point sliding window with a step size of 20 points is used for the mid-frequency sub-band to balance the preservation of feature peak details with computational efficiency; and a 50-point sliding window with a step size of 10 points is used for the high-frequency sub-band to quickly filter high-frequency noise and preserve vibrational features. This allows for the sequential extraction of data segments from the spectral data within each sub-band according to the sliding window order, forming a window-based spectral sequence. The resulting window-based spectral sequences for the three sub-bands are obtained, each sequence consisting of multiple window data segments. For example, the low-frequency sub-band sequence contains [(window 1 data), (window 2 data), ..., (window n data)], providing multi-scale input for subsequent weighted fitting.
[0041] Preferably, in step 2122, for the window-based spectral sequence (containing window data segments with different sub-bands), a weighting function is defined for each spectral point i within the window: w i The weight of the i-th spectral point; I i Let i be the intensity value of the i-th spectral point; This represents the average intensity of spectral points within the current window; coefficient 5 is used to adjust the steepness of the weight change, and 1.2 is the characteristic peak intensity threshold coefficient. Through this function, when I... i Higher than Time (i.e., the characteristic peak region), weight w i <0.2, suppressing its influence on the fitting; when I i Close to or below Time (i.e., background area), weight w i >0.8, enhances the background fitting weight.
[0042] Furthermore, for the spectral data within each window, a cubic polynomial B(λ) = a0 + a1λ + a2λ is used. 2 +a3λ 3 A fitting is performed, where λ is the wavelength and a0, a1, a2, a3 are polynomial coefficients. The weighted mean square error is minimized. Solve for the coefficients, where y i Let B(λ) be the measured intensity value of the i-th spectral point. i The value is the predicted value of the polynomial at that point. The fitting process is repeated three times, updating the weight matrix based on the new fitting residuals after each iteration to gradually approximate the true baseline, thus generating a preliminary baseline curve. This curve fits the background signal in different subbands. In the yeast extract spectrum, the 1650 cm⁻¹ value is considered optimal. -1 The baseline residual error of the characteristic peak of amide I was reduced to 12%.
[0043] Preferably, in step 2123, the temperature influence factor k is determined based on historical experimental data statistics for the initially estimated baseline curve. T = 1 + 0.01|T + 15|, where T is the actual temperature during spectral acquisition (unit: °C). The preliminary baseline curve is corrected: B corr (λ)=B(λ)·k T .
[0044] Among them B corr B(λ) is the temperature-compensated baseline curve, and B(λ) is the preliminary estimated baseline curve. This formula uses the temperature factor k... T The baseline is scaled to compensate for spectral baseline drift caused by thermal expansion and contraction. Then, the temperature-compensated baseline curve is further subtracted to obtain the baseline-de-baseline spectrum y. corr (λ)=y(λ)-B corr (λ), and calculate its smoothness index: S is the smoothness index, y corr,i Here, represents the intensity value of the i-th point in the baseline-removed spectrum, and N is the total number of spectral data points. This index reflects the degree of fluctuation in the baseline-removed spectrum; a smaller value indicates a smoother spectrum and better baseline subtraction.
[0045] Therefore, the temperature-compensated baseline curve B is obtained. corr (λ) and the corresponding smoothness index S are used to evaluate the baseline subtraction effect and guide the next step of processing.
[0046] In step 2124, for the temperature-compensated baseline curve B corr (λ) and the smoothness index S, based on a set smoothness threshold S thresh (If set to 0.5 based on historical data statistics), if the current smoothness index S > S thresh This indicates that the baseline curve has local fluctuations or outliers. Furthermore, regarding the temperature-compensated baseline curve B... corr (λ) Apply medium filtering with a window size of 5 points (adjustable as needed). For each point i, sort the baseline values of the two points before and after it (5 points in total), and use the median value as the filtered baseline value for that point. Finally, subtract the filtered baseline curve again to obtain the final baseline-drift-free spectrum y. final (λ)=y(λ)-B filtered (λ), where B filtered (λ) represents the baseline curve after median filtering, serving as the baseline-drift-free spectrum. In the yeast extract spectrum, adenine is present at 730 cm⁻¹. -1 The baseline residual error of the peaks decreased from the initial 18% to 6%, and the retention rate of characteristic peak regions exceeded 95%, providing high-quality data for subsequent spectral analysis.
[0047] Optionally, such as Figure 3 The image shown is a denoised and normalized spectral data plot. Step 22: Perform median filtering and normalization on the effective spectral data to obtain denoised and normalized spectral data, including the following steps:
[0048] Step 221: Detect noise points in the effective spectral data after removing fluorescence interference and generate a noise mask matrix;
[0049] Step 222: Based on the dynamic window mid-range filtering algorithm, combined with the noise mask matrix generated in step 221, noise suppression is performed on the effective spectral data for removing fluorescence interference to generate a denoised spectrum;
[0050] Step 223: Based on the multimodal normalization framework, the denoised spectrum generated in step 222 is subjected to intensity normalization processing to obtain denoised and normalized spectral data.
[0051] contrast Figure 2 visible, Figure 3 The fluctuations in the curves are more regular, and noise and other interference factors are effectively suppressed.
[0052] Optionally, step 221, based on the spectral feature peak identification results, performs noise point detection on the effective spectral data after removing fluorescence interference to generate a noise mask matrix, specifically including the following steps:
[0053] Step 2211: Perform pre-smoothing processing based on dynamic curvature on the effective spectral data to remove fluorescence interference, so as to obtain curvature-preserving smoothed spectra;
[0054] Step 2212: Perform local standard deviation statistics on the curvature-preserving smooth spectrum based on the constructed three-scale sliding window to generate a multi-scale standard deviation matrix;
[0055] Step 2213: Based on the set peak-valley separation factor, perform adaptive threshold discrimination processing on the multi-scale standard deviation matrix to generate a noise binary discrimination matrix;
[0056] Step 2214: Generate a noise mask matrix based on the noise binary discrimination matrix.
[0057] Optionally, step 2214, generating a noise mask matrix based on the noise binary discrimination matrix, specifically involves: marking the characteristic peak positions of the effective spectral data free from fluorescence interference based on the statistical values of the full width at half maximum (FWHM) of the characteristic peaks in the historical standard Raman spectral library and the local curvature characteristics of the current spectrum, to determine the characteristic peak protection interval; mapping the characteristic peak protection interval to the noise binary discrimination matrix to correct the noise binary discrimination matrix; and performing a "dilation-erosion" closing operation on the corrected noise binary discrimination matrix to generate the noise mask matrix.
[0058] Preferably, in step 2211, for the effective spectral data after removing fluorescence interference (baseline subtracted, characteristic bands clipped, such as 600–1800 cm⁻¹), -1 For each point i in the spectrum, the local curvature is calculated using second-order differences:
[0059]
[0060] Where: y i Δx represents the spectral intensity value at the i-th wavelength point; Δx is the wavelength sampling interval (unit: cm). -1 ), usually 1cm -1 C i It reflects the curvature of the spectrum at that point; the larger the value, the sharper the peak.
[0061] Furthermore, based on curvature C i Adaptive adjustment of smooth window size: If C i >mean(C)+2σ C (Characteristic peak region), using a 3-point moving average to preserve peak shape details; if C i≤mean(C) (flat region), using 11-point median filtering to suppress low-frequency noise, where mean(C) is the global curvature mean, σ C denoted as the standard deviation of curvature.
[0062] Based on the above processing, a curvature-preserving smooth spectrum is generated, making phenylalanine 1003 cm⁻¹ -1 The peak width at half maximum (FWHM) error is ≤2%, while the signal-to-noise ratio (SNR) is improved by 15%.
[0063] Preferably, in step 2212, a smooth spectrum (wavelength range such as 600–1800 cm⁻¹) is preserved for the curvature. -1 (Sampling points N = 1201), defining sliding window sizes of 11, 21, and 31 points to correspond to different noise scales: 11-point window: captures high-frequency small-amplitude noise (such as electron pulses); 21-point window: balances feature preservation and moderate noise suppression; 31-point window: suppresses low-frequency large-amplitude noise (such as temperature drift). For the spectral data within each window, calculate the standard deviation: n is the window size (11, 21, or 31); This represents the average spectral intensity within the window; the window slides in steps of 5 points to cover the entire spectral band.
[0064] Based on the above processing, a three-scale standard deviation matrix is generated. Each row corresponds to the standard deviation of a window of 11, 21, or 31 points, used for multi-scale noise localization.
[0065] Preferably, in step 2213, for each wavelength point i, the factor is calculated based on the three-scale standard deviation matrix Σ and the original spectrum y: baseline(i) is the baseline value after subtraction in step 212; γ i Reflecting the position of this point relative to the global peaks and valleys, γ i A value ≥0.3 is considered to be in the vicinity of the characteristic peak.
[0066] Furthermore, background region (γ) is generated through adaptive thresholding. i <0.3): Threshold set to μ+2σ; Characteristic peak region (γ i ≥0.3): The threshold is raised to μ+3σ; where μ and σ are the global mean and standard deviation of the three-scale standard deviations. If the standard deviation σ of a point at any scale is... w If (i) > threshold(i), then it is determined to be a noise point and assigned the value D. i =1, otherwise D i =0, generating a noisy binary discrimination matrix.
[0067] Based on the above processing, a binary noise discrimination matrix D is generated, which correctly identifies 98% of the impulse noise points in the yeast extract, with a characteristic peak misclassification rate of <5%.
[0068] Preferably, in step 2214, for the noise binary discrimination matrix D, the historical standard Raman spectrum library, and the current spectrum, the statistical values of the full width at half maximum (FWHM) of the characteristic peaks of the components are obtained from the historical library, such as 1003 cm⁻¹ for phenylalanine. -1 The peak's average FWHM value is 12 cm. -1 Combined with the current local curvature C of the spectrum i Dynamically calculate the protection radius:
[0069]
[0070] For each characteristic peak position λ peak The protection interval is [λ] peak -R,λ peak +R].
[0071] Furthermore, the protected interval is mapped to matrix D, and the elements within the interval are forcibly set to 0 (signal points), overwriting the original discrimination result. Finally, morphological closing operation is performed: structuring element: 3×3 rhombus matrix; operation steps: first dilation then erosion, i.e.: D erode =D dilate SE, thereby filling the internal pores of the noise block, smoothing the edges, and eliminating isolated noise blocks with an area of less than 5 points.
[0072] Based on the above processing, a noise mask matrix M is finally generated, where M i =1 represents a noise point, M i =0 represents the signal point, the characteristic peak protection rate is ≥95%, and the noise block identification accuracy reaches 97%.
[0073] Optionally, step 23, vectorizing the denoised and normalized spectral data to obtain Raman spectral feature vectors, includes the following steps:
[0074] Step 231: Based on the physicochemical properties of spectral characteristic peaks, perform multi-scale feature extraction on the denoised and standardized spectral data to generate a multi-dimensional feature matrix;
[0075] Step 232: Based on the component-spectral association knowledge base, perform semantic enhancement encoding on the multidimensional feature matrix to construct a feature-component association graph;
[0076] Step 233: Construct a model based on the biomolecular response tensor, vectorize the feature-component correlation diagram, and obtain the Raman spectral feature vector.
[0077] Preferably, in step 231, the noise-reduced and normalized spectral data (wavelength range such as 600–1800 cm⁻¹) are... -1 For peaks with SNR>10, the physicochemical properties of the characteristic peaks are analyzed: corresponding to the chemical bond vibration frequencies, the peak position (λ) is identified by using the second derivative method to identify local maxima, such as 1003 cm⁻¹. -1 The corresponding C-C stretching vibration of phenylalanine; the normalized intensity value is taken as the peak intensity (I), reflecting the relative content of the component, such as 1650 cm⁻¹. -1 The intensity of amide I band is positively correlated with protein concentration; the half-width at half-maximum (FWHM), symmetry, and area are calculated to form peak shape parameters, such as the characteristic peak FWHM of highly hydrolyzed yeast extract being 15% narrower than that of low-hydrolyzed samples.
[0078] Furthermore, differential features within a 5-point sliding window are extracted for each feature peak to perform local scale calculations, such as the first derivative. Second derivative ±50cm from the peak position -1 The spectral envelope features of the interval were analyzed, and the top 10 low-frequency coefficients were extracted using Discrete Cosine Transform (DCT) for mesoscale calculations. Finally, the statistical features of the entire band were used as the global scale, such as the mean μ, standard deviation σ, and information entropy H = -∑p. i log p i (p i (Normalized intensity probability).
[0079] Finally, the above features are concatenated along the "peak attribute-scale-feature type" dimension to form a multidimensional feature matrix. Where: N is the number of characteristic peaks (e.g., N = 12 key peaks in yeast extract); M is the dimension of a single peak feature (5 dimensions locally + 10 dimensions at the mesoscale + 3 dimensions globally = 18 dimensions). For example, the feature vector of the phenylalanine peak in the yeast sample is [1003, 0.85, 12.3, 0.21, -0.05, ...], which preserves the physicochemical properties and multi-scale distribution of the peak.
[0080] In step 232, knowledge base matching and association strength calculation are performed on the multidimensional feature matrix F and the component-spectral association knowledge base (containing spectral data of 200+ fermentation product components): for each feature peak f in the matrix i Search for matching components in the knowledge base and calculate the relevance using cosine similarity: c j The standard feature vector of component j in the knowledge base; s ≥ 0.8 indicates a strong correlation (e.g., 1003cm). -1 (s = 0.92 with phenylalanine).
[0081] Furthermore, through semantically enhanced encoding, for each feature peak f i If component c is matched j Then add the semantic tag L(f) i ) = c j And generate weight w i,j =s(f i ,c j For unmatched peaks (such as spurious peaks caused by noise), mark them as L(f). i = unknown, weight w i,j =0.1.
[0082] Finally, the node set is constructed using the feature-component correlation graph: feature peak node V F ={f1,f2,...,f N}, component node V C ={c1,c2,...,c M}; Edge set: If s(f i ,c j If f > 0.5, then add edge e(f) i ,c j ), with a weight of w i,j This forms a bipartite graph G = (V F ∪V C E), as a feature-component correlation graph G, such as 1003 cm in yeast samples. -1 The edge weight between the peak and the phenylalanine node is 0.92, 1650 cm⁻¹. -1 The edge weight between the peak and the protein node is 0.88, which visualizes the semantic relationship between the components and the spectral features.
[0083] Preferably, in step 233, the biomolecular response tensor is constructed for the feature-component correlation graph G and the biomolecular response tensor construction model: a third-order tensor is defined. N represents the number of characteristic peaks (e.g., 12), M represents the number of components (e.g., 20), and K represents the number of metabolic pathways (e.g., 5); element T i,j,k The response intensity of characteristic peak i to component j through metabolic pathway k is initialized from metabolic network data in the knowledge base: T i,j,k =w i,j ·p(k|c j ), p(k|c j ) is component c j The probability of participating in pathway k (e.g., the probability of phenylalanine participating in the tyrosine metabolic pathway is 0.7). Furthermore, through graph structure vectorization, a Graph Attention Network (GAT) is used to embed the associated graph G, generating node representation vectors: Where: N(i) is the set of neighboring nodes of node i; α i,jFor attention weights, defined by LeakyReLU(a T [Wh i ‖Wh j ]) Calculate; W is the trainable weight matrix, and σ is the activation function.
[0084] Finally, tensor decomposition and eigenvector generation are performed: the tensor T is decomposed using Tucker decomposition to obtain the core tensor G and factor matrices U, V, W, and a flattened representation of the core tensor is taken; the feature peak embedding vector generated by GAT is concatenated with the core tensor representation, and the dimensionality is reduced to 128 dimensions through a fully connected layer. in This indicates a splicing operation.
[0085] Based on the above processing, a 128-dimensional Raman spectral feature vector v is finally generated. For example, in the vector of the yeast sample, the weight of the phenylalanine-related dimension is 0.75 and the weight of the protein-related dimension is 0.68, which encodes the physicochemical properties and biological semantic information of the spectral features.
[0086] Optionally, see Figure 4 As shown, step 3, inputting the Raman spectral augmentation vector into the fermentation product quality prediction model for forward propagation to predict the quality of microbial fermentation products, specifically includes the following steps:
[0087] Step 31: Input the Raman spectral feature vector into the dynamic wavenumber sensing convolutional layer of the fermentation product quality prediction model for adaptive extraction of chemical bond vibration features to obtain a multi-scale feature map containing chemical bond specificity.
[0088] Step 32: Input the multi-scale feature map into the low-temperature noise suppression pooling layer for environmentally adaptive dimensionality reduction to generate a compressed feature map resistant to low-temperature noise;
[0089] Step 33: Input the compressed feature map into the metabolic pathway association attention layer for biomarker semantic enhancement processing to generate a weighted feature matrix with metabolic semantics;
[0090] Step 34: Input the weighted feature matrix into the dual-modal fusion prediction layer for task-specific intelligent decision-making to generate quality prediction results for microbial fermentation products.
[0091] Optionally, the dynamic wavenumber-sensing convolutional layer specifically includes: a wavenumber-sensing module, a weight generation module, and a convolution operation module. Step 31 involves inputting the Raman spectral feature vector into the dynamic wavenumber-sensing convolutional layer of the fermentation product quality prediction model for adaptive extraction of chemical bond vibration features to obtain a multi-scale feature map containing chemical bond specificity. This specifically includes the following steps:
[0092] Step 311: Input the Raman spectral feature vector into the wavenumber sensing module to extract the wavelength range features corresponding to different chemical bond vibrations based on the wavenumber information of the spectrum;
[0093] Step 312: Input the statistical information of the vibrational characteristics of the chemical bond in the historical spectral data into the weight generation module to generate feature-associative adaptive weights for the wavenumber interval features;
[0094] Step 313: The convolution operation module performs a convolution operation on the Raman spectrum feature vector based on feature correlation adaptive weights to extract multi-scale feature maps containing chemical bond specificity.
[0095] To implement a dynamic wavenumber-aware convolutional layer, the wavenumber-aware module, weight generation module, and convolution operation module are each designed as specific neural network structures, and their parameters are optimized through training. Steps 311-313 are supplemented in detail below from a technical implementation perspective:
[0096] Preferably, in step 31, the wavenumber sensing module processes Raman spectral feature vectors. (L represents the number of wavelength points) and the corresponding wavelength sequence The wavelength sensing module consists of a multilayer perceptron (MLP) and a convolutional layer, denoted as M_{wp}. The MLP is used to determine intervals based on wavelength information, and the convolutional layer extracts features from the corresponding intervals. The wavelength sequence λ is input into the MLP, whose structure is [L→N]. h →B], where N h B is the number of neurons in the hidden layer, and B is the number of chemical bond types. The output of the MLP is a probability matrix. Where P b,i Let represent the probability that the i-th wavelength point belongs to the characteristic region of chemical bond b. The calculation process is as follows:
[0097] h = σ1(W1λ + b1), P = σ2(W2h + b2), where W1 and b1 are the weights and biases of the first layer of the MLP, W2 and b2 are the weights and biases of the second layer, and σ1 and σ2 are the activation functions of the two layers (e.g., ReLU and Softmax). Based on the probability matrix P, for each chemical bond b, the weighted eigenvector R is calculated. b : The feature vector R of each chemical bond b By concatenating these components, the final wavelength range feature matrix is obtained.
[0098] Based on the above processing, a wavelength range feature matrix R containing different chemical bond characteristics is finally generated, providing a basis for subsequent weight generation.
[0099] In step 312, the weight generation module consists of a bidirectional long short-term memory network (Bi-LSTM) and a fully connected layer, denoted as M. wg Bi-LSTM is used to capture temporal information in the feature matrix, and fully connected layers are used to generate weights. Its processing object is the wavelength-range feature matrix. And statistical information (mean) of the vibrational characteristics of each chemical bond in historical spectral data. Standard deviation Specifically, R is input into the Bi-LSTM, and the hidden layer dimension of the Bi-LSTM is N. s The output is The calculation process is as follows: H = Bi - LSTM(R).
[0100] Furthermore, H is concatenated with historical statistical information μ and σ and then input into the fully connected layer. The structure of the fully connected layer is [(N s +2)×B→N f →B], where N f The number of neurons in the intermediate layer is [number], and the output is an adaptive weight vector for feature association. The calculation process is as follows: Z = [H; μ; σ]h w =σ3(W w Z+b w W = σ4(W) o h w +b o ), where W w ,b w W represents the weights and biases of the intermediate layers in the fully connected layer. o ,b o These are the weights and biases of the output layer, and σ3 and σ4 are the activation functions (such as ReLU and Sigmoid), respectively.
[0101] Based on the above processing, a feature association adaptive weight vector W is finally generated, which is used to adjust the intensity of the convolution operation.
[0102] Preferably, in step 313, the convolution operation module is composed of multiple convolutional neural networks (CNNs) of different sizes operating in parallel, denoted as M. co Each CNN corresponds to a specific convolutional kernel size, used to extract features at different scales. Step 313 processes the Raman spectral feature vector. Feature association adaptive weight vector Specifically,
[0103] For each chemical bond b, according to the weight W b For the preset basic convolution kernel (K is the kernel size) is weighted to obtain the actual kernel K used. b :
[0104] Convolution operations are performed on the spectral feature vector X using convolution kernels of different sizes (e.g., K1, K2, K3) to obtain feature maps F at different scales. b,k :
[0105] F b,k =X*K b,k , where * represents the convolution operation and k represents the index of the convolution kernel size.
[0106] Finally, all chemical bonds and feature maps at all scales are stitched together and fused to obtain the final multi-scale feature map F containing chemical bond specificity: This yields a multi-scale feature map F containing chemical bond specificity, which serves as input for subsequent models.
[0107] Optionally, the low-temperature noise suppression pooling layer includes a noise evaluation module, a pooling strategy selection module, and a pooling operation module. Step 32 involves inputting the multi-scale feature map into the low-temperature noise suppression pooling layer for environmentally adaptive dimensionality reduction to generate a compressed feature map resistant to low-temperature noise. This specifically includes the following steps:
[0108] Step 321: The noise assessment module analyzes the multi-scale feature map by calculating the degree of fluctuation of pixel values in different regions of the multi-scale feature map to assess the noise intensity distribution characteristics.
[0109] Step 322: The pooling strategy selection module adaptively selects a pooling strategy based on the noise intensity distribution characteristics;
[0110] Step 323: The pooling operation module performs pooling operations on the multi-scale feature map according to the selected pooling strategy to generate a compressed feature map resistant to low-temperature noise.
[0111] Preferably, the noise assessment module employs a convolutional neural network with a U-Net architecture, denoted as M. na It consists of an encoder and a decoder, and outputs a noise intensity distribution map. Feature encoding is based on the following formula: E l =Conv down (E l-1 ), l=1,...,L,E0=F is the input multi-scale feature map Conv down This represents a convolution plus downsampling operation (such as a convolution with a stride of 2); E L This is the encoded feature representation. Noise intensity mapping is performed based on the following formula: N = σ(Conv) up (E L Conv upThis represents a deconvolution plus upsampling operation; σ is the Sigmoid activation function, which normalizes the output to the [0,1] interval; This is a noise intensity distribution mapping.
[0112] Preferably, in step 321, for multi-scale feature maps Multi-scale feature extraction is performed as follows: Features at different scales are extracted through the encoder of U-Net. For example: the first convolutional layer uses a 3×3 kernel, 64 channels, and ReLU activation; downsampling uses max pooling with a stride of 2; and deep features are enhanced using residual blocks. During sound intensity estimation, the spatial resolution is restored through the decoder, ultimately outputting a noise intensity map N, where each pixel value N(x,y) represents the noise intensity at position (x,y). Then, N is divided into K regions Ω. k Calculate the average noise intensity for each region: This forms the noise intensity vector. For example, in a low-temperature environment, the noise intensity value of the feature map edge region is 0.8 (high noise), and the value of the central region is 0.2 (low noise).
[0113] For the Pooling Strategy Selection Module, a Gated Recurrent Unit (GRU) combined with an attention mechanism is used, denoted as M. ps Output the optimal pooling strategy weights for each region. Specifically, implement noisy feature sequence modeling: h t =GRU(n t ,h t-1 ),t=1,...,K n t Let h be the noise intensity of the t-th region. t This is the hidden state of the GRU, used to capture sequence dependencies.
[0114] Perform attention weight calculation: W h And v are trainable parameters; α t Let be the attention weight for the t-th region.
[0115] Implement pooling strategy for weight generation: p t =Softmax(W p [h t ;n t ]+b p ), The weights of the t-th region for the maximum, average, and median pooling strategies are [h] t ;n t ] indicates a vector concatenation operation.
[0116] Preferably, in step 322, the noise intensity distribution vector is... Sequence feature extraction is performed, treating the noise intensity vector n as a sequence input. A GRU is used to capture the dependencies between regions, outputting a hidden state sequence {h1, h2, ..., h...}. K Then, an attention mechanism enhancement is implemented to calculate the attention weight α for each region. t The importance of high-noise regions is emphasized; for example, in low-temperature environments, the attention weight of edge regions is significantly higher than that of the center region. Next, a policy weight mapping is implemented to concatenate the hidden state and noise intensity, and then a pooled policy weight vector p for each region is generated through a fully connected layer and Softmax activation. t For example: high noise region (n t >0.7): Tends towards median pooling (weight ≈ 0.9); low noise region (n t <0.3): Tends to max pooling (weight≈0.8).
[0117] Based on the above processing, a pooling strategy weight matrix is generated. Each line p t This represents the weight allocation of the three pooling strategies for the t-th region.
[0118] Preferably, the pooling operation module employs parallel pooling units combined with a dynamic routing mechanism, denoted as M. po Output a noise-resistant compressed feature map. During parallel pooling operations: F max =MaxPool(F),F avg =AvgPool(F),F med =MedianPool(F), These represent three different pooling results. When generating the region mask, Ω t′ The corresponding original region Ω in the feature map after pooling t Location; This is a binary mask. During weighted fusion, p t,i Let F be the weight of the t-th region for the i-th pooling strategy. i ∈{F max ,F avg ,F med}, This is the final compressed feature map.
[0119] Preferably, in step 323, for the multi-scale feature map F, the pooling strategy weight matrix P performs parallel pooling calculations to simultaneously execute three pooling operations: max pooling: preserves the peak information of the features, suitable for low-noise regions; average pooling: smooths the features, suitable for medium-noise regions; median pooling: suppresses impulse noise, suitable for high-noise regions. Further, region masking is applied to map the region divisions of the original feature map to the pooled feature map, generating K region masks {M1, M2, ..., M...}. K Finally, dynamic weighted fusion is implemented for each region Ω. t′ According to the strategy weight p t The three pooling results are weighted and fused. For example, in high-noise regions (such as edge regions), the median pooling result is mainly used; in low-noise regions (such as center regions), the max pooling result is mainly used.
[0120] Based on the above processing, a compressed feature map F' resistant to low-temperature noise is generated. In an environment of -20℃, compared with traditional max pooling, the signal-to-noise ratio of the feature map is improved by 4.5dB and the key peak retention rate is improved by 18%.
[0121] Optionally, the metabolic pathway association attention layer includes: a metabolic pathway database module, an attention calculation module, and a weighted operation module. Step 33: Input the compressed feature map into the metabolic pathway association attention layer for biomarker semantic enhancement processing to generate a weighted feature matrix with metabolic semantics, specifically including the following steps:
[0122] Step 331: The metabolic pathway database module associates the compressed feature map that resists low temperature noise with the corresponding metabolic pathway and biomarker information based on the known microbial metabolic pathways and related biomarker information to generate metabolic pathway features.
[0123] Step 332: The attention calculation module calculates the metabolic semantic weights of the low-temperature noise-resistant compressed feature map based on metabolic pathway features.
[0124] Step 333: The weighted operation module performs a weighted operation on the compressed feature map according to the metabolic semantic weights to generate a weighted feature matrix with metabolic semantics.
[0125] Preferably, the Metabolic Pathway Database Module (M) mp The architecture employs a graph neural network (GNN) combined with knowledge graph embedding, denoted as M. mp This module constructs a knowledge graph based on known microbial metabolic pathways and biomarker information, and uses graph convolution operations to achieve the association mapping between feature graphs and metabolic pathways.
[0126] Specifically, when constructing the knowledge graph, a directed graph G = (V, E) is constructed, where V is a set of nodes containing biomarker nodes v. bm (such as amino acids, enzymes, etc.) and metabolic pathway nodes v mp E is the edge set, and edge e ij Represents node v i With v j The relationships between nodes (e.g., "participation in metabolic pathways" and "product generation"). Node feature matrix. D represents the node feature dimension.
[0127] When performing feature map-path association mapping: Z = GCN(X,A), where GCN is a graph convolutional network operation; Let be the adjacency matrix of the graph; This is the updated node feature matrix, which contains semantic information about metabolic pathways.
[0128] When generating metabolic pathway characteristics, F mp =Aggregate(Z) mp Z mp This represents the set of feature vectors for metabolic pathway nodes; Aggregate is the aggregation function (such as mean pooling) that outputs the metabolic pathway features. N mp This represents the number of metabolic pathways.
[0129] Preferably, in step 331, the compression feature map for low-temperature noise resistance is... Given a knowledge graph of microbial metabolic pathways and biomarkers, feature graph node embedding is performed to flatten the compressed feature graph F′ into a vector. The features are mapped to biomarker nodes through fully connected layers and added to the knowledge graph G. Then, graph convolutional updates are performed to update the graph L layers using a graph convolutional network (GCN).
[0130] H (l) Let L be the feature matrix of the nodes in the l-th layer; The adjacency matrix with self-loops; for The degree matrix; W (l) σ is the trainable weight matrix, and σ is the activation function (such as ReLU).
[0131] Metabolic pathway feature extraction is performed to obtain metabolic pathway feature F by aggregating the updated metabolic pathway node features through mean pooling. mp .
[0132] Based on the above processing, the metabolic pathway feature F is finally generated. mpFor example, in yeast fermentation, compressed feature maps are associated with metabolic pathways such as glycolysis and the tricarboxylic acid cycle, and feature representation vectors of each pathway are output.
[0133] Preferably, the attention calculation module (M) ac The architecture employs a multi-head attention (MHA) mechanism combined with a gating mechanism, denoted as M. ac This module calculates the metabolic semantic weights of the compressed feature map based on metabolic pathway features. Specifically, during multi-head self-attention calculation: MultiHead(Q,K,V)=Concat(head1,...,head) h W O , For query, key, value matrix; head i =Attention(QW i Q ,KW i K VW i V ); W i Q W i K W i V W O This is a trainable weight matrix.
[0134] Next, metabolic semantic weights are generated: α = Gating(MultiHead(F mp ,F mp ,F mp Gating is a gate function (such as Sigmoid); This is the attention weight vector for metabolic pathways.
[0135] Preferably, in step 332, the metabolic pathway characteristic F is targeted. mp Perform multi-head self-attention computation to incorporate metabolic pathway features F mp As a query, key, and value matrix, different semantic associations are captured through parallel computation using h attention heads. Gated weights are generated to pass the multi-head self-attention output through a gating function (such as Sigmoid) to generate a metabolic semantic weight vector α. For example, in the glycolysis pathway, corresponding weight values are generated based on feature importance.
[0136] Based on the above processing, a metabolic semantic weight vector α is finally generated, which represents the importance of each metabolic pathway in the compressed feature map.
[0137] Preferably, the weighted operation module (M_{wo}) adopts a weighted fusion network architecture, denoted as M_{wo}. This module weights the compressed feature map using metabolic semantic weights to generate a weighted feature matrix with metabolic semantics. Specifically, when assigning pathway-feature map weights: W assign =α·1 T , It is a vector of all 1s. This is the path-feature map weight matrix. When generating the weighted feature matrix: W is the weighted characteristic matrix; assign (i) is a slice of the weight matrix of the i-th metabolic pathway.
[0138] Preferably, in step 333, for the low-temperature noise-resistant compressed feature map F′ and the metabolic semantic weight vector α, a weight matrix expansion is performed to expand the metabolic semantic weight vector α into a weight matrix W that matches the dimension of the compressed feature map. assign Additionally, weighted fusion is performed based on W. assign We perform a weighted summation on the compressed feature map F′ to highlight features related to important metabolic pathways.
[0139] Based on the above processing, a weighted feature matrix F with metabolic semantics is finally generated. weighted In the yeast fermentation scenario, the feature weights related to amino acid synthesis pathways are increased, thereby enhancing the accuracy of subsequent quality prediction.
[0140] Preferably, the dual-modal fusion prediction layer includes a modality fusion module and a decision model module. Step 34 involves inputting the weighted feature matrix into the dual-modality fusion prediction layer for task-specific intelligent decision-making to generate quality prediction results for microbial fermentation products. This specifically includes the following steps:
[0141] Step 341: The modal fusion module fuses the weighted feature matrix with the fermentation time-series data to obtain fused features;
[0142] Step 341: The decision model module performs a nonlinear transformation on the fused features to obtain a classification probability distribution map, thereby generating quality prediction results for microbial fermentation products.
[0143] Preferably, the modality fusion module consists of a bidirectional long short-term memory network (Bi-LSTM) and a gated fusion unit, denoted as M. mf This module achieves multimodal data fusion by capturing the dynamic features of time-series data and combining them with metabolic semantic information from a weighted feature matrix. Its implementation principle can be expressed as follows:
[0144] Let the fermentation time sequence data be Where N is the number of samples, T is the number of time steps, and D is the number of time steps. t For time-series data, features are extracted using Bi-LSTM. middle and The forward and backward hidden states are respectively, and the two are concatenated to obtain the temporal features. D h For hidden layer dimensions. Weighted feature matrix. First, global average pooling is performed to obtain...
[0145] Calculate the fusion gate vector g: g = σ(W) g [F pooled H t ]+b g ), where σ is the Sigmoid activation function, W g and b g These are trainable parameters. The final fused feature F fusion For: F fusion =g⊙F pooled +(1-g)⊙H t , where ⊙ represents element-wise multiplication.
[0146] Therefore, in the specific technical implementation, for fermentation time-series data (such as temperature, pH, dissolved oxygen, etc., changing over time), a Bi-LSTM network is used to capture the time-series dependencies and obtain the temporal features of each time step. Global average pooling is then applied to the weighted feature matrix to compress spatial dimensional information into channel-dimensional feature vectors. A gating mechanism is used to dynamically adjust the fusion ratio based on the importance of the two modalities, resulting in the fused feature F. fusion This allows for the organic integration of metabolic semantic information and temporal dynamic features, thereby generating fused features. It includes metabolic semantics of the weighted feature matrix and dynamic change information of fermentation time series data.
[0147] In step 342, the decision model module uses a multilayer perceptron (MLP) combined with the softmax activation function, denoted as M. dm This module performs a nonlinear transformation on the fused features, outputting the classification probability distribution of the microbial fermentation product quality. Fusion Feature F fusion After nonlinear mapping via a multilayer perceptron: y1 = ReLU(W1F) fusion +b1), y2=ReLU(W2y1+b2) where W1, b1, W2, b2 are trainable parameters, and ReLU is the activation function. Finally, the classification probability distribution map P is generated by the Softmax function: P=Softmax(W3y2+b3), where W3 and b3 are the output layer parameters. Nclass This represents the number of quality grade categories (e.g., Excellent, Good, Average, Poor).
[0148] In the specific technical implementation, the fused features are sequentially input into the hidden layers of a multilayer perceptron. The ReLU activation function is used to introduce nonlinearity, enhancing the model's ability to express complex relationships. At the output layer, the Softmax function is used to convert the feature vector into a probability distribution, where each element represents the probability that the fermentation product belongs to a corresponding quality level. This ultimately generates a quality prediction result for the microbial fermentation product, i.e., a classification probability distribution map P. For example, the probability of predicting that the yeast fermentation product is at the "excellent" level is 0.7, providing a quantitative basis for quality control during the fermentation process.
[0149] The fusion feature F output by the modal fusion module fusion As input to the decision model module, the nonlinear transformation and softmax activation of the multilayer perceptron ultimately generate the quality prediction probability distribution. This collaborative approach enables the model to both utilize metabolic semantic information in the weighted feature matrix to determine product component characteristics and dynamically adjust prediction results based on fermentation time-series data. In practical applications, compared to single-modal prediction, the dual-modal fusion prediction layer can improve the accuracy of microbial fermentation product quality prediction by 18%, effectively guiding production process optimization.
[0150] See Figures 5-6 The model's classification accuracy on the test set samples is illustrated using a confusion matrix diagram, showing the prediction accuracy and classification performance for qualified and unqualified samples, as well as the prediction results for the amino acid content of yeast extract. Figure 5 The confusion matrix diagram for the qualification classification results of yeast extract is shown in the figure. The horizontal axis (Predictedlabel) represents the sample category predicted by the model, which is divided into two categories: "Not Qualified" and "Qualified". The vertical axis (Truelabel) represents the true category of the sample, which is also divided into two categories: "Not Qualified" and "Qualified".
[0151] Taking the top-left cell as an example, the value 0.89 in the matrix indicates that among samples with the true category of "Not Qualified," the proportion correctly predicted as "Not Qualified" by the model is 0.89, meaning the prediction accuracy for unqualified samples is 89%. The top-right cell value 0.11 indicates that among samples with the true category of "Not Qualified," the proportion incorrectly predicted as "Qualified" by the model is 0.11, meaning the misclassification rate for unqualified samples is 11%. The bottom-left cell value 0.04 indicates that among samples with the true category of "Qualified," the proportion incorrectly predicted as "Not Qualified" by the model is 0.04, meaning the misclassification rate for qualified samples is 4%. The bottom-right cell value 0.96 indicates that among samples with the true category of "Qualified," the proportion correctly predicted as "Qualified" by the model in this application is 0.96, meaning the prediction accuracy for qualified samples is 96%.
[0152] "Accuracy: 0.92" indicates that the model's overall prediction accuracy for all samples is 0.92, meaning that 92% of the samples are correctly classified. "Macro F1-score: 0.92" indicates that the macro-average F1 score is 0.92. The F1 score is an evaluation metric that comprehensively considers precision and recall. The macro-average F1 score is calculated separately for each class and then averaged to reflect the model's overall performance across different classes. In the color bar on the right, the darkness of the color represents the magnitude of the prediction probability; the darker the color, the higher the prediction probability, used to visually display the probability strength corresponding to the values in the matrix.
[0153] Figure 6 The image shows the predicted amino acid content of yeast extract. The horizontal axis (Target Variables) represents different target variables, i.e., different types of amino acids, such as Aspartic acid, Threonine, and Serine. The vertical axis (RMSE) represents the root mean square error, which measures the deviation between the model's predicted values and the actual values. A smaller RMSE value indicates a more accurate model prediction.
[0154] The pink bars represent the RMSE error of the PLS (Partial Least Squares Regression) model on the validation set for the corresponding amino acid.
[0155] The orange bars represent the RMSE error of the CNN (Convolutional Neural Network) model on the validation set for the corresponding amino acid.
[0156] The light blue bars represent the RMSE error of the PLS-CNN Error-Min Fusion strategy model on the corresponding amino acid.
[0157] ω value: Represents the weight of CNN in the fusion model, used to show the performance of the fusion model under different weight configurations.
[0158] In the scatter plot on the right, the horizontal axis (Reference Value) represents the reference value of amino acid content, i.e., the true value. The vertical axis (Predicted Value) represents the model's predicted value of amino acid content. Among the dots of different colors, the pink cross represents the correspondence between the PLS model's predicted value and the reference value in the test set. The orange dots represent the correspondence between the traditional CNN model's predicted value and the reference value in the test set. The light blue triangle represents the correspondence between the model's predicted value and the reference value in the test set. The black dashed line represents the ideal situation where the predicted value and the reference value are completely equal, i.e., y = x. The closer the scatter points are to this line, the more accurate the model's prediction. By observing the degree of deviation of different colored dots from the black dashed line, the prediction accuracy of different models in the test set can be intuitively compared.
[0159] See also Figure 7 The chart compares the performance metrics of PLS (Partial Least Squares Regression), CNN (Convolutional Neural Network), and the model proposed in this application for predicting the content of hydrolyzed and free amino acids at a laser wavelength of 1064 nm. Performance metrics include the coefficient of determination (R²). 2 ), root mean square error (RMSE) and macro normalized root mean square error (Macro-NRMSE).
[0160] R 2 (Ratio of Determination): Reflects the goodness of fit of the model to the data, ranging from 0 to 1. The closer to 1, the stronger the model's explanatory power, meaning a higher degree of fit between the predicted and actual values. Taking hydrolyzed amino acids as an example, the R-value of the PLS model... 2 The R² value is 0.9868, indicating that the PLS model can explain 98.68% of the variation in the hydrolyzed amino acid data; the R² value of the CNN model is... 2 The R-value is 0.9900, indicating a higher degree of fit; the R-value of the model in this application is... 2 The value was 0.9909, indicating that it had the best fit for the hydrolyzed amino acid data among the three.
[0161] RMSE (Root Mean Square Error): Measures the degree of deviation between the model's predicted value and the true value. The smaller the RMSE value, the closer the model's predicted value is to the true value, and the higher the model's prediction accuracy. For hydrolyzed amino acids, the RMSE of the PLS model is 0.2941, the CNN model is 0.2567, and the model in this application is 0.2449. It can be seen that the prediction error of the model in this application is relatively small, and the prediction accuracy is higher.
[0162] Macro-NRMSE (Macro-Normalized Root Mean Square Error): This metric is a normalized version of the root mean square error and is also used to evaluate the prediction error of a model. A smaller value indicates better model performance. In the prediction of hydrolyzed amino acids, the PLS model has a Macro-NRMSE of 0.1920, the CNN model has 0.1992, and the model in this application has 0.1772. The model in this application performs best in terms of normalized error.
[0163] For 1064nm / hydrolyzed amino acids: Performance of three models in predicting hydrolyzed amino acid content. Considering all indicators, the model in this application performs best in R... 2 The PLS model showed slightly higher accuracy than other models, with the lowest RMSE and Macro-NRMSE, indicating that it has a good fit and high prediction accuracy in predicting the content of hydrolyzed amino acids. For 1064nm / free amino acids: the performance of the three models in predicting the content of free amino acids at this wavelength is presented. The PLS model R... 2 The R value is 0.9435, RMSE is 0.3029, and Macro-NRMSE is 0.2920; the CNN model R... 2 The R value was 0.9506, the RMSE was 0.2832, and the Macro-NRMSE was 0.3305; the model R in this application... 2 The R value was 0.9522, the RMSE was 0.2770, and the Macro-NRMSE was 0.2777. The model in this application performed well in R... 2 It has a slight advantage, with relatively small RMSE and Macro-NRMSE, and also shows good performance in predicting free amino acid content.
[0164] Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-mentioned technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, technical solutions formed by substituting the above-mentioned features with technical features disclosed in this application (but not limited to) that have similar functions.
Claims
1. A method for analyzing the quality of microbial fermentation products based on Raman spectroscopy and neural network algorithms, characterized in that, include: Step 1: Collect Raman spectral data of microbial fermentation products; Step 2: Vectorize the Raman spectral data to obtain the Raman spectral feature vectors; Step 3: Input the Raman spectral feature vector into the fermentation product quality prediction model for forward propagation to predict the quality of microbial fermentation products.
2. The method according to claim 1, characterized in that, Step 2: Vectorize the Raman spectral data to obtain Raman spectral feature vectors, including the following steps: Step 21: Perform baseline correction and band clipping on the Raman spectral data to obtain effective spectral data free from fluorescence interference; Step 22: Perform median filtering and normalization on the effective spectral data to remove fluorescence interference, so as to obtain denoised and standardized spectral data; Step 23: Vectorize the denoised and standardized spectral data to obtain Raman spectral feature vectors.
3. The method according to claim 2, characterized in that, Step 21: Perform baseline correction and band clipping on the Raman spectral data to obtain effective spectral data free of fluorescence interference, including the following steps: Step 211: Based on the distribution characteristics of Raman spectral characteristic peaks, pre-screen the characteristic intervals of the original Raman spectral data to generate the initial region of interest spectrum; Step 212: Based on the multi-scale dynamic weight iterative fitting model, perform dynamic baseline estimation and subtraction on the initial region of interest spectrum to generate a baseline-de-drift spectrum; Step 213: Based on the database of key component characteristic peaks of fermentation products, the baseline-drift spectrum is finely trimmed for characteristic bands to obtain effective spectral data free of fluorescence interference.
4. The method according to claim 3, characterized in that, Step 212: Based on the multi-scale dynamic weight iterative fitting model, perform dynamic baseline estimation and subtraction on the initial region of interest spectrum to generate a baseline-de-drift spectrum, specifically including: Step 2121: Divide the spectrum of the initial region of interest into subbands according to wavenumber, and perform sliding window decomposition on the divided subbands to obtain a spectral sequence based on window decomposition. Step 2122: Based on the defined spectral point weighting function, perform weighted polynomial fitting on the window-decomposition-based spectral sequence to generate a preliminary estimated baseline curve. Step 2123: Based on the statistical temperature factor, compensate for the thermal expansion and contraction effect of the initially estimated baseline curve to calculate the smoothness index of the spectrum after deducting the baseline. Step 2124: Based on the smoothness index of the spectrum after baseline subtraction, perform median filtering on the initially estimated baseline curve to generate a baseline-de-drift spectrum.
5. The method according to claim 2, characterized in that, Step 22: Perform median filtering and normalization on the effective spectral data to remove fluorescence interference, in order to obtain denoised and standardized spectral data, including the following steps: Step 221: Detect noise points in the effective spectral data after removing fluorescence interference and generate a noise mask matrix; Step 222: Based on the dynamic window mid-range filtering algorithm, combined with the noise mask matrix generated in step 221, noise suppression is performed on the effective spectral data for removing fluorescence interference to generate a denoised spectrum; Step 223: Based on the multimodal normalization framework, the denoised spectrum generated in step 222 is subjected to intensity normalization processing to obtain denoised and normalized spectral data.
6. The method according to claim 2, characterized in that, Step 23: Vectorize the denoised and standardized spectral data to obtain Raman spectral feature vectors, including the following steps: Step 231: Based on the physicochemical properties of spectral characteristic peaks, perform multi-scale feature extraction on the denoised and standardized spectral data to generate a multi-dimensional feature matrix; Step 232: Based on the component-spectral association knowledge base, perform semantic enhancement encoding on the multidimensional feature matrix to construct a feature-component association graph; Step 233: Construct a model based on the biomolecular response tensor, vectorize the feature-component correlation diagram, and obtain the Raman spectral feature vector.
7. The method according to claim 2, characterized in that, Step 3: Input the Raman spectral augmentation vector into the fermentation product quality prediction model for forward propagation to predict the quality of microbial fermentation products. This includes the following steps: Step 31: Input the Raman spectral feature vector into the dynamic wavelength sensing convolutional layer of the fermentation product quality prediction model for adaptive extraction of chemical bond vibration features to obtain a multi-scale feature map containing chemical bond specificity. Step 32: Input the multi-scale feature map into the low-temperature noise suppression pooling layer for environmentally adaptive dimensionality reduction to generate a compressed feature map resistant to low-temperature noise; Step 33: Input the compressed feature map into the metabolic pathway association attention layer for biomarker semantic enhancement processing to generate a weighted feature matrix with metabolic semantics; Step 34: Input the weighted feature matrix into the dual-modal fusion prediction layer for task-specific intelligent decision-making to generate quality prediction results for microbial fermentation products.
8. The method according to claim 7, characterized in that, The dynamic wavenumber-sensing convolutional layer specifically includes: a wavenumber-sensing module, a weight generation module, and a convolution operation module. Step 31 involves inputting the Raman spectral feature vector into the dynamic wavenumber-sensing convolutional layer of the fermentation product quality prediction model for adaptive extraction of chemical bond vibration features to obtain a multi-scale feature map containing chemical bond specificity. This specifically includes the following steps: The Raman spectral feature vector is input into the wavenumber sensing module to extract wavenumber range features corresponding to different chemical bond vibrations based on the wavenumber information of the spectrum; The input to the weight generation module is based on the statistical information of the vibrational characteristics of the chemical bond in historical spectral data to generate adaptive weights for the wavelength range characteristics; The convolution operation module performs convolution operations on Raman spectral feature vectors based on feature correlation adaptive weights to extract multi-scale feature maps containing chemical bond specificity.
9. The method according to claim 7, characterized in that, The low-temperature noise suppression pooling layer includes a noise assessment module, a pooling strategy selection module, and a pooling operation module. Step 32 involves inputting the multi-scale feature map into the low-temperature noise suppression pooling layer for environmentally adaptive dimensionality reduction to generate a compressed feature map resistant to low-temperature noise. This specifically includes the following steps: The noise assessment module analyzes multi-scale feature maps and assesses the noise intensity distribution characteristics by calculating the fluctuation of pixel values in different regions of the multi-scale feature maps. The pooling strategy selection module adaptively selects the pooling strategy based on the noise intensity distribution characteristics; The pooling operation module performs pooling operations on multi-scale feature maps according to the selected pooling strategy to generate compressed feature maps that are resistant to low-temperature noise.
10. The method according to claim 7, characterized in that, The metabolic pathway association attention layer includes a metabolic pathway database module, an attention calculation module, and a weighted operation module. Step 33 involves inputting the compressed feature map into the metabolic pathway association attention layer for biomarker semantic enhancement processing to generate a weighted feature matrix with metabolic semantics. This specifically includes the following steps: The metabolic pathway database module uses known microbial metabolic pathways and related biomarker information to associate compressed feature maps that resist low-temperature noise with the corresponding metabolic pathways and biomarker information to generate metabolic pathway features. The attention computation module calculates metabolic semantic weights of the compressed feature map, which is resistant to low-temperature noise, based on metabolic pathway features. The weighted operation module performs weighted operations on the compressed feature map according to the metabolic semantic weights to generate a weighted feature matrix with metabolic semantics.
Citation Information
Cited By
AMT signal denoising method and device based on adaptive multi-stage U-Net
CN121210852A