Rapid screening method of pigment high-yield strain

By combining Raman spectroscopy with deep learning models, the problems of strain screening efficiency and accuracy were solved, achieving high-throughput and accurate strain screening, especially for the identification and screening of unknown pigment types in mutant libraries, shortening the screening cycle and reducing costs.

CN121904751APending Publication Date: 2026-04-21QINGDAO SINGLE CELL BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO SINGLE CELL BIOTECH CO LTD
Filing Date
2025-12-29
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing strain screening technologies suffer from low screening efficiency and poor accuracy, especially in microbial fermentation processes where it is difficult to achieve high throughput and accurate identification of high-yield strains.

Method used

By combining Raman spectroscopy with a deep learning model, a deep learning model based on the out-of-distribution algorithm is constructed to identify pigment features within single cells, set sorting thresholds for high-throughput screening, and iteratively optimize the model to improve screening accuracy.

Benefits of technology

It achieves high-throughput and accurate strain screening, can identify potential high-yield strains, shorten the screening cycle, reduce manpower and reagent consumption, and is suitable for the detection of unknown pigment types in mutant libraries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904751A_ABST
    Figure CN121904751A_ABST
Patent Text Reader

Abstract

The invention discloses a rapid screening method of a pigment high-yield strain, and belongs to the field of pigment synthesis. According to the technical scheme, the method comprises the steps that Raman spectrums of a plurality of pigment high-yield strains serve as a training set, and a deep learning model based on an out of distribution algorithm is constructed; determining the category of a sorted target pigment, after the target pigment is identified by the classification model, setting a sorting threshold value by using the intensity difference value of Raman peaks, carrying out flow type Raman detection on a mutant library to be screened, calculating the Raman spectrum characteristic peak difference value of the Raman spectrum judged as the target pigment in real time, and comparing the Raman spectrum characteristic peak difference value with the sorting threshold value. The cells with the characteristic peak difference value larger than the threshold value are sorted and collected. The method is applied to rapid screening of the high-yield pigment strains, solves the problems of low screening efficiency and poor screening precision of the existing high-yield pigment strain screening technology, and has the characteristics of high throughput and high screening precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of pigment synthesis, and in particular relates to a rapid screening method for high-pigment-producing strains. Background Technology

[0002] In the field of synthetic biology, the microbial synthesis of pigments is a research hotspot, especially carotenoids, which are important natural pigments widely used in the food, health products, pharmaceuticals, and cosmetics industries. Yeast, due to its efficient metabolic capacity and ease of gene manipulation, is an ideal host for carotenoid production. However, wild-type strains typically have low yields, necessitating the construction of mutant libraries through mutagenesis or genetic engineering, and the screening of high-yielding strains. Currently, commonly used strain screening techniques mainly include dilution culture, microfluidic sorting, fluorescence flow cytometry (FACS), and optical tweezers. Dilution culture is simple to operate but inefficient, easily leading to missed screening of target strains; microfluidic technology, while enabling single-cell manipulation, suffers from throughput limitations and complex equipment; FACS offers high throughput but requires fluorescent labeling and may affect cell viability; and optical tweezers suffer from low efficiency.

[0003] The aforementioned techniques for strain screening each have significant limitations and are insufficient to meet the demands of industrial-grade high-throughput screening. Raman spectroscopy, due to its label-free and non-destructive detection characteristics, offers a new technological pathway for microbial screening. This technique can directly acquire information on intracellular biochemical components, making it particularly suitable for detecting metabolites such as carotenoids that lack natural fluorescence. Therefore, developing a new technology capable of accurately identifying carotenoid characteristic signals while possessing high-throughput screening capabilities is of great significance for promoting the development of the microbial manufacturing industry.

[0004] However, in the field of microbial fermentation, establishing yield prediction regression models using Raman spectroscopy to screen high-yielding strains faces numerous technical challenges. Microbial fermentation is a highly dynamic process, and its final yield is influenced by a variety of interrelated factors, including but not limited to culture medium composition, temperature, pH, dissolved oxygen, cell metabolic state, quorum sensing, and the induction of secondary metabolites. When attempting to establish a regression model between Raman spectroscopy and final yield, spectral data from a single or a few sampling points often only reflect the cell state and pigment content at a specific time point, making it difficult to comprehensively capture all the key dynamic variables affecting the final yield. The relationship between intracellular pigment accumulation and final yield is not a simple linear one. Cells may accumulate a large amount of pigment precursors in the early stages of fermentation, but the final yield is affected by subsequent growth and transformation efficiency. During fermentation, even within the same batch culture, significant heterogeneity exists between cells, including growth rate, metabolic activity, and pigment synthesis capacity. Significant differences exist in the metabolic state and pigment accumulation levels among different cells, and the heterogeneity of single-cell Raman spectra makes it difficult to establish a unified regression model. While single-cell Raman spectra can be acquired, corresponding yield values ​​at the single-cell level cannot be obtained. Correlating population-averaged Raman spectra with yield values ​​obtained by population-based high-performance liquid chromatography (HPLC) may mask the characteristics of a few high-yielding cells, making it difficult for yield prediction regression models to effectively distinguish and quantify such differences. The complex intracellular biological matrix and diverse precursor pigments generate substantial background Raman signals, interfering with the characteristic peaks of the target carotenoid. Even if the model performs well on certain batches, its predictive reliability will significantly decrease when faced with new batches of fermentation exhibiting different heterogeneous characteristics. Furthermore, the model may become completely unusable when screening a library of mutants with entirely unknown variations. Simultaneously, there are differences between the standard Raman spectra of pure substances and the Raman spectra of intracellularly embedded carotenoids, making it impossible to establish standard curves using the spectra of pure substances. Summary of the Invention

[0005] In view of the shortcomings of the existing technology, the technical problem to be solved by the present invention is to overcome the problems of low screening efficiency and poor screening accuracy of the existing high pigment-producing strain screening technology, and to propose a rapid screening method for high pigment-producing strains with high throughput and high screening accuracy.

[0006] To solve the aforementioned technical problem, the technical solution adopted by the present invention is as follows: This invention provides a rapid screening method for high-pigment-producing strains, comprising: using the Raman spectra of several screened high-pigment-producing strains as a training set to construct a deep learning model based on the out-of-distribution algorithm.

[0007] The above process established a strategy and identification model for screening a spectral feature library of "real intracellular environment" for pigments, emphasizing the importance of constructing intracellular spectral models based on representative high-yielding strains, and avoiding the misleading nature of pure substance spectra.

[0008] This invention provides a rapid screening method for high-pigment-producing strains, comprising: determining the target pigment category for sorting; after identifying the target pigment in a classification model, setting a sorting threshold using the intensity difference of Raman peaks; performing flow Raman detection on a library of mutants to be screened; calculating the Raman spectrum characteristic peak difference in real time for the Raman spectra of the identified target pigments; comparing the difference with the sorting threshold; and sorting and collecting cells with characteristic peak differences greater than the threshold.

[0009] The above process is the first to combine Raman spectroscopy with open-set detection deep learning algorithms for the identification of cytochromes (such as carotenoids). It solves the problem of limited target pigment quantity in modeling data and the "black box" problem of traditional classification models when facing novel and unknown pigment types that may appear in mutant libraries. In particular, it provides active detection and identification capabilities for unknown pigment types that may appear in mutant library screening, improving the accuracy and generalization ability of screening, and can discover potential superior mutant strains with novel pigment structures. A deep learning-driven pigment classification model is established. The architecture of this model can seamlessly integrate new pigment categories. By simply adding data of the new category, the entire model can be retrained, or it can be expanded through transfer learning or fine-tuning strategies.

[0010] This invention provides a rapid screening method for high-pigment-producing strains, comprising: fermenting the high-producing strains collected by sorting, quantitatively analyzing the target pigments in the fermentation products, adjusting the sorting threshold of the classification model based on the analysis results, and realizing iterative optimization of the classification model.

[0011] The above process enables a "one-stop" non-destructive screening process from single-cell Raman spectroscopy to pigment "type" and "yield". It has built a fully automated, high-throughput, and iteratively optimizable integrated screening platform of "Raman spectroscopy-deep learning-intelligent sorting", which greatly shortens the screening cycle and reduces labor costs and reagent consumption.

[0012] In some embodiments, the high-pigment-producing strain is a high-carotenoid-producing strain. Carotenoids include: α-carotene (α-Car), β-carotene (β-Car), lycopene (Lyc), zeaxanthin (Zea), astaxanthin (Ast), and canthaxanthin (Can). High-producing strains include, but are not limited to, yeasts, microalgae, *Escherichia coli*, *Corynebacterium glutamicum*, and *Bacillus*; yeasts include: *Yersinia lipolytica*, *Saccharomyces cerevisiae*, *Pichia pastoris*, *Rhodotorula glutinis*, *Kluyveromyces marxi*, *Kluyveromyces lactis*, *Hansenula polymorpha*, and *Pharbitis rubescens*; microalgae include: *Haematococcus pluvialis*, *Dunaliella salina*, *Chlamydomonas aeruginosa*, *Chlorella vulgaris*, *Spirulina*, *Phaeodactylum tricornutum*, and *Micrococcus pluvialis*.

[0013] In some embodiments, the Raman spectra of several high-pigment-producing strains selected are obtained through a dual screening process: matching the Raman peak shift of the C=C stretching vibration and matching the similarity with the standard spectrum.

[0014] The aforementioned dual screening strategy facilitates the precise selection of effective spectral data that can represent the target carotenoid (rather than its precursor or byproduct) from heterogeneous cell populations.

[0015] In some embodiments, the Raman spectra of high pigment-producing strains are obtained by injecting a single-cell suspension into a microfluidic chip and acquiring the Raman spectra of a single cell using a Raman spectroscopy acquisition system.

[0016] The above technical solution employs high-throughput single-cell Raman spectroscopy acquisition technology: combined with microfluidic technology, it achieves non-destructive, in-situ, efficient, and stable capture and precise acquisition of Raman spectra of single cells. Fine-tuning of laser power, integration time, and grating line count was performed to maximize the target signal and minimize background and light quenching. This overcomes the limitation of traditional macroscopic spectral analysis in single-cell identification, laying the foundation for high-precision, high-throughput screening.

[0017] In some embodiments, the rapid screening method for the above-mentioned high-pigment-producing strains further includes data preprocessing: after removing abnormal spectra from the Raman spectrum, selecting an effective Raman signal range of 700 cm⁻¹. -1 -1800 cm -1 The baseline correction algorithm is used to remove the smooth baseline drift in the spectrum, the Savitzky-Golay smoothing method is used to reduce the random noise in the spectrum, and multiple normalization strategies are used to normalize the spectrum.

[0018] The above process applies spectral preprocessing and feature engineering such as baseline correction, noise smoothing, and normalization to process the spectrum. It adopts a feature extraction series network architecture as a powerful deep feature extractor to perform in-depth analysis of Raman spectral data, extracting the molecular fingerprint of carotenoids from the high-dimensional and complex Raman spectral signals, and realizing high-precision identification and differentiation of different carotenoids.

[0019] In some embodiments, the rapid screening method for high-pigment-producing strains further includes model training and parameter optimization: Raman spectra of pigment samples of the corresponding category are collected in multiple batches, and the training set and validation set are divided in a 7:3 ratio. The scores of all known category samples are calculated on the training set, and the threshold is determined based on the results of the validation set and the sorting purpose.

[0020] In some embodiments, the ResNet34 deep learning model based on the open set detection idea is used as a feature extractor, and the classification model is established in combination with the out of distribution algorithm.

[0021] The out-of-distribution algorithm explicitly classifies novel pigment types outside the training set as "unknown / abnormal," rather than forcibly categorizing them into known classes. This is crucial for "accidental discoveries" in mutant library screening and addresses the issue of classification models having a limited number of categories and not being able to cover all categories. The model exhibits higher robustness to samples outside the training data distribution. A threshold can be set to further exclude non-pigment spectra from pigment samples in the training set.

[0022] In some embodiments, when using a classification model to screen the mutant library to be screened, the net intensity of the key Raman peak that is strongly correlated with the content of the target pigment is extracted, the difference is sorted in descending order, and the top one percent or a proportion set according to the actual situation is selected as the sorting threshold for the current batch.

[0023] The above process is a screening method based on relative strength within a batch, effectively avoiding the impact of differences in fermentation substrates between different batches and fluctuations in overall yield levels. It avoids directly establishing complex and unstable yield prediction regression models, instead focusing on finding relative indicators of "high yield".

[0024] In some embodiments, when collecting the Raman spectra of high-pigment-producing strains, a high-density grating of 600 lines / mm is used, the laser power is controlled to be less than 300 mW, and the integration time is optimized to be less than 1 second.

[0025] In the above process, a high-density grating of 600 lines / mm is used to obtain a diameter of less than 3 cm. -1The system achieves high spectral resolution, precisely capturing minute shifts in the vibrational modes of carotenoid molecules; it controls the laser power below 300 mW to avoid excessive laser power and integration time from photoquenching pigment signals and causing thermal damage to cells; and it optimizes the integration time to less than 1 second to maximize spectral acquisition throughput while ensuring signal-to-noise ratio (SNR), thus meeting the needs of large-scale single-cell screening.

[0026] In some embodiments, more than 3,000 high-quality single-cell Raman spectra were collected for each high-pigment-producing strain. This large-scale data acquisition strategy is the cornerstone for building robust and generalizable classification models that can fully cover intercellular heterogeneity and dynamic changes in pigment expression.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides a rapid screening method for high-pigment-producing bacterial strains, which has the following advantages: A strategy and identification model for screening a spectral feature library of "real intracellular environment" for pigments were established, emphasizing the importance of constructing intracellular spectral models based on representative high-yielding strains, thus avoiding the misleading effect of pure substance spectra. For the first time, Raman spectroscopy is combined with a deep learning algorithm for open-set detection for the identification of cytochrome (such as carotenoids). This solves the problem of the limited number of target pigments in the modeling data and the "black box" problem of traditional classification models when facing novel and unknown pigment types that may appear in the mutant library. In particular, it provides the ability to actively detect and identify unknown pigment types that may appear in the mutant library screening, improves the accuracy and generalization of screening, and can discover potential superior mutants with novel pigment structures. A deep learning-driven pigment classification model was established. The model's architecture can seamlessly integrate new pigment categories. By simply adding data from the new category, the entire model can be retrained, or it can be expanded through transfer learning or fine-tuning strategies. It has achieved a "one-stop" non-destructive screening from single-cell Raman spectroscopy to pigment "type" and "yield", and built a fully automated, high-throughput, and iteratively optimizable integrated screening platform of "Raman spectroscopy-deep learning-intelligent sorting", which greatly shortens the screening cycle and reduces labor costs and reagent consumption. Attached Figure Description

[0028] Figure 1 A flowchart illustrating the rapid screening method for high-pigment-producing strains provided by this invention; Figure 2 The original Raman spectrum collected is provided in the embodiment of the present invention; Figure 3 This is a statistical chart of the C=C peak shift for different samples provided in an embodiment of the present invention; Figure 4 The validation set classification recall and precision results of the classification model provided in this embodiment of the invention; Figure 5 This is an intensity difference distribution diagram of the C=C vibration Raman peaks provided in an embodiment of the present invention; Figure 6 The ranking of the top 10% of strength differences provided in the embodiments of the present invention; Figure 7 The results of β-carotene yield analysis provided in the embodiments of the present invention are shown. Detailed Implementation

[0029] The technical solutions in specific embodiments of the present invention will now be described in detail and completely with reference to the accompanying drawings. Obviously, the described embodiments are merely some specific implementations of the overall technical solution of the present invention, and not all implementations. Based on the overall concept of the present invention, all other embodiments obtained by those skilled in the art fall within the scope of protection of the present invention.

[0030] This invention provides a rapid screening method for high-pigment-producing bacterial strains. This method is based on flow cytometry Raman sorting technology, which identifies carotenoid types using a classification model and combines this with the characteristic peak intensity of the pigments to screen for high-pigment-producing bacterial strains. The process is as follows: Figure 1 As shown, it specifically includes: 1. Preparation of high-carotenoid-producing strains Recognizing the significant differences between the Raman spectra of pure substances and those of pigment molecules encapsulated within cells (e.g., microenvironmental effects, aggregation states, cell wall / membrane signal interference), constructing a highly selective intracellular carotenoid classification model cannot rely solely on the Raman spectra of pure substances. It requires the preparation of representative samples containing the target carotenoid. This invention focuses on five carotenoids of significant economic and biological value: β-carotene (β-Car), lycopene (Lyc), zeaxanthin (Zea), astaxanthin (Ast), and canthaxanthin (Can). For each target carotenoid, multiple (no fewer than 3-5) representative and well-validated high-yielding bacterial strains were carefully selected. These strains should cover different microbial categories (e.g., yeast, microalgae, Escherichia coli, Bacillus, and Corynebacterium glutamicum), different genetic backgrounds (e.g., different mutant lines and engineered strains), and cell states to fully capture the inherent spectral characteristics and variability of the target pigment in different cell communities.

[0031] Based on the growth characteristics of different microorganisms, select a suitable culture environment and control key parameters such as culture time (e.g., until pigment synthesis reaches its peak, approximately 72 hours), temperature, light (if applicable), pH, and dissolved oxygen to maximize the intracellular accumulation of the target pigment. Harvest cells using a gentle and efficient method (e.g., low-speed centrifugation). Thoroughly wash the cells with sterile buffer (e.g., phosphate-buffered saline, PBS) to completely remove residual culture medium and avoid interference with Raman signals. Resuspend the washed cells in sterile buffer (e.g., PBS) to adjust the cell concentration to 1×10⁻⁶. 6 ~1×10 7 Cells were collected at a density of cells / mL to prepare a cell suspension, which was then introduced into a microfluidic chip for single-cell Raman spectroscopy acquisition. More than 3000 high-quality single-cell Raman spectra were collected for each high-yielding strain. This large-scale data acquisition strategy is the cornerstone for establishing robust and generalizable classification models, fully covering intercellular heterogeneity and dynamic changes in pigment expression.

[0032] 2. Single-cell Raman spectroscopy acquisition: The prepared single-cell suspension is injected into a microfluidic chip. The chip's flow control system guides each single cell to the laser scanning area. A high-sensitivity Raman spectroscopy acquisition system is used to acquire Raman spectra of the single cells located at the laser excitation point. Appropriate parameters such as laser wavelength, grating lines, laser power, and integration time are selected to obtain high-quality Raman spectra.

[0033] The main differences among different carotenoids stem from the length of the conjugated chain and the structure of the terminal groups, which lead to wavenumber shifts in Raman peak positions and differences in C=C stretching vibration Raman peaks. For example, longer conjugated chains result in better electronic delocalization and lower vibrational frequencies (smaller wavenumbers). To reflect the spectral differences among different carotenoids, the grating line count is increased to provide spectral resolution. A high-density grating of 600 lines / mm is used to obtain a resolution less than 3 cm⁻¹. -1 The system achieves high spectral resolution, precisely capturing minute shifts in the vibrational modes of carotenoid molecules; it controls the laser power below 300 mW to avoid excessive laser power and integration time from photoquenching pigment signals and causing thermal damage to cells; and it optimizes the integration time to less than 1 second to maximize spectral acquisition throughput while ensuring signal-to-noise ratio (SNR), thus meeting the needs of large-scale single-cell screening.

[0034] 3. Screening of corresponding pigments by Raman spectroscopy Recognizing that even strains producing high levels of specific pigments may exhibit heterogeneity in individual cells (e.g., differences in the initiation stage of pigment synthesis, intermediate metabolites, trace expression of non-target pigments, and variations in cell health), the Raman spectra collected from each strain do not fully represent the specific pigment. Therefore, strategies must be developed to identify the "signature" spectra that truly represent the target carotenoid. For example, in astaxanthin-producing strains, some cells may have high levels of canthaxanthin, or the cells may not be in a state conducive to astaxanthin accumulation; in these cases, the intracellular pigments are mostly precursor pigments for astaxanthin.

[0035] First, Raman spectra of various pigment standards were collected, and key Raman peaks in these standards' spectra were identified, especially those related to C=C stretching vibrations (approximately 1500-1650 cm⁻¹). -1 The peaks are related to the length of the conjugated chain and the structure of the terminal group. For example, the longer the conjugated chain, the better the electron delocalization, the lower the C=C vibration frequency, and the smaller the wavenumber of the corresponding Raman peak.

[0036] The first screening is performed using the Raman shift of the C=C peak. A cosine similarity algorithm is then used to compare the spectra with those of known carotenoid standards. A cosine similarity threshold (e.g., greater than 0.7) is set for a second screening to identify spectra with high similarity to the standards. This dual screening strategy requires not only that the key Raman peak positions of the target carotenoid match the standards, but also that the overall spectra be similar. By selecting spectra from the Raman spectra of each strain that have high similarity to the standards and match the C=C stretching vibration Raman peak shift, a high-quality training set is provided for subsequent model construction.

[0037] 4. Classification Model Construction (1) Data preprocessing After removing anomalous spectra from the Raman spectrum, the effective Raman signal range is selected, typically within 700 cm⁻¹. -1 Up to 1800 cm -1 This range contains key information about the vibrations of the carotenoid skeletal structure. Advanced baseline correction algorithms, such as Iterative Adaptive Weighted Penalized Partial Least Squares (IAPLS), polynomial fitting, and Asymmetric Weighted Penalized Partial Least Squares (AWPLS), are employed to accurately remove smooth baseline drift in the spectrum and preserve effective Raman peak information to the greatest extent possible. Savitzky-Golay smoothing and other methods are used to reduce random noise in the spectrum. Various normalization strategies are employed, including but not limited to: Min-Max Normalization and normalization of the maximum value within a specified range (e.g., normalizing the 420 cm⁻¹ value). -1 up to 2400 cm -1The maximum value of any wavelength point or interval within the range is set to 1), vector normalization, and total intensity normalization are used to normalize the spectrum, eliminating the differences in spectral intensity caused by different acquisition batches, different instrument parameters, and changes in cell state. This allows the model to focus on the relative peak intensity and peak position information of the spectrum, enhancing the robustness of the model.

[0038] (2) Model selection This invention employs an advanced deep learning model based on the Open-Set Recognition (OSR) concept to address novel or rare pigment types that may appear outside the model (not included in the training set) during mutant library screening. Specifically, feature extraction networks can be selected, including but not limited to CNN architectures such as VGG, ResNet, and GoogLeNet, and transformer models like ViT and Swin Transformer as powerful deep feature extractors. These networks can learn rich, semantically meaningful high-dimensional spectral feature representations from preprocessed Raman spectral data (which can be considered a one-dimensional "pseudo-image" or directly input into the modified network structure). Combined with the out-of-distribution algorithm, a classification model is constructed that can effectively identify known pigment types and exclude unknown (outside the training set) pigment types. This algorithm achieves open-set detection by measuring the residual classification error (RCE) between the input sample and the known category "prototype" (or cluster center). A key threshold is set by analyzing known category and "unknown / abnormal" category samples in the training and validation sets.

[0039] The out-of-distribution algorithm explicitly classifies novel pigment types outside the training set as "unknown / abnormal," rather than forcibly categorizing them into known classes. This is crucial for "accidental discoveries" in mutant library screening, while also addressing the issue that classification models have a limited number of categories and cannot cover all categories. The model exhibits higher robustness to samples outside the training data distribution. A threshold can be set to further exclude non-pigment spectra from pigment samples in the training set. Threshold determination strategy: For sorting of known categories: If the main purpose is to select high-yield samples from five known pigments, the threshold should be set as low as possible to ensure the sensitivity of the classification model, for example, less than 0.8.

[0040] For discovering new categories: If the goal is to actively discover “novel” pigment types outside the training set, the threshold should be set as high as possible to ensure that it is only judged as unknown when the RCE is significantly increased, for example, greater than 0.95.

[0041] (3) Model training and parameter optimization Raman spectra of pigment samples of the corresponding categories were collected in multiple batches, and the training set and validation set were divided in a 7:3 ratio. The scores of all known category samples were calculated on the training set, and an appropriate threshold was determined based on the results of the validation set and the sorting purpose.

[0042] 5. Model application and screening of mutant libraries After determining the pigment category to be sorted, and after the model successfully identifies the target carotenoid (such as β-carotene), accurately locate its most important C=C stretching vibration Raman peak, which is most closely related to yield (e.g., 1520 cm⁻¹ for β-carotene). -1 The system requires a maximum of 500 effective target spectra. Based on the Lambert-Beer law, the intensity of the Raman spectral signal is positively correlated with the concentration of target molecules in the sample (i.e., the amount of pigment accumulation in the cells). Therefore, the intensity of the Raman peak can be used to set the sorting threshold. To improve accuracy, when calculating the peak intensity of the target peak, the background Raman signal from the nearby "quiet zone" or spectrally flat region needs to be subtracted. This can be achieved by calculating the net peak intensity difference between the maximum peak intensity of the C=C vibration Raman peak band and the peak intensity or average value of a certain point or segment in the "quiet zone" or spectrally flat region. These difference data are then sorted from high to low, and the top 1% or other values ​​are used as the sorting threshold. In the Raman spectral sorting system, a defined threshold is set to complete the subsequent cell sorting. Cells with characteristic peak differences greater than the threshold (high-yielding cells) are sorted and collected, while the remaining cells flow into the waste liquid outlet.

[0043] 6. Production Verification The high-yielding cells obtained through sorting were plated and cultured to obtain single-clone colonies. Subsequently, the pure single-clone strains were subjected to small-scale shake-flask fermentation. High-performance liquid chromatography (HPLC) was used to quantitatively analyze the shake-flask fermentation products, accurately verifying the carotenoid yield-enhancing effect of the sorted strains. Based on the HPLC validation results, the sorting threshold was adjusted, and through iterative optimization, the average yield of sorted cells was further improved, enabling the screening system to achieve higher efficiency and accuracy.

[0044] The rapid screening method for the above-mentioned high-pigment-producing strains of the present invention: (1) High-throughput single-cell Raman spectroscopy acquisition technology was adopted: combined with microfluidic technology, non-destructive, in-situ, efficient, and stable capture and precise acquisition of Raman spectra of single cells were achieved. The laser power, integration time, and number of grating lines were finely optimized to maximize the target signal and minimize background and light quenching. This overcomes the limitation of traditional macroscopic spectral analysis in identifying single cells and lays the foundation for high-precision and high-throughput screening.

[0045] (2) By using a dual screening strategy of matching the Raman peak shift of C=C stretching vibration and matching the similarity with standard spectra, effective spectral data that can represent the target carotenoid (rather than its precursor or byproduct) can be accurately selected from heterogeneous cell populations.

[0046] (3) Deep learning-driven carotenoid classification model: The spectrum is processed by baseline correction, noise smoothing, normalization and other spectral preprocessing and feature engineering. A series of feature extraction network architectures are used as powerful deep feature extractors to perform deep analysis on Raman spectral data. The molecular fingerprints of carotenoids are extracted from the high-dimensional and complex Raman spectral signals to achieve high-precision identification and differentiation of different carotenoids.

[0047] (4) The classification model introduces the idea of ​​open set detection and introduces the out of distribution algorithm: This enables the classification model to identify pigment types outside the training set (unknown), which greatly improves the model's adaptability and robustness in dealing with complex screening scenarios such as mutant libraries. It can actively discover mutants with novel pigment structures, rather than being limited to known categories, and also closes the problem of limited pigment types in the classification model.

[0048] (5) Setting the threshold for the intensity difference of Raman peaks in dynamic C=C stretching vibration: Extract the net intensity of key Raman peaks that are strongly correlated with the content of the target pigment, sort the differences in descending order, and select the top 1% or a proportion set according to the actual situation as the screening threshold for the current batch. This screening method based on relative intensity within a batch effectively avoids the influence of differences in fermentation substrates between different batches and fluctuations in overall yield levels. It avoids directly establishing complex and unstable yield prediction regression models, but focuses on finding relative indicators of "high yield".

[0049] (6) Intelligent screening and verification closed-loop system: The sorting threshold is dynamically set by combining the yield prediction value, and the HPLC verification results are used for feedback and iterative optimization to continuously improve the sorting efficiency and the success rate of obtaining high-yielding plants.

[0050] (7) Integrated high-yield cell screening process: The above technologies are organically integrated to form a complete automated process from single-cell spectral acquisition, species identification, yield prediction to final high-yield cell sorting. It realizes efficient, high-throughput, high-selectivity, and quantifiable screening of high-yield strains of cytochromes (especially carotenoids).

[0051] The rapid screening method for the above-mentioned high-pigment-producing strains has the following characteristics: (1) For the first time, Raman spectroscopy was combined with a deep learning algorithm for open-set detection for the identification of cytochrome (carotenoid) species: This solves the problem of the limited number of target pigments in the modeling data and the "black box" problem of traditional classification models when facing novel and unknown pigment types that may appear in the mutant library. In particular, it provides the ability to actively detect and identify unknown pigment types that may appear in the mutant library screening, improves the accuracy and generalization ability of screening, and can discover potential superior mutants with novel pigment structures.

[0052] (2) Establish a deep learning-driven carotenoid classification model. The architecture of this model can seamlessly integrate new carotenoid categories. By simply adding data of the new category, the entire model can be retrained, or it can be extended through transfer learning or fine-tuning strategies.

[0053] (3) The integration of metabolic engineering and artificial intelligence, based on the dynamic culture regulation technology of metabolic analysis and the deep integration of spectral analysis of deep learning, realizes the intelligent association of the whole chain of "species identification - product synthesis capability - sorting decision".

[0054] (4) A strategy for screening the spectral feature library of "real intracellular environment" for carotenoids was established. Identification model: The importance of constructing intracellular spectra based on representative high-yielding strains was emphasized to avoid the misleading effect of pure substance spectra. A dual screening strategy was proposed, which involves matching the Raman peak shift of C=C stretching vibration and matching the similarity with standard spectra.

[0055] (5) A dynamic C=C stretching vibration Raman peak intensity difference threshold setting method was established to avoid directly establishing a complex and unstable yield prediction regression model, and instead focus on finding a relative indicator of "high yield".

[0056] (6) A "one-stop" non-destructive screening process was achieved, from single-cell Raman spectroscopy to the "type" and "yield" of carotenoids. A fully automated, high-throughput, and iteratively optimizable integrated screening platform of "Raman spectroscopy-deep learning-intelligent sorting" was constructed, which greatly shortened the screening cycle and reduced labor costs and reagent consumption.

[0057] (7) It provides cutting-edge technical support for the breeding of new carotenoid varieties that combine “directed discovery” and “accidental discovery”; it combines the dual advantages of precise screening of known high-yielding strains (directed discovery) and keen capture of unknown pigment types (accidental discovery), which significantly accelerates the discovery and development of carotenoid engineered strains with excellent traits (such as high yield, novel structure, and enhanced function).

[0058] To provide a clearer and more detailed description of the rapid screening method for high-pigment-producing strains provided by this invention, specific embodiments will be described below.

[0059] Example 1 Screening of high β-carotene-producing strains of Yersinia lipolytica S1. Preparation of high-β-carotene-producing Yersinia lipolytic yeast samples Representative samples of the target carotenoids include: β-carotene (β-Car), lycopene (Lyc), zeaxanthin (Zea), astaxanthin (Ast), and canthaxanthin (Can). At least 3-5 strains known to produce the corresponding pigments in high quantities should be selected.

[0060] Sample preparation process: Four strains of *Yersinia lipophila* with different pigment yields were streaked onto YPD (1% Yeast Extract, 2% Peptone, 2% Dextrose) plates and cultured for 2 days (30℃). After obtaining single colonies, each colony was picked and transferred to 5 mL of YPD medium and cultured overnight (30℃, 200 rpm). A 1% inoculum was then inoculated into 300 mL shake flasks (100 mL YPD medium) and fermented for 72 hours (30℃, 200 rpm). 500 μL of the fermentation broth was transferred to an imported 1.5 mL centrifuge tube, centrifuged at 5000 rpm for 2 min, the supernatant was discarded, and the cells were washed with sterile water. The cells were centrifuged again at 5000 rpm for 2 min, the supernatant was discarded, and the cells were resuspended in sterile water and diluted to (2~3) × 10⁻⁶. 7 CFU / mL.

[0061] S2, Single-cell Raman spectroscopy acquisition Raman spectra were collected from samples with high yields of the corresponding pigments, with more than 3000 spectra collected for each sample. After removing outliers from the Raman spectra, the original spectra are as follows: Figure 2 As shown. Find the Raman shift corresponding to the C=C stretching vibration for each spectrum. The values ​​for each spectrum are taken as 1489-1550 cm⁻¹. -1 The Raman spectrum of a cell is determined by analyzing the Raman band and performing Gaussian fitting on that band. The Raman shift corresponding to the maximum intensity value of that band after Gaussian fitting is the C=C peak position. The statistical distribution of the Raman characteristic peaks for each sample is shown below. Figure 3 As shown. Figure 3Consistent with the patterns reflected in the Raman spectra of theoretical knowledge and material standards, the C=C peak position varies due to differences in the length and end structure of the conjugated double bond chain, resulting in red shift or blue shift phenomena.

[0062] Figure 3 It can be seen that the positions of the C=C peaks differ among different pigment samples, but there is also overlap. Lycopene differs significantly from other samples. Based on the results of this figure, spectral screening was performed. The proportions of β-Car, Lyc, Zea, Ast, and Can Raman spectra selected from this set of data were 56.52%, 53.45%, 69%, 36.77%, and 63.24%, respectively. The same method was used to screen Raman spectra collected from other strains. The cosine similarity of the screened Raman spectra with the corresponding standard Raman spectra was calculated, and spectra with significant differences from the standard spectra were further removed. After screening, these were used as modeling samples for the classification model.

[0063] S3, Model Establishment The filtered data underwent data preprocessing to extract the effective band range. In this case, the extracted band range was 700 cm. -1 up to 1800 cm -1 An adaptive weighted penalized partial least squares baseline correction algorithm is used to remove gentle baseline drift in the spectrum. The Savitzky-Golay smoothing method is employed to reduce random noise in the spectrum. The band at 980cm is used. -1 up to 1030 cm -1 The maximum value within the band range is normalized to eliminate intensity differences caused by variations in conditions. For the preprocessed data, an advanced deep learning model based on Open-Set Recognition (OSR) is used, with ResNet34 as the feature extractor and an out-of-distribution algorithm to build a classification model. To screen out target pigments as much as possible, a model threshold of 0.75 is set. The model results are as follows... Figure 4 As shown, the recall and precision of each pigment in the validation set were both above 99%. The model further identified that after screening, the Raman spectra of β-Car, Lyc, Ast, Zea, and Can were excluded at 32.7%, 15.40%, 22.53%, 29.13%, and 24.27%, respectively. This further illustrates that although the C=C characteristic peak and cosine similarity were used to screen the prepared samples, the screened spectra still contained spectra that differed significantly from the corresponding pigments and were judged as XXX by the model.

[0064] S4. Model application and screening of mutant libraries Import the established model, determine the pigment categories to be sorted, and collect cell Raman spectra using flow cytometry. The spectra are further analyzed using the classification model to determine their category, accumulating more than 500 effective target spectra. According to Beer-Lambert law, the peak intensity of a Raman spectrum is positively correlated with the corresponding substance concentration; therefore, the intensity difference of the Raman peaks can be used to set the sorting threshold. Statistical analysis of the C=C vibrational Raman peaks (1510-1530 cm⁻¹) of these 500 Raman spectra was performed. -1 The maximum peak intensity of the band is 1700 cm⁻¹ -1 The peak intensity difference, intensity distribution as follows Figure 5 As shown. The strength differences are sorted from high to low, as follows: Figure 6 As shown, the top 1% of values ​​are used as the sorting threshold, which is greater than 30,000. The mutant library to be screened is prepared into a single-cell suspension and subjected to flow cytometry Raman detection. The Raman spectrum characteristic peak difference value of the Raman spectrum determined to be the target pigment is calculated in real time and compared with the set threshold. Cells with characteristic peak difference values ​​greater than the threshold (high-yielding cells) are sorted and collected, and the remaining cells flow into the waste liquid outlet. The high-yielding cells are automatically sorted and processed, and tens of thousands of cells can be processed in a single experiment.

[0065] S5. Cell verification after sorting The high-yielding cells obtained through sorting were plate-cultured to obtain single-clone colonies. These single-clone strains were then subjected to shake-flask fermentation under the same conditions and time as normal fermentation. The cell suspension was then extracted, and the yield was verified by HPLC and compared with the original strain. Specifically, this included: The high-yielding cells obtained through sorting were plated and cultured (30℃). After obtaining single clones, each single clone was picked and transferred to 5 mL of YPD medium and cultured overnight (30℃, 200 rpm). The cells were then inoculated into 300 mL shake flasks (100 mL YPD medium) at a 1% inoculum rate and fermented for 72 hours (30℃, 200 rpm). Unmutated wild-type cells were used as a control and fermented simultaneously with the selected mutant strains for 72 hours.

[0066] Take 500 μL of fermentation broth into a 1.5 mL imported centrifuge tube, centrifuge at 12000 rpm for 2 min, discard the supernatant, then wash with distilled water, centrifuge at 12000 rpm for 2 min, aspirate the supernatant, add 1 mL of 3M HCl, resuspend the cells, incubate in boiling water for 2 min, immediately place on ice for 3 min, repeat 3 times, centrifuge at 12000 rpm for 2 min, aspirate the supernatant, wash twice with 1 mL of distilled water, add 1 mL of acetone (with 1% BHT) and a trace amount of quartz sand, and vortex thoroughly for 10 min. Centrifuge at 12000 rpm for 5 min, use a 1 mL syringe to draw the extract and filter through an organic filter membrane, take 200 μL and add to a vial, and determine the product concentration by HPLC.

[0067] Quantitative analysis using high-performance liquid chromatography (HPLC) employed dual-wavelength detection, with the measurement wavelength set at 450 / 470 nm and the integration wavelength at 470 nm. The column temperature was 25 °C, and the mobile phase flow rate was 1 mL / min. Mobile phase A consisted of pure methanol; mobile phase B consisted of acetonitrile and water in a ratio of 9:1; and mobile phase C consisted of methanol and isopropanol in a ratio of 3:2. The measurement conditions were as follows: 0–15 min, 0–90% °C; 15–30 min, 90% °C; 30–35 min, 90–0% °C; and 35–55 min, 0% °C. During the final 35–55 min period, the column was reequilibrated with mobile phase B to return it to its initial state; otherwise, a slow drift in peak elution time would occur.

[0068] HPLC analysis showed that 11 mutant strains exhibited increased β-carotene production, accounting for 64.7% of the total, with an average increase of 66.83% in cell content and a maximum increase of 121.3%, reaching 88.3 mg / g. Figure 7 ).

[0069] The above-mentioned technical solution of the present invention overcomes the problems of low throughput, strong label dependence, complex operation and difficulty in establishing yield regression models in existing strain screening technologies, and provides a highly efficient screening method based on single-cell Raman spectroscopy.

[0070] The beneficial effects of this invention are as follows: (1) Innovative model construction method: For the first time, Raman spectroscopy is combined with deep learning algorithm of open set detection for the identification of cytochrome (carotenoid) types, which solves the problem of limited target pigments in modeling data and the "black box" problem of traditional classification models when facing new and unknown pigment types in mutant libraries. Moreover, the feature extraction framework can be extended to many methods.

[0071] (2) Establish a deep learning-driven carotenoid classification model, whose architecture can seamlessly integrate new carotenoid categories.

[0072] (3) By setting a threshold based on the Raman peak intensity difference of the current batch dynamic C=C stretching vibration, the influence of differences in fermentation substrates between different batches and fluctuations in overall yield level is effectively avoided. It is verified that the characteristic peak intensity difference is significantly positively correlated with the strain yield, providing a reliable theoretical basis for strain screening based on Raman spectroscopy. Compared with traditional methods that rely on fluorescent labeling or culture phenotype, this model realizes direct and non-destructive detection of cell metabolites.

[0073] (4) The application of Raman spectroscopy data has been optimized: the "fingerprint recognition" characteristic of Raman spectroscopy is combined with "relative intensity comparison".

[0074] (5) Highly efficient and precise sorting system: Based on a flow cytometry Raman sorter, this invention establishes a "one-stop" non-destructive screening method for "species" and "yield". A single experiment can process tens of thousands of cells, and the screening throughput is significantly improved compared with the traditional HPLC method. Experimental results show that the β-carotene yield of the strains screened using this method is more than 30% higher than that of the starting strain, verifying the accuracy of the model and sorting.

[0075] (6) Significant technological advantages and application value 1) Label-free detection: No fluorescent labeling or genetic modification is required, maintaining the natural state of the cell; 2) Non-destructive analysis: Avoids the impact of traditional sorting techniques on cell viability; 3) Easy to operate: The entire process from spectral acquisition to sorting is automated; 4) High versatility: This model construction method can be extended to the screening of strains for other metabolites, such as astaxanthin and lycopene.

[0076] In summary, this invention achieves a major breakthrough in microbial strain screening technology by combining Raman spectroscopy with deep learning algorithms for open set detection for the first time. This breakthrough is achieved through a method for setting thresholds based on the intensity difference of Raman peaks of dynamic C=C stretching vibration in the current batch, combined with high-throughput sorting technology. This provides an efficient and reliable new method for strain improvement in the field of biomanufacturing.

Claims

1. A rapid screening method for high-pigment-producing bacterial strains, characterized in that, include: Raman spectra of several high-pigment-producing strains were used as the training set to construct a deep learning model based on the out-of-distribution algorithm. The target pigment category is determined. After the classification model identifies the target pigment, the sorting threshold is set by the intensity difference of the Raman peaks. The mutant library to be screened is subjected to flow Raman detection. The Raman spectrum characteristic peak difference of the Raman spectrum determined to be the target pigment is calculated in real time and compared with the sorting threshold. Cells with characteristic peak difference greater than the threshold are sorted and collected. The high-yielding strains collected by sorting were fermented, and the target pigments in the fermentation products were quantitatively analyzed. Based on the analysis results, the sorting threshold of the classification model was adjusted to achieve iterative optimization of the classification model.

2. The rapid screening method for high-pigment-producing strains according to claim 1, characterized in that, High-pigment-producing strains were selected from yeast, microalgae, Escherichia coli, Corynebacterium glutamicum, and Bacillus; yeasts were selected from Yeastia lipolytica, Saccharomyces cerevisiae, Pichia pastoris, Rhodotorula rubrum, Kluyveromyces marx, Kluyveromyces lactis, Hansenula polymorpha, and Pharfia rubrum; microalgae were selected from Haematococcus pluvialis, Dunaliella salina, Chlamydomonas aeruginosa, Chlorella vulgaris, Spirulina, Phaeodactylum tricornutum, and Micrococcus pluvialis.

3. The rapid screening method for high-pigment-producing strains according to claim 1 or 2, characterized in that, The Raman spectra of several high-pigment-producing strains were obtained through a dual screening process, which involved matching the Raman peak shifts of the C=C stretching vibration and matching the similarity with standard spectra.

4. The rapid screening method for high-pigment-producing strains according to claim 1 or 2, characterized in that, Raman spectra of high pigment-producing strains were obtained by injecting single-cell suspensions into a microfluidic chip and acquiring Raman spectra of individual cells using a Raman spectroscopy acquisition system.

5. The rapid screening method for high-pigment-producing strains according to claim 1 or 2, characterized in that, This also includes data preprocessing: after removing anomalous spectra from the Raman spectrum, the effective Raman signal range of 700 cm⁻¹ is selected. -1 -1800 cm -1 The baseline correction algorithm is used to remove the smooth baseline drift in the spectrum, the Savitzky-Golay smoothing method is used to reduce the random noise in the spectrum, and multiple normalization strategies are used to normalize the spectrum.

6. The rapid screening method for high-pigment-producing strains according to claim 1 or 2, characterized in that, It also includes model training and parameter optimization: Raman spectra of pigment samples of the corresponding categories are collected in multiple batches, and the training set and validation set are divided in a 7:3 ratio. The scores of all known category samples are calculated on the training set, and the threshold is determined based on the results of the validation set and the sorting purpose.

7. The rapid screening method for high-pigment-producing strains according to claim 1 or 2, characterized in that, We use ResNet34, a deep learning model based on open set detection, as the feature extractor and combine it with the out-of-distribution algorithm to build a classification model.

8. The rapid screening method for high-pigment-producing strains according to claim 1 or 2, characterized in that, When using a classification model to screen the mutant library, the net intensity of the key Raman peak that is strongly correlated with the content of the target pigment is extracted, the difference is sorted in descending order, and the top 1% or a proportion set according to the actual situation is selected as the sorting threshold for the current batch.

9. The rapid screening method for high-pigment-producing strains according to claim 4, characterized in that, When collecting Raman spectra of high-pigment-producing strains, a high-density grating of 600 lines / mm was used, the laser power was controlled to be below 300 mW, and the integration time was optimized to be less than 1 second.

10. The rapid screening method for high-pigment-producing strains according to claim 9, characterized in that, More than 3,000 high-quality single-cell Raman spectra were collected for each high-pigment-producing strain.