Prediction and screening method for two-to-two interactions between Lactobacillus bulgaricus and Streptococcus thermophilus

By constructing a genomic analysis and machine learning model of Lactobacillus Bulgaria and Streptococcus thermophilus, the problem of time-consuming and laborious screening of strain combinations in the prior art is solved, and efficient and accurate strain combination screening is achieved, and yogurt production efficiency is improved.

CN119889483BActive Publication Date: 2025-08-26INNER MONGOLIA AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510029908.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-08-26
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The prior art is time-consuming and labor-intensive to screen the combination of Lactobacillus Bulgaria and Streptococcus thermophilus, and has low prediction accuracy, making it difficult to efficiently screen out the combination of strains that can speed up the fermentation rate of yogurt production.

Method used

By calculating the whole genome k-mer data of Lactobacillus Bulgaria and Streptococcus thermophilus, the KEGG matrix was constructed, and the machine learning models such as GAN and feature selection methods were combined to screen out key features, a high-precision prediction model was constructed, and the optimal model was verified in combination with fermentation experiments.

Benefits of technology

A high-throughput and efficient screening of strain combinations that can speed up the fermentation rate of yogurt production has been achieved, which significantly improves the prediction accuracy and shortens the screening time.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention provides a method for predicting and screening two-to-two interactions between Lactobacillus bulgaricus and Streptococcus thermophilus, belonging to the field of food production technology. The method calculates the whole-genome k-mer data of two strains of Lactobacillus bulgaricus and two strains of Streptococcus thermophilus to form their respective feature vectors, and calculates the gene copy number to construct a KEGG matrix. By combining the KEGG features and the k-mer feature frequencies, a comprehensive feature vector is generated. The top 200 important features are screened from the real labeled samples using the chi-square test, gradient boosting, and variance analysis. Next, pseudo-label samples are generated using a generative adversarial network (GAN), and a machine learning model is constructed based on the real labeled samples to predict the interaction effect of the strain combination. Finally, the accuracy of the model prediction is verified through fermentation experiments, and the optimal model is selected. The present invention can efficiently predict the interactive symbiotic potential of the strain combination, thereby improving the efficiency and quality of yogurt production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of food production, specifically the technical field of yogurt production, and in particular to a method for predicting and screening a two-to-two interaction between Lactobacillus bulgaricus and Streptococcus thermophilus. Background Art

[0002] Yogurt is a curd-like product made by fermenting Lactobacillus bulgaricus and Streptococcus thermophilus in pasteurized or concentrated milk, with or without the addition of milk powder (or skim milk powder). The finished product contains a large number of active microorganisms. Yogurt is rich in nutrients such as calcium, protein, riboflavin, and vitamins. Yogurt has been shown to balance intestinal flora, enhance immunity, lower cholesterol, and slow aging. As more and more people consume yogurt on a daily basis, this places higher demands on the quality of yogurt production.

[0003] Existing technology combines two strains of Streptococcus thermophilus and two strains of Lactobacillus bulgaricus. These are randomly combined in biological experiments, and then phenotypic data such as acid production rate and proteolytic capacity are measured to determine whether these four strains can interact symbiotically to accelerate fermentation during yogurt production and improve fermentation properties such as viscosity and water retention. This method is time-consuming and labor-intensive, with low yields. Determining whether a group of bacteria interacts can take three to four months. Summary of the Invention

[0004] The purpose of the present invention is to provide a two-to-two interaction prediction and screening method for Lactobacillus bulgaricus and Streptococcus thermophilus, which can achieve high-throughput and efficient prediction while ensuring prediction accuracy.

[0005] To achieve the above objectives, the present invention provides a method for predicting and screening a two-to-two interaction between Lactobacillus bulgaricus and Streptococcus thermophilus, comprising the following steps:

[0006] Step S1: Calculate the k-mer data of the whole genomes of two strains of Lactobacillus bulgaricus and two strains of Streptococcus thermophilus respectively, and calculate their respective dimensional feature vector, calculate the gene copy number of each strain to form the KEGG matrix;

[0007] Step S2: The KEGG features of the four strains were merged according to the principle of adding the overlapping gene copy numbers and duplicating the non-overlapping gene copy numbers to obtain Features; the k-mer feature frequencies of the four strains are accumulated to obtain Features Features and Feature concatenation Features

[0008] Step S3: Set the number of true label samples to , according to this True label sample pairs The first 200 features in the feature importance ranking list were selected through three feature selection methods: chi-square test, gradient boosting and variance analysis.

[0009] Step S4: The GAN generator and discriminator work alternately to complete the iterative process of generating fake data and distinguishing true and false data, and finally generate pseudo-labeled samples;

[0010] Step S5: constructing a machine learning model based on the real label samples and the pseudo label samples, and then using the constructed machine learning model to predict Lactobacillus bulgaricus and Streptococcus thermophilus in a two-to-two combination;

[0011] Step S6: Select several combinations from the predicted results to conduct fermentation tests, comprehensively evaluate the fermentation effects of the strain combinations based on the fermentation characteristics, compare the experimental results with the prediction results of the machine learning model, and select the optimal model with the highest prediction accuracy.

[0012] Preferably, in step S1, k=5-9.

[0013] Preferably, in step S5, the machine learning model includes logistic regression (LR), support vector machine (SVM), random forest (RF), K-nearest neighbor (KNN) and Gaussian naive Bayes (GNB).

[0014] Therefore, the present invention adopts the above-mentioned two-to-two interaction prediction and screening method of Lactobacillus bulgaricus and Streptococcus thermophilus, and the beneficial technical effects are as follows:

[0015] By conducting in-depth genome analysis of Lactobacillus bulgaricus and Streptococcus thermophilus, and combining a series of operations including KEGG operations, k-mer feature extraction, refined feature selection, and GAN (Generative Adversarial Network) data augmentation, a high-precision prediction model for the interactions between two strains of Lactobacillus bulgaricus and two strains of Streptococcus thermophilus was successfully constructed. This model can efficiently predict in batches whether any combination of these four strains can achieve symbiotic interactions.

[0016] During feature selection, the most significant feature combinations affecting interactions were precisely selected, ensuring that the predictive model focused on the most critical information. Furthermore, data augmentation techniques were used to further enhance the effectiveness of machine learning modeling, improving both prediction efficiency and throughput while also ensuring the accuracy of prediction results. This series of optimization measures collectively boosts the potential for application of this invention in dairy fermentation and other related fields. DETAILED DESCRIPTION

[0017] The technical solution of the present invention is further illustrated by the following examples.

[0018] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0019] Example 1

[0020] A method for predicting and screening a two-to-two interaction between Lactobacillus bulgaricus and Streptococcus thermophilus comprises the following steps:

[0021] Step S1: Feature extraction of Lactobacillus bulgaricus and Streptococcus thermophilus.

[0022] Calculate the k-mer (k=5-9) data of the whole genome of 2 strains of Lactobacillus bulgaricus and 2 strains of Streptococcus thermophilus respectively, and calculate their respective The gene copy number of each strain was calculated using CENSOR, CNVnator and other software to form a KEGG matrix.

[0023] Step S2: characteristic combination of 2 strains of Lactobacillus bulgaricus and 2 strains of Streptococcus thermophilus.

[0024] The KEGG features of the four strains were merged according to the principle of adding the overlapping gene copy numbers and duplicating the non-overlapping gene copy numbers. Features.

[0025] The k-mer feature frequencies of the four bacterial strains are accumulated to obtain Features.

[0026] Will Features and Feature concatenation Features.

[0027] Step S3: Perform feature selection based on the existing small amount of labeled data.

[0028] If the number of labeled samples is , according to this labeled sample pairs The features are selected through three feature selection methods: chi-square test, gradient boosting and variance analysis, and the top 200 features are selected from the feature importance ranking list of the three methods.

[0029] Step S4: data enhancement.

[0030] right The iterative process of generating fake data and distinguishing true and false data is completed through the two alternating steps of the GAN generator and discriminator, and finally the Pseudo-labeled data.

[0031] Step S5: Model construction.

[0032] use samples ( is generated, The results are true labeled positive samples) using LR, SVM, RF, KNN, and GNB modeling to construct five machine learning models. They were used to predict 265,364,100 2:2 combinations consisting of 181 strains of Lactobacillus bulgaricus and 181 strains of Streptococcus thermophilus already in the laboratory. The model prediction results for all combinations were obtained (0 or 1, 0 indicates no interaction, 1 indicates interaction), which were then submitted to the laboratory for verification.

[0033] Step S6: Laboratory verification and optimal model determination.

[0034] 30 groups are randomly selected from step S5 to conduct fermentation experiments. Based on the fermentation characteristics such as fermentation time, viscosity, and water holding capacity obtained from the fermentation experiments, the fermentation labels of the 30 strain combinations are comprehensively evaluated to determine whether they are 0 or 1.

[0035] The results of laboratory verification were compared with the prediction results of five machine learning models, and the optimal model was selected as the logistic regression model.

[0036] In step S7, the optimal model obtained in step S6 was used to conduct a set of experiments to predict the interaction between two strains of Lactobacillus bulgaricus (IMAU20360 and IMAU20428) and two pairs of Streptococcus thermophilus (IMAU10630 and IMAU40145). The predicted result, "interaction," was output within 5 seconds. This demonstrates that the time required to screen a 2:2 starter strain using the present invention is significantly shorter than that required in a laboratory setting.

[0037] It is worth noting that the contents not elaborated in detail in the present invention are all prior art and are well known to those skilled in the art.

[0038] Therefore, the present invention adopts the above-mentioned two-to-two interaction prediction and screening method of Lactobacillus bulgaricus and Streptococcus thermophilus, which can achieve high-throughput and efficient prediction while ensuring the prediction accuracy.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for predicting and screening a two-to-two interaction between Lactobacillus bulgaricus and Streptococcus thermophilus, characterized in that: The following steps are involved: Step S1: Calculate the k-mer data of the whole genomes of two strains of Lactobacillus bulgaricus and two strains of Streptococcus thermophilus respectively, and calculate their respective dimensional feature vector, calculate the gene copy number of each strain to form the KEGG matrix; Step S2: The KEGG features of the four strains were merged according to the principle of adding the overlapping gene copy numbers and duplicating the non-overlapping gene copy numbers to obtain Features The k-mer feature frequencies of the four strains are accumulated to obtain Features Will Features and Feature concatenation Features Step S3: Set the number of true label samples to , according to this True label sample pairs The first 200 features in the feature importance ranking list were selected through three feature selection methods: chi-square test, gradient boosting and variance analysis. Step S4: The GAN generator and discriminator work alternately to complete the iterative process of generating fake data and distinguishing true and false data, and finally generate pseudo-labeled samples; Step S5: constructing a machine learning model based on the real label samples and the pseudo label samples, and then using the constructed machine learning model to predict Lactobacillus bulgaricus and Streptococcus thermophilus in a two-to-two combination; Step S6: Select several combinations from the predicted results to conduct fermentation tests, comprehensively evaluate the fermentation effects of the strain combinations based on the fermentation characteristics, compare the experimental results with the prediction results of the machine learning model, and select the optimal model with the highest prediction accuracy.

2. The method for predicting and screening a two-to-two interaction between Lactobacillus bulgaricus and Streptococcus thermophilus according to claim 1, wherein: In step S1, k=5-9.

3. The method for predicting and screening a two-to-two interaction between Lactobacillus bulgaricus and Streptococcus thermophilus according to claim 1, wherein: In step S5, the machine learning models include logistic regression, support vector machine, random forest, K-nearest neighbor and Gaussian naive Bayes.

Citation Information

Patent Citations

  • Cultures with improved phage resistance

    CN104531672A

  • Prediction method for interaction between lactobacillus bulgaricus and streptococcus thermophilus

    CN114999586A