Methods, apparatus, equipment and media for predicting the yield of metabolites
By constructing a prediction model based on k-mer sequences, the yield of metabolites can be predicted directly from genotype to phenotype, solving the labor-intensive problem of assessing and predicting the yield of strain metabolites in existing technologies. This achieves efficient and accurate yield prediction and supports the rapid screening and modification of microbial strains.
Patent Information
- Application Number
- CN202511398859.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Existing methods for assessing and predicting the yield of strain metabolites are labor-intensive and have low throughput, which cannot meet the needs of modern industry for rapid iteration and optimization of strains.
By constructing a prediction model based on k-mer sequences, machine learning algorithms are used to screen out feature sequences that are positively or negatively correlated with metabolites from the whole genome feature vector, enabling direct prediction from genotype to phenotype.
It enables rapid and accurate prediction of metabolite yield, avoiding the time-consuming and laborious process of traditional culture and detection, and provides an efficient and accurate computational tool for high-throughput screening and targeted modification of microbial strains, shortening the prediction cycle.
Smart Images

Figure CN120895107B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, equipment and medium for predicting the yield of metabolites. Background Technology
[0002] Current methods for assessing and predicting the yield of bacterial metabolites primarily rely on phenotypic screening. Phenotypic screening involves large-scale bacterial culture and fermentation, followed by analytical chemistry techniques such as high-performance liquid chromatography (HPLC) to analyze the fermentation broth of each strain individually, thereby determining the yield of metabolites.
[0003] Existing phenotypic screening methods are labor-intensive processes, consuming significant human, material, and time resources, and exhibiting extremely low throughput. Their screening efficiency cannot meet the demands of modern industry for rapid strain iteration and optimization. Therefore, achieving rapid and accurate prediction of metabolite yields is a crucial issue that urgently needs to be addressed in the industry. Summary of the Invention
[0004] This invention provides a method, apparatus, equipment, and medium for predicting the yield of metabolites, thereby enabling a rapid and accurate prediction process for the yield of metabolites.
[0005] This invention provides a method for predicting the yield of metabolites, comprising the following steps:
[0006] Obtain the feature set of the target strain; the feature set is constructed based on the k-mer sequences of all fixed-length nucleotide strings of the target strain that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite;
[0007] The feature set is input into the prediction model to obtain the predicted yield of the target metabolite output by the prediction model.
[0008] The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
[0009] According to a method for predicting the yield of metabolites provided by the present invention, the process of constructing the feature set includes:
[0010] Based on different k values in the k-value combinations, the whole genome sequence of the target strain is decomposed to obtain all k-mer sequences of the target strain;
[0011] Based on the frequency of occurrence of k-mer sequences of length k in the whole genome sequence of the target strain, a whole genome feature vector of the target strain is constructed;
[0012] Based on the whole genome feature vector, all k-mer sequences of the target strain are screened for features, and the feature set is constructed based on the screened k-mer sequences that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite.
[0013] According to a method for predicting the yield of metabolites provided by the present invention, the process of determining the combination of k values includes:
[0014] Obtain multiple associated gene sequences related to the synthesis of the target metabolite;
[0015] Based on the k value within the preset k value range, each associated gene sequence is decomposed to obtain the k-mer decomposed sequence corresponding to each k value within the preset k value range.
[0016] Based on the accuracy of predicting the target metabolite based on the k-mer decomposition sequence corresponding to each k value, the k values in the preset k value range are screened to determine the k value combination.
[0017] According to a method for predicting the yield of metabolites provided by the present invention, the method for constructing a whole-genome feature vector of the target strain based on the occurrence frequency of k-mer sequences of each k-value in the whole genome sequence of the target strain includes:
[0018] In the whole genome sequence of the target strain, the frequency of occurrence of k-mer sequences of each k value length was determined;
[0019] The occurrence frequencies of k-mer sequences with different k values were combined to obtain the whole genome feature vector of the target strain.
[0020] According to a method for predicting the yield of metabolites provided by the present invention, the method involves screening all k-mer sequences of the target strain based on the whole-genome feature vector, and constructing the feature set based on the screened k-mer sequences positively correlated with the target metabolite and k-mer sequences negatively correlated with the target metabolite, including:
[0021] Based on the whole genome feature vector, feature screening is performed on all k-mer sequences of the target strain to determine the k-mer sequences in the whole genome sequence of the target strain whose frequency is positively correlated with the predicted target metabolite and whose frequency is negatively correlated with the predicted target metabolite.
[0022] The feature set is constructed based on the k-mer sequences that predict positive correlation with the target metabolite and the k-mer sequences that predict negative correlation with the target metabolite.
[0023] According to the method for predicting the yield of a metabolite provided by the present invention, when the target metabolite is γ-aminobutyric acid (GABA), the associated gene sequences include the gadR gene sequence, gadA gene sequence, gadB gene sequence, gadC gene sequence, and gltX gene sequence.
[0024] The present invention also provides a device for predicting the yield of metabolites, comprising the following modules:
[0025] The data acquisition module is used to acquire the feature set of the target strain; the feature set is constructed based on the k-mer sequences of all fixed-length nucleotide strings of the target strain that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite.
[0026] The prediction module is used to input the feature set into the prediction model and obtain the yield prediction result of the target metabolite output by the prediction model.
[0027] The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
[0028] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the method for predicting the yield of metabolites as described above.
[0029] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for predicting the yield of metabolites as described above.
[0030] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a method for predicting the yield of metabolites as described above.
[0031] The present invention provides a method, apparatus, device, and medium for predicting metabolite yield. By constructing a prediction model that directly correlates k-mers (k-mers) of the genomic low-level sequence with metabolite yield, it achieves direct prediction from genotype to phenotype. This approach avoids the time-consuming and laborious experimental screening process that relies on strain culture, fermentation, and product detection. It provides an efficient and accurate computational tool for high-throughput screening and targeted modification of microbial strains, significantly shortening the metabolite prediction cycle and enabling rapid and accurate prediction of metabolite yield. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0033] Figure 1 This is a flowchart illustrating the method for predicting the yield of metabolites provided by the present invention.
[0034] Figure 2 This is a schematic diagram of the metabolic product yield prediction device provided by the present invention.
[0035] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0037] Figure 1 This is a schematic flowchart of the method for predicting the yield of metabolites provided by the present invention, as shown below. Figure 1 As shown, the method includes the following:
[0038] Step 110: Obtain the feature set of the target strain; the feature set is constructed based on the k-mer sequences that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite in all fixed-length nucleotide string k-mer sequences of the target strain;
[0039] Step 120: Input the feature set into the prediction model to obtain the predicted yield of the target metabolite output by the prediction model;
[0040] The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
[0041] The subject executing the method for predicting the yield of metabolites provided by this invention can be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. For example, a mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., while a non-mobile electronic device can be a server, network attached storage (NAS), or personal computer (PC), etc. This invention does not impose specific limitations.
[0042] The technical solution of this invention will be described in detail below using the method for predicting the yield of metabolites provided by this invention, which is executed by a computer.
[0043] In step 110, a feature set of the target strain is obtained; the feature set is constructed based on the k-mer sequences of all fixed-length nucleotide strings of the target strain that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite.
[0044] Specifically, a target strain refers to any microbial strain whose metabolite yield needs to be predicted. The target metabolite can be any compound synthesized by a microorganism that has economic or research value, such as γ-aminobutyric acid (GABA).
[0045] A feature set is a group of screened biomarkers that are strongly correlated with the yield of a target metabolite. In this invention, the feature set is constructed based on the k-mer sequences of all fixed-length nucleotide strings of the target strain that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite.
[0046] A k-mer is a continuous nucleotide substring of length k in a DNA sequence. The frequency of a k-mer sequence that is positively or negatively correlated with a target metabolite can be characterized by its occurrence in the genome.
[0047] Optionally, whole-genome sequencing can be performed on the target strain to obtain its complete genome sequence data. Based on the complete genome sequence data, all k-mers of the target strain can be identified. Feature selection for each k-mer sequence of the target strain can be performed using various feature selection algorithms. For example, statistical test-based methods, such as analysis of variance, can be used to select features by calculating the statistical significance between the frequency of each k-mer and the yield label. Model-based methods can also be used, such as utilizing the feature importance score inherent in the random forest algorithm, or using a model embedded with L1 regularization. L1 regularization compresses the coefficients of unimportant features to zero, thereby achieving automated feature selection.
[0048] By setting statistical significance or feature importance thresholds, a small subset of the most predictive core features can be selected from millions of original k-mer features of the target strain. These selected features are k-mer sequences whose frequency is positively correlated with the predicted target metabolite (e.g., their presence is associated with high yield) and k-mer sequences whose frequency is negatively correlated with the predicted target metabolite (e.g., their presence is associated with low yield). These selected positive and negative k-mer sequences together constitute the final feature set used for model training and prediction.
[0049] In step 120, the feature set is input into the prediction model to obtain the predicted yield of the target metabolite output by the prediction model.
[0050] A predictive model is a pre-trained machine learning or statistical model. During training, the predictive model learns the mapping relationship between specific k-mer sequence combinations in a feature sample set and the yield of metabolites.
[0051] By inputting the feature set of the target strain into the prediction model, the prediction model can calculate and output a predicted value based on the internally learned rules. The predicted value is the predicted result of the yield of the target metabolite of the target strain.
[0052] To obtain a prediction model, the initial prediction model needs to be trained. The training process is based on a feature sample set and the corresponding output labels of the feature sample set.
[0053] It should be noted that the sample strains are a group of strains used for model training whose actual yields of target metabolites have been determined experimentally using methods such as fermentation culture and high-performance liquid chromatography. The feature sample set is a set of features, consisting of positively correlated and negatively correlated k-mer sequences, constructed for each sample strain based on the whole genome sequence of these sample strains using the same method as constructing the feature set of the target strains. The yield label is the experimentally measured actual yield value that corresponds one-to-one with each sample strain.
[0054] The training process may include using a set of feature samples and corresponding output labels as input, and training the model using a supervised machine learning regression algorithm.
[0055] Optionally, to obtain the best predictive performance, a wide range of machine learning regression models can be evaluated and compared, and the model with the highest prediction accuracy can be selected as the prediction model. For example, models that can be evaluated and compared include: linear regression, ridge regression, Lasso regression, support vector regression (SVR), k-nearest neighbor regression (KNN), decision trees, random forests, gradient boosting machines (GBM, XGBoost, LightGBM), and multilayer perceptrons (MLP), etc.
[0056] By employing a cross-validation strategy, metrics such as the coefficient of determination (R²) and root mean square error (RMSE) are used to evaluate and compare the performance of different models. Ultimately, the model with the highest accuracy is selected as the final prediction model. For example, in constructing a prediction model for GABA yield, by comparing 15 different regression algorithms, the random forest regression model was found to significantly outperform the others, achieving a prediction accuracy (measured by R²) of 0.73 on the independent validation set, demonstrating the model's reliability. Therefore, the random forest regression model was ultimately selected as the prediction model.
[0057] The yield prediction method provided by this invention achieves direct prediction from genotype to phenotype by constructing a prediction model that directly links k-mers, a low-level genomic sequence feature, with the yield of metabolites. This approach avoids the time-consuming and laborious experimental screening process that relies on strain cultivation, fermentation, and product detection. It provides an efficient and accurate computational tool for high-throughput screening and targeted modification of microbial strains, greatly shortening the metabolite prediction cycle and enabling rapid and accurate prediction of metabolite yields.
[0058] In one embodiment, the process of constructing the feature set includes:
[0059] Based on different k values in the k-value combinations, the whole genome sequence of the target strain is decomposed to obtain all k-mer sequences of the target strain;
[0060] Based on the frequency of occurrence of k-mer sequences of length k in the whole genome sequence of the target strain, a whole genome feature vector of the target strain is constructed;
[0061] Based on the whole genome feature vector, all k-mer sequences of the target strain are screened for features, and the feature set is constructed based on the screened k-mer sequences that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite.
[0062] A combination of k values refers to a set of optimized k length values, such as {5, 6, 7}. Using multiple different k values allows for the capture of sequence pattern information at different scales. Decomposing the whole genome sequence involves using a sliding window to slide across the entire genome sequence with a step size of 1, thereby extracting all nucleotide substrings of length k.
[0063] Constructing a genome-wide feature vector refers to converting sequence information into a numerical form that can be processed by machine learning models. Specifically, for each k value in a combination of k values, we first count all possible k-mers of length k. For example, when k=5, we have... =1024 different k-mers. The frequency or number of times each appears in the whole genome sequence of the target strain is used to form a frequency vector corresponding to a k value.
[0064] Feature screening is necessary for each k-mer sequence because k-mer features across the entire genome are extremely high-dimensional, reaching millions of dimensions, and contain a large amount of redundant and noisy information unrelated to yield. Feature screening allows the identification of a small subset of core features that decisively influence yield from this massive dataset. The selected k-mer sequences, positively and negatively correlated with the target metabolite, together constitute the final, concise, and efficient feature set used for prediction.
[0065] The method for predicting the yield of metabolites provided by this invention, by clearly defining the process of extracting k-mer features at the whole genome scale and performing subsequent screening, ensures that the model can capture functional site information that may affect yield beyond the preset gene range, thereby improving the accuracy of prediction.
[0066] In one embodiment, the process of determining the combination of k values includes:
[0067] Obtain multiple associated gene sequences related to the synthesis of the target metabolite;
[0068] Based on the k value within the preset k value range, each associated gene sequence is decomposed to obtain the k-mer decomposed sequence corresponding to each k value within the preset k value range.
[0069] Based on the accuracy of predicting the target metabolite based on the k-mer decomposition sequence corresponding to each k value, the k values in the preset k value range are screened to determine the k value combination.
[0070] It should be noted that associated genes refer to genes identified through existing biological knowledge or literature mining that are related in the biosynthesis, regulation, or transport pathways of target metabolites. For example, these genes may be genes encoding key enzymes, regulatory proteins, or transport proteins.
[0071] The preset range of k values is a candidate range used to determine the optimal k value. For example, it can be set to an integer from 3 to 10. The preset range of k values is a preliminary range that can be determined in advance based on historical experience or expert experience.
[0072] For each candidate k-value, the k-mer sequences obtained from the associated genes are used as features to construct a preliminary, small-scale prediction model, and its prediction accuracy on the training set is evaluated. Selection based on prediction accuracy means comparing the performance of the preliminary model under different candidate k-values and selecting the few k-values that result in the highest prediction accuracy, forming the optimal combination of k-values. For example, in a study on GABA yield, by evaluating the impact of different k-values on model performance, it was empirically determined that the combination of k=5, 6, and 7 most effectively identifies the association with GABA yield, achieving the best preliminary prediction effect.
[0073] The metabolite yield prediction method provided by this invention ensures that the k-mer features extracted from the whole genome are extracted at the most informative scale by determining the optimal k-mer scale for whole-genome feature extraction, thus providing a foundation for building a high-performance final prediction model and further improving the accuracy of the overall prediction method.
[0074] In one embodiment, a whole-genome feature vector of the target strain is constructed based on the frequency of occurrence of k-mer sequences of length k in the whole genome sequence of the target strain, including:
[0075] In the whole genome sequence of the target strain, the frequency of occurrence of k-mer sequences of each k value length was determined;
[0076] The occurrence frequencies of k-mer sequences with different k values were combined to obtain the whole genome feature vector of the target strain.
[0077] Specifically, assuming the determined k-value combination is {5, 6, 7}, first, the frequency of all k-mers of length 5 in the entire genome of the target strain is calculated, forming a frequency vector V5. Then, similarly, the frequency of all k-mers of length 6 is calculated, forming a frequency vector V6. Next, the frequency of all k-mers of length 7 is calculated, forming a frequency vector V7. Combining the frequencies of k-mer sequences with different k-values means concatenating these vectors, i.e., V_combined = [V5, V6, V7], to form a higher-dimensional, comprehensive feature vector representing the overall genome sequence composition. This comprehensive feature vector simultaneously contains sequence pattern information at different scales.
[0078] This multi-scale feature fusion strategy enables the constructed genome-wide feature vector to simultaneously capture information from both function-related short sequence motifs and longer conserved sequence fragments. This provides richer and more comprehensive raw input data for subsequent feature selection and model training than single-scale k-mer, helping the model learn more complex sequence-yield relationships and thus improving the robustness and accuracy of predictions.
[0079] In one embodiment, the step of screening all k-mer sequences of the target strain based on the whole-genome feature vector, and constructing the feature set based on the screened k-mer sequences positively correlated with the target metabolite and k-mer sequences negatively correlated with the target metabolite, includes:
[0080] Based on the whole genome feature vector, feature screening is performed on all k-mer sequences of the target strain to determine the k-mer sequences in the whole genome sequence of the target strain whose frequency is positively correlated with the predicted target metabolite and whose frequency is negatively correlated with the predicted target metabolite.
[0081] The feature set is constructed based on the k-mer sequences that predict positive correlation with the target metabolite and the k-mer sequences that predict negative correlation with the target metabolite.
[0082] Various feature selection algorithms can be used to screen features for each k-mer sequence. For example, statistical test-based methods, such as Analysis of Variance (ANOVA), can be used to screen features by calculating the statistical significance between the frequency of each k-mer and the yield label. Model-based methods can also be used, such as utilizing the feature importance score inherent in the random forest algorithm, or using models embedded with L1 regularization. L1 regularization tends to compress the coefficients of unimportant features to zero, thereby achieving automated feature selection.
[0083] By setting statistical significance or feature importance thresholds, a small subset of the most predictive core features can be selected from millions of original k-mer features. These selected features are k-mer sequences whose frequency is positively correlated with the predicted target metabolite (e.g., their presence is associated with high yield) and k-mer sequences whose frequency is negatively correlated with the predicted target metabolite (e.g., their presence is associated with low yield). These selected positive and negative k-mer sequences together constitute a concise and efficient feature set for the final model training and prediction.
[0084] The method for predicting the yield of metabolites provided by this invention effectively solves the curse of dimensionality problem caused by whole-genome k-mer features by applying a feature selection algorithm. It not only significantly reduces computational complexity and the risk of model overfitting, but also accurately extracts sequence markers with strong biological associations with yield from massive amounts of data, greatly improving the predictive performance of the model and the interpretability of the results.
[0085] In one embodiment, when the target metabolite is γ-aminobutyric acid (GABA), the associated gene sequence includes the gadR gene sequence, gadA gene sequence, gadB gene sequence, gadC gene sequence, and gltX gene sequence.
[0086] The five genes (gadR, gadA, gadB, gadC, gltX) are a core gene family closely related to the microbial GABA biosynthesis pathway, which was precisely located through a systematic analysis of scientific literature related to GABA research.
[0087] When optimizing the k-value combination, this set of determined gene sequences is used as the exploration region. The optimal k-value combination is determined by evaluating the predictive ability of the features extracted from these genes by different k-values on GABA yield.
[0088] The following describes the metabolic product yield prediction device provided by the present invention. The metabolic product yield prediction device described below can be referred to in correspondence with the metabolic product yield prediction method described above.
[0089] like Figure 2 As shown, the device includes:
[0090] The data acquisition module 210 is used to acquire the feature set of the target strain; the feature set is constructed based on the k-mer sequences that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite in all fixed-length nucleotide string k-mer sequences of the target strain;
[0091] Prediction module 220 is used to input the feature set into the prediction model to obtain the yield prediction result of the target metabolite output by the prediction model;
[0092] The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
[0093] The yield prediction device provided by this invention achieves direct prediction from genotype to phenotype by constructing a prediction model that directly correlates k-mers (low-level genome sequence features) with metabolite yield. This approach avoids the time-consuming and laborious experimental screening process that relies on strain culture, fermentation, and product detection. It provides an efficient and accurate computational tool for high-throughput screening and targeted modification of microbial strains, greatly shortening the metabolite prediction cycle and enabling rapid and accurate prediction of metabolite yield.
[0094] In one embodiment, the data acquisition module 210 is specifically used for:
[0095] The process of constructing the feature set includes:
[0096] Based on different k values in the k-value combinations, the whole genome sequence of the target strain is decomposed to obtain all k-mer sequences of the target strain;
[0097] Based on the frequency of occurrence of k-mer sequences of length k in the whole genome sequence of the target strain, a whole genome feature vector of the target strain is constructed;
[0098] Based on the whole genome feature vector, all k-mer sequences of the target strain are screened for features, and the feature set is constructed based on the screened k-mer sequences that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite.
[0099] In one embodiment, the data acquisition module 210 is further configured to:
[0100] The process of determining the combination of k values includes:
[0101] Obtain multiple associated gene sequences related to the synthesis of the target metabolite;
[0102] Based on the k value within the preset k value range, each associated gene sequence is decomposed to obtain the k-mer decomposed sequence corresponding to each k value within the preset k value range.
[0103] Based on the accuracy of predicting the target metabolite based on the k-mer decomposition sequence corresponding to each k value, the k values in the preset k value range are screened to determine the k value combination.
[0104] In one embodiment, the data acquisition module 210 is further configured to:
[0105] The whole-genome feature vector of the target strain is constructed based on the frequency of occurrence of the k-mer sequence of each k-value in the whole genome sequence, including:
[0106] In the whole genome sequence of the target strain, the frequency of occurrence of k-mer sequences of each k value length was determined;
[0107] The occurrence frequencies of k-mer sequences with different k values were combined to obtain the whole genome feature vector of the target strain.
[0108] In one embodiment, the data acquisition module 210 is further configured to:
[0109] Based on the whole-genome feature vector, feature screening is performed on all k-mer sequences of the target strain, and the feature set is constructed based on the selected k-mer sequences positively correlated with the target metabolite and k-mer sequences negatively correlated with the target metabolite, including:
[0110] Based on the whole genome feature vector, feature screening is performed on all k-mer sequences of the target strain to determine the k-mer sequences in the whole genome sequence of the target strain whose frequency is positively correlated with the predicted target metabolite and whose frequency is negatively correlated with the predicted target metabolite.
[0111] The feature set is constructed based on the k-mer sequences that predict positive correlation with the target metabolite and the k-mer sequences that predict negative correlation with the target metabolite.
[0112] In one embodiment, the data acquisition module 210 is further configured to:
[0113] When the target metabolite is γ-aminobutyric acid (GABA), the associated gene sequences include the gadR gene sequence, gadA gene sequence, gadB gene sequence, gadC gene sequence, and gltX gene sequence.
[0114] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, communications interface 320, and memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute a method for predicting the yield of metabolites. This method includes: acquiring a feature set of a target strain; the feature set is constructed based on k-mer sequences positively correlated with the target metabolite and k-mer sequences negatively correlated with the target metabolite from all fixed-length nucleotide string k-mer sequences of the target strain.
[0115] The feature set is input into the prediction model to obtain the predicted yield of the target metabolite output by the prediction model.
[0116] The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
[0117] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0118] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is capable of executing the metabolite yield prediction method provided by the above methods, the method comprising: obtaining a feature set of a target strain; the feature set being constructed based on k-mer sequences positively correlated with the target metabolite and k-mer sequences negatively correlated with the target metabolite among all fixed-length nucleotide string k-mer sequences of the target strain;
[0119] The feature set is input into the prediction model to obtain the predicted yield of the target metabolite output by the prediction model.
[0120] The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
[0121] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for predicting the yield of metabolites provided by the methods described above, the method comprising: acquiring a feature set of a target strain; the feature set being constructed based on k-mer sequences positively correlated with the target metabolite and k-mer sequences negatively correlated with the target metabolite among all fixed-length nucleotide string k-mer sequences of the target strain;
[0122] The feature set is input into the prediction model to obtain the predicted yield of the target metabolite output by the prediction model.
[0123] The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0125] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting the yield of a metabolite, characterized in that, include: Obtain the feature set of the target strain; the feature set is constructed based on the k-mer sequences of all fixed-length nucleotide strings of the target strain that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite; The feature set is input into the prediction model to obtain the predicted yield of the target metabolite output by the prediction model. The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
2. The method for predicting the yield of metabolites according to claim 1, characterized in that, The process of constructing the feature set includes: Based on different k values in the k-value combinations, the whole genome sequence of the target strain is decomposed to obtain all k-mer sequences of the target strain; Based on the frequency of occurrence of k-mer sequences of length k in the whole genome sequence of the target strain, a whole genome feature vector of the target strain is constructed; Based on the whole genome feature vector, all k-mer sequences of the target strain are screened for features, and the feature set is constructed based on the screened k-mer sequences that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite.
3. The method for predicting the yield of metabolites according to claim 2, characterized in that, The process of determining the combination of k values includes: Obtain multiple associated gene sequences related to the synthesis of the target metabolite; Based on the k value within the preset k value range, each associated gene sequence is decomposed to obtain the k-mer decomposed sequence corresponding to each k value within the preset k value range. Based on the accuracy of predicting the target metabolite based on the k-mer decomposition sequence corresponding to each k value, the k values in the preset k value range are screened to determine the k value combination.
4. The method for predicting the yield of metabolites according to claim 3, characterized in that, The whole-genome feature vector of the target strain is constructed based on the frequency of occurrence of the k-mer sequence of each k-value in the whole genome sequence, including: In the whole genome sequence of the target strain, the frequency of occurrence of k-mer sequences of each k value length was determined; The occurrence frequencies of k-mer sequences with different k values were combined to obtain the whole genome feature vector of the target strain.
5. The method for predicting the yield of metabolites according to claim 2, characterized in that, Based on the whole-genome feature vector, all k-mer sequences of the target strain are screened for features. Based on the screened k-mer sequences positively correlated with the target metabolite and k-mer sequences negatively correlated with the target metabolite, the feature set is constructed, including: Based on the whole genome feature vector, feature screening is performed on all k-mer sequences of the target strain to determine the k-mer sequences in the whole genome sequence of the target strain whose frequency is positively correlated with the predicted target metabolite and whose frequency is negatively correlated with the predicted target metabolite. The feature set is constructed based on the k-mer sequences that predict positive correlation with the target metabolite and the k-mer sequences that predict negative correlation with the target metabolite.
6. The method for predicting the yield of metabolites according to claim 3, characterized in that, When the target metabolite is γ-aminobutyric acid (GABA), the associated gene sequences include the gadR gene sequence, gadA gene sequence, gadB gene sequence, gadC gene sequence, and gltX gene sequence.
7. A device for predicting the yield of metabolites, characterized in that, include: The data acquisition module is used to acquire the feature set of the target strain; the feature set is constructed based on the k-mer sequences of all fixed-length nucleotide strings of the target strain that are positively correlated with the target metabolite and the k-mer sequences that are negatively correlated with the target metabolite. The prediction module is used to input the feature set into the prediction model and obtain the yield prediction result of the target metabolite output by the prediction model. The prediction model is trained based on a feature sample set and the yield labels corresponding to the feature sample set. The feature sample set is constructed based on all k-mer sequences of the sample strain that are positively correlated with the target metabolite and negatively correlated with the target metabolite.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for predicting the yield of metabolites as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for predicting the yield of metabolites as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for predicting the yield of metabolites as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Screening method of biomarker and related application thereof
CN114974432A
High-added-value metabolite chassis created based on artificial intelligence
CN116312766A