Machine learning based nitrogenase activity prediction method and system
A nitrogenase activity prediction model constructed using machine learning methods, combined with the multidimensional features of nitrogenase samples, solves the problem of nitrogenase activity enhancement, achieving efficient and accurate prediction and optimization, and supporting research on the function of nitrogen fixation gene clusters.
Patent Information
- Application Number
- CN202411911037.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-20
- Filing Date
- 2024-12-24
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-12-24
AI Technical Summary
In existing technologies, factors such as nitrogenase expression levels and codon preferences, which affect nitrogen fixation efficiency, have not been fully integrated into the activity prediction and analysis framework, making it difficult to improve nitrogenase activity. Traditional methods are time-consuming and resource-intensive.
We used machine learning methods to obtain the fusion feature vector, expression level and codon preference of nitrogenase samples, and combined XGBoost and stacked SVR algorithms to build a prediction model to predict nitrogenase activity.
This paper presents an efficient and accurate method for predicting nitrogenase activity, which can quickly identify key genes and amino acid characteristics, optimize nitrogenase activity, and provide theoretical basis and practical guidance for genetic engineering modification.
Smart Images

Figure CN119832994B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biotechnology, in particular to a method and system for predicting nitrogenase activity based on machine learning. BACKGROUND
[0002] Nitrogen-fixing microorganisms convert atmospheric nitrogen into ammonia through nitrogenase, which plays a crucial role in global nitrogen cycling. Nitrogen is a major component of the Earth's atmosphere, but most organisms in nature cannot directly utilize nitrogen due to its very stable triple bond between molecules. Nitrogen-fixing microorganisms can convert atmospheric nitrogen into ammonia through nitrogenase systems, providing plants with a usable nitrogen source, thereby supporting agricultural production and promoting nitrogen cycling in ecosystems. Nitrogenase itself is a highly complex metal enzyme complex, mainly divided into three types: molybdenum nitrogenase, vanadium nitrogenase and iron nitrogenase, among which molybdenum nitrogenase is the most common and most thoroughly studied. Different types of nitrogenase depend on different metal elements as their active centers, and metals such as molybdenum, vanadium and iron play a crucial catalytic role.
[0003] The mechanism of nitrogenase is very complex, involving a series of electron transfer and chemical reaction processes. Although a large number of studies have been devoted to revealing the structure-function relationship of nitrogenase, its complete mechanism has not been fully elucidated, especially the fine regulation and structural changes required by nitrogenase in the process of catalyzing nitrogen reduction. Due to the limitations of nitrogenase in activity and stability, it faces many challenges in improving nitrogen fixation efficiency, applying to agriculture and industry, etc. Currently, the improvement of nitrogenase activity still depends on the in-depth understanding and research of its molecular mechanism, and how to optimize its performance through genetic engineering or other technical means to meet the demand for nitrogen source in agricultural production, and promote its wide application in environmental remediation, energy production, etc.
[0004] In recent years, the rapid development of artificial intelligence technology, especially machine learning, has provided new tools for the study of complex biological systems. Elnaggar A. et al. published pre-trained large models ProtT5 in IEEE transactions on pattern analysis and machine intelligence, which can extract high-level features from large-scale protein sequences and perform well in biological function prediction. However, there is still a lack of methods based on these new technologies to comprehensively predict and optimize nitrogenase activity. In addition, important factors such as the expression level of nitrogenase, codon bias, etc. that affect nitrogen fixation efficiency have not been fully integrated into the existing analysis framework.
[0005] Protein design is an exceptionally complex and challenging task, as the function of a protein depends on a complex structure that is formed by the amino acid sequence through highly nonlinear, dynamic interactions. Designing efficient nitrogen-fixing bacteria is particularly difficult. Nitrogenase is a highly complex and extremely oxygen-sensitive enzyme, whose activity is not only dependent on the specific amino acid composition, but is also significantly influenced by factors such as the expression level of nitrogenase, codon bias, and the like. In addition, nitrogenase requires a complex protection mechanism provided by the host cell, which makes it more cumbersome and unpredictable to engineer nitrogen-fixing bacteria to improve nitrogen fixation efficiency. Traditional experimental methods require a large amount of time and resources to screen and optimize the design, while our model significantly improves the design efficiency. Our method can quickly predict the activity of nitrogenase, identify key genes and amino acid features, and provide optimization directions combined with explanatory analysis.
[0006] Therefore, it is necessary to provide an innovative method based on machine learning, which comprehensively optimizes the prediction method of nitrogenase activity by combining the characteristics of nitrogenase sample protein sequences, the expression level of nitrogenase, and the codon bias, and provides a theoretical basis and practical guidance for the genetic engineering of nitrogen-fixing microorganisms. SUMMARY
[0007] The present application provides a machine learning-based nitrogenase activity prediction method and system, which can solve the technical problem that the expression level of nitrogenase, codon bias and other important factors affecting nitrogen fixation efficiency are not fully integrated into the existing nitrogenase activity prediction analysis framework.
[0008] In a first aspect, the present application provides a machine learning-based nitrogenase activity prediction method, comprising the following steps:
[0009] Obtaining nitrogenase sample activity data;
[0010] Extracting a fusion feature vector, an expression level of nitrogenase, a codon bias of a nitrogenase sample target strain, and a copy number of a nitrogenase sample nitrogen fixation-related gene from the nitrogenase sample;
[0011] Based on the fusion feature vector of the nitrogenase sample protein sequence, the expression level of the nitrogenase, the codon bias of the nitrogenase sample target strain, and the copy number of the nitrogenase sample nitrogen fixation-related gene, a nitrogenase activity prediction model is constructed based on machine learning;
[0012] Based on the constructed nitrogenase activity prediction model, the nitrogenase activity is predicted and obtained.
[0013] In combination with the first aspect, in an embodiment, the nitrogenase sample activity data is obtained, specifically comprising the following steps:
[0014] The activity value of the nitrogenase sample is measured by the acetylene reduction method;
[0015] According to the activity value of the nitrogenase sample obtained by measurement, the nitrogenase sample is distinguished as a positive sample and a negative sample.
[0016] In combination with the first aspect, in an implementation, the extracting the fusion feature vector, the expression level of the nitrogenase, the codon bias of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation related gene of the nitrogenase sample from the nitrogenase sample specifically includes the following steps:
[0017] obtaining a fusion feature vector of a protein sequence of the nitrogenase sample;
[0018] obtaining the expression level of the nitrogenase;
[0019] obtaining the codon bias of the target strain of the nitrogenase sample;
[0020] obtaining the copy number of the nitrogen fixation related gene of the nitrogenase sample.
[0021] In combination with the first aspect, in an implementation, the obtaining the fusion feature vector of the protein sequence of the nitrogenase sample specifically includes the following steps:
[0022] extracting a 1024-dimensional feature vector of the nifD, nifH, and nifK protein sequences from the nitrogenase sample by a pre-trained model ProtT5;
[0023] extracting a 343-dimensional feature vector of the nifD, nifH, and nifK protein sequences by calculating the triple amino acid encoding of the nitrogenase sample;
[0024] extracting a 400-dimensional feature vector of the nifD, nifH, and nifK protein sequences by calculating the dipeptide composition of the nitrogenase sample;
[0025] extracting a 50-dimensional feature vector of the nifD, nifH, and nifK protein sequences by calculating the pseudo-amino acid composition of the nitrogenase sample;
[0026] fusing the 1024-dimensional feature vector, the 343-dimensional feature vector, the 400-dimensional feature vector, and the 50-dimensional feature vector to obtain the fusion feature vector of the protein sequence of the nitrogenase sample.
[0027] In combination with the first aspect, in an implementation, the constructing the nitrogenase activity prediction model based on machine learning according to the fusion feature vector of the protein sequence of the nitrogenase sample, the expression level of the nitrogenase, the codon bias of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation related gene of the nitrogenase sample specifically includes the following steps:
[0028] dividing the obtained nitrogenase activity data and the fusion feature vector, the expression level of the nitrogenase, the gene copy number, and the codon bias nitrogen into a first training set and a first test set;
[0029] With the training set as the model input, an XGBoost algorithm is used to construct a classification model, and the classification model is used as the nitrogenase activity prediction model.
[0030] In combination with the first aspect, in an implementation, the nitrogenase activity prediction model is constructed based on machine learning according to the fusion feature vector of the nitrogenase sample protein sequence, the expression level of the nitrogenase, the codon bias of the nitrogenase sample target strain, and the copy number of the nitrogenase sample nitrogen fixation related gene, and specifically includes the following steps:
[0031] The obtained nitrogenase activity data and the fusion feature vector are divided into a first layer training set and a first layer test set;
[0032] The first layer training set is used as the input, and a support vector regression method is used to train to obtain a first layer model, and the prediction value of the first layer model is obtained;
[0033] The obtained nitrogenase activity data and the prediction value of the first layer model, the expression level of the nitrogenase, the gene copy number, and the codon bias are divided into a second layer training set and a second layer test set;
[0034] The second layer training set is used as the model input, and a stacking optimization method is used to obtain a second layer model, and the second layer model is used as the nitrogenase activity prediction model.
[0035] In combination with the first aspect, in an implementation, after the nitrogenase activity prediction model is constructed based on machine learning according to the fusion feature vector of the nitrogenase sample protein sequence, the expression level of the nitrogenase, the codon bias of the nitrogenase sample target strain, and the copy number of the nitrogenase sample nitrogen fixation related gene, the following steps are further included:
[0036] The obtained nitrogenase activity prediction model is verified for effectiveness based on the second layer test set, and an evaluation result of the model prediction performance is obtained.
[0037] In combination with the first aspect, in an implementation, after the nitrogenase activity prediction model is constructed based on machine learning according to the fusion feature vector of the nitrogenase sample protein sequence, the expression level of the nitrogenase, the codon bias of the nitrogenase sample target strain, and the copy number of the nitrogenase sample nitrogen fixation related gene, the following steps are further included:
[0038] The nitrogenase activity prediction model is analyzed based on the Python package SHAP, and the feature contribution value of the nif gene is calculated;
[0039] According to the feature contribution value of the nif gene, the smallest gene cluster is screened from the nitrogenase sample.
[0040] In the second aspect, the application provides a nitrogenase activity prediction system based on machine learning, which comprises:
[0041] a nitrogenase sample activity data acquisition module configured to acquire nitrogenase sample activity data;
[0042] a nitrogenase multi-dimensional feature acquisition module configured to extract a fusion feature vector, an expression level of nitrogenase, a codon bias of a target strain of the nitrogenase sample, and a copy number of nitrogenase-related genes of the nitrogenase sample from the nitrogenase sample;
[0043] a nitrogenase activity prediction model construction module in communication connection with the nitrogenase activity data and the nitrogenase multi-dimensional feature acquisition module, and configured to construct a nitrogenase activity prediction model based on machine learning according to the fusion feature vector of the protein sequence of the nitrogenase sample, the expression level of nitrogenase, the codon bias of the target strain of the nitrogenase sample, and the copy number of the nitrogenase-related genes of the nitrogenase sample;
[0044] a nitrogenase activity prediction model in communication connection with the nitrogenase activity prediction model construction module, and configured to predict nitrogenase activity based on the constructed nitrogenase activity prediction model.
[0045] In combination with the second aspect, in an implementation, the multi-dimensional feature acquisition module comprises:
[0046] a fusion feature vector acquisition unit configured to acquire a fusion feature vector of a protein sequence of the nitrogenase sample;
[0047] an expression level of nitrogenase acquisition unit configured to acquire an expression level of nitrogenase;
[0048] a strain codon bias acquisition unit configured to acquire a codon bias of a target strain of the nitrogenase sample;
[0049] a gene copy number acquisition unit configured to acquire a copy number of nitrogenase-related genes of the nitrogenase sample;
[0050] a feature vector fusion model in communication connection with the fusion feature vector acquisition unit, the expression level of nitrogenase acquisition unit, the strain codon bias acquisition unit, and the gene copy number acquisition unit, and configured to fuse a 1024-dimensional feature vector, a 343-dimensional feature vector, a 400-dimensional feature vector, and a 50-dimensional feature vector to acquire a fusion feature vector of a protein sequence of the nitrogenase sample.
[0051] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:
[0052] The application obtains activity data of nitrogenase samples, extracts a multi-dimensional feature including a fusion feature vector of a nitrogenase sample protein sequence, an expression level of the nitrogenase, a codon bias of a target strain of the nitrogenase sample, and a gene copy number related to nitrogen fixation, as model input, obtains a nitrogenase activity prediction model based on machine learning, and integrates important factors affecting nitrogenase efficiency into nitrogenase activity prediction analysis, thereby providing an efficient and accurate nitrogenase activity prediction method, which can provide support for the functional research of nitrogen fixation gene clusters. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 A flowchart of the nitrogenase activity prediction method based on machine learning provided by the embodiments of the application is shown.
[0054] Figure 2 A ROC curve diagram of the nitrogenase activity prediction model obtained by using the XGBoost machine learning algorithm provided by the embodiments of the application is shown.
[0055] Figure 3 A scatter plot of the prediction value of the nitrogenase activity prediction model obtained by using the stacked SVR machine learning algorithm provided by the embodiments of the application is shown.
[0056] Figure 4 An illustration of the influence of the gene copy number on the prediction result in the stacked SVR machine learning model by using SHAP analysis provided by the embodiments of the application is shown. DETAILED DESCRIPTION
[0057] In order to enable persons skilled in the art to better understand the schemes of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor fall within the scope of protection of the present application.
[0058] The terms "include" and "have" and any variations thereof in the specification and claims of the present application and the above-described drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include other steps or units not listed or can optionally include other steps or units inherent to the process, method, product or device. The terms "first", "second" and "third" and the like descriptions are used to distinguish different objects, and do not represent the order or limit the types of "first", "second" and "third".
[0059] In the description of the embodiments of the application, "exemplary", "for example", "for instance" or "like" are used to represent that the embodiments described are examples, illustrations or descriptions. Any embodiment or design solution described as "exemplary", "for example" or "for instance" in the embodiments of the application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. In fact, the use of "exemplary", "for example", "for instance" or the like is intended to present the relevant concept in a specific manner.
[0060] In the description of the embodiments of the application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text only represents a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the application, "multiple" means two or more than two.
[0061] In some processes described in the embodiments of the application, a plurality of operations or steps are included, which appear in a specific order, but it should be understood that these operations or steps can be executed or executed in parallel without the order in which they appear in the embodiments of the application. The serial number of the operation is only used to distinguish different operations, and the serial number itself does not represent any execution order. In addition, these processes can include more or fewer operations, and these operations or steps can be executed in sequence or in parallel, and these operations or steps can be combined.
[0062] In order to make the purpose, technical solutions and advantages of the application more clear, the embodiments of the application will be further described in detail below with reference to the drawings.
[0063] In a first aspect, referring to Figure 1 The embodiments of the application provide a method for predicting nitrogenase activity based on machine learning, comprising the following steps:
[0064] Step S1: Obtain nitrogenase sample activity data;
[0065] Step S2: Extract the fusion feature vector, the expression level of the nitrogenase, the codon preference of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation related gene of the nitrogenase sample from the nitrogenase sample;
[0066] Step S3: According to the fusion feature vector of the protein sequence of the nitrogenase sample, the expression level of the nitrogenase, the codon preference of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation related gene of the nitrogenase sample, a nitrogenase activity prediction model is constructed based on machine learning;
[0067] Step S5: Based on the constructed nitrogenase activity prediction model, the nitrogenase activity is predicted and obtained.
[0068] The application obtains activity data of nitrogenase samples, extracts a multi-dimensional feature including a fusion feature vector of a protein sequence of the nitrogenase sample, an expression level of the nitrogenase, a codon bias of a target strain of the nitrogenase sample, and a gene copy number related to nitrogen fixation, as model input, obtains a nitrogenase activity prediction model based on machine learning, and integrates important factors affecting nitrogenase efficiency into nitrogenase activity prediction analysis, thereby providing an efficient and accurate nitrogenase activity prediction method, which can provide support for the functional research of nitrogen fixation gene clusters.
[0069] In an embodiment, the obtaining step S1: obtaining activity data of nitrogenase samples, specifically includes the following steps:
[0070] Step S11: measuring the activity value of the nitrogenase sample by the acetylene reduction method, and standardizing the measured activity value of the nitrogenase sample, the unit of which is nmolC2H2 / mg protein / hour;
[0071] Step S12: distinguishing the nitrogenase sample as a positive sample and a negative sample according to the measured activity value of the nitrogenase sample, specifically, marking the activity of the nitrogenase sample higher than 50 nmolC2H2 / mg protein / hour as positive, and the activity of the nitrogenase sample lower than the value as negative.
[0072] In an embodiment, the step S2: extracting a fusion feature vector, an expression level of the nitrogenase, a codon bias of a target strain of the nitrogenase sample, and a copy number of a nitrogen fixation related gene of the nitrogenase sample from the nitrogenase sample, specifically includes the following steps:
[0073] Step S21: obtaining a fusion feature vector of a protein sequence of the nitrogenase sample;
[0074] Step S22: obtaining an expression level of the nitrogenase;
[0075] Step S23: obtaining a codon bias of a target strain of the nitrogenase sample;
[0076] Step S24: obtaining a copy number of a nitrogen fixation related gene of the nitrogenase sample.
[0077] In an embodiment, the step S21: obtaining a fusion feature vector of a protein sequence of the nitrogenase sample, specifically includes the following steps:
[0078] Step S211: extracting nifD, nifH, and nifK protein sequences from the nitrogenase sample using a pre-trained model ProtT5 (ProtT5-XL-UniRef50) downloaded from a website (https: / / github.com / agemagician / ProtTrans), generating a 1024-dimensional feature matrix for each, and taking an average value to generate a fixed-dimensional feature vector;
[0079] Step S212: Calculate the triad amino acid code (CT) of the nitrogenase sample using the "Conjoint Triad (CT) vector size=343" module from the website (https: / / bioinfo.usu.edu / profeatx / submit / ), extract the 343-dimensional feature vector of the nifD, nifH, nifK protein sequence;
[0080] Step S213: Calculate the dipeptide composition (DPC) using the "Dipeptide Composition (DPC), size=400" module from the website (https: / / bioinfo.usu.edu / profeatx / submit / ), extract the 400-dimensional feature vector of the nifD, nifH, nifK nitrogenase sequence;
[0081] Step S214: Calculate the pseudo-amino acid composition (PAAC) using the "Pseudo-Amino Acid Composition (PAAC), vector size=50" module from the website (https: / / bioinfo.usu.edu / profeatx / submit / ), extract the 50-dimensional feature vector of the nifD, nifH, nifK nitrogenase sequence;
[0082] Step S215: Fuse the 1024-dimensional feature vector, 343-dimensional feature vector, 400-dimensional feature vector and 50-dimensional feature vector to obtain the fusion feature vector of the protein sequence of the nitrogenase sample, which is implemented as follows:
[0083] Ensure that the obtained 1024-dimensional, 343-dimensional, 400-dimensional and 50-dimensional feature vectors are one-to-one corresponding at the sample level, that is, each nitrogenase sample protein sequence has a corresponding feature vector of these four different dimensions. Then these feature vectors are spliced and integrated in the dimension direction according to the order, for example, for a certain sequence, its 1024-dimensional feature vector, 343-dimensional feature vector, 400-dimensional feature vector and 50-dimensional feature vector are sequentially connected end to end to form a new fusion feature vector with a dimension of 1024+343+400+50=1817, and the fusion feature vector acquisition work of all nitrogenase sample protein sequences is completed, which is applied to the prediction task of sample enzyme activity.
[0084] In an embodiment, the step S22 of obtaining the expression level of the nitrogenase is implemented as follows:
[0085] Step S221: Calculate the E values and Fop values of the lnA, nifB, nifD, nifH, nifK, nifE, nifN, and nifX genes using the R package coRdon (https: / / bioconductor.org / packages / release / bioc / html / coRdon.html);
[0086] Step S222: Calculate the CAI values of the lnA, nifB, nifD, nifH, nifK, nifE, nifN, and nifX genes using the Python package CAI (https: / / cai.readthedocs.io / en / latest / );
[0087] The E values, Fop values, and CAI values obtained by calculation reflect the gene expression levels of the nitrogenase samples.
[0088] In an embodiment, the step S23: obtaining the codon bias of the target strain of the nitrogenase sample, is specifically implemented as:
[0089] Calculate the RSCU (Relative Synonymous Codon Usage) of each species genome using the Python package CAI (https: / / cai.readthedocs.io / en / latest / );
[0090] Calculate the ECD (Euclidean Distance) of each target strain genome relative to the non-codon-biased genome, and the calculation formula is as follows:
[0091]
[0092] Where n is the number of codon bias, x is the codon bias of different amino acids in the target strain, and n≥i≥1.
[0093] The codon bias of the target strain of the nitrogenase sample is reflected by RSCU and ECD.
[0094] In an embodiment, the step S24: obtaining the copy number of the nitrogen fixation related genes of the nitrogenase sample, is specifically implemented as:
[0095] The copy number of 34 genes related to nitrogen fixation process of nitrogenase (nifD, nifH, nifK, amtB, fixA, fixB, fixC, fixX, glnK, lnA, nifA, nifB, nifE, nifF, nifJ, nifL, nifM, nifN, nifP, nifQ, nifS, nifT, nifU, nifV, nifW, nifX, nifY, nifZ, rnfA, rnfB, rnfC, rnfD, rnfE, rnfG) is counted.
[0096] In an embodiment, the step S3 comprises the following steps:
[0097] Step S31A: the obtained nitrogenase activity data and the fusion feature vector, the expression level of nitrogenase, the gene copy number and the codon bias of nitrogenase are divided into a first training set and a first test set;
[0098] Step S32A: taking the training set as the model input, a classification model is constructed by using the XGBoost algorithm, and the classification model is taken as the nitrogenase activity prediction model.
[0099] In an embodiment, after the step S3, the method further comprises the following steps:
[0100] Step S4A: the prediction performance of the classification model is evaluated based on the first test set using five-fold cross-validation, the average AUC is 0.9365, and the F1 value is 0.85.
[0101] Through steps S31A-S32A and step S4A, the fusion feature vector of the protein sequence of the nitrogenase sample, the expression level of the nitrogenase, the codon bias of the target strain of the nitrogenase sample and the copy number of the nitrogenase related genes of the nitrogenase sample are taken as the model training set, the best classification model based on XGBoost realizes an average AUC of 0.9365 and an F1 score of 0.85 in five-fold cross-validation, an activity prediction model of nitrogenase with good prediction performance is obtained, and the ROC (Receiver Operating Characteristic) curve of the nitrogenase activity prediction model obtained by using the XGBoost machine learning algorithm is as shown in FIG. 1. Figure 2
[0102] In an embodiment, the step S3: based on the fusion feature vector of the nitrogenase sample protein sequence, the expression level of the nitrogenase, the codon bias of the nitrogenase sample target strain and the copy number of the nitrogenase sample nitrogen fixation related gene, a nitrogenase activity prediction model is constructed based on machine learning, specifically comprising the following steps:
[0103] Step S31B: the obtained nitrogenase activity data and the fusion feature vector are divided into a first layer training set and a first layer test set;
[0104] Step S32B: using the first layer training set as input, a first layer model is trained using a support vector regression method to obtain the prediction value of the first layer model;
[0105] Step S33B: the obtained nitrogenase activity data and the prediction value of the first layer model, the expression level of the nitrogenase, the gene copy number and the codon bias are divided into a second layer training set and a second layer test set;
[0106] Step S34B: using the second layer training set as model input, a second layer model (i.e. a stacked SVR machine learning model) is obtained by using a stacking optimization method, and the second layer model is used as a nitrogenase activity prediction model.
[0107] In an embodiment, after the step S3: based on the fusion feature vector of the nitrogenase sample protein sequence, the expression level of the nitrogenase, the codon bias of the nitrogenase sample target strain and the copy number of the nitrogenase sample nitrogen fixation related gene, a nitrogenase activity prediction model is constructed based on machine learning, the following steps are further included:
[0108] Step S4B: based on the second layer test set, the prediction performance of the second layer model is verified and evaluated by a five-fold cross-validation method, and the average R 2 is 0.5572 and the MAE is 0.2530; the prediction value scatter plot of the nitrogenase activity prediction model obtained by using the stacked SVR machine learning algorithm is as shown in Figure 3 .
[0109] Through steps S31B-S34B and step S4B, using the fusion feature vector of the nitrogenase sample protein sequence, the expression level of the nitrogenase, the codon bias of the nitrogenase sample target strain and the copy number of the nitrogenase sample nitrogen fixation related gene as the model training set, the regression model based on the stacking method of SVR which performs best is used as the nitrogenase activity prediction model, and the average R 2 is 0.5572 and the MAE is 0.2530, which achieves excellent results.
[0110] In an embodiment, after the step S5: based on the constructed nitrogenase activity prediction model, the nitrogenase activity is predicted, the following steps are further included:
[0111] Step S61: analyzing the nitrogenase activity prediction model based on the Python package SHAP (https: / / pypi.org / project / shap / ) to calculate the feature contribution value of the nif gene; using SHAP to analyze the influence of the copy number of each gene in the stacked SVR machine learning model on the prediction result as shown in Figure 4
[0112] Figure 4 Figure (a) is a relationship diagram of the copy number of nifA gene and SHAP value; Figure 4 Figure (b) is a relationship diagram of the copy number of nifB gene and SHAP value; Figure 4 Figure (c) is a relationship diagram of the copy number of nifD gene and SHAP value; Figure 4 Figure (d) is a relationship diagram of the copy number of nifE gene and SHAP value; Figure 4 Figure (e) is a relationship diagram of the copy number of nifH gene and SHAP value; Figure 4 Figure (f) is a relationship diagram of the copy number of nifK gene and SHAP value; Figure 4 Figure (g) is a relationship diagram of the copy number of nifN gene and SHAP value; Figure 4 Figure (h) is a relationship diagram of the copy number of nifQ gene and SHAP value; Figure 4 Figure (i) is a relationship diagram of the copy number of nifS gene and SHAP value; Figure 4 Figure (j) is a relationship diagram of the copy number of nifI gene and SHAP value; Figure 4 Figure (k) is a relationship diagram of the copy number of nifW gene and SHAP value; Figure 4 Figure (l) is a relationship diagram of the copy number of nifX gene and SHAP value; Figure 4 Figure (m) is a relationship diagram of the copy number of nifU gene and SHAP value; Figure 4 Figure (n) is a relationship diagram of the copy number of nifV gene and SHAP value; Figure 4 Figure (o) is a relationship diagram of the copy number of nifZ gene and SHAP value;
[0113] Step S62: screening the minimum gene cluster from the nitrogenase sample according to the feature contribution value of the nif gene; specifically, screening the nif gene with a feature contribution value higher than the feature contribution value threshold as the minimum gene cluster of the nitrogenase sample.
[0114] In a second aspect, the present application provides a machine learning-based nitrogenase activity prediction system, comprising:
[0115] a nitrogenase sample activity data acquisition module, configured to acquire nitrogenase sample activity data;
[0116] The nitrogenase multi-dimensional feature acquisition module is configured to extract a fusion feature vector, an expression level of the nitrogenase, a codon bias of the target strain of the nitrogenase sample, and a copy number of the nitrogen fixation-related gene of the nitrogenase sample from the nitrogenase sample.
[0117] The nitrogenase activity prediction model construction module is communicatively connected to the nitrogenase activity data and the nitrogenase multi-dimensional feature acquisition module, and is configured to construct a nitrogenase activity prediction model based on machine learning according to the fusion feature vector of the protein sequence of the nitrogenase sample, the expression level of the nitrogenase, the codon bias of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation-related gene of the nitrogenase sample.
[0118] The nitrogenase activity prediction model is communicatively connected to the nitrogenase activity prediction model construction module, and is configured to predict the nitrogenase activity based on the constructed nitrogenase activity prediction model.
[0119] In an embodiment, the multi-dimensional feature acquisition module includes:
[0120] The fusion feature vector acquisition unit is configured to acquire the fusion feature vector of the protein sequence of the nitrogenase sample.
[0121] The expression level acquisition unit is configured to acquire the expression level of the nitrogenase.
[0122] The strain codon bias acquisition unit is configured to acquire the codon bias of the target strain of the nitrogenase sample.
[0123] The gene copy number acquisition unit is configured to acquire the copy number of the nitrogen fixation-related gene of the nitrogenase sample.
[0124] The feature vector fusion model is communicatively connected to the fusion feature vector acquisition unit, the expression level acquisition unit, the strain codon bias acquisition unit, and the gene copy number acquisition unit, and is configured to fuse the 1024-dimensional feature vector, the 343-dimensional feature vector, the 400-dimensional feature vector, and the 50-dimensional feature vector to acquire the fusion feature vector of the protein sequence of the nitrogenase sample.
[0125] The functions of the modules in the above machine learning-based nitrogenase activity prediction system correspond to the steps in the above machine learning-based nitrogenase activity prediction method embodiment, and the functions and implementation processes will not be repeated here.
[0126] In a third aspect, the embodiments of the present application provide a machine learning-based nitrogenase activity prediction device. The machine learning-based nitrogenase activity prediction device can be a personal computer (PC), a notebook computer, a server, or other device with data processing function.
[0127] The communication interface includes an input / output (I / O) interface, a physical interface, and a logical interface, and the like, which are used to realize the interconnection of devices inside the machine learning based diazotase activity prediction device, and the interconnection of the machine learning based diazotase activity prediction device and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber interface, an ATM interface, and the like; and the user device can be a display, a keyboard, and the like.
[0128] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), and the like.
[0129] The processor can be a general-purpose processor, which can invoke the machine learning based diazotase activity prediction program stored in the memory and execute the machine learning based diazotase activity prediction method provided by the embodiments of the present application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed by the machine learning based diazotase activity prediction program when invoked can refer to various embodiments of the machine learning based diazotase activity prediction method of the present application, which will not be described here.
[0130] In a fourth aspect, the embodiments of the present application also provide a readable storage medium.
[0131] The readable storage medium of the present application stores a machine learning based diazotase activity prediction program, wherein the machine learning based diazotase activity prediction program is executed by the processor to implement the steps of the machine learning based diazotase activity prediction method as described above.
[0132] The method implemented by the machine learning based diazotase activity prediction program when executed can refer to various embodiments of the machine learning based diazotase activity prediction method of the present application, which will not be described here.
[0133] It should be noted that the above-mentioned sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0134] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a plurality of instructions for making a terminal device execute the method described in each embodiment of the present application.
[0135] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for predicting nitrogenase activity based on machine learning, characterized in that: The following steps are involved: Obtaining nitrogenase sample activity data, including: measuring the activity value of the nitrogenase sample by an acetylene reduction method; and distinguishing the nitrogenase sample as a positive sample or a negative sample based on the measured activity value of the nitrogenase sample; Extracting a fusion feature vector, an expression level of nitrogenase, a codon preference of a target strain of the nitrogenase sample, and a copy number of a nitrogen fixation-related gene of the nitrogenase sample from the nitrogenase sample, including: obtaining a fusion feature vector of a protein sequence of the nitrogenase sample; obtaining an expression level of nitrogenase; obtaining a codon preference of a target strain of the nitrogenase sample; and obtaining a copy number of a nitrogen fixation-related gene of the nitrogenase sample; The method for obtaining the fusion feature vector of the protein sequence of the nitrogenase sample comprises: extracting the fusion feature vector of the protein sequence of the nitrogenase sample by the pre-trained model ProtT5. nifD 、 nif 、 nifK 1024-dimensional feature vector of protein sequence; extracted by calculating the triplet amino acid code of nitrogenase sample nifD 、 nif 、 nifK 343-dimensional feature vector of protein sequence; extracted by calculating the dipeptide composition of nitrogenase sample nifD 、 nif 、 nifK 400-dimensional feature vector of protein sequence; extracted by calculating pseudo amino acid composition of nitrogenase sample nifD 、 nif 、 nifK A 50-dimensional feature vector of the protein sequence; fusing the 1024-dimensional feature vector, the 343-dimensional feature vector, the 400-dimensional feature vector, and the 50-dimensional feature vector to obtain a fused feature vector of the protein sequence of the nitrogenase sample; A machine learning-based prediction model for nitrogenase activity was constructed based on the fusion feature vector of the nitrogenase sample protein sequence, the expression level of nitrogenase, the codon preference of the target strain of the nitrogenase sample, and the copy number of nitrogen fixation-related genes in the nitrogenase sample. Based on the constructed nitrogenase activity prediction model, the nitrogenase activity is predicted.
2. The method for predicting nitrogenase activity based on machine learning according to claim 1, wherein: The method constructs a nitrogenase activity prediction model based on machine learning according to the fusion feature vector of the nitrogenase sample protein sequence, the expression level of nitrogenase, the codon preference of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation-related gene of the nitrogenase sample, and specifically includes the following steps: Divide the obtained nitrogenase activity data and the fusion feature vector, nitrogenase expression level, gene copy number and codon preference into the first training set and the first test set; The training set was used as the model input, and the XGBoost algorithm was used to build a classification model, which was used as the nitrogenase activity prediction model.
3. The method for predicting nitrogenase activity based on machine learning according to claim 1, wherein: The method constructs a nitrogenase activity prediction model based on machine learning according to the fusion feature vector of the nitrogenase sample protein sequence, the expression level of nitrogenase, the codon preference of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation-related gene of the nitrogenase sample, and specifically includes the following steps: Divide the obtained nitrogenase activity data and fusion feature vectors into the first-level training set and the first-level test set; Using the first-layer training set as input, the support vector regression method is used to train the first-layer model and obtain the predicted value of the first-layer model; The obtained nitrogenase activity data, the predicted values of the first-layer model, and the expression level, gene copy number, and codon preference of nitrogenase are divided into a second-layer training set and a second-layer test set; The second-layer training set was used as the model input, and the stacking optimization method was used to obtain the second-layer model, which was used as the nitrogenase activity prediction model.
4. The method for predicting nitrogenase activity based on machine learning according to claim 3, wherein: After constructing the nitrogenase activity prediction model based on machine learning according to the fusion feature vector of the nitrogenase sample protein sequence, the expression level of nitrogenase, the codon preference of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation-related gene of the nitrogenase sample, the following steps are further included: The obtained nitrogenase activity prediction model was validated based on the second-layer test set to obtain the evaluation results of the model prediction performance.
5. The method for predicting nitrogenase activity based on machine learning according to claim 1, wherein: After constructing the nitrogenase activity prediction model based on machine learning according to the fusion feature vector of the nitrogenase sample protein sequence, the expression level of nitrogenase, the codon preference of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation-related gene of the nitrogenase sample, the following steps are further included: The nitrogenase activity prediction model was analyzed based on the Python package SHAP, and the characteristic contribution value of the nif gene was calculated; Based on the characteristic contribution values of nif genes, the smallest gene cluster was screened from nitrogenase samples.
6. A nitrogenase activity prediction system based on machine learning, characterized in that: The nitrogenase activity prediction system is used to implement the method according to any one of claims 1 to 5, comprising: A nitrogenase sample activity data acquisition module is used to obtain nitrogenase sample activity data; The nitrogenase multidimensional feature acquisition module is used to extract fusion feature vectors, nitrogenase expression levels, codon preferences of target strains of nitrogenase samples, and copy numbers of nitrogen fixation-related genes from nitrogenase samples; a nitrogenase activity prediction model construction module, which is in communication with the nitrogenase activity data and the nitrogenase multidimensional feature acquisition module, and is used to construct a nitrogenase activity prediction model based on machine learning according to the fusion feature vector of the nitrogenase sample protein sequence, the expression level of nitrogenase, the codon preference of the target strain of the nitrogenase sample, and the copy number of the nitrogen fixation-related gene of the nitrogenase sample; The nitrogenase activity prediction model is communicatively connected to the nitrogenase activity prediction model construction module and is used to predict and obtain nitrogenase activity based on the constructed nitrogenase activity prediction model.
7. The nitrogenase activity prediction system based on machine learning according to claim 6, characterized in that: The multi-dimensional feature acquisition module includes: A fusion feature vector acquisition unit, used to acquire a fusion feature vector of the protein sequence of the nitrogenase sample; A nitrogenase expression level acquisition unit, used to acquire the expression level of nitrogenase; A strain codon preference acquisition unit is used to obtain the codon preference of the target strain of the nitrogenase sample; A gene copy number acquisition unit is used to obtain the copy number of nitrogen fixation-related genes in nitrogenase samples; The feature vector fusion model is communicatively connected to the fused feature vector acquisition unit, the nitrogenase expression level acquisition unit, the strain codon preference acquisition unit, and the gene copy number acquisition unit, and is used to fuse the 1024-dimensional feature vector, the 343-dimensional feature vector, the 400-dimensional feature vector, and the 50-dimensional feature vector to obtain a fused feature vector of the protein sequence of the nitrogenase sample.