Machine learning classification method based on functional gene sets of microbial virulence factors

By constructing a machine learning classification method based on the functional gene set of microbial virulence factors, the problem of insufficient utilization of microbial virulence factor information in existing technologies is solved, achieving efficient differentiation of disease states and improving the accuracy and robustness of the classification model.

CN120748514BActive Publication Date: 2025-11-07JILIN UNIV FIRST HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511232174.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-11-07
Estimated Expiration
2045-09-01

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively utilize microbial virulence factor information from metagenomic and metatranscriptomic data to differentiate between different disease states, resulting in insufficient accuracy in distinguishing between phenotypically similar diseases such as non-infectious respiratory syndromes and lower respiratory tract infections, particularly in differentiating pathogenic microorganisms.

Method used

We construct a machine learning classification method based on the functional gene set of microbial virulence factors. By extracting virulence factor-related information from metagenomic and metatranscriptomic data and combining it with the VFDB database, we use single-sample gene set enrichment analysis and XGBoost model for classification, which reduces the number of features and improves classification robustness.

Benefits of technology

The model improves the accuracy of disease state differentiation and demonstrates excellent classification performance on both the training and test sets, with an AUC greater than 0.90. The sensitivity, specificity, accuracy, precision, and recall are all ≥80%, and it supports the analysis of individual samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748514B_ABST
    Figure CN120748514B_ABST
Patent Text Reader

Abstract

The present application is suitable for the technical field of bioinformatics, and provides a machine learning classification method based on a microorganism virulence factor functional gene set. The method extracts virulence factor related information from metagenomic and metatranscriptomic data, and after quality control and quantitative matching, combines 14 virulence factor functional classes in the VFDB database, calculates the score of each sample on each functional class through single sample gene set enrichment analysis, forms a feature sample matrix, and then uses the XGBoost model for classification. The method considers the important role of virulence factors, reduces the number of features through the functional class scoring method, improves the classification robustness and generalization, and supports single sample analysis. The model verification result shows that it has better classification performance on the training set and the test set, with AUC greater than 0.90, and sensitivity, specificity, accuracy, precision and recall rate all greater than or equal to 80%, providing an effective means for classification based on microorganism features.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of bioinformatics, and particularly relates to a machine learning classification method based on a microorganism virulence factor functional gene set. BACKGROUND

[0002] Microbial communities will change in group composition under different environments (such as polluted and clean rivers) and different organs or disease states of the human body, which makes them potential indicators for distinguishing different types of environments or disease states. Currently, in clinical practice, for some phenotypically similar diseases, such as non-infectious respiratory syndrome and lower respiratory tract infections (LRTIs), it is difficult to distinguish them because both may have similar symptoms such as fever, which may lead to the fact that infected patients do not receive anti-infection treatment in time, or that anti-infection treatment is performed on non-infectious respiratory syndrome patients. In terms of microbial detection, most clinical medical institutions still mainly use methods based on microbial culture, but this method has obvious defects such as low sensitivity and long time consumption; serological detection and polymerase chain reaction (PCR) detection are also used to some extent, but they have limited types of detectable microorganisms, which cannot meet the needs of clinical diagnosis. In some large hospitals, metagenomic detection technology based on deep sequencing has been used to assist in judging infectious pathogens, however, metagenomic sequencing usually detects multiple microbial species, including non-pathogenic or opportunistic pathogenic colonizing bacteria, which greatly interferes with the accurate distinction of infection-related characteristics, making it difficult to directly extract effective features related to pathogenicity. Due to the strong heterogeneity of microbial composition, which is affected by many factors such as diet, it makes it difficult to distinguish different groups based on heterogeneous data. When distinguishing different groups based on metagenomic sequencing, it is difficult to find different groups of shared differential microorganisms, resulting in low efficiency of distinction. In addition, some microorganisms are conditionally pathogenic microorganisms, and even if the presence of the microorganism is detected, it cannot be directly determined that it is harmful, further increasing the difficulty of identification.

[0003] Microbial pathogenicity relies on a variety of virulence factors, which help microorganisms evade the immune system and invade host tissues. For example, the capsule is a crucial virulence factor for Streptococcus pneumoniae, preventing immune cells (such as macrophages and neutrophils) from phagocytosing the bacteria; different serotypes have different capsule structures, leading to variations in pathogenicity. Ply, a cholesterol-binding toxin, can perforate host cell membranes, causing cell death, and also promotes inflammatory responses, exacerbating tissue damage (such as lung infections and meningitis), and helps bacteria evade immune attacks by hijacking macrophage MRC1 receptors. Notably, different bacterial species may cause disease through the same or similar pathogenic mechanisms. For instance, both Streptococcus pneumoniae and Haemophilus influenzaetype b-Hib produce capsular polysaccharides, suggesting that even if different patients are infected with different bacteria, they may share common virulence factors. Based on these shared virulence factors, it is hoped that the ability to differentiate diseases can be improved. However, while metagenomic and metatranscriptomic sequencing data contain information on these virulence factors, they have not yet been effectively applied to classification. Furthermore, microbial genomes are highly complex, and different microorganisms may achieve pathogenicity by perturbing different genes within the same functional class. If virulence factor genes could be categorized into a limited set of functional classes and scored, not only could the number of features be reduced, but the robustness of detection could also be improved, thereby enhancing the generalizability of the model. Therefore, this invention proposes a machine learning classification method based on a set of functional genes representing microbial virulence factors. Summary of the Invention

[0004] The purpose of this invention is to provide a machine learning classification method based on the functional gene set of microbial virulence factors, aiming to solve the problems mentioned in the background art.

[0005] The objective of this invention is achieved through the following technical solution:

[0006] A machine learning classification method based on the functional gene set of microbial virulence factors includes the following steps:

[0007] Step 1: Extract microbial sequences and quantify virulence factors from metagenomic and metatranscriptomic data;

[0008] Microbial sequences were obtained from metagenomic and metatranscriptomic data; virulence factor sequences and functional annotations were obtained from a virulence factor database; quality-controlled metagenomic and metatranscriptomic reads were compared with the core dataset sequences of the virulence factor database to obtain read counts mapped to each virulence factor gene; virulence factors at the DNA and RNA levels were standardized and quantified, and the standardized quantitative values ​​of the two were added together to form a third-level feature.

[0009] Step 2: Single-sample enrichment analysis method is used to calculate the virulence factor functional class score of each sample;

[0010] Based on the quantitative results of three levels, the virulence factor functional class in the virulence factor database is taken as the functional gene set, and the single-sample gene set enrichment analysis method is used to calculate the score of each sample on each functional class to construct a feature sample matrix;

[0011] Step 3: Data set division and identification of significantly different functional classes;

[0012] The feature sample matrix is randomly divided into a training set and a test set according to a proportion; and the virulence factor functional class with significant difference is identified in the training set;

[0013] Step 4: XGBoost machine learning model construction;

[0014] The virulence factor functional class with significant difference is taken as a feature, and the XGBoost model is trained on the training set, and the optimal hyperparameter combination is obtained through cross-validation;

[0015] Step 5: Model verification and performance evaluation;

[0016] The trained model is verified on the test set, and the model performance is evaluated by using a preset index.

[0017] Further, the virulence factor database is VFDB database, and the virulence factors contained therein are summarized into 14 virulence factor functional classes.

[0018] Further, the standardized quantitative processing adopts log2(CPM+1), wherein the calculation formula of CPM is as follows:

[0019] ;

[0020] Wherein, is the number of reads per million for gene ; is the target virulence factor gene; is the number of sequencing reads mapped to gene ; Total counts is the total number of reads mapped to the microbial genome; 10 6 is a normalization coefficient, which scales the value to the "number of reads per million" scale.

[0021] Further, the single-sample gene set enrichment analysis is ssGSEA, comprising the following steps: sorting all genes of a sample in descending order of expression; constructing a weighted cumulative distribution function of genes in a gene set S and a cumulative distribution function of genes not in the gene set S for a target gene set S; and calculating the difference between the two cumulative distribution functions as an enrichment score.

[0022] Further, the calculation formula of the weighted cumulative distribution function of genes in the gene set S is:

[0023] ;

[0024] wherein, denotes the weighted cumulative distribution function of genes in the gene set S; is the jth gene arranged according to a pre-arrangement rule; is the expression of the gene ; and is the contribution of only the first i genes in the sorting; and alpha is a weight index; if alpha = 0, only the presence or absence of the gene is considered, and the expression is not considered.

[0025] The calculation formula of the cumulative distribution function of genes not in the gene set S is:

[0026] ;

[0027] wherein, denotes the cumulative distribution function of genes not in the gene set S; denotes a gene not belonging to the target gene set S; N is the total number of genes; and |S| is the total number of genes included in the target gene set S; denotes the count of genes not in the gene set S among the first i genes.

[0028] Further, the data set division ratio is 7:3, i.e., the training set accounts for 70%, and the test set accounts for 30%.

[0029] Further, the XGBoost model is determined by performing N times of 5-fold cross-validation to determine the optimal hyperparameter combination.

[0030] Further, the preset index includes AUC, sensitivity, specificity, accuracy, precision and recall rate.

[0031] Compared with the prior art, the present application has the following beneficial effects:

[0032] The application constructs a machine learning classification framework based on the scoring of the functional gene set of virulence factors in the metagenome and metatranscriptome. The framework extracts virulence factor related information from the metagenome and metatranscriptome data, and after quality control and matching quantification, combines the 14 virulence factor functional classes in the VFDB database, uses single sample gene set enrichment analysis to calculate the score of each sample in these functional classes, forms a feature sample matrix, and then uses the XGBoost model for classification. This framework uses the functional gene set of virulence factors in the metagenome and metatranscriptome as classification features, taking into account the important role of virulence factors in microbial related processes, while reducing the number of features through the functional class scoring method, improving the robustness and generalization of classification, and supporting the analysis of individual samples. The model verification results show that it performs better classification performance on the training set and the test set, with AUC greater than 0.90, and sensitivity, specificity, accuracy, precision and recall rate indicators are all ≥80%, providing an effective means for classifying based on microbial features. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 A flowchart of the method of the application.

[0034] Figure 2 The performance of classification based on the virulence factor feature gene set; wherein A is the AUC in the training set; B is the AUC in the test set; C is the sensitivity, specificity, accuracy, precision and recall rate. DETAILED DESCRIPTION

[0035] In order to have a clearer understanding of the technical features, objectives and beneficial effects of the application, the technical solutions of the application will be described in detail below, but it should not be understood as limiting the scope of the application.

[0036] The specific implementation of the application will be described in detail below in combination with specific embodiments.

[0037] One embodiment of the application provides a machine learning classification method based on the functional gene set of microbial virulence factors, and a flowchart thereof is shown as Figure 1 The method comprises the following steps:

[0038] Step 1: Extracting microbial sequences and virulence factor quantification from metagenome and metatranscriptome data;

[0039] This step aims to extract virulence factor related data from metagenome and metatranscriptome data, and the specific process is as follows:

[0040] Sample and virulence factor data acquisition: Sample information (microbial sequences) including 26 samples in group A and 18 samples in group B were extracted from the metagenomic and metatranscriptomic data of published materials (Integrating host response and unbiased microbe detection for lower respiratory tract infection diagnosis in critically ill adults, Proc Natl Acad Sci U S A, 2018 Dec 26; 115(52): E12353-E12362. doi: 10.1073 / pnas.1809700115. Epub 2018 Nov 27), and the virulence factor sequences and functional annotations of virulence factor genes were obtained based on the virulence factor database (VFDB, which contains 4368 virulence factors and has been summarized into 14 virulence factor functional classes (see Table 1). https: / / www.mgc.ac.cn / VFs / download.htm

[0041] Table 1 Virulence factor functional classes

[0042]

[0043] Quality control: For the original sequencing data, the reads containing adapters, indexes, and bases with an average quality score lower than 30 were removed to obtain high-quality cleaned data.

[0044] Metagenomic and metatranscriptomic data matching and output: The quality-controlled metagenomic reads were matched with DNA sequences from the VFDB core dataset using the BBMap v 39.06 tool (Bushnell, 2014, using default parameters). After the matching was completed, the output results contained the read count (the number of sequencing reads mapped to the target virulence factor gene) and the total read count of each virulence factor for each sample.

[0045] Virulence factor quantitative calculation:

[0046] The DNA (metagenomic) and RNA (metatranscriptomic) levels of virulence factors were quantified, and the present application adopted log2(CPM+1) standardization processing. CPM stands for Counts Per Million, that is, how many reads are mapped to a certain gene in every million sequencing reads.

[0047] ;

[0048] wherein, is the gene​ Readings per million; Target virulence factor genes (such as the capsular polysaccharide-encoding gene ply). To map to genes The number of sequencing reads; Total counts is the total number of reads mapped to the microbial genome; 10 6 To standardize the coefficients, the values ​​are scaled to the "per million readings" scale.

[0049] Furthermore, the standardized quantitative values ​​at the DNA and RNA levels are added together to form a third level of characteristics (DNA + RNA).

[0050] Step 2: Calculate the functional class score of virulence factors for each sample using single-sample enrichment analysis;

[0051] Based on the quantitative results of virulence factors at the three levels (DNA, RNA, and DNA+RNA), 14 functional classes of virulence factors in the VFDB database were used as functional gene sets. Single-sample gene set enrichment analysis (ssGSEA) was used to calculate the score of each sample on the 14 functional classes of virulence factors, and finally 42 gene set scores (14 functional classes of virulence factors × 3 levels) were obtained, which were used as feature sample matrices, i.e., input features of the machine learning model.

[0052] ssGSEA is a nonparametric, ordination-based method for calculating the enrichment score of a single sample on a specific gene set. It is implemented using the R package GSVA (version 1.50.5). Its core idea is similar to the classic GSEA, but while GSEA compares samples across groups, ssGSEA extends the analysis to a single sample. For a sample's gene expression vector E=(e1,e2,…,eN) and a given target gene set S (functional class), the steps of ssGSEA are as follows:

[0053] 1. Sequencing gene expression:

[0054] Sort all genes in the sample according to their expression levels from highest to lowest to obtain a sorted gene list G=(g1,g2,…,gN).

[0055] 2. Calculate the weighted empirical cumulative distribution:

[0056] ssGSEA constructs two cumulative distribution functions for each target gene set S:

[0057] Cumulative weight distribution of genes in the target gene set S:

[0058] ;

[0059] wherein, is the weighted cumulative distribution function of the gene in the gene set S, which is the weighted cumulative proportion of the effect value of the target gene set S gene in the first i genes; is the jth gene sorted by expression from large to small; is the expression of the gene ; is the contribution of only the first i genes in the sorting; α is the weight index, which is used to regulate the contribution of gene expression to the enrichment score; when α=0, only whether the gene belongs to the target gene set S (i.e. the existence of the gene) is considered, without distinguishing the high and low of the gene expression; when α=1, it represents linear weighting; when >1, the contribution of strong effect genes is amplified). In the present application, the default parameter α=0.25 is used for calculation.

[0060] Cumulative weighted distribution of genes not in the target gene set S:

[0061] ;

[0062] wherein, is the cumulative distribution function of the gene not in the gene set S; is the gene not belonging to the target gene set S; N is the total number of genes; |S| is the total number of genes contained in the target gene set S; is the count of non-gene set S genes in the first i genes.

[0063] 3. Calculate the enrichment score (ES):

[0064] The enrichment score of ssGSEA is the gap between the two cumulative distribution curves:

[0065] ;

[0066] wherein, ES is the enrichment score.

[0067] Step 3: Data set division and identification of significantly different functional classes;

[0068] The feature sample matrix (row is 42 gene set scores, column is 44 samples) constructed in step 2 is randomly divided into training set and test set according to the ratio of 7:3. In the training set, the differentially expressed gene set is identified, and the virulence factor functional class with significant difference is screened out.

[0069] Step 4: XGboost machine learning model construction;

[0070] The XGBoost model was trained in the training set with the function class of the virulence factor with significant difference as the characteristic, and the optimal super parameter combination was obtained by performing N times 5-fold cross-validation. XGBoost, which stands for eXtreme Gradient Boosting, is an optimized distributed gradient boosting library, which is a high-efficiency implementation of the gradient boosting algorithm. Its core is to combine multiple weak classifiers (decision trees) into a powerful classifier. Each new weak classifier is committed to correcting the errors of all previous weak classifiers, and through continuous iteration, the accuracy of the model is gradually improved. The XGBoost model can output the relative importance of the variables included, and has interpretability.

[0071] Step 5: Model verification and performance evaluation

[0072] The trained XGBoost model was verified on the test set, and the evaluation indexes included AUC (area under the curve), sensitivity, specificity, accuracy, precision and recall rate.

[0073] The results are shown in Figure 2 The model showed good classification performance on both the training set and the test set, with AUC greater than 0.90 (see Figure 2 A and B), and the sensitivity, specificity, accuracy, precision and recall rate were all ≥80% (see Figure 2 C).

[0074] The above is only the preferred embodiment of the present application, it should be pointed out that for those skilled in the art, without departing from the concept of the present application, can make a number of deformation and improvement, these should be regarded as the protection scope of the present application, these will not affect the effect and practicality of the patent of the present application.

Claims

1. Machine learning classification method based on functional gene sets of microbial virulence factors, characterized in that, The method comprises the following steps: Step 1: Extracting microbial sequences and quantifying virulence factors in metagenomic and metatranscriptomic data; Obtaining microbial sequences from metagenomic and metatranscriptomic data; obtaining virulence factor sequences and functional annotations based on a virulence factor database; aligning the quality-controlled metagenomic and metatranscriptomic reads with the core dataset sequences of the virulence factor database to obtain read counts mapped to each virulence factor gene; standardizing and quantifying the virulence factors at the DNA and RNA levels, and adding the standardized quantification values of the two to form the third level of features; Step 2: Calculating the virulence factor functional class scores of each sample by single-sample enrichment analysis method; Based on the quantification results of the three levels, taking the virulence factor functional classes in the virulence factor database as the functional gene set, the scores of each sample in each functional class are calculated by the single-sample gene set enrichment analysis method to construct a feature sample matrix; Step 3: Data set division and identification of significantly different functional classes; Randomly dividing the feature sample matrix into a training set and a test set according to a proportion; identifying significantly different virulence factor functional classes in the training set; Step 4: XGBoost machine learning model construction; Taking the virulence factor functional classes with significant differences as features, training the XGBoost model on the training set, and obtaining the optimal hyperparameter combination through cross-validation; Step 5: Model verification and performance evaluation; Verifying the trained model on the test set and evaluating the model performance by using preset indicators.

2. The machine learning classification method based on the functional gene sets of microbial virulence factors according to claim 1, characterized in that, The virulence factor database is the VFDB database, which contains 14 virulence factor functional classes.

3. The machine learning classification method based on the functional gene sets of microbial virulence factors according to claim 1, characterized in that, The standardized quantification processing adopts log2(CPM+1), wherein the calculation formula of CPM is as follows: ; wherein, is the number of reads per million reads for a gene ; is the target virulence factor gene; is the number of sequencing reads mapped to a gene ; Total counts is the total number of reads mapped to the microbial genome; 10 6 is the normalization coefficient to scale the values to the "per million reads" scale.

4. The machine learning classification method based on the functional gene sets of microbial virulence factors according to claim 1, characterized in that, The single-sample gene set enrichment analysis is ssGSEA, which comprises the following steps: sorting all genes of the sample in descending order of expression; constructing a weighted cumulative distribution function of genes in gene set S and a cumulative distribution function of genes not in gene set S for target gene set S; calculating the difference between the two cumulative distribution functions as the enrichment score.

5. The machine learning classification method based on the functional gene sets of microbial virulence factors according to claim 4, characterized in that, The calculation formula of the weighted cumulative distribution function of genes in gene set S is as follows: ; wherein, represents the weighted cumulative distribution function of a gene in the gene set S; is the jth gene arranged according to the preordering rule; is the expression level of the gene ; and is the contribution of only the first i genes in the ranking; a is the weight index; if a = 0, only the presence or absence of the gene is considered, not the expression level. The calculation formula of the cumulative distribution function of genes not in gene set S is as follows: ; wherein, denotes the cumulative distribution function of genes not in the gene set S; denotes genes not belonging to the target gene set S; N is the total number of genes; |S| is the total number of genes contained in the target gene set S; denotes the count of genes not in the gene set S among the first i genes.

6. The machine learning classification method based on the functional gene sets of microbial virulence factors according to claim 1, characterized in that, The data set division ratio is 7:3, that is, the training set accounts for 70% and the test set accounts for 30%.

7. The machine learning classification method based on the functional gene sets of microbial virulence factors according to claim 1, characterized in that, The XGBoost model determines the optimal hyperparameter combination by performing N times of 5-fold cross-validation.

8. The machine learning classification method based on the functional gene sets of microbial virulence factors according to claim 1, characterized in that, The preset indicators include AUC, sensitivity, specificity, accuracy, precision, and recall rate.

Citation Information

Patent Citations

  • Methods for detecting and treating low-virulence infections

    CN108138178A

  • Toxin gene abundance detection method based on metagenomics and annotation database construction method

    CN114621997A