Method for screening of biomarkers associated with respiratory tract infections based on macro-transcriptomics

By screening respiratory infection-related biomarkers using metatranscriptomics, combined with differential expression and co-expression analysis and machine learning algorithms, the problem of inaccurate pathogen identification in traditional diagnostic methods has been solved, enabling rapid and accurate identification of respiratory infections and efficient assessment of severe conditions.

CN120738336BActive Publication Date: 2026-07-21中国人民解放军总医院第八医学中心
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
中国人民解放军总医院第八医学中心
Filing Date
2025-05-21
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional diagnostic methods cannot accurately identify pathogens in respiratory infections, especially lower respiratory tract infections. They are slow to detect and have low sensitivity, and lack rapid and accurate biomarker screening methods.

Method used

Metatranscriptomics was used to screen for respiratory infection-related biomarkers. By combining metatranscriptomics sequencing, quality control, differential expression analysis, and co-expression analysis with machine learning algorithms, biomarkers such as IL6, CXCL8, TNF-α, HSP90AA1, and S100A9 were screened for the identification of respiratory infections and severe conditions.

Benefits of technology

It improved the accuracy and efficiency of respiratory infection diagnosis. The combination of biomarkers showed good clinical diagnostic performance, especially in differentiating severe cases with a diagnostic accuracy of 81.07%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120738336B_ABST
    Figure CN120738336B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of biological detection, and discloses a screening method of respiratory tract infection related biomarkers based on macro-transcriptomics. The application performs macro-transcriptome sequencing on respiratory tract infection samples with different clinical phenotypes, performs data quality control, alignment, transcript quantification, retains pathogen and host information, and then identifies genes stably expressed or significantly changed in different groups by combining differential expression analysis and co-expression analysis, obtains potential biomarkers, and obtains the biomarkers by taking the intersection genes of three machine learning algorithms of LASSO algorithm, random forest model and SVM model. The application provides a screening method of biomarkers for rapid and accurate identification of respiratory tract infection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biological detection technology and relates to a method for screening respiratory infection-related biomarkers based on metatranscriptomics. Background Technology

[0002] Respiratory tract infections (RTIs) are diseases caused by the invasion of the respiratory tract by pathogens, leading to an inflammatory response. They are the most common cause of human illness, frequently resulting in hospitalization and death. Individuals with weakened immune systems (weakened heart, lungs, and immune systems), the elderly, and infants are at risk of developing serious complications. Lower respiratory tract infections (LRTIs), in particular, cause more than 3 million deaths annually, and their diagnosis presents numerous challenges. Traditional diagnostic methods can only isolate a small number of pathogens, hindering further research on the large number of microorganisms that cannot be isolated and cultured. Furthermore, they suffer from slow detection speed and low sensitivity, leading to inaccurate pathogen identification in many cases. Conventional microbiological methods identify pathogens only in 30-40% of LRTI cases.

[0003] Metatranscriptomics is a research method that directly extracts all RNA from a microbial community in a specific environment to analyze its gene expression dynamics and regulatory mechanisms. It can comprehensively reveal the composition and function of the microbial community in the environment, as well as the host's immune response. Metatranscriptomics can not only detect the transcriptional activity of pathogens but also analyze the host's response to infection, thus providing more comprehensive information for diagnosis and treatment. Compared with traditional diagnostic methods, metatranscriptomics can provide more comprehensive pathogen detection and host response analysis, helping to improve the accuracy and efficiency of diagnosis.

[0004] Biomarkers are characteristics that can be objectively measured and assessed, reflecting physiological or pathological processes and their biological effects on exposure or treatment interventions. They are used to indicate normal biological processes, pathological processes, or responses to exposure or intervention. Exploring biomarkers for respiratory infections based on metatranscriptomics is of great significance for the rapid and accurate identification of respiratory infections. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a method for screening respiratory infection-related biomarkers based on metagenomics. Metatranscriptomic sequencing is performed on respiratory infection samples with different clinical phenotypes. Data quality control, alignment, and transcript quantification are conducted to obtain host transcriptomic information. Subsequently, differential expression analysis and co-expression analysis are combined to identify genes that are stably expressed or significantly altered in different groups, yielding potential biomarkers. The biomarkers are obtained by taking the intersection genes from three machine learning algorithms: LASSO, Random Forest, and SVM. This invention provides a rapid and accurate method for screening biomarkers for the identification of respiratory infections.

[0006] To achieve the technical objective of this invention, on one hand, this invention provides a method for screening respiratory infection-related biomarkers based on metagenomics, specifically including the following steps: Obtain samples from different groups and extract total RNA samples from the samples; Metatranscriptome sequencing was performed to obtain metatranscriptome data. The metatranscriptome data were quality controlled and compared, transcripts were quantified, and differential expression and co-expression analyses were combined to screen out differentially expressed and co-expressed genes in different groups to obtain potential biomarkers. The differentially expressed genes and co-expressed genes are selected using a machine learning algorithm, and the intersection genes are then used to obtain biomarkers. The diagnostic value of the biomarkers was verified.

[0007] Furthermore, the types of samples in this invention include bronchoalveolar lavage fluid and blood.

[0008] Furthermore, this invention divides patients with respiratory tract infections into an infected group and a non-infected group. Patients in the infected group meet any one of the following criteria: CRP ≥ 10 g / L, PCT ≥ 0.5 ng / mL, positive serological test (Streptococcus pneumoniae or Haemophilus influenzae), positive clinical culture result (analytical sample cultured to produce fungal or bacterial pathogens), G test result > 80 pg / mL, or positive mNGS result. Patients in the non-infected group have CRP < 10 g / L and PCT < 0.5 ng / mL, negative serological test result, negative clinical culture result, G test result ≤ 80 pg / mL, and negative mNGS result.

[0009] Furthermore, the present invention further divides the infected group into a severe pneumonia group and a mild pneumonia group. Patients in the severe pneumonia group meet any one of the following criteria: inflammatory marker CRP ≥ 30 g / L, inflammatory marker PCT ≥ 2 ng / mL, and G test result ≥ 200 pg / mL. Patients in the mild pneumonia group meet the following criteria: inflammatory marker 10 g / L ≤ CRP < 30 g / L, inflammatory marker 0.5 ng / mL ≤ PCT < 2 ng / mL, and 80 pg / mL < G test result < 200 pg / mL.

[0010] Furthermore, this invention collects bronchoalveolar lavage fluid and blood from patients with respiratory infections in different groups as analytical samples, and extracts total RNA samples from the analytical samples.

[0011] Furthermore, adapter sequences and low-quality bases were removed from the total RNA sample, rRNA was removed, and then the RNA was reverse transcribed into cDNA. The ends were repaired, tails were added, and sequencing adapters were ligated. The cDNA was then amplified by PCR to construct a cDNA library, and metagenomic sequencing was performed to obtain metagenomic data.

[0012] Furthermore, the metatranscriptome data underwent quality control and filtering using FastQC and Fastp, with both controls and filters set to default parameters. RNA quality was assessed using the RIN value, requiring RIN ≥ 7.0. Bowtie2 was used to map the quality-controlled reads data to the human reference genome GRCh38, generating a BAM file containing location information.

[0013] Furthermore, based on the BAM file, Cufflink was used to calculate the original read count for each gene to quantify transcripts.

[0014] Furthermore, the differential expression analysis compares the genes of the infected group and the non-infected group, as well as the severe pneumonia group and the mild pneumonia group, sets a threshold, and screens out the differentially expressed genes; The threshold is set to |log2FC|≥1 and FDR≤0.05.

[0015] Furthermore, the co-expression analysis includes constructing a weighted gene co-expression network, setting a soft threshold, and calculating the R value, where R is the Person correlation coefficient. 2 When the threshold is >0.8 and the average connectivity is high, the soft threshold is set to 7. The co-expression analysis also includes calculating the correlation between the module feature genes and the symptom phenotype, screening significantly correlated modules, and extracting the top 10% of genes with the highest connectivity within the modules from the significantly correlated modules to obtain the co-expressed genes; the symptom phenotype is an infection state or a severe state.

[0016] Specifically, the infection status is divided into infected and uninfected, which is determined by clinical culture results and mNGS results. If the clinical culture result or mNGS result is positive, it is determined to be infected; if both the clinical culture result and mNGS result are negative, it is determined to be uninfected.

[0017] Specifically, the severity of illness is divided into severe and mild cases, distinguished by clinical indicators, namely the SOFA score and the APACHE II score. A score of 2 < SOFA score ≤ 5 and a score of 15 < APACHE II score ≤ 24 is considered mild, while a score of SOFA score > 5 and an APACHE II score > 24 is considered severe.

[0018] Furthermore, the machine learning algorithms are LASSO algorithm, random forest model and SVM model.

[0019] Furthermore, the validation includes establishing ROC curves for biomarkers, with an AUC > 0.8 indicating diagnostic value.

[0020] Compared with the prior art, the technical solution provided by the present invention has at least the following beneficial effects or advantages: (1) This invention performs metagenomic sequencing on respiratory infection samples with different clinical phenotypes, removes non-informative sequences such as host RNA through quality control, retains pathogen and host information, and then identifies genes that are stably expressed or significantly changed in different groups by combining differential expression analysis and co-expression analysis to obtain potential biomarkers. The biomarkers are obtained by taking the intersection genes of three machine learning algorithms: LASSO algorithm, random forest model, and SVM model. This invention provides a method for screening biomarkers for rapid and accurate identification of respiratory infections.

[0021] (2) Based on the screening method for respiratory infection-related biomarkers provided by this invention, the biomarkers related to respiratory infection were screened as IL6, CXCL8, TNF-α, HSP90AA1, and S100A9. The diagnostic performance of the biomarkers was further verified by analyzing the ROC curves. The results showed that the area under the curve of the biomarker combination was 0.847, indicating good clinical diagnostic performance. Furthermore, the diagnostic accuracy of the above biomarkers reached 80.14%, demonstrating high discriminative ability in the diagnosis of respiratory infections. This invention uses the above screening method to screen biomarkers related to whether a disease is severe. The obtained biomarkers are IL6, CXCL8, and TNF-α. The diagnostic performance of the biomarkers was further verified by analyzing the ROC curves. The results showed that the area under the curve of the biomarker combination was 0.862, indicating good clinical diagnostic performance. Furthermore, the diagnostic accuracy of the above biomarkers reached 81.07%, demonstrating high discriminative ability in the diagnosis of whether a respiratory infection is severe. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention.

[0023] Figure 1 This is a flowchart of a method for screening respiratory infection-related biomarkers based on metatranscriptomics. Detailed Implementation

[0024] The technical solution of the present invention will now be described with reference to the embodiments. However, the present invention is not limited to the following embodiments. Unless otherwise specified, the experimental methods and detection methods described in the following embodiments are conventional methods; unless otherwise specified, the reagents and materials are commercially available.

[0025] Example 1 This embodiment provides a method for screening respiratory infection-related biomarkers based on metagenomics, as shown in the flowchart below. Figure 1 As shown, the screening of biomarkers for identifying infection status specifically includes the following steps: 1. Sample Data Acquisition: Biological samples, including bronchoalveolar lavage fluid and blood, were collected from patients with respiratory tract infections for analysis. Total RNA samples were extracted from the analytical samples using a tissue cell RNA extraction kit (Trizol extraction method) (purchased from Shanghai Shangbao Biotechnology Co., Ltd.), and the extracted RNA samples were stored at -80℃. 2. Library Construction: Cutadapt was used to remove adapter sequences and low-quality bases. The Ribo-Zero™ rRNARemoval Kit (Gram-Negative Bacteria) 1 Kit (purchased from Shanghai Cente Biotechnology) was used to remove rRNA. Eukaryotic mRNA was enriched using Oligo(dT) magnetic beads via AT complementary pairing with the polyA tail of mRNA. Fragmentation buffer was added to break the mRNA into short fragments. Then, using the RNA as a template, the first strand of cDNA was synthesized using six-base random primers. The second strand of cDNA was synthesized with buffer, dNTPs, and DNA polymerase I. The double-stranded cDNA was then purified using AMPure XP beads, end-repaired, tailed, and ligated with sequencing adapters. Finally, the cDNA was amplified by PCR to construct a cDNA library. The enriched cDNA was then subjected to metagenomic sequencing using NovaSeq on the Illumina platform, with a sequencing depth of approximately 10 Gb, yielding metagenomic data. 3. Data quality control and alignment: Metatranscriptome data underwent quality control and filtering using FastQC (v 0.11.9) and Fastp (v 0.20.1), with both quality control and filtering set to default parameters. RNA quality was assessed using RIN (RNA integrity number) values ​​(RIN≥7.0). Bowtie2 (v 2.3.5.1) was used to map the quality-controlled reads data to the human reference genome GRCh38, generating BAM files containing location information. 4. Transcript Quantification: Based on the alignment results (BAM file), the raw read counts for each gene are calculated using Cufflink. 5. Standardized entry of patient clinical information: Establish a structured clinical database covering basic patient information (age, immune status) and symptom characteristics. Group patients according to different symptoms into an infection group (meeting any one of the following criteria: CRP ≥ 10 g / L, PCT ≥ 0.5 ng / mL, positive serological test (Streptococcus pneumoniae or Haemophilus influenzae), positive clinical culture result (analytical sample cultured to find fungal or bacterial pathogens), G test result > 80 pg / mL, or positive mNGS result) and a non-infection group (CRP < 10 g / L and PCT < 0.5 ng / mL, negative serological test result, negative clinical culture result, G test result ≤ 80 pg / mL, and negative mNGS result). 6. Screening biomarkers: Combining differential expression analysis, co-expression analysis, and machine learning algorithms to deeply mine gene expression data from different groups in order to identify genes that are stably expressed or significantly changed in the infected group as biomarkers; (1) Differential expression analysis: pairwise comparisons (infected group VS non-infected group) were performed using the edgeR package. A threshold was set (|log2FC|≥1 and FDR≤0.05) to screen for genes that showed significant changes only in a certain symptom group; (2) Co-expression analysis: A weighted gene co-expression network was constructed, a soft threshold (power value) was set to ensure a scale-free topology, and the Person correlation coefficient (R) was calculated. When R... 2 When the correlation coefficient is >0.8 and the average connectivity is high, the soft threshold is set to 7; the Spearman correlation coefficient between module feature genes (MEs) and symptom phenotypes (infection status) is calculated. Infection status is divided into infected and uninfected, determined by clinical culture results and mNGS results. A positive clinical culture result or mNGS result indicates infection, while a negative clinical culture result or mNGS result indicates uninfected. Significantly correlated modules are screened ( P<0.05); extract hub genes from significant modules (top 10% of module connectivity), and screen candidate genes based on differential expression results; (3) Functional annotation: ClusterProfiler was used to perform functional annotation and enrichment analysis on differentially expressed genes and co-expressed genes; (4) Machine learning algorithm: Based on differentially expressed genes and co-expressed genes, redundant genes are compressed by cross-validation based on LASSO regression, and variables with high contribution to symptom classification (the top 10 genes with the highest contribution) are retained; the Random Forest model in R language is used to predict respiratory infection-related genes. The Random Forest model is constructed using the R package "randomForest". The top 10% of genes with the highest scores are selected according to the importance scores of the genes; the "rfe" function of the R package "caret" is used to gradually remove the features with the least contribution to the model to evaluate the importance of genes. In each iteration, the SVM model is used to sort the features. Each time, 20% of the least important features are removed. The performance of the feature subset is evaluated by cross-validation, and the feature combination with the largest AUC is selected as the final result; the results of LASSO, Random Forest and SVM model algorithms are integrated, and the intersection genes are taken as biomarkers. 7. Validation of screening results: The robustness of biomarkers was validated on independent datasets. ROC curves of biomarkers were established, and those with AUC > 0.8 were considered to have clinical value.

[0026] Based on the above screening method, the biomarkers associated with respiratory tract infections were identified as IL6, CXCL8, TNF-α, HSP90AA1, and S100A9. Further validation of the diagnostic performance of these biomarkers was conducted using ROC curve analysis. The results showed that the area under the curve for the biomarker combination was 0.847, indicating good clinical diagnostic performance. Furthermore, the diagnostic accuracy of these biomarkers reached 80.14%, demonstrating their high discriminative ability in the diagnosis of respiratory tract infections.

[0027] Example 2 This embodiment provides a method for screening respiratory infection-related biomarkers based on metagenomics, as shown in the flowchart below. Figure 1 As shown, the screening of biomarkers for identifying severe illness includes the following steps: 1. Sample Data Acquisition: Biological samples, including bronchoalveolar lavage fluid and blood, were collected from patients with respiratory tract infections for analysis. Total RNA samples were extracted from the analytical samples using a tissue cell RNA extraction kit (Trizol extraction method) (purchased from Shanghai Shangbao Biotechnology Co., Ltd.), and the extracted RNA samples were stored at -80℃. 2. Library Construction: Cutadapt was used to remove adapter sequences and low-quality bases. The Ribo-Zero™ rRNARemoval Kit (Gram-Negative Bacteria) 1 Kit (purchased from Shanghai Cente Biotechnology) was used to remove rRNA. Eukaryotic mRNA was enriched using Oligo(dT) magnetic beads via AT complementary pairing with the polyA tail of mRNA. Fragmentation buffer was added to break the mRNA into short fragments. Then, using the RNA as a template, the first strand of cDNA was synthesized using six-base random primers. The second strand of cDNA was synthesized with buffer, dNTPs, and DNA polymerase I. The double-stranded cDNA was then purified using AMPure XP beads, end-repaired, tailed, and ligated with sequencing adapters. Finally, the cDNA was amplified by PCR to construct a cDNA library. The enriched cDNA was then subjected to metagenomic sequencing using NovaSeq on the Illumina platform, with a sequencing depth of approximately 10 Gb, yielding metagenomic data. 3. Data quality control and alignment: Metatranscriptome data underwent quality control and filtering using FastQC (v 0.11.9) and Fastp (v 0.20.1), with both quality control and filtering set to default parameters. RNA quality was assessed using RIN (RNA integrity number) values ​​(RIN≥7.0). Bowtie2 (v 2.3.5.1) was used to map the quality-controlled reads data to the human reference genome GRCh38, generating BAM files containing location information. 4. Transcript Quantification: Based on the alignment results (BAM file), the raw read counts for each gene are calculated using Cufflink. 5. Standardized entry of patient clinical information: Establish a structured clinical database covering basic patient information (age, immune status) and symptom characteristics. Group patients according to different symptoms. The infection group (meeting any one of the following criteria: CRP ≥ 10 g / L, PCT ≥ 0.5 ng / mL, positive serological test (Streptococcus pneumoniae or Haemophilus influenzae), positive clinical culture result (analysis of sample yielding fungal or bacterial pathogens), G test result > 80 pg / mL, or positive mNGS result) is further divided into the severe pneumonia group (meeting any one of the following criteria: CRP ≥ 30 g / L, PCT ≥ 2 ng / mL, or G test result ≥ 200 pg / mL) and the mild pneumonia group (meeting any one of the following criteria: 10 g / L ≤ CRP < 30 g / L and 0.5 ng / mL ≤ PCT < 2 ng / mL, 80 pg / mL < G test result < 200 pg / mL). 6. Screening biomarkers: Combining differential expression analysis, co-expression analysis, and machine learning algorithms to deeply mine gene expression data from different groups in order to identify genes that are stably expressed or significantly changed in the severe pneumonia group as biomarkers; (1) Differential expression analysis: pairwise comparisons (severe pneumonia group vs. mild pneumonia group) were performed using the edgeR package. A threshold was set (|log2FC|≥1 and FDR≤0.05) to screen for genes that showed significant changes only in a certain symptom group; (2) Co-expression analysis: A weighted gene co-expression network was constructed, a soft threshold (power value) was set to ensure a scale-free topology, and the Person correlation coefficient (R) was calculated. When R... 2 When the correlation coefficient is >0.8 and the average connectivity is high, the soft threshold is set to 7; the Spearman correlation coefficient between module feature genes (MEs) and symptom phenotypes (severe illness status) is calculated, and the severe illness status is scored based on clinical indicators (SOFA score, APACHE II score). The scoring table is shown in Table 1. Significantly correlated modules are selected. P <0.05); extract hub genes from significant modules (top 10% of module connectivity), and screen candidate genes based on differential expression results; (3) Functional annotation: ClusterProfiler was used to perform functional annotation and enrichment analysis on differentially expressed genes and co-expressed genes; (4) Machine learning algorithm: Based on differentially expressed genes and co-expressed genes, redundant genes are compressed by cross-validation based on LASSO regression, and variables with high contribution to symptom classification (the top 10 genes with the highest contribution) are retained; the Random Forest model in R language is used to predict respiratory infection-related genes. The Random Forest model is constructed using the R package "randomForest". The top 10% of genes with the highest scores are selected according to the importance scores of the genes; the "rfe" function of the R package "caret" is used to gradually remove the features with the least contribution to the model to evaluate the importance of genes. In each iteration, the SVM model is used to sort the features. Each time, 20% of the least important features are removed. The performance of the feature subset is evaluated by cross-validation, and the feature combination with the largest AUC is selected as the final result; the results of LASSO, Random Forest and SVM model algorithms are integrated, and the intersection genes are taken as biomarkers. 7. Validation of screening results: The robustness of biomarkers was validated on independent datasets. ROC curves of biomarkers were established, and those with AUC > 0.8 were considered to have clinical value.

[0028] Table 1. Severe Illness Status Scoring Table

[0029] Based on the above screening method, the biomarkers associated with whether the illness is severe were identified as IL6, CXCL8, and TNF-α. To further verify the diagnostic performance of the biomarkers, ROC curve analysis was performed. The results showed that the area under the curve for the combination of biomarkers was 0.862, indicating good clinical diagnostic performance. Furthermore, the diagnostic accuracy of the above biomarkers reached 81.07%, demonstrating that the biomarkers have a high discriminative ability in diagnosing whether respiratory infections are severe.

[0030] The specification only describes preferred embodiments of the present invention. The present invention is not limited to the above embodiments. Various changes and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit and scope of the present invention should fall within the protection scope defined by the present invention.

Claims

1. The application of reagents for detecting the expression levels of biomarker genes in the preparation of diagnostic agents for detecting severe respiratory infections, characterized in that, The biomarkers are IL6, CXCL8, and TNF-α.