Metagenome sequencing data analysis method of sample, electronic equipment and system
By acquiring sample data in one sampling, comparing, splitting and analyzing, the problems of high workload and high cost in the prior art are solved, and efficient and accurate acquisition of microbiome and host genomic data, especially functional gene annotation of skin samples.
Patent Information
- Application Number
- CN202410156890.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-04
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art requires two samples to obtain microbiome and host genomic data, resulting in high workload and high cost, and invasive sampling methods are not popular with subjects.
The sequencing data of the samples were obtained by one sampling, and compared with the host reference genome, split into host genome data and microbiome data, and analyzed separately. Tools such as Qualimap and GATK were used for quality evaluation and mutation analysis, tools such as MetaPhlAn were used for species composition annotation, and functional gene annotation was used for iHSMGC reference gene set.
It realizes the acquisition of biomic data through one sampling, which reduces costs and improves efficiency, especially the functional gene annotation for skin samples is more accurate, supporting research on skin microorganisms and host interactions.
Smart Images

Figure CN120472992A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of molecular biology technology, and specifically relates to a method for analyzing metagenomic sequencing data of a sample, as well as electronic equipment and a system. Background Art
[0002] A vast array of microorganisms resides on the human skin, mouth, lungs, intestines, genitals, and other sites, forming a unique microbial community. This microbial community, beginning in infancy, is influenced by a variety of internal and external factors, including genes, age, diet, environment, and lifestyle, forming a complex symbiotic relationship with the human body, significantly influencing immune regulation, nutrition, neuroscience, metabolite breakdown, and microenvironmental homeostasis. However, further research is needed to determine the metabolic processes within these commensal bacteria that influence microenvironmental homeostasis, and the genomic variations that occur in host cells during microbiome-host interactions.
[0003] Metagenomes (meta) are genomes used to study environmental microbial samples. In recent years, with the rapid development of metagenomic sequencing technology, in-depth research on the interaction between the microbiome and the host has become possible. Unlike 16S rRNA gene sequencing, which relies on PCR amplification of specific genes, metagenomic sequencing technology can not only detect a large number of unculturable microorganisms in the sample, but also has a higher resolution for species identification, usually at the species or strain level. In addition, the deeper sequencing depth can also realize the mining of downstream functional genes, which is conducive to in-depth understanding of the interaction between skin microorganisms and the host at the level of microbial functional genes.
[0004] The bioinformatics analysis of off-line data is a key step in converting microbial genomic information into insights into their potential impact on the host. The core of this conversion process is to predict the potential biological functions of gene sequences generated by metagenomic sequencing data. This is achieved by aligning query sequences (of unknown function) with a database of reference sequences with known taxonomic and functional information. Based on sequence similarity, sequences with known functions are used to annotate gene sequences of unknown function, thereby predicting the functions of unknown genes.
[0005] Currently, the commonly used analysis process for metagenomic sequencing based on read length comparison is as follows: Figure 1 As shown, the main processes include quality control filtering of low-quality sequences and adapters in the raw sequencing data, removal of host-derived genomic sequences, and annotation of microbial species and functional genes. To obtain human genomic information, whole-genome sequencing of blood DNA is typically performed by drawing blood from a vein, but this requires two separate sampling steps, which is labor-intensive and costly. Summary of the Invention
[0006] The embodiments of the present invention provide a method, electronic device, and system for analyzing metagenomic sequencing data of a sample. The metagenomic sequencing data analysis method can obtain dual-omics information through a single collection, which is cost-effective and more efficient.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is:
[0008] A first aspect of the present invention provides a method for analyzing metagenomic sequencing data of a sample, comprising the following steps:
[0009] Obtaining sequencing data of the sample;
[0010] Comparing the sequencing data with the host reference genome to obtain host genome data and microbiome data;
[0011] The host genome data and the microbiome data are analyzed separately.
[0012] In some embodiments of the present invention, analyzing the host genome data includes performing at least one of quality integrity assessment and mutation analysis on the host genome data.
[0013] In some embodiments of the present invention, Qualimap is used for quality integrity assessment.
[0014] In some embodiments of the present invention, GATK is used for mutation analysis.
[0015] In some embodiments of the present invention, analyzing the microbiome data includes performing at least one of species composition annotation and functional gene annotation on the microbiome data.
[0016] In some embodiments of the present invention, at least one of MetaPhlAn, mOTUs2, Mothur, Kraken, Kraken2, KarkenUniq, Megan, Bracken, Centrifuge, CLARK, CLARK-S, k-SLAM, MegaBLAST, metaOthello, PathSeq, prophyle, and taxMaps is used for species composition annotation.
[0017] In some embodiments of the present invention, the reference gene set for functional gene annotation includes iHSMGC.
[0018] In some embodiments of the invention, the host is a human.
[0019] In some embodiments of the present invention, the sample is at least one of skin, saliva, lung, intestine, feces, and urine.
[0020] In some embodiments of the present invention, the host reference genome includes any one of hg18, hg19, and hg38.
[0021] In some embodiments of the present invention, the sequencing data further includes removing low-quality sequences and adapters before alignment.
[0022] According to a second aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the aforementioned metagenomic sequencing data analysis method.
[0023] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the aforementioned metagenomic sequencing data analysis method when running the computer program.
[0024] A fourth aspect of the embodiments of the present invention provides a sample metagenomic sequencing data analysis system, comprising:
[0025] An acquisition module, configured to acquire sequencing data of the sample;
[0026] An alignment module is used to align the sequencing data with a host reference genome to separate the host genome data and the microbiome data;
[0027] An analysis module is used to analyze the host genome data and the microbiome data respectively.
[0028] The beneficial effects of the embodiments of the present invention are:
[0029] The present invention aims to achieve the goal of simultaneously acquiring dual-omics data information, including microbiome data information and human genome data information, through only one sampling at the analytical technology level.
[0030] In addition, the analysis method of the present invention integrates the reference gene set iHSMGC for the human skin microbiome in the functional gene annotation analysis module applicable to microorganisms, which can achieve accurate annotation of functional genes specific to the skin microbial habitat. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flowchart of metagenomic sequencing data analysis in the existing technology.
[0032] Figure 2It is a flow chart of the metagenomic sequencing analysis of the present invention.
[0033] Figure 3 This is an example of a metagenomic sequencing analysis process for a sample in one embodiment of the present invention.
[0034] Figure 4 This is the coverage statistics result of human genome data derived from skin samples in one embodiment of the present invention.
[0035] Figure 5 This is a display of the relevant functional gene annotation results of the microorganisms of the skin sample in one embodiment of the present invention. DETAILED DESCRIPTION
[0036] In the description of the present invention, the terms "first," "second," and "third" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first," "second," or "third" may explicitly or implicitly include at least one of such features. In the description of the present invention, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0037] The first aspect of the present invention provides a method for analyzing metagenomic sequencing data of a sample, referring to Figure 2 , including the following steps S100 to S300.
[0038] S100: Obtain sequencing data of the sample.
[0039] Among them, metagenomics refers to the study of all microorganisms in a sample. Without separation and cultivation, the DNA of all culturable and unculturable microorganisms is directly extracted from the sample, and the total genetic information of the microorganisms is obtained through sequencing, so as to study the community structure, species classification, systematic evolution, gene function, metabolic network, etc. of the microorganisms, as well as their relationship with the surrounding environment.
[0040] Among them, microorganisms include organisms in the viral kingdom, prokaryotic kingdom, protist kingdom, and fungal kingdom. According to their cell structure, they can be divided into non-cellular structure microorganisms (such as viruses, subviruses, etc.), prokaryotic cell microorganisms (such as bacteria, chlamydia, mycoplasma, rickettsia, spirochetes, actinomycetes, cyanobacteria, archaea, etc.) and eukaryotic cell microorganisms (such as fungi, unicellular algae, protozoa, etc.).
[0041] Wherein, the sample refers to a sample obtained from a host individual, and the host individual can be, for example, a mammal. Mammals include, but are not limited to, animals such as Carnivora, Perissodactyla, Artiodactyla, Rodentia, Lagomorpha, Primates, and specifically any one of dogs, cats, pigs, sheep, cattle, horses, rabbits, mice, monkeys, orangutans, gorillas, chimpanzees, and humans. The sample can be a sample collected from a site where symbiotic microorganisms are optionally present, such as samples from the skin, lungs, intestines, epithelium (including oral cavity, genitalia, etc.), and body fluids or secretions such as saliva, sputum, alveolar lavage fluid, bronchial brushings, urine, and feces, intestinal contents, etc., and at least one sample.
[0042] Sequencing refers to the identification of DNA base sequences. In metagenomics, sequencing data can be obtained using any of the following methods: first-generation sequencing, second-generation sequencing, or third-generation sequencing. First-generation sequencing includes Maxam-Gilbert sequencing, Sanger dideoxy sequencing, fluorescent automated sequencing, and hybridization sequencing. Second-generation sequencing includes 454RocheGS FLX, Illumina, SOLiD, Ion Torrent, and BGISEQ. Third-generation sequencing includes HeliScope, PacBio HiFi, PacBio CLR, ONT, and Cyclone WT.
[0043] The general sequencing process typically includes constructing a sequencing library and loading the constructed sequencing library onto a sequencing platform. Depending on the selected sequencing platform, different types of sequencing libraries can be constructed, such as single-end or double-end, linear or circular. The sequencing data obtained after the sequencing library is unloaded from the sequencing platform typically consists of multiple reads. Therefore, in some embodiments of the present application, obtaining sequencing data for a sample refers to obtaining the read length of the sample obtained by sequencing.
[0044] In some embodiments, before the sequencing data is aligned, it is also filtered and quality controlled, such as removing low-quality sequences and adapters contained in the sequencing data. In some embodiments, the filtering and quality control of the sequencing data can be performed by existing methods, such as any one of SOAPnuke, fastp, seqtk, Trimmomatic, etc. In some embodiments, the low-quality sequence varies according to the requirements of different sequencing platforms and software. For example, for SOAPnuke, it can be a sequence in which the low-quality bases in the read length are above a certain threshold (such as the proportion of bases with a base quality less than 10 in the read length is not less than 50%, etc.), and reads with too low average quality values can be further removed. In some embodiments, removing adapters includes directly removing reads containing adapter sequences or removing adapter sequences in the read length. In addition, filtering and quality control also include removing reads containing a high proportion of N, such as reads with an N ratio greater than 5%.
[0045] S200: Align the sequencing data with the host reference genome and split it into host genome data and microbiome data.
[0046] For samples, which include microbial DNA and host DNA, generally speaking, in the process of metagenomic sequencing data analysis, in order to obtain microbiome data, after completing quality control, it is necessary to remove the host-derived genomic sequence through comparison, and only retain the single-omics data information of the microbiome. When it is necessary to further analyze the interaction between the host and the microorganism, in order to obtain the host's genomic information, taking humans as an example, it is usually necessary to use venous blood drawing to perform whole-genome sequencing of blood DNA. However, this method requires two separate samplings (microbiome sampling and human blood sample collection), which is labor-intensive and costly; and as far as invasive blood sample collection is concerned, subjects may be more inclined to choose non-invasive sampling methods.
[0047] To this end, in the embodiment of the present application, during the alignment process, the sequencing data aligned to the host reference genome are retained as host genome reads, thereby achieving the splitting of the sequencing data. The collection of host genome reads is used as host genome data, while other reads that failed to be aligned to the host reference genome are used as microbial genome reads, and the collection of microbial genome reads is used as microbiome data, and these two are directly used for subsequent analysis.
[0048] In some embodiments, alignment can be performed using any alignment tool known in the art, such as bwa, bowtie2, GATK, or the like. In some embodiments, during the alignment process, a read is allowed to have a maximum of k base mismatches. In some embodiments, k ≤ 3, for example, 1, 2, or 3. Therefore, if a read has more than k base mismatches, the read cannot be aligned to the host reference genome.
[0049] In some embodiments, the host is human, and thus the host reference genome is the entire human genome or a set range of the entire human genome, or can be the genome of one or more human individuals. In some embodiments, the reference genome includes, but is not limited to, any of GRCh36 (hg18), GRCh37 (hg19), and GRCh38 (hg38).
[0050] S300: Analyze host genome data and microbiome data separately.
[0051] In some embodiments, the analysis of the host genome data includes at least one of quality integrity assessment and mutation analysis of the host genome data. The integrity of the host genome in the sample can be confirmed by quality integrity assessment. For example, the quality integrity of the host genome can be assessed by evaluating at least one of the indicators such as the alignment of the host genome data in different chromosomes, sequencing depth, and alignment quality in the sample. In some embodiments, the quality integrity assessment can be performed by Qualimap or other tools known in the art. Mutation detection of the host genome data is achieved by mutation analysis, and the mutation information at the host genome level obtained can provide an analysis basis for downstream related whole genome association analysis. In some embodiments, mutation analysis can be performed by GATK or other tools known in the art, and at least one of the results such as the mutation detection file of the host gene, the mutation result annotation file, and the mutation detection evaluation file can be obtained.
[0052] In some embodiments, the analysis of host genomic data also includes predictive analysis of specific marker genes, thereby providing a reference for diagnosis, treatment, health care, and beauty. For example, predictive analysis of aging-related marker genes can be performed on human skin samples, providing a reference for individuals to understand their skin condition at the genomic level, facilitating genetic skincare and the development of personalized skincare plans. The aging-related marker gene can be at least one of inflammatory response, oxidative stress, apoptosis, cell proliferation, DNA damage repair, or other related genes known in the art.
[0053] Marker genes can be genes differentially expressed in one or more cell types or one or more individuals, for example, genes that are specifically up-regulated or down-regulated. Expression levels can be quantified using any of the following metrics: RPKM (Reads Per Kilobase of transcript per Million reads mapped), FPKM (Fragments Per Kilobase of transcript per Million reads mapped), or TPM (Transcripts Per Kilobase of exon model per Million mapped reads). In some embodiments, the quantification method is TPM. In some embodiments, the differential expression condition is that the TPM value of a gene is greater than or less than a predetermined fold change (FC) relative to the TPM value of the gene in the control group, thereby determining differentially expressed genes that are specifically up-regulated or down-regulated. In some embodiments, the fold change is calculated by transforming the fold change of the TPM value. In some specific embodiments, the transformation includes logarithmic transformation, for example, taking the logarithm of 2, e, or 10. Because some genes may not be expressed, the logarithm is prepended with one. In some specific embodiments, the fold change is calculated using the logarithmic transformation log2(TPM+1). It is understood that the differential expression condition also includes that the differential expression between the two is significant, such as a significant P value of less than 0.05. In some embodiments, it also includes a corrected P value of less than 0.1.
[0054] In some embodiments, analyzing the microbiome data includes annotating the microbiome data with at least one of species composition and functional gene annotation. In some embodiments, species composition annotation can be performed based on microbial markers or by comparison with a microbial reference genome. Specifically, at least one of MetaPhlAn, mOTUs2, Mothur, Kraken, Kraken2, KarkenUniq, Megan, Bracken, Centrifuge, CLARK, CLARK-S, k-SLAM, MegaBLAST, metaOthello, PathSeq, prophyle, and taxMaps can be selected. Microbial markers can, for example, be characteristic genes or fragments thereof at the genus, species, or strain level of a microorganism.
[0055] In some embodiments, the functional gene annotation of microorganisms uses the commonly used tool HUMAnN3, which relies on the ChocoPhlAn 3 reference database. This database is mainly based on the core data of the public database UniProt / UniRef and the microbial genomes and their annotated proteins / genes in the taxonomy and genome databases of NCBI. After processing, it is applied to metagenomic bacterial community classification, functional annotation, strain identification and phylogenetic analysis.
[0056] Although the public databases such as UniProt and NCBI that this comparison process relies on have a wide range of microbial species and high diversity, they may be underrepresented for some specific samples. For example, when the host is human and the sample type is skin, unlike other mucosal systems, human skin is a nutrient-poor, high-salt, acidic environment; and the skin surface is filled with various lipid substances not found in other parts of the body. Some of these lipid components have antimicrobial activity (such as saponin), while others can be metabolized by skin microorganisms into free fatty acids, diglycerides, and monoglycerides (such as triglycerides), which have biological activities such as resisting the colonization of other microorganisms or stimulating and regulating host cell metabolism.
[0057] To this end, in some embodiments of the present application, for human skin samples, when annotating functional genes, the reference gene set includes the reference gene set iHSMGC derived from the human skin microbiome. The use of this reference gene set will facilitate the accurate annotation of habitat-specific functional genes in skin microbiome samples, and provide a reference basis for the study of interactions between skin microorganisms and hosts, as well as the role of skin microorganisms in the occurrence, development, and evaluation of skin diseases.
[0058] In some embodiments of the present application, functional gene annotation is performed using BBmap.
[0059] In some embodiments of the present application, the functional gene annotation further includes quantification of these functional genes. In some embodiments, the quantitative method includes any one of RPKM (Reads Per Kilobase of transcript per Million reads mapped), FPKM (Fragments Per Kilobase of transcript per Millionreads mapped), TPM (Transcripts Per Kilobase of exon model per Million mappedreads), etc., by quantitatively removing the influence of sequencing depth and gene length. Different quantitative methods can be selected according to different sequencing data. For example, RPKM can be selected for single-end sequencing results, and FPKM can be selected for double-end sequencing results. In some embodiments, the quantitative method is TPM.
[0060] In some methods, the relative abundance of relevant genes in the skin-derived iHSMGC reference gene set in the sample is calculated based on the result file generated after BBMap alignment. In some embodiments, the relative abundance of relevant functional genes in the iHSMGC reference gene set in the metagenomic sequencing data is described by the TPM value, and the TPM value is calculated as follows:
[0061]
[0062] in,
[0063] A second aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions for enabling a computer to execute the aforementioned metagenomic sequencing data analysis method.
[0064] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the aforementioned metagenomic sequencing data analysis method when running the computer program.
[0065] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the aforementioned metagenomic sequencing data analysis method described in the embodiments of the present invention. The processor performs metagenomic sequencing data analysis on the sample by executing the non-transitory software program and instructions stored in the memory.
[0066] The memory may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function, while the data storage area may store and execute the aforementioned programs. Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.
[0067] In some embodiments, the memory may include a memory remotely located relative to the processor, and the remote memory may be connected to the processor via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0068] The non-transient software programs and instructions required to implement the above method are stored in the memory, and when executed by one or more processors, the above-mentioned metagenomic sequencing data analysis method is performed.
[0069] A fourth aspect of the embodiments of the present invention provides a sample metagenomic sequencing data analysis system, comprising:
[0070] An acquisition module is used to obtain sequencing data of samples;
[0071] The alignment module is used to align the sequencing data with the host reference genome and split it into host genome data and microbiome data;
[0072] Analysis module, the analysis module is used to analyze host genome data and microbiome data respectively.
[0073] In some embodiments, the microorganisms include organisms from the virus kingdom, prokaryotic kingdom, protist kingdom, and fungal kingdom, and can be divided according to their cell structure into non-cellular structure microorganisms (such as viruses, subviruses, etc.), prokaryotic cell microorganisms (such as bacteria, chlamydia, mycoplasma, rickettsia, spirochetes, actinomycetes, cyanobacteria, archaea, etc.) and eukaryotic cell microorganisms (such as fungi, unicellular algae, protozoa, etc.).
[0074] In some embodiments, the sample refers to a sample obtained from a host individual, and the host individual can be, for example, a mammal. Mammals include, but are not limited to, animals of the order Carnivora, Perissodactyla, Artiodactyla, Rodentia, Lagomorpha, Primates, and specifically any one of dogs, cats, pigs, sheep, cattle, horses, rabbits, mice, monkeys, orangutans, gorillas, chimpanzees, and humans. The sample can be a sample collected from a site where symbiotic microorganisms are optionally present, such as samples from the skin, lungs, intestines, epithelium (including oral cavity, genitalia, etc.), and body fluids or secretions such as saliva, sputum, alveolar lavage fluid, bronchial brushings, urine, and feces, intestinal contents, and at least one of the samples.
[0075] In some embodiments, the metagenomic sequencing data analysis system further includes a sequencing module for sequencing samples to obtain sequencing data, specifically, for sequencing sequencing libraries to obtain sequencing data. In some embodiments, the metagenomic sequencing data analysis system further includes a DNA extraction module for extracting DNA from samples. In some embodiments, the metagenomic sequencing data analysis system further includes a library construction module for constructing a library from the extracted DNA to obtain a sequencing library. In some embodiments, the sequencing library is a single-end or double-end, chain-shaped or circular sequencing library.
[0076] In some embodiments, the sequencing module uses any of first-generation sequencing, second-generation sequencing, or third-generation sequencing to obtain sequencing data. First-generation sequencing includes Maxam-Gilbert sequencing, Sanger dideoxy sequencing, fluorescent automated sequencing, and hybridization sequencing; second-generation sequencing includes 454 Roche GS FLX, Illumina, SOLiD, IonTorrent, and BGISEQ; and third-generation sequencing includes HeliScope, PacBio HiFi, PacBio CLR, ONT, and Cyclone WT.
[0077] In some embodiments, before the sequencing data is aligned, it is also included in the filtering and quality control, such as removing low-quality sequences and connectors contained in the sequencing data. In some embodiments, the filtering and quality control of the sequencing data can be performed by existing methods, such as any one of SOAPnuke, fastp, seqtk, Trimmomatic, etc. Taking SOAPnuke as an example, low-quality sequences can be sequences in which the low-quality bases in the read length are above a certain threshold (such as the proportion of bases with a base quality less than 10 in the read length is not less than 50%, etc.), and reads with too low average quality values can be further removed. In some embodiments, removing connectors includes directly removing reads containing connector sequences or removing connector sequences in reads. In addition, filtering and quality control also include removing reads containing a high proportion of N, such as reads with an N ratio greater than 5%.
[0078] In some embodiments, alignment can be performed using any alignment tool known in the art, such as bwa, bowtie2, GATK, or the like. In some embodiments, during the alignment process, a read is allowed to have a maximum of k base mismatches. In some embodiments, k ≤ 3, for example, 1, 2, or 3. Therefore, if a read has more than k base mismatches, the read cannot be aligned to the host reference genome.
[0079] In some embodiments, the host is human, and thus the host reference genome is the entire human genome or a set range of the entire human genome, or can be the genome of one or more human individuals. In some embodiments, the reference genome includes, but is not limited to, GRCh36 (hg18), GRCh37 (hg19), GRCh38 (hg38), or any of the following:
[0080] In some embodiments, the analysis of the host genome data includes at least one of quality integrity assessment and mutation analysis of the host genome data. The integrity of the host genome in the sample can be confirmed by quality integrity assessment. For example, the quality integrity of the host genome can be assessed by evaluating at least one of the indicators such as the alignment of the host genome data in different chromosomes, sequencing depth, and alignment quality in the sample. In some embodiments, the quality integrity assessment can be performed by Qualimap or other tools known in the art. Mutation detection of the host genome data is achieved by mutation analysis, and the mutation information at the host genome level obtained can provide an analysis basis for downstream related whole genome association analysis. In some embodiments, mutation analysis can be performed by GATK or other tools known in the art, and at least one of the results such as the mutation detection file of the host gene, the mutation result annotation file, and the mutation detection evaluation file can be obtained.
[0081] In some embodiments, the analysis of host genomic data also includes predictive analysis of specific marker genes, thereby providing a reference for diagnosis, treatment, health care, and beauty. For example, predictive analysis of aging-related marker genes can be performed on human skin samples, providing a reference for individuals to understand their skin condition at the genomic level, facilitating genetic skincare and the development of personalized skincare plans. The aging-related marker gene can be at least one of inflammatory response, oxidative stress, apoptosis, cell proliferation, DNA damage repair, or other related genes known in the art.
[0082] Marker genes can be genes differentially expressed in one or more cell types or one or more individuals, for example, genes that are specifically up-regulated or down-regulated. Expression levels can be quantified using any of RPKM, FPKM, and TPM. In some embodiments, the quantification method is TPM. In some embodiments, the differential expression condition is that the TPM value of a gene is greater than or less than a set fold change (FC) relative to the TPM value of the gene in the control group, resulting in differentially expressed genes that are specifically up-regulated or down-regulated. In some embodiments, the fold change is calculated by transforming the fold change in the TPM value. In some specific embodiments, the transformation includes logarithmic transformation, for example, taking the logarithm of 2, e, or 10. Because some genes may not be expressed, the logarithm is pre-added by one. In some specific embodiments, the fold change is calculated using the logarithmic transformation log2(TPM+1). It is understood that the differential expression condition also includes significant differential expression, for example, a significance P value of less than 0.05. In some embodiments, the corrected P value is also less than 0.1.
[0083] In some embodiments, analyzing the microbiome data includes annotating the microbiome data with at least one of species composition and functional gene annotation. In some embodiments, species composition annotation can be performed based on microbial markers or by comparison with a microbial reference genome. Specifically, at least one of MetaPhlAn, mOTUs2, Mothur, Kraken, Kraken2, KarkenUniq, Megan, Bracken, Centrifuge, CLARK, CLARK-S, k-SLAM, MegaBLAST, metaOthello, PathSeq, prophyle, and taxMaps can be selected. Microbial markers can, for example, be characteristic genes or fragments thereof at the genus, species, or strain level of a microorganism.
[0084] In some embodiments, the functional gene annotation of the microorganism is performed using the commonly used tools HUMAnN3 or iHSMGC.
[0085] In some embodiments of the present application, functional gene annotation is performed using BBmap.
[0086] In some embodiments of the present application, the functional genes are annotated and quantified. In some embodiments, the quantitative method includes any one of RPKM, FPKM, TPM, etc., which removes the influence of sequencing depth and gene length by quantitative analysis. Different quantitative methods can be selected according to different sequencing data. For example, RPKM can be selected for single-end sequencing results, and FPKM can be selected for double-end sequencing results. In some embodiments, the quantitative method is TPM.
[0087] In some methods, the relative abundance of relevant genes in the skin-derived iHSMGC reference gene set in the sample is calculated based on the result file generated after BBMap alignment. In some embodiments, the relative abundance of relevant functional genes in the iHSMGC reference gene set in the metagenomic sequencing data is described by the TPM value, and the TPM value is calculated as follows:
[0088]
[0089] in,
[0090] The device or system implementation described above is merely illustrative. The modules described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network elements. Some or all of these modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0091] It is understood that all or some of the steps disclosed above can be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). It is understood that computer storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include but are not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, disk storage or other magnetic storage device, or any other medium that can be used to store desired information and can be accessed by a computer.
[0092] Additionally, it will be appreciated that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0093] The present invention is further described in detail below through specific examples.
[0094] It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0095] The experimental methods in the following examples, where specific conditions are not specified, were generally performed under conventional conditions or the conditions recommended by the manufacturers. The materials and reagents used in these examples were commercially available unless otherwise specified.
[0096] The main reagents and instruments involved in the following examples are shown in Tables 1 and 2:
[0097] Table 1. Main reagents
[0098] Reagent material name factory Skin swab and preservation fluid sampling kit ESwab MGIEasy Genomic DNA Extraction Kit (Magnetic Bead Method) BGI MGIEasy Enzyme-Digested DNA Library Preparation Kit BGI Anhydrous ethanol Xilong Chemical Isopropyl alcohol Xilong Chemical RNase A Thermo Fisher 1.5ml microcentrifuge tube Crystalgen
[0099] Table 2. Main instruments
[0100] Instrument name factory Low-temperature centrifuge Eppendorf Comfortable constant temperature mixer Eppendorf MGISEQ-2000 sequencing platform BGI
[0101] Example 1: Sample Sequencing
[0102] refer to Figure 3 , skin samples were collected from volunteers, DNA was extracted, sequencing libraries were constructed and sequenced. The specific process is as follows:
[0103] 1. Collection of facial skin samples:
[0104] 1) Using sterile saline-soaked swabs, skin samples were collected from 15 volunteers;
[0105] 2) Keep the swab head in proper contact with the skin and wipe repeatedly 10-20 times, sampling an area of 4 cm × 4 cm. Rotate the swab during the wiping process to ensure that the entire swab can be combined with sufficient skin sample;
[0106] 3) Immerse the skin swab sample in preservation solution for preservation.
[0107] 2. Extraction of skin genomic DNA:
[0108] Extraction was performed according to the SOP of the MGIEasy Genomic DNA Extraction Kit (magnetic bead method).
[0109] 3. Construction of sequencing library:
[0110] The sequencing library was constructed according to the instructions of the MGIEasy enzyme digestion DNA library preparation reagent kit.
[0111] 4. Whole genome sequencing:
[0112] Whole-genome sequencing of skin-derived human genomic DNA was performed using the BGISEQ-2000 sequencing platform.
[0113] Example 2: Sequencing data analysis
[0114] refer to Figure 3 The sequencing results are pre-filtered for quality control, and then the sequencing data are compared with the host reference genome to split the dual-omics data, obtaining microbiome data and human genome data for the next step of analysis. The specific process is as follows:
[0115] 1. Quality control filtering of raw sequencing data: Use the filter function module of SOAPnuke software to remove low-quality sequences and adapter sequences from the raw sequencing data of skin samples (parameter settings: -l 20 -q 0.5 -n 0.1 -dm 5 -Q2 -G);
[0116] 2. Separation of human genome data from metagenomic sequencing data: Using the bowtie2 sequence alignment tool, the clean reads after quality control filtering were aligned with the human reference genome hg38 to separate the mixed omics data while retaining the microbiome data and human genome data. A sam file recording the alignment information was obtained (adding the parameter --al-conc can obtain a matching sequence file).
[0117] The ratio of human genome data to microbiome data separated in this step is shown in Table 3, where the human genome data volume (number of reads) accounts for an average of 69.62% of the total sequencing data, and the microbiome data accounts for an average of 30.38% of the total sequencing data.
[0118] Table 3. Proportion of human genome data and microbiome data obtained by data splitting
[0119]
[0120]
[0121] 3. Alignment statistics: The flagstat function and coverage function of the samtools software, as well as the Qualimap software, were used to perform statistics on the coverage, depth, and alignment quality of the chromosomes in the sequence alignment results. The results are as follows: Figure 4 As shown, on average 93.47% of the reference genome was covered to at least 1×.
[0122] 4. Mutation detection in human genome data: Use the HaplotypeCaller function module in the GATK analysis tool to perform SNP calling mutation analysis on human genome data. The specific steps are as follows:
[0123] 1) Use the gatk MarkDuplicates tool to mark duplicate sequences in the bam file;
[0124] 2) Use samtools software to build an index for the bam file marked with repeated sequences;
[0125] 3) Use the picard.jar AddOrReplaceReadGroups tool to assign a new read-group label to all reads in the file;
[0126] 4) Use the gatk BaseRecalibrator tool to recalibrate the base quality scores of the output results of step 3;
[0127] 5) Use the gatk HaplotypeCaller tool to perform mutation detection on the output of step 4;
[0128] 6) Use the gatk SelectVariants tool to split the SNP and INDEL mutation types in the mutation detection result file respectively;
[0129] 7) Use the gatk VariantRecalibrator tool to recalibrate the mutation quality scores of the SNP file and INDEL file obtained in step ⑥ to obtain the calibrated variant detection VCF file;
[0130] 8) Annotation of human genome mutation detection data: Merge the SNP file and INDEL file from step 7 and use the VEP tool to annotate the mutation file;
[0131] 9) Quality assessment of mutation detection in human genome data: Use the bcftools tool to perform quality assessment of mutation detection.
[0132] The statistical table of mutation detection analysis of the human genome of 15 volunteers is shown in Table 4. Using the metagenomic sequencing data analysis method provided in this embodiment, mutation detection analysis of human genome data found that an average of 93.47% of the reference genome was covered to at least 1×, the average number of SNPs was 3258286, and the average number of Indels was 1033317. In addition, the average value of Ti / Tv, an indicator used to evaluate the accuracy of variation detection, was approximately 2.09 (reference range 2.0-2.1), which is within a reasonable range. This shows that the skin sampling method of the present invention and the analysis method for metagenomic sequencing data of skin samples can realize routine analysis of human genome polymorphisms.
[0133] Table 4. Statistics of gene mutation detection in human genome data
[0134] Sample ID ≥1× coverage Number of SNPs Number of indels Transition / Transversion DP8400025478BL_L01_90 93.6% 3,407,994 985,375 2.06 DP8400025973TL_L01_66 92.53% 2,743,160 700,931 2.11 FP100003143TL_L01_66 93.85% 3,175,234 1,124,901 2.10 FP100003144TL_L01_15 93.55% 3,376,788 1,150,941 2.11 FP100003144TL_L01_16 93.01% 3,150,654 1,057,959 2.14 FP100003144TL_L01_50 93.33% 3,236,520 1,080,454 2.12 FP100003148TL_L01_122 93.66% 3,363,167 1,039,206 2.07 FP100003148TL_L01_95 93.55% 3,257,856 1,003,076 2.08 FP100003266BL_L01_26 93.61% 3,333,939 1,160,073 2.09 FP100003466BR_L01_45 93.71% 3,170,157 993,745 2.09 FP100003466BR_L01_46 93.56% 3,371,732 1,021,525 2.07 FP100003466BR_L01_48 93.63% 3,422,616 1,046,413 2.07 FP100003466BR_L01_57 93.46% 3,258,225 1,020,960 2.08 FP100003467BR_L01_67 93.35% 3,189,880 1,045,914 2.09 FP100003467BR_L01_75 93.63% 3,416,372 1,068,286 2.06
[0135] 5. Microbiome Data Analysis:
[0136] 1) Call MetaPhlAn 4 to analyze the skin microbial species composition and generate a skin microbial species composition file;
[0137] 2) Use the Seqtk tool to randomly sample skin microbial metagenomic data to speed up downstream data comparison;
[0138] 3) Using BBmap alignment software and the human skin microbiome reference gene set (iHSMGC) as a reference, perform skin microbiome-specific functional gene annotation;
[0139] 4) Use self-written scripts to quantify the abundance of skin microbial functional genes and obtain a quantitative table of related functional genes. The main contents are as follows Figure 5 As shown in the figure, it mainly contains species information, functional information, and TPM values representing the relative abundance of genes. More in-depth analysis can be performed based on the above results, which will not be repeated here.
[0140] As can be seen from the above examples, the present invention achieves the ability to simultaneously obtain dual-omics data (microbiome data and human genome data) through a single sampling process. The analysis method of the present invention enables simultaneous acquisition of dual-omics data (microbiome data and human genome data) from a single facial skin sample, eliminating the need for separate sampling, thus saving costs and increasing efficiency.
[0141] In addition, the workflow of the present invention specifically integrates the reference gene set iHSMGC for the human skin microbiome in the skin microbial functional gene annotation analysis module to achieve accurate annotation of skin microbial habitat-specific functional genes. Therefore, the downstream microbiome data analysis part can obtain statistics on the composition of microbial species and annotation results of microbial functional genes related to the skin microenvironment. In addition to obtaining an integrity assessment report on the quality of skin-derived human genome data, the human genome data analysis part can also obtain mutation detection files, mutation result annotation files, and mutation detection evaluation files of human genes. The analysis results of the above-mentioned dual-omics combination can provide a reference basis for the study of the interaction between skin microorganisms and hosts, as well as the study of skin microorganisms in the occurrence, development, and evaluation of skin diseases.
[0142] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A method for analyzing metagenomic sequencing data of a sample, characterized in that: The following steps are involved: Obtaining sequencing data of the sample; Comparing the sequencing data with the host reference genome to obtain host genome data and microbiome data; The host genome data and the microbiome data are analyzed separately.
2. The method for analyzing metagenomic sequencing data according to claim 1, wherein: Analyzing the host genome data includes performing at least one of quality integrity assessment and mutation analysis on the host genome data.
3. The method for analyzing metagenomic sequencing data according to claim 2, wherein: Quality integrity assessment was performed using Qualimap and / or mutation analysis was performed using GATK.
4. The method for analyzing metagenomic sequencing data according to claim 1, wherein: Analyzing the microbiome data includes performing at least one of species composition annotation and functional gene annotation on the microbiome data.
5. The method for analyzing metagenomic sequencing data according to claim 4, wherein: Species composition annotation was performed using at least one of MetaPhlAn, mOTUs2, Mothur, Kraken, Kraken2, KarkenUniq, Megan, Bracken, Centrifuge, CLARK, CLARK-S, k-SLAM, MegaBLAST, metaOthello, PathSeq, prophyle, and taxMaps; and / or, The reference gene set for functional gene annotation includes iHSMGC.
6. The method for analyzing metagenomic sequencing data according to claim 1, wherein: The host is a human, and / or the sample is at least one of skin, saliva, lung, intestine, feces, and urine.
7. The method for analyzing metagenomic sequencing data according to claim 6, wherein: The host reference genome includes any one of hg18, hg19, and hg38.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the metagenomic sequencing data analysis method according to any one of claims 1 to 7.
9. An electronic device, characterized in that The method comprises a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the metagenomic sequencing data analysis method according to any one of claims 1 to 7 when running the computer program.
10. A sample metagenomic sequencing data analysis system, characterized in that: include: An acquisition module, configured to acquire sequencing data of the sample; The comparison module is used to compare the sequencing data with the host reference genome and split the host primary genomic and microbiome data; An analysis module is used to analyze the host genome data and the microbiome data respectively.
Citation Information
Patent Citations
Analysis method and system of metagenome data
CN108334750A
Method and device for microbiological analysis of host sample
CN111009286A
Metagenome sequencing data automatic analysis method
CN112289375A
Method and system for detecting microorganisms and drug-resistant genes in sample
CN112530519A