Ancient metagenomic fragment identification method based on paired-end sequencing and its application
The dual-end sequencing method addresses the challenges of ancient DNA analysis by enhancing efficiency and accuracy in identifying and analyzing ancient microbial communities through advanced data processing and alignment techniques.
Patent Information
- Application Number
- CN202411206536.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-30
AI Technical Summary
The prior art is difficult to efficiently and accurately process and analyze metagenomic fragments of ancient microbials, especially in the face of degradation and interference from modern pollutants, and traditional methods have limitations in microbial sample identification.
Using a dual-ended sequencing method, including sample processing, sequence data preprocessing, alignment and damage assessment, use specific software tools such as kneaddata, trimmomatic, flash, pmdtool and pydamage, to remove modern contamination and evaluate DNA damage, ensuring the purity and accuracy of the data.
It improves the data processing and analysis efficiency of ancient metagenomic fragments, can accurately reveal the genetic information and evolutionary history of ancient biological groups, and provides reliable data support for archaeology, genetics and environmental science.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of gene detection. Specifically, it particularly relates to a method for identifying ancient metagenomic fragments based on paired-end sequencing and its application. Background Art
[0002] Ancient metagenomic research aims to analyze the DNA in ancient biological remains to reconstruct their genomic information and reveal their genetic background and evolutionary history. These studies are of great significance for understanding ancient biodiversity, ecosystem changes, and human migration and cultural exchanges. However, ancient DNA research faces many unique challenges. Ancient DNA is often highly degraded due to environmental influences and usually has a lower concentration compared to modern DNA. In addition, ancient DNA samples are vulnerable to interference from modern contaminants, which further increases the difficulty of analysis.
[0003] Traditional methods for identifying ancient DNA mainly rely on alignment with modern genomes. These methods assess DNA damage and determine the age through the alignment results. However, since these techniques are mainly designed for human DNA, they show obvious limitations when dealing with non-human samples, especially microbial samples. Microbial populations have a high degree of diversity and a rapid evolutionary rate, and cannot be aligned using a standard reference genome like the human genome. This limits the application of traditional methods in the study of ancient microbial DNA and makes it difficult to accurately identify and analyze the composition and function of ancient microbial populations.
[0004] Therefore, developing an efficient and accurate data processing and analysis method is of great significance for the identification of ancient metagenomic fragments. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for identifying ancient metagenomic fragments based on paired-end sequencing and its application, which is a method for efficiently and accurately processing and analyzing ancient metagenomic fragment data to solve the deficiencies in the existing technology during the data processing and analysis process.
[0006] To achieve the above object, the present invention is realized through the following technical solutions:
[0007] A method for identifying ancient metagenomic fragments based on paired-end sequencing, comprising the following steps:
[0008] (1) Sample collection and processing: Extract samples from ancient biological remains, perform pretreatment, and construct a library for paired-end 150bp sequencing;
[0009] (2) Preprocessing of sequence data: First, use quality control software for data correction, select DNA fragments in the sequencing data, use paired-end merging software to perform paired-end merging on the selected sequencing data, and take the successfully merged fragments for subsequent analysis;
[0010] (3) Sequence alignment: Use DNA fragments to splice into contigs, and use the contigs as a reference to compare with the target DNA fragments;
[0011] (4) Damage assessment: Use ancient DNA assessment software to evaluate the contrast between the target fragment and the contig, select the damaged fragments of the target DNA fragment, and mark the damage rate of the target DNA fragment;
[0012] (5) Selection of ancient DNA: Select fragments with different damage rates as ancient DNA according to the damage rates of different ages.
[0013] Furthermore, in the step (1), rinse the surface of the sample with 3% hydrogen peroxide to remove modern DNA contamination; since ancient DNA fragments are short and easily damaged, conventional cell disruption methods (magnetic beads and sonication) are not used; use the MGIEasy PCRFree DNA Library Preparation Reagent Kit to construct a paired-end 150bp library for sequencing.
[0014] Furthermore, the step (2) specifically includes
[0015] 1) Data correction: Use quality control software such as kneaddata or trimmomatic to remove low-quality sequences and adapter contamination in the sequencing data to ensure the purity and reliability of the data;
[0016] 2) Read length filtering: Select DNA fragments in the sequencing data so that they do not exceed 150bp to ensure the purity of the data source; tools such as fastqc or seqkit can be used.
[0017] 3) Removal of modern contamination: Use paired-end merging software such as flash or pandaseq to perform paired-end merging on the selected sequencing data, and take the successfully merged fragments for subsequent analysis.
[0018] Furthermore, in the step (4), the ancient DNA assessment software can be pmdtool or pydamage, etc.
[0019] Application of the above method in archaeology, genetics and environmental science.
[0020] Compared with the prior art, the beneficial effects of the present invention are:
[0021] The present invention includes steps such as preprocessing, alignment, annotation, and functional analysis of sequence data. By efficiently and accurately processing and analyzing ancient metagenomic fragments, it is possible to reveal the genetic information of ancient biological populations and their evolutionary history, providing important data support and technical guidance for fields such as archaeology, genetics, and environmental science.
[0022] The method provided by the present invention can effectively improve the efficiency and accuracy of data processing and analysis in the identification of ancient metagenomic fragments. Through phylogenetic analysis, the genetic information of ancient biological populations and their evolutionary history can be comprehensively revealed, providing reliable data support and technical guidance for research in related fields.
[0023] The present invention can not only overcome the limitations of traditional methods, adapt to different types of ancient samples, but also provide more comprehensive genetic information, promoting the in-depth development of ancient metagenomic research. Through the development and application of new methods, researchers will be able to more accurately analyze the structure and function of ancient microbial communities. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 The pmdtool software identifies the characteristics of ancient DNA, where a is the damage rate at the 5' end of the DNA fragment and b is the damage rate at the 3' end of the DNA fragment.
[0025] Figure 2 The pydamage identifies the characteristics of ancient DNA.
[0026] Figure 3 The DNA length distribution of the selected DNA.
[0027] Figure 4 The pmdtool software identifies the characteristics of ancient DNA, where a is the damage rate at the 5' end of the DNA fragment and b is the damage rate at the 3' end of the DNA fragment.
[0028] Figure 5 Functional annotation of ancient DNA fragments. DETAILED DESCRIPTION OF THE INVENTION
[0029] The following further describes and explains the technical solutions of the present invention in conjunction with embodiments and drawings.
[0030] Example 1:
[0031] A method for identifying ancient metagenomic fragments of dental calculus samples from a Han Dynasty ancient tomb in Shandong includes the following steps:
[0032] 1. Sample collection and processing:
[0033] Obtain dental calculus samples from a Han Dynasty ancient tomb in Shandong. To maximize the preservation of ancient DNA, the samples are immediately transported and stored in a low-temperature environment after collection.
[0034] To remove modern contamination, the surface of the sample was rinsed with 3% hydrogen peroxide, and then the surface layer was removed using a mechanical scraping technique under sterile conditions to ensure minimal DNA contamination in the collected sample.
[0035] 2. DNA Extraction and Sequencing:
[0036] Ancient DNA was extracted from the dental calculus samples using an ancient DNA extraction method based on organic solvents (such as phenol-chloroform).
[0037] The MGIEasy PCRFree DNA Library Preparation Reagent Kit was used for library construction and paired-end sequencing, with a read length of 150 bp. To reduce sample loss, the starting DNA amount for library construction was strictly controlled in this step.
[0038] 3. Preprocessing of Sequence Data:
[0039] The FastQC software was used to perform a preliminary quality check on the raw data, and then the low-quality sequences and adapters were trimmed using the Trimmomatic software.
[0040] Read length filtering was performed using the Seqkit tool to ensure that the length of all retained fragments was within 150 bp. At the same time, the Bowtie2 software was used to align the sequences with the modern human genome to remove human DNA contamination.
[0041] The paired-end merging software flash was used for data merging, and only the successfully merged high-quality sequences were retained for subsequent analysis.
[0042] 4. Sequence Alignment:
[0043] The SPAdes assembly tool was used to assemble short fragments into contigs, and these were aligned with the ancient microbial genome reference database to preliminarily screen out possible ancient-source DNA fragments.
[0044] 5. Damage Assessment
[0045] The pmdtool software was used to assess the damage of the aligned DNA fragments, with particular attention paid to the damage rates at the 5' and 3' ends of the DNA fragments. As Figure 1 shown. The 5' end ( Figure 1 a) and 3' end ( Figure 1 b) of the selected fragment had a 1-bp damage rate of 0.13 and a 5-bp damage rate of 0.02. The damage curve was smooth and the damage rate was reasonable, conforming to the characteristics of ancient DNA.
[0046] Another ancient DNA identification software, pydamage, was used for cross-validation ( Figure 2) The results showed that the damage rate of the 1st bp at the 5' end of the selected fragments was 0.12, the damage rate of the 5th bp was 0.02, and the p value < 0.01, which was credible. The damage curve was smooth and the damage rate was reasonable, conforming to the characteristics of ancient DNA.
[0047] 6. DNA Quality
[0048] The quality of the obtained DNA fragments was evaluated using the fastqc software, as Figure 3 shown, most of the lengths were in the range of 90 - 140 bp, with a relatively reasonable distribution and longer preserved fragments.
[0049] Example 2:
[0050] An ancient metagenomic fragment identification method for dental calculus samples from Jiangsu ancient tombs, comprising the following steps:
[0051] 1. Sample collection and processing:
[0052] Dental calculus samples were obtained from Jiangsu ancient tombs. After preliminary screening, dental calculus with a smooth surface and likely to preserve more ancient DNA was selected.
[0053] To remove surface contamination, ultrasonic cleaning technology was used on the sample surface to remove contaminants through high-frequency vibration, and hydrogen peroxide was used for secondary cleaning.
[0054] 2. DNA extraction and sequencing:
[0055] Commercial QIAamp DNA Micro Kit was used for ancient DNA extraction. This method is suitable for the extraction of trace DNA and can obtain relatively complete DNA fragments from low-concentration samples.
[0056] MGIEasy PCRFree DNA Library Preparation Reagent Kit was used for library construction, and paired-end sequencing was performed with a read length of 150 bp.
[0057] 3. Pretreatment of sequence data:
[0058] Sequence quality control and adapter removal were performed using the kneaddata software to ensure the purity of the sequencing data. Contamination removal was also carried out, especially the removal of known host DNA. Subsequently, the seqkit tool was used to process the sequences, and only high-quality short fragments (not exceeding 150 bp) were retained for subsequent sequence alignment.
[0059] 4. Sequence alignment:
[0060] MetaSPAdes was used for scaffolding and assembly to generate high-quality contigs. These contigs were then aligned with the ancient microbial genome reference database to identify fragments of possible ancient origin.
[0061] 5. Damage assessment:
[0062] Damage assessment was carried out using the pmdtool software, with particular attention paid to the terminal base damage of DNA fragments. The damage rates at the 5' and 3' ends of DNA fragments were specifically focused on. As Figure 3 shown, the 1bp damage rate at the 5' and 3' ends of the selected fragments was 0.16, the 5bp damage rate was 0.05, and the 10bp damage rate was close to 0. The damage curve was smooth and the damage rate was reasonable, meeting the characteristics of ancient DNA.
[0063] 6. Functional gene analysis:
[0064] Seqkit was used to select the identified ancient DNA fragments and conduct in-depth functional gene analysis. In particular, functional annotation was carried out using the kegg tool to analyze the potential functions of these genes in the ancient microbial community. As Figure 4 shown, a variety of metabolic genes could be successfully identified and compared, and their abundances could be determined.
[0065] The present invention provides an efficient and accurate data processing and analysis method for identifying ancient metagenomic fragments. Through reasonable step design and advanced technical means, it overcomes many challenges in ancient DNA analysis, providing reliable data support and technical guidance for research in fields such as archaeology, genetics, and environmental science. This method can not only improve the efficiency and accuracy of ancient DNA research, but also be widely applied in multiple fields such as ancient anthropology, ecology, and etiology, with broad application prospects and important academic value.
[0066] Finally, although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. An ancient metagenomic fragment identification method based on paired-end sequencing, characterized in that, It includes the following steps: (1) Sample collection and processing: Extract samples from ancient biological remains, perform preprocessing, and construct a library for sequencing with 150bp paired-ends. (2) Preprocessing of sequence data: First, use quality control software for data correction, select DNA fragments in the sequencing data, use paired-end merging software to merge the selected sequencing data in pairs, and take the successfully merged fragments for subsequent analysis. The specific steps of step (2) include: 1) Data correction: Use quality control software such as kneaddata or trimmomatic to remove low-quality sequences and adapter contamination in the sequencing data to ensure the purity and reliability of the data. 2) Read length filtering: Select DNA fragments in the sequencing data so that they do not exceed 150bp to ensure the purity of the data source, using tools such as fastqc or seqkit. 3) Removal of modern contamination: Use paired-end merging software such as flash or pandaseq to merge the selected sequencing data in pairs, and take the successfully merged fragments for subsequent analysis. (3) Sequence alignment: Use DNA fragments to assemble into contigs, and use the contigs as a reference to compare with the target DNA fragments. (4) Damage assessment: Use ancient DNA assessment software to evaluate the contrast between the target fragment and the contig, select the damaged fragments of the target DNA fragment to be identified, and mark the damage rate of the target DNA fragment. (5) Selection of ancient DNA: Select fragments with different damage rates as ancient DNA according to the damage rates of different ages.
2. The method for identifying ancient metagenomic fragments according to claim 1, wherein In step (1), the sample surface is rinsed with 3% hydrogen peroxide to remove modern DNA contamination. Use the MGIEasy PCRFree DNA Library Preparation Reagent Kit for paired-end 150bp library construction and sequencing.
3. The ancient metagenomic fragment identification method according to claim 1, wherein In step (4), the ancient DNA assessment software is pmdtool or pydamage.
4. Application of the ancient metagenomic fragment identification method according to claim 1 in archaeology, genetics, and environmental science.
Citation Information
Patent Citations
Method for filtering modern DNA pollution from ancient DNA data and application thereof
CN110970086A
Use of dilute hydrogen peroxide to remove DNA contamination
US20070289605A1