A method for detecting full-length 16S rRNA amplicon nanopore sequencing in human semen samples
By employing nested primer combinations and customized sequencing filtering strategies in human semen sample testing, the problems of poor amplification specificity and inaccurate data cleaning were solved, enabling high-resolution microbiome detection, especially high-resolution identification of closely related bacteria in semen samples.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KUNMING UNIV OF SCI & TECH
- Filing Date
- 2026-04-30
- Publication Date
- 2026-06-02
Smart Images

Figure CN122128414A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of microbiome detection technology, and specifically relates to a method for detecting full-length 16S rRNA amplicon nanopore sequencing of human semen sample microbiome. Background Technology
[0002] Research on the human semen microbiome is of great significance for male reproductive health, infertility diagnosis, and assisted reproductive technology. However, human semen samples are typical low-biomass reproductive tract samples with relatively low microbial load and extremely high host-derived DNA background. Furthermore, they are more susceptible to exogenous contamination during collection, transportation, extraction, and amplification. Therefore, semen microbiome testing places high demands on sample collection standardization, nucleic acid extraction stability, PCR amplification specificity, nanopore library construction quality, and post-analysis data cleaning strategies.
[0003] Currently, traditional short-read 16S rRNA sequencing (such as sequencing based on the V3-V4 region) suffers from insufficient taxonomic resolution, making it difficult to accurately identify closely related bacteria at the genus and species levels. Nanopore long-read sequencing can cover longer 16S rRNA amplification regions, and each sequence carries a higher amount of phylogenetic information, which helps improve taxonomic resolution at both the genus and species levels. It also has application potential in identifying closely related bacteria, resolving complex communities, and preserving effective information in low-biomass human semen samples. However, the actual performance of this type of detection method depends not only on the sequencing platform itself, but also on the full-length 16S rRNA amplicon or nested PCR amplification strategy, library preparation process, NanoPlot data quality control, length and quality filtering, primer trimming, feature clustering, non-target sequence filtering, and the degree of matching with the reference database.
[0004] For special samples like semen, current technologies face the following bottlenecks: (1) Poor amplification specificity: High background host DNA often leads to non-specific amplification, making it difficult to obtain high-quality target bacterial sequences.
[0005] (2) Inaccurate data cleaning: The data from low biomass samples contains a large number of non-target reads, and conventional filtering strategies are prone to loss of effective information or false positive results.
[0006] (3) Lack of process standardization: There is a lack of a complete closed-loop process from standardized pretreatment, specific amplification to customized bioinformatics analysis, including: nucleic acid extraction, full-length 16S rRNA amplicon amplification or nested PCR, nanopore library construction, NanoPlot quality control, QIIME 2 analysis to classification annotation, making it difficult to obtain interpretable human semen microbiome detection results stably.
[0007] Therefore, developing a highly specific, high-resolution full-length 16S rRNA amplicon nanopore sequencing detection method for human semen samples is a technical problem that urgently needs to be solved in the field of reproductive microecology research. Summary of the Invention
[0008] To address the technical problems of complex background, susceptibility to contamination, low effective sequence retention rate, and poor data interpretability in low-biomass samples such as human semen in existing technologies, this invention provides a method for detecting full-length 16S rRNA amplicon nanopore sequencing of the human semen sample microbiome.
[0009] The present inventor's method for detecting full-length 16S rRNA amplicon nanopore sequencing in semen samples is as follows: Step S1, Sample collection and preprocessing: Obtain liquefied human semen samples and set up a negative control throughout the preprocessing process. The negative control is used for subsequent background subtraction, starting from the start of the sampling consumables process and proceeding synchronously with the real sample. Step S2, DNA extraction: Extract microbial genomic DNA from the pre-treated sample; Step S3, long amplicon nested PCR amplification: nested PCR amplification is performed using extracted microbial genomic DNA as a template, followed by quality control; The first round of amplification used near full-length universal outer primers for bacterial 16S rRNA, and the outer primer sequences were 16S-27F: AGAGTTTGATCCTGGCTCAG (SEQ ID NO:1) and 16S-1492R: GGTTACCTTGTTACGACTT (SEQ ID NO:2). The second round of amplification used the amplification products from the first round as templates, and specific amplification was performed using inner primers. The inner primer sequences were 16S-340F: TCCTACGGGNGGCAGCAGT (SEQ ID NO:3) and 16S-1390R: CGACGGGCGGTGWGTRCA (SEQ ID NO:4), where R represents A / G; target long amplicons with a length between 1.0kb and 1.1kb were obtained. The reaction conditions for the nested PCR amplification are as follows: First round of PCR: 95℃ pre-denaturation for 5 min; 95℃ for 30 s, 50℃ for 30 s, 72℃ for 2 min, for 30 cycles; 72℃ extension for 10 min; Second round of PCR: 95℃ pre-denaturation for 5 min; 95℃ for 30 s, 53℃ for 30 s, 72℃ for 2 min, for 35 cycles; 72℃ extension for 10 min; Step S4, Library Construction and Nanopore Sequencing: Purify the target long amplicon, construct a nanopore sequencing library and perform single-molecule long read sequencing to output raw sequencing data; Step S5, Data Depth Filtering: The raw sequencing data is filtered for length and quality. The filtering threshold is set as follows: quality score Q≥20 and read length between 1000bp and 1700bp. Step S6, Feature generation and non-target sequence removal: Remove primer sequences from the filtered data, perform de novo clustering with 99% consistency to generate representative sequences and amplicon features, and remove archaea, chloroplasts, mitochondria and unannotated non-target feature sequences; Step S7, Taxonomic Annotation: Taxonomic annotation is performed on the preserved feature sequences based on the reference database to obtain human semen sample microbiome data.
[0010] The detection method of the present invention can be applied in any of the following aspects: (1) Application in the detection of microbiome in semen samples from infertile men; (2) Application in the analysis of the composition and differences of human semen microecology; (3) Application in high-resolution taxonomic annotation of semen samples at the genus and species levels.
[0011] The beneficial effects of this invention are as follows: Innovative amplification primer design (high specificity): Addressing the challenge of extremely high host DNA background in semen, this invention abandons conventional single-amplification and innovatively employs a specific nested primer combination (outer primer 27F / 1492R combined with inner primer 340F / 1390R), perfectly avoiding non-specific host amplification regions and stably obtaining high-quality long amplicon of 1.0kb-1.1kb, thus solving the technical bottleneck of difficulty in obtaining target genes from low biomass samples.
[0012] Customized sequencing filtering strategy (high fidelity): Based on the product characteristics of the above-mentioned inner primers, this invention sets a highly targeted dual filtering threshold (Q≥20 and 1000bp≤read length≤1700bp) at the bioinformatics analysis end. Combined with the forced removal of non-target sequences, it successfully intercepted more than 50% of host background interference and sequencing chimeras, completely solving the problem of poor interpretability of semen sample data.
[0013] Breakthrough taxonomic resolution: Relying on effective long amplicon lengths of over 1.0kb and high-fidelity data cleaning, this invention breaks through the limitation of traditional short-read sequencing, which can only annotate down to the family / genus level, and successfully achieves high-resolution identification of closely related bacteria in semen samples at both the genus and species levels. Attached Figure Description
[0014] Figure 1 This is a flowchart of the full-length 16S rRNA amplicon nanopore sequencing detection process for human semen samples according to an embodiment of the present invention. Figure 2 This is a nested PCR electrophoresis detection diagram of 16S rRNA long amplicon in an embodiment of the present invention. In the diagram, 1-10 are the amplification results of 10 samples, and N is the negative control. Figure 3 This is a NanoPlot raw read length histogram according to an embodiment of the present invention; Figure 4 This is a graph showing the relationship between NanoPlot read length and average quality value in an embodiment of the present invention. Figure 5 This is a statistical output graph of NanoPlot based on read segment length according to an embodiment of the present invention; Figure 6 This is a diagram illustrating the composition of the genus-level classification annotation in an embodiment of the present invention. Figure 7 This is a graph showing the total frequency of the main genus-level classification units in an embodiment of the present invention. Detailed Implementation
[0015] The technical solution of the present invention will be further described in detail below through embodiments, but the content of the present invention is not limited thereto. Unless otherwise specified, the methods in this embodiment are conventional methods, and the materials and reagents used are obtained from commercial sources or prepared according to conventional methods unless otherwise specified. Example 1: Design and Amplification Method of Nested PCR Primers for 16S rRNA Long Amplicon Adapted to the Low Biomass Characteristics of Human Semen This invention addresses the "dual dilemma" of low microbial load and extremely high host DNA background in human semen samples by designing a dedicated 16S rRNA long amplified nested PCR primer combination and amplification strategy. In this embodiment, 10 representative human semen samples were selected from infertile populations to establish a full-length 16S rRNA amplified nested PCR amplification procedure for the human semen sample microbiome. Figure 1 ).
[0016] 1. Nested primer design Nested PCR amplification primers were designed based on the full-length bacterial 16S rRNA sequence. Primer sequence information is as follows: Forward direction of outer primer (16S-27F): AGAGTTTGATCCTGGCTCAG (SEQ ID NO:1) Reverse outer primer (16S-1492R): GGTTACCTTGTTACGACTT (SEQ ID NO:2) Forward inner primer (16S-340F): TCTACGGGNGGCAGCAGT (SEQ ID NO:3) Inner primer reverse (16S-1390R): CGACGGGCGGTGWGTRCA (SEQ ID NO:4), (where R indicates A / G) 2. Nested PCR amplification process (1) Standardized sample collection and preprocessing Male subjects collected semen samples after abstinence for 2 to 7 days, following WHO semen analysis guidelines. Semen samples were collected via masturbation into sterile sampling cups, liquefied for 30 minutes at room temperature or 37°C, and then aliquoted. One portion was used for routine semen analysis, and the remainder for microbiome testing.
[0017] Contamination prevention details: Filter-equipped pipette tips are used throughout the entire pretreatment process. Tips are replaced promptly between samples, and zonal operations are implemented. Considering that low-biomass samples are susceptible to exogenous contamination, this embodiment strictly includes a negative control throughout the entire process. The negative control is used from the moment sampling consumables are opened or the preservation solution is dispensed, and it proceeds along with the real samples into subsequent extraction, amplification, and library construction steps. Samples that cannot be processed immediately are stored at -80°C, and repeated freeze-thaw cycles are avoided.
[0018] (2) Extraction and quality control of microbial genomic DNA Before extraction, equilibrate the semen sample to room temperature and mix thoroughly. Add 200 μL of the sample to a pre-loaded reagent deep-well plate, along with 20 μL of the accompanying Proteinase K. Place the deep-well plate and magnetic bead sleeve into the TGuide S32 automated extractor and run the preset program to complete lysis, magnetic bead binding, rinsing, and elution. Transfer the eluted product to a new nuclease-free centrifuge tube and store for short-term storage at -20°C and long-term storage at -80°C.
[0019] After DNA extraction, fluorescence quantification was performed using the Qubit dsDNA HS Assay Kit, and nucleic acid purity was assessed using NanoDrop. If necessary, a small amount of DNA was taken for agarose gel electrophoresis to observe nucleic acid integrity and fragment distribution.
[0020] (3) Nested PCR amplification First round of amplification: Using a near full-length universal primer pair for bacterial 16S rRNA (16S-27F and 16S-1492R), trace amounts of bacterial nucleic acid were specifically enriched from the massive amount of host DNA. First round of PCR: 95℃ pre-denaturation for 5 min; 30 cycles of 95℃ for 30 s, 50℃ for 30 s, and 72℃ for 2 min; extension at 72℃ for 10 min. Theoretical product length: approximately 1.5 kb.
[0021] Second-round amplification: Using the first-round product as a template, secondary specific amplification was performed using the inner primer pair (16S-340F and 16S-1390R). Second-round PCR conditions: Using the first-round product dilution as a template, pre-denaturation was performed at 95℃ for 5 min; 35 cycles were performed at 95℃ for 30 s, 53℃ for 30 s, and 72℃ for 2 min; followed by an extension at 72℃ for 10 min. This inner primer pair underwent rigorous screening, and its amplified products not only avoided non-specific banding regions easily generated in semen samples but also perfectly matched the read length preferences for nanopore sequencing. Effective target product length: approximately 1.0 kb to 1.1 kb.
[0022] (4) Results Analysis Nested PCR amplification results are as follows Figure 2 As shown, using the nested primers and amplification conditions of this invention, target long amplicon with a clean background, single band, and size between 1.0kb and 1.1kb was stably obtained from all 10 samples, completely breaking through the technical bottleneck of obtaining the target gene from semen samples.
[0023] Example 2: Nanopore sequencing method based on long amplicon sequencing and customized data cleaning strategy For the 1.0kb-1.1kb long amplicon obtained in Example 1, a highly adapted nanopore single-molecule long read sequencing method and a data cleaning strategy were established to ensure the authenticity and high fidelity of low biomass data.
[0024] 1. Purification of amplified products, construction of nanopore libraries, and sequencing The amplification products obtained in Example 1 were purified using magnetic beads for samples with single bands and low background; for semen samples with many impurities or complex background, they were recovered by gel excision before library construction. The purified PCR products were subjected to library quantification, and DNA repair, universal adapter ligation, post-ligation processing, magnetic bead purification, and sequencing complex preparation were performed according to the requirements of the universal sample preparation kit. The purified products were loaded onto the G-seq 500 nanopore sequencing platform, and sequencing was completed using the G-TR01 sequencing kit and G-MK02 microfluidic chip. Single-end FASTQ files were output after sequencing.
[0025] 2. NanoPlot data quality control NanoPlot was used to generate quality control charts from the original FASTQ files of 10 samples. The raw data characterization results are as follows: Figure 3 , Figure 4 and Figure 5NanoPlot quality control statistics (NanoStats) showed that the number of raw reads in the 10 samples was 66,507, the total number of bases was 82,617,566, the average read length was 1242.2 bp, and the N50 was 1133.0 bp. This N50 value is highly consistent with the 1.0kb-1.1kb target product designed with inner primers in Example 1, proving that the sequencing process accurately captured the target long fragment.
[0026] 3. Length filtering, quality filtering, and primer trimming NanoFilt was used for length and quality filtering, with strict thresholds set as follows: Q≥20 and 1000bp≤read length≤1700bp. The filtered sample files generated manifest.csv and were imported into the QIIME 2 platform in SingleEndFastqManifestPhred33 format.
[0027] Then, the amplification inner primer and its reverse complementary sequence were removed using q2-cutadapt trim-single with the parameters set as follows: error-rate=0.20, overlap=16, and discard-untrimmed.
[0028] The total number of reads after filtering was 25,444, with an average read count of 2,544.4, an average read length of 1,127.7 bp, and an average Q value that significantly increased to 44.4; the total number of reads after primer trimming was 25,375.
[0029] Technical Results: After the aforementioned customized sequencing and cleaning strategies, the total frequency of the final feature table stabilized at 12376, representing a median sequence length of 1077 bp. This indicates that the sequencing method successfully transformed the original noisy signal into high-purity, high-confidence semen microbiome characteristic data.
[0030] Example 3: Taxonomic annotation and application analysis of semen microecology based on long amplicon nanopore sequencing This embodiment demonstrates the specific bioinformatics implementation path and application results of the above core methods in actual clinical samples.
[0031] 1. Feature generation, non-target sequence filtering, and taxonomic annotation Continuing from the sequence of quality filtering and trimming in Example 2: (1) The filtered sequences are deduplicated using q2-vsearch dereplicate-sequences, and then de novo clustering is performed with q2-vsearch cluster-features-de-novo at 99% consistency to generate representative sequences and amplicon features.
[0032] (2) Using the classification annotation results, archaea, chloroplasts, mitochondria and unannotated features are removed by q2-taxa filter-table and q2-taxa filter-seqs to complete the background subtraction of non-target sequences.
[0033] (3) Taxonomic annotation was completed on the QIIME 2 platform. The reference database was constructed based on SILVA version 138. The reference sequence was subjected to abnormal sequence removal, length filtering and redundancy removal in combination with the RESCRIPT process. Then, the feature-classifier classify-sklearn was used to perform taxonomic annotation on the feature sequence.
[0034] 2. Application Analysis of Basic Data Output In this embodiment, all 10 samples were included in the final feature table. The total frequency of the final feature table was 12376, the median sequencing depth was 1194.0, the number of representative sequences was 5764, and the median length of the representative sequences was 1077 bp. This result provides high-quality underlying data for subsequent genus-level and species-level taxonomic annotation.
[0035] 3. Representation and Application of Classification Annotation Results Based on the level-6 (genus-level) classification annotation results exported from taxonomy-filtered-barplot.qzv, the feature frequencies of 10 samples were summarized and their relative abundances were calculated. The major genera with the highest total frequency were selected as the subjects of the illustration, and the rest were merged into other taxonomic units.
[0036] Analysis results as follows Figure 6 and Figure 7 As shown, the detection results of this invention clearly reconstruct the microecological composition of human semen samples: accurately identifying the core semen flora including Lactobacillus, Corynebacterium, Staphylococcus, etc., and intuitively demonstrating the significant differences in the abundance and composition of microorganisms among semen samples from different infertile patients.
[0037] The results of this application confirm that this method breaks through the limitations of traditional short-read sequencing and successfully achieves high-resolution identification of closely related bacteria at the genus and species levels in semen samples, providing strong technical support for comparative analysis of the reproductive tract microecology in infertile individuals.
Claims
1. A method for detecting full-length 16S rRNA amplicon nanopore sequencing of the microbiome in human semen samples, characterized in that, Includes the following steps: S1, Obtain a liquefied human semen sample and extract microbial genomic DNA from the sample; S2, using extracted microbial genomic DNA as a template, performs long amplicon nested PCR amplification, followed by quality control; The first round of amplification used the near full-length universal outer primers 16S-27F and 16S-1492R for bacterial 16S rRNA; the second round of amplification used the amplification product from the first round as a template and the inner primers 16S-340F and 16S-1390R for specific amplification to obtain target long amplicon with a length between 1.0kb and 1.1kb. S3, Library Construction and Nanopore Sequencing: Construct nanopore sequencing libraries from purified target long amplicones and perform single-molecule long read sequencing to output raw sequencing data; S4, Data Depth Filtering: Filter the raw sequencing data by length and quality. The filtering threshold is set as follows: quality score Q≥20 and read length between 1000bp and 1700bp. S5, Feature Generation and Taxonomic Annotation: Primer sequences are removed from the filtered data, feature clustering is performed and non-target feature sequences are eliminated. Then, taxonomic annotation is performed on the retained feature sequences based on the reference database to obtain human semen sample microbiome data.
2. The detection method according to claim 1, characterized in that: The sequence of the outer primer 16S-27F is shown in SEQ ID NO:1; the sequence of 16S-1492R is shown in SEQ ID NO:2; the sequence of the inner primer 16S-340F is shown in SEQ ID NO:3; and the sequence of 16S-1390R is shown in SEQ ID NO:
4.
3. The detection method according to claim 1, characterized in that: In step S1, a negative control is set up to be activated by the self-sampling consumables and operated synchronously with the real sample for subsequent background subtraction.
4. The detection method according to claim 1, characterized in that: In step S5, the feature clustering is de novo clustering with 99% consistency; non-target feature sequences include archaea, chloroplasts, mitochondria, and unannotated feature sequences.
5. The detection method according to claim 1, characterized in that: It is applied to the detection of microbiome in semen samples from infertile men, analysis of the composition and differences of human semen microecology, and high-resolution taxonomic annotation of semen samples at the genus and species levels.