Multiplexed pathogen identification method and device based on targeted nanopore sequencing

CN122879412APending Publication Date: 2026-10-09THE 1ST AFFILIATED HOSPITAL OF SHIHEZI UNIVERSITY +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611130371.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0003]针对现有技术存在的问题,本发明提供一种基于靶向纳米孔测序的多重病原体鉴定方法及装置,以解决现有宏基因组分类工具依赖k-mer算法,无法有效区分布鲁氏菌与苍白杆菌属等近缘病原体

Benefits of technology

[0015](1)本发明明确摒弃宏基因组分类工具,完全依赖靶标特异性比对结果进行物种判定,通过将扩增子引物序列软剪切标记后,利用标准深度统计工具计算靶标区域覆盖度,排除了引物区域对覆盖度评估的干扰,显著降低了布鲁氏菌与苍白杆菌属等近缘种、金黄色葡萄球菌与表皮葡萄球菌等常见环境菌的误报率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122879412A_ABST
    Figure CN122879412A_ABST
Patent Text Reader

Abstract

The application provides a kind of multiplex pathogen identification method and device based on targeted nanopore sequencing, constructs the Panel of multiplex PCR primer, carries out multiplex PCR amplification to multiple species sample based on the primer design parameter set of Panel, obtains original electrical signal data through nanopore sequencing after, the target sequence is obtained by pre-processing original electrical signal;The target sequence is compared to the target gene reference sequence set, and the bases of the two end soft cutting regions in the comparison result are cut off, the sequencing depth and coverage of each target region are calculated, and the read number is calculated according to the sequencing depth;In the case where any target region in the two target regions of the same species meets the first condition, the sample of the species is used as a preliminary screening sample;In the case where the two target regions of the same species in the preliminary screening sample both meet the second condition, the species is used as the identification result of multiple species sample.The application can well distinguish multiple species and improve the accuracy of multiple species identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics technology, and in particular to a method and apparatus for identifying multiple pathogens based on targeted nanopore sequencing. Background Technology

[0002] High-throughput targeted sequencing (tNGS) technology amplifies specific target regions via multiplex PCR, significantly reducing host interference and shortening the detection cycle. However, existing targeted detection protocols mostly design primers based on universal marker genes such as 16S rRNA or ITS, and combine them with drug resistance gene detection. This approach still cannot overcome the technical bottleneck of distinguishing closely related species. Furthermore, its analytical process relies on targeted sequence identification after primer alignment, unique read calculation, and batch contamination removal. Summary of the Invention

[0003] To address the problems existing in the prior art, this invention provides a method and apparatus for multiplex pathogen identification based on targeted nanopore sequencing, in order to solve the problem that existing metagenomic classification tools rely on the k-mer algorithm and cannot effectively distinguish between Brucella and closely related pathogens such as Paleobacterium.

[0004] This invention provides a method for multiplex pathogen identification based on targeted nanopore sequencing, comprising:

[0005] A panel of multiplex PCR primers was constructed. Based on the primer design parameter set of the panel, multiplex PCR amplification was performed on samples from multiple species. Raw electrical signal data was obtained by nanopore sequencing. The raw electrical signals were preprocessed to obtain the target sequence.

[0006] The target sequence is aligned to the target gene reference sequence set. After removing the soft-sheared bases at both ends of the alignment result, the sequencing depth and coverage of each target region are calculated, and the number of reads is calculated based on the sequencing depth.

[0007] If either target region of the same species meets the first condition in two target regions, the sample of that species is used as the initial screening sample; if both target regions of the same species in the initial screening sample meet the second condition, the species is used as the identification result of the multi-species sample.

[0008] The first condition is that the number of read segments is greater than or equal to a first preset threshold, and the coverage is greater than or equal to a second preset threshold; the second condition is that the number of read segments is greater than or equal to a third preset threshold, and the coverage is greater than or equal to a fourth preset threshold, wherein the third preset threshold is greater than the first preset threshold, and the fourth preset threshold is greater than or equal to the second preset threshold.

[0009] This invention also provides a multiplex pathogen identification device based on targeted nanopore sequencing, comprising:

[0010] The preprocessing module is used to construct a panel for multiplex PCR primers. Based on the primer design parameter set of the panel, multiplex PCR amplification is performed on samples of multiple species. Raw electrical signal data is obtained by nanopore sequencing. The raw electrical signal is then preprocessed to obtain the target sequence.

[0011] The calculation module is used to align the target sequence to the target gene reference sequence set, remove the soft-sheared bases at both ends of the alignment result, calculate the sequencing depth and coverage of each target region, and calculate the number of reads based on the sequencing depth.

[0012] The identification module is used to identify a sample of a species as a preliminary screening sample if either of the two target regions of the same species meets a first condition; and to identify the species as the identification result of the multi-species sample if both target regions of the same species in the preliminary screening sample meet a second condition.

[0013] The first condition is that the number of read segments is greater than or equal to a first preset threshold, and the coverage is greater than or equal to a second preset threshold; the second condition is that the number of read segments is greater than or equal to a third preset threshold, and the coverage is greater than or equal to a fourth preset threshold, wherein the third preset threshold is greater than the first preset threshold, and the fourth preset threshold is greater than or equal to the second preset threshold.

[0014] The method and apparatus for multiplex pathogen identification based on targeted nanopore sequencing provided by this invention have the following main advantages:

[0015] (1) This invention explicitly abandons metagenomic classification tools and relies entirely on target-specific alignment results for species determination. By soft-splitting and labeling the amplicon primer sequence, the target region coverage is calculated using standard depth statistics tools, eliminating the interference of primer regions on coverage assessment and significantly reducing the false alarm rate of Brucella and closely related species such as Paleobacterium, as well as common environmental bacteria such as Staphylococcus aureus and Staphylococcus epidermidis.

[0016] (2) The amount of targeted sequencing data in this invention is much lower than that of mNGS (only 0.1~0.5 M reads are needed, which is only one percent of the amount of conventional mNGS data). Under the same hardware conditions, the analysis time of the entire process can be controlled within a few minutes, which is significantly lower than the several hours of the mNGS process.

[0017] (3) This invention is the first to incorporate brucellosis arthritis and periprosthetic joint infection into the same targeted sequencing detection system. It simultaneously detects Brucella and four common PJI pathogens through a single tube of multiplex PCR reaction. The dual-target grading judgment strategy can complete the identification within a few hours, providing a high-confidence etiological basis for the accurate diagnosis and treatment of joint diseases in endemic areas. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the multiplex pathogen identification method based on targeted nanopore sequencing provided by the present invention.

[0020] Figure 2 This is a flowchart of sequencing data preprocessing, comparison analysis and judgment in the multiplex pathogen identification method based on targeted nanopore sequencing provided by the present invention;

[0021] Figure 3 This is a schematic diagram of the structure of the multiplex pathogen identification device based on targeted nanopore sequencing provided by the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0023] The following is combined with Figure 1 This invention describes a method for multiplex pathogen identification based on targeted nanopore sequencing, comprising:

[0024] Step 101: Construct a panel of multiplex PCR primers. Based on the primer design parameter set of the panel, perform multiplex PCR amplification on samples of multiple species, obtain raw electrical signal data by nanopore sequencing, and preprocess the raw electrical signal to obtain the target sequence.

[0025] Multi-species samples may include blood, joint fluid, and bone marrow aspiration fluid from patients with brucellosis arthritis, as well as joint fluid, periprosthetic tissue, or ultrasonically vibrated fluid from patients with periprosthetic joint infection.

[0026] Multiple species may include Brucella spp. and common pathogens of periprosthetic joint infections (Propionibacterium acnes, Pseudomonas aeruginosa, Staphylococcus aureus, and Staphylococcus epidermidis), and may also include other species.

[0027] Step 102: Align the target sequence to the target gene reference sequence set, remove the soft-sheared bases at both ends of the alignment result, calculate the sequencing depth and coverage of each target region, and calculate the number of reads based on the sequencing depth;

[0028] High-quality target sequences were aligned to the target gene reference sequence set using Minimap2 to suppress end seeding penalties and exclude minor alignments. After sorting and indexing the alignment results using SAMtools, BamClipper was used to remove soft-clipped bases from both ends of the alignment results to ensure that coverage calculations were based only on real biological alignment segments. SAMtoolsdepth was used to calculate the sequencing depth and base coverage ratio of each target region.

[0029] Step 103: If either of the two target regions of the same species meets the first condition, the sample of that species is used as the initial screening sample; if both target regions of the same species in the initial screening sample meet the second condition, the species is used as the identification result of the multi-species sample.

[0030] The first condition is that the number of read segments is greater than or equal to the first preset threshold (1 segment) and the coverage is greater than or equal to the second preset threshold (90%); the second condition is that the number of read segments is greater than or equal to the third preset threshold (5 segments) and the coverage is greater than or equal to the fourth preset threshold (90%), wherein the third preset threshold is greater than the first preset threshold and the fourth preset threshold is greater than or equal to the second preset threshold.

[0031] The flowchart for sequencing data preprocessing, alignment analysis, and judgment is as follows: Figure 2 As shown, species identification is based solely on the results of Minimap2 alignment to the target gene reference sequence set, without performing taxonomic annotation steps based on k-mer frequency, Markov models, or the least common ancestor algorithm. It relies solely on target-specific alignment results to distinguish closely related species.

[0032] Each target region is determined independently. If any target meets the above threshold, the species identification information for the corresponding species is output. Brucella species identification information is used for the etiological detection of Brucella infection, while Staphylococcus aureus, Staphylococcus epidermidis, Propionibacterium acnes, or Pseudomonas aeruginosa species identification information is used for the etiological detection of infection by the corresponding pathogens. A standardized detection report is finally output, covering species identification conclusions, sequencing depth and coverage ratio matrix for each target, key quality control parameters (total number of reads, target read ratio, etc.), and an identification display heatmap.

[0033] For example, a multiplex PCR primer panel was prepared based on the optimal primer design parameter set (panel_016). The following procedure was performed on the plasmid simulation sample and the corresponding results were obtained:

[0034] (1) Simulated sample preparation

[0035] Specific target fragments (approximately 400 bp each) of Brucella spp. (BCSP31 / IS711), Staphylococcus aureus (spa / femA), Staphylococcus epidermidis (tuf / sesC), Propionibacterium acnes (camp1 / sodA), and Pseudomonas aeruginosa (ecfX / gyrB) were cloned into the pUC57 plasmid vector, and the sequence correctness was verified by Sanger sequencing. High-purity plasmid DNA was extracted using a plasmid miniprep kit, and the concentration was determined by NanoDrop and converted to copy number.

[0036] Single-species simulated samples: each plasmid was individually diluted to 10-1 6 10 4 10 2 copies / μL;

[0037] Mixed simulated sample: Brucella, Aristolochic Acid (closely related genus control), Staphylococcus aureus and Staphylococcus epidermidis, equimolarly mixture of four plasmids (10 oz each). 4 (copies / μL)

[0038] Negative control: empty pUC57 plasmid (10 6 (copies / μL).

[0039] (2) Multiplex PCR amplification and nanopore sequencing

[0040] The validated panel was used to amplify the samples, which were then sequenced using a carbon nanopore sequencing platform. In this embodiment, the actual sequencing data per sample was 203,050 raw reads.

[0041] (3) Data preprocessing

[0042] After filtering with NanoFilt, sequences with a read length of 300-550 bp and Q≥7 were retained; chimeras were removed using VSEARCH (–chimeras_length_min 300). The actual quality control results are shown in Table 1.

[0043] Table 1. Actual Quality Control Results

[0044]

[0045] The targeted readings accounted for 88.97%, and the chimera ratio was <0.01%, meeting the quality control standards.

[0046] (4) Comparison and coverage analysis

[0047] Minimap2 was aligned to the target gene reference sequence set (-ax map-ont –end-seed-pen 5 –secondary=no), and BamClipper was used for soft excision followed by SAMtools calculation of coverage. The detection results for each target are shown in Table 2.

[0048] Table 2 Target Specificity Detection Results

[0049]

[0050] (5) Results of dual-target classification

[0051] According to Table 3:

[0052] Confirmation results: Brucella (gene 1: 318 reads / 100% coverage; gene 9: 15,037 reads / 100% coverage), Staphylococcus epidermidis (gene 2: 26,161 reads / 100% coverage; gene 5: 31,917 reads / 100% coverage), Propionibacterium acnes (gene 3: 16,627 reads / 100% coverage; gene 10: 12,706 reads / 100% coverage), Staphylococcus aureus (gene 6: 24,959 reads / 100% coverage; gene 7: 44,714 reads / 100% coverage), both targets reached the confirmation threshold (≥5 reads and ≥95% coverage), and confirmation identification information was output.

[0053] Suspected signal: Pseudomonas aeruginosa (gene 4: 8,222 reads / 100% coverage; gene 8: 0 reads / 0% coverage). Only one target meets the confirmation threshold, and the other target is negative. It is marked as a suspected signal and a retest is recommended.

[0054] Negative target: No reads were detected in gene 8 (Pseudomonas aeruginosa target 2), with 0% coverage.

[0055] Table 3 Species identification results

[0056]

[0057] Brucella, Staphylococcus epidermidis, Propionibacterium acnes, and Staphylococcus aureus were detected in the sample in this embodiment, with Pseudomonas aeruginosa as a suspected signal.

[0058] This embodiment abandons the metagenomic classifier and directly uses the Minimap2 target-specific alignment results as the sole criterion for judgment. It adopts a dual-target hierarchical strategy to improve the confidence of the judgment, thus fundamentally avoiding the problem of confusion between closely related species in the k-mer algorithm.

[0059] Based on the above embodiments, the preprocessing described in this embodiment includes one or more of base recognition, quality filtering, host sequence removal, and chimera removal;

[0060] The quality filtering includes one or more of the following: removing reads containing adapter sequences and primer dimers, eliminating sequences with an average quality value less than a fifth preset threshold (e.g., 7), and read length filtering.

[0061] Basecalling was performed using base recognition software (such as QPreasy), and FASTQ format sequences were output.

[0062] NanoFilt was used for quality filtering to remove reads containing adapter sequences and primer dimers, and to remove sequences with an average quality value Q less than 7. The read length filtering threshold was set to 300~550 bp based on the theoretical length range of the target amplicon.

[0063] Given that chimeric sequences are easily generated during multiplex PCR amplification, the –uchime_ref mode of VSEARCH was used for chimera detection and removal, with the parameter –chimeras_length_min set to 300 to obtain high-quality target sequences.

[0064] Based on the above embodiments, this embodiment, when either target region of the same species meets the first condition, after using the sample of that species as the initial screening sample, further includes:

[0065] If at least one target region of the same species in the initial screening sample does not meet the second condition, the species is marked as a suspected signal and a retest is requested.

[0066] Based on the above embodiments, this embodiment, after using the species as the identification result of the multi-species sample, further includes:

[0067] An identification report is generated based on the identification results, the target coverage matrix table, and the target coverage visualization heatmap.

[0068] Based on the dual-target grading results, a pre-set report template is automatically matched to generate a standardized identification report that includes species identification conclusions, a target coverage matrix table, and a visual heatmap of target coverage. The identification report outputs species identification information for confirmed positive samples, marks suspected signal samples with retest prompts, and is output in PDF or HTML format.

[0069] Based on the above embodiments, the panel for constructing multiplex PCR primers in this embodiment includes:

[0070] Genomes of multiple species were obtained from public genome databases, and a genome database was constructed after preprocessing the genomes.

[0071] Reference sequences of validated species-specific target genes for each species are retrieved from public nucleic acid databases. Homologous sequence searches are performed on the genome database using the reference sequences to obtain a set of homologous sequences. Multiple sequence alignment is then performed on the set of homologous sequences to obtain a target gene sequence database.

[0072] The target gene sequence database is subjected to CD-HIT-EST clustering. After merging the clusters, multiple sequence alignment is performed. The two ends of the multiple sequence alignment results are pruned to obtain the pruned alignment file.

[0073] The DegePrime software was used to perform degenerate primer scanning on the trimmed alignment file to generate a candidate degenerate primer set. The MFEprimer was used to perform a three-level thermodynamic evaluation on the candidate degenerate primer set to generate a risk primer pair set. The candidate degenerate primer set was then filtered using a three-rule combination.

[0074] The filtered candidate degenerate primer set was batch BLAST aligned to the NCBI RefSeq representative prokaryotic genome database. A two-level mismatch filtering was applied. Primer pairs that passed the two-level mismatch filtering were predicted. A four-level judgment was performed based on the taxonomic information of the hit species. Primer pairs were selected into the qualified pool according to the judgment results.

[0075] Based on the set of risky primer pairs, a greedy algorithm is used to select primer pairs from the qualified pool to construct a panel.

[0076] For Brucella spp., species-specific genes (such as BCSP31) validated in the literature were selected as target gene reference sequences. Using these reference sequences as query sequences, a homologous sequence search was performed in a constructed high-quality Brucella genome database (100 genomes) using MEGABLAST.

[0077] The parameters for homologous sequence search can be set as follows: sequence identity (pident) greater than or equal to 90%, query coverage (qcov_hsp_perc) greater than or equal to 85%, and expected value (evalue) less than or equal to 1e-6, thereby obtaining a set of highly similar homologous sequences. Multiple sequence alignment is then performed on the homologous sequence set using MAFFT or ClustalW algorithms to construct a target gene sequence database containing strain-level genetic diversity.

[0078] This strategy uses known specific target genes as anchor points to ensure species-level specificity of subsequent primer design; at the same time, by scanning conserved regions in the homologous sequence set, it fully preserves the natural variation sites between strains, making the designed primers both strain-level inclusive.

[0079] After obtaining the homologous sequence set, 46 non-redundant homologous sequences were obtained through CD-HIT-EST redundancy removal. Multiple sequence alignment was then performed using MAFFT (v7.490). The total length after alignment was 1064 bp (including inter-sequence insertion and deletion sites). When performing CD-HIT-EST clustering, a sequence consistency threshold of 0.95 within each cluster can be set to group highly similar sequences together, thereby reducing redundancy.

[0080] DegePrime automatically generates candidate degenerate primer sequences based on the site frequency matrix of multiple sequence alignments, under the constraints of coverage greater than or equal to 0.8 and degeneracy less than or equal to 16. The candidate degenerate primer set is then subjected to parallel scanning of multiple lengths, ranging from 18 to 24 bp, to obtain a multi-length candidate pool, providing ample candidate space for subsequent panel optimization.

[0081] DegePrime's scanning parameters also include a coverage depth greater than or equal to 3, skipping the boundary regions of 10 bp at both ends of the alignment to avoid primers falling into low-quality regions, and 100 iterations to ensure sufficient search.

[0082] The MFEprimer was used to perform a three-level thermodynamic evaluation of the candidate degenerate primer set, including Tm and ΔG calculation, heterodimer detection and hairpin structure detection, to generate a risk primer pair set.

[0083] Tm and ΔG calculations were performed using preset salt concentration conditions: 50 mmol / L for monovalent cations, 1.5 mmol / L for divalent cations, 0.25 mmol / L for dNTPs, and 50 nmol / L for primers. Heterodimer detection employed configurable score thresholds, mismatch thresholds, and dG thresholds, reporting only dimer pairs exceeding the thresholds. Hairpin structure detection utilized configurable score thresholds, dG thresholds, Tm thresholds, minimum hairpin stem length, minimum hairpin loop length, maximum hairpin loop length, and maximum hairpin gap.

[0084] Thermodynamic evaluation is compatible with degenerate base expansion variants. The IUPAC full mapping table is used to expand the degenerate bases R, Y, S, W, K, M, B, D, H, V, N into a definite base combination. Thermodynamic parameters are calculated for each variant, and the most unfavorable value is taken as the basis for risk assessment of the degenerate primer.

[0085] A three-rule combination filter was applied to the candidate degenerate primer set to reduce the risk of nonspecific extension caused by 3' end mismatch:

[0086] The first rule requires a definite GC clamp at the 3' end: the last base at the 3' end must be a definite G or C, degenerate bases are not allowed, and the stability of the extension initiation must be ensured.

[0087] The second rule is a potential GC density limit within a 5-base window: the number of potential GC bases in the last 5 bases at the 3' end is less than or equal to 3; the potential GC bases include G, C and all degenerate bases that may contribute to G / C (S, R, Y, K, M, B, D, H, V, N), and a conservative strategy is used for counting;

[0088] The third rule is the ban on repetitive sequences: four or more consecutive identical bases are prohibited, as are four or more consecutive dinucleotide repetitions (i.e., eight or more consecutive dinucleotide repetitions).

[0089] The filtered hit results were used for three-dimensional PCR product prediction based on orientation, coordinates, and length. Only results conforming to the F+R- or F-R+ orientation and with product lengths between 50 and 3000 bp were retained. A four-level judgment was performed based on the taxonomic information of the hit species: no hit (NO_TARGET), hit in a non-target genus (FAIL_GENUS), hit in a target genus but not a target species (PASS_GENUS), and hit in a target genus and target species (PASS_SPECIES). Only primer pairs at the PASS_GENUS and PASS_SPECIES levels were retained and entered into the qualified pool.

[0090] The two-level mismatch filtering includes: the sum of the total number of mismatched bases and the total number of gaps is less than or equal to the sum of the total number of gaps and is less than or equal to the third preset threshold (e.g., 3); the number of mismatched bases in the last 5 bases of the 3' end is less than or equal to the fourth preset threshold (e.g., 1) and the number of gaps is 0.

[0091] A four-level determination is performed based on the taxonomic information of the matched species, including:

[0092] FAIL_GENUS: The hit species belongs to a non-target genus, indicating that the primers have cross-genus cross-reactivity, and the primer pair is eliminated;

[0093] PASS_GENUS: The target species belongs to the target genus but not the target species, indicating that the primers are conserved at the genus level but not specific at the species level; primer pairs at this level are retained in the qualified pool, but information on closely related species is recorded for subsequent risk assessment, which is suitable for genus-level detection or scenarios where closely related species are difficult to distinguish;

[0094] PASS_SPECIES: The target species is a specific genus and species, indicating that the primers are species-specific; primer pairs of this level are given priority in the qualified pool.

[0095] NO_TARGET: No valid hit or only hits unclassified species within the target genus. Based on the genus-specific markers used in primer design, it is determined as PASS_GENUS or NO_RESULT.

[0096] Select primer pairs that hit the PASS_GENUS level for both the target genus and non-target species, and those that hit the PASS_SPECIES level for the target genus and target species, and enter them into the qualified pool.

[0097] Based on the set of risky primer pairs, a greedy algorithm is used to iteratively select primer pairs from the qualified pool of each target. Priority is given to combinations with zero heterodimer conflict. If this is not feasible, the combination with the fewest conflicts is selected until one pair is selected for each target and the total number of conflicts is less than or equal to a preset threshold.

[0098] After constructing a panel by selecting primer pairs from the qualified pool using a greedy algorithm, if the number of panels is insufficient to meet a preset value, adaptive supplementation sampling is initiated for the target with the smallest number of primer pairs in the qualified pool: the target with the fewest primer pairs in the qualified pool is identified, the evaluated primer pairs are excluded from the original candidate set, and batch BLAST sampling is performed again to expand the qualified pool for that target. The panel construction step is then repeated until the preset number of panels is met. Based on the constructed panels, detection kits containing those primer panels can be designed.

[0099] This embodiment employs a "known target gene anchoring + MEGABLAST homology expansion" strategy, using literature-verified specific target genes as seeds to expand strain-level diversity templates from a high-quality genome set. This avoids the missed detection risk of single reference sequence design and the cross-reaction of closely related species caused by whole-genome scanning of conserved regions. A two-layer clustering strategy of CD-HIT-EST and merging is used, providing high-quality alignment templates for primer design through a dynamic balance between intra-cluster redundancy removal and inter-cluster merging. A four-level specificity judgment algorithm, through two-level mismatch filtering, three-dimensional product prediction of direction-coordinate-length, and taxonomic grading, achieves a shift from a binary logic of "elimination upon hit" to a fine-grained management of "gradual assessment and controllable risk." A greedy algorithm is used to construct panels, improving panel construction efficiency and quality stability.

[0100] Based on the above embodiments, the preprocessing in this embodiment includes hierarchical sorting, deduplication, filtering, and adaptive equalization sampling;

[0101] The hierarchical sorting includes: arranging the NCBI reference genome first; sorting the genomes in descending order of assembly level; sorting the genomes in descending order of N50 value; sorting the genomes in descending order of CheckM integrity score; and sorting the genomes in ascending order of contamination rate.

[0102] The deduplication includes: if multiple genomes have the same assembly name, the genome with the GCF number recorded in the RefSeq database among the multiple genomes shall be retained first.

[0103] The filtering includes: removing genomes from the genome whose ANI verification status has failed;

[0104] The adaptive balanced sampling includes: determining the number of TaxIDs for each species; if the number of TaxIDs for each species is greater than a first preset number (e.g., 100), selecting the first preset number of genomes corresponding to the TaxIDs of each species from the globally hierarchically sorted genomes; otherwise, selecting the second preset number of genomes corresponding to the TaxIDs of each species from the genomes; calculating the difference between the first preset number and the second preset number (e.g., 1); and using the genomes with the first difference number in the genomes to complete the sampling, wherein the first preset number is greater than the second preset number.

[0105] It can perform multi-dimensional hierarchical ranking based on whether the genome is an NCBI reference genome, assembly level, N50 value, CheckM integrity score, and contamination rate. It can perform adaptive balanced sampling based on the number of TaxIDs within a species.

[0106] The hierarchical sorting rules are as follows: First, the NCBI reference genome is retained; then, it is sorted in descending order by Complete Genome, Chromosome, Scaffold, and Contig level (assigned scores of 4, 3, 2, and 1 respectively); finally, it is sorted in descending order by N50 value, descending order by CheckM integrity score, and ascending order by contamination rate. The final result is a high-quality genome set covering 100 different TaxIDs, with Complete Genomes accounting for more than 80%, average CheckM integrity greater than 98%, and average contamination rate less than 2%.

[0107] This embodiment constructs a high-quality genome database containing strain-level genetic diversity through hierarchical sorting, deduplication, filtering, and adaptive balanced sampling.

[0108] Based on the above embodiments, the cluster merging obtained from clustering in this embodiment includes:

[0109] An undirected graph is constructed using graph theory algorithms (such as NetworkX), with each cluster as a node and the ANI similarity between clusters (greater than or equal to a set value) as edges.

[0110] The average nucleic acid identity (ANI) between each cluster and other clusters is calculated using Mash. Small clusters are identified where the average nucleic acid identity is less than a first preset threshold (e.g., 0.95) and the number of sequences is less than a second preset threshold (e.g., 5). Small clusters refer to clusters where the number of sequences is less than the second preset threshold.

[0111] Traverse the neighbor nodes of the small cluster in the undirected graph, select the neighbor node with the largest sequence number as the merging target, and merge the small cluster into the merging target.

[0112] This strategy ensures both the purity of sequences within clusters and prevents rare variants from being discarded due to being clustered alone, thus ensuring that the templates for subsequent primer design are both conserved and diverse.

[0113] After processing with this two-layer clustering strategy, the original sequence was optimized from a chaotic mixed state into three high-quality alignment clusters. The sequence consistency within each cluster was greater than 95%, and key variant sites were preserved between clusters, providing a high-quality template for DegePrime scanning.

[0114] Based on the above embodiments, this embodiment includes the following primer pairs for predicting hit results through two-level mismatch filtering:

[0115] Determine the orientation, coordinates, and length of each primer pair;

[0116] If the coordinates of the 3' end of the forward primer are less than the coordinates of the 3' end of the reverse primer, the forward primer strand direction is positive and the reverse primer strand direction is negative, or the coordinates of the 3' end of the reverse primer are less than the coordinates of the 3' end of the forward primer, the reverse primer strand direction is positive and the forward primer strand direction is negative, then the direction of each primer pair is valid.

[0117] If the lengths of each primer pair with valid orientation are within a preset range, then each primer pair is determined as the final prediction result.

[0118] Three-dimensional PCR product prediction based on orientation, coordinates, and length was performed on the hit results obtained through two-level filtering.

[0119] First, the sstart and send of the BLAST output are standardized to start and end (the minimum value is start and the maximum value is end) to eliminate the influence of chain direction;

[0120] Next, calculate the 3' end coordinates F_3prime of the forward primer and R_3prime of the reverse primer: for the positive primer, F_3prime equals end; for the negative primer, R_3prime equals start.

[0121] Then verify the valid direction combinations: the F+R- combination requires F_3prime to be less than R_3prime and the forward primer direction to be positive and the reverse primer direction to be negative; the F-R+ combination requires R_3prime to be less than F_3prime and the reverse primer direction to be positive and the forward primer direction to be negative.

[0122] Finally, calculate the product length: in the F+R- direction, the product length is equal to end_r minus start_f plus 1; in the F-R+ direction, the product length is equal to end_f minus start_r plus 1.

[0123] Only results with product lengths in the range of 50-3000 bp are retained, excluding false positive products caused by large-scale genome rearrangement or alignment errors.

[0124] Based on the above embodiments, this embodiment, after constructing a panel by selecting primer pairs from the qualified pool using a greedy algorithm based on the risk primer pair set, further includes:

[0125] Calculate the sum of the degeneracy of all primers in each panel, and rank the panels in ascending order based on the sum of the degeneracy.

[0126] Calculate the number of risk pairs among all primer pairs within each panel, and rank the panels in ascending order based on the number of risk pairs.

[0127] Calculate the standard deviation of Tm values ​​for all primers within each panel, and rank the panels in ascending order based on the standard deviation of Tm values;

[0128] The rankings corresponding to each Panel are weighted and summed based on the weights of the sum of degeneracy, the weights of risk versus quantity, and the weights of the standard deviation of the Tm value.

[0129] Based on the weighted summation of all panels, the final panel is selected.

[0130] The obtained panel set is scored using a three-dimensional weighted ranking method, including the ranking of total degeneracy, the ranking of heterodimer conflict number, and the ranking of Tm standard deviation. The weighted ranking is calculated according to the preset weights, and the sorted primer design parameter set is output.

[0131] The resulting Panel set is scored using a three-dimensional weighted ranking method, specifically including:

[0132] The first dimension is the sum of degeneracy: calculate the sum of the degeneracy of all primers in the panel and rank them in ascending order (the smaller the value, the higher the ranking).

[0133] The second dimension is the heterodimer conflict number: count the number of risk pairs between all primer pairs in the panel and rank them in ascending order;

[0134] The third dimension is Tm balance: calculate the standard deviation of the Tm values ​​of all primers in the panel and rank them in ascending order.

[0135] The weighted ranking sum is calculated using preset weights: total degeneracy (0.4), heterodimer conflict number (0.4), and Tm standard deviation (0.2). A smaller weighted ranking sum indicates a better overall panel quality.

[0136] The output is a sorted set of primer design parameters, including the optimal primer pair ID, sequence, localization coordinates, annealing temperature, GC content, expected product length, degeneracy, and thermodynamic parameters for each target. Based on this set of primer design parameters, corresponding multiplex PCR primer panels can be prepared.

[0137] The following describes the multiplex pathogen identification device based on targeted nanopore sequencing provided by the present invention. The multiplex pathogen identification device based on targeted nanopore sequencing described below can be referred to in correspondence with the multiplex pathogen identification method based on targeted nanopore sequencing described above.

[0138] like Figure 3 As shown, the device includes:

[0139] The preprocessing module 301 is used to construct a panel of multiplex PCR primers. After multiplex PCR amplification of samples from multiple species based on the primer design parameter set of the panel, raw electrical signal data is obtained by nanopore sequencing. The raw electrical signal is then preprocessed to obtain the target sequence.

[0140] The calculation module 302 is used to align the target sequence to the target gene reference sequence set, remove the soft-sheared bases at both ends of the alignment result, calculate the sequencing depth and coverage of each target region, and calculate the number of reads based on the sequencing depth.

[0141] The identification module 303 is used to take a sample of a species as a preliminary screening sample if either of the two target regions of the same species meets a first condition; and to take the species as the identification result of the multi-species sample if both target regions of the same species in the preliminary screening sample meet a second condition.

[0142] The first condition is that the number of read segments is greater than or equal to a first preset threshold, and the coverage is greater than or equal to a second preset threshold; the second condition is that the number of read segments is greater than or equal to a third preset threshold, and the coverage is greater than or equal to a fourth preset threshold, wherein the third preset threshold is greater than the first preset threshold, and the fourth preset threshold is greater than or equal to the second preset threshold.

[0143] This embodiment abandons the metagenomic classifier and directly uses the Minimap2 target-specific alignment results as the sole criterion for judgment. It adopts a dual-target hierarchical strategy to improve the confidence of the judgment, thus fundamentally avoiding the problem of confusion between closely related species in the k-mer algorithm.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for multiplex pathogen identification based on targeted nanopore sequencing, characterized in that, include: A panel of multiplex PCR primers was constructed. Based on the primer design parameter set of the panel, multiplex PCR amplification was performed on samples from multiple species. Raw electrical signal data was obtained by nanopore sequencing. The raw electrical signals were preprocessed to obtain the target sequence. The target sequence is aligned to the target gene reference sequence set. After removing the soft-sheared bases at both ends of the alignment result, the sequencing depth and coverage of each target region are calculated, and the number of reads is calculated based on the sequencing depth. If either target region of the same species meets the first condition in two target regions, the sample of that species is used as the initial screening sample; if both target regions of the same species in the initial screening sample meet the second condition, the species is used as the identification result of the multi-species sample. The first condition is that the number of segments read is greater than or equal to a first preset threshold, and the coverage is greater than or equal to a second preset threshold; The second condition is that the number of segments read is greater than or equal to the third preset threshold, and the coverage is greater than or equal to the fourth preset threshold. The third preset threshold is greater than the first preset threshold, and the fourth preset threshold is greater than or equal to the second preset threshold.

2. The method for multiplex pathogen identification based on targeted nanopore sequencing according to claim 1, characterized in that, The preprocessing includes one or more of the following: base recognition, quality filtering, host sequence removal, and chimera removal; The quality filtering includes one or more of the following: removing reads containing adapter sequences and primer dimers, eliminating sequences with an average quality value less than a fifth preset threshold, and read length filtering.

3. The method for multiplex pathogen identification based on targeted nanopore sequencing according to claim 1, characterized in that, If either target region of the same species meets the first condition, after using the sample of that species as the initial screening sample, the process also includes: If at least one target region of the same species in the initial screening sample does not meet the second condition, the species is marked as a suspected signal and a retest is requested.

4. The method for multiplex pathogen identification based on targeted nanopore sequencing according to claim 1, characterized in that, Following the identification of this species as part of the multi-species sample, the following is also included: An identification report is generated based on the identification results, the target coverage matrix table, and the target coverage visualization heatmap.

5. The method for multiplex pathogen identification based on targeted nanopore sequencing according to claim 1, characterized in that, Panels for constructing multiplex PCR primers include: Genomes of multiple species were obtained from public genome databases, and a genome database was constructed after preprocessing the genomes. Reference sequences of validated species-specific target genes for each species are retrieved from public nucleic acid databases. Homologous sequence searches are performed on the genome database using the reference sequences to obtain a set of homologous sequences. Multiple sequence alignment is then performed on the set of homologous sequences to obtain a target gene sequence database. The target gene sequence database is subjected to CD-HIT-EST clustering. After merging the clusters, multiple sequence alignment is performed. The two ends of the multiple sequence alignment results are pruned to obtain the pruned alignment file. The DegePrime software was used to perform degenerate primer scanning on the trimmed alignment file to generate a candidate degenerate primer set. The MFEprimer was used to perform a three-level thermodynamic evaluation on the candidate degenerate primer set to generate a risk primer pair set. The candidate degenerate primer set was then filtered using a three-rule combination. The filtered candidate degenerate primer set was batch BLAST aligned to the NCBI RefSeq representative prokaryotic genome database. A two-level mismatch filtering was applied. Primer pairs that passed the two-level mismatch filtering were predicted. A four-level judgment was performed based on the taxonomic information of the hit species. Primer pairs were selected into the qualified pool according to the judgment results. Based on the set of risky primer pairs, a greedy algorithm is used to select primer pairs from the qualified pool to construct a panel.

6. The method for multiplex pathogen identification based on targeted nanopore sequencing according to claim 5, characterized in that, The preprocessing includes hierarchical sorting, deduplication, filtering, and adaptive equalization sampling; The hierarchical sorting includes: arranging the NCBI reference genome first; sorting the genomes in descending order of assembly level; sorting the genomes in descending order of N50 value; sorting the genomes in descending order of CheckM integrity score; and sorting the genomes in ascending order of contamination rate. The deduplication includes: if multiple genomes have the same assembly name, the genome with the GCF number recorded in the RefSeq database among the multiple genomes shall be retained first. The filtering includes: removing genomes from the genome whose ANI verification status has failed; The adaptive balanced sampling includes: determining the number of TaxIDs for each species; if the number of TaxIDs for each species is greater than a first preset number, selecting the first preset number of genomes corresponding to the TaxIDs of each species from the globally hierarchically sorted genomes; otherwise, selecting the second preset number of genomes corresponding to the TaxIDs of each species from the genomes; calculating the difference between the first preset number and the second preset number; and using the genomes with the first difference number in the genomes to complete the sampling, wherein the first preset number is greater than the second preset number.

7. The method for multiplex pathogen identification based on targeted nanopore sequencing according to claim 5, characterized in that, Cluster merging from clustering results includes: An undirected graph is constructed using a graph theory algorithm, with each cluster as a node and the ANI similarity between clusters as edges. Calculate the average nucleic acid consistency between each cluster and other clusters, and determine the small clusters whose average nucleic acid consistency is less than a first preset threshold and whose sequence number is less than a second preset threshold; The small cluster is selected from the neighbor nodes with the largest sequence number in the undirected graph as the merging target, and the small cluster is merged into the merging target.

8. The method for multiplex pathogen identification based on targeted nanopore sequencing according to claim 5, characterized in that, Predicted primer pairs for hit results obtained through two-level mismatch filtering include: Determine the orientation, coordinates, and length of each primer pair; If the coordinates of the 3' end of the forward primer are less than the coordinates of the 3' end of the reverse primer, the forward primer strand direction is positive and the reverse primer strand direction is negative, or the coordinates of the 3' end of the reverse primer are less than the coordinates of the 3' end of the forward primer, the reverse primer strand direction is positive and the forward primer strand direction is negative, then the direction of each primer pair is valid. If the lengths of each primer pair with valid orientation are within a preset range, then each primer pair is determined as the final prediction result.

9. The method for multiplex pathogen identification based on targeted nanopore sequencing according to claim 5, characterized in that, After constructing the panel by selecting primer pairs from the qualified pool using a greedy algorithm based on the set of risky primer pairs, the method further includes: Calculate the sum of the degeneracy of all primers in each panel, and rank the panels in ascending order based on the sum of the degeneracy. Calculate the number of risk pairs among all primer pairs within each panel, and rank the panels in ascending order based on the number of risk pairs. Calculate the standard deviation of Tm values ​​for all primers within each panel, and rank the panels in ascending order based on the standard deviation of Tm values; The rankings corresponding to each Panel are weighted and summed based on the weights of the sum of degeneracy, the weights of risk versus quantity, and the weights of the standard deviation of the Tm value. Based on the weighted summation of all panels, the final panel is selected.

10. A multiplex pathogen identification device based on targeted nanopore sequencing, characterized in that, include: The preprocessing module is used to construct a panel for multiplex PCR primers. Based on the primer design parameter set of the panel, multiplex PCR amplification is performed on samples of multiple species. Raw electrical signal data is obtained by nanopore sequencing. The raw electrical signal is then preprocessed to obtain the target sequence. The calculation module is used to align the target sequence to the target gene reference sequence set, remove the soft-sheared bases at both ends of the alignment result, calculate the sequencing depth and coverage of each target region, and calculate the number of reads based on the sequencing depth. The identification module is used to identify a sample of a species as a preliminary screening sample if either of the two target regions of the same species meets a first condition; and to identify the species as the identification result of the multi-species sample if both target regions of the same species in the preliminary screening sample meet a second condition. The first condition is that the number of segments read is greater than or equal to a first preset threshold, and the coverage is greater than or equal to a second preset threshold; The second condition is that the number of segments read is greater than or equal to the third preset threshold, and the coverage is greater than or equal to the fourth preset threshold. The third preset threshold is greater than the first preset threshold, and the fourth preset threshold is greater than or equal to the second preset threshold.