A method of assessing universal primer species / typing level identification capability

By constructing a microbial database and processing sequencing data, the species/typing level identification capabilities of universal primers were evaluated, which solved the problem of lack of evaluation methods in the existing technology, achieved accurate identification and screening of universal primers, and improved the reliability of pathogenic microorganism detection.

CN118748039BActive Publication Date: 2025-10-17HUGOBIOTECH BEIJING CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410918254.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2025-10-17
Estimated Expiration
2044-07-10

AI Technical Summary

Technical Problem

Existing technologies lack effective evaluation methods to assess the typing-level effectiveness of universal primers in pathogen identification, resulting in reliance on experimental testing in clinical applications, which lacks accuracy and reliability.

Method used

A microbial database was constructed, and universal primers were aligned to the microbial database to obtain targeted fragments and perform sequencing data processing. Through alignment and deduplication statistics, the species/typing level identification ability of universal primers was calculated, providing a method for evaluating the species/typing level identification of universal primers.

Benefits of technology

It achieves accurate evaluation of the species/typing level identification capability of universal primers, improves the basis for primer screening and the ability to predict results, and ensures the effectiveness of universal primers in the detection of pathogenic microorganisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118748039B_ABST
    Figure CN118748039B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of molecular biology, and discloses a method for evaluating the species / typing level identification capability of universal primers. The method comprises the following steps: constructing a microorganism database, aligning the universal primers to the microorganism database to obtain target fragments and species / typing information corresponding to the universal primers, obtaining sequencing data by cutting the target fragments from the F end to a specific length, aligning the sequencing data to the microorganism database for species classification and identification after deduplication and statistics, calculating the sequence return rate of the universal primers for a specific species / typing, and evaluating the identification capability of the universal primers for the specific species / typing level. The application provides a method for evaluating the species level identification capability of universal primers for pathogenic microorganism identification, and evaluates the performance of the universal primers for identifying species / typing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of molecular biology, and relates to a method for evaluating the species / typing level identification capability of universal primers. BACKGROUND

[0002] Viral or bacterial infections pose a threat to human health worldwide, and timely and accurate etiological diagnosis is crucial for clinical treatment and reduction of antibiotic misuse. Traditional detection methods such as microscopic observation, isolation culture and PCR have the disadvantages of long detection time and small detection target. Multiplex PCR targeted sequencing technology, also known as amplicon targeted sequencing technology, is a targeted sequencing technology that combines multiplex PCR technology with second-generation sequencing technology, which can rapidly and efficiently detect a variety of pathogens, including viruses, bacteria, fungi and parasites. Multiplex PCR targeted sequencing technology has become one of the important tools for clinical pathogen detection, which can accurately and rapidly diagnose and treat diseases. Second-generation sequencing technology (NGS) enriches specific genomic regions by specifically amplifying specific fragments of target species and sequencing verification.

[0003] Targeted next-generation sequencing technology (tNGS) combines multiplex PCR amplification with sequencing technology, and the core technology is the design of target species primers. For some complex typing microorganisms, genus universal primers can be designed, which can reduce the number of primers in the primer pool, reduce the interaction between primers, and also reduce production costs and operational complexity. The core of tNGS pathogen detection in clinical application is to design primer pools for multiple pathogenic microorganisms in clinical focus. The universal primers in the primer pool can usually identify species level or typing level, which is meaningful. Whether the universal primers can achieve ideal typing efficiency needs a specific evaluation method. In clinical pathogen detection, universal primers can usually capture multiple target sequences within the range, and the existing technology does not provide a specific evaluation method for distinguishing, qualitatively analyzing and reverse deducing the typing efficiency of universal primers. In actual application, the typing efficiency of universal primers is more dependent on experimental detection for judgment. Therefore, it is of great significance to provide an analysis method for evaluating the identification capability of universal primers for specific species. SUMMARY

[0004] To develop a method capable of evaluating the performance of universal primer typing, the present application provides a method for evaluating the species / typing level identification ability of universal primer. It includes: constructing a microbial database, aligning the universal primer to the microbial database to obtain the targeted fragments and the species / typing information corresponding to the universal primer, obtaining sequencing data by cutting the targeted fragments from the F end to a specific length, after deduplication and statistics, aligning the sequencing data to the microbial database for species classification identification, calculating the sequence return ratio of the universal primer for a specific species / typing, and evaluating the identification ability of the universal primer for a specific species / typing level. The present application provides a method for evaluating the species level identification ability of universal primer for pathogenic microorganism identification, and evaluates the performance of universal primer for species / typing identification.

[0005] A species is also called a species, which is the basic unit of biological classification. At the species level, it can be further divided into subtypes and serotypes according to antigen or genomic characteristics, i.e. the typing referred to in the present application. Not all species have typing.

[0006] To achieve the technical purpose of the present application, in one aspect, the present application provides a method for evaluating the species / typing level identification ability of universal primer. It includes: constructing a microbial database based on a public database, aligning the universal primer sequence to the microbial database to obtain the targeted fragments and the species / typing information corresponding to the universal primer sequence; obtaining sequencing data by cutting the targeted fragments from the F end to a length consistent with the sequencing read length, performing deduplication processing on the sequencing data, and simultaneously recording the total number of sequence data corresponding to the species / typing; aligning the sequencing data to the microbial database for species classification identification, and counting the number of sequence data returned to the target species / typing; and obtaining the identification ability of the universal primer for the species / typing level according to the total number of sequence data corresponding to the species / typing and the number of sequence data returned to the target species / typing.

[0007] Further, in the method for evaluating the species level identification ability of universal primer provided by the present application, the information in the microbial database includes the species / typing name, the genomic sequence of the species / typing, the sequence number and the biological taxonomy ID; the deduplication processing includes the deduplication processing of the species information and the deduplication processing of the gene sequence. The species classification identification selects the sequencing data with a coverage of more than 95%, a specificity of more than 95%, and an alignment result that is unique to the species / typing and is the target species / typing as the sequencing data returned to the target species / typing.

[0008] Exemplarily, the method for constructing the microbial database provided by the present application includes extracting the microbial database including archaea, bacteria, viruses, fungi and protists from the public database, and constructing the microbial database including the genomic sequence and the species / typing name corresponding to the sequence, the sequence number and the biological taxonomy ID.

[0009] The present application does not make specific limitations on the length of the sequencing data, and those skilled in the art can adjust the sequencing read length.

[0010] The present application does not make specific limitations on the type of public database, and those skilled in the art can use NT database or other microbial professional database such as GISAID as a public database.

[0011] Further, in the method for evaluating the species / typing level identification ability of the universal primer provided by the present application, the calculation formula of the return ratio is R=u / U; R is the return ratio; u is the number of sequencing data returned to the target species / typing; and U is the total number of sequencing data corresponding to the species / typing.

[0012] In another aspect, the present application claims a computer device / system / equipment, comprising a memory, a processor, and a computer program stored in the memory, wherein the memory executes the computer program to realize the above-mentioned method for evaluating the species / typing level identification ability of the universal primer.

[0013] In another aspect, the present application claims a computer readable storage medium, which stores a computer program / instruction, wherein the computer program / instruction is executed by a computer to realize the above-mentioned method for evaluating the species / typing level identification ability of the universal primer.

[0014] In addition, the present application claims the application of the above-mentioned method for evaluating the species / typing level identification ability of the universal primer in evaluating the species / typing level identification ability of the universal primer.

[0015] Exemplarily, the universal primer comprises an enterovirus universal primer, the F-terminal nucleotide sequence of the enterovirus universal primer is TGAGTCCTCCGGCCCCTGAATG, and the R-terminal nucleotide sequence of the enterovirus universal primer is ATATATTGTCACCATAAGCAGAT; the typing comprises clinical key enterovirus typing, and the clinical key enterovirus typing comprises enterovirus 71 type, enterovirus D68 type, echovirus 30 type, coxsackievirus 3 type, and coxsackievirus A16 type.

[0016] The application evaluates the identification ability of the enterovirus universal primer on the clinical key typing enterovirus 71 (Enterovirus A71, EV-A71), enterovirus D68 (Enterovirus D68, EV-D68), echovirus 30 (Echovirus 30, E30), coxsackievirus B3 (Coxsackievirus B3, CVB3) and coxsackievirus A16 (Coxsackievirus A16, CoxA16) by the above-mentioned method, and it is found that the return ratio of the enterovirus universal primer on the typing Enterovirus A71 is 91.453%, and the return ratio of the primer on the typing Echovirus type 30 is 35.714%.

[0017] Further, the application verifies the identification ability of the enterovirus universal primer on the above-mentioned clinical key typing in different test samples, and it is found that the corresponding enterovirus can be detected in different test samples, which indicates that the method for evaluating the species / type level identification ability of the universal primer provided by the application can better identify the typing efficiency of the universal primer, and the return ratio of the universal primer on the specific species / type is >0%, which indicates that the primer can be used for identifying the corresponding species / type.

[0018] Compared with the prior art, the technical scheme provided by the application at least has the following beneficial effects or advantages:

[0019] The evaluation method provided by the application can evaluate the universal primer species / type identification capability, can increase the basis for primer screening and the result prediction capability. The method provided by the application evaluates the identification capability of the enterovirus universal primer on the clinical key types of enterovirus 71 (Enterovirus A71, EV-A71), enterovirus D68 (Enterovirus D68, EV-D68), echovirus 30 (Echovirus 30, E30), coxsackievirus B3 (Coxsackievirus B3, CVB3) and coxsackievirus A16 (Coxsackievirus A16, CoxA16), and it is known that the return ratio of the enterovirus universal primer to the type Enterovirus A71 is 91.453%, and the return ratio of the primer to the type Echovirus type 30 is 35.714%. Further, the application verifies the identification capability of the enterovirus universal primer on the above-mentioned clinical key types in different test samples, and it is found that the corresponding enterovirus can be detected in different test samples, which shows that the method for evaluating the universal primer species / type identification capability provided by the application can better identify the type efficiency of the universal primer, and the return ratio of the universal primer to the specific species / type is >0%, which shows that the primer can be used to identify the corresponding species / type. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only some embodiments of the application.

[0021] Figure 1 The flow chart for evaluating the universal primer species / type identification capability. DETAILED DESCRIPTION

[0022] The technical solutions of the application will be described below in combination with the embodiments, but the application is not limited to the following embodiments. The experimental methods and detection methods described in each embodiment are all conventional methods unless otherwise specified; and the reagents and materials described are all commercially available unless otherwise specified.

[0023] Embodiment 1

[0024] The embodiment provides a method for evaluating the universal primer species / type identification capability, and the evaluation process is as shown in Figure 1 The method comprises the following steps.

[0025] A microbial database is constructed based on a public database, and the microbial database includes archaea, bacteria, virus, fungi, and protists. The constructed microbial database includes genome sequences, species / subtype names corresponding to the sequences, sequence numbers (Accession ID), and biological taxonomy ID (Taxonomy ID).

[0026] The universal primer sequence is aligned to the microbial database to obtain a target fragment and species / subtype information corresponding to the universal primer sequence.

[0027] The obtained target fragment is cut from the F end to a length consistent with the sequencing read length (75 bp), and is used as sequencing data (reads). The reads are subjected to species information and fragment sequence deduplication processing (second-generation sequencing platform SE75), and the total number of species corresponding information is recorded.

[0028] The deduplicated reads are aligned to the microbial database, and the reads are subjected to species classification identification through an internal identification process. The internal identification process includes: the reads are aligned to the microbial database using software blast, and the alignment results of the reads with a coverage of more than 95%, a specificity of more than 95%, a unique alignment species, and a target species are selected.

[0029] The identification ability of the universal primer at the species level is calculated according to the total number of species corresponding information and the number of reads aligned to the target species. The calculation formula is: R = u / U.

[0030] In the formula, R is the recall rate, u is the number of reads aligned to the target species, and U is the total number of sequencing data corresponding to the species.

[0031] Exemplarily, the microbial database is constructed through the NT database.

[0032] Example 2

[0033] The present embodiment provides a method for evaluating the species / typing level identification ability of universal primers, and the application and verification of the method in evaluating enterovirus universal primers. The enterovirus universal primer is specifically numbered EV529, and the primer sequence is F end TGAGTCCTCCGGCCCCTGAATG and R end ATATATTGTCACCATAAGCAGAT. The enterovirus universal primer can cover more than 95% of 5319 enterovirus complete genomes from GenBank, including 249 sub-classifications. The present embodiment takes the clinical key typing enterovirus 71 (EV-A71), enterovirus D68 (EV-D68), echovirus 30 (E30), coxsackievirus B3 (CVB3) and coxsackievirus A16 (CoxA16) as target classifications, and evaluates the return ratio of enterovirus universal primers.

[0034] 1. Evaluation of enterovirus universal primer identification ability

[0035] Step 1: Download the NT database and extract the microbial database including archaea, bacteria, virus, fungi, and protists, and the extracted information includes genome sequence, species / typing name corresponding to the sequence, sequence number and biological classification ID;

[0036] Step 2: Align the primer sequence to the microbial database to obtain the targeted fragment and the corresponding species / typing information;

[0037] Step 3: Take 75bp from the F end of the obtained targeted fragment as the second-generation sequencing platform SE75 sequencing data (reads), and perform species information and fragment sequence deduplication processing on the reads, while recording the total number of species corresponding information;

[0038] Step 4: Align the deduplicated reads to the microbial database, and identify the species classification of the reads through the internal identification process, while counting the number of reads returned to the target species / typing. The internal identification process includes: using software blast to align the reads to the microbial database, selecting reads with coverage of more than 95%, specificity of more than 95%, unique alignment species and target species alignment results;

[0039] Step 5: Calculate the species / typing unique sequence return ratio, the calculation formula is: R=u / U.

[0040] The return ratio of the enterovirus universal primer in the target classification calculated by the above method is shown in Table 1.

[0041] Table 1: Return ratio of enterovirus universal primer in target classification

[0042] Phenotyping Total number of reads / strip Number of reads / strip Ratio of reads / % Enterovirus A71 117 107 91.453 Coxsackievirus B3 25 21 84.000 Coxsackievirus A16 60 50 83.333 Enterovirus D68 38 31 81.579 Echovirus 30 14 5 35.714

[0043] As shown in Table 1, the return ratio of the enterovirus universal primer for different enterovirus types is different, and the return ratio of the primer for Enterovirus A71 is 91.453%, and the return ratio of the primer for Echovirus type 30 is 35.714%.

[0044] 2. Experimental verification of the identification ability of the enterovirus universal primer for Enterovirus A71, Enterovirus D68, Echovirus 30, Coxsackievirus B3, and Coxsackievirus A16

[0045] The test sample types include MOCK (virus / bacterial standard strain) and clinical sample alveolar lavage fluid, cerebrospinal fluid, blood and sputum, and the identification ability of the enterovirus universal primer is verified based on the target next-generation sequencing method (tNGS), and the verification results are shown in Table 2.

[0046] Table 2: Detection results of enterovirus universal primer for different types of samples

[0047] Mock Bronchoalveolar lavage fluid Blood Sputum Cerebrospinal fluid Coxsackievirus A16 1 / / / / Coxsackievirus B3 1 / / / / Echovirus 30 / / / / 1 Enterovirus A71 1 / / / / Enterovirus D68 / 2 1 1 /

[0048] The numbers in Table 2 represent the number of samples tested, and the test samples in Table 2 can all detect the corresponding enterovirus. It is shown that the method for evaluating the species level identification ability of the universal primer provided by the present application can better identify the typing efficiency of the primer, and the return ratio of the universal primer for a specific species / type is >0%, indicating that the primer can be used to identify the corresponding species / type.

[0049] The above-described embodiments are part of the embodiments of the present application, rather than all the embodiments. The detailed description of the embodiments of the present application is not intended to limit the scope of the claimed application, but only represents selected embodiments of the present application. All other embodiments obtained by related deduction and replacement made by those skilled in the art under the condition of the concept of the present application, without making creative efforts, fall within the scope of protection of the present application.

Claims

1. A method for evaluating the species / typing level identification capability of universal primers, characterized in that: include: Build a microbial database including archaea, bacteria, viruses, fungi, and protists based on public databases; Comparing the universal primer sequence to the microbial database to obtain the targeted fragment and species / typing information corresponding to the universal primer sequence; The targeted fragment is cut from the F end to a length consistent with the sequencing read length to obtain sequencing data, the sequencing data is deduplicated, and the total number of sequencing data corresponding to the species / typing is statistically recorded; Aligning the targeted fragments to the microbial database for species classification and identification, and counting the number of sequencing data matched to the target species / type; The identification capability of the universal primer at the species / genotype level is obtained based on the total number of sequencing data corresponding to the species / genotype and the number of sequencing data matched back to the target species / genotype; The calculation formula of the return ratio is R=u / U; R is the ratio of the back-matching ratio; u is the number of sequencing data corresponding to the target species / genotype; and U is the total number of sequencing data corresponding to the species / genotype.

2. The method for evaluating species / typing level identification capability of universal primers according to claim 1, wherein: The public database includes the NT database and the GISAID database; The information in the microbial database includes the name of the species / type, the genome sequence of the species / type, the sequence number and the biological classification ID.

3. The method for evaluating species / typing level identification capability of universal primers according to claim 1, wherein: The deduplication processing includes deduplication processing of species / typing information and deduplication processing of gene sequences.

4. The method for evaluating species / typing level identification capability of universal primers according to claim 1, wherein: The species classification identification selects sequencing data with a coverage of more than 95%, a specificity of more than 95%, and a matching result that is unique to the target species / type as the sequencing data matched back to the target species / type.

5. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the method according to any one of claims 1 to 4.

6. A computer-readable storage medium storing a computer program / instruction, characterized in that: When the computer program or instruction is executed by a computer, the method according to any one of claims 1 to 4 is implemented.

7. Use of the method according to any one of claims 1 to 4 in evaluating the species / typing level identification capability of universal primers.

8. The use according to claim 7, characterized in that The universal primer includes an enterovirus universal primer, the F-terminal nucleotide sequence of the enterovirus universal primer is TGAGTCCTCGGCCCCTGAATG, and the R-terminal nucleotide sequence of the enterovirus universal primer is ATATATTGTCACCATAAGCAGAT.

9. The use according to claim 7, characterized in that The typing includes clinically important enterovirus typing, and the clinically important enterovirus typing includes: enterovirus 71, enterovirus D68, echovirus 30, coxsackievirus 3, and coxsackievirus A16.

Citation Information

Patent Citations

  • Multi-target detection primer group and design method thereof

    CN114974427A

  • Alignment method based on targeted high-throughput sequencing sequence and application thereof

    CN116959580A