A method and device for evaluating the effectiveness of a reagent or reagent combination for detecting pathogens

By comparing and analyzing pathogen genome sequences and evaluating the effectiveness of pathogen detection reagents, the problems of false negatives caused by insufficient sensitivity and mutations in new coronavirus detection were solved, and rapid adaptation and accuracy improvement of virus detection methods were achieved.

CN114627966BActive Publication Date: 2025-09-12XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011450416.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-11
Publication Date
2025-09-12
Estimated Expiration
2040-12-11

AI Technical Summary

Technical Problem

In existing technologies, the new coronavirus detection methods have problems such as insufficient sensitivity and false negative results caused by viral mutations. There is a lack of effective computational analysis tools to evaluate and verify molecular diagnostic methods such as RT-PCR, making it difficult to deal with the risk of transmission of viral mutations.

Method used

A method and analysis software have been developed to evaluate the effectiveness of pathogen detection reagents by obtaining the genome sequence information of pathogens and using database comparison and sequence analysis, including the design of primers and probes, combined with sequence comparison and mutation analysis to evaluate their targeting and coverage of viruses.

Benefits of technology

It can quickly and accurately evaluate the effectiveness of pathogen detection reagents, adapt to virus mutations, improve the sensitivity and accuracy of detection, and assist in the development and updating of diagnostic reagents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627966B_ABST
    Figure CN114627966B_ABST
Patent Text Reader

Abstract

The present application relates to a method for evaluating the effectiveness of a reagent for detecting pathogens, and an apparatus for evaluating the effectiveness of a reagent or combination of reagents for detecting pathogens. The method and apparatus of the present application can evaluate the effectiveness, coverage, etc. of newly designed nucleic acid sequence-based pathogen (particularly virus) detection reagents (e.g., primers and probes), thereby assisting in the development of diagnostic reagents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of molecular biology, and in particular to a method for evaluating the effectiveness of a reagent for detecting pathogens, and a device for evaluating the effectiveness of a reagent or a combination of reagents for detecting pathogens. Background Art

[0002] Coronavirus disease 2019 (COVID-19) is a newly emerging acute respiratory infectious disease. Coronaviruses (CoVs) are a class of enveloped, linear, single-stranded, positive-sense RNA viruses that are widespread in nature. Coronaviruses are divided into four genera: alpha, beta, gamma, and delta. Alpha and beta coronaviruses can infect mammals, while gamma and delta coronaviruses primarily infect birds. Six coronaviruses have been previously identified as infecting humans: alpha coronaviruses (HCoV-229E, HCoV-NL63) and beta coronaviruses (HCoV-HKU1, HCoV-OC43, MERS-CoV, and SARS-CoV). Patients present with symptoms ranging from the common cold to severe lung infections. The novel coronavirus (severe acute respiratory syndrome coronavirus 2, SARS-CoV-2) is a beta coronavirus not previously identified in humans. It is highly contagious and spreads primarily through respiratory droplets, respiratory secretions, and direct contact. COVID-19 patients often present with specific, similar symptoms, such as fever, malaise, and cough. Most patients experience mild flu-like symptoms and have a good prognosis, but a small number of patients are critically ill and rapidly develop acute respiratory distress syndrome, respiratory failure, multiple organ failure, and even death.

[0003] The novel coronavirus is highly contagious, the general population is susceptible, and it has an incubation period, during which it is contagious. Therefore, effective control of the spread of the epidemic requires prompt identification of infected individuals, control of the source of infection, and severing of transmission routes. The prerequisite for all containment measures is rapid, accurate, and sensitive diagnosis.

[0004] Currently, many diagnostic methods for the novel coronavirus (COVID-19) available domestically and internationally primarily utilize real-time fluorescence PCR technology to specifically detect novel coronavirus nucleic acid in specimens such as throat swabs, sputum, other lower respiratory secretions, blood, feces, and urine. Clinically, nucleic acid testing is used as one of the diagnostic criteria. However, due to inherent sensitivity issues with the reagents and variations in the target viral sequence, false negative results can occur. This problem has become particularly prominent as the virus continues to spread and mutate.

[0005] Many existing web tools and software only provide methods for case statistics and mutation analysis, without addressing how to use this information to guide clinical practice, thus lacking practical application. Simple modeling and predictive analysis often differ significantly from actual results, making them of limited value. The "gold standard" for detecting the novel coronavirus remains RT-PCR, and whether RT-PCR can produce highly sensitive, reliable results is often the determining factor in the risk of epidemic spread. Computational analysis tools are currently unavailable to evaluate and validate molecular diagnostic methods like RT-PCR. The effectiveness of RT-PCR and its continued application in the future and to a wider range of regions will need to be assessed based on the evolution of the virus.

[0006] Therefore, an analytical method and analytical tool that can integrate SARS-CoV-2 mutation analysis and RT-PCR assay validation analysis is particularly important. Summary of the Invention

[0007] In the present invention, unless otherwise specified, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. At the same time, in order to better understand the present invention, the definitions and explanations of relevant terms are provided below.

[0008] As used herein, the term "database" refers to a database containing one or more (e.g., at least 10, at least 50, at least 10 2 At least 5*10 2 At least 10 3 At least 5*10 3 At least 10 4 At least 5*10 4 In the present application, the database can be manually established. For example, one or more (e.g., at least 10, at least 50, at least 10) pathogens can be obtained and entered. 2 Types, at least 5*10 2 Types, at least 10 3 Types, at least 5*10 3 Types, at least 10 4 Types, at least 5*10 4The database can be established by using the genome sequences of multiple pathogens (e.g., one or more strains). In addition, the database can also be a public database containing multiple pathogen genome sequences. Such public databases include, but are not limited to, databases of the National Center for Bioinformation (CNCB) / National Genome Science Data Center (NGDC), the EMBL nucleotide sequence database, DDBJ (DNA Data Bank of Japan), GenBank (GenBank), MIM (Online Mendelian Inheritance in Man (OMIM), MGD (Mouse genome database), ECO2DBASE (Escherichia coli gene-protein database (2D gelspots), TIGR (The bacterial database(s) of 'The Institute of Genome Research'), TubercuList (Mycobacterium tuberculosis H37Rv genome database), and HIV (HIV sequence database), etc. In certain preferred embodiments of the present application, the genomic sequence information in the database has a unified data format. In certain preferred embodiments, the genomic sequence information in the database is stored in the fasta format.

[0009] As used herein, “severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)”, formerly known as “new coronavirus” or “2019-nCov”, belongs to the genus Beta coronavirus and is an enveloped single-stranded positive-sense RNA virus. The genome sequence of SARS-CoV-2 is known to those skilled in the art, and can be found, for example, in GenBank: MN908947. SARS-CoV-2 contains at least three membrane proteins, including surface spike protein (S), integral membrane protein (M) and membrane protein (E). The current detection method for the new coronavirus mainly targets the open reading frame lab (ORFlab), nucleocapsid protein (N) and envelope protein (E) in its genome.

[0010] As used herein, the terms "new coronavirus pneumonia" and "COVID-19" refer to pneumonia caused by infection with SARS-CoV-2. The two have the same meaning and can be used interchangeably.

[0011] As used herein, the term "alignment" refers to comparing and matching the sequences of two or more polypeptides or two or more nucleic acids. When a position in the two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of the two DNA molecules is occupied by adenine, or a position in each of the two polypeptides is occupied by lysine), then the molecules are aligned identically at that position. Typically, comparison is made when the two sequences are aligned for maximum identity. Such an alignment can be achieved using, for example, the method of Needleman et al. (1970) J. Mol. Biol. 48: 443-453, which can be conveniently performed using a computer program such as the Align program (DNAstar, Inc.). The percent identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)), which has been incorporated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4.

[0012] As used herein, the term "nucleotide sequence targeting or specific for a pathogen" refers to a nucleotide sequence that is capable of specifically hybridizing or annealing to the genome of a pathogen or a fragment thereof under conditions that allow nucleic acid hybridization, annealing, or amplification. In certain embodiments, a "nucleotide sequence targeting or specific for a pathogen" may comprise a nucleotide sequence that is complementary (e.g., partially complementary or fully complementary) to the genome of a pathogen or a fragment thereof.

[0013] As used herein, the term "gene chip" refers to a substrate on the surface of which an array of target nucleotides of known sequence is immobilized.

[0014] As used herein, the term "SNP (single nucleotide polymorphism)" refers to a nucleic acid sequence polymorphism caused by a variation of a single nucleotide at the genomic level. A site in a genome that has a single nucleotide polymorphism is referred to as a "SNP site."

[0015] The inventors of the present application have developed a set of methods and analysis software (devices) that can evaluate the effectiveness of reagents (eg, primers or probes) for detecting pathogens (eg, SARS-CoV-2).

[0016] In one aspect, the present invention provides a method for evaluating the effectiveness of an agent or combination of agents for detecting a pathogen, wherein the agent comprises a nucleotide sequence that targets or is specific for the pathogen, the method comprising:

[0017] (1) obtaining the nucleotide sequence of the agent and a database containing one or more (e.g., at least 10, at least 50, at least 10 2 At least 5*10 2 At least 10 3 At least 5*10 3 At least 10 4 At least 5*10 4 or more) genome sequences of the pathogen;

[0018] (2) performing sequence alignment on the nucleotide sequence with all or part of the genome sequences of the pathogen in the database to obtain a sequence alignment result;

[0019] (3) Determine the effectiveness of the reagent in detecting the pathogen based on the sequence comparison results.

[0020] In certain embodiments, the pathogen is selected from bacteria (e.g., Mycobacterium tuberculosis, Staphylococcus, Streptococcus, Pneumococcus, Escherichia coli, Shigella, Salmonella), fungi (e.g., Trichophyton, Epidermophyton, Candida albicans), viruses (e.g., coronavirus, Ebola virus, hepatitis B virus, human immunodeficiency virus 1, rabies virus, human papillomavirus), and any combination thereof. In certain embodiments, the pathogen is a coronavirus, such as a beta coronavirus (HCoV-HKU1, HCoV-OC43, MERS-CoV, SARS-CoV, SARS-CoV-2). In certain embodiments, the pathogen is SARS-CoV-2.

[0021] In certain embodiments, the reagent detects the pathogen by sequencing, PCR (eg, RT-PCR, real-time fluorescence quantitative PCR), isothermal amplification, melting curve analysis, nucleic acid blot analysis, or any combination thereof.

[0022] In certain embodiments, the reagent is selected from a primer, a probe, a gene chip, or any combination thereof. In certain embodiments, the reagent comprises a forward primer, a reverse primer and a probe.

[0023] In certain embodiments, the agent is designed based on a designated pathogen genomic sequence; for example, the agent is designed to contain a nucleotide sequence that is partially complementary or fully complementary to the designated pathogen genomic sequence.

[0024] In certain embodiments, the database is a public database containing a plurality of pathogen genome sequences. Such public databases include, but are not limited to, databases of the National Center for Bioinformation (CNCB) / National Genome Science Data Center (NGDC), EMBL nucleotide sequence database, DDBJ (DNA Data Bank of Japan), GenBank (GenBank), Escherichia coli ECO2DBASE (Escherichia coli gene-protein database (2D gelspots), TIGR (The bacterial database(s) of 'The Institute of Genome Research'), Mycobacterium tuberculosis TubercuList (Mycobacterium tuberculosis H37Rv genome database), and HIV (HIV sequence database).

[0025] In some embodiments, the database is an artificially created database. In some embodiments, the database is obtained and entered into one or more (e.g., at least 10, at least 50, at least 10) of the pathogens. 2 Types, at least 5*10 2 Types, at least 10 3 Types, at least 5*10 3 Types, at least 10 4 Types, at least 5*10 4 In certain embodiments, the genome sequence is preprocessed before entering the genome sequence into the database. Such preprocessing includes, but is not limited to: converting the genome sequence into a specified data format (e.g., FASTA format); removing incomplete genome sequences; removing genome sequences that do not meet the requirements of the pathogen's target host; removing special characters that are not bases (e.g., A, T, C, G) in the genome sequence; or any combination thereof.

[0026] In certain embodiments, the genomic sequence information in the database has a unified data format. In certain preferred embodiments, the genomic sequence information in the database is stored in fasta format.

[0027] In certain embodiments, the one or more genomic sequences are derived from the same or different genus, species, subspecies, subtype, subgroup, subpopulation, or strain of the pathogen.

[0028] In certain embodiments, the database further contains annotation information for the genome sequences of one or more of the pathogens.

[0029] In certain embodiments, the annotation information is selected from the following: genus information, species information, subspecies information, subtype information, subgroup information, subpopulation information, strain information, regional information, and age information of the pathogen from which the genome sequence is derived; functional regions, mutation regions, mutation sites, mutation types, mutation frequencies, conserved regions, and conserved sites of the genome sequence; or any combination thereof.

[0030] In certain embodiments, the method further comprises, before performing step (2), annotating and / or analyzing the genome sequence of one or more of the pathogens.

[0031] In certain embodiments, for one or more genome sequences of the pathogens, the annotation is selected from the following information: genus information, species information, subspecies information, subtype information, subgroup information, subpopulation information, strain information, regional information, age information, or any combination thereof of the pathogen from which the genome sequence is derived.

[0032] In certain embodiments, sequence analysis is performed on one or more of the pathogen's genomic sequences; for example, a sequence alignment is performed on two or more of the pathogen's genomic sequences. Preferably, after sequence alignment, information selected from the following is determined and annotated for the pathogen's genomic sequence: functional regions, mutation regions, mutation sites, mutation types, mutation frequencies, conserved regions, conserved sites, or any combination thereof. In certain embodiments, one or more of the pathogen's genomic sequences are analyzed by sequence alignment with a reference genomic sequence.

[0033] In certain embodiments, in step (2), the nucleotide sequence is compared with the genome sequences of all the pathogens in the database.

[0034] In certain embodiments, in step (2), a portion of the genome sequence of the pathogen is selected based on the specified information, and then the nucleotide sequence is aligned with the selected genome sequence. Preferably, the specified information is selected from the annotation information described above or any combination thereof, such as genus information, species information, subspecies information, subtype information, subgroup information, subpopulation information, strain information, regional information, and age information of the pathogen; functional regions, mutation regions, mutation sites, mutation types, mutation frequencies, conserved regions, and conserved sites of the genome sequence; or any combination thereof.

[0035] In the present application, the nucleotide sequence can be aligned with the genome sequence of the pathogen using various alignment algorithms. Such alignment algorithms include, but are not limited to, the algorithm proposed by Needleman et al. (J. Mol. Biol. 48: 443-453 (1970)), the algorithm proposed by E. Meyers and W. Miller (Comput. Appl Biosci., 4: 11-17 (1988)), and can be conveniently performed using computer programs (e.g., Align program, VectorNTI program, NUCleotide MUMmer program, etc.).

[0036] In certain embodiments, wherein, in step (3), based on the sequence alignment results, the number or ratio of pathogen genomic sequences that the reagent can target, the region of the pathogen genomic sequence targeted by the reagent, the number of mutation sites contained in the region, the mutation type and / or the frequency of the mutation type of the mutation site, the mutation frequency of the mutation site, the position of the mutation site in the region, or any combination thereof are determined to evaluate the effectiveness of the reagent in detecting the pathogen.

[0037] Without being limited by theory, the inventors of the present application believe that the more pathogen genomic sequences that the agent can target, or the higher the ratio, or the fewer mutation sites contained in the targeted region, or the lower the mutation frequency of the mutation sites contained, or the farther the mutation sites contained are from the 3' end of the nucleotide sequence of the agent, the higher the effectiveness of the agent. Conversely, the fewer pathogen genomic sequences that the agent can target, or the lower the ratio, or the more mutation sites contained in the targeted region, or the higher the mutation frequency of the mutation sites contained, or the closer the mutation sites contained are to the 3' end of the nucleotide sequence of the agent, the lower the effectiveness of the agent. In addition, the method of the present application can be used to simultaneously evaluate the effectiveness of multiple agents and determine the agent or combination of agents with the highest effectiveness.

[0038] Therefore, in yet another aspect of the present application, there is provided a device for evaluating the effectiveness of a reagent or a combination of reagents for detecting pathogens, characterized in that it comprises:

[0039] Memory; and

[0040] A processor is coupled to the memory, and the processor is configured to execute the aforementioned method based on instructions stored in the memory.

[0041] In certain embodiments, the memory stores a computer program, which implements the aforementioned method when executed by the processor.

[0042] In certain embodiments, the memory further stores the database; or, the memory stores a computer program, which, when executed by the processor, can connect to a server storing the database and obtain database information.

[0043] In another aspect, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described above is implemented.

[0044] In another aspect of the present application, a kit for detecting SARS-CoV-2 is provided, wherein the kit comprises a first reagent set and a second reagent set, wherein:

[0045] The first reagent set comprises a first primer set and a first detection probe, and the first primer set and the first detection probe can target the nucleotide sequence encoding the ORF1ab protein in the SARS-CoV-2 genome; and the second reagent set comprises a second primer set and a second detection probe, and the second primer set and the second detection probe can target the nucleotide sequence encoding the N protein in the SARS-CoV-2 genome.

[0046] In certain embodiments, the first primer set comprises a first upstream primer and a first downstream primer, and the first upstream primer comprises a sequence complementary to positions 13342 to 13362 of the nucleotide sequence of the genome of SARS-CoV-2 (e.g., GenBank: MN908947).

[0047] In certain embodiments, the first primer set comprises a first upstream primer and a first downstream primer, and the first downstream primer comprises a sequence complementary to positions 13442 to 13460 of the nucleotide sequence of the genome of SARS-CoV-2 (e.g., GenBank: MN908947).

[0048] In certain embodiments, the first primer set comprises a first upstream primer and a first downstream primer, the first upstream primer comprises a sequence complementary to positions 13342 to 13362 of the nucleotide sequence of the genome of SARS-CoV-2 (e.g., GenBank: MN908947), and the first downstream primer comprises a sequence complementary to positions 13442 to 13460 of the nucleotide sequence of the genome of SARS-CoV-2 (e.g., GenBank: MN908947).

[0049] In certain embodiments, the second primer set comprises a second upstream primer and a second downstream primer, and the second upstream primer comprises a sequence complementary to positions 29145 to 29166 of the nucleotide sequence of the genome of SARS-CoV-2 (e.g., GenBank: MN908947).

[0050] In certain embodiments, the second primer set comprises a second upstream primer and a second downstream primer, and the second downstream primer comprises a sequence complementary to positions 13442 to 13460 of the nucleotide sequence of the genome of SARS-CoV-2 (e.g., GenBank: MN908947).

[0051] In certain embodiments, the second primer set comprises a second upstream primer and a second downstream primer, the second upstream primer comprises a sequence complementary to positions 29145 to 29166 of the nucleotide sequence of the genome of SARS-CoV-2 (e.g., GenBank: MN908947), and the second downstream primer comprises a sequence complementary to positions 13442 to 13460 of the nucleotide sequence of the genome of SARS-CoV-2 (e.g., GenBank: MN908947).

[0052] In certain embodiments, the first detection probe comprises a sequence complementary to positions 13377 to 13404 of the nucleotide sequence of the genome of SARS-CoV-2 (eg, GenBank: MN908947).

[0053] In certain embodiments, the second detection probe comprises a sequence complementary to positions 29179 to 29198 of the nucleotide sequence of the genome of SARS-CoV-2 (eg, GenBank: MN908947).

[0054] In certain embodiments, the complementarity is perfect complementarity.

[0055] In certain embodiments, the first primer set comprises a first upstream primer and a first downstream primer, and the first upstream primer has a nucleotide sequence as shown in SEQ ID NO: 1.

[0056] In certain embodiments, the first primer set comprises a first upstream primer and a first downstream primer, and the first downstream primer has a nucleotide sequence as shown in SEQ ID NO: 2.

[0057] In certain embodiments, the first primer set comprises a first upstream primer and a first downstream primer, the first upstream primer having a nucleotide sequence as shown in SEQ ID NO: 1, and the first downstream primer having a nucleotide sequence as shown in SEQ ID NO: 2.

[0058] In certain embodiments, the second primer set comprises a second upstream primer and a second downstream primer, and the second upstream primer has a nucleotide sequence as shown in SEQ ID NO:13.

[0059] In certain embodiments, the second primer set comprises a second upstream primer and a second downstream primer, and the second downstream primer has a nucleotide sequence as shown in SEQ ID NO:14.

[0060] In certain embodiments, the second primer set comprises a second upstream primer and a second downstream primer, the second upstream primer having a nucleotide sequence as shown in SEQ ID NO:13, and the second downstream primer having a nucleotide sequence as shown in SEQ ID NO:14.

[0061] In certain embodiments, the first detection probe has a nucleotide sequence as shown in SEQ ID NO:3.

[0062] In certain embodiments, the second detection probe has a nucleotide sequence as shown in SEQ ID NO:15.

[0063] Advantageous Effects of the Invention

[0064] The genome sequences of pathogens (especially viruses) mutate rapidly. Accordingly, reagents used to detect pathogens (especially viruses) need to be evaluated and updated promptly and quickly to obtain the most accurate detection results. The technical solution of the present invention has one or more of the following advantages over the prior art:

[0065] (1) Ability to judge, analyze, and predict the effectiveness and universality of existing nucleic acid sequence-based pathogen (especially virus) detection reagents (e.g., RT-PCR-based novel coronavirus detection reagents);

[0066] (2) The methods and devices of the present application can evaluate the effectiveness, coverage, etc. of newly designed nucleic acid sequence-based pathogen (especially virus) detection reagents (such as primers and probes), and assist in the development of diagnostic reagents.

[0067] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples, but it will be understood by those skilled in the art that the following drawings and examples are intended only to illustrate the present invention and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art based on the following detailed description of the accompanying drawings and preferred embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 A flow chart showing an exemplary embodiment of the method and apparatus of the present application is shown.

[0069] Figure 2 An interactive interface showing the software output.

[0070] Figure 3 The mutation types and mutation frequencies of the viral genome regions bound by the detection reagents (primers and probes) of USACDC-N1 are shown.

[0071] Figure 4 The results of the evaluation of different detection reagents (primers and probes) by the method of the present application are shown.

[0072] Figure 5 The mutation types and frequencies of the viral binding sites corresponding to the 5 bases at the 3' end of the primers used in the RT-PCR detection scheme are shown.

[0073] Figure 6 The effectiveness of the simultaneous use of two sets of detection reagents is evaluated by the method of the present application, wherein: Figure 6 A is the mutation type and frequency of the viral genome region bound by the ChinaCDC-ORF1ab group of reagents. Figure 6 B is the mutation type and mutation frequency of the viral genome region bound by the HKU-N group of reagents, Figure 6 C is the detection result that mutations exist in the viral genome regions bound by both groups of detection reagents, among which assay1 is the ChinaCDC-ORF1ab group reagent and assay2 is the HKU-N group reagent.

[0074] Sequence information

[0075] Information on the partial sequences involved in the present invention is provided in Table 1 below.

[0076] Table 1: Description of sequences

[0077] SEQ ID NO: describe sequence 1 ChinaCDC-ORF1ab upstream primer CCCTGTGGGTTTTACACTTAA 2 ChinaCDC-ORF1ab downstream primers ACGATTGTGCATCAGCTGA 3 ChinaCDC-ORF1ab probe FAM-CCGTCTGCGGTATGTGGAAAGGTTATGG-BHQ1 4 Charite-RdRP upstream primer GTGARATGGTCATGTGTGGCGG 5 Charite-RdRP downstream primer CARATGTTAAASACACTATTAGCATA 6 Charite-RdRP probe FAM-CCAGGTTGGWACRTCATCMGGTGATGC-BHQ1 7 Charite-E upstream primer ACAGGTACGTTAATAGTTAATAGCGT 8 Charite-E downstream primer ATATTGCAGCAGTACGCACACA 9 Charite-E probe FAM-ACACTAGCCATCCTTACTGCGCTTCG-BHQ1 10 HKU-ORF1b upstream primer TGGGGYTTTACRGGTAACCT 11 HKU-ORF1b downstream primers AACRCGCTTAACAAAGCACTC 12 HKU-ORF1b probe FAM / ZEN-TAGTTGTGATGCWATCATGACTAG-IBFQ 13 HKU-N upstream primer TAATCAGACAAGGAACTGATTA 14 HKU-N downstream primers CGAAGGTGTGACTTCCATG 15 HKU-N probe FAM / ZEN-GCAAATTGTGCAATTTGCGG-IBFQ 16 USACDC-N1 upstream primer GACCCCAAAATCAGCGAAAT 17 USACDC-N1 downstream primer TCTGGTTACTGCCAGTTGAATCTG 18 USACDC-N1 probe FAM-ACCCCGCATTACGTTTGGTGGACC-BHQ1 19 USACDC-N2 upstream primer TTACAAACATTGGCCGCAAA 20 USACDC-N2 downstream primer GCGCGACATTCCGAAGAA 21 USACDC-N2 probe FAM-ACAATTTGCCCCCAGCGTTAG-BHQ1 22 USACDC-N3 upstream primer GGGAGCCTTGAATACACCAAAA 23 USACDC-N3 downstream primer TGTAGCACGATTGCAGCATTG 24 USACDC-N3 probe FAM-AYCACATTGGCACCCGCAATCCTG-BHQ1 25 Thailand-N upstream primer CGTTTGGTGGACCCTCAGAT 26 Thailand-N downstream primer AATGGAGAACGCAGTGGGG 27 Thailand-N probe FAM-CAACTGGCAGTAACCA-BQH1 DETAILED DESCRIPTION

[0078] The present invention will now be described with reference to the following examples which are intended to illustrate the present invention (but not to limit the present invention). It will be appreciated by those skilled in the art that the examples are provided to illustrate the present invention by way of example and are not intended to limit the scope of the present invention.

[0079] Example 1. Process of the exemplary method of the present application

[0080] The flow chart of this embodiment is as follows Figure 1 As shown, the details are as follows:

[0081] 1. Processing of input information

[0082] 1.1 Input pathogen original genome sequence

[0083] This example analyzes the SARS-CoV-2 genome sequence in FASTA format (SARS-CoV-2 genome sequence data obtained from the GISAID database as of August 29, 2020). Before entering the database, the FASTA format file was preprocessed and filtered, including:

[0084] (1) Confirm the integrity of the fasta file (i.e., determine whether it completely covers the SARS-CoV-2 genome sequence) and remove incomplete genome sequences;

[0085] (2) Confirm the target host of the virus (for example, limit the target host to humans) and remove genome sequences that do not meet the target host requirements;

[0086] (3) Confirm whether each sequence corresponds to a sample and remove samples with duplicate fasta IDs;

[0087] (4) Remove special characters that are not bases (such as A, T, C, G) in the sample sequence.

[0088] The pre-processed and filtered genome sequences were input into the database, with a total of 58,162 viral genome sequences input.

[0089] 1.2 Input primer and probe sequences

[0090] In this example, the user-designed sequences of primers and probes for detecting SARS-CoV-2 by RT-PCR, as well as the specific locations of the genomic sequences targeted by the primers and probes, are prepared into an Excel spreadsheet or text file for input.

[0091] 2. Analysis and Annotation

[0092] 2.1 Sequence alignment

[0093] The input viral genome sequence was compared with the viral sequences available in the NCBI database (https: / / www.ncbi.nlm.nih.gov / nuccore / NC_045512.2 / ) using the software NUCMER (NUCleotideMUMmer) version 3.1.

[0094] 2.2 Mutation Annotation

[0095] By comparison, the mutation information in the input viral genome sequence is determined and obtained, including the type of mutation (such as SNP, deletion and insertion), position and sequence information (such as nucleotide information of the SNP site, sequence of the inserted or deleted nucleotide fragment), and these mutation information are further organized into NUCMER objects (for input of R language scripts).

[0096] The mutation information is annotated and integrated using R language functions. Users can adjust the slider on the interactive interface to select the number of observations displayed in the output image, or open a drop-down box to select different countries, genes, and image display methods. The output can be adjusted and downloaded locally.

[0097] 2.3 Notes on other information

[0098] Annotate other information of the input viral genome sequence, such other information includes: virus genus information, species information, strain information, isolation region information (country), isolation time information.

[0099] 3. Output information

[0100] 3.1 Analysis of viral genome sequences

[0101] The genome sequence information of the virus is analyzed to obtain a description of the basic situation of the sample mutation, including the single sample mutation rate, mutation type frequency, global and local single nucleotide polymorphisms and protein mutation maps, single gene mutation situation, mutation hotspots, etc.

[0102] 3.2 RT-PCR primer evaluation

[0103] Analyzing the mutation patterns and mutation rates in the viral regions targeted by designed primers and probes provides a preliminary indication of their coverage and effectiveness for detecting viral samples identified by genomic sequences in the database. As viruses evolve and viral sequence data continues to update, frequent mutations in the target regions of these reagents can often lead to high false-negative rates in detection reagents (primers and probes, etc.), reducing the reliability of molecular diagnosis of COVID-19. As the gold standard for RT-PCR detection, designing optimal primers is crucial, and the mutation rate in the targeted region is a useful evaluation method. The mutation rate obtained by the user can be viewed as an evaluation score; a higher score indicates lower primer effectiveness. Furthermore, the location of the mutation site within the primer is an important consideration. When the mutation site is located at the 3' end of the primer, the virus detection rate by RT-PCR is significantly reduced, resulting in a high number of false-negative results. This significantly reduces the effectiveness of the primer.

[0104] In addition, when users input sequence information for multiple detection reagents simultaneously, they can compare the effectiveness of different detection reagents in detecting viruses and select the best combination of detection reagents to achieve the highest detection effectiveness.

[0105] 3.3 Other aspects

[0106] In addition, users can adjust the slider on the interactive interface to select the number of observations displayed in the output image, or open the drop-down box to select different countries, genes, image display methods, etc. The output results can be adjusted and downloaded locally. Users can select input and output parameters by themselves. With just a click of the mouse, the analysis results are displayed and the download function is available. The interactive interface of the software operation is as follows Figure 2 shown.

[0107] Example 2. Evaluation of the effectiveness of reagents

[0108] According to the method described in Example 1, a database was established based on the SARS-CoV-2 genome data obtained from the GISAID database as of August 29, 2020, and annotated and analyzed. Then, according to the method described in Example 1, the mutations in the viral genomic regions bound by the primers and probes used in the USACDC-N1 assay (US Centers for Disease Control and Prevention) were analyzed to evaluate their effectiveness. The specific sequences of the primers and probes are shown in SEQ ID NOs: 16-18.

[0109] The test results are as follows Figure 3 As shown, Figure 3As indicated by the middle arrow, nearly all viruses in the database have a three-base mutation at this location: 28311:C→T. This means that the region of the viral genome bound by the primers and probe used in the USACDC-N1 assay (US Centers for Disease Control and Prevention) almost all harbor this mutation, confirming the low detection effectiveness of these primers and probes. Therefore, the method of this application can accurately assess the effectiveness of a reagent.

[0110] Example 3. Evaluation of the effectiveness of reagents

[0111] Taking the 9 sets of reagents (primers and probes) for detecting SARS-CoV-2 by RT-PCR as an example, according to the method described in Example 1, a database was established based on the SARS-CoV-2 genome data obtained from the GISAID database as of August 29, 2020. The sequence information of the 9 sets of reagents (primers and probes) was input into the program (the specific sequences of the 9 sets of primers and probes are shown in SEQ ID NO: 1-27 in Table 1) for analysis and evaluation. The specific analysis results are shown in Figure 1. Figure 4 shown.

[0112] from Figure 4 The results show that the genomic region targeted by the Charite-E group of detection reagents (SEQ ID NO: SEQ ID NO: 7-9) has the lowest mutation rate (0.29057%), and thus this group of reagents has the highest detection effectiveness. On the other hand, the genomic region targeted by the USACDC-N1 group of detection reagents (SEQ ID NO: 16-18) has the highest mutation rate (3.1653%), and thus this group of reagents has the lowest detection effectiveness. Therefore, the method of the present application can evaluate the effectiveness of multiple groups of reagents.

[0113] Example 4. Primer effectiveness evaluation

[0114] This example evaluates the mutation of the viral genome binding site corresponding to the 5 bases at the 3' end of the primers used for SARS-CoV-2 detection by RT-PCR. According to the method described in Example 1, a database was established based on the SARS-CoV-2 genome data obtained from the GISAID database as of August 29, 2020. The sequence information of 9 sets of primers was input into the program (the specific sequences of the 9 sets of primers are shown in SEQ ID NO: 1-27 in Table 1) for analysis and evaluation. The specific analysis results are shown in Figure 1. Figure 5 shown.

[0115] from Figure 5The results show that primers targeting the SARS-CoV-2 N gene have poor effectiveness, with high mutation rates in their binding regions. The Thailand-N primer set had a mutation rate of 0.22007%, the USACDC-N2 primer set had a mutation rate of 0.2132%, the USACDC-N1 primer set had a mutation rate of 0.20632%, and the USACDC-N3 primer set had a mutation rate of 0.09456%. The primers targeting the SARS-CoV-2 ORF1ab gene, namely the ChinaCDC-ORF1ab primer set, had the lowest mutation rate in their binding regions, at only 0.01719%.

[0116] Example 5. Evaluation of the effectiveness of reagents

[0117] This example evaluates the effectiveness of using two sets of reagents to detect SARS-CoV-2. According to the method described in Example 1, a database was established based on the SARS-CoV-2 genome data obtained from the GISAID database as of August 29, 2020. The sequence information of the ChinaCDC-ORF1ab and HKU-N reagents was input into the program (the specific sequences of the primers and probes of ChinaCDC-ORF1ab are shown in SEQ ID NO: 1-3, and the specific sequences of the primers and probes of HKU-N are shown in SEQ ID NO: 13-15) for analysis and evaluation. The specific analysis results are shown in Figure 1. Figure 6 shown.

[0118] from Figure 6 The results show that the two sets of reagents, ChinaCDC-ORF1ab and HKU-N, have good complementarity in detecting SARS-CoV-2. Figure 6 A shows the mutation of the virus region targeted by the ChinaCDC-ORF1ab group of reagents. Figure 6 Figure C shows that only 291 strains of viruses mutated in the region targeted by the ChinaCDC-ORF1ab group of reagents. Figure 6B shows the mutations in the viral region targeted by the HKU-N group of reagents. Figure 6 C shows that only 598 strains had mutations in the region targeted by the HKU-N test kit, while only 2 strains had mutations in the regions targeted by both the ChinaCDC-ORF1ab and HKU-N test kits. This means that the false negative rate of tests using both test kits is significantly reduced, and the effectiveness of using both sets of test kits is relatively high.

[0119] In summary, the above examples demonstrate that the method or device of the present application can effectively evaluate reagents (e.g., primers and probes) used to detect pathogens, and can simultaneously evaluate multiple groups of reagents, and the evaluation results are accurate, convenient, and universal.

[0120] Although the specific embodiments of the present invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the teachings disclosed, and these changes are within the scope of protection of the present invention. The full scope of the present invention is given by the appended claims and any equivalents thereof.

Claims

1. A method for evaluating the effectiveness of multiple reagents for detecting pathogens, wherein: The plurality of reagents all contain nucleotide sequences targeting the pathogen, and the method comprises: (1) obtaining nucleotide sequences of the plurality of agents and a database containing genome sequences of one or more of the pathogens; Furthermore, the genome sequence of the pathogen is converted into the fasta format, incomplete genome sequences are removed, genome sequences that do not meet the requirements of the pathogen's target host are removed, and non-base special characters in the genome sequence are removed; (2) annotating and / or analyzing the genome sequences of one or more of the pathogens; wherein the annotation is selected from the following information: genus information, species information, subspecies information, subtype information, subgroup information, subpopulation information, strain information, region information, age information, or any combination thereof of the pathogen from which the genome sequence is derived; Then, the nucleotide sequence is aligned with all or part of the genome sequences of the pathogen in the database, and, after the sequence alignment, information selected from the following in the genome sequence of the pathogen is determined and annotated: functional region, mutation region, mutation site, mutation type, mutation frequency, conserved region, conserved site, or any combination thereof; Then, analyzing the genome sequence of one or more of the pathogens by performing sequence alignment with a reference genome sequence, and obtaining a sequence alignment result; (3) Determining, based on the sequence comparison results, the number or ratio of pathogen genomic sequences that can be targeted by the multiple reagents, the region of the pathogen genomic sequence targeted by the reagents, the number of mutation sites contained in the region, the mutation type and / or frequency of the mutation type of the mutation site, the mutation frequency of the mutation site, the position of the mutation site in the region, or any combination thereof, thereby evaluating the effectiveness of the reagent in detecting the pathogen; Wherein, the reagent is a primer and / or a probe.

2. The method according to claim 1, wherein The pathogen is selected from bacteria, fungi, viruses, and any combination thereof.

3. The method according to claim 2, wherein: The method has one or more characteristics selected from the following: (1) The pathogen is selected from the group consisting of Mycobacterium tuberculosis, Staphylococcus, Streptococcus, Pneumococcus, Escherichia coli, Shigella, Salmonella, and any combination thereof; (2) The pathogen is selected from the group consisting of Trichophyton, Epidermophyton, Candida albicans, and any combination thereof; (3) The pathogen is selected from coronavirus, Ebola virus, hepatitis B virus, human immunodeficiency virus 1, rabies virus, human papillomavirus, and any combination thereof.

4. The method according to claim 2, wherein: The pathogen is SARS-CoV-2.

5. The method according to claim 1 or 2, wherein: The reagent detects the pathogen by sequencing, PCR, isothermal amplification, melting curve analysis, nucleic acid blot analysis, or any combination thereof.

6. The method according to claim 1 or 2, wherein: The method has one or more characteristics selected from the following: (1) The reagent is designed based on the genome sequence of a specified pathogen; (2) The reagent is designed to contain a nucleotide sequence that is partially complementary or completely complementary to the genome sequence of a designated pathogen; (3) The reagents include a forward primer, a reverse primer and a probe.

7. The method according to claim 1, wherein The database is established by obtaining and entering the genome sequences of one or more virus strains of the pathogen.

8. The method according to claim 7, wherein: The method has one or more characteristics selected from the following: (1) The database is a public database containing multiple pathogen genome sequences; (2) The database is an artificially created database; (3) The database is established by obtaining and entering the genome sequence of one or more strains of the pathogen; (4) The genome sequence information in the database has a unified data format; (5) The genome sequence information in the database is stored in FASTA format.

9. The method of claim 1, wherein the plurality of genomic sequences are derived from the same or different genus, species, subspecies, subtype, subgroup, subpopulation or strain of the pathogen.

10. The method according to claim 1, wherein sequence alignment is performed on the genome sequences of two or more pathogens.

11. The method according to claim 1, wherein in step (2), the nucleotide sequence is compared with the genome sequences of all the pathogens in the database.

12. The method according to claim 1, wherein in step (2), a portion of the genome sequence of the pathogen is selected based on the specified information, and then the nucleotide sequence is compared with the selected genome sequence.

13. A device for evaluating the effectiveness of a reagent or combination of reagents for detecting pathogens, characterized in that include: Memory; as well as A processor coupled to the memory, the processor being configured to execute the method of any one of claims 1 to 12 based on instructions stored in the memory.

14. The apparatus according to claim 13, wherein the memory stores a computer program, and when the program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

15. The device according to claim 13, wherein the memory further stores the database; or the memory stores a computer program, which, when executed by a processor, can connect to a server storing the database and obtain database information.

16. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Fusion type Taq DNA polymerase as well as preparation method and application thereof

    CN111690626A

  • Virus database method

    JP2009131242A