Primer design methods, systems, and computer storage media based on nucleic acid identity

By classifying and designing representative primers based on target nucleic acid sequences and establishing a candidate primer pool, the problems of low coverage and host contamination in multiplex PCR primer design were solved, achieving efficient pathogen detection.

CN116206685BActive Publication Date: 2026-05-05HUGOBIOTECH BEIJING CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUGOBIOTECH BEIJING CO LTD
Filing Date
2023-02-28
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing multiplex PCR primer designs suffer from low coverage and are prone to host contamination, especially when faced with a large pathogen library and viral mutations, making it difficult to design efficient, dimeric-free primer pools.

Method used

By classifying the target nucleic acid sequences, candidate primers representing the sequences are designed, and a candidate primer pool is established to exclude non-specific amplification and dimers, thus forming the final primer pool.

Benefits of technology

It achieves high-coverage primer design for a large pathogen library, reduces host contamination rate, and is suitable for complex pathogen detection scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206685B_ABST
    Figure CN116206685B_ABST
Patent Text Reader

Abstract

This application discloses a primer design method, system, and computer storage medium based on nucleic acid consistency. The method includes: obtaining a target nucleic acid sequence; classifying the target nucleic acid sequence to obtain multiple taxonomic units; using at least one target nucleic acid sequence from each taxonomic unit as a representative sequence; designing candidate primers for the representative sequence; sequentially selecting at most one pair of candidate primers from each taxonomic unit to establish a candidate primer pool; and excluding candidate primers in the candidate primer pool that exhibit non-specific amplification to obtain the final primer pool. This application is simple to operate and can handle large nucleic acid libraries, making it particularly suitable for universal primer design for pathogen libraries. Compared to similar software, the method in this application is more systematic and comprehensive, resulting in a primer pool with high coverage and low host contamination rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technology, and in particular to primer design methods, systems and computer storage media based on nucleic acid consistency. Background Technology

[0002] For conventional PCR (Polymerase Chain Reaction), a good primer pair is crucial for the success of the reaction; therefore, primer design can be considered the foundation of PCR. Many factors need to be considered when designing primers for conventional PCR, and many software programs can now perform this task. Multiplex PCR, developed from conventional PCR, is a method that performs multiple PCR reactions within a single PCR reaction system. Utilizing this characteristic, multiplex PCR can efficiently and systematically classify diseases and detect pathogens.

[0003] However, primer design for multiplex PCR is more complex than that for conventional PCR because it deals with a series of target nucleic acid sequences. Currently, there are two common methods: one is to first perform multiple sequence alignment and then use software such as DegePrime or Prider to select candidate primers, followed by using conventional PCR primer design software to verify their suitability. This method is computationally intensive and has very low primer coverage. The second method is to design primers for each target nucleic acid individually and then mix the designed primers; a representative software for this is MPprimer. This method can lead to primer dimers forming between primers.

[0004] Clinical pathogen detection generally falls into two categories: first, the pathogen to be detected is a known type; second, the type of pathogen to be detected cannot be determined, nor whether it is a mixed infection. For the first case, only the virus needs to be detected. However, if the virus has a high degree of mutation, many genomic types, and generally low nucleic acid sequence consistency, the aforementioned software is often unsuitable for designing universal primers. The second case is more complex, requiring detection based on a library of potentially infecting pathogens, making primer design even more challenging. It can be seen that the difficulty in primer design for clinical pathogen detection lies not only in the vast number of primers to design but also in the need for multifaceted considerations and the elimination of other contaminants. Therefore, designing a primer pool containing the minimum number of primers, without special primer structures, and without off-target effects has become crucial for the clinical application of multiplex PCR. Research in this area has extremely high application value but has not yet been reported. Summary of the Invention

[0005] This application provides a primer design method, system, and computer storage medium based on nucleic acid consistency to solve the problems of low coverage and easy host contamination in the prior art of multiplex PCR primer design.

[0006] On the one hand, embodiments of this application provide a primer design method based on nucleic acid consistency, including:

[0007] Obtain the target nucleic acid sequence;

[0008] The target nucleic acid sequence is classified to obtain multiple taxonomic units;

[0009] At least one target nucleic acid sequence in each taxonomic unit is used as a representative sequence;

[0010] Design candidate primers to represent the sequence;

[0011] Select at most one pair of candidate primers for each taxonomic unit to establish a candidate primer pool;

[0012] Candidate primers that exhibit non-specific amplification are excluded from the candidate primer pool to obtain the final primer pool.

[0013] On the other hand, embodiments of this application also provide a primer design system based on nucleic acid consistency, including:

[0014] Sequence acquisition module, used to acquire target nucleic acid sequences;

[0015] The sequence classification module is used to classify the target nucleic acid sequence and obtain multiple classification units;

[0016] The representative sequence lookup module is used to select at least one target nucleic acid sequence in each taxonomic unit as a representative sequence.

[0017] The candidate primer design module is used to design candidate primers representing the sequence.

[0018] The candidate primer pool creation module is used to sequentially select at most one pair of candidate primers for each taxonomic unit to create a candidate primer pool;

[0019] The primer pool filtering module is used to exclude candidate primers in the candidate primer pool that have non-specific amplification, and obtain the final primer pool.

[0020] On the other hand, embodiments of this application also provide a computer storage medium storing a plurality of computer instructions for causing a computer to execute the above-described method.

[0021] The primer design method, system, and computer storage medium based on nucleic acid consistency in this application have the following advantages:

[0022] It is easy to operate and can handle large nucleic acid libraries, making it particularly suitable for designing universal primers for pathogen libraries. Compared to similar software, the method in this application is more systematic and comprehensive, resulting in primer pools with high coverage and low host contamination rate. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating the primer design method based on nucleic acid consistency provided in this application embodiment;

[0025] Figure 2 A schematic diagram of the sequence weight of the core classification units provided in the embodiments of this application;

[0026] Figure 3 A schematic diagram illustrating the coverage of the total primer pool and the core primer pool provided for embodiments of this application;

[0027] Figure 4 A schematic diagram of primer pool coverage for five types of viruses provided in the embodiments of this application;

[0028] Figure 5 This is a schematic diagram illustrating the sequence consistency of five types of viruses provided in the embodiments of this application;

[0029] Figure 6 A schematic diagram illustrating the number of classification units for the five types of viruses provided in this application embodiment;

[0030] Figure 7 A schematic diagram of the off-target locations of the universal primers designed for degePrime in the experiment. Detailed Implementation

[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] The applicant found in their research that among commonly used multiplex PCR primer design methods, the first method—which involves multiple sequence alignment followed by degePrime (Hugerth et al., 2014) or Prider (Smolander et al., 2022)—has significant limitations due to its reliance on multiple sequence alignment. The computational complexity of the multiple sequence alignment process increases exponentially with the number of sequences. When the number of sequences exceeds 1000, most alignment software (such as MUSCLE (Edgar et al., 2004)) cannot run. Moreover, if the number of sequences is too large, the consistency between sequences will inevitably be poor, and even multiple alignment cannot obtain conserved regions, thus failing to obtain primers with high coverage. The second method—designing primers for each target nucleic acid and then mixing the designed primers—is only suitable for situations with a small number of samples to be tested. This is because as the number of target nucleic acids increases, the number of primers also increases, increasing the chance of primer dimer formation. When the number of primers reaches a certain level, dimers will inevitably form, leading to primer amplification during PCR while the target sequence remains unamplified. In clinical testing, primer design presents numerous challenges when it's impossible to determine the specific pathogen being tested or whether a co-infection is occurring. Firstly, the viral reservoir contains a large variety of viruses, each with numerous nucleic acid types (e.g., human coronaviruses have hundreds of thousands of sequences recorded in NCBI). Secondly, to detect viruses, universal primers need to be designed to maximize the effectiveness of each virus in the reservoir, increasing the risk of primer dimer formation. Furthermore, both of these situations face the problem of host genome contamination. Nucleic acids extracted during pathogen detection cannot avoid host contamination, necessitating the prediction and exclusion of off-target scenarios during primer design. Therefore, it's evident that existing primer design methods suffer from low primer coverage and high host contamination rates.

[0033] To address the problems in the prior art, this application first classifies the target nucleic acid sequences, and then designs corresponding candidate primers for the representative sequences in each classification unit. Since the target nucleic acid sequences in the same classification unit have high consistency, the designed candidate primers will also have high coverage. Next, the candidate primers in each classification unit are added to the candidate primer pool according to the principle of non-interaction. Finally, the candidate primers in the candidate primer pool are tested for off-target effects to reduce primer contamination of the host.

[0034] Figure 1 A flowchart illustrating the primer design method based on nucleic acid identity provided in this application embodiment. This application embodiment provides a primer design method based on nucleic acid identity, including:

[0035] S100, obtain the target nucleic acid sequence.

[0036] For example, a large number of nucleic acid sequences can be downloaded using EFetch, and then the nucleic acid sequences can be screened according to the requirements to obtain the target nucleic acid sequences for primer design. Finally, the target nucleic acid sequences need to be preprocessed, including deduplication.

[0037] S110 classifies the target nucleic acid sequence to obtain multiple classification units.

[0038] For example, a consistent clustering approach can be used to divide a large number of target nucleic acid sequences into multiple classification units based on consistency, so that the target nucleic acid sequences in each classification unit have high consistency.

[0039] S120 uses at least one target nucleic acid sequence in each taxonomic unit as a representative sequence.

[0040] For example, S120 specifically includes: taking the target nucleic acid sequence with the required length in the classification unit as the representative sequence; determining the similarity between other target nucleic acid sequences in the classification unit and the representative sequence, and taking the target nucleic acid sequence with the similarity reaching the similarity threshold as the representative sequence as well.

[0041] Specifically, multiple target nucleic acid sequences that meet the length requirement in each classification unit can be screened first, and then one of them can be randomly selected as a representative sequence, or the longest target nucleic acid sequence in the classification unit can be directly selected as the representative sequence. After the selection of representative sequences is completed, it is also determined whether the number of currently selected representative sequences exceeds a threshold N. This threshold N is a threshold set by the user. When the number of representative sequences exceeds the threshold, a certain number of representative sequences, such as N or a positive integer less than N, are randomly selected as the final representative sequences.

[0042] S130, design candidate primers to represent the sequence.

[0043] For example, since the target nucleic acid sequences are classified using consistent clustering in S110, the target nucleic acid sequences within the same classification unit have high consistency. In this case, existing primer design software, such as degePrime, can be used to design candidate primers. This method of classifying first and then designing candidate primers not only ensures that the designed candidate primers have high coverage of the target nucleic acid sequences within the classification unit, but also allows for the parallel execution of candidate primer design for multiple classification units, significantly reducing computational load and processing time.

[0044] Furthermore, to ensure coverage, after designing the candidate primers, the candidate primers are sorted in descending order of primer efficiency according to their respective taxonomic units.

[0045] S140, select at most one pair of candidate primers for each taxonomic unit in sequence to establish a candidate primer pool.

[0046] For example, at most one pair of candidate primers for each taxonomic unit can be selected sequentially according to the order of the sorted candidate primers to establish a candidate primer pool.

[0047] In the embodiments of this application, there are two methods for establishing the candidate primer pool. The first method is to determine whether the candidate primer to be added interacts with the candidate primers already added to the candidate primer pool. If there is no interaction, the primer is added to the candidate primer pool; if there is interaction, it is not added. When no candidate primers are added, the candidate primer pool is empty. Therefore, the first candidate primer added does not need to be judged for interaction and can be added directly. However, starting from the second candidate primer, interaction judgment is required. Only candidate primers that do not interact can be added to the candidate primer pool. Moreover, if a classification unit has already added a pair of candidate primers to the candidate primer pool, other candidate primers are no longer judged, and the judgment of candidate primers in the next classification unit is directly initiated. Only when a candidate primer in a classification unit cannot be added to the candidate primer pool due to interaction will the interaction judgment of the next pair of candidate primers in the current classification unit continue until a pair of candidate primers is added to the candidate primer pool, or all candidate primers cannot be added to the candidate primer pool due to interaction issues.

[0048] In the first method, if a taxon has no candidate primers added to its candidate primer pool, the taxon is filtered out. After establishing the candidate primer pool, it is re-established based on the candidate primers of the filtered taxon. Using this method, the candidate primer pool established based on the taxon before filtering is used as the first candidate primer pool, and the candidate primer pool established based on the taxon after filtering is used as the second candidate primer pool. This filtering and re-establishment of candidate primer pools is repeated until all taxon units have a pair of candidate primers added to their corresponding candidate primer pools.

[0049] The second method is similar to the first, but adds a backtracking step. The backtracking process is as follows: if all candidate primers in the current classification unit cannot be added to the candidate primer pool due to interaction issues, return to the previous classification unit and re-evaluate whether other candidate primers in the previous unit, excluding those already added to the candidate primer pool (if any), can be added. If such candidate primers exist, add a pair of candidate primers from the previous classification unit to the candidate primer pool and move to the next classification unit. Repeat this process until all classification units have been evaluated, resulting in the final candidate primer pool.

[0050] S150, candidate primers in the candidate primer pool that have non-specific amplification are excluded to obtain the final primer pool.

[0051] For example, once candidate primers that cause non-specific amplification are excluded from the candidate primer pool, the candidate primer pool can be called a primer pool.

[0052] In the embodiments of this application, if any two pairs of candidate primers in the candidate primer pool simultaneously match a non-target nucleic acid sequence (i.e., on the host) at the end of 9 bp or more, and the product is less than 2000, it is considered that non-specific amplification exists.

[0053] After filtering the candidate primer pools, the presence of hairpin structures and dimers is checked. If present, these primers are excluded. This application has already considered and excluded hairpin structures and dimers during primer pool design. To assess the efficiency of this exclusion process, MFEprimer can be used to verify the candidate primer pools. If the verification results show no hairpin structures or dimers, the primer pools are considered accurate and usable.

[0054] Experimental instructions

[0055] This application randomly selected five major virus classes from the NCBI (National Center for Biotechnology Information) database: human coronavirus (429 full-length genome sequences), human respiratory syncytial virus (HRSV) (1166 full-length genome sequences), rhinovirus (1384 full-length genome sequences), enterovirus (2522 full-length genome sequences), and influenza virus (23615 Infulenza A (HA) gene sequences and 9129 Infulenza B (NA) gene sequences). Among these, 5501 full-length genome sequences and 32744 CDS sequences, totaling 38666 nucleic acid sequences, were used to form a microvirus library as test data.

[0056] Then, universal primers for the aforementioned microviral database were designed using three methods: Prider, DegePrime, and the multiPrime method described in this application. Considering that the first two software programs cannot handle the large number of microviral databases, rhinovirus was chosen as the representative virus, and the efficiency of primer design for rhinovirus using Prider and DegePrime was compared with that of primer design for microviral databases using multiPrime. The comparison results showed that the total primer coverage for microviral databases designed by multiPrime reached 84.7%, with the total coverage for rhinovirus exceeding 75%, higher than the coverage of universal primers designed by Prider (<50%) and DegePrime (60%) for rhinovirus. Subsequent third-generation sequencing tests on candidate primers designed by multiPrime and candidate rhinovirus using degePrime revealed that the PCR results from multiPrime did not contain host fragments, while the results from degePrime were mostly host DNA. This indicates that the universal primers designed by multiPrime not only have the highest coverage but also the lowest host contamination rate, making them very suitable for universal primer design for large numbers of targeted nucleic acid sequences.

[0057] I. Universal Primer Design

[0058] 1. Using Prideer to design universal primers for rhinovirus. Prideer is fast, completing universal primer design in about 10 minutes. The largest primer set is group80, with 13 candidate sequences, covering 19.7% (266 / 1348) of the rhinovirus sequence. The second largest primer set is group247, with 15 candidate sequences, covering 9.6% (130 / 1348) of the rhinovirus sequence. The intersection of the two largest primer sets was used to form the pre- and post-PCR primers. This primer set only covers a maximum of 9.6% (130 / 1348) of the rhinovirus sequence. The cumulative coverage of the first 5 candidate primer sets is less than 50%. This is because the primers designed by Prideer do not contain degenerate bases, which is also why it is the fastest.

[0059] Table 1 Results of universal primer design for rhinovirus -- Prider

[0060]

[0061] 2. Universal primers for rhinovirus were designed using degePrime. DegePrime was initially designed for designing universal primers for bacteria; however, the screening primer section after the main program is not suitable for screening pathogen libraries. Therefore, this application randomly selected a universal primer with high coverage. It is worth noting that before running the main program, multiple sequence alignment software was used to align 1348 rhinovirus sequences, a step that consumed a significant amount of time (10 hours and 50 minutes). The final selected primer pair was AGGAAGGATGCAATGCT, CAAGAGTGGAGCAGAAT. This primer pair could identify 60% (827 / 1348) of the rhinovirus sequences.

[0062] 3. Universal primers were designed for the entire microvirus database using multiPrime. multiPrime has a moderate running speed; with a representative sequence threshold N of 500, the running time is approximately 10 hours, and with a representative sequence threshold N of 100, the running time is approximately 2 hours. Considering the possibility of running overnight on a server, this application set a threshold of 500 for overnight running. The results of multiPrime are in primer pool mode, divided into a total primer pool and a core primer pool. The total primer pool contains 35 primer pairs, and the core primer pool contains 22 primer pairs. Statistical results are as follows: Figure 2-6 As shown, the rhinovirus portion achieved a total coverage of over 70%, the highest detection rate compared to the previous two software programs. Furthermore, since multiPrime uses degePrime, this application searched for universal primers designed by degePrime in the total primer pool and core primer pool designed by multiPrime. The results showed that this primer was not included in the primer pool of this application. The reason for this is that this primer may form a dimer with other primers in the primer pool or cause off-target effects.

[0063] II. Universal Primer Detection

[0064] In universal primer design, the highest coverage universal primers designed with degePrime were not found in the multiPrime primer pool. To find the reason, this application used both degePrime universal primers and the multiPrime primer pool to test clinical samples. The results showed that the degePrime universal primers yielded 19,633 raw reads, of which 18,890 were clean reads. However, unfortunately, 98.44% of these reads were host sequences. In stark contrast, the multiPrime primer pool, although containing more primers, yielded over 90% rhinovirus reads. This indicates that universal primers designed solely with degePrime are prone to off-target effects if not screened.

[0065] To analyze the causes of off-target effects, this application also used the primers to predict off-target locations on the T2T version of the human genome. Figure 7 As shown, predictions revealed that this primer could form an off-target reaction with the most common primer, CP068277.2:179158299-179160349. The off-target reaction occurred because the last 12 bp of primer F matched the end of this position, while the last 9 bp of primer R matched the front end, resulting in a 783 bp PCR fragment. Although other positions were not predicted, this did not prevent the primer from being excluded, which is why this primer pair was not included in the multiPrime results.

[0066] This application also provides a primer design system based on nucleic acid consistency, the system comprising:

[0067] Sequence acquisition module, used to acquire target nucleic acid sequences;

[0068] The sequence classification module is used to classify the target nucleic acid sequence and obtain multiple classification units;

[0069] The representative sequence lookup module is used to select at least one target nucleic acid sequence in each taxonomic unit as a representative sequence.

[0070] The candidate primer design module is used to design candidate primers representing the sequence.

[0071] The candidate primer pool creation module is used to sequentially select at most one pair of candidate primers for each taxonomic unit to create a candidate primer pool;

[0072] The primer pool filtering module is used to exclude candidate primers in the candidate primer pool that have non-specific amplification, and obtain the final primer pool.

[0073] This application also provides a computer storage medium storing a plurality of computer instructions for causing a computer to execute the above-described method.

[0074] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0075] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A primer design method based on nucleic acid identity, characterized in that, include: Obtain the target nucleic acid sequence; The target nucleic acid sequence is classified to obtain multiple classification units; At least one target nucleic acid sequence in each of the aforementioned taxonomic units is used as a representative sequence; Design candidate primers for the representative sequence; Select at most one pair of candidate primers for each of the aforementioned taxonomic units in turn to establish a candidate primer pool; Candidate primers that exhibit non-specific amplification are excluded from the candidate primer pool to obtain the final primer pool; In this process, after the candidate primers are designed and obtained, the candidate primers are sorted in order of primer efficiency from high to low according to the classification unit in which they belong. After sorting, at most one pair of candidate primers for each classification unit is selected in the order of the candidate primers to establish the candidate primer pool. When selecting candidate primers for the classification unit to be added to the candidate primer pool, it is determined whether the candidate primer to be added interacts with the candidate primers already added to the candidate primer pool. If there is no interaction, it is added to the candidate primer pool; if there is an interaction, it is not added to the candidate primer pool. If no candidate primers are added to the candidate primer pool in a given classification unit, the classification unit is filtered out. After the candidate primer pool is established, the candidate primer pool is re-established based on the candidate primers of the filtered classification units. If no candidate primers are added to the candidate primer pool in the current classification unit, the system returns to the previous classification unit and re-evaluates whether the candidate primers in the previous classification unit can be added to the candidate primer pool. If such candidate primers exist, a pair of candidate primers from the previous classification unit is added to the candidate primer pool, and the system jumps to the next classification unit.

2. The primer design method based on nucleic acid consistency according to claim 1, characterized in that, The step of using at least one target nucleic acid sequence from each of the classification units as a representative sequence includes: The target nucleic acid sequence that meets the length requirement in the classification unit is used as the representative sequence; The similarity between other target nucleic acid sequences in the classification unit and the representative sequence is determined, and the target nucleic acid sequences whose similarity reaches the similarity threshold are also used as the representative sequences.

3. The primer design method based on nucleic acid consistency according to claim 2, characterized in that, Also includes: When the number of representative sequences exceeds a threshold, a certain number of representative sequences are randomly selected as the final representative sequences.

4. The primer design method based on nucleic acid identity according to claim 1, characterized in that, Also includes: The primer pool is checked to see if hairpin structures and dimers exist. If they do, the primers containing hairpin structures and dimers are excluded.

5. A primer design system based on nucleic acid identity, wherein the system applies the method described in claim 1, characterized in that, include: Sequence acquisition module, used to acquire target nucleic acid sequences; A sequence classification module is used to classify the target nucleic acid sequence to obtain multiple classification units; A representative sequence query module is used to select at least one target nucleic acid sequence in each of the classification units as a representative sequence; A candidate primer design module is used to design candidate primers for the representative sequence; The candidate primer pool establishment module is used to sequentially select at most one pair of candidate primers for each of the classification units to establish a candidate primer pool; The primer pool filtering module is used to exclude candidate primers in the candidate primer pool that have non-specific amplification, so as to obtain the final primer pool.

6. A computer storage medium, characterized in that, The computer storage medium stores a plurality of computer instructions, which are used to cause the computer to perform the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Zooplankton rrnL gene amplification primers and screening method, application and application method thereof

    CN104404154A

  • Method for simultaneously detecting multiple viruses, system for detecting multiple viruses and application of method and system

    CN115261508A