Method, device and equipment for simulating amplicon of primer on genome and storage medium
By developing a new method of simulated primers to amplify amplicons on the genome, the problems of low accuracy of amplification results in the prior art, inability to support degenerate primers and low processing efficiency are solved, and a simulated amplicon treatment with high accuracy, reliability and high efficiency are achieved.
Patent Information
- Application Number
- CN202510217721.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing simulated amplicon software has low accuracy in amplification results under high mismatch parameter settings, cannot support degenerate primers, and is not efficient when processing large-scale genome sets.
Develop a new method to simulate primers amplicons on the genome, supports high degenerate primers, allows for flexible setting of mismatch parameters, improve processing efficiency using multi-threading technology, and is developed based on Go language, which can run stably in Linux and Windows systems.
It significantly improves the accuracy and reliability of simulated amplicons, improves processing efficiency, and can effectively deal with the processing tasks of large-scale genome sets, and comprehensively make up for the shortcomings of the existing technology.
Smart Images

Figure BDA0005287898510000061 
Figure BDA0005287898510000071
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bioinformatics and relates to a method, device, equipment and storage medium for simulating primers to amplify amplicons on a genome. Background Art
[0002] In the field of gene detection, targeted sequencing has wide applications in aspects such as viral Sanger typing, qPCR quantification, ddPCR quantification, and tNGS pathogen detection. For the primer design link in the sequencing process, inclusive features are particularly crucial. Usually, degenerate primers are used to enhance the sensitivity of microbial amplification so that the primers can cover more genomes within a species or within a genotype. Primer-simulated amplicons refer to the process of simulating the amplification of target DNA fragments through computer software to optimize the efficiency and specificity of PCR or other amplification technologies. For example, before performing wet experiments, computer simulation can be used to predict inclusiveness (i.e., coverage), which can exclude some unqualified primers in advance, thereby reducing the number of wet experiment reactions and reducing time and money costs.
[0003] Currently, for software for primer-simulated amplicons, such as SnapGene, seqkit amplicon, isPCR, etc., potential target region sequences for amplification can be extracted based on the positions of upstream and downstream primers on the genome. However, the above software also has some deficiencies. For example, when the high mismatch parameter is set in seqkit amplicon, the accuracy of its amplification results is greatly reduced. When there are multiple matching sites for a pair of primers on the genome, this software only outputs the longest amplicon. However, according to the principle of competitive amplification, short amplicons are the main body consuming primers, which is obviously contrary to the actual amplification situation; isPCR does not support degenerate primers. When the current virus mutation frequency is high and degenerate primers are widely used, isPCR cannot meet the usage requirements under complex primer conditions; SnapGene software is a Windows desktop version, which is unable to handle scenarios of hundreds or thousands of genome sets and its closed-source nature limits further functional expansion and optimization, which is not conducive to wide application and collaborative development.
[0004] In view of the many defects of existing software for simulated amplicons, developing a simulated amplicon method that allows flexible selection of mismatch parameters, supports degenerate primers, and can perform high-throughput processing is of great significance for the primer design field. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art and actual needs, the present invention provides a method, device, equipment and storage medium for simulating primers to amplify amplicons on a genome, effectively improving the accuracy, reliability and processing efficiency of simulated amplicons.
[0006] To achieve this purpose, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for simulating amplicons of primers on a genome, the method comprising the following steps:
[0008] (1) Obtain the primer sequence to be simulated and the target genome sequence; if the primer to be simulated is a degenerate primer, extend the degenerate primer into all primer sequences without degenerate bases according to the degenerate base composition (replace the degenerate bases with the corresponding nucleic acid bases);
[0009] (2) Slide-window cut the target genome sequence to obtain the sequence within the window;
[0010] (3) Calculate the edit distance between the primer sequence to be simulated and the sequence within each window, and obtain the positions of the sequences within the windows whose corresponding edit distances are not greater than the mismatch number threshold on the genome;
[0011] (4) Write the positions on the genome obtained in step (3) into the lists of the upstream primer and the downstream primer respectively, where the starting position of the sliding window on the genome is the upstream primer and the ending position is the downstream primer;
[0012] (5) According to the upstream primer and the downstream primer obtained in step (4), select the two closest positions of the upstream primer and the downstream primer as the final amplicon positions, and extract the amplicon sequences.
[0013] The present invention develops a brand-new method for simulating amplicons, which focuses on supporting primers with a high degeneracy number, can accurately handle complex primer situations, allows flexible setting of the mismatch number parameter, thereby effectively improving the accuracy and reliability of amplification. It can significantly improve the operation efficiency by means of multi-threading technology to handle the processing tasks of large-scale genome sets. The program can be developed and run based on the Go language and can run stably in mainstream systems such as Linux and Windows, achieving a perfect balance between running speed and stability, comprehensively making up for the deficiencies of the existing technology, providing strong technical support and guarantee for the primer simulation amplification work related to targeted sequencing in the field of gene detection, and promoting the further development and innovation of the technology in this field.
[0014] Preferably, the method further comprises repeating steps (2)-(5) for different target genome sequences to obtain the simulated amplicons of the primer to be simulated on all target genome sequences.
[0015] Preferably, the length of the sliding cut is the length of the primer to be simulated, and the step size is 1-3 bp.
[0016] Preferably, the mismatch number threshold is 3.
[0017] Preferably, the sequence alignment process can be implemented using Go language channels and multi-threading.
[0018] As a preferred technical solution, the method for amplifying amplicons of the simulation primer on the genome includes the following steps:
[0019] (1) Obtain the sequence of the primer to be simulated and the target genome sequence; if the primer to be simulated is a degenerate primer, extend the degenerate primer according to the degenerate composition into all primer sequences without degenerate bases;
[0020] (2) Slide-window cut the target genome sequence, with the window size being the length of the primer to be simulated and the step size being 1-3 bp, to obtain the sequence within the window;
[0021] (3) Calculate the edit distance between the primer sequence and the sequence within each window, and obtain the positions on the genome of the sequences within the window with the corresponding edit distance not greater than 3;
[0022] (4) Write the positions on the genome obtained in step (3) into the lists of the upstream primer and the downstream primer respectively, with the starting position of the sliding window on the genome being the upstream primer and the ending position being the downstream primer;
[0023] (5) According to the upstream primer and the downstream primer obtained in step (4), select the two closest positions of the upstream primer and the downstream primer as the final amplicon positions, and extract the amplicon sequences;
[0024] (6) Repeat steps (2)-(5) for different target genome sequences to obtain the simulated amplicons of the primer to be simulated on all target genome sequences.
[0025] In a second aspect, the present invention provides a device for amplifying amplicons of a simulation primer on a genome, and the device includes:
[0026] A sequence acquisition unit for performing operations including obtaining the sequence of the primer to be simulated and the target genome sequence; if the primer to be simulated is a degenerate primer, extending the degenerate primer according to the degenerate composition into all primer sequences without degenerate bases;
[0027] A sequence cutting unit for performing operations including slide-window cutting the target genome sequence to obtain the sequence within the window;
[0028] A calculation unit for performing operations including calculating the edit distance between the primer sequence and the sequence within each window, and obtaining the positions on the genome of the sequences within the window with the corresponding edit distance not greater than the mismatch number threshold;
[0029] A matching unit for performing operations including writing the positions on the genome obtained by the calculation unit into the lists of the upstream primer and the downstream primer respectively, with the starting position of the sliding window on the genome being the upstream primer and the ending position being the downstream primer;
[0030] An extraction unit for performing operations including selecting two positions closest to the upstream primer and the downstream primer from the upstream primer and the downstream primer obtained in the matching unit as the final amplicon positions, and extracting the amplicon sequences.
[0031] Preferably, the method further includes operating on the repetitive sequence cutting unit, the calculation unit, the matching unit, and the extraction unit for different target genomic sequences to obtain the simulated amplicons of the primer to be simulated in all target genomic sequences.
[0032] Preferably, the length of the sliding cut is the length of the primer to be simulated, and the step size is 1-3 bp.
[0033] Preferably, the mismatch number threshold is 3.
[0034] In a third aspect, the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for simulating the amplicon of the primer on the genome as described in the first aspect are implemented.
[0035] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method for simulating the amplicon of the primer on the genome as described in the first aspect are implemented.
[0036] Compared with the prior art, the present invention has at least the following beneficial effects:
[0037] The present invention designs a brand-new method for simulating the amplicon of the primer on the genome, which focuses on supporting primers with a high degeneracy number, can accurately handle complex primer situations, allows flexible setting of the mismatch number parameter, thereby effectively improving the accuracy and reliability of amplification, can significantly improve the operation efficiency by means of multi-threading technology to handle the processing tasks of large-scale genomic sets, is developed based on the Go language, can run stably in mainstream systems such as Linux and Windows, achieves a perfect balance between running speed and stability, comprehensively makes up for the deficiencies of the prior art, provides strong technical support and guarantee for the primer simulation amplification work related to targeted sequencing in the field of gene detection, and promotes the further development and innovation of the technology in this field. Specific Embodiments
[0038] The technical solutions of the present invention will be further described below through specific embodiments. However, the following examples are only simple examples of the present invention and do not represent or limit the scope of the protection of the rights of the present invention. The scope of protection of the present invention shall be subject to the claims.
[0039] For those without specific technical or conditions noted in the examples, follow the techniques or conditions described in the literature in this field or the product specifications. For reagents or instruments without the manufacturer noted, they are all conventional products that can be obtained through regular channels.
[0040] In the specific embodiments of the present invention, the seqkit amplicon software is a cross-platform and ultra-fast toolkit developed based on the MIT protocol by Shen Wei of Chongqing Medical University, and is specifically used for FAST / Q file operations. Among them, amplicon is a subroutine of seqkit, and its function is to extract amplicons (or specific regions around them) through primers. When using this software, the following points need to be noted:
[0041] 1. The software only outputs the longest matching position;
[0042] 2. Mismatches are allowed, and the user can adjust the number of threads and the parameters for allowing mismatches according to needs;
[0043] 3. The software supports degenerate bases, but they are not allowed to be used in regular expressions.
[0044] The isPCR software is a PCR simulation software developed by Jim Kent of UCSC (University of California, Santa Cruz). When using this software, a sequence database and a primer input file containing three columns of information (name, forward primer, reverse primer) need to be provided. Both of the above-mentioned software are obtained from BIOCONDA.
[0045] Example 1
[0046] This example provides a method for simulating amplicons.
[0047] (1) Parameters
[0048] Required parameters: upstream primer sequence, downstream primer sequence; reference genome set FASTA file;
[0049] Optional parameters: maximum number of mismatches, default is 3, the number of mismatches is the maximum number of different bases allowed between the primer and the template; maximum number of concurrent threads, default is 1; output file path name.
[0050] (2) Calculation method
[0051] Open the multi-threaded sequence channel, iterate through the genomes and write them into the sequence channel one by one, and allocate the number of concurrent extraction tasks at the same time according to the thread limit. The extraction method is as follows:
[0052] 1) If the primer is a degenerate primer, the degenerate base primer needs to be extended according to the degenerate composition into all non-degenerate primer sequences and stored in a list.
[0053] The degeneracy rule table is shown in Table 1.
[0054] Table 1
[0055]
[0056]
[0057] For example, a degenerate primer sequence is:
[0058] CG(Y)TGGATGCG(N)TTCATGA.
[0059] The extended primer list is:
[0060] CG(C)TGGATGCG(A)TTCATGA;
[0061] CG(C)TGGATGCG(T)TTCATGA;
[0062] CG(C)TGGATGCG(G)TTCATGA;
[0063] CG(C)TGGATGCG(C)TTCATGA;
[0064] CG(T)TGGATGCG(A)TTCATGA;
[0065] CG(T)TGGATGCG(T)TTCATGA;
[0066] CG(T)TGGATGCG(G)TTCATGA;
[0067] CG(T)TGGATGCG(C)TTCATGA.
[0068] 2) Slide a window to cut the genomic sequence. The window size is the primer length, and the step size is 1 to obtain the sequence set within the window;
[0069] For example, the genomic sequence is:
[0070] GTGAATGAAGATGGCGTCTAACGACGCTGCCACTGCGACCGCTGG;
[0071] It can be cut into:
[0072] GTGAATGAAGATGGCGTCTA
[0073] TGAATGAAGATGGCGTCTAA
[0074] GAATGAAGATGGCGTCTAAC
[0075] AATGAAGATGGCGTCTAACG
[0076] ATGAAGATGGCGTCTAACGA
[0077] ……。
[0078] 3) Iterative primer list, align the primers with the window sequence set, calculate the edit distance between the primers and each sequence within the window, and obtain the genomic positions of the sequences within the window whose edit distance is within the mismatch number threshold.
[0079] 4) Write the genomic positions obtained in step 3) into the lists of upstream and downstream primers respectively. The starting position of the sliding window on the genome is the upstream primer, and the ending position is the downstream primer.
[0080] 5) According to the principle of PCR competitive amplification, select the two closest positions of the upstream and downstream primers as the final simulated amplicon positions, that is, the shortest amplicon, and extract the corresponding sequences and add them to the final result list.
[0081] 6) Continue to iterate the genome set, repeat steps 2)-5) above, simulate the amplicons of the primers in all target genomes. This step can specify the number of threads to open channels for concurrency to improve the analysis speed; write the final result list into the result document.
[0082] Example 2
[0083] In this example, the sequence of SEQ ID NO.1 is taken as an example to further verify the method of the present invention.
[0084] SEQ ID NO.1:
[0085] GTGAATGAAGATGGCGTCTAACGACGCTGCCACTGCGACCGCTGGCACTACACCTTTTGCTGTTCTAACGACGCTGCCACTGCTGGGTCTTAGACGCCAGTTGTTCTCGGTGGTATTGGCATGGTCCTTGGTTTCACCAAAGAGAGGATTGGCCGACTACTGAGCAGGATTAGATACTATGTCATTCAGCTGCGCGTAGA GCCAGGATTAGATACTATGTCAAGTGTTTCCAAGAAATGATCTACTCACT.
[0086] The degree of virus diversity is high, and it is difficult to find conserved regions. Sometimes primers have to be designed at multiple homologous sequence positions within the genome. The specific situation is as follows. The upstream primer sequence is: TCTAACGACGCTGCCACTGC, and the downstream primer sequence is: TGACATAGTATCTAATCCTG.
[0087] Refer to the method in Example 1 to simulate primer amplicons, and the output result is SEQ ID NO.2:
[0088] TCTAACGACGCTGCCACTGCGACCGCTGGCACTACACCTTTTGCTGTTCTAACGACGCTGCCACTGCTGGGTCTTAGACGCCAGTTGTTCTCGGTGGTATTGGCATGGTCCTTGGTTTCACCAAAGAGAGGATTGGCCGACTACTGAGCAGGATTAGATACTATGTCA.
[0089] Use the seqkit amplicon software to simulate amplicons. Set the parameter -P to search only on the positive strand (since all are single-stranded RNA viruses in the examples); the -m 3 parameter allows a maximum of 3 mismatched bases, which is consistent with the method of the present invention; the -j4 parameter uses 4 threads, and the output result is SEQ ID NO.3:
[0090] TCTAACGACGCTGCCACTGCGACCGCTGGCACTACACCTTTTGCTGTTCTAACGACGCTGCCACTGCTGGGTCTTAGACGCCAGTTGTTCTCGGTGGTATTGGCATGGTCCTTGGTTTCACCAAAGAGAGGATTGGCCGACTACTGAGCAGGATTAGATACTATGTCATTCAGCTGCGCGTAGAGCCAGGATTAGATACTATGTCA.
[0091] It can be seen that the method of the present invention selects the final amplicon length of 168bp according to the minimum amplicon length of the upstream and downstream primers. While seqkit amplicon selects according to the longest amplicon, and the final length is 206bp. If the middle fragment is longer, seqkit amplicon will obtain a longer amplicon, which is obviously an incorrect amplification product.
[0092] Example 3
[0093] This example simulates the primer amplicons of Echovirus type 11.
[0094] The upstream primer sequence is: RCARTCCTCHTGYGTRYTRTGYA, and the downstream primer sequence is: CYTTYARYARYCGGACRGAGAA. In the NCBI Virus database (reference website: https: / / www.ncbi.nlm.nih.gov / labs / virus / vssi / # / virus?SeqType_s=Nucleotide), 99 complete genomes of Echovirus 11 were screened according to the species number (taxid) of "12078", the genome completeness (Completeness_s) of "complete", and the host (HostLineage_ss) of Homo sapiens (human).
[0095] The amplicon extraction was performed using the method of the present invention and seqkit amplicon respectively. Due to the existence of degenerate primers, isPCR could not be used.
[0096] The sensitivity is the number of genomes successfully simulated and amplified, and the accuracy is determined based on the length of the target region and the position of the BLAST alignment to the reference genome. The results are shown in Table 2, indicating that the method of the present invention has high sensitivity and accuracy.
[0097] Table 2
[0098] Experimental group Sensitivity Accuracy The method of the present invention 99(100%) 99(100%) seqkit amplicon 85(85.85%) 85(100%)
[0099] Example 4
[0100] This example simulates the primer amplicons of Norovirus GII.
[0101] The upstream primer sequence is: CARGARBCNATGTTYAGRTGGATGAG, and the downstream primer sequence is: CCRCCNGCATRHCCRTTRTACAT. In the NCBI Virus database, 605 complete genomes of GII genotyped Norovirus were screened according to the species number (taxid) of "122929", the genome completeness (Completeness_s) of "complete", and the host (HostLineage_ss) of Homo sapiens (human).
[0102] The amplicon extraction was performed using the method of the present invention and seqkit amplicon respectively. Due to the existence of degenerate primers, isPCR could not be used. The results are shown in Table 3, indicating that the method of the present invention has high sensitivity and accuracy.
[0103] Table 3
[0104] Experimental group Sensitivity Accuracy The method of the present invention 605(100%) 605(100%) seqkit amplicon 575(95.04%) 575(100%)
[0105] Example 5
[0106] This example simulates the primer amplicons of the novel coronavirus.
[0107] The upstream primer sequence is CCCTGTGGGTTTTACACTTAA, and the downstream primer sequence is ACGATTGTGCATCAGCTGA. According to the species number (taxid) "2697049", genome integrity (Completeness_s) "complete", and host (HostLineage_ss) Homo sapiens (human) in the NCBI Virus database, 1000 complete genomes of the novel coronavirus + reference genome NC_045512.2 were randomly downloaded.
[0108] The method of the present invention, seqkit amplicon, and isPCR (set the parameter -maxSize = 2000 to allow the output of the amplicon sequence with a maximum length of 2000 bp, which is consistent with the method of the present invention, and other parameters are default) were respectively used for amplicon extraction. The results are shown in Table 4, indicating that the method of the present invention has high sensitivity and accuracy.
[0109] Table 4
[0110] Sensitivity Accuracy The method of the present invention 1001(100%) 1001(100%) seqkit amplicon 1001(100%) 1001(100%) isPCR 999(99.80%) 999(100%)
[0111] Example 6
[0112] The present invention provides a device for simulating primer amplicons on a genome, and the device includes:
[0113] A sequence acquisition unit for performing operations including obtaining the primer sequence to be simulated and the target genome sequence; if the primer to be simulated is a degenerate primer, then extending the degenerate primer according to the degenerate composition into all primer sequences without degenerate bases;
[0114] A sequence cutting unit for performing operations including sliding window cutting of the target genome sequence to obtain the sequence within the window;
[0115] A calculation unit for performing operations including calculating the edit distance between the primer sequence and the sequence within each window, and obtaining the positions of the sequences within the windows with the corresponding edit distances not greater than the mismatch number threshold on the genome;
[0116] A matching unit for performing operations including writing the positions on the genome obtained by the calculation unit into the lists of the upstream primer and the downstream primer respectively, with the starting position of the sliding window on the genome being the upstream primer and the ending position being the downstream primer;
[0117] An extraction unit for performing operations including selecting, based on the upstream primer and the downstream primer obtained from the matching unit, the two positions closest to the upstream primer and the downstream primer as the final amplicon positions, and extracting the amplicon sequences.
[0118] Example 7
[0119] This embodiment provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for simulating primer amplification of amplicons on a genome as described in Example 1 are implemented.
[0120] Those skilled in the art should understand that various forms of hardware, software, firmware, dedicated processors, or combinations thereof can be used to obtain the device of the present invention.
[0121] Example 8
[0122] This embodiment provides a computer-readable storage medium storing a computer program with program code that, when running in a corresponding processor, controller, computing device, or terminal, implements the steps of the method for simulating primer amplification of amplicons on a genome as described in Example 1.
[0123] In summary, the present invention designs a brand-new method for simulating primer amplification of amplicons on a genome, which focuses on supporting primers with a high degeneracy number, can accurately handle complex primer situations, allows flexible setting of the mismatch number parameter, thereby effectively improving the accuracy and reliability of amplification. It can significantly improve the running efficiency by means of multi-threading technology to handle the processing tasks of large-scale genome sets. Developed based on the Go language, it can run stably in mainstream systems such as Linux and Windows, achieving a perfect balance between running speed and stability, comprehensively making up for the deficiencies of the existing technology, providing strong technical support and guarantee for the primer simulation amplification work related to targeted sequencing in the field of gene detection, and promoting the further development and innovation of the technology in this field.
[0124] The applicant declares that the above description is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily thought of within the technical scope disclosed by the present invention fall within the protection scope and the disclosure scope of the present invention.
Claims
1. A method for simulating primers to amplify a genome, characterized in that: The method comprises the following steps: (1) obtaining a primer sequence to be simulated and a target genome sequence; if the primer to be simulated is a degenerate primer, extending the degenerate primer into all primer sequences without degenerate bases according to the degenerate base composition; (2) Sliding window cutting the target genome sequence to obtain the sequence within the window; (3) calculating the edit distance between the primer sequence to be simulated and the sequence in each window, and obtaining the position of the sequence in the window on the genome whose edit distance is not greater than the mismatch number threshold; (4) writing the positions on the genome obtained in step (3) into the lists of upstream primers and downstream primers respectively, with the starting position of the sliding window on the genome being the upstream primer and the ending position being the downstream primer; (5) According to the upstream primer and the downstream primer obtained in step (4), the two closest positions of the upstream primer and the downstream primer are selected as the final amplicon positions, and the amplicon sequence is extracted.
2. The method for amplifying a genome using a simulated primer according to claim 1, characterized in that: The method further comprises repeating steps (2) to (5) for different target genome sequences to obtain simulated amplicons of the primers to be simulated in all target genome sequences.
3. The method for amplifying a genome using a simulated primer according to claim 1 or 2, characterized in that: The length of the sliding cut is the length of the primer to be simulated, and the step length is 1 to 3 bp.
4. The method for amplifying a genome using a simulated primer according to any one of claims 1 to 3, characterized in that: The mismatch number threshold is 3.
5. The method for amplifying a genome using a simulated primer according to any one of claims 1 to 4, characterized in that: The method comprises the following steps: (1) obtaining a primer sequence to be simulated and a target genome sequence; if the primer to be simulated is a degenerate primer, extending the degenerate primer into a primer sequence of all non-degenerate bases according to the degenerate composition; (2) Sliding window cutting of the target genome sequence, with the window size being the length of the primer to be simulated and the step length being 1 to 3 bp, to obtain the sequence within the window; (3) calculating the edit distance between the primer sequence and the sequence in each window, and obtaining the position of the sequence in the window on the genome whose edit distance is not greater than 3; (4) writing the positions on the genome obtained in step (3) into the lists of upstream primers and downstream primers respectively, with the starting position of the sliding window on the genome being the upstream primer and the ending position being the downstream primer; (5) According to the upstream primer and the downstream primer obtained in step (4), the two closest positions of the upstream primer and the downstream primer are selected as the final amplicon positions, and the amplicon sequence is extracted; (6) Repeat steps (2) to (5) for different target genome sequences to obtain simulated amplicons of the primers to be simulated in all target genome sequences.
6. A device for simulating primers to amplify a genome, characterized in that: The device comprises: A sequence acquisition unit, used to execute the steps including acquiring a primer sequence to be simulated and a target genome sequence; if the primer to be simulated is a degenerate primer, extending the degenerate primer into a primer sequence of all non-degenerate bases according to the degenerate composition; A sequence cutting unit, used for executing sliding window cutting of the target genome sequence to obtain a sequence within the window; A calculation unit, configured to perform operations including calculating the edit distance between the primer sequence and each sequence in the window, and obtaining the position of the sequence in the window on the genome whose corresponding edit distance is not greater than a mismatch number threshold; A matching unit, used for executing the steps including writing the positions on the genome acquired by the calculation unit into the lists of upstream primers and downstream primers respectively, wherein the starting position of the sliding window on the genome is the upstream primer, and the ending position is the downstream primer; The extraction unit is used to perform the steps of selecting the two closest positions of the upstream primer and the downstream primer as the final amplicon positions according to the upstream primer and the downstream primer obtained in the matching unit, and extracting the amplicon sequence.
7. The device for simulating primers to amplify a genome according to claim 6, characterized in that: The method further comprises operating a repeating sequence cutting unit, a calculating unit, a matching unit and an extracting unit on different target genome sequences to obtain simulated amplicons of the primers to be simulated in all target genome sequences.
8. The device for simulating primers to amplify a genome according to claim 6 or 7, characterized in that: The length of the sliding cut is the length of the primer to be simulated, and the step length is 1 to 3 bp; Preferably, the mismatch number threshold is 3.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method for simulating primers to amplify a genome as described in any one of claims 1 to 5 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for simulating primers to amplify a genome as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Design of primers for amplicon sequencing and construction method of amplicon sequencing library
CN110957005A
Primer design method and system based on k-mer algorithm
CN111326210A
Quality control method of library tag primer and application thereof
CN114807309A
Degenerate primer design method and related equipment thereof
CN119446268A
Methods of detecting differences in genomic sequence representation
US20040185477A1