System, method and device for designing target species-specific primers based on information entropy

Entropy analysis in primer design optimizes PCR by selecting conservative binding sites and avoiding high-entropy regions, addressing SNP-induced misalignment issues and enhancing primer specificity and amplification efficiency.

CN119495361BActive Publication Date: 2025-07-15HUGOBIOTECH BEIJING CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510080362.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-07-15
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Existing primer design methods are difficult to accurately quantify the effects of single nucleotide polymorphisms (SNPs) in conserved regions, especially in evaluating the specific effects of each base site on primer performance, resulting in false negative results of PCR detection and bias in amplification products.

Method used

Introducing information entropy analysis, by calculating the entropy value of the genomic sequence, screening out conserved regions as primer binding sites, and avoiding high entropy sites, especially the 3' end of the primer, primers with higher specificity and amplification efficiency are designed.

Benefits of technology

It significantly improves the design quality of primers and the reliability of PCR detection, reduces the risk of primer mismatch, and ensures amplification efficiency and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119495361B_ABST
    Figure CN119495361B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of biotechnological primer design, and particularly to a system, method and device for designing target species-specific primers based on information entropy. The system includes a data acquisition module, a data screening module and a data optimization module. The data screening module is used to obtain short sequence fragment data with identity and coverage greater than or equal to 80%. The data optimization module is used to process the short sequence fragment data with identity and coverage greater than or equal to 80%, screen out short sequence fragments covering 90% of the genome, obtain the entropy value of each base in the short sequence fragment, quantitatively score the short sequence fragment according to the entropy value and coverage information, screen out the short sequence fragments with higher scores, and replace the bases with the most conservative bases to obtain optimized short sequence fragments, and design primers based on these short sequence fragments. The introduction of entropy value analysis in the primer design process by this system significantly improves the design quality and application efficiency of primers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology primer design, and particularly to a system, method, and device for designing target species-specific primers based on information entropy. Background Art

[0002] Since its invention in the 1980s, polymerase chain reaction (PCR) has become an important tool in the detection of pathogenic microorganisms. The core advantage of this technology lies in its high sensitivity and specificity, enabling researchers to detect extremely low concentrations of microbial DNA in samples. Therefore, PCR has been widely applied in fields such as infectious disease diagnosis, environmental monitoring, and food safety detection. The key to the success of PCR lies in primer design. Primers need to be highly specific to ensure binding only to the DNA of the target pathogen. If the binding site of the primer contains single nucleotide polymorphisms (SNPs), especially SNPs located at the 3' end of the primer, it may lead to a decrease in the binding efficiency of the primer to the target sequence, affecting the amplification efficiency and resulting in false negative results. In addition, amplification product bias may also lead to an incorrect estimation of the target DNA concentration.

[0003] Primer design usually preferentially selects conserved regions in the target sequence as binding sites to avoid the influence of SNPs as much as possible. However, existing design methods are difficult to accurately quantify the influence of SNPs in conserved regions, especially in evaluating the specific influence of each base site in these regions on primer performance. Summary of the Invention

[0004] Based on the above technical problems, the present invention innovatively introduces entropy value analysis in the primer design process, significantly improving the design quality and application efficiency of primers, and providing a more reliable solution for PCR-related fields.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] In a first aspect, the present invention provides a system for designing target species-specific primers based on information entropy, the system comprising the following modules:

[0007] A data acquisition module for obtaining the genomic sequences of the target species and its related species;

[0008] A data screening module for processing the reference genomic sequence of the target species to obtain short sequence fragment data with identity and coverage greater than or equal to 80%;

[0009] A data optimization module for processing short sequence fragment data with identity and coverage greater than or equal to 80%, screening out short sequence fragment data covering 90% of the genome, obtaining the entropy value of each base in the screened short sequence fragments, quantitatively scoring the short sequence fragments covering 90% of the genome according to the entropy value and coverage information, screening the short sequence fragments covering 90% of the genome with the top scores, and replacing the bases with the most conservative bases to obtain optimized short sequence fragment data;

[0010] A primer design module for designing primers based on the optimized short sequence fragment data and ensuring that the 3' end of the primers avoids high-entropy sites to obtain candidate primers;

[0011] A primer quality assessment module for assessing the quality of the candidate primers to obtain candidate primers meeting the assessment conditions.

[0012] Preferably, the data acquisition module includes the following units:

[0013] A first data acquisition unit for processing all genomic sequences of the target species to obtain the reference genomic sequence of the target species and other genomic sequences of the target species;

[0014] A second data acquisition unit for obtaining the reference genomic sequence of a closely related species of the target species.

[0015] Preferably, the data screening module includes the following units:

[0016] A splitting unit for splitting the reference genomic sequence of the target species into short sequence fragments;

[0017] A first alignment unit for aligning the short sequence fragments to the reference genomic sequence of a closely related species of the target species and screening out the unaligned short sequence fragments as candidate fragments;

[0018] A second alignment unit for aligning the candidate fragments to other genomic sequences of the target species and screening out the aligned short sequence fragments as target fragments;

[0019] A calculation unit for calculating the identity and coverage of the target fragments in other genomic sequences of the target species and screening out the target fragments with identity and coverage greater than or equal to 80% as optimized fragments.

[0020] Preferably, the data optimization module includes the following units:

[0021] A first calculation unit for counting the coverage of the optimized fragments in all genomic sequences of the target species, screening out short sequence fragments covering 90% of the genome, and obtaining all aligned sequences;

[0022] A comparison unit, configured to perform multiple sequence alignment on the target fragment sequences screened by the first calculation unit and all the aligned sequences in other genomes of the target species, generate an alignment sequence with position alignment, and obtain an alignment matrix;

[0023] A second calculation unit, configured to calculate the entropy value of each base site in the alignment matrix to obtain the entropy value of the optimized fragment;

[0024] A scoring unit, configured to perform quantitative scoring on the optimized fragments according to the entropy value of the optimized fragments and the coverage information of the optimized fragments, and screen the optimized fragments with higher scores;

[0025] An optimization unit, configured to optimize the bases in the optimized fragments with higher scores, replace the bases with the most conservative bases in the multiple sequence alignment, and mark the base sites with high entropy values.

[0026] Preferably, the quality assessment includes one or more of primer dimer analysis, primer conservation and genome coverage assessment, primer off-target analysis, and amplification product species specificity assessment.

[0027] Preferably, the target species is a pathogenic microorganism species, and the pathogenic microorganism species is bacteria, fungi or viruses.

[0028] More preferably, the pathogenic microorganism species is bacteria.

[0029] Further preferably, the pathogenic microorganism species is Streptococcus pneumoniae.

[0030] In a second aspect, the present invention provides a method for designing target species-specific primers based on information entropy, the method comprising the following steps:

[0031] Obtain the genomic sequences of the target species and its related species;

[0032] Process the reference genomic sequence of the target species to obtain short sequence fragment data with identity and coverage greater than or equal to 80%;

[0033] Process the short sequence fragment data with identity and coverage greater than or equal to 80%, screen the short sequence fragment data covering 90% of the genome, obtain the entropy value of each base in the short sequence fragments covering 90% of the genome, perform quantitative scoring on the short sequence fragments covering 90% of the genome according to the entropy value and coverage information, screen the short sequence fragments covering 90% of the genome with higher scores, and replace the bases with the most conservative bases to obtain optimized short sequence fragment data;

[0034] Design primers according to the optimized short sequence fragment data, and ensure that the 3' end of the primers avoids high entropy value sites to obtain candidate primers;

[0035] Quality assessment is performed on the candidate primers to obtain candidate primers that meet the assessment conditions.

[0036] Preferably, the target species is a pathogenic microorganism species, and the pathogenic microorganism species is a bacterium, a fungus or a virus.

[0037] More preferably, the pathogenic microorganism species is a bacterium.

[0038] Even more preferably, the pathogenic microorganism species is Streptococcus pneumoniae.

[0039] In a third aspect, an apparatus for designing target species-specific primers based on information entropy is provided, and the apparatus includes the system as described in the present invention.

[0040] In a fourth aspect, a readable storage medium is provided, and a plurality of computer programs are stored in the readable storage medium, and when the computer programs run on a processor, they execute the method as described in the present invention.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] (1) The introduction of the concept of information entropy in the present invention provides strong support for the selection of conserved regions and primer binding sites. The entropy in the genome of the present invention borrows the concept of entropy in information theory, which is a method for quantifying the complexity or information uncertainty of sequences in the genome and reflects the degree of sequence diversity. In the genome, regions with lower entropy values are often conserved regions, where the base composition is stable and the mutation rate is low, which are suitable as primer binding sites. Regions with higher entropy values usually have stronger variability, such as non-coding regions or highly variable regions, which are not suitable as binding sites. By analyzing the entropy value distribution of the genomic sequence, primer binding sites can be more scientifically screened, primer design can be optimized, and thus the reliability and accuracy of PCR detection can be improved.

[0043] (2) The method for designing target species-specific primers based on information entropy in the present invention pre-computes the conserved fragments of the target species genome and the entropy value of each base site, evaluates the conservation of the fragments based on information entropy and sorts them, and at the same time avoids high-entropy sites in primer design, thereby improving the specificity and amplification efficiency of the primers.

[0044] (3) Traditional primer design methods mainly focus on the conservation and specificity of the primer region, but do not effectively evaluate the number of SNPs in the region and the primer mismatches that may be caused. The present invention accurately screens out high-SNP regions and avoids them by calculating the entropy value, especially avoiding the 3' end of the primer being located at high-entropy sites.

[0045] (4) If the primer is located at the SNP site and the amplified target sequence does not match the base of the primer at this site, it will significantly affect the amplification efficiency. Through entropy value screening and fragment optimization, the present invention effectively reduces the risk of primer mismatch and ensures that the designed primers reach a higher level in terms of conservativeness and amplification efficiency.

[0046] (5) The present invention provides a complete and scientific optimization process from genome screening to primer design, evaluation and verification, which can quickly design specific primers with strong conservativeness and high amplification efficiency.

[0047] In summary, the present invention innovatively introduces entropy value analysis in the primer design process, significantly improving the design quality and application efficiency of primers, and providing a more reliable solution for the PCR-related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of the working process of the system for designing target species-specific primers based on information entropy provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0049] The present invention will be specifically described below in combination with the specific embodiments and examples, and the advantages and various effects of the present invention will be presented more clearly therefrom. Those skilled in the art should understand that these specific embodiments and examples are used to illustrate the present invention, rather than limiting the present invention.

[0050] Example 1

[0051] This example provides a system for designing target species-specific primers based on information entropy, and the system includes the following modules:

[0052] A data acquisition module for obtaining the genomic sequences of the target species and its related species;

[0053] A data screening module for processing the reference genomic sequence of the target species to obtain short sequence fragment data with identity and coverage greater than or equal to 80%;

[0054] A data optimization module for processing the short sequence fragment data with identity and coverage greater than or equal to 80%, screening out the short sequence fragment data covering 90% of the genome, obtaining the entropy value of each base in the short sequence fragment data covering 90% of the genome, quantitatively scoring the short sequence fragment data covering 90% of the genome according to the entropy value and coverage information, screening out the short sequence fragment data covering 90% of the genome with the top scores, and replacing the bases with the most conservative bases to obtain the optimized short sequence fragment data;

[0055] A primer design module for designing primers according to the optimized short sequence fragment data and ensuring that the 3'-end of the primer avoids high-entropy sites to obtain candidate primers;

[0056] A primer quality assessment module for assessing the quality of candidate primers to obtain candidate primers that meet the assessment criteria.

[0057] In this embodiment, the data acquisition module includes the following units:

[0058] A first data acquisition unit for processing all genomic sequences of the target species to obtain a reference genomic sequence of the target species and other genomic sequences of the target species;

[0059] A second data acquisition unit for obtaining a reference genomic sequence of a related species of the target species.

[0060] In this embodiment, the data screening module includes the following units:

[0061] A splitting unit for splitting the reference genomic sequence of the target species into short sequence fragments;

[0062] A first alignment unit for aligning the short sequence fragments to the reference genomic sequence of a related species of the target species and screening the unaligned short sequence fragments as candidate fragments;

[0063] A second alignment unit for aligning the candidate fragments to other genomic sequences of the target species and screening the aligned short sequence fragments as target fragments;

[0064] A calculation unit for calculating the identity and coverage of the target fragments in other genomic sequences of the target species and screening the target fragments with an identity and coverage greater than or equal to 80% as optimized fragments.

[0065] In this embodiment, the data optimization module includes the following units:

[0066] A first calculation unit for counting the coverage of the optimized fragments in all genomic sequences of the target species, screening the short sequence fragments with a coverage greater than 90% to obtain all aligned sequences;

[0067] An alignment unit for performing multiple sequence alignment on the target fragment sequences screened by the first calculation unit and all aligned sequences in other genomes of the target species to generate a position-aligned alignment sequence and obtain an alignment matrix;

[0068] A second calculation unit for calculating the entropy value of each base site in the alignment matrix to obtain the entropy value of the optimized fragment;

[0069] A scoring unit for quantitatively scoring the optimized fragments according to the entropy value of the optimized fragments and the coverage information of the optimized fragments, and screening the optimized fragments with higher scores;

[0070] Optimization unit, which is used to optimize the bases in the optimized fragments with higher scores, replace the bases with the most conserved bases in the multiple sequence alignment, and mark the base sites with high entropy values.

[0071] In this embodiment, the quality assessment includes one or more of primer dimer analysis, primer conservation and genome coverage assessment, primer off-target analysis, and amplification product species specificity assessment.

[0072] In this embodiment, the target species is Streptococcus pneumoniae.

[0073] Example 2

[0074] This embodiment provides a method for designing target species-specific primers based on information entropy. The method includes the following steps:

[0075] Obtain the genomic sequences of the target species and its related species;

[0076] Process the reference genomic sequence of the target species to obtain short sequence fragment data with identity and coverage greater than or equal to 80%;

[0077] Process the short sequence fragment data with identity and coverage greater than or equal to 80%, screen the short sequence fragment data covering 90% of the genome, obtain the entropy value of each base in the short sequence fragments covering 90% of the genome, quantitatively score the short sequence fragments covering 90% of the genome according to the entropy value and coverage information, screen the short sequence fragments covering 90% of the genome with higher scores, and replace the bases with the most conserved bases to obtain optimized short sequence fragment data;

[0078] Design primers based on the optimized short sequence fragment data, and ensure that the 3'-end of the primers avoids high entropy value sites to obtain candidate primers;

[0079] Conduct quality assessment on the candidate primers to obtain candidate primers that meet the assessment conditions.

[0080] In this embodiment, the target species is Streptococcus pneumoniae.

[0081] Example 3

[0082] This embodiment provides a method for designing target species-specific primers based on information entropy. The method includes the following steps:

[0083] Obtain the genomic sequences of the target species and its related species.

[0084] On the one hand, obtain all the genomic sequences of the target species (Streptococcus pneumoniae) from the NCBI GenBank or RefSeq public genomic databases, and obtain the reference genome number of Streptococcus pneumoniae from the NCBI Genome database; if there are more than 1000, 5000, or 10000 genomic sequences of Streptococcus pneumoniae, redundancy removal is required, and the dRep software is used to remove highly similar genomes with an ANI value greater than 0.995. On the other hand, obtain the reference genomic sequences of 5 - 10 closely related species of Streptococcus pneumoniae from the GenBank or RefSeq public genomic databases.

[0085] Split and align the closely related species.

[0086] Select the reference genomic sequence of Streptococcus pneumoniae and split it into fragments of length n, for example, split it into short sequence fragments of length 50, 100, 200, or 500 bp; use an alignment tool (such as bowtie2) to align these short sequence fragments to the reference genomic sequences of its closely related species, and only retain the short sequence fragments that are not aligned to its closely related species.

[0087] Screen the specific fragments of Streptococcus pneumoniae.

[0088] Use alignment tools such as blastn to align the above - retained short sequence fragments to all the redundancy - removed genomic sequences (other genomic sequences) of Streptococcus pneumoniae, and screen out the fragments that meet the following criteria:

[0089] Identity (identity) ≥ 80%, 90%, or 95%; Coverage (coverage) ≥ 80%, 90%, or 95%.

[0090] Perform coverage filtering and multiple sequence alignment.

[0091] According to the screening results of the previous step, calculate the coverage of the fragments in all the genomic sequences of Streptococcus pneumoniae, and screen out the fragments with a genomic coverage greater than or equal to 90%; collect all the aligned sequences of these fragments in the Streptococcus pneumoniae genome, and use tools such as Muscle to perform multiple sequence alignment on the above - screened target fragment sequences and all the aligned sequences in other genomes of the target species to generate position - aligned alignment sequences and obtain an alignment matrix.

[0092] Calculate the entropy values of the sites and fragments.

[0093] According to the multiple sequence alignment results, calculate the entropy value of each base site in the alignment matrix to measure the uncertainty of the base at that position; a high entropy value indicates a large variability of the base at that site (SNP site), and the risk of primer mismatch at this site is relatively high. The entropy of a single site is defined as follows:

[0094]

[0095] Among them, is the frequency of occurrence of bases A, T, C, G and the deletion signal, H is the entropy of a single site, and i is a natural number from 1 to 5.

[0096] Fragment quality score.

[0097] Define the entropy value of a fragment as the sum of the entropy values of all its sites, subtract the coverage of the fragment from the entropy value of the fragment to obtain the quality score of the fragment; sort the fragments based on the quality score, and retain the optimized fragments with higher scores as candidate fragments for primer design.

[0098] Optimize the base sequence of the candidate fragments for primer design.

[0099] Since certain sites in the reference genome may not necessarily represent the sequences of most genomes of Streptococcus pneumoniae, it is necessary to optimize the bases of the candidate fragments: according to the multiple sequence alignment results, replace the base at each site with the base with the highest occurrence probability (the most conservative base); mark the sites with entropy values higher than 0.5 to avoid these high-entropy sites when designing primers.

[0100] Primer design.

[0101] Use Primer3 software to design primers based on the optimized candidate fragment sequences, set the parameter PRIMER_LOWERCASE_MASKING = 1 to ensure that the primers avoid high-entropy sites, and further ensure that the 3'-ends of the primers avoid high-entropy sites. Design 100 - 10,000 pairs of primers for each optimized candidate fragment to obtain candidate primers.

[0102] Primer quality assessment.

[0103] Conduct a comprehensive quality assessment on the designed candidate primers, including the following aspects: primer dimer analysis, primer conservation and genome coverage assessment, primer off-target analysis, and species specificity assessment of amplification products. Only retain the candidate primers that meet the assessment conditions.

[0104] Experimental verification.

[0105] Conduct experimental verification on the candidate primers that pass the assessment to finally obtain reliable target species-specific primers.

[0106] In addition, the embodiments of the present invention also provide a device for designing target species-specific primers based on information entropy, and the device includes the system as described in Embodiment 1 of the present invention.

[0107] Furthermore, an embodiment of the present invention also provides a readable storage medium, in which multiple computer programs are stored, and when the computer programs run on a processor, they execute the method described in Embodiment 2 of the present invention.

[0108] It should be understood that the present invention disclosed herein is not limited to the particular methods, schemes, and materials described, as these may vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the present invention, which is limited only by the appended claims.

[0109] Those skilled in the art will also recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. These equivalents are also encompassed by the appended claims.

Claims

1. A system for designing target species-specific primers based on information entropy, characterized in that, The system includes the following modules: A data acquisition module, which is used to obtain the genomic sequences of the target species and its related species; The data acquisition module includes the following units: A first data acquisition unit, which is used to process all genomic sequences of the target species to obtain the reference genomic sequence of the target species and other genomic sequences of the target species; A second data acquisition unit, which is used to obtain the reference genomic sequence of the related species of the target species; A data screening module, which is used to process the reference genomic sequence of the target species to obtain short sequence fragment data with identity and coverage greater than or equal to 80%; The data screening module includes the following units: A splitting unit, which is used to split the reference genomic sequence of the target species into short sequence fragments; A first alignment unit, which is used to align the short sequence fragments to the reference genomic sequence of the related species of the target species, and screen the unaligned short sequence fragments as candidate fragments; A second alignment unit, which is used to align the candidate fragments to other genomic sequences of the target species, and screen the aligned short sequence fragments as target fragments; A calculation unit, which is used to calculate the identity and coverage of the target fragments in other genomic sequences of the target species, and screen the target fragments with identity and coverage greater than or equal to 80% as optimized fragments; A data optimization module, which is used to process the short sequence fragment data with identity and coverage greater than or equal to 80%, screen out the short sequence fragment data covering 90% of the genome, obtain the entropy value of each base in the screened short sequence fragments, perform quantitative scoring on the short sequence fragments covering 90% of the genome according to the entropy value and coverage information, screen the short sequence fragments covering 90% of the genome with the top scores, and replace the bases with the most conservative bases to obtain the optimized short sequence fragment data; The data optimization module includes the following units: A first calculation unit, which is used to count the coverage of the optimized fragments in all genomic sequences of the target species, screen out the short sequence fragments covering 90% of the genome, and obtain all aligned sequences; An alignment unit, which is used to perform multiple sequence alignment on the target fragment sequences screened by the first calculation unit and all aligned sequences in other genomes of the target species to generate a position-aligned alignment sequence and obtain an alignment matrix; A second calculation unit, which is used to calculate the entropy value of each base site in the alignment matrix to obtain the entropy value of the optimized fragment; A scoring unit, which is used to perform quantitative scoring on the optimized fragments according to the entropy value of the optimized fragments and the coverage information of the optimized fragments, and screen the optimized fragments with the top scores; Specifically: define the entropy value of the optimized fragment as the sum of the entropy values of all its sites, subtract the coverage of the optimized fragment from the entropy value of the optimized fragment to obtain the quality score of the optimized fragment; sort the optimized fragments based on the quality score, and retain the optimized fragments with the top scores as candidate fragments for designing primers; An optimization unit, which is used to optimize the bases in the optimized fragments with the top scores, replace the bases with the most conservative bases in the multiple sequence alignment, and mark the base sites with high entropy values, and the most conservative base is the base with the highest occurrence probability in the multiple sequence alignment; The primer design module is used to design primers according to the optimized short sequence fragment data, mark the sites with entropy values higher than 0.5, and ensure that the 3'-ends of the primers avoid the high-entropy sites with entropy values higher than 0.5 to obtain candidate primers. The primer quality assessment module is used to evaluate the quality of the candidate primers to obtain candidate primers that meet the assessment conditions.

2. The system according to claim 1, wherein The quality assessment includes one or more of primer dimer analysis, primer conservation and genome coverage assessment, primer off-target analysis, and species specificity assessment of amplification products.

3. The system according to claim 1, wherein The target species is a pathogenic microorganism species, and the pathogenic microorganism species is a bacterium, a fungus, or a virus.

4. A method for designing target species-specific primers based on information entropy, characterized in that, The method includes the following steps: Obtain the genomic sequences of the target species and its related species; specifically: It is used to process all the genomic sequences of the target species to obtain the reference genomic sequence of the target species and other genomic sequences of the target species. It is used to obtain the reference genomic sequence of the related species of the target species. Process the reference genomic sequence of the target species to obtain short sequence fragment data with identity and coverage greater than or equal to 80%; specifically: It is used to split the reference genomic sequence of the target species into short sequence fragments. It is used to align the short sequence fragments to the reference genomic sequence of the related species of the target species, and screen the unaligned short sequence fragments as candidate fragments. It is used to align the candidate fragments to other genomic sequences of the target species, and screen the aligned short sequence fragments as target fragments. It is used to calculate the identity and coverage of the target fragments in other genomic sequences of the target species, and screen the target fragments with identity and coverage greater than or equal to 80% as optimized fragments. Process the short sequence fragment data with identity and coverage greater than or equal to 80%, screen the short sequence fragment data covering 90% of the genome, obtain the entropy value of each base in the short sequence fragments covering 90% of the genome, quantitatively score the short sequence fragments covering 90% of the genome according to the entropy value and coverage information, screen the short sequence fragments covering 90% of the genome with the top scores, and replace the bases with the most conservative bases to obtain the optimized short sequence fragment data; specifically: The first calculation unit statistics the coverage of the optimized fragments in all the genomic sequences of the target species, screens the short sequence fragments covering 90% of the genome, and obtains all the aligned sequences. It is used to perform multiple sequence alignment on the target fragment sequences screened by the first calculation unit and all the aligned sequences in other genomes of the target species to generate aligned alignment sequences and obtain an alignment matrix. It is used to calculate the entropy value of each base site in the alignment matrix to obtain the entropy value of the optimized fragment. It is used to quantitatively score the optimized fragments according to the entropy value of the optimized fragments and the coverage information of the optimized fragments, and screen the optimized fragments with the top scores; specifically: define the entropy value of the optimized fragment as the sum of the entropy values of all its sites, subtract the coverage of the optimized fragment from the entropy value of the optimized fragment to obtain the quality score of the optimized fragment; sort the optimized fragments based on the quality score, and retain the optimized fragments with the top scores as candidate fragments for primer design. Optimize the bases in the optimized fragments with top scores, replace the bases with the most conservative bases in the multiple sequence alignment, and mark the base sites with high entropy values. The most conservative base is the base with the highest occurrence probability in the multiple sequence alignment; Design primers based on the optimized short sequence fragment data, mark the sites with entropy values higher than 0.5 as restricted, and ensure that the 3'-end of the primer avoids the high entropy value sites with entropy values higher than 0.5 to obtain candidate primers; Conduct quality assessment on the candidate primers to obtain candidate primers that meet the assessment conditions.

5. The method according to claim 4, wherein The target species is a pathogenic microorganism species, and the pathogenic microorganism species is bacteria, fungi or viruses.

6. An apparatus for designing target species-specific primers based on information entropy, characterized in that, The device includes the system according to any one of claims 1-3.

7. A readable storage medium, characterized in that, Multiple computer programs are stored in the readable storage medium, and when the computer programs run on the processor, they execute the method according to claim 4.

Citation Information

Patent Citations

  • Method and device for obtaining species-specific consensus sequences of microorganisms and application

    CN111477276A

  • Multi-sequence conservative interval detection method, degenerate primer design method, related device and electronic equipment

    CN117116347A

  • Fungus species identification method based on analysis of whole-genome, and target nucleotide, primer pair, kit and use thereof

    WO2024208276A1