Multi-target pathogen detection primer group design method, equipment, medium and program product

The design method of multi-target pathogen detection primer group optimization through three-layer verification and greedy algorithms solves the limitations of target region screening and primer optimization in the existing technology, and achieves efficient and accurate multi-target pathogen detection.

CN120126567APending Publication Date: 2025-06-10BEIJING CAPITALBIO MEDLAB CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510203961.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing multi-target detection technology has limitations in target area screening strategies and primer optimization methods, which affects the sensitivity and accuracy of the detection results, making it difficult to meet the needs of high-throughput and multi-pathogenic accurate detection.

Method used

Through the synergy of three-layer verification, a multi-target pathogen detection primer set is designed, including single-remain primer design, two-layer specificity verification and greedy algorithm optimization, ensuring the specificity and efficiency of the primer set.

Benefits of technology

It realizes the full process optimization from primer design to sequencing applications, improves the specificity and sensitivity of the detection results, reduces false positive results, shortens the time cost of the design process, and provides reliable support for high-throughput pathogen detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126567A_ABST
    Figure CN120126567A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent medical treatment, and particularly relates to a multi-target pathogen detection primer group design method, equipment, a medium and a program product. The method comprises the following steps: acquiring a target gene bank of various pathogens; performing single primer design on the basis of the target gene bank to obtain a single primer set, extracting amplicons of primer pairs in the primer set, and screening to obtain a first layer of primer pair set after specific verification consisting of primer pairs corresponding to the amplicons conforming to species specificity; extracting a sequencing product of the primer pair subjected to the first-layer specific verification, and screening to obtain a second-layer primer pair set subjected to the specific verification, wherein the second-layer primer pair set consists of primer pairs corresponding to the sequencing product conforming to the species specificity of the pathogen; and screening the primer pair set subjected to the second-layer specific verification to obtain an optimized primer group, wherein Panel specificity is taken as a screening condition in the screening process. According to the invention, through the synergistic effect of three-layer verification, better optimization of the design of the multi-target multiple primer is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medicine, and more particularly, to a method, device, medium and program product for designing a primer set for multi-target pathogen detection. Background Art

[0002] Lower respiratory tract infection (LRTI) is a common and serious infectious disease in clinical practice. Its pathogenic pathogens are diverse, including bacteria, viruses, fungi, and parasites, etc. Effectively and accurately identifying pathogens is crucial for guiding treatment and improving the prognosis of patients.

[0003] Currently, the research on multi-target detection mainly focuses on the following technologies:

[0004] qPCR (Real-time Fluorescent Quantitative PCR): Commonly used for detecting single or a small number of targets, but its throughput is low and it is difficult to meet the diagnostic needs of complex infections.

[0005] Multiplex PCR: The detection throughput is improved by simultaneously amplifying multiple targets, but the amplification efficiency may be reduced due to interference between primers.

[0006] Metagenomic sequencing (mNGS): It can achieve high-throughput and multi-pathogen detection, but its high cost and complex data analysis limit its popularization in clinical practice.

[0007] Multiplex PCR targeted sequencing technology (also known as amplicon targeted sequencing technology) combines multiplex PCR technology with next-generation sequencing (NGS) technology, becoming an efficient and rapid pathogen detection method. By specifically amplifying or probe capturing specific fragments of target pathogens, the targeted region is enriched and sequenced, which is widely used in the detection of various pathogens such as bacteria, fungi, viruses, and parasites; among them, targeted next-generation sequencing technology (tNGS) is based on multiplex PCR amplification and combines sequencing analysis. The technical barrier lies in the design and evaluation of ultra-multiplex target primers.

[0008] Currently, the multi-target detection technology based on multiplex PCR combines some characteristics of metagenomic sequencing (mNGS) and multiplex PCR targeted amplicon sequencing technology (tNGS) for the detection of various pathogens, but there are still certain limitations.

[0009] 1 Characteristics and limitations of metagenomic next-generation sequencing (mNGS): As an unbiased detection technology, metagenomic next-generation sequencing can achieve comprehensive detection of all genomic sequences in a sample, and its advantages lie in high throughput and broad spectrum. However, there are the following problems in the practical application of metagenomic next-generation sequencing: 1) Insufficient targeting: mNGS adopts a random amplification strategy without specific targets. Although it can cover all possible pathogens, it also generates a large amount of host background noise, reducing the detection sensitivity; 2) Analysis complexity: mNGS requires a large amount of bioinformatics analysis to screen out pathogenic-related sequences from massive data, with a long analysis time and high professional requirements for analysts; 3) High cost: Due to the involvement of high-throughput sequencers and complex data processing, the cost of mNGS is relatively high, making it difficult to be popularized in routine clinical tests; 4) Limited quantification ability: mNGS usually cannot provide accurate pathogen load information, limiting its application in the assessment of infection severity.

[0010] 2 Characteristics and technical barriers of targeted amplicon sequencing by multiplex PCR (tNGS): The targeted sequencing technology by multiplex PCR combines multiplex PCR with next-generation sequencing technology through primer design, making up for the lack of targeting of mNGS. This technology targets specific fragments of clinically relevant pathogens, enriches the target region through PCR amplification or probe capture, and conducts high-throughput sequencing. Its advantages are reflected in: 1) Clear target: Based on the designed primer pool, it can accurately capture the specific region of the target pathogen, improving the detection specificity and sensitivity; 2) High throughput and quantification: It can achieve simultaneous detection of multiple pathogens and provide relatively accurate quantitative information, which is helpful for clinical disease diagnosis and severity assessment; 3) Fast and efficient: It has a lower cost and a simpler operation process compared with metagenomic sequencing, and is suitable for rapid clinical diagnosis.

[0011] However, the multiplex PCR targeted sequencing technology still faces the following technical barriers: (1) Limitations of target region screening methods: Current target region screening is mostly based on the K-mer fragmentation strategy, which extracts pathogen genomic data from public databases to screen species-specific regions. However, this method has limitations when dealing with complex pathogen communities or highly variable genomes. When the K-mer method fails, it will lead to incomplete screening results and missing target regions, and other methods (such as machine learning or other bioinformatics analyses) need to be combined for integrated screening to improve the screening efficiency and comprehensiveness. (2) Limitations of primer set optimization methods: The core of primer design is to design a multi-target primer set through an optimization algorithm. Existing technologies usually use the greedy algorithm to gradually combine primers, but this strategy is prone to falling into local optimal solutions, resulting in poor performance of the final generated primer pool. Problems such as dimer effects and non-specific amplification between primers may be ignored, thus affecting the sensitivity and accuracy of detection results. Improving the optimization algorithm (such as improving the consideration of dimer effects and non-specific amplification in the algorithm process, and trying heuristic algorithms or artificial intelligence-based optimization models) is a feasible solution to improve the quality of primer design; (3) Insufficient specific verification of primer pairs; In summary, although the existing multi-target detection technology based on multiplex PCR has certain advantages in pathogen detection, its limitations in target region screening strategies and primer optimization methods limit its further application in high-throughput and multi-pathogen precise detection. These problems provide an important research direction and technical innovation space for the present invention. Summary of the Invention

[0012] In view of the above problems, the present invention provides a method for designing a primer set for multi-target pathogen detection of lower respiratory tract infections, which screens target pathogens of various respiratory tract infection pathogens and realizes the detection of multi-target pathogens in clinical samples.

[0013] The present application (in the first aspect) discloses a method for designing a primer set for multi-target pathogen detection, including:

[0014] S1: Obtain a target gene library of various pathogens;

[0015] S2: Based on the target gene library, perform single-primer design to obtain a single-primer set, and based on the single-primer set, obtain a first primer pair set.

[0016] S3: Based on the primer pairs in the first primer pair set, obtain amplicons, and screen to obtain a first-layer specifically verified primer pair set composed of primer pairs corresponding to the amplicons that meet the species specificity of the pathogens.

[0017] S4: Obtain sequencing products based on the primer pairs in the first-layer specifically verified primer pair set, and screen to obtain a second-layer specifically verified primer pair set composed of primer pairs corresponding to the sequencing products that meet the species specificity of the pathogen;

[0018] S5: Use an algorithm to screen the second-layer specifically verified primer pair set to obtain an optimized primer group, and use the Panel specificity as the screening condition during the screening process.

[0019] Furthermore, use the greedy algorithm constrained by Panel specificity verification to screen to obtain an optimized primer group. During the screening process of the greedy algorithm constrained by Panel specificity verification, when looking for the next local optimum primer based on the local optimum primer pair combinations in multiple steps of the greedy algorithm, discard the primer pairs that lead to non-Panel specificity;

[0020] Optionally, the discrimination method for non-Panel specific pairing includes: a primer pair combination includes N primer pairs, and the forward primer of the i-th primer pair is denoted as Pf i , and the reverse primer is denoted as Pr i . If any forward primer Pf i forms a pair with the reverse primer Pr j of a non-owned primer pair, it is non-Panel specific pairing, where i≠j;

[0021] Optionally, if any forward primer of the current local optimum primer pair combination of the greedy algorithm forms a pair with the reverse primer of the next primer pair to be optimized, it is non-Panel specific pairing;

[0022] Optionally, the forward primer of the i-th primer pair in the current local optimum primer pair combination is denoted as Pf i , and the reverse primer is denoted as Pr i . The value range of i is from 1 to the number of primer pairs in the current local optimum primer pair combination. If the forward primer Pf k of the local optimum solution in the k-th step currently being optimized forms a pair with the reverse primer Pr j of a non-owned primer pair, or the reverse primer Pr k forms a pair with the forward primer Pf i of a non-owned primer pair, it is non-Panel specific pairing.

[0023] 1. Furthermore, the method for screening to obtain primer pairs corresponding to amplicons that meet the species specificity of the pathogen is: obtain amplicons based on the forward primer and reverse primer of the primer pair, perform sequence alignment of the amplicon sequence with the reference genome, and obtain the result of whether the amplicon meets species specificity according to the sequence alignment result.

[0024] 2. Further, the method for screening primer pairs corresponding to sequencing products that conform to the species specificity of the pathogen is as follows: obtaining the corresponding sequencing products based on the forward primer and the reverse primer of the primer pair, aligning the sequencing products with the reference genome, and obtaining whether the sequencing products conform to the species specificity result according to the sequence alignment result;

[0025] No result that conforms to species specificity;

[0026] Optionally, the alignment uses the blastn tool;

[0027] Optionally, the reference genome includes any one or more of the following: host genome, common pathogen reference genome, common laboratory contaminant microbial reference genome, reference genome of closely related species of the detection index, reference genome of the detection Panel, plasmid reference genome.

[0028] 3. Further, a greedy algorithm with Panel-specific verification constraints is used to screen an optimized primer set. The optimization objective of the heuristic algorithm is to minimize the likelihood loss value of primer dimers in the optimized primer pair set, and at the same time, the primer pair combinations that conform to Panel specificity are preferred during the optimization process.

[0029] 4. Further, the first primer pair set contains N primer pairs, and the optimized primer pair set contains M

[0030] primer pairs, where N > M. The optimization objective of the heuristic algorithm is to minimize the likelihood loss of primer sequences of M primer pairs selected from N primer pairs for M target pathogens to form dimers

[0031] value;

[0032] Optionally, the likelihood loss value is expressed as:

[0033]

[0034] where Loss represents the likelihood loss value, k ∈ [1, M], M represents the number of primer pairs in the optimized primer pair set, represents the dimer penalty of the k-th primer pair;

[0035] Optionally, the way to obtain the dimer penalty of the k-th primer pair is: slicing the primer sequence of the k-th primer pair to obtain sequence fragments, traversing the primer sequence of the k-th primer pair using a sliding window, and calculating the dimer penalty of the matching fragment according to the matching fragment that matches the sequence fragment obtained by traversing;

[0036] Optionally, the dimer penalty of the matching fragment of the k-th primer pair is expressed as:

[0037]

[0038] where \(i\in[1,N match \), \(N match is the number of matching segments that match the sequence segment obtained by traversal, \(\alpha\) represents the penalty coefficient, and score seq represents the dimer penalty of the matching segment, and Num seqGc represents the GC frequency of the sequence segment, and distance to3′ represents the distance from the sequence segment to the 3′ end;

[0039] Optionally, the dimer penalty of the matching segment of the \(k\)th primer pair is expressed as:

[0040]

[0041] where End 4 , End 5 and End 6 represent the penalty coefficients of sequence segments with lengths of 4bp, 5bp, and 6bp respectively, and Md 4 , Md 5 and Md 6 represent the maximum distances from sequence segments with lengths of 4bp, 5bp, and 6bp to the 3′ end respectively, and GCcount4, GCcount5, and GCcount6 represent the number of GC bases in sequence segments with lengths of 4bp, 5bp, and 6bp respectively;

[0042] Optionally, the heuristic algorithm includes one or more of the following: simulated annealing algorithm, genetic algorithm, particle swarm algorithm, dynamic programming algorithm.

[0043] Furthermore, the target gene library is obtained by collecting pathogen genomic sequences related to lower respiratory tract infections from a public database, and the target gene library includes different pathogens and corresponding targets;

[0044] Optionally, the types of the pathogens include one or more of the following: viruses, bacteria, fungi, and parasites;

[0045] Optionally, the method further includes: performing conservative screening on the target gene library to obtain a conservative fragment set, and performing single primer design based on the conservative fragment set to obtain a single primer set;

[0046] Optionally, the conservative fragment set includes an in-species conservative region of CDS and an in-species conservative region of K-ner;

[0047] Optionally, the step of performing conservative screening on the target gene library includes: extracting CDS sequences for the target regions of bacteria based on a CFF annotation file, and obtaining the CDS intraspecies conservative regions through multiple sequence alignment of the CDS sequences;

[0048] Optionally, the step of performing conservative screening on the target gene library further includes: screening the target regions of viruses, fungi, and parasites based on sequence K-mer fragmentation, counting the coverage of the fragment regions within different genomic strains of the same species, and obtaining the K-mer intraspecies conservative regions according to the coverage;

[0049] Optionally, the method further includes: filtering the single primer set to obtain a filtered single primer set, and obtaining the first primer pair set based on the filtered single primer set; optionally, the filtering includes any one or more of the following: primer3 penalty threshold filtering, 3'-end sequence composition filtering, simple repeat base filtering, simple repeat sequence filtering, 3'-end GC Clamp filtering, and inverted repeat sequence filtering;

[0050] Optionally, Primer3 is used for the single primer design.

[0051] The second aspect of the present application discloses a multi-target pathogen detection primer set design system, including:

[0052] An acquisition module 201: used to acquire the target gene libraries of multiple pathogens;

[0053] A single primer design module 202: used to perform single primer design based on the target gene library to obtain a single primer set, and obtain a first primer pair set based on the single primer set,

[0054] A first-layer specificity verification and screening module 203: used to extract the amplicons of the primer pairs in the first primer pair set, and screen to obtain a first-layer specifically verified primer pair set composed of the primer pairs corresponding to the amplicons that meet the species specificity of the pathogen;

[0055] A second-layer specificity verification and screening module 204: used to extract the sequencing products of the primer pairs in the first-layer specifically verified primer pair set, and screen to obtain a second-layer specifically verified primer pair set composed of the primer pairs corresponding to the sequencing products that meet the species specificity of the pathogen;

[0056] A third-layer specificity verification and screening module 205: used to screen the second-layer specifically verified primer pair set to obtain an optimized primer set, and the screening process uses the Panel specificity as the screening condition.

[0057] The third aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, it is used to execute the steps of the above method.

[0058] The fourth aspect of the present application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above method.

[0059] The fifth aspect of the present application discloses a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the above method.

[0060] The present application has the following beneficial effects:

[0061] (1) Through the synergistic effect of three-layer verification: by combining a three-level specific verification system, the whole process from primer design to sequencing application is optimized:

[0062] (2) Based on three-layer verification and performing Panel-specific verification during the greedy algorithm process, compared with a single greedy algorithm, we can achieve obvious gains and are less likely to fall into the trap of local optimal solutions;

[0063] (3) The present application introduces an algorithm into multiple primer design and optimizes key parameters, the influencing parameters for calculating Dimer penalties, and the distance to the 3' end;

[0064] (4) The present application realizes the specific verification of the primer set inside and outside the Panel and in related species. The criteria for selecting related species references are: in the species classification database, select the nodes 2 to 3 distances upward from the detection index as the related species range of the detection index. The reference genome composition of related species includes, but is not limited to, the representative genomes within the related species range;

[0065] (5) The present application can and has realized a primer set for efficient multi-target detection of pathogens in lower respiratory tract infections. Description of the Drawings

[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0067] Figure 1 It is a schematic flowchart of the method provided in the first aspect of the embodiments of the present invention;

[0068] Figure 2 It is a schematic diagram of a program product provided in the second aspect of the embodiment of the present invention;

[0069] Figure 3 It is a schematic diagram of a computer device provided in the embodiment of the present invention;

[0070] Figure 4 It is a schematic diagram of the architecture of an exemplary computing device provided in the embodiment of the present invention;

[0071] Figure 5 It is a schematic diagram of a storage medium provided in the embodiment of the present invention;

[0072] Figure 6 It is a schematic diagram of a method for verifying the non-specificity of a Panel provided in the embodiment of the present invention;

[0073] Figure 7 It is an ER diagram of a species relationship database provided in the embodiment of the present invention. Detailed implementation manners

[0074] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0075] In some processes described in the specification, claims and above-mentioned drawings of the present invention, a plurality of operations that appear in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0076] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0077] Figure 1 It is a schematic diagram of a method for designing a primer set for detecting multiple target pathogens based on evaluation provided in the embodiment of the present invention. Specifically, the method includes the following steps:

[0078] S101: Obtain the target gene libraries of multiple pathogens;

[0079] S102: Design single primers based on the target gene libraries to obtain a set of single primers, and obtain a first set of primer pairs based on the set of single primers.

[0080] S103: Obtain amplicons based on the primer pairs in the first set of primer pairs, and screen to obtain a first-layer specifically verified set of primer pairs composed of the primer pairs corresponding to the amplicons that meet the species specificity of the pathogens.

[0081] S104: Obtain sequencing products based on the primer pairs in the first-layer specifically verified set of primer pairs, and screen to obtain a second-layer specifically verified set of primer pairs composed of the primer pairs corresponding to the sequencing products that meet the species specificity of the pathogens.

[0082] S105: Screen the second-layer specifically verified set of primer pairs to obtain an optimized primer group, and use the Panel specificity as the screening condition during the screening process.

[0083] First, we screened the target regions through a dual strategy of characteristic gene CDS screening and sequence fragmentation, and adopted separate or integrated strategies for different species. In the selection of the two strategies, we preferentially used the strategy of characteristic gene screening combined with multiple sequence alignment in the screening process of bacterial target regions, and preferentially used the sequence fragmentation strategy in the screening process of viral, fungal, and parasitic target regions. For some indicators, we used both strategies for design and processing. Compared with simply using the sequence fragmentation strategy, the strategy based on characteristic gene screening can effectively increase the number of potential target regions. The reasons are as follows: In actual design, the number of genomes is often large. When using the sequence fragmentation strategy, under the premise of given limited computing resources, there is a certain probability that potential conserved regions will not be processed at a certain cutting length and cutting step. However, based on characteristic genes, this potential drawback can be overcome on a relatively large-scale conserved sequence. Through the mutual assistance of the two methods, we can achieve rapid design and screening of target regions covering multiple pathogen categories such as bacteria, viruses, fungi, and parasites, as the data source for downstream primer design.

[0084] We applied the heuristic simulated annealing algorithm to the combinatorial optimization of the multi-target primer pool. This combination needs to overcome many challenges. The addition of multiple pairs of primer pairs often leads to an exponential increase in primer dimers. There are few reports in actual Panel design. We obtained better screening effects through an optimization function based on dimer penalty scores and evaluating the specificity of the candidate primer set. At the same time, we integrated the use of the greedy algorithm. Compared with a single greedy algorithm, we can achieve obvious gains and are less likely to fall into the trap of local optimal solutions.

[0085] At present, there are still deficiencies in the control of specificity in multiple PCR primer design schemes, and there is a lack of systematic solutions. If only the specificity of primer pairs is concerned, it may lead to non-specific amplification in multiple PCR, resulting in a decrease in the efficiency of the reaction system and the generation of unexpected detection results, ultimately affecting the reliability of the detection results. To address this issue, this patent constructs a multi-round specificity verification method system, which can effectively improve the accuracy of primer design and the detection performance of the detection panel.

[0086] The present invention provides a method for designing primers for multi-target pathogen detection of lower respiratory tract infections and its application, mainly including the following technical steps:

[0087] The first step: Construction of a pathogen database

[0088] Collect the genomic sequences of common and rare pathogens related to lower respiratory tract infections from public databases (NCBI Refseq or Genbank) to construct a target gene database (a reference library for the scope of the test panel, a reference library for related species, and a common pathogen database).

[0089] In some embodiments, the pathogens collected include, but are not limited to, the following pathogens:

[0090] Acinetobacter baumannii, Klebsiella aerogenes, Klebsiella oxytoca, Escherichia coli, Haemophilus ducreyi, Klebsiella pneumoniae, Enterococcus faecalis, Staphylococcus aureus, Morganella morganii, Staphylococcus haemolyticus, Enterococcus faecium, Stenotrophomonas maltophilia, Streptococcus agalactiae, Staphylococcus epidermidis, Mycobacterium tuberculosis complex, Mycobacterium kansasii, Mycobacterium avium, Mycobacterium abscessus, Pseudomonas aeruginosa,

[0091] Human herpesvirus 1, Human herpesvirus 2, Human polyomavirus 1, Human polyomavirus 2, Human herpesvirus 5, Human herpesvirus 4 (EB), Human herpesvirus 3,

[0092] Candida glabrata, Aspergillus fumigatus, Candida albicans

[0093] In some embodiments, the data volume of the collected target gene database is: 267 index pathogens;

[0094] The second step: Screening and optimization of targets (regions):

[0095] The efficient selection of genomic target regions is crucial for improving the specificity of primer design. We screen the target regions based on the principles of interspecies specificity and intraspecies conservation:

[0096] Conservative screening: Screen for conserved signature gene fragments of specific pathogens. The dual strategy includes Strategy 1 and Strategy 2. Strategy 1: Extract CDS regions based on genomic annotation file information; extract CDS sequences based on the annotated GFF annotation file, and then obtain the intra-species conserved regions through multiple sequence alignment (MAFFT).

[0097] Strategy 2: Screen for conserved regions based on the K-mer sequence fragmentation strategy. Based on sequence K-mer fragmentation, the coverage of its fragment regions within different genomic strains of the same species is statistically analyzed as its intra-species conserved fragment set. The conserved fragment sets (the length range of target region amplicons: 100 - 2000) screened by the above two strategies are used for downstream analysis.

[0098] In the selection of the two strategies, Strategy 1 is preferentially used in the screening process of bacterial target regions, and Strategy 2 is preferentially used in the screening process of viral, fungal, and parasite target regions; by integrating the two strategies for design and processing, we can achieve rapid design and screening of target regions covering multiple pathogen categories such as bacteria, viruses, fungi, and parasites, serving as the data source for downstream primer design.

[0099] Furthermore, based on the specificity of the target region, the target regions screened by integrating Strategy 1 and Strategy 2 are further screened: Use a specificity-verified reference library, including the human hg19 + common pathogens + species within the detection Panel range + plasmids + related species database. Before primer design, use bioinformatics alignment tools to analyze the specificity of the conserved sequences in the target region and exclude and filter out potential non-specific regions.

[0100] Step 3: Obtain candidate primers:

[0101] 3.1 Single primer design

[0102] Primer3 is currently the most widely used and industry-recognized gold standard for PCR primer design software. Use Primer3 for single primer design, using custom primer design parameters such as amplicon size, primer length, GC content, Tm value, and secondary structure. Primer length range: 20 - 22, TM range: 58 - 62, optimal 60, GC range: 40 - 54, amplicon length range: 170 - 220

[0103] 3.2 Primer filtering

[0104] To obtain primer sequences of better quality, the primer pairs obtained from primer3 are also filtered according to the set filtering rules, including but not limited to: primer3 penalty threshold filtering, 3'-end sequence composition filtering, simple repeat base filtering, simple repeat sequence filtering, 3'-end GC Clamp filtering, and inverted repeat sequence filtering.

[0105] 3.3 Primer pair specificity verification (first - layer specificity verification)

[0106] After completing the primer custom filtering, we used a method based on sequence alignment and species - non - specific amplicon discovery to verify the specificity of primer pairs.

[0107] The specific steps are as follows:

[0108] (1) Primer sequence alignment

[0109] Use the btools tool to align the primer sequences to multiple reference genomes, including:

[0110] Host genome, reference genomes of common pathogens, reference genomes of common laboratory contaminant microorganisms, reference genomes of closely related species of detection indicators, reference genomes of detection Panels, plasmid reference genomes

[0111] (2) Collection of alignment results and primer - pair coverage detection

[0112] Collect all alignment results to identify all amplicons that can be formed between the forward primer and the reverse primer.

[0113] (3) Species - specificity judgment

[0114] Rely on a self - built species classification relationship database to judge the species - specificity of amplicons.

[0115] (4) Screening of primer pairs

[0116] If primer pairs with species non - specificity are found, discard them to ensure the accuracy and effectiveness of subsequent experiments.

[0117] Functions and benefits of specificity verification:

[0118] This round of specificity verification aims to reduce the non - specificity of amplicons covered by primer - pair combinations and ensure the specificity of single - primer and single - Agent: By verifying the specificity of primers, it is ensured that amplicons are only targeted at the target species, thereby reducing false - positive results. On the other hand, this round of verification will reduce the total number of single primers in the downstream primer pool to a certain extent and shorten the overall design running time cost.

[0119] (2) Since the species - specificity of our amplicon recognition, and the sequencing products of forward and reverse primers all have species non - specificity and need to be compared with the self - built species database, in order to clarify the disclosure of this application, this application discloses the ER diagram of the species relationship database, Figure 7 as shown, and describes the method steps for judging species - specificity based on the species relationship database;

[0120] The species classification database is used in the first target region sequence specificity, the second primer pair specificity, and the third Panel specificity, and the usage methods are similar and the content is the same. After aligning the primer sequence as the query sequence to the specific reference sequence database, the specificity is judged through the following process:

[0121] Ⅰ. Acquisition of query sequence species information: Obtain the species TaxID of the reference sequence from the sequence number and TaxID relationship cache model and the sequence number and TaxID relationship model of the species classification database according to the accession number of the query sequence. Obtain the complete species classification information of the query sequence from the species classification relationship model of the species classification database according to the species TaxID of the query sequence.

[0122] Ⅱ. Annotation of reference sequence species information: Obtain the species TaxID of the reference sequence from the sequence number and TaxID relationship cache model and the sequence number and TaxID relationship model of the species classification database according to the number of the reference sequence.

[0123] Ⅲ. Ambiguous species annotation: Obtain the ambiguous species annotation information from the ambiguous species annotation model of the species classification database according to the TaxIDs of the query sequence and the reference sequence in 1 and 2.

[0124] Ⅳ. Specificity judgment logic: 1) If the TaxIDs of the query sequence and the reference sequence are the same, the result is accurately specific. 2) If the TaxID of the query sequence belongs to the TaxID of the reference sequence, the result is ambiguous and is default specific. 3) If the TaxID of the reference sequence belongs to the TaxID of the query sequence, the result is accurately specific. 4) If the TaxID of the query sequence or the TaxID of the reference sequence involves ambiguous species information, find the nearest common ancestor TaxID in the classification relationship between the TaxID of the query sequence and the TaxID of the reference sequence (using the LCA algorithm: instead of processing the entire species internal classification tree, start with the TaxIDs of the query sequence and the reference sequence and traverse the parent nodes upward to find the nearest common ancestor TaxID), and then calculate the distance from the ambiguous parent TaxID of the ambiguous species TaxID to the nearest common ancestor TaxID. If the distance is less than or equal to 1, the result is ambiguous and is default specific; if the distance is greater than 1, discuss the distance from the other TaxID to the nearest common ancestor TaxID. When it is equal to 0, the result is accurately specific. When it is equal to 1, the result is ambiguous and is default specific. When it is greater than 1, the result is accurately non-specific.

[0125] 3.4 Specificity verification of sequencing products (second-layer specificity verification)

[0126] While using single-end sequencing technology of 75bp (SE75) or single-end 50bp (SE50) to reduce the detection cost, to ensure the specificity of detection, it is necessary to ensure that the sequencing products corresponding to the forward primer and the reverse primer have a certain specificity. The specific methods are as follows:

[0127] (1) Sequencing product extraction

[0128] Extract the sequencing products corresponding to the forward primer and the reverse primer (such as 50bp).

[0129] (2) Alignment analysis

[0130] Use the blastn tool to align the extracted sequencing products to multiple reference genomes, including:

[0131] Host genome

[0132] Reference genomes of common pathogens

[0133] Reference genomes of common laboratory contaminating microorganisms

[0134] Reference genomes of closely related species of detection indicators

[0135] Reference genome of detection Panel

[0136] Plasmid reference genome

[0137] (3) Collect alignment results

[0138] Collect all alignment results to evaluate the specificity of the sequencing products, including: retaining the alignment results using the blastn metablast mode with default alignment similarity, coverage, and evalue thresholds.

[0139] (4) Species specificity judgment

[0140] Rely on the self-built species classification relationship database to judge the species specificity of the sequencing products.

[0141] (5) Primer pair screening

[0142] If the sequencing products of both the forward primer and the reverse primer show species non-specificity, discard this primer pair to ensure the accuracy and reliability of subsequent experiments.

[0143] Functions and benefits of specificity verification:

[0144] This round of specificity verification aims to ensure the specificity of the sequencing products of SE75 and SE50 in the actual sequencing process. By controlling the specificity of the single-end sequencing products (forward or reverse), it is ensured that any short sequence of 50bp obtained by single-end sequencing is species-specific, thereby reducing false positive results. On the other hand, this round of verification will reduce the total number of single primers in the downstream primer pool to a certain extent and shorten the overall design running time cost.

[0145] Step 4: Optimization of the multi-target primer set:

[0146] The complexity of ultra-multiplex PCR. Compared with single PCR, the design complexity of ultra-multiplex PCR has increased significantly. In N-fold PCR, the number of primer dimer combinations can reach C(2N,2). Therefore, there are at least M 2N combinations for the primers of M targets. The optimal solution to this problem cannot be obtained by enumeration. In view of this, the multi-primer design (mpd) problem is considered a high-dimensional problem of a highly non-convex fitness landscape in mathematics, and standard convex optimization algorithms are not applicable.

[0147] First, we hope to define the problem to be solved by the algorithm:

[0148] As mentioned above, we have screened N pairs of primers. These N pairs of primers contain the corresponding primers on multiple targets of M detected pathogens. Now, we hope to screen N to obtain K pairs of primers, and the K pairs of primers meet the following conditions:

[0149] ① The K pairs of primers should be able to detect all M pathogens, N >= K >= M

[0150] ② And the possibility of the K pairs of primers forming dimers should be small.

[0151] ③ And there should be low non-specific amplification between the K pairs of primers, and they should be able to detect all M pathogens well.

[0152] For the optimization of this technical problem, we further used different algorithms and corresponding strategies in two different embodiments to solve the problem we defined. Among them, for the condition that the possibility of the K pairs of primers forming dimers in ② is small, in the implementation of the greedy algorithm, we screened primer combinations with a small possibility of forming dimers through a series of restrictive conditions; in the implementation of the heuristic algorithm, we designed a loss function based on dimer penalty scores and made this loss function as small as possible to control the possibility of forming dimers between the K-fold primers.

[0153] Method 1: Adopt the greedy algorithm

[0154] The greedy algorithm optimizes the primer set, mainly by evaluating the possibility of forming dimers with 4 - 8 bp at the 3' end. The method we use to calculate the loss value is based on the complementary pairing situation between primers, paying special attention to the number of complementary pairing sequences, the number of bases from the 3' end, and the GC content in the complementary pairing sequences. Research reports have shown that complementary pairing within 3 bases does not significantly affect the efficiency of the PCR reaction. Therefore, we finally choose 4 - 8 bp as the threshold.

[0155] In the greedy algorithm, the specificity of the primer set is used as the key condition for primer pair screening, and primer pairs that do not meet the conditions will be discarded.

[0156] Pre - processing: Align all candidate primer sequences to the reference genome, collect all alignment results, and annotate species information for the alignment results.

[0157] Sort the alignment results according to the species information, and construct a fast - access index with the primer sequences as keys.

[0158] Calculation: During the process of constructing the primer set, the alignment results of the primer sequences to be selected are mixed with the existing primer set.

[0159] The solution steps of the greedy algorithm:

[0160] First, we initialize an empty primer set table; then use the greedy algorithm to screen and obtain the final K pairs of primers according to the following method steps:

[0161] 1. Sort the data of the forward primer sequences and reverse primer sequences described above according to the average penalty score, so that primer pairs with a smaller average penalty score have a higher priority. Here, the penalty score is the score of the primer sequence, such as the penalty score of Primer3 during primer design;

[0162] 2. Prepare the greedy algorithm loop, sort according to the number of target regions of the pathogen index, so that the pathogen loop with a smaller number of target regions has a higher priority;

[0163] That is, the greedy algorithm loop is a loop with a loop count of N, and the loop traversal of pathogens with a smaller number of target regions is on the outer side; in the traversal of this step, the number of loop iterations of each nested loop is the number of target regions corresponding to the current layer;

[0164] 3. Since there are multiple targets in the target regions traversed by the loop, during the loop process, only one primer pair is selected from the multiple targets in the target regions of each pathogen index;

[0165] 4. In order to control the number K of primers in our final primer group, limit each pathogen index to only select a certain number of target regions;

[0166] In one embodiment, we have 9 target pathogens, and the limited number here is 3; then there are 27 primer combinations screened in this step. Among them, some do not meet the specificity within the Panel. Therefore, after removing the primers that do not meet the Panel specificity, the final number of primer sets is less than 27, which is 10.

[0167] 5. Limit the frequency of the 6bp sequences at the 3' ends of the forward and reverse primer sequences of the currently selected primer pair binding to the existing primer set at the 3' end not to exceed a certain value;

[0168] 6. Limit the frequency of the 5bp sequences at the 3' ends of the forward and reverse primer sequences of the currently selected primer pair binding to the existing primer set at the 3' end not to exceed a certain value;

[0169] 7. Limit the frequency of the 4bp sequences at the 3' ends of the forward and reverse primer sequences of the currently selected primer pair binding to the existing primer set at the 3' end not to exceed a certain value;

[0170] 8. Limit that no non-Panel-specific amplification is formed between any primer sequence of the currently selected primer pair and any primer sequence of the existing primer set; for the comparison result combination of the forward and reverse primer sequences of the currently selected primer and the existing primers, if Panel non-specificity is found, discard the current primer pair to ensure that the selected primer pairs have high specificity. As Figure 6 shown, Pf1 and Pr1 are the first pair of primer pairs designed, and Pfi and Prj are any other forward or reverse primers. If non-specific pairing such as Pf1 and Prj or Pfi and Pr1 in the figure is formed, delete the current primer pair to be added, and iterate in this way until all primer pair sets that meet the requirements are found.

[0171] The better effects obtained by the limitation of the constraint conditions are: (1) quickly obtaining a local optimal solution from a large amount of data; (2) taking care of the pathogen detection indicators with less data to ensure that rare results are not lost; (3) the results and processes are repeatable.

[0172] Method 2: Introduction of primer set optimization by heuristic algorithm

[0173] In the primer optimization of the heuristic algorithm, specificity is a secondary optimization problem, and the optimization process mainly considers the optimal primer dimer penalty.

[0174] At this time, we can set the value of K to be equal to M corresponding to the number of target pathogen species to obtain the fewest primer set combinations; or set K to M + 1; M + 2; to make the number of final primer sets as small as possible.

[0175] When K = M, that is, when selecting one primer pair for each target pathogen, the size of the screening space can be expressed as:

[0176]

[0177] Among them, C1, C2, …, CM respectively represent the number of primers corresponding to the 1st, 2nd, …, Mth target pathogen species included in the N-fold primers, and C1 + C2 + … + CM = N.

[0178] It can be seen that the larger N is and the larger M is, the more complex this technical problem will be. For this screening problem, we use a heuristic algorithm to select the K primer pair combinations that minimize the objective function from the screening space.

[0179] The specific steps are as follows:

[0180] Pretreatment: Align all candidate primer sequences to the reference genome, collect all alignment results to discover all amplicons that can be formed between the forward primer and the reverse primer.

[0181] Identify and record all Panel-nonspecific amplicons.

[0182] The loss function to be optimized is defined as the likelihood function for calculating the formation of primer dimers by primers. The penalty points include the loss of heterologous primer dimers and the loss value of homologous primer dimers formed by the same primer. The likelihood function is defined as the regression function of the primers p1 - pn in the primer pool. Among them, the base distance distance parameter to the 3'-end is an important factor affecting the penalty function. In actual tests, the 4 - 8 bp of the distance to the 3'-end has the greatest impact.

[0183] For the current primer set, calculate the measure value of Panel non-specificity. The secondary specificity consideration is added as a constraint condition to the fitness function. The non-specific amplicons identified from the candidate primer data of all indicators are used as the data basis. It is required that the number of non-specific amplicons in the primer set generated by the heuristic algorithm is lower than a certain value, otherwise the primer set is determined to be a primer set with unqualified specificity.

[0184] If this measure value exceeds the Panel non-specificity tolerance threshold set by the program, a higher fitness penalty score is given to reduce the priority of this primer set. Iterate until the result converges or the preset conditions are met and stop.

[0185] The likelihood function is defined as the regression function of the K primer pairs in the primer pool. Among them, the base distance di parameter to the 3'-end is an important factor affecting the penalty function. In actual tests, the 4 - 8 bp of the distance to the 3'-end has the greatest impact.

[0186] Design of the loss function:

[0187] The optimized objective function of the heuristic algorithm is expressed as:

[0188]

[0189] N match represents the number of matching sequence fragments in the primer set, that is, the number of penalty fragments (short sequences) found in the resulting primer set, score seq represents the penalty of the i-th matching sequence, E represents the penalty coefficient, and the coefficients for matching sequences of different lengths are different, GC count represents the GC frequency of the matching sequence, distance to3′ represents the distance from the matching sequence to the 3'-end of the primer sequence.

[0190] It can also be expressed as:

[0191]

[0192] Score k represents the fitness function of the k-th primer pair among the K primer pairs in the optimized primer set. This fitness function represents the dimer penalty of the matching fragments of the k-th primer pair and measures the possibility of the k-th primer pair forming primer dimers in the current primer set;

[0193] The objective function measures the sum of the fitness functions of all K primer pairs. The optimization objective of the heuristic algorithm is to minimize the objective function of the primer set S; traverse the primer set composed of K primers, take sequence fragments with lengths of 4bp, 5bp, and 6bp at the 3'-end, and count the frequency Num of the sequence fragments seq , and the Score of each sequence fragment seq is equal to Num seq

[0194] Among them, the steps to obtain the fitness function of the k-th primer pair are as follows:

[0195] Use a sliding window to traverse the primer sequence of the k-th primer pair, find matching fragments that are consistent with any sequence fragment (where the sequence fragments are sequence fragments of different lengths, such as 4bp, 5bp, and 6bp length sequence fragments), and calculate the score of the i-th matching fragment according to the following formula:

[0196]

[0197] Score match(i) represents the dimer penalty of the i-th matching fragment obtained by sliding window matching, Score seq represents the dimer penalty of the sequence fragment corresponding to the matching fragment; E is a constant representing the penalty coefficient, and the coefficients for different length fragments are different; GC count represents the GC frequency of the matching fragment; distanceto3' Indicates the distance from the matching segment to the 3'-end;

[0198] Finally, the sliding window traverses to obtain n match matching segments. Then, the dimer penalty of the k-th primer pair's matching segment is expressed as:

[0199]

[0200] Score match(i) Indicates the dimer penalty of the i-th matching segment obtained by sliding window matching, up to n match Indicates the number of matching segments in the k-th primer pair;

[0201] The optimization goal is to find the primer set with the lowest Loss in the optimized primer set S;

[0202] Among them, S represents the optimized primer set containing K primer pairs; argmin represents finding the parameter that minimizes the dimer likelihood function;

[0203] In some other embodiments, the following method is used to calculate the primer set penalty calculation formula of the simulated annealing algorithm:

[0204] Step 1, define the penalty coefficients for sequence segments with lengths of 4bp, 5bp, and 6bp as End 4 , End 5 and End 6 .

[0205] Step 2, represent the maximum distance from sequence segments with lengths of 4bp, 5bp, and 6bp to the 3'-end as Md 4 , Md 5 and Md 6 .

[0206] Step 3, represent the number of GC bases in sequence segments with lengths of 4bp, 5bp, and 6bp as GCcount.

[0207] Step 4, represent the primer set as PS, where the number of primers in the primer set is M, and the penalty for the 4bp-length at the 3'-end of the primer set is DS 4 , the penalty for the 5bp-length is DS 5 , and the penalty for the 6bp-length is DS 6 .

[0208] Step 5, count the frequencies Freq 4 , Freq 5 and Freq 6 .

[0209] Step 6, calculate the penalty score for the 4-bp length sequence fragment at the 3'-end of the primer sequences in the primer set:

[0210]

[0211] Md4 represents the maximum distance from the 4-bp length sequence fragment used for calculating the penalty score on the primer sequence to the 3'-end of the primer sequence. If the distance exceeds this value, the calculation of the penalty score on this sequence will stop;

[0212] Step 7, calculate the penalty score for the 4-bp length sequence fragment at the 3'-end of the primer sequences in the primer set:

[0213]

[0214] Step 8, calculate the penalty score for the 4-bp length sequence fragment at the 3'-end of the primer sequences in the primer set:

[0215]

[0216] Step 9, total dimer penalty score of the primer set:

[0217] DS ps = DS 4 + DS 5 + DS 6

[0218] After transformation, the total penalty score, which is also the primer dimer likelihood loss value function, is expressed as:

[0219] DS ps = DS 4 + DS 5 + DS 6

[0220]

[0221] That is to say, the penalty / moderateness value of the k-th primer pair is equal to the sum of its 4 / 5 / 6-bp penalty scores in the current primer set, expressed as:

[0222]

[0223] That is to say, the moderateness value of the k-th primer is calculated from the sequence fragments with lengths of 4bp, 5bp, and 6bp;

[0224] Through multiple iterations, we can find a primer set close to the global optimum and finally select the multi-target super-multiplex primer group.

[0225] In some embodiments, we take the simulated annealing algorithm in heuristic algorithms as an example. Although the simulated annealing algorithm has achieved success in global optimization problems in multiple fields, effectively applying it to the combinatorial optimization of a multi-target primer pool requires overcoming many challenges.

[0226] The addition of multiple primer pairs often leads to an exponential increase in primer dimers. We evaluate the specificity of the candidate primer set through a fitness function based on dimer penalties.

[0227] Specifically, during the optimization process, a detailed analysis of the dimerization risk that each pair of primers may form is required, and the fitness value is calculated accordingly. This method requires us to accurately identify the potential dimerization risk during the primer screening stage and take effective measures to control it, thereby ensuring the specificity of the primer set and the reliability of the design.

[0228] The objective function to be optimized is designed as an evaluation of the primer dimer likelihood loss value function. And the calculation of this objective function is the most resource-consuming and rate-limiting step in the simulated annealing algorithm.

[0229] Its influencing factors are based on such assumptions and experiences: if the number of complementary pairing sequences between two primers is more, and they are closer to the 3' end and have a higher GC content, then the loss value of this primer pair is higher, and the possibility of forming primer dimers is greater. Conversely, the lower the loss value, the smaller the possibility of forming primer dimers.

[0230] Finally, a primer set is obtained through the above steps, where the primer set contains K kinds of primers.

[0231] Function: The purpose of using closely related species for specificity verification is to replace the extremely large NT / NR database, reduce the computational cost, and improve the design efficiency; through this round of specificity verification, the possibility of non-specific amplification in the set of multiplex primer pairs within the Panel is greatly reduced.

[0232] Synergistic effect of three-layer verification: Through the combination of a three-level specificity verification system, the entire process from primer design to sequencing application is optimized, with the following beneficial effects:

[0233] (1) Comprehensively ensure specificity: The first layer of verification reduces the risk of non-specific amplification in the preliminary primer design; the second layer of verification ensures the species specificity of the actual sequencing results through the comparison and screening of sequencing products; the third layer of verification ensures the overall efficiency of the primer set from the perspectives of global and local optimization.

[0234] (2) Reduce false positive results: The three-layer verification complements each other, ensuring the accuracy of the amplification products and sequencing results, and minimizing false positive results to the greatest extent.

[0235] (3) Reduced the time cost of the design process: At the same time, through the hierarchical optimization strategy, the number of redundant primers and low-quality primers was reduced, significantly shortening the time cost of the design process.

[0236] (4) Provided reliable support: In response to the complexity of the design for a large number of targets, the three-layer verification system effectively improved the optimization effect of the primer set, providing reliable support for high-throughput pathogen detection.

[0237] Generally speaking, this verification system not only improved the application efficiency of the multiplex PCR targeted sequencing technology, but also significantly improved the reliability and repeatability of the experimental results.

[0238] By adopting the method described in this solution, a primer set applicable to multi-target respiratory pathogens was obtained, and the effectiveness of the primer set screened by the method of this solution was verified through experiments.

[0239] In some embodiments, the greedy algorithm and the SA algorithm screened out the same primer set, and we constructed a multiplex PCR system based on the screened primer set for verification.

[0240] 4 Construction of the multiplex PCR system:

[0241] The multi-target amplification system was optimized through experiments, including parameters such as reaction conditions and primer concentrations.

[0242] Working concentration of primers: 0.5 uM / reaction

[0243] Storage concentration of primers: 0.125 uM / reaction

[0244] The amplification system and reaction conditions are shown in the following table

[0245]

[0246]

[0247] 5 Verification of the detection system:

[0248] The sensitivity, specificity and repeatability of the detection system were evaluated using standard samples and clinical samples.

[0249] Application and promotion:

[0250] The optimized detection system was applied to the enterprise standard products mixed, and the detection limit results of the detection system are as follows.

[0251]

[0252]

[0253] The samples composed of the following pathogen indicators are used as specific analysis samples for testing, and the results are shown in the following table:

[0254]

[0255] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the present invention. As Figure 3 shown, the device may include: one or more processors, and one or more memories; wherein, computer-readable code is stored in the memory, and when the computer-readable code is run by the one or more processors, the above-described method can be executed.

[0256] The processor in this embodiment may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.

[0257] Generally speaking, the various example embodiments of the present disclosure may be implemented in hardware or special circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor or other computing devices. When the aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, special circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0258] For example, the method or device according to the embodiments of the present disclosure may also be implemented by means of Figure 4 the architecture of the computing device 3000 shown. As Figure 4 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course,Figure 4 The architecture shown is only exemplary. When implementing different devices, one or more components in the computing device shown may be omitted according to actual needs. Figure 4

[0259] An embodiment of the present invention also provides a computer-readable storage medium. As Figure 5 shown, it is a schematic diagram of the storage medium provided by the embodiment of the present invention. Computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the methods according to the embodiments of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories of the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories. It should be noted that the memories of the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0260] An embodiment of the present disclosure also provides a computer program product or a computer program. When the computer program is executed by a processor, the steps of the above method are implemented. As Figure 2 shown, the computer program product or the computer program includes:

[0261] An acquisition module 201: configured to acquire a target gene library of multiple pathogens;

[0262] A single primer design module 202: configured to perform single primer design based on the target gene library to obtain a set of single primers, and obtain a first set of primer pairs based on the set of single primers.

[0263] ​The first-layer specificity verification and screening module 203: It is used to extract the amplicons of the primer pairs in the first primer pair set, and screen out the primer pairs corresponding to the amplicons that meet the species specificity of the pathogen to form the first-layer specifically verified primer pair set.

[0264] The second-layer specificity verification and screening module 204: It is used to extract the sequencing products of the primer pairs in the first-layer specifically verified primer pair set, and screen out the primer pairs corresponding to the sequencing products that meet the species specificity of the pathogen to form the second-layer specifically verified primer pair set.

[0265] The third-layer specificity verification and screening module 205: It is used to screen out the optimized primer group from the second-layer specifically verified primer pair set, and the screening process takes the Panel specificity as the screening condition.

[0266] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0267] Generally speaking, the various example embodiments of the present disclosure can be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0268] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0269] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.

[0270] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0271] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0272] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for designing a multi-target pathogen detection primer set, characterized in that: The method comprises: S1: Obtain target gene libraries of multiple pathogens; S2: Designing a single primer based on the target gene library to obtain a single primer set, and obtaining a first primer pair set based on the single primer set, S3: extracting amplicons of primer pairs in the first primer pair set, and screening to obtain a first-layer specifically verified primer pair set consisting of primer pairs corresponding to amplicons that meet the species specificity of the pathogen; S4: extracting the sequencing products of the primer pairs in the first-layer specifically verified primer pair set, and screening to obtain the second-layer specifically verified primer pair set consisting of primer pairs corresponding to the sequencing products that meet the species specificity of the pathogen; S5: The second-layer specifically verified primer pair set is screened to obtain an optimized primer set, and the panel specificity is used as the screening condition in the screening process.

2. The method for designing a multi-target pathogen detection primer set according to claim 1, characterized in that: The optimized primer set is obtained by screening using a greedy algorithm with panel-specific verification constraints. During the screening process of the greedy algorithm with panel-specific verification constraints, the primer pair that results in non-panel specificity is discarded when searching for the next local optimal primer based on the combination of local optimal primer pairs in multiple steps of the greedy algorithm. Optionally, the method for distinguishing non-Panel specific pairing includes: a primer pair combination includes N primer pairs, and the forward primer of the i-th primer pair is represented by Pf i , the backward primer is denoted as Pr i , if any forward primer Pf i The backward primer Pr j The pairing formed by the comparison result is a non-Panel specific pairing, where i≠j; Optionally, if any forward primer of the current local optimal primer pair combination of the greedy algorithm forms a pair with the backward primer of the next primer pair to be optimized, it is a non-Panel specific pairing; Optionally, the forward primer of the i-th primer pair in the current local optimal primer pair combination is denoted as Pf i , the backward primer is denoted as Pr i , i ranges from 1 to the number of primer pairs in the current local optimal primer pair combination. If the forward primer Pf of the local optimal solution of the kth step currently being optimized is k The backward primer Pr j The comparison results form a pair, or the backward primer Pr k The front primer Pf of the non-primer pair i The pairing formed by the alignment result is a non-panel specific pairing.

3. The method for designing a multi-target pathogen detection primer set according to claim 1, characterized in that: The method for screening a primer pair corresponding to an amplicon that meets the species specificity of the pathogen is as follows: obtaining an amplicon based on a forward primer and a backward primer of the primer pair, comparing the sequence of the amplicon with a reference genome, and obtaining a result of whether the amplicon meets the species specificity based on the sequence comparison result.

4. The method for designing a multi-target pathogen detection primer set according to claim 1, characterized in that: The method for screening the primer pair corresponding to the sequencing product that meets the species specificity of the pathogen is as follows: obtaining the corresponding sequencing product based on the forward primer and the backward primer of the primer pair, performing sequence alignment on the sequencing product and the reference genome, and obtaining a result of whether the sequencing product meets the species specificity according to the sequence alignment result; Optionally, the alignment uses a blastn tool; Optionally, the reference genome includes any one or more of the following: host genome, common pathogen reference genome, common laboratory contamination microorganism reference genome, related species reference genome of detection indicators, reference genome of detection panel, plasmid reference genome.

5. The method for designing a multi-target pathogen detection primer set according to claim 1, characterized in that: The optimized primer set is screened using a greedy algorithm with panel-specific verification constraints. The optimization goal of the heuristic algorithm is to minimize the primer dimer likelihood loss value of the optimized primer pair set, while giving priority to the primer pair combination that meets the panel specificity during the optimization process.

6. The method for designing a multi-target pathogen detection primer set according to claim 5, characterized in that: The first primer pair set includes N primer pairs, the optimized primer pair set includes M primer pairs, wherein N>M, and the optimization goal of the heuristic algorithm is to minimize the likelihood loss value of dimers formed by primer sequences of the M primer pairs screened from the N primer pairs for the M target pathogens; Optionally, the likelihood loss value is expressed as: Among them, Loss represents the likelihood loss value, k∈[1,K], K represents the number of primer pairs in the optimized primer pair set, N≥K≥M, score k represents the fitness function of the kth primer pair; Optionally, the dimer penalty of the kth primer pair is obtained by slicing the primer sequence of the kth primer pair to obtain a sequence fragment, using a sliding window to traverse the primer sequence of the kth primer pair, and calculating the dimer penalty of the matching fragment according to the matching fragment obtained by the traversal that matches the sequence fragment; Optionally, the fitness function of the kth primer pair is expressed as: Among them, i∈[1,n match ],n match is the number of matching fragments traversed on the kth pair of primers, score match(i) represents the dimer penalty of the i-th matching fragment; Optionally, the dimer penalty for the i-th matching fragment is expressed as: Among them, score seq Indicates the dimer penalty of the sequence fragment corresponding to the matching fragment. E is a constant, indicating the penalty coefficient. The coefficients of fragments of different lengths are different. count Indicates the GC frequency of the matching fragment; distance to3' Indicates the distance from the matching fragment to the 3' end; Optionally, the fitness function of the kth primer pair is expressed as: Among them, End4, End5 and End6 represent the penalty coefficients of sequence fragments of 4bp, 5bp and 6bp in length, respectively; Md4, Md5 and Md6 represent the maximum distance from the sequence fragments of 4bp, 5bp and 6bp in length to the 3' end, respectively; GCcount4, GCcount5 and GCcount6 ​​represent the number of GC bases in the sequence fragments of 4bp, 5bp and 6bp in length, respectively; Optionally, the heuristic algorithm includes one or more of the following: simulated annealing algorithm, genetic algorithm, particle swarm algorithm, dynamic programming algorithm.

7. The method for designing a multi-target pathogen detection primer set according to claim 1, characterized in that: Collecting pathogen genome sequences related to lower respiratory tract infection from public databases to obtain the target gene library, wherein the target gene library includes different pathogens and targets corresponding to different pathogens; Optionally, the types of pathogens include one or more of the following: viruses, bacteria, fungi and parasites; Optionally, the method further comprises: performing conservative screening on the target gene library to obtain a conservative fragment collection, and performing single primer design based on the conservative fragment collection to obtain a single primer set; Optionally, the conservative fragment collection includes CDS intra-species conservative regions and K-ner intra-species conservative regions; Optionally, the step of performing conservative screening on the target gene library includes: extracting CDS sequences from the target region of bacteria based on the CFF annotation file, and obtaining CDS intraspecies conserved regions by multiple sequence alignment of the CDS sequences; Optionally, the step of conservative screening of the target gene library further includes: screening the target regions of viruses, fungi, and parasites based on sequence K-mer fragmentation, counting the coverage of the fragment regions in different genome strains of the same species, and obtaining the K-mer intraspecies conservative regions according to the coverage; Optionally, the method further comprises: performing primer filtering on the single primer set to obtain a filtered single primer set, and obtaining the first primer pair set based on the filtered single primer set; Optionally, the filtering comprises any one or more of the following: primer3 penalty threshold filtering, 3' end sequence component filtering, simple repeat base filtering, simple repeat sequence filtering, 3' end GC Clamp filtering, and inverted repeat sequence filtering; Optionally, Primer3 is used for the single primer design.

8. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • A sliding window-based method for evaluating primer pool dimerization in targeted PCR.

    CN122575484A