Method and apparatus for predicting whether amplification reaction primer is applicable, and storage medium

By using deep learning models to predict the suitability of amplification reaction primers, this approach solves the problems of complex RPA primer design and insufficient prediction of PCR/RPA amplification reactions in existing technologies, and achieves rapid and accurate primer design guidance.

WO2025245657A1PCT designated stage Publication Date: 2025-12-04BOE TECHNOLOGY GROUP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/095466
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing software cannot effectively predict the suitability of recombinase polymerase amplification (RPA) primers, requiring users to perform multiple screenings and optimizations. Furthermore, existing technologies cannot simultaneously predict the success of PCR and RPA amplification reactions.

Method used

A deep learning model, including word embedding, coding and classification layers, is used to generate feature values ​​by converting target and primer sequences into multiple tokens. The model is then fine-tuned using a pre-trained DNABert model to predict whether primers are suitable for RPA and PCR amplification reactions.

Benefits of technology

It provides accurate judgment and can be used for DNA template and primer design for both PCR and RPA amplification reactions, improving the efficiency and accuracy of design and solving the problems of complexity and multiple screening in primer design in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024095466_04122025_PF_FP_ABST
    Figure CN2024095466_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for predicting whether an amplification reaction primer is applicable, and a storage medium. The method comprises: acquiring a target sequence and a primer sequence, converting the target sequence and the primer sequence into a plurality of tokens, and generating one or more first feature values on the basis of the primer sequence and / or a parameter of an amplification reaction; and inputting the plurality of tokens and the one or more first feature values into a first deep learning model to obtain a prediction result of whether the target sequence and the primer sequence are applicable to an RPA amplification reaction and a PCR amplification reaction at the same time, wherein the first deep learning model comprises a token embedding layer, an encoding layer, and a classification layer, the token embedding layer is configured to generate a total embedding representation vector on the basis of the plurality of tokens, the encoding layer is configured to encode the total embedding representation vector to obtain an encoded representation vector, and a first classification layer is configured to extract nonlinear features from the encoded representation vector and the one or more first feature values, and perform classification prediction on the basis of the extracted nonlinear features.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, and storage media for predicting the suitability of amplification reaction primers Technical Field

[0001] This disclosure relates to, but is not limited to, the field of biotechnology, and in particular to a method, apparatus, and storage medium for predicting the suitability of amplification reaction primers. Background Technology

[0002] Polymerase chain reaction (PCR) amplifies a DNA fragment and is widely used in biology, molecular biology, biochemistry, biomedicine, and medical molecular diagnostics. The basic reagents for a PCR reaction include: a double-stranded DNA template (the DNA fragment to be amplified, at least the sequences at both ends must be known), DNA polymerase, four mononucleotides (ATP, CTP, GTP, and TTP), two DNA primers located at both ends of the template, and a reaction buffer. These reagents are mixed in appropriate proportions and placed in a temperature-controlled cycler adapted to the DNA primers. The DNA to be amplified is then amplified at an exponential rate.

[0003] Recombinase polymerase amplification (RPA) is considered a nucleic acid detection technique that can replace PCR. RPA utilizes a recombinase protein to form a complex with primers. This complex promotes the binding of the primers to homologous target sequences of double-stranded DNA, allowing the polymerase to perform subsequent synthesis. The entire process only requires a reaction time of 20-30 minutes at 37-42°C. Compared to PCR, the entire process does not require high-temperature denaturation and low-temperature annealing steps, is simpler to operate, and does not require expensive equipment. RPA primers are longer than typical PCR primers, usually requiring 30-35 bases, and the amplification product is less than 300 bp.

[0004] The key to RPA analysis lies in the design of amplification primers and probes. Most PCR primers are not suitable for RPA reactions because RPA primers are longer than regular PCR primers, typically requiring 30-38 bases. Primers that are too short will reduce recombination rates, affecting amplification speed and detection sensitivity. Currently, there is no dedicated software available for users to design primers required for RPA reactions; users need to perform multiple screening and optimization processes to achieve the desired experimental results.

[0005] Summary of the Invention

[0006] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0007] This disclosure provides a method for predicting the suitability of amplification reaction primers, including:

[0008] Obtain the target sequence and primer sequence, convert the target sequence and primer sequence into multiple tokens, and generate one or more first feature values ​​based on the primer sequence and / or the parameters of the amplification reaction;

[0009] Multiple tokens and one or more first feature values ​​are input into a first deep learning model to obtain a prediction result on whether the target sequence and primer sequence are simultaneously applicable to RPA amplification and PCR amplification reactions. The first deep learning model includes a word embedding layer, an encoding layer, and a first classification layer. The word embedding layer is configured to generate a total embedding representation vector based on multiple tokens; the encoding layer is configured to encode the total embedding representation vector to obtain an encoded representation vector; the first classification layer is configured to extract nonlinear features from the encoded representation vector and one or more first feature values, and perform classification prediction based on the extracted nonlinear features.

[0010] This disclosure also provides an apparatus for predicting the suitability of amplification reaction primers, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the method for predicting the suitability of amplification reaction primers according to any embodiment of this disclosure based on the instructions stored in the memory.

[0011] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for predicting whether amplification reaction primers are suitable as described in any embodiment of this disclosure.

[0012] This disclosure also provides a program product including instructions that, when executed by a computer, perform a method for predicting the suitability of amplification reaction primers as described in any embodiment of this disclosure.

[0013] This disclosure also provides an apparatus for predicting the suitability of amplification reaction primers, comprising: a data preprocessing module and a prediction module, wherein:

[0014] The data preprocessing module is configured to acquire target sequences and primer sequences, convert the target sequences and primer sequences into multiple tokens, and generate one or more first feature values ​​based on the primer sequences and / or parameters of the amplification reaction.

[0015] The prediction module is configured to input multiple tokens and one or more of the first feature values ​​into a first deep learning model to obtain a prediction result of whether the target sequence and primer sequence are simultaneously applicable to RPA amplification and PCR amplification reactions. The first deep learning model includes a word embedding layer, an encoding layer, and a first classification layer. The word embedding layer is configured to generate a total embedding representation vector based on multiple tokens. The encoding layer is configured to encode the total embedding representation vector to obtain an encoded representation vector. The first classification layer is configured to extract nonlinear features from the encoded representation vector and one or more of the first feature values, and perform classification prediction based on the extracted nonlinear features.

[0016] After reading and understanding the accompanying diagrams and detailed descriptions, other aspects can be understood.

[0017] Overview of the attached figures

[0018] The accompanying drawings are provided to further illustrate the technical solutions of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure. The shapes and sizes of the components in the drawings do not reflect actual proportions and are only intended to illustrate the content of this disclosure.

[0019] Figure 1 is a flowchart illustrating a method for predicting the suitability of amplification reaction primers according to an exemplary embodiment of this disclosure;

[0020] Figure 2A is a schematic diagram of a hairpin-like secondary structure;

[0021] Figure 2B shows an example of a primer dimer;

[0022] Figure 3A is a schematic diagram of the structure of a first deep learning model provided by an exemplary embodiment of the present disclosure;

[0023] Figure 3B is a schematic diagram of the structure of another first deep learning model provided by an exemplary embodiment of this disclosure;

[0024] Figure 4 is a schematic diagram of a device for predicting the suitability of amplification reaction primers according to an exemplary embodiment of the present disclosure;

[0025] Figure 5 is a schematic diagram of another device for predicting the suitability of amplification reaction primers provided by an exemplary embodiment of this disclosure.

[0026] Detailed Explanation

[0027] To make the objectives, technical solutions, and advantages of this disclosure clearer, the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be arbitrarily combined with each other.

[0028] Unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" indicate that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, but do not exclude other elements or objects.

[0029] As shown in Figure 1, this disclosure provides a method for predicting the suitability of amplification reaction primers, including:

[0030] Step 101: Obtain the target sequence and primer sequence, convert the target sequence and primer sequence into multiple tokens, and generate one or more first feature values ​​based on the primer sequence and / or the parameters of the amplification reaction;

[0031] Step 102: Input multiple tokens and one or more first feature values ​​into the first deep learning model to obtain the prediction result of whether the target sequence and primer sequence are simultaneously applicable to RPA amplification reaction and PCR amplification reaction. The first deep learning model includes a word embedding layer, an encoding layer and a first classification layer. The word embedding layer is configured to generate a total embedding representation vector based on multiple tokens; the encoding layer is configured to encode the total embedding representation vector to obtain an encoded representation vector; the first classification layer is configured to extract nonlinear features from the encoded representation vector and one or more first feature values, and perform classification prediction based on the extracted nonlinear features.

[0032] The method for predicting the applicability of primers for amplification reactions provided in this disclosure converts the target sequence and primer sequence into multiple tokens, and generates one or more first feature values ​​based on the primer sequence and / or amplification reaction parameters. The multiple tokens and one or more first feature values ​​are input into a first deep learning model to obtain a prediction result on whether the target sequence and primer sequence are simultaneously applicable to RPA and PCR amplification reactions. This method can provide an accurate judgment on whether a new primer pair-template DNA can be simultaneously applied to RPA and PCR amplification reactions, thereby providing effective guidance for the design engineering of DNA templates and primers that are applicable to both PCR and RPA reactions.

[0033] In this embodiment of the disclosure, amplifying a DNA target requires a pair of primers (F, R), wherein primer F is a forward primer that binds to the target strand; and primer R is a reverse primer that binds to the complementary strand of the target strand.

[0034] In some exemplary embodiments, the method further includes:

[0035] The k-mer method was used to segment the DNA sequence to build a token vocabulary.

[0036] Genetic regulatory codes are highly complex due to their polysemy and long-range semantic relationships, often proving difficult to capture using conventional informatics methods, especially in situations where data is scarce. The k-mer features of a set of DNA / genome sequences are crucial for revealing hidden patterns within those sequences. Here, k is an integer constant ranging from 2 to tens, depending on the specific application requirements.

[0037] In some exemplary implementations, the token input to the first deep learning model is generated from the target sequence and primer sequence using the k-mer method, where k is a natural number between 3 and 6.

[0038] K-mer notation, by linking each deoxynucleotide base to the base following it, contains richer semantic information. Generally, reads of length L can be divided into L-K+1 k-mers. For example, the DNA sequence “ATGGCTA” can be divided into (7-3+1) = 5 3-mer sequences: {ATG, TGG, GGC, GCT, CTA}, or into (7-5+1) = 3 5-mer sequences: {ATGGC, TGGCT, GGCTA}, or into (7-6+1) = 2 6-mer sequences: {ATGGCT, TGGCTA}.

[0039] After converting all sequences to k-mer sequences, regardless of the original sequence length, we only have a finite number of possible combinations of ATCG. For example, if we use 6-mer sequences, we have 4... 6 The different base combinations form our token vocabulary. Our token vocabulary consists of all permutations of k-mer and the following 5 special tokens:

[0040] [CLS] stands for Category Token, used to indicate the beginning and end of a conversation or text;

[0041] [UNK] indicates an unknown token, used to indicate a token that does not exist in the vocabulary;

[0042] [SEP] stands for Separate Token, used to separate two sentences;

[0043] [PAD] indicates that a token is being filled;

[0044] [MASK] stands for Mask Token, used to extend the length of the input sequence or to mark parts of the sequence that do not need to be processed.

[0045] Therefore, if we use 6-mer, our token vocabulary has a total of (4 6 +5) = 4101 words.

[0046] In some exemplary implementations, the word embedding layer generates a total embedding representation vector based on multiple tokens, including:

[0047] Convert multiple tokens into word embedding vectors;

[0048] Generate a position embedding vector based on the position of each token in the sequence;

[0049] Generate a word type vector based on whether each token belongs to the target sequence or the primer sequence;

[0050] The word embedding vector, position embedding vector, and word type vector are added element-wise to obtain the total embedding representation vector.

[0051] In this embodiment of the disclosure, the embedding representation vectors used by the model (including word embedding vectors, position embedding vectors, word type vectors, and total embedding representation vectors) can be represented as (B, S, D). h ), where B represents the number of sequences given to the model at one time (a sequence consists of a target sequence and a pair of tokens corresponding to primer sequences); S represents the maximum sequence length. If the length of a sequence is less than S, we will pad it with padding characters to make it S in length; D h This represents the hidden dimension; each character is converted to a length of D. h The vector representation of . For example, B = 1, S = 512, D h =768, however, this disclosure does not limit it.

[0052] In this embodiment of the disclosure, after converting the target sequence and primer sequence into multiple k-mer sequence type tokens respectively, a special token [CLS] is added at the beginning of the sequence (indicating the start of the entire sequence), a special token [SEP] is added at the end of the sequence (indicating the end of the sequence), and [SEP] (indicating the separation of sequences) is added between the target sequence and primer sequence and between the forward primer sequence and the reverse primer sequence.

[0053] Word embedding vectors are the embedding representation vectors of these tokens. We assign a sequence number to each 6-mer in the token vocabulary, converting the original letters into numbers. Suppose AGC, GCA, CAC, and ACT are the 11th, 14th, 100th, and 24th words in the token vocabulary, respectively. Then, the sequence "AGCACT" becomes a word embedding vector [11, 14, 100, 24] after k-mer tokenization, which can then be received by the neural network.

[0054] The position embedding vector is the embedding representation vector of the position of each token in the sequence. Therefore, the position embedding vector can be generated based on the position of each token in the sequence.

[0055] The word type embedding vector indicates whether the token belongs to the target chain sequence or the primer pair sequence. Therefore, a word type vector is generated based on whether each token belongs to the target sequence or the primer sequence.

[0056] The word embedding vector, position embedding vector, and word type embedding vector are added element-wise to obtain the total embedding representation vector for each sequence. This total embedding representation vector is then input into the coding layer for encoding.

[0057] In some exemplary embodiments, the encoding layer includes multiple stacked Transformer layers, each Transformer layer including multiple attention heads. The outputs of the multiple attention heads are concatenated and input into a linear layer for linear transformation to obtain multi-head attention output. The last matrix output by the last Transformer layer among the multiple Transformer layers is the encoded representation vector.

[0058] In this embodiment of the disclosure, the encoding layer is composed of multiple stacked Transformer layers, which can process long sequences in a short time. For example, the Transformer layer includes 12 layers, each Transformer layer has 12 attention heads, the outputs of the 12 attention heads are concatenated and input into a linear layer to obtain the multi-head attention output, and the last matrix output by the last Transformer layer is the final encoded representation vector.

[0059] In some exemplary implementations, the packet loss rate of the attention head can be 0.1 to avoid overlearning.

[0060] The formula for calculating self-attention is as follows: Q = XW Q K = XW K V = XW V ; MultiHead(Q,K,V)=Concat(head1,…,head h W O ;

[0061] Among them, W Q W K W V W O Both are weight matrices, head i Let be the i-th attention head. The outputs of multiple attention heads are concatenated and then fed into a linear layer to obtain the final output of multi-head attention. Multiple Transformer layers are stacked, and the last matrix output of the last layer is used as the final representation (i.e., the encoded representation vector) of all tokens.

[0062] In some exemplary embodiments, the first feature value generated based on the primer sequence includes at least one of the following:

[0063] Number of consecutive Gs in the first 5 bp of the 5' end, number of Gs or Cs in the first 5 bp of the 3' end, percentage of GC content, number of primer dimer structures, and number of primer hairpin structures.

[0064] Generally, when designing primers, the 5' end should avoid consecutive Gs and is preferably cytosine, which can promote fragment recombination; the 3' end should ideally be G or C to improve amplification performance; the GC content percentage should ideally be between 30-70%; and primers should avoid sequences that easily form secondary structures or hairpin structures to reduce primer dimer formation. Taking the R1 sequence of HPV18P1 as an example, in the sequence TTTTTGCCTTTTTCTGCCCACTATTTAAAG, CTTT and ARG will complementarily pair to form a hairpin-like secondary structure, as shown in Figure 2A. The more types of potential hairpin structures a primer can form, the greater the negative impact on amplification success. Similarly, the F1 sequence of HPV16P1 can form two different dimers, and the potential primer dimer site information is shown in Figure 2B.

[0065] In some exemplary embodiments, the first feature value generated based on the parameters of the amplification reaction includes the target length.

[0066] Generally, the longer the target (i.e. amplification product) is, the slower the amplification speed. Too long a target length may not meet the needs of rapid testing.

[0067] This disclosure enhances the accuracy of the prediction results of a first deep learning model by introducing one or more first feature values.

[0068] In some exemplary embodiments, the method further includes: standardizing one or more first feature values.

[0069] This embodiment of the disclosure can accelerate the convergence speed of the model by preprocessing the first feature values ​​such as target length using the StandardScaler method.

[0070] The calculation formula for the standardization method can be expressed as: x′=x-mean / std, where x represents a first feature value of a sample, mean represents the mean of the first feature value of all samples, and std represents the standard deviation of the first feature value of all samples.

[0071] Currently, some technologies propose predicting the success of PCR amplification for a specific primer set and DNA template based on the relationship between primer sequences and templates. However, this technology has the following limitations:

[0072] 1) It can only predict whether PCR amplification will be successful, but cannot predict whether PCR and RPA amplification will be successful.

[0073] 2) The method of converting DNA sequences into tokens in this technology is relatively complex and time-consuming;

[0074] 3) It requires a large amount of labeled data, which leads to limited model performance and applicability in scenarios where data is scarce; and during training, there are problems such as gradient vanishing and inefficiency.

[0075] In some exemplary embodiments, as shown in Figures 3A and 3B, a first deep learning model includes one or more second sub-deep learning models and two fully connected network layers, wherein each second sub-deep learning model is configured to process multiple tokens to obtain a first nonlinear feature among the multiple tokens; a fully connected network layer is configured to process one or more first feature values ​​to obtain a second nonlinear feature among the one or more first feature values; and another fully connected network layer is configured to perform classification prediction based on the extracted first and second nonlinear features.

[0076] In some exemplary implementations, as shown in Figures 3A and 3B, each second sub-deep learning model includes a word embedding layer, an encoding layer, and a bidirectional LSTM layer, wherein the bidirectional LSTM layer is configured to extract bidirectional long-range dependencies in the encoded representation vector to obtain a latent feature vector.

[0077] In some exemplary embodiments, the structure of the word embedding layer, encoding layer, and bidirectional LSTM layer in each second sub-deep learning model may be the same as the structure of the word embedding layer, encoding layer, and bidirectional LSTM layer in the DNABert model.

[0078] DNABert (DNA Bidirectional Encoder Representations from Transformers) is a model used in the field of natural language processing. It is an extension of the BERT (Bidirectional Encoder Representations from Transformers) model, which uses the bidirectional encoding of DNA sequences to interpret and analyze DNA sequences.

[0079] This disclosure embodiment uses a pre-trained DNABert model and fine-tunes it, which greatly reduces the number of training samples required, solves the problem of limited performance and applicability of large natural language models in data-scarce scenarios, and experiments show that by using the DNABert model, high model performance can be obtained with fewer iterations, thereby saving computational steps and improving processing speed.

[0080] In some exemplary embodiments, as shown in Figures 3A and 3B, the first classification layer may include a bidirectional LSTM layer, a first fully connected network layer, and a second fully connected network layer, wherein:

[0081] The bidirectional LSTM layer is configured to extract bidirectional long-range dependencies in the encoded representation vector to obtain the latent feature vector.

[0082] The first fully connected network layer is configured to extract nonlinear features from one or more first feature values ​​to obtain a first nonlinear feature vector.

[0083] The second fully connected network layer is configured to perform classification prediction based on the hidden feature vector and the first nonlinear feature vector.

[0084] The method for predicting the applicability of amplification reaction primers in this embodiment of the present disclosure can learn the bidirectional long-range dependency between the target sequence and the primer sequence by designing a bidirectional LSTM layer in the first classification layer, thereby avoiding the problems of gradient vanishing and low efficiency.

[0085] In this embodiment of the disclosure, the dimension of the first fully connected network layer can be 32-dimensional or 64-dimensional. By inputting one or more first feature values ​​into the 32-dimensional or 64-dimensional first fully connected network layer for learning, a high-dimensional nonlinear feature vector can be obtained.

[0086] In some exemplary embodiments, as shown in Figures 3A and 3B, the second fully connected network layer includes a first splicing layer, a feature extraction layer, and a classification prediction layer, wherein:

[0087] The first concatenation layer is configured to concatenate the latent feature vector with the first nonlinear feature vector to obtain the concatenated feature vector.

[0088] The feature extraction layer is configured to extract non-linear features from the concatenated feature vector to obtain a second non-linear feature vector.

[0089] The classification prediction layer is configured to perform classification prediction based on a second nonlinear feature vector.

[0090] In some exemplary embodiments, as shown in Figure 3B, when the first deep learning model includes multiple second sub-deep learning models, each of the multiple second sub-deep learning models includes its own word embedding layer, its own encoding layer, and its own bidirectional LSTM layer. The tokens input to the multiple second sub-deep learning models are transformed from the target sequence and the primer sequence according to different methods. The first classification layer also includes a second concatenation layer, which is set between the bidirectional LSTM layer and the second fully connected network layer. It is configured to concatenate the hidden feature vectors output by the multiple sub-deep learning models to obtain a concatenated hidden feature vector.

[0091] In this embodiment of the disclosure, by using multiple second sub-deep learning models, the sequence semantic features of tokens corresponding to various sequence partitions can be supplemented, thereby enriching the representation of features.

[0092] In some exemplary implementations, the tokens input to multiple second sub-deep learning models are transformed from target sequences and primer sequences using different k-mer methods, with different k values ​​corresponding to different k-mer methods.

[0093] For example, suppose there are four second sub-deep learning models, and the tokens input to the four second sub-deep learning models are transformed from target sequences and primer sequences using the 3-mer, 4-mer, 5-mer, and 6-mer methods, respectively. In actual use, this disclosure does not limit the number of second sub-deep learning models or the k-mer method corresponding to the token input to each second sub-deep learning model.

[0094] In some exemplary embodiments, the bidirectional LSTM layer includes a forward LSTM layer, a backward LSTM layer, and a third splicing layer, wherein:

[0095] The forward LSTM layer is configured to perform forward computation on the encoded representation vector to obtain the forward hidden vector;

[0096] The inverse LSTM layer is configured to perform reverse computation on the encoded representation vector to obtain the inverse hidden vector;

[0097] The third concatenation layer is configured to concatenate the forward latent vector with the backward latent vector to obtain the latent feature vector.

[0098] The calculation formula for the forward LSTM layer is as follows:

[0099] The calculation formula for the inverse LSTM layer is as follows:

[0100] Where i is the input gate, f is the forget gate, o is the output gate, and c is the memory unit. W is a temporary memory unit. i W f W c W o U i U f U c U o Both are weight matrices, b i b f b c b o Both are bias vectors, and t represents the t-th time.

[0101] Substitute the encoded representation vector output by the encoding layer into the x value in the calculation formula of the forward LSTM layer. t In the above calculation process, the positive hidden vector is obtained. Similarly, the encoded representation vector output by the encoding layer is substituted into the y-value in the inverse LSTM layer calculation formula. t In this process, the reverse implicit vector is obtained based on the above calculation. Will and By concatenating the two, we can obtain the latent feature vector H, as shown in the following formula:

[0102] In some exemplary embodiments, the method further includes:

[0103] Obtain a pre-trained second sub-deep learning model, which includes a word embedding layer, an encoding layer, and a bidirectional LSTM layer;

[0104] After the bidirectional LSTM layer, a first fully connected network layer and a second fully connected network layer are added to obtain the first deep learning model.

[0105] In this embodiment of the disclosure, the second sub-deep learning model may be derived from the DNABert model.

[0106] In some exemplary embodiments, after obtaining the first deep learning model, the method further includes:

[0107] The parameters of the first deep learning model were fine-tuned using a first dataset, which included multiple pairs of target sequences and primer sequence sets, as well as tag data on whether each pair of target sequences and primer sequence sets was suitable for both RPA and PCR amplification reactions.

[0108] In one example, the first dataset used in this disclosure includes 94 primer pairs. These primer sequences have been tested in biological experiments to determine their suitability for both PCR and RPA reactions to amplify 17 HPV viral DNA targets. Of these, 24 primer pairs are suitable for both PCR and RPA amplification reactions, while 70 primer pairs are not. Table 1 shows example data of some HPV target chains and primer pairs in the first dataset.

[0109] Table 1

[0110] Amplifying a DNA target requires a pair of primers (F, R). Primer F is the forward primer, which binds to the target strand, and primer R is the reverse primer, which binds to the complementary strand of the target strand. As shown in Table 1, HPV18 has multiple target sequences, each corresponding to a primer pair. Tags include 0 and 1; 0 indicates that the DNA target strand-primer pair cannot be used simultaneously for PCR and RPA reactions, while tag 1 indicates that the DNA target strand-primer pair can be used simultaneously for PCR and RPA reactions.

[0111] We segmented both the target sequence and primer sequence using the k-mer method, and all the segmented tokens can be regarded as words in natural language. As shown in Table 2, in this example, the maximum length of a sentence composed of the target sequence and primer pair sequence is 487 tokens, the minimum length of a sentence is 32 tokens, and the average length of a sentence is 101 tokens.

[0112] Table 2

[0113] In some exemplary implementations, fine-tuning the parameters of the first deep learning model using a first dataset includes:

[0114] The first dataset is divided into a first training dataset and a first test dataset. The ratio of positive to negative samples in the first training dataset is close to that in the first test dataset.

[0115] The first deep learning model was trained iteratively multiple times using the first training dataset and the first test dataset. During each iteration, the weights of the encoding layer remained unchanged, while the weights of the first classification layer were adjusted according to a preset loss function.

[0116] In this example, there are 94 pairs of DNA targets and primer sets. These data are randomly divided into a first training dataset and a test set in an 8:2 ratio. The first training dataset includes 84 pairs of DNA targets and primer sets, containing 21 positive examples and 63 negative examples, with a positive-to-negative ratio of 1:3. The test set also includes 10 pairs of DNA targets and primer sets, containing 2 positive examples and 8 negative examples, with a positive-to-negative ratio of 1:4. Table 3 shows the dataset size after random partitioning and an example of input data.

[0117] Table 3

[0118] In this embodiment of the disclosure, when fine-tuning the parameters of the first deep learning model using the first dataset, multiple first feature values ​​affecting primer-template amplification can be generated based on the first dataset. These multiple first feature values ​​include, but are not limited to: amplification product length (i.e., target length), the number of repeating Gs in the first 5 bp of the 5' end, the number of Gs or Cs in the first 5 bp of the 3' end, the GC content percentage, the number of primer dimer structures, and the number of hairpin structures within the primer. Generally, the longer the amplification product (i.e., the target), the slower the amplification speed, which cannot meet the requirements for rapid testing. The number of repeating Gs at the 5' end of the primer should be as small as possible, and the number of Gs or Cs in the first 5 bp of the 3' end should be as large as possible. Excessively high (>70%) or excessively low (<30%) GC content may be detrimental. Base pairing interactions within and between primers may contribute to the generation of primer dimers and hairpins, thereby affecting the amplification reaction. Taking the R1 sequence of HPV18P1 as an example, CTTT and AAAAG in TTTTTGCCTTTTTCTGCCCACTATTTAAAG will pair complementaryly to form a hairpin-like secondary structure, as shown in Figure 2A. The more types of potential hairpin structures a primer can form, the greater the negative impact on amplification success. Similarly, the F1 sequence of HPV16P1 can form two different dimers, and the potential primer dimer site information is shown in Figure 2B.

[0119] The number of consecutive Gs in the first 5 bp of the 5' end, the number of Gs or Cs in the first 5 bp of the 3' end, the percentage of GC content, the number of hairpin structures within the primer, the number of dimers formed by the primer itself, and the number of dimers formed between primers can be calculated from the primer sequence. The target length can be obtained according to the requirements of the amplification reaction settings. Table 4 shows some of the first characteristic values ​​corresponding to some target sequences and primer pair sequences.

[0120] Table 4

[0121] For multiple primary feature values, such as target length, a standardization (StandardScaler) method is used for preprocessing. The standardization calculation formula is as follows: x′=x-mean / std, where x represents the feature value of a sample, mean represents the mean of that feature across all samples, and std represents the standard deviation of that feature across all samples. Using standardization for preprocessing can accelerate the convergence speed of the model.

[0122] In this embodiment of the disclosure, the pre-trained second sub-deep learning model can be derived from the DNABert model. The DNABert model consists of three layers: a word embedding layer, an encoding layer, and a second classification layer. Replacing the second classification layer with the first classification layer yields the first deep learning model of this embodiment. The word embedding layer is configured to generate a total embedding representation vector based on multiple tokens. The total embedding representation vector consists of a word embedding vector (1, 512, 768), a position embedding vector (1, 512, 768), and a word type embedding vector (1, 512, 768). Here, 1 represents one sequence input to the model at a time, 512 represents the maximum sequence length, and 768 represents the size of the hidden dimension. We segment the DNA sequence into k-mers-type tokens and add a special token [CLS] at the beginning of the sequence (indicating the start of the entire sequence) and a special token [SEP] at the end (indicating the end of the sequence). We also add [SEP] (sequence separator) between the target sequence and primer sequences, and between the forward and reverse primer sequences. The word embedding vector is the embedding representation vector of these tokens. The position embedding vector is the embedding representation vector of each token's position in the sequence. The word type embedding vector indicates whether the token belongs to the target strand sequence or the primer pair sequence. We element-wise sum the word embedding vector, position embedding vector, and word type embedding vector to obtain the total embedding representation vector x(1, 512, 768) for each sequence, and input this vector into the coding layer for mapping.

[0123] The encoding layer consists of multiple Transformer layers. For example, the encoding layer can be composed of 12 stacked Transformer layers, each with 12 attention heads. The packet loss rate of each attention head is 0.1 to avoid overlearning. The formula for calculating self-attention is as follows: Q = XW Q K = XW K V = XW V ; MultiHead(Q,K,V)=Concat(head1,…,head h W O ;

[0124] Among them, W Q W K W V W O Both are weight matrices, head i This is the i-th head. The outputs of multiple self-attention heads are concatenated together and then fed into a linear layer to obtain the final output of multi-head attention. Multiple transformer layers are stacked, and the last matrix output by the last transformer layer is used as the final representation (i.e., the encoded representation vector) of all tokens.

[0125] Because our sample size is small but the data are very similar, consisting of 6-mer DNA token sequences, we aim to predict the correlation between targets and primer pairs. Therefore, during training, we keep the weight matrix of the multi-layer transformer unchanged (the same as the pre-training parameters), and only fine-tune the weight parameters of the classification layer. We then extract high-level semantic features between tokens at different distances through the second classification layer for task prediction.

[0126] The first classification layer consists of a bidirectional LSTM layer and two fully connected network layers. For example, two bidirectional LSTM layers can be used, employing dropout to avoid overlearning, resulting in a packet loss rate of 0.1. Since the BERT encoding layer has already captured contextual and semantic information, the final representation of all tokens is input into the bidirectional LSTM layer to capture long-range dependencies. The calculation formula for the forward LSTM layer is as follows:

[0127] The calculation formula for the inverse LSTM layer is as follows:

[0128] Where i is the input gate, f is the forget gate, o is the output gate, and c is the memory unit. W is a temporary memory unit. i W f W c W o U i U f U c U o Both are weight matrices, b i b f b c b o All are bias vectors, where t represents the t-th time step. The weight matrix of the classification layer needs to be learned and updated to adapt the above weight parameters to our dataset.

[0129] Substitute the final output token embedding vector (i.e., the encoded representation vector) of the encoding layer into the calculation formula of the forward LSTM layer. t In the above calculation, the positive hidden vector is obtained. Similarly, the reverse implicit vector can also be obtained. By concatenating the two, we can obtain the latent feature vector H of all tokens, as shown in the following formula:

[0130] Before inputting the hidden feature vector H into the second fully connected network layer, we obtain several preprocessed first feature values ​​{target length, number of repeated Gs in the first 5 bp of the 5' end, number of Gs or Cs in the first 5 bp of the 3' end, GC content percentage, number of primer dimer structures, and number of primer hairpin structures}. We then extract nonlinear features from one or more of these first feature values ​​through the first fully connected network layer to obtain the first nonlinear feature vector Z. Next, we concatenate the first nonlinear feature vector Z and the hidden feature vector H through the second fully connected network layer to obtain a concatenated feature vector. We then extract the nonlinear features from this concatenated feature vector to obtain the second nonlinear feature vector, and perform classification prediction based on this second nonlinear feature vector. The calculation formula is as follows:

[0131] y=sigmoid(W*concat(H,Z)+b).

[0132] During fine-tuning, the cross-entropy loss between the predicted and true values ​​is calculated using the following loss function:

[0133] Where N represents the number of samples, y i It is the true label (0 or 1) of the i-th sample. It is the predicted probability value of the i-th sample (the probability that the model predicts it to be class 1).

[0134] We used AdamW with fixed weight decay as the optimizer for learning and updating the classification layer weight matrix. Other hyperparameters are shown in Table 5. Experimental results show that after 5 iterations, the model achieves good results on the first test dataset, with an accuracy of 90%, an F1 score of 0.67, an MCC coefficient of 0.67, an AUC target of 0.625, a precision of 1, and a recall of 0.5, as shown in Table 6.

[0135] Table 5

[0136] Table 6

[0137] Next, we will compare the Siamese RNN model with the first deep learning model mentioned above to obtain the advantages and disadvantages of the two methods in learning the potential semantic relationship between templates and primer pairs.

[0138] First, through vocabulary mapping, each token is associated with a unique numeric ID, which can be represented as a vector, for example, using one-hot encoding or the more common word embedding representation. For instance, we use word embedding to randomly initialize each token as a continuous vector, where each dimension represents a semantic feature. This representation enables the model to process and understand text / sequences and to acquire more semantic information by learning the relationships between tokens. Therefore, we obtain two matrices: the target word embedding matrix x(512, 768) and the primer pair word embedding matrix y(512, 768), where 512 represents the maximum sentence length and 768 represents the dimension of each token's word embedding vector.

[0139] Then, the relationship between the target and primer pairs is learned through a GRU network (a type of recurrent neural network). The target word embedding matrix x and the primer pair word embedding matrix y are input into two bidirectional dynamic GRU networks, which share weights. The bidirectional dynamic GRU network consists of a forward GRU layer and a backward GRU layer, and the GPU mainly uses the update gate z. t and reset door r t The components, and their calculation formulas are as follows:

[0140] Where, xt Enter h for the current position t-1 This is the hidden state of the previous position. h is a temporary hidden state. t W represents the implicit state at the current position. z W r U z U r U is the weight matrix, b z b r b are bias vectors, and t represents the time step t. Inputting the target word embedding matrix x into the above formula yields the positive latent matrix of x. t is a positive integer greater than 1.

[0141] Similarly, input x into the following formula:

[0142] Obtain the reverse implicit matrix of x t is a positive integer greater than 1. The implicit semantic vector of x is obtained by concatenating the forward and reverse implicit matrices.

[0143] Similarly, the primer pair word embedding matrix y is also input into the bidirectional dynamic GRU network described above to obtain the hidden semantic vector H of y. prim Specifically, the primer-word embedding matrix is ​​input into the formula above, using the same weights, to obtain the hidden matrix.

[0144] Then, calculate the implicit matrix H. prim and H temp The L1 distance between the two matrices is used to measure the difference between them in different dimensions. The calculation formula is as follows:

[0145] L1_distance = H prim -H temp .

[0146] Next, the calculated L1 distance is input into the fully connected layer (Dense layer) to predict the relationship between the two input matrices based on the L1 distance; Sigmoid is the activation function of this layer, and after outputting the predicted value, its log loss with the true value is calculated according to the following formula:

[0147] Where N represents the number of samples, y i It is the true label (0 or 1) of the i-th sample. It is the predicted probability value of the i-th sample (the probability that the model predicts it to be class 1).

[0148] Since the GRU network with target word embedding vector input shares the same weights and biases as the GRU network with primer pair input, this model can be called a Siamese RNN model. The hidden vector output by the unidirectional GRU layer has a size of 200, the optimizer is lazyadam, the learning rate is 0.0005, the decay rate is 1800, the decay exponent is 0.95, the maximum number of iterations is 300, the batch size is 8, and the packet loss rate is 0, as shown in Table 7.

[0149] Table 7

[0150] Due to the addition of an early stopping mechanism, the model typically does not require 300 iterations and can be trained in 19 epochs. On the validation set, the accuracy is 88% and the F1 score is 0.8. The trained model achieves 90% accuracy, an F1 score of 0.67, an MCC coefficient of 0.67, an AUC of 0.5, a precision of 1, and a recall of 0.5 on the test set, as shown in Table 8.

[0151] Table 8

[0152] Comparing the test results of the first deep learning model and the Siamese RNN model reveals that their accuracy and other metrics are almost identical, differing only in AUC. The AUC value of the first deep learning model is higher than that of the Siamese RNN model. (AUC (Area Under Curve) is defined as the area under the ROC curve and the coordinate axis, calculated based on the true label and the probability of a positive prediction; the closer the AUC is to 1.0, the higher the realism of the detection method). Furthermore, the first deep learning model only requires 5 iterations for fine-tuning to achieve good performance; therefore, using the first deep learning model is more efficient.

[0153] This disclosure provides a method for predicting the applicability of primers for amplification reactions. Based on a set of target sequences and primer sequences, it predicts whether the set of target and primer sequences can be simultaneously applied to RPA and PCR amplification reactions. This disclosure uses a first deep learning model to learn the sequence semantic features between the target and primer sequences; it combines multiple first feature values, such as target length, GC content, and secondary structure features, with the sequence semantic features to optimize the classification layer of the pre-trained model, training a large-scale deep learning model adapted to our task, providing an accurate judgment for predicting whether a new primer pair-template DNA can be simultaneously applied to RPA and PCR amplification reactions. This disclosure can learn a predictive model on a small dataset, effectively guiding the design and engineering of DNA templates and primers applicable to both PCR and RPA reactions.

[0154] The method for predicting the applicability of amplification reaction primers in this disclosure introduces deep learning to help design primers suitable for both RPA and PCR reactions. By optimizing and fine-tuning the pre-trained DNABert model, good classification performance can be obtained even with only a small amount of training data, and the prediction of RPA and PCR primer amplification results can be achieved more quickly and efficiently. By introducing multiple first feature values ​​such as target length, primer GC content, and potential secondary structures, and extracting multiple first feature values ​​and nonlinear features in the encoding representation vector, these features are concatenated and input into a fully connected network for prediction, thereby improving the model's prediction performance.

[0155] As shown in Figure 4, this embodiment of the present disclosure also provides an apparatus for predicting whether amplification reaction primers are suitable, including a data preprocessing module 410 and a prediction module 420, wherein:

[0156] The data preprocessing module 410 is configured to acquire the target sequence and primer sequence, convert the target sequence and primer sequence into multiple tokens, and generate one or more first feature values ​​based on the primer sequence and / or parameters of the amplification reaction.

[0157] The prediction module 420 is configured to input multiple tokens and one or more of the first feature values ​​into a first deep learning model to obtain a prediction result of whether the target sequence and primer sequence are simultaneously applicable to RPA amplification and PCR amplification reactions. The first deep learning model includes a word embedding layer, an encoding layer, and a first classification layer. The word embedding layer is configured to generate a total embedding representation vector based on multiple tokens. The encoding layer is configured to encode the total embedding representation vector to obtain an encoded representation vector. The first classification layer is configured to extract nonlinear features from the encoded representation vector and one or more of the first feature values, and perform classification prediction based on the extracted nonlinear features.

[0158] In some exemplary embodiments, the first feature value generated based on the primer sequence includes at least one of the following: the number of consecutive Gs in the first 5 bp of the 5' end, the number of Gs or Cs in the first 5 bp of the 3' end, the percentage of GC content, the number of primer dimer structures, and the number of primer hairpin structures.

[0159] In some exemplary embodiments, the data preprocessing module 410 converts the target sequence and the primer sequence into multiple tokens according to the k-mer method, where k is a natural number between 3 and 6.

[0160] In some exemplary embodiments, the first feature value generated based on the parameters of the amplification reaction includes: target length.

[0161] In some exemplary embodiments, the data preprocessing module 410 is further configured to: standardize one or more of the first feature values.

[0162] In some exemplary embodiments, the data preprocessing module 410 is further configured to: use the k-mer method to cut the DNA sequence to build a token vocabulary.

[0163] In some exemplary embodiments, the word embedding layer generates a total embedding representation vector based on multiple tokens, including:

[0164] Convert multiple tokens into word embedding vectors;

[0165] Generate a position embedding vector based on the position of each token in the sequence;

[0166] Generate a word type vector based on whether each token belongs to the target sequence or the primer sequence;

[0167] The word embedding vector, position embedding vector, and word type vector are summed element by element to obtain the total embedding representation vector.

[0168] In some exemplary embodiments, the encoding layer includes multiple stacked Transformer layers, each Transformer layer including multiple attention heads. The outputs of the multiple attention heads are concatenated and input into a linear layer for linear transformation to obtain a multi-head attention output. The last matrix output by the last Transformer layer among the multiple Transformer layers is the encoded representation vector.

[0169] In some exemplary embodiments, the first classification layer includes a bidirectional LSTM layer, a first fully connected network layer, and a second fully connected network layer, wherein:

[0170] The bidirectional LSTM layer is configured to extract bidirectional long-range dependencies in the encoded representation vector to obtain a latent feature vector;

[0171] The first fully connected network layer is configured to extract nonlinear features from one or more of the first feature values ​​to obtain a first nonlinear feature vector.

[0172] The second fully connected network layer is configured to perform classification prediction based on the hidden feature vector and the first nonlinear feature vector.

[0173] In some exemplary embodiments, the second fully connected network layer includes a first splicing layer, a feature extraction layer, and a classification prediction layer, wherein:

[0174] The first splicing layer is configured to splice the hidden feature vector with the first nonlinear feature vector to obtain a spliced ​​feature vector.

[0175] The feature extraction layer is configured to extract nonlinear features from the concatenated feature vector to obtain a second nonlinear feature vector;

[0176] The classification prediction layer is configured to perform classification prediction based on the second nonlinear feature vector.

[0177] In some exemplary embodiments, the first deep learning model includes multiple second sub-deep learning models, each of which includes its own word embedding layer, its own encoding layer, and its own bidirectional LSTM layer. The tokens input to the multiple second sub-deep learning models are transformed from the target sequence and the primer sequence according to different methods. The first classification layer also includes a second concatenation layer, which is disposed between the bidirectional LSTM layer and the second fully connected network layer and is configured to concatenate the latent feature vectors output by the multiple sub-deep learning models to obtain a concatenated latent feature vector.

[0178] In some exemplary embodiments, the tokens input to the plurality of second sub-deep learning models are transformed from the target sequence and the primer sequence according to different k-mer methods, and the k values ​​corresponding to different k-mer methods are different.

[0179] In some exemplary embodiments, the bidirectional LSTM layer includes a forward LSTM layer, a reverse LSTM layer, and a third splicing layer, wherein:

[0180] The forward LSTM layer is configured to perform forward computation on the encoded representation vector to obtain the forward hidden vector;

[0181] The reverse LSTM layer is configured to perform reverse computation on the encoded representation vector to obtain the reverse hidden vector;

[0182] The third concatenation layer is configured to concatenate the forward latent vector with the reverse latent vector to obtain a latent feature vector.

[0183] In some exemplary embodiments, the first deep learning model includes one or more second sub-deep learning models and a fully connected network, wherein each second sub-deep learning model is configured to process multiple tokens to obtain a first nonlinear feature among the multiple tokens; the fully connected network is configured to process one or more first feature values ​​to obtain a second nonlinear feature among the one or more first feature values, and to perform classification prediction based on the extracted first nonlinear feature and second nonlinear feature.

[0184] In some exemplary embodiments, the apparatus further includes a fine-tuning module, wherein the fine-tuning module is configured to: acquire a pre-trained second sub-deep learning model, the second sub-deep learning model including the word embedding layer, the encoding layer and the bidirectional LSTM layer; add the first fully connected network layer and the second fully connected network layer after the bidirectional LSTM layer to obtain the first deep learning model; and fine-tune the parameters of the first deep learning model using a first dataset, the first dataset including multiple pairs of target sequences and primer sequence sets, and label data on whether each pair of target sequences and primer sequence sets is simultaneously applicable to RPA amplification and PCR amplification reactions.

[0185] In some exemplary embodiments, the fine-tuning module uses a first dataset to fine-tune the parameters of a first deep learning model, including: dividing the first dataset into a first training dataset and a first test dataset, wherein the proportion of positive and negative samples in the first training dataset is close to the proportion of positive and negative samples in the first test dataset; iteratively training the pre-trained first deep learning model multiple times using the first training dataset and the first test dataset, keeping the weights of the encoding layer unchanged during each iteration, and adjusting the weights of the first classification layer according to a preset loss function.

[0186] This disclosure also provides an apparatus for predicting the suitability of amplification reaction primers, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of a method for predicting the suitability of amplification reaction primers as described in any embodiment of this disclosure, based on the instructions stored in the memory.

[0187] As shown in Figure 5, in one example, the device for predicting whether the amplification reaction primers are suitable may include: a processor 510, a memory 520, a bus system 530, and a transceiver 540, wherein the processor 510, the memory 520, and the transceiver 540 are connected via the bus system 530, the memory 520 is used to store instructions, and the processor 510 is used to execute the instructions stored in the memory 520 to control the transceiver 540 to transmit and receive signals. Specifically, transceiver 540 can acquire target sequence and primer sequence under the control of processor 510. Processor 510 converts the target sequence and primer sequence into multiple tokens and generates one or more first feature values ​​based on the primer sequence and / or parameters of the amplification reaction. The multiple tokens and one or more first feature values ​​are input into a first deep learning model to obtain a prediction result of whether the target sequence and primer sequence are simultaneously applicable to RPA amplification reaction and PCR amplification reaction. The first deep learning model includes a word embedding layer, an encoding layer, and a first classification layer. The word embedding layer is configured to generate a total embedding representation vector based on multiple tokens. The encoding layer is configured to encode the total embedding representation vector to obtain an encoded representation vector. The first classification layer is configured to extract nonlinear features from the encoded representation vector and one or more first feature values, and perform classification prediction based on the extracted nonlinear features.

[0188] It should be understood that processor 510 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0189] Memory 520 may include read-only memory and random access memory, and provides instructions and data to processor 510. A portion of memory 520 may also include non-volatile random access memory. For example, memory 520 may also store device type information.

[0190] In addition to the data bus, the bus system 530 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 530 in Figure 5.

[0191] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the processor 510 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in memory 520, and the processor 510 reads information from memory 520 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.

[0192] This disclosure also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method for predicting the suitability of amplification reaction primers as described in any embodiment of this disclosure. The method for predicting the suitability of amplification reaction primers by executing executable instructions is substantially the same as the method for predicting the suitability of amplification reaction primers provided in the above embodiments of this disclosure, and will not be described in detail here.

[0193] In some possible implementations, various aspects of the method for predicting the suitability of amplification reaction primers provided in this disclosure can also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the method for predicting the suitability of amplification reaction primers according to various exemplary embodiments of this disclosure as described above. For example, the computer device can perform the method for predicting the suitability of amplification reaction primers as described in embodiments of this disclosure.

[0194] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0195] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0196] It should be noted that the above embodiments or implementation methods are merely exemplary and not restrictive. Therefore, this disclosure is not limited to the content specifically shown and described herein. Various modifications, substitutions, or omissions can be made to the form and details of the implementations without departing from the scope of this disclosure.

Claims

1. A method for predicting the suitability of primers for an amplification reaction, comprising: Obtain the target sequence and primer sequence, convert the target sequence and primer sequence into multiple tokens, and generate one or more first feature values ​​based on the primer sequence and / or the parameters of the amplification reaction; Multiple tokens and one or more of the first feature values ​​are input into a first deep learning model to obtain a prediction result on whether the target sequence and primer sequence are simultaneously applicable to RPA amplification and PCR amplification reactions. The first deep learning model includes a word embedding layer, an encoding layer, and a first classification layer. The word embedding layer is configured to generate a total embedding representation vector based on multiple tokens. The encoding layer is configured to encode the total embedding representation vector to obtain an encoded representation vector. The first classification layer is configured to extract nonlinear features from the encoded representation vector and one or more of the first feature values, and perform classification prediction based on the extracted nonlinear features.

2. The method according to claim 1, wherein, The first feature value generated based on the primer sequence includes at least one of the following: The number of consecutive Gs in the first 5 bp of the 5' end, the number of Gs or Cs in the first 5 bp of the 3' end, the percentage of GC content, the number of primer dimer structures, and the number of primer hairpin structures.

3. The method according to claim 1, wherein, The first feature value generated based on the parameters of the amplification reaction includes: target length.

4. The method according to claim 1, further comprising: Standardize one or more of the first feature values.

5. The method according to claim 1, wherein, The word embedding layer generates a total embedding representation vector based on multiple tokens, including: Convert multiple tokens into word embedding vectors; Generate a position embedding vector based on the position of each token in the sequence; Generate a word type vector based on whether each token belongs to the target sequence or the primer sequence; The word embedding vector, the position embedding vector, and the word type vector are added element by element to obtain the total embedding representation vector.

6. The method according to claim 1, wherein, The encoding layer includes multiple stacked Transformer layers, each Transformer layer includes multiple attention heads, the outputs of the multiple attention heads are concatenated and input into a linear layer for linear transformation to obtain multi-head attention output, and the last matrix output by the last Transformer layer among the multiple Transformer layers is the encoding representation vector.

7. The method according to claim 1, wherein, The first classification layer includes a bidirectional LSTM layer, a first fully connected network layer, and a second fully connected network layer; The bidirectional LSTM layer is configured to extract bidirectional long-range dependencies in the encoded representation vector to obtain a latent feature vector; The first fully connected network layer is configured to extract nonlinear features from one or more of the first feature values ​​to obtain a first nonlinear feature vector. The second fully connected network layer is configured to perform classification prediction based on the hidden feature vector and the first nonlinear feature vector.

8. The method according to claim 7, wherein, The second fully connected network layer includes a first splicing layer, a feature extraction layer, and a classification prediction layer, wherein: The first splicing layer is configured to splice the hidden feature vector with the first nonlinear feature vector to obtain a spliced ​​feature vector. The feature extraction layer is configured to extract nonlinear features from the concatenated feature vector to obtain a second nonlinear feature vector; The classification prediction layer is configured to perform classification prediction based on the second nonlinear feature vector.

9. The method according to claim 7, wherein, The first deep learning model includes multiple second sub-deep learning models, each of which includes its own word embedding layer, its own encoding layer, and its own bidirectional LSTM layer. The token input to the multiple second sub-deep learning models is generated by converting the target sequence and the primer sequence according to different methods. The first classification layer further includes a second concatenation layer, which is located between the bidirectional LSTM layer and the second fully connected network layer. It is configured to concatenate the latent feature vectors output by the multiple sub-deep learning models to obtain a concatenated latent feature vector.

10. The method according to claim 9, wherein, The tokens input to the multiple second sub-deep learning models are transformed from the target sequence and the primer sequence using different k-mer methods, with different k values ​​corresponding to different k-mer methods.

11. The method according to claim 7, wherein, The bidirectional LSTM layer includes a forward LSTM layer, a reverse LSTM layer, and a third splicing layer; The forward LSTM layer is configured to perform forward computation on the encoded representation vector to obtain the forward hidden vector; The reverse LSTM layer is configured to perform reverse computation on the encoded representation vector to obtain the reverse hidden vector; The third concatenation layer is configured to concatenate the forward latent vector with the reverse latent vector to obtain a latent feature vector.

12. The method according to claim 7, further comprising, prior to the method: Obtain a pre-trained second sub-deep learning model, the second sub-deep learning model including the word embedding layer, the encoding layer and the bidirectional LSTM layer; After the bidirectional LSTM layer, the first fully connected network layer and the second fully connected network layer are added to obtain the first deep learning model.

13. The method according to claim 12, further comprising, after obtaining the first deep learning model: The parameters of the first deep learning model are fine-tuned using a first dataset, which includes multiple pairs of target sequences and primer sequence sets, as well as tag data on whether each pair of target sequences and primer sequence sets is applicable to both RPA amplification and PCR amplification reactions.

14. The method according to claim 13, wherein, The step of fine-tuning the parameters of the first deep learning model using the first dataset includes: The first dataset is divided into a first training dataset and a first test dataset, and the ratio of positive to negative samples in the first training dataset is close to that in the first test dataset. The first deep learning model is trained iteratively multiple times using the first training dataset and the first test dataset. During each iteration, the weights of the encoding layer remain unchanged, and the weights of the first classification layer are adjusted according to a preset loss function.

15. The method according to claim 1, wherein, The token input to the first deep learning model is generated by converting the target sequence and the primer sequence using the k-mer method, where k is a natural number between 3 and 6.

16. The method according to claim 1, further comprising: The k-mer method was used to segment the DNA sequence to build a token vocabulary.

17. An apparatus for predicting the suitability of amplification reaction primers, comprising a memory; and a processor connected to the memory, the memory for storing instructions, the processor being configured to perform the steps of the method for predicting the suitability of amplification reaction primers as claimed in any one of claims 1 to 16 based on the instructions stored in the memory.

18. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for predicting whether an amplification reaction primer is applicable as described in any one of claims 1 to 16.

19. A computer program product comprising instructions that, when executed by a computer, perform a method for predicting whether an amplification reaction primer is applicable as claimed in any one of claims 1 to 16.

20. An apparatus for predicting the suitability of primers for an amplification reaction, comprising a data preprocessing module and a prediction module, wherein: The data preprocessing module is configured to acquire target sequences and primer sequences, convert the target sequences and primer sequences into multiple tokens, and generate one or more first feature values ​​based on the primer sequences and / or parameters of the amplification reaction. The prediction module is configured to input multiple tokens and one or more of the first feature values ​​into a first deep learning model to obtain a prediction result of whether the target sequence and primer sequence are simultaneously applicable to RPA amplification and PCR amplification reactions. The first deep learning model includes a word embedding layer, an encoding layer, and a first classification layer. The word embedding layer is configured to generate a total embedding representation vector based on multiple tokens. The encoding layer is configured to encode the total embedding representation vector to obtain an encoded representation vector. The first classification layer is configured to extract nonlinear features from the encoded representation vector and one or more of the first feature values, and perform classification prediction based on the extracted nonlinear features.

Citation Information

Patent Citations

  • Process for aligning targeted nucleic acid sequencing data

    CN110692101A

  • Primer design method and system based on k-mer algorithm

    CN111326210A

  • RNA-protein binding site prediction method for context deep characterization

    CN117524307A

  • Systems and methods for off-target sequence detection

    US20180075186A1