A method for generating candidate nucleic acid aptamers based on Poisson flow conditional generation model
Through a method based on the Poisson flow condition generation model, combined with SELEX technology and protein-ligand comparison pre-training model, high-quality nucleic acid aptamer sequences for specific proteins are generated, which solves the problem of generating high-quality aptamers in the prior art and achieves rapid and effective aptamer generation.
Patent Information
- Application Number
- CN202411066717.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-08-06
AI Technical Summary
The prior art has difficulty in efficiently generating high-quality nucleic acid aptamers, especially in the sequence space, where there is a challenge in selecting aptamers that can specifically bind to the target protein, and existing models cannot be conditionally generated for specific binding pockets.
The candidate nucleic acid aptamer generation method based on the Poisson flow condition generation model is adopted. By obtaining the protein data of the target protein, the candidate aptamer sequence is screened using SELEX technology, and the protein-ligand comparison pre-trained model and the noise reduction prediction model are combined to generate aptamer sequences for specific proteins.
The rapid and efficient generation of high-quality aptamer sequences is achieved, and targeted generation can be carried out for specific binding pockets, overcoming the problems of insufficient sequence space sampling and data limitation in the prior art.
Smart Images

Figure CN119108019B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical fields of artificial intelligence, computational biology and nucleic acid aptamer design, and in particular to a method for generating candidate nucleic acid aptamers based on a Poisson flow conditional generation model. Background Art
[0002] As a single-stranded oligonucleotide, aptamers have the advantages of short screening cycle, suitability for low immunogenic antigens, wide binding targets, low synthesis cost and thermal stability compared to traditional antibodies. However, high-quality aptamers are extremely rare in sequence space. The traditional systematic evolution of ligands by exponential enrichment (SELEX) technology can only sample a small part of the sequence space for screening, and various candidate aptamers are limited by the actual sequencing capacity in the experiment. Therefore, it is very challenging to select aptamers that can specifically bind to the target protein from the sequence space.
[0003] In recent years, artificial intelligence-driven drug discovery (AIDD) has received widespread attention from researchers. Many artificial intelligence technologies have been successfully applied to various tasks in drug discovery. One application method is to generate aptamers using sequences screened by SELEX technology. Most existing models are trained based on SELEX data of a single protein, or optimized based on existing sequence information. They have strong limitations and cannot generate conditions for specific binding pockets. In addition, the amount of data on protein-aptamer complexes in existing public data is very small. How to make full use of these limited data to guide new drug development is also an urgent problem to be solved.
[0004] Glossary:
[0005] Transformer: is a deep learning model architecture used for natural language processing (NLP) and other sequence-to-sequence tasks;
[0006] U-Net: A deep learning model based on convolutional neural networks;
[0007] ESM-2 model: a Transformer-based deep learning model that takes the amino acid sequence of a protein as input to generate protein representations;
[0008] MASSA model: A deep learning model based on graph neural network that takes the amino acid sequence, properties and structure of proteins as input to generate protein representation;
[0009] Levenshtein distance: is an algorithm used to measure the similarity between two strings;
[0010] GC base content: the ratio of guanine and cytosine among the four bases in deoxyribonucleic acid.
[0011] SELEX technology: Systematic Evolution of Ligands by Exponential Enrichment, is a technology that selects nucleic acid aptamers with high affinity to target substances from a random single-stranded nucleic acid sequence library. This technology first synthesizes a single-stranded oligonucleotide library in vitro and then mixes it with the target substance. By washing away the nucleic acid that is not bound to the target substance, the nucleic acid molecules bound to the target substance are separated and used as a template for PCR amplification to carry out the next round of screening. Through repeated screening and amplification, DNA or RNA molecules with high affinity to the target substance are finally separated from the random library. Summary of the invention
[0012] In order to solve the above technical problems, the present invention discloses a method for generating candidate nucleic acid aptamers based on a Poisson flow conditional generation model.
[0013] The technical solution of the present invention is as follows:
[0014] A method for generating candidate nucleic acid aptamers based on a Poisson flow conditional generation model comprises the following steps:
[0015] Step 1: obtaining protein data P of a plurality of target proteins, wherein the protein data includes at least one of amino acid sequence data and protein structure data;
[0016] Step 2: Screening out candidate aptamer nucleotide sequence data X corresponding to the target protein by SELEX technology to obtain a training data set {X, P};
[0017] Step 3: Extract features of protein data P through the protein pre-training model to obtain the feature vector ∈ p ;
[0018] Step 4: Input the aptamer nucleotide sequence X corresponding to the target protein into the pre-trained model for extracting aptamer representation, and obtain the representation of the aptamer corresponding to the target protein in the latent space ∈ a , construct a protein-ligand comparison pre-training model, and transform the feature vector ∈ p Input the protein-ligand comparison pre-trained model to obtain a new protein representation of the target protein in the latent space where the aptamer is located The protein-ligand comparison pre-training model is trained until the loss function loss is minimized or the preset number of cycles is reached, thereby obtaining a trained protein-ligand comparison pre-training model;
[0019] The loss function loss is as follows:
[0020]
[0021] in, is the cosine similarity between protein representation and aptamer representation, i represents the i-th training data, |||| represents the norm, and N represents the number of training data; t i is the correspondence between the i-th aptamer and the selected target protein. If the i-th aptamer is obtained by SELEX screening of the selected target protein, then t i =1, otherwise t i =0;
[0022] Step 5: New protein representation based on protein data P As a condition, the aptamer nucleotide sequence data X is used to train the noise reduction prediction model to obtain a trained noise reduction prediction model;
[0023] Step 6: Input the protein data of the protein for which an aptamer needs to be generated into the trained protein-ligand comparison pre-training model to obtain the corresponding new protein representation, and then input the new protein representation into the trained noise reduction prediction model to generate the corresponding aptamer.
[0024] As a further improvement, in step three, feature extraction is performed on the protein data P using the MASSA model.
[0025] As a further improvement, in step 4, the protein-ligand comparison pre-training model is a deep learning model CPAP.
[0026] For further improvement, the specific steps of step 5 are as follows:
[0027] Step 5.1: Randomly sample a batch of data from the training set {X, P} Where B is the number of sampled data, X i is the aptamer sequence in the i-th training data, P i is the protein sequence in the i-th training data;
[0028] Step 5.2: The nucleotide sequence in the aptamer contains four bases: adenine A, thymine T, guanine G, and cytosine C. i The base sequence is one-hot encoded to obtain the one-hot encoding representation The representation is converted to continuous space by the following method to obtain the new aptamer representation z i :
[0029]
[0030] in is the one-hot encoding representation of the i-th aptamer, l is the length of the aptamer sequence, u is a random variable that obeys a uniform distribution, and U represents uniform distribution;
[0031] Step 5.3: z i Standardize and get the standardized representation let The value of each dimension is within (-1,1), and the formula is as follows:
[0032]
[0033] Step 5.4: Initialize the high-dimensional Poisson flow aptamer conditional generative model parameters: First, sample the standard deviation from the log-normal distribution in a, b are hyperparameters, B is the number of data in a batch, is a Gaussian distribution, σ i represents a random variable that follows an exponential Gaussian distribution for the i-th data sample, Represents Gaussian distribution; then randomly select sampling point r i : Where D is the augmented dimension of the Poisson flow model; the perturbation radius R corresponding to the i-th training data sampled i : in D data is the dimension of the training data, R1 represents a random variable that obeys the beta distribution, Beta() represents the beta distribution, α and β are two parameters that determine the beta distribution; the sampling perturbation angle in I is the identity matrix; u i represents a random variable that follows a Gaussian distribution, v i represents the i-th disturbance angle;
[0034] Step 5.5: Perturb the aptamer sequence data into the enhanced space to obtain the perturbed aptamer sequence data represents the sequence data of the perturbed aptamer;
[0035] Step 5.6: Transform the feature vector ∈ p Input the protein-ligand comparison pre-trained model to obtain a new protein representation of the target protein in the latent space where the aptamer is located As label input, a high-dimensional Poisson flow adapter conditional generative model is constructed;
[0036] Step 5.7: Using the perturbed aptamer sequence data The denoising prediction model fθ() is calculated, and the loss function of the denoising prediction model is Where fθ() is a deep learning model based on Transformer architecture or U-Net architecture, λ(σ i ) is the loss weight, λ(σ i )=1 / c out (σ i ) 2 ;
[0037] in Fθ() is a deep learning model based on Transformer architecture, c in (σ i ) is the control input The function of weight, c out (σ i ) is the control F θ () Output weight function, c skip (σ i ) is used to control the jump connection The function of weight, c noise (σ i ) is the function that controls the noise weight, σ data is the standard deviation of the training data;
[0038] Step 5.8: Update the noise reduction prediction model parameters, and repeat steps 5.1-5.7 until the preset number of cycles or the noise reduction prediction model parameters converge to obtain the trained noise reduction prediction model.
[0039] For further improvement, the specific steps of step six are as follows:
[0040] Step 6.1: Input the protein for which the aptamer needs to be generated into the protein pre-training model to obtain the feature vector ∈ p , feature vector ∈ p Input the protein-ligand comparison pre-training model to obtain the new protein representation of the protein that needs to generate the aptamer in the latent space where the aptamer is located The protein for which the aptamer needs to be generated is a new protein representation in the latent space where the aptamer is located Conditional generation model of high-dimensional Poisson flow aptamer trained for conditional input;
[0041] Step 6.2: Initialize parameters and set the farthest sampling position where σ max is the parameter to control the sampling starting point, D is the augmented dimension of the Poisson flow model, and the sampling perturbation radius in D data is the dimension of the training data, and the sampling perturbation angle in I is the identity matrix;
[0042] Step 6.3: Randomly generate an initial vector with the same dimension as the aptamer representation It is equivalent to sampling at infinity on the hyperplane as the initial input of the generative model;
[0043] Step 6.4: As a condition, the trained denoising prediction model As the denoising function, the second-order Huon method sampling is constructed so that Falling back to the initial hyperplane, we get the corresponding Then, the target aptamer X0 is generated by obtaining the maximum dimension among the four dimensions corresponding to each base.
[0044] The effective effects of the present invention are:
[0045] 1. The present invention uses a Poisson flow generation model, taking into account both the training speed and training effect of the generation model.
[0046] 2. The present invention uses protein information as a generation condition, and can generate aptamers in a targeted manner. Different from the previous use of data from a single protein, by using the SELEX screening results of multiple proteins as prior knowledge to learn the relationship between proteins and aptamers, compared to existing models that can only learn aptamer distribution characteristics from the SELEX data of a single protein, this model can further optimize the aptamer sequence for a certain protein based on its unsatisfactory SELEX screening results.
[0047] 3. The present invention proposes a CPAP pre-training model that can achieve zero-sample generation. Through the CPAP model, the protein representation is matched with the representation of the corresponding aptamer sequence screened by SELEX. Compared with directly using the protein representation as a condition, the aptamer sequence generated by the present invention for proteins not included in the data set has better properties. Therefore, the present application effectively overcomes the various shortcomings of the prior art and has a high industrial utilization value. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic diagram of a flow chart of a method for generating an aptamer provided in an embodiment of the present invention.
[0049] Figure 2 Schematic diagram of the pre-trained model for protein-ligand comparison.
[0050] Figure 3Schematic diagram of the model generated for the Poisson flow aptamer;
[0051] Figure 4 This is a schematic diagram comparing the Levenshtein distance, GC base content distribution, and minimum free energy distribution of the data generated by the present invention with the experimental data, the data generated by the unconditional model, and the randomly generated data. DETAILED DESCRIPTION
[0052] The following examples are provided to further illustrate the present invention, but the present invention is not limited thereto.
[0053] The present invention provides a method for generating aptamer conditions based on a Poisson flow generation model, such as Figure 1 As shown, the method includes:
[0054] Step 1: Obtain sequence data and / or structural data of the target protein: In this experimental example, preferably, we obtained 127 sets of candidate nucleic acid sequences for high-throughput screening from the laboratory and recorded the uniport numbers of the corresponding proteins. We obtained the PDB files of the proteins in the AlphaFold protein structure database, obtained their sequence and structural data, and divided the data into training sets and validation sets, of which 100 sets of data were used as training sets and 27 sets of data were used as validation sets;
[0055] Step 2: Obtain the aptamer sequence data X corresponding to the protein screened by the SELEX technology, and perform preprocessing to obtain a training data set {X, P}: In this embodiment, preferably, the specific steps of step 2 include: firstly, count each sequence X i The number of enriched entries after high-throughput screening i , retain the sequences with more than 10 enrichment numbers, and duplicate lnp for each sequence i Second, save and high-throughput screening of candidate nucleic acid sequences X i Corresponding protein information P i , get the final training data {X, P};
[0056] Step 3: Using the sequence data and structure data P of the protein, feature extraction is performed through the protein pre-training model to obtain the feature vector ∈ p : Preferably, we use the MASSA model, taking the sequence and structure data in the obtained PDB file as input, and obtain the corresponding representation ∈ p ;
[0057] Step 4: Figure 2 As shown, a protein-ligand comparison pre-training (CPAP) model is constructed and trained, and the protein feature vector ∈ pMapping to the latent space used to characterize the aptamer, so that the new protein feature vector is located near the characterization vector of the aptamer corresponding to the protein in the latent space;
[0058] In this embodiment, preferably, the specific steps of step 4 include:
[0059] Step 4.1: Using the protein feature vector ∈ p As input, it is input into the deep learning model CPAP() used to map the vector to the aptamer representation space, and the protein representation in the latent space where the aptamer is located is obtained. Among them, CPAP() consists of Transformer and 5 fully connected layers;
[0060] Step 4.2: Take the corresponding aptamer sequence X as input and input it into the pre-trained model RNA-FM for extracting aptamer representation to obtain the representation of the aptamer in the latent space ∈ a ;
[0061] Step 4.3: Train the CPAP model described in step 4.1 and transform the protein representation vector and the corresponding aptamer characterization vector ∈ a Align in the latent space and calculate With ∈ a The cosine similarity of protein and corresponding aptamer is close to 1, and the similarity with non-corresponding aptamer is close to 0. The cross entropy loss function is used as the loss function of the model. The formula is as follows:
[0062]
[0063] in is the cosine similarity between protein representation and aptamer representation, t i is the correspondence between the selected aptamer and the selected protein. If the selected aptamer is obtained by SELEX screening of the selected protein, then t i =1, otherwise t i = 0. The model is trained using this loss function for 450 epochs to obtain a trained CPAP model;
[0064] Step 4.4: Through the trained CPAP model, the protein representation behind the aptamer in the latent space and the corresponding aptamer is obtained. As a condition for the generation of high-dimensional Poisson flow aptamer models.
[0065] Step 5: Figure 3 As shown, taking the protein feature vector as a condition, the training data obtained in step 2 is used to construct and train a high-dimensional Poisson flow aptamer conditional generation model;
[0066] In this example, preferably, in step 5, the protein feature vector As a condition, the aptamer sequence data X is used to construct and train a high-dimensional Poisson flow aptamer conditional generation model, specifically including:
[0067] Step 5.1: Randomly sample a batch of aptamer data from the training set X (aptamer sequence data) Where B is the number of data in a batch, B = 1024;
[0068] Step 5.2: The nucleotide sequence in the aptamer contains four bases: adenine A, thymine T, guanine G, and cytosine C. i The base sequence is one-hot encoded to obtain the one-hot encoding representation The representation is converted to continuous space by the following method to obtain the new aptamer representation z i :
[0069]
[0070] in is the one-hot encoding representation of the i-th aptamer, l is the length of the aptamer sequence, and u is a random variable that follows a uniform distribution. Step 5.3: Set z i Standardize and get the standardized representation Let the value of each dimension be within (-1,1), the formula is as follows:
[0071]
[0072] Step 5.4: Initialize the high-dimensional Poisson flow aptamer conditional generation model parameters. Specifically, first sample the standard deviation from the log-normal distribution in B is the number of data in a batch, B = 1024; then randomly select sampling point r i : Where D is the augmented dimension of the Poisson flow model, D = 2048; the sampling perturbation radius R i : in N is the dimension of training data, N=144; sampling perturbation angle where u i ~N(0,I), I is the identity matrix;
[0073] Step 5.5: Perturb the aptamer sequence data into the enhanced space to obtain the perturbed aptamer structure data
[0074] Step 5.6: The protein sequence data and / or structure data are used to obtain the feature vector of the corresponding protein using the feature vector extracted in step 4. Input to the network as labels;
[0075] Step 5.7: Calculate the noise reduction prediction model using the perturbed aptamer sequence data:
[0076]
[0077] in is the noise reduction prediction model, F θ () is a deep learning model based on the Transformer architecture, λ(σ i ) is the loss weight, λ(σ i )=1 / c out (σ i ) 2 , σ data =0.5;
[0078] Step 5.8: Update the network parameters through the Adam optimizer, and loop through steps 5.1-5.5 to obtain the denoising prediction model and finally obtain the Poisson flow generation model.
[0079] Step 6: Generate new aptamers based on the trained Poisson flow generation model and using the required protein sequence data and / or structure data as conditions.
[0080] In this example, preferably, in step 6, generating a new aptamer according to the Poisson flow generation model specifically includes:
[0081] Step 6.1: The protein sequence data and / or structure data are used to obtain the feature vector of the corresponding protein using the feature vector extracted in step 4. Enter the network as a condition;
[0082] Step 6.2: Initialize parameters and set the farthest sampling position where σ max To control the parameters of the sampling starting point, D is the augmented dimension of the Poisson flow model, D = 2048, and the sampling perturbation radius in N is the dimension of training data, N=144, sampling perturbation angle in I is the identity matrix;
[0083] Step 6.3: Randomly generate an initial vector with the same dimension as the aptamer representation It is equivalent to sampling at infinity on the hyperplane as the initial input of the generative model;
[0084] Step 6.4: As a condition, the trained denoising prediction model As the denoising function, the second-order Huon method sampling is constructed so that Falling back to the initial hyperplane, we get the corresponding Then, by obtaining the maximum dimension among the four dimensions corresponding to each base, the target aptamer X0 can be generated.
[0085] In this example, preferably, the control protein is used as a conditional input to generate the network, and the generated sequence is compared and analyzed with the control sequence, the unconditional generated sequence, and the randomly generated sequence in terms of Levenshtein distance, GC base content, and minimum free energy distribution. Some of the results are as follows: Figure 4 As shown, compared with the sequences generated using the unconditional version of the model and the randomly generated sequences, the aptamer sequences generated by the present invention are closer in Levenshtein distance and GC base content to the control group sequences, and the distribution of minimum free energy is also more similar, indicating that the sequences generated by the present invention are of better quality.
[0086] The present invention is not limited to the above-mentioned implementation method. Anyone should know that any technical solution identical or similar to the present invention made under the inspiration of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for generating candidate nucleic acid aptamers based on a Poisson flow conditional generation model, characterized in that: The steps include: Step 1: obtaining protein data P of a plurality of target proteins, wherein the protein data includes at least one of amino acid sequence data and protein structure data; Step 2: Screening out candidate aptamer nucleotide sequence data X corresponding to the target protein by SELEX technology to obtain a training data set {X, P}; Step 3: Extract features of protein data P through the protein pre-training model to obtain the feature vector ∈ p ; Step 4: Input the aptamer nucleotide sequence X corresponding to the target protein into the pre-trained model for extracting aptamer representation, and obtain the representation of the aptamer corresponding to the target protein in the latent space ∈ a , construct a protein-ligand comparison pre-training model, and transform the feature vector ∈ p Input the protein-ligand comparison pre-trained model to obtain a new protein representation of the target protein in the latent space where the aptamer is located The protein-ligand comparison pre-training model is trained until the loss function loss is minimized or the preset number of cycles is reached, thereby obtaining a trained protein-ligand comparison pre-training model; The loss function loss is as follows: in, is the cosine similarity between protein representation and aptamer representation, i represents the i-th training data, || || represents the norm, and N represents the number of training data; t i is the correspondence between the i-th aptamer and the selected target protein. If the i-th aptamer is obtained by SELEX screening of the selected target protein, then t i =1, otherwise t i =0; Step 5: New protein representation based on protein data P As a condition, the aptamer nucleotide sequence data X is used to train the noise reduction prediction model to obtain a trained noise reduction prediction model; Step 5.1: Randomly sample a batch of data from the training set {X, P} Where B is the number of sampled data, X i is the aptamer sequence in the i-th training data, P i is the protein sequence in the i-th training data; Step 5.2: The nucleotide sequence in the aptamer contains four bases: adenine A, thymine T, guanine G, and cytosine C. i The base sequence is one-hot encoded to obtain the one-hot encoding representation The representation is converted to continuous space by the following method to obtain the new aptamer representation z i : in is the one-hot encoding representation of the i-th aptamer, l is the length of the aptamer sequence, u is a random variable that obeys a uniform distribution, and U represents uniform distribution; Step 5.3: z i Standardize and get the standardized representation let The value of each dimension is within (-1,1), and the formula is as follows: Step 5.4: Initialize the high-dimensional Poisson flow aptamer conditional generative model parameters: First, sample the standard deviation from the log-normal distribution in a, b are hyperparameters, B is the number of data in a batch, is a Gaussian distribution, σ i represents a random variable that follows an exponential Gaussian distribution for the i-th data sample, Represents Gaussian distribution; then randomly select sampling point r i : Where D is the augmented dimension of the Poisson flow model; the perturbation radius corresponding to the i-th training data is in D data is the dimension of the training data, R1 represents a random variable that obeys the beta distribution, Beta() represents the beta distribution, α and β are two parameters that determine the beta distribution; the sampling perturbation angle in I is the identity matrix; u i represents a random variable that follows a Gaussian distribution, v i represents the i-th disturbance angle; Step 5.5: Perturb the aptamer sequence data into the enhanced space to obtain the perturbed aptamer sequence data represents the sequence data of the perturbed aptamer; Step 5.6: Transform the feature vector ∈ p Input the protein-ligand comparison pre-trained model to obtain a new protein representation of the target protein in the latent space where the aptamer is located As label input, a high-dimensional Poisson flow adapter conditional generative model is constructed; Step 5.7: Using the perturbed aptamer sequence data The noise reduction prediction model f is calculated θ (), the loss function of the denoising prediction model is where f θ () is a deep learning model based on Transformer architecture or U-Net architecture, λ(σ i ) is the loss weight, λ(σ i )=1 / c out (σ i ) 2 ; in F θ () is a deep learning model based on the Transformer architecture, c in (σ i ) is the control input The function of weight, To control F θ () Output weight function, c skip (σ i ) is used to control the jump connection The function of weight, c noise (σ i ) is the function that controls the noise weight, σ data is the standard deviation of the training data; Step 5.8: Update the noise reduction prediction model parameters, and repeat steps 5.1-5.7 until the preset number of cycles or the noise reduction prediction model parameters converge to obtain a trained noise reduction prediction model; Step 6: Input the protein data of the protein for which an aptamer needs to be generated into the trained protein-ligand comparison pre-training model to obtain the corresponding new protein representation, and then input the new protein representation into the trained noise reduction prediction model to generate the corresponding aptamer.
2. The method for generating candidate nucleic acid aptamers based on a Poisson flow conditional generation model according to claim 1, characterized in that: In the step three, feature extraction is performed on the protein data P using the MASSA model.
3. The method for generating candidate nucleic acid aptamers based on a Poisson flow conditional generation model according to claim 1, characterized in that: In the step 4, the protein-ligand comparison pre-training model is a deep learning model CPAP.
4. The method for generating candidate nucleic acid aptamers based on a Poisson flow conditional generation model according to claim 1, characterized in that: The specific steps of step six are as follows: Step 6.1: Input the protein for which the aptamer needs to be generated into the protein pre-training model to obtain the feature vector ∈ p , feature vector ∈ p Input the protein-ligand comparison pre-training model to obtain the new protein representation of the protein that needs to generate the aptamer in the latent space where the aptamer is located The protein for which the aptamer needs to be generated is a new protein representation in the latent space where the aptamer is located A trained high-dimensional Poisson flow adapter conditional generative model for conditional input; Step 6.2: Initialize parameters and set the farthest sampling position where σ max is the parameter to control the sampling starting point, D is the augmented dimension of the Poisson flow model, and the sampling perturbation radius in D data is the dimension of the training data, and the sampling perturbation angle in I is the identity matrix; Step 6.3: Randomly generate an initial vector with the same dimension as the aptamer representation It is equivalent to sampling at infinity on the hyperplane as the initial input of the generative model; Step 6.4: As a condition, the trained denoising prediction model As the denoising function, the second-order Huon method sampling is constructed so that Falling back to the initial hyperplane, we get the corresponding Then, the target aptamer X0 is generated by obtaining the maximum dimension among the four dimensions corresponding to each base.
Citation Information
Patent Citations
Aptamer generation method based on conditional discrete diffusion model
CN116631499A
KR20200019404A