Activity-guided deep implicit evolution-based aptamer design method and application thereof
Through a deep implicit evolution method based on activity-guided, GM-VAE training and hidden space genetic evolution strategy optimize aptamer sequences are solved, and the problem of low aptamer screening efficiency and poor accuracy in traditional methods is achieved, achieving the effect of efficient discovery of highly active aptamer.
Patent Information
- Application Number
- CN202510211665.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Traditional aptamer screening methods have problems such as low data processing efficiency, poor screening accuracy and difficulty in effectively using high-throughput data, which makes it difficult to improve the activity and selectivity of aptamer.
A deep implicit evolution method based on activity guidance was used to train the aptamer sequences through Gaussian hybrid variable autosegment encoder (GM-VAE), capture their potential structural characteristics and family distribution, and optimize the aptamer sequences through hidden space genetic evolution strategies.
The screening efficiency of aptamer is improved, and the discovery efficiency and accuracy of high-active aptamers are significantly improved, and the shortcomings of traditional methods in data processing and screening efficiency are overcome.
Smart Images

Figure BDA0005286169720000041 
Figure BDA0005286169720000121 
Figure BDA0005286169720000123
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of molecular biology, and specifically, relates to a novel design method and application of aptamers based on activity-guided deep implicit evolution. Background Art
[0002] Aptamers are usually obtained by systematic evolution of ligands by exponential enrichment (SELEX), but traditional aptamer screening and design methods have many limitations. In SELEX, aptamer selection based on read counts is susceptible to interference from polymerase chain reaction (PCR) amplification bias, resulting in the fact that high-abundance sequences do not necessarily correspond to high affinity. In addition, early clustering techniques, such as AptaCluster and FASTAptamer, although attempting to reveal the commonalities of aptamer sequences, ignored the secondary structure of the sequences. Similarly, motif finding methods such as MEMRIS and Aptamotif, although focusing on high-affinity structural patterns, also did not fully consider the influence of the overall secondary structure. More importantly, these methods are less efficient when dealing with large HT-SELEX libraries and are difficult to quickly and accurately screen out high-quality aptamers from massive data.
[0003] With the development of deep sequencing technology, the amount of data generated by aptamer screening has increased rapidly, but traditional methods are difficult to effectively utilize this data to promote the progress of aptamer research. Research shows that the bias generated during PCR amplification results in the fact that high-abundance sequences do not always reflect their affinity for the target molecule, further reducing the accuracy of screening. In addition, traditional clustering methods are time-consuming in the calculation process and unable to efficiently capture aptamer sequences with high affinity when facing large-scale data. The strategy of screening based on high-abundance sequences, although some sequences can be discovered in the short term, fails to effectively improve the activity and selectivity of the final aptamers. Therefore, there is an urgent need for new technical solutions to overcome the bottlenecks of existing methods in data processing and screening efficiency.
[0004] In actual screening, although next-generation sequencing technology is used to sort and verify the copy numbers of nucleic acid sequences, due to the lack of a direct positive correlation between copy numbers, fluorescence activation multiples and affinity, the discovery efficiency of aptamers in the early stage is still low. Even more complicated is that in the early rounds of screening, there are often a large number of disordered sequences; while in the later stage of screening, aptamer sequences have clustered into families, and it becomes more difficult to identify the active sequences among them.
[0005] These problems further highlight the deficiencies of traditional methods in dealing with large-scale data. Therefore, there is an urgent need in this field to develop more advanced technical solutions to improve the screening efficiency and effectively enhance the activity and applicability of aptamers. Summary of the Invention
[0006] The present invention provides a method for improving the efficiency of aptamer screening and effectively discovering highly active aptamers.
[0007] In a first aspect of the present invention, there is provided a method for aptamer design and optimization based on activity-guided deep implicit evolution, comprising the following steps:
[0008] (s1) Provide a high-throughput screening aptamer sequence dataset, which includes nucleic acid sequence information and copy number information of aptamers; and randomly divide the aptamer sequence dataset into a training set and a test set according to a certain ratio;
[0009] (s2) Data processing: Convert each nucleotide character in the nucleic acid sequence in the dataset into a corresponding numerical index, so as to obtain the numerical form of each nucleic acid sequence; and expand the numerical index into a vector, so as to obtain a nucleic acid sequence representation composed of vectors;
[0010] Among them, the vector representation of nucleotides in the nucleic acid sequence in the training set is the first vector, and the vector representation of nucleotides in the nucleic acid sequence in the test set is the second vector;
[0011] (s3) Encoder-decoder training: Input the nucleic acid sequence representation composed of the first vectors into a Gaussian mixture variational autoencoder (GM-VAE) for training, so that the trained encoder can capture the potential structural features and aptamer family distributions in the aptamer nucleic acid sequence, and map these features to a low-dimensional latent space; and enable the decoder to generate the original input sequence; then use the nucleic acid sequence representation composed of the second vectors for testing;
[0012] (s4) Data dimensionality reduction and clustering: Use the trained encoder in (s3) to perform embedded dimensionality reduction on the nucleic acid sequences of the target round, so as to map the nucleic acid sequences into a low-dimensional latent space, capture the potential similarities and structural features between nucleic acid sequences; and obtain different clusters (or families) of nucleic acid sequences according to the potential similarities and structural features;
[0013] (s5) Selection of active sequence regions and deep implicit evolution: Select the clusters corresponding to nucleic acid sequences with high copy numbers and define them as active sequence regions; and perform a latent space genetic evolution strategy on the active sequence regions;
[0014] (s6) Obtain the optimized aptamer sequence: After the evolution in (s5), obtain the designed and optimized aptamer sequence; and
[0015] (s7) Decoder output: Use the decoder to output the designed and optimized aptamer sequence obtained in (s6).
[0016] In another preferred example, the different clusters of the nucleic acid sequences refer to the family distribution of the aptamers corresponding to the nucleic acid sequences.
[0017] In another preferred example, the different clusters of the nucleic acid sequences are the aptamer family distributions.
[0018] In another preferred example, the method further includes the following steps:
[0019] (s0) Provide the nucleic acid aptamer sequence information and related experimental data obtained by SELEX technology (including HT-SELEX) screening, where the experimental data includes copy number information; and preprocess the data to obtain a high-throughput screening aptamer sequence data set;
[0020] Among them, the preprocessing includes: removing redundant and measured error sequences from the data.
[0021] In another preferred example, the SELEX technology includes HT-SELEX technology.
[0022] In another preferred example, in step (s1), it includes using the "train_test_split" function in the "sklearn.model_selection" module to randomly divide the aptamer sequence data set into a training set and a test set according to a ratio of 9:1.
[0023] In another preferred example, in step (s2), it includes the following sub-steps:
[0024] (s2a) Convert each character in the nucleic acid sequence in the data set into a corresponding numerical index, that is, map nucleotides A, T / U, G, C to 0, 1, 2, 3 respectively;
[0025] (s2b) Use the "nn.Embedding" module in the PyTorch framework (version 1.5.0) to expand each numerical index into a 32-dimensional vector, so as to obtain the vector representation of each nucleotide; thus obtain the nucleic acid sequence representation composed of vectors; among them, the vector representation of the nucleotides in the nucleic acid sequence in the training set is the first vector, and the vector representation of the nucleotides in the nucleic acid sequence in the test set is the second vector.
[0026] In another preferred example, in step (s3), the data information input into the Gaussian mixture variational autoencoder includes: the number of nucleic acid sequences, the dimension of the nucleic acid sequences, and the length of the nucleic acid sequences.
[0027] In another preferred example, the Gaussian mixture variational autoencoder (GM-VAE) includes an encoder and a decoder.
[0028] In another preferred example, the Gaussian mixture variational autoencoder (GM-VAE) further includes a Gaussian mixture model (GMM).
[0029] In another preferred example, the encoder is a convolutional neural network (CNN) architecture including multiple skip connection layers.
[0030] In another preferred example, in each of the skip connection layers, each layer is composed of a convolutional layer, a batch normalization layer (BatchNormalization Layer), and a Leaky ReLU non-linear activation function.
[0031] In another preferred example, the input latent space sampling vector z first passes through the first fully connected layer FC D,32 (D represents the dimension of z), and sequentially passes through batch normalization and the Leaky ReLU non-linear activation function to obtain the feature vector X1; subsequently, X1 is used as the input to pass through the second fully connected layer FC 22,64 , and also passes through batch normalization and the activation function to obtain the feature vector X2; subsequently, the feature vector X2 passes through the third fully connected layer FC 64,32 , and undergoes a similar process to obtain the feature vector X3. It should be noted that, in order to further fuse and extract the feature information of the sequence, the present invention performs an addition operation on the feature vectors X1 and X3, and passes through the fourth fully connected layer to obtain the feature vector X4. Next, the feature vector X4 依 passes through three transposed convolutional layers for
[0032] In another preferred example, the loss function of the variational autoencoder combined with the Gaussian mixture model of the present invention is:[[]]END]]
[0033]
[0034] In another preferred example, in step (s5), the genetic evolution strategy protects the active sequence design and the active sequence optimization.
[0035] In another preferred example, in step (s5), the genetic evolution strategy includes elitism, crossover, and mutation.
[0036] In another preferred example, in step (s5), the elitism is used for the active sequence design, and the crossover and mutation are used for the active sequence optimization.
[0037] In another preferred example, the elitism includes the following two algorithms: the Gaussian mixture model (GMM) and the K-means clustering algorithm (K-means).
[0038] In another preferred example, the elitism is implemented through two packages, namely sklearn.mixture.GaussianMixture and sklearn.cluster.KMeans in the Scikit-learn library (version 1.0.2).
[0039] In another preferred example, the elitism also introduces the Silhouette Coefficient to evaluate the reasonable number of aptamer families calculated by the two algorithms.
[0040] In another preferred example, the crossover includes the following steps: generating a random number between 0 and 1, and when the value is lower than the preset crossover rate threshold, performing a Uniform Order Crossover (UOX) operation.
[0041] In another preferred example, the crossover further includes the following steps: creating a random boolean mask with the same length as the parent sequence, and based on the mask information, swapping the elements at the corresponding positions between the parent sequences to generate two new offspring sequences.
[0042] In another preferred example, the mutation includes: performing mutation on the selected offspring; replacing the random index of the latent variable z of the offspring with a random normal distribution value.
[0043] In another preferred example, in step (s3), it further includes optimizing the trained encoder, including optimizing the encoder using a standard optimization algorithm (Adam optimizer).
[0044] In the second aspect of the present invention, there is provided a device for aptamer design and optimization, including:
[0045] (a) A data input module, which is configured to input a dataset of aptamer sequences obtained from high-throughput screening, where the dataset includes nucleic acid sequence information and copy number information of the aptamers; and randomly divide the aptamer sequence dataset into a training set and a test set according to a certain ratio;
[0046] (b) A data processing module, which is configured to convert each nucleotide character in the nucleic acid sequence in the dataset into a corresponding numerical index, thereby obtaining the numerical form of each nucleic acid sequence; and expand the numerical index into a first vector, thereby obtaining a nucleic acid sequence representation composed of the first vectors;
[0047] Among them, the vector representation of the nucleotides in the nucleic acid sequence in the training set is the first vector, and the vector representation of the nucleotides in the nucleic acid sequence in the test set is the second vector;
[0048] (c) Analysis module, the analysis module is configured to perform the following operations:
[0049] (i) Encoder-decoder training: Input the nucleic acid sequence representation composed of the first vector into a Gaussian mixture variational autoencoder (GM-VAE) for training, so that the trained encoder can capture the latent structural features and aptamer family distribution in the aptamer nucleic acid sequence, and map these features to a low-dimensional latent space; and enable the decoder to generate the original input sequence; then use the nucleic acid sequence representation composed of the second vector for testing;
[0050] (ii) Data dimensionality reduction and clustering: Use the trained encoder to perform embedded dimensionality reduction on the nucleic acid sequences of the target rounds, so as to map the nucleic acid sequences into a low-dimensional latent space, capture the potential similarities and structural features between nucleic acid sequences; and obtain different clusters (or families) of nucleic acid sequences according to the potential similarities and structural features;
[0051] (iii) Selection of active sequence regions and deep implicit evolution: Select the clusters corresponding to nucleic acid sequences with high copy numbers, and define them as active sequence regions; and perform an implicit space genetic evolution strategy on the active sequence regions;
[0052] (iv) Obtain the optimized aptamer sequence: After the evolution in (iii), obtain the designed and optimized aptamer sequence;
[0053] (d) Output module: The output module is configured to output the designed and optimized aptamer sequence. a
[0054] In another preferred example, in another preferred example, the different clusters of the nucleic acid sequences refer to the family distribution of the aptamers corresponding to the nucleic acid sequences.
[0055] In another preferred example, the different clusters of the nucleic acid sequences are the aptamer family distribution.
[0056] In another preferred example, the device further includes (a0) a data preprocessing module, and the data preprocessing module is configured to: preprocess the nucleic acid aptamer sequence information and related experimental data obtained by SELEX technology screening for a specific target, so as to obtain a high-throughput screening aptamer sequence data set;
[0057] Wherein, the experimental data includes copy number information; the preprocessing includes: removing redundant and measured error sequences in the data.
[0058] In another preferred example, the following process is performed in the data processing module:
[0059] (b1) Convert each character in the nucleic acid sequences in the dataset into a corresponding numerical index, that is, map nucleotides A, T / U, G, and C to 0, 1, 2, and 3 respectively;
[0060] (b2) Use the "nn.Embedding" module in the PyTorch framework (version 1.5.0) to expand each of the numerical indices into a 32-dimensional vector, thereby obtaining the vector representation of each nucleotide; thereby obtaining the nucleic acid sequence representation composed of vectors; wherein, the vector representation of nucleotides in the nucleic acid sequences in the training set is the first vector, and the vector representation of nucleotides in the nucleic acid sequences in the test set is the second vector.
[0061] In another preferred example, the Gaussian mixture variational autoencoder (GM-VAE) includes an encoder and a decoder.
[0062] In another preferred example, the genetic evolutionary strategy includes elitism, crossover, and mutation.
[0063] In another preferred example, the elitism includes the following two algorithms: Gaussian mixture model (GMM) and K-means clustering algorithm (K-means).
[0064] In another preferred example, the elitism is implemented through two packages, sklearn.mixture.GaussianMixture and sklearn.cluster.KMeans in the Scikit-learn library (version 1.0.2).
[0065] In another preferred example, the elitism also introduces the Silhouette Coefficient to evaluate the reasonable number of aptamer families calculated by the two algorithms.
[0066] In another preferred example, the crossover includes the following steps: generate a random number between 0 and 1, and when the value is lower than the preset crossover rate threshold, perform a Uniform Order Crossover (UOX) operation.
[0067] In another preferred example, the crossover further includes the following steps: create a random boolean mask with the same length as the parent sequence, and based on the mask information, exchange the elements at the corresponding positions between the parent sequences, thereby generating two new offspring sequences.
[0068] In another preferred example, the mutation includes: performing mutation on the selected offspring; replacing the random index of the latent variable z of the offspring with a random normal distribution value.
[0069] In another preferred example, during the training of the encoder-decoder, it also includes optimizing the trained encoder, including optimizing the encoder using a standard optimization algorithm (Adam optimizer).
[0070] In the third aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect of the present invention.
[0071] It should be understood that within the scope of the present invention, the above technical features of the present invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be repeated one by one here. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 Shows the overall flowchart of a novel aptamer design method based on activity-guided deep implicit evolution.
[0073] Figure 2 Shows that after training the model with the public database dataset DRA009383 (Dataset 1), the encoder is used to visualize the sequence space, and the distribution of sequence families is visualized based on two elitist algorithms.
[0074] Figure 3 Shows that after training the model with the unpublished dataset (Dataset 2) provided by the collaborator, the encoder is used to visualize the sequence space, and the distribution of sequence families is visualized based on two elitist algorithms.
[0075] Figure 4 Shows the novel aptamer sequences with potential activity and secondary structure information designed by our inventive method after training the model with Dataset 1.
[0076] Figure 5 Shows the experimental verification results of the novel aptamer sequences with potential activity designed by our inventive method after training the model with Dataset 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] The present inventors have conducted extensive and in-depth research and for the first time developed a method for designing and optimizing aptamers based on activity-guided deep implicit evolution. The present invention uses unsupervised machine learning and generative models to map nucleic acid sequences to a low-dimensional latent space, thereby deeply learning the features of existing aptamer sequences and active sequences, and generating new aptamer sequences with high activity and affinity.
[0078] Specifically, the present invention uses the nucleic acid aptamer sequence information and experimental data (including copy number information) of one or more rounds of high-throughput sequencing of SELEX technology (including HT-SELEX) to train a binding Gaussian mixture variational autoencoder (GM-VAE), enabling the encoder to map the nucleic acid aptamer sequence into a low-dimensional latent space and learn the potential similarities and structural features between nucleic acid sequences; and selects the active sequence region for guidance, thereby performing a genetic evolution strategy on the active sequence region, and finally obtaining an optimized brand-new aptamer sequence. On this basis, the present invention is completed.
[0079] Term
[0080] To more easily understand the present disclosure, certain terms are first defined. As used in this application, unless otherwise expressly specified herein, each of the following terms shall have the meaning given below. Other definitions are set forth throughout the application.
[0081] As used herein, the term "comprising" or "including" can be open-ended, semi-closed, and closed. In other words, the term also includes "consisting essentially of" or "consisting of".
[0082] As used herein, unless otherwise specified, any concentration range, percentage range, ratio range, or integer range shall be understood to include any integer value within the range and, where appropriate, fractional values thereof (e.g., one-tenth and one-hundredth of an integer).
[0083] As used herein, the term "and / or" relates to and encompasses any and all possible combinations of one or more of the related listed items.
[0084] As used herein, the term "Gaussian mixture variational autoencoder (GM-VAE)" is an extended model of the variational autoencoder (VAE, Variational Autoencoder), which can better model the multi-modal data distribution by introducing a Gaussian mixture distribution in the latent space.
[0085] As used herein, SELEX (Systematic Evolution of Ligands by EXponential enrichment) technology is a method for obtaining high-affinity nucleic acid aptamers through in vitro screening. HT-SELEX (High-Throughput SELEX) is a high-throughput version of SELEX technology.
[0086] The main advantages of the present invention include:
[0087] (a) The method of the present invention overcomes the deficiencies of the prior art in dealing with complex sequence distributions and the aptamer optimization process.
[0088] (b) The method of the present invention improves the discovery efficiency of aptamers: By combining a deep variational autoencoder (VAE) with a Gaussian mixture model (GMM), the present invention can accurately learn the low-dimensional latent space distribution characteristics of aptamer sequences. This method not only solves the inefficiency problem caused by data complexity and dimensionality issues in traditional methods but also can efficiently mine potential high-activity aptamer sequences.
[0089] (c) The method of the present invention avoids the interference of early invalid sequences: Through the active region guidance and latent space genetic evolution strategy, the present invention avoids the interference of disordered sequences and low-affinity sequences in the early screening, ensuring a more accurate and efficient screening process. This technology can accurately identify active sequences at the early screening stage, thereby improving the accuracy and efficiency of screening.
[0090] (d) The method of the present invention ensures dataset compatibility and universality: The present invention can not only adapt to different sequencing datasets (including single-round or multi-round sequencing data) but also has broad application compatibility. This enables the technology to play a role in a variety of biological research and applications, whether in cancer screening, antibody development, or other fields that require the screening of high-affinity molecules.
[0091] (e) The method of the present invention promotes the application of aptamers in multiple fields: The present invention provides a new and efficient technical means for aptamer design and screening, providing a strong technical support for aptamer research and being able to accelerate the application and development of aptamers in multiple fields such as biomedicine, environmental monitoring, and drug development.
[0092] The following further elaborates the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions noted in the following embodiments are usually carried out under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are by weight.
[0093] Example 1
[0094] 1.1 Provide experimental datasets and data processing
[0095] The first step of the present invention is to provide a dataset of aptamer sequences for high-throughput screening. This dataset is usually the nucleic acid aptamer sequence information and experimental data (such as copy number information) obtained by high-throughput sequencing using the HT-SELEX technique, either single-round or multi-round.
[0096] Contain the sequence information of single-round or multi-round nucleic acid aptamers obtained through HT-SELEX technology and related experimental data (such as copy number information). Specifically, the experimental data set not only includes sequence information, but also contains enrichment information (copy number) in the current sequencing round. This data set provides sufficient input information for the model, making subsequent sequence analysis and optimization more accurate.
[0097] Data processing includes the following steps:
[0098] (1) Data preprocessing: To ensure the accuracy and availability of the data, it is first necessary to preprocess the original data. The preprocessing steps include: removing redundant and mismeasured sequences from the original sequence data. By removing inaccurate sequencing data (such as characters other than ACGU) and duplicate data, the credibility of the screening results is guaranteed.
[0099] (2) Dataset division: In dataset division, the preprocessed sequence data is randomly divided into a training set and a test set in a ratio of 9:1 using the "train_test_split" function in the "sklearn.model_selection" module for model training and testing respectively. The model with the minimum test loss is selected through iteration. The purpose of the training set is to help the model learn the distribution and aggregation characteristics of the sequences with a large amount of data, and the test set with a small amount of data verifies the results of model training and helps adjust the learning method and training parameters of this invention. This data division method ensures that the model can obtain a sufficient sample size during training while reserving a part of the data to verify the performance of the model, thereby improving the generalization ability and prediction accuracy of the model.
[0100] (3) Sequence representation: The nucleic acid sequence is first converted into a One-hot encoding form suitable for machine learning algorithms. First, each character in the nucleotide sequence is converted into a corresponding numerical index. Specifically, nucleotides A, T / U, G, C are mapped to 0, 1, 2, 3 respectively.
[0101] This conversion process converts nucleotide characters into numerical forms, and then through the "nn.Embedding" module in the PyTorch framework (version 1.5.0), the numerical indexes are expanded into 32-dimensional vectors and input into the model for training.
[0102] 1.2 Construct a pre-trained encoder-decoder network architecture
[0103] The core model framework of the present invention consists of an encoder and a decoder of a Convolutional Neural Network (CNN), and an implicit space that captures the distribution characteristics of the aptamer sequence family in single-round or multi-round sequencing data. The three together form a Variational Autoencoder (VAE). Among them, the implicit space establishes a probability model through a parameterized Gaussian distribution, and its dimensional characteristics are learned and generated by the inference network in the encoder, which can effectively capture the evolutionary law of the aptamer sequence family in multi-round SELEX screening. The decoder network parameterizes the latent variable z into the conditional probability distribution of nucleotides through a conditional generation mechanism, and reconstructs the nucleic acid sequence using a transposed convolution operation. Its random sampling process can effectively generate new sequence variants with both family conservatism and reasonable variability. In the vector generation and distribution training stage in the latent space, in order to capture the multi-modal distribution of sequences, the present invention combines a Gaussian Mixture Model (GMM).
[0104] Specifically, the Gaussian Mixture Model enhances the expression ability of the potential structure in sequence data of different rounds by representing the implicit space as a mixture of multiple Gaussian distributions, making the distribution of aptamer sequencing data in the low-dimensional space more intuitive and clear. The construction and implementation of this model can effectively capture the complex structures and features in the sequences.
[0105] The specific construction process is as follows:
[0106] (1) Encoder network structure: The encoder network is a Convolutional Neural Network (CNN) architecture that includes multiple Skip Connection Layers. First, the unique encoding form of each nucleic acid sequence is embedded as a 32-dimensional vector, and these vectors serve as the input to the model. Subsequently, the input data passes through up to 6 skip connection layers to learn the structure and interaction information of the sequences.
[0107] Specifically, in the design of the skip connection layer, each layer consists of a convolutional layer, a Batch Normalization Layer, and a Leaky ReLU non-linear activation function.
[0108] First, the input latent space sampling vector z first passes through the first fully connected layer FC D,32 (D represents the dimension of z), and sequentially passes through batch normalization and the Leaky ReLU non-linear activation function to obtain the feature vector X1; subsequently, X1 serves as the input and passes through the second fully connected layer FC 32,64, also through batch normalization and activation function, the feature vector X2 is obtained; subsequently, the feature vector X2 passes through the third fully connected layer PC 64,32 , and similar processing is performed to obtain the feature vector X3. It should be noted that, in order to further fuse and extract the feature information of the sequence, the present invention performs an addition operation on the feature vectors X1 and X3, and through the fourth fully connected layer, the feature vector X4 is obtained. Next, the feature vector X4 passes through three transposed convolutional layers in sequence, and each layer includes batch normalization and Leaky ReLU activation function processing, and finally the output is obtained.
[0109] (2) Decoder network structure: The decoder network is a multi-layer structure composed of fully connected layers and transposed convolutional layers, and the final output is the nucleic acid probability of each position in the sequence in this round of SELEX sequencing, that is, a categorical distribution is assigned to each position of the sequence.
[0110] (3) Introduction of Gaussian mixture model: A binary vector c ∈ {0, 1} k×1 is introduced to represent which Gaussian component the latent variable z belongs to. Finally, the present invention introduces a new distribution to approximate the posterior distribution p θ (z, c|x). Where c is the categorical variable of the aptamer family to which each nucleic acid sequence belongs, and its probability is the discrete distribution p(c|φ). Where and z is the latent variable.
[0111] Therefore, the overall deep variational autoencoder is specifically manifested as an encoder with skip connections that converts the sequence X into a latent distribution in the implicit space where the visualization of z is to simulate the distribution of the high-dimensional original dataset p(x). The decoder network architecture reconstructs z in the latent space as the original input data by learning p e (x, c|z). p(z|c) is a Gaussian mixture distribution parameterized by μ c and σ z under the condition of category c.
[0112] Assuming that x and c are conditional on z, the joint probability p(x, z, c) can be decomposed into:
[0113] p(x, z, c) = p(x|z)p(z|c)p(c) (Equation 1)
[0114] Among them, each probability model is defined respectively as:
[0115]
[0116] p(x|z) = Ber(x|μ x ) (Equation 4)
[0117] In addition, the present invention assumes that can be decomposed into:
[0118]
[0119] Then the present invention defines the following formula to describe the specific probability distribution:
[0120]
[0121] Finally, the loss function of the variational autoencoder combined with the Gaussian mixture model in the present invention is:
[0122]
[0123] It can not only effectively learn the distribution of aptamer sequencing data in different rounds, but also maximize the log-likelihood of the current round of sequencing data and ensure that the reconstructed estimated data is similar to the input nucleic acid sequence data. This model can adaptively adjust the weights to process sequencing data with different aggregation distribution characteristics, making the distribution of aptamer sequencing data in different rounds clearer, thereby improving the generalization ability and prediction performance of the model. At the same time, using the Kullback-Leibler divergence (KL divergence) for regularization helps prevent the model from overfitting and ensures a reasonable distribution of latent variables.
[0124] (4) Training and fine-tuning strategies of the model: The training and fine-tuning strategies of the method of the present invention adopt the strategy of training round by round and fine-tuning with data of the target round. This strategy aims to improve the encoding and decoding capabilities of the model for nucleic acid sequences, imitate the process of SELEX screening to gradually learn the data distribution in each round, and optimize the performance of the model under a specific data distribution in the fine-tuning stage. Specifically, this strategy includes the following two main steps:
[0125] 1) Pre-training with a dataset increasing round by round: First, train the model with a dataset increasing by round (such as Round 1, Round 2, Round 3). The purpose of this process is to enable the model to gradually learn and adapt to the distribution characteristics of different datasets, and improve the encoding and decoding capabilities of the model for nucleic acid sequences. This helps the model further improve its generalization ability and in-depth understanding of nucleic acid sequences.
[0126] 2) Fine-tuning with a small sample of high-copy data of the target round: In the fine-tuning stage, use the high-copy sequence dataset of the target round to be studied (such as Round 3) (usually the first 5000 - 10000 sequences) to further fine-tune the model. Enable the model to better adapt to and generalize the distribution of this specific round of dataset. The fine-tuning process optimizes the model's response ability to the target data by adjusting the model parameters, while ensuring that the model can have the best performance when dealing with the data of the target round to be studied.
[0127] 1.3 Activity-guided aptamer discovery strategy
[0128] Since amplification bias and non-specific adsorption may lead to the inclusion of sequences unrelated to the target molecule in the later enriched sequence data, and these unrelated sequences may form different clusters in the sequence distribution just like the active aptamer family. Therefore, the core hypothesis of this strategy of the present invention is that the sequence aggregation region mapped by the high-copy data represents the active sequence region. Among these active sequence regions, there are sequences with high activity and high fluorescence activation multiples.
[0129] In the present invention, first, the aptamer sequence data is trained by a Gaussian mixture variational autoencoder (GM-VAE) to learn the low-dimensional latent space representation of the sequences. The trained encoder can effectively capture the potential structural features in the aptamer sequences and map these features to a low-dimensional latent space.
[0130] Subsequently, the active sequence region is defined by mapping the "aptamer family" where the high-copy sequences are located, so as to avoid the interference of false-positive sequences on the screening as much as possible. Subsequently, genetic evolution research is focused on the aptamer families in the active sequence region to design aptamers with high activity and high affinity.
[0131] The core advantage of this strategy is that through the guidance of activity data, intelligent screening and optimization can be carried out during the initial discovery process of aptamers. Compared with the traditional screening methods that rely on single copy number and sequence matching, the activity-guided strategy can significantly improve the experimental success rate and practical application value of aptamers, ensuring that the aptamers obtained in practical applications have better performance and reliability.
[0132] Specifically, it includes the following steps:
[0133] (1) Sequence space mapping: First, the sequences of about 5000-10000 in the SELEX sequencing are mapped into the latent space by the trained encoder.
[0134] (2) Aptamer family classification: Based on the Gaussian mixture model (GMM) and the K-means clustering algorithm (K-means), the sequences in the latent space are divided into different aptamer families.
[0135] (3) Definition of active sequence region: According to the high-copy number sequences and their corresponding aptamer family distributions, the active sequence region is defined.
[0136] (4) Deep implicit evolution of the active sequence region: Genetic evolution research is carried out on the aptamer families in the active sequence region, and the sequence design is optimized through operations such as crossover and mutation to further improve the activity and affinity of the aptamers.
[0137] 1.4 Deep implicit evolution of the active sequence region
[0138] Based on the active region-guided aptamer discovery strategy, after determining the active sequence regions for different rounds of data, the present invention performs an implicit spatial genetic evolution strategy on these active regions.
[0139] This strategy mainly involves three Darwinian evolutionary operations: elitism, crossover and mutation, aiming to efficiently screen out high-quality aptamer sequences with high activity and high fluorescence activation multiples through deep evolution strategies.
[0140] The specific steps include:
[0141] (1) The first evolutionary operation is elitism. It is implemented by the sklearn.mixture.GaussianMixture and sklearn.cluster.KMeans packages in the Scikit-learn library (version 1.0.2). The relevant parameters are shown in Table 1 below:
[0142] Table 1
[0143]
[0144] And introduce a parameter for rationally evaluating the clustering effect - Silhouette Coefficient to evaluate the reasonable number of aptamer families calculated by the two clustering algorithms. The principle of the silhouette coefficient is to calculate the average intra-cluster distance (a) and the average nearest cluster distance (b) of each nucleic acid sequence, measure the closeness of the sequence to its cluster and the distance from other clusters, so as to quantify the pros and cons of the clustering effect. The cluster represents the distribution of aptamer families calculated and predicted by the present invention. The average intra-cluster distance (a) evaluates the average distance between the nucleic acid sequence and other sequences in the same cluster, which measures the closeness of the nucleic acid sequence to the sequence of the cluster to which it belongs. The average nearest cluster distance (b) evaluates the average distance between the nucleic acid sequence data and all nucleic acid sequences of the closest different clusters. The final comprehensive evaluation, its range is between [-1,1]. The larger the value, the better the current clustering effect and the more reasonable the distribution of aptamer families.
[0145]
[0146] Subsequently, based on two clustering algorithms, namely the Gaussian Mixture Model (GMM) and the K-Means clustering algorithm (K-means), the clustering parameters "n_components" and "n_clusters" were calculated in a loop for values ranging from 5 to 15, i.e., each method was run 11 times. Additionally, to avoid possible errors when fitting the model, each evaluation was repeated 200 times and the optimal clustering model was retained, and the average silhouette_score calculated by the optimal model at the current number of clusters was output to rationally evaluate the classification results of the "aptamer family" under the current clustering parameters.
[0147] Finally, based on the silhouette_scores calculated from running 11 times, the optimal model was determined by scoring according to the maximum silhouette_score, and the optimal model was selected to predict the distribution of the aptamer family and the elite sequences of each aptamer family. These elite sequences will play a key role in subsequent genetic evolution. Through genetic operations such as crossover and mutation, potential highly active aptamers in a larger sequence space are explored. The specific implementation is achieved through sklearn.metrics.silhouette_score in the Scikit-learn library (version 1.0.2).
[0148] (2) Uniform order crossover operation determined by the crossover rate: Generate a random number between 0 and 1. When this value is lower than the preset crossover rate threshold, perform the Uniform Order Crossover (UOX) operation. The UOX operation ensures that gene segments of the parental sequences are exchanged in a uniform random manner, retaining the characteristics of the parental sequences while introducing new mutations.
[0149] Random boolean mask-assisted sequence recombination: First, create a random boolean mask of the same length as the parental sequences. Subsequently, based on the mask information, elements at corresponding positions are exchanged between the parental sequences, thereby generating two new offspring sequences. This process increases the diversity of the aptamer sequences. This crossover fusion strategy enables the transfer of the characteristics of excellent parents to the offspring and increases the diversity of the aptamer sequences, thus accelerating the evolutionary process of the entire aptamer family.
[0150] (3) Set the probability of mutation and perform mutation on the selected offspring; randomly index the latent variable z of the offspring and replace it with a random normal distribution value. The mutation rate should be very small, considering that a larger mutation rate will lead to a random search by the algorithm, thus disrupting the overall genetic evolution. Therefore, the mutation rate is set very small.
[0151] 1.5 Model Training and Optimization
[0152] The present invention uses a standard optimization algorithm (such as the Adam optimizer) to train the model, and selects a loss function suitable for the objectives of the present invention (as shown in Equation 8) as the objective function. The training process adopts common learning rate adjustment strategies and introduces an early stopping mechanism to prevent overfitting.
[0153] In the pre-training stage, the maximum number of iterations is 2000, and a suitable learning rate adjustment method is used. In the fine-tuning stage, the learning rate is further reduced, and a more stringent early stopping criterion is adopted. The training terminates when the loss of the model does not change significantly for 50 consecutive epochs. This method effectively avoids overfitting during the training process and improves the generalization ability of the model.
[0154] Specifically, it includes the following steps:
[0155] (1) Taking the log evidence lower bound (ELBO) loss as the optimization objective, which consists of a reconstruction term and a regularization term. The reconstruction term uses the cross-entropy loss function to measure the ability of the model to reconstruct the original input data, and the regularization term regularizes the latent variables into the GMM manifold through the Kullback-Leibler divergence (KL divergence).
[0156] In the pre-training stage, the maximum number of iterations is set to 2000, the Adam optimizer is used, the learning rate is 1e-3, the training batch size is 512, and an early termination mechanism is introduced (stop if there is no obvious improvement in the performance of the validation set for 10 consecutive epochs); in the fine-tuning stage, the training batch size is still 512, the learning rate is lowered to 1e-4, and the early termination criterion is adjusted to no significant improvement in the test performance for 15 consecutive epochs. The training and fine-tuning are implemented based on the GPU version of the PyTorch framework and the Scikit-learn library.
[0157] Example 2
[0158] The aptamer design method based on activity-guided deep implicit evolution proposed by the present invention aims to improve the efficiency and accuracy of aptamer screening through a unique combination of deep learning and genetic algorithms. This method uses a pre-training strategy to learn the distribution of sequences in the low-dimensional latent space and generates efficient aptamer sequences through genetic evolution operations. The implementation process is as Figure 1 shown. To verify the effectiveness and feasibility of the method of the present invention, the present invention uses the dataset DRA009383 in the public database DDBJ for experiments, demonstrating the application effect of this method in real data.
[0159] The specific implementation steps are as follows:
[0160] (1) Data preparation: Two datasets were used in this case study. Dataset 1 selected the dataset DRA009383 from the public database DDBJ; Dataset 2 selected the internal unpublished SELEX high-throughput sequencing dataset provided by the cooperation group as the data for validating the method of the present invention. The dataset contains the high-throughput screening and sequencing data of a certain round during the aptamer SELEX screening, and only contains sequence information.
[0161] Among them, Dataset 1 is the sequence sequencing data targeting human transglutaminase 2, with a total of 80,846 sequences, saved in the FASTQ format. Dataset 2 is the sequence sequencing data targeting a certain fluorophore (unpublished), with a total of 1,048,576 sequences, saved in the FASTA format.
[0162] (2) Data preprocessing and feature extraction: Based on the overall processing flow of the method of the present invention, first, the sequences in the dataset were automatically preprocessed, duplicate sequences were removed, inaccurate sequencing sequence information was removed, and the relevant features of each sequence were extracted. After processing, there were 38,513 sequences in Dataset 1 and 485,735 sequences in Dataset 2. Subsequently, the data was input into the Gaussian mixture-based deep variational autoencoder (GM-VAE) model for training, and the trained encoder was used to perform embedded dimensionality reduction on these sequences, mapping the sequences into a low-dimensional latent space to capture the potential similarities and structural features between the sequences.
[0163] (3) Genetic algorithm evolution operation: Using the latent space genetic evolution strategy based on the Gaussian mixture model, three-step genetic evolution was performed on the sequence spaces of Dataset 1 ( Figure 2 ) and Dataset 2 ( Figure 3 ): elitism, crossover fusion, and mutation. And the active region can be further determined according to the experimental results, so as to perform the genetic evolution of the active sequence region. Among them, two clustering algorithms were used to achieve the division of sequence families and the mining of central sequences in elitism.
[0164] (4) Generate highly efficient aptamer sequences: After multiple rounds of evolution, finally, potential new aptamer sequences with high activity and rich structural diversity were generated.
[0165] Experimental results and analysis:
[0166] The aptamer sequences designed by the method of the present invention were compared with traditional methods, and the specific experimental results are as follows:
[0167] Dataset 1: When using the public dataset DRA009383 for validation, the method of the present invention successfully designed 12 aptamer sequences with potential high activity and performed secondary structure visualization ( Figure 4), and they all represent the active sequences in each aptamer family. In addition, sequence similarity calculation of the generated sequences shows that all the generated sequences are completely new sequences and do not appear in the training set. This indicates that the method of the present invention can design potentially highly active new aptamer sequences.
[0168] Dataset 2: When using the unpublished dataset provided by the collaborating research group for verification, the method of the present invention successfully designed 30 potentially highly active aptamer sequences. Experimental verification showed that the fluorescence activation multiples of 27 sequences were about 400 times that of the control sequence, and the fluorescence activation multiples of some sequences reached about 600 times. This indicates that 86.7% of the sequences designed by the present invention have high fluorescence activation multiples through experimental verification ( Figure 5 ). In addition, sequence similarity calculation of the generated sequences shows that all the generated sequences are completely new sequences and do not appear in the training set. This indicates that the method of the present invention can design potentially highly active new aptamer sequences.
[0169] Discussion
[0170] The present invention proposes a completely new aptamer design method - activity-guided deep implicit evolution. The invention provides a solution to the above problems. By learning the low-dimensional latent space distribution of sequences through a pre-training strategy and combining genetic algorithm evolution operations, a solution that can efficiently capture the complex distribution characteristics of aptamer sequencing data is provided. On the one hand, through a unique deep variational autoencoder and Gaussian mixture model, the complex distribution characteristics of aptamer sequencing data are accurately captured; on the other hand, by using the active region guidance and latent space genetic evolution strategy, the mining of highly active aptamers is focused. Finally, the efficient design of nucleic acid sequences with high activity and high fluorescence activation multiples is realized, the aptamer discovery efficiency is improved, and it shows excellent performance in dataset compatibility and design universality, effectively filling the deficiencies of traditional methods.
[0171] The aptamer design method of the present invention overcomes the limitations of traditional aptamer design methods by combining Gaussian mixture variational autoencoder (GM-VAE) with genetic algorithm. This method improves the screening efficiency and accuracy of aptamers by learning the low-dimensional latent space representation of aptamer sequences. The core innovation of this technical solution lies in learning the latent features of aptamers in the latent space through the autoencoder and combining genetic algorithm for optimization, thereby generating aptamer sequences with high affinity.
[0172] In the present invention, the latent variables in the latent space of high-copy sequences are obtained through the trained encoder, and the individuals of the initial population suitable for genetic evolution research are constructed. The innovation lies in encoding and learning the latent space to ensure that the genetic algorithm can generate a higher-quality initial population based on the actual screening data, thereby improving the efficiency of the subsequent evolution process. This method of constructing the initial population based on the latent space effectively avoids the deficiency of the traditional method that relies on high-copy number screening.
[0173] The present invention proposes an activity-guided aptamer discovery strategy, which optimizes the aptamer screening process by combining the experimental activity data of aptamers (such as affinity, fluorescence signal, etc.). This method not only relies on sequence information but also integrates the experimental data of aptamers, making the finally screened aptamers have higher affinity and experimental verification success rate. The innovation lies in introducing the activity data into the early stage of aptamer design, greatly improving the screening efficiency and accuracy.
[0174] The present invention innovatively applies the genetic algorithm to the optimization process of aptamers, and uses operations such as selection, crossover, and mutation to evolve the initial population. These evolutionary operations can perform global search in the low-dimensional latent space to discover potential aptamer sequences with high affinity. The genetic algorithm has significant advantages in optimizing the specificity and affinity of aptamers, especially when facing complex multi-round screening data, it can effectively improve the recognition and optimization efficiency of aptamers.
[0175] The technical solution proposed by the present invention combines high-throughput screening data and deep learning technology, and uses a machine learning model to optimize the aptamer sequence. Especially in high-throughput screening, by using a deep learning-based model to predict and optimize the aptamer sequence, this method improves the efficiency of aptamer discovery and reduces the time and cost required for experimental screening. The innovation lies in the combination of deep learning and high-throughput data, enhancing the processing ability of complex data and the prediction accuracy of aptamers.
[0176] The present invention learns the latent feature space of aptamer sequences through a pre-trained Gaussian mixture variational autoencoder (GM-VAE), and uses this model to generate candidate sequences of aptamers. This method innovatively uses deep learning technology to model and predict the latent space of aptamers, optimizing the aptamer generation process based on manual design or random screening in the traditional method. The key to this technology lies in effectively capturing the structural information of the sequence in a lower-dimensional space through the pre-trained model, thereby providing more accurate and efficient input data for the genetic algorithm.
[0177] All documents mentioned in this invention are cited herein by reference as if each individual document was cited by reference. In addition, it should be understood that after reading the above teachings of this invention, those skilled in the art can make various changes or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
Claims
1. An aptamer design and optimization method based on activity-guided deep implicit evolution, characterized in that: The following steps are involved: (s1) providing an aptamer sequence data set for high-throughput screening, wherein the data set includes nucleic acid sequence information and copy number information of the aptamer; and randomly dividing the aptamer sequence data set into a training set and a test set according to a certain ratio; (s2) data processing: converting each nucleotide character in the nucleic acid sequence in the data set into a corresponding numerical index, thereby obtaining a numerical form of each nucleic acid sequence; and expanding the numerical index into a vector, thereby obtaining a nucleic acid sequence representation composed of the vector; The vector of nucleotides in the nucleic acid sequence in the training set is represented as a first vector, and the vector of nucleotides in the nucleic acid sequence in the test set is represented as a second vector; (s3) Encoder-decoder training: The nucleic acid sequence representation formed by the first vector is input into a Gaussian mixture variational autoencoder (GM-VAE) for training, so that the trained encoder can capture the potential structural features and aptamer family distribution in the aptamer nucleic acid sequence and map these features to a low-dimensional latent space; and the decoder generates the original input sequence; and then the nucleic acid sequence representation formed by the second vector is used for testing; (s4) Data dimensionality reduction and clustering: using the encoder trained in (s3) to perform embedded dimensionality reduction on the target number of rounds of nucleic acid sequences, thereby mapping the nucleic acid sequences into a low-dimensional latent space to capture the potential similarities and structural features between nucleic acid sequences; and obtaining different clusters (or families) of nucleic acid sequences based on the potential similarities and structural features; (s5) Selection of active sequence regions and deep implicit evolution: Select clusters corresponding to nucleic acid sequences with high copy numbers and define them as active sequence regions; and perform latent space genetic evolution strategy on the active sequence regions; (s6) obtaining an optimized aptamer sequence: after the evolution in (s5), obtaining a designed and optimized aptamer sequence; and (s7) Decoder output: using the designed and optimized aptamer sequence obtained in the decoder output (s6).
2. The method according to claim 1, characterized in that The method further comprises the following steps: (s0) providing single-round and / or multi-round nucleic acid aptamer sequence information and related experimental data obtained by SELEX technology screening, wherein the experimental data includes copy number information; and preprocessing the data to obtain a high-throughput screening aptamer sequence data set; Wherein, the preprocessing includes: removing redundancy in the data and determining erroneous sequences.
3. The method according to claim 1, characterized in that In step (s1), the aptamer sequence dataset is randomly divided into a training set and a test set in a ratio of 9:1 using the "train_test_split" function in the "sklearn.model_selection" module.
4. The method according to claim 1, characterized in that In step (s2), the following sub-steps are included: (s2a) converting each character in the nucleic acid sequence in the data set into a corresponding numerical index, that is, mapping nucleotides A, T / U, G, and C to 0, 1, 2, and 3, respectively; (s2b) Using the "nn.Embedding" module in the PyTorch framework (version 1.5.0), each of the numerical indexes is expanded into a 32-dimensional vector to obtain a vector representation of each nucleotide; thereby obtaining a nucleic acid sequence representation composed of a vector; wherein the vector representation of the nucleotides in the nucleic acid sequence in the training set is a first vector, and the vector representation of the nucleotides in the nucleic acid sequence in the test set is a second vector.
5. The method according to claim 1, characterized in that The Gaussian mixture variational autoencoder (GM-VAE) includes an encoder and a decoder.
6. The method according to claim 1, characterized in that In step (s5), the genetic evolution strategy includes elitism, crossover and mutation.
7. The method according to claim 1, characterized in that The elitism includes the following two algorithms: Gaussian mixture model (GMM) and K-means clustering algorithm (K-means).
8. A device for aptamer design and optimization, characterized in that: include: (a) a data input module, wherein the data input module is configured to input a high-throughput screened aptamer sequence data set, wherein the data set includes aptamer nucleic acid sequence information and copy number information; and randomly divide the aptamer sequence data set into a training set and a test set according to a certain ratio; (b) a data processing module, the data processing module being configured to convert each nucleotide character in the nucleic acid sequence in the data set into a corresponding numerical index, thereby obtaining a numerical form of each nucleic acid sequence; and expand the numerical index into a first vector, thereby obtaining a nucleic acid sequence representation composed of the first vector; The vector of nucleotides in the nucleic acid sequence in the training set is represented as a first vector, and the vector of nucleotides in the nucleic acid sequence in the test set is represented as a second vector; (c) an analysis module, the analysis module being configured to perform the following operations: (i) Encoder-decoder training: The nucleic acid sequence representation composed of the first vector is input into a Gaussian mixture variational autoencoder (GM-VAE) for training, so that the trained encoder can capture the potential structural features and aptamer family distribution in the aptamer nucleic acid sequence and map these features to a low-dimensional latent space; and the decoder generates the original input sequence; and then the nucleic acid sequence representation composed of the second vector is used for testing; (ii) Data dimensionality reduction and clustering: using the trained encoder to perform embedded dimensionality reduction on the target number of rounds of nucleic acid sequences, thereby mapping the nucleic acid sequences into a low-dimensional latent space to capture the potential similarities and structural features between nucleic acid sequences; and obtaining different clusters (or families) of nucleic acid sequences based on the potential similarities and structural features; (iii) Selection of active sequence regions and deep implicit evolution: clusters corresponding to nucleic acid sequences with high copy numbers are selected and defined as active sequence regions; and a latent space genetic evolution strategy is performed on the active sequence regions; (iv) obtaining an optimized aptamer sequence: after the evolution in (iii), obtaining a designed and optimized aptamer sequence; (d) Output module: The output module is configured to output the designed and optimized aptamer sequence.
9. The device according to claim 8, characterized in that The device further comprises (a0) a data preprocessing module, wherein the data preprocessing module is configured to: for a specific target, preprocess the single-round and / or multiple-round nucleic acid aptamer sequence information and related experimental data obtained by SELEX technology screening, thereby obtaining a high-throughput screening aptamer sequence data set; Wherein, the experimental data includes copy number information; the preprocessing includes: removing redundancy in the data and removing sequences with measurement errors.
10. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the method according to claim 1 is implemented.
Citation Information
Patent Citations
Aptamer generation method based on conditional discrete diffusion model
CN116631499A
Aptamer candidate library generation method, system and medium
CN118116468A
Candidate nucleic acid aptamer generation method based on Poisson flow condition generation model
CN119108019A
Facilitation of aptamer sequence design using encoding efficiency to guide choice of generative models
US20240087682A1