Nucleic acid aptamer generation and screening method and device, electronic equipment and program product
By using a trained model generation and screening system, and leveraging the three-dimensional structural information of the target protein and reference nucleic acid aptamers, high-affinity nucleic acid aptamers can be generated and screened efficiently. This solves the problems of long cycle time, high cost and low efficiency in existing technologies, and improves the binding specificity and efficiency of nucleic acid aptamers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN UNIV
- Filing Date
- 2025-12-23
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies are time-consuming, complex, costly, and inefficient in nucleic acid aptamer screening, making it difficult to cover the vast nucleic acid sequence space and resulting in the omission of high-affinity sequences.
By using a trained protein structure encoder, conditional diffusion model, nucleic acid encoder, and semantic space alignment module, high-affinity nucleic acid aptamers are generated and screened. The three-dimensional structural information of the target protein and the reference nucleic acid aptamer are used as guiding conditions, combined with the guiding strength coefficient, to conduct targeted exploration and optimization, thereby achieving efficient generation and screening.
This method enables efficient generation and screening of high-affinity nucleic acid aptamers, reducing the blind spots and workload of experimental screening, and improving the binding specificity and efficiency of the generated nucleic acid aptamers.
Smart Images

Figure CN121963853A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of deep learning technology, and in particular relates to a method, apparatus, electronic device and program product for generating and screening nucleic acid aptamers. Background Technology
[0002] Nucleic acids are important macromolecules in living organisms that carry and transmit genetic information. They mainly include deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). RNA is a single-stranded nucleic acid molecule with ribose as its backbone, and its bases are composed of adenine (A), cytosine (C), guanine (G), and uracil (U). RNA molecules can form stable secondary and tertiary structures through intramolecular base pairing, thereby acquiring specific spatial conformations and molecular recognition capabilities.
[0003] Nucleic acid aptamers are oligonucleotide molecules composed of single-stranded DNA or RNA that can recognize target molecules, including proteins, small molecules, metal ions, and cells, with high affinity and specificity. These molecules are typically obtained through in vitro screening techniques—Systematic Evolution of Ligands by Exponential Enrichment (SELEX). Because nucleic acid aptamers possess molecular recognition capabilities similar to antibodies, controllable chemical synthesis, and excellent physicochemical stability, they are known as "chemical antibodies" and show broad application prospects in fields such as biological detection, disease diagnosis, and targeted therapy.
[0004] Traditional nucleic acid aptamer screening mainly relies on SELEX technology, which gradually enriches sequences with high affinity for the target through multiple rounds of in vitro screening and amplification. However, this process usually requires more than ten rounds of screening and sequencing, which is time-consuming, complex, and costly. Furthermore, the success rate of screening is affected by various factors such as experimental conditions, sample diversity, and enrichment efficiency, resulting in low overall efficiency.
[0005] While the introduction of high-throughput sequencing technology has improved the efficiency of sequence data acquisition to some extent, its capacity still cannot cover the theoretically vast nucleic acid sequence space. As sequence length increases, the number of possible sequence combinations grows exponentially, limiting experimental screening to an extremely limited sequence subspace, leading to the omission of potentially high-affinity sequences. Therefore, how to achieve efficient generation and screening of high-affinity nucleic acid aptamers is a pressing technical problem that needs to be solved. Summary of the Invention
[0006] This application provides a method, apparatus, electronic device, and program product for generating and screening nucleic acid aptamers, which can achieve efficient generation and screening of high-affinity nucleic acid aptamers.
[0007] In a first aspect, embodiments of this application provide a method for generating and screening nucleic acid aptamers, including: The target protein is input into a trained protein structure encoder to extract the first protein feature of the target protein; The guiding conditions and guiding strength coefficient are input into the trained conditional diffusion model to generate at least one candidate nucleic acid aptamer; the guiding conditions include the first protein feature and the reference nucleic acid aptamer, and the guiding strength coefficient is used to control the degree to which the generation process of the trained conditional diffusion model depends on the guiding conditions. Each candidate nucleic acid aptamer is input into the trained nucleic acid encoder to extract the first nucleic acid sequence features of each candidate nucleic acid aptamer; The first nucleic acid sequence features and the first protein features of each candidate nucleic acid aptamer are input into a trained semantic space alignment module to calculate the prediction score of each candidate nucleic acid aptamer; the prediction score of a candidate nucleic acid aptamer characterizes the degree of fit between the candidate nucleic acid aptamer and the target protein in terms of binding properties. Each candidate nucleic acid aptamer and its predicted score are added to the nucleic acid candidate pool; The candidate nucleic acid aptamer with the highest predicted score in the nucleic acid candidate pool is determined as the nucleic acid aptamer with the best binding ability to the target protein.
[0008] In this embodiment, by inputting the target protein into a trained protein structure encoder, the first protein feature of the target protein can be extracted. By inputting the guidance conditions and guidance strength coefficient into a trained conditional diffusion model, at least one candidate nucleic acid aptamer can be generated. By inputting each candidate nucleic acid aptamer into a trained nucleic acid encoder, the first nucleic acid sequence feature of each candidate nucleic acid aptamer can be extracted. By inputting the first nucleic acid sequence feature and the first protein feature of each candidate nucleic acid aptamer into a trained semantic space alignment module, the prediction score of each candidate nucleic acid aptamer can be calculated. This prediction score represents the degree of fit between the corresponding candidate nucleic acid aptamer and the target protein in terms of binding characteristics. By adding each candidate nucleic acid aptamer and its prediction score to the nucleic acid candidate pool, the candidate nucleic acid aptamer with the highest prediction score in the nucleic acid candidate pool can be determined as the nucleic acid aptamer with the best binding ability to the target protein, thereby obtaining a high-affinity nucleic acid aptamer. This scheme uses the first protein feature containing the three-dimensional structural information of the target protein and the current optimal reference nucleic acid aptamer as guiding conditions, and combines the guiding strength coefficient with a trained conditional diffusion model. This allows for deep integration of the protein's structural context during the generation process, and intelligent targeted exploration and optimization in the neighborhood space of high-quality sequences. This results in the batch generation of candidate nucleic acid aptamers with potential high binding affinity. Then, a semantic space alignment model is used to quantitatively predict and score each candidate nucleic acid aptamer, and the candidate nucleic acid aptamers in the candidate nucleic acid pool are screened based on the prediction scores. This enables effective evaluation and screening of each generated candidate nucleic acid aptamer, and the selection of high-affinity nucleic acid aptamers. Overall, this scheme achieves efficient generation and screening of high-affinity nucleic acid aptamers.
[0009] Secondly, embodiments of this application provide a nucleic acid aptamer generation and screening apparatus, comprising: The first extraction module is used to input the target protein into a trained protein structure encoder to extract the first protein feature of the target protein. A candidate generation module is used to input guiding conditions and guiding strength coefficients into a trained conditional diffusion model to generate at least one candidate nucleic acid aptamer; the guiding conditions include the first protein feature and a reference nucleic acid aptamer, and the guiding strength coefficient is used to control the degree of dependence of the generation process of the trained conditional diffusion model on the guiding conditions; The second extraction module is used to input each candidate nucleic acid aptamer into the trained nucleic acid encoder to extract the first nucleic acid sequence features of each candidate nucleic acid aptamer. The scoring calculation module is used to input the first nucleic acid sequence features and the first protein features of each candidate nucleic acid aptamer into a trained semantic space alignment module to calculate the predicted score of each candidate nucleic acid aptamer; the predicted score of a candidate nucleic acid aptamer characterizes the degree of fit between the candidate nucleic acid aptamer and the target protein in terms of binding characteristics. The candidate addition module is used to add each candidate nucleic acid aptamer and its predicted score to the nucleic acid candidate pool. The target determination module is used to identify the candidate nucleic acid aptamer with the highest predicted score in the nucleic acid candidate pool as the nucleic acid aptamer with the best binding ability to the target protein.
[0010] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device enables the nucleic acid aptamer generation and screening method as described in the first aspect above.
[0011] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a computer, implements the nucleic acid aptamer generation and screening method described in the first aspect above.
[0012] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when run, causes the nucleic acid aptamer generation and screening method described in the first aspect above to be executed.
[0013] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating the nucleic acid aptamer generation and screening method provided in the embodiments of this application; Figure 2 This is another flowchart illustrating the nucleic acid aptamer generation and screening method provided in the embodiments of this application; Figure 3 This is an example diagram of semantic space alignment of protein and nucleic acid sequences provided in the embodiments of this application; Figure 4 This is an example diagram of the pre-training stage based on various non-coding ribonucleic acid sequences provided in the embodiments of this application; Figure 5 This is an example diagram of the protein-guided fine-tuning stage provided in the embodiments of this application; Figure 6 This is an example diagram of the nucleic acid aptamer iterative screening stage provided in the embodiments of this application; Figure 7 This is a result experimental diagram provided in an embodiment of this application; Figure 8 This is another experimental result diagram provided in the embodiments of this application; Figure 9 This is another experimental result diagram provided in the embodiments of this application; Figure 10 This is a schematic diagram of the structure of the nucleic acid aptamer generation and screening device provided in the embodiments of this application; Figure 11 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0016] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0017] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0018] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0019] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0020] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0021] In recent years, the rapid development of artificial intelligence (AI) technology has provided new ideas for the design of nucleic acid aptamers. Generative models based on deep learning (such as Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and Autoregressive Language Modes) have shown great potential in biological sequence modeling. These models can automatically learn the implicit rules between sequence and function from existing aptamer sequence data and generate new sequences with potential binding activity in a high-dimensional latent space. This significantly expands the aptamer design space, reduces the blind spots and workload of experimental screening, and provides a new computational strategy for efficiently obtaining high-affinity aptamers.
[0022] In 2022, Natsuki Iwano et al. proposed the RaptGen model, which is based on VAEs and Hidden Markov Models (HMMs) to generate candidate aptamers by exploring the diversity of nucleic acid sequences in the latent space. However, this method relies on SELEX screening data, and the model's generality and scalability are limited, making it difficult to extend to the design of aptamers for other targets.
[0023] Subsequently, Incheol Shin et al. proposed the AptaTrans method, a model capable of predicting the binding activity between nucleic acid aptamers and protein sequences, and further using the prediction results to guide the AptaMCTS model for sequence design. The interaction between nucleic acid aptamers and target proteins is primarily determined by their three-dimensional conformation. The RnaFlow model proposed by Divya Nori et al. incorporates three-dimensional protein structural information during the generation process, thereby improving the structural rationality and biological relevance of RNA sequence design.
[0024] In summary, existing deep generative models in the field of nucleic acid aptamer design mainly present two approaches: first, aptamer design based on protein sequence information, but lacking sufficient modeling of the three-dimensional structure of proteins and their binding site characteristics; second, sequence generation based on protein structure, but lacking effective screening and optimization mechanisms, making it difficult to guarantee the binding activity and structural stability of the generated sequences.
[0025] To address the aforementioned issues, this application provides a method, apparatus, electronic device, and program product for generating and screening nucleic acid aptamers. This solution uses first protein features containing the three-dimensional structural information of the target protein and the currently optimal reference nucleic acid aptamer as guiding conditions, combined with a guiding strength coefficient input into a trained conditional diffusion model. This allows for deep integration of the protein's structural context during generation and intelligent targeted exploration and optimization within the neighborhood of high-quality sequences. This enables the batch generation of candidate nucleic acid aptamers with potentially high binding affinity. Furthermore, a semantic space alignment model is used to quantitatively predict and score each candidate nucleic acid aptamer, and the candidate nucleic acid aptamers in the pool are screened based on these prediction scores. This achieves effective evaluation and screening of the generated candidate nucleic acid aptamers, identifying high-affinity nucleic acid aptamers and achieving efficient generation and screening of high-affinity nucleic acid aptamers overall.
[0026] The nucleic acid aptamer generation and screening method provided in this application embodiment can be applied to electronic devices such as mobile phones, tablets, wearable devices, augmented reality (AR) / virtual reality (VR) devices, desktop computers, servers, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application embodiment does not impose any restrictions on the specific type of electronic device.
[0027] To illustrate the technical solution of this application, specific embodiments are described below.
[0028] Please see Figure 1 , Figure 1 The flowchart illustrating the nucleic acid aptamer generation and screening method provided in this application embodiment is shown as an example and not a limitation. The method includes the following steps: Step 101: Input the target protein into the trained protein structure encoder to extract the first protein feature of the target protein.
[0029] Optionally, the target protein mentioned above can refer to any given protein for which nucleic acid aptamers need to be generated. Based on this, this embodiment can achieve refined generation and structural optimization of nucleic acid aptamers for any given protein (i.e., target protein).
[0030] Step 102: Input the guiding conditions and guiding strength coefficients into the trained conditional diffusion model to generate at least one candidate nucleic acid aptamer.
[0031] The guiding conditions include the first protein characteristic and the reference nucleic acid aptamer. The guiding strength coefficient is used to control the dependence of the trained conditional diffusion model's generation process on the guiding conditions. The aforementioned reference nucleic acid aptamer can refer to a pre-selected nucleic acid aptamer with a high degree of compatibility with the target protein's binding characteristics; it can also be called a high-quality nucleic acid aptamer.
[0032] The conditional diffusion model constructed in this embodiment can design nucleic acid aptamers for specific protein structures within a vast nucleic acid sequence space.
[0033] In this embodiment, by using the first protein feature containing the three-dimensional structural information of the target protein and the current optimal reference nucleic acid aptamer as guiding conditions, and combining the guiding strength coefficient with the trained conditional diffusion model, it is possible to not only deeply integrate the structural context of the protein during the generation process, but also to intelligently explore and optimize in the neighborhood space of high-quality sequences, thereby generating a batch of candidate nucleic acid aptamers with potentially high binding affinity.
[0034] Step 103: Input each candidate nucleic acid aptamer into the trained nucleic acid encoder to extract the first nucleic acid sequence features of each candidate nucleic acid aptamer.
[0035] In this embodiment, for each candidate nucleic acid aptamer, in order to capture the sequence characteristics and potential functional information of the candidate nucleic acid aptamer, the candidate nucleic acid aptamer can be input into a trained nucleic acid encoder, and the nucleic acid sequence features (i.e., the first nucleic acid sequence features) of the candidate nucleic acid aptamer can be extracted by the trained nucleic acid encoder.
[0036] Step 104: Input the first nucleic acid sequence feature and the first protein feature of each candidate nucleic acid aptamer into the trained semantic space alignment module to calculate the prediction score of each candidate nucleic acid aptamer.
[0037] Among them, the prediction score of a candidate nucleic acid aptamer characterizes the degree of fit between the candidate nucleic acid aptamer and the target protein in terms of binding properties, which can also be referred to as the degree of matching.
[0038] In some embodiments, for the first nucleic acid sequence feature of each candidate nucleic acid aptamer, the cosine similarity between the first nucleic acid sequence feature and the first protein feature can be calculated using a trained semantic space alignment model, and the cosine similarity can be determined as the predicted score of the candidate nucleic acid aptamer.
[0039] In this embodiment, by introducing a semantic space alignment module, an intrinsic screening mechanism based on the alignment features of protein and nucleic acid aptamers can be implemented to effectively evaluate and screen the generated candidate nucleic acid aptamers, thereby significantly improving the binding affinity and specificity of the generated nucleic acid aptamers. Furthermore, high-quality candidate nucleic acid aptamers can be obtained without relying on extensive experimental screening, offering advantages such as high throughput, low cost, and good scalability.
[0040] Step 105: Add each candidate nucleic acid aptamer and its predicted score to the nucleic acid candidate pool.
[0041] In this embodiment, after each candidate nucleic acid aptamer is generated, each generated candidate nucleic acid aptamer and its predicted score are added to the nucleic acid candidate pool, which can dynamically update the nucleic acid candidate pool and realize dynamic management and resource optimization of the nucleic acid candidate pool.
[0042] Step 106: The candidate nucleic acid aptamer with the highest predicted score in the nucleic acid candidate pool is identified as the nucleic acid aptamer with the best binding ability to the target protein.
[0043] Since the prediction score of a candidate nucleic acid aptamer characterizes the degree of fit between the candidate nucleic acid aptamer and the target protein, the higher the prediction score of a candidate nucleic acid aptamer, the higher the degree of fit between its binding characteristics and the target protein. Therefore, the candidate nucleic acid aptamer with the highest prediction score is identified as the nucleic acid aptamer with the best binding ability to the target protein, and high-affinity nucleic acid aptamers can be screened out.
[0044] In this embodiment, by using the first protein feature containing the three-dimensional structural information of the target protein and the current optimal reference nucleic acid aptamer as guiding conditions, and combining the guiding strength coefficient with the input to the trained conditional diffusion model, it is possible to not only deeply integrate the structural context of the protein during the generation process, but also to intelligently explore and optimize in the neighborhood space of high-quality sequences, thereby generating a batch of candidate nucleic acid aptamers with potential high binding affinity. Then, a semantic space alignment model is used to quantitatively predict and score each candidate nucleic acid aptamer, and the candidate nucleic acid aptamers in the candidate nucleic acid pool are screened based on the prediction scores. This enables effective evaluation and screening of each generated candidate nucleic acid aptamer, and the selection of high-affinity nucleic acid aptamers, thus achieving efficient generation and screening of high-affinity nucleic acid aptamers.
[0045] In some embodiments of this application, in Figure 1 Before using the protein structure encoder, nucleic acid encoder, and semantic space alignment module in the illustrated embodiment, these modules need to be trained. This training enables the protein structure encoder to transform complex three-dimensional protein structural information into high-dimensional feature vectors (i.e., protein features), the nucleic acid encoder to transform the rich sequence features and structural information of nucleic acid sequences in the high-dimensional semantic space into high-dimensional feature vectors (i.e., nucleic acid sequence features), and the semantic space alignment module to establish a mapping relationship between protein and nucleic acid sequences in the semantic space, achieving semantic consistency alignment between protein and nucleic acid sequences. The training process for the protein structure encoder, nucleic acid encoder, and semantic space alignment module includes the following steps: Figure 2 Steps 201 to 206 are shown.
[0046] Step 201: Obtain samples of each first protein-nucleic acid complex.
[0047] In some embodiments, at least one protein-nucleic acid complex can be extracted from a protein data bank (PDB) as a first protein-nucleic acid complex sample. For example, more than 20,000 protein-nucleic acid complexes from the protein database can all be used as the first protein-nucleic acid complex sample. The protein-nucleic acid complex includes protein and nucleic acid sequences.
[0048] Step 202: Based on the first protein sample and the first nucleic acid sequence sample in each first protein-nucleic acid complex sample, generate each first training sample and the true score of each first training sample.
[0049] A first training sample includes a first protein sample and a first nucleic acid sequence sample. The true score is determined by whether the first protein sample and the first nucleic acid sequence sample in the corresponding first training sample are located in the same first protein-nucleic acid complex. For example, if the first protein sample and the first nucleic acid sequence sample are located in the same first protein-nucleic acid complex, the true score of the first training sample composed of the first protein sample and the first nucleic acid sequence sample can be determined as 1; if the first protein sample and the first nucleic acid sequence sample are not located in the same first protein-nucleic acid complex, the true score of the first training sample composed of the first protein sample and the first nucleic acid sequence sample can be determined as 0. Based on this, the range of the predicted score can be defined as [0,1].
[0050] In some embodiments, a training sample set can be constructed by iterating through all first protein samples and all first nucleic acid sequence samples, pairing them together. Each first training sample in the training sample set consists of one first protein sample and one first nucleic acid sequence sample, which may originate from the same first protein-nucleic acid complex or from different first protein-nucleic acid complexes.
[0051] Step 203: For any first training sample, input the first protein sample in the first training sample into the protein structure encoder to be trained to extract the second protein feature of the first protein sample.
[0052] In some embodiments, a protein structure encoder may include a protein structure language model and an adapter that enables the protein structure language model to be better adapted to nucleic acid aptamer generation tasks. By way of example and not limitation, the protein structure language model may be SaProt.
[0053] In some embodiments, for any protein, the protein feature extraction process may include: extracting the sequence representation of the protein using Foldseek. Foldseek maps the three-dimensional structure of the protein to a structural sequence, fusing the sequence information and structural information of the protein to provide richer input for characterization extraction. SaProt further encodes the structural sequence into a numerical vector, comprehensively capturing the protein's local structural patterns, conserved domains, and physicochemical features. Assuming the three-dimensional structure of the protein is... Inputting it into Foldseek yields the corresponding structure sequence. The length is L (Foldseek uses amino acids as the basic unit, converting the atomic-level three-dimensional structure corresponding to each amino acid into a structural letter). , Belongs to the structural vocabulary Foldseek maps each amino acid to a structural sequence. :
[0054]
[0055] in, Indicates the size of the structural vocabulary, each Represents a discretized structural letter, This indicates Foldseek. Represents the specific structural letters in the structural vocabulary.
[0056] Input the structural sequence S into SaProt to extract protein features. Take the feature vector cls of its global representation, where one dimension of the vocabulary is... Protein characteristics It can be represented as follows:
[0057] in, It represents SaProt.
[0058] Step 204: Input the first nucleic acid sequence sample from the first training sample into the nucleic acid encoder to be trained to extract the second nucleic acid sequence features of the first nucleic acid sequence sample.
[0059] In some embodiments, the nucleic acid encoder may include a nucleic acid language model and an adapter that enables the nucleic acid language model to be better adapted to the nucleic acid aptamer generation task. As an example and not a limitation, the nucleic acid language model may be RiNALMo.
[0060] In some embodiments, for any nucleic acid sequence, the process of extracting nucleic acid sequence features may include: Assuming nucleic acid sequence It can be represented as follows:
[0061] in, Indicates the length of the nucleic acid sequence. This represents the i-th base in the nucleic acid sequence. Indicates adenine, It represents cytosine. It represents guanine. It represents thymine.
[0062] To capture the sequence characteristics and potential functional information of nucleic acid sequences, the nucleic acid sequence is input into a nucleic acid language model to generate a feature representation of each base. :
[0063] in, It represents RiNALMo.
[0064] Can The global feature representation is determined as the nucleic acid sequence feature of the nucleic acid sequence. It can be represented as follows:
[0065] in, Indicates will Chinese correspondence The feature vector of the marker is determined to be a nucleic acid sequence feature.
[0066] Step 205: Input the second protein feature and the second nucleic acid sequence feature into the semantic space alignment module to be trained to calculate the prediction score of the first training sample.
[0067] In some embodiments, the semantic space alignment module can calculate the similarity between protein features and nucleic acid sequence features to assess their degree of matching in the vector space and determine this similarity as a prediction score. Specifically, cosine similarity can be used as a constraint; the cosine similarity between protein features and nucleic acid sequence features... It can be represented as follows:
[0068] in, This represents the vector dimension of protein features and nucleic acid sequence features. It should be noted that the cosine similarity calculation formula above... These are protein features fine-tuned by the adapter in the protein structure encoder, and the cosine similarity calculation formula mentioned above... It is the nucleic acid sequence feature after being fine-tuned by the adapter in the nucleic acid encoder.
[0069] Step 206: Based on the real scores and predicted scores of each first training sample, train the protein structure encoder, the nucleic acid encoder, and the semantic space alignment module to be trained, to obtain the trained protein structure encoder, the trained nucleic acid encoder, and the trained semantic space alignment module.
[0070] In some embodiments, based on the true scores and predicted scores of each first training sample, a loss value can be calculated using a preset loss function, and the protein structure encoder, nucleic acid encoder, and semantic space alignment module to be trained can be trained based on this loss value. Optionally, a preset loss function can be set according to actual needs; this application does not limit the specific type of the preset loss function.
[0071] like Figure 3 The diagram shown is an example of semantic space alignment of protein and nucleic acid sequences provided in an embodiment of this application. Figure 3 In this process, the protein structure (i.e., the protein itself) is input into a protein structure encoder, which can extract the protein structure. , , , , By analyzing protein characteristics, the nucleic acid sequence is input into a nucleic acid encoder, which can extract proteins. , , , , By inputting protein features and nucleic acid sequence features into the semantic space alignment module, cosine similarity between protein features and nucleic acid sequence features can be calculated. Specifically, when training the protein structure encoder, nucleic acid encoder, and semantic space alignment module, the parameters of the protein structure language model and nucleic acid language model can be frozen, and only the parameters in the adapters of the protein structure encoder, nucleic acid encoder, and semantic space alignment module can be trained, thus achieving high-efficiency and low-cost training.
[0072] In some embodiments of this application, in Figure 1 Before using the conditional diffusion model in the illustrated embodiments, it is necessary to train the conditional diffusion model to improve the quality and diversity of candidate nucleic acid aptamers in the nucleic acid aptamer generation task. The training process of the above-mentioned conditional diffusion model may include the following steps: Obtain the non-coding ribonucleic acid sequences that meet the length requirements; Based on each non-coding ribonucleic acid sequence, the conditional diffusion model to be trained is pre-trained to obtain the pre-trained conditional diffusion model; Obtain samples of each second protein-nucleic acid complex; Based on each second protein nucleic acid complex sample, the pre-trained conditional diffusion model was fine-tuned to obtain a well-trained conditional diffusion model.
[0073] In some embodiments, one or more non-coding ribonucleic acid (NCRNA) sequences that meet the length requirements can be screened from the RNACentral database. Since NCRNA sequences are obtained from the RNACentral database and are naturally occurring in organisms, they can also be referred to as natural ribonucleic acid sequences or natural non-coding ribonucleic acid sequences.
[0074] Optionally, the above length requirement can be set based on the length of the nucleic acid aptamer. For example, the length of a nucleic acid aptamer is usually 20 to 100 bases, so the length requirement can be set to 20 to 100 bases.
[0075] In this embodiment, during the pre-training stage based on each non-coding ribonucleic acid (NCRNA) sequence, each NCRNA sequence serves as both the input condition for the conditional diffusion model to be trained and the target sample for the model to learn. Through self-supervised pre-training, the conditional diffusion model can fully learn the inherent distribution characteristics and diversity of NCRNA sequences, thereby improving the generalization ability and generation diversity of the conditional diffusion model.
[0076] In some embodiments, at least one protein-nucleic acid complex can be extracted from the PDB as a second protein-nucleic acid complex sample. It should be noted that each of the aforementioned second protein-nucleic acid complexes can be completely identical to each of the first protein-nucleic acid complexes (e.g., multiple protein-nucleic acid complexes can be extracted and used simultaneously as the first and second protein-nucleic acid complexes), partially identical, or completely different (e.g., different protein-nucleic acid complexes can be extracted and used as the first and second protein-nucleic acid complexes respectively).
[0077] In this embodiment, during the fine-tuning stage guided by the second protein-nucleic acid complex, the training in this stage is based on the pre-trained conditional diffusion model described above, and further learns the interaction pattern between protein structure and nucleic acid sequence, thereby achieving refined generation and structural optimization of nucleic acid aptamers for any given target protein.
[0078] In this embodiment, the above two-stage training strategy enables the trained conditional diffusion model to not only have the ability to generate reasonable and diverse nucleic acid sequences on a macroscopic level, but also to capture the precise binding rules between protein structures and nucleic acid sequences on a microscopic level, thereby significantly improving the specificity and reliability of nucleic acid aptamer generation.
[0079] In some embodiments of this application, pre-training of the conditional diffusion model to be trained is performed based on each non-coding ribonucleic acid sequence, including: Each non-coding ribonucleic acid (NCA) sequence is input into a trained nucleic acid encoder to extract the third nucleic acid sequence features of each NCA sequence. The conditional diffusion model to be trained is pre-trained by using the third nucleic acid sequence features of each non-coding ribonucleic acid sequence as the second training sample and the first denoising condition. Based on each second protein-nucleic acid complex sample, the pre-trained conditional propagation model was fine-tuned, including: The second protein sample in each second protein nucleic acid complex sample is input into the trained protein structure encoder to extract the third protein features of each second protein sample. The second nucleic acid sequence samples from each second protein-nucleic acid complex sample are input into a trained nucleic acid encoder to extract the fourth nucleic acid sequence features of each second nucleic acid sequence sample. For each second protein sample, the pre-trained conditional diffusion model is fine-tuned by using the fourth nucleic acid sequence feature of the second protein sample as the third training sample and the third protein feature of the second protein sample as the second denoising condition.
[0080] In some embodiments, during the pre-training and fine-tuning phases, the conditional diffusion model recovers the original data samples by progressively adding noise (e.g., adding noise for a total of 1000 steps) to the original training samples (e.g., the second and third training samples) and using a deep learning model to predict the noise. During the pre-training and fine-tuning processes of the conditional diffusion model, the model learns a transformation method that maps the noise distribution to the data distribution. The noise addition and denoising processes during pre-training and fine-tuning are as follows: Noise addition process: Define the initial nucleic acid sample The nucleic acid sequence characteristics (i.e., the original training samples) are Let the noise sample of the t-step noise addition be denoted as . When given the original training samples Adding 1000 steps of noise yields ,at this time A data point belonging to a normal distribution is represented, and the complex distribution of the entire dataset is transformed into a standard normal distribution, as shown in the following expression:
[0081] in, Indicates from the normal distribution The noise obtained from the sampling, Indicates control of noise intensity. The change in t is monotonically decreasing.
[0082] Generation process (denoising): Define denoising conditions (third nucleic acid sequence features in the pre-training stage, third protein features in the fine-tuning stage), conditional diffusion model. The noise added to the sample is predicted, thereby learning the relationship between the sample distributions at each time step. The specific expression is as follows:
[0083] in, This represents the sample reconstructed by the conditional diffusion model. Indicates sample The denoising step t is represented by c, which represents the denoising condition and is used to guide the conditional diffusion model G to remove noise in the correct direction.
[0084] Through the restored sample Convert it into a nucleic acid sequence and calculate its relationship with the initial nucleic acid sequence. The differences guide the denoising direction of the conditional diffusion model G, and the specific loss function. The design is as follows:
[0085] in, This represents the noise predicted by the conditional diffusion model. This represents the cross-entropy loss function, used to calculate the reconstructed sample. The converted nucleic acid sequence and the original nucleic acid sequence The differences.
[0086] During the pre-training phase, the loss value can be calculated based on the aforementioned loss function, and the conditional diffusion model to be trained can be pre-trained based on this loss value.
[0087] During the fine-tuning phase, the loss value can be calculated based on the aforementioned loss function, and the pre-trained conditional diffusion model can be fine-tuned based on this loss value.
[0088] like Figure 4 The diagram shown is an example of the pre-training stage based on various non-coding ribonucleic acid sequences provided in this application embodiment. Figure 4 In this context, 'c' represents the first noise-adding condition. , This represents the posterior probability of the generation process during the pre-training phase. , This represents the posterior probability of the noise-adding process during the pre-training phase. Figure 4 The cubes in the diagram represent nucleic acid sequence features from the pre-training phase. For example... Figure 5 The figure shown is an example diagram of the protein-guided fine-tuning stage provided in an embodiment of this application. Figure 5 In this context, 'c' represents the second noise-adding condition. , This represents the posterior probability of the generation process during the fine-tuning phase. , This represents the posterior probability of the noise addition process during the fine-tuning phase. Figure 5 The cube in the diagram represents the nucleic acid sequence characteristics during the fine-tuning phase.
[0089] To generate high-affinity nucleic acid aptamers based on the characteristics of the target protein, some embodiments of this application employ an iterative generation algorithm based on a conditional diffusion model. This algorithm continuously optimizes candidate nucleic acid aptamers through multiple rounds of generation and screening, enabling the generated nucleic acid aptamers to bind more effectively to the target protein. Specifically, before determining the candidate nucleic acid aptamer with the highest predicted score in the nucleic acid candidate pool as the nucleic acid aptamer with the optimal binding affinity to the target protein, the algorithm further includes: Calculate the difference between the predicted score of the reference nucleic acid aptamer and the highest predicted score in the nucleic acid candidate pool; The guidance strength coefficient is updated based on the score differences and the predicted score of the reference nucleic acid aptamer. The reference nucleic acid aptamer in the guiding conditions is updated to the candidate nucleic acid aptamer corresponding to the highest predicted score in the nucleic acid candidate pool; Based on the updated guidance strength coefficient and the updated guidance condition, return to execute the step of inputting the guidance condition and guidance strength coefficient into the trained conditional diffusion model, as well as subsequent steps, until the number of times the return execution is reached reaches the preset number of returns.
[0090] The aforementioned score difference can refer to the absolute value of the difference between the predicted score of the reference nucleic acid aptamer and the highest predicted score in the nucleic acid candidate pool.
[0091] The update formula for the above guiding strength coefficient is as follows:
[0092] in, This represents the updated guidance strength coefficient. This indicates the predicted score based on the reference nucleic acid aptamer. This represents the highest predicted score in the nucleic acid candidate pool. This represents a constant. Optionally, this constant can be set according to actual needs or empirical values to increase the guiding strength coefficient to a reasonable range.
[0093] In this embodiment, the guidance strength coefficient is updated based on the score difference and the predicted score of the reference nucleic acid aptamer, which enables dynamic adjustment of the guidance strength coefficient. The higher the guidance strength coefficient, the more the trained conditional diffusion model tends to explore in the neighborhood space of the reference nucleic acid aptamer, while the lower guidance strength coefficient encourages the trained conditional diffusion model to conduct a wider search.
[0094] In this embodiment, by updating the guidance strength coefficient and guidance conditions, and then returning to execute the steps of inputting the guidance conditions and guidance strength coefficients into the trained conditional diffusion model, as well as subsequent steps, until the number of return executions reaches a preset number, condition generation under positive and negative prompts can be achieved, thereby realizing an iterative generation algorithm based on the conditional diffusion model. Optionally, the preset number of return executions can be set according to actual needs or empirical values. It should be noted that each time the process returns to execute the step of inputting the guiding conditions and guiding strength coefficients into the trained conditional diffusion model, as well as subsequent steps, candidate nucleic acid aptamers will be generated and the nucleic acid candidate pool will be updated. That is, each return execution is an iteration, and in each iteration, a new set of candidate nucleic acid aptamers is generated using the trained conditional diffusion model under the dual conditions of the first protein feature and the reference nucleic acid aptamer.
[0095] In this embodiment, during the iteration phase, a dynamic guiding strength coefficient and a reference nucleic acid aptamer are introduced. The generated candidate nucleic acid aptamers are optimized and updated through a multi-round iteration mechanism, which can continuously improve the binding affinity and specificity of the nucleic acid aptamers.
[0096] In some embodiments of this application, before inputting the guiding conditions and guiding strength coefficients into the trained conditional diffusion model, the method further includes: If a nucleic acid aptamer exists that binds to the target protein, then that nucleic acid aptamer is designated as the reference nucleic acid aptamer. If no nucleic acid aptamer binds to the target protein, the first protein feature and the guiding strength coefficient are input into the trained conditional diffusion model to generate at least one candidate nucleic acid aptamer. The reference nucleic acid aptamer is then determined from the candidate nucleic acid aptamers generated this time.
[0097] The aforementioned nucleic acid aptamer that binds to the target protein can refer to a nucleic acid sequence that originates from the same protein-nucleic acid complex as the target protein.
[0098] It should be noted that the process of inputting the first protein feature and the guidance strength coefficient into the trained conditional diffusion model to generate at least one candidate nucleic acid aptamer can be called the first iteration. After determining the reference nucleic acid aptamer, the process of inputting the guidance conditions and guidance strength coefficient into the trained conditional diffusion model is called the iteration after the first iteration. For example, it can be the second iteration, the third iteration, etc., until the number of iterations reaches the preset number of iterations (i.e., the preset number of returns plus 2).
[0099] For each candidate nucleic acid aptamer generated in this iteration, i.e. each candidate nucleic acid aptamer generated in the first iteration, steps 103 to 105 are also executed to add each candidate nucleic acid aptamer generated in the first iteration and its predicted score to the nucleic acid candidate pool, and the candidate nucleic acid aptamer with the highest predicted score among the candidate nucleic acid aptamers generated in the first iteration is determined as the reference nucleic acid aptamer.
[0100] To maintain the diversity of candidate nucleic acid aptamers, in some embodiments of this application, before adding each candidate nucleic acid aptamer and its predicted score to the nucleic acid candidate pool, the following steps are also included: Calculate the similarity between each candidate nucleic acid aptamer and the reference nucleic acid aptamer; Each candidate nucleic acid aptamer and its predicted score are added to the nucleic acid candidate pool, including: Candidate nucleic acid aptamers whose predicted scores are greater than the score threshold and whose similarity is greater than the similarity threshold, along with their predicted scores, are added to the nucleic acid candidate pool.
[0101] Optionally, the similarity calculation method between the candidate nucleic acid aptamer and the reference nucleic acid aptamer can be selected according to actual needs. For example, Euclidean distance, cosine similarity, or radial basis function (RBF) kernel can be used to calculate the similarity.
[0102] Optionally, a scoring threshold can be set based on actual needs or experience.
[0103] Optionally, a similarity threshold can be set according to actual needs or empirical values, or it can be dynamically set according to the average similarity and standard deviation of all candidate nucleic acid aptamers (for example, calculating the cosine similarity between the generated candidate nucleic acid aptamers and the reference nucleic acid aptamers, statistically analyzing the corresponding values, and removing candidate nucleic acid aptamers that are far from the sample center (e.g., the sample center defined by the mean and variance) to filter out candidate nucleic acid aptamers that are not very similar to the reference nucleic acid aptamers.
[0104] In this embodiment, candidate nucleic acid aptamers can be screened using scoring and similarity thresholds. Selected aptamers are then added to the nucleic acid candidate pool, allowing for pool updates. When the pool is not full, new aptamers are added directly. If the pool is full, the aptamer with the lowest predicted score is removed, and a new aptamer with a higher predicted score is added. This mechanism ensures the pool continuously evolves towards higher quality and greater diversity during iteration. Optionally, the capacity of the nucleic acid candidate pool can be set based on actual needs or empirical values to determine whether the pool is full.
[0105] like Figure 6 The diagram illustrates an example of the iterative screening stage for nucleic acid aptamers provided in this embodiment. This embodiment utilizes a closed-loop iterative process of "generation—evaluation—screening" to progressively optimize candidate nucleic acid aptamers guided by protein features and reference nucleic acid aptamers. A dynamic guidance intensity coefficient adjustment mechanism enables the conditional diffusion model to achieve a balance between exploration and utilization; while the candidate screening strategy based on predicted scores and similarity can simulate the SELEX screening stage, yielding nucleic acid aptamers with better scores. Overall, this framework achieves adaptive search of the nucleic acid aptamer sequence space, providing a generalizable generative method for data-driven molecular recognition design.
[0106] To verify the effectiveness of the nucleic acid aptamer generation and screening method provided in this application (which can be referred to as the Nucleic Acid aptamer Generator (NAGen) model), a dataset S1 containing 13,083 protein-nucleic acid complexes was constructed based on PDB. 67,398 non-coding ribonucleic acid sequences were screened from the RNACentral database to construct a pre-training dataset S2. Dataset S1 was used for protein-nucleic acid semantic alignment training, and S2 was used for pre-training of the conditional diffusion generation model. Dataset S1 was then used to fine-tune the conditional diffusion model to achieve protein-guided nucleic acid aptamer design.
[0107] The effectiveness of the proposed method was validated using a dataset constructed using relevant technologies as a test set. The Plddt (confidence score per residue) index of the protein-nucleic acid complex was calculated using RosettaFoldNA to evaluate the method's performance. The results showed (e.g.) Figure 7 (As shown). The proposed method performs only one round of generation, and a quarter of the selected nucleic acid aptamers have a better structural Plddt than the original nucleic acid aptamers (i.e., nucleic acid sequences located in the same protein-nucleic acid complex as the protein). After iterative optimization with positive and negative sample hints, more than half of the final obtained nucleic acid aptamers are better than the original nucleic acid aptamers (e.g., ...). Figure 8 (As shown).
[0108] The binding free energy (ΔG) of the protein-nucleic acid complex was further calculated using the CoPRA model. This indicator can be used to measure the activity and binding strength of the protein-nucleic acid complex and is an important basis for screening high-affinity nucleic acid aptamers; the lower the ΔG, the stronger the activity. To evaluate the effectiveness of different screening strategies, the original nucleic acid aptamers were used as a baseline (i.e., Figure 9 The proposed method (using the src parameter) was used to optimize nucleic acid aptamers. Five control groups were set up: no optimization (NAGen), optimization based on Euclidean distance (NAGen w_evo_edu), optimization based on cosine similarity (NAGen w_evo_cos), optimization based on RBF kernel (NAGen w_evo_rbf), and optimization only once without iterative generation (NAGen w_one-shot). The results are shown in Figure 9. The median ΔG of the nucleic acid aptamers generated in one step (NAGen w_one-shot) was lower than that of the original nucleic acid aptamers, indicating that the generated aptamers generally had stronger binding activity. The nucleic acid aptamers generated iteratively showed even higher binding activity, verifying the feasibility and effectiveness of the iterative generation strategy.
[0109] To verify the effectiveness of the proposed method, it was compared with existing nucleic acid aptamer design methods, and the results are shown in Table 1. The average Plddt of nucleic acid aptamers generated by NAGen in a single run was slightly lower than that of the other four existing methods. However, after iterative optimization and screening, the Plddt of the nucleic acid aptamers generated significantly outperformed existing methods, further demonstrating the effectiveness of the proposed method.
[0110] Table 1
[0111] Among them, RNAFlow, RNA-BAnG, Apta-MCTS, and AptaTrans in Table 1 are four existing methods.
[0112] The method proposed in this application significantly improves the effectiveness, generalization ability, and target protein adaptability of nucleic acid aptamer generation, providing a new technical approach for the intelligent and efficient design of nucleic acid aptamers.
[0113] This application provides a method for generating and screening nucleic acid aptamers, which is an iterative generation method based on protein structure guidance. This method uses the protein characteristics of the target protein as input conditions, combining a conditional diffusion model with two stages: pre-training of non-coding ribonucleic acid sequences and protein-guided fine-tuning. Through a closed-loop iterative mechanism of "generation-evaluation-screening," it achieves efficient nucleic acid sequence design starting from protein structure. In each iteration, the model adaptively adjusts the guidance strength coefficient based on the predicted scores of candidate nuclide aptamers, and combines a similarity constraint mechanism to balance sequence diversity and affinity, thereby generating high-quality nucleic acid aptamer sequences through continuous optimization of the candidate pool.
[0114] In summary, the protein structure-guided iterative generation framework for nucleic acid aptamers proposed in this application significantly outperforms traditional design methods in terms of generation diversity, structural reliability, specific recognition, and generalization ability. This method provides an efficient, interpretable, and theoretically supported new approach for the automated design of nucleic acid aptamers based on protein structure information.
[0115] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0116] Corresponding to the nucleic acid aptamer generation and screening method described in the above embodiments, Figure 10 A schematic diagram of the structure of the nucleic acid aptamer generation and screening device provided in the embodiments of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0117] Reference Figure 10 The device includes: The first extraction module 1001 is used to input the target protein into a trained protein structure encoder to extract the first protein feature of the target protein. The candidate generation module 1002 is used to input guiding conditions and guiding strength coefficients into a trained conditional diffusion model to generate at least one candidate nucleic acid aptamer; the guiding conditions include the first protein feature and a reference nucleic acid aptamer, and the guiding strength coefficient is used to control the degree of dependence of the generation process of the trained conditional diffusion model on the guiding conditions; The second extraction module 1003 is used to input each candidate nucleic acid aptamer into the trained nucleic acid encoder to extract the first nucleic acid sequence features of each candidate nucleic acid aptamer. The scoring calculation module 1004 is used to input the first nucleic acid sequence features and the first protein features of each candidate nucleic acid aptamer into a trained semantic space alignment module to calculate the predicted score of each candidate nucleic acid aptamer; the predicted score of a candidate nucleic acid aptamer characterizes the degree of fit between the candidate nucleic acid aptamer and the target protein in terms of binding characteristics. The candidate addition module 1005 is used to add each candidate nucleic acid aptamer and its predicted score to the nucleic acid candidate pool. The target determination module 1006 is used to determine the candidate nucleic acid aptamer with the highest predicted score in the nucleic acid candidate pool as the nucleic acid aptamer with the best binding ability to the target protein.
[0118] In some embodiments, the above-described apparatus further includes: The difference calculation module is used to calculate the difference between the predicted score of the reference nucleic acid aptamer and the highest predicted score in the nucleic acid candidate pool; The first update module is used to update the guidance strength coefficient based on the score difference and the predicted score of the reference nucleic acid aptamer; The second update module is used to update the reference nucleic acid aptamer in the guidance conditions to the candidate nucleic acid aptamer corresponding to the highest predicted score in the nucleic acid candidate pool; The return execution module is used to return to the execution of the steps of inputting the guidance conditions and guidance strength coefficients into the trained conditional diffusion model, as well as subsequent steps, based on the updated guidance strength coefficients and the updated guidance conditions, until the number of return executions reaches the preset number of return executions.
[0119] In some embodiments, the above-described apparatus further includes: The first determining module is configured to determine the reference nucleic acid aptamer if a nucleic acid aptamer exists that binds to the target protein. The second determining module is used to input the first protein characteristics and the guiding strength coefficient into the trained conditional diffusion model if no nucleic acid aptamer binds to the target protein, generate at least one candidate nucleic acid aptamer, and determine the reference nucleic acid aptamer from the candidate nucleic acid aptamers generated this time.
[0120] In some embodiments, the above-described apparatus further includes: A similarity calculation module is used to calculate the similarity between each candidate nucleic acid aptamer and the reference nucleic acid aptamer. The aforementioned candidate addition module 1005 is specifically used for: The candidate nucleic acid aptamers whose predicted scores are greater than the score threshold and whose similarity is greater than the similarity threshold, along with their predicted scores, are added to the nucleic acid candidate pool.
[0121] In some embodiments, the above apparatus further includes a first training module, the first training module being configured to: Obtain samples of each first protein-nucleic acid complex; Based on the first protein sample and the first nucleic acid sequence sample in each of the first protein-nucleic acid complex samples, each first training sample and the true score of each first training sample are generated; a first training sample includes a first protein sample and a first nucleic acid sequence sample, and the true score is determined by whether the first protein sample and the first nucleic acid aptamer in the corresponding first training sample are located in the same first protein-nucleic acid complex. For any of the first training samples, the first protein sample in the first training sample is input into the protein structure encoder to be trained to extract the second protein feature of the first protein sample; the first nucleic acid sequence sample in the first training sample is input into the nucleic acid encoder to be trained to extract the second nucleic acid sequence feature of the first nucleic acid sequence sample. The second protein feature and the second nucleic acid sequence feature are input into the semantic space alignment module to be trained to calculate the prediction score of the first training sample. Based on the actual scores and predicted scores of each of the first training samples, the protein structure encoder, the nucleic acid encoder, and the semantic space alignment module to be trained are trained to obtain the trained protein structure encoder, the trained nucleic acid encoder, and the trained semantic space alignment module.
[0122] In some embodiments, the above apparatus further includes a second training module, the second training module being used for: Obtain the non-coding ribonucleic acid sequences that meet the length requirements; Based on the aforementioned non-coding ribonucleic acid sequences, the conditional diffusion model to be trained is pre-trained to obtain the pre-trained conditional diffusion model. Obtain samples of each second protein-nucleic acid complex; Based on the samples of each second protein-nucleic acid complex, the pre-trained conditional diffusion model is fine-tuned to obtain the trained conditional diffusion model.
[0123] In some embodiments, the second training module is specifically used for: Each non-coding ribonucleic acid sequence is input into the trained nucleic acid encoder to extract the third nucleic acid sequence features of each non-coding ribonucleic acid sequence; The conditional diffusion model to be trained is pre-trained by using the third nucleic acid sequence features of each non-coding ribonucleic acid sequence as the second training sample and the first denoising condition. The second protein sample in each second protein nucleic acid complex sample is input into the trained protein structure encoder to extract the third protein feature of each second protein sample. The second nucleic acid sequence samples from each of the second protein-nucleic acid complex samples are input into the trained nucleic acid encoder to extract the fourth nucleic acid sequence features of each second nucleic acid sequence sample; For each of the second protein samples, the pre-trained conditional diffusion model is fine-tuned by using the fourth nucleic acid sequence feature of the second protein sample as the third training sample and the third protein feature of the second protein sample as the second denoising condition.
[0124] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0125] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 11 As shown, the electronic device 11 of this embodiment includes: at least one processor 1100 ( Figure 11 (Only one is shown in the diagram), memory 1101, and computer program 1102 stored in said memory 1101 and executable on said at least one processor 1100, which, when executing said computer program 1102, implements the steps in any of the above method embodiments.
[0126] The electronic device may include, but is not limited to, a processor 1100 and a memory 1101. Those skilled in the art will understand that... Figure 11 This is merely an example of electronic device 11 and does not constitute a limitation on electronic device 11. It may include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0127] The processor 1100 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0128] In some embodiments, the memory 1101 may be an internal storage unit of the electronic device 11, such as a hard disk or memory of the electronic device 11. In other embodiments, the memory 1101 may be an external storage device of the electronic device 11, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 11. Furthermore, the memory 1101 may include both internal and external storage units of the electronic device 11. The memory 1101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 1101 can also be used to temporarily store data that has been output or will be output.
[0129] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0130] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0131] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0132] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0133] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0134] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0135] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for generating and screening nucleic acid aptamers, characterized in that, include: The target protein is input into a trained protein structure encoder to extract the first protein feature of the target protein; The guiding conditions and guiding strength coefficient are input into the trained conditional diffusion model to generate at least one candidate nucleic acid aptamer; the guiding conditions include the first protein feature and the reference nucleic acid aptamer, and the guiding strength coefficient is used to control the degree to which the generation process of the trained conditional diffusion model depends on the guiding conditions. Each candidate nucleic acid aptamer is input into the trained nucleic acid encoder to extract the first nucleic acid sequence features of each candidate nucleic acid aptamer; The first nucleic acid sequence features and the first protein features of each candidate nucleic acid aptamer are input into a trained semantic space alignment module to calculate the prediction score of each candidate nucleic acid aptamer. A prediction score for a candidate nucleic acid aptamer characterizes the degree of fit between the candidate nucleic acid aptamer and the target protein in terms of binding properties; Each candidate nucleic acid aptamer and its predicted score are added to the nucleic acid candidate pool; The candidate nucleic acid aptamer with the highest predicted score in the nucleic acid candidate pool is determined as the nucleic acid aptamer with the best binding ability to the target protein.
2. The method for generating and screening nucleic acid aptamers according to claim 1, characterized in that, Before determining the candidate nucleic acid aptamer with the highest predicted score in the nucleic acid candidate pool as the nucleic acid aptamer with the best binding ability to the target protein, the process also includes: Calculate the difference between the predicted score of the reference nucleic acid aptamer and the highest predicted score in the nucleic acid candidate pool; The guidance strength coefficient is updated based on the score difference and the predicted score of the reference nucleic acid aptamer. The reference nucleic acid aptamer in the guidance conditions is updated to the candidate nucleic acid aptamer corresponding to the highest predicted score in the nucleic acid candidate pool; Based on the updated guidance strength coefficient and the updated guidance condition, return to the step of inputting the guidance condition and guidance strength coefficient into the trained conditional diffusion model and subsequent steps, until the number of times the return execution is reached reaches the preset number of returns.
3. The method for generating and screening nucleic acid aptamers according to claim 1, characterized in that, Before inputting the guiding conditions and guiding strength coefficients into the trained conditional diffusion model, the following steps are also included; If a nucleic acid aptamer exists that binds to the target protein, then that nucleic acid aptamer is designated as the reference nucleic acid aptamer. If no nucleic acid aptamer binds to the target protein, the first protein characteristics and the guiding strength coefficient are input into the trained conditional diffusion model to generate at least one candidate nucleic acid aptamer, and the reference nucleic acid aptamer is determined from the candidate nucleic acid aptamers generated this time.
4. The method for generating and screening nucleic acid aptamers according to claim 1, characterized in that, Before adding each candidate nucleic acid aptamer and its predicted score to the nucleic acid candidate pool, the following steps are also included: Calculate the similarity between each candidate nucleic acid aptamer and the reference nucleic acid aptamer; The step of adding each candidate nucleic acid aptamer and its predicted score to the nucleic acid candidate pool includes: The candidate nucleic acid aptamers whose predicted scores are greater than the score threshold and whose similarity is greater than the similarity threshold, along with their predicted scores, are added to the nucleic acid candidate pool.
5. The method for generating and screening nucleic acid aptamers according to any one of claims 1 to 4, characterized in that, The method for generating and screening nucleic acid aptamers also includes: Obtain samples of each first protein-nucleic acid complex; Based on the first protein sample and the first nucleic acid sequence sample in each of the first protein-nucleic acid complex samples, each first training sample and the true score of each first training sample are generated; a first training sample includes a first protein sample and a first nucleic acid sequence sample, and the true score is determined by whether the first protein sample and the first nucleic acid aptamer in the corresponding first training sample are located in the same first protein-nucleic acid complex. For any of the first training samples, the first protein sample in the first training sample is input into the protein structure encoder to be trained to extract the second protein feature of the first protein sample; the first nucleic acid sequence sample in the first training sample is input into the nucleic acid encoder to be trained to extract the second nucleic acid sequence feature of the first nucleic acid sequence sample. The second protein feature and the second nucleic acid sequence feature are input into the semantic space alignment module to be trained to calculate the prediction score of the first training sample. Based on the actual scores and predicted scores of each of the first training samples, the protein structure encoder, the nucleic acid encoder, and the semantic space alignment module to be trained are trained to obtain the trained protein structure encoder, the trained nucleic acid encoder, and the trained semantic space alignment module.
6. The method for generating and screening nucleic acid aptamers according to any one of claims 1 to 4, characterized in that, The method for generating and screening nucleic acid aptamers also includes: Obtain the non-coding ribonucleic acid sequences that meet the length requirements; Based on the aforementioned non-coding ribonucleic acid sequences, the conditional diffusion model to be trained is pre-trained to obtain the pre-trained conditional diffusion model. Obtain samples of each second protein-nucleic acid complex; Based on the samples of each second protein-nucleic acid complex, the pre-trained conditional diffusion model is fine-tuned to obtain the trained conditional diffusion model.
7. The method for generating and screening nucleic acid aptamers according to claim 6, characterized in that, The pre-training of the conditional diffusion model to be trained based on the aforementioned non-coding ribonucleic acid sequences includes: Each non-coding ribonucleic acid sequence is input into the trained nucleic acid encoder to extract the third nucleic acid sequence features of each non-coding ribonucleic acid sequence; The conditional diffusion model to be trained is pre-trained by using the third nucleic acid sequence features of each non-coding ribonucleic acid sequence as the second training sample and the first denoising condition. The fine-tuning of the pre-trained conditional propagation model based on each of the second protein-nucleic acid complex samples includes: The second protein sample in each second protein nucleic acid complex sample is input into the trained protein structure encoder to extract the third protein feature of each second protein sample. The second nucleic acid sequence samples from each of the second protein-nucleic acid complex samples are input into the trained nucleic acid encoder to extract the fourth nucleic acid sequence features of each second nucleic acid sequence sample; For each of the second protein samples, the pre-trained conditional diffusion model is fine-tuned by using the fourth nucleic acid sequence feature of the second protein sample as the third training sample and the third protein feature of the second protein sample as the second denoising condition.
8. A nucleic acid aptamer generation and screening device, characterized in that, include: The first extraction module is used to input the target protein into a trained protein structure encoder to extract the first protein feature of the target protein. A candidate generation module is used to input guiding conditions and guiding strength coefficients into a trained conditional diffusion model to generate at least one candidate nucleic acid aptamer; the guiding conditions include the first protein feature and a reference nucleic acid aptamer, and the guiding strength coefficient is used to control the degree of dependence of the generation process of the trained conditional diffusion model on the guiding conditions; The second extraction module is used to input each candidate nucleic acid aptamer into the trained nucleic acid encoder to extract the first nucleic acid sequence features of each candidate nucleic acid aptamer. The scoring calculation module is used to input the first nucleic acid sequence features and the first protein features of each candidate nucleic acid aptamer into the trained semantic space alignment module to calculate the predicted score of each candidate nucleic acid aptamer. A prediction score for a candidate nucleic acid aptamer characterizes the degree of fit between the candidate nucleic acid aptamer and the target protein in terms of binding properties; The candidate addition module is used to add each candidate nucleic acid aptamer and its predicted score to the nucleic acid candidate pool. The target determination module is used to identify the candidate nucleic acid aptamer with the highest predicted score in the nucleic acid candidate pool as the nucleic acid aptamer with the best binding ability to the target protein.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to implement the nucleic acid aptamer generation and screening method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, Includes a computer program, which, when run, causes the nucleic acid aptamer generation and screening method as described in any one of claims 1 to 7 to be performed.