Nucleic acid aptamers specifically combined with CD8 molecules and application of nucleic acid aptamers
Through machine learning, optimized Y-shaped nucleic acid aptamer structure, the complex and cost-effective problems of existing CD8 cell detection methods are solved, efficient and economical CD8 cell recognition and combination, and the diagnosis and prognosis capabilities are improved.
Patent Information
- Application Number
- CN202510247919.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
The existing CD8 cell detection methods have the problem of complex antibody preparation and high cost and limited tracer use, making it difficult to achieve efficient and economical CD8 cell recognition and binding.
Through machine learning, the secondary structural rules were analyzed, the CD8 nucleic acid aptamer candidate library was optimized, and a nucleic acid aptamer with high affinity and high specificity binding to CD8 protein was obtained. The Y-shaped core structure was adopted, including two stem loops, a multi-branch loop and the main stem, specifically binding to CD8 molecules.
It achieves efficient and economical specific identification and binding of CD8 molecules or CD8 positive cells, improves diagnostic and prognostic capabilities, and reduces detection costs.
Smart Images

Figure BDA0005296102590000051 
Figure BDA0005296102590000061 
Figure BDA0005296102590000091
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and specifically relates to a group of nucleic acid aptamers that specifically bind to CD8 molecules and their applications. Background Art
[0002] CD8 is a glycoprotein expressed on the surface of all cytotoxic T lymphocytes (CTLs) and various non-specific immune cells (such as DCs, macrophages, monocytes, and NK cells), and belongs to the members of the Immunoglobulins superfamily (IgSF) in cell adhesion molecules. It is a homodimer composed of two CD8α chains or a heterodimer composed of CD8α and CD8β chains. CD8 can bind to the α3 domain of the heavy chain of the Major histocompatibility complex class I molecule (MHC-I), assist the T cell receptor (TCR) to recognize the antigen presented by it, can increase the stability of the ternary complex structure of the antigen, TCR, and MHC molecules, and then participate in T cell activation by binding to MHC-I molecules. In addition, the CxCP motif in the cytoplasmic region of CD8 can bind to the Src family tyrosine protein kinase p56 lck Complete the signal transduction of T lymphocytes in immune regulation, ultimately leading to the production of lymphokines, the movement, adhesion, and activation of CTLs. This mechanism enables CTLs to recognize and eliminate infected cells or tumor cells.
[0003] A large number of studies have shown that inhibiting CD8 can significantly reduce the immune response, making it a viable target for developing biotherapeutic drugs. Using a reasonable immunotherapy strategy, CD8 + Cytotoxic lymphocytes can be exogenously reactivated and / or induced, and can be used to treat diseases such as autoimmune disorders, allergies, and various cancers. CD8 infiltrating into the tumor microenvironment + Cytotoxic lymphocytes can selectively recognize and eliminate cancer cells. However, as the tumor progresses continuously, CD8 + Cytotoxic lymphocytes will differentiate into a low-vitality state with inactivated or weakened killing ability. Therefore, CD8 + Cytotoxic lymphocytes have diagnostic and prognostic significance in various cancers. At present, antibodies are mainly used to detect CD8 in peripheral blood or related tissues +Cytotoxic lymphocytes are detected, and positron emission tomography (PET) using a radiolabeled tracer is also a current popular detection method. However, there are still many deficiencies: the preparation of antibodies is complex and costly, and the use of tracers is limited by the half-life of radioisotopes and cell division, resulting in in vivo probe dilution.
[0004] Aptamers are single-stranded oligonucleotides that specifically bind to targets by forming three-dimensional structures. Aptamers have the advantages of rapid in vitro screening and low immunogenicity, and have various types of binding targets, including metal ions, organic small molecules, polypeptides, proteins, viruses, bacteria, cells, tissues, etc. They are similar to antibodies and have a variety of applications, including therapy, biosensors, and diagnostics. Therefore, it is of great significance to obtain the core structure by analyzing the secondary structure rules through machine learning, optimize the CD8 aptamer candidate library using the core structure, and further obtain aptamers that specifically bind to CD8 protein with high affinity and high specificity in the treatment, diagnosis, and prognosis of cancer. Summary of the Invention
[0005] The object of the present invention is to provide a group of aptamers that specifically bind to CD8 molecules and their applications.
[0006] In a first aspect, the present invention claims aptamers targeting CD8 molecules.
[0007] The aptamers targeting CD8 molecules claimed by the present invention comprise a core structure; the core structure is similar to a Y shape and consists of two stem-loops, a multi-branched loop, and a main stem; the multi-branched loop is located between the two stem-loops;
[0008] Among the two stem-loops, the stem-loop near the 5'-end is denoted as stem-loop 1, which is composed of a random base sequence, and the stem-loop near the 3'-end is denoted as stem-loop 2. The stem region of stem-loop 2 is composed of GC-paired bases, and its loop structure is formed by the fixed sequence AGCTTGAAAT;
[0009] The multi-branched loop has the fixed base GTGA; on the multi-branched loop, there is a branch at the 5'-end adjacent to the fixed base GTGA, denoted as branch 1, and a branch at the 3'-end adjacent to the fixed base GTGA, denoted as branch 2. There is also a branch 3 at the position opposite to the fixed base GTGA between branch 1 and branch 2;
[0010] Branch 1 is connected to the end of the stem away from the loop in stem-loop 1; branch 2 is connected to the end of the stem away from the loop in stem-loop 2; branch 3 is connected to the main stem;
[0011] The main stem is composed of random base pairing.
[0012] Furthermore, the aptamer nucleic acid contains a core sequence; the core sequence is GTGA-NNN-AGCTTGAAA or GTGA-NN-AGCTTGAAA from the 5' end to the 3' end; wherein, N is A or T or C or G.
[0013] Even further, the aptamer nucleic acid targeting the CD8 molecule can specifically be any of the following:
[0014] (A1) A single-stranded DNA molecule shown in any of SEQ ID No.1 to SEQ ID No.35;
[0015] The aptamer nucleic acid sequences of the present invention (any of SEQ ID No.1 to SEQ ID No.35) all have the core structure and the core sequence described above.
[0016] (A2) An aptamer nucleic acid with the same function obtained by deleting or adding one or several nucleotides to the aptamer nucleic acid shown in (A1);
[0017] Even further, other sequences that have individual base modifications with the aptamer nucleic acid sequences of the present invention (any of SEQ ID No.1 to SEQ ID No.35), have a similarity of more than 80%, have the core structure described above, have the same or extremely similar applications as the aptamer nucleic acid of the present invention, and have no impact on the overall function should also be considered to fall within the protection scope of the present invention.
[0018] (A3) An aptamer nucleic acid with the same function obtained by substituting or modifying nucleotides of the aptamer nucleic acid shown in (A1);
[0019] Even further, the modification can be phosphorylation, methylation, amination, thiolation, isotopic labeling, etc.
[0020] (A4) An aptamer nucleic acid with the same function obtained from the RNA molecule encoded by the aptamer nucleic acid shown in (A1).
[0021] In the second aspect, the present invention claims to protect aptamer nucleic acid derivatives.
[0022] The aptamer nucleic acid derivatives claimed to be protected by the present invention are any of the following:
[0023] (B1) An aptamer nucleic acid derivative with the same function as the aptamer nucleic acid obtained by modifying the backbone of the aptamer nucleic acid described in the first aspect above into a phosphorothioate backbone;
[0024] (B2) A peptide nucleic acid encoded by the nucleic acid aptamer described in the first aspect above, to obtain a derivative of the nucleic acid aptamer having the same function as the nucleic acid aptamer;
[0025] (B3) Connect a chemical group and / or fluorescein and / or anti-tumor drug and / or radioactive element and / or biological enzyme and / or biotin and / or nanomaterial to one end or the middle (connecting at one position or multiple positions) of the nucleic acid aptamer described in the first aspect above, to obtain a derivative of the nucleic acid aptamer having the same function as the nucleic acid aptamer.
[0026] In the third aspect, the present invention claims the application of the nucleic acid aptamer described in the first aspect above or the derivative of the nucleic acid aptamer described in the second aspect above in any of the following:
[0027] (C1) Preparing a product for specifically recognizing CD8 molecules;
[0028] (C2) Specifically recognizing CD8 molecules.
[0029] In the fourth aspect, the present invention claims the application of the nucleic acid aptamer described in the first aspect above or the derivative of the nucleic acid aptamer described in the second aspect above in any of the following:
[0030] (D1) Preparing a product for specifically binding to CD8-positive cells;
[0031] (D2) Specifically binding to CD8-positive cells.
[0032] In the fifth aspect, the present invention claims the application of the nucleic acid aptamer described in the first aspect above or the derivative of the nucleic acid aptamer described in the second aspect above in any of the following:
[0033] (E1) Preparing a product for detecting CD8 molecules or CD8-positive cells;
[0034] (E2) Detecting CD8 molecules or CD8-positive cells.
[0035] In the sixth aspect, the present invention claims the application of the nucleic acid aptamer described in the first aspect above or the derivative of the nucleic acid aptamer described in the second aspect above in any of the following:
[0036] (F1) Preparing a product for capturing or enriching or purifying CD8 molecules or CD8-positive cells;
[0037] (F2) Capturing or enriching or purifying CD8 molecules or CD8-positive cells.
[0038] In the seventh aspect, the present invention claims a product.
[0039] The active ingredient of the product claimed in the present invention is (or contains) the nucleic acid aptamer described in the first aspect above or the nucleic acid aptamer derivative described in the second aspect above;
[0040] The product has any one of the following functions:
[0041] (G1) Specifically recognize CD8 molecules;
[0042] (G2) Specifically bind to CD8-positive cells;
[0043] (G3) Detect CD8 molecules or CD8-positive cells;
[0044] (G4) Capture, enrich or purify CD8 molecules or CD8-positive cells.
[0045] In each of the above relevant aspects, the product can be a kit.
[0046] In the eighth aspect, the present invention claims a method for obtaining the nucleic acid aptamer described in the first aspect above.
[0047] The method for obtaining the nucleic acid aptamer described in the first aspect above claimed by the present invention may include the following steps: using CD8 as the target protein, screening to obtain a nucleic acid aptamer library; then using the nucleic acid aptamer library to perform nucleic acid aptamer structure calculation according to the method including the following steps S1-S10, and further obtaining the nucleic acid aptamer described in the first aspect above according to the calculation result.
[0048] S1. Perform high-throughput sequencing on the nucleic acid aptamer library to obtain sequencing data;
[0049] S2. Preprocess the sequencing data to obtain nucleic acid aptamer data;
[0050] S3. Input the nucleic acid aptamer data of multiple nucleic acid aptamers that meet the preset abundance requirements in the nucleic acid aptamer data into a deep learning model for training to map the sequences of multiple nucleic acid aptamers to a low-dimensional representation in the latent space;
[0051] S4. Based on the trained deep learning model to be trained, use the encoder to remap the sequences of multiple nucleic acid aptamers to the latent space, and use the Gaussian mixture model to cluster the representations in the latent space, so as to divide the sequences of multiple nucleic acid aptamers into different families;
[0052] S5. According to the characteristics of the sequences of nucleic acid aptamers in each family, obtain the sequence identification map of each family, wherein the high-frequency base sequence in the sequence identification map of each family is the common rule of the family;
[0053] S6. Calculate the proportion of the sequence rules for each family according to the size of each family and its common rules, and obtain the high-proportion common rules for each family as the core sequences related to binding in each family.
[0054] S7. Use a nucleic acid secondary structure prediction model to perform secondary structure prediction on the sequences of the nucleic acid aptamers of each family, and obtain a secondary structure text file for each family containing sequences, secondary structures, and folding scores.
[0055] S8. For each family, according to its secondary structure text file, save the sequences and secondary structures into the first list respectively, calculate the regions of the core sequences in each nucleic acid aptamer, calculate the smallest substructures in the secondary structures that contain these regions, and save the smallest substructures and their sequences into the second list.
[0056] S9. Calculate the number of secondary structure elements of each substructure in the second list through the first recursive algorithm, count the most frequently occurring structural combinations as the core structures, and calculate the base frequencies at each position in the sequence list of each family that contains the core sequences through the second recursive algorithm to obtain the final structural representation of each family representing high-frequency bases.
[0057] S10. Use the final structural representation of each family representing high-frequency bases to optimize the nucleic acid aptamer sequences of each family.
[0058] In order to make full use of the rich aptamer candidate library generated in single-round screening and collect more information for aptamer truncation and optimization, the present invention adopts a machine learning clustering method to identify the core regions in each aptamer family. Subsequently, the present invention applies machine learning to the secondary structure analysis guided by core sequences and successfully discovers a common structure (i.e., the core structure) that can specifically bind to the CD8 protein. Using the common structure simplifies the truncation and optimization process of the aptamer candidate library and also promotes the design of new aptamers that can bind to the CD8 protein with high specificity and high affinity. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 Analysis of the results of single-round and multi-round screening of nucleic acid aptamers for CD8 protein. a shows the analysis of the binding of enriched libraries in different rounds to CD8-immobilized microspheres by flow cytometry. b shows the frequency distribution of different copy numbers of aptamers binding to target microspheres in enriched libraries in different rounds. c shows the correlation analysis of the nucleic acid aptamer copy numbers between different rounds (p < 0.001). d shows the correlation analysis of the 6-mer frequencies between the top 1,000 aptamers in the first round of screening and all sequences in the fifth round of screening.
[0060] Figure 2It is the UAE-Clustering model architecture. a shows that the model uses an autoencoder model to learn sequence information. The encoder part uses UNet to extract information, while the decoder part uses CNN to reconstruct data. The input of the model consists of the outer product of the matrix generated by one-hot encoding the sequence and its transposed matrix, and the pairing matrix representing the base pairing strength. By using the cross-entropy loss function to calculate the error between the original sequence and the reconstructed sequence, the model gradually learns and captures the deep information of the sequence after iterative training. b shows the clustering results of the top 1000 aptamers in the low-dimensional space before single-round screening. c shows the multiple sequence alignment after clustering the first-round sequencing results and visualizing the classification based on the local feature patterns of the sequences using WebLogo for the deep learning clustering method.
[0061] Figure 3 To analyze the results of the CD8 protein Re-SILEX screening. a shows that seqLogo illustrates the conserved sequences of the aptamers in the families selected by Re-SILEX and the distribution of adjacent bases. b shows the secondary structure feature learning method guided by the core sequence. c shows the proportions of different secondary structures formed by the core sequence. d shows the general core structure model of the CD8-targeting aptamers. The core structure is determined by calculating the number of substructures in the smallest structure containing the fixed sequence, determining the lengths of the most common stems and loops, and analyzing the base frequencies at each position in these substructures. e and f show the truncation and mutation optimization of the CD8 protein aptamers based on the core structure.
[0062] Figure 4 It is the flow cytometry binding analysis of the truncated aptamers from the Re-SILEX screening.
[0063] Figure 5 It is the SPR characterization of the truncated and optimized aptamers targeting the CD8 protein.
[0064] Figure 6 They are the truncated and optimized aptamers targeting the CD8 protein. a shows the optimization of the nucleic acid aptamer CD8ReSI-3a to CD8ReSI-3b. b shows the optimization of the nucleic acid aptamer CD8ReSI-5a to CD8ReSI-5b.
[0065] Figure 7Secondary structure analysis of CD8 protein aptamers selected by single-round SILEX screening. a is the statistical analysis of the distribution of various secondary structures formed by the core sequence. b is the calculation of the core structure obtained from the base pattern around the core sequence in Structure I, accounting for 40.6%. c is the calculation of the core structure obtained from the base pattern around the core sequence in Structure II, accounting for 3.8%. d is the statistical analysis of the base pattern at common positions around the core sequence in the common structure. e and f are the molecular optimization based on the core structure. g is the analysis and design using the core structure as a template. h is the cleavage of the CD8ReSI-5a aptamer previously obtained from Re-SILEX into two segments, and it is found that the mixture of these two aptamers can bind to the CD8 protein.
[0066] Figure 8 Manual folding of the structures of six original sequences in single-round screening. a-f are CD8SI-1, CD8SI-2, CD8SI-3, CD8SI-4, CD8SI-5, and CD8SI-6 respectively.
[0067] Figure 9 Truncated and optimized aptamers (a-e) from single-round screening and de novo aptamer CD8DES2 (f) designed based on the core structure.
[0068] Figure 10 Flow cytometry binding analysis of six truncated aptamers from single-round screening with the positive target CD8 protein (a) and the negative control EGF protein (b).
[0069] Figure 11 SPR characterization of truncated and optimized aptamers targeting the CD8 protein.
[0070] Figure 12 Flow cytometry characterization of the binding ability of truncated aptamers to the CD8 protein.
[0071] Figure 13 For the truncated aptamer and CD8 + Flow cytometry characterization of the binding ability of truncated aptamers to CD8 Specific embodiments
[0072] The present invention will be further described in detail below in conjunction with specific embodiments. The provided embodiments are only for clarifying the present invention and not for limiting the scope of the present invention. The following provided embodiments can be used as a guide for those of ordinary skill in the art to make further improvements and do not constitute any limitation to the present invention in any way.
[0073] In the experimental methods of the following examples, unless otherwise specified, they are all conventional methods, carried out according to the techniques or conditions described in the literature in this field or according to the product instructions. The materials, reagents, etc. used in the following examples, unless otherwise specified, can all be obtained from commercial channels.
[0074] CD8 usually exists in the form of a dimer formed by α and β chains (CD8αβ) or two α chains (CD8αα). The CD8 protein mentioned in the following experiments is all CD8α protein.
[0075] The Re-SILEX screened CD8 aptamer sequences involved in the following examples are shown in Table 1, the SILEX screened CD8 aptamer sequences are shown in Table 2, and the sequences of the novel CD8 nucleic acid aptamers CD8DES1 and CD8DES2 designed according to the core structure obtained by the machine learning method are shown in Table 3.
[0076] Table 1. Re-SILEX screened CD8 aptamer sequences
[0077]
[0078] Table 2. SILEX screened CD8 aptamer sequences
[0079]
[0080] Table 3. Sequences of the novel CD8 nucleic acid aptamers CD8DES1 and CD8DES2 designed according to the core structure obtained by the machine learning method
[0081] Name Sequence (5’-3’) CD8DES1 TGGGGGCGGAAAACCGCGTGACCCAGCTTGAAATGGGCCCC(SEQ ID No.34) CD8DES2 TGGAGGAGGGTAACCTCGTGAGCGAGCTTGAAATCGCCTCC(SEQ ID No.35)
[0082] Example 1. Establishment of the machine learning method for the single-round screening sequence of the nucleic acid aptamer of the present invention and screening of the nucleic acid aptamer specifically binding to the CD8 molecule
[0083] 1. Experimental methods
[0084] 1. Single-round and multi-round screening
[0085] This part involves the following step S1:
[0086] S1. High-throughput sequencing is performed on the nucleic acid aptamer library obtained through single-round / multi-round screening to obtain sequencing data (S1).
[0087] In the present invention, a total of two screenings, SILEX and Re-SILEX, are involved. The main difference between the two screenings lies in the library and the primers used in amplification. Therefore, the screening methods are described uniformly.
[0088] The specific differences are as follows:
[0089] 1) Different libraries: The two screenings are divided into SILEX screening and Re - SILEX screening. The libraries used are LibZJ: 5'-ACCGACCGTGCTGGACTCtAt-N45-aCTATGAGCGAGCCTGGCGt-3' (N45 represents 45 consecutive Ns); and LibCN1: 5’-AAGGAGCAGCGTGGAGGATANNNNNNNNNNNNNNNNNNNNNAGCTTGAAANNNNNNNNNNNNNNNNNGCTTTAAGGCCGGTCCTAGCA-3’.
[0090] Among them, N is A or T or C or G.
[0091] 2) Different primers
[0092] SILEX screening:
[0093] PZJS(5’-3’): FAM - ACCGACCGTGCTGGACTCTA;
[0094] PZJAB(5’-3’): biotin - ACGCCAGGCTCGCTCATAGT.
[0095] Re - SILEX screening:
[0096] PTBS(5’-3’): FAM - AAGGAGCAGCGTGGAGGATA;
[0097] Cn - biotion(5’-3’): biotin - TGCTAGGACCGGCCTTAAAGC.
[0098] The target protein CD8 (SinoBiological, 10980 - H02H) and the control protein EGF (SinoBiological, 10605 - H01H) were conjugated to sugar beads (Cytiva, 17071601) according to the instructions for subsequent target protein screening and characterization.
[0099] The screening steps of the protein are roughly as follows: First, the initial random library is denatured and renatured (95 °C, 5 min; on ice, 5 min; room temperature, 15 min), and then incubated with the microspheres conjugated with the target molecule at room temperature for 60 min with shaking during the incubation. The incubated sample is passed through a precast column and washed twice with 200 μl of Washing Buffer (DPBS solution containing 5 mM MgCl2 and 1% Tween-80. The same applies hereinafter), and the protein glycan beads that have bound to the library are recovered using ddH2O. PCR amplification is performed using the recovered ssDNA as a template to prepare the next round of ssDNA library. Repeat the above screening steps, and add a negative screening step with the control protein EGF, gradually reduce the positive screening incubation time, and increase the number of washing steps before the positive screening of the target protein CD8α for screening.
[0100] The single-round library of the obtained nucleic acid aptamers is incubated with the target microspheres for 30 min, and the unbound sequences are removed by washing with Washing Buffer. Referring to our previous method [Wu, X., et al., Efficient Strategy to Discover DNA Aptamers Against Low Abundance Cell Surface Proteins in Scarce Samples. Journal of the American Chemical Society, 2024. 146(39): p. 26667-26675.], a different molecular identity card is linked to each of the single-round libraries of nucleic acid aptamers that bind to the target (Molecular identity card: CCTGTCTTGTCTGCCT XXXX nnnnnnnnnn nnnnnnnnnn nnnnnnnnnn ACCTCTCAGAATTCGCACCA, the underlined part is the tag part for labeling different cells. n is a or t or c or g. Different molecular identity cards label different libraries, as shown in Table 4) for high-throughput sequencing. Finally, the types and quantities of nucleic acid sequences that bind to the target molecule in each round are obtained using an analysis program based on molecular identity tags.
[0101] Table 4. Different libraries are labeled with different molecular identity cards
[0102] Name Label Library PTBHTS-9 cggT SILEX-R1 PTBHTS-10 aagT SILEX-R2、Re-SILEX-R1 PTBHTS-13 ttcA SILEX-R3、Re-SILEX-R2 PTBHTS-15 ttaG SILEX-R4、Re-SILEX-R3 PTBHTS-16 gcgC SILEX-R5、Re-SILEX-R4 PTBHTS-18 gtcT Re-SILEX-R5
[0103] 2. Raw data analysis
[0104] This part involves the following step S2:
[0105] S2. Preprocess the sequencing data (such as removing abnormal samples) to obtain nucleic acid aptamer data.
[0106] Among them, S2 includes:
[0107] S21. Calculate the abundance of each nucleic acid aptamer in different samples through molecular identity tags, exclude nucleic acid aptamers whose length does not meet the preset requirements (for example, the length is less than or equal to 20), and use the dynamic programming algorithm to remove the influence of primer dimers;
[0108] S22. Merge the nucleic acid aptamer data of each sample after filtering. If the nucleic acid aptamer does not exist in a certain sample, it is counted as 0, and the original data similar to genomic differential analysis is obtained. The genomic data is the number of genes corresponding to different samples, while the nucleic acid aptamer data is the abundance of nucleic acid aptamers corresponding to different samples;
[0109] S23. Save the result of S22 as an excel file, with each column corresponding to the nucleic acid aptamer, aptamer length, and aptamer abundance corresponding to each sample.
[0110] Primer dimers may produce a small number of base mutations during amplification. All dimers cannot be accurately found directly by the string matching method. Therefore, in this embodiment, the dynamic programming algorithm is used to calculate the longest common subsequence between the aptamer and the primer and remove it as the primer dimer. Among them, S21 includes:
[0111] S211. Establish a two-dimensional array (matrix) L, where L[i][j] represents the length of the longest common subsequence of the first i characters of the first sequence and the first j characters of the second sequence;
[0112] S212. Initialize the first row and the first column to 0 because the length of the longest common subsequence of any sequence and an empty sequence is 0;
[0113] S213. Traverse the bases of the two sequences. If the characters are the same (X[i - 1] == Y[j - 1]), then L[i][j] = L[i - 1][j - 1] + 1. If the characters are different, take L[i][j] = max(L[i - 1][j], L[i][j - 1]);
[0114] S214. The value in the lower right corner L[m][n] of the matrix is the length of the longest common subsequence of the two sequences;
[0115] S215. Calculate the ratio of the longest common subsequence to the primer length. The preset threshold in this embodiment is 0.9. If the ratio of the longest common subsequence to the primer length is greater than 0.9, it is removed as the primer dimer.
[0116] This step can be performed on high-throughput sequencing data: (a) The sequencing data is used to deduplicate the library based on the unique molecular identifiers carried by each aptamer. (b) First, the forward and reverse primers used in sequencing are removed. Then, the library sequence to which it belongs is determined according to the primers for screening the library. Finally, the abundance of aptamers in each round of the library is determined according to the UMI sequences used in each round of enriched library. (c) Filter aptamer sequences with abnormal lengths, such as primer dimers and overly short sequences that only contain the forward primer. (d) After filtering, in order to show the changes in aptamers in each round, the aptamer abundance and aptamer motifs are used respectively to demonstrate the efficiency of screening. Then, nucleic acid aptamer structure analysis is carried out. Secondary structure prediction is performed on a large number of aptamers obtained by Re-SILEX to obtain the positions of the core sequence and core structure in the original sequence, and the number of structures such as stem-loops in the core structure is calculated respectively. Further calculate the base distribution frequencies of various circular structures and paired positions in the complete structure containing the core structure. Finally, the secondary structure of the nucleic acid aptamer with the highest probability, that is, the most likely to bind to the target, is obtained from a large number of structures.
[0117] 3. UAE-Clustering Dimensionality Reduction and Clustering
[0118] This part involves the following steps S3 and S4:
[0119] S3: Input the aptamer data of multiple nucleic acid aptamers (in this embodiment, the first 1000 nucleic acid aptamers with descending abundance in the nucleic acid aptamer data) that meet the preset abundance requirements in the nucleic acid aptamer data into a deep learning model for training to map the sequences of multiple nucleic acid aptamers to a low-dimensional representation in the latent space.
[0120] S4: Based on the trained deep learning model to be used, use the encoder to remap the sequences of multiple nucleic acid aptamers to the latent space, and use the Gaussian mixture model to cluster the representations in the latent space, so as to divide the sequences of multiple nucleic acid aptamers into different families. In this embodiment, 10 families with different distribution characteristics are finally identified.
[0121] Among them, S3 includes:
[0122] S31: Map the 'ATGC' bases of each nucleic acid aptamer in the nucleic acid aptamer data to a digital sequence represented by '0123', with a size of 1000*L. At the same time, divide the matrix into a training set and a test set in a ratio of 6:4 for training and verification of the deep learning model;
[0123] S32: The model input is a digital sequence. The digital sequence is encoded into an L * 4 matrix using one-hot encoding, and the outer product of this matrix and its transposed matrix is taken to obtain an image representation with 16 channels and a size of L * L. This matrix can describe all possible base pairing patterns at each position in the sequence. To avoid the sparsity caused by converting the sequence into 16 channels, the method of splicing the base pairing score matrix is used for optimization, and finally an output matrix with 17 channels is obtained. Among them, the base pairing score represents the pairing strength between different bases. In this embodiment, the GC pairing score is set to 3, the AT pairing score is set to 2, the weakest TG pairing is set to 0.8, and the rest are set to 0;
[0124] S33. The deep learning model uses an autoencoder to learn the sequence rules by reducing the cross-entropy loss between the reconstructed sequence and the real sequence. The regularization expression is:
[0125]
[0126] In the formula, L represents the length of the sequence, N represents the number of base types (for example, in DNA, it is 4, representing A, T, C, and G respectively), X ij represents the one-hot encoding value (0 or 1) of the base j at the i-th position in the original sequence, represents the predicted probability value of the base j at the i-th position in the reconstructed sequence.
[0127] Among them, S33 includes:
[0128] S331. The encoder of the autoencoder is a U-Net. Taking the 17-channel output matrix as the encoder input, it undergoes four downsamplings and upsamplings in the U-Net to extract local features, and the output is a two-dimensional matrix containing two features representing the position distribution of the sequence in the latent space;
[0129] S332. The decoder of the autoencoder uses a residual network to reconstruct the sequence in the latent space. The decoder output is an L * 4 matrix to facilitate the calculation of the cross-entropy between the model and the input digital sequence.
[0130] When the Gaussian Mixture Model (GMM) is used for clustering, it considers the possibility of data points belonging to each cluster based on a probability framework. During the clustering process, each cluster is assumed to be a Gaussian distribution, and the entire data set is composed of the mixture of these Gaussian distributions. Each Gaussian distribution defines a cluster, with its own mean (center) and covariance (shape, direction, size), as well as a mixing weight representing the proportion of this distribution in the overall data. To cluster the data, the present invention uses the Gaussian Mixture Model to perform 100 iterations on the encoder sample data, thereby obtaining the optimal model parameters. Based on these parameters, the present invention predicts the category to which each data belongs.
[0131] 4. Nucleic acid aptamer structure analysis
[0132] This part involves the following steps S5 - S10:
[0133] S5. According to the characteristics of the sequences of nucleic acid aptamers in each family, obtain the sequence logo graph of each family. Among them, the high - frequency base sequence in the sequence logo graph of each family is the common rule of that family.
[0134] S6. According to the size of each family and its common rule, calculate the proportion of the sequence rule of each family, and obtain the high - proportion common rule of each family as the core sequence related to binding in each family.
[0135] S7. Use a nucleic acid secondary structure prediction model (such as the MXfold2 model) to predict the secondary structure of the sequences of nucleic acid aptamers in each family, and obtain a secondary structure text file for each family containing the sequence, secondary structure, and folding score.
[0136] S8. For each family, according to its secondary structure text file, save the sequence and secondary structure into the first list respectively, calculate the region of the core sequence in each nucleic acid aptamer, calculate the smallest sub - structure in the secondary structure that contains this region, and save the smallest sub - structure and its sequence into the second list.
[0137] S9. Calculate the number of secondary structure elements (such as hairpin, multi - branch, bulge, internal loop, etc. structures) of each sub - structure in the second list through the first recursive algorithm, count the most frequently occurring structure combination as the core structure, and calculate the base frequency of each position containing the core sequence in the sequence list of each family through the second recursive algorithm to obtain the final structure representation of each family representing high - frequency bases.
[0138] S10. Use the final structure representation of each family representing high - frequency bases to optimize the nucleic acid aptamer sequences of each family. For example, it can be used to guide sequence truncation, mutation with sequences that do not contain this structure, or design of new aptamers, etc.
[0139] Among them, S5 includes:
[0140] Use a sequence logo graph generation tool (such as WebLogo3) to obtain the sequence logo graph of each family;
[0141] When the number of families is not enough to complete the classification of all nucleic acid aptamer sequences, combine the multiple sequence alignment method to align the families that cannot be classified.
[0142] Among them, S7 includes:
[0143] The folding scores calculated by the deep neural network are combined with Turner's nearest neighbor free energy parameters through a nucleic acid secondary structure prediction model, and thermodynamic regularization is used to prevent significant deviation of the folding scores of the secondary structure from the thermodynamic parameters.
[0144] Specifically, a large number of aptamers obtained by Re-SILEX need to use a large-scale secondary structure prediction tool to analyze the distribution pattern of the core structure. The MXfold2 secondary structure prediction model uses a deep neural network to calculate four folding scores for each pair of nucleotides, and these scores are used to evaluate the scores of the nearest neighbor loops. Similar to MXfold [Akiyama, M., K. Sato, and Y. Sakakibara, A max-margin training of RNA secondary structure prediction integrated with the thermodynamic model. Journal of Bioinformatics and Computational Biology, 2018. 16(06): p. 1840025.], the model combines the folding scores calculated by the deep neural network with Turner's nearest neighbor free energy parameters. Then, a Zuker-style dynamic programming (DP) algorithm is used to predict the optimal secondary structure to maximize the sum of the nearest neighbor loop scores. The deep neural network is trained through a maximum margin framework (structured support vector machine, SSVM) to minimize the structured hinge loss function in a thermodynamically regularized manner, thereby preventing significant deviation of the folding scores of the secondary structure from the free energy of the thermodynamic parameters.
[0145] After obtaining the secondary structure rules, a recursive algorithm is further used to calculate the number of hairpin, multi-branch, bulge, and internal loop structures of each sub-structure in the new structure list. Subsequently, the base distribution frequencies of various loop structures and pairing positions in the complete structure containing the core are calculated. Finally, based on the results with the highest probability among the analyzed numerous structures, the secondary structure of the nucleic acid aptamer most likely to bind to the target is identified.
[0146] 5. Flow cytometry characterization of the binding of nucleic acid aptamers to microglobulin
[0147] In a 100 μl reaction system, microspheres conjugated with target molecules (3 μg of the target) are incubated with a 200 nM nucleic acid aptamer enrichment library or a nucleic acid aptamer labeled with a fluorescent molecule at room temperature for 30 min, washed twice with 200 μl of Washing Buffer, and then the samples resuspended in Washing Buffer are tested using a CytoFLEX LX flow cytometer (Beckman), and the experimental data are processed using the software FlowJo_V10.
[0148] 6. Flow Cytometry Characterization of the Binding of Nucleic Acid Aptamers to Cells
[0149] Jurkat cells: A cell line derived from human T lymphocyte leukemia cells, belonging to the helper T cell subset of T cells, and serving as a CD8-negative cell model.
[0150] The test target cells are the following two types: CD8-positive Jurkat cells (transgenic cells expressing CD8 obtained by introducing the CD8 protein-encoding gene into Jurkat cells) and CD8-negative Jurkat cells (untransfected Jurkat cells).
[0151] Wash the target cells 2 - 3 times with DPBS. Use enzyme-free digestion solution to digest the cells from the culture dish and wash them twice with Washing Buffer. Finally, resuspend them in Binding Buffer (20 μL of 5 mg / mL HSDNA; 100 μL of 10 mg / mL BSA; 880 μL of Washing Buffer). Incubate the target cells with 200 nM of nucleic acid aptamers labeled with fluorescent molecules in a 100 μl reaction system at 4°C for 30 min. Wash twice with 200 μl of Washing Buffer. Subsequently, test the sample resuspended in 200 μl of Washing Buffer using a CytoFLEX LX flow cytometer (Beckman), and process the experimental data using the software FlowJo_V10.
[0152] 7. Surface Plasmon Resonance (SPR) Test Experiment
[0153] According to the instructions, couple the target protein and the control protein to a CM5 chip (Cytiva, 29104988). Then pass 500 nM of nucleic acid aptamers through the chip, with a binding time of 120 s, a dissociation time of 180 s, and regeneration with 1.5 M NaCl at a flow rate of 30 μL / min.
[0154] To measure the affinity of the nucleic acid aptamers, perform serial dilutions with SPR buffer (DPBS solution containing 5 mM MgCl2) at concentrations of 800 nM, 600 nM, 500 nM, 400 nM, 200 nM, 100 nM, and 50 nM in sequence. Inject the samples and detect the binding to the target protein using a surface plasmon resonance instrument. Fit the data using the Langmuir isothermal adsorption model.
[0155]
[0156] where R is the detected response, [L] is the free ligand concentration, R max is the maximum response, and K dis the dissociation constant.
[0157] II. Results and Analysis
[0158] 1. Single-round and multi-round screening of CD8 protein nucleic acid aptamers based on the molecular identity tag strategy
[0159] First, in order to study the variation law of aptamers in different screening rounds, the present invention used a single-stranded DNA library (LibZJ) containing 2 nmol of a 45-base random region (about 10 15 variants) to perform five rounds of screening to obtain aptamers targeting human recombinant CD8 protein. The screening pressure was gradually increased during the five rounds of screening, and the screening process was monitored by flow cytometry ( Figure 1 in a). As Figure 1 shown in a, there was an obvious binding between the library of the third round of screening and CD8 protein, and there was a large displacement between the enriched library of the 5th round and CD8. The enriched libraries of each round did not bind to the control protein EGF. Subsequently, the enriched libraries from each round of screening and the library bound to the target microspheres were subjected to high-throughput sequencing using the molecular identity (UMI) tag strategy.
[0160] By analyzing the result data, we found that the sequencing of the enriched library binding to the microsphere target protein had fewer primer dimers than the sequencing of directly binding to the protein under the same conditions, which might be due to the generation of primer dimers during the PCR amplification process. Since primer dimers do not bind to the microsphere target molecules, the enriched library binding to the microsphere target molecules can more truly reflect the situation of potential nucleic acid aptamers. We analyzed the aptamer copy number frequency of the enriched libraries of different rounds binding to the target microspheres. As Figure 1 shown in b, with the increase of the screening rounds, the copy number of aptamers binding to the target microspheres increased, while the variety of unique aptamers decreased.
[0161] We analyzed the correlation between the copy number change and the K-mer frequency distribution in different rounds, as Figure 1As shown in Figure c, a significant positive correlation was found between the number of aptamer copies per round (P value < 0.001). According to previous studies [Song, J., et al., A Sequential Multidimensional Analysis Algorithm for Aptamer Identification based on Structure Analysis and Machine Learning. Analytical Chemistry, 2020. 92(4): p. 3307-3314.] and the analysis of different K-mers, 6-mers were superior to other lengths in the identification of target-binding aptamers. We further analyzed the characteristics of 1000 aptamer sequences screened in the first round and all aptamer sequences in the fifth-round enriched library using 6-mers. As Figure 1 shown in Figure d, the characteristics of the two screening libraries were highly similar (correlation: 0.99), indicating that the first-round screening library labeled with UMI already contained information from the multi-round screening library, which was consistent with the results we obtained in previous cell experiments [Zhang, D., et al., Streamlining RNA Aptamer Selection via Unique Molecular Identifiers and High-Throughput Sequencing. Analytical Chemistry, 2024. 96(42): p. 16686-16694. Wu, X., et al., Efficient Strategy to Discover DNA Aptamers Against Low Abundance Cell Surface Proteins in Scarce Samples. Journal of the American Chemical Society, 2024. 146(39): p. 26667-26675.].
[0162] 2. Sequence family analysis of the CD8α protein nucleic acid aptamer library by single-round SILEX screening
[0163] To determine the core sequences in the single-round selection library, the present invention attempts to use deep learning methods to learn the distribution law of the single-round SILEX aptamer family, and classifies the aptamer library into multiple families with similar characteristics rather than just sequence similarity. Past DNA / RNA sequence deep learning [Alipanahi, B., et al., Predicting the sequence specificities of DNA- and RNA-binding proteins by deep learning. Nature Biotechnology, 2015. 33(8): p. 831-838. Smith, C.J., et al., Enabling large-scale genome editing at repetitive elements by reducing DNA nicking. Nucleic Acids Res, 2020. 48(9): p. 5183-5195. Angermueller, C., et al., DeepCpG: accurate prediction of single-cell DNA methylation states using deep learning. Genome Biology, 2017. 18(1): p. 67.] methods used one-hot encoding as input. We adopted the encoding method in UFold [Fu, L., et al., UFold: fast and accurate RNA secondary structure prediction with deep learning. Nucleic Acids Res, 2022. 50(3): p. e14.], and used more detailed base pairing information as the model input, which fully considered all possible distal base interactions. To utilize this information to classify sequences with similar structures, we proposed a deep learning method UAE-Clustering( Figure 2 as shown in a) of Figure 2 ). Similar to the RaptGen [Iwano, N., et al., Generative aptamer discovery using RaptGen. Nat Comput Sci, 2022. 2(6): p. 378-386.] model that uses low-dimensional information in the sequence latent variable space for sequence generation, UAE-Clustering learns the core sequence information of the first-round selection sequences and uses a Gaussian mixture model in the latent variable space to achieve sequence classification(
[0164] Direct multiple sequence alignment cannot directly obtain its homologous sequences or related information, indicating that the homology analysis of nucleic acid sequences cannot directly use the existing multiple sequence alignment methods in bioinformatics. We further preliminarily classified all sequences into ten different families, used the Clustal Omega [Madeira,F.,et al.,Search and sequence analysistools services from EMBL-EBI in 2022.Nucleic acids research,2022.50(W1):p.W276-W279.] multiple sequence alignment method to obtain detailed information on each homologous family, and visually displayed the information through WebLogo [Schneider,T.D.and R.M.Stephens,Sequence logos:a new way to display consensussequences.Nucleic Acids Res,1990.18(20):p.6097-100.Crooks,G.E.,et al.,WebLogo:a sequence logo generator.Genome Res,2004.14(6):p.1188-90.], that is, the position distribution of the core sequence GTGAGGAGCTTGAAA in each family ( Figure 2 c) in. The same core sequence position distribution as in the fifth round was already present in the first round, which further indicates that the relevant information of the nucleic acid aptamer sequence binding to the target molecule already exists in one round.
[0165] Due to the diversity of aptamer sequences, simple clustering analysis of motif sequence similarity is not sufficient to reveal its deep patterns. Therefore, for the first time, we mined potential features through a deep learning model to achieve family classification based on local sequence patterns.
[0166] 3. Verification of the core sequence using Re-SILEX
[0167] To explore whether the common region of single-round SILEX is the core sequence affecting target binding, the present invention designed a random library (LibCN1) containing a partial core region (AGCTTGAAA) for single-round and multi-round screening.
[0168] The screening steps were the same as before, and the binding of each round of aptamer pool was monitored by flow cytometry and gel electrophoresis. It was found that significant displacement occurred in the third round of CD8 microbead screening, while no binding was observed on the control EGF protein. We performed sequencing analysis on aptamer pools from different rounds using a unique molecular identity tagging strategy. Different from the data obtained from previous screening, the raw data of the new library contained a large number of unique aptamers (252,650 sequences, while 272,462 sequences were selected previously), the aptamer enrichment of the fifth-round library was lower (633 copy numbers, while 44,865 copy numbers in the previous screening), and the highest abundance was only in the hundreds, but the flow cytometry results showed better binding than the first screening. This phenomenon may be due to the fact that in the LibCN1 screening library, due to the presence of a core sequence, the initial library (about 10 15 different molecules) contained more sequences with binding ability, while in the completely random screening library (about 10 15 different molecules), the initial sequences containing the core binding sequence were limited.
[0169] 4. Machine learning of the secondary structure of single-round Re-SILEX aptamers mediated by the core sequence
[0170] To obtain the regulatory region around the fixed sequence in Re-SELEX screening, the present invention applied the same method as the above family sequence analysis. First, the sequences were aligned, and then the related family sequences at the same position as the core sequence were obtained ([[]] Figure 3 Figure 3 a) in). The finally obtained regulatory region sequence was CGTGA-NNN-AGCTTGAAA, which was highly similar to the core sequence obtained by single-round screening.
[0171] The binding of the target depends not only on the direct interaction between the core sequence and the molecular surface, but also involves specific structures formed by other bases through mechanisms such as base pairing. These structures help to maintain the specific interaction between the aptamer and the target. However, due to the diversity of nucleic acid secondary structures, it is difficult to accurately predict the specific secondary structure when binding to the target molecule by traditional methods. To analyze this process more precisely, the present invention developed a machine learning algorithm for the secondary structure of single-round aptamers mediated by the core sequence. The present invention used MXfold2 [Sato, K., M. Akiyama, and Y. Sakakibara, RNA secondary structure prediction using deep learning with thermodynamic integration. Nature Communications, 2021. 12(1): p. 941.] to perform large-scale structure prediction on Re-SILEX aptamers, and further used a structure analysis algorithm to learn the characteristics of the local secondary structure of the core region ([[]] Figure 3Figure 3 In b), it includes stem-loops and multi-branched loops. The formation of these specific conformations may play a crucial regulatory role in the structural stability of the core sequence and its binding specificity to the target. As Figure 3 shown in c), the number of sequences with a stem-loop secondary structure in the fixed region AGCTTGAAA is 24,867, accounting for 62.4% of the aptamers (39,758). The remaining 37.6% of the aptamer fixed regions cannot form stem-loops. Further analyzing the structures formed by other core regions in the aptamers that form stem-loops in the fixed base region, it is found that the number of GTGA in multi-branches is 13,711, accounting for 55.2%, while the number in the stem structure is 11,301. According to the most prevalent structure (multi-branches), we counted the length of each stem-loop in this structure and analyzed the base frequencies at each position. Therefore, by applying machine learning to the secondary structures of 39,758 different sequences, we hypothesized that they could form a common secondary structure ( Figure 3 shown in d).
[0172] 5. Secondary Structure Analysis and Truncation Optimization of CD8 Protein Re-SILEX Nucleic Acid Aptamers
[0173] To verify the correctness of the general structure, we analyzed the secondary structures of representative aptamers and performed truncation and optimization according to the predicted structures. According to the sequence copy number analysis, five aptamer candidates with the highest copy numbers obtained from Re-SILEX, CD8ReSI-1, CD8ReSI-2, CD8ReSI-3, CD8ReSI-4, CD8ReSI-5 (sequences are shown in Table 1), and the relatively later sequences CD8ReSI-38, CD8ReSI-43 (sequences are shown in Table 1) were subjected to secondary structure prediction. Among the three structures obtained for each aptamer, only CD8ReSI-4, CD8ReSI-5, and CD8ReSI-43 had the Figure 3 predicted core structure shown in d), while the remaining aptamers could not obtain the expected structure. We hypothesized that the function of the aptamer is closely related to the core secondary structure it forms. To verify the correctness of our above hypothesis, we forced it to form the secondary structure we predicted and then performed truncation optimization. Subsequently, the aptamers were truncated according to the predicted structures and FAM-labeled at the 5′ end. Then, flow cytometry or surface plasmon resonance (SPR) was used to detect the binding performance. The analysis by flow cytometry showed that the aptamers CD8ReSI-1a, CD8ReSI-2a, CD8ReSI-3a, CD8ReSI-3b, CD8ReSI-4a, CD8ReSI-5a, CD8ReSI-5b, CD8ReSI-38a, CD8ReSI-43a (sequences are shown in Table 1) could all bind to the CD8α protein ( Figure 4)。We further used surface plasmon resonance (SPR) technology to measure the binding and dissociation constants of the aptamers. The results showed that the truncated structures could bind to CD8 protein, and the affinities were all at the nanomolar level ( Figure 5 )。
[0174] To further verify the accuracy of the predicted secondary structure. For example, in CD8ReSI-2a (21.5 ± 0.08 nM, Figure 3 in e), the bases of A 11 T 12 A 13 were deleted to extend the backbone and truncated after positions A8 and C 52 . This optimization increased the affinity of the modified aptamer CD8ReSI-2b (sequence see Table 1) to 2.7 ± 0.02 nM. In CD8ReSI-3a (6.5 ± 0.05 nM, Figure 6 in a), the guanine at position G 42 was mutated to cytosine while retaining five pairs of bases in the backbone. This optimization increased the affinity of the modified aptamer CD8ReSI-3b (sequence see Table 1) to 3.2 ± 0.04 nM. In CD8ReSI-5a (2.3 ± 0.02 nM, Figure 6 in b), the stem I was truncated by three pairs of bases to obtain the aptamer CD8ReSI-5b (sequence see Table 1), and its affinity remained at 1.8 ± 0.02 nM.
[0175] At the same time, CD8ReSI-38 could not form the three-branch structure of the core structure whether using the structure prediction tool [Lorenz, R., et al., ViennaRNA Package 2.0. Algorithms for Molecular Biology, 2011. 6(1): p. 26.][Markham, N.R. and M. Zuker, UNAFold, in Bioinformatics: Structure, Function and Applications, J.M. Keith, Editor. 2008, Humana Press: Totowa, NJ. p. 3 - 31.] or manual folding. We truncated a section of the aptamer at the 3' end in an attempt to break up the pairing of the redundant branches on the right side and force the formation of a three-branch ( Figure 3 CD8ReSI-38a in f, sequence see Table 1), but the result was different from the expectation. The right part of the core sequence still underwent self-folding, and the GTGA part was still not unraveled. Flow cytometry and SPR results both showed weak binding ability (K D = 38.2 ± 1.56 nM, Figure 5In CD8ReSI-38a). Then, the aptamer CD8ReSI-38a (sequence see Table 1) by removing A1G2C3A4G5C6 and G 57 T 58 G 59 C 60 T 61 the bases at the positions, truncating the loop after the positions of G 51 and C 46 as the new 5' and 3' ends, and connecting G 56 to G7 to optimize its core structure. This modification restored the multi-branched structure of the GTGA region ( Figure 3 in CD8ReSI-38c in f, sequence see Table 1). After further experimental verification, the binding affinity was improved. In addition, changing the T4 and T5 bases in CD8-ReSI38b (sequence see Table 1) to enhance the structural stability further improved the affinity, from 12.7 ± 0.08 nM ( Figure 5 in CD8ReSI-38b) to 3.1 ± 0.02 nM ( Figure 5 in CD8ReSI-38c, sequence see Table 1). The above results indicate that some predicted aptamer stable secondary structures may not be the same as the real structures when binding to the target. Therefore, relying solely on the secondary structure prediction of a few aptamers cannot accurately reflect the real structural information.
[0176] 6. Machine learning of the secondary structure of single-round aptamers based on the core sequence
[0177] To verify whether the same structures as those obtained in Re-SILEX also exist in the sequences obtained by single-round screening, the present invention performed secondary structure prediction on the nucleic acid aptamers obtained by single-round SILEX screening. Due to the strong randomness of the library, some bases in the core sequence may not be optimal. The present invention grouped the aptamers with mutations in individual bases into one category, and after filtering the first 1000 aptamers in the first round of analysis, 770 aptamers with core sequences were obtained. As Figure 7 shown in a, there are 342 aptamers whose core sequences can fold into multi-branched and stem-loop structures. Also due to the random library, the positions of the core sequences are at different positions in the aptamers, resulting in two forms of secondary structure folding. Through machine learning, we found that Structure I ( Figure 7 in b) is very similar to the Re-SILEX structure. At the same time, although its core sequence is closer to the 5' end, Structure II ( Figure 7 in c) can still form a similar secondary structure. Subsequently, we analyzed the base pairing in the backbone regions of the two structures, as well as the base frequency distribution at each position in the loop structure of the core region, and formed a common structure ( Figure 7 in d).
[0178] To prove the correctness of the predicted structure, we manually folded the top 6 aptamers CD8SI-1, CD8SI-2, CD8SI-3, CD8SI-4, CD8SI-5, CD8SI-6 (sequences are shown in Table 2) from single-round screening ( Figure 8 ), further truncated them (for the sequences of the aptamers obtained after truncation, see Table 2, and the secondary structures are as shown in e and f in Figure 7 and a - e in Figure 9 ) and verified their binding specificities using flow cytometry or SPR. The experimental results showed that they all specifically bound to CD8 protein but not to the control protein ( Figure 10 ) and the affinities obtained by SPR were all at the nanomolar level ( Figure 11 ).
[0179] Based on the structural information, we modified CD8SI-1a to obtain CD8SI-1b by excising the bases at positions A 43 and C 44 , connecting the starting base A1 with the terminating base T 52 , and mutating the base at T50. The results showed that the affinity of CD8SI-1b (2.3 ± 0.01 nM) was significantly higher than that of CD8SI-1a (10.1 ± 0.04 nM) ( Figure 11 ). Using a similar method, by deleting the bases at positions A 32 C 33 G 34 T 35 C 36 T 37 G 38 G 39 and connecting the base at T1 with the base A 52 , we reshaped CD8SI-3a to obtain CD8SI-3b. The affinity of CD8SI-3b (6.5 ± 0.05 nM) was better than that of CD8SI-3a (12.8 ± 0.14 nM) ( Figure 11 ). This indicates that the aptamers binding to CD8 rely on the backbone - loop and multi - branched structures formed by the core sequence, and the complementary pairing in the backbone region is crucial for maintaining the structural stability. The binding experimental results of the single - round screening aptamers show that the factors affecting their binding are not only the core sequence but also the core structure formed jointly with the surrounding bases. Combining with the structural analysis algorithm of machine learning, the conformational distribution law of binding to the target can be extracted from single - round screening, so as to optimize the aptamers and improve their affinities.
[0180] Inspired by Structure I and Structure II, the CD8ReSI-5a nucleic acid aptamer was cleaved into CD8ReSI-5aa and CD8ReSI-5ab( Figure 7In part h), it was found that they could not bind to the CD8 protein when they existed alone, and only when both of them existed could they bind to the target molecule ( Figure 12 ), and this method is expected to be used to construct a biosensor for detecting the CD8 protein. Combining the structural rules obtained from the above calculations, we redesigned two aptamers, CD8DES1 and CD8DES2, which had never appeared in the screening library (the sequences are shown in Table 3) ( Figure 7 in part g, Figure 9 in part f), and used surface plasmon resonance (SPR) to detect their binding ability. The results showed that both of these aptamers could bind to the CD8 protein, and the KD values were 26.4 ± 0.12 nM and 23 ± 0.08 nM respectively ( Figure 11 ). This finding further demonstrated that the secondary structures obtained from our Re-SILEX and SILEX screens (as shown in Figure 3 part d) were indeed effective structures for binding to the CD8 protein, verifying the accuracy and effectiveness of our structural analysis and design methods.
[0181] So far, the core structure contained in the nucleic acid aptamer targeting the CD8 molecule required to be protected by the present invention has been obtained; the core structure is similar to a Y shape and is composed of two stem-loops, a multi-branched loop, and a main stem; the multi-branched loop is located between the two stem-loops;
[0182] Among the two stem-loops, the stem-loop close to the 5'-end is denoted as stem-loop 1, which is composed of a random base sequence, and the stem-loop close to the 3'-end is denoted as stem-loop 2. The stem region of the stem-loop 2 is composed of GC paired bases, and its loop structure is formed by the fixed sequence AGCTTGAAAT;
[0183] The multi-branched loop has the fixed base GTGA; on the multi-branched loop, there is a branch at the 5'-end adjacent to the fixed base GTGA, denoted as branch 1, and there is a branch at the 3'-end adjacent to the fixed base GTGA, denoted as branch 2. There is also a branch 3 at the position opposite to the fixed base GTGA between branch 1 and branch 2;
[0184] Branch 1 is connected to the end of the stem away from the loop in stem-loop 1; branch 2 is connected to the end of the stem away from the loop in stem-loop 2; branch 3 is connected to the main stem;
[0185] The main stem is composed of random base pairs.
[0186] 7. Verification of the CD8 aptamer at the cellular level
[0187] To investigate the binding ability of the optimized nucleic acid aptamers at the cellular level, the present invention evaluated their binding ability to CD8-positive expressing cells. Compared with recombinant proteins, the physicochemical environment on the cell membrane surface is more complex, posing higher requirements for the specificity and affinity of nucleic acid aptamers to target molecules. The present invention detected the binding of aptamers to cells by flow cytometry. The results showed that all nucleic acid aptamers that bind to CD8 protein could specifically bind to CD8-positive T cells and would not bind to CD8-negative cells ( Figure 13 in a, b). On this basis, we selected three (CD8SI-1b, CD8ReSI-38b, CD8DES2) from the SILEX library, the Re-SILEX library, and the manually designed aptamers that have been verified to bind to proteins, and further used flow cytometry to verify the binding affinity of these aptamers to CD8-positive T cells. The results showed that the affinities of CD8SI-1b, CD8ReSI-38b, and CD8DES2 were 0.20±0.02 nM, 2.99±0.20 nM, and 3.18±0.32 nM respectively ( Figure 13 in c-e), and they all had strong affinity for CD8-positive T cells. This experiment further proved the correctness of the aptamer secondary structure obtained by machine learning.
[0188] The present invention has been described in detail above. For those skilled in the art, without departing from the spirit and scope of the present invention and without unnecessary experiments, the present invention can be implemented within a wide range under equivalent parameters, concentrations, and conditions. Although specific embodiments of the present invention are given, it should be understood that the present invention can be further improved. In short, according to the principle of the present invention, this application intends to cover any modifications, uses, or improvements of the present invention, including changes made using conventional techniques known in the art that depart from the scope disclosed in this application.
Claims
1. A nucleic acid aptamer targeting CD8 molecules, characterized in that: The nucleic acid aptamer comprises a core structure; The core structure is composed of two stem loops, a multi-branched loop and a main stem; the multi-branched loop is located between the two stem loops; Of the two stem loops, the stem loop near the 5' end is recorded as stem loop 1, which is composed of a random base sequence, and the stem loop near the 3' end is recorded as stem loop 2, the stem region of the stem loop 2 is composed of GC paired bases, and its loop structure is formed by a fixed sequence AGCTTGAAAT; The multi-branched ring has a fixed base GTGA; on the multi-branched ring, there is a branch at the 5' end adjacent to the fixed base GTGA, recorded as branch 1, and there is a branch at the 3' end adjacent to the fixed base GTGA, recorded as branch 2, and there is a branch 3 located between the branch 1 and the branch 2 and opposite to the fixed base GTGA; The branch 1 is connected to the end of the stem away from the ring in the stem loop 1; the branch 2 is connected to the end of the stem away from the ring in the stem loop 2; the branch 3 is connected to the main stem; The main stem consists of random base pairing.
2. The nucleic acid aptamer according to claim 1, characterized in that: The nucleic acid aptamer comprises a core sequence; the sequence of the core sequence from the 5' end to the 3' end is GTGA-NNN-AGCTTGAAA or GTGA-NN-AGCTTGAAA; wherein N is A or T or C or G.
3. A nucleic acid aptamer targeting CD8 molecules, characterized in that: The nucleic acid aptamer is any of the following: (A1) a single-stranded DNA molecule shown in any one of SEQ ID No. 1 to SEQ ID No. 35; (A2) deleting or adding one or more nucleotides to the nucleic acid aptamer shown in (A1) to obtain a nucleic acid aptamer with the same function; (A3) Substituting or modifying the nucleotides of the nucleic acid aptamer shown in (A1) to obtain a nucleic acid aptamer having the same function; (A4) An aptamer having the same function as that obtained from an RNA molecule encoded by the aptamer shown in (A1).
4. A nucleic acid aptamer derivative, characterized in that: The nucleic acid aptamer derivative is any of the following: (B1) transforming the backbone of the nucleic acid aptamer described in any one of claims 1 to 3 into a thiophosphate backbone to obtain a derivative of the nucleic acid aptamer having the same function as the nucleic acid aptamer; (B2) a peptide nucleic acid encoded by the nucleic acid aptamer according to any one of claims 1 to 3, to obtain a derivative of the nucleic acid aptamer having the same function as the nucleic acid aptamer; (B3) Connecting a chemical group and / or fluorescein and / or anti-tumor drug and / or radioactive element and / or biological enzyme and / or biotin and / or nanomaterial to one end or the middle of the nucleic acid aptamer described in any one of claims 1 to 3 to obtain a derivative of the nucleic acid aptamer having the same function as the nucleic acid aptamer.
5. Use of the nucleic acid aptamer according to any one of claims 1 to 3 or the nucleic acid aptamer derivative according to claim 4 in any of the following: (C1) preparing a product for specifically recognizing CD8 molecules; (C2) specifically recognizes the CD8 molecule.
6. Use of the nucleic acid aptamer according to any one of claims 1 to 3 or the nucleic acid aptamer derivative according to claim 4 in any of the following: (D1) preparing a product for specifically binding to CD8 positive cells; (D2) Specific binding to CD8 positive cells.
7. Use of the nucleic acid aptamer according to any one of claims 1 to 3 or the nucleic acid aptamer derivative according to claim 4 in any of the following: (E1) preparing a product for detecting CD8 molecules or CD8 positive cells; (E2) Detection of CD8 molecules or CD8 positive cells.
8. Use of the nucleic acid aptamer according to any one of claims 1 to 3 or the nucleic acid aptamer derivative according to claim 4 in any of the following: (F1) preparing a product for capturing, enriching or purifying CD8 molecules or CD8-positive cells; (F2) Capturing, enriching or purifying CD8 molecules or CD8 positive cells.
9. A product, wherein the active ingredient is the nucleic acid aptamer according to any one of claims 1 to 3 or the nucleic acid aptamer derivative according to claim 4; The product has any of the following functions: (G1) specifically recognizes CD8 molecules; (G2) specifically binds to CD8-positive cells; (G3) Detection of CD8 molecules or CD8 positive cells; (G4) Capturing, enriching or purifying CD8 molecules or CD8 positive cells.
10. A method for obtaining a nucleic acid aptamer according to any one of claims 1 to 3, comprising the following steps: using CD8 as a target protein, obtaining a nucleic acid aptamer library by screening; then using the nucleic acid aptamer library to calculate the nucleic acid aptamer structure according to the method comprising the following steps S1 to S10, and then obtaining the nucleic acid aptamer according to any one of claims 1 to 3 according to the calculation results; S1. performing high-throughput sequencing on the nucleic acid aptamer library to obtain sequencing data; S2, preprocessing the sequencing data to obtain nucleic acid aptamer data; S3, inputting the nucleic acid aptamer data of multiple nucleic acid aptamers that meet the preset abundance requirements into the deep learning model for training, so as to map the sequences of the multiple nucleic acid aptamers to a low-dimensional representation in the latent space; S4. Based on the trained deep learning model, the encoder is used to remap the sequences of multiple nucleic acid aptamers to the latent space, and the Gaussian mixture model is used to cluster the representation of the latent space, so as to classify the sequences of multiple nucleic acid aptamers into different families. S5. According to the characteristics of the sequences of the nucleic acid aptamers of each family, a sequence logo of each family is obtained, wherein the high-frequency base sequences in the sequence logo of each family are the common rules of the family; S6. According to the size of each family and its common rules, the proportion of sequence rules of each family is calculated, and the high-proportion common rules of each family are obtained as the core sequence related to binding in each family; S7, using a nucleic acid secondary structure prediction model to predict the secondary structure of the nucleic acid aptamer sequence of each family, and obtaining a secondary structure text file containing the sequence, secondary structure and folding score of each family; S8. For each family, according to its secondary structure text file, the sequence and secondary structure are saved in the first list respectively, the region of the core sequence in each nucleic acid aptamer is calculated, and the minimum substructure containing the region in the secondary structure is calculated, and the minimum substructure and its sequence are saved in the second list; S9, calculating the number of secondary structure elements of each substructure in the second list by the first recursive algorithm, counting the structural combination with the highest occurrence frequency as the core structure, and calculating the base frequency of each position containing the core sequence in the sequence list of each family by the second recursive algorithm, and obtaining the final structural representation of the high-frequency bases representing each family; S10. Optimize the nucleic acid aptamer sequence of each family using the final structure representation of the high-frequency bases representing each family.