A design method employing computer-aided aptamer screening and tailoring

CN122575458APending Publication Date: 2026-08-14JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本发明待解决的技术问题为:其一,现有适配体筛选方法普遍需要多轮迭代富集,流程耗时、劳动强度高,且对靶标固定化与人工操作依赖较强,导致筛选效率与重复性易受实验条件波动影响

Benefits of technology

第一,针对现有筛选方法效率低、人工依赖强、结果不稳定的问题,本发明通过构建靶标分子的共有结构模型作为统一虚拟筛选模板,并结合多软件分子对接与共识评分策略,实现了高效、低偏倚的初步筛选。首先,对多种同系物靶标分子进行结构对齐与共有骨架提取,所构建的共有结构模型能够代表该类靶标的共同识别特征,从而在一次虚拟筛选中即可覆盖多个结构相近的靶标分子,避免了针对每个同系物单独筛选的繁琐,极大提升了筛选效率的潜力,并为获得广谱识别能力的适配体奠定了基础。其次,采用至少两种不同的分子对接软件进行模拟,并对不同软件的结果进行标准化处理与综合评分排序,这种多算法交叉验证与共识决策机制,有效降低了单一对接软件或评分模型因算法原理与参数差异带来的随机误差与系统偏倚,从而显著提高了虚拟筛选结果的鲁棒性、可靠性与可重复性,减少了因算法局限性导致的“假阳性”或“假阴性”筛选结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575458A_ABST
    Figure CN122575458A_ABST
Patent Text Reader

Abstract

This invention discloses a computer-aided aptamer screening and pruning design method, belonging to the fields of bioinformatics and molecular design. Targeting polychlorinated biphenyls (PCBs), this method constructs a common structural model as a unified virtual screening template; it uses single-base mutations to construct an initial library and pre-screens stable sequences based on thermodynamic parameters; it employs parallel simulations using at least two molecular docking software programs, standardizes and comprehensively scores the results to screen high-affinity candidate aptamers; finally, based on the identified key binding sites, it uses a fitness function combining binding length penalty and structural constraints, and achieves automated and rational pruning of aptamer sequences through a cascade optimization of genetic algorithms, particle swarm optimization, and simulated annealing algorithms. This invention significantly improves the screening efficiency and pruning accuracy of broad-spectrum aptamers, yielding short-sequence, high-affinity nucleic acid recognition elements suitable for rapid detection of environmental pollutants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics and molecular design, specifically relating to a design method that uses computer-aided aptamer screening and tailoring. Background Technology

[0002] Aptamers are a class of single-stranded nucleic acid molecules, typically obtained through phylogenetic index enrichment (SELEX) screening. Due to their high affinity, specificity, chemical stability, and ease of synthesis and modification, aptamers have shown promising applications in analytical detection and biosensing. Currently, various SELEX strategies have been developed for different targets, including GO-SELEX based on graphene oxide, MB-SELEX based on magnetic beads, and capture SELEX. However, both traditional SELEX and improved strategies generally suffer from problems such as multiple screening rounds, long cycles, and high labor intensity. Furthermore, target immobilization processes are often cumbersome, with low tolerance for human error and significant impact on repeatability from fluctuations in experimental conditions, thus limiting screening efficiency and stability. Therefore, there is an urgent need for an aptamer screening method or system that can reduce reliance on manual labor, shorten the screening cycle, and improve screening reliability to supplement or improve the traditional SELEX process.

[0003] With the development of bioinformatics, computational molecular docking can be used to predict the conformation and binding trends of molecular interactions, significantly reducing experimental trial-and-error costs and becoming an important tool in molecular recognition research. Given that aptamer-target recognition is essentially a molecular interaction, computational screening has also been used in recent years for aptamer screening and optimization, improving screening efficiency through virtual mutation, structural filtering, and molecular dynamics based on existing aptamer sequences. However, existing computer-aided SELEX methods largely rely on known aptamers as a starting point and the construction of random mutation libraries, making them unsuitable for targets lacking prior aptamer information. Furthermore, when relying on a single docking software or a single scoring model, the prediction results lack robustness. On the other hand, the original sequences obtained by SELEX are often long and structurally redundant; even after mutation optimization, non-essential regions may still be retained, hindering subsequent synthesis, conformational stabilization, and application development. In addition, aptamer functional truncation strategies have been used in recent years to remove non-functional regions, reduce steric hindrance, and improve recognition performance. However, traditional truncation relies heavily on empirical trial-and-error and low-throughput validation, resulting in high workload, low efficiency, and a lack of clear constraints and reproducible evaluation criteria. Therefore, it is still necessary to combine computational evaluation with experimental verification, and to achieve a more efficient truncation optimization process under structural and site constraints.

[0004] Based on this, the present invention proposes an aptamer screening method, which combines computer simulation analysis of the common structure of seven polychlorinated biphenyls, introduces multi-docking algorithm cross-validation in the screening stage to improve prediction robustness, further verifies the conformational stability and binding ability of the preferred sequence using molecular dynamics simulation and affinity experiment, and conducts unsupervised learning truncation based on the binding site to obtain shorter aptamers with better affinity. Summary of the Invention

[0005] The technical problems to be solved by this invention are as follows: First, existing aptamer screening methods generally require multiple rounds of iterative enrichment, which is time-consuming, labor-intensive, and heavily reliant on target immobilization and manual operation, making the screening efficiency and repeatability susceptible to fluctuations in experimental conditions. Second, existing computer-aided screening methods mostly start with the aptamer sequences of known targets to construct random mutation libraries, which is difficult to apply to targets lacking prior aptamer information; at the same time, when relying on a single docking software or a single scoring model, the prediction results are sensitive to the algorithm and parameters, resulting in insufficient robustness of the screening conclusions. Third, the original aptamer sequences obtained by SELEX are usually long and may contain structural redundancy. Existing optimization strategies do not adequately constrain key binding sites and conformational stability, and functional truncation is still mainly based on empirical trial and error, resulting in low throughput and long cycles, making it difficult to efficiently obtain aptamers with "short sequences, strong affinity, and stable conformations" for engineering applications.

[0006] This invention provides a design method for computer-aided aptamer screening and pruning, which enables computer-aided screening, robust evaluation, and targeted truncation optimization of candidate aptamers while reducing the reliance on traditional SELEX iterations and manual intervention. The method mainly includes the following steps: We performed structural alignment and common backbone extraction on target molecules of multiple homologs to construct a common structure model for virtual screening. Known aptamer sequences that are structurally similar to the target molecules of the homologues are selected as seed sequences, and an initial candidate aptamer library is constructed through iterative random single-base mutation and sequence deduplication strategies. The secondary structure of candidate sequences was predicted using a fold free energy assessment tool. Based on the fold free energy, GC base pair ratio, and number of inner loops, clustering was performed to screen out candidate sequences with strong stability. A tertiary structure model of the selected candidate aptamers is constructed, and at least two different molecular docking software programs are used to simulate molecular docking between the tertiary structure model and the common structure model to obtain binding energy data. The results output by different molecular docking software are standardized and ranked by comprehensive scoring, and candidate aptamers are selected based on the ranking results. Based on molecular docking results, key binding sites for target binding in candidate aptamers and their location information in secondary structures are identified. Based on the identified binding site information, truncation optimization is performed: a fitness function containing a length penalty term and binding site retention constraints is defined, and an optimization algorithm is used to prune and optimize the candidate aptamer sequence to obtain the truncated aptamer sequence.

[0007] Optionally, after selecting candidate aptamers based on the sorting results, the method further includes the steps of: performing molecular dynamics simulations on the selected candidate aptamers to evaluate their conformational stability, and verifying their affinity and specificity through isothermal titration calorimetry experiments.

[0008] Optionally, after obtaining the truncated aptamer sequence, the method further includes the steps of: performing molecular docking and molecular dynamics simulations on the truncated aptamer sequence to evaluate its conformational stability, and verifying its affinity and specificity through isothermal titration calorimetry experiments.

[0009] Optionally, the target molecules of the multiple homologues are polychlorinated biphenyls (PCBs).

[0010] Optionally, the at least two different molecular docking software programs include AutoDock-GPU, HDOCK, and Patchdock.

[0011] Optionally, the fitness function may further include a penalty term for penalizing changes in the local structural environment of key binding sites.

[0012] Optionally, the step of using optimization algorithms to prune and optimize candidate aptamer sequences specifically includes: firstly, using a genetic algorithm for preliminary optimization, and then, based on the optimization results of the genetic algorithm, further using a particle swarm optimization algorithm and a simulated annealing algorithm for optimization.

[0013] Optionally, the clustering based on the indicators of folding free energy, GC base pair ratio, and number of inner loops specifically involves calculating the indicators and screening candidate sequences with low free energy, high GC base pair ratio, and few inner loops through cluster analysis.

[0014] Optionally, the truncation optimization specifically includes: converting the candidate aptamer sequence into a binary mask, the binary mask being used to characterize whether each base in the sequence is retained, and the optimization algorithm iteratively optimizing the binary mask to minimize the fitness function.

[0015] The present invention also provides an aptamer, which is obtained by screening or tailoring optimization using the design method described above.

[0016] Compared with the prior art, the present invention has the following beneficial effects: First, addressing the issues of low efficiency, heavy reliance on manual intervention, and unstable results in existing screening methods, this invention constructs a shared structural model of target molecules as a unified virtual screening template. Combined with multi-software molecular docking and consensus scoring strategies, it achieves efficient and low-biased preliminary screening. Firstly, structural alignment and shared backbone extraction are performed on multiple homologous target molecules. The constructed shared structural model represents the common recognition features of this type of target, thus covering multiple structurally similar target molecules in a single virtual screening. This avoids the tedious process of screening each homolog individually, greatly enhancing the potential for screening efficiency and laying the foundation for obtaining aptamers with broad-spectrum recognition capabilities. Secondly, simulations are performed using at least two different molecular docking software programs, and the results from different software are standardized and comprehensively scored and ranked. This multi-algorithm cross-validation and consensus decision-making mechanism effectively reduces random errors and system biases caused by differences in algorithm principles and parameters of a single docking software or scoring model. This significantly improves the robustness, reliability, and repeatability of the virtual screening results, reducing false positives or false negatives caused by algorithm limitations.

[0017] Second, addressing the limitation of existing computer-aided screening methods in the absence of prior aptamer information, this invention constructs a complete computational design process starting from seed sequence expansion. This method does not necessarily rely on known high-affinity aptamers as a starting point. Instead, it utilizes iterative random single-base mutations and sequence deduplication strategies to construct a sizable and diverse initial candidate aptamer library from one or more initial seed sequences (including known related sequences or rationally designed sequences). Subsequently, through clustering pre-screening based on stability indicators such as folding free energy, GC content, and the number of inner loops, structurally unstable or difficult-to-fold sequences can be quickly eliminated, concentrating computational resources on more promising candidates. This process, from library construction to structural pre-screening, reduces reliance on specific prior knowledge, expands the method's applicability, and makes de novo computational design possible for novel targets or targets lacking known aptamers.

[0018] Third, addressing the problems of existing aptamer truncation optimization strategies relying on empirical trial and error, resulting in low efficiency and difficulty in maintaining binding activity, this invention proposes a machine learning-based truncation optimization method guided by binding site information, realizing a transformation from manual experience-based decision-making to automated intelligent truncation. This method first accurately identifies key binding sites and their secondary structural environment information in candidate aptamers based on molecular docking results. On this basis, by defining a fitness function that includes a length penalty term and binding site preservation constraints, the two core objectives of "maintaining key interactions" and "controlling sequence length" are quantified, and an optimization algorithm is used for automatic search. In particular, the use of genetic algorithms, particle swarm optimization, and simulated annealing algorithms for iterative optimization of the sequence or its binary mask enables efficient global optimization within a vast sequence truncation space, quickly finding an optimized scheme with shorter sequence length and more stable conformation while maintaining key binding sites and local structural environment. This process completely eliminates the inefficient traditional manual truncation and trial-and-error model, achieving targeted and rapid aptamer simplification. While significantly shortening the sequence length, reducing synthesis costs, and improving conformational stability, it retains the original binding affinity and specificity to the maximum extent, thus efficiently obtaining "short sequence, strong affinity, and stable conformation" aptamers that are more suitable for practical engineering applications.

[0019] In summary, this invention combines structure-guided multi-software consensus screening with structure-guided machine learning truncation optimization to form a closed-loop design optimization system. This method not only significantly improves the efficiency and success rate of aptamer screening and optimization but also enhances the automation and reliability of the entire process. It provides a powerful computational design tool for rapidly obtaining high-performance nucleic acid aptamers, especially for broad-spectrum identification of complex homologues such as polychlorinated biphenyls (PCBs) in fields like food safety monitoring. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the computer-aided screening process for broad-spectrum aptamers provided by the present invention.

[0022] Figure 2 This is a schematic diagram of the structure-guided machine learning truncation process provided by the present invention.

[0023] Figure 3 This is a graph showing the change in adaptability during the cutting process provided by the present invention. Detailed Implementation

[0024] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] This invention provides a design method for computer-aided aptamer screening and trimming. Traditional aptamer screening techniques, such as SELEX, suffer from time-consuming processes, strong reliance on manual intervention, and unstable results. While existing computer-aided screening can reduce the amount of experimentation, it still has drawbacks such as strong dependence on prior sequences, insufficient robustness of single-algorithm predictions, and a lack of effective means for efficient and viable trimming optimization of long sequences. This invention integrates computational biology and machine learning to construct a complete closed-loop process from target analysis to intelligent sequence trimming, aiming to efficiently and directionally obtain short-sequence, high-affinity, and conformationally stable aptamers.

[0026] Example 1: A computer-aided method for aptamer screening and tailoring design This embodiment uses polychlorinated biphenyl (PCB) contaminants, commonly found in the field of food safety, as a target example to illustrate the implementation process of the method of the present invention. PCBs comprise numerous chlorinated biphenyl homologues, which are structurally highly similar but exhibit subtle differences, making it particularly difficult to screen aptamers with broad-spectrum recognition capabilities. For example... Figure 1 As shown, the method of the present invention systematically solves this problem through the following steps: Step 1: First, seven representative polychlorinated biphenyl (PCB) indicator monomers were selected. In this embodiment, the selected PCB indicator monomers include PCB28, PCB138, PCB180, PCB52, PCB153, PCB101, and PCB118. Their precise three-dimensional structures were obtained using molecular simulation software (Schrödinger Maestro software was used in this embodiment). Then, the three-dimensional structures of these homologues were spatially superimposed and aligned, using their shared biphenyl backbone as the core, ensuring that key functional groups, such as chlorine atom substitution sites, overlapped as much as possible in space. By calculating and analyzing the common volume and spatial distribution of key atoms between the superimposed molecules, the common three-dimensional pharmacophore features and spatial backbone of all homologues were extracted. Specifically, the two-dimensional structures of PCBs were first drawn in ChemDraw and exported as CDX files, and then imported into the molecular simulation software. Using the OPLS4 force field, the introduced homologous molecules were aligned with a common framework, generating dominant conformations within an energy window ≤10 kcal / mol and spatially superimposed. Based on the common framework with topological consistency ≥85%, their common core structural features were determined. Based on this common feature, in The software constructs a pharmacophore model composed of feature points such as hydrophobic centers, aromatic ring centers, hydrogen bond interaction sites, and charged centers. This pharmacophore model quantifies the key interaction features shared by this class of target molecules. Subsequently, using this pharmacophore model as a spatial constraint, an averaged standard PDB format 3D coordinate file representing the common structural features of this series of homologues is constructed as a unified template for subsequent virtual screening. This shared structural model represents the most core and common identifiable structural features of this target family, rather than a specific molecule.

[0027] This step transforms the challenge of screening multiple structurally similar targets into a screening problem for a unified virtual target, achieving the efficient goal of "one virtual screening covering multiple targets," and laying the core structural foundation for obtaining aptamers with broad-spectrum recognition capabilities in the future.

[0028] Step 2: In this embodiment, publicly reported aptamer sequences targeting PCB72, PCB77, and PCB106, respectively, are selected as seed sequences or parent sequences. These seed sequences target PCB72, PCB77, and PCB106, which are structurally similar to the seven PCB indicator monomers (PCB28, PCB138, PCB180, PCB52, PCB153, PCB101, and PCB118) selected in Step 1 of this invention. All of them possess a biphenyl core and chlorine atom substitutions in varying numbers and positions. Therefore, based on these seed sequences, mutation expansion is expected to obtain candidate sequences capable of recognizing a wider range of PCB homologues. This embodiment uses a Python script to perform continuous iterative random single-base mutations on these seed sequences. Specifically, assuming the length of the seed sequence is N, the single-point mutation probability is set to p = 1 / N, with each base having an equal and independent probability of mutation. In one round of mutation, a fixed number of consecutive mutation operations are performed (3N times in this embodiment): First, a random single-base mutation is performed using the original seed sequence as a template to obtain the first mutant; then, the previous mutant is used as a template for the next mutation, and so on, until 3N mutations are completed, thereby generating a series of continuously accumulating mutated sequences. To avoid unlimited expansion of the library size and to eliminate redundant information, after each round of mutation, all generated sequences undergo strict deduplication, i.e., identical sequences are removed. Through such mutation and deduplication operations, an initial candidate aptamer sequence library with controllable size and sufficient sequence diversity is finally constructed.

[0029] This step overcomes the limitation of existing computer-aided screening that relies too heavily on a single known high-affinity sequence as the starting point. Through controllable evolutionary simulation and information compression, it effectively expands the search space for candidate sequences and increases the possibility of discovering new high-potential binding sequences.

[0030] Step 3: The initial library may contain a large number of sequences with unstable or difficult-to-fold secondary structures. Performing time-consuming 3D docking simulations on all of them would be inefficient. Therefore, this step implements rapid thermodynamic stability pre-screening. A batch folding free energy assessment tool is used; in this embodiment, the RNAfold program from the ViennaRNA software package is employed to predict the most stable secondary structure and its corresponding minimum folding free energy for each candidate sequence. Simultaneously, two key structural stability auxiliary indicators are automatically extracted from the prediction results: the ratio of GC base pairs to the total number of bases, and the number of inner rings in the secondary structure. Typically, a more negative indicator... Numerical values, higher GC ratios, and fewer inner loops collectively characterize a more stable and less dissociable nucleic acid secondary structure. Subsequently, clustering analysis methods (such as the K-means algorithm) were employed based on... The three indicators—GC ratio, number of inner loops, and GC ratio—are used to cluster all candidate sequences. In a specific embodiment of the invention, since these three indicators independently characterize the stability of sequences, and their original data distribution itself contains clear physicochemical meaning, it is unnecessary to perform GC analysis before clustering. The sequences were preprocessed using standardization based on GC ratio and number of inner loops. The purpose of cluster analysis is to group sequences into different groups based on their relative merit. Specifically, this is done according to the merit of each metric (e.g., GC ratio, number of inner loops). The three indicators (more negative values, higher GC ratio, fewer inner loops) together constitute eight mutually exclusive combinations of good and bad. The clustering process involves assigning all candidate sequences to the clusters corresponding to these eight combinations. Subsequently, the sequence clusters corresponding to the unique globally optimal combination of "low free energy, high GC ratio, and few inner loops" (i.e., all three indicators being 'good') are selected as structurally stable and easily foldable candidate subsets and fed into subsequent steps.

[0031] This step acts as a highly efficient pre-filter, eliminating thermodynamically unfavorable sequences before investing significant computational resources in three-dimensional simulations. This significantly improves the computational efficiency of the overall screening process and ensures that the starting sequences for subsequent analyses have a sound structural foundation.

[0032] Step 4: For candidate sequences that pass the pre-screening, they need to be converted from one-dimensional sequences into three-dimensional spatial models that can be used for fine intermolecular interaction analysis. First, using RNA tertiary structure modeling tools, this example uses the RNAComposer online server. The sequence and its predicted secondary structures are input to generate its three-dimensional spatial structure model. If the target aptamer is a single-stranded DNA, molecular visualization and processing software such as PyMOL can be used to replace the corresponding uracil bases in the RNA model with thymine, and a brief energy minimization optimization can be performed to obtain a reasonable single-stranded DNA conformation.

[0033] Next, the core virtual affinity assessment stage begins. The tertiary structure model of each candidate aptamer and the shared structure model of PCBs constructed in the first step are imported into at least two different molecular docking software programs for parallel simulation. In a preferred embodiment, three molecular docking software programs are used: AutoDock-GPU, HDOCK, and PatchDock. Each software program, based on its own unique search algorithm and scoring function, independently simulates the binding process between the aptamer and the target model, and outputs the predicted optimal binding conformation and its corresponding binding energy or overall score.

[0034] The key innovation of this step lies in the parallel docking and cross-validation across multiple platforms. Since the algorithmic principles, conformational search strategies, and scoring functions of any single docking software have inherent limitations, they may lead to prediction bias or false positives. Simultaneously employing multiple software programs with different principles allows for a multi-faceted evaluation of the binding potential of candidate sequences, effectively offsetting the systematic bias of a single algorithm and greatly improving the robustness, reliability, and repeatability of the virtual screening results.

[0035] Step 5: This step aims to achieve highly robust screening. The results output by different molecular docking software in Step 4 have different dimensions and numerical distribution ranges. For example, AutoDock outputs negative binding energies, while HDOCK outputs positive scores, making direct horizontal comparison and ranking impossible. To solve this problem, data standardization is necessary. In this embodiment, for each docking software used, the original scores of all candidate sequences are processed using methods such as Z-score standardization to transform them into a standard normal distribution with a mean of 0 and a standard deviation of 1. Then, for each candidate sequence, the sum of its standardized scores across all docking software used is calculated, resulting in a comprehensive score T. All candidate sequences are ranked according to the comprehensive score T. A higher comprehensive score T indicates that the sequence is consistently predicted to have a good binding trend under different algorithm systems, and its prediction results have higher consensus and robustness. The top-ranked sequences by comprehensive score are selected as the final preferred candidate aptamers. This step integrates heterogeneous data from different algorithms into a unified and comparable evaluation index through data standardization and multi-source scoring fusion, enabling fair comparison and consensus decision-making across platforms. It is the core step for obtaining highly reliable and low-biased screening conclusions, effectively solving the problem of insufficient robustness of prediction results when relying on a single scoring model.

[0036] To confirm the reliability of the screening results, this embodiment further performs molecular dynamics simulations on the selected candidate aptamers to evaluate the conformational dynamic stability of their complexes bound to the target over the simulated timescale. Simultaneously, the dissociation constants of the aptamers with actual polychlorinated biphenyl (PCB) molecules are directly determined using isothermal titration calorimetry (ITC) experiments to verify their affinity and specificity, ensuring the effectiveness of the preliminary calculations and screening.

[0037] Step 6: The preferred candidate aptamers selected in Step 5 typically have long sequences, requiring further analysis of their predicted binding modes to guide precise trimming. This embodiment focuses on analyzing the key regions where they form stable interactions with the target model in molecular docking simulations. Specifically, the aptamer nucleotide bases that significantly interact with the target molecule are extracted from the DLG result file output by the AutoDock-GPU software. This is achieved by extracting the optimal binding conformation of the ligand from the DLG file (based on binding free energy), calculating the spatial distance between the ligand atoms and the RNA atoms in the acceptor PDB structure, and so on. To determine the presence of an interaction, a threshold is set. The numbers of all receptor residues satisfying the distance threshold are recorded, mapped to the RNA sequence index, and the type of interaction (including hydrogen bonds, hydrophobic contacts, van der Waals forces, etc.) is identified. These bases are defined as key binding sites. Simultaneously, a secondary structure prediction tool (ViennaRNA in this example) is used again to analyze the local microenvironment of these key site bases within the aptamer's own secondary structure, such as whether they are located in a stable double-stranded stem region, a flexible single-stranded loop region, or a convex loop region, and their position index in the primary sequence is precisely recorded. The determination of the local microenvironment relies entirely on the output of the secondary structure prediction software. Specifically, by calling RNA.ptable to parse the bracket notation of secondary structures, stem regions and various types of loop regions are clearly defined, including hairpin loops, internal loops, multi-branched loops, and unpaired regions.

[0038] This step, by precisely locating its core functional region and the structural context upon which that region depends, provides a dual constraint for subsequent machine learning pruning: a functional anchor point and a structural environment that must be maintained at all costs. This ensures the directionality of the pruning.

[0039] Step 7: This step is the core innovation of this invention, aiming to automatically and intelligently trim long-sequence aptamers into shorter, more stable, and active optimized sequences. The complete process is as follows: Figure 2 As shown. The specific implementation is as follows: 1. Problem Modeling and Encoding: The original candidate aptamer sequence to be pruned is converted into a binary mask vector as an optimization variable. The length of this vector is equal to the length of the original sequence. Each bit in the vector corresponds to a base position in the sequence, with a value of 1 indicating that the base is retained in the final truncated sequence and a value of 0 indicating that it is deleted. This binary mask is used to characterize whether each base in the sequence is retained. For example, for an 8-base sequence, the mask "11011001" indicates that the 3rd, 4th, and 7th bases are deleted. In this invention, this encoding method does not impose any positional dependency constraints, so the binary mask allows for consecutive 0s of arbitrary length, that is, it allows for the deletion of consecutive base segments of arbitrary length in the original sequence. Whether in initialization, crossover, mutation, or subsequent particle swarm optimization and simulated annealing operations, there are no specific restrictions on the length or distribution pattern of consecutive 0s, so the algorithm can freely explore solutions containing consecutive deletions of arbitrary length in the search space. Furthermore, this invention does not explicitly define hard constraints on the minimum retained sequence length or the maximum consecutive deletion length.

[0040] 2. Define the fitness function for quantitative evaluation: A fitness function F is defined to accurately quantify the quality of the mask corresponding to any pruning scheme. The design of function F directly embeds all the objectives and constraints of pruning optimization, including a length penalty term and binding site preservation constraints. (1) Sequence length penalty term Where L is the number of 1s in the mask, i.e., the length of the preserved sequence; A positive weighting coefficient. This directly incentivizes the generation of shorter sequences, reducing synthesis costs and application spatial steric hindrance.

[0041] (2) Binding energy trend term :in The change in binding energy between the truncated sequence and the target is estimated based on empirical rules or simplified models. It is based on the structural differences in the retained binding site regions to estimate the fitness of sequence optimization. This is a weighting coefficient. This parameter aims to maintain or enhance the affinity of the truncated aptamer by constraining the perturbation of key functional areas during the pruning process. The specific calculation steps are as follows: Binding sites were extracted: aptamer residues that significantly interacted with the target model were extracted from the molecular docking results (DLG file output by AutoDock-GPU) and receptor structure file (PDB file) and defined as key binding sites.

[0042] Parsing the original secondary structure domains: The minimum free energy secondary structure of the original aptamer sequence is predicted using a secondary structure prediction tool (RNAfold program in ViennaRNA package), and the secondary structure is parsed into different domain types according to the dotted bracket notation, including but not limited to hairpin loops, internal loops, multi-branched loops, unpaired regions and stem regions.

[0043] Constructing a binary mask: Create a binary mask vector of equal length for the original aptamer sequence to represent the pruning scheme.

[0044] Computational structural differences: The specific numerical definition is the degree of difference between the predicted secondary structure of all key binding sites in the preserved sequence fragment and the corresponding secondary structure in the original sequence, under a given pruning scheme represented by a binary mask. This difference is quantified by calculating the Levenshtein Distance. The smaller the Levenshtein Distance, the closer the secondary structure of the key regions after pruning is to the original state, which is more conducive to maintaining the original binding conformation and affinity.

[0045] (3) Mandatory retention penalty for key sites This is a hard constraint. If any bit of the mask corresponding to a certain clipping scheme has a binary value of 0 corresponding to the key binding site identified in step six, then... The item is assigned a very large penalty value, which is preferably set in this embodiment. This results in an extremely poor fitness function value for the proposed scheme, effectively eliminating it during the optimization process. This ensures that key binding sites are not deleted during the pruning process.

[0046] (4) Local structural environment maintenance penalty item This is an important soft constraint used to penalize changes in the local structural environment of critical binding sites. If the local secondary structure type of the critical binding site in the trimmed new sequence changes in a way that is detrimental to maintaining the original binding conformation, such as changing from a functionally necessary flexible loop region to a rigid double-stranded region, then... The item is assigned a large penalty value, which is preferably set in this embodiment. This guides the optimization process to tend to maintain the original functional conformation environment.

[0047] In summary, the fitness function can be expressed as: The values ​​of each penalty term are not arbitrarily set, but are determined based on statistical analysis and response surface optimization of binding site structure loss under random pruning schemes. Specifically: (1) (Mandatory retention penalty for key sites): Statistical analysis of binding site structural loss under random pruning schemes (measured by Levenshtein distance) revealed a mean of approximately 15-30, with a maximum value typically not exceeding 80. To ensure that "binding sites cannot be deleted" becomes a truly hard constraint, the penalty value must be at least an order of magnitude higher than the normal structural loss, thereby rapidly eliminating any individuals violating this constraint during population evolution. Response surface methodology validation has confirmed that this embodiment preferably sets... =1000, this value is the minimum effective penalty value to ensure stable convergence of the algorithm.

[0048] (2) (Local structural environment preservation penalty): This penalty term aims to guide the optimization process to prioritize preserving the original structural environment of key binding sites. It is set to... = 500, which is significantly higher than the normal structural loss (ensuring priority preservation of the original structural type), but lower than... (To avoid shrinking the algorithm's effective search space due to excessive restrictions). This setting reflects a trade-off in constraint priorities: binding site integrity takes precedence over structural environment preservation.

[0049] (3) (Binding site mapping validity penalty): If the pruned sequence index range cannot effectively map the original binding site (e.g., the index exceeds the sequence range), it is considered a physically infeasible pruning scheme. This embodiment preferably sets... = 500, its penalty value is equal to the penalty for structural type change, reflecting that the priority of this constraint is equal to structural environment preservation, but lower than binding site integrity. Through this priority system ( The fitness function can flexibly explore the sequence pruning space while ensuring that the binding core is not destroyed, and finally obtain an optimized aptamer that is both concise and maintains high affinity.

[0050] Weighting coefficient and These are used to balance the penalty imposed on sequence length and the requirement for structural conservation of key binding sites, respectively. Specifically: Relative size relationship: The value should usually be much larger than ,Right now This is because preserving the structural integrity of key binding sites is fundamental to maintaining aptamer affinity; damage to this region will lead to loss of function, and therefore should be given the highest optimization priority. Shortening sequence length is a secondary optimization objective, and its penalty should not excessively interfere with the protection of core functional regions.

[0051] The principle of dynamic adjustment: The coefficient used as the length penalty term can be dynamically adjusted based on the expected pruning results. If, in a certain optimization, the truncated sequence is too long and fails to achieve the desired compression effect, it can be appropriately increased. The value of is adjusted to enhance the algorithm's drive to shorten the sequence; conversely, if the optimized sequence is too short, excessive pruning may affect the stability or integrity of the structure, in which case the value can be appropriately reduced. The value of is adjusted to relax the constraint on length, allowing the algorithm to focus more on maintaining structural conservatism. In a typical optimization process, It can be set to a small initial value (such as 0.1 or 0.01) and fine-tuned based on the iteration results.

[0052] With the above settings, those skilled in the art can flexibly adjust the weighting coefficients according to specific application requirements and optimization goals, thereby obtaining the optimal aptamer sequence that combines high affinity with appropriate length.

[0053] The optimization objective is to search for the binary mask that minimizes the fitness function F. This function successfully transforms four complex and interdependent biological objectives—shortening the sequence, maintaining affinity, preserving key sites, and maintaining the structural environment—into a computable mathematical optimization problem. It is the theoretical core of intelligent pruning, and its combination of site preservation constraints and structural environment constraints significantly outperforms traditional empirical truncation.

[0054] 3. Efficient Search Using a Serial Optimization Algorithm: Since the number of possible pruning schemes increases exponentially with the sequence length, an efficient optimization algorithm is needed to iteratively optimize the binary mask to minimize the fitness function. This invention employs a serial hybrid strategy of genetic algorithm (GA), particle swarm optimization (PSO), and simulated annealing (SA) (i.e., " This strategy, through a three-stage collaborative process, balances global exploration capabilities with local convergence accuracy, thereby stably searching for the optimal pruning solution. The specific implementation is as follows: (1) Preliminary optimization of the genetic algorithm: The binary mask is regarded as a "chromosome", and a population is randomly initialized. Iterative optimization is performed by simulating genetic operations such as selection, crossover, and mutation in biological evolution. The goal is to quickly perform a coarse search in the global scope, locate high-quality regions, and generate a diverse candidate population. The parameter settings of the GA algorithm are as follows: Population size: pop_size = min(200, max(50, seq_len / / 3)) (i.e., dynamically adjusted according to sequence length); Maximum number of generations: max_generations = 200; Crossover probability: cxpb = 0.5; Mutation probability: mutpb = 0.1; Choose the pressure (tournament size): 4.

[0055] The termination conditions for the GA algorithm are as follows (it stops when any one of them is met): Reaching the maximum number of generations (200 generations); Low population diversity: When the normalized entropy of the population is below 0.1, it indicates that the individuals are highly homogeneous, and it is difficult to generate new solutions through continued evolution. Therefore, the population should be terminated early to save computational resources. No significant improvement for 20 consecutive generations: If the optimal fitness decreases by less than 0.001 for 20 consecutive generations, it is considered convergent and terminated early.

[0056] (2) Particle Swarm Optimization (PSO): After the genetic algorithm terminates, the first n_particles individuals in its final population are directly used as the initial particle swarm of the PSO algorithm. At the same time, the global optimal solution of the genetic algorithm is used as the initial global optimal guide for the particle swarm. The information interaction between particles is used to perform fine search and accelerate convergence to the local optimal region.

[0057] The parameters for the PSO algorithm are set as follows: Number of particles: n_particles = 30; Maximum number of iterations: max_iterations = 500; Inertia weight: w = 0.7; Cognitive coefficient: c1 = 2.0; Social coefficient: c2 = 2.0.

[0058] The PSO algorithm terminates when the maximum number of iterations (500) is reached. Since PSO occurs after GA, the solution is close to the optimal region at this point, eliminating the need for an additional early stopping mechanism. Furthermore, the fixed number of iterations ensures that the computational load for each run is controllable and facilitates cross-sequence comparisons.

[0059] (3) Simulated Annealing (SA) Global Optimization: After the Particle Swarm Optimization (PSO) algorithm terminates, its global optimal position gbest_position is used as the initial solution for the Simulated Annealing (SA) algorithm, and the fitness value of the individual is retained as the initial state. By using random perturbation and the Metropolis acceptance criterion, the solution can escape local optima with a certain probability, thereby further improving the quality of the solution and ensuring the global optimality of the final result.

[0060] The parameters for the SA algorithm are set as follows: Maximum number of iterations: max_iterations = 3000; Initial temperature: initial_temp = 500.0; Cooling rate: cooling_rate = 0.99; Neighborhood mutation probability: indpb = 0.2.

[0061] The SA algorithm terminates when the maximum number of iterations (3000) is reached. The core of simulated annealing lies in the temperature decay process, which must undergo a complete cooling curve to ensure solution quality. At the same time, since SA is the final stage, the computational cost is relatively controllable, so early stopping is not set.

[0062] To verify the effectiveness of this concatenation strategy, this embodiment extracts the fitness score trend during the pruning process. For example... Figure 3 As shown in the experimental data, after the genetic algorithm, particle swarm optimization (PSO) algorithm, and simulated annealing algorithm were applied sequentially, the fitness scores all showed a significant decrease. This indicates that each algorithm fulfilled its design function: the genetic algorithm quickly approximates the high-quality region, the PSO algorithm finely searches to the local optimum, and the simulated annealing algorithm further improves the quality of the solution through probabilistic jumps, ultimately outputting a stable binary mask that optimizes the fitness function F value. This sequential strategy of "first a global coarse search, then a local fine search, and then escaping the trap" effectively balances the breadth and depth of the search, avoids a single algorithm getting stuck in local optima, and significantly improves the success rate and stability of the pruning optimization.

[0063] 4. Generate the final truncated sequence: Based on the optimal binary mask obtained through optimization, extract all the bases marked as "retained" from the original aptamer sequence and connect them in their original order to obtain the final truncated aptamer sequence that has been calculated and optimized.

[0064] The entire seventh step above completely revolutionizes the traditional truncation model that relies on manual experience and repeated trial and error. By quantifying binding site information, structural constraints, and length targets into a computable fitness function, and using an efficient cascade optimization algorithm for automatic search, the aptamer pruning process is made directional, automated, and high-throughput quantized. This is a key technological guarantee for obtaining the ideal aptamer with "short sequence, strong affinity, and stable conformation".

[0065] To comprehensively evaluate the pruning effect, after obtaining the truncated aptamer sequence, this embodiment re-performed molecular docking and molecular dynamics simulations on the new sequence, comparing and analyzing the changes in binding mode, predicted binding energy, and dynamic stability with the original long sequence from a computational perspective. Furthermore, isothermal titration calorimetry (ITC) experiments were used to directly and quantitatively determine the affinity and specificity of the aptamers before and after pruning for various polychlorinated biphenyl homologues, thereby demonstrating the success of the pruning optimization.

[0066] Example 2: An aptamer obtained by screening or trimming and optimizing using the design method described in Example 1.

[0067] Using the computer-aided design method described in Example 1, this embodiment successfully obtained a truncated and optimized aptamer. Specifically, using a reported aptamer sequence for polychlorinated biphenyls as a seed sequence, a truncated aptamer with higher affinity was obtained after screening and truncating optimization using the method of this invention.

[0068] The original aptamer sequence before trimming and optimization was: 5'-AGGGTAGGTTGTAGATCAGCGGGTTCTCTATGAAGCTGCTCTCACCACCTCTGTGTACAGTCCCTAACCTAGTC-3' (SEQ ID NO. 1, sequence length 74 nucleotides). Verification using isothermal titration calorimetry (ITC) showed that the binding dissociation constant (Kd) of this original sequence with the polychlorinated biphenyl (PCB) mixed standard was... This indicates that its binding force is weak and the experimental repeatability is poor.

[0069] After pruning and optimizing the original sequence using the method described in Embodiment 1 of this invention, a truncated aptamer sequence was obtained, the sequence of which is as follows: 5'-CAGGGCACCCTTCTAATAGGAATTAGGT-3' (SEQ ID NO. 2, sequence length 28 nucleotides). The binding dissociation constant (Kd) of this truncated aptamer to a mixed standard of polychlorinated biphenyls, as verified by the ITC, is [value missing]. .

[0070] Experimental results show that, after trimming and optimization using the method of this invention, the aptamer sequence length was shortened from 74 nucleotides to 28 nucleotides, a reduction of approximately 62%. Simultaneously, its binding affinity to the target increased by approximately 3.3 times, and the stability (standard deviation) of the binding data was significantly improved. This fully demonstrates that the method of this invention can efficiently and directionally obtain aptamers with shorter sequences, higher affinity, and more stable conformation. This aptamer can be used as a recognition element to develop highly sensitive and selective biosensors for detecting polychlorinated biphenyl (PCB) residues or as a sample pretreatment enrichment material, fully demonstrating the powerful capabilities and significant application value of the method of this invention in the rational design and rapid optimization of functional nucleic acid molecules.

[0071] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A computer-aided method for aptamer screening and tailoring design, characterized in that, Includes the following steps: We performed structural alignment and common backbone extraction on target molecules of multiple homologs to construct a common structure model for virtual screening. Known aptamer sequences that are structurally similar to the target molecules of the homologues are selected as seed sequences, and an initial candidate aptamer library is constructed through iterative random single-base mutation and sequence deduplication strategies. The secondary structure of candidate sequences was predicted using a fold free energy assessment tool. Based on the fold free energy, GC base pair ratio, and number of inner loops, clustering was performed to screen out candidate sequences with strong stability. A tertiary structure model of the selected candidate aptamers is constructed, and at least two different molecular docking software programs are used to simulate molecular docking between the tertiary structure model and the common structure model to obtain binding energy data. The results output by different molecular docking software are standardized and ranked by comprehensive scoring, and candidate aptamers are selected based on the ranking results. Based on molecular docking results, key binding sites for target binding in candidate aptamers and their location information in secondary structures are identified. Based on the identified binding site information, truncation optimization is performed: a fitness function containing a length penalty term and binding site retention constraints is defined, and an optimization algorithm is used to prune and optimize the candidate aptamer sequence to obtain the truncated aptamer sequence.

2. The method according to claim 1, characterized in that, After selecting candidate aptamers based on the sorting results, the process further includes the following steps: performing molecular dynamics simulations on the selected candidate aptamers to evaluate their conformational stability, and verifying their affinity and specificity through isothermal titration calorimetry experiments.

3. The method according to claim 1, characterized in that, After obtaining the truncated aptamer sequence, the method further includes the steps of performing molecular docking and molecular dynamics simulations on the truncated aptamer sequence to evaluate its conformational stability, and verifying its affinity and specificity through isothermal titration calorimetry experiments.

4. The method according to claim 1, characterized in that, The target molecules of the various homologues are polychlorinated biphenyls.

5. The method according to claim 1, characterized in that, The at least two different molecular docking software programs include AutoDock-GPU, HDOCK, and Patchdock.

6. The method according to claim 1, characterized in that, The fitness function also includes a penalty term for penalizing changes in the local structural environment of key binding sites.

7. The method according to claim 1, characterized in that, The optimization of candidate aptamer sequences using optimization algorithms specifically includes: firstly, using a genetic algorithm for preliminary optimization, and then, based on the optimization results of the genetic algorithm, further using a particle swarm optimization algorithm and a simulated annealing algorithm for optimization.

8. The method according to claim 1, characterized in that, The clustering based on indicators such as folding free energy, GC base pair ratio, and number of inner loops specifically involves calculating the indicators and screening candidate sequences with low free energy, high GC base pair ratio, and few inner loops through cluster analysis.

9. The method according to claim 1, characterized in that, The truncation optimization specifically includes: converting the candidate aptamer sequence into a binary mask, the binary mask being used to characterize whether each base in the sequence is retained, and the optimization algorithm iteratively optimizing the binary mask to minimize the fitness function.

10. An aptamer, characterized in that, The aptamer is obtained by screening or trimming and optimization using the method described in any one of claims 1 to 9, and the nucleotide sequence is shown in SEQ ID NO. 1 or SEQ ID NO. 2.