A method for ranking homogenous RNA
Patent Information
- Application Number
- PCT/SG2026/050181
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-24
- Publication Date
- 2026-10-01
Smart Images

Figure SG2026050181_01102026_PF_FP_ABST
Abstract
Description
[0001] A Method for Ranking Homogenous RNA
[0002] Technical Field
[0003] The present invention relates, in general terms, to a method and system for ranking homogenous ribonucleic acid (RNA) based on a similarity between a structure of the RNA and a native structure of the same sequence.
[0004] Background
[0005] The reference in this specification to any prior publication (or information derived from it), or to any matter which is known, is not, and should not be taken as an acknowledgment or admission or any form of suggestion that that prior publication (or information derived from it) or known matter forms part of the common general knowledge in the field of endeavour to which this specification relates.
[0006] RNA molecules play central roles in cellular functions, such as gene expression, protein synthesis, and catalytic activities. Resolving RNA structures is critical for understanding their functions. However, experimental methods like X-ray crystallography, nuclear magnetic resonance (NMR) spectroscopy, and cryogenic electron microscopy (Cryo-EM) are costly and time-consuming. Computational prediction of RNA 3-dimensional (3D) structure is therefore highly significant.
[0007] Unlike proteins, which have seen landmark achievements like AlphaFold and RosettaFold, RNA 3D structure prediction has yet to reach such promising accuracy. This is due to the scarcity of RNA data - experimentally resolved RNA structures account for less than 1% of resolved protein structures, and the inherent structural complexity and instability - their folding pathways are highly sensitive to environmental conditions, exhibiting multiple transient energy landscapes. These limitations result in RNA predictive methods typically generating multiple conformations for a specific sequence through different parameter initializations or Monte Carlo sampling, for capturing dynamic behaviours of RNAs. Although this helps explore the potential solution space, it complicates the selection of the optimal conformation. Therefore, evaluating RNA 3D structures, especially learning structural characteristics, and generalizing structural knowledge to assessthe accuracy of unseen conformations, is difficult despite being important.
[0008] Traditional RNA 3D structure evaluation methods rely on knowledge-based statistical potentials, benchmarking candidates against the reference state that represents an idealized baseline conformation or energy state. RASP and 3dRNAscore build all-atom distance potentials based on averaged reference states. e-SCORE focuses on relative arrangements and orientations of the bases. DFIRE-RNA introduces a distance-scaled statistical potential based on finite-ideal-gas reference state. Rosetta scoring function incorporates a weighted combination of knowledge- and physics-based potentials, such as van der Waals interactions, base-pairing, and other statistical terms. rsRNASP refines statistical potentials by distinguishing short- and long-range residue interactions. The recent cgRNASP enhances rsRNASP by further refining the short-range interactions. Despite these advancements, a key limitation of this type of evaluation lies in the inherent difficulty of accurately modelling the reference state.
[0009] Machine learning evaluation methods have more recently been proposed. However, such methods generally assess the likelihood that each individual conformation is the native conformation for a particular RNA sequence. These methods fail to capture relative structural differences between conformations, experience optimization imbalances due to RMSD variation, where longer-chain RNAs typically exhibit broader ranges, which can lead to training errors dominating the optimization, and are sensitive to outliers, degrading robustness.
[0010] It would be desirable to overcome or ameliorate at least one of the abovedescribed problems, or at least to provide a useful alternative.
[0011] Summary
[0012] Disclosed is a computer-implemented method for training a ranking model to rank ribonucleic acid (RNA) structures based on an input RNA sequence, comprising:
[0013] receiving an RNA conformation data set comprising a plurality of non- redundant single-chain RNA sequences, and, for each sequence, a plurality of structures including a native single-chain structure andone or more computationally simulated homogeneous RNA structures (CSHRS);
[0014] pairwise aligning the RNA sequences;
[0015] computing a similarity score for each pair of the RNA sequences; clustering the RNA sequences into a plurality of clusters of RNA sequences based on the similarity scores;
[0016] for each RNA sequence in a first said cluster:
[0017] generating a plurality of lists, each list being a permutation of the plurality of structures corresponding to the respective RNA sequence;
[0018] applying a ranking model to predict a ranking distribution over each list, corresponding to a probability that an order of the structures in the respective list is accurate; and
[0019] updating the ranking model based on a difference between the predicted ranking distribution and a ground-truth ranking distribution; and
[0020] outputting the trained ranking model, the trained ranking model configured to output a Top-n structures from a top ranked said list.
[0021] Also disclosed is a method for generating a candidate structure for a specific RNA sequence, the method comprising:
[0022] receiving the specific RNA sequence;
[0023] receiving or generating a plurality of new conformations corresponding to the specific RNA sequence;
[0024] generating a new list, the new list being a permutation of the plurality of the new conformations;
[0025] applying a ranking model trained according to the method of claim 1, to the new list, to predict a ranking distribution over the new list; and based on the ranking distribution, outputting a Top-1 said new conformation as a most likely candidate structure for the further RNA sequence.
[0026] Also disclosed is a system for training a ranking model to rank ribonucleic acid (RNA) structures based on an input RNA sequence, comprising:
[0027] memory;at least one processor;
[0028] one or more receivers;
[0029] an alignment module;
[0030] a clustering module;
[0031] a ranking model module comprising a ranking model; and
[0032] an output;
[0033] wherein the memory stores instructions that, when executed by the at least one processor, cause the system to:
[0034] receive, at the one or more receivers, an RNA conformation data set comprising a plurality of non-redundant single-chain RNA sequences, and, for each sequence, a plurality of structures including a native single-chain structure and one or more computationally simulated homogeneous RNA structures (CSHRS);
[0035] pairwise aligning the RNA sequences at the alignment module; computing a similarity score, at the similarity assessor, for each pair of the RNA sequences;
[0036] using the clustering module to cluster the RNA sequences into a plurality of clusters of RNA sequences based on the similarity scores; for each RNA sequence in a first said cluster, applying the ranking module to:
[0037] generate a plurality of lists, each list being a permutation of the plurality of structures corresponding to the respective RNA sequence;
[0038] apply the ranking model to predict a ranking distribution over each list, corresponding to a probability that an order of the structures in the respective list is accurate; and
[0039] update the ranking model based on a difference between the predicted ranking distribution and a ground-truth ranking distribution; and
[0040] output the trained ranking model at the output, the trained ranking model configured to output a Top-n structures from a top ranked said list.
[0041] Brief description of the drawingsEmbodiments of the present invention will now be described, by way of nonlimiting example, with reference to the drawings in which:
[0042] Figure 1: data collection workflow, involving unsupervised clustering and rigorous splitting.
[0043] Figure 2: a partial similarity map (similarity dendrogram) used to generate clusters of RNA sequences.
[0044] Figure 3: a simplified comparison of the present, list ranking method, compared with the traditional state-of-the-art Atomic Rotationally Equivariant Scorer (ARES).
[0045] Figure 4: a schematic illustration of a method for training a ranking model to identify structures corresponding to RNA sequences.
[0046] Figure 5: results showing performance of a model trained according to the method of Figure 4, in identifying near-native candidates (RMSD ≤ 2 Å) on the RNA Puzzles benchmark.
[0047] Figure 6: results showing performance of a model trained according to the method of Figure 4, in identifying near-native candidates (RMSD ≤ 2 Å) on the TestCLS-1 Benchmark.
[0048] Figure 7: results of a case study comparing the present list ranking model performance against the performance of an alternative model, HomRank-PT, trained in accordance with present teachings, ARES+, trained on a dataset curated per the process of Figure 1, and ARES.
[0049] Figure 8: ranking accuracy versus ΔRMSD.
[0050] Figure 9: demonstration of the present ranking model capturing finer-grained structural characteristics without prior specification.
[0051] Figure 10: a block diagram of a system for training a ranking model and thereafter using the trained model for RNA sequence structure prediction.
[0052] Detailed description
[0053] Described herein are two significant advances over known methods, for improving the accuracy of RNA structure prediction. The first advance is a method for curating a training dataset with rigorous preprocessing to ensure model generalization for predicting native RNA structures. Using this training dataset, the ARES model is retrained, generating ARES+. ARES+ has significantly higherprediction accuracy. The second advance is a framework for boosting near-native structure identification. In particular, a homogeneous RNA ranking method is described that directly optimizes the selection of near-native conformations from candidate sets of conformations, resembling expert evaluation strategies. This framework achieves 95% and 100% top-1 and top-5 success rates, respectively, significantly outperforming ARES+. These results demonstrate that carefully designed datasets and the expert-like selection paradigm can substantially improve the accuracy and robustness of RNA 3D structure assessment, offering a promising direction for deep learning-based RNA evaluation and near-native conformation selection.
[0054] The first main advance involves assembling a new dataset from the Protein Data Bank (PDB) or other archive of experimentally determined 3D biological macromolecules and their complexes. The new dataset comprises singlechain, non-redundant RNAs. To facilitate model generalization on unseen RNAs, the assembled RNA molecules are aligned. In some embodiments, this alignment step involves sequence alignment, such as pairwise sequence alignment. Sequence alignment may involve pairwise alignment of recognised sequence segments. A similarity score is then calculated between the pairs of sequences. Pairwise similarity scoring may use matches, mismatches and apply affine gap penalties. Alignment may be achieved using the Bio. Align module from the BioPython library to align RNA sequences and calculate their similarity. In other embodiments, pairwise sequence alignment involves universal structural alignment, where universal structural alignment involves automatically comparing 3-dimensional RNA and protein shapes. In such embodiments, relative alignment uses a length-independent scoring function known and a template modelling score, where a higher score indicates closer alignment, and is calculated using a known process.
[0055] After alignment, the RNAs are clustered into a plurality of clusters of RNA, based on the similarity scores. Clustering is hierarchical and unsupervised. In some embodiments, hierarchical unsupervised clustering uses known methods to group similar RNAs into a nested cluster structure. The structure may take the form of a dendrogram or similarity map. During clustering, RNA labels are disregarded, enabling the clustering method to be applied to RNAPuzzles as well as labelled RNAs. Moreover, automatically identifying clusters avoids manual splitting or splitting based on publication years of RNAs.
[0056] Figure 1 shows the data processing pipeline, including data collection, clustering, and splitting. In experiments, a comprehensive set of 190 non-redundant and single-chain RNA molecules was assembled, including 151 non-redundant RNAs from the PDB database, along with RNA Puzzles data (21 RNAs) and those from ARES: training (14 RNAs) and validation (4 RNAs).
[0057] Pairwise sequence alignment was performed for 190 RNAs and their similarities calculated. Using the resulting similarity map, hierarchical clustering was used to categorize RNAs into three distinct groups: a larger cluster (CLS-0) comprising 146 RNAs, and two smaller clusters (CLS-1 contains 21 RNAs, CLS-2 includes 23 RNAs). A partial clustering map (also referred to as a sequence similarity map) is shown in Figure 2, in which the branching indicates where groups of RNAs differ in some sequence feature and thereby causes cluster formation. A root of the dendrogram if off-page, to the right. In other embodiments, the RNA sequences are clustered by a known method, such as one of K-means clustering, density-based clustering, graphbased clustering, or hierarchical clustering.
[0058] CLS-0 is designated as the training set due to its larger size than the other clusters, while CLS-1 and CLS-2 are assigned as benchmarks, referred to as TestCLS-1 (21 RNAs) and TestCLS-2 (23 RNAs), respectively. As shown in Figure 1, the 15 RNA Puzzles within CLS-0 are excluded from the training data. The RNA Puzzles are commonly recognized as standard benchmarks. Consequently, the training set comprises 131 RNAs and the RNA Puzzles (21 RNAs) are set as an independent benchmark. The data split is shown in Table 1.
[0059] Dataset | Training | TestCLS-1 | TestCLS-2 | RNA Puzzles RNA Molecules | 131 | 21 | 23 | 21 Sequence Length Range | 22-144 | 61-257 | 12-49 | 41-188 Structural Candidates | 131,000 | 152,917 | 23,000 | 451,629Table 1: Details of dataset splitting.
[0060] Notably, 5 training and 3 validation RNAs from ARES, along with 15 RNA Puzzles, are clustered into CLS-O. Similarly, CLS-1 contains 9 training and 1 validation RNAs from ARES. These findings indicate a critical limitation of conventional splitting methods, specifically manual partitioning or reliance on publication years of RNA data, may unintentionally introduce data leakage between training and testing data, thereby impairing the reliability of model evaluation.
[0061] For each RNA, a predetermined number of conformations is generated. Conformation generation may involve any known method, such as molecular dynamics simulation. In some embodiments, the one or more computationally simulated structures are generated using a domain-knowledge model specifying lowest energy base pairs. For present purposes, 1,000 candidates (conformations) are generated using molecular dynamics simulations for each of the 151 new RNA molecules. The structural candidates from ARES (1,000 conformations for each of 18 RNAs) and RNA Puzzles benchmark (total 451,630 for 21 RNAs) from remain unchanged.
[0062] To evaluate the retrieval capacity for near-native structures (RMSD ≤ 2°Å with respect to the experimentally resolved native structure), approximately 1% of the candidates for each RNA in the benchmarks are near-natives, while the rest are decoy structures (RMSD > 2°Å), as shown in Tables 2 to 4.
[0063] ID I 3 3 I S » T S § S£s IS
[0064] SIT w 332 243,;'. D 332 133 52 -SD
[0065] Ill 12 S3 Tib S4i S3 ST > S IS 32 21
[0066]
[0067] ■ADD <2A:3IA 3U1 IDS 62!® 172 43 S* 3*2 Table 2: number of near-native candidates in RNA Puzzles benchmark
[0068]
[0069] Table 3: number of near-native candidates in TestCLS-1 benchmark. The subscript is the entity identifier.
[0070]
[0071] Table 4: number of near-native candidates in TestCLS-2 benchmark. The subscript is the entity identifier.
[0072] Experiments showing the model agnostic improved accuracy of using the curated dataset are described below, with reference to a framework for predicting RNA structures.
[0073] Conventional methods, such as ARES, train an RMSD regressor to individually predict RMSD score for each candidate. This neglects the relative structural differences among RNA conformations. To counter this limitation, the present methods employ a listwise ranking approach to explicitly model the relative structural difference among homogeneous candidates, closely aligning with expert-driven evaluation paradigm. The difference between the two concepts is reflected in Figure 3.
[0074] On the left-hand side of Figure 3, a conventional RMSD-based regression approach is shown. Such methods include ARES in which RMSD regressors are trained to score each candidate independently, failing to capture relative structural differences among conformations. On the right-hand side of Figure 3, the present framework is shown, in which listwise ranking explicitly modelsrelative structural differences across structurally similar candidates, closely aligning with expert-driven evaluation practices.
[0075] Using the above concept of formulating a curated dataset, a method (400) is described with reference to Figure 4, for training a listwise ranking model to rank conformations for RNA sequences, thereby facilitating efficient RNA structure prediction.
[0076] To implement the method (400), the first step involves receiving a dataset (402). The dataset is an RNA conformation data set comprising a plurality of non-redundant single-chain RNA sequences. During the training phase, for each sequence, the dataset also includes a plurality of structures including a native single-chain structure and one or more computationally simulated homogeneous RNA structures (CSHRS). The native single-chain structure is the ground truth structure sought to be identified by the ranking model, once trained. The dataset may be "received" insofar as it is retrieved from external memory or an external database, received as a signal over a network or any other transmission mechanism. Alternatively, part of the dataset may be retrieved from an external database, such as the PDB, with part of the dataset being generated. In some embodiments, on receipt of the RNA sequences and native single-chain structure, the method involves generating computationally simulated homogeneous RNA structures (CSHRS - i.e., conformations) for each RNA sequence. Simulation may involve generating conformations using molecular dynamics simulations.
[0077] After receiving and / or generating the dataset, the RNA sequences are clustered using the method described with reference to Figure 1. Specifically, the RNA sequences are pairwise aligned (404), a similarity score is calculated for each pair of the RNA sequences (406) and the RNA sequences are then clustered based on the similarity scores (408).
[0078] The largest cluster is then selected as the first cluster (CLS-0 in the example of Figure 1). It is from this list that the training data is manipulated to implement the listwise ranking algorithm. The listwise ranking algorithm ranks lists of conformations, based on structural differences between conformations in each list. In this manner, it compares structures directly,thereby learning structural relationships an expert might incorporate into manual structure selection, along with finer grained features a manual assessment might not identify. This is to be contrasted with independent assessment of the suitability of each conformation using root mean square error (RMSE) regression techniques.
[0079] Conventional deep learning-based scoring models typically use RMSD as a proxy for structural accuracy, training a regressor to predict scores that closely approximate the ground-truth RMSD. RMSD quantifies the structural deviation of candidates from the native conformation, with lower RMSD indicating closer alignment. Let X = {x=1be the set of training data, and let y = {y=1denote the corresponding ground-truth RMSD values. The training objective is to minimize the Mean Squared Error (MSE) loss, defined as:
[0080] n ^MSE= —nZ_i (yi ~ yd2, (1)
[0081]
[0082] i=l
[0083] where y is the ground-truth RMSD for the i-th instance, and y = s(xi;0) is the predicted RMSD. Here, s - 9) denotes the regression score function modeled by a neural network parameterized by 9, which is optimized during training.
[0084] In contrast, the method (400) inherently incorporates domain-knowledge into the model, by ranking lists of conformations by considering each list as a single instance to be ranked. Given a set of candidates for a specific sequence, i.e., homogeneous candidates, denoted as c = {c -=1, biological experts leverage their domain knowledge (e.g., base pairing patterns, tertiary interactions, and overall stability) to identify candidates with favourable characteristics, such as energetically stable conformations. Let s(-) represent a scoring function that encapsulates experts' assessments, which reflects their preferences for these candidates. Such an expert-driven evaluation can be formulated as:
[0085] Cl > c2> c3> ••• > chi (2
[0086]
[0087] f s(Ci) > s(c2) > s(c3) > ••• > S(Q), )where q > c2denotes the candidate q is superior to c2, and I is the length of minilist, which can be adjusted flexibly based on specific requirements or preference. To align with expert-driven evaluation, a list-wise ranking technique, such as ListNet, is used to optimize the rankings of homogeneous candidates. In this context, s(x,;0) represents the ranking score function modeled by a neural network parameterized by 6, where higher scores indicate a closer alignment with the native conformation.
[0088] To implement the listwise ranking model, given I homogeneous candidates for a particular RNA sequence, a plurality of lists are generated. Each list is a permutation of the plurality of structures corresponding to an RNA sequence (recall: the structures include the native RNA structure and 1,000 simulated conformations for each RNA sequence). In various embodiments, each list is a permutation of all structures associated with RNAs in a cluster, all structures for a particular RNA sequence the native structure of which is sought to be identified, a mini-batch (i.e., selection of a proper subset) of all structures associated with RNAs in a cluster, or a mini-batch of structures for a particular RNA sequence.
[0089] After generating the plurality of lists, the ranking model is applied to predict a ranking distribution over each list. The ranking distribution corresponds to a probability that an order of the structures in the respective list is accurate (i.e., that the order of structures in the list corresponds to the likelihood that the structures are the actual, or native, structure of the RNA sequence). For this purpose, each permutation n e gt(representing a mini-list consisting of I candidates) is associated with a probability as follows,
[0090] = n
[0091]
[0092] Li Tj=t< P
[0093] where gtis the set of all permutations of I candidates, ntis the t-th element of n, and <)(•) is a strictly positive, monotonically increasing transformation, which often uses the softmax function d>( ) =exp (7yr)where r = [n.r?, —,rt] is a list of
[0094]
[0095] scalar scores and T is the temperature coefficient. The softmax function is used to map ranking scores to a probability distribution. This results in the rank modellingthe likelihood of different permutations. Additionally, exponential transformation amplifies score differences, enabling higher scores to have a stronger influence on the ranking probabilities.
[0096] Notably, the model can be trained on the full permutations of the entire set of conformations. However, in the experimental embodiment, there are 1,000 simulated conformations for each sequence in the training dataset, with a greater number of candidate conformations for each sequence in CLS-1 and CLS-2. Even for the training dataset, the number of permutations of the full set of candidates would be prohibitively large - 1,0001, being a number with nearly 2,600 digits, permutations for each RNA sequence. Therefore, a mini-batch of Z candidates is specified, to substantially reduce the permutation space and computation time. In particular, for each RNA sequence in the first cluster, the plurality of lists is generated by randomly sampling structures from the respective plurality of structures. The randomly sampled structures form a mini-batch that is permuted to generate the plurality of lists for the respective RNA sequence.
[0097] It is evident that Ps(n) > 0 and
[0098]
[0099] = 1 / confirming that Ps(n) defines a valid probability distribution over the set gt. However, as mentioned above, since there are ll permutations, the computation becomes intractable for large I. To improve efficiency, the ranking model determines a probability that an order of a Top-structures in each list is accurate. To that end, the top- / c probability is considered, where the top- / c candidates in the permutation n are specified (denoted
[0100]
[0101] as e1). This method reduces the computation from O(Z!) to o(
[0102]
[0103] \—k!(il—- k—)\ )J. Eq. 3 is transformed into
[0104]
[0105] Also, Ps(7rw) > 0 and = 1, ensuring that the top- k probabilities
[0106]
[0107] form a valid probability distribution over the set
[0108]
[0109] In some embodiments, k = 1. This prioritises the Top-1 candidate (i.e., most likely candidate in a list), as the goal of RNA 3D structure evaluation is to maximize the likelihood of the mostaccurate candidate being ranked first. Applying the ranking model only to the Top-1 structure, computation is reduced to O(Z).
[0110] The above describes the construction of the predicted ranking distribution. For the ground-truth ranking distribution, a similar technique is applied. The difference is that the ground-truth RMSD values are not directly used but are instead converted into ranking scores (e.g., 1,1 / 2, •••,1 / 1 ), through a reciprocal transformation of the ranking order determined by the ground-truth RMSD, where smaller RMSD values correspond to higher rankings. Additionally, candidates with A RMSD < 6A are assigned the same ranking order to account for the negligible structural deviations, where <5 is a tunable hyper-parameter.
[0111] The different between the predicted ranking distribution and the ground-truth ranking distribution is then used to update the ranking model. In some embodiments, focal loss or hinge loss may be used. Presently, loss is defined as the cross-entropy loss. The cross-entropy loss measures the alignment between the ground-truth and predicted ranking distributions, which is expressed as:
[0112] ^HomRank = - Py(7T(1))logPs(7T(1)), M c £;(1), |M| = m, (5)
[0113]
[0114] where
[0115]
[0116] Py(7r(1)) jSthe top-1 ground-truth ranking distribution,
[0117]
[0118] is the corresponding predicted ranking distribution, M is a subset of g^\ consisting of m mini-lists 7r(1)randomly sampled from g^ within each mini-batch.
[0119] The ranking model can then be updated for each RNA sequence, and outputted. Once trained, the ranking model is configured to output a Top-n structures (usually, n - 1 as the most likely structure is desired) from a top ranked list -typically, a list in which the Top- candidates are ranked in order, correctly.
[0120] Since the ranking model ranks homogeneous candidates from the same sequence, it aligns closely with expert-driven evaluation and the inference task in RNA 3D structure evaluation. The list-wise ranking approach prioritizes the relative ordering of homogeneous candidates, rather than predicting absolute RMSDscores for each candidate independently, enabling it to effectively distinguish finegrained structural differences between high-fidelity conformations from low-fidelity ones. Since the process emphasizes rankings rather than absolute scores, it inherently reduces sensitivity to outliers.
[0121] Table 5 summarizes the comparisons of RMSD regression and HomRank.
[0122] Design RMSD regression IhnrtRank i'a s' Send ag hewnganenus RNAs Ranking hc»m«gme<m RNA D&ss Ms; LRtNet making ke» Rpnt data m mini-batch lb - IA Hmuogejwsnns RNA k: RMSD Nksdtl OURSSt RMSD (exact wine } Ranking tsrdee Relative valve)
[0123]
[0124] Table 5: RMSD Regression vs. HomRank.
[0125] The architecture of the present ranking model employs an atomic rotationally equivalent scorer backbone, with key modifications including the design of a homogeneous RNA sampler to ensure homogeneous candidates within a minibatch and the integration of the list-based ranking technique to predict their ranking scores. The neural network is a sequential model comprising an atomic embedding layer to initialize atom representation, a plurality of blocks (in the embodiment of the experiments, three blocks) to capture structural equivariance and invariance, and a number of fully-connected layers corresponding to the number of blocks, to fuse representation and generate the ranking scores - e.g., an E3NN architecture. Each block includes a self-interaction layer, an equivariant convolution layer, a pointwise normalization layer, followed by a self-interaction layer, and a non-linearity layer.
[0126] The present ranking models are trained on four Nvidia A100 GPUs, each equipped with 80GB of memory. The training leverages Distributed Data Parallel (DDP) to ensure synchronous gradient updates across GPUs. The batch size for each GPU is 32, with each mini-batch comprising 32 randomly sampled mini-lists n, each containing 10 candidates. The optimization employs the Adam optimizer with a learning rate 5e-3. The temperature coefficient T for the softmax transformation is fixed at 0.1. The threshold 5 for ARMSD is set to 0.1 A. Consistent with the settingsin ARES, the model is trained for a single epoch. For Hom Rank- PT, the model is optimized using the MSE loss during the pre-training stage, with batch size and learning rate aligning with. In the training stage, the parameter configurations remain consistent with HomRank.
[0127] The models were tested using a variety of evaluation metrics:
[0128] (i) Minimum RMSD: Measures the minimum RMSD among the top- k evaluations:
[0129] Minimum RMSD = min Rh(6) iei, 2, -,k
[0130] where {R1, R2,---, Rk represent the ground-truth RMSD corresponding to the top- k best-scoring candidates.
[0131] (ii) Mean RMSD: Measures the arithmetic mean RMSD among the top- k evaluations:
[0132] k Mean RMSD = -^ (7)
[0133]
[0134] i=l
[0135] (iii) Success Rate (SRate): The success rate SRate, for the i-th RNA, is a binary indicator defined as follows,
[0136] >_ > fl, if the top- k candidates include near-natives ( RMSD < 2A),
[0137] j — j ff\O /
[0138] [0, otherwise.
[0139] >
[0140] Average Success Rate ( SRate ): The average success rate is the arithmetic >
[0141] mean of SRate i, across m RNAs, i.e., SRate = -^ ™iSRate,.
[0142] (iv) Retrieval Rate (RRate): The retrieval rate RRate i, for the i-th RNA, is the proportion of near-native conformations retrieved within the top- k evaluations:
[0143] hi
[0144] RRate, = (9) min( / c,r,) 'where htdenotes the number of near-native candidates retrieved within the top- k best-scoring candidates, r, is the total number of near-natives, and min(k,r,) accounts for cases where r, fewer than k.
[0145] Average Retrieval Rate ( RRate): The average retrieval rate is the arithmetic mean of RRate it across m RNAs, i.e., RRate =
[0146]
[0147] RRate,.
[0148] (v) Spearman's rank correlation coefficient ( p ): The Spearman's rank correlation coefficient, pirfor the i-th RNA is a nonparametric measure of rankbased correlation, measuring the alignment between the ranking of candidates sorted by predicted score and the real ranking sorted by ground-truth RMSD. Let n, be the total number of candidates for the i-th RNA, and let
[0149]
[0150] be the difference in ranks for the j-th candidate. The coefficient p is calculated by,
[0151]
[0152] where a higher pt(upper bound is +1 ) indicates stronger alignment between the predicted ranking and the ground-truth ranking, while a lower value (lower bound is -1) implies a strong inverse relationship, and 0 denotes no correlation.
[0153] Average Spearman's rank correlation coefficient ( p ): The average Spearman correlation is the arithmetic mean of, across m RNAs, i.e., p =
[0154]
[0155] The present listwise ranking model was experimentally evaluated on the RNA Puzzles benchmark. Figure 5, images (a) to (c) illustrate the overall comparison on the RNA Puzzles benchmark, which consists of 21 RNAs from the RNA Puzzles structure prediction challenge. The structural candidates were generated by Rosetta FARFAR2 sampling method. Each RNA in this benchmark comprises at least 1,500 structural candidates, including one native conformation, 1% near-native structures (RMSD
[0156]
[0157] 2°A), and the remaining as decoy candidates (RMSD > 2°A). Figure3(a) depicts the minimum RMSD among the top 1, 10, 100 best-scoring candidates. Thepresent ranking model (termed HomRank in Figure 5) achieves the strongest performance in identifying near-native candidates, followed by HomRank-PT (essentially the same model as HomRank, but implementing a pre-training phase involving mean-squared error-based RMSD loss per ARES) and ARES +. The improvements of ARES+ over ARES indicate the benefits conferred by new training dataset. Among the five baseline methods, cgRNASP emerges as the most competitive, whereas RNA3DCNN exhibits the worst performance.
[0158] Figure 5, image (b) presents the average success rate of identifying nearnatives in the top-k evaluations. The present ranking model identifies 95% near-natives in the top-1 evaluations, outperforming HomRank-PT and cgRNASP (both at 90%), as well as ARES+ (86%), ARES (71%), lociPARSE (52%), Rosetta (48%), and RNA3DCNN (5%). Furthermore, the present ranking model is the only method to achieve a 100% average success rate within the top-5 evaluations, surpassing ARES+ and cgRNASP, which require top-20 evaluations, as well as HomRank-PT, which needs top-100 evaluations to reach a comparable success rate. In contrast, ARES, Rosetta, RNA3DCNN, lociPARSE fail to achieve 100% even when the search range is expanded to the top-1000 candidates, with their metrics remaining between 71% and 95%, indicating their limited scoring capacity.
[0159] In the case study of RNA Puzzle 11 shown in Figure 5, image (c), the present ranking model is compared with ARES and two leading competitors, cgRNASP and lociPARSE. The present ranking model accurately identifies the native conformation as the best-scoring candidate, while the baselines erroneously select decoy candidates with significantly higher RMSD values: ARES (RMSD = 10.95 A), cgRNASP (RMSD = 18.77 A), and lociPARSE (RMSD=8.18 A).
[0160] The present ranking model was further evaluated on theTestCLS-1 Benchmark as shown in Figure 6, images (a) to (c). Relevantly, model generalization is crucial for the RNA 3D structure evaluation task. While the present ranking model was trained on homogenous conformations, that does not of itself ensure generalisability. Similarly, existing deep learning-based scoring methods mayperform well on in-distribution data, yet may exhibit degraded accuracy on new RNAs, hindering practical usage of such tools. The TestCLS-1 benchmark comprises 21 RNAs belonging to CLS-1 - a cluster distinct from the training data, providing a rigorous evaluation of model generalization on unseen structures.
[0161] Figure 6(a) presents the minimum RMSD for the top 1, 10, 100 best-scoring candidates. The present ranking model, HomRank-PT, and ARES+ consistently outperform the compared baselines. ARES+ demonstrates significant improvements over ARES, highlighting the contributions of the new curated dataset, described with reference to Figure 1, in enhancing model generalization. HomRank-PT achieves further improvements, validating the present technical contribution that homogeneous RNA ranking effectively enhances top-k retrieval performance. For the compared baselines, cgRNASP is the strongest competitor. However, its performance on TestCLS-1 benchmark is significantly inferior to that on RNA Puzzles benchmark, indicating its limited generalization and adaptability on unseen RNAs.
[0162] Figure 6(b) quantifies the average success rate of identifying near-natives within the top-k evaluations, assessing the capacity to identify near-native conformations. The present ranking model, HomRank-PT, and ARES+ identify 86%, 81%, and 71% of near-native molecules in the top-1 evaluations, significantly surpassing cgRNASP (43%), ARES (33%), lociPARSE (33%), Rosetta (5%), and RNA3DCNN (0%). Besides, the present ranking model, HomRank-PT, and ARES+ achieve 100% average success rate within the top-100, top-50, and top-100 evaluations, respectively. However, even when the search range is extended to the top-1000 candidates, the metric for the five baselines ranges from 29% to 95%.
[0163] Figure 6(c) depicts the average retrieval rate of near-natives, measuring their proportion within the top-k evaluations. It can be observed that the present ranking model, HomRank-PT, and ARES+ consistently exceed the baselines by substantial margins, demonstrating superior retrieval accuracy and reliability.
[0164] In summary, the two knowledge-based baselines and three recent deep learningbaselines struggle to generalize on the challenging TestCLS-1 benchmark. ARES+ demonstrates notable improvement over ARES, highlighting the enhanced performance powered by the new dataset. Furthermore, the present ranking model and its loss-modified similar model, HomRank-PT, show promising capacity in retrieving near-native candidates, validating the effectiveness of homogeneous RNA ranking technique.
[0165] Figure 7 showcases the best-scoring candidates identified by each method. For the first case, 5A2Q_2, the present ranking model, HomRank-PT, and ARES+ successfully identify the native molecule, while the baselines mistakenly select decoy structures with significant RMSD deviations ranging from 15.70 A to 23.21 A. In the second case, 1GID_1, the present ranking model is the only model to accurately pinpoint the native conformation, showing the most competitive performance. Although HomRank-PT fails to discern the native one, it selects a higher-fidelity candidate (RMSD = 5.44 A) than ARES+ (RMSD=6.52 A) and other baseline methods, whose RMSD spanning 5.45 A to 7.50 A. In summary, the two knowledge-based baselines and three recent deep learning baselines struggle to generalize on the challenging TestCLS-1 benchmark. ARES+ demonstrates notable improvement over ARES, highlighting the enhanced performance powered by the new dataset. Furthermore, the present ranking model and HomRank-PT show promising capacity in retrieving near-native candidates, validating the effectiveness of homogeneous RNA ranking technique.
[0166] Further experiments investigated model performance in ranking pairs with varying structural similarity measured by ARMSD. These experiments mirror the practical scenario faced by structural biologists: when comparing pairs with large ARMSD, distinguishing the candidate closer to the native conformation is relatively straightforward. However, as ARMSD decreases, structural similarities increase, making the judgment more challenging and raising the likelihood of mistakes.
[0167] Twelve RNAs were selected from the TestCLS-1 and TestCLS-2 and six initial ARMSD intervals were defined: (0 A, 2 A], (2 A, 4 A], (4 A, 6 A], (6 A, 8 A], (8 A, 10 A], and > 10 A. These intervals are adaptively adjusted based on theground-truth RMSD values of RNA candidates. For each interval, a random sampling of 1,000 pairs was made and accuracy calculated, i.e., the proportion of correct pairwise rankings.
[0168] Figure 8 shows that, for most models, the ranking accuracy improves as ΔRMSD increases, similar to the evaluation paradigm of domain experts. The present ranking model, HomRank-PT (also the present ranking model, but with modified pre-training and loss function), and ARES+ consistently achieve higher accuracy than other baselines in most cases. An interesting finding is that Rosetta and lociPARSE show a negative correlation in the case of 2WW9_4, with accuracy dropping from approximately 50% for ΔRMSD ∈ (0 Å, 2 Å] to around 20% for RMSD ≥ 10 Å. Similarly, for 4U7U_6, the accuracy of cgRNASP and ARES declines sharply, from 40% for ΔRMSD ∈ (0 Å, 2 Å] to below 10% for ΔRMSD ≥ 10 Å. These results suggest that Rosetta, lociPARSE, cgRNASP, and ARES have limited capacity to generalize to cases with significant structural discrepancies, likely due to biases in the training data or inherent deficiency in capturing critical structural features.
[0169] Another observation is that incorporating homogeneous RNA ranking into the training process effectively mitigates negative correlation and improves ranking accuracy - in the case of 4U7U_6, ARES exhibited a steep drop in accuracy as ΔRMSD increases, ARES+ improves the performance over ARES but still retains a negative correlation. In contrast, HomRank-PT not only surpasses ARES and ARES+ with higher accuracy but also reverses the trend, achieving a positive correlation. These results demonstrate the effectiveness of homogeneous RNA ranking in capturing key structural features that differentiate high-fidelity candidates from the low-fidelity ones, enabling models to reliably identify the candidate closer to the native conformation.
[0170] A follow-up experiment sought to determine whether the present ranking model could identify key biological characteristics without explicit specification in the model. To that end, RNA double helices were introduced with varying inter-strand distances, shown in Figure 9, image (a). Figure 9, image (b) compares the results of the present ranking model and ARES. For convenience, the negative rankingscore is plotted for the present ranking model. Both models accurately assign lowest score when the inter-strand distance of the RNA double helices approaches the experimentally determined optimal distance (vertical line), corresponding to ideal base pairing. As the inter-strand distance decreases below the equilibrium distance (2nd point), both the present ranking model and ARES assign higher scores. This behaviour is attributed to significantly increased repulsive interactions at short-range distance. Conversely, when the distance exceeds the optimal one, the present ranking model exhibits distinct behaviour depending on the distance of separation. Specifically, its negative score firstly increases, reaching a peak (3rd point), then decreases and eventually stabilizes at a distant separation (4th point). The rising phase can be explained by the growing dominance of attractive forces, whereas the subsequent falling phase occurs as the strands move too far apart for significant interactions, resulting in a rapid decline of attractive forces. In the stable phase, the distance is sufficiently large to render interactions negligible. However, these finer-grained characteristics, particularly the transitions within the curve, are not captured by ARES. These findings indicate the present ranking model has superior sensitivity to the intricate interplay of structural and interaction dynamics in RNA double helices across varying inter-strand distance.
[0171] The skilled person will appreciate that the methods described above can be implemented in a computing system, an example of which is shown in Figure 10. That system 1000 may comprise hardware, firmware or software components, or any combination thereof, for implementing the method for training a ranking model to RNA structures based on an input RNA sequence. The system 1000 includes non-transitory memory 1010 for storing instructions the execution of which dictates the behaviour of the system 1000. The system 1000 comprises at least one processor 1008, which may be housed in a single server or distributed in servers 1006 across a network 1002.
[0172] One or more receivers 1012 (which also includes transceivers and other devices capable of receiving data) of the system 1000 receive the RNA conformation data set. The computationally simulated structures may be generated externally, or by using a structure simulator 1014 of the system to generate each computationally simulated structure by simulating RNA structures using a domain-knowledgemodel specifying lowest energy base pairs. The structure simulator 1014 will then transmit the one or more computationally simulated structures to the one or more receivers 1012.
[0173] An alignment module 1016 of the system 1000 performs the pairwise alignment step and a clustering module 1018 of the system 1000 then performs the clustering step. The system 1000 also includes a ranking model module 1020 comprising a ranking model 1022, that performs the list generation, ranking and updating steps, after which the output 1024 of the system 1000 outputs the trained ranking model 1022, the trained ranking model 1022 being configured to output the Top-n structures from a top ranked said list.
[0174] It will be appreciated that many further modifications and permutations of various aspects of the described embodiments are possible. Accordingly, the described aspects are intended to embrace all such alterations, modifications, and variations that fall within the spirit and scope of the appended claims.
[0175] Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.
[0176] The reference in this specification to any prior publication (or information derived from it), or to any matter which is known, is not, and should not be taken as an acknowledgment or admission or any form of suggestion that that prior publication (or information derived from it) or known matter forms part of the common general knowledge in the field of endeavour to which this specification relates.
Claims
Claims1. A computer-implemented method for training a ranking model to rank ribonucleic acid (RNA) structures based on an input RNA sequence, comprising:receiving an RNA conformation data set comprising a plurality of non- redundant single-chain RNA sequences, and, for each sequence, a plurality of structures including a native single-chain structure and one or more computationally simulated homogeneous RNA structures (CSHRS);pairwise aligning the RNA sequences;computing a similarity score for each pair of the RNA sequences; clustering the RNA sequences into a plurality of clusters of RNA sequences based on the similarity scores;for each RNA sequence in a first said cluster:generating a plurality of lists, each list being a permutation of the plurality of structures corresponding to the respective RNA sequence;applying a ranking model to predict a ranking distribution over each list, corresponding to a probability that an order of the structures in the respective list is accurate; andupdating the ranking model based on a difference between the predicted ranking distribution and a ground-truth ranking distribution; andoutputting the trained ranking model, the trained ranking model configured to output a Top-n structures from a top ranked said list.
2. The method of claim 1, wherein applying the ranking model to predict a ranking distribution over each list, comprises determining a probability that an order of a Top-k said structures in the respective list is accurate.
3. The method of claim 2, wherein k = 1.
4. The method of any one of claims 1 to 3, comprising:receiving a further RNA sequence;receiving or generating a plurality of further conformations for the further RNA sequence;generating a further list, the further list being a permutation of the plurality of the further conformations;applying the ranking model to predict a ranking distribution over the further list; andbased on the ranking distribution, outputting a Top-1 said further conformation as a most likely candidate structure for the further RNA sequence.
5. The method of any preceding claim, wherein updating the ranking model based on a difference between the predicted ranking distribution and a ground-truth ranking distribution comprises computing a cross-entropy loss between probabilities associated with individual structures in the list and corresponding probabilities in the ground-truth ranking distribution, and updating based on the cross-entropy loss.
6. The method of any preceding claim, wherein computing the similarity scores comprises generating a sequence similarity map and clustering structures based on branches of the sequence similarity map.
7. The method of any preceding claim, wherein clustering the RNA sequences into a plurality of groups of RNA sequences comprises one of K-means clustering, density-based clustering, graph-based clustering, or hierarchical clustering.
8. The method of any preceding claim, wherein the one or more computationally simulated structures are generated by simulating RNA structures using a domain-knowledge model specifying lowest energy base pairs.
9. The method of any preceding claim, wherein receiving the one or more CSHRS comprises generating the CSHRS by molecular dynamics (MD) simulations.
10. The method of any preceding claim, wherein, for each RNA sequence in the first cluster, the plurality of lists is generated by randomly sampling structures from the respective plurality of structures, to generate a minibatch, and generating the plurality of lists by permuting the mini-batch.
11. The method of any preceding claim, wherein pairwise sequence alignment uses universal structural alignment, wherein the structural similarity score is a template modelling score.
12. A method for generating a candidate structure for a specific RNA sequence, the method comprising:receiving the specific RNA sequence;receiving or generating a plurality of new conformations corresponding to the specific RNA sequence;generating a new list, the new list being a permutation of the plurality of the new conformations;applying a ranking model trained according to the method of claim 1, to the new list, to predict a ranking distribution over the new list; and based on the ranking distribution, outputting a Top-1 said new conformation as a most likely candidate structure for the further RNA sequence.
13. A system for training a ranking model to rank ribonucleic acid (RNA) structures based on an input RNA sequence, comprising:memory;at least one processor;one or more receivers;an alignment module;a clustering module;a ranking model module comprising a ranking model; andan output;wherein the memory stores instructions that, when executed by the at least one processor, cause the system to:receive, at the one or more receivers, an RNA conformation data set comprising a plurality of non-redundant single-chain RNA sequences, and, for each sequence, a plurality of structures including a nativesingle-chain structure and one or more computationally simulated homogeneous RNA structures (CSHRS);pairwise aligning the RNA sequences at the alignment module; computing a similarity score, at the similarity assessor, for each pair of the RNA sequences;using the clustering module to cluster the RNA sequences into a plurality of clusters of RNA sequences based on the similarity scores; for each RNA sequence in a first said cluster, applying the ranking module to:generate a plurality of lists, each list being a permutation of the plurality of structures corresponding to the respective RNA sequence;apply the ranking model to predict a ranking distribution over each list, corresponding to a probability that an order of the structures in the respective list is accurate; and update the ranking model based on a difference between the predicted ranking distribution and a ground-truth ranking distribution; andoutput the trained ranking model at the output, the trained ranking model configured to output a Top-n structures from a top ranked said list.
14. The system of claim 13, wherein applying the ranking model to predict a ranking distribution over each list, comprises determining a probability that an order of a Top-k said structures in the respective list is accurate.
15. The system of claim 13 or 14, wherein the ranking module updates the ranking model by computing a cross-entropy loss between probabilities associated with individual structures in the list and corresponding probabilities in the ground-truth ranking distribution.
16. The system of any one of claims 13 to 15, wherein the clustering module computes generates a sequence similarity map and clusters structures based on branches of the sequence similarity map.
17. The system of any one of claims 13 to 16, comprising a structure simulator for generating the one or more computationally simulated structures by simulating RNA structures using a domain-knowledge model specifying lowest energy base pairs, and transmits the one or more computationally simulated structures to the one or more receivers.
18. The system of any one of claims 13 to 16, wherein receiving the one or more CSHRS comprises generating, by the one or more processors, the CSHRS by molecular dynamics (MD) simulations.
19. The system of any one of claims 13 to 18, wherein, for each RNA sequence in the first cluster, the plurality of lists is generated by randomly sampling structures from the respective plurality of structures, to generate a minibatch, and generating the plurality of lists by permuting the mini-batch.
20. The system of any one of claims 13 to 19, wherein pairwise sequence alignment uses universal structural alignment, wherein the structural similarity score is a template modelling score.