Method and system for sampling datasets for identification of druggable sites in RNA

The method and system address the limitations of conventional RNA druggability prediction by using spatial proximity and non-redundant dataset generation to train ML models, improving the accuracy of RNA-target prioritization and drug discovery.

WO2026013692A1PCT designated stage Publication Date: 2026-01-15INDIAN INST OF TECH MADRAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IN2025/050992
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-11
Filing Date
2025-07-04
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Conventional computational approaches for druggability prediction are inadequate for RNA-ligand interactions due to structural, electrostatic, and conformational differences, failing to capture subtle variations in RNA binding site geometry and dynamics, and lacking capability to resolve binding site-level predictions.

Method used

A method and system for sampling datasets that identify druggable sites in RNA structures by determining binding sites based on spatial proximity, generating a non-redundant negative dataset, and training a Machine Learning model using structural features to improve prediction accuracy.

Benefits of technology

Enables precise identification and prediction of druggable sites in RNA structures, enhancing the reliability of RNA-target prioritization and supporting drug discovery by providing balanced and non-redundant training data for ML models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2025050992_15012026_PF_FP_ABST
    Figure IN2025050992_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to method and system (104) for sampling datasets for identification of druggable sites in Ribonucleic Acids (RNAs). The method comprises receiving, by a system, a dataset of RNA structures bound to small molecules, from one or more sources. Further, the method comprises determining a plurality of binding sites from the dataset, based on a spatial proximity threshold between the RNA structures and ligand atoms. A set of structural features are determined from the plurality of binding sites to obtain a positive dataset of druggable sites. A plurality of negative binding sites are predicted based on the positive dataset of druggable sites. A non-redundant negative dataset is generated by resolving inter-source and intra source redundancy among the plurality of negative binding sites, wherein the non-redundant negative dataset is used with the positive dataset to train an ML model for identification of the druggable sites in the RNA structures.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR SAMPLING DATASETS FOR IDENTIFICATION OF DRUGGABLE SITES IN RNATECHNICAL FIELD

[0001] The present disclosure generally relates to the field of data analytics. More particularly, the present disclosure relates to a method and a system for sampling datasets for identification of druggable sites in Ribonucleic Acids (RNAs).BACKGROUND

[0002] Ribonucleic Acids (RNAs) play a critical role in a wide array of cellular processes, ranging from gene regulation and transcriptional control to enzymatic catalysis. Dysregulation of both coding and non-coding RNA species has been implicated in a variety of pathological conditions, including cancer, viral infections, and genetic disorders.

[0003] The increasing prevalence of diseases where traditional protein targets are inaccessible, ineffective, or insufficient has led to the exploration of RNA molecules as alternative therapeutic targets. However, the identification and prioritization of druggable binding sites in RNA structures remain a significant technical challenge.

[0004] Conventional computational approaches for druggability prediction have been largely optimized for protein targets and rely on structural characteristics unique to protein-ligand interfaces. As a result, these models exhibit limited or no applicability when repurposed for RNA-ligand interactions due to fundamental differences in structural, electrostatic, and conformational properties of RNA binding sites.

[0005] Furthermore, existing RNA-targeted computational tools are either limited to wholemolecule target assessment or lack the capability to resolve binding site-level predictions. In particular, current methodologies are inadequate in handling the high degree of conformational flexibility exhibited by RNA structures, including changes between apo (unbound), and holo (ligand-bound) forms, or variations introduced by point mutations.

[0006] Moreover, existing methods are unable to capture subtle, yet biologically significant, variations in local binding site geometry, surface characteristics, and dynamic behaviour across RNA conformers, thereby posing a technical challenge in prediction of RNA druggability.

[0007] Accordingly, there exists a need for improved approaches for sampling datasets for identification of druggable sites in RNA.

[0008] The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the invention and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art.SUMMARY

[0009] In an embodiment, the present disclosure discloses a method of sampling datasets for identification of druggable sites in Ribonucleic Acids (RNAs). The method comprises receiving a dataset of RNA structures bound to small molecules, from one or more sources. Further, the method comprises determining a plurality of binding sites from the dataset, based on a spatial proximity threshold between the RNA structures and ligand atoms. A set of structural features are determined from the plurality of binding sites to obtain a positive dataset of druggable sites. A plurality of negative binding sites are predicted based on the positive dataset of druggable sites. A non-redundant negative dataset is generated by resolving intersource and intra source redundancy among the plurality of negative binding sites, wherein the non-redundant negative dataset is used with the positive dataset to train a Machine Learning (ML) model for identification of the druggable sites in the RNA structures.

[0010] In an embodiment, the present disclosure discloses a system for sampling datasets for identification of druggable sites in RNAs. The system comprises one or more processors and a memory. The one or more processors are functionally coupled to the memory, the one or more processors being responsive to computer-executable instructions contained in the program code and operative to receive a dataset of RNA structures bound to small molecules, from one or more sources. Further, the one or more processors are configured to determine a plurality of binding sites from the dataset, based on a spatial proximity threshold between the RNA structures and ligand atoms. A set of structural features are determined from the plurality ofbinding sites to obtain a positive dataset of druggable sites. A plurality of negative binding sites are predicted based on the positive dataset of druggable sites. A non-redundant negative dataset is generated by resolving inter-source and intra source redundancy among the plurality of negative binding sites, wherein the non-redundant negative dataset is used with the positive dataset to train a Machine Learning (ML) model for identification of the druggable sites in the RNA structures.

[0011] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWINGS

[0012] The novel features and characteristics of the disclosure are set forth in the appended claims. The disclosure itself, however, as well as a preferred mode of use, further objectives, and advantages thereof, will best be understood by reference to the following detailed description of an illustrative embodiment when read in conjunction with the accompanying figures. One or more embodiments are now described, by way of example only, with reference to the accompanying figures wherein like reference numerals represent like elements and in which:

[0013] Figure 1 illustrates an exemplary environment for sampling datasets for identification of druggable sites in Ribonucleic Acids (RNAs), in accordance with some embodiments of the present disclosure;

[0014] Figure 2 illustrates a detailed diagram of a system for sampling datasets for identification of druggable sites in RNAs, in accordance with some embodiments of the present disclosure;

[0015] Figure 3 illustrates an exemplary user interface of an application for the prediction of druggable binding sites in RNA using the trained ML model, in accordance with an embodiment of the present disclosure;

[0016] Figure 4 shows an exemplary flow chart illustrating a method for sampling datasets for identification of druggable sites in RNAs, in accordance with some embodiments of the present disclosure; and

[0017] Figure 5 shows a block diagram of a general-purpose computing system for sampling datasets for identification of druggable sites in RNAs, in accordance with embodiments of the present disclosure.

[0018] It should be appreciated by those skilled in the art that any block diagram herein represents conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown.DETAILED DESCRIPTION

[0019] In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.

[0020] While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternatives falling within the scope of the disclosure.

[0021] The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises. . . a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or apparatus.

[0022] Machine Learning (ML) models are increasingly used in computational biology applications, including molecular screening, target prediction, and structural analysis. However, the inherent complexity of biological systems, such as structural variability in RNA molecules, differences in ligand binding conformations, and limited labelled datasets, poses significant challenges for effectiveness of such models. The present disclosure relates to a method and a system for sampling datasets for identification of druggable sites in RNAs to obtain both a positive dataset and a non-redundant negative dataset of RNA-small molecule binding sites. The datasets are sampled by identifying binding site regions based on spatial proximity thresholds and by resolving structural redundancies across data sources.

[0023] Figure 1 illustrates an exemplary environment 100 for sampling datasets for identification of druggable sites in Ribonucleic Acids (RNAs), in accordance with some embodiments of the present disclosure. The exemplary environment 100 comprises one or more data sources 102, a system 104, a Machine Learning (ML) model 110, and a network 112.

[0024] The ML model 110, is a computing system with interconnected nodes. The ML model 110 recognizes hidden patterns and correlations in data, regions and classifies them over time for continuous learning and improvement. In an embodiment, the ML model 110 may comprise a support vector machine. In some non -limiting embodiments, the ML model 110 may comprise at least one of a discriminant analyses, including Linear Discriminant Analysis (LDA) and Quadratic Discriminant Analysis (QDA), gaussian processes, tree-based methods, boosting methods, a combination thereof, and the like. A person skilled in the art will appreciate that the present disclosure is applicable to any ML model other than the above-mentioned ML models.

[0025] The network 112 may include one or more wired and / or wireless networks. For example, the network 112 may include a cellular network (e.g., a Long-Term Evolution (LTE) network, a third generation (3G) network, a fourth generation (4G) network, a Code Division Multiple Access (CDMA) network, and / or the like), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., the Public Switched Telephone Network (PSTN)), a private network (e.g., a private network associated with a transaction service provider), an ad hoc network, an intranet, the Internet, a fiber optic-based network, a cloud computing network, and / or the like, and / or a combination of these or other types of networks.

[0026] The system 104 receives a dataset of RNA structures bound to small molecules, from the one or more data sources 102. The system 104 determines a plurality of binding sites from the dataset, based on a spatial proximity threshold between the RNA structures and ligand atoms. The system 104 determines a set of structural features from the plurality of binding sites to obtain a positive dataset of druggable sites. The system 104 predicts a plurality of negative binding sites, based on the positive dataset of druggable sites.

[0027] The system 104 generates a non-redundant negative dataset by resolving inter-source and intra source redundancy among the plurality of negative binding sites, wherein the non- redundant negative dataset is used with the positive dataset to train a Machine Learning (ML) model 110 for identification of the druggable sites in the RNA structures. The system 104 may include Input / Output (VO) interface 106, a memory 107, and Central Processing Units 105 (also referred as “CPUs” or “one or more processors 105”. In some embodiments, the memory 107 may be communicatively coupled to the one or more processors 105. The memory 107 stores instructions executable by the one or more processors 105. The one or more processors 105 may comprise at least one data processor for executing program components for executing user or system-generated requests. The memory 107 may be communicatively coupled to the one or more processors 105. The memory 107 stores instructions, executable by the one or more processors 105, which, on execution, may cause the one or more processors 105 to sample datasets for identification of druggable sites in RNAs. The I / O interface 106 is coupled with the one or more processors 105 through which an input signal or / and an output signal is communicated. For example, the dataset of RNA structures may be received from a user via the I / O interface 106. In an embodiment, the system 104 may be implemented in a variety of computing systems, such as a laptop computer, a desktop computer, a Personal Computer (PC), a notebook, a smartphone, a tablet, a server, a network server, a cloud-based server, and the like. In some non-limiting embodiments, the system 104 may be implemented in an Interactive Development Environment (IDE) or an interactive code and documentation platform.

[0028] The system 104 is of advantage that the system 104 enables identification and prediction of druggable sites by generating a balanced and non-redundant dataset comprising both positive and negative RNA-ligand binding sites. The system 104 utilizes a structured approach for non- redundant negative dataset generation, utilizing techniques such as, but not limited to,backbone motif search, flexible blind docking, and binding site prediction methods, thereby providing a biologically effective training data for the ML model 110.

[0029] Further, the system 104 integrates feature-level representations of binding sites using specific structural categories, such as atomic composition, surface area, shape, torsional characteristics and pharmacophore content to improve the precision of the ML model 110. The trained machine learning model results in improved performance metrics on blind test datasets and also enables quantification of druggability score changes due to mutations or structural transitions. The system 104 further improves reliability of RNA-target prioritization and supports downstream applications in RNA-targeted drug discovery.

[0030] Figure 2 illustrates a detailed diagram 200 of the system 104 for sampling datasets for identification of druggable sites in RNAs, in accordance with some embodiments of the present disclosure. In an embodiment, the memory 107 may include one or more modules 202 and data 201. The one or more modules 202 may be configured to perform the steps of the present disclosure using the data 201, to sample datasets for identification of druggable sites in RNAs. In an embodiment, each of the one or more modules 202 may be a hardware unit which may be outside the memory 107 and coupled with the system 104. As used herein, the term modules 202 refers to an Application Specific Integrated Circuit (ASIC), an electronic circuit, a Field- Programmable Gate Arrays (FPGA), Programmable System-on-Chip (PSoC), a combinational logic circuit, and / or other suitable components that provide the described functionality. The one or more modules 202 when configured with the described functionality defined in the present disclosure will result in a novel hardware.

[0031] In one implementation, the one or more modules 202 may include, for example, a data input module 203, a binding site determination module 204, a positive dataset generation module 205, a negative binding site prediction module 206, and a redundancy resolution module 207. It will be appreciated that such aforementioned modules 202 may be represented as a single module or a combination of different modules. In one implementation, the data 201 may include, for example, the binding sites 101, the positive dataset 108, the negative binding sites 113, and non-redundant negative dataset 114.

[0032] In an embodiment, the data input module 203 may be configured to receive a dataset of RNA structures bound to small molecules from the one or more data sources 102. The one or more data sources 102 may include structural repositories such as Protein Data Bank (PDB).In some non-limiting embodiments, the data input module 203 may perform optional filtering of the acquired dataset to exclude synthetic aptamers or structures lacking biologically relevant ligands.

[0033] In an embodiment, the binding site determination module 204 is configured to determine a plurality of binding sites from the dataset based on a spatial proximity threshold between the RNA structures and ligand atoms. The binding site determination module 204 may identify binding site residues using a predefined distance criterion (e.g., 6.5 A) from ligand atoms, accounting for multiple conformers in the case of Nuclear Magnetic Resonance (NMR) structures, to handle RNA conformational dynamics.

[0034] In an embodiment, the positive dataset generation module 205 is configured to determine a set of structural features from the plurality of binding sites to obtain a positive dataset of druggable sites. In some non-limiting embodiments, the positive dataset generation module 205 extracts a plurality of structure-based descriptors spanning categories such as atomic composition, pharmacophore properties, surface area, spatial geometry and torsional strain, which are computed for each binding site.

[0035] In some non-limiting embodiments, the method further includes pruning the positive dataset comprising the RNA-small molecule complexes by filtering structures that: (a) lack biologically relevant small molecules or (b) contain synthetic RNA aptamers.

[0036] In an embodiment, the negative binding site prediction module 206 is configured to predict a plurality of negative binding sites, based on the positive dataset of druggable sites. The negative binding site prediction module 206 may employ at least one of a backbone motif search, blind docking simulations, or known pocket prediction methods to identify structurally plausible yet non-druggable binding regions across RNA structures. A person skilled in the art will appreciate that the present disclosure is applicable to any computational structural biology methods other than the above-mentioned methods. In some non-limiting embodiments, the negative binding site prediction module 206 predicts a plurality of non-redundant negative binding sites and thereby obtain the non-redundant negative dataset. The non-redundant negative dataset is used with the positive dataset to train a Machine Learning (ML) model 110 for identification of the druggable sites in the RNA structures.

[0037] In an embodiment, the redundancy resolution module 207 is configured to generate a non-redundant negative dataset by resolving inter-source and intra source redundancy among the plurality of negative binding sites, wherein the non-redundant negative dataset is used with the positive dataset to train a Machine Learning (ML) model 110 for identification of the druggable sites in the RNA structures.

[0038] In some non-limiting embodiments, a feature category of the set of structural features is selected from at least one of an atomic composition category, a pharmacophore category, a surface area category, a shape category, a roughness category, a torsion angles category and a puckering category.

[0039] In some non-limiting embodiments, the plurality of negative binding sites are predicted using at least one of a backbone motif search method, a flexible blind docking method, and a binding site prediction method.

[0040] In some non-limiting embodiments, the ML model 110 is a support vector machine (SVM) with a sigmoid kernel.

[0041] In some non-limiting embodiments, the method further includes (a) receiving an input comprising information associated with an RNA structure, from a user, (b) assigning a score to each of the identified druggable sites, based on the input, and (c) determine a set of positive binding sites and negative binding sites, based on the score.

[0042] In some non-limiting embodiments, the method further includes (a) receiving a representation of apo and holo RNA structures associated with an RNA target, and (b) assigning a score to each of the identified binding sites using the ML model 110.

[0043] In some non-limiting embodiments, the method further includes determining changes in the binding sites caused by single-point mutations by (a) generating representations of a plurality of mutated RNA structures upon introduction of the single-point mutations in the RNA structures associated with the positive dataset, and (b) determining a difference in druggability score between wild-type and mutated RNA structures based on the ML model 110 for determining changes in the binding sites caused by the single-point mutations.

[0044] Figure 3 illustrates an exemplary user interface 300 of an application for the prediction of druggable binding sites in RNA using the trained ML model 110, in accordance with an embodiment of the present disclosure. The user interface 300 enables input of an RNA structure file, selection of NMR model (if applicable), and specification of binding site residues or a ligand identifier.

[0045] A prediction mode may be selected, and relevant structural inputs may be provided, that are further processed by the ML model 110 to assign a druggability score to the binding site. The user interface 300 further enables uploading of pre-annotated binding site data (list of residues or a PDB identifier of the bound ligand) and includes controls to initiate or reset predictions. In some non-limiting embodiments, NMR structures are used for the prediction. The user interface 300 may enable a selection of a specific NMR model.

[0046] Figure 4 shows an exemplary flow chart illustrating a method for sampling datasets for identification of druggable sites in RNAs, in accordance with some embodiments of the present disclosure. As illustrated in Figure 4, the method 400 may comprise one or more steps. The method 400 may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform particular functions or implement particular abstract data types.

[0047] The order in which the method 400 is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the scope of the subject matter described herein. Furthermore, the method can be implemented in any suitable hardware, software, firmware, or combination thereof.

[0048] At step 402, the method receives a dataset of RNA structures bound to small molecules, from one or more data sources. At step 404, the method determines a plurality of binding sites from the dataset, based on a spatial proximity threshold between the RNA structures and ligand atoms. At step 406, the method determines a set of structural features from the plurality of binding sites to obtain a positive dataset of druggable sites. At step 408, the method predicts a plurality of negative binding sites, based on the positive dataset of druggable sites. At step 410,the method generates a non-redundant negative dataset by resolving inter-source and intra source redundancy among the plurality of negative binding sites, wherein the non-redundant negative dataset is used with the positive dataset to train a Machine Learning (ML) model for identification of the druggable sites in the RNA structures.

[0049] The method is of advantage that the method enables identification and prediction of druggable sites by generating a balanced and non-redundant dataset comprising both positive and negative RNA-ligand binding sites. The method utilizes a structured approach for non- redundant negative dataset generation, utilizing backbone motif search, flexible blind docking, and binding site prediction methods, thereby providing a biologically effective training data for the ML model 110.

[0050] Further, the method integrates feature-level representations of binding sites using specific structural categories, such as atomic composition, surface shape, torsional characteristics and pharmacophore content to improve the precision of the ML model 110. The machine learning model 110 results in improved performance metrics on blind test datasets and also enables quantification of druggability score changes due to mutations or structural transitions. The method further improves reliability of RNA-target prioritization and supports downstream applications in RNA-targeted drug discovery.

[0051] In some non-limiting embodiments, the method further includes pruning the dataset comprising the RNA-small molecule complexes by filtering structures that: (a) lack biologically relevant small molecules or (b) contain synthetic RNA aptamers.

[0052] In some non-limiting embodiments, a feature category of the set of structural features is selected from at least one of an atomic composition category, a pharmacophore category, a surface area category, a shape category, a roughness category, a torsion angles category and a puckering category.

[0053] In some non-limiting embodiments, the plurality of negative binding sites are predicted using at least one of a backbone motif search method, a flexible blind docking method, and a binding site prediction method.

[0054] In some non-limiting embodiments, the ML model 110 is a support vector machine (SVM) with a sigmoid kernel.

[0055] In some non-limiting embodiments, the method further includes (a) receiving an input comprising information associated with an RNA structure, from a user, (b) assigning a score to each of the identified druggable sites, based on the input, and (c) determine a set of positive binding sites and negative binding sites, based on the score.

[0056] In some non-limiting embodiments, the method further includes (a) receiving a representation of apo and holo RNA structures associated with an RNA target, and (b) assigning a score to each of the identified binding sites using the ML model 110.

[0057] In some non-limiting embodiments, the method further includes determining changes in the binding sites caused by single-point mutations by (a) generating representations of a plurality of mutated RNA structures upon introduction of the single-point mutations in the RNA structures associated with the positive dataset, and (b) determining a difference in druggability score between wild-type and mutated RNA structures based on the ML model 110 for determining changes in the binding sites caused by the single-point mutations.COMPUTER SYSTEM

[0058] Figure 5 illustrates a block diagram of an exemplary computer system 500 for implementing embodiments consistent with the present disclosure. In an embodiment, the computer system 500 may be the system 104. Thus, the computer system 500 may be used to sample datasets for identification of druggable sites in RNA. The computer system 500 may comprise a Central Processing Unit 502 (also referred as “CPU” or “processor”). The processor 502 may comprise at least one data processor. The processor 502 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc. The processor 502 may be used to realize the processor 105 described in Figure 1.

[0059] The processor 502 may be disposed in communication with one or more input / output (I / O) devices (not shown) via I / O interface 501. The I / O interface 501 may employ communication protocols / methods such as, without limitation, audio, analog, digital, monoaural, RCA, stereo, IEEE (Institute of Electrical and Electronics Engineers) -1394, serial bus, universal serial bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, digital visual interface (DVI), high-definition multimedia interface (HDMI), Radio Frequency (RF)antennas, S-Video, VGA, IEEE 802. n / b / g / n / x, Bluetooth, cellular (e.g., code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, or the like), etc.

[0060] Using the I / O interface 501, the computer system 500 may communicate with one or more I / O devices. For example, the input device 510 may be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, fax machine, dongle, biometric reader, microphone, touch screen, touchpad, trackball, stylus, scanner, storage device, transceiver, video device / source, etc. The output device 511 may be a printer, fax machine, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, Plasma display panel (PDP), Organic light-emitting diode display (OLED) or the like), audio speaker, etc. The I / O interface 501 may be used to realize the I / O interface 106 described in Figure 1.

[0061] The processor 502 may be disposed in communication with the communication network 509 via a network interface 503. The network interface 503 may communicate with the communication network 509. The network interface 503 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc. The communication network 509 may include, without limitation, a direct interconnection, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, etc. The network interface 503 may employ connection protocols include, but not limited to, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc.

[0062] The communication network 509 includes, but is not limited to, a direct interconnection, an e-commerce network, a peer to peer (P2P) network, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, WiFi, and such. The first network and the second network may either be a dedicated network or a shared network, which represents an association of the different types of networks that use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / internet Protocol (TCP / IP), Wireless Application Protocol (WAP), etc., to communicate with each other. Further, the first network and the second network may include a variety of network devices, including routers, bridges, servers, computing devices, storagedevices, etc. The computer system 500 may receive the input image 101 from a user, over a communication network 509.

[0063] In some embodiments, the processor 502 may be disposed in communication with a memory 505 (e.g., RAM, ROM, etc. not shown in Figure 5) via a storage interface 504. The storage interface 504 may connect to memory 505 including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as serial advanced technology attachment (SATA), Integrated Drive Electronics (IDE), IEEE- 1394, Universal Serial Bus (USB), fiber channel, Small Computer Systems Interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory devices, solid-state drives, etc.

[0064] The memory 505 may store a collection of program or database components, including, without limitation, user interface 506, an operating system 507, web browser 508 etc. In some embodiments, computer system 500 may store user / application data, such as, the data, variables, records, etc., as described in this disclosure. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle® or Sybase®. The memory 505 may be used to realize the memory 107 described in Figure 1. The memory 505 may be communicatively coupled to the processor 502. The memory 505 stores instructions, executable by the one or more processors 502, which, on execution, may cause the processor 502 to sample datasets for identification of druggable sites in RNA.\

[0065] The operating system 507 may facilitate resource management and operation of the computer system 500. Examples of operating systems include, without limitation, APPLE MACINTOSH® OS X, UNIX®, UNIX-like system distributions (E.G, BERKELEY SOFTWARE DISTRIBUTION™ (BSD), FREEBSD™, NETBSD™, OPENBSD™, etc ), LINUX DISTRIBUTIONS™ (E.G, RED HAT™, UBUNTU™, KUBUNTU™, etc ), IBM™ OS / 2, MICROSOFT™ WINDOWS™ (XP™, VISTA™ / 7 / 8, 10 etc ), APPLE® IOS™, GOOGLE® ANDROID™, BLACKBERRY® OS, or the like.

[0066] In some embodiments, the computer system (500) may implement the web browser (508) stored program component. The web browser (508) may be a hypertext viewing application, for example MICROSOFT® INTERNET EXPLORER™, GOOGLE® CHROME™, MOZILLA® FIREFOX™, APPLE® SAFARI™, etc. Secure web browsing maybe provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. Web browsers (508) may utilize facilities such as AJAX™, DHTML™, ADOBE® FLASH™, JAVASCRIPT™, JAVA™, Application Programming Interfaces (APIs), etc. In some embodiments, the computer system (500) may implement a mail server (not shown in Figure) stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as ASP™, ACTIVEX™, ANSI™ C++ / C#, MICROSOFT®, NET™, CGI SCRIPTS™, JAVA™, JAVASCRIPT™, PERL™, PHP™, PYTHON™, WEBOBJECTS™, etc. The mail server may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, the computer system (500) may implement a mail client stored program component. The mail client (not shown in Figure) may be a mail viewing application, such as APPLE® MAIL™, MICROSOFT® ENTOURAGE™, MICROSOFT® OUTLOOK™, MOZILLA® THUNDERBIRD™, etc.

[0067] .

[0068] Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., be non-transitory. Examples include Random Access Memory (RAM), Read-Only Memory (ROM), volatile memory, non-volatile memory, hard drives, Compact Disc Read-Only Memory (CD ROMs), Digital Video Disc (DVDs), flash drives, disks, and any other known physical storage media.

[0069] The terms "an embodiment", "embodiment", "embodiments", "the embodiment", "the embodiments", "one or more embodiments", "some embodiments", and "one embodiment" mean "one or more (but not all) embodiments of the invention(s)" unless expressly specified otherwise.

[0070] The terms "including", "comprising", “having” and variations thereof mean "including but not limited to", unless expressly specified otherwise.

[0071] The enumerated listing of items does not imply that any or all of the items are mutually exclusive, unless expressly specified otherwise. The terms "a", "an" and "the" mean "one or more", unless expressly specified otherwise.

[0072] A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary a variety of optional components are described to illustrate the wide variety of possible embodiments of the invention.

[0073] When a single device or article is described herein, it will be readily apparent that more than one device / article (whether or not they cooperate) may be used in place of a single device / article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it will be readily apparent that a single device / article may be used in place of the more than one device or article, or a different number of devices / articles may be used instead of the shown number of devices or programs. The functionality and / or the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality / features. Thus, other embodiments of the invention need not include the device itself.

[0074] The illustrated operations of Figures 4 and 5 show certain events occurring in a certain order. In alternative embodiments, certain operations may be performed in a different order, modified, or removed. Moreover, steps may be added to the above-described logic and still conform to the described embodiments. Further, operations described herein may occur sequentially or certain operations may be processed in parallel. Yet further, operations may be performed by a single processing unit or by distributed processing units. In some non-limiting embodiments, functionality of the ML model may be implemented in parallel using multi-core CPUs.

[0075] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an applicationbased here on. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.

[0076] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

We claim:

1. A method of sampling datasets for identification of druggable sites in Ribonucleic Acids(RNAs), the method comprising: receiving, by a system, a dataset of RNA structures bound to small molecules, from one or more sources; determining, by the system, a plurality of binding sites from the dataset, based on a spatial proximity threshold between the RNA structures and ligand atoms; determining, by the system, a set of structural features from the plurality of binding sites to obtain a positive dataset of druggable sites; predicting, by the system, a plurality of negative binding sites, based on the positive dataset of druggable sites; and generating, by the system, a non-redundant negative dataset by resolving inter-source and intra source redundancy among the plurality of negative binding sites, wherein the non- redundant negative dataset is used with the positive dataset to train a Machine Learning (ML) model for identification of the druggable sites in the RNA structures.

2. The method as claimed in claim 1, further comprising: pruning the dataset comprising the RNA-small molecule complexes by filtering structures that: (a) lack biologically relevant small molecules or (b) contain synthetic RNA aptamers.

3. The method as claimed in claim 1, wherein a feature category of the set of structural features is selected from at least one of an atomic composition category, a pharmacophore category, a surface area category, a shape category, a roughness category, a torsion angles category and a puckering category.

4. The method as claimed in claim 1, wherein the plurality of negative binding sites are predicted using at least one of a backbone motif search method, a flexible blind docking method, and a binding site prediction method.

5. The method as claimed in claim 1, wherein the ML model is a support vector machine with a sigmoid kernel.

6. The method as claimed in claim 1, further comprising:receiving an input comprising information associated with an RNA structure, from a user; assigning a score to each of the identified binding sites, based on the input; and determine a set of positive binding sites and negative binding sites, based on the score.

7. The method as claimed in claim 1, further comprising: receiving a representation of apo and holo RNA structures associated with an RNA target; and assigning a score to each of the identified druggable sites using the ML model.

8. The method as claimed in claim 1, further comprising determining changes in the binding sites caused by single-point mutations by: generating representations of a plurality of mutated RNA structures upon introduction of the single-point mutations in the RNA structures associated with the positive dataset; and determining a difference in druggability score between wild-type and mutated RNA structures based on the ML model for determining changes in the binding sites caused by the single-point mutations.

9. A system for sampling datasets for identification of druggable sites in Ribonucleic Acids(RNAs), the system comprises: a memory for storing executable program code; and a processor, functionally coupled to the memory, the processor being responsive to computer-executable instructions contained in the program code and operative to: receive a dataset of RNA structures bound to small molecules, from one or more sources; determine a plurality of binding sites from the dataset, based on a spatial proximity threshold between the RNA and ligand atoms; determine a set of structural features from the plurality of binding sites to obtain a positive dataset of druggable sites; predict a plurality of negative binding sites, based on the positive dataset of druggable sites; and generate a non-redundant negative dataset by resolving inter-source and intra source redundancy among the plurality of negative binding sites, wherein the non-redundant negativedataset is used with the positive dataset to train a Machine Learning (ML) model for identification of the druggable sites in the RNA structures.

10. The system as claimed in to claim 9, wherein the processor is further operative to: receive an input comprising information associated with an RNA structure, from a user; assign a score to each of the identified binding sites, based on the input; and determine a set of positive binding sites and negative binding sites, based on the score.