Methods, systems, and media methods for applying machine learning to chemical mapping data for RNA tertiary structure prediction

A machine learning-based method predicts RNA tertiary structure using chemical mapping data, overcoming the limitations of experimental methods by integrating chemical mapping directly into the prediction process, enhancing accuracy and efficiency.

JP2025529832APending Publication Date: 2025-09-09ATOMIC AI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025510292
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-19
Filing Date
2023-08-18
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Experimental methods for determining RNA tertiary structure, such as crystallography-based diffraction imaging and electron microscopy, are expensive and slow, and the flexibility of RNA molecules often prevents accurate structure determination.

Method used

A computer-implemented method using machine learning algorithms trained on chemical mapping data and tertiary structure data to predict RNA tertiary structure, bypassing the need for secondary structure determination and integrating chemical mapping data directly into the prediction process.

Benefits of technology

Enables efficient and cost-effective prediction of RNA tertiary structure, improving accuracy by directly utilizing chemical mapping data, even for RNA molecules with low sequence identity or unavailable data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025529832000001_ABST
    Figure 2025529832000001_ABST
Patent Text Reader

Abstract

Disclosed herein are methods, systems, and media for predicting the tertiary structure of a target RNA molecule, the methods, systems, and media including: creating a training dataset including chemical mapping data for one or more of a first plurality of RNA molecules and tertiary structure data for one or more of a second plurality of RNA molecules; training a machine learning algorithm using the training dataset; applying the trained learning algorithm to predict the tertiary structure of the RNA molecule of interest; and outputting the predicted tertiary structure of the RNA molecule of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 371,983, filed August 19, 2022, which is incorporated by reference in its entirety. [Background technology]

[0002] The 3D structure of biomolecules plays an important role in determining their function. Nucleic acids, such as RNA molecules, are a class of biomolecules. Despite this, experimental determination of RNA tertiary structure by crystallography-based diffraction imaging, electron microscopy, or nuclear magnetic resonance spectroscopy is expensive and slow. Furthermore, the flexibility of RNA molecules often makes it impossible to determine their structure using these experimental methods. Therefore, there remains an urgent need for an efficient and inexpensive method for determining the structure of RNA molecules. Summary of the Invention

[0003] In some aspects, the present disclosure provides a computer-implemented method for predicting a tertiary structure of an RNA molecule of interest, the method comprising: creating a training dataset, the training dataset comprising chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; training a machine learning algorithm using the training dataset; predicting the tertiary structure of the RNA molecule of interest by applying the trained machine learning algorithm; and outputting the predicted tertiary structure of the RNA molecule of interest.

[0004] In some embodiments, the present disclosure provides a computer-implemented method for predicting a tertiary structure of an RNA molecule of interest, the method comprising: obtaining a machine learning model, wherein the machine learning model is trained by a process comprising creating a training dataset comprising chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules, and training the machine learning model using the training dataset; applying the machine learning model to predict the tertiary structure of the RNA molecule of interest; and outputting the predicted tertiary structure of the RNA molecule of interest.

[0005] In some embodiments, the chemical mapping data is generated by a process comprising contacting an RNA molecule with a chemical probing agent, optionally the RNA molecule being at least one of the first plurality of RNA molecules or the RNA molecule of interest.

[0006] In some embodiments, the chemical probing agent comprises dimethyl sulfate (DMS).

[0007] In some embodiments, the chemical probing agent comprises a SHAPE (selective 2'-hydroxyl acylation and primer extension) reagent.

[0008] In some embodiments, the SHAPE reagent is 1-methyl-7-nitroisatoic anhydride (1M7), 1-methyl-6-nitroisatoic anhydride (1M6), 5-nitroisatoic anhydride (5NIA), or N-methyl-nitroisatoic anhydride (NMIA).

[0009] In some embodiments, the chemical probing agent comprises 2A3 ((2-aminopyridin-3-yl)(1H-imidazol-1-yl)methanone).

[0010] In some embodiments, the RNA molecule of interest comprises a portion of the transcriptome.

[0011] In some embodiments, the transcriptome is a human transcriptome.

[0012] In some embodiments, the training set includes chemical mapping data for at least about 10, 100, 500, 1,000, 10,000, or more than 10,000 sequences.

[0013] In some embodiments, the training set includes chemical mapping data for at most about 10, 100, 500, 1,000, or 10,000 sequences.

[0014] In some embodiments, the chemical mapping data is for sequences that occur in different abundances than in natural systems.

[0015] In some embodiments, the chemical mapping data was collected from an in vitro source.

[0016] In some embodiments, the method further comprises, prior to applying the machine learning model, tuning the machine learning algorithm based on chemical mapping data of the RNA molecule of interest.

[0017] In some embodiments, the machine learning algorithm comprises one or more artificial neural networks (ANNs).

[0018] In some embodiments, the method further comprises training an ANN to predict chemical mapping data for an RNA molecule of interest.

[0019] In some embodiments, the method further comprises predicting said chemical mapping data for the RNA molecule of interest from a predicted tertiary structure of the RNA molecule of interest.

[0020] In some embodiments, the method further comprises predicting a predicted tertiary structure of the RNA molecule of interest based on chemical mapping data for the RNA molecule of interest.

[0021] In some embodiments, the method further comprises using the same embedding to predict chemical mapping data for the RNA molecule of interest and a predicted tertiary structure of the RNA molecule of interest.

[0022] In some embodiments, the tertiary structure comprises the 3D coordinates of multiple atoms that make up the RNA molecule of interest.

[0023] In some embodiments, the tertiary structure comprises the 3D coordinates of each atom that makes up the RNA molecule of interest.

[0024] In some embodiments, the tertiary structure comprises one or more 3D coordinates of a plurality of nucleotides that make up the RNA molecule of interest.

[0025] In some embodiments, the tertiary structure comprises one or more 3D coordinates of each nucleotide that makes up the RNA molecule of interest.

[0026] In some embodiments, the tertiary structure of an RNA molecule of interest is parameterized based on a distance map.

[0027] In some embodiments, the tertiary structure of an RNA molecule of interest is parameterized based on distance maps and angles.

[0028] In some embodiments, the methods do not require the determination or prediction of the secondary structure of the RNA molecule of interest.

[0029] In some embodiments, the methods predict aspects of the tertiary structure of a target RNA that are not captured by the base-pairing prediction of the target RNA.

[0030] In some embodiments, the predicted tertiary structure comprises one or more of a pseudoknot, a multi-way junction, a coaxial stack, an a-minor motif, a kissing stem loop, a ribose zipper, or a tetraloop / tetraloop acceptor.

[0031] In some embodiments, the chemical mapping data comprises multidimensional chemical mapping data for one or more RNA molecules of the first plurality of RNA molecules.

[0032] In some embodiments, the predicted tertiary structure is a target for a pharmaceutical drug.

[0033] In some embodiments, the method further comprises determining a target region or subsequence of the RNA molecule of interest that the pharmaceutical agent targets based on the predicted tertiary structure.

[0034] In some embodiments, the method further comprises formulating a pharmaceutical agent based on the predicted tertiary structure.

[0035] In some embodiments, the training dataset further comprises a multiple sequence alignment of a third plurality of RNA molecules.

[0036] In some embodiments, the first plurality of RNA molecules and the second plurality of RNA molecules are the same or different.

[0037] In some embodiments, the first plurality of RNA molecules and the third plurality of molecules are the same.

[0038] In some embodiments, the first plurality of RNA molecules is unrelated to the RNA molecule of interest.

[0039] In some embodiments, the second plurality of RNA molecules is unrelated to the RNA molecule of interest.

[0040] In some embodiments, the third plurality of RNA molecules is unrelated to the RNA molecule of interest.

[0041] In some embodiments, one RNA molecule of the first plurality of RNA molecules has about 80%, 70%, 60%, 50%, 40%, 30%, 20%, or less sequence identity to the RNA molecule of interest.

[0042] In some aspects, the present disclosure provides a computer-implemented system for predicting the tertiary structure of an RNA molecule of interest, the system including a computing device including at least one processor and instructions executable by the at least one processor to perform operations, the operations including creating a training dataset, the training dataset including one or more of chemical mapping data for a first plurality of RNA molecules, tertiary structure data for a second plurality of RNA molecules, training a machine learning algorithm using the training dataset, predicting the tertiary structure of the RNA molecule of interest by applying the trained machine learning algorithm, and outputting the predicted tertiary structure of the RNA molecule of interest.

[0043] In some aspects, the present disclosure provides a non-transitory computer-readable storage medium executable by one or more processors and encoded with instructions for providing an application for predicting a tertiary structure of an RNA molecule of interest, the application including: a training dataset module configured to create a training dataset including chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; a training module configured to train a machine learning algorithm using the training dataset; an inference module configured to predict the tertiary structure of the RNA molecule of interest by applying the trained machine learning algorithm; and an output module configured to report the predicted tertiary structure of the RNA molecule of interest.

[0044] In some aspects, the present disclosure provides a non-transitory computer-readable medium comprising: accessing an RNA tertiary structure prediction system, the RNA tertiary structure prediction system manufactured by a process comprising creating a training dataset including chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules, training a machine learning model using the training dataset, and storing the trained machine learning model in a non-transitory computer-readable medium; and computer program code that, when executed by a computing system, causes the computing system to perform operations including predicting a tertiary structure of an RNA molecule of interest using the RNA tertiary structure prediction system and outputting the predicted tertiary structure of the RNA molecule of interest.

[0045] In some aspects, the present disclosure provides a computer-implemented method for predicting a tertiary structure of an RNA molecule of interest, the method comprising: sending a query to predict the tertiary structure of the RNA molecule of interest to a computer comprising a machine learning model, wherein the machine learning algorithm generates the tertiary structure, and the machine learning model has been trained by a process comprising creating a training dataset comprising chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules, and training the machine learning model using the training dataset; and receiving the predicted tertiary structure of the RNA molecule of interest from the computer.

[0046] Therefore, embodiments of the present disclosure include a method and system for RNA tertiary structure prediction by creating a training set including chemical mapping data for one or more RNA molecules of a first plurality of RNA molecules and tertiary structure data for one or more RNA molecules of a second plurality of RNA molecules.The training data may optionally include other types of data, such as multiple sequence alignments or predicted or observed secondary structures.The training set may include different types of data for each training item, and the chemical mapping data may be from multiple experimental methods or from the use of various parameters.The training set may then be used to train a computational method, including a machine learning algorithm, and the trained machine learning algorithm may predict the tertiary structure of one or more target RNA molecules, and then output the predicted tertiary structure of the target RNA molecule.

[0047] One aspect of the present disclosure provides a computer-implemented method for predicting the tertiary structure of an RNA molecule of interest, the method comprising: (a) creating a training dataset, the training dataset comprising (i) chemical mapping data for a first plurality of RNA molecules and (ii) tertiary structure data for a second plurality of RNA molecules; (b) training a machine learning algorithm using the training dataset; (c) predicting the tertiary structure of the RNA molecule of interest by applying the trained machine learning algorithm; and (d) outputting the predicted tertiary structure of the RNA molecule of interest. In some embodiments, the chemical mapping data is generated by a process comprising contacting the RNA molecule with a chemical probing agent. In some embodiments, the chemical probing agent comprises dimethyl sulfate (DMS). In some embodiments, the chemical probing agent comprises SHAPE (selective 2'-hydroxyl acylation and primer extension) reagent. In some embodiments, the SHAPE reagent is 1-methyl-7-nitroisatoic anhydride (1M7), 1-methyl-6-nitroisatoic anhydride (1M6), 5-nitroisatoic anhydride (5NIA), or N-methyl-nitroisatoic anhydride (NMIA). In some embodiments, the chemical probing agent comprises 2A3 ((2-aminopyridin-3-yl)(1H-imidazol-1-yl)methanone). In some embodiments, the RNA of interest comprises a portion of a transcriptome. In some embodiments, the transcriptome is a human transcriptome. In some embodiments, the training set comprises chemical mapping data for at least about 10, 100, 500, 1,000, 10,000, or more than 10,000 sequences. In some embodiments, the training set comprises chemical mapping data for at most about 10, 100, 500, 1,000, or 10,000 sequences. In some embodiments, the chemical mapping data is for sequences that occur at different abundances than in natural systems. In some embodiments, the chemical mapping data is collected from an in vitro source. In some embodiments, the machine learning algorithm comprises one or more artificial neural networks (ANNs).In some embodiments, the method further comprises predicting chemical mapping data for the RNA of interest by training an ANN. In some embodiments, the method further comprises predicting the chemical mapping data for the RNA of interest from the predicted tertiary structure of the RNA of interest. In some embodiments, the method further comprises outputting features of the RNA molecule of interest generated by the machine learning algorithm. In some embodiments, the tertiary structure comprises 3D coordinates of a plurality of atoms comprising the RNA molecule of interest. In some embodiments, the tertiary structure comprises 3D coordinates of each atom comprising the RNA molecule of interest. In some embodiments, the tertiary structure comprises one or more 3D coordinates of a plurality of nucleotides comprising the RNA molecule of interest. In some embodiments, the tertiary structure comprises one or more 3D coordinates of each nucleotide comprising the RNA molecule of interest. In some embodiments, the tertiary structure of the RNA molecule of interest is parameterized based on a distance map. In some embodiments, the tertiary structure of the RNA molecule of interest is parameterized based on a distance map and an angle. In some embodiments, the method does not require determining or predicting a secondary structure of the RNA molecule of interest. In some embodiments, the method predicts aspects of the tertiary structure of the target RNA that are not captured by the base pairing prediction of the target RNA. In some embodiments, the predicted tertiary structure includes one or more of a pseudoknot, a multiway junction, a coaxial stack, an a-minor motif, a kissing stem loop, a ribose zipper, or a tetraloop / tetraloop acceptor. In some embodiments, the chemical mapping data includes multidimensional chemical mapping data for one or more RNA molecules of the first plurality of RNA molecules. In some embodiments, the tertiary structure is a pharmaceutical target. In some embodiments, the training dataset further includes a multiple sequence alignment of a third plurality of RNA molecules. In some embodiments, the first plurality of RNA molecules and the second plurality of RNA molecules are the same. In some embodiments, the first plurality of RNA molecules and the third plurality of RNA molecules are the same. In some embodiments, the first plurality of RNA molecules and the second plurality of RNA molecules are different.In some embodiments, the first plurality of RNA molecules and the third plurality of molecules are different. In some embodiments, the first plurality of RNA molecules and the second plurality of RNA molecules comprise a mutually exclusive set of RNA molecules. In some embodiments, the first plurality of RNA molecules and the third plurality of molecules comprise a mutually exclusive set of RNA molecules. In some embodiments, the first plurality of RNA molecules and the second plurality of RNA molecules comprise at least one different RNA molecule. In some embodiments, the first plurality of RNA molecules and the third plurality of molecules comprise at least one different RNA molecule. In some embodiments, the first plurality of RNA molecules and the second plurality of RNA molecules comprise at least one identical RNA molecule. In some embodiments, the first plurality of RNA molecules and the third plurality of molecules comprise at least one identical RNA molecule. In some embodiments, the first plurality of RNA molecules are unrelated to the RNA molecule of interest. In some embodiments, the second plurality of RNA molecules are unrelated to the RNA molecule of interest. In some embodiments, the third plurality of RNA molecules are unrelated to the RNA molecule of interest. In some embodiments, one RNA molecule of the first plurality of RNA molecules has about 80%, 70%, 60%, 50%, 40%, 30%, 20%, or less sequence identity to the RNA molecule of interest.

[0048] Another aspect of the present disclosure provides a computer-implemented system for predicting the tertiary structure of an RNA molecule of interest, the system including a computing device including at least one processor and instructions executable by the at least one processor to perform operations, the operations including: (a) creating a training dataset, the training dataset including one or more of: (i) chemical mapping data for a first plurality of RNA molecules; (ii) structural data for a second plurality of RNA molecules; (b) training a machine learning algorithm using the training dataset; (c) predicting the tertiary structure of the RNA molecule of interest by applying the trained machine learning algorithm; and (d) outputting the predicted tertiary structure of the RNA molecule of interest.

[0049] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium executable by one or more processors and encoded with instructions for providing an application for predicting a tertiary structure of an RNA molecule of interest, the application including: (a) a training dataset module configured to create a training dataset including chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; (b) a training module configured to train a machine learning algorithm using the training dataset; (c) an inference module configured to predict the tertiary structure of the RNA molecule of interest by applying the trained machine learning algorithm; and (d) an output module configured to report the predicted tertiary structure of the RNA molecule of interest.

[0050] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]

[0051] The novel features of the invention are set forth with particularity in the appended claims. The features and advantages of the present invention will be better understood by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings.

[0052] [Figure 1] 1 illustrates a non-limiting example of a computing device, in this example the device has one or more processors, memory, storage, and network interfaces. [Figure 2] FIG. 1 shows a non-limiting example of a system for predicting RNA structure from sequence information. [Figure 3A] FIG. 1 illustrates a non-limiting example of a system including the machine learning algorithms described herein. [Figure 3B] FIG. 1 illustrates a non-limiting example of information flow during training of a machine learning algorithm included in the systems and methods described herein. [Figure 3C] FIG. 1 illustrates a non-limiting example of information flow during training of a machine learning algorithm included in the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0053] Unless otherwise specified, all technical terms, notations, and other technical and scientific terms or terminology used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. In some embodiments, terms having commonly understood meanings are defined herein for clarity and / or ease of reference, and the inclusion of such definitions herein should not necessarily be construed as representing a substantial difference from what is commonly understood in the art.

[0054] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the present disclosure. Thus, the description of a range should be considered to have all possible subranges specifically disclosed, as well as individual numerical values ​​within that range. For example, the description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0055] The terminology used herein is for the purpose of describing specific instances only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Furthermore, the terms "including," "includes," "having," "has," "with," or variations thereof, when used in either the detailed description and / or claims, are intended to be as inclusive as the term "comprising."

[0056] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" may mean ±10%, as is customary in the art. Alternatively, "about" may mean within ±20%, ±10%, ±5%, or ±1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term may mean within an order of magnitude, within five-fold, or within two-fold of a value. When a particular value is described in this application and claims, unless otherwise specified, it should be assumed that the term "about" means within an acceptable error range for the particular value. Also, when ranges and / or subranges of values ​​are provided, the ranges and / or subranges may include the endpoints of the ranges and / or subranges.

[0057] The term "substantially" as used herein generally refers to a value approaching 100% of a given value. For example, a peptide "substantially localized" in an organ may indicate that about 90% by weight of the peptide, salt, or metabolite is present in the organ, relative to the total amount of the peptide, salt, or metabolite. In some cases, this term may refer to an amount that may be at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 99.99% of the total amount. In some cases, this term may refer to an amount that may be about 100% of the total amount.

[0058] The term "nucleic acid" generally refers to a polymeric form of nucleotides of any length, i.e., either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-, double-, or multi-stranded form. Nucleic acids can be exogenous or endogenous to a cell. Nucleic acids can exist in a cell-free environment. Nucleic acids can be genes or fragments thereof. Nucleic acids can be DNA. Nucleic acids can be RNA. Nucleic acids can have any three-dimensional structure and perform any function. Nucleic acids can contain one or more analogs (e.g., modified backbones, sugars, or nucleobases).

[0059] "In vitro" is defined as a biological process, reaction, or experiment that is made to occur in isolation from a whole organism, e.g., in a test tube, artificial environment, or culture. "In vitro" is defined as a biological process, reaction, or experiment that occurs within an organism.

[0060] "Position" generally refers to a particular nucleotide by its index relative to position zero depending on the context, eg, the first nucleic acid at the 5' end of the molecule.

[0061] A "region" generally refers to a portion of a nucleic acid, said portion being less than or equal to the entire nucleic acid.

[0062] The "secondary structure" of a nucleic acid (e.g., an RNA molecule) generally refers to the pattern of base pairing predicted or observed in the molecule. Base pairs can include canonical (Watson-Crick), Hoogsteen, wobble, sugar edge, or non-canonical base pairs. Base pairs can be between any edge (Watson-Crick, Hoogsteen, sugar, or CH) of two nucleotides. The secondary structure can include one or more secondary structure motifs. A secondary structure motif generally refers to a well-defined pattern of paired or unpaired nucleotides that repeats across multiple RNA molecules. Non-limiting examples of secondary structure motifs include helices, hairpin or stem loops, internal loops or bulges, junction loops or multiway junctions, terminal mismatches, and single-nucleotide overhangs. The secondary structure is predicted by a computational algorithm. Non-limiting examples of packages that implement such algorithms include ViennaRNA, NUPACK, RNAstructure, RNAsoft, CONTRAfold, CycleFold, LeamToFold, MXfold, and SPOT-RNA. Secondary structure can be obtained from the atomic (3D) coordinates of RNA molecules. Secondary structure can be represented, for example, by a base pair matrix or a "dot bracket" string.

[0063] The "tertiary structure" and "3D structure" of a nucleic acid (e.g., an RNA molecule) are generally used interchangeably herein to refer to the predicted or observed atomic coordinates of a nucleic acid. Certain elements or regions of an RNA molecule may be referred to herein as "tertiary motifs." Tertiary motifs generally refer to well-defined 3D motifs that repeat across multiple RNA molecules. Non-limiting examples of tertiary motifs include pseudoknots, multiway junctions, coaxial stacks, a-minor motifs, kissing stem loops, ribose zippers, and tetraloop / tetraloop receptors. Tertiary structure can be determined from experimental methods such as, for example, X-ray crystallography, electron microscopy, and nuclear magnetic resonance spectroscopy. Tertiary structure can be predicted from the sequence of an RNA molecule and other data or representations by methods such as the systems disclosed herein.

[0064] The term " sequence identity " or " percent identity " in the context of two or more nucleic acid or polypeptide sequences generally refers to two (for example, in pairwise alignment) or more (for example, in multiple sequence alignment) sequences that are the same or have a certain percentage of the same amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured by a sequence comparison algorithm.Suitable sequence comparison algorithms for nucleic acid sequences include, for example, BLASTN, which uses the following parameters: word size (W) of 28, expectation threshold (E) of 0.05, and reward / penalty ratio of 1 / -2 (these are the default parameters of BLASTN in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov).

[0065] As used herein, "chemical mapping" generally refers to one or more techniques used to probe solvent accessibility or conformational flexibility, interactions between nucleotide pairs in a sequence, or any other mechanism that results in a probing signal that depends in part on the 3D structure of at least a portion of a molecule (e.g., a nucleic acid such as an RNA molecule or portion thereof). Chemical mapping can include contacting a molecule with a chemical probing agent and measuring a signal indicative of a reaction with the chemical probing agent at a position or region (e.g., a nucleotide or atom) of the molecule. Non-limiting examples of chemical probing agents include methylating agents, such as dimethyl sulfate (DMS), and acylating agents. Optionally, the acylating agent includes selective 2'-hydroxyl acylation and primer extension (SHAPE) reagents. Optionally, the SHAPE reagent comprises 1-methyl-7-nitroisatoic anhydride (1M7), 1-methyl-6-nitroisatoic anhydride (1M6), 5-nitroisatoic anhydride (5NIA), or N-methyl-nitroisatoic anhydride (NMIA). In some embodiments, the chemical probing agent comprises 2A3. Optionally, measuring the signal can include using a reverse transcriptase with an increased propensity for mismatches and / or terminations at the modified positions, and then sequencing the resulting DNA to infer the chemical mapping signal.

[0066] overview Chemical mapping experiments have been used to complement experimental and structural biological approaches for RNA secondary structure prediction. Information about the secondary structure can be used to improve the accuracy of tertiary structure prediction. In chemical mapping experiments, RNA molecules are exposed to chemicals that induce small changes in the RNA molecule. After further experimental processing, these changes can be measured, for example, by DNA sequencing, an approach that is relatively inexpensive and amenable to high-throughput experiments. However, the resulting data are noisy, and the detailed process leading from the 3D structure and its exposure to chemical agents to the experimental data is resistant to predictive understanding. Instead, these data are used to guide secondary structure predictions using a heuristic based on the observation that Watson-Crick base-paired bases are less sensitive to chemical mapping than unpaired bases.

[0067] In parallel, computational methods for predicting RNA tertiary structure from nucleotide sequences have been developed. These include template-based modeling approaches, approaches that sample candidate structures and then select among them, and recent deep learning-based approaches. Although these approaches may also use secondary structures inferred from chemical mapping data, they do not integrate chemical mapping data to improve the accuracy of tertiary structure prediction, for example, when the RNA for which chemical mapping data is available and the RNA for which tertiary structure prediction is desired are of substantially different sequences. Furthermore, no method / algorithm for tertiary structure prediction that teaches the incorporation of chemical mapping data and uses it to directly improve tertiary structure prediction has been described in the art. This may be due to the lack of available chemical mapping data at scale and the lack of utility / achievement in collecting and using such data in the past. For this reason, it is not conventional to use chemical mapping data to directly predict RNA tertiary structure (e.g., without relying on intermediate secondary structure prediction), and the use of machine learning techniques that enable this has not been possible until now.

[0068] This paper describes a method and system for predicting the tertiary structure of nucleic acid (for example, RNA).This method and system can include one or more machine learning algorithms that are trained on the data that characterizes one or more reference RNA molecules.Then, the trained machine learning algorithm can be applied to the target RNA molecule to predict the tertiary structure of the target RNA molecule.

[0069] The machine learning algorithm may be trained at least in part based on data including chemical mapping data. The chemical mapping data may include data indicating which portions (e.g., nucleotides or atoms) of a nucleic acid (e.g., RNA) molecule are protected from attack by a chemical modifying agent (e.g., a chemical probing agent). In some embodiments, the chemical modifying agent comprises dimethyl sulfate (DMS). In some embodiments, the chemical modifying agent comprises a selective 2'-hydroxyl acylation and primer extension (SHAPE) reagent. In some embodiments, the SHAPE reagent comprises an acylating agent such as 1-methyl-7-nitroisatoic anhydride (1M7), 1-methyl-6-nitroisatoic anhydride (1M6), 5-nitroisatoic anhydride (5NIA), or N-methyl-nitroisatoic anhydride (NMIA). In some embodiments, the chemical probing agent comprises 2A3.

[0070] The disclosed methods and systems may include one or more machine learning algorithms. The one or more machine learning algorithms may include one or more artificial neural networks (ANNs). ANNs with different architectures may be combined to process or predict data of one or more modalities indicative of the tertiary structure of nucleic acids. For example, a recurrent neural network (RNN), transformer, or other attention network architecture may be used to process sequence data, while a graph neural network may be used to predict or refine 3D structures. In particular, the disclosed ANNs may include layers equivalent to rigid body rotations and translations in 3D, making them particularly suitable for learning and predicting molecular structures.

[0071] Compared to other methods for predicting RNA structure, the disclosed method may offer certain benefits and advantages. The disclosed method may be configured to predict tertiary structure directly from chemical mapping data (e.g., without calculating or accepting as input the secondary structure of the target molecule). Determining secondary structure from chemical mapping data (i.e., requiring such a determination to predict tertiary structure) may also result in incorrect prediction of base-paired nucleotides if they are actually involved in a higher-order (e.g., tertiary) structural motif, or if they are not base-paired but their chemical mapping signal is influenced by other nearby atoms. These errors may arise because residues that provide high chemical mapping signals are generally predicted to be base-paired in other ways. However, any region of an RNA molecule with reduced solvent accessibility or conformational flexibility, for example, due to involvement in tertiary structure motif interactions, may exhibit a relatively low chemical mapping signal, even if the region is not base-paired. Furthermore, the disclosed method may be configured to predict the tertiary structure of RNA molecules for which chemical mapping data is unavailable, even if the region has low sequence identity compared to the molecules used to train the machine learning algorithm. In some embodiments, the method can include predicting a predicted tertiary structure of an RNA molecule of interest based on chemical mapping data for the RNA molecule of interest. In some embodiments, the method can include using the same embedding (e.g., as output from another machine learning algorithm that processes the RNA sequence) to predict the chemical mapping data of the RNA of interest and the predicted tertiary structure of the RNA of interest.

[0072] An exemplary system 200 is shown in Figure 2. System 200 is configured to process one or more inputs 201 of the same or different modalities. Input 201 may include one or more RNA sequences having a structure predicted by system 200. Input 201 may further include other information indicative of the structure of the RNA molecules. For example, input 201 may include one or more of chemical mapping data for one or more of the input RNA sequences, sequence alignment or conservation information for one or more of the input RNA sequences, secondary structure information for one or more of the input RNA sequences, or one or more candidate or reference (e.g., experimentally determined) tertiary structures.

[0073] In some embodiments, the training set may include sequences with data characterizing the secondary structure of the sequences. This data may be used in various forms, such as in the form of a base pair probability matrix, either as pairwise input features for the algorithm or as additional prediction targets (e.g., similar to the case of 2D chemical mapping data described elsewhere herein). In some embodiments, the secondary structure data may or may not be used to optimize structural modules (e.g., to improve tertiary structure prediction capabilities) during training, as described below. Various secondary structures may be characterized for one or more elements in the sequences in the training set, for example, annotations such as stems, hairpin loops, pseudoknots, bulges, internal loops, multi-loops, etc.

[0074] The input 201 is processed by a computational algorithm 205 to generate a predicted 3D structure 210 corresponding to the input. The computational algorithm 205 may include one or more artificial neural networks (ANNs) trained on multiple RNA sequences and tertiary structures. The computational algorithm 205 may include multiple machine learning algorithms (or modules) configured to process data of different modalities. Furthermore, the input of one machine learning algorithm may include the output of one or more other machine learning algorithms, and the output of one machine learning algorithm may constitute the input for one or more other machine learning algorithms. The computational algorithm may perform additional predictions on the RNA sequence and generate corresponding outputs. In the example illustrated in FIG. 3A, the computational algorithm 205 includes a transformer module 220, a structure module 230, and a chemistry module 240. The transformer module 220 is configured to receive an input RNA sequence. The transformer module 220 may receive an encoding (e.g., one-hot encoding) or embedding of the RNA sequence. The system may include one or more additional modules (e.g., deep neural networks) (not shown) configured to generate the encodings or embeddings.

[0075] The Transformer module 220 may include an attention network including an attention mechanism (e.g., one or more attention layers) configured to process the RNA sequence (or a representation thereof, such as an embedding). The Transformer 220 may revise the sequence embedding as well as a second pairwise embedding of residues that encodes information about the relationship or proximity between pairs of residues. The Transformer 220 may be configured to revise the pairwise embedding based at least in part on the sequence embedding. The Transformer 220 may output the revised sequence embedding and pairwise embedding. The output revised sequence embedding and pairwise embedding may be used by other machine learning algorithms that comprise the computational algorithm 205.

[0076] In one example, the sequence embedding and pairwise embedding outputs from the transformer 220 may subsequently be input to the structure module 230, as further illustrated in FIG. 3A. The structure module 230 may include an attention mechanism. The structure module 230 may include a machine learning algorithm (e.g., an ANN) employing geometry recognition and equivariant attention operations, as described elsewhere herein, to generate a representation of the 3D structure of the input RNA. The structure module 230 may predict the 3D coordinates of the RNA molecule from the sequence embedding and pairwise embedding. In one example, the structure module 230 generates a representation of the predicted RNA tertiary structure by predicting rotations and translations in the coordinates of at least some atoms (e.g., C4', C1', N1 / N9) from a local residue frame to a global molecular frame (e.g., main frame) for at least a subset of the residues of the RNA molecule. In some embodiments, the structure module 230 may predict the coordinates of all atoms in the target RNA. Alternatively, the structure module 230 may predict the coordinates of only a subset of atoms in the target RNA. The subset of atoms may include at least one atom from every nucleotide of the target RNA, or may include atoms from only certain regions (e.g., nucleotides) of the target RNA. The computational algorithm 205 may output, as at least part of its output, a predicted 3D structure 210 of the target sequence.

[0077] In some embodiments, the tertiary structure representation includes different (sub)sets of atoms. The mainframe representation may include various atoms, but may not be identical at every nucleotide in the RNA molecule. The mainframe representation may range from zero atoms to a complete enumeration of atoms at the level of individual nucleotides. The structure representation may not involve explicit embedding in terms of 3D coordinates at the level of a distance map, which is represented as part of pairwise features within a structure module. This distance map representation may be extended by angles, dihedra, or both.

[0078] In one example, the computational algorithm 205 may further include a chemistry module 240, as illustrated in FIG. 3A. The chemistry module may be configured to predict chemical mapping data 211 for an input RNA molecule based at least in part on the predicted 3D structure (or a representation or embedding thereof). In one example, the chemistry module 240 includes a message-passing graph neural network (MPGNN) that accepts the predicted RNA structure as input. In another example, the MPGNN for the chemistry module may be replaced with a transformer. The chemistry module 240 may operate on a graph representation of the predicted RNA structure in which at least a subset of atoms (e.g., C4', C1', and N1 / N9) are represented by graph nodes, and the graph edges are drawn between atoms in close Euclidean space (e.g., within a certain cutoff, such as 15 Å). The chemistry model 240 may revise the graph through a series of message-passing steps to generate the predicted chemical mapping data 211. Additionally or alternatively, the chemistry module 240 may include an equivariant neural network architecture, such as a point convolution architecture or an equivariant message passing architecture, as described elsewhere herein.

[0079] FIG. 3B illustrates the flow of information during training of the example system illustrated in FIG. 3A. Training the system 205 may include providing a training dataset including one or more reference RNA molecules with known chemical mapping data. The training data (or a subset thereof) may be fed through the (untrained) system in a forward direction (indicated by the solid arrows) to generate predicted outputs. The discrepancy between the predicted and known outputs may be quantified by a loss function. The choice of loss function may be based in part on the type of output data. In the example illustrated in FIG. 3B, an L2 function may be used to quantify the error between the predicted and observed chemical mapping data for the reference RNA molecules. Based on the quantified error, gradients for one or more parameters (e.g., weights, biases, thresholds) of a computational algorithm (e.g., a machine learning algorithm) may be calculated, for example, by backpropagation to revise the parameters of the neural network (indicated by the dashed arrows) so that the output (predicted) values ​​match the known values.

[0080] In another example shown in FIG. 3C, the system is trained by quantifying errors in the tertiary structure of RNA molecules predicted in a training dataset (or a subset thereof). In this example, the loss function may include, for example, the root mean square deviation between the predicted reference structure and the superimposed reference structure, a global distance test (GDT) score, an inter-residue contact area distance (CAD) score, or a local difference distance test (LDDT) score, or another score sufficiently suitable for assessing the structural similarity of biomolecular structures. The loss function may include a regression loss function. The loss function may include a logistic loss function. The loss function may include a variational loss. The loss function may include a prior. The loss function may include a Gaussian prior. The loss function may include a non-Gaussian prior. The loss function may include an adversarial loss. The loss function may include a reconstruction loss. The loss function may quantify the difference between the predicted reference structure and the superimposed reference structure in the pairwise distance between nucleotides i and j for each pair of nucleotides (e.g., via one or more atoms of the nucleotide, the center of mass of the nucleotide, etc.). The loss function may quantify the difference between the predicted reference structure and the superimposed reference structure in the angle formed by a sequence of three nucleotides i, j. The loss function may quantify the difference between the predicted reference structure and the superimposed reference structure in the dihedron (torsion angle) formed by a sequence of four nucleotides i, j. The loss function may quantify the difference between the predicted reference structure and the superimposed reference structure in the incorrect dihedron formed by a sequence of four nucleotides i, j, j, where three nucleotides are bound to the central nucleotide.

[0081] Apart from the loss function, training may also include, for example, selecting a set of initial parameter values ​​(or a method for generating them) and an optimization scheme. In some embodiments, the initialization includes Xavier initialization. In some embodiments, the optimization includes Adam optimization.

[0082] The methods and systems described herein can be used with any nucleic acid sequence. In some embodiments, the target nucleic acid sequence is an RNA sequence. In some embodiments, the target RNA is a coding RNA. In some embodiments, the target RNA is a non-coding RNA (ncRNA). In some embodiments, the target RNA is a circular RNA. In some embodiments, the target RNA is a transfer RNA (tRNA), ribosomal RNA (rRNA), messenger RNA (mRNA), small nuclear RNA (snRNA), small nucleolar RNA (snRNA), microRNA (miRNA), short hairpin RNA (shRNA), small interfering RNA (siRNA), Y RNA, vault RNA, antisense RNA, transcription initiator RNA (tiRNA), transcription start site-associated RNA (TSSa-RNA), piwi-interacting RNA (piRNA), guide RNA (gRNA), or ribozyme. The target RNA can include single-stranded RNA (ssRNA) or double-stranded RNA (dsRNA). In some embodiments, the target RNA comprises a transcriptome or a portion thereof. In some embodiments, the transcriptome comprises a transcriptome from a eukaryote, a prokaryote, or an archaea. In some embodiments, the transcriptome comprises a transcriptome from a fungus, a bacterium, a virus, a protist, an algae, a plant, or an animal. In some embodiments, the transcriptome comprises a human transcriptome. In some embodiments, the target RNA comprises an RNA that is synthetic or not normally found in nature. In some embodiments, the target RNA comprises a naturally occurring RNA sequence. In some embodiments, fine-tuning the machine learning model can generate a more accurate tertiary structure of the RNA molecule of interest. After the machine learning model is trained on an initial dataset, chemical mapping data of the RNA molecule of interest can be used to fine-tune the machine learning algorithm. By revising the machine learning parameters based on the chemical mapping data of the RNA molecule, the machine learning model can output a more accurate tertiary structure of the RNA molecule of interest.In some embodiments, the machine learning model can be fine-tuned using chemical mapping data from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or any number of RNA molecules of interest. Because the RNA molecules of interest may be related sequences with relatively high sequence similarity, including more RNA molecules of interest can further improve the accuracy of the machine learning algorithm.

[0083] In some aspects, the present disclosure provides a method for determining target region or subsequence in target RNA molecule.Target region or subsequence can be the part of target RNA molecule, which can disrupt the structure of target RNA molecule when pharmaceutical agent is bound, and enhance or destroy its biological function.Target can be identified based on predicted tertiary structure.In some embodiments, pharmaceutical agent that binds to target RNA molecule can be formulated based on predicted tertiary structure.

[0084] Chemical Mapping Data The methods and systems described herein may include, import, operate on, or output one or more datasets containing chemical mapping data. A chemical mapping dataset may include data indicating the reactivity of regions (e.g., nucleotides or atoms) of a molecule (e.g., RNA) to chemical probing agents. A chemical mapping dataset may include indicators of conformational flexibility or solvent accessibility of one or more nucleotides or subsequences of a sequence. A chemical mapping dataset may include indicators of interactions between nucleotide pairs in a sequence. A chemical mapping dataset may include molecular binding data (biomolecules, as well as more generally organic and inorganic molecules). In some embodiments, the chemical probing agent includes a methylating agent such as dimethyl sulfate. In some embodiments, the chemical probing agent includes an acylating agent such as selective 2'-hydroxyl acylation and primer extension (SHAPE) reagents. In some embodiments, the chemical mapping data includes experimentally observed reactivity. In some embodiments, the chemical mapping data includes predicted reactivity.

[0085] In some embodiments, the chemical mapping data includes dimethyl sulfate (DMS) reactivity data. DMS can methylate unpaired cytosine and adenine residues in RNA molecules, but has reduced reactivity (or otherwise reduced solvent accessibility or conformational flexibility) toward base-paired cytosine and adenine residues. A DMS reactivity profile of an RNA molecule (e.g., characterizing the reactivity of one or more residues or other regions toward DMS) can be obtained by treating the RNA molecule with DMS, reverse transcribing the methylated RNA with reverse transcriptase, and sequencing the resulting DNA. Reverse transcriptase frequently incorporates incorrect DNA nucleotides, inserts additional nucleotides, deletes nucleotides, and / or aborts termination when it encounters a methylated RNA residue. Thus, the relative observed frequency of these mutations at a given position correlates with the conformational flexibility or solvent accessibility of a given nucleotide overall. DMS reactivity profiles can also be determined by direct RNA sequencing, for example, using a nanopore device.

[0086] In some embodiments, the chemical mapping data includes selective 2'-hydroxyl acylation and primer extension (SHAPE) reactivity data. SHAPE reagents are generally acylating agents that can acylate the 2'-hydroxyl of unpaired nucleotides but have reduced reactivity toward base-paired nucleotides (or otherwise reduced solvent accessibility or conformational flexibility). A SHAPE reactivity profile of an RNA molecule (e.g., characterizing the reactivity of one or more residues or other regions toward a SHAPE reagent) can be obtained by treating the RNA molecule with a SHAPE reagent and reverse transcribing the modified RNA to generate cDNA. Reverse transcriptase frequently incorporates incorrect DNA nucleotides, inserts additional nucleotides, deletes nucleotides, and / or aborts termination when it encounters an acylated RNA residue. Thus, quantification of the length or mutations of cDNAs in a cDNA pool provides a readout of which regions (e.g., nucleotides) of an RNA molecule exhibit the greatest conformational flexibility or solvent accessibility. SHAPE reactivity profiles can also be determined by direct RNA sequencing, for example, using a nanopore device.

[0087] In some embodiments, the chemical mapping data includes reactivity data of an RNA molecule and one or more RNA molecules derived from the RNA molecule to one or more chemical probing agents. The derived RNA molecule may contain point mutations in the RNA molecule. Such mutations may be generated randomly, such as by using an error-prone polymerase, or the derivative RNA may be rationally designed with specific point mutations. One example of a method for generating such data is mutation and map readout by next-generation sequencing (M2-seq). M2-seq and other multidimensional chemical mapping experiments can show which nucleotides respond to perturbations (e.g., chemical modification by chemical probing agents) on a nucleotide-by-nucleotide basis, allowing for inference of which paired nucleotides interact in the RNA structure. The machine learning algorithms (e.g., ANNs) disclosed herein may be configured to incorporate, operate on, output, or predict multidimensional chemical mapping data.

[0088] The chemical mapping data may include chemical reactivity data characterizing the reactivity of a region of an RNA molecule to one chemical probing agent. Alternatively, the chemical mapping data may include chemical reactivity data characterizing the reactivity of a region of an RNA molecule to more than one chemical probing agent. For example, the chemical mapping data includes DMS reactivity data for multiple RNA molecules. In another example, the chemical mapping data includes SHAPE reactivity data for multiple RNA molecules. In yet another example, the chemical mapping data includes DMS and SHAPE reactivity data for multiple RNA molecules.

[0089] Data for training the machine learning algorithms described herein can include chemical mapping data for one or more reference molecules (e.g., RNA molecules). In some embodiments, the training data includes chemical mapping data for at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 5,000, 10,000, 100,000, 1,000,000, 10,000,000 or more reference molecules. In some embodiments, the training data includes chemical mapping data for at most about 1,000,000, 10,000,000, 100,000, 10,000, 5,000, 2,000, 1,000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or fewer reference molecules. In some embodiments, the chemical mapping data includes chemical mapping data for target RNA molecules. In some embodiments, the chemical mapping data does not include chemical mapping data for the target molecule. In some embodiments, the chemical mapping data includes multidimensional chemical mapping data.

[0090] In some embodiments, one or more reference RNA molecules are not related to the target RNA molecule.The reference RNA molecule can comprise at most about 80%, 70%, 60%, 50%, 40%, 30%, 20% or less identity with the target RNA molecule.Alternatively, one or more reference RNA molecules can be related to the target molecule.The reference RNA molecule can comprise at least about 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity with the target RNA molecule.

[0091] In some embodiments, the chemical mapping data is collected from in vitro experiments. In some embodiments, the chemical mapping data is collected from in vivo experiments. In some embodiments, the one or more reference RNA molecules are synthetic RNA. In some embodiments, the one or more reference RNA molecules are natural RNA. In some embodiments, the one or more reference RNA molecules comprise a mixture of synthetic RNA and natural RNA.

[0092] Tertiary structure prediction The disclosed methods and systems may include one or more machine learning algorithms. The one or more machine learning algorithms may include one or more artificial neural networks (ANNs). ANNs with different architectures may be combined to process or predict data of one or more modalities indicative of nucleic acid tertiary structure. For example, a recurrent neural network (RNN), transformer, or other attention network architecture may be used to process sequence data, while a graph neural network may be used to predict or refine 3D structures. In particular, the disclosed ANNs may include layers equivalent to rigid-body rotations and translations in 3D, making them particularly suitable for learning and predicting molecular structures. The disclosed methods may be configured to directly predict tertiary structure from chemical mapping data (e.g., without calculating or accepting as input the secondary structure of the target molecule).

[0093] The structural model of the machine learning algorithm can output a predicted tertiary structure. The tertiary structure can be iteratively refined based on multiple passes through the structural module. For example, the output from the structural module can be recursively provided as input to the structural model to output a revised tertiary structure. The machine learning algorithm can be configured to receive the tertiary structure and predict a set of one or more revisions that can be added to the tertiary structure. The revisions can be to 3D coordinates, angles, dihedrals, distance histograms, or any other measure of the 3D structure.

[0094] The iterative process can terminate after meeting a convergence criterion. The convergence criterion can be, for example, when the revised tertiary structure is substantially identical to the previous tertiary structure. Such a criterion can be based on the magnitude of the difference between the revised tertiary structure and the previous tertiary structure. For example, if pairwise distances are used for the convergence criterion, the convergence criterion can be formulated as when the revision is less than 0.1 angstroms per pair on average (various other thresholds can be used). Similarly, the convergence criterion can be formulated in terms of angles, dihedrals, 3D coordinates, or any other measure of 3D structure. In some embodiments, the machine learning algorithm can be configured to output a measure of its accuracy or confidence. For example, the machine learning algorithm can be designed to output an index of accuracy or confidence based on examples seen in the training dataset. In some embodiments, the machine learning algorithm can output the local accuracy of the structure prediction (e.g., which can be measured based on the metric of the predicted local distance difference test (pLDDT)). The local accuracy can be used as a convergence criterion, in which case the algorithm will revise the tertiary structure until a local accuracy threshold is met or until the threshold no longer substantially improves. The ability of the algorithm to predict the quality of tertiary structure predictions on its own can be trained based on ground truth tertiary structure data provided during training and an LDDT metric that can be calculated in comparing the tertiary structure data with the tertiary structure predictions.

[0095] Machine Learning Algorithms The methods and systems described herein may include one or more machine learning algorithms. The machine learning algorithm may include an unsupervised machine learning algorithm. The trained algorithm may include a supervised machine learning algorithm. The machine learning algorithm may include a self-supervised machine learning algorithm.

[0096] In some embodiments, the machine learning algorithms of the methods and systems described herein utilize one or more artificial neural networks (ANNs). An ANN may be a machine learning algorithm that can be trained to map an input data set to an output data set, where the ANN includes an interconnected group of nodes organized into multiple node layers. For example, an ANN architecture may include at least one input layer, one or more hidden layers, and an output layer. An ANN may include any total number of layers and any number of hidden layers, where the hidden layers function as trainable feature extractors that enable mapping a set of input data to an output value or set of output values. As used herein, a deep learning algorithm (such as a deep neural network (DNN)) is an ANN that includes multiple hidden layers, e.g., two or more hidden layers. Each layer of a neural network may include several nodes (or "neurons"). A node receives inputs that result directly from either the input data or the node outputs in the previous layer and performs a specific operation (e.g., a summation operation). Connections from inputs to nodes are associated with weights (or weight coefficients). A node may sum the products of all pairs of inputs and their associated weights. The weighted sum may be offset with a bias. The output of a node or neuron may be gated using a threshold or activation function. The activation function may be a linear or nonlinear function. The activation function may be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions, such as a saturated hyperbolic tangent function, an identity function, a binary step function, a logistic function, an arctangent function, a soft sine function, a parametric rectified linear unit function, an exponential linear unit function, a soft plus function, a bent identity function, a soft exponential function, a sinusoidal function, a sinc function, a Gaussian function, or a sigmoidal function, or any combination thereof.

[0097] Non-limiting examples of structural components of the machine learning algorithms described herein include convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTMs), attention networks, and transformers, graph neural networks (GNNs), message passing neural networks (MPNNs), and combinations or variations thereof.

[0098] In some embodiments, a neural network includes a series of layers, each containing individual units called "neurons." In some embodiments, a neural network includes an input layer, where data is presented, one or more internal or "hidden" layers, and an output layer. Neurons may be connected to neurons in other layers via connections with weights, which are parameters that control the strength of the connections. The number of neurons in each layer may be related to the complexity of the problem being solved. The minimum number of neurons required in a layer may be determined, for example, by the complexity of the problem, and the maximum number may be limited, for example, by the neural network's ability to generalize. Input neurons may receive the data being presented and then send the data to a first hidden layer via connection weights that are modified during training. The first hidden layer may process the data and send the results to the next layer via a second set of weighted connections. Each subsequent layer may "pool" the results from the previous layer into more complex relationships. Furthermore, while traditional software programs require specific instructions to be written in order to perform a function, neural networks are programmed by training them with a known set of examples and allowing them to modify themselves during (and after) training to provide a desired output, such as an output value. After training, when presented with new input data, the neural network is configured to generalize what it "learned" during training and apply what it learned from training to new input data it has not seen before, thereby generating an output related to that input.

[0099] In some embodiments, the machine learning algorithm includes a CNN. The CNN may be a deep or feed-forward ANN. The CNN may be applicable to analyzing sequence data. The CNN may include an input layer, an output layer, and multiple hidden layers. The hidden layers of the CNN may include a convolutional layer, a pooling layer, a fully connected layer, and a normalization layer.

[0100] A convolutional layer may apply a convolution operation to the input and pass the result of the convolution operation to the next layer. The convolution operation may reduce the number of free parameters, allowing the network to be deeper with fewer parameters. In a neural network, each neuron may receive input from several locations in the previous layer. In a convolutional layer, neurons may receive input only from a restricted subregion of the previous layer. The parameters of a convolutional layer may include a set of learnable filters (or kernels) containing one or more learnable weights. The learnable filters may have small receptive fields and extend across the entire depth of the input volume. During a forward pass, each filter may be convolved across the width and height of the input volume, computing a dot product between the filter's ingress and the input to generate a two-dimensional activation map for that filter. As a result, the network may learn filters that activate when detecting certain types of features at certain spatial locations in the input.

[0101] In some embodiments, the machine learning algorithm includes an RNN. An RNN is a neural network with cyclical connections that can encode and process sequence data, such as sequences of RNA molecules. The RNN can include an input layer configured to receive a sequence of inputs. The RNN can further include one or more hidden recurrent layers that maintain state. At each step, each hidden recurrent layer can calculate its output and next state. The next state can depend on the previous state and the current input. The state can be maintained across steps to capture dependencies in the input sequence.

[0102] The RNN can be a long short-term memory (LSTM) network. The LSTM network may be composed of LSTM units. The LSTM units may include cells, input gates, output gates, and forget gates. The cells may be responsible for maintaining dependency tracking between elements in an input sequence. The input gates can control the extent to which new values ​​flow into the cells, the forget gates can control the extent to which values ​​remain in the cells, and the output gates can control the extent to which values ​​in the cells are used to calculate the output activation of the LSTM unit.

[0103] Alternatively, the machine learning algorithm may include a transformer. The transformer may be a model without recurrent connections. Instead, it may rely on an attention mechanism. The attention mechanism may focus or "pay attention" to certain input regions while ignoring others. This may increase the performance of the model, as certain input regions may be less relevant. At each step, the attention unit may, among other operations, compute a dot product of the context vector and the input at that step. The output of the attention unit may define where the most relevant information is located within the input sequence.

[0104] In some embodiments, the machine learning algorithm comprises a graph neural network (GNN). A GNN is a neural network specifically designed to perform inference on graph-based data. The GNN used in the method and system herein may be a graph convolutional network (GCN), a graph attention network (GAT), a message-passing neural network, or other GNN that performs permutation-invariant pooling and aggregation.

[0105] In some embodiments, the ANNs described herein may include one or more equivariant neural networks, such as point convolutional architectures or equivariant graph neural networks. These and related architectures may include neural network layers that are equivariant to translations and rotations and are therefore well suited to learning from or predicting molecular structure (e.g., RNA tertiary structure) data.

[0106] The weight coefficients, bias values, and thresholds, or other computational parameters of a neural network, may be "taught" or "learned" in a training phase using one or more sets of training data. Training a neural network may involve providing inputs to an untrained neural network to generate predicted outputs, comparing the predicted outputs to expected outputs, and revising the neural network's parameters to account for the difference between the predicted and expected outputs. A loss function may be used to quantify the difference between the predicted and expected outputs. Based on the calculated difference, a gradient for each parameter may be calculated by backpropagation to revise the neural network's parameters so that the output values ​​calculated by the ANN are consistent with the examples contained in the training data set. This process may be repeated for a certain number of iterations or until some stopping criterion is met.

[0107] The choice of loss function for a particular neural network may be based in part on the type of data the neural network is configured to process. For example, a neural network (such as MPGNN) configured to predict chemical mapping data for input tertiary RNA structures may be trained to optimize an L2 (squared error) loss. In another example, a neural network may be trained to optimize an L1 (absolute error) or cross-entropy loss. In yet another example, a neural network may be trained to optimize a function or score that quantifies the structural similarity between two or more molecules. The loss function may include the root mean square deviation between the predicted and superimposed reference structures, a global distance test (GDT) score, an inter-residue contact area distance (CAD) score, or a local difference distance test (LDDT) score.

[0108] The systems and methods described herein may use more than one machine learning algorithm to determine output (e.g., chemical mapping data or tertiary structure of an RNA molecule). The systems and methods may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more machine learning algorithms. One machine learning algorithm among the one or more machine learning algorithms may be trained on a specific type of data (e.g., chemical mapping data, alignment or storage data, tertiary structure data). Alternatively, a machine learning algorithm may be trained on more than one type of data. The input of one machine learning algorithm may include the output of one or more other machine learning algorithms. Furthermore, a machine learning algorithm may receive the output of one or more machine learning algorithms as its input.

[0109] Computing Systems Referring to FIG. 1 , a block diagram is shown illustrating an exemplary machine comprising a computer system 100 (e.g., a processing or computing system) capable of executing a set of instructions to cause a device to perform or implement any one or more of the aspects and / or methods for static code scheduling of the present disclosure. The computing system can be configured to send and receive a request to predict a tertiary structure of an RNA molecule of interest and / or to send and receive a predicted tertiary structure of the RNA molecule of interest. In some embodiments, the computing system can perform a method including sending a query to predict the tertiary structure of the RNA molecule of interest to a computer including a machine learning model. The machine learning algorithm can be configured to generate the tertiary structure. The machine learning model can be trained by a process including: (i) creating a training dataset including chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; (ii) training the machine learning model using the training dataset; or both. The computer-implemented method includes receiving the predicted tertiary structure of the RNA molecule of interest from the computer. In some embodiments, the computing system can perform a method including receiving a query to predict the tertiary structure of the RNA molecule of interest using the machine learning model. The machine learning algorithm can be configured to generate the tertiary structure. The machine learning model can be trained by a process including: (i) creating a training dataset including chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; (ii) training the machine learning model using the training dataset; or both. The computer-implemented method includes generating a predicted tertiary structure of the RNA molecule of interest from a computer. The computer-implemented method includes transmitting the predicted tertiary structure of the RNA molecule of interest from a computer.

[0110] The components of FIG. 1 are merely examples and do not limit the scope of use or functionality of any hardware, software, embedded logic components, or combinations of two or more such components, implementing a particular embodiment.

[0111] Computer system 100 may include one or more processors 101, memory 103, and storage 108, which communicate with each other and with other components via a bus 140. Bus 140 may also link to a display 132, one or more input devices 133 (which may include, for example, a keypad, keyboard, mouse, stylus, etc.), one or more output devices 134, one or more storage devices 135, and various tangible storage media 136. All of these elements may interface with bus 140 directly or through one or more interfaces or adapters. For example, various tangible storage media 136 may interface with bus 140 through storage media interface 126. Computer system 100 may take any suitable physical form, including, but not limited to, one or more integrated circuits (ICs), a printed circuit board (PCB), a mobile handheld device (such as a cell phone or PDA), a laptop or notebook computer, a distributed computer system, a computing grid, or a server.

[0112] Computer system 100 includes one or more processors 101 (e.g., a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), or a quantum processing unit (QPU)) that perform functions. Processor 101 optionally includes a cache memory unit 102 for temporary local storage of instructions, data, or computer addresses. Processor 101 is configured to assist in the execution of computer-readable instructions. Computer system 100 may provide the functionality of the components shown in FIG. 1 as a result of processor 101 executing non-transitory processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 103, storage 108, storage device 135, and / or storage medium 136. The computer-readable media may store software implementing particular embodiments, and processor 101 may execute the software. Memory 103 may read software from one or more other computer-readable media (e.g., mass storage devices 135, 136) or from one or more other sources via a suitable interface, such as network interface 120. The software may also cause the processor 101 to perform one or more processes, or one or more steps of one or more processes, described or illustrated herein. Performing such a process or step may include defining data structures stored in memory 103 and modifying the data structures as directed by the software.

[0113] Memory 103 may include various components (e.g., machine-readable media), including, but not limited to, random access memory components (e.g., RAM 104) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM®), phase change random access memory (PRAM), etc.), read-only memory components (e.g., ROM 105), and any combination thereof. ROM 105 may function to communicate data and instructions unidirectionally to processor 101, while RAM 104 may function to communicate data and instructions bidirectionally with processor 101. ROM 105 and RAM 104 may include any suitable tangible computer-readable media, as described below. In one example, a basic input / output system 106 (BIOS), containing the basic routines that help to transfer information between elements within computer system 100, such as during start-up, may be stored in memory 103.

[0114] Persistent storage 108 is bidirectionally connected to processor 101, optionally via storage control unit 107. Persistent storage 108 provides additional data storage capacity and may include any suitable tangible computer-readable media described herein. Storage 108 may be used to store operating system 109, executables 110, data 111, applications 112 (application programs), etc. Storage 108 may also include an optical disk drive, a solid-state memory device (e.g., a flash-based system), or any combination of the above. Information in storage 108 may, where appropriate, be incorporated into memory 103 as virtual memory.

[0115] In one example, storage device 135 may removably interface with computer system 100 (e.g., via an external port connector (not shown)) via storage device interface 125. In particular, storage device 135 and associated machine-readable media may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for computer system 100. In one example, software may reside, completely or partially, within the machine-readable media on storage device 135. In another example, software may reside, completely or partially, within processor 101.

[0116] Bus 140 connects a wide variety of subsystems. As used herein, reference to a bus may, where appropriate, encompass one or more digital signal lines that perform a common function. Bus 140 may be any of several types of bus structures, including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof, using any of a variety of bus architectures. By way of non-limiting example, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, a HyperTransport (HTX) bus, a serial advanced technology attachment (SATA) bus, and any combination thereof.

[0117] Computer system 100 may also include input devices 133. In one example, a user of computer system 100 may input commands and / or other information into computer system 100 via input devices 133. Examples of input devices 133 include, but are not limited to, an alphanumeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touchscreen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combination thereof. In some embodiments, input devices are Kinect, Leap Motion, etc. Input devices 133 may interface to bus 140 via any of a variety of input interfaces 123 (e.g., input interface 123), including, but not limited to, serial, parallel, gameport, USB, FIREWIRE®, THUNDERBOLT®, or any combination of the above.

[0118] In particular embodiments, when computer system 100 is connected to network 130, computer system 100 may communicate with other devices connected to network 130, particularly mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, etc. Communications between computer system 100 may be sent through network interface 120. For example, network interface 120 may receive incoming communications (e.g., requests or responses from other devices) in the form of one or more packets (e.g., Internet Protocol (IP) packets) from network 130, and computer system 100 may store the incoming communications in memory 103 for processing. Computer system 100 may similarly store outgoing communications (e.g., requests or responses to other devices) in the form of one or more packets in memory 103, which may be communicated from network interface 120 to network 130. Processor 101 may access these communication packets stored in memory 103 for processing.

[0119] Examples of network interface 120 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of network 130 or network segment 130 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, a corporate network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus, or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combination thereof. A network such as network 130 may employ wired and / or wireless communication modes. In general, any network topology may be used.

[0120] Information and data can be displayed via display 132. Examples of display 132 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED), such as a passive matrix OLED (PMOLED) or active matrix OLED (AMOLED) display, a plasma display, and any combination thereof. Display 132 can interface with processor 101, memory 103, and fixed storage 108, as well as other devices, such as input device 133, via bus 140. Display 132 is linked to bus 140 via video interface 122, and data transfer between display 132 and bus 140 can be controlled via graphics control unit 121. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD), such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting example, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headsets, etc. In still further embodiments, the display is a combination of devices such as those disclosed herein.

[0121] In addition to the display 132, computer system 100 may include one or more other peripheral output devices 134, including, but not limited to, audio speakers, printers, storage devices, and any combination thereof. Such peripheral output devices may be connected to bus 140 via output interface 124. Examples of output interface 124 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE® port, a THUNDERBOLT® port, and any combination thereof.

[0122] Additionally, or alternatively, computer system 100 may provide functionality as a result of logic hardwired or otherwise implemented in circuitry, which may operate in place of or in conjunction with software to perform one or more processes, or one or more steps of one or more processes, described or illustrated herein. References to software in this disclosure may encompass logic, and references to logic may encompass software. Furthermore, references to computer-readable media may encompass, where appropriate, circuitry (such as an IC) that stores software for execution, circuitry that implements logic for execution, or both. This disclosure encompasses any suitable combination of hardware, software, or both.

[0123] Those skilled in the art will appreciate that the various illustrative logic blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of such hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.

[0124] The various illustrative logic blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a DSP combined with a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0125] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processors, or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a user terminal.

[0126] In accordance with the description herein, suitable computing devices include, by way of non-limiting example, distributed and cloud computing platforms, server clusters, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, and netbook computers.

[0127] In some embodiments, a computing device includes an operating system configured to execute executable instructions. An operating system is software, including, for example, programs and data, that manages the device's hardware and provides services for application execution. Those skilled in the art will recognize that suitable server operating systems include, by way of non-limiting example, FreeBSD, OpenBSD, NetBSD®, Linux®, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those skilled in the art will recognize that suitable personal computer operating systems include, by way of non-limiting example, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those skilled in the art will also recognize that suitable mobile smartphone operating systems include, by way of non-limiting example, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.

[0128] Non-transitory computer-readable storage medium In some embodiments, the platforms, systems, media, and methods disclosed herein comprise one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of a networked computing device. In further embodiments, the computer-readable storage medium is a tangible component of the computing device. In still further embodiments, the computer-readable storage medium is optionally removable from the computing device. In some embodiments, computer-readable storage media include, by way of non-limiting example, CD-ROMs, DVDs, flash memory devices, solid-state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems such as cloud computing systems and services, and the like. In some cases, the programs and instructions are encoded permanently, nearly permanently, semi-permanently, or non-transitory on the medium.

[0129] computer program In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or the use thereof. A computer program includes a sequence of instructions that is executable by one or more processors of a computing device's CPU and that is written to perform specified tasks. The computer-readable instructions may be implemented as program modules, such as functions, objects, application programming interfaces (APIs), computing data structures, etc., that perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those skilled in the art will recognize that computer programs may be written in various languages ​​and in various versions.

[0130] The functionality of the computer-readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program includes one instruction sequence. In some embodiments, a computer program includes multiple instruction sequences. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from multiple locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more stand-alone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.

[0131] Web Applications In some embodiments, the computer program comprises a web application. In light of the disclosure provided herein, those skilled in the art will recognize that web applications, in various embodiments, utilize one or more software frameworks and one or more database systems. In some embodiments, the web application is built on a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, the web application utilizes one or more database systems, including, by way of non-limiting example, relational, non-relational, object-oriented, associative, XML, and document-oriented database systems. In further embodiments, suitable relational database systems include, by way of non-limiting example, Microsoft® SQL Server, mySQL™, and Oracle®. Those skilled in the art will also recognize that web applications may be written in one or more versions of one or more languages. Web applications may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written in part in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or eXtensible Markup Language (XML). In some embodiments, a web application is written in part in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written in part in client-side scripting such as Asynchronous JavaScript and XML (AJAX), Flash® ActionScript, JavaScript, or Silverlight®.In some embodiments, the web application is written in part in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tcl, Smalltalk, WebDNA®, or Groovy. In some embodiments, the web application is written in part in a database query language such as Structured Query Language (SQL). In some embodiments, the web application integrates an enterprise server product such as IBM® Lotus Domino®. In some embodiments, the web application includes a media player element. In various further embodiments, the media player element utilizes one or more of many suitable multimedia technology products, including, by way of non-limiting examples, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft® Silverlight®, Java®, and Unity®.

[0132] Mobile Applications In some embodiments, the computer program comprises a mobile application provided to the mobile computing device. In some embodiments, the mobile application is provided to the mobile computing device at the time of manufacture. In other embodiments, the mobile application is provided to the mobile computing device via a computer network as described herein.

[0133] Given the disclosure provided herein, mobile applications are created using hardware, languages, and development environments known in the art and techniques known to those skilled in the art. Those skilled in the art will recognize that mobile applications are written in a variety of languages. Suitable programming languages ​​include, by way of non-limiting example, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.

[0134] Suitable mobile application development environments are available from several sources. Commercially available development environments include, but are not limited to, Airplay SDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available free of charge, but are not limited to, Lazarus, MobiFlex, MoSync, and PhoneGap. Mobile device manufacturers also distribute software development kits, including, but not limited to, the iPhone® and iPad® (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.

[0135] Standalone Applications In some embodiments, the computer program comprises a stand-alone application, which is a program that runs as an independent computer process rather than as an add-on, e.g., a plug-in, to an existing process. Those skilled in the art will recognize that stand-alone applications are often compiled. A compiler is a computer program that converts source code written in a programming language into binary object code, such as assembly language or machine code. Suitable compiled programming languages ​​include, but are not limited to, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB.NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, the computer program comprises one or more executable compiled applications.

[0136] Web browser plugin In some embodiments, the computer program includes a web browser plug-in (e.g., an extension). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Software application manufacturers support plug-ins to allow third-party developers to extend the application, facilitate the easy addition of new features, and reduce the application's size. When supported, plug-ins allow customization of the software application's functionality. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display specific file types. Those skilled in the art will be familiar with several web browser plug-ins, including Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In some embodiments, the toolbar includes one or more web browser extensions, add-ins, or add-ons. In some embodiments, the toolbar includes one or more explorer bars, tool bands, or desk bands.

[0137] In view of the disclosure provided herein, one of ordinary skill in the art will recognize that several plug-in frameworks are available that allow for the development of plug-ins in a variety of programming languages, including, by way of non-limiting example, C++, Delphi, Java™, PHP, Python™, and VB.NET, or combinations thereof.

[0138] A web browser (also called an Internet browser) is a software application designed for use with networked computing devices to search, present, and traverse information resources on the World Wide Web. Suitable web browsers include, by way of non-limiting example, Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, and KDE Konqueror. In some embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, minibrowsers, and wireless browsers) are designed for use on mobile computing devices, including, by way of non-limiting example, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non-limiting example, Google® Android® Browser, RIM BlackBerry® Browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® Browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony® PSP™ Browser.

[0139] Software Module In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or the use thereof. In light of the disclosure provided herein, software modules are created by techniques known to those skilled in the art using machines, software, and languages ​​known in the art. The software modules disclosed herein are implemented in numerous ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or a combination thereof. In further various embodiments, a software module comprises multiple files, multiple sections of code, multiple programming objects, multiple programming structures, multiple distributed computing resources, multiple cloud computing resources, or a combination thereof. In various embodiments, one or more software modules include, by way of non-limiting examples, web applications, mobile applications, standalone applications, and distributed or cloud computing applications. In some embodiments, a software module is within one computer program or application. In other embodiments, a software module is within more than one computer program or application. In some embodiments, a software module is hosted on one machine. In other embodiments, a software module is hosted on multiple machines. In further embodiments, the software modules are hosted on a distributed computing platform, such as a cloud computing platform. In some embodiments, the software modules are hosted on one or more machines in one location. In other embodiments, the software modules are hosted on one or more machines in more than one location.

[0140] Database In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or the use thereof. Given the disclosure provided herein, those skilled in the art will recognize that many databases are suitable for storing and retrieving nucleic acid (e.g., RNA) structure, sequence, and chemical mapping information. In various embodiments, suitable databases include, by way of non-limiting example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, XML databases, document-oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL®, Oracle, DB2, Sybase, and MongoDB. In some embodiments, the database is internet-based. In further embodiments, the database is web-based. In still further embodiments, the database is cloud computing-based. In certain embodiments, the database is a distributed database. In other embodiments, the database is based on one or more local computer storage devices. [Example]

[0141] Example 1 - System for predicting RNA tertiary structure from chemical mapping data The training set was created based on i) 500 RNA nucleotide sequences and experimentally determined corresponding tertiary structures obtained from the Protein Data Bank (PDB) and ii) 10,000 RNA sequences for which 1D chemical mapping data were available in the form of site-specific reactivity via dimethyl sulfate (DMS) exposure of the corresponding RNA molecules, reverse transcription to DNA, and DNA sequencing.

[0142] The training set is used to train a machine learning algorithm for RNA tertiary structure prediction, and the information flow for this is shown in Figures 3A and 3B. The numerical representation of a given RNA nucleotide sequence is processed by two subparts of the algorithm, called the transformer module and the structural module, respectively. The combined output of the modules is the predicted RNA tertiary structure. This output is the input for another part of the algorithm, called the chemical module. The output of this module is the predicted chemical mapping data.

[0143] In this embodiment of the methods and systems described herein, the neural network architecture includes a transformer and a structure module. Nucleotide sequences are encoded using a one-hot vector representation that passes through a dense neural network layer to generate a numerical embedding. This embedding is input to a transformer module, which includes a series of gated attention layers that iteratively refine the sequence embedding at the nucleotide level and a second pairwise embedding. The output sequence and pairwise representations are input to a structure module, which predicts rotation and translation matrices for mainframes defined by the nucleotide atoms (e.g., C4', C1', N1 / N9), forming a representation of the RNA tertiary structure. The prediction of the rotation and translation matrices is based on geometry-aware attention operations and invariant-point attention. Applying the predicted rotation and translation matrices to individual mainframes yields a predicted tertiary structure.

[0144] The chemistry module includes a message-passing graph neural network that receives the predicted RNA tertiary structure as input. Atoms (C4', C1', N1 / N9) represent graph nodes, and edges are represented between each node and nearby nodes in Euclidean space based on the tertiary structure (e.g., within a specific cutoff, such as 15 Å). Euclidean distances correspond to edge features. Node features include atom types (e.g., C4', C1', N1 / N9) and nucleotide types, with each atom coded as one-hot. The node and edge features are revised through a series of message-passing steps. Predicted chemical mapping data is output based on the revised node and edge feature values, such as the average over all node features corresponding to nucleotides in the case of 1D chemical mapping data.

[0145] The neural network architecture, from RNA nucleotide sequence to RNA tertiary structure prediction to chemical mapping data, is discriminative. Therefore, the network parameters of all modules can be optimized based on training signals received by comparing predicted chemical mapping data with actual chemical mapping data for the portion of the training set that includes nucleotide sequence and chemical mapping data. Here, the Adam optimizer and L2 loss are used for the chemical mapping data. Training of this algorithm is shown in Figure 3B. This first training step is alternated with training on data in which we have nucleotide sequence and the corresponding experimentally determined tertiary structure. Training of this algorithm is shown in Figure 3C. Here, the loss (in terms of the LDDT metric) is calculated at the tertiary structure level, and the loss signal is backpropagated through the transformer and structure modules to revise their parameters.

[0146] The tertiary structure prediction algorithm can be trained using chemical mapping data that does not include any tertiary structures for the RNA molecules in the chemical mapping dataset. Furthermore, after training, the chemical mapping data is not required and the currently optimized algorithm can be used to make tertiary structure predictions for new RNA sequences.

[0147] The parameters of the neural network architecture and the described training procedures are optimized using methods described elsewhere herein (e.g., by withholding a subset of training data encompassing RNA sequences with known tertiary structures and evaluating the accuracy of the tertiary structure prediction algorithm for a given set of parameters based on this withheld data). These parameters may include the number and specifications of different neural network layers, the dimensionality of the learned embedding, the choice of learning rate, the initialization of the algorithm's learnable variables (e.g., Xavier initialization), the optimizer (e.g., Adam), the loss function (e.g., L2 norm) and their relative weights in the optimization, the use of dropout layers, and the number of message-passing steps and nearest neighbors (e.g., for a chemistry module that may include a graph neural network).

[0148] After the algorithm is trained, one or more RNA sequences of interest are provided to the algorithm, which predicts and outputs RNA tertiary structure and chemical mapping data.

[0149] Example 2 - Iterative refinement of predicted RNA tertiary structure The system is trained and deployed as described in Example 1 to predict the tertiary structure of a target RNA molecule. The predicted tertiary structure is iteratively improved based on multiple passes through the structure module, where pairwise nucleotide representations derived from the current tertiary structure prediction are provided as new inputs to the structure module. This iterative process ends when the algorithm's own prediction of the local accuracy of the structure prediction (measured based on the predicted local distance difference test (pLDDT) metric) converges. The algorithm's ability to predict the quality of its own tertiary structure prediction is trained based on ground truth tertiary structure data provided during training and the LDDT metric, which can be calculated by comparing the tertiary structure data with the tertiary structure prediction.

[0150] Example 3 - Multitask setup for prediction of RNA tertiary structure and chemical mapping data In some embodiments, chemical mapping data prediction does not need to precede tertiary structure prediction, but is implemented as a multitask prediction problem in which at least part of the machine learning algorithm is shared between tasks. For example, one can start with a one-hot vector embedding of the sequence (e.g., of length 4 per nucleotide) and a second pairwise representation in which the relative displacements between nucleotides (at the sequence level) for each nucleotide pair are one-hot encoded and clipped at a predetermined length (e.g., 65). Both the one-hot sequence representation and the pairwise representation can be further processed through a linear layer to obtain an initial embedding, in the case of the pairwise representation after concatenating the substitution embedding for each nucleotide pair with the one-hot code of nucleotide identity. These embeddings can be input to the Transformer module described in Example 1. The sequence and pairwise outputs of this module can then be passed through a task-specific head, such as the structure module described for tertiary structure prediction or an example of chemical mapping data prediction, a linear network layer, followed by a sigmoid nonlinearity that predicts the probability of mutation for each position in the input sequence. The parameters of the transformer module can be optimized based on training data containing tuples of both (sequence, chemical mapping) and (sequence, RNA tertiary structure) information, a procedure that can result in improved accuracy in both tasks compared to training independent predictors for each task.

[0151] Example 4 - Chemical mapping data and sequential training of models for RNA tertiary structure prediction First, a machine learning algorithm is trained to predict chemical mapping data from RNA sequences using an initial encoding and a Transformer module similar to that described in Example 3, but without the multitasking setting. For a given target RNA sequence, the single and pairwise representations from the pre-trained Transformer module can be subsequently extracted and processed through a linear layer to obtain the desired dimensionality, which can be used as input for a machine learning algorithm designed to predict tertiary structure. This algorithm can include the Transformer and structure modules described in the previous examples. The parameters of the pre-trained ML algorithm that makes predictions on chemical mapping data can be kept fixed or adjusted as part of training the algorithm that predicts tertiary structure from the sequence.

[0152] In some embodiments, the tertiary structure representation includes different (sub)sets of atoms. The mainframe representation may include a variety of atoms, but may not be identical at every nucleotide in the RNA molecule. The mainframe representation may range from zero atoms to a complete enumeration of atoms at the level of individual nucleotides. The structure representation may not involve explicit embedding, in terms of 3D coordinates at the level of a distance map that is represented as part of pairwise features within the structure module and on which the chemistry module subsequently operates. This distance map representation may be extended by angles.

[0153] In some embodiments, the training set may further include sequences with data on the secondary structure of the sequences. This data may be used in various forms, such as in the form of a base pair probability matrix, either as pairwise input features for the algorithm or as additional prediction targets (similar to the case of 2D chemical mapping data described, for example). In the case of additional prediction embodiments, the secondary structure data may or may not be used to optimize the structural module during training (and thus improve tertiary structure prediction capabilities).

[0154] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. 1. A computer-implemented method for predicting the tertiary structure of an RNA molecule of interest, comprising: (a) creating a training data set, said training data set comprising: (i) chemical mapping data for a first plurality of RNA molecules; and (ii) including tertiary structure data for a second plurality of RNA molecules; (b) training a machine learning algorithm using the training dataset; (c) applying the trained machine learning algorithm to predict the tertiary structure of the RNA molecule of interest; (d) outputting the predicted tertiary structure of the target RNA molecule; A method comprising:

2. 1. A computer-implemented method for predicting the tertiary structure of an RNA molecule of interest, comprising: (a) obtaining a machine learning model, the machine learning model comprising: (i) generating a training dataset comprising chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; and (ii) training the machine learning model using the training dataset; (b) applying the machine learning model to predict the tertiary structure of the RNA molecule of interest; (c) outputting the predicted tertiary structure of the target RNA molecule; A method comprising:

3. 3. The method of claim 1 or 2, wherein the chemical mapping data is generated by a process comprising contacting an RNA molecule with a chemical probing agent, and optionally the RNA molecule is at least one of the first plurality of RNA molecules or the RNA molecule of interest.

4. The method of claim 3 , wherein the chemical probing agent comprises dimethyl sulfate (DMS).

5. The method of claim 3, wherein the chemical probing agent comprises a SHAPE (selective 2'-hydroxyl acylation and primer extension) reagent.

6. 6. The method of claim 5, wherein the SHAPE reagent is 1-methyl-7-nitroisatoic anhydride (1M7), 1-methyl-6-nitroisatoic anhydride (1M6), 5-nitroisatoic anhydride (5NIA), or N-methyl-nitroisatoic anhydride (NMIA).

7. 4. The method of claim 3, wherein the chemical probing agent comprises 2A3 ((2-aminopyridin-3-yl)(1H-imidazol-1-yl)methanone).

8. The method of any one of claims 1 to 7, wherein the RNA molecule of interest comprises a portion of the transcriptome.

9. The method of claim 8, wherein the transcriptome is a human transcriptome.

10. 10. The method of any one of claims 1 to 9, wherein the training set comprises chemical mapping data for at least about 10, 100, 500, 1,000, 10,000, or more than 10,000 sequences.

11. 11. The method of any one of claims 1 to 10, wherein the training set comprises chemical mapping data for at most about 10, 100, 500, 1,000, or 10,000 sequences.

12. The method of any one of claims 1 to 11, wherein the chemical mapping data relates to sequences that occur in different abundance than in natural systems.

13. The method of any one of claims 1 to 12, wherein the chemical mapping data is collected from an in vitro source.

14. 14. The method of claim 1, further comprising the step of tuning the machine learning algorithm based on chemical mapping data of the RNA molecule of interest before applying the machine learning model.

15. The method of any one of claims 1 to 14, wherein the machine learning algorithm comprises one or more artificial neural networks (ANN).

16. 16. The method of claim 15, further comprising training the ANN to predict chemical mapping data for the RNA molecule of interest.

17. 17. The method of claim 16, further comprising predicting the chemical mapping data for the RNA molecule of interest from a predicted tertiary structure of the RNA molecule of interest.

18. 17. The method of claim 16, further comprising predicting a predicted tertiary structure of the RNA molecule of interest based on the chemical mapping data for the RNA molecule of interest.

19. 17. The method of claim 16, further comprising predicting the chemical mapping data for the RNA molecule of interest and a predicted tertiary structure of the RNA molecule of interest using the same embedding.

20. The method of any one of claims 1 to 19, wherein the tertiary structure comprises the 3D coordinates of a plurality of atoms that make up the RNA molecule of interest.

21. 21. The method of claim 20, wherein the tertiary structure comprises the 3D coordinates of each atom that constitutes the RNA molecule of interest.

22. 22. The method of any one of claims 1 to 21, wherein the tertiary structure comprises one or more 3D coordinates of a plurality of nucleotides that make up the RNA molecule of interest.

23. 23. The method of claim 22, wherein the tertiary structure comprises one or more 3D coordinates of each nucleotide that constitutes the RNA molecule of interest.

24. The method of any one of claims 1 to 23, wherein the tertiary structure of the RNA molecule of interest is parameterized based on a distance map.

25. The method of any one of claims 1 to 24, wherein the tertiary structure of the RNA molecule of interest is parameterized based on a distance map and angles.

26. The method of any one of claims 1 to 25, which does not require determination or prediction of the secondary structure of the RNA molecule of interest.

27. 27. The method of any one of claims 1 to 26, which predicts aspects of the tertiary structure of a target RNA that are not captured by base pairing prediction of the target RNA.

28. 28. The method of any one of claims 1 to 27, wherein the predicted tertiary structure comprises one or more of a pseudoknot, a multi-way junction, a coaxial stack, an a-minor motif, a kissing stem loop, a ribose zipper, or a tetraloop / tetraloop acceptor.

29. 29. The method of any one of claims 1 to 28, wherein the chemical mapping data comprises multidimensional chemical mapping data for one or more RNA molecules of the first plurality of RNA molecules.

30. The method of any one of claims 1 to 29, wherein the predicted tertiary structure is a target for a pharmaceutical drug.

31. The method of any one of claims 1 to 29, further comprising determining a target region or subsequence of the RNA molecule of interest that is targeted by a pharmaceutical agent based on the predicted tertiary structure.

32. 32. The method of any one of claims 1 to 31, further comprising formulating a pharmaceutical product based on the predicted tertiary structure.

33. 33. The method of any one of claims 1 to 32, wherein the training dataset further comprises a multiple sequence alignment of a third plurality of RNA molecules.

34. 34. The method of Claim 33, wherein the first plurality of RNA molecules and the second plurality of RNA molecules are the same or different.

35. 35. The method of claim 33 or 34, wherein the first plurality of RNA molecules and the third plurality of molecules are the same.

36. 36. The method of any one of claims 1 to 35, wherein the first plurality of RNA molecules is unrelated to the RNA molecule of interest.

37. The second plurality of RNA molecules is unrelated to the RNA molecule of interest.

36. The method according to any one of claims 1 to 35.

38. 38. The method of any one of claims 33 to 37, wherein the third plurality of RNA molecules is unrelated to the RNA molecule of interest.

39. 39. The method of any one of claims 1-38, wherein one RNA molecule of the first plurality of RNA molecules has about 80%, 70%, 60%, 50%, 40%, 30%, 20% or less sequence identity to the RNA molecule of interest.

40. 1. A computer-implemented system for predicting the tertiary structure of an RNA molecule of interest, comprising: a computing device comprising at least one processor; and instructions executable by said at least one processor for performing operations, said operations comprising: a) creating a training dataset, said training dataset comprising: i) chemical mapping data for a first plurality of RNA molecules; ii) generating tertiary structure data for a second plurality of RNA molecules, including one or more of the tertiary structure data; b) training a machine learning algorithm using the training dataset; c) applying the trained machine learning algorithm to predict the tertiary structure of the RNA molecule of interest; and d) outputting the predicted tertiary structure of the RNA molecule of interest; 1. A computer-implemented system comprising:

41. 1. A non-transitory computer-readable storage medium executable by one or more processors and encoded with instructions for providing an application for predicting the tertiary structure of an RNA molecule of interest, the application comprising: a) a training dataset module configured to generate a training dataset including chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; b) a training module configured to train a machine learning algorithm using the training dataset; and c) an inference module configured to predict the tertiary structure of the RNA molecule of interest by applying the trained machine learning algorithm; d) an output module configured to report the predicted tertiary structure of the RNA molecule of interest; 1. A non-transitory computer-readable storage medium comprising:

42. accessing an RNA tertiary structure prediction system, said RNA tertiary structure prediction system comprising: generating a training dataset comprising chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; training the machine learning model using the training dataset; and accessing the trained machine learning model produced by a process including storing the trained machine learning model on the non-transitory computer-readable medium; When executed by a computing system, the computing system: predicting the tertiary structure of the RNA molecule of interest using the RNA tertiary structure prediction system; and outputting a predicted tertiary structure of the RNA molecule of interest; 1. A non-transitory computer-readable medium comprising:

43. 1. A computer-implemented method for predicting the tertiary structure of an RNA molecule of interest, comprising: (a) sending a query for predicting a tertiary structure of the RNA molecule of interest to a computer comprising a machine learning model, wherein the machine learning algorithm generates the tertiary structure, and the machine learning model: (i) generating a training dataset comprising chemical mapping data for a first plurality of RNA molecules and tertiary structure data for a second plurality of RNA molecules; and (ii) trained by a process comprising training the machine learning model using the training dataset; (b) receiving from the computer a predicted tertiary structure of the RNA molecule of interest; A method comprising: