A physical model-based computational framework and system for designing nucleic acids for therapeutics and research

The tBOND-G and tBOND-L computational framework addresses inefficiencies in nucleic acid design by predicting optimal guide RNAs and oligonucleotides for structured ssRNA, achieving precise cleavage and selective inhibition, enhancing therapeutic and diagnostic efficacy.

WO2026136966A1PCT designated stage Publication Date: 2026-06-25CALIFORNIA INST OF TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CALIFORNIA INST OF TECH
Filing Date
2025-12-19
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Existing methods for designing functional nucleic acids, such as guide RNAs and antisense oligonucleotides, are inefficient and often result in poor outcomes due to the complex secondary structures of single-stranded RNA (ssRNA) blocking target access and the inability to distinguish between specific RNA fragments and their highly similar parent molecules, leading to inefficient processing and unintended off-target effects.

Method used

A computational framework, tBOND-G and tBOND-L, which uses machine learning models to predict optimal guide RNA and complementary oligonucleotide sequences for precise cleavage and binding, respectively, by analyzing structural accessibility and thermodynamic stability, and simulating competitive binding environments to ensure high specificity and efficiency.

Benefits of technology

The framework enables high-efficiency and high-specificity design of nucleic acids, allowing for precise cleavage of structured ssRNA into target fragments and selective inhibition of pathogenic fragments without disrupting essential cellular functions, reducing development time and costs for therapeutics and diagnostics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025060785_25062026_PF_FP_ABST
    Figure US2025060785_25062026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented framework for designing high-efficiency and high-specificity nucleic acid molecules targeting transfer RNA (tRNA) and its derivatives (tDRs). The framework comprises two primary algorithms. The first, tBOND-G, designs guide RNAs (gRNAs) for Cas13-mediated tRNA cleavage by calculating physical parameters, including target site accessibility and binding energy, and processing them through a Support Vector Machine (SVM) model to predict cleavage efficiency. The second algorithm, tBOND-L, designs therapeutic Locked Nucleic Acid-modified antisense oligonucleotides (LNA-ASOs) that specifically target a tDR without binding to its parent tRNA. This method utilizes a processor to derive an efficiency score based on relative binding affinity and a specificity score based on simulated competitive binding environments.
Need to check novelty before this filing date? Find Prior Art

Description

" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT A PHYSICAL MODEL-BASED COMPUTATIONAL FRAMEWORK AND SYSTEM FOR DESIGNING NUCLEIC ACIDS FOR THERAPEUTICS AND RESEA CHCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to US Prov. App. No. 63 / 736,977 filed on December 20, 2024, the contents of which are incorporated herein by reference in their entirety.STATEMENT OF GOVERNMENT INTEREST

[0002] This invention was made with government support under Grants No. R35 HL150807 and R01 HL174709 awarded by the National Institutes of Health. The government has certain rights in the invention.INCORPORATION BY REFERENCE STATEMENT FOR SEQUENCE LISTING

[0003] Further, the computer readable form of the sequence listing of the XML file P3304-PCT-ST26 created on December 19, 2025, and having size 20,523 bytes measured on Windows Server 2019 Datacenter is incorporated herein by reference in its entirety.FIELD OF THE DISCLOSURE

[0004] The present disclosure relates generally to the field of computational biochemistry and bioinformatics. More specifically, it pertains to computer-implemented methods and systems for the design and optimization of functional nucleic acid molecules.BACKGROUND

[0005] Functional nucleic acids, such as guide RNAs and antisense oligonucleotides (ASOs), have become essential tools for regulating cellular processes and developing molecular therapeutics. Historically, designing these molecules relied on inefficient trial-and-error screening, which is resource-intensive and often yields poor results. Consequently, the field has shifted toward computational frameworks that utilize biophysical principles, such as thermodynamic stability and structural accessibility, to rationally design more effective candidates." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0006] Despite these advancements, achieving precise cleavage of single-stranded ribonucleic acid (ssRNA) remains a significant technical challenge. The complex secondary structures of ssRNA can block access to target sites, and existing methods often struggle to distinguish between specific RNA fragments and their highly similar parent molecules, leading to inefficient processing or unintended off-target effects.SUMMARY

[0007] The present disclosure provides a framework comprising computer-implemented methods and systems for the high-efficiency, high-specificity design of nucleic acids targeting structured single- stranded RNAs (ssRNAs) and their derivatives.

[0008] In particular, in accordance with the disclosure, two distinct but complementary processes are presented: Target Binding and Optimized Nucleotide Design with guide RNAs (tBOND-G) and Target Binding and Optimized Nucleotide Design with complementary oligonucleotides (tBOND-L). These processes transform the design process from one of empirical chance to one of predictive, data-driven engineering.

[0009] The first aspect of the disclosure is the tBOND-G method and system, which is configured to design optimal guide RNA (gRNA) sequences for use with an RNA-guided RNA-targeting nuclease (e.g., a Casl3 family nuclease). This system enables the cleavage of a parent structured single- stranded RNA to endogenously generate specific target fragments (e.g., tRNA-derived fragments (tDRs) from a full-length tRNA) for functional studies.

[0010] The tBOND-G system integrates multiple data types into a predictive machine learning model. A processor is configured to receive a parent structured single- stranded RNA (ssRNA) sequence (e.g., a target tRNA sequence), generate a library of candidate gRNAs, and calculate two physical parameters for each candidate: (1) an accessibility score, derived from a secondary structure model of the ssRNA, which quantifies how physically exposed a target site is for binding, and (2) a binding energy ratio (Delta G Ratio), which quantifies the thermodynamic stability of the complex formed between the gRNA and the parent ssRNA.

[0011] These calculated parameters are then used as features in a pre-trained Support Vector Machine (SVM) model, which has learned the relationship between these physical properties and" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT actual experimental cleavage outcomes.

[0015] The SVM model outputs a predicted cleavage efficiency score, allowing for the identification of optimal gRNAs. The system can further predict the exact cleavage sites and generate a virtual gel visualization of the expected target fragments (e.g., tDR fragments), providing an invaluable guide for subsequent experiments.

[0013]

[0012] Furthermore, the tBOND-G system can be configured to perform an off-target analysis, wherein a selected gRNA sequence is computationally screened against a transcriptome database to identify and score potential unintended binding sites, ensuring high specificity for the final selected candidate.

[0013] The second aspect of the disclosure is the tBOND-L method and system, which is configured to design optimal complementary oligonucleotide sequences (e.g., antisense oligonucleotides (ASOs)) for the specific binding to (and potential inhibition of) target fragments (e.g., disease-associated tDRs). This system deals with the challenge of designing a complementary oligonucleotide that can specifically bind to a target fragment without crossreacting with the nearly identical parent structured ssRNA.

[0014] The tBOND-L system performs a systematic analysis of all possible complementary oligonucleotide candidates for a given target fragment region. For each candidate, a processor is configured to calculate two distinct performance metrics: (1) Efficiency, a measure of the oligonucleotide's binding strength to its intended target fragment relative to its binding to competing sequences within the same molecule, and (2) Specificity, a novel metric calculated by simulating a multi-component reaction mixture (containing the oligonucleotide, the target fragment, competing fragments, and the parent structured ssRNA) to determine the probability of the oligonucleotide binding exclusively to its intended target in a competitive cellular-like environment.

[0015] The system outputs a ranked list of complementary oligonucleotide candidates based on these metrics, streamlining the selection process for potent and specific agents. Together, the tBOND-G and tBOND-L systems provide a structured and validated computational framework that improves the precision, efficiency, and success rate of designing nucleic acids for structured ssRNA-related research and therapeutics." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0016] According to a third aspect, a method is described for the functional characterization and validation of a specific target fragment (e.g., a tRNA-derived fragment). The method comprises: identifying a specific target fragment sequence derived from a parent structured single-stranded RNA (ssRNA) molecule, and operating an integrated computational platform in a first gain-of-function mode to design a guide RNA (gRNA) that targets the parent structured ssRNA. This first mode involves executing a machine learning model to select a gRNA sequence based on the structural accessibility of the parent structured ssRNA and the thermodynamic binding energy of the gRNA-parent ssRNA complex to efficiently induce the cleavage of the parent structured ssRNA into the specific target fragment.

[0017] The method further comprises operating the platform in a second loss-of-function mode to design a complementary oligonucleotide inhibitor specific to the same target fragment generated in the first mode. This second mode involves simulating a competitive binding environment to select a complementary oligonucleotide sequence based on a calculated efficiency score and a calculated specificity score, ensuring the complementary oligonucleotide binds the target fragment with high affinity while exhibiting negligible binding to the parent structured ssRNA.

[0018] The method also comprises coordinating the first and second modes to establish a functional feedback loop, wherein the designed gRNA is utilized to induce the expression of the target fragment to observe a specific biological phenotype, and the designed complementary oligonucleotide is subsequently utilized to inhibit the target fragment to reverse said phenotype, thereby verifying the causal role of the target fragment in a biological pathway.

[0019] According to a fourth aspect, a method is described for treating a disease associated with the aberrant accumulation of a pathogenic target fragment (e.g., a tRNA-derived fragment (tDR)). The method comprises: identifying a pathogenic target fragment associated with a condition in an individual, wherein the pathogenic target fragment is comprised within a full-length parent structured single-stranded RNA (ssRNA) molecule (e.g., a parent tRNA), and generating a library of candidate complementary oligonucleotide sequences complementary to the identified pathogenic target fragment.

[0020] The method further comprises calculating an efficiency score for each candidate in the library, representing the binding affinity of the complementary oligonucleotide to the pathogenic" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT target fragment relative to potential off-target sequences located within the parent structured ssRNA molecule and the non-pathogenic remnants from parent structured ssRNA, and simulating a competitive binding environment to derive a specificity score. The simulation models the equilibrium concentrations of the candidate complementary oligonucleotide, the pathogenic target fragment, the non-pathogenic remnants from parent structed ssRNA and the parent structured ssRNA to predict the probability of exclusive binding to the fragment.

[0021] The method also comprises selecting a specific complementary oligonucleotide sequence from the library that exhibits high efficiency and high specificity for the fragment while minimizing binding to the parent structured ssRNA, and administering a therapeutically effective amount of the selected complementary oligonucleotide to the individual having the condition to reduce the accumulation of the pathogenic target fragment without disrupting the essential biological functions (e.g., protein synthesis) of the parent structured ssRNA.

[0022] According to a fifth aspect, a method is described for modulating cellular function by inducing the formation of a bioactive target fragment (e.g., a bioactive tRNA-derived fragment). The method comprises identifying a parent structured single- stranded RNA (ssRNA) molecule (e.g., a structured precursor tRNA) capable of being processed into a bioactive target fragment involved in gene expression regulation, translation repression, or stress response, and delivering to a target cell population a CRISPR-associated RNA-targeting nuclease system (e.g., a Casl3 family nuclease) and a guide RNA (gRNA). The spacer sequence of the gRNA is selected to target a specific region of the parent structured ssRNA.

[0023] The method further comprises employing the tBOND-G algorithm to predict cleavage efficiency for the gRNA. The tBOND-G algorithm utilizes a machine learning model configured to analyze biophysical parameters, specifically relying on a numerical accessibility score derived from a secondary structure model of the parent structured ssRNA and a thermodynamic binding energy value of the gRNA-parent structured ssRNA complex to identify sequences that target accessible loops within the folded structure.

[0024] The method utilizes a CRISPR-associated RNA-targeting nuclease selected from RNA-targeting variants (e.g., Cas13 variants) to induce the controlled generation of the bioactive target fragment within the target cell population, thereby modulating said cellular function." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0025] According to a sixth aspect, a method is described for detecting aberrant accumulation of a target fragment biomarker (e.g., a tRNA-derived fragment (tDR)) in an individual. The method comprises performing the tBOND-L algorithm of the present disclosure to design a complementary oligonucleotide specific for the target fragment biomarker.

[0026] The method further comprises obtaining a biological sample from the individual and contacting the sample with a detectable probe comprising the designed complementary oligonucleotide. The sequence of the complementary oligonucleotide probe is selected to specifically hybridize to the target fragment biomarker, wherein the target fragment biomarker is comprised within a full-length parent structured single-stranded RNA (ssRNA) molecule. The contacting is performed to allow formation of a target-probe complex.

[0027] The method also comprises detecting the presence or quantifying the level of the targetprobe complex in the sample. In some embodiments, detection of a level of the complex above a predetermined threshold indicates the presence of a condition or a specific stage of condition progression, enabling diagnosis based on detection of the target fragment biomarker.

[0028] According to a seventh aspect, a method of molecular engineering is described for regulating cellular function (e.g.. protein translation efficiency or stress response pathways) in a eukaryotic cell. The method comprises introducing into the cell a programmable CRISPR-associated RNA-targeting nuclease system and a specifically designed guide RNA (gRNA). The gRNA is engineered to target a precise region of a specific parent structured single-stranded RNA (ssRNA) transcript to induce the controlled formation of a bioactive target fragment (e.g., a tDR) capable of inhibiting translation initiation or promoting stress granule formation.

[0029] The method further comprises selecting the gRNA sequence using the tBOND-G computational framework. This framework optimizes cleavage efficiency by analyzing the structural accessibility of the parent structured ssRNA loops and the thermodynamic stability of the gRNA-ssRNA interaction, thereby allowing for tunable control over the cell’s translational state without abolishing the entire parent ssRNA pool.

[0030] In some embodiments, the method of molecular engineering further comprises verifying functional decoupling in the RNA regulatory network by contacting the cell with a high-affinity" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT complementary oligonucleotide reagent. This oligonucleotide is designed using the tBOND-L algorithm to bind the specific cleaved target fragment while discriminating against the full-length parent structured ssRNA, thereby selectively inhibiting the regulatory function of the fragment to validate the engineered phenotype while maintaining the canonical metabolic function of the parent.

[0031] According to an eighth aspect, a method is described for the design of a nucleic acid reagent (e.g., a therapeutic or diagnostic oligonucleotide) specific for a target environment. The method comprises:(a) obtaining a biological sample from the target environment (e.g., a specific tissue, cell line, or patient biopsy characterized by a unique molecular landscape);(b) experimentally measuring a concentration ratio between a target fragment (e.g.. a pathogenic tDR) and its parent structured ssRNA (e.g., the parent tRNA) within said sample;(c) inputting said concentration ratio into the tBOND-L computational framework as a boundary condition for the competitive binding simulation; and(d) calculating a specificity score specific to the target environment to select a complementary oligonucleotide that is optimized to distinguish the target from the parent under the specific physiological conditions of said environment.

[0032] In the context of gain-of-function studies, the computer-implemented methods and systems for the design of nucleic acids targeting structured single-stranded RNAs (ssRNAs) and their derivatives, and related methods and systems herein described, in some embodiments allow for the efficient and specific endogenous generation of target fragments (e.g., tDRs or tRNA halves). In particular embodiments, the systems enable the production of biologically relevant target fragments within the cellular environment using Cas-mediated cleavage by selecting guide RNAs (gRNAs) that target physically accessible regions of the parent ssRNA secondary structure. Furthermore, the systems of the disclosure allow researchers to visualize and predict the precise cleavage sites and resulting fragment sizes via virtual gel simulations prior to experimentation, ensuring that the induced fragments match the intended biological targets." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0033] In the context of loss-of-function and therapeutic applications, the computer-implemented methods and systems for the design of nucleic acids targeting structured ssRNAs and their derivatives, and related methods and systems herein described, allow for the selective inhibition of pathogenic target fragments (e.g., pathogenic tDRs) without disrupting the essential cellular functions of the parent structured ssRNA (e.g.. the full-length tRNA). These embodiments address the significant challenge of sequence homology between fragments and their parents by utilizing a multi-component competitive binding simulation. This enables the design of high-affinity complementary oligonucleotide reagents (e.g., ASOs, LNAs, or PNAs) that exhibit high specificity for the fragment, ensuring that the housekeeping role of the parent molecule (e.g., in protein synthesis) is preserved while the disease-driving fragment is neutralized.

[0034] In the context of context- specific design, the methods and systems herein described allow for the personalization of nucleic acid reagents based on the specific molecular landscape of a target environment. By incorporating experimental measurements of the concentration ratio between a target fragment and its parent ssRNA as input parameters, the algorithms perform accurate stoichiometric calculations. This capability allows for the design of reagents that are tuned to the specific abundance levels found in a particular disease state or tissue type, thereby maximizing efficacy while minimizing off-target binding in vivo.

[0035] The computer-implemented methods and systems for the design of nucleic acids targeting structured ssRNAs and their derivatives, and related methods and systems herein described, allow in some embodiments for the precise modulation of critical cellular processes regulated by RNA fragments, including apoptosis, autophagy, stress granule formation, and translation repression. This capability facilitates both the functional characterization of these molecules in basic research and the development of targeted therapies for diseases driven by RNA dysregulation, such as cancer, metabolic disorders, cardiovascular diseases, renal diseases, and neurological conditions. The enhanced efficiency (binding favorability to on target species) and specificity (less off-target binding to the parental strand) provided by these systems will markedly decrease the cost and time required to develop novel activators and inhibitors of bioactive RNA fragments.

[0036] The computer-implemented methods and systems for the design of nucleic acids targeting structured ssRNAs and their derivatives, and related methods and systems herein described, allow" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT in some embodiments achieving high-efficiency design of nucleic acids, wherein the design selection utilizes a machine learning model characterized by a Receiver Operating Characteristic Area Under the Curve (ROC AUC) of at least 0.90. In some embodiments, the selected nucleic acids exhibit a predicted cleavage efficiency score of at least 0.6, or in preferred embodiments, at least 0.7 on a normalized scale. Accordingly, the computational selection results in an experimental validation rate for the designed nucleic acids of at least 80% and up to 100% with respect to the generation of specific RNA fragments.

[0037] The computer-implemented methods and systems for the design of nucleic acids targeting structured ssRNAs and their derivatives, and related methods and systems herein described, allow for the high- specificity design of nucleic acids, particularly complementary oligonucleotides (e.g., ASOs or LNAs), wherein the selected nucleic acid candidates exhibit a calculated specificity score of at least 75, and in some embodiments at least 80, based on a competitive binding simulation. In certain embodiments, these high-specificity candidates simultaneously exhibit a calculated efficiency score of at least 100, representing a binding affinity superior to that of the native interaction context.

[0038] The computer-implemented methods and systems for the design of nucleic acids targeting structured ssRNAs and their derivatives, and related methods and systems herein described, can be used in connection with any applications wherein high specificity and / or high selectivity RNA targeting is desired. Exemplary applications comprise application in a variety of industrial and medical fields, such as biotherapeutics, medical drug development, and clinical applications. Specific implementations extend to diagnostic applications, including in-vitro diagnostics, cancer diagnostics, and prenatal diagnostics, where precise RNA profiling is required. Furthermore, the technology is applicable in broader sectors such as biotechnology, agricultural biotechnology, and food testing, as well as specialized fields like bioanalysis, genetic testing, and immunology. The scope of the disclosure thus encompasses these domains and any additional applications identifiable by a skilled person upon reading the present specification.

[0039] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT BRIEF DESCRIPTION OF DRAWINGS

[0040] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate one or more embodiments of the present disclosure and, together with the description of example embodiments, serve to explain the principles and implementations of the disclosure.

[0041] FIG. 1 is an algorithmic workflow diagram illustrating the overall process of the tBOND-G system for designing guide RNAs (gRNAs) for RNA-guided RNA-targeting nucleases (e.g., Casl3), from user input to final prediction outputs. In the illustrated embodiment, the user input (step 105) is exemplified by the distinct tRNA sequence set forth in SEQ ID NO: 21 (Homo sapiens tRNA-Ala-CGC), which serves as a representative input string for initiating the library generation process.

[0042] FIG. 2 illustrates the physical principles behind the tBOND-G algorithm, showing how structured single- stranded RNA (ssRNA) secondary structure (exemplified here by tRNA) is used to calculate accessibility and how binding energy is calculated based on thermodynamic principles.

[0043] FIG. 3 shows the 2D cluster mapping used to train the Support Vector Machine (SVM) model, plotting accessibility against the Delta G ratio and visualizing the decision boundary that separates high- and low-performing gRNAs.

[0044] FIG. 4 depicts a molecular model of an exemplary pspCas13b-gRNA-tRNA complex, derived from molecular dynamics simulations, utilized to identify the putative binding pockets responsible for specific target fragment (e.g., 5' and 3' tDR) cleavage.

[0045] FIG. 5 presents experimental validation results, comparing the virtual gel predictions of the tBOND-G system with actual Northern blot data for gRNAs selected from a high-accessibility cluster.

[0046] FIG. 6 presents further experimental validation, comparing virtual gel predictions with Northern blot data for gRNAs selected from a high-binding-energy (AG) cluster.

[0047] FIG.7 presents additional experimental validation for gRNAs selected from a cluster with balanced accessibility and binding energy, demonstrating the model's ability to predict low-" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT efficiency outcomes correctly.

[0048] FIG. 8 is directed to the tBOND-G model by presenting validation results for a different target, tRNA-Asp-GTC, confirming its applicability across different ssRNA sequences.

[0049] FIG. 9 is a flowchart illustrating the pseudocode and logical steps of the tBOND-L algorithm for designing complementary oligonucleotide sequences (e.g., ASOs).

[0050] FIG. 10 is a diagram illustrating the four-component interaction system modeled by the tBOND-L algorithm to calculate the specificity of a complementary oligonucleotide in a competitive binding environment.

[0051] FIG. 11 is an exemplary output plot from the tBOND-L system, visualizing all designed complementary oligonucleotides on a 2D graph of Specificity versus Efficiency to facilitate the selection of optimal candidates.

[0052] FIG.12 is a block diagram illustrating a generic computer system architecture, comprising a CPU, memory, storage devices, and input / output interfaces, suitable for implementing the computational methods and algorithms described herein, including tBOND-G and tBOND-L.DETAILED DESCRIPTION

[0053] Described herein are computer-implemented methods and systems for design of nucleic acids targeting structured single stranded RNAs and their derivatives. In particular described herein are apparatuses, systems, non-transitory computer- readable media, and computer-implemented methods for the in silico design, identification, and engineering of nucleic acids targeting structured single-stranded RNAs and their derivatives.

[0054] The terms “nucleic acid” “NA” or “polynucleotide” as used herein indicates an organic polymer composed of two or more monomers including nucleotides, nucleosides or analogs thereof. The term “nucleotide” refers to any of several compounds that consist of a ribose or deoxyribose sugar joined to a purine or pyrimidine base and to a phosphate group and that is the basic structural unit of nucleic acids. The term “nucleoside” refers to a compound (such as" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT guanosine or adenosine) that consists of a purine or pyrimidine base combined with deoxyribose or ribose and is found especially in nucleic acids. The term “nucleotide analog” or “nucleoside analog” refers respectively to a nucleotide or nucleoside in which one or more individual atoms have been replaced with a different atom or a with a different functional group. Exemplary functional groups that can be comprised in an analog include methyl groups and hydroxyl groups and additional groups identifiable by a skilled person.

[0055] Exemplary monomers of a polynucleotide comprise deoxyribonucleotide, ribonucleotides, and modified nucleotide monomers (such as LNA nucleotides and PNA nucleotides). The term “deoxyribonucleotide” refers to the monomer, or single unit, of DNA, or deoxyribonucleic acid. Each deoxyribonucleotide comprises three parts: a nitrogenous base, a deoxyribose sugar, and one or more phosphate groups. The nitrogenous base is typically bonded to the T carbon of the deoxyribose, which is distinguished from ribose by the presence of a proton on the 2' carbon rather than an -OH group. The phosphate group is typically bound to the 5' carbon of the sugar.

[0056] The term “DNA” or “deoxyribonucleic acid” as used herein indicates a polynucleotide composed of deoxyribonucleotide bases or an analog thereof to form an organic polymer. The term “deoxyribonucleotide” refers to any compounds that consist of a deoxyribose (deoxyribonucleotide) sugar joined to a purine or pyrimidine base and to a phosphate group, and that are the basic structural units of a deoxyribonucleic acid, typically adenine (A), cytosine (C), guanine (G), and thymine (T). In DNA adjacent ribose nucleotide bases are chemically attached to one another in a chain typically via phosphodiester bonds. The term “deoxyribonucleotide analog” refers to a deoxyribonucleotide in which one or more individual atoms have been replaced with a different atom with a different functional group. For example, deoxyribonucleotide analogues include chemically modified deoxyribonucleotides, such as methylation hydroxymethylation glycosylation and additional modifications identifiable by a skilled person.

[0057] The term “modified nucleotides” refers to a nucleic acid monomer that is not the standard DNA or RNA nucleotide or nucleoside. In particular, modified nucleotides comprise nucleotide analogs presenting one or more individual atoms which have been replaced with a different atom or with a different functional group. Exemplary functional groups that can be comprised in an analog include methyl groups and hydroxyl groups and additional groups identifiable by a skilled" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT person.

[0058] The term “ribonucleotide” refers to the monomer, or single unit, of RNA, or ribonucleic acid. Ribonucleotides have one, two, or three phosphate groups attached to the ribose sugar. The term “RNA” or “ribonucleic acid” as used herein indicates a polynucleotide composed of ribonucleotide bases or an analog thereof linked to form an organic polymer. The term “ribonucleotide” refers to any compounds that consist of a ribose (ribonucleotide) sugar joined to a purine or pyrimidine base and to a phosphate group, and that are the basic structural units of a ribonucleic acid, typically adenine (A), cytosine (C), guanine (G), and uracil (U). In an RNA adjacent ribose nucleotide bases are chemically attached to one another in a chain typically via phosphodiester bonds.

[0059] The term “locked nucleic acids” (LNA) as used herein indicates a modified RNA nucleotide. The ribose moiety of an LNA nucleotide is modified with an extra bridge connecting the 2' and 4' carbons. The bridge "locks" the ribose in the 3'-endo structural conformation, which is often found in the A-form of DNA or RNA. LNA nucleotides can be mixed with DNA or RNA bases in the oligonucleotide whenever desired. The locked ribose conformation enhances base stacking and backbone pre-organization. This significantly increases the thermal stability (melting temperature) of oligonucleotides. LNA oligonucleotides display unprecedented hybridization affinity toward complementary single-stranded RNA and complementary single- or doublestranded DNA. Structural studies have shown that LNA oligonucleotides induce A-type (RNA-like) duplex conformations as will be understood by a skilled person.

[0060] The term “polyamide polynucleotide”, “peptide nucleic acid” or “PNA” as used herein indicates a type of artificially synthesized polymer composed of monomers linked to form a backbone composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The various purine and pyrimidine bases are linked to the backbone by methylene carbonyl bonds. Since the backbone of PNA contains no charged phosphate groups, the binding between PNA / DNA strands is stronger than between DNA / DNA strands due to the lack of electrostatic repulsion. PNA oligomers also show greater specificity in binding to complementary DNAs, with a PNA / DNA base mismatch being more destabilizing than a similar mismatch in a DNA / DNA duplex. This binding strength and specificity also applies to PNA / RNA duplexes. PNAs are not" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT easily recognized by either nucleases or proteases, making them resistant to enzyme degradation. PNAs are also stable over a wide pH range. In some embodiments, polynucleotides can comprise one or more non-nucleotidic or non nucleosidic monomers identifiable by a skilled person.

[0061] Accordingly, the term “polynucleotide” includes nucleic acids of any length, and in particular DNA, RNA, analogs thereof, such as LNA and PNA, and fragments thereof, possibly including non-nucleotidic or non-nucleosidic monomers, each of which can be isolated from natural sources, recombinantly produced, or artificially synthesized. A “nucleotidic oligomer” or “oligonucleotide” as used herein refers to a polynucleotide of three or more but equal to or less than 300 nucleotides. Polynucleotides can typically be provided in single-stranded form or doublestranded form (herein also duplex form, or duplex).

[0062] A “single-stranded polynucleotide” refers to an individual string of monomers linked together through an alternating sugar phosphate backbone. In particular, the sugar of one nucleotide is bound to the phosphate of the next adjacent nucleotide by a phosphodiester bond. Depending on the sequence of the nucleotides, a single- stranded polynucleotide can have various secondary structures, such as the stem-loop or hairpin structure, through intramolecular self-base-paring. A hairpin loop or stem loop structure occurs when two regions of the same strand, usually complementary in nucleotide sequence when read in opposite directions, base-pairs to form a double helix that ends in an unpaired loop. The resulting lollipop-shaped structure is a key building block of many RNA secondary structures as will be understood by a skilled person. The term “small hairpin RNA” or “short hairpin RNA” or “shRNA” as used herein indicate a sequence of RNA that makes a tight hairpin turn and can be used to silence gene expression via RNAi.

[0063] A single strand polynucleotide has a 5’ end and a 3’ end. The terms “5’ end” and “3’ end” of a single stranded polynucleotide indicate the terminal residues of the single strand polynucleotide and are distinguished based on the nature of the free group on each extremity. The 5 '-end of a single strand polynucleotide designates the terminal residue of the single strand polynucleotide that has the fifth carbon in the sugar-ring of the deoxyribose or ribose at its terminus (5’ terminus). The 3'-end of a single strand polynucleotide designates the residue terminating at the hydroxyl group of the third carbon in the sugar-ring of the nucleotide or nucleoside at its terminus (3’ terminus). The 5’ end and 3’ end terminus in various cases can be modified" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT chemically or biologically e.g. by the addition of functional groups or other compounds as will be understood by the skilled person.

[0064] A “double-stranded polynucleotide” or “duplex polynucleotide” refers to two singlestranded polynucleotides bound to each other through complementarily binding. The duplex typically has a helical structure, such as a double- stranded DNA (dsDNA) molecule or a double stranded RNA, which is maintained largely by non-covalent bonding of base pairs between the strands and by base stacking interactions. The term “5’-3’ terminal base pair” with reference to a duplex polynucleotide refers to the base pair positioned at an end of the duplex polynucleotide that is formed by the 5’ end of one single strand of the two single strand forming the duplex polynucleotide base-paired with the 3’ end of the single strand forming the duplex polynucleotide complementary to the one single strand. Accordingly a duplex polynucleotide formed by a first single strand complementarily bound to a second single strand, has two opposite ends: a first end of the duplex polynucleotide having a “5 ’-3’ terminal base pair” formed by the 5’ end of the first single strand and the 3’ end of the second single strand, and a second end of the duplex polynucleotide opposite to the first formed by the 3’ end of the first single strand and the 5’ end of the second single strand.

[0065] The term “complementary” as used herein indicates a property of single stranded polynucleotides in which the sequence of the constituent monomers on one strand chemically matches the sequence on another other strand to form a double stranded polynucleotide. Chemical matching indicates that the base pairs between the monomers of the single strand can be non-covalently connected via two or three hydrogen bonds with corresponding monomers in the another strand. In particular, in this application, when two polynucleotide strands, sequences or segments are noted to be complementary, this indicates that they have a sufficient number of complementary bases to form a thermodynamically stable double- stranded duplex. Double stranded of complementary single stranded polynucleotides include dsDNA, dsRNA, DNA: RNA duplexes as well as intramolecular base paring duplexes formed by complementary sequences of a single polynucleotide strand (e.g., hairpin loop).

[0066] The terms “complementary bind”, “base pair”, and “complementary base pair” as used herein with respect to nucleic acids indicates the two nucleotides on opposite polynucleotide" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT strands or sequences that are connected via hydrogen bonds. For example, in the canonical Watson-Crick DNA base pairing, adenine (A) forms a base pair with thymine (T) and guanine (G) forms a base pair with cytosine (C). In RNA base paring, adenine (A) forms a base pair with uracil (U) and guanine (G) forms a base pair with cytosine (C). Accordingly, the term “base pairing” as used herein indicates formation of hydrogen bonds between base pairs on opposite complementary polynucleotide strands or sequences following the Watson-Crick base pairing rule as will be applied by a skilled person to provide duplex polynucleotides. Accordingly, when two polynucleotide strands, sequences or segments are noted to be binding to each other through complementarily binding or complementarily bind to each other, this indicate that a sufficient number of bases pairs forms between the two strands, sequences or segments to form a thermodynamically stable double- stranded duplex, although the duplex can contain mismatches, bulges and / or wobble base pairs as will be understood by a skilled person.

[0067] The term "thermodynamic stability" as used herein indicates a lowest energy state of a chemical system. Thermodynamic stability can be used in connection with description of two chemical entities (e.g., two molecules or portions thereof) to compare the relative energies of the chemical entities. For example, when a chemical entity is a polynucleotide, thermodynamic stability can be used in absolute terms to indicate a conformation that is at a lowest energy state, or in relative terms to describe conformations of the polynucleotide or portions thereof to identify the prevailing conformation as a result of the prevailing conformation being in a lower energy state. Thermodynamic stability can be detected using methods and techniques identifiable by a skilled person. For example, for polynucleotides thermodynamic stability can be determined based on measurement of melting temperature Tm, among other methods, wherein a higher Tmcan be associated with a more thermodynamically stable chemical entity as will be understood by a skilled person. Contributors to thermodynamic stability can comprise chemical compositions, base compositions, neighboring chemical compositions, and geometry of the chemical entity.

[0068] Nucleic acid sequences can be designed to ensure complementary and specific binding of nucleic acids to form thermodynamically stable duplexes.

[0069] The wording “specific” “specifically” or “specificity” as used herein with reference to the binding of a first molecule to second molecule refers to the recognition, contact and formation of" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT a stable complex between the first molecule and the second molecule, together with substantially less to no recognition, contact and formation of a stable complex between each of the first molecule and the second molecule with other molecules that may be present. Exemplary specific bindings are antibody-antigen interaction, cellular receptor-ligand interactions, polynucleotide hybridization, enzyme substrate interactions and additional interactions identifiable by a skilled person.

[0070] Accordingly, the wording “specific binding” in nucleic acids refers to the precise molecular recognition between two sequences that form a thermodynamically stable interaction through complementary base pairing while avoiding unwanted cross-interactions with other sequences, as will be understood by a skilled person.

[0071] A polynucleotide in the sense of the disclosure can have a primary, secondary or a tertiary structure, as will be understood by a skilled person.

[0072] The term "primary structure" as used herein with reference to a polynucleotide indicates the specific linear sequence of nucleotide monomers linked together via the phosphodiester backbone, typically read in the 5' to 3' direction. This sequence determines the genetic or functional information encoded by the molecule and dictates the potential for higher-order folding based on the complementarity of the constituent bases.

[0073] The term "secondary structure" refers to the recurring local structural patterns formed by the intramolecular or intermolecular hydrogen bonding between nucleotide bases. In the context of single- stranded RNA, secondary structure is characterized by the formation of double-helical regions known as stems, where complementary sequences base-pair, employing both canonical Watson-Crick pairs and non-canonical wobble pairs, and single-stranded regions such as hairpin loops, internal loops, bulges, and junctions where base pairing is absent or disrupted. The stability of a given secondary structure is generally governed by thermodynamic principles, aiming to minimize the free energy of the system.

[0074] The term "tertiary structure" refers to the overall three-dimensional geometric shape of the polynucleotide, resulting from the folding and packing of secondary structural elements into a compact, globular, or extended conformation. Tertiary structure is stabilized by long-range" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT interactions, including coaxial stacking of helices, base triples, and metal ion interactions, and is critical for the biological function and molecular recognition properties of the RNA, such as the specific L-shaped conformation adopted by a mature tRNA molecule.

[0075] Various computational tools are available that can be used to design sequences configured to form stable complexes through specific and complementary binding of nucleic acids. In the context of the present disclosure, particular emphasis is placed on software capable of modeling thermodynamic equilibria in multi-component systems. Exemplary software for designing stable nucleic acid duplexes comprises NUPACK (nupack.org), [2] which is utilized in embodiments described herein to perform structure prediction, stability analysis, and complex thermodynamic calculations in a test tube mixture setting. Additional tools that provide essential features for calculating thermodynamic parameters and predicting secondary structures include OligoAnalyzer (idtdna.com / calc / analyzer), DNAforge (dnaforge.org) which enables automated sequence design with molecular dynamics simulation capabilities, and PFRED, available through GitHub, which offers open-source solutions for oligonucleotide design with emphasis on stability criteria. Furthermore, the computational frameworks described herein may integrate these physics-based engines with scientific computing libraries, such as Pandas and NumPy, to process large sequence libraries and perform the molecular weight calculations, melting temperature predictions, secondary structure evaluations, hybridization predictions, and cross-reactivity analyses required to manufacture stable, long-lasting complexes.

[0076] The present disclosure relates to computer-implemented methods and systems for the design of nucleic acids targeting structured single- stranded RNAs and their derivatives.

[0077] The term “structured single-stranded RNA” or “structured ssRNA” as used herein refers to a ribonucleic acid molecule, or a region thereof, which, despite being composed of a single polynucleotide chain, exhibits a high propensity to fold upon itself to form stable secondary and tertiary structures through intramolecular base pairing. Unlike unstructured or "random coil" RNA regions, structured ssRNAs are characterized by defined geometric conformations maintained by thermodynamic equilibria. These structures typically comprise recurring motifs including, but not limited to, hairpin loops, stem-loops, internal loops, bulges, pseudoknots, and multi-way junctions. These defined structural states provide the antecedent basis for the physical parameters evaluated" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT by the computational algorithms described herein, wherein the secondary configuration is analyzed to calculate nucleotide accessibility scores and the tertiary structure is utilized to model molecular binding interactions.

[0078] Accordingly, the structured single- stranded RNA targeted by the methods herein may comprise various classes of functional and regulatory RNAs. In addition to tRNA, exemplary structured ssRNAs include, but are not limited to: ribosomal RNA (rRNA); small nuclear RNA (snRNA); small nucleolar RNA (snoRNA); microRNA precursors (pre-miRNA); long non-coding RNAs (IncRNAs) (which often contain modular structural domains); ribozymes; riboswitches; and viral RNA genomes or fragments thereof (e.g., HIV TAR or RRE elements). Furthermore, the term encompasses structured regions within messenger RNA (mRNA), such as the 5' untranslated region (5' UTR), the 3' untranslated region (3' UTR), and Internal Ribosome Entry Sites (IRES), which adopt complex folds to regulate translation and stability.

[0079] In some embodiments herein described, the structured single- stranded RNA targeted by the methods and systems described herein can comprise various classes of functional and regulatory RNAs beyond transfer RNA. In various embodiments, the target is a non-coding RNA or a regulatory element within a coding RNA that relies on specific folding for biological function. Exemplary structured ssRNAs comprise: ribosomal RNA (rRNA) and small nuclear RNA (snRNA), which form the catalytic cores of the ribosome and spliceosome, respectively; microRNA precursors (pre-miRNA), which adopt characteristic hairpin structures; and long noncoding RNAs (IncRNAs) (e.g., MALAT1 or HOTAIR), which contain modular structural domains for chromatin remodeling or protein scaffolding. Additional targets include riboswitches and ribozymes, which undergo specific conformational changes to regulate gene expression. Furthermore, the structured target may comprise viral RNA elements, such as Internal Ribosome Entry Sites (IRES), viral tRNA-like structures (TLS), or the HIV Trans-Activation Response (TAR) element, as well as structured regions within messenger RNA (mRNA) untranslated regions (UTRs), all of which present defined thermodynamic stability profiles suitable for the computational modeling described herein as will be understood by a skilled person upon reading of the present disclosure.

[0080] The structured single-stranded RNA molecules suitable for targeting and modulation by" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT the disclosed methods may vary significantly in length, provided they possess sufficient sequence length to form stable secondary structures (e.g., stems and loops) capable of being modeled thermodynamically. In various embodiments, the structured parent nucleic acid comprises a nucleotide sequence length ranging from about 30 nucleotides to about 10,000 nucleotides. In preferred embodiments where the target is a distinct regulatory RNA, such as a tRNA, snRNA, or pre-miRNA, the length typically ranges from about 50 nucleotides to about 200 nucleotides. In embodiments targeting larger molecules, such as long non-coding RNAs (IncRNAs) or messenger RNAs (mRNAs), the system may process the full-length sequence with limited accuracy or, more preferably, focus on specific structured domains or 'folding windows' having a length of approximately 50 to 200 nucleotides, which serve as the effective thermodynamic unit for the accessibility and binding energy calculations described herein.

[0081] In preferred embodiments, the structured single-stranded RNA is a transfer RNA. The term "transfer RNA" or "tRNA" as used herein refers to a constructive non-coding RNA molecule that serves as a fundamental component of the cellular protein synthesis machinery across all domains of life. Functionally, tRNA acts as an adaptor molecule, facilitating the translation of genetic information encoded in messenger RNA (mRNA) into amino acid sequences during polypeptide chain formation. Structurally, while specific nucleotide sequences and lengths may vary among different species — typically ranging from 73 to 93 nucleotides — tRNAs share a highly conserved "secondary configuration" and a stable "tertiary structure" regardless of the organism of origin. This conserved secondary configuration generally adopts a "cloverleaf" pattern composed of basepaired stems and single- stranded loops, including the D-loop. T-loop, and variable loop. This configuration further folds into the tertiary structure, which generally presents as a compact L-shaped three-dimensional conformation. These defined structural states provide the antecedent basis for the physical parameters evaluated by the computational algorithms described herein, wherein the secondary configuration is analyzed to calculate nucleotide accessibility scores and the tertiary structure is utilized to model molecular binding interactions.

[0082] Accordingly, the term "tRNA" encompasses molecules derived from any organism, including but not limited to humans, non-human mammals, plants, fungi, bacteria, archaea, and viruses, as well as specific variants found within different individuals of a species. With respect to viral origins, the term expressly includes both viral-encoded tRNA molecules, utilized by certain" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT large DNA viruses and bacteriophages to supplement host translation, and viral tRNA-like structures (TLS) found in the genomes of RNA viruses which adopt the characteristic cloverleaf secondary configuration and L-shaped tertiary structure to mimic host tRNAs and interact with cellular machinery. In preferred embodiments, however, the target tRNA is a mammalian tRNA. Mammalian tRNAs, typically ranging in length from 76 to 90 nucleotides, exhibit highly predictable secondary structures supported by extensive available data at the tertiary level, making them particularly suitable substrates for the precise accessibility calculations and cleavage site predictions described in the present disclosure.

[0013]

[0083] While the detailed description and examples provided herein utilize transfer RNA (tRNA) as a representative substrate, it will be appreciated that the computational frameworks disclosed (including the tBOND-G and tBOND-L algorithms) are broadly applicable to the design of nucleic acids targeting any structured single-stranded RNA (ssRNA). The core physical parameters utilized by the algorithms — specifically thermodynamic binding energy and structural accessibility derived from secondary folding — are universal biophysical properties of RNA, independent of the molecule's specific biological classification. Accordingly, in various embodiments, the target RNA molecule is selected from the group consisting of messenger RNA (mRNA) (particularly structured regions such as 5' or 3' UTRs), long non-coding RNA (IncRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), and viral RNA genomes. In specific embodiments, the target is a viral tRNA-like structure (TLS) or a viral regulatory element that mimics the secondary structure of a host RNA to hijack cellular machinery. The system is configured to receive these non-tRNA sequences as inputs, determine their secondary structure using the folding algorithms described herein (e.g., NUPACK), and identify accessible cleavage sites or binding motifs in the same manner as described for tRNA.

[0084] Furthermore, the structured single-stranded RNAs described herein may be advantageously fragmented. In the context of the present disclosure, a "fragment", herein indicated also as a derivative, refers to any portion of the full-length RNA sequence (herein referred to as the "parent nucleic acid") that retains a defined secondary structure. Accordingly, the specific sequence region within the parent nucleic acid that shares identity with the fragment is referred to as the "corresponding segment." In some embodiments, these fragments possess functional significance (e.g., as regulatory molecules or catalytic cores), while in other embodiments, they may simply" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT represent stable structural intermediates. The computational methods described herein are particularly configured to address fragmentation in two advantageous contexts: (1) Therapeutic Cleavage, wherein the designed nucleic acid agent induces the specific fragmentation of the parent nucleic acid (e.g., via RNase H recruitment or RNAi mechanisms) to degrade the structured RNA; and (2) Targeting of Stable Fragments, wherein the system designs agents to bind naturally occurring, stable fragments (such as tRNA-derived fragments or tRFs) which may possess independent regulatory roles. Additionally, the term “derivative” further encompasses RNA molecules subjected to post-transcriptional modification (e.g., methylation, pseudouridylation) or chemical alteration. As these modifications can stabilize specific fragment conformations or alter thermodynamic accessibility, the algorithms herein are configured to account for such derivatives to ensure precise targeting of any structured RNA moiety.

[0085] To further contextualize the computational challenges addressed by the present disclosure, the following section details the properties of Transfer RNA (tRNA). As a molecule defined by its rigid, thermodynamically stable secondary and tertiary structures, tRNA serves as the primary exemplary substrate for the methods described herein, representing the archetypal structured target against which the efficacy of the present design algorithms is demonstrated.

[0086] Transfer RNA (tRNA) is a molecule in biology, well-known for its canonical role in protein synthesis, where it acts as an adaptor molecule that decodes the genetic information in messenger RNA (mRNA) and transfers the corresponding amino acid to a growing polypeptide chain. For decades, this was considered its primary function.

[0087] However, recent discoveries have revealed that tRNAs are also a source of a diverse class of small non-coding RNAs. produced when the full-length tRNA molecule is cleaved by specific enzymes (ribonucleases) like Angiogenin or DICER. These smaller molecules, herein indicated as derivatives, are known as tRNA-derived fragments (tRFs) or tRNA-derived small RNAs (tDRs). Far from being random degradation products, tDRs are now recognized as critical regulatory molecules involved in a wide array of cellular processes, including RNA silencing, translation regulation, stress granule formation, and epigenetic inheritance. Their dysregulation has been implicated in numerous human diseases, including various cancers, metabolic disorders, and neurologic conditions." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0088] Applications requiring study and manipulation of tDRs, however, face two technical hurdles: i) Studying Endogenous tDRs: Gain-of-function studies relying on introducing synthetic tDR molecules into cells may not accurately replicate the function of naturally produced tDRs, which often possess specific chemical modifications that are critical for their activity; and ii) Therapeutic Targeting of tDRs: For diseases driven by the over-expression of a specific tDR, a promising therapeutic strategy is to inhibit or degrade that tDR. The challenge lies in the high degree of sequence homology between a tDR and its parent tRNA. A therapeutic agent, such as an antisense oligonucleotide, must be designed with high specificity to target only the pathogenic tDR without disrupting the essential pool of the full-length parent tRNA, which is vital for protein synthesis. Achieving this level of specificity is a significant design challenge.

[0089] The present disclosure provides a rational, predictive, and efficient framework to overcome these challenges and unlock the full research, diagnostic, engineering, and therapeutic application of the tRNA / tDR system encompassing tRNA and all its derivatives.

[0090] As used herein, the term "derivative of a nucleic acid" refers to a polynucleotide molecule that is structurally distinct from but sequence-related to a larger precursor molecule, referred to herein as the "parent nucleic acid." Derivatives encompass both (i) fragments (wherein the derivative consists of a nucleotide sequence that aligns with a specific subsequence or region within the full-length parent nucleic acid, termed the "corresponding segment"), and (ii) modified variants (wherein the derivative retains the sequence of the parent or segment thereof but comprises post-transcriptional chemical modifications).

[0091] The term "parent nucleic acid" is defined as the full-length, naturally occurring or synthetic precursor molecule, such as a mature transfer RNA (tRNA), messenger RNA (mRNA), or viral RNA genome, which serves as the substrate for the generation of the smaller derivative molecule through biological or synthetic processing events, including but not limited to enzymatic cleavage or ribonuclease processing.

[0092] The term "corresponding segment" refers specifically to the continuous string of nucleotides within the parent nucleic acid structure that is homologous to the sequence of the derivative." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0093] As understood by those skilled in the art, determination of percentage of similarity between any two sequences can be accomplished using a mathematical algorithm. Non-limiting examples of such mathematical algorithms are the algorithm of Myers and Miller [7] [Ref Myers, E. W. and W. Miller, Optimal alignments in linear space. Computer applications in the biosciences: CAB IOS. 1988. 4(1): p. 11-17.], the local homology algorithm of Smith et al.

[0014] [Smith, T. F. and M. S. Waterman, Comparison of biosequences. Advances in applied mathematics, 1981. 2(4): p.482-489];the homology alignment algorithm of Needleman and Wunsch [8] [Ref Needleman, S. B. and C. D. Wunsch, A general method applicable to the search for similarities in the amino acid sequence of two proteins. Journal of molecular biology, 1970. 48(3): p. 443-453]; the search-for-similarity-method of Pearson and Lipman

[0010] [Ref Pearson, W. R. and D. J. Lipman, Improved tools for biological sequence comparison. Proceedings of the National Academy of Sciences, 1988.85(8): p. 2444-2448.]; the algorithm of Karlin and Altschul [5] [Ref Karlin, S. and S. F. Altschul, Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes. Proceedings of the National Academy of Sciences, 1990. 87(6); p. 2264-2268.], modified as in Karlin and Altschul [6] [Ref Karlin, S. and S. F. Altschul, Applications and statistics for multiple high-scoring segments in molecular sequences. Proceedings of the National Academy of Sciences, 1993. 90(12): p. 5873-5877.]. Computer implementations of these mathematical algorithms can be utilized for comparison of sequences to determine sequence identity. Such implementations include, but are not limited to: CLUSTAL in the PC / Gene program (available from Intelligenetics, Mountain View, Calif.); the ALIGN program (Version 2.0) and GAP, BESTFIT, BLAST, FASTA

[0010] [Ref Pearson, W. R. and D. J. Lipman, Improved tools for biological sequence comparison. Proceedings of the National Academy of Sciences. 1988. 85(8): p. 2444-2448.] and TFASTA in the Wisconsin Genetics Software Package, Version 8 (available from Genetics Computer Group (GCG), 575 Science Drive, Madison, Wis., USA). Alignments using these programs can be performed using the default parameters.

[0094] To determine if the candidate constitutes a derivative, the percentage of sequence identity is calculated over the full length of the candidate sequence relative to the identified corresponding segment of the parent. A segment is confirmed as a derivative if this calculated identity meets the specified threshold (e.g., at least 90%, 95%, or 100%), indicating that the molecule originated from or corresponds to that specific locus within the parent structure." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0095] In particular, a nucleic acid is classified as a derivative of a parent nucleic acid if it exhibits a sequence identity of at least 90%, preferably at least 95%, and most preferably 100% to its corresponding segment located within the parent molecule, or if it exhibits the requisite sequence identity but differs via one or more chemical modifications (e.g., methylation, pseudouridylation). This quantitative threshold reflects the derivative's origin as a physical fragment or variant of the parent, ensuring that the derivative retains the specific genetic or structural information of the source region. In the specific context of the embodiments described herein, the parent nucleic acid is a full-length tRNA and the derivative is a tRNA-derived fragment (tRF) or tRNA-derived small RNA (tDR), which is produced by site-specific cleavage of the parent tRNA and functions as an independent regulatory molecule distinct from the full-length precursor.

[0096] Accordingly, as used herein, the term "derivative of tRNA" refers to a small non-coding RNA molecule that is structurally derived from a full-length parent transfer RNA (tRNA) molecule through a process of specific, regimented cleavage. These derivatives include molecules expressly termed "tRNA-derived fragments" (tRFs) and "tRNA-derived small RNAs" (tDRs).

[0097] Structurally, a tRNA derivative consists of a specific subsequence of the parent tRNA, exhibiting at least 90%, preferably at least 95%, and most preferably 100% sequence identity to the corresponding region (e.g., the 5’ end or 3' end) of the full-length precursor. Biologically, these derivatives are produced by the stereotypical cleavage of the full-length tRNA by specific stress-activated ribonucleases, including but not limited to Angiogenin (ANG), DICER, ELAC2, and RNASE 1, which may be followed by helicase-dependent unwinding to produce the functional small RNA.

[0098] Accordingly, the term tRNA derivatives encompasses various specific classes of fragments, such as "tRNA halves" produced by cleavage in the anticodon loop, and smaller fragments generated from the 5' or 3' ends, all of which function as distinct regulatory molecules in cellular processes rather than random degradation products. tRNA derivatives are categorized into distinct classes based on the specific cleavage site within the full-length parent molecule. These classes include "tRNA halves" (also known as tiRNAs), which are produced by specific cleavage in the anticodon loop of the mature tRNA, typically by the ribonuclease Angiogenin. Structurally, these halves correspond to the 5' or 3' portion of the tRNA and are approximately 30-" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT 35 nucleotides in length. tRNA derivatives also comprise smaller "tRNA-derived fragments" (tRFs), which are typically 14-30 nucleotides in length and are generated by cleavages at the 5' end (5'-tRFs) or 3' end (3'-tRFs) of the mature tRNA or its precursor, often mediated by enzymes such as Dicer or Angiogenin. The term "tRNA-derived small RNA" (tDR) is utilized herein as a comprehensive umbrella term encompassing both tRNA halves, tRFs, and other functional small RNAs resulting from the regimented processing of parent tRNAs. These tDRs are distinguished from random degradation products by their precise 5' and 3' termini, stable secondary structures, and specific regulatory roles in cellular physiology as will be understood by a skilled person.

[0099] The present disclosure describes a comprehensive computational framework designed to facilitate the high-precision modulation of RNA regulatory networks. As used herein, the term "cleavage" refers to the specific enzymatic hydrolysis of the phosphodiester backbone of a nucleic acid molecule, resulting in the fragmentation of a full-length precursor sequence into smaller, distinct oligonucleotide products. This process is catalyzed by "nucleases," a class of enzymes capable of cleaving the phosphodiester bonds between the nucleotide subunits of nucleic acids. The embodiments described herein facilitate cleavage via multiple mechanisms, including the recruitment of endogenous nucleases (e.g., RNase H recruitment by antisense oligonucleotides) and the utilization of engineered RNA-guided nucleases, specifically those of the CRISPR-associated (Cas) system, to perform programmable cleavage.

[0100] As used herein, the term 'targeting moiety' refers broadly to any oligonucleotide or nucleic acid analogue capable of hybridizing to a specific genomic or transcriptomic sequence to direct an effector function to that site. While in preferred embodiments described herein the targeting moiety is a guide RNA (gRNA) utilized by a Cas nuclease, the term encompasses other guide sequences utilized by diverse RNA-targeting systems. These include, but are not limited to. short interfering RNAs (siRNAs) or short hairpin RNAs (shRNAs) that guide the RISC complex, gapmer antisense oligonucleotides that guide RNase H, and crRNAs utilized by Class 1 CRISPR systems (e.g., Cas6, Csy4). In all such embodiments, the tBOND-G algorithm described herein is applicable for optimizing the sequence of the targeting moiety based on the accessibility of the target site and the thermodynamic stability of the resulting [Targeting Moiety: Target] duplex.

[0101] In embodiments utilizing CRISPR, the term " CRISPR-Cas system" as used herein refers" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT to a class of adaptive immune systems found in bacteria and archaea, characterized by the use of CRISPR-associated (Cas) nucleases to detect and cleave foreign genetic material. In the context of the computer-implemented methods described herein, the Cas system is modeled as a programmable ribonucleoprotein (RNP) complex comprising two primary functional components: a protein effector (the Cas nuclease) and a nucleic acid guide (the guide RNA or gRNA).

[0102] As used in the present disclosure, gRNA (guide RNA) is a short synthetic RNA molecule composed of a "spacer" region that is complementary to a target nucleic acid sequence and a "scaffold" region that binds to a Cas protein. The gRNA "guides" the Cas protein to the target. The guide RNA utilized in CRISPR-Cas systems is generally composed of two distinct functional regions: a "scaffold" sequence and a "spacer" sequence. The scaffold is a conserved structural domain that adopts a specific secondary configuration, typically involving stem-loops, which is recognized and bound by the Cas protein, thereby anchoring the RNA to the effector nuclease. The spacer is a variable sequence that serves as the targeting component, designed to be complementary to a specific "protospacer" sequence on the target nucleic acid. Functionally, the Cas system relies on the thermodynamic hybridization between this spacer and the target; the ribonucleoprotein complex scans potential targets and, upon finding a complementary match, undergoes a conformational change that activates the nuclease domains, resulting in the enzymatic cleavage of the target substrate.

[0103] In such embodiments, the Cas nuclease serves as the enzymatic machinery responsible for the cleavage of the phosphodiester bonds in the target nucleic acid, while the gRNA provides the sequence- specific instruction that directs the nuclease to the precise target site. In this context, "nucleases" are defined as a broad class of enzymes that catalyze the hydrolysis of the phosphodiester backbone of nucleic acids, effectively cutting DNA or RNA into smaller fragments. While this enzymatic activity is fundamental to many biotechnological applications, including gene editing and RNA processing, the practical deployment of nucleases often suffers from a lack of sufficient specificity. Despite the sequence guidance provided by a guide RNA (gRNA) or other targeting moiety, off-target cleavage events — where the nuclease cuts unintended genomic or transcriptomic sites with similar sequences — remain a persistent challenge, particularly in complex cellular environments where high sequence homology exists between the target and essential non-target molecules. The computational framework of the present disclosure" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT is specifically intended to address this limitation by integrating thermodynamic modeling and structural accessibility data to predict and optimize the specificity of these nuclease-mediated events before experimental application.

[0104] In particular, the present disclosure provides a computational framework composed of two core algorithms, tBOND-G and tBOND-L, which are implemented as methods on a computer system to provide — in addition to or in place of inefficient, trial-and-error experimental approaches — a predictive, physics-based design process for creating nucleic acid tools to modulate structured single- stranded RNAs and their derivatives.

[0105] A common thread between both algorithms is their rooting in biophysical principles. Rather than treating nucleic acid sequences as abstract strings of letters, the embodiments of the present disclosure model their real-world physical properties. Both algorithms adopt thermodynamic calculations to predict how molecules will fold (secondary structure) and how strongly they will bind to each other (binding energy). This physics-based approach provides a more generalizable predictive power than purely sequence-based or statistical methods, allowing for the creation of highly specific nucleic acid tools to modulate structured RNA targets.

[0106] In embodiments utilizing the tBOND-G algorithm (directed to RNA-guided nucleases), central to the computational design is the structure of the guide RNA, which is generally composed of two distinct functional regions: a "scaffold" sequence and a "spacer" sequence. The scaffold is a conserved structural domain that adopts a specific secondary configuration (typically involving stem-loops) recognized and bound by the Cas protein, thereby anchoring the RNA to the effector. The spacer is a variable sequence, customized by the user or the algorithm, which is complementary to the "protospacer" sequence on the target nucleic acid. The functioning of the Cas system relies on the thermodynamic hybridization between this spacer and the target; the complex scans potential targets and, upon finding a complementary match, undergoes a conformational change that activates the nuclease domains (such as the HEPN domains in Cas 13), resulting in cleavage.

[0107] Accordingly, the computer-implemented methods of the present disclosure rely on the structural and functional similarities shared among various RNA-targeting Cas proteins to generalize results obtained for specific variants, such as pspCas13b. While specific Cas variants" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT (e.g., Cas13a, Cas13b, Cas13c, Cas13d, Cas13x, Cas13y) may differ in their primary amino acid sequence or the specific nucleotide sequence of their gRNA scaffolds, they share a unified mechanism of action dependent on two physical parameters modeled by the tBOND-G algorithm: the thermodynamic stability of the gRNA-target duplex and the structural accessibility of the target site. Because all Cas nucleases require the gRNA spacer to physically access and hybridize with the target RNA to initiate catalysis, the computational logic of prioritizing accessible "breathing" loops and maximizing binding energy is universally applicable across the family. Consequently, the system allows for the substitution of different Cas variants simply by updating the scaffold sequence definition in the library generation step, while the core predictive physics regarding target recognition remain constant.

[0108] Exemplary CRISPR-associated (Cas) nucleases suitable for use in the methods and systems described herein include members of the Cas 13 family, which are distinct from DNA-targeting enzymes in their specific ability to target and cleave single-stranded RNA (ssRNA) substrates. In preferred embodiments, the Cas nuclease is selected from the group consisting of Cas13a (also known as C2c2), Cas13b, Cas13c, Cas13d (including specific orthologs such as RfxCas13d), Cas13x, and Cas13y. The disclosure specifically contemplates the use of the pspCas13b ortholog derived from Prevotella sp. P5-125 as a representative effector. Furthermore, the term " Cas nuclease" as used herein broadly encompasses any naturally occurring orthologs, engineered variants, or functional equivalents of these enzymes — such as high-fidelity or catalytically enhanced mutants — provided they retain the programmable, RNA-guided ribonuclease activity capable of recognizing a specific guide RNA scaffold and cleaving a target RNA sequence based on the thermodynamic accessibility principles described by the tBOND-G algorithm.

[0109] In some embodiments, the computer-implemented methods and systems of the disclosure are based on a tBOND-G algorithm. The tBOND-G algorithm is a tool designed to solve the problem of reliably generating specific fragments (e.g., tRFs or other derived small RNAs) from structured single-stranded RNA targets (such as tRNA). It is implemented as a computer-executable workflow that takes a target RNA sequence (e.g., a tRNA) and desired targeting moiety length (e.g., gRNA length) as input. The algorithm uses a machine learning model (e.g., an SVM) that is trained not just on sequence data, but on a combination of calculated physical parameters" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT and actual experimental results.

[0110] As used herein, the term "machine learning model" refers to a computational algorithm or statistical framework configured to identify patterns within a dataset and make predictions on new data without being explicitly programmed for each specific instance. In the context of the present disclosure, the model functions as a supervised learning system that maps a set of input variables, known as features, to a predicted output variable. Specifically, the input features comprise the calculated biophysical parameters — such as thermodynamic binding energy ratios and structural accessibility scores — while the output variable corresponds to a quantitative measure of biological activity, such as cleavage efficiency or binding affinity. During a training phase, the model processes a dataset of experimentally verified performances to establish a mathematical relationship, such as a decision boundary or hyperplane, that differentiates between high-performing and low-performing candidates. A machine learning model encompasses Support Vector Machine (SVM) and additional predictive architectures such as Artificial Neural Networks (ANNs), Random Forests, Decision Trees, and Gradient Boosting algorithms capable of regression or classification based on biophysical feature inputs, as will be understood by a skilled person.

[0111] As used herein, the term " Support Vector Machine" or " SVM" refers to a supervised machine learning algorithm that analyzes data for classification and regression analysis. In the context of the present disclosure, the SVM operates by constructing a hyperplane or set of hyperplanes in a high-dimensional space to separate data points into distinct classes. When used as a classifier for targeting moiety design (e.g., gRNA), the SVM distinguishes between high-efficiency and low-efficiency candidates by finding an optimal decision boundary that maximizes the margin between these classes based on input feature vectors comprising biophysical parameters such as structural accessibility and thermodynamic binding energy. In preferred embodiments, the SVM utilizes a kernel function, such as a Radial Basis Function (RBF), to efficiently map inputs into high-dimensional feature spaces, thereby allowing for the separation of complex, non-linear datasets.

[0015]

[0112] As used herein, the term " Artificial Neural Network" or " ANN" refers to a computational model inspired by the biological neural networks of animal brains. An ANN comprises a collection of connected units or nodes called artificial neurons, which are typically organized into layers" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT including an input layer, one or more hidden layers, and an output layer. Each connection between neurons can transmit a signal, and the receiving neuron processes the signal and signals downstream neurons connected to it. In the context of the disclosed methods, an ANN can be trained to learn the non-linear mapping between the input biophysical parameters of a nucleic acid sequence and its resulting biological activity, adjusting the weights of the connections during a training phase to minimize prediction error.

[0113] \As used herein, the term " Decision Tree" refers to a non-parametric supervised learning method used for classification and regression. The model predicts the value of a target variable by learning simple decision rules inferred from the data features. Structurally, it represents a flowchart-like tree structure where an internal node represents a test on a feature or attribute (e.g., whether an accessibility score exceeds a certain threshold), each branch represents the outcome of the test, and each leaf node represents a class label or decision. In the present framework, a decision tree may be utilized to hierarchically split the library of candidate sequences based on biophysical metrics to arrive at a classification of cleavage efficiency.

[0114] As used herein, the term " Random Forest" refers to an ensemble learning method that operates by constructing a multitude of decision trees at training time. For classification tasks, the output of the random forest is the class selected by the majority of trees (the mode); for regression tasks, it is the mean prediction of the individual trees. This method corrects for the habit of individual decision trees to overfit to their training set. Within the scope of the present disclosure, a random forest algorithm aggregates the predictions of multiple decision trees derived from random subsets of the training data to provide a robust assessment of gRNA or complementary oligonucleotide performance metrics.

[0115] As used herein, the term " Gradient Boosting" refers to a machine learning technique for regression and classification problems that produces a prediction model in the form of an ensemble of weak prediction models, typically decision trees. Unlike random forests which build trees in parallel, gradient boosting builds the model in a stage-wise fashion, adding new models to correct the errors made by existing models. This technique optimizes a differentiable loss function and is particularly effective for handling the complex, non-linear relationships between the multiple biophysical features utilized in the algorithms described herein." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0116] In computer-implemented methods of the disclosure comprising the tBOND-G algorithm, the computer system first calculates two features for every possible candidate: Accessibility and Binding Energy.

[0117] As used herein, the term " Accessibility" refers to a quantitative measure representing the degree to which a specific nucleotide sequence within a folded RNA molecule is physically exposed and available for binding to another molecule. It is calculated as the proportion or percentage of unstructured bases (i.e., bases not involved in base-pairing) within a specific target region relative to the total number of bases in that region. This metric is derived from the secondary structure of the RNA, where regions such as loops are considered accessible, while doublestranded stems are considered inaccessible.

[0118] As used herein, the term " Binding Energy" (often represented as a Delta G Ratio) refers to a normalized measure of the change in Gibbs free energy that occurs when a nucleic acid strand, such as a gRNA, hybridizes to its target sequence. It quantifies the thermodynamic stability of the resulting complex formed between the two molecules. In the computational methods described herein, the binding energy is typically calculated by simulating the reaction between reactants (e.g., gRNA and target RNA) to form a product complex and is normalized against a theoretical maximum binding energy (such as that of a continuous GC-rich strand) to provide a standardized ratio for comparison across different sequences.

[0119] In embodiments of the disclosure based on tBOND-G, Accessibility and Binding Energy capture the two requirements for a successful targeting moiety: the target site must be physically reachable, and the moiety must bind to it stably. The machine learning model (e.g. SVM model) learns the complex, non-linear relationship between these two parameters and the likelihood of successful cleavage or binding.

[0120] In certain embodiments, the secondary structure of the structured single-stranded RNA utilized for these calculations is known and characterized. For example, for highly conserved targets such as transfer RNA (tRNA) or ribosomal RNA (rRNA), the system is configured to retrieve the determined structure (e.g., in dot-bracket notation) directly from an external database such as RNA Central or GtRNAdb [1]

[0016] . However, in other embodiments — such as when targeting specific messenger RNA (mRNA) untranslated regions (UTRs), long non-coding RNAs" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT (IncRNAs) with unresolved folding patterns, or rapidly evolving viral RNA genomes — the precise secondary structure may not be known a priori. In such instances, the step of providing the parent sequence further comprises a step of determining the secondary structure prior to the execution of the accessibility algorithms. This structural determination is performed by the processor utilizing thermodynamic folding simulations (e.g., calculating the Minimum Free Energy (MFE) structure via algorithms such as RNAfold or NUPACK) [2] or by incorporating experimental probing data (e.g., SHAPE-MaP reactivity profiles) to constrain the folding prediction, thereby ensuring that the subsequent accessibility scores reflect the most probable physiological conformation of the target.

[0121] In particular, to determine Accessibility, the system retrieves secondary structure data (e.g., dot-bracket notation) for the target RNA (e.g., tRNA) and quantifies the proportion of unstructured bases within the specific target window, thereby assessing the physical exposure of the site for effector binding (e.g., Cas or RNase H recruitment). Simultaneously, the system calculates Binding Energy by employing a nucleic acid thermodynamics package (e.g., NUPACK) to simulate the hybridization of the targeting sequence (e.g., gRNA spacer) to the target RNA, outputting a normalized " Delta G Ratio" that reflects the thermodynamic stability of the resulting complex relative to a theoretical maximum.

[0122] The final output is not merely a sequence, but a ranked list of candidates with a predicted efficiency score and a virtual gel that simulates the experimental outcome, thereby providing a complete and actionable guide for the researcher.

[0123] These calculated features are then input into a machine learning model configured to execute a supervised learning algorithm. Regardless of the specific topology, the model operates by establishing a mathematical mapping — such as a non-linear decision boundary or a regression function — within a high-dimensional feature space. This mapping represents the learned complex relationship between the input physical parameters and the actual cleavage outcomes derived from an experimental training set. The final output extends beyond a simple sequence list; it provides a ranked dataset of candidates with predicted cleavage efficiency scores and generates a "virtual gel" simulation. This simulation utilizes molecular dynamics modeling of the nuclease's (e.g., Cas protein) binding pockets to predict precise 5' and 3' fragment sizes, thereby providing a complete" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT and actionable experimental guide.

[0124] In some embodiments, the machine learning model is a Support Vector Machine (SVM) utilizing a Radial Basis Function (RBF) kernel. The model constructs a non-linear decision boundary in high-dimensional space, having "learned" the complex relationship between these physical parameters and actual cleavage outcomes from an experimental training set.

[0125] In preferred implementations of the machine learning model, specifically when utilizing a Support Vector Machine (SVM), the system employs a Radial Basis Function (RBF) kernel. The RBF kernel is particularly advantageous for this application as it maps the non-linear relationship between the thermodynamic stability (Delta G) and the structural accessibility scores into a higherdimensional feature space, allowing for the precise definition of the decision boundary between functional and non-functional targeting sequences.

[0126] The final output extends beyond a simple sequence list; it provides a ranked dataset of candidates with predicted cleavage efficiency scores and generates a "virtual gel" simulation. This simulation utilizes molecular dynamics modeling of the effector's (e.g., Cas protein) binding pockets to predict precise 5’ and 3’ fragment sizes, thereby providing a complete and actionable experimental guide.

[0127] In other embodiments, the model can alternatively comprise an Artificial Neural Network (ANN), a Random Forest, a Gradient Boosting machine, or a Deep Learning architecture as will be understood by a skilled person upon reading of the present disclosure.

[0128] In some embodiments, after identifying high-efficiency candidates (e.g., gRNA) via the Support Vector Machine (SVM) model, the tBOND-G system executes an automated off-target analysis module to validate the specificity of the selected sequences. This secondary screening step is critical for mitigating potential off-target effects, such as collateral cleavage or non-specific knockdown of essential transcripts, which could confound experimental data or induce cytotoxicity in a cellular environment. In operation, the processor is configured to extract the targeting sequence (e.g., spacer) of a high-ranking candidate and perform a comprehensive homology search, utilizing alignment tools such as a BLAST-like algorithm or Bowtie, against a pre-defined, curated sequence database. While this database typically comprises a reference" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT transcriptome (e.g., RefSeq or Ensembl for human mRNA and non-coding RNA) to identify relevant RNA transcripts that may be inadvertently targeted, it may in some embodiments also include genomic sequences to predict potential interactions with nascent pre-mRNA transcripts.

[0129] The system computationally identifies all potential off-target loci that exhibit a sequence similarity exceeding a predetermined threshold relative to the targeting sequence. For each identified potential off-target, the system calculates a quantitative off-target score. This score is computed using a weighted algorithm that accounts for both the total number of nucleotide mismatches and their specific spatial distribution along the duplex. Mismatches located within a critical "seed" region of the targeting sequence — typically a segment essential for nucleating the hybridization and activating the conformational change of the effector nuclease — are penalized more heavily than mismatches in less critical peripheral regions. Based on this aggregate scoring, the system applies a filtering logic to the candidate library, strictly prioritizing sequences that possess a dual characteristic: a high predicted on-target cleavage efficiency score and a low off-target potential score.

[0130] This prioritization utilizes a predetermined safety threshold for the off-target score. In some embodiments, this safety threshold is a numerical value representing the maximum allowable sequence identity or thermodynamic affinity to a non-target transcript. If a candidate possesses an off-target score exceeding this safety threshold, it is automatically rejected from the ranked list, regardless of its predicted on-target efficiency, to ensure experimental safety and specificity.

[0131] This filtration ensures that the final output nucleic acid is indeed highly specific, directing the cleavage mechanism (e.g., Cas enzyme) exclusively to the intended target while minimizing thermodynamic affinity for unintended transcripts with partial homology. Furthermore, while the computational frameworks described herein provide high-confidence predictions of cleavage efficiency and binding specificity, in all embodiments, the method may further comprise an optional step of validating the computational output through wet-bench experimentation. Following the in-silico generation of the ranked candidate list or the virtual gel simulation, the user or the system may select a subset of top-ranking candidates for physical synthesis and biological testing. Exemplary validation assays comprise, Northern blotting to visualize specific fragment generation (as simulated by the virtual gel), RT-qPCR to quantify target knockdown or fragment" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT induction, and Western blotting to assess downstream effects on protein translation. This experimental feedback not only verifies the efficacy of the specific designed reagents but can also be utilized to update the training dataset of the machine learning model, thereby refining the decision boundaries for future designs.

[0132] Further features of the tBOND-G algorithm are further illustrated in connections with the exemplary illustrations of FIG. 1 to FIG. 4.

[0133] FIG. 1 shows an exemplary workflow (100) of the tBOND-G algorithm as implemented on a computer system. The process starts with a user providing a target RNA sequence (e.g., a tRNA) (step 105). The system then performs two parallel tasks: (1) it generates a comprehensive targeting library (e.g., gRNA) by creating all possible subsequences of a specified length (110), and (2) it retrieves the known secondary structure of the target RNA from an external database like RNA Central [161(110). From these inputs, the system calculates the physical parameters of Accessibility (120) and Reaction Energy (125) (Delta G Ratio). These parameters, along with a set of known experimental data (130), are fed into a pre-trained model (135), e.g., an SVM model. The model outputs a Targeting Efficacy Prediction (e.g., Cleavage Efficiency) (140). The system then uses molecular modeling to predict the specific cleavage sites (145, 150) and generates a Virtual Gel Visualization (155), which simulates the expected experimental result, alongside an output targeting sequence (160).

[0134] FIG.2 shows the biophysical basis for the parameter calculations in FIG. 1. The top panel (200A) shows how the known, complex folded structure of a target RNA (205) is used to determine Accessibility. A target site on the RNA (205) is considered accessible if its bases are in singlestranded regions (loops) rather than base-paired regions (stems). The bottom panel (200B) illustrates the Binding Energy Calculation. The system models the gRNA / targeting moiety (210) and target RNA (205) as reactants and the complex (215) as the product. Using a thermodynamics package (e.g. NUPACK ®), it calculates the change in Gibbs free energy for this binding reaction and normalizes it to create the Delta G Ratio, a standardized measure of binding stability.

[0135] FIG. 3 details the machine learning component of the tBOND-G system. The left panel (300A) is a scatter plot where each point represents a gRNA from an experimental training set, plotted according to its calculated Delta G Ratio (x-axis) and Accessibility (y-axis). The size of" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT the dot corresponds to its measured experimental performance. The right panel (300B) shows the output of the SVM model trained on this data. The model has generated a non-linear decision boundary (the transition from a first type of region to a second type of region) that separates the feature space into zones of predicted high performance (first type) and low performance (second type). This trained model can now classify new, unseen gRNA candidates.

[0136] FIG.4 provides a structural rationale for the cleavage site prediction. It shows a 3D model (405) of the pspCas13b protein complexed with its gRNA and a target RNA strand. This model was generated using molecular dynamics simulations. The simulations identified two distinct binding pockets within the protein that interact with the gRNA at specific positions. The model hypothesizes that interaction at the first pocket (involving residues 592R and 583R) leads to cleavage that produces the 3' fragment (e.g., 3' tDR), while interaction at the second pocket (involving residue 812K) leads to cleavage that produces the 5' fragment (e.g., 5' tDR). This allows the system to predict the likely cleavage outcome.

[0137] Experimental validations of the tBOND-G systems were performed and are illustrated in Examples 1 to 4, which collectively demonstrate the predictive accuracy and robustness of the computational framework in a biological setting. To verify the output of the machine learning model, the system’s " Virtual Gel" visualizations — which simulate the expected size and migration of 5’ and 3’ fragments — were directly compared against physical Northern Blot results obtained from cellular RNA extracts. These validations were conducted using HEK293 cells expressing the pspCas13b nuclease, into which specific gRNA candidates selected by the algorithm were introduced. The examples are structured to rigorously test distinct performance clusters identified by the Support Vector Machine (SVM) decision boundary, ensuring that the model correctly weights the contributions of structural accessibility and thermodynamic stability.

[0138] Specifically, the validation strategy assessed guide RNAs selected from an "accessibilitydominant" cluster, where the target site is located within highly unstructured loops of the target secondary structure (see Example 1). This testing confirms the biophysical premise that Casl3 nucleases require physical access to single-stranded regions to initiate effective cleavage. Further validations examined guide RNAs from a "binding-energy-dominant" cluster, characterized by high Delta G ratios, to verify that strong thermodynamic stability can drive cleavage efficiency" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT even in regions of moderate structural complexity (see Example 2). By confirming activity in both clusters, the disclosure establishes that the algorithm successfully integrates multiparametric data to identify diverse pathways to successful cleavage.

[0139] In addition to verifying positive hits, the experimental validations included a "negative control" assessment using guide RNAs selected from a balanced cluster with moderate accessibility and energy scores, for which the model predicted low cleavage efficiency (Example 3). The agreement between the predicted lack of cleavage bands in the virtual gel and the absence of bands in the physical Northern Blot confirms the model’s high specificity and its ability to effectively screen out non-functional sequences, thereby reducing false positives. Additionally, the generalizability of the framework was validated by applying the algorithm to a distinct target sequence, tRNA-Asp-GTC, demonstrating that the physical principles encoded in the software are universally applicable across different structured RNA targets and are not overfitted to a specific training target (see Example 4).

[0140] In some embodiments, the computer-implemented methods and systems of the disclosure are based on a tBOND-L algorithm. The tBOND-L algorithm is a design tool engineered to solve the problem of specificity in targeting fragments of structured single-stranded RNAs (e.g., tDRs). Its main feature is a two-parameter scoring system that simultaneously optimizes for both binding strength (Efficiency) and target selectivity (Specificity).

[0141] In this context, " Efficiency" is defined as a calculated metric assessing the thermodynamic binding affinity of the candidate nucleic acid (e.g., ASO) to its intended target sequence relative to the energy cost of disrupting secondary structures within the target. It essentially quantifies the capability of the candidate to invade the structured RNA and form a stable duplex.

[0142] In this context, " Specificity" as used herein in connection with the tBOND-L algorithm is defined as a metric derived from a computational simulation of a competitive binding environment. It is calculated by determining the equilibrium concentration of the desired [candidate:target] complex relative to the initial concentration of reactants in a simulated mixture containing the candidate, the target fragment, and the full-length parent molecule. Therefore, Specificity within tBOND-L quantifies the probability of the agent binding exclusively to its intended fragment target (e.g., a tDR) while avoiding the structurally related full-length parent (e.g., the mature tRNA)." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0143] The tBOND-L algorithm is implemented on a computer system that systematically evaluates every possible candidate sequence for a target fragment to identify optimal binders. The system operates by modeling a multi-component "virtual test tube" environment containing the candidate agent and all relevant molecular competitors — specifically the target fragment, the complementary fragment half (e.g., the 5' or 3’ sibling), and the full-length parent RNA — at defined initial concentrations. By calculating the thermodynamic equilibrium concentrations of all possible duplexes (candidate:target vs. candidate:parent), the system determines the specific probability that the agent will bind its intended target rather than the full-length precursor. This dualoptimization approach ensures that the selected agent is not only potent but also highly specific, minimizing potential off-target effects (such as the depletion of essential full-length tRNAs) that could compromise safety and efficacy.

[0144] While the tBOND-L algorithm is described in certain embodiments with respect to Antisense Oligonucleotides (ASOs), the scope of the disclosure extends to the design of any high-affinity nucleic acid analog capable of strand invasion. As used herein, "high-affinity nucleic acid analog" comprises Locked Nucleic Acids (LNA), Peptide Nucleic Acids (PNA), Phosphorodiamidate Morpholino Oligomers (PMO), 2'-O-methyl RNA, and other chemically modified backbones that exhibit enhanced thermal stability compared to DNA or RNA. The computational system is configured to adjust the thermodynamic parameters (e.g., nearest-neighbor energy rules) within the simulation module to account for the specific hybridization enthalpy and entropy of these alternative backbones, thereby calculating accurate Efficiency and Specificity scores for any chemically modified candidate.

[0145] Further features of the tBOND-L algorithm are further illustrated in connection with the exemplary illustrations of FIG.9 to FIG. 11.

[0146] FIG. 9 is a flowchart outlining the computer-implemented steps of the tBOND-L algorithm. The process involves importing a target parent RNA sequence (905), generating a library (910) of all possible candidate oligonucleotide sequences of a specified length, and then iterating (915) through each candidate.

[0147] To ensure exhaustive coverage of the target sequence, the step of generating the candidate library (910) may be performed using a sliding window algorithm. The processor is configured to" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT traverse the target RNA sequence using a window size corresponding to the desired oligonucleotide length (e.g., 15 to 20 nucleotides) with a stride of one nucleotide. This iteration generates every possible overlapping subsequence within the target region, ensuring that the optimal thermodynamic binding site is not missed due to arbitrary segmentation.

[0148] For each candidate, the system calculates its GC content (920), simulates a multicomponent test tube analysis (925) (e.g., utilizing a thermodynamics package such as NUPACK®), and calculates the final Efficiency (930) and Specificity (935) ratios based on the equilibrium concentrations derived from that simulation.

[0149] FIG. 10 illustrates the novel concept behind the Specificity calculation in the tBOND-L algorithm. It diagrams the network of potential interactions in the simulated competitive binding system. The system models the binding competition between the candidate agent (e.g., ASO) (1005), the target fragment (e.g., 3' tDR) (1010), the non-target sibling fragment (e.g., 5' tDR) (1015), and the Full-Length Parent Nucleic Acid (e.g., tRNA) (1020). By calculating the equilibrium concentration of the desired [CandidateiTarget Fragment] complex (1025) relative to all other possible complexes, the algorithm derives a realistic measure of specificity.

[0150] FIG. 11 shows a sample output visualization from the tBOND-L system. Each designed candidate is plotted as a point on a 2D graph with Efficiency on the x-axis and Specificity on the y-axis. This allows a user to quickly and intuitively identify the most promising candidates, which are located in the top-right " Target Range" quadrant, representing molecules that are both highly potent and highly specific.

[0151] The computational systems of the present disclosure produce structured, data-rich outputs that are designed to be directly actionable for researchers and application developers. The following tables provide examples of the data generated by the tBOND-L and tBOND-G algorithms, respectively.

[0152] Table 1 shows a sample output from the tBOND-L system (directed to targeting fragments / derivatives). This data would typically be generated as a comma-separated values (CSV) file, allowing for easy sorting and analysis. Each row represents a single candidate targeting moiety (e.g.. ASO) sequence that was designed and evaluated by the algorithm. Notably, toA Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT facilitate the specificity calculations described herein, the system aligns the candidate not only against the target fragment but also against the full-length parent nucleic acid.Table 1Candida Target Parent Nucleic ASO GC Effie Spec!te ASO Fragment Acid (e.g., Homo Length Cont ienc ficity Percentage of (e.g., 3' sapiens tRNA- ent y ASO binding to tDR) Asp-GTC) Full Length RNA TCCTCGTTAGTATAG 0.65GAATCG GCGGGAGA TGGTGAGTATCCCC AACCCC CCGGGGTT GCCTGTCACGCGG GGTCTC CGATTCCC GAGACCGGGGTTC CC (SEQ CGACGGGG GATTCCCCGACGGGID NO: A (SEQ ID GAGCCA (SEQ ID 81.5D NO: 2) NO: 3) 20 1 26.22 3.11TCCTCGTTAGTATAG 0.63GCGGGAGA TGGTGAGTATCCCC AATCGAA CCGGGGTT GCCTGTCACGCGG CCCCGG CGATTCCC GAGACCGGGGTTC TCTCCC CGACGGGG GATTCCCCGACGGG(SEQ ID A (SEQ ID GAGCCA (SEQ ID 28.6NO: 4) NO: 5) NO: 6) 19 4 24.48 1.17TCCTCGTTAGTATAG 0.67GCGGGAGA TGGTGAGTATCCCC ATCGAA CCGGGGTT GCCTGTCACGCGG CCCCGG CGATTCCC GAGACCGGGGTTC TCTCCC CGACGGGG GATTCCCCGACGGG(SEQ ID A (SEQ ID GAGCCA (SEQ ID 21.8NO: 7) NO: 8) NO: 9) 18 9 18.57 1.18TCCTCGTTAGTATAG 0.65AATCGAA GCGGGAGA TGGTGAGTATCCCC CCCCGG CCGGGGTT GCCTGTCACGCGG TCTCCC CGATTCCC GAGACCGGGGTTCG (SEQ CGACGGGG GATTCCCCGACGGGID NO: A (SEQ ID GAGCCA (SEQ ID 26.710) NO: 11 ) NO: 12) 20 4 17.10 1.56TCCTCGTTAGTATAG 0.7ATCGAA GCGGGAGA TGGTGAGTATCCCC CCCCGG CCGGGGTT GCCTGTCACGCGG TCTCCC CGATTCCC GAGACCGGGGTTC GC (SEQ CGACGGGG GATTCCCCGACGGGID NO: A (SEQ ID GAGCCA (SEQ ID 20.913) NO: 14) NO: 15) 20 9 14.98 1.40 Note: Sequences in this table are displayed using standard computational notation where Thymine (T) represents Uracil (U).

[0153] As shown in Table 1, the output includes: the designed ASO sequence ready for synthesis;" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT the full Target Fragment (tDR) sequence for reference; the Parent Nucleic Acid sequence (used to calculate competitive binding); the ASO Length and GC Content, which are relevant physical properties; and the two calculated metrics. Efficiency and Specificity. A user can sort this table to prioritize candidates, such as the first and ninth entries, which exhibit both high efficiency and high specificity, making them excellent candidates for therapeutic development.

[0154] Table 2 provides a detailed example of the output generated by the tBOND-G system for a single candidate targeting moiety (specifically a gRNA for a Casl3 effector). This comprehensive output provides all the necessary information for a researcher to understand the rationale behind the design and to plan their experiments. While the system is applicable to various RNA-guided nucleases, such as Casl3a, Casl3b, Casl3c, Casl3d, Casl3x, or Casl3y. the specific data presented below was generated using a model validated for pspCas13b targeting a structured tRNA.Table 2Data Field Value / SequencegRNA Barcode 1gRNA TACCACTGAGCTACACCCCCGTTGTGGAA GGTCCAGTTTTT GAGGGGCTATTACAAC (SEQ ID NO: 16)tRNA (target subsequence) GGGGGTGTAGGTCAGTGGTA (SEQ ID NO: 17)tRNA target (parent polynucleotide) GGGGGTGTAGCTCAGTGGTAGAGCGCGT GCTTAGCATGCACGAGGCCCCGGGTTCA ATCCGGGGCACCTCCACCA (SEQ ID NO: 18)tRNA Structure (((((((■■((((. ))))■(((((. ))))). (((((. ))))))))))))....gRNA structure U20D7(U1 D2(U1 D3(U8)U1 )U1 )Max Energy -108.68tRNA Target Name tRNA-Ala-AgcDelta G Ratio 0.37685816" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT Table 2Data Field Value / SequenceProduct Concentration Ratio 1Accessibility 0.403846155 PFS TT3 PFS GASpacer Length 20tDR Halves 3' tDRtDR sequence GGGGGTGT (SEQ ID NO:19)+CAGTGGTAGAGCGCGTGCTTAGCAT GCACGAGGCCCCGGGTTCAATCCCCGGC ACCTGGACCA (SEQ ID NO: 20)5'tDR Length 83'tDR Length 64Cleavage Efficiency 0.60549922Note: Sequences in this table are displayed using standard computational notation where Thymine (T) represents Uracil (U)."

[0155] As shown in Table 2, the output provides a complete characterization of the candidate. It includes unique identifiers (gRNA Barcode, Parent Name) and, critically, distinguishes between the Parent Nucleic Acid (the full-length input sequence) and the Target Subsequence (the specific region to which the gRNA is designed to hybridize). The output further details the predicted secondary structures in dot-bracket notation and the key calculated physical parameters utilized by the algorithm, specifically the Delta G Ratio and Accessibility Score. Finally, the table presents the Cleavage Efficiency Score predicted by the SVM model (0.605 in this instance, indicating a high likelihood of efficacy) and the precise characteristics of the resulting cleavage products: the predicted fragment type (e.g., 3' tDR), the exact cut site indicated by a plus sign ('+') within the sequence, and the resulting fragment lengths. This level of granular detail allows a researcher to make a highly informed decision about which candidates to advance to experimental validation. Additionally, in some embodiments, this output is supplemented with an 'Off-Target Score' or a list of potential off-target transcripts identified during the homology search, providing a complete" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT specificity profile for the candidate.

[0156] It is noted that the alignment of the 5' tDR fragment (SEQ ID NO: 19) and the 3' tDR fragment (SEQ ID NO: 20) against the parent tRNA polynucleotide (SEQ ID NO: 18) reveals a gap of four nucleotides ('AGCT' SEQ ID NO: 22). In the context of the present disclosure, this internal deletion represents the specific cleavage site or loop region (e.g., the D-loop or anticodon loop) that is excised, unwound, or biologically processed during the Cas-mediated fragmentation event, resulting in the generation of the distinct bioactive tDR species.

[0157] In some embodiments, the tBOND-G algorithm and tBOND-L algorithm are provided as components of a tBOND-Integrated computational platform configured for the functional characterization and validation of specific RNA fragments (e.g., tRFs) derived from structured single-stranded RNA parents. Computer-implemented methods and systems of this set of embodiments address a fundamental challenge in RNA biology: establishing a causal link between a specific RNA fragment and a cellular phenotype, distinct from the function of its full-length parent molecule. The integrated system coordinates the operation of two distinct computational modules to facilitate a rigorous experimental feedback loop comprising both gain-of-function (generation) and loss-of-function (inhibition) validations within the same biological system. By combining these modalities, the platform allows for the precise dissection of RNA regulatory networks, ensuring that observed biological effects are attributable to the specific fragment rather than artifacts of precursor depletion or off-target interference.

[0158] In a first operational mode, the integrated platform functions as a gain-of-function design tool utilizing the tBOND-G algorithmic principles described herein. A processor receives the sequence of a full-length parent nucleic acid (e.g., tRNA) and identifies a target fragment region to be generated. The system then executes a machine learning model to design a targeting moiety (such as a gRNA) optimized for a programmable nuclease (e.g., a Casl3 variant) to specifically cleave the parent into the desired fragment. Unlike traditional overexpression vectors that introduce synthetic oligonucleotides, this approach utilizes the cell’s own machinery (or introduced effectors) to process the endogenous RNA pool into the specific fragment. The selection of the targeting moiety is driven by biophysical parameters, specifically the structural accessibility of the cleavage site and the thermodynamic stability of the agent-target complex. The" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT machine learning model prioritizes sequences that target "breathing" loops or single-stranded regions within the secondary structure of the parent, ensuring efficient intracellular processing. This induced generation allows researchers to observe the emergence of a specific biological phenotype in a biologically relevant context where endogenous base modifications are preserved.

[0159] In a second operational mode, the integrated platform functions as a loss-of-function design tool utilizing the tBOND-L competitive simulation framework. To confirm that the phenotype observed in the first mode is indeed caused by the fragment and not an artifact of parent depletion, the system designs a high-affinity inhibitor specifically targeting the induced fragment. While Locked Nucleic Acids (LNAs) are utilized in preferred embodiments due to their enhanced thermal stability, the system is adaptable to design other high -affinity nucleic acid analogs, such as Peptide Nucleic Acids (PNA) or Phosphorodiamidate Morpholino Oligomers (PMO). The processor employs a multi-component competitive binding simulation to identify inhibitor sequences based on two calculated metrics: an efficiency score and a specificity score. The efficiency score ensures robust binding to the target fragment, while the specificity score ensures negligible binding to the abundant full-length parent. This simulation models the thermodynamic equilibrium of the inhibitor in a mixture containing the target fragment, the complementary fragment strand, and the full-length parent molecule, thereby selecting reagents that can discriminate between the fragment and the precursor with high precision.

[0160] In some embodiments, the method of operating the tBOND-Integrated platform further comprises coordinating the first and second modes to establish a functional feedback loop or " Toggle" experimental strategy. In this workflow, a researcher utilizes the designed tBOND-G agent to induce the expression of the target fragment, thereby "turning on” a specific cellular phenotype. Subsequently, or in parallel experimental setups, the researcher utilizes the designed tBOND-L inhibitor to bind and neutralize that specific fragment, thereby "turning off" the signal. If the phenotype induced by the generation step is successfully reversed by the specific inhibitor, the causal role of the fragment is biologically validated. This integrated workflow provides a robust method for dissecting the functional independence of RNA fragments from their precursors, guiding the development of precise therapeutic interventions. In various embodiments, this coordination is managed via a unified user interface that allows for the simultaneous design of paired Generation (G) and Locking (L) sets for any given structured RNA input." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0161] The teachings of the present disclosure represent a practical and specific computer-implemented tool that transforms data representing a biological molecule through a series of concrete computational steps to produce a useful, tangible output, i.e„ a set of optimized nucleic acid sequences for synthesis and experimental use.

[0162] The methods are performed by a general-purpose computer system comprising at least one processor, a memory (e.g., RAM), non-volatile storage (e.g., a hard drive), a user input device (e.g., keyboard), and a display. The memory stores a set of computer-executable instructions that, when executed by the processor, configure the system to perform the steps of the tBOND-G and tBOND-L algorithms.

[0163] The system may be implemented within a Python (e.g., version 3.9 or higher) computing environment, leveraging several specialized software libraries that are integral to its function. The NUPACK® [2] library is used to perform the complex thermodynamic calculations of nucleic acid secondary structure and binding energies, which are central to the invention's physics-based approach. The Pandas and NumPy libraries are used for efficient data manipulation and numerical operations, particularly for managing the large datasets of candidate sequences and their calculated parameters. The Matplotlib library is used to generate the data visualizations, such as the SVM map (FIG. 3) and the ASO / Inhibitor performance plot (FIG. 11), which are relevant for user interpretation of the results.

[0164] The process is a specific, ordered series of transformations performed by the processor. For tBOND-G, the processor:1. Receives a digital string representing a target parent RNA sequence (e.g., a tRNA).2. Executes instructions to algorithmically generate a new dataset of related but distinct targeting moiety sequence strings (e.g., gRNAs).3. Accesses an external database (e.g., RNA Central) to retrieve a dot- bracket notation string representing the parent RNA's structure.4. Executes the NUPACK software module to transform the sequence and structure strings into numerical values representing accessibility and binding energy.5. Loads a pre-trained SVM model file from memory and uses it to classify the calculated" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT numerical values, outputting a new numerical value representing targeting efficacy (e.g., cleavage efficiency).6. Executes molecular dynamics modeling instructions to identify specific nucleotide positions (cleavage sites) and calculates the resulting fragment lengths.7. Generates a digital image file (the virtual gel) representing these fragment lengths for output on the display.

[0165] This constitutes a concrete application of computing technology to solve a specific technical problem in the field of biotechnology, producing a practical result that guides real-world laboratory work.

[0166] The teachings of the present disclosure provide technical improvements over existing methods for designing nucleic acid tools for structured RNA targeting.1. Improved Efficiency and Success Rate: Traditional methods rely on screening large libraries of candidate sequences, a process with a notoriously low success rate. The tBOND- G and tBOND-L systems replace this with a predictive model that enriches for high- performing candidates. By pre-screening candidates computationally, the invention dramatically increases the "hit rate" of sequences that are successful in experiments, saving significant time, labor, and material costs.2. Solves the Technical Problem of Specificity: The tBOND-L algorithm provides a novel and effective solution to the specific technical challenge of targeting a fragment (e.g., tDR) without affecting its parent nucleic acid (e.g., tRNA). The unique four-component simulation and calculation of a " Specificity" score is a specific, unconventional approach that directly addresses a major roadblock in the development of fragment-based therapeutics.3. Enables New Research Capabilities: The tBOND-G system provides, for the first time, a reliable method for inducing the endogenous generation of specific RNA fragments. This is a technical improvement that allows researchers to study the function of these molecules in a more biologically relevant context, complete with their native chemical modifications, which was not previously feasible.4. Provides Improved Output: The process in accordance with the present disclosure does not" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT simply output a list of sequences. The tBOND-G system provides a predicted efficiency score and a virtual gel visualization. The tBOND-L system provides a 2D plot of Efficiency vs. Specificity. These outputs are technical improvements that provide a richer, more intuitive, and more actionable guide for the end-user compared to raw sequence data, reducing the cognitive burden on the researcher and facilitating better decision-making.5. Reduces Reliance on Physical Experimentation: By providing a highly accurate predictive framework, the present disclosure shifts a significant portion of the design and optimization process from the physical laboratory to a computational environment. This reduces the number of required physical experiments, accelerating the pace of research and development.

[0167] FIG. 12 illustrates a generic computer system architecture suitable for implementing the various methods and systems described herein, particularly the tBOND-G and tBOND-L algorithms. The computer system includes a Central Processing Unit (CPU) (1200), executing instructions and performing calculations for the computational steps of the tBOND algorithms. For instance, the CPU (1200) performs sequence analysis, binding affinity predictions, and optimization routines as described in detail for tBOND-G and tBOND-L. Input devices (1205), such as a keyboard or mouse, allow users to input parameters, target parent sequences, or initiate the execution of the tBOND algorithms. The CPU (1200) processes this input and, upon completing the computational analysis, may send results to output devices (1210), such as a display monitor or printer, to present the predicted off-targets, binding scores, virtual gel simulations, or optimized targeting moiety designs.

[0168] The system also includes system memory (1215), which can comprise both volatile memory (RAM) and non-volatile memory (ROM). RAM is particularly crucial for temporarily storing the large datasets of reference genomes, structural databases (e.g., RNA secondary structures), candidate libraries, and intermediate computational results generated during the execution of the tBOND-G and tBOND-L alignment and scoring phases. Storage devices (1220), such as Solid State Drives (SSD) or Hard Disk Drives (HDD), provide long-term data storage for the installed tBOND software, reference databases, and user-defined sequence libraries. A data / address bus (1225) facilitates high-speed communication between the CPU (1200), system memory (1215), and storage devices (1220), ensuring efficient data flow for complex calculations. Furthermore, the system may include various input / output (VO) and network interfaces (1230)," A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT such as network interface cards (NICs) (1235) for connecting to external databases or cloud computing resources, or USB / peripheral ports (1240) for importing input files or exporting results.

[0169] The various methods and algorithms described herein, including the tBOND-G and tBOND-L algorithms, may be embodied as computer-executable instructions stored on one or more non-transitory computer-readable media. Such non-transitory computer-readable media can include, but are not limited to, system memory (1215), storage devices (1220) such as solid-state drives or hard disk drives, CD-ROMs, DVDs, flash memory devices, or any other tangible medium capable of storing instructions that, when executed by one or more processors (e.g., CPU (1200)), cause the processors to perform the steps of the disclosed methods. These instructions enable the computer system to receive input data, perform the sequence analysis, binding score calculations, and optimization routines, and generate the described outputs (e.g., ranked candidate lists, specificity plots, and virtual gels) as detailed throughout this specification.

[0170] The computational frameworks detailed in the preceding aspects provide a robust engine for generating high-precision nucleic acid tools. However, the utility of these systems extends beyond mere sequence prediction; they serve as a foundational platform for addressing complex challenges in translational medicine and biotechnology. By transforming the physical parameters of RNA folding and binding into predictive design criteria, the described methods enable the creation of reagents that function reliably in the complex milieu of a living cell. This transition from in silico modeling to in vitro and more notably in vivo application opens a wide array of opportunities for modulating the tRNA-tDR regulatory axis, as well as distinct regulatory networks involving other structured single-stranded RNAs.

[0171] For example, in the field of therapeutics and diagnostics, the ability to distinguish between a functional fragment and its essential precursor is paramount. The tDRs generated or targeted by these systems are implicated in a broad spectrum of pathologies, ranging from the suppression of tumor suppressors in cancer to the modulation of synaptic function in neurological disorders. Consequently, the nucleic acids designed by these frameworks find immediate application in developing precision therapies that minimize off-target toxicity and in creating diagnostic probes capable of detecting specific RNA processing signatures associated with disease states. In additional embodiments, these principles are applied to other structured RNA targets where" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT distinguishing a specific structural motif or fragment is critical. For example, the framework may be utilized to design agents targeting conserved secondary structures within viral RNA genomes (e.g., HIV, SARS-CoV-2, Influenza) to inhibit viral replication without affecting host transcripts, or to target specific domains of long non-coding RNAs (IncRNAs) that act as structural scaffolds for chromatin-modifying complexes.

[0172] Furthermore, the versatility of the disclosed framework supports advanced applications in biological engineering and functional genomics. Whether the goal is to dissect the regulatory logic of RNA networks, engineer stress-resilient cell lines for biomanufacturing, or develop programmable genetic circuits, the methods described herein provide the necessary precision to manipulate endogenous RNA pools on demand. Beyond natural endogenous targets, the system is further configured in some embodiments to design tools for synthetic structured RNAs, such as riboswitches, aptamers, or catalytic ribozymes used in synthetic biology circuits. Accordingly, the following aspects detail specific methods of using these computationally designed molecules to achieve tangible biological outcomes, demonstrating the practical integration of the tBOND platform into molecular medicine and bioengineering workflows.

[0173] Exemplary applications utilizing the disclosed framework comprise a wide range of disciplines including biomedical research, therapeutic development, and molecular engineering. In the therapeutic domain, the methods are particularly valuable for developing precision interventions for diseases driven by RNA dysregulation, such as cancer, neurological disorders, and metabolic conditions. By allowing for the specific neutralization of pathogenic fragments without inducing off-target cytotoxicity, the system overcomes the limitations of traditional antisense therapies. While particularly advantageous for the high-homology tRNA / tDR system, the disclosure further contemplates the application of these methods to other classes of structured non-coding RNAs, including but not limited to small nucleolar RNAs (snoRNAs), precursor microRNAs (pre-miRNAs), and circular RNAs (circRNAs). In the realm of basic research and functional genomics, these systems serve as critical tools for dissecting complex RNA regulatory networks, enabling the "decoupling" of a fragment's function from its parent molecule to validate biological causality. Furthermore, in synthetic biology and biotechnology, these methods find application in the programmable control of cellular states, allowing engineers to modulate protein translation efficiency and stress responses by precisely regulating the processing of endogenous" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT tRNA pools or other essential structural RNAs. Finally, the high-specificity reagents generated by these frameworks are applicable to diagnostic platforms requiring the resolution of specific RNA isoforms or processing intermediates as biomarkers for disease stratification.

[0174] In some embodiments, the framework can be used in connection with a method for treating a condition associated with the aberrant accumulation of a target RNA fragment, such as a tRNA-derived fragment (tDR) or a fragment derived from another structured single- stranded RNA.

[0175] As used herein, the term "condition" indicates a physical status of the body of an individual (as a whole or as one or more of its parts), that does not conform to a standard physical status associated with a state of complete physical, mental and social well-being for the individual. In the present disclosure, conditions herein described comprise disorders and diseases driven by RNA dysregulation. The term "disorder" indicates a condition of the living individual that is associated with a functional abnormality of the body, such as the fragment-mediated repression of protein translation (e.g., by tDRs) or stress granule formation. The term "disease" indicates a condition of the living individual that impairs normal functioning of the body, such as cancer metastasis or neurodegeneration, which is typically manifested by distinguishing signs and symptoms associated with the pathogenic fragment. In additional embodiments, the condition may relate to viral infections, wherein the pathogenic fragment is a stable intermediate of a viral RNA genome (e.g., a flavivirus structural element), or disorders linked to long non-coding RNAs (IncRNAs), wherein the fragment represents a dysregulated structural domain.

[0176] The method represents a direct therapeutic application of the computational design principles described herein. The method initiates with the identification of a target fragment (e.g., a tDR) associated with the condition in an individual, wherein the target fragment sequence is comprised entirely within a full-length target parent nucleic acid (e.g., a mature tRNA).

[0177] The term "individual" as used herein in the context of treatment includes a single biological organism, including but not limited to, animals and in particular higher animals and in particular vertebrates such as mammals and in particular human beings, in which the RNA processing and structural mechanisms described herein are conserved. Following identification, the system generates a comprehensive library of candidate complementary oligonucleotide sequences, specifically utilizing high-affinity nucleic acid analogs (such as Locked Nucleic Acids" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT (LNAs), Peptide Nucleic Acids (PNAs), or other chemically modified backbones) complementary to the identified target region.

[0178] The method proceeds by subjecting the library to a computational analysis to calculate an efficiency score and a specificity score through the simulation of a competitive binding environment. Based on these calculated metrics, a specific complementary oligonucleotide sequence is selected that exhibits both high efficiency and high specificity. The method concludes with the administration of a therapeutically effective amount of the selected oligonucleotide agent to the individual to effect treatment or prevention.

[0179] The term "treatment" as used herein indicates any activity that is part of a medical care for, or deals with, the condition, medically or surgically, specifically encompassing the administration of the designed complementary oligonucleotide to neutralize the pathogenic fragment.

[0180] The term "prevention" as used herein indicates any activity which reduces the burden of mortality or morbidity from the condition in the individual. In the context of fragment-associated pathologies, this takes place at primary, secondary and / or tertiary prevention levels, wherein: a) primary prevention avoids the development of a disease by inhibiting fragment accumulation prior to phenotypic onset; b) secondary prevention activities are aimed at early disease treatment, such as targeting fragments detected in early diagnostic screenings, thereby increasing opportunities for interventions to prevent progression of the disease and emergence of symptoms; and c) tertiary prevention reduces the negative impact of an already established disease by restoring function — specifically by restoring the functional pool of the parent nucleic acid (e.g., tRNA) for protein synthesis — and reducing disease-related complications caused by the fragment.

[0181] In some embodiments, the framework of the present disclosure is used in connection with a method for treating a condition associated with the aberrant accumulation of a target tRNA-derived fragment (tDR) or other structured RNA fragment in an individual. This method utilizes the tBOND-L computational framework to address the specific therapeutic challenge of distinguishing a pathogenic RNA fragment from its essential, full-length precursor.

[0182] In embodiments herein described, the method for treating a condition associated with the" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT aberrant accumulation of a target fragment in an individual comprises identifying a target fragment associated with a condition in an individual. This step involves defining the precise nucleotide sequence of the pathogenic fragment, which is characterized by being comprised entirely within the sequence of a full-length target parent nucleic acid, thereby creating a high degree of sequence homology that complicates traditional targeting. Following identification, the method comprises generating a library of candidate sequences.

[0183] In this context, the term 'pathogenic fragment' or 'pathogenic target fragment' refers to a specific subset of RNA fragments (e.g., tDRs) whose aberrant accumulation is causally linked to the etiology or progression of the condition. Unlike 'bystander' degradation products that may accumulate passively during cell death, a pathogenic fragment actively drives cellular dysfunction, for example, by competitively binding to RNA-binding proteins (e.g., YBX1), displacing translational initiation factors (e.g., eIF4G), or nucleating stress granules in a manner that arrests essential protein synthesis. Accordingly, the identification step described herein distinguishes these bioactive, disease-driving fragments from inert RNA debris based on functional assays or transcriptomic signatures associated with the disease state.

[0184] In the context of the methods described herein, the step of generating a library refers to a computer-implemented process wherein a processor constructs a digital dataset comprising a plurality of nucleotide sequence strings. This library is not a physical collection of biological molecules, but rather a virtual repository of candidate sequences that serves as the input for subsequent thermodynamic simulations. The generation process initiates when the processor receives a primary input string representing the nucleotide sequence of the target, such as the full-length parent nucleic acid or a specific fragment.

[0185] The library is generated using a "sliding window" algorithm or iteration function. The processor is configured to systematically scan the digital input string, extracting subsequences of a user-defined length (e.g., 20-30 nucleotides for Casl3 gRNA spacers, or 15-20 nucleotides for complementary oligonucleotides). The algorithm typically operates with a one-nucleotide stride, moving the selection window one base at a time from the 5' end to the 3’ end of the target sequence. This exhaustive enumeration ensures that every possible overlapping subsequence within the target region is captured and evaluated, preventing the accidental exclusion of potentially high-" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT performing candidates that might be missed by heuristic or non-systematic selection methods.

[0186] Furthermore, the generation step includes a transformation of the raw extracted subsequences into their functional counterparts. Because the therapeutic or guide reagents must hybridize to the target, the processor computationally generates the reverse complement of each extracted window. For a gRNA library, these reverse complement strings represent the "spacer" sequences; for an inhibitor library, they represent the complementary oligonucleotide sequences. These generated strings are then stored in the system memory as a structured digital library, ready to be iterated through the thermodynamic assessment modules (e.g., NUPACK) to calculate the efficiency and specificity scores described in the preceding aspects.

[0187] In embodiments herein described, the method for treating a condition associated with the aberrant accumulation of a target RNA fragment (e.g., a tDR) in an individual further comprises performing a computational assessment of this library to calculate an efficiency score for each candidate.

[0188] The step of performing a computational assessment to calculate an efficiency score for each candidate in the digital library is executed by the processor using a thermodynamic modeling engine. This assessment is not a simple sequence match but a quantitative prediction of physical binding behavior. The process begins by retrieving the sequence of the candidate (e.g., the complementary oligonucleotide) and the sequence of its intended target (the specific fragment). The processor then utilizes a nucleic acid thermodynamics software package, such as NUPACK®, [2] to calculate the Gibbs free energy (Delta G) of hybridization for the formation of the Candidate: Target duplex. This value represents the energy released when the candidate successfully binds to the correct fragment.

[0189] To derive the final efficiency score, the system does not look at this binding energy in isolation. Instead, it performs a comparative analysis. The processor simultaneously calculates the binding energy of the candidate against potential off-target structures. Crucially, in the context of structured RNA targeting (e.g.. tRNA), these off-target structures are often intramolecular secondary structures (stems and loops) within the target sequence or the parent molecule itself that might compete for the candidate's binding site or sequester the target sequence in a double- stranded conformation that blocks access. The efficiency score is computed as a ratio or normalized value" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT comparing the Candidate: Target binding energy against these competing energetic states. A higher score indicates that the candidate has sufficient affinity to thermodynamically invade and unfold the native secondary structure of the target to form a stable duplex, effectively outcompeting the natural folding propensity of the RNA.

[0190] In embodiments herein described, the method for treating a condition associated with the aberrant accumulation of a target RNA fragment (e.g., tDR) in an individual further comprises simulating a competitive binding environment to derive a specificity score. In this step, the processor executes a multi-component reaction simulation that models the thermodynamic equilibrium of a mixture containing specific initial concentrations of the candidate complementary oligonucleotide, the target fragment, and the abundant full-length parent nucleic acid. By calculating the partition function for this system, the algorithm predicts the equilibrium concentrations of all possible complexes — specifically distinguishing between the concentration of the desired [01igonucleotide: Fragment] complex and the undesired | Oligonucleotide: Parent | complex. The specificity score is derived from this ratio, providing a probabilistic measure of exclusive binding in a physiological context.

[0191] In particular, the processor initializes a virtual "test tube" environment comprising four distinct nucleic acid species: the candidate complementary oligonucleotide sequence, the specific target fragment (e.g., the 3' tRF), the complementary fragment (e.g., the 5' tRF, representing the non-target remainder of the parent), and the full-length parent nucleic acid (e.g., tRNA). The simulation utilizes a nucleic acid thermodynamics software package, such as NUPACK®[2], to calculate the partition function for the system, taking into account all possible interactions between these species, including the formation of the desired [Oligonucleotide: Target] duplex, the undesired [01igonucleotide: Parent] duplex, and other potential secondary structures or homodimers (e.g., setting the complex size limit to maxsize ~ 3).

[0192] Based on these partition functions and the defined initial concentrations (e.g., 1 pM for each species), the processor solves for the equilibrium concentration of each complex. The specificity score is then derived by calculating the ratio of the equilibrium concentration of the desired [Oligonucleotide: Target] complex to the total initial concentration of the oligonucleotide (or the sum of all oligonucleotide-containing complexes). This quantitative ratio represents the" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT probability that a given complementary oligonucleotide molecule will bind to the specific pathogenic fragment rather than being sequestered by the abundant parent nucleic acid pool, thereby ensuring high target selectivity in a competitive physiological environment.

[0193] In embodiments herein described, the method for treating a condition associated with the aberrant accumulation of a target fragment in an individual also comprises selecting a specific complementary oligonucleotide sequence from the library based on its position within a target performance range. This step is performed by the processor executing a multi-parametric filtering or ranking algorithm on the processed library. This effectively transforms the raw numerical data generated in the previous calculation steps into an actionable design decision. Specifically, the processor maps each candidate sequence onto a two-dimensional data structure or visualization plane, wherein a first dimension corresponds to the calculated Efficiency Score and a second dimension corresponds to the calculated Specificity Score.

[0194] The "target performance range" is defined as a specific quadrant or bounded region within this two-dimensional space, delineated by pre-determined numerical thresholds. For example, the system may define the target range as the subset of candidates exhibiting an Efficiency Score exceeding a first value (e.g., > 80, indicating robust binding) and a Specificity Score exceeding a second value (e.g., > 75, indicating high discrimination against the parent). The processor filters the library to exclude any candidate failing to meet these dual criteria. From the remaining subset of high-performing candidates located within this Target Range, the system selects the optimal sequence, potentially applying a secondary sorting logic to prioritize either maximal potency (highest efficiency) or maximal safety (highest specificity) depending on the specific therapeutic requirements of the condition being treated.

[0195] In embodiments herein described, the method for treating a condition associated with the aberrant accumulation of a target fragment in an individual additionally comprises administering a therapeutically effective amount of this selected complementary oligonucleotide to the individual having the condition. In such embodiments, the administering is performed to reduce and preferably minimize the accumulation of the pathogenic fragment, thereby treating the condition, while preserving the functional pool of the parent nucleic acid to maintain essential cellular functions (e.g., protein synthesis). The step of administering is performed by delivering the" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT pharmaceutical composition comprising the oligonucleotide to the target tissue or systemic circulation, wherein the specific mode of administration is selected based on the physicochemical properties of the oligonucleotide agent.

[0196] In some embodiments, the administering is performed via parenteral routes, including but not limited to intravenous (IV), subcutaneous (SC), or intramuscular (IM) injection. For systemic delivery, particularly for targeting the liver (e.g., in metabolic conditions associated with tRNA-derived fragments), the complementary oligonucleotide may be conjugated to a targeting ligand such as N-acetylgalactosamine (GalNAc) to enhance uptake by hepatocytes. Alternatively, for targeting extra-hepatic tissues, the oligonucleotide can be formulated in lipid nanoparticles (LNPs) or conjugated to cell-penetrating peptides (CPPs) or lipids (e.g., cholesterol) to facilitate traversal of cell membranes and endosomal escape.

[0197] In embodiments where the condition affects the central nervous system (CNS), such as neurodegenerative disorders (e.g., Alzheimer’s disease, Parkinson’s disease) driven by pathogenic fragments, the administering is preferably performed via intrathecal (IT) or intracerebroventricular (1CV) injection to bypass the blood-brain barrier. For localized conditions, direct injection into the target tissue (e.g., intravitreal for eye disorders or intratumoral for cancer) may be employed to maximize local concentration while minimizing systemic exposure. Accordingly, the route of administration is specifically selected based on the pathology of the condition: systemic parenteral injection is utilized for widely disseminated or metabolic targets, while specialized local delivery is utilized for CNS or solid tumor targets to ensure the complementary oligonucleotide reaches the specific cellular pool of the pathogenic fragment.

[0198] A "therapeutically effective amount" is determined as the quantity of complementary oligonucleotide sufficient to reduce the level of the pathogenic target fragment to a non-pathogenic range or to ameliorate symptoms of the condition, without inducing significant toxicity or off-target effects (such as the depletion of the parent nucleic acid). This amount is established through dose-ranging studies that monitor biomarkers of fragment reduction and clinical endpoints. The dosing regimen may involve a loading phase followed by maintenance doses, adjusted based on the tissue half-life of the specific modified oligonucleotide chemistry employed.

[0199] In some embodiments, the framework of the present disclosure can be used in connection" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT with a method for modulating cellular function by inducing the formation of a bioactive RNA fragment (e.g., a tDR). In those embodiments, the method provides a "gain-of-function" molecular engineering strategy that utilizes the cell's endogenous transcriptional output as a substrate for generating regulatory molecules.

[0200] In embodiments herein described, the method comprises identifying a structured parent nucleic acid (e.g., a precursor tRNA) known or predicted to be processed into a bioactive fragment. These fragments are integral to diverse biological pathways, including the regulation of gene expression via RNA interference-like mechanisms, the repression of translation initiation, or the nucleation of stress granules during cellular response to environmental stimuli.

[0201] Following target identification, the method further comprises delivering to a target cell population a system comprising a CRISPR-associated (Cas) RNA-guided nuclease and a specifically engineered targeting moiety (e.g., gRNA). The delivery mechanism may involve viral vectors (e.g., AAV, Lentivirus), ribonucleoprotein (RNP) complexes, or lipid nanoparticle (LNP) formulations, depending on the specific cell type and the desired duration of effect.

[0202] In embodiments herein described, the method also comprises selecting the spacer sequence of the gRNA to target a specific region of the parent RNA, distinct from regions protected by tight tertiary folding (such as the anticodon stem in tRNAs). This selecting step comprises executing the tBOND-G algorithm to generate a cleavage efficiency prediction for the gRNA, thereby ensuring that the delivered complex effectively processes the target. The executing step involves applying a machine learning model, such as a Support Vector Machine (SVM), configured to process critical biophysical parameters. Specifically, the processing comprises calculating a numerical accessibility score, which is derived from the dot-bracket secondary structure representation of the parent. This score quantifies the proportion of unpaired nucleotides within the target window, prioritizing "breathing" loops or single- stranded regions that are physically available for Cas interaction. Simultaneously, the processing comprises determining a thermodynamic binding energy value (Delta G Ratio) of the [gRNA: Parent] complex to ensure the hybridization is energetically favorable enough to displace local secondary structures. By integrating these parameters, the algorithm operates to filter a library of potential spacers to identify those that target accessible loops within the folded structure, thereby rejecting candidates" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT likely to fail due to steric hindrance.

[0203] In embodiments herein described, the method additionally comprises inducing the controlled generation of the bioactive fragment within the target cell population. This inducing step utilizes a Cas nuclease selected from RNA-targeting variants — including but not limited to Casl3a, Cas 13b, Cas 13c, Cas 13d, Casl3x, or Casl3y — which, upon binding to the accessible target site identified by the tBOND-G algorithm, undergoes a conformational change that activates its catalytic domains. This activation results in the site-specific hydrolysis of the phosphodiester backbone of the parent RNA. This programmed cleavage event releases the specific 5' or 3' fragment into the cytoplasm, thereby increasing its intracellular concentration and modulating said cellular function. In preferred embodiments, the selection of the gRNA is further refined by simulating a virtual gel result that predicts the exact migration pattern of the cleaved fragments, allowing the user to verify that the induced molecular species corresponds to the biologically active natural fragment.

[0204] In some embodiments, the framework of the present disclosure can be used in connection with a method for detecting aberrant accumulation of a target RNA fragment (e.g., tDR) biomarker in an individual. This method relies on the high- specificity molecular recognition capabilities enabled by the computational design framework described herein.

[0205] In embodiments herein described, the method comprises the step of performing the tBOND-L algorithm to design a complementary oligonucleotide probe (e.g., comprising Locked Nucleic Acid (LNA) modifications) specific for the target fragment biomarker. This design step is critical for diagnostic accuracy because the target fragment is comprised entirely within the sequence of a full-length target parent nucleic acid (e.g.. tRNA), which is typically present in the sample at high abundance. Standard probe design would likely result in cross -hybridization with the parent, leading to false-positive signals. The tBOND-L algorithm overcomes this by executing a competitive binding simulation that identifies a probe sequence capable of thermodynamically distinguishing the free fragment from the same sequence embedded within the folded parent molecule.

[0206] In embodiments herein described, the method further comprises obtaining a biological sample from the individual. The biological sample may be selected from fluids such as blood," A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT plasma, serum, urine, or cerebrospinal fluid (liquid biopsy), or from solid tissue samples obtained via biopsy or surgical resection.

[0207] In embodiments herein described, the method also comprises contacting the sample with a detectable probe comprising the designed complementary oligonucleotide. The probe is typically modified with a detectable moiety, such as a fluorophore, a radioactive isotope, an enzyme (e.g., horseradish peroxidase), or a high-affinity tag (e.g., biotin). The contacting step is performed under hybridization conditions — defined by specific temperature, ionic strength, and denaturing agents — that allow the probe to anneal to its complementary sequence. Due to the computational selection of the probe sequence, these conditions favor the formation of a stable [Probe: Fragment] complex while thermodynamically destabilizing potential interactions with the full-length parent nucleic acid.

[0208] In embodiments herein described, the method additionally comprises detecting the presence or quantifying the level of the target [Probe: Fragment] complex in the sample. Detection may be achieved through various analytical platforms, including but not limited to Northern blotting, quantitative PCR (qPCR), microarray analysis, flow cytometry, or in situ hybridization (ISH). The signal intensity generated by the detectable moiety is directly proportional to the amount of the specific complex formed. In determining a diagnosis, the method involves comparing the quantified level of the complex against a predetermined threshold. This threshold is typically established based on a reference range derived from healthy individuals. Detection of a level of the complex above this predetermined threshold indicates the aberrant accumulation of the fragment, thereby signaling the presence of a condition (such as a specific cancer subtype or neurodegenerative state) or a specific stage of condition progression. This enables a diagnosis based strictly on the detection of the pathological fragment biomarker, unconfounded by the levels of the housekeeping parent nucleic acid.

[0209] In some embodiments, the framework of the present disclosure can be used in connection with a method of molecular engineering for regulating protein translation efficiency or stress response pathways in a eukaryotic cell. This method leverages the programmable nature of CRISPR-associated nucleases to effectuate precise "molecular surgery" on the endogenous RNA pool." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0210] In embodiments herein described, the method comprises introducing into the eukaryotic cell a programmable Cas RNA-guided nuclease system, which may be selected from RNA-targeting variants such as Cas 13a, Cas 13b, Cas 13c, Cas 13d, Casl3x, or Casl3y, alongside a specifically designed targeting moiety (e.g., gRNA). The gRNA is engineered with a spacer sequence that targets a precise region of a specific structured parent RNA (e.g., a transfer RNA (tRNA) transcript). Unlike standard knockdown approaches intended to eliminate the target, this targeting is designed to induce the controlled formation of specific bioactive cleavage products, such as tRNA-derived stress-induced RNAs (tiRNAs) or tRNA fragments (tRFs). These induced fragments are functionally capable of inhibiting translation initiation or promoting the assembly of stress granules, thereby allowing the engineer to modulate the cell's metabolic state.

[0211] In embodiments herein described, the method further comprises selecting the targeting moiety (e.g., gRNA) sequence utilizing the tBOND-G computational framework. This selection process is critical for achieving "tunable control" over the cellular phenotype. The framework optimizes cleavage efficiency by analyzing two distinct biophysical parameters: the structural accessibility of the parent RNA loops (identifying regions that are not sterically hindered by the rigid tertiary fold) and the thermodynamic stability of the [Agent: Parent] interaction. By selecting agents with specific predicted efficiency scores, the method allows the engineer to titrate the rate of fragmentation. This ensures that a sufficient pool of the full-length parent nucleic acid remains intact to sustain essential housekeeping functions (e.g., protein synthesis), while generating a sufficient concentration of the regulatory fragment to trigger the desired stress response or translational pausing.

[0212] In embodiments herein described, the method also comprises verifying functional decoupling in the RNA regulatory network. This step is performed by contacting the cell with a high-affinity complementary oligonucleotide (e.g., an LNA or ASO reagent) designed to selectively neutralize the induced signal. This inhibitor is designed using the tBOND-L algorithm, which employs a competitive binding simulation to identify sequences that bind the specific cleaved fragment with high affinity while discriminating against the full-length parent nucleic acid. By administering this specific inhibitor, the engineer can selectively inhibit the regulatory function of the fragment (e.g., dispersing stress granules) without affecting the canonical function of the precursor. This validates the engineered phenotype, confirming that the observed physiological" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT change is driven by the fragment's activity rather than the depletion of the parent pool. This methodology establishes a platform for functional decoupling in RNA regulatory networks: if the phenotype induced by the Cas-mediated cleavage is reversed by the administration of the fragment-specific inhibitor (which does not bind the parent), the causal role of the fragment is confirmed.

[0213] In some embodiments, the framework of the present disclosure can be used in connection with a method for the design of a nucleic acid reagent (e.g., a therapeutic or diagnostic oligonucleotide) specific for a target environment. This aspect addresses the variability of RNA expression profiles across different tissues, disease states, or individual patients, enabling a "precision medicine" approach to reagent design. Unlike static design methods that assume standard stoichiometric ratios, this method tunes the reagent to the specific molecular landscape where it is intended to function.

[0214] In embodiments herein described, the method comprises obtaining a biological sample from the target environment. The target environment may be a specific tissue type (e.g.. neuronal tissue vs. hepatic tissue), a cell line, or a patient biopsy characterized by a unique disease pathology. The method further comprises experimentally measuring a concentration ratio between the specific target fragment (e.g., a pathogenic tDR) and its structured parent nucleic acid (e.g., the parent tRNA) within said sample. This quantification may be performed utilizing techniques such as quantitative PCR (qPCR), RNA-sequencing (RNA-seq), or Northern blotting to establish the precise abundance levels of the competitor molecules.

[0215] Following this measurement, the method comprises inputting said concentration ratio into the tBOND-L computational framework as a simulated physiological condition for the competitive binding calculation. The processor is configured to adjust the initial concentration variables (1:50 tDR:parent RNA, measured with existing experimental conditions) in the equilibrium equations to reflect the empirically measured ratio rather than a theoretical default. For example, if a specific tumor type exhibits a 10-fold upregulation of the parent tRNA relative to the fragment, the simulation is adjusted to penalize off-target binding more heavily. The method concludes by calculating a specificity score specific to the target environment to select a complementary oligonucleotide that is optimized to distinguish the target from the parent under those specific" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT physiological conditions. This capability allows researchers and clinicians to design reagents that are chemically tuned to the specific abundance levels found in a particular disease state (e.g., oncological, neurological, or renal conditions), thereby maximizing efficacy while minimizing off-target binding in vivo.

[0216] In embodiments herein described wherein the framework of the disclosure is used in connection with practical applications in fields such as therapeutics, diagnostics, biological engineering, and functional genomics, the methods of the disclosure can be performed in combination with corresponding systems identifiable by a skilled person.

[0217] Systems to implement methods of the disclosure wherein the framework is used in connection with practical applications can be provided in the form of kits of parts configured to facilitate the modulation of RNA regulatory networks. In a kit of parts, the engineered nucleic acid components — specifically the targeting moieties (e.g., gRNAs) optimized for nuclease-mediated cleavage and the complementary oligonucleotides (e.g., ASOs or LNAs) optimized for fragmentspecific binding — along with the necessary enzymatic reagents to perform the cleavage or binding reactions, can be comprised in the kit independently. These components can be included in one or more compositions, and each construct or component can be in a composition together with a suitable vehicle.

[0218] The term "vehicle" as used herein indicates any of various media acting usually as solvents, carriers, binders, or diluents for the functionalized nucleic acid monomers and related protein components (such as the Cas nuclease) that are comprised in the composition as an active ingredient.

[0219] The term 'pharmaceutically acceptable carrier' or 'excipient' as used herein refers to a nontoxic, inert solid, semi-solid or liquid filler, diluent, encapsulating material or formulation auxiliary of any type. Exemplary pharmaceutically acceptable carriers include, but are not limited to, sugars such as lactose, glucose and sucrose; starches such as corn starch and potato starch; cellulose and its derivatives such as sodium carboxymethyl cellulose, ethyl cellulose and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients such as cocoa butter and suppository waxes; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, com oil and soybean oil; glycols such as propylene glycol; esters such as ethyl oleate and ethyl laurate; agar; buffering" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT agents such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer's solution; ethyl alcohol, and phosphate buffer solutions, as well as other non-toxic compatible lubricants such as sodium lauryl sulfate and magnesium stearate. Furthermore, releasing agents, coating agents, sweetening, flavoring and perfuming agents, preservatives and antioxidants can also be present in the composition, according to the judgment of the formulator.

[0220] In particular, the composition including the designed targeting moieties, complementary oligonucleotides, or Cas ribonucleoproteins can be used in one of the methods or systems herein described for therapeutic administration, diagnostic analysis, or molecular engineering.

[0221] In some embodiments, systems to implement methods of the disclosure can comprise a library of computationally selected targeting moieties (e.g., gRNAs), a supply of RNA-guided nuclease (provided as purified protein, mRNA, or expression vectors), and related functionalized high-affinity nucleic acid analogs (such as LNAs). The systems can further comprise agents for facilitating the delivery of these components into target cells, such as lipid nanoparticles or transfection reagents, and devices for highlighting successful target modulation.

[0222] In some embodiments, systems to implement methods of the disclosure can comprise the functionalized complementary oligonucleotide or targeting moiety herein described, specific reaction buffers, and for each of the targeted fragments (e.g., tDRs), a positive control reagent (such as a synthetic fragment mimic) and a negative control reagent (such as a scrambled sequence). The system may optionally include agents for attaching ligands or detectable moieties to the nucleic acid strands, and optional devices to indicate or measure the localization of the nuclease or the inhibitor to the specific cellular compartment containing the target parent nucleic acid or fragment.

[0223] In some embodiments of the systems, the "reaction buffers" can comprise specific counterions and salts, such as magnesium chloride, which can be used for example in embodiments where the biomolecular targets comprise Cas nucleases that require specific ionic conditions for catalytic activation and conformational stability during the cleavage event.

[0224] In some embodiments of the systems, hybridization or annealing agents can be present in" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT the system when the system is directed to perform diagnostic detection steps, such as Northern blotting or in situ hybridization. These agents facilitate the specific annealing of the designed complementary oligonucleotide probe to the target fragment while maintaining thermodynamic stringency to prevent cross-hybridization with the parent nucleic acid (e.g., tRNA), as will be understood by a skilled person.

[0225] Additional components of the systems can include labeled polynucleotides serving as internal standards, labeled antibodies for detecting Cas protein expression, labels for modifying the complementary oligonucleotide probes, reference standards comprising purified parent RNA or fragments, and additional components identifiable by a skilled person upon reading of the present disclosure.

[0226] The terms "label" and "labeled molecule" as used herein refer to a molecule capable of detection, including but not limited to radioactive isotopes, fluorophores, chemiluminescent dyes, chromophores, enzymes, enzyme substrates, enzyme cofactors, enzyme inhibitors, dyes, metal ions, nanoparticles, metal sols, ligands (such as biotin, avidin, streptavidin or haptens) and the like. The term "fluorophore" refers to a substance or a portion thereof which is capable of exhibiting fluorescence in a detectable image. As a consequence, the wording "labeling signal" as used herein indicates the signal emitted from the label that allows detection of the label, including but not limited to radioactivity, fluorescence, chemiluminescence, production of a compound in outcome of an enzymatic reaction and the like.

[0227] In embodiments herein described of the systems, the components of the kit can be provided, with suitable instructions and other necessary reagents, in order to perform the methods here disclosed. The kit will normally contain the compositions in separate containers. Instructions, for example written or audio instructions, on paper or electronic support such as tapes, CD-ROMs, flash drives, or by indication of a Uniform Resource Locator (URL) which contains access to the computational algorithms (tBOND-G and tBOND-L) or a PDF copy of the instructions for carrying out the assay and interpretation of the virtual gel simulations, will usually be included in the kit. The kit can also contain, depending on the particular method used, other packaged reagents and materials (i.e., wash buffers, lysis buffers, and the like).

[0228] To ensure reproducibility of the predictive capabilities described herein, specific" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT parameters of the machine learning architecture utilized in the tBOND-G algorithm are defined. In the preferred embodiment, the classifier is built using a Support Vector Machine (SVM) architecture. Unlike linear classifiers, the system utilizes a Radial Basis Function (RBF) kernel to capture the non-linear decision boundaries inherent in RNA thermodynamic behavior. The model is configured with specific hyperparameters, including a regularization parameter (C) set to 1.0 and a kernel coefficient (gamma) set to 'scale' (wherein gamma is calculated as 1 / (n_features * X.varQ)). These specific settings were empirically determined to provide the optimal balance between bias and variance, allowing the model to generalize effectively across different structured parent RNA isotypes without overfitting to the training data.

[0229] The training regimen for the model involved a supervised learning approach using a curated dataset of experimentally verified cleavage events. The data was partitioned using a stratified split, designating 80% of the dataset for training the model weights and 20% for validation testing to assess predictive accuracy. The feature vectors serving as input for the SVM comprised the two primary biophysical components calculated by the thermodynamic engine: the Accessibility Score (representing structural exposure) and the Delta G Ratio (representing binding energy relative to the theoretical maximum). By mapping these features into a high-dimensional space, the RBF kernel enables the segregation of "functional" vs. "non-functional" targeting moieties with a precision that simple linear thresholding cannot achieve.

[0230] Furthermore, the computational efficiency of the framework allows for its integration into high-throughput discovery pipelines. For the tBOND-G algorithm (cleavage induction), which requires complex molecular dynamics simulations to generate the " Virtual Gel," the system processes approximately 200 candidate sequences within 20 minutes on a standard workstation, averaging a processing delay of roughly 1 second per step. In contrast, the tBOND-L algorithm (specificity scoring), which solves a defined set of equilibrium equations, operates at a significantly higher velocity, capable of screening approximately 1,000 candidate complementary oligonucleotides per minute. This differential in processing speed reflects the distinct computational burdens of modeling kinetic cleavage events versus thermodynamic binding equilibria, yet both remain sufficiently rapid to enable the real-time design of reagents for genomewide applications." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0231] While the preferred embodiment described above utilizes a Support Vector Machine (SVM) to classify candidate sequences, the teachings of the present disclosure are not limited to this specific architecture. It is contemplated that other machine learning frameworks may be employed to execute the tBOND-G and tBOND-L algorithms, provided they are trained on the same foundational biophysical features (Accessibility Scores and Thermodynamic Binding Ratios) defined herein.

[0232] In one alternative embodiment, the predictive engine is implemented using a Random Forest classifier.

[0233] This ensemble learning method constructs a multitude of decision trees during training (e.g., n_estimators = 500) and outputs the class that is the mode of the classes of the individual trees. This architecture is particularly advantageous for handling non-linear interactions between the accessibility and binding energy parameters without requiring the extensive hyperparameter tuning often associated with kernel-based methods. In this configuration, the feature importance metrics generated by the forest can be used to dynamically weight the input parameters, potentially assigning higher predictive value to "seed region" accessibility in specific parent RNA contexts.

[0234] In another alternative embodiment, the system utilizes a Artificial Neural Network (ANN) or a Deep Learning framework. Specifically, a Multi-Layer Perceptron (MLP) architecture comprising an input layer for the biophysical feature vectors, multiple hidden layers (e.g., three layers with 64, 32, and 16 neurons respectively) utilizing Rectified Linear Unit (ReLU) activation functions, and a final sigmoid output layer can be employed. This approach allows for the modeling of highly complex, high-order dependencies between the thermodynamic stability of the [Targeting Moiety: Fragment] complex and the local secondary structure of the parent nucleic acid. For large-scale genomic applications, a Convolutional Neural Network (CNN) may be utilized, wherein the input is treated as a one-hot encoded sequence matrix, allowing the model to learn local motif-based features (such as specific nucleotide preferences at the cleavage site) directly from the raw sequence data in conjunction with the pre-calculated thermodynamic scores.

[0235] Furthermore, Gradient Boosting algorithms, such as XGBoost or LightGBM, may be utilized to optimize the classification boundary. These iterative techniques build strong predictive models from an ensemble of weak learners by optimizing a differentiable loss function. In the" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT context of the present disclosure, a gradient boosting framework is particularly effective for maximizing the " Specificity" score in the tBOND-L algorithm, as it can iteratively penalize falsepositive predictions (i.e., candidates that bind the parent RNA) during the training phase, thereby fine-tuning the decision boundary to prioritize safety and strictly exclude off-target binders.

[0236] In some embodiments, the computational frameworks described herein are not static but operate within an Active Learning ecosystem. The system is configured to execute an iterative feedback loop wherein the physical outputs of the experimental validation (e.g., the specific migration bands observed on the Northern Blot or the quantitative Ctvalues from RT-qPCR) are fed back into the training dataset.

[0237] Specifically, upon the completion of an experimental cycle (as described in Examples 1-4), the " Virtual Gel" predictions are computationally compared against the digitized experimental images. Discrepancies — such as a predicted cleavage event that failed to occur (false positive) or an unpredicted fragment that appeared (false negative) — are tagged as "high-value" training examples. The processor then triggers a re-training module that adjusts the weights of the Support Vector Machine (SVM) or updates the decision nodes of the Random Forest. This process allows the algorithm to autonomously refine its understanding of the " Accessibility" and " Binding Energy" thresholds specific to different cell types or distinct classes of structured parent nucleic acids, thereby progressively increasing the predictive accuracy of the tool over successive design cycles.

[0238] In embodiments utilizing the tBOND-L algorithm for the design of complementary oligonucleotides (e.g., inhibitors or probes), the processor is further configured to optimize the chemical modification topology of the candidate sequence. The system recognizes that the placement of high-affinity analogs — such as Locked Nucleic Acids (LNA), 2’-O-Methyl (2'OMe), or Phosphorothioate (PS) bonds — dramatically alters the thermodynamic parameters of hybridization.

[0239] Instead of outputting a simple nucleotide string (e.g., " A-G-C-T"), the algorithm generates a specific " Gapmer" or " Mixmer" pattern. The processor iterates through various modification permutations (e.g., placing LNA bases at the 5' and 3' wings while leaving a central DNA gap for RNase H recruitment). For each permutation, the thermodynamic engine recalculates the" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT Specificity Score, specifically penalizing patterns that increase affinity for the parent nucleic acid duplex (off-target) while rewarding patterns that lock the oligonucleotide into a conformation favorable for binding the single-stranded fragment. This results in a "chemically aware" design output that provides the user with the exact synthesis recipe required to achieve the predicted biological effect.

[0240] While the primary embodiments described herein utilize equilibrium thermodynamics (Gibbs Free Energy, AG) to predict binding and cleavage, alternative embodiments of the tBOND framework incorporate Kinetic Modeling parameters. The system acknowledges that in the crowded intracellular environment, the rate of binding (kon) and dissociation (koff) can be as critical as the final stability.

[0241] In this embodiment, the processor is configured to calculate the energy barrier of strand displacement. When targeting a fragment that is sequestered within a parent RNA's secondary structure, the targeting moiety must first nucleate a transient interaction (often at a "toehold" region) and then sequentially displace the native stem. The algorithm estimates the activation energy required for this invasion step. Candidates that exhibit a high equilibrium affinity (AG) but an insurmountable kinetic barrier (e.g., no accessible toehold to initiate binding) are penalized or discarded. This kinetic filter is particularly valuable for designing diagnostic probes, ensuring that the readout occurs within a relevant timescale (minutes) rather than requiring hours to reach thermodynamic equilibrium.

[0242] In certain embodiments, the computational methods described herein are implemented within a client-server architecture or a cloud computing environment (SaaS). In this configuration, the computational calculations (specifically the thermodynamic ensemble calculations performed by the NUPACK module and the iterative "sliding window" scoring) may be offloaded to a remote high-performance computing (HPC) cluster. [2]

[0243] The end-user interacts with the system via a lightweight web-based interface on a client device (e.g., laptop, tablet). The user inputs the target parent RNA sequence or accession number, and the request is transmitted over a network to the server. The server executes the tBOND-G or tBOND-L algorithms and transmits the processed results (specifically the " Virtual Gel" images and the " Efficiency vs. Specificity" scatter plots) back to the client device for rendering. This" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT distributed approach ensures that the sophisticated biophysical modeling is accessible to researchers without requiring specialized local hardware or software installation.

[0244] The system further comprises a specialized Graphical User Interface (GUI) configured to facilitate the intuitive selection of optimal candidates. The GUI presents the " Target Performance Range" (as described in FIG. 11) not as a static image, but as an interactive data visualization.

[0245] Users can hover over specific data points (representing individual complementary oligonucleotides or targeting moieties) to view instantaneous "tooltips" displaying the detailed thermodynamic parameters, sequence composition, and predicted off-target risks. Furthermore, the GUI includes a " Lasso Selection Tool" allowing the user to draw a boundary around a specific cluster of high-efficiency candidates. Upon selection, the system automatically retrieves the corresponding sequences and generates a downloadable " Synthesis Manifest" file (e.g., in CSV or FASTA format) formatted for direct submission to a nucleic acid synthesis provider.

[0246] In a further embodiment representing a fully integrated " Design-to-Build" workflow, the output of the tBOND algorithms is directly coupled to an Automated Nucleic Acid Synthesizer.

[0247] Upon the user's confirmation of the selected library, the processor generates machine-readable instructions (e.g., a phosphoramidite dispensing protocol) and transmits these signals to a connected synthesis platform. This physical integration transforms the digital sequence strings generated by the model directly into tangible chemical matter (the physical gRNA or LNA-ASO molecules) without manual transcription. This embodiment explicitly bridges the gap between in silica computation and physical manufacturing, underscoring the practical, industrial application of the invention.

[0248] The broad capabilities of the tBOND-Integrated platform and its constituent algorithms (tBOND-G and tBOND-L) are further illustrated in the following Examples. These examples comprise experimental validation of the computational principles described herein, demonstrating the transformation of theoretical biophysical models into functional biological tools. While the following experiments specifically utilize transfer RNA (tRNA) as the model structured parent nucleic acid and Cast 3b as the model effector nuclease, it should be understood that these selections serve as a representative "stress test" for the system. tRNAs are among the most" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT structurally complex, heavily modified, and thermodynamically stable RNAs in the cell; therefore, success in this challenging context validates the system's applicability to other classes of structured single-stranded RNAs (e.g.. viral genomes, IncRNAs, riboswitches) as described in the preceding embodiments.EXAMPLES

[0249] The computational frameworks and methods for designing high- specificity nucleic acids targeting tRNAs and their derivatives are further illustrated in the following examples, which are provided by way of illustration and are not intended to be limiting.

[0250] A skilled person will be able to identify additional implementations of the tBOND algorithms herein described in view of the content of the present disclosure. The following specific examples are given to illustrate the practice of the invention, demonstrating both the experimental validation of the predictive models and the practical application of the framework for molecular engineering, but are not to be considered as limiting the invention in any way.

[0251] In particular, exemplary gRNA and complementary oligonucleotide designs and related products, compositions, methods and systems for the modulation of tRNA biology are described in connection with specific experimental tests and procedures. A skilled person will be able to understand and identify the modifications required to adapt the results illustrated in the exemplary embodiments of this section — such as the successful cleavage of mammalian tRNAs or the specific inhibition of pathogenic fragments — to additional embodiments of RNA targeting within the scope of the present disclosure.

[0252] The broad capabilities of the tBOND-Integrated platform and its constituent algorithms (tBOND-G and tBOND-L) are further illustrated in the following Examples. These examples provide a rigorous experimental validation of the computational principles described herein, demonstrating the transformation of theoretical biophysical models into functional biological tools. While the following experiments specifically utilize transfer RNA (tRNA) as the model structured parent nucleic acid and Cast 3b as the model effector nuclease, it should be understood that these selections serve as a representative "stress test" for the system. tRNAs are among the most structurally complex, heavily modified, and thermodynamically stable RNAs in the cell;" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT therefore, success in this challenging context validates the system's applicability to other classes of structured single-stranded RNAs (e.g., viral genomes, IncRNAs, riboswitches) as described in the preceding embodiments.

[0253] The experimental validation strategy detailed below was designed to isolate and verify the individual contributions of the key parameters driving the machine learning models. Example 1 investigates the " Accessibility" parameter, testing the hypothesis that targeting moieties directed at single-stranded loops are significantly more effective than those targeting rigid stems. Example 2 interrogates the " Binding Energy" parameter, confirming that high thermodynamic affinity can compensate for moderate structural barriers. Example 3 serves as a critical negative control, validating the specificity of the decision boundary by demonstrating that candidates falling outside the predicted performance range fail to induce cleavage. Finally, Example 4 demonstrates the generalizability of the framework by applying the same algorithmic logic to a distinct target sequence, confirming that the learned physical rules are universal rather than target- specific.

[0254] All validations were performed in a relevant biological environment using human HEK293 cells. The readouts compare the system's " Virtual Gel" predictions (generated entirely in silico prior to experimentation) against actual physical Northern Blot results. The high degree of concordance observed between the predicted and actual fragment sizes and intensities serves as a reduction to practice of the claimed computational methods. The following examples are intended to illustrate the invention and are not to be construed as limiting the scope of the appended claims.Example 1: Validation of tBOND-G gRNAs Selected from a Region of High Structural Accessibility

[0255] This example establishes the validity of the tBOND-G computational framework by testing the hypothesis that structural accessibility is a primary determinant of Casl3 cleavage efficiency. The experiment was designed to verify whether guide RNAs (gRNAs) selected solely from the "accessibility-dominant" cluster — characterized by target sites located in unstructured loops of the tRNA — would consistently yield cleavage products, thereby validating the machine learning model's reliance on secondary structure inputs.

[0256] Accordingly, this example shows the validation of the " Accessibility" parameter within" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT the tBOND-G algorithm. The objective was to confirm that targeting moieties (gRNAs) directed against predicted single- stranded loops in the target RNA structure induce efficient cleavage, whereas those targeting rigid stems show reduced efficacy, and to verify the predictive accuracy of the " Virtual Gel" simulation.

[0257] To ensure uniform background activity for the experiment, stable expression of the effector nuclease was first established. Specifically, HEK293 cells were transduced with lentiviral particles encoding the pspCas13b nuclease. Following antibiotic selection to isolate successful integrants, a stable Casl3-expressing HEK.293 cell line was established and maintained in standard growth medium comprising DMEM supplemented with 10% FBS. Concurrently, a specific endogenous transfer RNA (tRNA) was selected as the structured parent nucleic acid target. The tBOND-G algorithm was executed to generate a library of candidate gRNAs targeting this tRNA to prepare for experimental validation.

[0258] The transfection protocol commenced by seeding the Casl3-expressing HEK293 cells into 6-well culture plates, where they were grown to approximately 70% confluency. Cells were then transfected with plasmids encoding specific candidate gRNAs selected by the algorithm, representing distinct " Accessibility-Dominant" and " Low-Accessibility" clusters to test the algorithmic predictions. The transfection was performed using a dosage of 1 pg of gRNA plasmid DNA per well. Subsequently, the transfected cells were incubated at 37°C with 5% CO2 for a period of 48 hours to allow sufficient time for gRNA expression, complex formation, and the subsequent Cas-mediated cleavage of the endogenous target.

[0259] At the conclusion of the 48-hour incubation period, the cells were harvested, and total RNA was extracted utilizing a standard phenol-chloroform (Trizol) or column-based extraction method. The abundance of the full-length parent tRNA and the generation of specific fragments (tDRs) were analyzed via Northern Blotting using radiolabeled probes specific to the 5' and 3' ends of the target tRNA. To quantify the success of the model, the physical migration bands observed on the Northern Blot were directly compared against the " Virtual Gel" predictions generated by the tBOND-G software. A prediction was considered accurate if the experimentally observed fragment size matched the computer- predicted size within a tolerance of ±2 nucleotides, thereby confirming the algorithm's ability to precisely map the cleavage site based on the binding pocket" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT geometry.

[0260] The results of this validation are illustrated in FIG. 5. As shown in the figure, the experimental validation of the tBOND-G system focuses on candidates predicted to be highly accessible. The SVM model's prediction map is presented in panel (a), with the selected gRNAs explicitly located in a boxed region representing the high predicted accessibility cluster. Panel (b) provides a direct comparison between the system's virtual gel predictions (schematic bars at the top of each gel image) and the actual Northern blot experimental results (bands at the bottom of each gel image) for both the resulting 3' tDRs and 5' tDRs.

[0261] Therefore. Northern Blot data (see FIG. 5) confirmed that gRNAs targeting accessible loops (as identified by the high accessibility scores in the tBOND-G output) resulted in distinct, strong bands corresponding to the predicted fragment sizes. In contrast, candidates targeting inaccessible stems showed little to no fragmentation. The close alignment (within ±2 nt) between the in silica Virtual Gel and the in vitro Northern Blot confirms the predictive power of the model.

[0262] The data presented in FIG. 5 shows a high degree of correspondence between the computational predictions and the biological outcomes. Specifically, the results indicate that 15 gRNAs were selected from the accessibility-dominant cluster out of a total library of 965 candidates. Remarkably, 100% of these selected gRNAs produced observable cleavage bands on the Northern blot, confirming that the algorithm successfully identified functional sequences.

[0263] Furthermore, the accuracy of the fragment size prediction confirms the mechanistic validity of the model. For the 3' tRNA-derived fragments (tDRs), 13 out of 15 experimental bands matched the virtual gel prediction, yielding an accuracy rate of approximately 81%. For the 5' tDRs, 11 out of 15 bands matched the prediction, yielding an accuracy of 63%. These results support several key conclusions: (1) structural accessibility is indeed a critical physical parameter for Casl3 activity; (2) the tBOND-G SVM model can effectively filter non-functional candidates; and (3) the " Virtual Gel" feature provides a reliable, actionable guide for researchers to identify specific cleavage products prior to experimentation.Example 2: Validation of tBOND-G gRNAs Selected from a Region of High Thermodynamic Binding Energy" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0264] This example provides further validation of the tBOND-G computational framework, specifically testing the predictive power of thermodynamic binding stability as a primary driver of Casl3 cleavage efficiency. While Example 1 focused on structural accessibility, this experiment selected guide RNAs (gRNAs) from a distinct "binding-energy-dominant" cluster identified by the Support Vector Machine (SVM) model. The objective was to verify whether gRNAs with high predicted binding energy (represented by a high Delta G Ratio) could effectively drive cleavage even in regions where structural accessibility might be less than optimal or moderate. This test effectively asks whether strong thermodynamic affinity can overcome the steric hindrance of the tRNA tertiary structure.

[0265] This example was designed to interrogate the " Binding Energy" parameter of the tBOND-G algorithm, specifically testing the hypothesis that high thermodynamic affinity can drive effective cleavage even when structural accessibility is moderate. To ensure a rigorous comparison with the accessibility-driven results observed in the previous experiment, the experimental setup utilized transfection conditions identical to those described in Example 1. Specifically, the same plasmid concentrations (1 pg per well) and incubation duration (48 hours) were employed, with the only variable being the specific nucleotide sequence of the gRNA seed region, which was computationally designed by the tBOND-G algorithm to target distinct regions of the structured parent nucleic acid.

[0266] The study focused on a specific "high-energy" cluster of candidates identified by the Support Vector Machine (SVM) model. This cluster was defined as a dense population of gRNA candidates exhibiting a calculated Delta G Ratio greater than 0.6, representing sequences with exceptional thermodynamic stability relative to the theoretical maximum. The objective was to determine if this enhanced binding energy could compensate for the energy penalty associated with invading more structured regions of the tRNA target.

[0267] The results of this validation are illustrated in FIG. 6. Similar to the presentation in the previous example, panel (a) displays the SVM model’s prediction map. In this instance, however, the overlay indicates the location of the selected gRNAs within the shaded region corresponding to high predicted binding energy. Panel (b) presents the direct comparison between the computational prediction and biological reality, showing the virtual gel predictions (schematic bars" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT at the top) aligned with the experimental Northern blot results (bands at the bottom) for both the resulting 3' and 5' tDR fragments.

[0268] Therefore, the Northern Blot analysis confirmed that candidates from this high-energy cluster successfully induced cleavage, generating fragment bands that matched the " Virtual Gel" predictions. Furthermore, the analysis addressed potential concerns regarding specificity trade-offs associated with such high-affinity binders. Critically, the data did not reveal any correlation between these high-affinity binders and increased off-target effects; specificity was primarily maintained through the precise matching of the seed region to the target, confirming that the algorithm can optimize for potency (Delta G) without compromising the safety profile of the reagent.

[0269] The comparison in FIG.6 reveals a strong agreement between the virtual simulations and the physical blot results. Green check marks in the figure indicate that the selected gRNAs successfully directed the Cast 3 nuclease to the target, resulting in effective fragmentation at the predicted sites. The alignment of the experimental bands with the predicted sizes in the virtual gel confirms the accuracy of the model for energy-driven cleavage mechanisms. This result is biologically significant as it demonstrates that high thermodynamic affinity can effectively compensate for structural barriers, allowing the tBOND-G system to identify functional gRNAs even in regions of the tRNA molecule that are not perfectly accessible. This complements the findings of Example 1, establishing that the SVM model correctly integrates and weights both accessibility and binding energy to predict cleavage success across diverse biochemical contexts.Example 3: Experimental Validation of tBOND-G gRNAs Selected from a Region of Medium Accessibility and Medium Binding Energy

[0270] This example evaluates the discriminatory power of the tBOND-G computational framework, specifically testing its ability to predict experimental failure. While Examples 1 and 2 focused on identifying positive "hits," a robust predictive model must also accurately identify and filter out non-functional sequences to prevent wasted experimental effort. To assess this capability, guide RNAs (gRNAs) were selected from a "balanced" cluster characterized by medium structural accessibility and medium thermodynamic binding energy. In this intermediate zone, the Support Vector Machine (SVM) model predicted a low cleavage efficiency (approximately 0.1 on the" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT normalized scale), hypothesizing that neither parameter was sufficient on its own or in combination to drive the conformational activation of the Casl3 nuclease.

[0271] This example served as a negative control assessment designed to validate the specificity of the Support Vector Machine (SVM) decision boundary. The experimental strategy involved selecting a "balanced" cluster of gRNA candidates that the model predicted would be nonfunctional. Geographically, these candidates were located in the low-performance region of the SVM feature map (represented as the lighter shaded zone in the right panel of FIG. 3). Quantitatively, this cluster was defined by candidates falling within a specific "zig-zag" shaped interaction zone characterized by a moderate Delta G Ratio of 0.4 to 0.6 and a moderate Accessibility Score of 0.3 to 0.6.

[0272] The objective was to confirm that sequences possessing these intermediate biophysical characteristics lack the necessary drive to effectuate cleavage. From a theoretical perspective, candidates in this range face a thermodynamic barrier: the binding energy available from the gRNA-target interaction is insufficient to overcome the stability of the native secondary structure (e.g., the tRNA duplex stems) of the structured parent nucleic acid. The Northern Blot results largely confirmed this hypothesis, as the majority of candidates selected from this cluster failed to produce detectable cleavage fragments, thereby validating the model's ability to effectively screen out non-functional sequences and reduce false positives.

[0273] The results of this negative control validation are illustrated in FIG. 7. Panel (a) depicts the SVM prediction map, with the selected gRNAs located in a region distinct from the high-performance clusters of Examples 1 and 2. Panel (b) displays the virtual gel predictions compared against the experimental Northern blot results. Unlike the previous examples, the virtual gel here predicts a lack of significant cleavage products for the majority of the barcodes.

[0274] The experimental data in FIG. 7 strongly confirms the negative predictive value of the model. Out of the 11 gRNAs selected from this balanced region, only one produced a visible cleavage band on the Northern blot. This outcome aligns closely with the model's low predicted cleavage efficiency score of approximately 0.1. By correctly predicting that the majority of these candidates would fail to induce cleavage, the tBOND-G system demonstrates its utility not just as a discovery tool, but as a screening filter. This capability to correctly identify and exclude poor" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT candidates is critical for high-throughput applications, ensuring that resources are focused only on sequences with the highest probability of biological activity.

[0275] Therefore, while the general trend strongly supported the model's predictions, a single gRNA candidate within this cluster did produce a visible cleavage band, representing a false negative result. Although no specific post-hoc analysis was performed to investigate the unique molecular features of this single outlier, its presence highlights the inherent probabilistic nature of biological modeling. However, the absence of activity in the vast majority of this cluster confirms that the defined numerical thresholds (Delta G < 0.6 combined with Accessibility < 0.6) represent a robust exclusion zone for minimizing experimental failure.Example 4: Experimental Validation of the tBOND-G Framework on a Distinct Target (tRNA-Asp-GTC)

[0276] This example establishes the generalizability and robustness of the tBOND-G computational framework. A critical concern in machine learning applications is "overfitting," where a model learns to predict the idiosyncrasies of a specific training dataset rather than the underlying physical rules. To demonstrate that the tBOND-G algorithm is not overfitted to the initial tRNA targets used for training, the entire validation process was repeated for a completely distinct RNA species, tRNA-Asp-GTC. This target possesses a unique nucleotide sequence and secondary structure profile, thereby serving as an independent test case to verify whether the learned biophysical parameters (accessibility and binding energy) remain predictive across diverse RNA substrates.

[0277] This example was conducted to validate the universal applicability of the tBOND-G computational framework. Specifically, the algorithm was applied to a distinct target sequence, tRNA-Asp-GTC, to determine if the physical principles encoded in the software could successfully predict cleavage outcomes across different RNA species. Crucially, this validation employed the exact same pre-trained Support Vector Machine (SVM) model utilized in Examples 1 through 3, without any re-calibration or retraining on Asp-GTC specific data. This approach was chosen to demonstrate the model's compatibility across various tRNA variants (and by extension, other structured RNAs) that share fundamental thermodynamic and structural properties despite possessing significant sequence variations." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT

[0278] To ensure a balanced assessment of the model's predictive power, a cohort of 16 gRNA candidates was selected for testing based on their specific cluster classifications. The selection was stratified to include 8 candidates from the "high-energy" cluster (defined by a Delta G Ratio > 0.6) and 8 candidates from the "high-accessibility" cluster (defined by a Delta G Ratio < 0.45 combined with an Accessibility Score > 0.5). This selection strategy allowed for a simultaneous reverification of both major performance drivers identified in the previous examples.

[0279] The results of this cross-validation are illustrated in FIG. 8. Panel (a) displays the SVM model heatmap generated specifically for the tRNA-Asp-GTC sequence. The overlay highlights the selected gRNA candidates, which are distributed across two distinct high-performance clusters: one characterized by high structural accessibility scores and another by high thermodynamic binding energy (Delta G) values. Both clusters exhibited an average predicted cleavage efficiency score of approximately 0.7. Panel (b) presents the comparative analysis between the Virtual Gel predictions and the experimental Northern Blot results for both the 3' and 5' cleavage fragments.

[0280] The data in FIG. 8 confirms that the predictive power of the tBOND-G model is broadly applicable. The experimental results demonstrated a 100% accuracy rate for predicting cleavage events; every gRNA selected by the model successfully induced cleavage of the tRNA-Asp-GTC target. Furthermore, the model achieved exceptional precision in predicting the specific fragment sizes. For the 3' tDRs, 14 out of 16 observed bands matched the virtual gel predictions, resulting in an 88% accuracy rate. For the 5' tDRs, the model achieved 100% accuracy, with all predicted bands aligning perfectly with the experimental observations. This high level of fidelity across a novel target sequence demonstrates that the algorithm relies on fundamental biophysical principles of RNA-Casl3 interactions rather than sequence- specific artifacts, validating its utility as a universal design tool for tRNA modulation.

[0281] The experimental results discussed in this example thus provided further insights into the biological dynamics of the generated fragments. In this specific case, it was observed that the 5' fragment (5' tDR) of tRNA-Asp-GTC generally exhibits higher intracellular abundance than the 3' fragment under most stress conditions. Despite this biological variance in fragment stability, the system did not observe any significant difference in detection limits compared to the previous examples, confirming that the " Virtual Gel" and the underlying cleavage predictions remain robust" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT even when applied to targets with differing in vivo processing profiles.Example 5: Validation of tBOND-L Specificity in a Competitive Environment (Prophetic)

[0282] This prophetic example describes the validation of the " Specificity" scoring metric of the tBOND-L algorithm. The primary objective is to demonstrate that a computationally selected complementary oligonucleotide (e.g., an LNA-ASO) can selectively bind a free target fragment (tDR) while exhibiting negligible binding to the full-length parent nucleic acid (tRNA), despite the target sequence being 100% homologous to the parent. To achieve this, the experimental design utilizes two synthetic RNA oligonucleotides: one representing the 20-nucleotide 3' tDR of tRNA-Asp-GTC (Target) and one representing the full-length 75-nucleotide tRNA-Asp-GTC (Parent).

[0283] The tBOND-L algorithm is executed to design two distinct LNA-ASO candidates for comparison. " Candidate A" is selected from the " Target Range" quadrant (High Efficiency / High Specificity) and is predicted to bind only the open 3' end of the tDR. In contrast, " Candidate B" is selected from a low-specificity cluster and is predicted to cross-react with the parent due to targeting an exposed loop in the full tRNA structure. To quantify the interaction, an electrophoretic gel shift assay is utilized, wherein the LNA-ASO candidates are 5'-labeled with Biotin. The protocol involves incubating the Biotin-labeled ASO with increasing concentrations of the Target tDR to determine the dissociation constant ( KD), followed by incubation with increasing concentrations of the Full-Length Parent tRNA to assess off-target affinity.

[0284] It is anticipated that Candidate A (High Specificity) will exhibit a strong binding curve (low KD) with the Target tDR but a flat line (no binding) with the Parent tRNA, even at high concentrations. This result would confirm that the thermodynamic penalty calculated by the algorithm — preventing invasion of the parent's stable stem — effectively blocks off-target binding. Conversely, Candidate B is expected to show binding to both the Target tDR and the Parent tRNA, confirming the algorithm's prediction that this sequence lacks specificity and justifying its exclusion from the selected library.Example 6: Therapeutic Rescue of Translational Suppression via tBOND-L (Prophetic)

[0285] This prophetic example illustrates the therapeutic utility of the tBOND-L platform, aiming to demonstrate that an optimized LNA inhibitor can reverse a cellular phenotype caused by a" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT pathogenic fragment without disrupting the essential function of the parent tRNA. In this experimental model, U2OS cells are treated with Sodium Arsenite (0.5 mM for 1 hour) to induce oxidative stress, a condition known to trigger the endogenous cleavage of tRNAs into 5' tiRNAs (tRNA-derived stress-induced RNAs) that repress global protein translation.

[0286] Following the stress induction, cells are transfected with either a control " Scrambled LNA" or a specific "tBOND-L Inhibitor," the latter being an LNA designed to specifically target the 5' tiRNA-Ala-AGC fragment with a high Specificity Score. The phenotypic rescue is subsequently measured using an O-propargyl-puromycin (OP-Puro) incorporation assay followed by flow cytometry to quantify global protein synthesis rates. Simultaneously, the levels of the mature tRNA-Ala-AGC parent are quantified via northern blot to ensure the inhibitor specifically silences the 5’tiRNA-Ala-AGC but does not degrade the essential parent pool.

[0287] The study anticipates that cells treated with the tBOND-L Inhibitor will show a significant recovery of protein synthesis rates compared to the Control group, indicating successful neutralization of the repressive fragment. Crucially, the levels of the mature tRNA-Ala-AGC parent are expected to remain stable (comparable to non-stressed baseline). This outcome would confirm that the inhibitor achieved "functional decoupling," effectively neutralizing the pathogenic fragment without depleting the housekeeping precursor necessary for cell viability.Example 7: Diagnostic Detection of a Viral RNA Fragment (Prophetic)

[0288] This example extends the tBOND-L framework to a non-tRNA target, specifically distinguishing a bioactive viral RNA fragment (sfRNA) from the full viral genome. The study focuses on Flaviviruses (e.g., Dengue, Zika), which produce a stable subgenomic flaviviral RNA (sfRNA) that accumulates in infected cells and is essential for pathogenicity. The sequence of the Dengue Virus (DENV) 3' UTR is utilized as the " Parent" input, and the known sfRNA sequence is input as the " Fragment".

[0289] The tBOND-L algorithm is employed to design a " Molecular Beacon" probe, filtering for sequences that can hybridize to the sfRNA but are sequestered within the complex secondary structure of the full-length viral 3' UTR. A "virtual test tube" simulation is then executed to compare the signal-to-noise ratio of the tBOND-L designed beacon against a standard linear probe" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT designed by conventional software.

[0290] The results are expected to demonstrate that the Standard Probe generates a high background signal due to significant cross -hybridization with the full-length viral genome (Parent). In contrast, the tBOND-L Probe is predicted to exhibit high specificity, generating a fluorescence signal only in the presence of the free sfRNA fragment. This validation would demonstrate the algorithm's broader utility for developing precision diagnostics for infectious diseases where distinguishing processed viral intermediates from genomic RNA is critical.

[0291] The examples set forth above are provided to give those of ordinary skill in the art a complete disclosure and description of how to make and use the embodiments of framework for the high-efficiency, high- specificity design of nucleic acids targeting tRNAs and their derivatives, and related compositions, devices, methods and systems of the disclosure, and are not intended to limit the scope of what the Applicants regard as their disclosure. Modifications of the abovedescribed modes for carrying out the disclosure can be used by persons of skill in the art and are intended to be within the scope of the following claims.

[0292] The entire disclosure of each document cited (including patents, patent applications, journal articles including related supplemental and / or supporting information sections, abstracts, laboratory manuals, books, or other disclosures) in the Background, Summary, Detailed Description, and Examples is hereby incorporated herein by reference. All references cited in this disclosure (comprising 1-17 at the end of the present specification, the numbers of some of which are also referred to within brackets throughout the present specification) are incorporated by reference to the same extent as if each reference had been incorporated by reference in its entirety individually. However, if any inconsistency arises between a cited reference and the present disclosure, the present disclosure takes precedence.

[0293] The terms and expressions which have been employed herein are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the disclosure claimed. Thus, it should be understood that although the disclosure has been specifically disclosed by preferred embodiments, exemplary embodiments and optional features, modification and variation" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this disclosure as defined by the appended claims.

[0294] Sequence Representation: The nucleotide sequences disclosed herein are presented in accordance with WIPO Standard ST.26. Accordingly, the symbol "t" is used to represent thymine in DNA and uracil in RNA. It is expressly understood that while the computational algorithms described herein (e.g.. tBOND-G and tBOND-L) process and output sequence data utilizing the standard DNA alphabet (A, C, G, T) for software compatibility, the physical molecules designed and claimed herein are RNA or modified RNA species. Therefore, in the context of the present disclosure, any appearance of the symbol 't' within a designated RNA sequence (e.g., a gRNA, tRNA, tDR, or LNA) refers to and denotes 'uracil' or 'u'

[0295] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. The term "plurality" includes two or more referents unless the content clearly dictates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure pertains.

[0296] When a Markush group or other grouping is used herein, all individual members of the group and all combinations and possible sub-combinations of the group are intended to be individually included in the disclosure. Every combination of components or materials described or exemplified herein can be used to practice the disclosure, unless otherwise stated. One of ordinary skill in the art will appreciate that methods, device elements, and materials other than those specifically exemplified can be employed in the practice of the disclosure without resort to undue experimentation. All art-known functional equivalents, of any such methods, device elements, and materials are intended to be included in this disclosure.

[0297] Whenever a range is given in the specification, for example, a temperature range, a frequency range, a time range, or a composition range, all intermediate ranges and all subranges, as well as, all individual values included in the ranges given are intended to be included in the" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT disclosure. Any one or more individual members of a range or group disclosed herein can be excluded from a claim of this disclosure. The disclosure illustratively described herein suitably can be practiced in the absence of any element or elements, limitation or limitations, which is not specifically disclosed herein.

[0298] " Optional" or "optionally" means that the subsequently described circumstance can or cannot occur, so that the description includes instances where the circumstance occurs and instances where it does not according to the guidance provided in the present disclosure. Combinations envisioned can be identified in view of the desired features of the device in view of the present disclosure, and in view of the features that result in the formation.

[0299] A number of embodiments of the disclosure have been described. The specific embodiments provided herein are examples of useful embodiments of the disclosure and it will be apparent to one skilled in the art that the disclosure can be carried out using a large number of variations of the devices, device components, methods steps set forth in the present description. As will be obvious to one of skill in the art, methods and devices useful for the present methods can include a large number of optional composition and processing elements and steps.

[0300] In particular, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claim.REFERENCES1. Chan, Patrick P., and Todd M. Lowe. “GtRNAdb 2.0: An Expanded Database of Transfer RNA Genes Identified in Complete and Draft Genomes.” Nucleic Acids Research, vol. 44, no. DI, 2016, pp. D184-D189, Oxford University Press.2. Fornace, Michael E., et al “NUPACK: Analysis and Design of Nucleic Acid Structures, Devices, and Systems.” NUPACK, 2022.3. Fu, Miao, et al.“Emerging Roles of tRNA-Derived Fragments in Cancer.” Molecular Cancer, vol. 22, no. 1, 2023, p. 30, BioMed Central.4. Huang, J., Rauscher, S., Nawrocki, G., Ran, T., Feig, M., De Groot, B. L., Grubmüller, H., and MacKerell Jr., A. D. “CHARMM36m: An Improved Force Field for Folded and Intrinsically" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT Disordered Proteins.” Nature Methods, vol. 14, no. 1, 2017, pp. 71-73, Nature Publishing Group.5. Karlin, S., and Altschul, S. F. “Applications and Statistics for Multiple High-Scoring Segments in Molecular Sequences.” Proceedings of the National Academy of Sciences, vol. 90, no. 12, 1993. pp. 5873-5877.6. Karlin, S., and Altschul, S. F. “Methods for Assessing the Statistical Significance of Molecular Sequence Features by Using General Scoring Schemes.” Proceedings of the National Academy of Sciences, vol. 87, no. 6, 1990, pp. 2264-2268.7. Myers, E. W., and Miller, W. “Optimal Alignments in Linear Space.” Computer Applications in the Biosciences (CABIOS), vol. 4, no. 1, 1988, pp. 11-17.8. Needleman, S. B., and Wunsch, C. D. “A General Method Applicable to the Search for Similarities in the Amino Acid Sequence of Two Proteins.” Journal of Molecular Biology, vol. 48, no. 3. 1970, pp. 443-453.9. Pandey, K. K., et al. “Regulatory Roles of tRNA-Derived RNA Fragments in Human Pathophysiology.” Molecular Therapy-Nucleic Acids, vol. 26, 2021, pp. 161-173, Cell Press. 10. Pearson, W. R., and Lipman, D. J. “Improved Tools for Biological Sequence Comparison.” Proceedings of the National Academy of Sciences, vol. 85, no. 8, 1988, pp. 2444-2448.11. Ruff, K. M., and Pappu, R. V. “AlphaFold and Implications for Intrinsically Disordered Proteins.” Journal of Molecular Biology, vol. 433, no. 20, 2021, p. 167208, Elsevier.12. Saikia, Moushumi, and Miltos Hatzoglou.“The Many Virtues of tRNA-Derived Stress-Induced RNAs (tiRNAs): Discovering Novel Mechanisms of Stress Response and Effect on Human Health.” Journal of Biological Chemistry, vol. 290, no. 50, 2015, pp. 29761-29768, American Society for Biochemistry and Molecular Biology.13. Slaymaker, Ian M., et al. “High-Resolution Structure of Casl3b and Biochemical Characterization of RNA Targeting and Cleavage.” Cell Reports, vol. 26, no. 13. 2019, pp. 3741-3751, Elsevier.14. Smith, T. F., and Waterman, M. S. “Comparison of Biosequences.” Advances in Applied Mathematics, vol. 2, no. 4, 1981, pp. 482-489.15. Soman, K. P., Loganathan, R., and Ajay, V. Machine Learning with SVM and Other Kernel" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT Methods. PHI Learning Pvt. Ltd., 2009.16. The RNAcentral Consortium. “RNAcentral: A Hub of Information for Non-Coding RNA Sequences.” Nucleic Acids Research, vol. 47, no. DI, 2019, pp. D221-D229, Oxford University Press.17. Zhao, Yu, et al. “The Function of tRNA-Derived Small RNAs in Cardiovascular Diseases.” Molecular Therapy-Nucleic Acids, vol. 35, no. 1, 2024, Cell Press.

Claims

" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT CLAIMS1. A computer-implemented method for designing an engineered targeting moiety for modulating a structured parent nucleic acid, the method comprising:(a) receiving, by a processor of a computer system, a digital data string representing a nucleotide sequence of the structured parent nucleic acid;(b) generating, by the processor, a library of digital data strings, each representing a candidate targeting moiety sequence derived from the structured parent nucleic acid sequence;(c) for each candidate targeting moiety sequence in the library, performing the steps of:(i) retrieving or predicting a secondary structure of the structured parent nucleic acid and calculating, by the processor, a numerical accessibility score for a target site corresponding to the candidate targeting moiety sequence, wherein the accessibility score represents a proportion of unstructured bases at the target site; and(ii) calculating, by the processor, a numerical binding energy value representing a thermodynamic stability of a complex formed by the candidate targeting moiety sequence and the structured parent nucleic acid;(d) inputting the calculated accessibility scores and binding energy values for the library of candidate targeting moiety sequences into a pre-trained Support Vector Machine (SVM) model stored in a memory of the computer system;(e) executing the SVM model by the processor to generate a predicted efficiency score for each candidate targeting moiety sequence; and(f) outputting, on a display device associated with the computer system, a ranked list of candidate targeting moiety sequences based on their predicted efficiency scores.

2. The method of claim 1, wherein the structured parent nucleic acid comprises a nucleotide sequence length ranging from 30 to 10,000 nucleotides, or wherein the structured parent nucleic acid is a defined structural domain within a messenger RNA (mRNA) or long non-coding RNA (IncRNA) having a length of 50 to 500 nucleotides.

3. The method of claim 1, further comprising:" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT (a) identifying, by the processor using a molecular dynamics model, one or more putative cleavage sites on the structured parent nucleic acid for a selected candidate targeting moiety sequence:(b) calculating the lengths of 5' and 3' fragments that would result from cleavage at the identified sites; and(c) generating and displaying a digital image simulating a gel electrophoresis result, wherein the digital image comprises virtual bands corresponding to the calculated lengths of the fragments.

4. The method of claim 1, wherein the SVM model is trained using a dataset comprising experimentally determined modulation results correlated with calculated accessibility scores and binding energy values for a plurality of training sequences.

5. The method of claim 1, wherein the binding energy value is a normalized Delta G ratio calculated using a nucleic acid thermodynamics software package.

6. The method of claim 1, further comprising, for a candidate targeting moiety sequence with a predicted efficiency score above a predetermined threshold:(a) performing, by the processor, a homology search of the candidate sequence against a transcriptome database to identify a set of potential off-target sequences;(b) calculating, for each potential off-target sequence, an off-target score based on a number and location of nucleotide mismatches with the candidate sequence; and(c) including the off-target score in the outputted ranked list.

7. A system for designing an engineered targeting moiety, comprising:(a) a processor; and(b) a memory communicatively coupled to the processor, the memory storing computer-executable instructions that, when executed by the processor, cause the system to perform the method of claim 1.

8. A computer-implemented method for designing a complementary oligonucleotide for specific binding to a target RNA fragment derived from a structured parent nucleic acid, the method" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT comprising:(a) receiving, by a processor of a computer system, a digital data string representing a nucleotide sequence of the structured parent nucleic acid from which the target RNA fragment is derived;(b) generating, by the processor, a library of digital data strings, each representing a candidate complementary oligonucleotide sequence complementary to a region of the target RNA fragment;(c) for each candidate complementary oligonucleotide sequence in the library, performing the steps of:(i) calculating, by the processor, a numerical efficiency score representing a binding affinity of the candidate to the target RNA fragment relative to a binding affinity of the candidate to at least one off-target sequence within the structured parent nucleic acid; and (ii) calculating, by the processor, a numerical specificity score by computationally simulating a competitive binding reaction comprising the candidate, the target RNA fragment, the structured parent nucleic acid, and at least one non-target fragment, wherein the specificity score represents a predicted equilibrium concentration of a complex formed by the candidate and the target RNA fragment relative to all other complexes;(d) ranking, by the processor, the candidate complementary oligonucleotide sequences based on their calculated efficiency scores and specificity scores; and(e) outputting, on a display device, a ranked list of the candidate complementary oligonucleotide sequences.

9. The method of claim 8, wherein the step of calculating the numerical specificity score is performed using a nucleic acid thermodynamics software package that calculates equilibrium concentrations for all molecular species in the simulated competitive binding reaction.

10. The method of claim 8, further comprising generating and displaying a two-dimensional plot, wherein each candidate complementary oligonucleotide sequence is represented as a point plotted according to its calculated efficiency score on a first axis and its calculated specificity score on a second axis." A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT 11. A system for designing a complementary oligonucleotide, comprising:(a) a processor; and(b) a memory communicatively coupled to the processor, the memory storing computer-executable instructions that, when executed by the processor, cause the system to perform the method of claim 8.

12. A computer-implemented method for designing a guide RNA (gRNA) for cleaving a target ribonucleic acid (RNA) molecule, the method comprising:(a) receiving, by a processor of a computer system, a digital data string representing a nucleotide sequence of the target RNA molecule;(b) generating, by the processor, a library of digital data strings, each representing a candidate gRNA sequence derived from the nucleotide sequence of the target RNA molecule;(c) for each candidate gRNA sequence in the library, performing the steps of:(i) determining a secondary structure of the target RNA and calculating, by the processor, a numerical accessibility score for a target site corresponding to the candidate gRNA sequence, wherein the accessibility score represents a proportion of unstructured bases at the target site; and(ii) calculating, by the processor, a numerical binding energy value representing a thermodynamic stability of a complex formed by the candidate gRNA sequence and the target RNA molecule sequence;(d) inputting the calculated accessibility scores and binding energy values for the library of candidate gRNA sequences into a pre-trained machine learning model stored in a memory of the computer system;(e) executing the machine learning model by the processor to generate a predicted cleavage efficiency score for each candidate gRNA sequence; and(f) outputting, on a display device associated with the computer system, a ranked list of candidate gRNA sequences based on their predicted cleavage efficiency scores.

13. The method of claim 12, wherein the target RNA molecule is a structured single- stranded RNA" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT (ssRNA).

14. The method of claim 12 or 13, wherein the target RNA molecule is a transfer RNA (tRNA).

15. The method of claim 14, wherein the target tRNA is a mammalian tRNA having a length of 76 to 90 nucleotides.

16. The method of claim 14, wherein the target tRNA is selected from the group consisting of a viral-encoded tRNA and a viral tRNA-like structure (TLS).

17. The method of any one of claims 12-16, wherein the machine learning model is a Support Vector Machine (SVM) utilizing a Radial Basis Function (RBF) kernel.

18. The method of claim 17. wherein the SVM is trained using a dataset comprising experimentally determined cleavage results correlated with calculated accessibility scores and binding energy values for a plurality of training gRNAs.

19. The method of any one of claims 12-16, wherein the machine learning model is selected from the group consisting of an Artificial Neural Network (ANN), a Random Forest, a Decision Tree, and a Gradient Boosting algorithm.

20. The method of any one of claims 12-19, further comprising:(a) identifying, by the processor using a molecular dynamics model, one or more putative cleavage sites on the target RNA for a selected candidate gRNA sequence;(b) calculating the lengths of 5’ and 3' RNA fragments that would result from cleavage at the identified sites; and(c) generating and displaying a digital image simulating a gel electrophoresis result, wherein the digital image comprises virtual bands corresponding to the calculated lengths of the fragments.

21. The method of any one of claims 12-20, wherein the numerical binding energy value is a normalized Delta G ratio calculated using a nucleic acid thermodynamics software package.

22. The method of any one of claims 12-21, further comprising an off-target analysis step for a candidate gRNA sequence having a predicted cleavage efficiency score above a predetermined" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT threshold, the step comprising:(a) performing, by the processor, a homology search of the candidate gRNA sequence against a transcriptome database to identify a set of potential off- target sequences;(b) calculating, for each potential off-target sequence, an off-target score based on a number and location of nucleotide mismatches with the candidate gRNA sequence; and(c) filtering the candidate gRNA sequence from the ranked list if the off-target score exceeds a safety threshold.

23. The method of any one of claims 12-22, wherein the candidate gRNA sequence comprises a spacer sequence and a scaffold sequence configured for use with a CRISPR-associated (Cas) nuclease selected from the group consisting of Cas 13a. Cas 13b. Cas 13c. Cas 13d, Casl3x, and Casl3y.

24. A computer-implemented method for designing a high-affinity nucleic acid analog for specific binding to a target RNA fragment derived from a structured parent RNA molecule, the method comprising:(a) receiving, by a processor of a computer system, a digital data string representing a nucleotide sequence of the structured parent RNA molecule from which the target RNA fragment is derived; (b) generating, by the processor, a library of digital data strings, each representing a candidate high-affinity nucleic acid analog sequence complementary to a region of the target RNA fragment; (c) for each candidate sequence in the library, performing the steps of:(i) calculating, by the processor, a numerical efficiency score representing a binding affinity of the candidate to the target RNA fragment relative to a binding affinity of the candidate to at least one off-target sequence within the structured parent RNA molecule; and(ii) calculating, by the processor, a numerical specificity score by computationally simulating a competitive binding reaction comprising the candidate, the target RNA fragment, the structured parent RNA molecule, and at least one non-target RNA fragment, wherein the specificity score represents a predicted equilibrium concentration of a complex formed by the candidate and the target RNA fragment relative to all other complexes;" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT (d) ranking, by the processor, the candidate sequences based on their calculated efficiency scores and specificity scores; and(e) outputting, on a display device, a ranked list of the candidate sequences.

25. The method of claim 24, wherein the structured parent RNA molecule is a transfer RNA (tRNA) and the target RNA fragment is a tRNA-derived fragment (tDR).

26. The method of claim 24 or 25, wherein the high-affinity nucleic acid analog is selected from the group consisting of a Locked Nucleic Acid (LNA) oligonucleotide, a Peptide Nucleic Acid (PNA), a Phosphorodiamidate Morpholino Oligomer (PMO), and a 2'-O-methyl modified oligonucleotide.

27. The method of any one of claims 24-26, wherein the step of calculating the numerical specificity score utilizes a nucleic acid thermodynamics software package to calculate partition functions for all molecular species in the simulated competitive binding reaction.

28. The method of any one of claims 24-27, wherein the library generation step comprises applying a sliding window algorithm to a sequence of the target RNA fragment with a one-nucleotide stride.

29. The method of any one of claims 24-28, further comprising generating and displaying a two-dimensional plot, wherein each candidate sequence is represented as a point plotted according to its calculated efficiency score on a first axis and its calculated specificity score on a second axis, and identifying a subset of candidates located within a target performance quadrant.

30. The method of claim 24, further comprising optimizing the candidate sequence for a specific target environment by: (a) obtaining a biological sample from the target environment; (b) experimentally measuring a concentration ratio between the target RNA fragment and the structured parent RNA molecule within said sample; and (c) inputting said concentration ratio into a competitive binding simulation as a boundary condition to calculate a specificity score tuned to the physiological conditions of the target environment.

31. A method for the functional characterization and validation of a specific RNA fragment, the method comprising:(a) identifying a specific target fragment sequence derived from a full-length structured parent" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT nucleic acid;(b) operating an integrated computational platform in a first gain-of-function mode to design a guide RNA (gRNA) that targets the parent nucleic acid, comprising executing the method of any one of claims 12-23 to select a gRNA sequence based on structural accessibility and thermodynamic binding energy;(c) operating the platform in a second loss-of-function mode to design a high-affinity inhibitor specific to the same target fragment, comprising executing the method of any one of claims 24-30 to select an inhibitor sequence based on a calculated efficiency score and a calculated specificity score; and(d) coordinating the first and second modes to establish a functional feedback loop for verifying the causal role of the fragment in a biological pathway.

32. A method for treating a condition associated with the aberrant accumulation of a target RNA fragment, the method comprising:(a) identifying a target fragment associated with a condition in an individual, wherein the target fragment is comprised within a full-length structured parent nucleic acid;(b) generating a library of candidate high-affinity nucleic acid analog sequences complementary to the identified target fragment;(c) calculating an efficiency score for each candidate in the library, representing the binding affinity of the candidate to the target fragment relative to potential off-target sequences located within the full-length structured parent nucleic acid;(d) simulating a competitive binding environment to derive a specificity score, wherein a simulation models the equilibrium concentrations of the candidate, the target fragment, and the full-length structured parent nucleic acid to predict the probability of exclusive binding to the fragment;(e) selecting a specific candidate sequence from the library that exhibits high efficiency and high specificity for the fragment while minimizing binding to the parent nucleic acid; and (f) administering a therapeutically effective amount of the selected specific candidate sequence to the individual having the condition to reduce the accumulation of the target fragment without" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT disrupting the function of the parent nucleic acid.

33. The method of claim 32, wherein the condition is selected from the group consisting of cancer, a neurological disorder, a metabolic disorder, and a viral infection.

34. The method of claim 32 or 33. wherein the high-affinity nucleic acid analog is a Locked Nucleic Acid (LNA) oligonucleotide or an Antisense Oligonucleotide (ASO).

35. The method of any one of claims 32-34, wherein the administering is performed via parenteral injection, intrathecal injection, intracerebroventricular infusion, or direct tissue injection.

36. A method for modulating cellular function by inducing the formation of a bioactive RNA fragment, the method comprising:(a) identifying a structured parent nucleic acid capable of being processed into a bioactive fragment;(b) selecting a spacer sequence for a guide RNA (gRNA) to target a specific accessible region of the parent nucleic acid using the method of any one of claims 12-23;(c) delivering to a target cell population a CRIS PR- associated (Cas) RNA-guided nuclease and the selected gRNA; and(d) inducing the controlled generation of the bioactive fragment within the target cell population to modulate a cellular function selected from gene expression regulation, translation repression, or stress response.

37. The method of claim 36, wherein the Cas nuclease is selected from the group consisting of Casl3a, Casl3b, Casl3c, Casl3d, Casl3x, and Casl3y.

38. A method for detecting aberrant accumulation of a target RNA fragment biomarker in an individual, the method comprising:(a) performing the method of any one of claims 24-30 to design a complementary oligonucleotide probe specific for the target fragment biomarker;(b) obtaining a biological sample from the individual;(c) contacting the sample with a detectable probe comprising the designed complementary" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT oligonucleotide probe under conditions that allow formation of a probe-fragment complex while minimizing cross -hybridization with a full-length structured parent nucleic acid; and(d) detecting the presence or quantifying the level of the probe-fragment complex in the sample to diagnose a condition.

39. A method of molecular engineering for regulating protein translation efficiency or stress response pathways in a eukaryotic cell, the method comprising:(a) introducing into the cell a programmable Cas RNA-guided nuclease system and a specifically designed guide RNA (gRNA);(b) wherein the gRNA is selected using the method of any one of claims 1-10 to target a precise region of a specific tRNA transcript; and(c) inducing the controlled formation of a tRNA-derived fragment (tDR) capable of inhibiting translation initiation or promoting stress granule formation.

40. The method of claim 39, further comprising verifying functional decoupling in the RNA regulatory network by contacting the cell with a high-affinity complementary oligonucleotide reagent selected using the method of any one of claims 24-30 to bind the tDR while discriminating against the specific tRNA transcript.

41. A system for designing nucleic acids targeting structured RNAs, comprising:(a) a processor; and(b) a memory communicatively coupled to the processor, the memory storing computer-executable instructions that, when executed by the processor, cause the system to perform the method of any one of claims 12-30.

42. A kit for modulating RNA regulatory networks, comprising:(a) a guide RNA (gRNA) comprising a spacer sequence selected according to the method of any one of claims 12-23;(b) a high-affinity nucleic acid analog inhibitor comprising a sequence selected according to the method of any one of claims 24-30; and" A Physical Model-Based Computational Framework..."Inventors: Shu wen Lei et al. Docket No.: P3304-PCT (c) instructions for using the gRNA and the inhibitor to perform the method of claim 31.