Method for predicting immunogenic double-stranded RNA loading

By determining immunogenic dsRNA levels through genotyping and calculating an IDS, the method addresses the unclear contribution of RNA editing to autoimmune diseases, facilitating precise disease stratification and targeted therapeutic interventions.

JP2025522326APending Publication Date: 2025-07-15THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024570663
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-07
Filing Date
2023-06-07
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The extent to which common genetic differences in RNA editing contribute to autoimmune and inflammatory diseases remains unclear, and existing methods do not effectively quantify immunogenic double-stranded RNA (dsRNA) levels to inform disease stratification and therapeutic interventions.

Method used

A method is developed to determine immunogenic dsRNA levels in individuals based on their genotype, using tissue-specific subsets and genotyping for risk alleles associated with dsRNA editing, allowing for the calculation of an immunogenic dsRNA score (IDS) to predict disease risk and therapeutic responses.

Benefits of technology

The method provides a precise assessment of dsRNA levels, enabling effective disease stratification and personalized therapeutic interventions by predicting the risk of autoimmune and inflammatory diseases and guiding therapies targeting the dsRNA-sensing pathway.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522326000029
    Figure 2025522326000029
  • Figure 2025522326000030
    Figure 2025522326000030
  • Figure 2025522326000031
    Figure 2025522326000031
Patent Text Reader

Abstract

Compositions and methods are provided for determining immunogenic dsRNA levels in an individual based on the genotype of the individual. This information is useful for disease stratification and for selecting appropriate therapeutic interventions, including, for example, inhibition of MDA5 and inhibition of type I interferon signaling. In some embodiments, a tissue-specific subset of immunogenic dsRNA is analyzed, and the tissue is associated with an immune-related disease of interest.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Background Genome-wide association studies (GWAS) have discovered hundreds of thousands of risk variants involved in traits and etiologies, but understanding of these molecular functions remains an ongoing challenge. Quantitative trait locus (QTL) studies, most notably exemplified by gene expression QTL (eQTL), have succeeded in connecting GWAS variants to their molecular mechanisms. Alternative splicing QTL (sQTL) have further broadened the discovery of these mechanisms. However, other post-transcriptional processes, such as RNA editing, remain largely unexplored despite the growing recognition of their important functions in health and disease.

[0002] One of the most abundant RNA modifications is adenosine-to-inosine (A-to-I) RNA editing, which is catalyzed by RNA-activated adenosine deaminases (ADAR) that bind to double-stranded RNA (dsRNA) substrates and convert adenosine to inosine. Since inosine is recognized as guanosine, unlike most other RNA modifications, RNA editing events can be accurately identified and quantified by standard RNA sequencing. Previous studies have identified millions of RNA editing sites in humans, more than 99% of which are located in inverted repeat Alu (IRAlu) that form dsRNA substrates.

[0003] What is important for editing in mammals are the two enzymatically active ADAR proteins, ADAR1 and ADAR2, which have distinct physiological functions in vivo. ADAR1 is widely expressed among human tissues and plays a crucial role in suppressing dsRNA sensing mediated by MDA5, a cytosolic sensor of "non-self" dsRNA (Figure 1a). Adar1 editing-deficient mice are embryonically lethal due to the elevated innate immune response indicated by the induction of interferon-induced genes (ISGs), but can be rescued to full lifespan when MDA5 is knocked out.

[0004] Immunogenic double-stranded RNA (dsRNA) structures are endogenous dsRNA structures similar to viral RNA and can trigger inappropriate activation of the innate immune response, which can cause severe damage to host cells. Adenosine-to-inosine (A-to-I) RNA editing is a common post-transcriptional modification abundant within repetitive elements of all metazoans. An important function of A-to-I RNA editing by ADAR1 is to suppress the immunogenic response by endogenous dsRNA.

[0005] In humans, loss-of-function mutations in ADAR1 and gain-of-function mutations in MDA5 have been identified in rare autoimmune diseases such as Aicardi-Goutières syndrome (AGS), further demonstrating the ADAR1-dsRNA-MDA5 axis as an underlying mechanism in immune diseases (Figure 1a). Protective loss-of-function alleles in MDA5 have also been found in genome-wide association studies (GWAS) of common inflammatory diseases such as type 1 diabetes (T1D), psoriasis, inflammatory bowel disease, vitiligo, vitamin B12 deficiency anemia, hypothyroidism, and coronary artery disease. Additionally, abnormal editing has been reported in several common autoimmune diseases including psoriasis, rheumatoid arthritis (RA), systemic lupus erythematosus (SLE), and multiple sclerosis (MS). However, the extent to which common genetic differences in RNA editing contribute to immune and inflammatory diseases has not yet been clarified. Summary of the Invention

[0006] Summary Compositions and methods are provided for determining immunogenic dsRNA levels in an individual based on the genotype of the individual. This information is useful for disease stratification and for selecting appropriate therapeutic interventions, including, for example, inhibition of MDA5 and inhibition of type I interferon signaling. In some embodiments, a tissue-specific subset of immunogenic dsRNA is analyzed, and the tissue is associated with an immune-related disease of interest.

[0007] An increase in immunogenic dsRNA levels is strongly associated with a high risk of autoimmune and inflammatory diseases. It is shown herein that ADAR-mediated adenosine-to-inosine (A-to-I) RNA editing is a major mechanism underlying the association of certain genetic variants with common autoimmune and immune-related diseases. cis RNA editing QTLs are referred to herein as "edQTLs" and are identified herein across multiple human tissues. An exemplary list of such edQTLs is shown in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577. These edQTLs are significantly enriched in GWAS signals for autoimmune and immune-mediated diseases.

[0008] DsRNA loci associated with disease are also identified, which can be defined as dsRNAs for which sufficient editing is important for suppressing autoimmunity. By determining the co-localization of signals between edQTLs found in 49 human tissues and reported genetic variants obtained from GWAS studies, co-localization events were identified that link immune-related diseases to genes that express putative immunogenic dsRNAs. The majority of these immunogenic dsRNAs are located within exons, particularly within UTRs where long dsRNA structures are often formed.

[0009] Immunogenic dsRNA acts in an aggregated state to induce cellular immunogenicity. These risk variants for immune-related diseases exhibit a directional effect, where, across multiple autoimmune and immune-related diseases, there is an overwhelmingly negative effect, i.e., the risk GWAS variants decrease overall dsRNA editing. The directional effect of RNA editing is even more important when tested in tissue types relevant to the disease. Associations are shown, for example, for autoimmune diseases including autoimmune thyroid disease (ATD), celiac disease, inflammatory bowel disease (IBD), primary biliary cirrhosis (PBC), systemic lupus erythematosus (SLE), multiple sclerosis, vitiligo, psoriasis, type 1 diabetes (T1D), rheumatoid arthritis, atopic conditions such as asthma and atopic dermatitis, and other conditions with a significant inflammatory component, such as amyotrophic lateral sclerosis (ALS), coronary artery disease, Parkinson's disease, systemic sclerosis, schizophrenia, Alzheimer's disease, and levels of high-density lipoprotein and low-density lipoprotein and triglycerides.

[0010] In some aspects, methods are provided for determining an individual's immunogenic dsRNA score (IDS), which is correlated with immune-related diseases and may be capable of predicting response to therapies, such as therapies directed at reducing unwanted activity of clinical sequelae in the dsRNA-sensing pathway. The IDS can also be used in stratifying individuals for clinical trials and for identifying individuals who will respond to therapies that reduce unwanted activity of clinical sequelae in the dsRNA-sensing pathway. Such unwanted activity can include, but is not limited to, a decrease in ADAR-mediated RNA editing in the whole body or in the target tissue; an increase in MDA5 activity in the whole body or in the target tissue; and an increase in type I interferon expression in the whole body or in the target tissue. These activities indicate points of therapeutic intervention for individuals determined to be at risk. In some aspects, the individual has previously been diagnosed as having an immune-related disease of interest.

[0011] In some embodiments, a method of determining an individual's immunogenic dsRNA score (IDS) includes genotyping the subject at a plurality of risk alleles associated with unwanted activity in the dsRNA sensing pathway, and creating a double-stranded RNA (dsRNA) burden score by weighting the number of risk alleles that decrease the RNA editing level by (a) the effect size of the association with the disease and (b) the effect size of its association with the RNA editing level. The risk alleles may be selected from the SNPs listed in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577. This information is also shown in Table 2 of priority provisional application US63 / 473,678, which is specifically incorporated herein by reference. A subset of the SNP list is shown in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577. This information is also shown in Table 3 of priority provisional application US63 / 473,678, which is specifically incorporated herein by reference.

[0012] In some embodiments, determining the immunogenic dsRNA score further includes genotyping the individual at one or more alleles associated with MDA5 or ADAR1. In some embodiments, the immunogenic dsRNA score further includes determining the expression level of dsRNA in a tissue associated with the disease or cells derived therefrom.

[0013] In some embodiments, the immunogenic dsRNA score further includes inputting additional clinical features including, but not limited to, one or more of age, gender, ethnic group, genetic background, family history, age of onset of the disease, duration of the disease, etc.

[0014] For example, genotyping of single nucleotide polymorphisms (SNPs) can be performed on DNA samples from individuals for single nucleotide polymorphisms by sequencing, hybridization, etc. known in the art. The methods disclosed herein identify a subset of SNP loci that predict immunogenic dsRNA levels by quantitative trait locus (QTL) mapping in multiple individuals, as shown in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577, which is specifically incorporated herein by reference. Using DNA sequencing, SNP loci with p-value scores less than a given threshold are classified by the directional effect of increasing immunogenic dsRNA levels. Using training on genome-wide association study (GWAS) data, the distribution of IDS between cases and controls for a given autoimmune / inflammatory disease is evaluated, and an IDS threshold for making predictions regarding high risk for the disease is determined. The method may include genotyping at least about 100, at least 200, at least about 500, at least about 1000, at least about 5000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000 SNPs, and may genotype substantially all of the edQTL SNPs listed in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577.

[0015] The immunogenic dsRNA score (IDS) can be calculated based on factors as shown above. For a given individual, a predicted IDS for a tissue / cell type associated with a disease that is greater than the IDS threshold for the disease is considered to indicate high risk. In some embodiments, the IDS (Ψ) is given by the formula: TIFF2025522326000001.tif9128, where TIFF2025522326000002.tif5128 is the effect size of the j-th SNP for the disease, TIFF2025522326000003.tif5128 is the effect size of the i-th dsRNA having the j-th SNP for the edQTL in the k-th tissue / cell type, TIFF2025522326000004.tif5128 is the expression level of the i-th dsRNA in the k-th tissue / cell type, which can be measured by, for example, qPCR, RNA-seq, scRNA-seq, etc., and optionally further includes clinical characteristics as described above.

[0016] Many diseases have underlying inflammatory components that contribute to the onset and / or progression of the disease. In some embodiments, the method includes treating an individual according to an IDS determination. Diseases of interest include, for example, autoimmune diseases such as rheumatoid arthritis (RA), T1D, systemic lupus erythematosus (SLE), multiple sclerosis (MS), autoimmune hepatitis; degenerative diseases such as osteoarthritis (OA), Alzheimer's disease (AD), and macular degeneration; metabolic diseases including type 2 diabetes, metabolic syndrome, non-alcoholic steatohepatitis (NASH), and alcoholic steatohepatitis; cardiovascular diseases such as atherosclerosis; cancers that may arise from and may induce inflammation; and other diseases with inflammatory components such as Parkinson's disease, Alzheimer's disease, etc.

[0017] The methods disclosed herein include a data analysis step, which may be provided as a program of computer-executable instructions and may be executed by a software component loaded on a computer. Such methods include one or more of the steps of inputting genotyping data, e.g., sequence information about SNP loci of interest; inputting the effect size of SNPs for disease association; inputting the effect size of SNPs for RNA editing level association; inputting the expression level of dsRNA in a tissue of interest. The method may also include determining a threshold level for the disease by training on GWAS data used to evaluate the distribution of IDS between cases and controls for a given disease and determining an IDS threshold to create a prediction of high risk for the disease. The method may include, for these inputs, the step of solving TIFF2025522326000005.tif9128. Other bioinformatics methods are provided for determining and quantifying when IDS is physiologically relevant. The method may further include the step of providing a computer-generated report including the analysis and determination of IDS.

Brief Description of the Drawings

[0018] The present invention is best understood from the following detailed description when read in conjunction with the accompanying drawings. It is emphasized that, in accordance with the practice, the various features of the drawings are not to scale. Instead, the dimensions of the various features are enlarged or reduced as appropriate for clarity of illustration. The drawings include the following figures.

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9-1

Figure 9-2

Figure 10

Figure 11

Figure 12

Figure 13-1

Figure 13-2

Figure 14

BEST MODE FOR CARRYING OUT THE INVENTION

[0020] Detailed Description Before describing the methods and compositions of the present invention, it must be understood that the present invention is not limited to the specific methods or compositions described, and thus, of course, may vary. It must also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting, as the scope of the present invention is limited only by the appended claims.

[0021] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range is specifically disclosed, unless the context clearly dictates otherwise. Each narrower range between any stated value or intervening value within the stated range and any other stated value or intervening value within the stated range is encompassed by the present invention. The upper and lower limits of these narrower ranges may independently be included in or excluded from the range, and each range that includes one or neither or both of the limits of the narrower range is also encompassed by the present invention, subject to any specifically excluded limit clearly defined in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of the included limits are also included in the present invention.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, some potential and preferred methods and materials are now described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials set forth in the cited publications. It is understood that the present disclosure supersedes any disclosure of the incorporated publications to the extent of any conflict.

[0023] It should be noted that, as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, references to "a cell" include a plurality of such cells, and references to "the peptide" include references to one or more peptides and their equivalents known to those of ordinary skill in the art, such as polypeptides.

[0024] The publications discussed herein are provided solely for the disclosure of publications prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publications by virtue of prior invention. Further, the dates of the provided publications may be different from the actual publication dates and it may be necessary to independently verify the actual publication dates.

[0025] A SNP (single nucleotide polymorphism) is a germline substitution of a single nucleotide at a specific position in the genome. SNPs within a population can be assigned a minor allele frequency - the lowest allele frequency observed at a locus in a particular population. Association studies can be used to determine whether a genetic variant is associated with a disease or trait. One of the major contributions of SNPs in clinical research is genome-wide association studies (GWAS). Genome-wide genetic data can be generated by multiple techniques, including SNP arrays and whole-genome sequencing. GWAS have been commonly used in the identification of SNPs associated with diseases or clinical phenotypes or traits. Since GWAS are genome-wide assessments, large sample sites are required to obtain sufficient statistical power to detect all possible associations. Some SNPs have relatively small effects on diseases or clinical phenotypes or traits. It is necessary to consider genetic models of diseases, such as dominant, recessive, or additive effects, to estimate the power of a study. SNP bioinformatics databases include, for example, dbSNP from the National Center for Biotechnology Information (NCBI); the Human Gene Mutation Database; the International HapMap Project, GWAS Central, and others.

[0026] GWAS In genome-wide association studies (GWAS), allele frequency differences in genetic variants between individuals with similar ancestry but different phenotypes are widely tested. GWAS can take into account copy number variants or sequence variations in the human genome, but the most commonly studied genetic variants in GWAS are single nucleotide polymorphisms (SNPs). GWAS typically report blocks of correlated SNPs that all show a statistically significant association with a trait of interest, known as a genomic risk locus. It is known from GWAS that most traits are affected by thousands of causal variants that each confer a very small risk, are often associated with many other traits, and are correlated with physically close causal and non-causal variants as a result of linkage disequilibrium.

[0027] The GWAS workflow involves collection of DNA and phenotype information from groups of individuals; genotyping of each individual using available GWAS arrays or sequencing strategies; quality control; imputation of unclassified variants using haplotype phasing and reference populations; and performance of statistical tests for association, with multiple post-GWAS analyses to interpret the results.

[0028] QTL Quantitative trait locus analysis is a statistical method for associating phenotype data with genotype data. Several types of markers can be used, including single nucleotide polymorphisms (SNPs), simple sequence repeats (SSRs, or microsatellites), restriction fragment length polymorphisms (RFLPs), and transposable element positions. Markers that are genetically associated with QTLs that affect a trait of interest segregate frequently with the value of the trait, whereas unassociated markers do not show a significant association with the phenotype.

[0029] In QTL analysis, public databases can be used. For example, the Genotype-Tissue Expression (GTEx) project includes data from samples collected from 54 non-diseased tissue sites across approximately 1000 individuals for molecular assays mainly including WGS, WES, and RNA-Seq. The remaining samples are available from the GTEx Biobank. The GTEx Portal provides open access to data including gene expression, QTL, and histological images. GTEx data can be mapped in conjunction with genotype data, such as SNP data.

[0030] An eQTL is a specific genomic region associated with variation in gene expression levels. eQTL analysis aims to identify and map the loci that affect the expression levels of genes. These loci can be found within or near genes and may be gene variants, such as single nucleotide polymorphisms (SNPs), that affect gene regulation or activity. By studying eQTLs, researchers can gain insights into the genetic factors contributing to differences in gene expression between individuals and understand how these variations can affect complex traits and diseases.

[0031] In the present disclosure, QLT analysis is used to explain the RNA editing level, defined as the ratio of edited (''G'') transcripts to all (''A'' and ''G'') transcripts at a single nucleotide across a catalog of sites. Principal component analysis (PCA) may be performed to identify and account for potential confounding factors in the editing level measurements.

[0032] In identifying edQTLs in tissue types, it may be useful to utilize the QTL mapping pipeline adopted by the GTEx Consortium, as is known in the art. See, for example, Liang et al. (2021) Nature Communications volume 12, article 1424; Wang et al. (2021) BMC Bioinformatics. 22(Suppl 9): 403; THE GTEX CONSORTIUM (2020) Science 369 (6509):1318-1330, each of which is hereby specifically incorporated by reference.

[0033] Genotyping A number of methods are used to determine the presence of variants in an individual. Genomic DNA is isolated from the individual to be tested. DNA can be isolated from any nucleated cell source, such as blood, hair shafts, saliva, mucus, biopsy material, feces, etc. Methods using PCR amplification on DNA derived from single cells can be performed, but it is convenient to use at least about 10 5 individual cells. If a large amount of DNA is available, genomic DNA can be used directly.

[0034] The methods of the present disclosure may involve sequencing of target loci and analysis of sequence data. Various methods and protocols for DNA sequencing and analysis are well known in the art and are described herein. Sequencing can be accomplished using high-throughput systems. In some cases, using high-throughput sequencing, at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, or at least 500,000 sequence reads are generated per hour, and each read is at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, or at least 150 bases per read. Sequencing can be performed using nucleic acids described herein, such as genomic DNA, cDNA derived from RNA transcripts, or RNA, as templates. Sequencing may include ultra-parallel sequencing.

[0035] For example, DNA sequencing can be accomplished using high-throughput DNA sequencing methods. Examples of next-generation sequencing and high-throughput sequencing include, for example, massively parallel signature sequencing, polony sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing with HiSeq, MiSeq, and other platforms, SOLiD sequencing, ion semiconductor sequencing (Ion Torrent), DNA nanoball sequencing, heliscope single-molecule sequencing, single-molecule real-time (SMRT) sequencing, MassARRAY®, and Digital Analysis of Selected Regions (DANSR™). For example, Stein RA (1 September 2008). 「Next-Generation Sequencing Update」. Genetic Engineering & Biotechnology News 28 (15); Quail, Michael; Smith, Miriam E; Coupland, Paul; Otto, Thomas D; Harris, Simon R; Connor, Thomas R; Bertoni, Anna; Swerdlow, Harold P; Gu, Yong (1 January 2012). 「A tale of three next generation sequencing platforms: comparison of Ion torrent, pacific biosciences and illumina MiSeq sequencers」.BMC Genomics 13 (1): 341; Liu, Lin; Li, Yinhu; Li, Siliang; Hu, Ni; He, Yimin; Pong, Ray; Lin, Danni; Lu, Lihua; Law, Maggie (January 1, 2012). 「Comparison of Next-Generation Sequencing Systems」. Journal of Biomedicine and Biotechnology 2012: 1-11; Qualitative and quantitative genotyping using single base primer extension coupled with matrix-assisted laser desorption / ionization time-of -flight mass spectrometry (MassARRAY(trademark)). Methods Mal Biol.2009;578:307-43; Chu T, Bunce K, Hogge WA, Peters DG. A novel approach toward the challenge of accurately quantifying fetal DNA in maternal plasma. Prenat Diagn 2010;30: 1226-9; and Suzuki N, Kamataki A, Yamaki J, Homma Y. Characterization of circulating DNA in healthy human plasma. Clinica chimica acta; international journal of clinical chemistry 2008;387:55-8). See also. Similarly, software programs for primary and secondary analysis of sequence data are well known in the art.

[0036] In some embodiments, high-throughput sequencing involves the use of technologies available from Helicos BioSciences Corporation (Cambridge, Massachusetts), such as the Single Molecule Sequencing by Synthesis (SMSS) method. In some embodiments, high-throughput sequencing involves the use of technologies available from 454 Lifesciences, Inc. (Branford, Connecticut), such as the use of a Pico Titer Plate device that includes an optical fiber plate that conveys chemiluminescent signals generated by a sequencing reaction recorded by a CCD camera in the instrument. Using such optical fibers enables the detection of at least 20 million base pairs in 4.5 hours.

[0037] In some embodiments, high-throughput sequencing is performed using reversible terminator chemistry in a Clonal Single Molecule Array (Solexa, Inc.) or sequencing-by-synthesis (SBS). These technologies are described in part in U.S. Patent Nos. 6,969,488; 6,897,023; 6,833,246; 6,787,308; and U.S. Patent Application Publication Nos. 20040106130; 20030064398; 20030022207; and Constans, A, The Scientist 2003, 17(13):36.

[0038] In one aspect, high-throughput sequencing of RNA or DNA can be performed using AnyDot.chips (Genovoxx, Germany) that enable monitoring of biological processes (e.g., miRNA expression or allelic variability (SNP detection)). In particular, using AnyDot-chips can enhance nucleotide fluorescence signal detection by 10x to 50x. Other high-throughput sequencing systems include those disclosed in Venter, J., et al. Science 16 February 2001; Adams, M. et al, Science 24 March 2000; and M. J, Levene, et al. Science 299:682-686, January 2003; as well as U.S. Patent Application Publication Nos. 20030044781 and 2006 / 0078937. The growth of the nucleic acid strand and the identification of the added nucleotide analogs can be repeated so that the nucleic acid strand is further extended and the sequence of the target nucleic acid is determined.

[0039] The methods disclosed herein may include DNA amplification. The amplification may include PCR-based amplification. Or, the amplification may include non-PCR-based amplification. DNA amplification may include using fiber optic detection after bead amplification as described in Marguiles et al. "Genome sequencing in microfabricated high-density pricolitre reactors", Nature, doi: 10.1038 / nature03959; as well as U.S. Patent Application Publication Nos. 20020012930; 20030058629; 20030100102; 20030148344; 20040248161; 20050079510; 20050124022; and 20060078909.

[0040] Nucleic acid amplification may involve the use of one or more polymerases. The polymerase may be a DNA polymerase. The polymerase may be an RNA polymerase. The polymerase may be a high fidelity polymerase. The polymerase may be KAPA HiFi DNA polymerase. The polymerase may be Phusion DNA polymerase.

[0041] Of interest is the use of polymerase chain reaction (PCR) to amplify DNA between two specific primers. The use of polymerase chain reaction is described in Saiki et al. (1985) Science 239:487, and a review of current techniques can be found in McPherson et al. (2000) PCR (Basics: From Background to Bench) Springer Verlag; ISBN: 0387916008. A detectable label may be included in the amplification reaction. Suitable labels include fluorescent dyes such as fluorescein isothiocyanate (FITC), rhodamine, Texas red, phycoerythrin, allophycocyanin, 6-carboxyfluorescein (6-FAM), 2',7'-dimethoxy-4',5'-dichloro-6-carboxyfluorescein (JOE), 6-carboxy-X-rhodamine (ROX), 6-carboxy-2',4',7',4,7-hexachlorofluorescein (HEX), 5-carboxyfluorescein (5-FAM), or N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), radioactive labels such as 32 P, 35 S, 3 H, etc. The label may be in a two-step system where the amplified DNA is conjugated to a biotin, hapten, etc. having a high affinity binding partner such as avidin, specific antibody, etc., and the binding partner is conjugated to a detectable label. The label may be conjugated to one or both of the primers. Alternatively, the pool of nucleotides used in the amplification is labeled to incorporate the label into the amplification product.

[0042] The primer pair is selected from the genomic sequence using conventional criteria for selection. The pair of primers hybridizes to opposite strands and collectively flanks the region of interest. The primers hybridize to complementary regions under stringent conditions and generally are at least about 16 nt in length, and may be 20, 25, or 30 nucleotides in length. The primers are selected to amplify a specific region suspected of containing a predisposing mutation. Typically, the length of the amplified fragment is selected so as to distinguish 3 to 7 units of repeats. For simultaneous analysis of multiple exons, multiplex amplification may be performed in which several primer sets are combined within the same reaction tube. Each primer may be conjugated to a different label.

[0043] The exact composition of the primer sequence is not critical to the present invention, but it must hybridize to the adjacent sequence under stringent conditions. Criteria for selecting amplification primers have been discussed previously. To maximize the resolution of size differences at the locus, it is preferred to select a primer sequence close to the SNP sequence such that the entire amplification product ranges from at least about 30, more usually at least about 50, preferably at least about 100 or 200 nucleotides in length, which can vary with the number of repeats present, to less than about 500 nucleotides in length. The number of repeats has been found to be polymorphic as previously described, whereby an individual difference occurs in the length of DNA between the amplification primers. Conveniently, a detectable label is included in the amplification reaction. Multiplex amplification may be performed in which several primer sets are combined within the same reaction tube. This is particularly advantageous when only limited amounts of sample DNA are available for analysis. Conveniently, each of the primer sets is labeled with a different fluorescent dye.

[0044] Determination of the effect size In the case of GWAS data, typically, the effect size is reported as summary statistics and is represented by either the odds ratio or the percentage of the phenotypic variance attributable to the locus (in the case of quantitative traits such as weight and height). The effect size of an edQTL for RNA editing is defined as the slope of the linear regression between the genotype dosage and the RNA editing level and is calculated as the effect of the alternative allele compared to the reference allele (the allele reported in the human genome reference sequence). The determination of the effect size of an edQTL may be performed, for example, by a statistical package for QTL mapping using FastQTL to perform a permutation pass following a nominal pass to determine all associations between all genetic variants within a particular cis window (e.g., 100 kb) and the editing level of the RNA editing site at the center of the cis window.

[0045] Expression level analysis The methods of the present disclosure may include the analysis of the expression level of transcripts of interest, for example, the analysis of the expression of immunogenic dsRNA. Any convenient method can be used to determine the level of a specific RNA in a sample. Methods of interest include, but are not limited to, microarray analysis, RNAseq, qPCR, and the like.

[0046] A microarray hybridizes an RNA population of interest or cDNA derived therefrom to a probe set known as a probe to determine the relative abundance of mRNA at the target.

[0047] RNA-Seq, also known as whole transcriptome shotgun sequencing, attempts to perform the same function that DNA microarrays have been used to perform in the past, but with greater resolution. In particular, DNA microarrays utilize specific probes, and the creation of these probes is necessarily influenced by prior knowledge of the genome and the size of the array being generated. In RNA-seq, these constraints are removed by simply sequencing all of the cDNA generated in a microarray experiment. This is made possible by next-generation sequencing technology. This technique has been rapidly adopted in the study of diseases such as cancer [4]. Subsequently, data from RNA-seq are analyzed by clustering in the same way that data from microarrays are typically analyzed.

[0048] RNAseq is a sequencing technique that uses next-generation sequencing (NGS) to reveal the presence and amount of RNA in a biological sample. In a typical workflow, RNA is isolated from multiple samples, converted into a cDNA library, sequenced into a computer-readable format, aligned to a reference, and quantified for downstream analysis such as differential expression. Single-cell RNA sequencing (scRNA-Seq) provides the expression profiles of individual cells. Gene expression patterns can be identified by gene clustering analysis.

[0049] In real-time polymerase chain reaction (qPCR), sequence amplification is carried out in the presence of a fluorescent dye, the fluorescence is measured after each cycle, and the intensity of the fluorescence signal reflects the instantaneous amount of DNA amplicon present in the sample at a particular time. In the first cycle, the fluorescence is too weak to be distinguished from the background. However, the point at which the fluorescence intensity increases above a detectable level corresponds in a proportional relationship to the initial number of template DNA molecules present in the sample. This point is called the quantification cycle and enables the determination of the absolute amount of target DNA in the sample according to a calibration curve generated with serially diluted standard samples of known concentration or copy number (usually decimal dilutions). Strategies for the real-time visualization of amplified DNA fragments include non-specific fluorescent DNA dyes and fluorescently labeled oligonucleotide probes.

[0050] The terms "subject", "individual", and "patient" are used interchangeably herein to refer to a mammal being evaluated for and / or being treated for a condition. In some embodiments, the mammal is a human. The terms "subject", "individual", and "patient" include, but are not limited to, an individual having a disease. The subject may be human, but may also include other mammals, particularly mammals useful as experimental models of human disease, such as mice, rats, etc.

[0051] For a patient, the term "sample" includes blood and other liquid samples of biological origin, solid tissue samples such as biopsy specimens or tissue cultures or cells derived therefrom, and their progeny. The term also includes samples that have been manipulated by any means after their acquisition, such as by treatment with reagents; washing; or concentration of a predetermined cell population such as diseased cells. This definition also includes samples that have been enriched for a particular type of molecule, such as nucleic acids, polypeptides, etc. The term "biological sample" includes clinical samples and also includes tissues obtained by surgical resection, tissues obtained by biopsy, cells in culture, cell supernatants, cell lysates, tissue samples, organs, bone marrow, blood, plasma, serum, etc. A "biological sample" includes a sample obtained from a patient's diseased cells, such as a sample containing polynucleotides and / or polypeptides obtained from a patient's diseased cells (e.g., a cell lysate or other cell extract containing polynucleotides and / or polypeptides); and a sample containing diseased cells from a patient. A biological sample containing diseased cells from a patient may include non-diseased cells.

[0052] The term "diagnosis" is used herein to refer to the identification of a molecular or pathological state, disease or condition in a subject, individual, or patient.

[0053] The term "prognosis" is used herein to refer to the prediction of the likelihood of death, or of disease progression including recurrence, spread, and drug resistance, in a subject, individual, or patient. The term "prediction" is used herein to refer to the act of foretelling or estimating, based on observation, experience, or scientific inference, the likelihood that a subject, individual, or patient will experience a particular event or clinical outcome. In one example, a physician may attempt to predict the likelihood that a patient will survive.

[0054] As used herein, terms such as "treatment", "the step of treating", etc. refer to the step of administering a drug or performing a procedure on or in a subject, individual, or patient for the purpose of obtaining an effect. The effect may be preventive in that it completely or partially prevents a disease or its symptoms, and / or may be therapeutic in that it brings about partial or complete cure of the disease and / or the symptoms of the disease. "Treatment" as used herein may include the treatment of cancer in mammals, particularly humans, and includes (a) the step of inhibiting the disease, i.e., the step of stopping the onset of the disease, and (b) the step of alleviating the disease or its symptoms, i.e., the step of causing regression of the disease or its symptoms.

[0055] The step of treating may refer to any sign of success in the treatment, remission, or prevention of a disease, including any objective or subjective parameter, such as relief; well-being; reducing symptoms or making the disease state tolerable to the patient; slowing the rate of degeneration or decline; or reducing the debilitating nature of the end point of degeneration. Treatment or remission of symptoms may be based on objective or subjective parameters, including the results of an examination by a physician.

[0056] As used herein, "therapeutically effective amount" refers to an amount of a therapeutic agent sufficient to treat or manage a disease or disorder. A therapeutically effective amount may refer to an amount of a therapeutic agent sufficient to delay or minimize the onset of a disease, e.g., sufficient to delay or minimize the growth and spread of cancer. A therapeutically effective amount may also refer to an amount of a therapeutic agent that provides a therapeutic benefit in the treatment or management of a disease. Further, the therapeutically effective amount of a therapeutic agent according to the present invention means an amount of the therapeutic agent alone, or in combination with other therapies, that provides a therapeutic benefit in the treatment or management of a disease.

[0057] As used herein, the term "administration schedule" typically refers to a series (typically multiple) of unit doses that are individually administered to a subject over a plurality of periods. In some embodiments, a given therapeutic agent has a recommended administration schedule, which may involve one or more doses. In some embodiments, an administration schedule includes multiple doses, each dose optionally separated from the others by a period of the same length. In some embodiments, an administration schedule includes multiple doses and at least two different periods separating the individual doses. In some embodiments, all doses within an administration schedule are of the same unit dose amount. In some embodiments, separate doses within an administration schedule are of different amounts. In some embodiments, an administration schedule includes a first dose of a first dose amount and one or more additional doses of a second dose amount different from the first dose amount that follow it. In some embodiments, an administration schedule includes a first dose of a first dose amount and one or more additional doses of a second dose amount the same as the first dose amount that follow it. In some embodiments, an administration schedule is correlated with a desirable or beneficial outcome (i.e., is a therapeutic administration schedule) when administered across a relevant population.

[0058] "In combination with," "combination therapy," and "combination product" refer, in certain embodiments, to co-administering to a patient the engineered proteins and cells described herein in combination with an additional therapy, such as surgery, radiation, chemotherapy, etc. When administered in combination, each component may be administered at the same time or may be administered sequentially at different times in any order. Thus, each component can be administered separately, but can be administered at times close enough together to produce a desired therapeutic effect.

[0059] "Combined administration" means administering one or more components, such as engineered proteins and cells, known therapeutic agents, etc., at a time such that the combination will have a therapeutic effect. Such combined administration may involve simultaneous (i.e., at the same time), prior, or subsequent administration of the components. One of ordinary skill in the art will not have difficulty determining the appropriate timing, sequence, and dosage of administration.

[0060] The use of the term "in combination" does not limit the order in which a prophylactic and / or therapeutic agent is administered to a subject having a disorder. The first prophylactic or therapeutic agent can be administered before (e.g., 5 minutes, 15 minutes, 30 minutes, 45 minutes, 1 hour, 2 hours, 4 hours, 6 hours, 12 hours, 24 hours, 48 hours, 72 hours, 96 hours, 1 week, 2 weeks, 3 weeks, 4 weeks, 5 weeks, 6 weeks, 8 weeks, or 12 weeks before), simultaneously with, or after (e.g., 5 minutes, 15 minutes, 30 minutes, 45 minutes, 1 hour, 2 hours, 4 hours, 6 hours, 12 hours, 24 hours, 48 hours, 72 hours, 96 hours, 1 week, 2 weeks, 3 weeks, 4 weeks, 5 weeks, 6 weeks, 8 weeks, or 12 weeks after) administration of the second prophylactic or therapeutic agent to a subject having a disorder.

[0061] As used herein with respect to DNA sequences, the term "sequence identity" refers to sequence identity between two molecules. If a position in both molecules is occupied by the same monomeric nucleotide, the molecules are identical at that position. Similarity between two nucleotide sequences is a linear function of the number of identical positions. Generally, the sequences are aligned so as to obtain the highest level of matches. If necessary, publicly available techniques and widely available computer programs, such as the GCS program package (Devereux et al., Nucleic Acids Res. 12:387, 1984), BLASTP, BLASTN, FASTA (Atschul et al., J. Molecular Biol. 215:403, 1990) can be used to calculate identity.

[0062] The term "isolated" refers to a molecule that is substantially free of its natural environment. For example, an isolated protein is substantially free of cellular material and other proteins from the cell source and tissue source from which it is derived. This term refers to a preparation that is pure enough for administration as a therapeutic composition, or at least 70% - 80% (w / w) pure, more preferably at least 80% - 90% (w / w) pure, even more preferably 90 - 95% pure, and most preferably at least 95%, 96%, 97%, 98%, 99%, or 100% (w / w) pure. A "separated" compound refers to a compound that has been removed from at least 90% of at least one component of the sample from which the compound was obtained. Any compound described herein can be provided as an isolated compound or a separated compound.

[0063] A number of inflammatory conditions can be associated with IDS for prognosis, stratification, and treatment. Many diseases have an underlying inflammatory component that contributes to the initiation and / or progression of the disease. Thus, the range of inflammatory diseases and disorders associated with inflammation is broad and includes autoimmune diseases such as rheumatoid arthritis (RA), systemic lupus erythematosus (SLE), multiple sclerosis (MS), and autoimmune hepatitis; degenerative diseases such as osteoarthritis (OA), Alzheimer's disease (AD), and macular degeneration; metabolic diseases including type II diabetes, metabolic syndrome, non-alcoholic steatohepatitis (NASH), and alcoholic steatohepatitis; cardiovascular diseases such as atherosclerosis; neurological conditions such as amyotrophic lateral sclerosis (ALS), Parkinson's disease, Alzheimer's disease, etc.; and other diseases with an inflammatory component.

[0064] Rheumatoid arthritis (RA) is a chronic syndrome characterized by symmetrical peripheral joint inflammation that can cause progressive destruction of joint and periarticular structures, usually with or without systemic symptoms (Firestein (2003) Nature 423(6937):356-61; Mcinnes and Schett. (2011) N Engl J Med. 365(23):2205-19). The cause is unknown. Genetic predisposition has been identified and in some populations, it has been mapped to the pentapeptide at the HLA-DRβ1 locus of the class II major histocompatibility complex gene. Environmental factors may also play a role. For example, smoking increases the risk of developing RA by about 10- to 20-fold in individuals with HLA-DR4, which includes the "shared epitope" polymorphism. Smoking is thought to induce an anti-citrullinated protein antibody (ACPA) response measured using commercially available cyclic citrullinated peptide (CCP) assays (Klareskog et al. (2006) Arthritis Rheum. 54(1):38-46). Additionally, periodontal inflammation and P. gingivalis infection may also play a role in initiating the autoimmune response that causes RA (Rutger and Persson. 2012, J Oral Microbiol. 4). The immunological changes may be initiated by multiple factors. Approximately 0.6% of the general population is affected, and the frequency is 2- to 3-fold higher in women than in men. Onset can occur at any age, but most often it is between 25 and 50 years old.

[0065] Systemic lupus erythematosus (SLE) is a systemic autoimmune disease characterized by malar rash, oral ulcers, photosensitivity, serositis, seizures, low white blood cell count, low platelet count, seizures, positive antinuclear antibody (ANA) test, and other positive autoantibodies. SLE is an autoimmune disease characterized by polyclonal B cell activation that results in immune complexes that contribute to tissue damage and various anti-protein and non-protein autoantibodies that cause inflammation (see, e.g., Kotzin et al., 1996, Cell 85:303-06 for a review of the disease). SLE has a variable course characterized by exacerbations and remissions and is difficult to study. For example, some patients may predominantly exhibit skin rashes and joint pain, experience spontaneous remission, and require little drug therapy. At the other end of the spectrum are patients who present with severe and progressive kidney complications (nephritis and encephalitis) that require therapy with high-dose steroids and cytotoxic drugs such as cyclophosphamide. Hydroxychloroquine slows the progression of SLE and is central to the treatment regimen for SLE management.

[0066] Multiple sclerosis (MS) is a debilitating inflammatory neurological disease characterized by demyelination of the central nervous system. This disease primarily affects young adults and has a higher incidence in women. Symptoms of the disease include fatigue, numbness, tremors, tingling, paresthesia, visual disturbances, dizziness, cognitive impairment, urinary dysfunction, reduced mobility, and depression. Four types: relapsing-remitting, secondary progressive, primary progressive, and progressive relapsing classify the clinical patterns of the disease (S. L. Hauser and D. E. Goodkin, Multiple Sclerosis and Other Demyelinating Diseases in Harrison's Principles of Internal Medicine 14th Edition, vol. 2, Mc Graw-Hill, 1998, pp. 2409-19).

[0067] Inflammatory bowel diseases include Crohn's disease and ulcerative colitis and are accompanied by autoimmune attacks on the intestine. These diseases cause chronic diarrhea often mixed with blood as well as symptoms of colonic dysfunction.

[0068] Systemic sclerosis (SSc, or scleroderma) is an autoimmune disease characterized by fibrosis of the skin and internal organs and extensive vasculopathy. Patients with SSc are classified according to the degree of skin sclerosis. That is, patients with limited SSc have skin thickening on the face, neck, and distal extremities, whereas patients with diffuse SSc also have complications in the trunk, abdomen, and proximal extremities. Visceral complications tend to occur earlier in the course of the disease in patients with diffuse disease compared to those with limited disease (Laing et al. (1997) Arthritis. Rheum. 40:734 - 42). The majority of patients with diffuse SSc who develop severe visceral complications do so within the first 3 years after diagnosis, and at the same time, the skin progresses to fibrosis (Steen and Medsger (2000) Arthritis Rheum.43:2437 - 44.). Common manifestations of diffuse SSc, which are a significant cause of morbidity and mortality, include interstitial lung disease (ILD), Raynaud's phenomenon and digital ulceration, pulmonary arterial hypertension (PAH) (Trad et al. (2006) Arthritis. Rheum. 54:184 - 91.), musculoskeletal symptoms, and cardiac and renal complications (Ostojic and Damjanov (2006) Clin. Rheumatol. 25:453 - 7). Current therapies focus on treating specific symptoms, but there is a lack of disease-modifying drugs that target the underlying etiology.

[0069] Autoimmune hepatitis is a disease in which the body's immune system attacks liver cells. This immune response causes inflammation of the liver, which is also called hepatitis. Researchers believe that genetic factors may make some people more susceptible to autoimmune diseases. Approximately 70 percent of people with autoimmune hepatitis are women. This disease is usually quite severe and, if untreated, will worsen over time. Autoimmune hepatitis is typically chronic, meaning that it can last for years and may lead to liver cirrhosis - scarring and hardening - which can eventually result in liver failure.

[0070] Coronary artery disease (CAD) is the narrowing or blockage of the arteries and blood vessels that supply oxygen and nutrients to the heart. Coronary artery disease (CAD) is caused by atherosclerosis, in which fatty materials build up on the inner lining of the arteries. The resulting blockage restricts blood flow to the heart. The result when blood flow is completely blocked is a heart attack. CAD is the leading cause of death in the United States for both men and women. As used herein, atherosclerosis (also called arteriosclerosis, atherosclerotic vascular disease, and arterial occlusive disease) refers to a cardiovascular disease characterized by plaque buildup on the blood vessel walls and vascular inflammation. Plaques consist of accumulations of intracellular and extracellular lipids, smooth muscle cells, connective tissue, inflammatory cells, and glycosaminoglycans. Inflammation occurs in combination with lipid accumulation within the blood vessel walls, and vascular inflammation is a prominent feature of the atherosclerotic disease process.

[0071] Myocardial infarction is usually ischemic myocardial necrosis caused by a sudden reduction in coronary blood flow to a part of the myocardium. In the majority of acute MI patients, an acute thrombus is often associated with plaque rupture, occluding the artery supplying the damaged area. Plaque rupture generally occurs in a blood vessel that has been previously partially occluded by an atherosclerotic lesion rich in inflammatory cells. Changes in platelet function induced by endothelial dysfunction and vascular inflammation in atherosclerotic lesions probably contribute to thrombosis. Myocardial infarction can be classified into ST-elevation MI and non-ST-elevation MI (also called unstable angina). Both types of myocardial infarction involve myocardial necrosis. ST-elevation myocardial infarction has transmural myocardial injury leading to ST elevation on the electrocardiogram. In non-ST-elevation myocardial infarction, the injury is subendocardial and not associated with ST-segment elevation on the electrocardiogram. Myocardial infarction (both ST-elevation and non-ST-elevation) represents an unstable atherosclerotic cardiovascular disease. Acute coronary syndrome encompasses all types of unstable coronary diseases. Heart failure may occur as a result of myocardial dysfunction caused by myocardial infarction.

[0072] The presence of inflammation in a condition can be detected by various approaches including medical history, physical examination, clinical tests, histological analysis of tissues, analysis of biomarkers, and imaging. Clinical features and physical examination markers of inflammation include or are pathologically related to swelling, exudation, edema, erythema, warmth, pain, or the influx of inflammatory cells or the production of inflammatory mediators. Clinical tests and / or histological markers when an increase in the number of inflammatory cells is demonstrated are abnormal. Inflammatory markers can include molecular markers. Examples of molecular markers include C-reactive protein, cytokines, antibodies, DNA sequences, RNA sequences, cartilage markers, metabolic markers, bone markers, or combinations thereof. Imaging can reveal findings including tissue enhancement, tissue edema and swelling, as well as other findings indicating inflammation. Examples of imaging markers of inflammation can include imaging markers measured using magnetic resonance imaging, ultrasound, computed tomography, angiography, and combinations thereof.

[0073] In one aspect, the method of the present invention includes a step of identifying an individual "at risk" of developing an inflammatory disease or an individual in the "early stage" of an inflammatory disease. "At risk" of developing an inflammatory disease includes (1) an individual with a high risk of developing an inflammatory disease and (2) an individual showing a "preclinical" disease state but not meeting the diagnostic criteria for an inflammatory disease (and thus not officially considered to have an inflammatory disease).

[0074] An individual with a "high risk" (also referred to as "at risk") of developing an inflammatory disease is an individual who has a higher likelihood of developing an inflammatory disease or a disease related to inflammation compared to the general population. Such an individual can be identified based on the IDS defined herein and may include, but is not limited to, the following characteristics: a family history of inflammatory disease; the presence of certain gene variants (genes) or combinations of gene variants that make the individual susceptible to such inflammatory diseases; the presence of physical findings, clinical test results, imaging findings, marker test results related to the development of inflammatory diseases (also referred to as "biomarker" test results), or marker test results related to the development of metabolic diseases; the presence of clinical signs related to inflammatory diseases; the presence of certain symptoms related to inflammatory diseases (although the individual is often asymptomatic); the presence of markers of inflammation (also referred to as "biomarkers"); and other findings indicating that the individual has a high likelihood of developing an inflammatory disease or a disease related to inflammation throughout their life. Most individuals at high risk of developing an inflammatory disease or a disease related to inflammation are asymptomatic and do not experience any symptoms related to the disease with a high risk of development.

[0075] Determination of immunogenic dsRNA score (IDS) Immunogenic dsRNA acts in the form of aggregates to induce cellular immunogenicity. These risk variants for immune-related diseases show a directional effect, where, across multiple autoimmune and immune-related diseases, there is an overwhelmingly negative effect, i.e., the risk GWAS variants reduce overall dsRNA editing. The directional effect of RNA editing is even more important when tested in tissue types relevant to the disease. Associations are shown, for example, for autoimmune diseases including autoimmune thyroid disease, celiac disease, inflammatory bowel disease, primary biliary cirrhosis, systemic lupus erythematosus, multiple sclerosis, psoriasis, type 1 diabetes (T1D), rheumatoid arthritis; atopic conditions such as asthma and atopic dermatitis; and other conditions with a significant inflammatory component such as amyotrophic lateral sclerosis (ALS), coronary artery disease, Parkinson's disease, systemic sclerosis, Alzheimer's disease, as well as levels of high-density lipoprotein and low-density lipoprotein and triglycerides.

[0076] IDS is correlated with immune-related diseases and can predict response to therapies, such as therapies directed at reducing unwanted activity of clinical sequelae in the dsRNA-sensing pathway. IDS can also be used for stratifying individuals for clinical trials and for identifying individuals who will respond to therapies that reduce unwanted activity of clinical sequelae in the dsRNA-sensing pathway. Such unwanted activity can include, but is not limited to, a decrease in ADAR-mediated RNA editing in the whole body, or in a target tissue or cell type; an increase in MDA5 activity in the whole body, or in a target tissue or cell type; and an increase in type I interferon expression in the whole body, or in a target tissue or cell type. These activities indicate a point of therapeutic intervention for individuals determined to be at risk. In some aspects, the individual has previously been diagnosed as having an immune-related disease of interest.

[0077] In some embodiments, a method of determining an individual's immunogenic dsRNA score (IDS) comprises genotyping the subject at a plurality of risk alleles associated with unwanted activity in the dsRNA sensing pathway, and creating a double-stranded RNA (dsRNA) burden score by weighting the number of risk alleles that decrease the RNA editing level by (a) the effect size of the association with the disease and (b) the effect size of the association of the RNA editing level with the risk alleles. The risk alleles can be selected from the SNPs listed in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577. In some embodiments, the determination of the immunogenic dsRNA score further comprises genotyping the individual at one or more alleles associated with a change in the RNA editing level.

[0078] In some embodiments, the immunogenic dsRNA score further comprises determining a specific dsRNA (referred to as immunogenic dsRNA) that is required to have an RNA editing level sufficient to maintain self-tolerance and avoid MDA5 activation. The immunogenic dsRNA may be selected from the edited genomic regions listed in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577. For example, as disclosed in Example 1 herein, SNPs computationally co-occur at GWAS loci of diseases of interest associated with the immune response and can provide a set of immunogenic dsRNA candidates.

[0079] Alternatively, immunogenic dsRNA can be experimentally obtained in a cell or tissue of interest by newly identifying an editing cluster defined as having (1) at least five editing sites and (2) any two adjacent sites spaced less than 120 nt apart. Identifying the editing cluster allows all sites in the dsRNA to be grouped to quantify the overall editing level defined as the cluster editing index. In a comparative analysis, search for editing clusters that have a significantly higher editing index in ADAR1 p150-complemented cells than in ADAR1 p110-complemented cells. Editing clusters similarly edited by p150 and p110 are likely to be less immunogenic. Clusters preferably edited by p150 are immunogenic dsRNA candidates. In some embodiments, the determination of immunogenic dsRNA further includes the dsRNAs listed in Tables S1 and S3 of Sun et al., bioRxiv 2022.08.29.505707, "A Small Subset of Cytosolic dsRNAs Must Be Edited by ADAR1 to Evade MDA5-Mediated Autoimmunity". See also Example 2.

[0080] In some embodiments, the immunogenic dsRNA score further comprises determining a disease-associated tissue or cell and the expression level of immunogenic dsRNA derived therefrom. The disease-associated tissue or cell can be selected by: (a) determining the largest tissue or cell by measuring the interferon (IFN) score by the aggregated expression level of interferon-stimulated genes (ISGs) (the ISGs can be selected from the list of 331 genes in Table 2), and (b) measuring the aggregated expression level of immunogenic dsRNA in the disease tissue scRNA-seq data using the single-cell disease relevance score (scDRS) method described in Zhang et al. (2022) Nature Genetics 54:1572-1580. The expression level of dsRNA can be further weighted to explain the immunogenicity of dsRNA. The weight is determined by the genetic effect of dsRNA in the increased disease risk measured in GWAS of inflammatory diseases. In the future, when it becomes possible to experimentally determine the efficacy of dsRNA in inducing MDA5 activation, this information can also be used as a weight.

[0081] In some embodiments, the immunogenic dsRNA score further comprises inputting one or more additional clinical features including, but not limited to, age, gender, family history, age of onset of the disease, duration of the disease, etc.

[0082] The immunogenic dsRNA score (IDS) can be calculated based on factors as shown above. For a given individual, a predicted IDS for a disease-related tissue / cell type that is greater than the IDS threshold for a given disease is considered to indicate a high risk. The IDS threshold can be experimentally determined by performing an MDA5 knockout assay in disease-related cell lines derived from human donors that include both patients and healthy controls. The cellular IFN response level is used as the readout for MDA5 knockout. For cells in which the IFN response is significantly reduced upon MDA5 knockout, the corresponding donor will likely be considered high risk. Conversely, for cells in which the IFN response does not change significantly upon MDA5 knockout, the corresponding donor will likely be considered low risk. It should be noted that high / low risk is a relative term with respect to the risk of MDA5-mediated disease. Once the IDS threshold has been determined and validated using a sufficient number of cell lines derived from human donors, no further experiments will be required for individual-level prediction in the future. The determination and validation of the IDS threshold need to be done for each different inflammatory disease. In some embodiments, the IDS (Ψ) is determined by the formula: TIFF2025522326000006.tif9128, where in the formula, TIFF2025522326000007.tif5128 is the effect size of the jth SNP for the disease in the kth tissue / cell type, TIFF2025522326000008.tif5128 is the effect size of the ith dsRNA having the jth SNP for the RNA editing level in the kth tissue / cell type, TIFF2025522326000009.tif5128 is the expression level of the ith dsRNA in the kth tissue / cell type, and optionally further includes clinical features as described above.

[0083] For example, genotyping of single nucleotide polymorphisms (SNPs) can be performed on DNA samples from individuals for a single nucleotide polymorphism by sequencing, hybridization, etc. known in the art. The methods disclosed herein identify a subset of SNP loci that predict immunogenic dsRNA levels by quantitative trait locus (QTL) mapping in multiple individuals, as shown in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577. Using DNA sequencing, SNP loci with p-value scores below a given threshold are classified by the directional effect of increasing immunogenic dsRNA levels. Using training on genome-wide association study (GWAS) data, the distribution of IDS between cases and controls is evaluated for a given autoimmune / inflammatory disease, and an IDS threshold is determined to make predictions regarding high risk for the disease. The methods may include genotyping at least about 100, at least 200, at least about 500, at least about 1000, at least about 5000, at least about 10,000, at least about 15,000, at least about 20,000, at least about 25,000, at least about 30,000 SNPs, and may genotype substantially all of the SNPs listed in Supplementary Table 3 of Li et al. (2022) Nature 608:569-577.

[0084] Treatment In some embodiments, the IDS is used to predict response to a therapy, e.g., a therapy directed at reducing unwanted activity of clinical sequelae in the dsRNA sensing pathway. In this case, a high-risk IDS score indicates a predicted response to a therapy directed at reducing unwanted activity of clinical sequelae in the dsRNA sensing pathway. Such unwanted activity can include, but is not limited to, a decrease in ADAR-mediated RNA editing in the whole body or target tissue; an increase in MDA5 activity in the whole body or target tissue; and an increase in type I interferon expression in the whole body or target tissue. These activities indicate a point of therapeutic intervention for an individual determined to be at risk.

[0085] In some embodiments, the IDS score is used to stratify individuals for clinical trials aimed at reducing unwanted activity of clinical sequelae in the dsRNA sensing pathway and to validate clinical trial results. Patient stratification, also known as patient segmentation or patient profiling, is a process used in clinical trials to group patients into distinct subpopulations based on specific characteristics or criteria. The purpose of patient stratification is to identify homogeneous patient groups that are likely to respond similarly to a particular treatment or intervention being tested in a clinical trial. By stratifying patients, researchers can enhance the efficiency and effectiveness of clinical trials, which leads to more accurate results and potentially accelerates the development of new therapies. Patient stratification helps identify specific patient subgroups within a heterogeneous disease population, enabling researchers to better understand treatment efficacy and safety profiles for different patient profiles. Biomarkers such as the IDS score can be used to identify patient subgroups that are more likely to benefit from a particular intervention, allowing for more targeted enrollment in clinical trials. Patient stratification can be done prospectively or retrospectively. Prospective stratification involves predefining patient subgroups before the start of a clinical trial based on prior knowledge or hypotheses. On the other hand, retrospective stratification involves analyzing patient data after the trial has started to identify subgroups based on observed responses. Various statistical techniques are used to identify and validate patient subgroups in clinical trials. These include clustering algorithms, classification models, regression analysis, and other machine learning methods. The choice of appropriate statistical techniques depends on the type and amount of data available for analysis.

[0086] To reduce the undesirable activity of clinical sequelae in the dsRNA sensing pathway, drugs that inhibit MDA5 activity; drugs that inhibit type I interferon activity; drugs that increase ADAR1 activity, or drugs that reduce the formation or expression of immunogenic dsRNA can be used as therapeutic agents. In some embodiments, the drug is an MDA5 inhibitor, such as a small molecule inhibitor, antisense RNA, or RNAi agent. Drugs that increase ADAR1 activity may include small molecules and coding constructs that increase expression in the target tissue.

[0087] The dosage and frequency may vary depending on the half-life of the drug in the patient. Such guidelines are understood by those skilled in the art to be adjusted according to the molecular weight of the active drug, clearance from the blood, route of administration, and other pharmacokinetic parameters. The dosage may also be varied for local administration, such as intranasal inhalation, or for systemic administration, such as i.m., i.p., i.v., oral, etc.

[0088] The active drug can be administered by any suitable means, including locally, orally, parenterally, intralung, and intranasally. Parenteral injection includes intramuscular, intravenous (bolus or slow drip), intraarterial, intraperitoneal, intrathecal, or subcutaneous administration. The drug can be administered in any medically acceptable manner. This may include injection by parenteral routes, such as intravenous, intravascular, intraarterial, subcutaneous, intramuscular, intratumoral, intraperitoneal, intracerebroventricular, intradermal, etc., as well as oral, nasal, ocular, rectal, or topical. Sustained release administration is also particularly included in the present disclosure by means such as depot injection or erodible implants.

[0089] As described above, the agent can be formulated together with a pharmaceutically acceptable carrier (one or more natural or synthetic organic or inorganic components combined with the agent of interest to facilitate its application). Suitable carriers include sterile saline, but other aqueous and non-aqueous isotonic sterile solutions and sterile suspensions known to be pharmaceutically acceptable are known to those skilled in the art. "Effective amount" refers to an amount that can alleviate or delay the progression of a diseased, altered, or damaged state. The effective amount can be determined on an individual basis and is based in part on the considerations of the symptoms to be treated and the desired outcome. The effective amount can be determined by those skilled in the art using such factors and no more than routine experimentation.

[0090] The agent can be administered as a pharmaceutical composition comprising a pharmaceutically acceptable excipient. The preferred form depends on the intended mode of administration and therapeutic use. The composition may also include a pharmaceutically acceptable non-toxic carrier or diluent, defined as a vehicle commonly used to formulate pharmaceutical compositions for animal or human administration, depending on the desired formulation. The diluent is selected so as not to affect the biological activity of the combination. Examples of such diluents are distilled water, physiological phosphate buffered saline, Ringer's solution, dextrose solution, and Hank's solution. Furthermore, the pharmaceutical composition or formulation may also include other carriers, adjuvants, or non-toxic, non-therapeutic, non-immunogenic stabilizers, etc.

[0091] As used herein, "commercially available" compounds are from Acros Organics (Pittsburgh PA), Aldrich Chemical (Milwaukee Wl.may be obtained from commercial suppliers including, but not limited to, Sigma Chemical and Fluka, Apin Chemicals Ltd. (Milton Park UK), Avocado Research (Lancashire U.K.), BDH Inc. (Toronto, Canada), Bionet (Cornwall, U.K.), Chemservice Inc. (West Chester PA), Crescent Chemical Co. (Hauppauge NY), Eastman Organic Chemicals, Eastman Kodak Company (Rochester NY), Fisher Scientific Co. (Pittsburgh PA), Fisons Chemicals (Leicestershire UK), Frontier Scientific (Logan UT), ICN Biomedicals, Inc. (Costa Mesa CA), Key Organics (Cornwall U.K.), Lancaster Synthesis (Windham NH), Maybridge Chemical Co. Ltd. (Cornwall U.K.), Parish Chemical Co. (Orem UT), Pfaltz & Bauer, Inc. (Waterbury CN), Polyorganix (Houston TX), Pierce Chemical Co. (Rockford IL), Riedel de Haen AG (Hannover, Germany), Spectrum Quality Product, Inc. (New Brunswick, NJ), TCI America (Portland OR), Trans World Chemicals, Inc. (Rockville MD), Wako Chemicals USA, Inc. (Richmond VA), Novabiochem and Argonaut Technology.

[0092] The compounds can also be prepared by methods known to those skilled in the art. The "methods known to those skilled in the art" used herein can be identified by various reference books and databases. Appropriate reference books and papers that detail the synthesis of reactants useful in the preparation of the compounds of the present invention or mention the preparation include, for example, "Synthetic Organic Chemistry", John Wiley & Sons, Inc., New York; S. R. Sandler et al., "Organic Functional Group Preparations", 2nd Ed., Academic Press, New York, 1983; H. O. House, "Modern Synthetic Reactions", 2nd Ed., W. A. Benjamin, Inc. Menlo Park, Calif. 1972; T. L. Gilchrist, "Heterocyclic Chemistry", 2nd Ed., John Wiley & Sons, New York, 1992; J. March, "Advanced Organic Chemistry: Reactions, Mechanisms and Structure", 4th Ed., Wiley-Interscience, New York, 1992. Specific and similar reactants can also be identified by indexes of known chemical substances prepared by the Chemical Abstract Service of the American Chemical Society, which are available in most public and university libraries and through online databases (for details, you may be able to inquire at American Chemical Society, Washington, D.C., ww.acs.org). Chemical substances that are known but not commercially available in catalogs may be prepared by custom chemical synthesis companies, and many of the standard chemical supply companies (e.g., those listed above) offer custom synthesis services.

[0093] The active agent of the present invention and / or the compound administered together with the active agent of the present invention are incorporated into various formulations for therapeutic administration. In one aspect, the agent is formulated into a pharmaceutical composition by combining it with a suitable, pharmaceutically acceptable carrier or diluent, and formulated into a preparation in solid, semi-solid, liquid, or gaseous form such as tablets, capsules, powders, granules, ointments, solutions, suppositories, injections, inhalants, gels, microspheres, and aerosols. Thus, administration of the active agent and / or other compounds can be accomplished in various ways, usually by oral administration. The active agent and / or other compounds may be systemic after administration, or may be local, thanks to the formulation or by using a graft that serves to retain an active dose at the transplant site.

[0094] Depending on the patient and condition being treated, as well as the route of administration, the active agent may be administered at a dose of 0.01 mg to 500 mg / kg body weight / day, for example, in the case of an average person, at a dose of about 20 mg / day. The dosage is appropriately adjusted for pediatric formulations.

[0095] The toxicity of the active agent can be determined by standard pharmaceutical procedures in cell culture or experimental animals, for example, by determining the LD50 (the dose that kills 50% of the population) or LD100 (the dose that kills 100% of the population). The dose ratio of the toxic effect to the therapeutic effect is the therapeutic index. The data obtained from these cell culture assays and animal tests can be used in the further optimization and / or definition of the therapeutic dosage range and / or dosage range below the therapeutic dose (for example, for human use). The exact formulation, route of administration, and dosage can be selected by the individual physician in view of the patient's condition.

[0096] Computer method The methods disclosed herein include a data analysis step, which is provided as a program of computer-executable instructions and can be executed by a software component loaded into a computer. Such methods include the step of inputting genotyping data, for example, sequence information about SNP loci of interest; the step of inputting the effect size of SNPs for disease relevance, the step of inputting the effect size of SNPs for RNA editing level relevance, and one or more of the step of inputting the expression level of dsRNA in the tissue of interest. The method may also include determining a threshold level for the disease by training on GWAS data used to evaluate the distribution of IDS between cases and controls for a given disease and determining an IDS threshold to create a prediction of high risk for the disease. The method may also include, for these inputs, the step of solving TIFF2025522326000010.tif9128. Other bioinformatics methods are provided for determining and quantifying when an IDS is physiologically relevant. The method may further include the step of providing a computer-generated report including the analysis and determination of the IDS.

[0097] The computer system may include a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") that may be a single-core or multi-core processor, which may be a plurality of processors for parallel processing. The system may also include a memory (e.g., random access memory, read-only memory, flash memory), an electronic storage device (e.g., hard disk), a communication interface (e.g., network adapter) for communicating with one or more other systems, and peripheral devices such as a cache, other memory, data storage, and / or an electronic display adapter. The memory, storage device, interface, and peripheral devices communicate with the CPU via a communication bus, such as a motherboard. The storage device may be a data storage device (or data repository) for storing data. The system is operatively connected to a computer network with the aid of the communication interface. The network may be the Internet, the Internet and / or an extranet, an intranet and / or an extranet that communicates with the Internet. The network may in some cases be a telecommunications and / or data network. The network may include one or more computer servers, thereby enabling distributed computing, such as cloud computing. The network may in some cases be able to execute a peer-to-peer network with the aid of the system, whereby the devices connected to the system may be able to act as clients or servers.

[0098] The system communicates with a processing system. The processing system may be configured to execute the methods disclosed herein. In some examples, the processing system is a nucleic acid sequencing system such as, for example, a next-generation sequencing system (e.g., an Illumina sequencer, an Ion Torrent sequencer, a Pacific Biosciences sequencer). The processing system may communicate with the system via a network or may communicate with the system by a direct (e.g., wired, wireless) connection. The processing system may be configured for analysis such as IDS analysis.

[0099] The methods described herein can be implemented via machine (or computer processor) executable code (or software) stored in an electronic storage location of the system, such as, for example, in a memory or an electronic storage device. During use, the processor can execute the code. In some examples, the code can be retrieved from the storage device and stored in a memory for rapid access by the processor. In certain situations, the electronic storage device can be excluded and the machine executable instructions are stored in a memory.

[0100] A computer-implemented system may comprise (a) a digital processing device including an operating system and a memory device configured to perform executable instructions, and (b) a computer program including executable instructions by the digital processing device, the computer program comprising: (i) a first software module configured to receive data related to DNA sequencing; (ii) a second software module configured to create an IDS by associating sequencing data; and (iii) a third software module configured to calculate a relative risk.

[0101] Computer-executable logic can operate on any of a variety of types of general-purpose computers, such as personal computers, network servers, workstations, or any computer of any other computer platform currently or later developed. In some aspects, a computer program product is described that includes a computer-usable medium having computer-executable logic (a computer software program including program code) stored thereon. The computer-executable logic can be executed by a processor, whereby the processor performs the functions described herein. In other aspects, some functions are implemented primarily in hardware, for example, using a hardware state machine. The implementation of a hardware state machine to perform the functions described herein will be apparent to those of ordinary skill in the relevant art.

[0102] Analysis and database storage may be implemented in hardware, in software, or in a combination of both. In one aspect of the invention, a machine-readable storage medium is provided, the medium comprising machine-readable data encoded data storage material that can display any data set and data comparison of the present invention when used with a machine programmed with instructions for using the data. Such data can be used for various purposes, such as patient monitoring, initial diagnosis, and the like. Preferably, the present invention is implemented in a computer program executed on a programmable computer comprising a processor, a data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The program code is applied to input data to perform the above-described functions, resulting in output information. The output information is applied to one or more output devices in a known manner. The computer can be, for example, a conventionally designed personal computer, microcomputer, or workstation.

[0103] Each program is preferably implemented in a high-level procedural programming language or an object-oriented programming language for communicating with a computer system. However, the program may be implemented in assembly language or machine language if desired. In any case, the language may be a compiler-type language or an interpreter-type language. Each such computer program, when read by a computer from a storage medium or device for the purpose of performing the procedures described herein, configures and drives a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., ROM or magnetic disk) for configuring and driving a computer. The system may also be regarded as being implemented as a computer-readable storage medium configured using a computer program. In this case, when using such a configured storage medium, the computer operates in a specific and predefined manner for performing the functions described herein.

[0104] Information can be input and output in the computer-based system of the present invention using various structural formats for input means and output means. One format for the output means tests data sets having various degrees of similarity to a reliable profile. Such a presentation provides those skilled in the art with a ranking of similarities and identifies the degree of similarity included in the test pattern.

Example

[0105] Experiment The following examples are presented to fully disclose and describe to those skilled in the art how to make and use the present invention, and are not intended to limit the scope of what the inventors regard as their invention, nor are the following experiments intended to show that they are all or the only experiments to be performed. Efforts have been made to ensure accuracy with respect to the numerical values used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should be accounted for. Unless otherwise specified, parts are by weight, molecular weights are weight-average molecular weights, temperatures are in degrees Celsius, and pressures are at or near atmospheric pressure.

[0106] Example 1 RNA editing plays a major role in common autoimmune and immune-related diseases. A major challenge in human genetics is to identify the molecular mechanisms of variants associated with traits and diseases. To achieve this, quantitative trait locus (QTL) mapping of gene variants with intermediate molecular phenotypes such as gene expression and splicing has been widely adopted. However, despite some success, the molecular basis of a substantial proportion of variants associated with traits and diseases remains unknown. Here, the inventors show that ADAR-mediated adenosine-to-inosine (A-to-I) RNA editing, a post-transcriptional event crucial for suppressing the cellular double-stranded RNA (dsRNA)-mediated innate immune interferon response, is a major mechanism underlying gene variants associated with common autoimmune and immune-related diseases. The inventors identified and characterized 30,319 cisRNA-editing QTLs (edQTLs) across 49 human tissues. These edQTLs were significantly enriched in GWAS signals for autoimmune and immune-mediated diseases, which exceeded the effects of expression and splicing QTLs. Co-localization analysis of edQTLs and disease risk loci further identified the precise locations of important, presumably immunogenic dsRNAs. Furthermore, aggregated inflammatory disease risk variants were associated with reduced nearby dsRNA editing and induction of the interferon response in inflammatory diseases. This unique directional effect is fully consistent with the established mechanism that the lack of RNA editing by ADAR1 leads to specific activation of MDA5, a dsRNA sensor, and subsequent interferon response and inflammation. The inventors' findings reveal dsRNA editing and sensing as previously unrecognized mechanisms underlying the genetic risk of common inflammatory diseases and present an immediately available therapeutic approach for treating patients with inflammatory diseases by MDA5 antagonism.

[0107] Identification of cis RNA-editing QTLs in human tissues In this study, the inventors aimed to obtain a systematic understanding of the role of A-to-I RNA editing in common and complex traits and diseases. The inventors mapped cis-editing QTLs (cis-edQTLs) using the substantially improved GTEx V8 RNA-seq and genotype data compared to previous efforts, such as larger sample sizes and more comprehensive RNA editing site annotations (Figure 1b; Methods). First, the inventors measured RNA editing levels defined as the ratio of edited (“G”) transcripts to the total of (“A” and “G”) transcripts at one nucleotide (“site”) across a catalog of >2.8 million sites and obtained reliable editing level quantification for 14,993–60,581 sites per tissue type (Figure 8a–b; Methods). To identify and account for potential confounding factors in the editing level measurements, the inventors performed principal component analysis (PCA) and found that ADAR1 expression levels were correlated with the top PCs explaining 8% of the total editing level variance (Figure 8c–d; Methods). Interestingly, moderately (40–60%) edited sites showed high levels of variation between individuals, whereas little variation was observed for sites edited less (<10%) or more (>90%) (Figure 8e). Similar observations were made for DNA methylation levels and RNA splicing ratios, for which competing models between splicing isoforms have been proposed to explain how local genetic perturbation affects splicing ratios.

[0108] Next, the inventors identified cis-edQTLs (hereinafter referred to as edQTLs) in each GTEx tissue type using the QTL mapping pipeline widely adopted by the GTEx Consortium. The inventors considered proximal common variants (MAF > 5%) within + / - 100 kb of the editing site and the editing site observed in at least 60 samples of each tissue type (method). Among all 287,965 editing sites tested, the inventors identified edQTLs for 30,319 sites (10.6%, FDR < 5%, permutation-based; method). Hereinafter, the inventors refer to editing sites with edQTLs as edSites. Furthermore, edQTLs were present in 32% (7,165) of all edited genes. Hereinafter, the inventors refer to edited genes as edGenes (Figure 9a). Within edGenes, the majority (about 60%) of edQTLs affected multiple editing sites simultaneously (Figure 9b). This was expected because editing sites tend to be located in close proximity. Furthermore, there were multiple independent edQTLs in nearly 30% of edGenes (Figure 9c). This indicates that editing sites can be co-controlled by multiple independent SNPs (Figure 9d). The number of edSites per tissue type was generally influenced by the sample size and the overall editing level (defined as the transcriptome-wide fraction of edited transcripts across all transcripts) (Figure 1c). Compared to recent studies using smaller datasets, the inventors identified approximately 9-fold more edSites (Figure 9e).

[0109] Characterization of cis RNA Editing QTLs in Human Tissues The inventors compared the effects of edQTLs across tissues by performing a meta-analysis (method). The inventors observed that the cis-genetic effects on RNA editing were highly consistent across tissues (Figure 1d; Figures 9f-g'). More specifically, 87.6% of cis-edQTLs had consistent effect directions across all tissues (Figure 1d; Figure 9f), whereas only 538 edQTLs showed tissue-specific effects (difference in effect size ≥2-fold; method) (Figure 9h). Furthermore, 31.1% of edQTLs were found in only one tissue due to tissue-specific gene expression. From the results of the inventors, it can be seen that the genetic landscape of RNA editing in humans is complex and requires multi-tissue data to be fully explained.

[0110] Next, the inventors compared edQTLs with expression QTLs (eQTLs) and splicing QTLs (sQTLs) identified in the same GTEx dataset. Unlike eQTLs that are enriched near the transcription start site (TSS), edQTLs were enriched in the 3'-UTR. This reflects that a large proportion (43%) of edSites are located in the 3'-UTR (Figure 1e, upper panel). Consistent with previous reports, sQTLs were enriched near splicing junctions, but edQTLs were not (Figure 1e, middle panel). As expected, edQTLs were strongly enriched near the editing sites (Figure 1e, lower panel). The slight enrichment observed for eQTLs and sQTLs near the editing sites suggests potential overlap between edQTLs and eQTLs and sQTLs. Indeed, 18.7% and 21.5% of edQTLs were also eQTLs or sQTLs, respectively (Figure 9i; Methods). This is consistent with smaller-scale recent studies. For genes with both edQTLs and eQTLs, most of the lead SNPs of edQTLs and eQTLs were >10 kb apart from each other and were enriched in genomic elements with distinct functional annotations (Figures 9j - k). The inventors' findings emphasize that most edQTLs are not detected as significant eQTLs or sQTLs in the GTEx dataset and potentially represent independent gene regulatory effects.

[0111] The inventors' edQTL maps also provide a unique opportunity to examine how RNA sequence and / or structural changes affect editing levels in human tissues. The inventors observed that edQTLs near editing sites generally show high statistical significance (Figure 10a). This highlights the effect of proximal regulatory elements, such as ADAR binding sites, RNA sequence motifs, and RNA secondary structures. Subsequently, the inventors performed a meta-analysis focused on Alu elements where >99% of human editing sites are found due to ADAR binding preference (Figure 1f; Methods). The inventors observed that the effect size of edQTLs is significantly correlated with the estimated ADAR1 binding strength (Figure 1f'). This suggests that genetic variants can alter editing levels by affecting ADAR1 binding on RNA. Next, the inventors evaluated how RNA sequence changes caused by genetic variants can affect editing levels at neighboring sites (Methods). The inventors found that an "AU A GG" sequence motif centered on edited "A" is preferred for high editing levels (Figures 10b - d). This is consistent with the previously reported "U A G" motif preferred by ADAR 41 . Furthermore, the inventors observed that edQTL SNPs are 13 - 54% more enriched in RNA secondary structures recognized by ADAR compared to neighboring non-edQTL SNPs (Figures 10e - f).

[0112] Genetic regulation of RNA editing contributes to disease heritability To evaluate the potential role of A-to-I dsRNA editing in common genetic diseases and traits, we assessed the enrichment of edQTLs in GWAS signals across multiple studies. A similar approach has been successfully applied in eQTL and sQTL studies to uncover the large contributions of gene expression and RNA splicing to complex diseases and traits. We found that edQTLs were highly enriched in GWAS signals, with more edQTLs than randomly sampled control SNPs matched in terms of linkage disequilibrium (LD), allele frequency, and gene density for autoimmune (exemplified by inflammatory bowel disease (IBD), lupus, multiple sclerosis (MS), and rheumatoid arthritis (RA)) as well as immune-related diseases (exemplified by coronary artery disease (CAD)) (Figure 2a; Methods). Consistent with previous studies, eQTLs and sQTLs were also enriched, but edQTLs had a larger effect size than eQTLs and sQTLs (Figure 2a). This observation held even when considering only eGenes and sGenes that were expressed to a similar extent as edGenes, which were generally more highly expressed (Figure 11a). Furthermore, we confirmed that these results, by quantitatively assessing QTL enrichment, explained the heritability of GWAS for the complex diseases and traits listed in Table 1. In 8 out of 9 autoimmune diseases tested, edQTLs were more enriched than eQTLs and sQTLs in terms of heritability (Figure 2b). Additionally, edQTLs were enriched in amyotrophic lateral sclerosis, CAD, triglycerides, and low-density lipoprotein (Figure 11b), all of which are associated with immune function.

[0113] (Table 1) Complex diseases and traits TIFF2025522326000011.tif19849

[0114] Using multi-tissue GTEx data, the inventors were also able to test the tissue-specific contribution of edQTLs in common diseases. In particular, using a recently published approach aimed at distinguishing directional, mediated effects from non-directional pleiotropic and linkage effects, the inventors estimated the proportion of heritability mediated by edQTLs in 24 diseases and traits.

[0115] This also showed that a higher overall proportion of heritability was mediated by edQTLs (0.18 ± 0.04) than by eQTLs (0.11 ± 0.02, estimated using the same method applied to GTEx V8) in autoimmune and immune-related diseases, but not in traits without an associated immune contribution (0.033 ± 0.02). Notably, edQTLs in immune system tissues (lymphocytes, spleen, and whole blood) collectively explained the largest proportion of heritability in most of the autoimmune and immune-related diseases tested (Figure 2c; Figure 11c). The heritability mediated by edQTLs was also enriched in known tissues associated with diseases such as digestive tissues (small intestine and colon) for IBD and celiac disease, brain tissues for ALS and Parkinson's disease, cardiovascular tissues for CAD, and pancreas for T1D (Figure 2c; Figure 11c').

[0116] In addition to GWAS of autoimmune and immune-related diseases, the inventors evaluated edQTL enrichment in GWAS signals of 33 highly heritable immune traits defined in recent studies. The inventors found that edQTLs were more enriched in GWAS of interferon response-related immune traits than in other immune traits tested (Figure 12). This is consistent with gene data from humans and mice showing that the lack of ADAR1 RNA editing induces MDA5-mediated innate immune interferon. Collectively, the inventors' data provide compelling evidence that edQTLs significantly contribute to the heritability of autoimmune and immune-related diseases, perhaps by inducing an interferon response mediated by dsRNA editing and sensing.

[0117] Identification and characterization of putatively immunogenic dsRNAs with disease relevance Genome-wide significant enrichment of edQTLs in the heritability of autoimmune and immune-related diseases has prompted us to precisely identify the positions of specific dsRNA loci associated with the diseases. We show these dsRNAs as putatively immunogenic dsRNAs where sufficient editing is important to suppress autoimmunity, perhaps by escaping MDA5 activation. We identified putatively immunogenic dsRNAs by systematically examining signal co-occurrence between edQTLs found in 49 human tissues and previously reported genetic variants obtained from 24 GWAS studies including 17 diseases and traits related to immune function (method). Overall, we identified 1,974 co-occurrence events (locus x edGene) linking 17 immune-related diseases to 194 genes that express putatively immunogenic dsRNAs (Figure 3a). Since dsRNA must be sensed by cytosolic MDA5 to induce immunity, we hypothesized that dsRNA must predominantly localize to exons / UTRs rather than introns to be present in the cytosol. Indeed, 178 out of 194 (92%) putatively immunogenic dsRNAs were located in exons, particularly in UTRs where long dsRNA structures are often formed (Figure 3b). In contrast, only 2,967 out of 15,620 (19%) total dsRNAs associated with edQTLs identified in this study were located in exons / UTRs, while 10,465 (67%) were located in introns where most IRAlu ddsRNAs are present (Figure 3b').

[0118] Next, the inventors characterized 194 putatively immunogenic dsRNAs. The majority of these (130, 67%) were located in IRAlu. This seemed considerably less than expected since in humans, almost all long dsRNAs are thought to be formed by IRAlu (see below). Forty-two (22%) of the 194 dsRNAs were common between at least two diseases (Figure 3c). This suggests that the immunogenicity of dsRNAs could serve as a general cause of susceptibility to multiple diseases. For example, the top candidate dsRNA found in TNFRSF14 was common to five diseases / traits.

[0119] Verification of putatively immunogenic dsRNAs To experimentally verify the immunogenicity of dsRNAs, first, the inventors evaluated whether the MDA5 protein could form filaments in vitro. The inventors tested three types of dsRNAs using negative stain electron microscopy (Methods). Each pair of dsRNAs formed MDA5 filaments of various lengths proportional to the dsRNA length, but these ssRNA controls did not (Figure 3h;). Furthermore, when dsRNAs were edited in vitro by ADAR1 before MDA5 incubation, the filament length was significantly shorter (Figure 3h–h’). This suggests that immunogenic dsRNAs need to be edited to escape MDA5 sensing.

[0120] Next, the inventors verified the immunogenicity of dsRNA in human cells using multiple lines of evidence. First, in addition to indicative hyperediting, the inventors searched for evidence of dsRNA formation. By using publicly available structure mapping and ADAR1 binding data in human cells, the inventors found that 29 out of 38 dsRNAs expressed in the corresponding cell lines showed high-confidence RNA structure signals for both strands in the overlapping region (normalized icSHAPE score ≥ 0.7). This indicates that a double-stranded structure was formed between the sense and antisense transcripts (Extended Data Figs. 6e–f). Second, the inventors evaluated the ability of dsRNA to induce MDA5-dependent immunogenicity by overexpressing candidates in human cells. To ensure that exogenous dsRNA was not edited to reduce its immunogenicity, the inventors expressed the CTSA:PLTP pair as a proof of concept in ADAR1-editing-deficient cells with inducible MDA5 (Methods). By transfection of a plasmid expressing sense and antisense transcripts in the overlapping region to mimic dsRNA formation, the inventors observed an increase in the immune response as indicated by the induction of three representative ISGs when MDA5 was expressed (Fig. 3i, unedited vs. ssRNA control). Third, the inventors examined the immunogenicity of this dsRNA in cells with WT ADAR1. As expected, the immune response was significantly reduced compared to ADAR1-editing-deficient cells (Fig. 3i, unedited vs. edited). This suggests that RNA editing of dsRNA by ADAR1 suppresses the immunogenicity of dsRNA. Fourth, the inventors hypothesized that the immunogenicity of dsRNA is determined by the structure of long dsRNA rather than the sequence. To test this, the inventors generated a scrambled version of dsRNA with randomized overlapping sequences while maintaining base-pair complementarity TIFF2025522326000012.tif4128. The inventors found indistinguishable immune induction between the scramble and the WT version (Figure 3i, unedited versus scramble). This confirms the inventors' hypothesis. In summary, from the inventors' analysis, it can be seen that long dsRNA becomes a very potent ligand for inducing the innate immune response by MDA5 only if it is not edited by ADAR1. This reflects the over-representation and importance of long dsRNA in inflammatory diseases.

[0121] Reduction of dsRNA editing underlies the increased genetic risk of inflammatory diseases From the above analysis of the inventors, the functional role of putatively immunogenic dsRNA in autoimmune and immune-related diseases is suggested. In the well-established ADAR1-dsRNA-MDA5 mechanism, the absence of ADAR1-mediated immunogenic dsRNA editing induces MDA5 activation and subsequent immune responses. First of all, these dsRNAs definitely act to aggregate and induce cellular immunogenicity and do not act alone. Therefore, the inventors judged that the risk variants of these diseases described above should show a directional effect to collectively reduce the neighboring dsRNA editing level. By reducing dsRNA editing, a better ligand for the host dsRNA sensor MDA5 will be generated, and thus the interferon response will increase (Figure 4a).

[0122] To test this model, we applied signed linkage disequilibrium profile (SLDP) regression to 24 complex traits and diseases (listed in Table 1) to assess the direction of the genome-wide collective effect of disease risk variants on editing levels, taking into account LD structure, allele frequencies, gene expression, and other potential systematic biases (Methods). Across multiple autoimmune and immune-related diseases, we detected an overwhelmingly negative direction of effect (i.e., risk GWAS variants decrease overall dsRNA editing). This supports our hypothesis that decreased dsRNA editing levels collectively increase disease risk (Figure 4b; Figure 13). The directional effect of RNA editing was even more pronounced when tested in tissue types relevant to the disease, as exemplified by the digestive tissue for IBD, cardiovascular tissue for CAD, pancreas for T1D, and brain tissue for Parkinson's disease. In contrast, such directional effects were not observed for eQTLs or control non-directional SNPs (Figure 4c; Figure 13c).

[0123] The inventors further tested the directional effect using RNA-seq data from patient samples of four immune-related diseases. The challenge with such analyses is that reduced immunogenic dsRNA editing leads to an interferon response, which in turn can induce ADAR1 expression and overall editing levels. Thus, the initial editing reduction driven by risk variants could potentially be masked by the eventual editing increase in the disease state. To overcome this, the inventors examined allele-specific editing (ASED) levels. This enabled measurement of the editing levels associated with risk alleles versus those associated with protective alleles. Overall, the inventors analyzed 152 synovial tissue samples from rheumatoid arthritis patients, 72 white matter samples from multiple sclerosis patients, 20 peripheral blood mononuclear cell (PBMC) samples from SLE patients, and 81 coronary artery samples from CAD patients. For each disease cohort, the inventors observed significantly reduced editing levels associated with risk alleles compared to protective alleles (Figure 4d, 33.7 ± 15.6% relative reduction, paired t-test, p < 2.7x10 -12 ; Figure 14). Furthermore, the degree of editing level reduction showed a positive correlation with high IFN response when measured by the IFN score (Methods), but not with other immune signatures (Figure 4e). This finding is in good agreement with the mechanism by which ADAR1-mediated RNA editing reduction leads to a dsRNA-mediated IFN response, providing strong evidence for the causal effect of dsRNA editing underlying common autoimmune and immune-related diseases.

[0124] The main functions of ADAR1-mediated dsRNA editing on cellular transcripts are to evade MDA5-mediated dsRNA sensing and autoimmunity. Mutations in ADAR1 and MDA5 can induce a strong immune response via dsRNA, leading to extremely rare autoimmune diseases. In this study, the inventors aimed to understand how the editing status of dsRNA can contribute to common human diseases. The inventors found that common genetic variants associated with RNA editing levels were significantly enriched in GWAS signals of common autoimmune and immune-related diseases, accounting for a significant proportion of disease heritability. Furthermore, the inventors showed a directional effect that GWAS-related genetic risk variants aggregate to generally reduce the neighboring dsRNA editing levels, especially in tissues related to the disease. Less-edited dsRNA serves as a good MDA5 substrate for inducing an interferon response. This is consistent with the well-established ADAR1-dsRNA-MDA5 axis (Figure 4f). From the inventors' findings, it is suggested that the enrichment of risk variants results in MDA5-dependent interferon responses and inflammation by collectively reducing the editing levels of associated dsRNAs (which are likely to be in the hundreds). This is further supported by previous findings of MDA5 loss-of-protection alleles in GWAS studies on several inflammatory diseases. In summary, the inventors' study, built on human genetics and well-established mechanisms, shows a feasible therapeutic approach to treat patients with inflammatory diseases by antagonizing MDA5.

[0125] In summary, the inventors' study presents RNA editing as an important mechanism underlying a large number of autoimmune and immune-related diseases, shedding light on the manipulation of dsRNA editing and sensing pathways for potential therapeutic agent development.

[0126] Method Quantification of editing levels in GTEx samples The GTEx gene expression data used in this study were obtained from the GTEx portal (GTEx Analysis V8 release) and measured in transcripts per million (TPM). Editing levels were quantified for GTEx Analysis Release V8 (dbGaP accession: phs000424.v8.p2), which consists of a total of 17,382 RNA-seq samples sequenced in 76bp paired-end reads. First, the inventors curated a list of reference editing sites for quantification. By incorporating known sites in the RADAR database, tissue-specific sites identified in GTEx V6p, and recently published high-editing sites, the inventors completed a list consisting of 2,802,572 human editing sites. To quantify the editing level, the inventors calculated the ratio of G reads to the sum of A and G reads at each site. The inventors included 15,201 RNA-seq samples from 838 donors with matching genotypes in 49 tissues with a sample size of ≧70 for editing level quantification and downstream analysis. For duplicate reads tagged during the RNA-seq mapping process, the inventors chose to retain the read with the highest base quality (random reads were selected if the base quality was the same). The inventors required that in each tissue, editing sites had to be covered by ≧20 non-duplicate reads in ≧60 samples considered testable in downstream analysis, and that the variation between samples was non-zero for cis-edQTL mapping. After applying the above filters, the inventors obtained 14,993 - 60,581 sites that differed between tissues for downstream analysis (Figure 8a). The code used for editing level quantification in the GTExV8 data and the reference editing site list is available at https: / / hub.docker.com / r / vanessa / mpileup / .

[0127] cis-edQTL mapping First, the inventors used principal component analysis (PCA) to identify potential confounding factors in the editing level measurements. Usually, editing level measurements are less confounded than gene expression. The inventors found that the top 10 PCs collectively contributed approximately 20% to the overall editing level variance (Figure 8c), and in agreement with previous observations, PC1 had a high correlation with ADAR1 expression levels. The number of PCs regressively estimated from the editing level measurements (see below) was selected to maximize the number of detected cis-edQTLs. ≤5 PCs were required in all tissues (the inventors tested 0 - 10 PCs).

[0128] To map edQTLs, the inventors considered all SNPs with MAF ≥ 0.05 that were within ±100 kb of the editing site. Variant call files (VCFs) of the genotype data of 838 GTEx donors with matching RNA-seq data were obtained from dbGAP (accession ID: phs000424.v8) based on the GRCh38 / hg38 reference.

[0129] The inventors used FastQTL for edQTL mapping. Before regressively estimating PCs, the raw editing level measurements were logit-transformed and then normalized to an N(0, 1) distribution among individuals within each tissue. The top 3 genotype PCs and gender and age were used as covariates for edQTL mapping. For each editing site, the adaptive permutation mode was set - used together with --permute 1000 10000. Trait-level q-values were calculated at a fixed p-value interval (λ = 0.85) for estimating π0 using the beta distribution-extrapolated empirical p-values from FastQTL. A false discovery rate (FDR) threshold of ≤0.05 was applied to identify editing sites with at least one significant edQTL.

[0130] To identify the list of all significant variant sites related to cis-edSites, the genome-wide empirical p-value threshold was defined as the empirical p-value of the site closest to the FDR threshold of 0.05. Subsequently, the nominal p-value threshold was calculated for each editing site based on the beta distribution model of the minimum p-value distribution obtained from gene permutations (from FastQTL). For each editing site, variants with nominal p-values smaller than the threshold were considered significant and included in the final list of variant-site pairs.

[0131] The inventors implemented a pipeline that strictly removed gene expression levels to account for potential confounding effects of gene expression in edQTL mapping. More specifically, for each editing site, the inventors included the expression level of its host gene (measured by TPM value) as an additional covariate, along with the genotype and phenotype covariates regression-estimated from the editing level measurements, and used the residuals for edQTL mapping.

[0132] edQTL tissue sharing The inventors applied multivariate adaptive shrinkage implemented in MashR to compare the edQTL effect sizes across tissues. To fit the MashR model, the inventors learned the MashR prior using a set of approximately 4,000 edQTLs common among 20 major tissue types, and then fit the MashR model using 40,000 randomly selected variant-trait pairs for the same set of edSites.

[0133] The inventors learned data-driven MashR priors by 1) PCA with a number of PCs = 3; 2) the empirical covariance of the observed Z scores. The data-driven covariance was further denoised by calling cov_ed in MashR. Furthermore, the inventors included a set of canonical covariances as additional MashR priors as described in "cis-edQTL mapping". The inventors fit the MashR model using a set of randomly selected variant-trait pairs and the error correlation estimated by applying the estimate_null_correlation function of MashR and the above priors. Using the resulting MashR model, the posterior mean, standard deviation, and local false sign rate (LFSR) of a given variant-trait pair were calculated. The effect size estimates and LFSR output by MashR were used as metrics for the scale and activity of edQTLs, respectively.

[0134] Comparative analysis among edQTL, eQTL, and sQTL The inventors obtained eQTLs and sQTLs mapped by the GTEx Analysis Working Group using GTEx V8 data (accession: phs000424.v8.p2) from dbGaP. The inventors evaluated the sharing of edQTLs with eQTLs and sQTLs by applying the stray π1. More specifically, the inventors identified significant SNP-editing pairs in a specific tissue and then used the distribution of the P values of these pairs, testing for expression levels or splicing ratios, to estimate π1, which is the ratio of non-null associations.

[0135] For meta-gene analysis, the inventors considered genes with all three types of QTLs mapped in GTEx tissues. For each gene, the corresponding cis signal was represented using the lead SNP of each QTL (in the case of edQTLs, the lead SNP of each site was used and compared with the lead SNPs of edQTLs and sQTLs of the same gene). The inventors used a gene-level model based on the GENCODE 68 V26 transcript annotation. In this case, to calculate the distribution of QTL SNPs, isoforms were collapsed into one "transcript" per gene with reference. To calculate the SNP density, in addition to all genes, the ±2kb sequences were collapsed into one meta-gene and further divided into 50 equal bins. For density calculation, in addition to splice junctions, the ±1kb sequences and for editing sites, the ±1kb sequences were treated similarly.

[0136] ADAR1 CLIP-seq analysis ADAR1 CLIP-seq was obtained from U87MG cells. Standard QC and filtering were performed using trim_galore (https: / / www.bioinformatics.babraham.ac.uk / projects / trim_galore / ) to filter high-quality sequencing reads excluding adapter sequences. Reads shorter than 15 nt were discarded after trimming. To characterize the binding profile of ADAR1 in Alu repeats, first, all reads were mapped to RefSeq genes using STAR with default settings. Then, we used BLASTN to align the mapped reads to the Alu consensus sequence and kept only the best hit for each read (parameters: -evalue 1e-10 -best_hit_score_edge 0.05 -best_hit_overhang 0.25 -perc_identity 50 -strand plus). RNA-seq data of U87MG cells were used as a control set. We processed the control RNA-seq reads in the same way as the CLIP-seq reads. Reads passing QC were mapped to the U87MG reference genome sequence using STAR with default settings. Then, the last mapped reads were aligned to the Alu consensus sequence as described above. To account for the non-uniform read coverage in Alu, we calculated the read density level per base within the Alu consensus sequence using the control RNA-seq, and then calculated the normalization factor per base by dividing the density level by the average density level of all Alu. Then, for the CLIP-seq data, read enrichment was calculated by multiplying the read count by the normalization factor.

[0137] RNA sequence motif Assuming that ADAR recognizes the triplet "UAG" motif in a position-dependent manner, the inventors attempted to test whether a specific type of nucleotide change at a particular position has a significantly strong effect on the alternating editing levels. For each edSite, the inventors considered all significantly associated SNPs located within a range of 50 nt upstream and 50 nt downstream of the editing site (a total of 100 positions). The inventors edited the ribonucleotide changes on RNA according to the gene's SNP and strand annotation (for example, if the SNP is a mutation from A to C on the reverse strand, it is interpreted as U to G on RNA). For each of the 100 positions, the inventors evaluated the average effect of all 12 types of ribonucleotide changes (Figure 10c). Notably, not all symmetric nucleotide changes show the same opposite effect. For example, the change from C to A at the -1 position has an average effect size of -0.8, while the change from A to C at the same position has an average effect size of +0.6. In some cases, the effects of changes such as A to U vs. U to A at the +1 position (-0.3 vs. -0.2), G to U vs. U to G at the +1 position (-0.4 vs. -0.1), etc. are in the same direction. To simplify the signal and reduce measurement noise, the inventors further grouped the data and showed the final effect of each of the four types of trinucleotides for the alternative alleles. The inventors plotted the sequence motif using ggseqlogo and the averaged effect sizes used to adjust the weights. Overall, a strong preference for A and U at the -2 and -1 positions, and in addition, a preference for G at the +1 and +2 positions was observed. This indicates the "AUAGG" motif.

[0138] RNA secondary structure To understand how mutations affect the editing level by changing the RNA secondary structure, the inventors predicted the edSites and the local RNA structures containing the associated SNPs. Since the computational prediction of long RNA molecules is technically difficult, the inventors restricted the prediction window to ±800 bp around each edSite (1601 bp in total) and considered only the SNPs that fell within that window. Furthermore, the inventors restricted their analysis to non-Alu editing sites. Overall, the inventors predicted the secondary structures of 8,043 editing sites and then annotated them with structural features using bpRNA (Figures 10e–f). For each of the five structural features (pseudoknots and dangling ends were not considered due to insufficient data), the inventors used the Mann–Whitney U test to compare the annotation of edVariants with SNPs that were found in the same local structure but not associated with the editing level.

[0139] Enrichment analysis of QTLs in GWAS signals To create a Q–Q plot of GWAS signals annotated with QTL information, the inventors used PLINK (--clump-r2 0.4 --clump-kb 250) to clump significant QTLs and obtain independent signals. To account for the gene expression levels of edGenes, which were generally higher than those of eGenes and sGenes, the inventors matched eGenes and sGenes to edGenes by the median expression level across tissues. For the purpose of creating a negative control set, the inventors considered four features to match a control SNP set to edVariants: 1) MAF distribution. All edVariants were divided into 50 equal bins by allele frequency, and control SNPs with matching allele frequencies were selected from the EUR set of 1000 Genomes using the median MAF of each bin; 2) LD (LD “buddies,” r 2The number of proxy SNPs at r < 0.7). Similar to MAF filtering, LD buddies were sampled by matching to the median in each bin of the edVariant distribution; 3) 3'-UTR density. edQTLs are strongly enriched around the 3'-UTR (average enrichment = 2.1, compared to genome-wide). Therefore, the inventors matched the number of 3'-UTRs in the locus around the control SNP using LD (r 2 > 0.7) and physical distance (250 kb) (enrichment = 2.0); 4) Distance to the nearest transcription termination site (TTS). The inventors sampled the distance from the SNP to the TTS (measured from the upstream site of the TTS so that the control SNP is located within the expression region) such that it was within the same range of deviation estimated from the distribution of distances from the edQTL to the TTS.

[0140] Heritability estimation and enrichment testing in GWAS The inventors used the Mediated Expression Score Regression (MESC) pipeline for heritability analysis. First, the inventors estimated the total editing score from the individual-level editing levels quantified in each GTEx V8 tissue with matching genotype information. Five editing level PCs (described above in the "cis-edQTL mapping" section) were used as covariates. For meta-analysis across tissues, an editing score was created using the edQTL effect sizes estimated using LASSO. Next, the inventors evaluated the heritability mediated by editing (h2med) using the editing scores from both individual tissues and from tissue groups. GWAS summary statistic data for 24 traits listed in Table 1 were obtained and converted to the.sumstats file format. LD scores calculated from 1000 Genomes Phase 3, stratified over a modified version of the baseline LD model v2.0, were downloaded from the Broad institute. The inventors took SNPs known in HapMap 3 as the SNP "universe".

[0141] To test for edQTL enrichment in immune-related traits, the inventors used GWAS from Sayaman et al. Overall, out of 139 detailed immune traits measured from approximately 9000 cancer patients registered in TCGA, 33 immune traits were highly heritable. Subsequently, GWAS was performed on these 33 immune traits. GWAS summary statistic data for these 33 immune traits, including six IFN response-related immune traits, can be publicly accessed via https: / / figshare.com / articles / dataset / Sayaman_et_al_TCGA_Germline-Immune_GWAS_Summary_Statistics / 13077920.

[0142] Estimation of the directional effect of RNA editing on complex traits and diseases The inventors estimated the directional effects by applying the signed LD profile (SLDP) regression method as described in the original paper. The inventors created a reference panel for SLDP regression using the 1000 Genomes Phase 3 European genotype. The LD scores of the same population were downloaded for the reference panel from https: / / data.broadinstitute.org / alkesgroup / SLDP / LDscore.tar.gz. The inventors transformed the GWAS summary statistic file as described above in the "Heritability estimation and enrichment tests in GWAS" section. The inventors processed the reference panel file by calculating the truncated singular value decomposition (SVD) for each LD block in the reference panel, which was later used to weight the regression performed by SLDP regression. The signed effect sizes of edQTLs were used as functional annotations for creating the signed LD profiles. For SNPs associated with multiple editing sites, the inventors combined the effect sizes across the associated sites within the closest editing cluster into one using the Stouffer's Z-score combination method. For the purpose of explicitly accounting for the potential signed effects of gene expression, the inventors also used the effect sizes of cis-eQTLs from the same gene set called with edQTLs to create another set of signed LD profiles used as a signed background model. To account for the systematic signed effects of minor alleles that can arise from population stratification or negative selection, the directional effects of minor alleles in five MAF bins of the same size were included in the signed background model as described in the original paper.

[0143] The inventors obtained publicly available patient-derived disease samples for allele-specific editing (ASED) analysis. Overall, the inventors tested 152 synovial tissue samples from rheumatoid arthritis patients (GSE89408), 72 white matter samples from multiple sclerosis patients (GSE138614), 20 peripheral blood mononuclear cell (PBMC) samples from SLE patients (GSE122459), and 81 coronary artery samples from CAD patients (GTEx). For each sample, the inventors performed stepwise mapping of the mapped RNA-seq reads near the corresponding disease GWAS risk variants and assigned each read to belong to either the risk haplotype or the protective haplotype. Then, the inventors quantified the editing level of the neighboring editing sites (found with the same paired-end reads as the variant or its LD buddies. r 2 >0.4) and calculated the total editing level of the sites near the risk allele and the protective allele in each sample.

[0144] The IFN score was calculated using mRNA expression data (measured in TPM) normalized by the median of the absolute deviation modified Z-score (ZMAD). This score was defined as the median of the ZMAD values of all signature genes in each sample. IFN signature genes were determined according to previous studies in inflammatory diseases.

[0145] GWAS Download and Preparation The inventors downloaded public GWAS summary statistics from various information sources and always reformatted them using tools freely available from https: / / github.com / mikegloudemans / gwas-download ("download" and "munge" modules).

[0146] SNP Selection for Coexistence Analysis The total number of GWAS trait x GWAS locus x QTL tissue x cis-QTL features is very large, so it is computationally difficult and unnecessary to run all possible combinations. The inventors ran the "overlap" module distributed at https: / / github.com / mikegloudemans / gwas-download to create a list of all trait / locus / tissue / feature combinations where the lead GWAS SNP for that trait at a given locus has a p-value < 5e-8 and overlaps with an eQTL having a p-value < 1e-5 for a given QTL feature in a given tissue. Each of these combinations represented one co-localization test to be performed. The inventors performed this process for editing, splicing, and expression QTLs to create an exhaustive list of 375,000 tests to run (26,000 edQTL; 183,000 sQTL 165,000 eQTL).

[0147] Co-localization analysis For the set of tests determined in the inventors' previous steps, the inventors performed co-localization analysis using the tool COLOC with default parameter settings and estimated allele frequencies from the complete set of 1000 Genomes individuals. For each test, the inventors obtained the H4 posterior probability (H4PP), an estimate of the probability that the GWAS and QTL studies share a common causal variant. For later analysis, the inventors considered a "co-localization" test if H4PP > 0.9 unless otherwise specified. This threshold indicates a high level of support for co-localization. The implementation of the wrapper pipeline the inventors used to perform the COLOC analysis is available at https: / / github.com / mikegloudemans / ensemble-colocalization-pipeline.

[0148] For the locuszoom plots shown in FIGS. 3 and 4, the inventors created plots comparing signal overlap between GWAS and edQTL, sQTL, and eQTL at the inventors' loci of interest using the publicly available R package LocusCompareR.

[0149] Protein Expression and Purification of MDA5 and ADAR1 Human MDA5 protein (residues 298 - 1025) was expressed from pET-50b(+) in Escherichia coli (E. coli) C41 cells and induced by addition of 0.2 mM IPTG. After incubation at 18 °C for 20 h with shaking, the cells were collected by centrifugation and resuspended in buffer containing 20 mM Tris-HCl, 500 mM NaCl, 5% glycerol, 20 mM imidazole, 0.5 mM PMSF, pH 8.0 and lysed by high-pressure homogenization. The protein was purified to homogeneity using Ni-NTA affinity, cation exchange, a second Ni-NTA affinity, and size exclusion chromatography (in this order). HRV3C protease was added for tagged His6-NusA cleavage after the first Ni-NTA affinity chromatography.

[0150] Human ADAR1 p110 isoform was expressed in SF9 cells as a recombinant protein with an N-terminal Twin-Strep-tag and purified by Strep-affinity chromatography. The cells were suspended and lysed in lysis buffer (20 mM Tris-HCl pH 8.0, 500 M NaCl, 1 mM TCEP, 0.55% Triton-X100, 1 mM PMSF, protease inhibitor (Sangon Biotech), and 100 ng / mL RNase A) and purified by Strep-affinity chromatography. The protein was eluted with 50 mM Tris-HCl, 200 mM KCl, 10% glycerol, 1 mM TCEP, 2.5 mM D-desthiobiotin.

[0151] dsRNA Preparation and In Vitro dsRNA Editing All dsRNAs were transcribed in vitro using T7 RNA polymerase. The two complementary strands were transcribed and purified separately. The pUC19 plasmid containing the target sequence was linearized with EcoRI, extracted with phenol chloroform, and precipitated with isopropanol. The in vitro transcription reaction was carried out at 37 °C for 4 h in 100 mM HEPES-K (pH 7.9), 10 mM MgCl2, 10 mM DTT, 6 mM NTPs each, 2 mM spermidine, 200 μg / mL linearized plasmid, and 100 μg / mL T7 RNA polymerase. The DNA template was digested with DNase I after the reaction. The transcripts were purified by 8% denaturing urea PAGE, extracted from the gel slices with 0.3 M sodium acetate, and precipitated with isopropanol. The RNAs of both complementary strands were mixed at a 1:1 molar ratio in annealing buffer (20 mM Tris-HCl pH 7.5, 50 mM NaCl, 1 mM EDTA), heated to 90 °C for 3 min, and then slowly cooled to room temperature. For in vitro ADAR1 dsRNA editing, 0.64 μg of annealed dsRNA was diluted and mixed with 0.4 nmol of purified hADAR1-p110 to a total volume of 100 μL in a buffer containing 50 mM Tris-HCl pH 7.5, 60 mM KCl, 4% glycerol, 0.002% NP-40, 1 mM DTT, 1 mM EDTA, 1 mg / mL BSA, and 0.4 U / μL recombinant RNase inhibitor (Takara). The reaction was incubated at 37 °C for 60 min, and the edited RNA was purified using the Absolutely RNA Nanoprep Kit (Agilent).

[0152] Negative Staining Electron Microscopy Samples containing 0.38 μM HsMDA5(298 - 1025) and 3.6 ng / μL dsRNA (regardless of length) were incubated with 1 mM ADP·AlF4 on ice for 60 minutes in a buffer containing 20 mM HEPES, 100 mM NaCl, 2 mM MgCl2, 2 mM DTT, and pH 7.5. 5 μl of the sample was applied to a glow-discharged 300-mesh carbon-coated copper grid (Beijing Zhongjingkeyi Technology), stained with 0.75% uranyl formate, and air-dried. Data were collected using a Talos L120C transmission electron microscope equipped with a 4Kx4K CETA CCD camera (FEI). Images were recorded at a nominal magnification of 45000x corresponding to a pixel size of 3.17 Å per pixel. Filament lengths were measured using ImageJ.

[0153] Cell line HEK293T-ADAR1-E912A-iMDA5-mCherry-pIFN-Lucia cells were maintained in DMEM (Gibco, 11995-065) supplemented with 10% FBS (Gibco, 16140-071). This cell line was generated by transducing the HEK293T-ADAR1-E912A ADAR1-editing deficient cell line with three lentiviral vectors encoding the rtTA reverse transactivator, a doxycycline-inducible construct encoding MDA5 linked to mCherry, and secreted luciferase Lucia (InvivoGen) under the control of an IFN-responsive promoter.

[0154] Plasmid construction pK-mC3-CTSA-PLTP-mR3 was constructed using gene fragments synthesized by Twist Bioscience (San Francisco, California) and standard molecular cloning methods. This plasmid contains mClover3 cassette and mRuby3 cassette facing each other, both under the control of separate EF-1α promoter and CAG enhancer. To mimic the endogenous overlapping state of these genes in the human genome, the 3'-UTR of CTSA is attached to mClover3, and the 3'-UTR of PLTP is attached to mRuby3. A scrambled version was assembled on the same backbone as the WT with a randomized 208bp overlapping region sequence while maintaining complementarity between the sense and antisense strands. pKER-mClover3 containing an mClover3 expression cassette driven by the EF-1α promoter and CAG enhancer and a rabbit β-globin terminator was cloned into the same vector backbone and used as an ssRNA control.

[0155] Overexpression assay of immunogenic dsRNA candidates HEK293T-ADAR1-E912A-iMDA5-mCherry-pIFN-Lucia cells were plated at 100,000 cells / well in a poly-L-lysine coating (Millipore Sigma, A-005-C) 24-well plate (BD Falcon, 353047). After 24 hours, cells were transfected with 500 ng / well of pK-mC3-CTSA-PLTP-mR3, pKER-mClover3, or Lipofectamine alone using Lipofectamine 3000 (Thermo-Fisher Scientific, L30001) according to the manufacturer's instructions. Twenty-four hours after transfection, the medium was aspirated and replaced with 500 μL of DMEM containing 0 μg / mL or 0.1 μg / mL of doxycycline. Twenty-four hours after adding doxycycline, the cells were washed with PBS and harvested with 0.5% trypsin. RNA was isolated from the cells using the Monarch Total RNA miniprep kit (NEB, T2010S). cDNA was generated from 500 ng of RNA using the iScript Advanced cDNA synthesis kit (Bio-Rad, 1725038). The cDNA was diluted approximately two-fold with nuclease-free water and subjected to qPCR using Kapa SYBR FAST 2x qPCR Master Mix and primers derived from Primer Bank (https: / / pga.mgh.harvard.edu / primerbank / ). 200 nM primers were used for each reaction and run on a Bio-Rad CFX96 according to the manufacturer's instructions for Kapa SYBR Fast. Data were analyzed using the ΔΔCt method. Each assay was performed as three separate biological replicates.

[0156] RT-PCR verification of plasmid expression To determine whether both cassettes on pK-mC3-CTSA-PLTP-mR3 were expressed, the isolated RNA was subjected to Turbo DNase (Life Technologies, AM2238) treatment to digest residual plasmids, purified with Monarch RNA Cleanup Kit (NEB, T2040L), and then subjected to cDNA synthesis with iScript Advanced cDNA synthesis kit. The cDNA was diluted approximately two-fold with nuclease-free water. Then, 1 μL of the diluted cDNA was amplified by PCR using OneTaq Quick-Load 2X Master Mix, Standard Buffer (NEB, M0486S), and primers designed to amplify from mClover3 to the CTSA 3'-UTR or primers designed to amplify from mRuby3 to the PLTP 3'-UTR. The resulting PCR products were subjected to electrophoresis on a 1% agarose gel.

[0157] Example 2 Here, the inventors developed novel genetic and computational approaches to identify immunogenic dsRNAs in various cell types. Using HEK293T cells and NPCs, the inventors identified common immunogenic dsRNA candidates and cell type-specific immunogenic dsRNA candidates. In both cell types, the inventors found, in contrast to previous findings, that only a small fraction of dsRNAs are immunogenic if not edited. As expected, in addition to identifying Alu repeats, the inventors discovered a new class of dsRNAs formed by overlapping genes transcribed in opposite directions. The inventors further characterized these immunogenic dsRNAs and identified features that distinguish immunogenic dsRNAs from non-immunogenic dsRNAs. The inventors' study provides a framework for revealing immunogenic dsRNAs in the ADAR1-dsRNA-MDA5 axis, thereby enabling future attempts to identify important dsRNAs with abnormal editing that may lead to autoimmune or immune-related diseases.

[0158] The inventors developed a strategy for comparing RNA editing at the total dsRNA level. Instead of relying on annotations such as Alu repeats, to identify dsRNAs in an unbiased manner, the inventors newly identified editing clusters as surrogates for long dsRNAs. To define an editing cluster, the inventors required that (1) there be at least five editing sites in the editing cluster and (2) the interval between any two adjacent sites be less than 120 nt. Using this new approach, the inventors identified 205,652 clusters in HEK293T cells, about 97% of which were in Alu and other repeats. The remaining approximately 3% of the clusters in non-repetitive regions had not been previously annotated as dsRNA candidates. By identifying editing clusters, the inventors were able to group all sites in dsRNAs and quantify the total editing level defined as the cluster editing index.

[0159] The inventors conducted a comparative analysis to search for editing clusters in which the editing index of ADAR1 p150-complemented cells was significantly larger than that of ADAR1 p110-complemented cells. The inventors found that most (about 96%) of the editing clusters were similarly edited by p150 and p110. This suggests that most of the dsRNAs have no or low immunogenicity. Among 39,079 testable editing clusters with sufficient coverage, the inventors identified 420 (1.1%) clusters located in 316 genes as putative immunogenic dsRNAs, and most of these were in untranslated regions (UTRs) or intergenic regions and not in introns. This is consistent with the prediction that immunogenic dsRNAs recognized by cytosolic MDA5 are embedded in processed RNAs exported to the cytoplasm. Some of the editing clusters preferably edited by p150 are edited to a lesser extent by p110, while others are specifically edited only by p150. In contrast, editing clusters highly edited by p110 were mainly present in introns. This is probably the result of overexpression of p110 in the nucleus. Since these intronic dsRNAs are removed during RNA splicing, they cannot be sensed by cytosolic MDA5 and will thus be non-immunogenic.

[0160] The inventors made similar findings when comparing the cluster editing index between ADAR2-complemented cells and ADAR2Δ NLS -complemented cells. The list of immunogenic dsRNAs identified from the comparison of p150 and p110 or ADAR2 and ADAR2Δ NLS was highly reproducible, especially for the immunogenic dsRNAs identified with high statistical significance.

[0161] Furthermore, the inventors compared the clustered editing index between WT and ADAR1p150− / − cells to identify immunogenic dsRNAs that are significantly less edited in ADAR1p150− / − cells. The inventors found that clusters preferred by p150 and clusters preferred by p110 are preferentially enriched in the cytoplasm and nucleus, respectively. These data support the inventors' discovery that immunogenic dsRNAs are preferentially edited in the cytoplasm.

[0162] The inventors examined whether the immunogenic dsRNAs identified in this study were enriched in the RNase protection assay. The inventors quantified the enrichment by the fold change in dsRNA expression in RNase-treated samples compared to untreated samples (in ADAR1 KO). The inventors found that immunogenic dsRNAs were significantly enriched for MDA5 binding compared to control dsRNAs (randomly selected non-immunogenic dsRNAs located in the 3'UTR at an equivalent expression level to the immunogenic dsRNAs).

[0163] The majority of the potentially immunogenic dsRNAs identified by the inventors were inverted repeats (90%) from the Alu family and other types of repetitive elements (8%). Intronic Alus constitute the majority of all Alus but are mostly excluded from the immunogenic group. The enrichment of immunogenic dsRNAs in mature mRNA was fully expected since they are sensed by MDA5 in the cytoplasm. Second, the inventors quantitatively compared the contributions of various dsRNA features for immunogenic Alu repeats versus non-immunogenic Alu repeats. The inventors found no significant differences in Alu length, percent base pairing, or expression level between the two groups. Instead, the prominent feature was the distance between the arms of the inverted Alu pair (i.e., loop size). The loop size of immunogenic inverted Alus was significantly smaller than that of their non-immunogenic counterparts.

[0164] To verify immunogenic dsRNA as an MDA5 substrate, the inventors examined the ability of immunogenic dsRNA to form MDA5 filaments in vitro using negative stain electron microscopy. The inventors observed a correlation between dsRNA length and MDA5 filament length. The inventors then chose to verify candidate dsRNAs that are inverted Alu, endogenous retrovirus (ERV) repeats, as well as cis-NATs. In all cases, the inventors observed filaments of various lengths depending on the dsRNA length.

[0165] Identification of Immunogenic dsRNA in Human Neural Progenitor Cells Evidence from several sources suggests that different cell types, tissues, and organisms have different repertoires of immunogenic dsRNAs. Neural progenitor cells were used as a prototype. To identify immunogenic dsRNAs in NPCs, the inventors performed a comparative editing analysis similar to that described above by the inventors. The inventors stably integrated the ADAR1 p150 or p110 isoform into ADAR1− / − NPCs immediately after differentiation. As expected, p150 suppressed ISG induction, while p110 did not. The inventors identified 950 (2.4%) immunogenic dsRNA candidates that were significantly more edited by p150 than by p110.

[0166] Only a subset (165) of immunogenic dsRNAs was found in HEK293T and NPCs. When the inventors restricted their analysis to genes expressed in both cell types, the sets of immunogenic dsRNAs almost completely overlapped. This suggests that the immunogenicity of dsRNA is intrinsically determined by its sequence and structural features, and its effect is influenced by host gene expression, which can vary across cell and tissue types. This means that as long as the identified immunogenic dsRNAs are expressed in any cell type, they will exhibit immunogenicity in other cell types.

[0167] Immunogenic dsRNAs Are Enriched in GWAS Signals of Common Inflammatory Diseases The inventors asked whether the immunogenic dsRNA candidates they identified were enriched in GWAS signals of common autoimmune and immune-related diseases. Indeed, the inventors found that immunogenic dsRNA candidates showed strong enrichment in GWAS signals of inflammatory diseases such as autoimmune thyroid disease, rheumatoid arthritis, type 1 diabetes, inflammatory bowel disease, etc., while non-immunogenic dsRNA candidates did not.

[0168] Example 1 above discloses hundreds of immunogenic dsRNA candidates that coexist at GWAS loci of autoimmune and immune-related diseases. This computational approach was successful in identifying disease-associated dsRNAs, but may be complemented by an unbiased search for immunogenic dsRNAs in a given cell type.

[0169] Method Novel identification of A-to-I RNA editing sites Novel identification of editing sites was performed on the pooled mapping results of HEK293T and NPC samples. The inventors called variants from the mapped RNA-seq reads using the UnifiedGenotyper tool from the Genome Analysis Toolkit (GATK) 80 and then, as previously described 77-78, parameters and filters were applied. Briefly, human variants in non-repetitive non-Alu and repetitive non-Alu sites were required to be supported by at least three mismatch reads. Variants in the human Alu region were required to be supported by one mismatch read. This set of variant candidates required several filtering steps to increase the accuracy of editing site discovery. The inventors also removed all known SNPs present in dbSNP (excluding SNPs of molecular type "cDNA"; database version 135) from the 1000 Genomes Project and the University of Washington Exome Sequencing Project. The inventors also performed high-editing specificity to recover editing sites from unmapped reads. Finally, the inventors combined the known editing sites from RADAR with the newly identified sites from each cell line to obtain 3,949,224 editing sites in HEK293T and 2,982,889 editing sites in NPC.

[0170] Identification of A-to-I RNA editing clusters Without relying on any previous annotations, to identify long dsRNAs that are potential MDA5 ligands, first, the inventors determined that all of these long dsRNAs should be highly edited to suppress their immunogenicity. The inventors obtained the location information of the editing sites within the genome as a proxy for dsRNA and applied a density-based clustering algorithm (DBSCAN) to identify genomic regions where many editing sites are close, which the inventors call editing clusters. DBSCAN required two parameters to define clusters: the minimum number of points (editing sites) required to form a cluster (minPts) and the farthest distance (eps) between two adjacent sites within the same cluster. For the first parameter (minPts), as expected, the inventors observed that the total number of clusters decreased as the minimum number of editing sites per cluster increased. The inventors set minPts = 5 as the minimum number of mismatches (edits). For the second parameter (eps), by plotting the distances between adjacent sites classified from shortest to longest, the inventors observed a "knee" point on the curve. At this point, eps = 120. Since this "knee" point can include >90% of the editing sites in the editing clusters while excluding distant outliers, the inventors concluded that eps = 120 was the optimal value. By applying DBSCAN clustering, the inventors identified 205,652 and 175,655 editing clusters in HEK293T and NPC cell lines, respectively. Furthermore, after applying a coverage cutoff (≥20 reads), the inventors obtained 39,079 and 40,182 editing clusters in HEK293T and NPC cell lines, respectively, for downstream analysis.

[0171] Quantification and comparison of the editing index The inventors adopted the concept of "editing index" as a measure of the editing level of editing clusters. For each cluster containing n editing sites, its editing index was Calculated as TIFF2025522326000013.tif11128. To ensure high reliability of the measured values of the editing index, the inventors required that it be TIFF2025522326000014.tif5128. For each editing cluster that met the above threshold under both conditions in the comparison, a Fisher's direct test was applied to the number of G-reads and A-reads to calculate the p-value for the significance of the unequal editing index. The multiple-testing p-values were adjusted using the Benjamini-Hochberg method.

[0172] (Table 2) TIFF2025522326000015.tif218147TIFF2025522326000016.tif218148TIFF2025522326000017.tif21865

[0173] References TIFF2025522326000018.tif212161TIFF2025522326000019.tif212161TIFF2025522326000020.tif212161TIFF2025522326000021.tif219161TIFF2025522326000022.tif212161TIFF2025522326000023.tif212161TIFF2025522326000024.tif11161

[0174] The foregoing are merely examples of the principles of the present invention. It will be understood by those skilled in the art that, although not explicitly described or shown herein, they can embody the principles of the present invention and devise various arrangements that are within the spirit and scope of the present invention. Furthermore, all of the examples and conditional language described herein are primarily intended to assist the reader in understanding the principles of the present invention and the concepts contributed by the inventors to the development of the art, and should be construed as being such specific examples and conditions, but not limited thereto. Moreover, all of the language in this specification that describes the principles, aspects, and embodiments of the present invention, as well as its specific examples, is intended to encompass both structural and functional equivalents thereof. Further, such equivalents are intended to include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function regardless of structure. Accordingly, the scope of the present invention is not intended to be limited to the exemplary aspects shown and described herein. More precisely, the scope and spirit of the present invention are embodied by the appended claims.

Claims

**Claim 1** A method for determining an individual's immunogenic double-stranded RNA (dsRNA) load score, comprising: Genotyping a subject at a plurality of risk alleles associated with changes in RNA editing levels related to an immune-related disease of interest; To create a dsRNA load score, (1) the effect size of the risk allele association with the disease of interest, (2) the effect size of the risk allele association with the RNA editing level, and (3) the expression level of immunogenic dsRNA in a tissue or cell associated with the disease are combined to create an immunogenic dsRNA load score by weighting the number of risk alleles that decrease the RNA editing level. comprising: The dsRNA load score is associated with an individual's predisposition to the disease of interest due to unwanted activation of the IFN immune response and responsiveness to therapies that reduce the clinical sequelae of unwanted activity in the dsRNA sensing pathway. The method. **Claim 2** The method of claim 1, wherein at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000 SNPs are genotyped. **Claim 3** The method of claim 1 or claim 2, further comprising genotyping the individual at one or more alleles associated with MDA5 or ADAR1. **Claim 4** The method according to any one of claims 1 to 3, wherein the disease of interest affects a specific tissue or cell type of interest. **Claim 5** The method according to any one of claims 1 to 4, further comprising determining the IFN response level and the expression level of dsRNA in the tissue of interest or cells derived therefrom. **Claim 6** The method according to any one of claims 1 to 5, wherein the dsRNA is selected by computational co-occurrence at GWAS loci of the disease of interest. **Claim 7** The dsRNA is identifying an editing cluster having at least 5 editing sites, wherein any two adjacent sites are spaced less than 120 nt apart; Determining editing clusters having an editing index significantly greater in ADAR1 p150-complemented cells than in ADAR1 p110-complemented cells, wherein editing clusters preferably edited by p150 are selected as immunogenic dsRNA candidates The method according to any one of claims 1 to 6, as determined by

8. wherein the immunogenic dsRNA score is Determining the expression level of a tissue or cell associated with a disease and immunogenic dsRNA derived therefrom The method according to any one of claims 1 to 7, further comprising

9. The tissue or cell associated with the disease is (a) Measuring an interferon (IFN) score based on the aggregated expression level of interferon-induced genes (ISGs) to determine the tissue or cell with the highest level, and (b) Measuring the aggregated expression level of immunogenic dsRNA in disease tissue scRNA-seq The method according to claim 8, selected by

10. The method according to any one of claims 1 to 9, wherein the expression levels of at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000 dsRNA loci are determined

11. The method according to any one of claims 1 to 10, further comprising the step of inputting additional clinical characteristics, age, gender, family history, age of onset of the disease, duration of the disease

12. A predetermined threshold of the score is set by training in a genome-wide association study for evaluating the distribution of IDS between cases and controls for the disease to determine an IDS threshold for predicting a high risk for the disease of interest to the subject

13. The method according to any one of claims 1 to 12, wherein the individual has previously been diagnosed with the disease of interest

14. The method according to any one of claims 1 to 13, wherein the disease of interest is an autoimmune disease

15. The method according to any one of claims 1 to 14, wherein the disease of interest is an inflammatory disease

16. The method according to any one of claims 1 to 15, wherein the disease of interest is autoimmune thyroid disease, celiac disease, inflammatory bowel disease, systemic lupus erythematosus, multiple sclerosis, primary biliary cirrhosis, psoriasis, type 1 diabetes (T1D), rheumatoid arthritis, vitiligo; atopic conditions such as asthma and atopic dermatitis; and other conditions with a significant inflammatory component, such as amyotrophic lateral sclerosis (ALS), coronary artery disease, type 2 diabetes, Parkinson's disease, systemic sclerosis, schizophrenia, Alzheimer's disease, and selected from the levels of high-density lipoprotein and low-density lipoprotein and triglycerides.

17. The method according to any one of claims 1 to 16, wherein the determination of the dsRNA load score is performed as a program of executable instructions by a computer and is performed by a software component loaded on the computer.

18. The method according to any one of claims 1 to 17, further comprising the step of treating the individual according to the determined dsRNA load score.

19. The immunogenic dsRNA load score (Ψ) is given by the formula wherein is the effect size of the jth SNP for the disease in the kth tissue / cell type, is the effect size of the ith dsRNA having the jth SNP for the RNA editing level in the kth tissue / cell type, is the expression level of the ith dsRNA in the kth tissue / cell type, the method according to any one of claims 1 to 18.

20. The method according to claim 19, wherein the dsRNA load score further comprises one or more clinical components.

21. The method according to any one of claims 1 to 20, wherein the individual is identified as responsive to MDA5 inhibition therapy by the dsRNA load score.

22. The method according to any one of claims 1 to 21, wherein the dsRNA load score identifies individuals having an increased MDA5-mediated immune response due to a lack of dsRNA editing in the tissue of interest.