Methods and systems for decoding gene expression in individual cell types and disease states

A high-resolution machine learning model using single-cell RNA sequencing data predicts cell type-specific gene expression, addressing the limitations of existing models by improving predictive accuracy and enabling advanced biotechnology and pharmaceutical applications.

WO2026030709A1PCT designated stage Publication Date: 2026-02-05GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/040340
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-07
Filing Date
2025-08-01
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing machine learning models for predicting gene expression are trained on low-resolution data from healthy human tissues, failing to distinguish between different cell types and predict how gene behavior is affected by disease.

Method used

A machine learning-based approach that processes long nucleic acid sequences at high resolution using single-cell RNA sequencing data from over 22 million cells, enabling cell type-specific and cell state-specific gene expression predictions, and incorporates a directed evolution process for in silico design and optimization of DNA regulatory sequences.

Benefits of technology

The model achieves improved predictive performance, allowing for the identification of cis-regulatory mechanisms, prediction of non-coding variant effects, and design of context-specific regulatory DNA elements, enhancing biotechnology and pharmaceutical applications such as gene therapies and CRISPR-based drug delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025040340_05022026_PF_FP_ABST
    Figure US2025040340_05022026_PF_FP_ABST
Patent Text Reader

Abstract

This present disclosure relates to machine learning-based methods and systems for making cell type- and cell state-specific predictions of gene expression for a given input DNA sequence, and for designing and / or optimizing DNA regulatory elements. In some instances, for example, the disclosed methods for predicting gene expression levels from genomic sequence data can comprise: receiving genomic sequence data comprising a DNA sequence for a specified gene; inputting the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: (i) trained on single cell gene expression data for a plurality of cell types and cell states, and (ii) configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific gene expression level; and outputting a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene.
Need to check novelty before this filing date? Find Prior Art

Description

ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 1 of 126 METHODS AND SYSTEMS FOR DECODING GENE EXPRESSION IN INDIVIDUAL CELL TYPES AND DISEASE STATES CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the priority benefit of United States Provisional Patent Application Serial No.63 / 678,900, filed August 2, 2024, of United States Provisional Patent Application Serial No.63 / 689,473, filed August 30, 2024, of United States Provisional Patent Application Serial No. 63 / 703,434, filed October 4, 2024, and of United States Provisional Patent Application Serial No.63 / 742,779, filed January 7, 2025, the contents of each of which are incorporated herein by reference in their entirety. FIELD

[0002] This present disclosure generally relates to machine learning-based methods and systems for the analysis of genomic data, and more specifically to machine learning models configured to predict the activity of a given gene in a variety of cell types and disease conditions based on its DNA sequence. BACKGROUND

[0003] The human genome contains over 20,000 genes, which encode protein and RNA molecules that give rise to the building blocks of human cells. However, the human body contains many types of cells, with dramatically different structure and function. This diversity arises because each cell type exhibits a different pattern of gene expression (the level of activity of each gene). Further, these expression patterns may be affected by disease, giving rise to unhealthy cell states with compromised function.

[0004] The expression pattern of a gene is determined in large part by the DNA sequence of the gene and its neighboring regions. The genomic DNA sequence thus contains features that determine every cell state in the human body, as well as features that influence how these cell states will be altered in disease conditions. Therefore, a critical problem in biology is how to predict the expression pattern of a gene from its surrounding DNA sequence.

[0005] DNA sequence-based machine learning models that predict gene expression for input genomic DNA sequences have proven valuable for many biological tasks, including understanding cis-regulatory syntax and interpreting non-coding genetic variants. However,MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 2 of 126 existing state-of-the-art models for prediction of gene expression have been trained on low- resolution gene expression data from healthy human tissues. As a result they cannot distinguish between different types of cells in the same tissue, nor can they predict how the behavior of a gene will be affected by disease. SUMMARY

[0006] Disclosed herein are machine learning-based methods, systems, and programming for decoding gene expression using one or more machine-learning models that collectively are configured to process an input DNA sequence (e.g., an input gene sequence and its flanking sequence regions) and output a prediction of cell type- and cell state-specific gene expression.

[0007] As noted above, current state-of-the-art models for prediction of gene expression have been trained largely on bulk expression profiles from healthy tissues or cell lines. Thus, they lack the ability to perform gene expression prediction tasks at the resolution of specific cell types or cell states (e.g., disease states) across a diverse set of tissue and disease contexts. The disclosed machine learning-based methods and model address this gap through a combination of: (i) a model architecture that enables processing of long nucleic acid sequences at high resolution, and (ii) the use of a training data set comprising single-cell or single-nucleus RNA sequencing data from over 22 million cells. The trained model successfully predicts the cell type-specific expression of previously unseen gene sequences based on their DNA sequence alone. A significant advantage of the presently disclosed model over bulk-trained models is the possibility of linking variants to phenotypes via cell type-specific effects

[0008] Also disclosed herein are machine learning-based methods for performing in silico design and / or optimization of cell type-specific and / or cell state-specific DNA regulatory sequences using long-context DNA sequence-based prediction models and a directed evolution process. The disclosed methods are based on iterative in silico mutagenesis of a starting sequence (e.g., a random sequence or candidate sequence) and scoring of regulatory sequence (e.g., promoter sequence) efficacy based on predicted gene expression level. In some instances, the disclosed machine learning-based methods and models may be configured to process one or more input DNA sequences (e.g., in the context of a construct comprising the regulatory sequence and a cargo gene sequence) using one or more long-context DNA sequence-based prediction models (e.g., one or more different DNA sequence-prediction models configured to predict gene expression), while also evaluating one or more genomic insertion sites as part ofMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 3 of 126 the design and / or optimization of a cell type-specific and / or cell state-specific DNA regulatory sequence.

[0009] The methods and systems described herein provide numerous technical advantages in terms of improved predictive and / or computational performance that derive from the model architecture and training approach used. For example, the model architecture is scalable to allow processing of very large data sets. The disclosed models are trained and / or fine-tuned on single-cell gene expression data (e.g., single cell RNA-seq data comprising sequence read counts as a function of genomic locus, and aggregated by cell type to create pseudobulk samples) using a two-component loss function (applied across the pseudobulk axis for a single gene rather than within individual pseudobulks) that consists of a Poisson loss applied to the total expression value of the gene across all pseudobulks, and a multinomial loss function that compares the true and predicted distribution of expression values across pseudobulks for the same gene, resulting in dramatically improved cell type- and / or cell state-specific prediction performance. The model output comprises a vector of predicted gene expression values for a specified gene in each of a specified list of cell types.

[0010] Due to their improved predictive performance, the disclosed machine learning-based methods and models can be used in a variety of important biotechnology and pharmaceutical industry applications. For example, the model output may be used to identify the cis-regulatory mechanisms driving cell type-specific gene expression and their changes in disease (e.g., to identify and / or validate new drug targets and therapeutic mechanisms of action), to predict the effect of non-coding variants at cell type resolution (e.g., to further our understanding of the regulation of gene expression and how changes to regulation of gene expression contribute to different disease states by affecting gene expression, protein production, and / or cellular function), and to design regulatory DNA elements with precisely tuned, context-specific functions (e.g., to manipulate gene expression in vivo). These capabilities can, in turn, improve the efficacy of downstream development and manufacturing processes in the biotechnology and pharmaceutical industries (e.g., by facilitating the design and development of effective gene therapies (e.g., gene therapies directed to turning off expression of mutated genes, or to promoting expression of beneficial genes), CRISPR-based drug delivery mechanisms (e.g., to improve the efficacy and precision of gene editing tools), and protein expression systems for biologics manufacturing (e.g., to maximize yields of expressed proteins), etc.).MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 4 of 126

[0011] Disclosed herein are methods for predicting gene expression levels from genomic sequence data, the methods comprising: receiving, at one or more processors, genomic sequence data comprising a DNA sequence for a specified gene; inputting, using the one or more processors, the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell- type-specific and / or a cell-state-specific gene expression level; and outputting, using the one or more processors, a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene. In some embodiments, the trained machine learning model is trained using a two-component loss function comprising a Poisson loss component and a multinomial negative log-likelihood component.

[0012] In some embodiments, the method further comprises using the trained machine learning model to predict an impact of a gene sequence variant on gene expression level.

[0013] In some embodiments, the method further comprises using the trained machine learning model to predict gene expression patterns for a plurality of specified genes.

[0014] In some embodiments, the method further comprises using the trained machine learning model to predict differences in gene expression patterns for a plurality of specified genes in healthy versus diseased tissues.

[0015] In some embodiments, the method further comprises: (i) using the trained machine learning model to identify and compare gene sequences that are highly expressed in a given cell type, and (ii) clustering sequence features in the highly expressed gene sequences to identify transcription factor binding motifs.

[0016] In some embodiments, the method further comprises using the trained machine learning model to design sequences for gene therapy, wherein the designed sequences are predicted to be actively expressed only in specific disease states.

[0017] In some embodiments, the genomic sequence data comprises the DNA sequence for the specified gene and a flanking sequence for at least one of a 3’ end or a 5’ end of the specified gene sequence. In some embodiments, the genomic sequence data comprises a DNA sequence ranging in length from about 300,000 nucleotides to about 700,000 nucleotides. In some embodiments, the genomic sequence data comprises a DNA sequence of about 500,000MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 5 of 126 nucleotides in length. In some embodiments, the genomic sequence data is provided as input to the machine learning model as a one-hot encoded matrix.

[0018] In some embodiments, the trained machine learning model comprises a deep neural network architecture. In some embodiments, the deep neural network architecture comprises at least one convolution layer. In some embodiments, the deep neural network architecture comprises at least one downsampling layer. In some embodiments, the deep neural network architecture comprises at least one self-attention layer. In some embodiments, the deep neural network architecture comprises at least one deconvolution layer with matched U-net connections. In some embodiments, the deep neural network architecture comprises at least one global mean pooling layer. In some embodiments, the deep neural network architecture comprises a linear output layer.

[0019] In some embodiments, the trained machine learning model is pre-trained on bulk gene expression profile data and / or epigenomic assay data. In some embodiments, the bulk gene expression profile data comprises RNA-seq data and / or CAGE-seq data. In some embodiments, the epigenomic assay data comprises DNase-seq data, ATAC-seq data, and / or ChIP-seq data.

[0020] In some embodiments, the single cell gene expression data comprises single cell RNA sequencing data.

[0021] In some embodiments, the plurality of cell types and cell states comprises at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 different types of cell.

[0022] In some embodiments, the plurality of cell types and cell states comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130 different cell states. In some embodiments, the different cell states comprise different development states, different metabolic states, different disease states, or any combination thereof.

[0023] In some embodiments, the output comprises a prediction of at least 10, at least 25, at least 50, at least 75, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000,MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 6 of 126 at least 8,000, at least 9,000, or at least 10,000 cell-type-specific and / or cell-state-specific gene expression levels for the specified gene.

[0024] Also disclosed are systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive genomic sequence data comprising a DNA sequence for a specified gene; input the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type-specific and / or a cell- state-specific gene expression level; and output a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene.

[0025] Also disclosed herein are non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive genomic sequence data comprising a DNA sequence for a specified gene; input the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific gene expression level; and output a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene.

[0026] Disclosed herein are methods for designing a DNA regulatory sequence efficacy, the methods comprising: receiving, at one or more processors, at least one candidate DNA regulatory sequence; receiving, at the one or more processors, a cargo gene sequence; generating, using the one or more processors, at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; inputting, using the one or more processors, the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level; determining, using the one or more processors, an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and selecting,MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 7 of 126 based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence.

[0027] In some embodiments, the at least one candidate DNA regulatory sequence comprises at least one promoter sequence.

[0028] In some embodiments, the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and the method further comprises: for each of a plurality of iterations: generating a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; inputting the construct sequence into the trained machine learning model; determining an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: updating, using the one or more processors, the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or outputting, using the one or more processors, an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence.

[0029] In some embodiments, the method further comprises repeating the method for each of a plurality of candidate insertion sites. In some embodiments, the method further comprises repeating the method using at least one additional trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state- specific gene expression level. In some embodiments, the method further comprises using the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence to develop a gene therapy or personalized anti-cancer vaccine.

[0030] In some embodiments, the initial candidate DNA regulatory sequence comprises a random DNA sequence.

[0031] In some embodiments, updating the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence comprises performing at least one single nucleotide substitution, single nucleotide insertion, or single nucleotide deletion. In some embodiments, updating initial candidate DNA regulatory sequence or the candidate DNAMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 8 of 126 regulatory sequence comprises performing at least one insertion, deletion, or substitution of a plurality of nucleotides. In some embodiments, updating the initial candidate DNA regulatory sequence or candidate DNA regulatory sequence comprises inserting nucleotides corresponding to at least one known regulatory sequence motif.

[0032] In some embodiments, the improved DNA regulatory sequence comprises a promoter sequence.

[0033] In some embodiments, the initial candidate DNA regulatory sequence is between 100 and 1,000 nucleotides in length.

[0034] In some embodiments, the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence is between 100 and 1,000 nucleotides in length.

[0035] In some embodiments, the prediction of cell-type-specific and / or a cell-state-specific gene expression level is made at single nucleotide resolution.

[0036] In some embodiments, the efficacy score for the at least one candidate DNA regulatory sequence is based on a determination of a maximum predicted cell-type-specific and / or a cell-state-specific gene expression level, a mean predicted cell-type-specific and / or a cell-state-specific gene expression level, a median predicted cell-type-specific and / or a cell- state-specific gene expression level, or an integrated predicted cell-type-specific and / or a cell- state-specific gene expression level across the cargo gene sequence, or any combination thereof. In some embodiments, the efficacy score for the at least one DNA regulatory sequence is further based on a determination of predicted cell-type-specific and / or a cell-state-specific off-target gene expression levels.

[0037] In some embodiments, the at least one candidate insertion site comprises a safe harbor locus. In some embodiments, the genome is the human genome.

[0038] In some embodiments, the trained machine learning model comprises a long-context DNA sequence prediction model. In some embodiments, the trained machine learning model comprises a Borzoi model, an Enformer model, or a Decima model.

[0039] In some embodiments, the at least one construct sequence comprising a candidate DNA regulatory sequence and the specified cargo gene is between 100 kilobases and 1 megabase in length.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 9 of 126

[0040] In some embodiments, sequence data for the at least on candidate DNA regulatory sequence is provided as input to the trained machine learning model as a one-hot encoded matrix.

[0041] In some embodiments, the trained machine learning model comprises a deep neural network architecture. In some embodiments, the deep neural network architecture comprises at least one convolution layer. In some embodiments, the deep neural network architecture comprises at least one downsampling layer. In some embodiments, the deep neural network architecture comprises at least one self-attention layer. In some embodiments, the deep neural network architecture comprises at least one deconvolution layer with matched U-net connections. In some embodiments, the deep neural network architecture comprises at least one global mean pooling layer. In some embodiments, the deep neural network architecture comprises a linear output layer.

[0042] In some embodiments, the machine learning model is pre-trained on bulk gene expression profile data and / or epigenomic assay data. In some embodiments, the bulk gene expression profile data comprises RNA-seq data and / or CAGE-seq data. In some embodiments, the epigenomic assay data comprises DNase-seq data, ATAC-seq data, and / or ChIP-seq data.

[0043] In some embodiments, the trained machine learning model is trained on single cell gene expression data for a plurality of cell types and cell states.

[0044] In some embodiments, the single cell gene expression data comprises single cell RNA sequencing data.

[0045] In some embodiments, the plurality of cell types and cell states comprises at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 different types of cell.

[0046] In some embodiments, the plurality of cell types and cell states comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130 different cell states. In some embodiments, the different cell states comprise different development states, different metabolic states, different disease states, or any combination thereof.

[0047] Disclosed herein are systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructionsMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 10 of 126 that, when executed by the one or more processors, cause the system to: receive at least one candidate DNA regulatory sequence; receive a cargo gene sequence; generate at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; input the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level; determine an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and select, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state- specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence.

[0048] In some embodiments, the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and the system further comprises instructions that, when executed by the one or more processors, cause the system to: for each of a plurality of iterations: generate a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; input the construct sequence into the trained machine learning model; determine an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: update the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or output an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence.

[0049] Disclosed herein are non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive at least one candidate DNA regulatory sequence; receive a cargo gene sequence; generate at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; input the at least one construct sequenceMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 11 of 126 into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level; determine an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and select, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence.

[0050] In some embodiments, the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and computer readable storage medium further comprises instructions that, when executed by the one or more processors, cause the system to: for each of a plurality of iterations: generate a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; input the construct sequence into the trained machine learning model; determine an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: update the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or output an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence.

[0051] Also disclosed herein are methods for predicting mRNA properties from genomic sequence data, the methods comprising: receiving, at one or more processors, genomic sequence data comprising a DNA sequence for a specified gene; inputting, using the one or more processors, the genomic sequence data to a machine learning model, wherein the machine learning model is configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific property of an mRNA molecule corresponding to a specified gene; and outputting, using the one or more processors, a prediction of a cell-type-specific and / or a cell-state-specific property of an mRNA molecule corresponding to the specified gene, wherein the predicted cell-type-specific and / or a cell-state-specific property comprises at least one of: (i) a predicted mRNA stability, or (ii) a predicted mRNA translation efficiency.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 12 of 126

[0052] In some embodiments, the method further comprises using the predicted mRNA stability and / or predicted mRNA translation efficiency to predict an expression level for a protein corresponding to the specified gene.

[0053] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims. INCORPORATION BY REFERENCE

[0054] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0056] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:

[0057] FIG. 1 provides a block diagram of an example prediction system, in accordance with some embodiments of the present disclosure.

[0058] FIG.2 provides a non-limiting, schematic illustration of a machine learning model architecture, in accordance with some embodiments of the present disclosure.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 13 of 126

[0059] FIG. 3 provides a non-limiting example of a flowchart for a process for training a machine learning model to predict cell type-specific and / or cell state-specific gene expression levels based on an input DNA sequence, in accordance with some embodiments of the present disclosure.

[0060] FIG.4 provides a non-limiting example of a flowchart for a process for predicting cell type-specific and / or cell state-specific gene expression levels, in accordance with some embodiments of the present disclosure.

[0061] FIG. 5 provides a non-limiting example of a process flowchart for using a long- context DNA sequence-based prediction model to perform automated DNA regulatory sequence design and optimization, in accordance with some embodiments of the present disclosure.

[0062] FIG. 6 provides a non-limiting, schematic illustration of a process for training a machine learning model to predict cell type-specific and / or cell state-specific gene expression levels based on an input DNA sequence, in accordance with some embodiments of the present disclosure.

[0063] FIG.7 provides a non-limiting example of box plots showing the distribution of the Pearson correlation between measured and predicted expression values for test set genes, as predicted by the disclosed model for pseudobulks from different human organs.

[0064] FIG.8 provides a nonlimiting example of box plots showing the distribution of the Pearson correlation between measured and predicted expression values for test set genes, for each pseudobulk annotated with a disease state.

[0065] FIG.9 provides non-limiting examples of precision-recall curve for classification of high-confidence sc-eQTLs from negative control variants, using either the Decima model’s predictions in the matched cell type (Decima) or using other model predictions.

[0066] FIGS. 10A-10C provide non-limiting examples of scatter plots showing comparisons of AUPRC data for classification of high-confidence sc-eQTLs from negative control variants in each of 21 cell types. Each point represents a cell type. FIG.10A: Decima predictions versus Borzoi (whole blood) predictions. FIG. 10B: Decima predictions versus Borzoi (matched) predictions. FIG.10C: Decima versus Distance-based predictions.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 14 of 126

[0067] FIG.11 provides a non-limiting schematic illustration of a machine learning model configured to predict cell type-specific and / or cell state-specific gene expression levels, in accordance with some embodiments of the present disclosure.

[0068] FIG. 12 provides a non-limiting schematic illustration of the use of a machine learning model to reconstruct the expression profiles of unseen genes, in accordance with some embodiments of the present disclosure.

[0069] FIG.13 provides a non-limiting example of data illustrating a histogram of Pearson correlation coefficients between measured and predicted gene expression vectors for each of a set of pseudobulk samples over a test set of genes.

[0070] FIG.14 provides a non-limiting example of data illustrating a histogram of Pearson correlation coefficients between measured and predicted per-gene expression vectors for each of a test set of genes.

[0071] FIG. 15 provides a non-limiting schematic illustration of the use of a machine learning model configured to predict cell type-specific and / or cell state-specific gene expression levels to identify sequence determinants of cell type specificity, in accordance with some embodiments of the present disclosure.

[0072] FIG.16 provides a non-limiting example of data illustrating a histogram of showing the Area Under the Receiver Operator Characteristic (AUROC) for classification of specific versus nonspecific genes in different cell types based on the disclosed machine learning model’s predictions in the same cell type.

[0073] FIG.17 provides a non-limiting example of a scatter plot showing the measured and predicted expression of the FABP1 gene in a set of pseudobulk samples.

[0074] FIG. 18 provides a non-limiting example of boxplots showing the measured and predicted expression of the FABP1 gene in pseudobulk samples representing enterocytes, compared to all other cell types.

[0075] FIG.19 provides a non-limiting example of a scatter plot showing the measured and predicted expression of the DNAH6 gene in a set of pseudobulk samples.

[0076] FIG. 20 provides a non-limiting example of boxplots showing the measured and predicted expression of the DNAH6 gene in pseudobulk samples representing ependymal cells, ciliated cells, or other cell types.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 15 of 126

[0077] FIG.21 provides a non-limiting example of a scatter plot showing the measured and predicted expression of the SPI1 gene in a set of pseudobulk samples.

[0078] FIG.22 provides a non-limiting examples of boxplots showing the expression of the SPI1 gene in pseudobulk samples representing specific cell types.

[0079] FIG. 23 provides a non-limiting example of data for per-nucleotide attribution scores averaged across different sequence elements in the gene body and promoter, for all test set genes.

[0080] FIG. 24 provides a non-limiting example of data for average per-nucleotide attribution scores within different distance ranges from the gene body, for all test set genes.

[0081] FIG. 25 provides a non-limiting example of data for pseudobulk scATAC-seq coverage and smoothed per-nucleotide attribution scores for predicted FABP1 expression. Upper: pseudobulk scATAC-seq coverage over a 64 kb genomic region surrounding the FABP1 gene. Lower: smoothed per-nucleotide attribution scores for predicted FABP1 expression over the same genomic region.

[0082] FIG.26 provides a non-limiting example of data for per-nucleotide Input x Gradient attribution scores over the 50 nucleotides upstream of the FABP1 transcription start site.

[0083] FIG.27 provides a non-limiting example of data for per-nucleotide Input x Gradient attribution scores over a distal CRE.

[0084] FIG. 28 provides a non-limiting example of box plots showing the measured expression of the CEBPA, CDX2 and RXRA genes in pseudobulk samples representing enterocytes and hepatocytes, compared to pseudobulk samples representing other cell types in the gut and liver.

[0085] FIG. 29 provides a non-limiting schematic illustration of the use of a machine learning model configured to predict cell type-specific and / or cell state-specific gene expression levels to identify drivers of cell type identity, in accordance with some embodiments of the present disclosure.

[0086] FIG. 30 provides a non-limiting example of an RFX binding motif that was repeatedly identified in ciliated cells.

[0087] FIG. 31 provides a non-limiting example of data that illustrates elevated TF expression in ciliated cells relative to that in other lung cell types.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 16 of 126

[0088] FIG. 32 provides a non-limiting example of box plots showing the measured expression of the six TFs with the most cell type- or lineage-specific contribution weights across seven epithelial cell types in the human lung.

[0089] FIG.33 provides a non-limiting example of a scatterplot showing the measured and predicted log fold change in expression of 1,811 test set genes in pseudobulks representing neurons vs. pseudobulks representing other cell types in the brain.

[0090] FIG. 34 provides a non-limiting example of box plots showing the measured expression of the MYT1L gene in different cell types in the brain, along with the MYT1L motif found by TF-MoDISco based on the model’s differential attributions,

[0091] FIG. 35 provides a non-limiting example of box plots showing the measured expression of the REST gene in different cell types in the brain, along with the REST motif found by TF-MoDISco based on the model’s differential attributions

[0092] FIG.36 provides a non-limiting example of a scatter plot showing the measured and predicted log fold change in expression of a test set genes in cycling regulatory T cells (Treg cycling) vs. non-cycling regulatory T cells (Treg) in skin.

[0093] FIG. 37 provides a non-limiting example of the top two motifs identified by TF- MoDISco based on differential attribution scores for genes upregulated in Treg cycling vs. Treg in skin.

[0094] FIG.38 provides a non-limiting example of a scatterplot showing the measured and predicted log fold change in expression of a test set genes in cardiac fibroblasts vs. fibroblasts in other tissues.

[0095] FIG. 39 provides a non-limiting example of the top two motifs identified by TF- MoDISco based on differential attribution scores for genes upregulated in cardiac fibroblasts.

[0096] FIG. 40 provides a non-limiting schematic illustration of the use of a machine learning model configured to predict cell type-specific and / or cell state-specific gene expression levels to predict the effect of DNA sequence variants, in accordance with some embodiments of the present disclosure.

[0097] FIG.41 provides a non-limiting example of a bar plot showing model performance in sc-eQTL classification in each of 21 blood cell types in the OneK1K dataset.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 17 of 126

[0098] FIG. 42 provides a non-limiting example of a scatter plot showing the model’s predicted effect size for high-confidence sc-eQTLs compared to their measured effect size (beta value) in the same cell type.

[0099] FIG. 43 provides a non-limiting example of the model’s predicted absolute effect size for high-confidence sc-eQTL variants in cell types where the variant is identified as an sc- eQTL with PIP > 0.9, compared to other cell types in which the variant had no significant effect (PIP < 0.1) but its corresponding eGene is still expressed.

[0100] FIG.44 provides a non-limiting example of the model’s predicted effect sizes for a monocyte-specific sc-eQTL in various blood cell types.

[0101] FIG. 45 provides a non-limiting example of the model’s attributions for JAZF1 expression and nucleotide-resolution Input x Gradient attribution scores. Upper panel: the model’s attributions for JAZF1 expression in the sequences containing the reference and alternate allele (upper panel), showing that the variant (highlighted in blue) reduces the attribution over a distal enhancer. Lower panel: Nucleotide-resolution Input x Gradient attribution scores computed with respect to predicted expression of JAZF1 in monocytes, for the 160 bp sequence surrounding the reference and alternate alleles of the variant in E-F), highlighting a C / EBP motif disrupted by the variant.

[0102] FIG.46 provides a non-limiting example of the measured expression of the CEBPA gene over the same cell types shown in the upper panel, highlighting its elevated expression in monocytes.

[0103] FIG. 47 provides a non-limiting example of box plots showing the distribution of the absolute log fold changes predicted by the disclosed sequence-based gene expression model (“Decima”) for 837 high-confidence GWAS variants and 8370 matched negative control variants.

[0104] FIG. 48 provides a non-limiting example of a precision-recall curves for classification of GWAS variants from matched negative control variants, based on either the predictions shown in FIG.39, or on the distance of the variant from the gene transcription start site (TSS).

[0105] FIG. 49 provides a non-limiting example of a heat map showing the model’s predicted log fold change caused by GWAS variants for 6 traits or categories (x-axis) in selected cell types (y-axis).MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 18 of 126

[0106] FIG. 50 provides a non-limiting example of a scatter plot showing the measured gene expression (log(CPM)+1) for the FES gene across cell types (x-axis) and the model’s predicted absolute log fold change in this gene’s expression due to the rs138682554 variant (y- axis).

[0107] FIG. 51 provides a non-limiting example of the model’s attributions for the sequence surrounding the variant in the highlighted cell types (FIG. 28A) along with the matched motif from HOCOMOCO v12.

[0108] FIG. 52 provides a non-limiting example of a scatter plot showing the measured gene expression (log(CPM)+1) for the SLC1A5 gene across cell types (x-axis) and the model’s predicted absolute log fold change (background-subtracted) in this gene’s expression due to the rs8105903 variant (y-axis).

[0109] FIG. 53 provides a non-limiting example of the model’s attributions for the sequence surrounding the variant in the highlighted cell types (FIG. 29A) along with the matched motif from HOCOMOCO v12.

[0110] FIG. 54 provides a non-limiting example of a scatter plot showing the measured gene expression (log(CPM)+1) for the NFE2 gene across cell types (x-axis) and the model’s predicted absolute log fold change (background-subtracted) in this gene’s expression due to the rs79755767 variant (y-axis).

[0111] FIG. 55 provides a non-limiting example of the model’s attributions for the sequence surrounding the variant in the highlighted cell types (FIG.46) along with the matched motif from HOCOMOCO v12.

[0112] FIG. 56 provides a non-limiting schematic illustration of the use of a machine learning model configured to predict cell type-specific and / or cell state-specific gene expression levels to predict disease vs. healthy differential gene expression, in accordance with some embodiments of the present disclosure.

[0113] FIG.57 provides a nonlimiting example of a scatter plot showing the measured and predicted log fold changes for a test gene set between matched healthy and disease pseudobulk samples for fibroblast cells and a Crohn’s Disease pseudobulk sample.

[0114] FIG. 58 provides a non-limiting schematic illustration of the use of a machine learning model configured to predict cell type-specific and / or cell state-specific geneMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 19 of 126 expression levels to identify DNA sequence features that drive disease vs. healthy differential gene expression, in accordance with some embodiments of the present disclosure.

[0115] FIG. 59 provides a non-limiting example of a histogram of Pearson correlations between measured and predicted log fold changes between matched healthy and disease pseudobulk samples for the same cell type, calculated over a test set of genes.

[0116] FIG. 60A provides a nonlimiting example of a scatter plot showing the measured and predicted log fold changes for a test gene set between matched healthy and disease pseudobulk samples for fibroblast cells and a Crohn’s Disease pseudobulk sample (same plot as that shown in FIG.57).

[0117] FIG. 60B provides a nonlimiting example of a scatter plot showing the measured and predicted log fold changes for a test gene set between matched healthy and disease pseudobulk samples for fibroblast cells and a scleroderma pseudobulk sample, a chronic autoimmune disease.

[0118] FIG. 60C provides a nonlimiting example of a scatter plot showing the measured and predicted log fold changes for a test gene set between matched healthy and disease pseudobulk samples for enterocyte cells and a Crohn’s Disease pseudobulk sample.

[0119] FIG. 60D provides a nonlimiting example of a scatter plot showing the measured and predicted log fold changes for a test gene set between matched healthy and disease pseudobulk samples for Type II pneumocyte cells and a scleroderma pseudobulk sample.

[0120] FIG.60E provides a non-limiting example of a boxplot showing the predicted log fold changes for genes that are upregulated or downregulated (measured absolute log fold change >= 1) in fibroblasts in a Crohn’s Disease pseudobulk sample.

[0121] FIG.60F provides a non-limiting example of a boxplot showing the predicted log fold changes for genes that are upregulated or downregulated (measured absolute log fold change >= 1) in fibroblasts in a scleroderma pseudobulk sample.

[0122] FIG.60G provides a non-limiting example of a boxplot showing the predicted log fold changes for genes that are upregulated or downregulated (measured absolute log fold change >= 1) in enterocytes in a Crohn’s Disease pseudobulk sample.

[0123] FIG.60H provides a non-limiting example of a boxplot showing the predicted log fold changes for genes that are upregulated or downregulated (measured absolute log fold change >= 1) in Type II pneumocytes in a scleroderma pseudobulk sample.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 20 of 126

[0124] FIG. 60I provides a non-limiting example of motifs found by TF-MoDISco based on predicted differential attribution scores (disease vs. matched healthy pseudobulk samples) across multiple cell types and diseases.

[0125] FIG. 60J provides a non-limiting example of a dimeric motif identified by TF- MoDISco in attributions for fibroblasts in multiple fibrotic diseases.

[0126] FIG. 60K provides a non-limiting example of data for measured log fold change (disease vs. healthy) in TWIST1 expression in various diseases and cell types, separated by whether the TWIST1 dimeric motif is found by TF-MoDISco based on the disease vs. healthy differential attribution scores in that disease and cell type.

[0127] FIG. 60L provides a non-limiting example of an NF-kB motif identified by TF- MoDISco in attributions for fibroblasts in Crohn’s Disease.

[0128] FIG. 61 provides a non-limiting schematic illustration of the use of a machine learning model configured to predict cell type-specific and / or cell state-specific gene expression levels for designing promoter sequences through a process of directed evolution, in accordance with some embodiments of the present disclosure.

[0129] FIG.62 provides a non-limiting example of data for predicted expression of a cargo gene across healthy and diseased fibroblast and non-fibroblast cells over 100 rounds of directed evolution.

[0130] FIG.63 provides a non-limiting example of data for predicted specificity of cargo gene expression in fibroblasts and disease fibroblasts, which were improved in design in rounds 0-50 for cell-type specificity and in rounds 50-100 for disease-state specificity.

[0131] FIG.64 provides a non-limiting example of data for in silico mutagenesis (ISM) of the synthetic regulatory element that reveals key sequence motifs whose perturbation is predicted to uniquely affect expression of the cargo gene in fibroblasts.

[0132] FIG. 65 provides a non-limiting example of data for in silico mutagenesis (ISM) with respect to disease fibroblast expression that identifies key motifs generated in the design process for the fibroblast evolved sequences, including a TWIST1, C / EBP, and IRF motif, which are implicated in fibroblast-specific and immune specific regulation.

[0133] FIG. 66 provides a non-limiting example of data for in silico mutagenesis (ISM) with respect to disease fibroblast expression that identifies key motifs generated in the designMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 21 of 126 process for the disease-fibroblast evolved sequences, including a TWIST1, C / EBP, and IRF motif, which are implicated in fibroblast-specific and immune specific regulation.

[0134] FIG. 67 provides a non-limiting example of data illustrating that the identified motifs (FIG.65 – FIG.66) match HOCOMOCO v12 motifs.

[0135] FIG.68 provides a non-limiting schematic illustration of an approach for predicting regulatory element activity across tissue and cell types, in accordance with some embodiments of the present disclosure.

[0136] FIG. 69 provides a non-limiting schematic illustration of the use of long-context DNA sequence models to predict functional genomics properties.

[0137] FIG.70 provides a non-limiting schematic illustration of the use of a long-context DNA sequence model to process an input construct comprising, e.g., a candidate promoter sequence and cargo gene sequence, and score the efficacy of the promoter sequence for gene expression based on one or more prediction functional genomics property profiles.

[0138] FIG. 71 provides a non-limiting schematic illustration of a computational framework for using one or more long-context DNA models to process a plurality of construct sequences comprising, e.g., a plurality of candidate promoter sequences and a cargo gene sequence inserted at variety of different candidate genomic insertion sites, and score the combinations of candidate promoter sequences, cargo gene sequence, and candidate genomic insertion sites for efficacy based on their predicted gene expression levels and / or one or more predicted genomic property profiles.

[0139] FIG.72 provides a non-limiting example of predicted gene expression data used to score a designed promoter and cargo gene sequence inserted at a specified genomic integration site during a directed evolution process.

[0140] FIG. 73 provides a non-limiting example of data illustrating the distribution of designed promoter efficacy scores for an hSyn insertion into prescreened responsive safe- harbor sites versus insertion into other test sequence sites.

[0141] FIG.74 provides a non-limiting example of validation of an approach for prediction of regulatory element activity in any cell type against experimental massively parallel reporter assay (MPRA) data for HepG2 activity in hepatocyte expression.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 22 of 126

[0142] FIG.75 provides a non-limiting example of validation of an approach for prediction of regulatory element activity in any cell type against experimental massively parallel reporter assay (MPRA) data for K562 activity in blood expression.

[0143] FIG.76 provides a non-limiting example of data for predicted hSyn activity.

[0144] FIG. 77 provides a non-limiting example of data for hSyn expression predictions across brain cell types.

[0145] FIG. 78 provides a non-limiting example of data for predicted promoter activity across tissue types.

[0146] FIG.79 provides a non-limiting example of a block diagram of a computer system, in accordance with some embodiments of the present disclosure.

[0147] FIG.80 provides a non-limiting example of a block diagram of an example artificial intelligence (AI) architecture included as part of an example computing system, in accordance with some embodiments of the present disclosure.

[0148] In the appended figures, similar components and / or features can have the same reference label. Further, various components of the same type can be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label. DETAILED DESCRIPTION

[0149] Sequence-to-function models have become pivotal in understanding how the genome sequence is interpreted and executed in various cellular contexts. These models are essential for generating mechanistic hypotheses for non-coding variants and designing synthetic regulatory sequences. Although recent models have expanded their biological coverage to include various experiments and tissues, their reliance on bulk samples from healthy individuals and cell lines limits the investigation of regulatory circuits specific to cell types or diseases. Single-cell RNA sequencing (scRNA-seq), for which substantial datasets are already publicly available, can overcome these limitations by providing extensive coverage of diverse cell states and diseases, thereby enriching our understanding of complex regulatory rules.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 23 of 126

[0150] However, the field of single-cell genomics faces challenges in inferring regulatory activity in individual cell types and states when chromatin accessibility data is unavailable. This limitation also significantly restricts the generation of hypotheses regarding the regulatory elements driving differential expression between cell types and cellular states. The disclosed model (“Decima”) bridges this gap by applying the utility of sequence-to-function models to single-cell datasets, enabling regulatory interpretation in various cell states and diseases, sc- eQTL effect predictions, and the design of regulatory sequences with cell type- and disease- enriched activity. The examples presented herein demonstrate that the disclosed model can learn regulatory rules related to cell type specificity, tissue residency, and disease-dependent cell states by replicating known regulatory syntax from the literature.

[0151] Machine learning-based methods, systems, and programming for decoding gene expression are described. The disclosed methods and systems use one or more machine- learning models that collectively are configured to process an input DNA sequence (e.g., an input gene sequence and its flanking sequence regions) and output a prediction of cell type- and cell state-specific gene expression.

[0152] The disclosed machine learning-based methods and models are capable of predicting gene expression at the resolution of specific cell types or cell states, and achieve this capability through a combination of: (i) a model architecture that enables processing of long nucleic acid sequences at high resolution, and (ii) the use of a training data set comprising single-cell or single-nucleus RNA sequencing data from over 22 million cells. The trained model successfully predicts the cell type-specific expression of previously unseen gene sequences based on their DNA sequence alone. A significant advantage of the presently disclosed model over bulk-trained models is the possibility of linking variants to phenotypes via cell type- specific effects

[0153] In one aspect, for example, the disclosed methods for predicting gene expression levels may comprise receiving (e.g., at one or more processors of a system configured to perform the method) genomic sequence data comprising a DNA sequence for a specified gene; inputting the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type- specific and / or a cell-state-specific gene expression level; and outputting, using the one or moreMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 24 of 126 processors, a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene.

[0154] Also disclosed herein are machine learning-based methods for performing in silico design and / or optimization of cell type-specific and / or cell state-specific DNA regulatory sequences using long-context DNA sequence-based prediction models and a directed evolution process. The disclosed methods are based on iterative in silico mutagenesis of a starting sequence (e.g., a random sequence or candidate sequence) and scoring of regulatory sequence (e.g., promoter sequence) efficacy based on predicted gene expression level. In some instances, the disclosed machine learning-based methods and models may be configured to process one or more input DNA sequences (e.g., in the context of a construct comprising the regulatory sequence and a cargo gene sequence) using one or more long-context DNA sequence-based prediction models (e.g., one or more DNA sequence-prediction models configured to predict gene expression), while also evaluating one or more genomic insertion sites as part of the design and / or optimization of a cell type-specific and / or cell state-specific DNA regulatory sequence.

[0155] Thus, in another aspect, the disclosed methods for designing a DNA regulatory sequence may comprise, for example, receiving (e.g., at one or more processors of a system configured to perform the method) at least one candidate DNA regulatory sequence; receiving a cargo gene sequence; generating at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; inputting the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type- specific and / or a cell-state-specific gene expression level; determining an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type- specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and selecting, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence.

[0156] The methods and systems described herein provide numerous technical advantages. For example, due to their improved predictive performance, the disclosed machine learning- based methods and models can be used to identify the cis-regulatory mechanisms driving cellMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 25 of 126 type-specific gene expression and their changes in disease, to predict the effect of non-coding variants at cell type resolution, and to design regulatory DNA elements with precisely tuned, context-specific functions. These capabilities can, in turn, improve the efficacy of downstream development and manufacturing processes (e.g., by facilitating the design and development of effective gene therapies, CRISPR-based drug delivery mechanisms, and protein expression systems for biologics manufacturing, etc.).

[0157] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the disclosed embodiments, and is provided in the context of a patent application and its requirements. I. Sequence to Function Deep Learning Models

[0158] As noted above, sequence-to-function deep learning models that predict gene expression from genomic DNA sequences have proven valuable for many biological tasks, including understanding cis-regulatory mechanisms and interpreting non-coding genetic variants. Recently, modeling of genomic DNA sequences of up to hundreds of kilobases in length in a genomics context has enabled effective prediction of, e.g., bulk cap analysis of gene expression (CAGE-seq) and RNA-seq profiles. However, current state-of-the-art models have been trained largely on bulk expression profiles from healthy tissues or cell lines. Thus, they cannot predict expression nor reveal regulatory mechanisms that act in specific cell types. Moreover, due to training primarily on samples from healthy donors, they cannot directly model pathological expression changes occurring in the context of disease. Therefore they fail to capture and learn the extensive biological information embedded in the rapidly accumulating single-cell RNA sequencing (scRNA-seq) or single-nucleus RNA sequencing (snRNA-seq) datasets.

[0159] Studying regulatory mechanisms using scRNA-seq data has been challenging but holds great potential. Current approaches typically rely on integrating single-cell ATAC-seq (scATAC) data, or analyzing the expression of transcription factors (TFs) or their known target genes, or analyzing the co-expression of TFs and targets. While effective, these approaches are limited by the availability of matched scATAC-seq or prior knowledge of TF-target relationships. However, the genomic sequence itself offers a rich and largely untapped data modality for uncovering regulatory mechanisms. By focusing on the genomic sequence, oneMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 26 of 126 can move beyond the constraints of known TF-target pairs and explore the regulatory landscape of biological settings for which chromatin accessibility data are unavailable, opening the door to a more comprehensive understanding of gene regulation across diverse cell types and conditions.

[0160] As a result, there is a need for sequence-based prediction models that can elucidate complex regulatory mechanisms at play within cell populations across different cell state and disease conditions. This gap has been widely recognized, leading to recent attempts to predict gene expression from sequence data in single cells. However, this approach is difficult to scale to atlas-level single-cell datasets that frequently include millions of cells, and is consequently limited to analyses of single datasets comprising a small number of cell types or states.

[0161] Disclosed herein is a machine learning model configured to predict the cell type- and condition-specific expression of a gene sequence from its surrounding genomic sequence. The model is trained on scRNA-seq and snRNA-seq count data aggregated by cell type and state to form pseudobulk samples or profiles, thereby allowing it to learn from a massive training corpus comprising data from over 22 million cells and representing 201 distinct cell types in 271 tissues and 82 disease states. The model is capable of accurately predicting pseudobulk gene expression, including expression of highly cell type-specific genes, under a variety of conditions. Additionally, the model is capable of identifying potential regulatory elements driving cell type- and cell state-specific gene expression, as well as tissue residency and disease-associated changes in gene expression. The model can be used to predict the effects of non-coding genetic variants in individual cell types, and can also be used to facilitate the design of disease-biased regulatory elements. II. Example Prediction System

[0162] FIG.1 is a block diagram of an example prediction system, in accordance with some embodiments. Prediction system 100 can be used to determine a predicted DNA property, such as a predicted expression level, a predicted functional property, and / or predicted impact of a sequence variant on gene expression for a given DNA sequence, e.g., a gene sequence. The prediction system 100 can include, e.g., computing platform 102, data store 104, and display system 106. Computing platform 102 may take any of a variety of forms. In some embodiments, for example, the computing platform 102 can include a single computer (orMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 27 of 126 computer system), or multiple computers in communication with each other. In some embodiments, the computing platform 102 can be a cloud computing platform.

[0163] Data store 104 and display system 106 are each in communication with computing platform 102. In some examples, one or more of data store 104 and / or display system 106 can be considered part of, or otherwise integrated with, computing platform 102. Thus, in some examples, computing platform 102, data store 104, and display system 106 can be separate components in communication with each other, but in other examples, some combination of these components can be integrated together. Communication between the different components can be implemented using any number of wired communications links, wireless communications links, optical communications links, or a combination thereof.

[0164] The prediction system 100 can include a sequence analyzer 108, which can be implemented using hardware, software, firmware, or a combination thereof. In some embodiments, the sequence analyzer 108 is implemented as part of computing platform 102. The sequence analyzer 108 receives sequence data 110 for processing. For example, sequence data 110 can be provided as input into the sequence analyzer 108, can be retrieved from data store 104 or some other type of storage (e.g., cloud storage), can be accessed from cloud storage, or can be obtained in some other manner. In some cases, sequence data 110 can be retrieved from data store 104 in response to receiving user input entered by a user via an input device.

[0165] In some embodiments, the sequence data 110 can be generated by processing a set of samples 112. The set of samples 112 may take the form of one or more biological samples from one or more subjects (e.g., a diseased sample, a healthy sample, or a combination thereof). In some embodiments, the set of samples 112 may include a sample obtained from a tumor of a subject. The tumor can be a manifestation of, for example, lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myelogenous leukemia, chronic myelogenous leukemia, chronic lymphocytic leukemia, T cell lymphocytic leukemia, non-small cell lung cancer, small-cell lung cancer, or a combination thereof. In some embodiments, the set of samples 112 may include a sample obtained from a subject diagnosed with a disease.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 28 of 126

[0166] In some embodiments, a sample in the set of samples 112 may include, for example, various nucleic acid molecules, e.g., various DNA molecules, various RNA molecules (e.g., various mRNA molecules, various rRNA molecules, or various tRNA molecules), or a combination thereof. When the set of samples 112 includes a diseased sample, the nucleic acid molecules may include one or more mutant nucleic acid sequences (e.g., mutant gene sequences, mutant regulatory sequences, mutant promoter sequences, mutant mRNA sequences, etc.).

[0167] In some embodiments, the set of samples 112 may comprise biological samples that are processed to extract nucleic acid molecules and generate sequence data 110. In some embodiments, multiple samples in the set of samples 112 can be processed at the same or at different times. In some embodiments, the prediction system 100 includes a sample analyzer that is used in processing the set of samples 112 to generate the sequence data 110.

[0168] In some embodiments, a sample in the set of samples 112 may include synthetic nucleic acid molecules, e.g., synthetic DNA molecules, synthetic RNA molecules (e.g., synthetic mRNA molecules, synthetic rRNA molecules, or synthetic tRNA molecules), or a combination thereof.

[0169] In some embodiments, a sample in the set of samples 112 may include designed or in silico nucleic acid sequences, e.g., designed or in silico DNA sequences, designed or in silico RNA sequences (e.g., designed or in silico mRNA molecules, designed or in silico rRNA molecules, or designed or in silico tRNA molecules), or a combination thereof.

[0170] In some embodiments, a sample in the set of samples 112 may include, e.g., an extracted, synthetic, designed, or in silico DNA sequence 129. In some embodiments, the DNA sequence 129 may correspond to a gene or a sequence that comprises a 5’ end subsequence 120, a central subsequence 118, and / or a 3’ end subsequence 122. In some embodiments, the 5’ end subsequence 120 may correspond to a 5’ (upstream or start) sequence 128. In some embodiments, the central subsequence 118 may correspond to a gene coding sequence (CDS) 126. In some embodiments, the 3’ end subsequence may correspond to a 3’ (downstream or stop) sequence 130. One or more sub-sequences of, e.g., DNA sequence 129 (e.g., 5’ sequence 128, coding sequence 126, or 3’ sequence 130) can be processed separately or as a single sequence.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 29 of 126

[0171] In some embodiments, DNA sequence 129 encodes for at least a portion of a corresponding protein. In some embodiments, DNA sequences (e.g., gene sequences), mRNA sequences, and / or protein sequences, or mutants thereof, can be identified by comparison to a corresponding reference sequence and / or by performing a search in an appropriate database (e.g., the RefSeq NCBI Reference Sequence Database, etc.) based on the DNA sequence, mRNA sequence, and / or protein sequence data obtained from the sample.

[0172] Sequence analyzer 108 receives the sequence data 110 as input for processing. The sequence analyzer 108 can include one or more machine-learning model(s) 132 that process the sequence data 110. In some embodiments, the sequence data 110 is sent directly into the one or more machine-learning model(s) 132 for processing. In some embodiments, the sequence analyzer 108 preprocesses the sequence data 110 prior to sending the sequence data 110 to the one or more machine-learning model(s) 132 for processing. Pre-processing the sequence data 110 may include, e.g., one-hot encoding of the sequence data and addition of a gene mask.

[0173] The one or more machine-learning model(s) 132 can be implemented in any of a number of different ways. In some embodiments, a machine learning model of the one or more machine-learning model(s) 132 can be any type of model that uses a set of element-focused scores that represent properties of a set of nucleic acid sequence representations. The one or more machine-learning model(s) 132 can be trained a training mode or deployed in a prediction mode. In the training mode, the one or more machine-learning model(s) 132 are trained using training data 133. Examples of the training data 133 are described in more detail below. The machine-learning model 132 is trained such that it can be deployed and used in the prediction mode.

[0174] The one or more machine-learning model(s) 132 process the sequence data 110, e.g., DNA sequence 129, via, a sequence processing engine, e.g., a DNA processing engine 139. In some embodiments the sequence processing engine may comprise one or more modules, e.g., a CDS processing engine 136, a 5’ sequence processing engine 138, and / or a 3’ sequence processing engine 140. In some embodiments, the separate processing engines for the coding sequence, the 5’ sequence, and / or the 3’ sequence enable improved predictive performance. Non-limiting examples of implementations for these different processing engines are described in greater detail below.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 30 of 126

[0175] As used herein, the terms “processing engine” and “engine” identify at least one software component and / or a combination of at least one software component and at least one hardware component which are designed / programmed / configured to interact and / or communicate data to other software and / or hardware components including but not limited to other processing engines.

[0176] The one or more machine-learning model(s) 132 process the sequence data 110 to generate an output that can be used, e.g., to generate a report 144. The report 144 may include the exact output of the one or more machine-learning model(s) 132, a transformed or filtered version of the output, or both. In some cases, the report 144 may include notifications, recommendations, alerts, or other information generated by the sequence analyzer 108 based on the output of the one or more machine-learning model(s) 132.

[0177] The report 144 can be an output that includes, for example, predicted nucleic acid sequence properties, e.g., a predicted DNA property, such as a predicted gene expression level, a predicted gene functional property, and / or a predicted impact of a sequence variant on gene expression for a given DNA sequence, e.g., a gene sequence, with respect to one or more input sequences.

[0178] In some embodiments, a report 144 can be displayed on a graphical user interface (GUI) 150 on the display system 106. A user may view and / or interact with the report 144 via the graphical user interface 150. In some embodiments, the user may use the report 144 to make decisions about, e.g., the development of a gene-based therapy, the development of a CRISPR- based delivery system, or the treatment of a subject from which at least one of the set of samples 112 was obtained (or collected).

[0179] In some embodiments, the prediction system 100 sends the report 144 to the remote system 152 (e.g., wirelessly). The remote system 152 can be a cloud computing platform, cloud storage, another computer system, a user device (e.g., a smartphone, a tablet, a laptop, etc.), or some other type of platform. In some embodiments, the remote system 152 can be a treatment manufacturing system (or machine) or a portion thereof. III. Example Model Architecture

[0180] The machine-learning model 132 of the embodiments described herein may include multiple subsystems (or subnetworks). Each of the multiple subsystems can include, for example, an encoder, a transformer encoder, and / or one or more processing layers.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 31 of 126

[0181] Machine-learning model 132 may include one or more encoders configured to, for example, transform an element of the input sequences (e.g., an amino acid sequence, a nucleic acid sequence, a codon sequence, etc.) based on the other elements of the input sequences. An encoder can be a transformer encoder.

[0182] The machine-learning model 132 may include one or more processing layers such as self-attention layers or convolution layers, or a neural network such as a long-short term memory unit (LSTM), recurrent structure, or recurrent component. The machine-learning model 132 can implement, for example, one or more self-attention layers. Machine-learning model 132 can use a self-attention mechanism, global attention mechanism, soft attention mechanism, local attention mechanism, and / or hard attention mechanism. In some instances, the machine-learning model 132 does not include any convolutional layer, any recurrent structure, any LSTM unit, and / or any recurrent component. In some instances, the machine learning model 132 is not a recurrent machine-learning model and / or does not include a recurrent neural network. In some instances, the machine-learning model 132 includes a recurrent neural network and / or may use positional encoding to handle sliding windows of sequence elements across one or more sequences. In some instances, the machine-learning model 132 is not a convolutional machine-learning model and / or does not include a convolutional neural network.

[0183] The machine-learning model 132 may include processing blocks, such one or more first processing blocks used to process, e.g., one or more nucleic acid sequence representations independent from a second processing block used to process, e.g., a representation of a functional genomics assay data profile. In some embodiments, the second processing block may process part or all of, e.g., a functional genomics assay data profile. The independence of these processing blocks can facilitate parallel processing when using the machine-learning model 132. Further, the independence may improve the performance (e.g., accuracy of predictions) of the machine learning model 132.

[0184] The machine-learning model 132 can be configured such that an output value at any given layer depends not only on a corresponding input value but also on one or more (e.g., all) other input values. Thus, the machine-learning model 132, a loss function, and / or an optimization function can be configured to improve an output corresponding to a single position representing, e.g., a level of gene expression for a given input DNA sequence for aMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 32 of 126 specified gene. In some instances, the loss function can comprise supervised loss function such as binary cross entropy or unsupervised loss functions. In some instances, the unsupervised loss functions can include a contrastive loss or regularization losses (e.g., L1 / L2 losses). In some instances, the loss function can comprise auxiliary loss functions. In such instances, the auxiliary loss function can be used alongside a main loss function to train the machine learning model 132. Accordingly, in some instances, the auxiliary loss function can improve the learning process by adding additional information or constraints. In some instances, any of a plurality of outputs of transformer encoders may represent such an occurrence probability. The machine-learning model 132 can be trained accordingly. In some instances, an endpoint (e.g., surplus endpoint) may represent (in response to training) a gene expression level and / or a functional genomic property probability or likelihood. Aggregated outputs can be, for example, fed to another layer, subsystem, or processing block (e.g., that includes one or more of: a processing layer such as a self-attention layer, or an encoder such as a transformer encoder).

[0185] In some instances, one, two, or all dimensions of an output from another layer and / or another subsystem or processing block is the same size as the input fed to the other layer and / or other subsystem or processing block. In some instances, an input fed to this other layer and / or other subsystem or processing block has a length along one axis that is greater than or equal to a sum of one or more of, for example, a number of nucleotides in a gene sequence, a number of nucleotides in an 5’ end flanking sequence, or a number of nucleotides in a 3’ end flanking sequence. In some instances, the length of the input is one longer than the total number of nucleotides. The length of the input along the one axis may exceed the summed count of nucleotides when, for example, an additional feature vector is appended to the nucleotide- specific feature values. Another dimension of the input can include a number of features (e.g., determined via a hyperparameter). An output generated by the other layer and / or other subsystem or processing block may have the same size as the input.

[0186] A subset of values of the output generated by the other layer and / or other subnetwork can be processed by another neural network (e.g., a fully connected feedforward network). The subset of values may include a 1-dimensional vector of values that may correspond to one set of feature values.

[0187] In some embodiments, a neural network within the machine-learning model 132 can be configured to output one or more results. The one or more results can include, for example,MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 33 of 126 a numeric result, binary result, and / or categorical result. Each of the one or more results can predict whether and / or an extent to which a nucleic acid sequence (e.g., a gene sequence) is likely to be expressed in a specified cell type and / or a specified cell state. The machine-learning model 132 may include one or more activation layers to produce an intermediate result (e.g., to transform a real-number interim value into a binary and / or categorical output). The machine- learning model 132 can be trained to generate multiple types of predictions (e.g., gene expression predictions and / or functional genomic property predictions). In some instances, a prediction can be binary or categorical. Other predictions can be non-binary or non-categorical. For example, a prediction can be scalar.

[0188] Machine-learning model 132 may include and / or can be included within an ensemble model. The ensemble model may include multiple (e.g., identical) sub-models that can be trained using different portions of the training data set.

[0189] FIG.2 provides a non-limiting, schematic illustration of a machine learning model architecture, in accordance with some embodiments of the present disclosure. The machine learning model architecture depicted in FIG. 2 can be, in some examples, based on a foundational model (e.g., the Borzoi model for prediction of RNA-seq coverage (Linder et al. (2023), “Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation”, bioRxiv preprint, https: / / doi.org / 10.1101 / 2023.08.30.555582) or the Enformer model (Avsec et al. (2021), “Effective gene expression prediction from sequence by integrating long-range interactions”, Nature Methods 18(10):1196–1203). The Borzoi model is based on a neural network architecture comprising a stack of convolution and downsampling layers, followed by a series of self-attention layers with relative positional encodings operating at 128 bp resolution. Repeated application of the convolution block achieves a 2-fold reduction of the sequence length and extracts local sequence patterns until each position in the sequence represents 128 bp. Then, repeated application of the self-attention (or transformer) block enables long-range interaction and exchange between every pair of sequence positions. The output is then up-sampled through a number of deconvolutional layers with matched U-net connections.

[0190] For the model disclosed herein, the model input is configured differently from that used in prior art models (e.g., the Borzoi model). Specifically, in addition to an input DNA sequence, the model input also comprises a binary gene mask to indicate the transcription startMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 34 of 126 and stop sites for the gene, as described elsewhere herein. Accordingly, an additional input channel was added to the first convolutional layer of the model.

[0191] For the model disclosed herein, the final layer of, e.g., the Borzoi architecture, was also replaced with a global mean pooling layer along the length axis (to average gene expression predictions across entire input DNA sequence for a given cell type), followed by a linear ‘head’ layer that outputs a vector of 8,856 values for each gene, corresponding to the predicted log(CPM+1) expression values in each cell type pseudobulk sample. Unlike prior art models, e.g., Borzoi, which output predictions for specified genomic positions comprising binned loci (e.g., 32 nucleotides per bin, of 16 nucleotides per bin), the model disclosed herein can be configured to output predictions at single-nucleotide resolution. IV. Example Model Training Process

[0192] The machine learning model 132 can be, in some instances, trained on atlas-scale single-cell expression data. In some instances, the machine learning model 132 may be based on the Borzoi model (Linder et al. (2023)), and may be initially trained to predict genome-wide sequencing coverage in bulk genomic datasets. Unlike existing models, machine learning model 132 can then fine-tuned on pseudo-bulked scRNA-seq data derived from over 22 million cells, thereby enabling predictions for specific cell types and conditions. For example, for a given 500 kb input sequence surrounding and comprising the DNA sequence for a gene not previously seen by the model, the model is able to predict the expression level of that gene in hundreds of different cell types across diverse cellular states, tissues, and disease conditions.

[0193] FIG.3 provides a non-limiting example of a flowchart for a process 300 for training a machine learning model (e.g., the “Decima” model) to predict cell type-specific and / or cell state-specific gene expression levels based on an input DNA sequence, in accordance with some embodiments of the present disclosure. Process 300 can be performed using the prediction system 100 in FIG.1 to train machine-learning model 132. In some instances, part or all of process 300 can be performed at a remote computing system that is remote relative to a user device and / or laboratory. The remote computing system can be a cloud computing system.

[0194] At step 302 in FIG.3, training data comprising gene sequences and corresponding aggregated scRNA-seq data is input to the model.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 35 of 126

[0195] In some instances, scRNA-seq data for use in training data may be retrieved from any of a variety of public and / or private sources (e.g., databases, technical publications, etc.) and compiled. Examples of scRNA-seq data sources include, but are not limited to, SCimilarity (Heimberg et al. (2023) “Scalable Querying of Human Cell Atlases via a Foundational Model Reveals Commonalities across Fibrosis-Associated Macrophages.” bioRxiv, https: / / doi.org / 10.1101 / 2023.07.18.549537), the Brain Cell Atlas, the Retinal Cell Atlas (Li, et al. (2023) “Integrated Multi-Omics Single Cell Atlas of the Human Retina”. bioRxiv, https: / / doi.org / 10.1101 / 2023.11.07.566105), the Skin Cell Atlas (Fiskin et al. (2023) “Multi- Modal Skin Atlas Identifies a Multicellular Immune-Stromal Community Associated with Altered Cornification and Specific T Cell Expansion in Atopic Dermatitis.” bioRxiv. https: / / doi.org / 10.1101 / 2023.10.29.563503.), the Broad Institute’s Single Cell Portal, and the EMBL’s Single Cell Expression Atlas, etc.

[0196] In some instances, the training data may comprise gene sequences and corresponding aggregated scRNA-seq data for at least 106cell types, cell states (e.g., metabolic states, disease states, etc.), and / or studies. In some instances, the training data may comprise gene sequences and corresponding aggregated scRNA-seq data for at least 2.5 x 106, at least 5 x 106, at least 7.5 x 106, at least 107, at least 2.5 x 107, at least 5 x 107, at least 7.5 x 107, or at least 108cell types, cell states (e.g., metabolic states, disease states, etc.), and / or studies.

[0197] In some instances, the scRNA-seq counts may be summed for all cells belonging to the same combination of cell type, tissue, disease state, and / or study. In some instances, the scRNA data may be filtered, e.g., to remove data for specified cell types, to remove data for unannotated cell types, and / or to remove data for specified disease states, etc.

[0198] In some instances, the scRNA-seq data (i.e., gene expression data) may be normalized and transformed using a log(1+x) transformation.

[0199] In some instances, the resulting normalized, log transformed gene expression data may be aggregated (e.g., by pseudobulk sample) to create a pseudobulk gene expression matrix, as described elsewhere herein and illustrated in FIG.6.

[0200] In some instances, the input DNA sequence data (i.e., comprising sequences from the pseudobulk gene expression matrix) is one-hot encoded. In some instances, the input DNA sequence data can further comprise a binary gene mask that indicates the location of the gene sequence (e.g., the transcription start and transcription stop sites) within the input genomicMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 36 of 126 DNA sequence. In some instances, the one-hot encoded, masked input DNA sequence is paired with a corresponding gene expression value from the pseudobulk gene expression matrix.

[0201] In some instances, the pseudobulk gene expression data may be split into training, validation, and test datasets (e.g., using an 80%, 10%, 10% split, respectively, of the initial pseudobulk gene expression dataset, or any other split ratio known to those of skill in the art).

[0202] At step 304 in FIG. 3, iterative training of the model is performed, where the iterative training process includes steps of: (i) evaluating a two-component loss function (i.e., Poisson loss + multinomial negative log-likelihood function) based on the measured and a current prediction for gene expression data across pseudobulk samples, (ii) updating the model weights and biases, and (iii) evaluating a termination condition.

[0203] Prior to beginning the iterative training process, model hyperparameters (i.e., settings such as the learning rate in neural networks, the number of hidden layers, or regularization parameters, that are set to control the model's structure, behavior, and learning process) and / or model weights and / or biases (i.e., learned parameters) may be initialized. In some instances, model weights and / or biases can be initialized from scratch, e.g., using a random set of values. In some instance, model weights and / or biases can be initialized using the weights and / or biases from a trained foundational model (e.g., a trained Borzoi model). In some instances, all or a portion of the model (e.g., all or a portion of the layers in the model) may be trained and / or fine-tuned, (e.g., by training the model as a whole on a first set of training data (e.g., by training a Borzoi foundation model on functional genomics property data), and then fine-tuning specified layers using a second set of training data (e.g., by fine-tuning on single cell gene expression data)). In some instances, model hyperparameters may be improved by repeating several rounds of the iterative model training process, where the hyperparameters are adjusted in one or more rounds.

[0204] In some instances (e.g., as illustrated in FIG.6 below), computing the multinomial negative log-likelihood component of the two-component loss function during iterative model training can be based on the distributions of measured and predicted gene expression across pseudobulk samples.

[0205] In some instances (e.g., as illustrated in FIG. 6 below), computing the Poisson component of the two-component loss function during iterative model training can be based on multiplying the distributions of measured and predicted gene expression values acrossMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 37 of 126 pseudobulk samples by the corresponding sums of the distributions of measured and predicted gene expression across pseudobulk samples.

[0206] In some instances, (e.g., as illustrated in FIG. 6 below), computing the two- component loss function during iterative model training can comprise multiplying the Poisson component by the corresponding weight, and adding the value of the multinomial negative log- likelihood.

[0207] The iterative model training process can be repeated until, e.g., a specified termination condition is met. Examples of suitable termination conditions include, but are not limited to, an early stopping criterion (e.g., where the model’s performance on a validation data set starts to degrade or ceases to improve significantly for a predefined number of consecutive iterations), reaching a fixed number of iterations, reaching a point where improvement in the loss function becomes insignificant between iterations, or the change in the weight vector (comprising values of each learned weight) between iterations falls below a predefined threshold.

[0208] At step 306 in FIG. 3, the improved values of the model weights and biases are output. In some instances, the model’s state (i.e., the set of weights and biases) at the point of improved performance on a validation data set (i.e., before decline or stagnation of performance sets in) can be restored and those values used as the final configuration of the model.

[0209] In one non-limiting example, as noted above and described in more detail in Example 1 below, the model can be initialized from a foundational model, e.g., Borzoi, and then fine-tuned. Single cell RNA-seq (scRNA-seq) datasets were combined and processed by summing counts from all cells belonging to the same combination of cell type, tissue, disease state, and study. The counts were then normalized and a log(1+x) transformation was applied. For each gene, an input window containing 524,288 bp of genomic sequence surrounding the gene was created. In addition to the one-hot encoded DNA sequence in the input window, the model was supplied with a binary mask (i.e., a gene mask) representing the location of the gene sequence within the input genomic sequence. The model was trained using a two-component loss function comprising a Poisson loss on the total expression of the gene across all pseudobulks, and a multinomial loss along the pseudobulk axis. Novel components of this training procedure (compared to, e.g., training of the Borzoi model) include the use of a binaryMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 38 of 126 gene mask to indicate the start and stop positions of the gene sequence, and the use of the aforementioned two-component loss function.

[0210] The weighting factors, bias values, and threshold values, or other computational parameters of the neural network (or other machine learning architecture), can be "taught" or "learned" in a training phase using one or more sets of training data (e.g., 1, 2, 3, 4, 5, or more than 5 sets of training data) and a specified training approach configured to solve, e.g., minimize, a loss function. A loss function can use an error term (e.g., mean squared error or median squared error) and / or an entropy term (e.g., cross entropy or binary cross entropy). The adjustable parameters for the neural network (e.g., deep learning model) may be determined based on input data from a training data set using an iterative solver (such as a gradient-based method, e.g., backpropagation), so that the output value(s) that the neural network computes (e.g., a prediction of gene expression level in a given cell type) are consistent with the examples included in the training data set. In some instances, multitask learning can be used, such that the model is simultaneously trained to predict each of two different types of results (e.g., gene expression level and mRNA stability). A static or non-static learning rate can be used. For example, learning rate annealing (e.g., using stepwise annealing or cosine annealing) can be used to reduce the learning rate over iterations. Validation-data assessment can be used to potentially terminate training early (e.g., upon determining that a performance target has been met).

[0211] In some instances, the training data set (e.g., paired sets of genomic sequence data and corresponding functional genomics properties) can be randomly parsed, shuffled, and / or divided to train various models within an ensemble of models.

[0212] In some instances, the disclosed models and methods may comprise retraining (or “fine-tuning”) a previously trained machine learning models (e.g., by iteratively retraining a previously trained model using one or more training data sets that differ from those used to train the model initially). In some instances, retraining the machine learning model may comprise using a continuous, e.g., online, machine learning model, i.e., where the model is periodically or continuously updated or retrained based on new training data. The new training data may be provided by, e.g., a single deployed local operational system, a plurality of deployed local operational systems, or a plurality of deployed, geographically-distributed operational systems. In some instances, the disclosed methods may employ, for example, pre-MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 39 of 126 trained neural networks, and the pre-trained neural networks can be fine-tuned using additional dataset(s) to tune the pre-trained neural network for the same, or a different, prediction objective.

[0213] The training of the model (i.e., determination of the adjustable parameters of the model using an iterative solver) may or may not be performed using the same hardware as that used for deployment of the trained model. V. Cell Type- and Cell State-Specific Prediction of Gene Expression Levels

[0214] FIG.4 provides a non-limiting example of a flowchart for a machine learning-based process 400 for predicting a cell type-specific and / or a cell state-specific gene expression level for a specified gene, in accordance with embodiments of the present disclosure. Process 400 can be implemented using, for example, the prediction system 100 described in FIG. 1. For example, process 400 can be implemented using the sequence analyzer 108 and the machine- learning model(s) 132 shown in FIG.1.

[0215] At step 402 in FIG. 4, genomic sequence data comprising a DNA sequence for a specified gene is received (e.g., by one or more processors of a system configured to perform process 400). Examples of the processing of DNA sequences to generate model inputs is described in more detail in Example 2 below (see, e.g., the method section on Creating Inputs for the Model).

[0216] In some instances, the genomic sequence data can comprise, e.g., the DNA sequence for the specified gene and a flanking sequence for at least one of a 3’ end or a 5’ end of the specified gene sequence (e.g., any human gene, any non-human mammalian gene, etc.).

[0217] In some instances, the genomic sequence data can comprise a DNA sequence ranging in length from about 300,000 nucleotides (or base pairs (bp)) to about 700,000 nucleotides (or base pairs (bp)) (with or without including the flanking sequence for at least one of the 3’ end or the 5’ end of the specified gene sequence). For example, in some instances, the genomic sequence data can comprise a DNA sequence having a length of at least 300,000, at least 350,000, at least 400,000, at least 450,000, at least 500,000, at least 550,000, at least 600,000, at least 650,000, or at least 700,000 nucleotides. In some instances, the genomic sequence data can comprise a DNA sequence having a length of at most 700,000, at most 650,000, at most 600,000, at most 550,000, at most 500,000, at most 450,000, at most 400,000, at most 350,000, or at most 300,000 nucleotides. Any of the lower and upper values describedMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 40 of 126 in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the genomic sequence data can comprise a DNA sequence having a length ranging from about 350,000 to about 550,000 nucleotides. Those of skill in the art will recognize that the genomic sequence data can comprise a DNA sequence having a length of any value within this range, e.g., about 564,000 nucleotides. In some instances, the genomic sequence data can comprise a DNA sequence of about 500,000 nucleotides in length.

[0218] In some instances, the genomic sequence data can be provided as input to the machine learning model as a one-hot encoded matrix. One-hot encoding is a technique used in machine learning to convert categorical data (e.g., colors, or the bases in a DNA sequence) into a binary vector format that can be processed by a machine learning algorithms. Each category is represented as a binary vector where one bit has a value of “1” (hot) and the remaining bits have a value of “0” (cold). In some instances, additional layers of information related to the DNA sequence can be added to the one-hot encoding including, but not limited to, a gene mask, positions of known regulatory elements, conservation scores, etc.

[0219] At step 404 in FIG. 4, the genomic sequence data is input to a trained machine learning model, where the trained machine learning model is: (i) trained on single cell gene expression data for a plurality of cell types and cell states, and (ii) configured to process genomic sequence data and predict a cell type-specific and / or a cell state-specific gene expression level. Examples of training a model on single cell gene expression data for a plurality of cell types and cell states is described in more detail in Examples 1 and 2 below (see, e.g., the method sections on Aggregation and Filtering of Pseudobulk Samples, Gene Filtering and Annotation, Normalization, Model Architecture, Training the Model, and Classification of Cell Type-Specific Genes in Example 2).

[0220] As noted elsewhere herein, in some instance, the trained machine learning model can comprise a deep neural network architecture. In some instances, the deep neural network architecture can comprise at least one convolution layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 convolutional layers).

[0221] In some instances, the deep neural network architecture can comprise at least one downsampling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 downsampling layers).

[0222] In some instances, the deep neural network architecture can comprise at least one self-attention layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 self-attention layers).MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 41 of 126

[0223] In some instances, the deep neural network architecture can comprise at least one deconvolution layer with matched U-net connections (at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 deconvolution layers with matched U-net connections).

[0224] In some instances, the deep neural network architecture may not include a global mean pooling layer. In some instances, the deep neural network architecture can comprise a single global mean pooling layer. In some instances, the deep neural network architecture can comprise at least one global mean pooling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 global mean pooling layers).

[0225] In some instances, the deep neural network architecture can comprise a linear output layer.

[0226] In some instances, the trained machine learning model is pre-trained on bulk gene expression profile data and / or epigenomic assay data. In some instances the bulk gene expression profile data may comprise, e.g., RNA-seq data and / or CAGE-seq data. In some instances, the epigenomic assay data may comprise, e.g., DNase-seq data, ATAC-seq data, and / or ChIP-seq data. In some instances, the single cell gene expression data can comprise single cell RNA sequencing data.

[0227] In some instances, the plurality of cell types and cell states can comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, or at least 200 different types of cell.

[0228] In some instances, the plurality of cell types and cell states can comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, or at least 150 different cell states. In some instances, for example, the different cell states can comprise different development states, different metabolic states, different disease states, or any combination thereof.

[0229] At step 406 in FIG.4, a prediction of a cell type-specific and / or a cell state-specific gene expression level for the specified gene is output (e.g., by the one or more processors of a system configured to perform process 200). Examples of the use of the model to predict cell type-specific and / or a cell state-specific gene expression levels is described in more detail in Example 2 below (see, e.g., the results section on Prediction of Pseudobulk gene Expression from DNA Sequence).MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 42 of 126

[0230] In some instances, the output can comprise a prediction of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 25, at least 50, at least 75, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000 cell-type-specific and / or cell-state-specific gene expression levels for the specified gene.

[0231] In some instances, process 400 may further comprise using the trained machine learning model to predict an impact of a gene sequence variant on gene expression level.

[0232] In some instances, process 400 may further comprise using the trained machine learning model to predict gene expression patterns for a plurality of specified genes.

[0233] In some instances, process 400 may further comprise using the trained machine learning model to predict differences in gene expression patterns for a plurality of specified genes in healthy versus diseased tissues.

[0234] In some instances, process 400 may further comprise: (i) using the trained machine learning model to identify and compare gene sequences that are highly expressed in a given cell type, and (ii) clustering sequence features in the highly expressed gene sequences to identify transcription factor binding motifs.

[0235] In some instances, process 400 may further comprise using the trained machine learning model to design sequences for gene therapy, wherein the designed sequences are predicted to be actively expressed only in specific disease states.

[0236] In some instances, the output of process 400 may comprise a report based on the prediction of a cell type-specific and / or a cell state-specific gene expression level(s) for the specified gene.

[0237] In some instances, the disclosed machine learning model may be configured to predict mRNA properties (e.g., mRNA stability and / or mRNA translation efficiency) from genomic sequence data for specific cell types and / or cell states. For example, disclosed herein are methods for or predicting mRNA properties from genomic sequence data comprising: receiving (e.g., at one or more processors of a system configured to perform the method) genomic sequence data comprising a DNA sequence for a specified gene; inputting (e.g., using the one or more processors) the genomic sequence data to a machine learning model, wherein the machine learning model is configured to process genomic sequence data and predict a cell-MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 43 of 126 type-specific and / or a cell-state-specific property of an mRNA molecule corresponding to a specified gene; and outputting (e.g., using the one or more processors) a prediction of a cell- type-specific and / or a cell-state-specific property of an mRNA molecule corresponding to the specified gene, wherein the predicted cell-type-specific and / or a cell-state-specific property comprises at least one of: (i) a predicted mRNA stability, or (ii) a predicted mRNA translation efficiency. In some instances, the method further comprises using the predicted mRNA stability and / or predicted mRNA translation efficiency to predict an expression level for a protein corresponding to the specified gene.

[0238] In order to configure the disclosed model to predict, e.g., mRNA translation efficiency, the model was trained on training data comprising one-hot encoded genomic DNA sequence data (with an associated binary gene mask) and corresponding translation efficiency (TE) data for a plurality of genes and sample types. In one instance, for example, the model was trained on expression data for 12,700 genes, with a validation data set comprising data for another 1,700 genes, and a test data set comprising data for another 1,500 genes. The trained model converged on a predicted translation efficiency for a given input genomic DNA sequence in about 2 epochs, with a Pearson correlation coefficient of 0.654 for a comparison of predicted TE to measured TE across 10 samples.

[0239] In another instance, an alternative multi-modal formulation of the model was trained to predict RNA-seq and Ribo-seq logTPM (transcripts per million) counts. TE was then inferred from the individual predictions. Predictions on a set of 1,522 test genes yielded a median Pearson correlation coefficient of 0.73 for comparison of predicted versus measured RNA-seq logTPM, and a median Pearson correlation coefficient of 0.74 for comparison of predicted versus measured Ribo-seq logTPM across 10 samples. With translation efficiency defined as the log ratio of Ribo-seq to RNA-seq counts, comparison of predicted TE to measured TE yielded a median Pearson correlation coefficient of 0.57 (or a Pearson correlation coefficient of 0.6 when averaged across cell types). VI. Computational Framework for Designing Regulatory Sequences

[0240] As noted above, in some instances, the disclosed machine learning-based models may be used for performing in silico design and / or optimization of cell type-specific and / or cell state-specific DNA regulatory sequences using a combination of long-context DNA sequence-based prediction models and a directed evolution process. The disclosed methods areMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 44 of 126 based on iterative in silico mutagenesis of a starting sequence (e.g., a random sequence or candidate sequence) and scoring of regulatory sequence (e.g., promoter sequence) efficacy based on predicted gene expression level. In some instances, the disclosed machine learning- based methods and models may be configured to process one or more input DNA sequences (e.g., in the context of a construct comprising the regulatory sequence and a cargo gene sequence) using one or more long-context DNA sequence-based prediction models (e.g., one or more different DNA sequence-prediction models configured to predict gene expression), while also evaluating one or more genomic insertion sites as part of the design and / or optimization of a cell type-specific and / or cell state-specific DNA regulatory sequence. In some instances, the long-context DNA sequence-based model(s) used with the disclosed framework for performing in silico design and / or improvement of cell type-specific and / or cell state- specific DNA regulatory sequences may be configured to output a vector comprising scalar values of predicted gene expression level for the input sequence for each cell type included in the training dataset. In some instances, the framework may be configured to output a predicted gene expression level (or regulatory sequence efficacy score derived therefrom) for the input sequence for a specified subset of cell types.

[0241] FIG.5 provides a non-limiting example of a flowchart for a process 500 for using a long-context DNA sequence-based prediction model to perform automated DNA regulatory sequence design and optimization. Process 500 can be implemented using, for example, the prediction system 100 described in FIG.1. For example, process 500 can be implemented using the sequence analyzer 108 and the machine-learning model(s) 132 shown in FIG.1.

[0242] At step 502 in FIG.5, at least one candidate DNA regulatory sequence is received (e.g., by one or more processors of a system configured to perform process 200D). Examples of the processing of DNA sequences to generate model inputs is described in more detail in Example 2 below (see, e.g., the method section on Creating Inputs for the Model).

[0243] In some instances, the at least one candidate DNA regulatory sequence comprises at least one promoter sequence, at least one enhancer sequence, at least one silencer sequence, at least one cis-regulatory sequence, at least one trans-regulatory sequence, or any combination thereof.

[0244] In some instances, the at least one candidate DNA regulatory sequences can comprise DNA sequence ranging in length from about 10 nucleotides (or base pairs (bp)) toMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 45 of 126 about 10,000 nucleotides (or base pairs (bp)), or longer. In some instances, the length of the at least one candidate DNA regulatory sequence can be at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 3,500, at least 4,000, at least 4,500, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000 nucleotides (or bp) in length. In some instances, the length of the at least one candidate DNA regulatory sequence can be at most 10,000, at most 9,000, at most 8,000, at most 7,000, at most 6,000, at most 5,000, at most 4,500, at most 4,000, at most 3,500, at most, 3,000, at most 2,500, at most 2,000, at most 1,500, at most 1,000, at most 900, at most 800, at most 700, at most 600, at most 500, at most 400, at most 300, at most 200, at most 100, at most 90, at most 80, at most 70, at most 60, at most 50, at most 40, at most 30, at most 20, or at most 10 nucleotides (or bp) in length. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the length of the at least one candidate DNA regulatory sequence can have a length ranging from about 200 to about 5,000 nucleotides. Those of skill in the art will recognize that the at least one candidate DNA regulatory sequence can have a length of any value within this range, e.g., about 232 nucleotides.

[0245] In some instances, the at least one candidate DNA regulatory sequences can comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1,000 candidate DNA regulatory sequences.

[0246] At step 504 in FIG.5, a cargo gene sequence is received (e.g., by the one or more processors of a system configured to perform process 200D).

[0247] In some instances, the cargo gene sequence may comprise the sequence for any specified gene (e.g., any human gene, any non-human mammalian gene, etc.).

[0248] At step 506 in FIG.5, at least one construct sequence is generated by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a specified genome at at least one candidate insertion site. Examples of generating constructs comprisingMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 46 of 126 cargo gene sequence are described below in Example 2 (see, e.g., the methods section on Regulatory Element Design), and in Examples 7 and 8.

[0249] In some instances, the genome is the human genome. In some instances, the genome may be a non-human genome, e.g., a non-human mammalian genome.

[0250] In some instances, the at least one candidate insertion site comprises a safe-harbor locus (e.g., a genomic locus where endogenous native gene expression level is minimal).

[0251] In some instances, the at least one construct sequence (e.g., a construct sequence comprising a candidate DNA regulatory sequence, the specified cargo gene sequence, and optionally, additional regulatory and / or non-regulatory sequence information (e.g., UTR sequences)) can be between about 200 bases and about 50 kilobases in length. In some instances, for example, the length of the construct can be at least 200 bases, at least 400 bases, at least 600 bases, at least 800 bases, at least 1,000 bases, at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb, at least 20 kb, at least 30 kb, at least 40 kb, or at least 50 kb. In some instances, the length of the construct can be at most 50 kb, at most 40 kb, at most 30 kb, at most 20 kb, at most 10 kb, at most 9 kb, at most 8 kb, at most 7 kb, at most 6 kb, at most 5 kb, at most 4 kb, at most 3 kb, at most 2 kb, at most 1,000 bases, at most 800 bases, at most 600 bases, at most 400 bases, or at most 200 bases. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the length of the construct sequence can have a length ranging from about 400 bases to about 4 kb. Those of skill in the art will recognize that the construct sequence can have a length of any value within this range, e.g., about 12.4 kb.

[0252] In some instances, the at least one construct sequence (e.g., a full construct sequence) can further comprise flanking regions of genomic sequence, and can be between about 100 kilobases and about 1 megabase in length. In some instances, for example, the length of the construct sequence (e.g., the full construct sequence) can be at least 100 kb, at least 200 kb, at least 300 kb, at least 400 kb, at least 500 kb, at least 600 kb, at least 700 kb, at least 800 kb, at least 900 kb, at least 1,000 kb. In some instances, the length of the construct sequence (e.g., the full construct sequence) can be at most 1,000 kb, at most 900 kb, at most 800 kb, at most 700 kb, at most 600 kb, at most 500 kb, at most 400 kb, at most 300 kb, at most 200 kb, at most 100 kb. Any of the lower and upper values described in this paragraph may be combined toMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 47 of 126 form a range included within the present disclosure. For example, in some instances the length of the construct sequence can have a length ranging from about 100 kb to about 600 kb. Those of skill in the art will recognize that the construct sequence can have a length of any value within this range, e.g., about 565 kb.

[0253] At step 508 in FIG. 5, the at least one construct sequence is input into a trained machine learning model configured to process an input construct sequence and predict a cell- type-specific and / or a cell-state-specific gene expression level. Examples of inputting a construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level are described in more detail in Examples 7 and 8.

[0254] In some instances, sequence data for the at least on candidate DNA regulatory sequence can be provided as input to the trained machine learning model in a format that includes the sequence data as a one-hot encoded matrix. In some instances, the sequence data for the at least on candidate DNA regulatory sequence can be provided as input to the trained machine learning model in a format that further comprises a gene mask (e.g., a binary mask as described elsewhere herein). In some instances, the input matrix provided as input to the model can comprise 5 dimensions (e.g., 4 for the one-hot encoded DNA sequence, and 1 for the mask).

[0255] In some instances, the trained machine learning model comprises a deep neural network architecture. In some instances, the deep neural network architecture comprises at least one convolution layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 convolutional layers), at least one downsampling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 downsampling layers), at least one self-attention layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 self-attention layers), at least one deconvolution layer with matched U-net connections (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 deconvolution layer with matched U-net connections), at least one global mean pooling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 global mean pooling layers), or any combination thereof. In some instances, the deep neural network architecture comprises a linear output layer.

[0256] In some instances, the trained machine learning model comprises a long-context DNA sequence prediction model, for example, a Borzoi model, an Enformer model, or a Decima model.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 48 of 126

[0257] In some instances, the machine learning model is pre-trained on bulk gene expression profile data (e.g., RNA-seq data and / or CAGE-seq data) and / or epigenomic assay data (e.g., DNase-seq data, ATAC-seq data, and / or ChIP-seq data).

[0258] In some instances, the trained machine learning model can be trained on single cell gene expression data for a plurality of cell types and cell states. In some instances, the single cell gene expression data comprises single cell RNA sequencing data.

[0259] In some instances, the plurality of cell types and cell states comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, or at least 200 different types of cell.

[0260] In some instances, the plurality of cell types and cell states comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, or at least 150 different cell states. In some instances, the different cell states comprise different cellular development states, different cellular metabolic states, different disease states, or any combination thereof.

[0261] In some instances, the prediction of cell-type-specific and / or a cell-state-specific gene expression level is made at single nucleotide resolution.

[0262] At step 510 in FIG.5, an efficacy score is determined for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state- specific gene expression level for the at least one construct sequence. Examples of determining an efficacy score for at least one candidate DNA regulatory sequence are described in Example 8.

[0263] In some instances, the efficacy score (or activity score) for the at least one candidate DNA regulatory sequence can be based on metrics such as a maximum predicted cell-type- specific and / or a cell-state-specific gene expression level, a mean predicted cell-type-specific and / or a cell-state-specific gene expression level, a median predicted cell-type-specific and / or a cell-state-specific gene expression level, or an integrated predicted cell-type-specific and / or a cell-state-specific gene expression level across the cargo gene sequence (or a specified genomic window), or any combination thereof. In some instances, the efficacy score (or activity score) can be based on a different metric (or different combination of metrics) for different types of functional genomic property profile data.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 49 of 126

[0264] In some instances, the efficacy score (or activity score) for the at least one DNA regulatory sequence can be further based on a determination of predicted cell-type-specific and / or a cell-state-specific off-target gene expression levels, i.e., expression levels in non- targeted cell types that are preferably kept as low as possible. For example, in a fibroblast specific regulatory element, off-target gene expression would correspond to gene expression in any cell type that is not a fibroblast. Typically, one tries to maximize on-target expression while minimizing off-target expression. A regulatory element with high cell-type “specificity” has high on-target and low off-target expression.

[0265] At step 512 in FIG.5, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence is selected. Examples of selecting at least one candidate DNA regulatory based, at least in part, on the efficacy score for the candidate are described in Example 8.

[0266] In some instances, the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and process 200D further comprises: for each of a plurality of iterations: generating a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; inputting the construct sequence into the trained machine learning model; determining an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: updating, using the one or more processors, the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or outputting, using the one or more processors, an improved cell- type-specific and / or a cell-state-specific DNA regulatory sequence. In some instances, the initial candidate DNA regulatory sequence comprises a random DNA sequence.

[0267] In some instances, process 500 further comprises repeating the method for each of a plurality of candidate insertion sites.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 50 of 126

[0268] In some instances, process 500 further comprises repeating the method using at least one additional trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level.

[0269] In some instances, process 500 further comprises using the improved cell-type- specific and / or a cell-state-specific DNA regulatory sequence to develop a gene therapy or personalized anti-cancer vaccine.

[0270] In some instances, updating the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence comprises performing at least one single nucleotide substitution, single nucleotide insertion, or single nucleotide deletion.

[0271] In some instances, updating initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence comprises performing at least one insertion, deletion, or substitution of a plurality of nucleotides.

[0272] In some instances, updating the initial candidate DNA regulatory sequence or candidate DNA regulatory sequence comprises inserting nucleotides corresponding to at least one known regulatory sequence motif.

[0273] In some instances, the improved DNA regulatory sequence comprises a promoter sequence, an enhancer sequence, a silencer sequence, a cis-regulatory sequence, or a trans- regulatory sequence.

[0274] In some instances, the initial candidate DNA regulatory sequence is between about 10 and about 1,000 nucleotides in length. In some instances, the length of the initial candidate DNA regulatory sequence can be at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1,000 nucleotides (or bp) in length. In some instances, the length of the initial candidate DNA regulatory sequence can be at most 1,000, at most 900, at most 800, at most 700, at most 600, at most 500, at most 400, at most 300, at most 200, at most 100, at most 90, at most 80, at most 70, at most 60, at most 50, at most 40, at most 30, at most 20, or at most 10 nucleotides (or bp) in length. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the length of the initial candidate DNA regulatory sequence can have a length ranging from about 20 to about 200MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 51 of 126 nucleotides. Those of skill in the art will recognize that the initial candidate DNA regulatory sequence can have a length of any value within this range, e.g., about 65 nucleotides.

[0275] In some instances, the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence is between about 10 and about 1,000 nucleotides in length. In some instances, the length of the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence can be at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1,000 nucleotides (or bp) in length. In some instances, the length of the improved cell-type-specific and / or a cell-state- specific DNA regulatory sequence can be at most 1,000, at most 900, at most 800, at most 700, at most 600, at most 500, at most 400, at most 300, at most 200, at most 100, at most 90, at most 80, at most 70, at most 60, at most 50, at most 40, at most 30, at most 20, or at most 10 nucleotides (or bp) in length. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the length of the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence can have a length ranging from about 30 to about 300 nucleotides. Those of skill in the art will recognize that the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence can have a length of any value within this range, e.g., about 56 nucleotides. VII. Examples Example 1 – Model Training Performance Data

[0276] This non-limiting example describes model prediction performance data for the training of a new machine learning model (“Decima”) to predict the cell type-specific and cell state-specific expression profiles of genes previously not seen by the model.

[0277] A non-limiting example of the model architecture is shown in FIG.2 and has been described elsewhere herein. In brief, the model disclosed herein used the Borzoi model (Linder et al. (2023)) as a foundational model, but the terminal linear layer of the Borzoi model was removed and a mean pooling layer along the length axis was added, followed by a linear layer that takes in a vector of size 1920 for each input and returns a vector of size 8,848. Since theMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 52 of 126 new model takes inputs of size (5 x 524,288), whereas the Borzoi model takes inputs of size (4, 524,288), an additional input channel was also added to the first convolutional layer.

[0278] The process used to train the model is illustrated in FIG. 6, and has also been described elsewhere herein. In brief, the model was trained using a two-component loss function applied across the pseudobulk axis for a single gene. The loss function consisted of two terms: (i) a Poisson loss applied to the total value of the gene across all pseudobulks, and (ii) a multinomial loss function that compares the true and predicted distribution of expression values across pseudobulks for the same gene.

[0279] Table 1 provides a non-limiting example of model prediction performance data comparing different model initialization strategies. Average Pearson correlation coefficients (per-pseudobulk and per-gene) were computed over held-out test genes for a single replicate of the Decima model trained using the indicated model initialization strategies. Borzoi replicate 0 was used for initialization. The method used to initialize the Decima model used in the examples below is highlighted in bold. Initializing the model from Borzoi, followed by tuning all layers yielded the best prediction performance, as indicated by the Pearson correlation coefficients. Table 1. Comparison of model performance data for different model initialization methods.

[0280] Table 2 provides a non-limiting example of model prediction performance data comparing different weights for the Poisson loss component within the two-component loss function. Average Pearson correlation coefficients (per-pseudobulk and per-gene) were computed over held-out test genes for single replicates of the Decima model trained with varying weights for the Poisson loss component within the two-component loss function. Borzoi replicate 0 was used for initialization. The value used to train the Decima model usedMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 53 of 126 in the examples below is highlighted in bold. Training the model using a Poisson weight of 1e- 4 yielded the best prediction performance, as indicated by the Pearson correlation coefficients. Table 2. Comparison of model performance data for models trained using different Poisson weights in the two-component loss function.

[0281] Table 3 provides a non-limiting example of model prediction performance data comparing different strategies for masking the input sequence DNA data. Average Pearson correlation coefficients (per-pseudobulk and per-gene) were computed over held-out test genes for single replicates of the Decima model trained with and without a binary gene mask. As an alternative to the binary gene mask, a version of Decima was also trained with embedding cropping, where the embeddings produced by the model are cropped to the boundaries of the gene of interest immediately before being passed to the mean pooling layer. Table 3. Comparison of model performance data for models trained using different input sequence data masking strategies.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 54 of 126

[0282] FIG.7 provides a non-limiting example of box plots showing the distribution of the Pearson correlation between measured and predicted expression values for test set genes, as predicted by the Decima model for pseudobulks from different human organs. The dashed line shows the overall median correlation.

[0283] FIG.8 provides a nonlimiting example of box plots showing the distribution of the Pearson correlation between measured and predicted expression values for test set genes, for each pseudobulk annotated with a disease state. Pseudobulks are split based on their annotated disease state and box plots are colored based on the data source. The dashed line shows the overall median correlation.

[0284] FIG.9 provides non-limiting examples of precision-recall curve for classification of high-confidence sc-eQTLs from negative control variants, using either the Decima model’s predictions in the matched cell type (Decima), Borzoi predictions in the whole blood GTeX track (Borzoi whole blood), Borzoi predictions in RNA-seq or CAGE tracks manually matched to the individual cell types in the OneK1K dataset (Borzoi matched; see the Methods section in Example 2), or distance of the variant from the gene TSS (Distance).

[0285] FIGS. 10A-10C provide non-limiting examples of scatter plots showing AUPRC data for classification of high-confidence sc-eQTLs from negative control variants in each of 21 cell types. Each point represents a cell type. FIG. 10A: plot of the AUPRC based on predictions from the Decima model versus predictions from the Borzoi model for the GTEx whole blood track. FIG.10B: plot of the AUPRC based on predictions from the Decima model versus predictions from the Borzoi model for RNA-seq or CAGE tracks manually matched to the individual cell types in the OneK1K dataset (see the Methods section in Example 2). FIG. 10C: plot of the AUPRC based on predictions from the Decima model versus the AUPRC based on distance of the variant from the gene TSS.

[0286] Table 4 provides a non-limiting example of model prediction performance data comparing the performance of the Decima model to that of the Borzoi model at separating 837 GWAS variants from 8,370 matched negative control variants based on different performance metrics (AUPRC: Area Under the Precision-Recall Curve; AUROC: Area Under the Receiver Operating Characteristic (ROC) curve).MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 55 of 126 Table 4. Comparison of model performance data for identifying GWAS variants.Example 2 - Prediction of Pseudobulk Gene Expression from DNA Sequence

[0287] This non-limiting example describes the training and use of a new machine learning model (“Decima”) to predict the cell type-specific and cell state-specific expression profiles of genes previously not seen by the model. As illustrated in FIG. 11, single-cell datasets were combined, filtered, and aggregated to create a pseudobulk gene expression matrix. Each row of the matrix corresponds to a unique combination of cell type, tissue, disease, and study, and each column corresponds to a single gene. The new model accepts the DNA sequence surrounding a single gene as input and predicts the corresponding column of the expression matrix.

[0288] Methods - Data Sources: Publicly-available single cell (sc) and single nucleotide (sn) RNA-seq count matrices were downloaded from the following sources: • SCimilarity (sc / snRNA-seq): Individual datasets were downloaded and prepared as described in Heimberg et al. (2023). •• Human skin atlas (scRNA-seq): •

[0289] Methods - Aggregation and Filtering of Pseudobulk Samples: The count vectors for all cells corresponding to the same cell type, tissue, disease and study were summed into aMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 56 of 126 single pseudobulk, producing a pseudobulk x gene count matrix using the scanpy.get.aggregate (adata, metadata_columns, func='sum') function in scanpy version 1.10.2, where ‘adata’ represents the cell-level count matrix and ‘metadata_columns’ denotes the metadata columns identifying each pseudobulk. Note that in this process, cells from different individuals in the same study are summed into a single pseudobulk.

[0290] Pseudobulks corresponding to cell lines, cancers, organoids, or unannotated cell types were discarded. From the SCimilarity dataset, pseudobulks corresponding to samples from the brain, skin, and retina were also discarded, as these tissues were better represented and annotated in the other atlases. Since cell types in the SCimilarity dataset were automatically annotated using an embedding-based nearest-neighbors approach, cell type annotations were also manually examined, and cell types that were likely misannotated (e.g. cell type annotations inconsistent with the tissue) were removed. Finally, extremely low quality / sparse pseudobulks that were composed of fewer than 50 cells, and which were also in the lowest 10% of pseudobulks in terms of the number of genes measured and total counts, were also removed.

[0291] Methods - Gene Filtering and Annotation: Genes and pseudobulks for which > 33% of expression values were missing were discarded. Genes were then annotated with their start and end coordinates in the hg38 genome using genome annotation files obtained from CellRanger (https: / / www.10xgenomics.com / support / software / cell-ranger / latest) and the NCBI gene database (https: / / www.ncbi.nlm.nih.gov / gene). All genes except those in autosomes or Chromosome X were removed.

[0292] Methods - Normalization: Each pseudobulk was normalized such that the total counts across all genes equaled 1 million, followed by applying a log(1+x) transform to all counts. Thus, the final values in the matrix represent log(1+CPM) values.

[0293] Methods - Creating Inputs for the Model: For each gene, a 524,288 bp genomic interval starting at 163,840 upstream of the TSS and extending through the gene body was created. The gene start was placed at this position because the Borzoi architecture lacks information to make fully accurate predictions for bases closer to the input boundary than this (Linder et al. (2023)). In cases where the interval extended near or beyond the end of the chromosome, the interval was shifted to be at least 10 kb away from the chromosome end. For each interval, the corresponding sequence from the hg38 genome assembly was extracted and intervals where > 40% of the bases were not resolved (e.g., because the interval was notMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 57 of 126 sequenced experimentally, or because the interval was not mappable so sequence reads could not be assigned to it) were dropped for genes on the negative strand. In all but 643 of the remaining 18,457 genes, the interval extended beyond the gene end.

[0294] To create inputs for the new model, the genomic sequence for the interval corresponding to each gene were taken, and the corresponding sequence on the negative strand was determined by reverse complementarity. Each gene sequence was then one-hot encoded to obtain a matrix of dimensions (4 x 524,288). To this matrix, a fifth row containing a binary mask that highlighted the location of the gene of interest was added, i.e. the value of the mask was 1 for all positions between the gene start and end, and 0 for all other positions. Combining the one-hot encoded sequence and the gene mask resulted in a matrix of shape (5 x 524,288) representing each gene.

[0295] Since the new model was initialized using the weights of the trained Borzoi model, the gene sequences were split into training, validation, and test sets based on their overlap with the genomic regions used for these sets in Borzoi. As a result, the new model’s test set of 1,811 genes do not overlap with any genomic region used for training or validation of Borzoi. For genes in the training set, the one-hot encoded sequence for 5000 additional bases were also extracted and added at the start and end of the interval. This was done to allow data augmentation by shifting the input window during training, as described below.

[0296] Methods – Model Architecture: The PyTorch version of the Borzoi model (Linder et al. (2023)) was downloaded from gReLU (Lal et al. (2024)). The terminal linear layer of this model was removed and an average pooling layer along the length axis was added, followed by a linear layer that takes in a vector of size 1920 for each input and returns a vector of size 8,848.

[0297] Since the model takes inputs of size (5 x 524,288), whereas the Borzoi model takes inputs of size (4, 524,288), an additional input channel to the first convolutional layer of the model was also added.

[0298] Methods - Training the Model: The new model’s loss function is based on the two- component loss function used in Borzoi (Linder et al. (2023)). However, it is applied across the pseudobulk axis for a single gene rather than within individual pseudobulks. Specifically, the loss function consists of two terms: a Poisson loss applied to the total value of the gene across all pseudobulks, and a multinomial loss function that compares the true and predictedMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 58 of 126 distribution of expression values across pseudobulks for the same gene. During training, a weight of 10-4was applied to the Poisson component. This down-weighting encourages the model to learn to predict differences in expression of the same gene between pseudobulks.

[0299] The entire model was trained for 15 epochs with the Adam optimizer, a learning rate of 3x10-5, and a batch size of 4 with gradient accumulation over every 5 batches. Data augmentation was performed during training by randomly shifting the input sequence by up to 5,000 bases in either direction along the genome. The validation loss was measured after each epoch and the model with the lowest validation loss was saved. The training procedure was performed on a single NVIDIA A100 GPU and took approximately one day. This was repeated 4 times, each time using a different replicate (i.e., a copy having the identical architecture, trained using the same training data and training method, but initialized with a different random set of initialization weights) of the Borzoi model (Linder et al. (2023)).

[0300] Methods - Classification of Cell Type-Specific Genes: To determine cell-type specific genes, a procedure similar to that described by Eraslan et al. (2022) was followed at the pseudobulk level. First, the expression of each gene was z-score normalized across pseudobulks. Then a mean z-score per cell type was computed for each gene by averaging across all pseudobulks annotated as corresponding to the cell type. For each cell type, the genes which had mean z-score greater than 1 in that cell type were then labeled as being specific to that cell-type.

[0301] The z-score normalization and averaging for the predicted expression values was repeated, and the ranking implied by the predicted mean z-scores was used to classify whether a gene was specific for a particular cell-type (i.e., whether it had an observed mean z-score > 1). The number of specific genes (and thus the implied class balance) will vary by cell type. For this reason, AUROC was reported as the main metric since it is insensitive to differences in class-balance, and the performance of a random classifier will be 0.5 in every cell-type. For the purpose of computing the AUROC metric, only test set genes were considered.

[0302] Methods - Comparing Attributions Between Sequence Regions: Coordinates of all annotated cis-regulatory elements (CREs) in the hg38 genome were downloaded from the ENCODE database (ENCODE Project Consortium et al. (2020)). Coordinates of exons for all genes were taken from the genome annotation file downloaded from CellRanger (https: / / www.10xgenomics.com / support / software / cell-ranger / latest).MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 59 of 126

[0303] For each gene in the test set, the predicted expression was averaged over all pseudobulks where the gene was expressed above a threshold (log(CPM+1) > 1.5). The Input x Gradient method was used to compute the attribution of this average expression (i.e., the relative importance of each input feature value on the average expression) for each nucleotide in the 524,288 bp input sequence. For each gene, the average of the absolute value of this attribution was computed over all nucleotides overlapping with the annotated exons of the gene, the promoter (+ / - 100 bp from the gene transcription start site (TSS)), and exon / intron junctions (+ / -10 bp surrounding each junction). For nucleotides in introns, separate averages were computed for those nucleotides that overlapped with an annotated cis-regulatory element (CRE) (“Intronic CREs”) and those that did not. Nucleotides outside the gene body were divided into categories based on their distance from the gene body (<1000bp, 1000-10,000bp, 10,000 - 100,000 bp and >=100,000 bp) and each category was then further divided into nucleotides that overlapped with an annotated CRE and those that did not. The average of the absolute value of the Input x Gradient attributions over all nucleotides was then calculated in each category.

[0304] Methods - Cell Type Specificity Interpretation: To find lineage and cell type-specific motifs in, for example, lung epithelium, the following procedure was used. First, the cell type specificity of expression was computed for each test set gene using the z-score method as done previously, except that the underlying gene-expression matrix was restricted only to lung / airway-related tissues. For each of the main epithelial cell types (ciliated cell, club cell, secretory cell, goblet cell, respiratory basal cell, type I pneumocyte and type II pneumocyte), the top 250 most specific genes (i.e., the 250 genes with the highest mean z-score for this cell- type) was then selected. For each of these genes, the gradient of the difference in gene expression between the cell type for which the gene was specific and the remaining epithelial cell types was computed with respect to each nucleotide in the input sequence. This gives a ‘differential attribution’ which highlights the parts of the sequence which the new model considers relevant for driving cell type-specific expression of genes in each individual epithelial cell type (Lal et al. (2024)). TF-MoDISco (Shrikumar et al. (2018)) was then used to assemble these attributions into motifs. To limit the considerable runtime of TF-MoDISco, the clustering was restricted to the region + / - 10kb around the annotated TSS of the gene, and the maximum seqlets was restricted to 10,000. Otherwise, default parameters were used.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 60 of 126

[0305] After clustering using TF-MoDISco, the resulting motifs were matched to HOCOMOCO v12 motifs (Kulakovskiy et al. (2018)) using Tomtom (Gupta et al. (2007)). Since many TFs share similar motifs, the HOCOMOCO motifs were clustered using the GimmeMotifs software (Bruse and van Heeringen (2018)). Each motif hit from TF-MoDISco was then associated to the respective motif cluster, and aggregated hits for the same cluster. Some motifs were excluded, specifically those corresponding to certain Zinc Finger TFs, because their low information content can lead to highly significant but low-quality Tomtom matches.

[0306] For each detected motif cluster, an aggregate “specificity contribution score” was computed based on TF-MoDISco as follows: the total number of positive seqlets which could be matched to any non-excluded motif with q-value < 0.05 was counted, and then the ratio of the number of seqlets found for this motif to the total number of seqlets was computed. If a motif cluster was not detected in a particular cell-type, the corresponding score is thus zero. Negative seqlets and negative motifs were excluded since, in such a differential analysis, they may correspond to motifs of other cell-types rather than cell-type-specific repressors. The 4 motifs where this score was most strongly enriched for a specific cell type or lineage were selected to highlight.

[0307] To find motifs which differentiate individual cell types, cell states or the same cell type in different tissues, the following procedure was used. Firstly, the tracks corresponding to the ‘positive’ condition (e.g., for a fibroblast cell type, this could be all fibroblast pseudobulks with organ = ‘heart’) and those corresponding to the ‘negative’ condition (e.g., all other fibroblast tracks) were selected. On the test set genes, the correlation between predicted and observed fold change between the positive and negative pseudobulks was computed. From all test genes, the 300 genes which the model predicts to be most upregulated in the ‘positive’ condition were then selected, and the gradient of the difference in predicted expression between the ‘positive’ and ‘negative’ condition was computed. TF-MoDISco clustering was then performed as described above. Highlighted motifs were found by manual inspection of the results.

[0308] Note that for these analyses, all predictions and attributions were averaged across all four model replicates. Additionally, the mean attribution across nucleotides of the gradient was subtracted, as described in Majdandzic, Rajesh, and Koo (2023)).MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 61 of 126

[0309] Methods - Processing sc-eQTL Data: The reprocessed OneK1K sc-eQTL (single cell Expression Quantitative Trait Loci) data (Yazar et al. (2022); Kerimov et al. (2021)) was downloaded from the EBI eQTL catalogue FTP server (https: / / ftp.ebi.ac.uk / pub / databases / spot / eQTL / susie / QTS000038 / ). All variants which could not be scored by the new model were filtered out because they are too far from their target gene. Variants in non-standard chromosomes and ENCODE blacklist regions were also filtered out. Additionally, because indels can change the length of the underlying sequence making comparisons more difficult, only single nucleotide variants were analyzed.

[0310] For each cell type in the OneK1K data, the SuSiE (G. Wang et al. (2020)) fine- mapped eQTL variants which achieved posterior inclusion probability (PIP) greater than 0.9 for some gene were then collected. These variants are very likely to be causally impacting the expression of their target gene and thus form the positive set for our eQTL analyses. For each pair of high Posterior Inclusion Probability (high-PIP) eQTL variant and target gene, 20 variants which were tested for the same gene but did not achieve statistical significance, did not form part of any SuSiE credible set for any gene in any cell type, and which had an allele frequency > 5% were then extracted. Where more than 20 such variants were available, the ones closest to the TSS of the target gene of the corresponding high-PIP variant were prioritized (as many causal variants tend to be close to the TSS). These variants formed the negative set, since the data provides no evidence that they have a measurable impact on gene expression. Accordingly, this procedure produced a class-balance of 1:20.

[0311] Methods - sc-eQTL Variant Effect Prediction: New model cell types were matched to OneK1K cell-types. Note that the matching was restricted only to pseudobulks coming from tissue “blood”, i.e., liver macrophages, for example, were not included. Gamma-delta and double-negative T-cells were not matched, since they were not annotated in the training data. For these cell types, the average predicted variant effect was computed across all blood cell types.

[0312] To compute the predicted variant effect for the new model, the following procedure was used. For each variant-gene pair, the genomic interval around the gene was taken as the reference sequence, as defined previously. To construct the alternate sequence, the reference allele was replaced with the eQTL alternate allele. Model predictions were then computed for both the reference and alternate sequence, and the difference in prediction recorded as theMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 62 of 126 variant effect (illustrated in FIG.40). This represents a log fold change as the model’s predicted values are in log1p scale. Note that the model will predict one such variant effect per pseudobulk. To get a consolidated score, the effect was averaged across all pseudobulks which matched the sc-eQTL cell-type of interest, using the matching defined above.

[0313] To test the cell type specificity of the new model’s eQTL predictions, cell type specific variants were first identified. For this, fine-mapped variants which achieved a PIP > 0.9 for a particular expression gene (eGene) in some cell-type were selected. For each such variant-gene pair, all the cell types where the variant achieved PIP > 0.9 as “on-target” cell types were denoted. All cell types where the variant achieved PIP < 0.1 were considered “off- target” cell types. Cell-types where the variant had an intermediate PIP were excluded. Moreover, off-target cell-types where the gene was not expressed, i.e. had an average log1p expression of less than 1.5, were excluded. This threshold was chosen as it separates the two modes of the distribution of average expression (across pseudobulks) of all genes. The on- target and off-target cell-types were then matched to Decima pseudobulks using the same procedure as described above. In some cases the resolution of the OneK1K data exceeds that of the new model, e.g. in the case of subtypes of CD4 T-cells. In this case, if a variant-gene pair is “on-target” for, e.g., one type of CD4 T-cell, then it was considered on-target for CD4 T-cells altogether. After matching, Decima variant effect predictions per cell type were averaged (to account for the fact that some cell-types are represented by more pseudobulks than others), and then the predictions for on-target and off-target cell-types, respectively, averaged.

[0314] Methods - Processing Genome-Wide Association Studies (GWAS) Data: To identify a set of high probability disease causing variants, fine mapping using summary statistics from 39 traits that had “gold-standard” gene sets defined by Zhang et al. (2022) was conducted. Our fine-mapping pipeline consists of three steps under the assumption of a maximum of 5 causal loci in the 1,000,000 bases surrounding each lead SNP. In the first step, DENTIST (W. Chen et al. (2021)) was used to identify and remove SNPs that are discordant with patterns of linkage disequilibrium (LD) in a defined reference panel, in this case, 50,000 individuals from the 380k UKBiobank reference panel (s3: / / broad-alkesgroup-ukbb-ld / UKBB_LD / ). In the second step, PolyFun (Weissbrod et al.2020) was used to compute functional priors for fine mapping using the baseline annotations (baselineLF2.2.UKB) provided by Gazal et al (2018) and L2-MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 63 of 126 regularized S-LDSC (stratified LD score regression). In the third step, the resulting priors were fed to SuSiE (G. Wang et al. (2020)) to generate a credible interval around each lead SNP and to assign posterior inclusion probabilities (PIP) to individual SNPs.

[0315] From the fine-mapping results, high-confidence (PIP > 0.9 and p-value < 10-6) GWAS SNPs within 100 kb of any annotated TSS were selected. Further, the subset of these SNPs that were annotated as regulatory variants in gnomAD (Karczewski et al. (2020)), and did not overlap with ENCODE blacklist regions, were selected. These 837 variants formed the positive set for our GWAS analyses.

[0316] Since the gene that is affected by these variants is unknown, the effect of each positive variant on all genes for which the variant overlapped with Decima’s prediction interval was predicted. Each gene was matched to the variant with the highest predicted effect (mean across all output cell types / tracks).

[0317] For each of these ‘positive’ GWAS variants, 10 SNPs which were annotated as regulatory and had MAF > 1% according to gnomAD, did not have PIP > 0.01 for any trait in the fine-mapped GWAS datasets, did not overlap with blacklist regions, and were between 10- 120,000 bp from the TSS of the matched gene were then extracted. Where more than 10 such variants were available, the 10 variants whose distance from the TSS was most similar to that of the GWAS variant was selected. These variant-gene pairs formed the negative set. Variant effect predictions for these variant-gene pairs were made using the new model (“Decima”) and Borzoi as described above.

[0318] To analyze variant effect predictions at the cell type level, Decima’s predicted variant effect scores were averaged for all tracks corresponding to the same cell type. To identify which high-PIP GWAS variants are also known eQTLs, fine-mapped eQTL data provided by Open Targets using the September 2022 release (https: / / ftp.ebi.ac.uk / pub / databases / opentargets / genetics / 22.09 / ) was used. Specifically, all data in the “v2d_credset” analysis directory was downloaded and the results filtered for type “eQTL” (Expression Quantitative Trait Locus) and a posterior probability (PIP) of at least 0.9.

[0319] Methods - Cell Type Specificity Analysis for GWAS Variants: GWAS variants which were predicted to have significantly higher impact (z-test p < 0.01) than their 10 matched control variants according to Decima were selected. Due to the limited number of such variants known for individual traits, variants corresponding to similar traits were grouped into generalMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 64 of 126 categories. Specifically, Inflammatory Bowel Disease, Crohn’s disease, asthma, eczema, autoimmune disease, lupus, and multiple sclerosis were combined into the category ‘Autoimmune and Inflammatory Diseases’. Mean corpuscular hemoglobin, RBC width, RBC count, and platelet count were combined into ‘Blood-related traits’.

[0320] For each remaining variant, the difference between Decima’s absolute predicted log- fold change (VEP score) for the variant and for its 10 matched negative control variants were computed in each cell type. These were then converted into z-scores and the per-cell type z- scores averaged for all variants corresponding to the same trait (or category).

[0321] Methods - Disease Interpretation: All pseudobulks from disease samples in the data which could be matched to a healthy control pseudobulk were first selected using the following criteria: • A corresponding healthy pseudobulk from the same tissue, cell type and study was available. • Both the disease and the corresponding healthy pseudobulk were supported by at least 500 cells each. • The absolute log-difference in size factor (sum of expression values for all genes) between the disease and the corresponding healthy pseudobulk was at most 0.5.

[0322] This resulted in a robust set of pairs of matched disease and healthy pseudobulks, with each pair corresponding to a unique “quadruplet” of cell type, tissue, study and disease. For these pairs, the observed and predicted log fold change in gene expression between the disease state and the corresponding healthy pseudobulk were then computed (FIG.47). Using only test set genes, the correlation between the observed and predicted log fold changes was computed.

[0323] For a subset of case studies, selected from the diseases where the new model performs best on test-genes, the 300 genes where Decima predicted the highest upregulation in disease were selected, and the gradients for the differential gene expression in disease vs. healthy states was computed. TF-MoDISco clustering was then used to find motifs, as described above.

[0324] Methods - Regulatory Element Design: In the example described below, the synthetic element includes a 200 bp regulatory sequence appended to the EBFP cargo gene (blue emission fluorescent protein). The cargo gene sequence was obtained from the GenBankMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 65 of 126 database (GenBank ID: MH450096.1) (Grajevskaja et al. (2018)). The synthetic element was inserted in a genomic safe-harbor locus located at chr22:29316378-29840666 in silico. This locus was selected due to its lack of endogenous gene expression. All in silico testing was evaluated using this placement location to ensure consistency across iterations. This construct was introduced 5120*32 base pairs (163,840 bp) upstream of the gene start location within the 524,288 bp input sequence, which matched the gene location observed by the model during training. At each round of directed evolution, in silico mutagenesis was performed across the 200bp regulatory element, and the single nucleotide mutation that maximized the difference in predicted expression between two target groups selected.

[0325] Fibroblast and non-fibroblast pseudobulks from a previously published Crohn’s disease study which formed part of the new model’s training set (Collection ID: 17481d16- ee44-49e5-bcf0-28c0780d8c4a) (Elmentaite et al. (2020)) were selected. Specifically, 1 healthy fibroblast sample, 1 UC fibroblast sample, 22 non-fibroblast healthy samples, and 22 non-fibroblast UC samples were selected. In the first 50 rounds of evolution, the regulatory elements were improved to drive expression specifically in fibroblasts relative to non-fibroblast cells, i.e. the difference in average expression between the 2 fibroblast pseudobulks and all 44 non-fibroblast pseudobulks was maximized. During the remaining 50 rounds, the design was refined to enhance expression in Crohn’s-derived fibroblasts compared to healthy fibroblasts, optimizing the sequence to maximize differential expression between the Crohn’s-fibroblast pseudobulk and the healthy fibroblast pseudobulk.

[0326] Motifs were identified using the HOCOMOCOv12 database (Kulakovskiy et al. (2018)) and gReLU (Lal et al. (2024)).

[0327] Results – Prediction of Pseudobulk gene Expression from DNA Sequence: As noted above, annotated sc / snRNA-seq atlases was obtained from SCimilarity (Heimberg et al. (2023)), the human brain atlas (X. Chen et al. (2024)), the skin atlas (Fiskin et al. (2023)), and the adult human retina atlas (Jin Li et al. (2023)). After processing and filtering, the training corpus comprised data from over 22 million cells.

[0328] The RNA-seq counts from all cells belonging to the same combination of cell type, tissue, disease state, and study was summed, resulting in a final count matrix consisting of 8,856 pseudobulk expression vectors and 18,457 genes. This matrix includes pseudobulks representing 201 distinct cell types in 271 tissues and 82 diseases; there are 6,481 uniqueMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 66 of 126 combinations of cell type, tissue, and disease state, and 4,478 unique combinations of cell type and tissue annotations. The counts in each pseudobulk were normalized to CPM (counts per million) values and applied a log(1+x) transformation for variance stabilization.

[0329] Each example seen by the model was focused on a single gene. For each gene, an input window containing 524,288 bp of genomic sequence surrounding the gene was created. This input window contained the gene transcription start site (TSS) and a minimum of 163,840 bp upstream of the TSS, with the remaining sequence covering the gene body and downstream regions. In addition to the one-hot encoded DNA sequence in the input window, the model was also supplied with a binary mask representing the location of the gene body in the input window (FIG.12).

[0330] The model was initialized with the recent Borzoi model (Linder et al. (2023)), which was trained to predict bulk expression profiles (RNA-seq and CAGE-seq) as well as epigenomic profiles (DNase-seq, ATAC-seq, and ChIP-seq) from the genome sequence. This allowed the model to benefit from learned epigenetic regulatory mechanisms, which are more easily learned from these assays than from scRNA-seq data, as well as from learned rules related to gene expression at bulk resolution. The final layer of the Borzoi model was replaced with a global mean pooling layer along the length axis, followed by a linear layer that outputs a vector of 8,856 values for each gene, corresponding to the predicted log(CPM+1) values in each pseudobulk (FIG.11). The 18,457 genes in the pseudobulk matrix were split into groups for training, validation, and testing based on their overlap with the genomic regions used to train Borzoi, such that the sequences in the test set were also in the test set of Borzoi.

[0331] The entire model was trained using a novel two-component loss function comprising a Poisson loss on the total expression of the gene across all pseudobulks, and a multinomial loss along the pseudobulk axis. This is based on the two-component loss introduced in (Avsec et al. (2021)) and adapted for Borzoi (Linder et al. (2023)); however, in these studies, the multinomial loss was calculated across sequence positions, encouraging the model to focus on between-region differences in biological activity. In contrast, our loss function encourages the model to correctly predict differences in expression between pseudobulks for a single gene.

[0332] The model was evaluated by predicting the expression of 1,811 held-out genes in all pseudobulks (FIG. 12). Four Decima models were trained, each initialized with a different replicate, and averaged their predictions to increase robustness. Considering only held-outMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 67 of 126 genes, the mean Pearson correlation between the measured and predicted expression vectors for each pseudobulk was 0.80, indicating that the model performs well at predicting differences in expression between unseen genes in the same condition (FIG.13, showing a histogram of Pearson correlation coefficients between the measured and predicted expression vectors for each of the 8,856 pseudobulks (the rows of the matrix in FIG.12), over the test set of 1,811 genes). This performance was largely consistent across the wide variety of datasets, cell types, and disease states included in the dataset. Furthermore, the mean correlation between the measured and predicted expression vectors for each held-out gene was 0.58, indicating that the model can also predict differences in expression of the same gene between different populations of cells (FIG.14, showing a histogram of Pearson correlation coefficients between the measured and predicted per-gene expression vectors (the columns of the matrix in FIG.12) for each of the 1,811 test set genes). This analysis thus validated that the model successfully learned determinants of both between-gene and between-cell type expression changes.

[0333] This example demonstrates that the disclosed model accurately predicts gene expression values (for each of a plurality of cell types and / or cell states (e.g., disease states)) based on an input comprising the genomic DNA sequence for a single gene and portions of adjacent sequence. Example 3 – Identification of Sequence Determinants of Cell Type Specificity

[0334] This non-limiting example describes the training and use of a machine learning model (“Decima”) to identify sequence determinants of cell type specificity. As illustrated in FIG. 15, the model was tested to determine whether it could successfully predict cell type- specific patterns of gene expression from sequence alone.

[0335] Test set genes specific to each cell type were first identified in the dataset based on their z-score normalized expression values (FIG. 15). Specifically, a gene was labeled as “specific” to a cell type if its average z-score in that cell type based on its measured expression was >=1. For each cell type, the z-scores based on Decima’s predicted expression matrix were tested to determine whether they could be used to classify the specific and nonspecific genes for that cell type (Eraslan et al. (2022)) (FIG. 15). An average AUROC of 0.82 across cell types was obtained (FIG. 16). This analysis suggested that the new model (“Decima”) has learned sequence features that determine the expression patterns of even cell type-specific genes.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 68 of 126

[0336] This was verified by examining the model’s predictions for individual genes with highly specific expression in diverse cell types. For example, the FABP1 gene encodes a protein critical for the uptake and intracellular transport of fatty acids, and is highly expressed in the enterocytes of the intestinal lining and hepatocytes in the liver. The model correctly predicts its elevated expression in both these cell types, distinguishing them from other cell types in the same tissues (FIG.17, providing an example of a scatter plot showing the measured and predicted expression (log (CPM+1)) of the FABP1 gene in all pseudobulks; FIG. 18, providing examples of boxplots showing the measured and predicted expression of the FABP1 gene in pseudobulks representing enterocytes and hepatocytes, compared to pseudobulks representing other cell types in the gut, other cell types in the liver, and all remaining cell types). Similarly, the model correctly predicts the expression of DNAH6, a gene that encodes a component of cilia and is consequently highly expressed in the ependymal and choroid plexus cells of the brain and in the ciliated cells of the lung (FIG.19, providing an example of a scatter plot showing the measured and predicted expression of the DNAH6 gene in all pseudobulks; FIG. 20, providing examples of boxplots showing the expression of the DNAH6 gene in pseudobulks representing ependymal cells, choroid plexus cells, ciliated cells, other cell types in the central nervous system, other cell types in the lung, and all remaining cell types); and SPI1, which encodes a transcriptional regulator of myeloid cell development and is highly expressed in cells of the myeloid lineage (FIG. 21, providing an example of a scatter plot showing the true and predicted expression of the SPI1 gene in all pseudobulks; FIG. 22, providing examples of Boxplots showing the expression of the SPI1 gene in pseudobulks representing monocytes, macrophages, microglia, other cell types in the blood, other cell types in the CNS, and all remaining cell types). In all of these boxplots, the lower and upper hinges correspond to the first and third quartiles, whiskers extend to 1.5 * IQR (inter-quartile range), and the remaining points are plotted individually.

[0337] The model was next evaluated for its ability to identify genomic features relevant to expression, including intron / exon boundaries and cis-regulatory elements. For each gene in the test set, the new model’s attributions were calculated across all nucleotides in the input genomic interval, using the Input x Gradient method (Shrikumar et al. (2017)). Gradients were calculated with respect to the average expression of the gene across all pseudobulks where the gene was strongly expressed (log(CPM + 1) > 1.5). The absolute value of these attributions,MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 69 of 126 averaged across a genomic region, were used as a measure of the importance of that region to the predicted expression of the gene. On average, the magnitude of the attributions was highest within the gene promoter, followed by the exon / intron junctions and exons of the gene. In comparison, introns of the same gene had lower attributions. However, intronic cis-regulatory elements (CREs) had higher attributions than the rest of the intron (FIG. 23, providing an example of box plots of per-nucleotide attribution scores averaged across different sequence elements in the gene body and promoter, for all test set genes; values plotted represent the absolute value of the input x gradient score, averaged across all nucleotides of the corresponding type of element for each gene). Outside the gene body, attribution magnitude decayed with distance from the gene (FIG.24, providing examples of box plots of average per- nucleotide attribution scores within different distance ranges from the gene body, for all test set genes; within each distance range, nucleotides within annotated cis-regulatory elements (CREs) are separated from other nucleotides; values plotted represent the absolute value of the input x gradient score, averaged across all nucleotides in the corresponding distance range from each gene). However, the model assigned higher attributions to nucleotides within annotated cis-regulatory elements (ENCODE Project Consortium et al. (2020)), even at elements far away (> 100 kb) from the predicted gene. This indicates that the model has learned to distinguish various classes of regulatory sequence elements, and uses these to inform its predictions.

[0338] To investigate whether the model could identify sequence features driving cell type- specific gene expression, its attributions in the genomic interval surrounding the previously mentioned FABP1 gene were examined (FIG. 18). In both enterocytes and hepatocytes, the model’s attributions for this gene highlighted several regions corresponding to accessible chromatin (K. Zhang et al. (2021)), including regions as far as > 50 kb from the transcription start site (FIG.25, Upper: pseudobulk scATAC-seq coverage (K. Zhang et al. (2021)) over a 64 kb genomic region surrounding the FABP1 gene; coverage tracks and peaks are shown for enterocytes (red) and hepatocytes (blue). Lower: smoothed per-nucleotide attribution scores for predicted FABP1 expression over the same genomic region in enterocytes (red) and hepatocytes (blue); values plotted represent the absolute value of the Input x Gradient score, smoothed with a gaussian filter of width 11 bp. scATAC-seq peaks are highlighted).MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 70 of 126

[0339] Closer examination of regions with the highest attribution revealed known regulatory motifs, such as a TATA box in the FABP1 promoter, which had a positive contribution to FABP1 expression in both cell types (FIG.26, providing plots of per-nucleotide Input x Gradient attribution scores over the 50 nucleotides upstream of the FABP1 transcription start site, in enterocytes (upper) and hepatocytes (lower); a TATA box is highlighted in yellow). In a distal CRE over 50 kb from the TSS which is accessible in both cell types (FIG.25), the model’s attributions highlighted C / EBP, CDX, and RXR transcription factor (TF) binding motifs in enterocytes; however, only the C / EBP and RXR motifs showed high attributions in hepatocytes (FIG.27, showing examples of plots of per-nucleotide Input x Gradient attribution scores over a distal CRE (also highlighted in pink in FIG. 25) in enterocytes (upper) and hepatocytes (lower); significant matches to HOCOMOCO v12 motifs are highlighted in yellow). This may be explained by the elevated expression of the CEBPA and RXRA genes encoding these TFs in both enterocytes and hepatocytes, whereas the CDX2 gene shows elevated expression only in enterocytes (FIG.28, providing examples of box plots showing the measured expression of the CEBPA, CDX2 and RXRA genes in pseudobulks representing enterocytes and hepatocytes, compared to pseudobulks representing other cell types in the gut and liver), consistent with literature (Grainger et al. (2013)).

[0340] This example demonstrates that the new model (“Decima”) has learned not only general sequence features such as gene and exon / intron boundaries, but also motifs with cell type-specific regulatory activity, and that the model’s attributions can identify the sequence features driving cell type-specific expression patterns. The model can thus be used to accurately predict cell type-specific patterns of gene expression from genomic DNA sequence alone. Example 4 – Model Attributions Reveal Drivers of Cell Type Identity, Cell State, and Tissue Residency

[0341] Given that the model can predict the cell type specificity of individual genes, a determination of whether examining the sequence features with high attribution scores for all genes specific to a cell type would reveal consistent motifs for transcription factors underlying cell type identity was performed. The procedure for this analysis is illustrated in FIG. 29 (schematic showing the method used to identify transcription factor (TF) motifs driving cell type identity based on the model’s attributions). The predicted log fold change in gene expression between the cell type of interest (‘positive’) and other ‘negative’ cell types in theMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 71 of 126 same tissue was calculated. Genes with the highest cell type-specificity in the ‘positive’ cell type were identified, and for each of these genes, the Input x gradient method was used to compute the attribution of its predicted log fold change with respect to each nucleotide in the input DNA sequence. This score was referred to as the “differential attribution” for each nucleotide. TF-MoDISco (Shrikumar et al. (2018)) was then applied to cluster these differential attributions and identify motifs with consistently high contributions. The contribution of each motif with an aggregate contribution score representing its contribution to cell type- specific gene expression was quantified in the ‘positive’ cell type.

[0342] This analysis was performed on seven epithelial cell types in the human lung, calculating differential attributions for each cell type using the remaining six as ‘negative’ cell types. After clustering with TF-MoDISco and assigning specificity contribution scores to motifs, the motifs that had high contribution scores in only one of the seven cell types, or in a single lineage, were identified. The results revealed strong and specific contribution scores for motifs of known lineage-determining TFs. RFX motifs were repeatedly identified in ciliated cells (FIG.30), and the expression of RFX TFs was also elevated in ciliated cells relative to other lung cell types (FIG. 31). The top four cell type-specific motifs were: RFX motifs in ciliated cells (Piasecki et al. (2010)), TEAD motifs in Type I pneumocytes (Little et al. (2021)), p63 motifs in basal cells (Daniely et al. (2004)), and GATA6 in type I pneumocytes (Yang et al. (2002)). Two motifs were also identified which contribute specifically to the airway epithelial lineage of cells: SOX2 (Shiraishi et al. (2024)) and a motif corresponding to the GRHL family (Kersbergen et al. (2018)). In all cases, the role of the TF was further supported by the elevated expression of the TF itself in the corresponding cell type (FIG.32, providing examples of box plots showing the measured expression of the six TFs with the most cell type- or lineage- specific contribution weights across seven epithelial cell types in the human lung; box plots are colored based on the contribution weight for the corresponding motif in each cell type; for each of the six TFs, the corresponding motif identified by TF-MoDISco is shown on the right).

[0343] Since the model predicted the difference in gene expression between neurons and non-neuronal cells in the brain with high accuracy (FIG. 33, providing an example of a scatterplot showing the measured and predicted log fold change in expression of 1,811 test set genes in pseudobulks representing neurons vs. pseudobulks representing other cell types in theMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 72 of 126 brain), it was tested to determine whether it could identify motifs driving a distinct neuronal identity. TF-MoDISco on differential attribution scores between neuronal and non-neuronal cell types in the brain revealed as the top result a motif with strong negative contributions, i.e., the motif decreases the expression of a gene in neurons relative to non-neuronal cells. This corresponded to the MYT1L transcription factor binding motif (FIG.34, providing an example of box plots showing the measured expression of the MYT1L gene in different cell types in the brain, along with the MYT1L motif found by TF-MoDISco based on the model’s differential attributions (attributions of the difference in predicted expression between neurons and non- neurons)). Indeed, MYT1L is a pan-neuronal TF that establishes neuronal identity by repressing gene programs for other cell types (Mall et al. (2017)). In contrast, TF-MoDISco analysis performed on differential attribution scores between non-neurons and neurons identified a negative contribution for the REST motif, i.e., the REST motif decreases predicted gene expression in non-neurons relative to neurons (FIG.35, providing box plots showing the measured expression of the REST gene in different cell types in the brain, along with the REST motif found by TF-MoDISco based on the model’s differential attributions (attributions of the difference in predicted expression between non-neurons and neurons); in FIG. 34 and FIG. 35, cell types are colored by their functional categories; BBB = Blood-brain barrier). REST is known to repress neuronal genes in non-neuronal cell types (Qureshi et al. (2010)). These examples show that the model’s attributions can identify lineage-specific repressors in addition to activators.

[0344] Having found that the new model successfully predicts differences between cell types, specific cases were examined to test whether it can predict and interpret the finer differences between the same cell type in different tissues or states. For instance, the model predicted the differential expression of test set genes in cycling vs. resting regulatory T cells, with a Pearson correlation of 0.51 between measured and predicted log fold-changes in expression (cycling vs. non-cycling; FIG.36, providing an example of a scatter plot showing the measured and predicted log fold change in expression of 1,811 test set genes in cycling regulatory T cells (Treg cycling) vs. non-cycling regulatory T cells (Treg) in skin, along with the top two motifs identified by TF-MoDISco (FIG.37) based on differential attribution scores for genes upregulated in Treg cycling vs. Treg). For each gene upregulated in the cycling state, the differential attribution (Treg cycling vs. Treg) was computed for each input nucleotide. TF-MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 73 of 126 MoDISco was run on these differential attributions, and the top two results matched motifs for the cell cycle related E2F family of TFs (FIG.37). Similarly, the model predicted differences in the expression of test set genes between cardiac fibroblasts and fibroblasts in other tissues, with a Pearson correlation of 0.53 (FIG.38, providing an example of a scatterplot showing the measured and predicted log fold change in expression of 1,811 test set genes in cardiac fibroblasts vs. fibroblasts in other tissues), and differential attribution scores for genes upregulated in cardiac fibroblasts consistently highlighted motifs for GATA4 and GATA6, which are known signature TFs for cardiac fibroblasts (Eraslan et al. (2022); Masuda et al. (2023)), as well as TEAD1, an essential regulator of cardiac fibroblast activation (Song et al. (2024); Burgos Villar et al. (2022)). FIG. 39 illustrates the top two motifs identified by TF- MoDISco based on differential attribution scores for genes upregulated in cardiac fibroblasts. All TF-MoDISco motifs are accompanied by the most similar motif in the HOCOMOCO v12 database, identified using Tomtom (Gupta et al. (2007)).

[0345] This example demonstrates that the disclosed model can be used to identify transcription factor binding motifs that drive cell type identity (e.g., by examining sequence features with high attribution scores for all genes specific to a cell type). Example 5 – Predicting the Impact of Non-Coding Variants at Cell Type Resolution

[0346] One of the major use cases for sequence-to-expression models is predicting the impact of non-coding genetic variation. In particular, due to linkage disequilibrium, it is often difficult to identify the causal variant for a particular phenotype among several variants in a locus identified by genome-wide association studies (GWAS), and to link these variants to the causal genes that they influence. In addition, it is often difficult to identify the causal cell type or condition in which a variant acts to influence the phenotype. This makes the new model disclosed herein (“Decima”) particularly promising for variant effect prediction, owing to its success at predicting cell type- specific expression.

[0347] The model was first evaluated to determine whether it can prioritize variants that significantly impact gene expression in specific cell types. This was tested using a dataset of fine-mapped single-cell expression quantitative trait loci (sc-eQTLs) in PBMCs based on data from the OneK1K consortium (Yazar et al. (2022)) and reprocessed by the eQTL catalog (Kerimov et al. (2021)). All 984 high-confidence (PIP >= 0.9) fine-mapped sc-eQTLs for which the variant fell within Decima’s pre-defined input window for the target gene wereMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 74 of 126 extracted, and each was matched one to 20 negative control variants. The model was used to predict variant effect scores (log fold change in gene expression caused by the variant) for all sc-eQTLs and controls in the matched cell type (FIG. 40, providing a schematic illustration showing the procedure for variant effect prediction using the disclosed machine learning model (“Decima”). In each cell type, the model’s scores were tested to determine whether they distinguished the fine-mapped sc-eQTLs in that cell type from matched controls (FIG. 41, providing an example of a bar plot showing the model’s performance in sc-eQTL classification in each of 21 blood cell types in the OneK1K dataset; performance was measured by the Area Under the Precision Recall Curve (AUPRC) for classification of sc-eQTLs vs. matched control variants in each cell type). Since the training set of the original Borzoi model also included RNA-seq data from FACS-sorted PBMCs, this analysis was repeated using Borzoi and it was found that the new model’s performance exceeded that of Borzoi in 19 of 21 cell types as well as in the whole dataset. It also exceeded the performance of a baseline metric (distance of the variant from the TSS).

[0348] Next, the model was assessed to determine whether it could correctly identify the direction of a variant’s effect on gene expression. For high-confidence (PIP > 0.9) sc-eQTLs, a significant positive correlation (Pearson’s rho = 0.42; p-value = 3 x 10-43) between the eQTL beta value and the model’s predicted variant effect in the corresponding cell type was found (FIG.42, providing an example of a scatter plot showing the model’s predicted effect size for high-confidence sc-eQTLs compared to their measured effect size (beta value) in the same cell type). Limiting this analysis to the sc-eQTL variants for which the model predicted any effect (predicted log fold change in expression > 0.01), the correlation coefficient increased to 0.58 (p-value = 2 x 10-41), with the direction of effect being predicted correctly in 87% of cases. Thus, for sc-eQTLs that were correctly identified by the new model, the model also correctly predicted the direction of effect in most cases.

[0349] In contrast to previous benchmarks of deep learning-based variant effect prediction that have focused on bulk eQTLs, our sc-eQTL benchmark offers the opportunity to evaluate whether the model can not only distinguish functional variants but also identify the causal cell type for a non-coding variant. Focusing on 513 variants that were found to be high-confidence sc-eQTLs in only a subset of the 21 cell types in the OneK1K study, it was found that the model’s predicted variant effect scores were higher in the sc-eQTL matched cell type comparedMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 75 of 126 to other cell types, even compared to the subset of other cell types in which the target gene is expressed (Wilcoxon test p-value = 1.1 x 10-7; FIG.43, providing an example of the model’s predicted absolute effect size for high-confidence sc-eQTL variants in cell types where the variant is identified as an sc-eQTL with PIP > 0.9, compared to other cell types in which the variant had no significant effect (PIP < 0.1) but its corresponding eGene is still expressed).

[0350] For instance, in the rs2158799 variant, which is a distal (~57 kb upstream of the TSS) sc-eQTL for JAZF1 in monocytes, the model correctly predicted that the strongest effect of the variant on this gene occurs in monocytes (FIG.44, providing an example of the model’s predicted effect sizes for a monocyte-specific sc-eQTL in various blood cell types). This effect was far stronger than that of the 20 matched negative control variants. The model’s attributions for the surrounding sequence suggest that the variant overlaps with a distal regulatory element that contributes strongly to JAZF1 expression in monocytes (FIG.45, upper panel, providing an example of the model’s attributions for JAZF1 expression in the sequences containing the reference and alternate allele, showing that the variant (highlighted in blue) reduces the attribution over a distal enhancer; values plotted represent the absolute input x gradient scores smoothed with a gaussian filter of width 11 bp) but not in other cell types. The attributions specifically highlight a C / EBP motif which is disrupted by the alternate allele (FIG.45, lower panel, providing an example of nucleotide-resolution Input x Gradient attribution scores computed with respect to predicted expression of JAZF1 in monocytes for the 160 bp sequence surrounding the reference and alternate alleles of the variant in FIG. 44 and FIG. 45, upper panel), highlighting a C / EBP motif disrupted by the variant). C / EBP factors are known master regulators of the myeloid lineage including monocytes (Rosenbauer and Tenen (2007)), suggesting that this variant acts by disrupting the accessibility of a cell type-specific enhancer. FIG.46 provides a non-limiting example of the measured expression of the CEBPA gene over the same cell types shown in the upper panel, highlighting its elevated expression in monocytes.

[0351] Given the model’s success at identifying eQTLs and predicting their impact at cell type-resolution, the model was evaluated to determine whether it could also be useful in interpreting GWAS (Genome-Wide Association Study) variants. This is a significantly harder task for several reasons: (1) only a subset of GWAS variants alter gene expression; (2) GWAS variants may have weaker effects on expression than high-confidence eQTLs; (3) for GWASMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 76 of 126 variants, the causal gene or cell type(s) was unknown; (4) GWAS variants may impact expression in cell states that are not represented in the dataset.

[0352] 837 high-confidence (PIP > 0.9, p-value < 10-6) fine-mapped GWAS SNPs were selected for 39 phenotypic traits. Using the model, the absolute log fold change in gene expression caused by each variant was predicted, on all genes for which the variant fell within the model’s pre-defined input window. Since the causal gene is unknown, each variant was matched to the gene where it has the strongest predicted impact on expression according to the model. Each GWAS variant was further matched to 10 negative control variants, and their impact on expression of the matched gene was predicted using the same procedure.

[0353] Initially, GWAS variants were compared to controls by assigning each variant a score corresponding to the average predicted absolute log fold change in gene expression caused by the variant, across all cell types. Overall, the model predicted that GWAS variants caused significantly larger changes in expression than their matched controls (FIG. 47, providing examples of box plots showing the distribution of the absolute log fold changes predicted by the model for 837 high-confidence GWAS variants and 8370 matched negative control variants). Based on these scores, GWAS variants could be classified with an AUPRC of 0.23, outperforming the baseline metric of distance from the variant to the gene TSS (FIG. 48, AUPRC 0.14). 40% of GWAS variants had significantly higher scores (z-test p-value < 0.05) than their matched controls, and 35% had higher scores than all 10 of their matched control variants.

[0354] It is expected that only a subset of GWAS variants that alter gene expression would be detectable in eQTL studies (Mostafavi et al. (2023)), highlighting the necessity of finding alternate methods to predict variant impact on gene expression. Indeed, only 95 of the 837 selected GWAS variants were also high-confidence eQTLs (11%). As expected, the model performed especially well at classifying GWAS variants that were also eQTLs, successfully assigning 62% of them significantly higher scores than their matched negatives. This represents a significant enrichment compared to the overall success rate of 39% (Fisher’s exact test p- value 10-5) and increases our confidence in the model’s ability to identify expression-altering GWAS variants. Identification of GWAS variants and controls was also attempted using RNA or CAGE predictions from Borzoi, and it was found that, while both models performed similarly, the new model outperformed Borzoi at this task.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 77 of 126

[0355] A significant advantage of the presently disclosed model over bulk-trained models is the possibility of linking variants to phenotypes via cell type-specific effects. Taking 334 variants which were successfully distinguished from their matched negative variants (z-test p- value < 0.01) by the model, the distribution of their predicted effect on expression across cell types was examined. For each variant, the mean effect (absolute log fold change in expression) of its 10 matched negatives was subtracted, and the resulting differences were then z-scored across cell types. The distribution of these z-scores across cell types was examined for variants matched to different higher-level phenotypic categories (FIG.49, providing an example of a heat map showing the model’s predicted log fold change caused by GWAS variants for 6 traits or categories (x-axis) in selected cell types (y-axis); for each GWAS variant, the average predicted log fold change for all 10 matched negative variants was subtracted from the predicted log fold change for the GWAS variant; the resulting background-subtracted log fold changes across 201 cell types were converted into z-scores; values in the heatmap are the average z-score for GWAS variants matched to the trait or category; only GWAS variants that could be distinguished from their matched controls based on the model’s predictions were included). Consistent with previous studies, the model predicted that, on average, variants associated with autoimmune and inflammatory diseases (asthma, eczema, autoimmune disease, Crohn’s disease, inflammatory bowel disease, multiple sclerosis, or lupus) had the strongest effect in T and NK cells (M. J. Zhang et al. (2022)). In contrast, variants associated with blood cell related traits (mean corpuscular hemoglobin, RBC count, RBC width, and platelet count) had the strongest predicted effect in erythroid progenitors and megakaryocytes, the precursors of RBCs and platelets respectively. Variants associated with height had the strongest predicted effect in fibroblasts and vascular cells, in line with previous analyses (K. Zhang et al. (2021)). Variants associated with blood triglyceride levels were predicted to most strongly alter gene expression in enterocytes and hepatocytes, which are responsible for the absorption and synthesis of triglycerides respectively (Alves-Bezerra and Cohen (2017)). Variants associated with neurological and psychiatric disorders (Alzheimer’s disease, Schizophrenia, Bipolar disorder, or neuroticism) were predicted to have their strongest impact in oligodendrocytes, oligodendrocyte precursors, and neurons. Finally, variants associated with respiratory conditions had the strongest predicted effect in lymphoid cells as well as lung epithelial cells such as secretory cells and club cells.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 78 of 126

[0356] Specific variants were examined to determine whether the model could additionally suggest mechanistic hypotheses underlying functional variants. For example, the rs138682554 variant is associated with hypertension and is also an eQTL in macrophages and monocytes, which are critical for the progression of hypertension (Wenzel (2018)). Consistent with the eQTL data, the model predicts that this variant increases expression of the FES gene, particularly in macrophages, monocytes and dendritic cells, and the attributions in these cell types reveal that the variant creates a potential binding site for the myeloid-specific TF SPI1 (FIG. 50, providing an example of a scatter plot showing the measured gene expression (log(CPM)+1) for the FES gene across cell types (x-axis) and the model’s predicted absolute log fold change in this gene’s expression due to the rs138682554 variant (y-axis); the values on the y-axis represent background-subtracted log fold changes, i.e., the mean absolute log fold change for 10 matched negative control variants was subtracted; cell types where the variant is predicted to have the strongest impact are highlighted; the model’s attributions for the sequence surrounding the variant in the highlighted cell types are shown along with the matched motif from HOCOMOCO v12 in FIG. 51). Another example is the rs8105903 variant which is associated with height and is also an eQTL in fibroblasts and monocytes. Consistent with the eQTL data, the model predicts that this variant increases expression of SLC1A5, particularly in fibroblasts, monocytes, and macrophages, and the attributions in these cell types reveal that the variant creates a potential binding site for the repressive TF ZEB2 (FIG.52, providing an example of a scatter plot showing the measured gene expression (log(CPM)+1) for the SLC1A5 gene across cell types (x-axis) and the model’s predicted absolute log fold change (background- subtracted) in this gene’s expression due to the rs8105903 variant (y-axis); cell types where the variant is predicted to have the strongest impact are highlighted; the model’s attributions for the sequence surrounding the variant in the highlighted cell types are shown along with the matched motif from HOCOMOCO v12 in FIG.53).

[0357] Given the model’s success at predicting cell types consistent with disease biology and eQTL studies, the model can be used to propose cell type-specific mechanisms for the large fraction of GWAS variants that have no associated cell type-specific eQTL. For example, the model predicted that the rs79755767 variant, associated with RBC width and platelet count, has its strongest effect on the NFE2 gene in megakaryocytes, erythroid progenitors, and hematopoietic stem cells (FIG. 54, providing an example of a scatter plot showing theMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 79 of 126 measured gene expression (log(CPM)+1) for the NFE2 gene across cell types (x-axis) and the model’s predicted absolute log fold change (background-subtracted) in this gene’s expression due to the rs79755767 variant (y-axis); cell types where the variant is predicted to have the strongest impact are highlighted; the model’s attributions for the sequence surrounding the variant in the highlighted cell types are shown along with the matched motif from HOCOMOCO v12 in FIG. 55). This gene encodes a major regulator of erythroid and megakaryocytic gene expression, and the model’s attributions suggest that the variant disrupts a potential RUNX2 motif (FIG.55).

[0358] These results demonstrate that the model can identify a subset of non-coding GWAS variants that act by mediating significant expression changes, including in contexts that may not be captured by eQTL studies. Further, in the cases where it is successful, it can be a valuable tool for revealing the underlying cell type-specific mechanisms of action. It is worth noting that the model provides a prediction of cell type-specific variant impact that is distinct from the expression level of the target gene. For example, the FES, SLC1A5 and NFE2 genes are all highly expressed in several unrelated cell types, in which Decima predicts only weak effects on gene expression (FIG.50 – FIG.55).

[0359] This example demonstrates that the disclosed model (“Decima”) can be used for accurate variant effect prediction. Furthermore, as described above, the model can be used to suggest and / or identify cell type-specific mechanisms that underlie variant-driven functional effects (e.g., through the creation of potential transcription factor binding sites that impact gene expression). The model can be used to identify a subset of non-coding GWAS variants that act by mediating significant expression changes (including in contexts that may not be captured by eQTL studies), and to determine the underlying cell type-specific mechanisms of action. Example 6 – Identification of Sequence Determinants of Disease-Specific Cell States

[0360] FIG. 56 provides a non-limiting schematic illustration of the use of the disclosed machine learning model (configured to predict cell type-specific and / or cell state-specific gene expression levels) to predict disease vs. healthy differential gene expression. Annotated single cell RNA seq atlases were used to provide ground-truth data for a gene expression matrix, which in turn was used to measure a true value for the log fold change in gene expression between disease and healthy cell states for a specified cell type. The ground truth values for log fold change in gene expression between disease and healthy cell states for a specified cellMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 80 of 126 type could then be compared to corresponding prediction made by the long-context, DNA sequence-based gene expression prediction model (i.e., the “Decima” model) described herein. FIG.57 provides a nonlimiting example of a scatter plot showing the measured and predicted log fold changes for a test gene set between matched healthy and disease pseudobulk samples for fibroblast cells and a Crohn’s Disease pseudobulk sample.

[0361] The model’s training dataset contains pseudobulks representing 82 disease conditions. Overall, the variation in gene expression between diseased and healthy states of the same cell type was much smaller than the variation between healthy cell types. Nevertheless, the model was tested to determine if it can predict differences in gene expression between matched healthy and disease samples for the same cell type. Pairs of matched healthy and disease pseudobulks for the same cell type in the same tissue and study were selected, and the measured and predicted log fold changes in gene expression (disease vs. healthy) for each pair was computed (FIG.58, providing a schematic illustration showing the method used to identify both shared and disease-specific driver TFs for various cell types and diseases based on the model’s attributions). Across all 565 selected disease / healthy pairs, an average Pearson correlation of 0.24 was found between the measured and predicted log fold changes for test set genes (FIG. 59, providing an example of a histogram of Pearson correlations between measured and predicted log fold changes between matched healthy and disease pseudobulks for the same cell type, calculated over 1,811 test set genes).

[0362] For example, the model predicts the change in expression of test set genes between fibroblasts in the ileum of Crohn’s disease patients compared to fibroblasts in the healthy ileum, with a Pearson correlation of 0.52 (see FIG. 60A, FIG. 60B, FIG. 60C, and FIG. 60D, providing examples of scatter plots showing the measured and predicted log fold changes for 1,811 test genes between matched healthy and disease pseudobulks for four cell type / disease combinations). The model correctly predicted the direction of effect for genes that are upregulated (log fold change > 1) or downregulated (log fold change < -1) by Crohn’s disease in fibroblasts, both in cases where the gene is affected by disease specifically in fibroblasts, and in cases where the expression of the gene is altered by disease across multiple cell types in the ileum (see FIG.60E, FIG.60F, FIG.60G, and FIG.60H, providing examples of boxplots showing the model’s predicted log fold changes for genes that are upregulated or downregulated (measured absolute log fold change >= 1) in disease; genes were split byMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 81 of 126 whether they are affected by the disease (measured absolute log fold change >= 1) specifically in the cell type of interest, or in multiple cell types; asterisks indicate groups in which the mean predicted log fold change is significantly higher or lower than 0 (Wilcoxon test p-value < 0.05)). Similarly, the model was shown to predict the change in expression of test set genes between enterocytes in the ileum of Crohn’s disease patients compared to the healthy ileum, fibroblasts in the lung of Scleroderma patients and fibroblasts in healthy lungs, and between type II pneumocytes in the lungs of Scleroderma patients and type II pneumocytes in healthy lungs (FIG.60A – FIG.60H).

[0363] Given these successful predictions, it follows that the model can potentially be a powerful tool to identify sequence features and motifs that drive cell type-specific disease responses. Accordingly, the model was evaluated to determine if it can identify shared factors contributing to multiple diseases. To identify TFs that potentially drive differences in gene expression between matched disease and healthy cell populations, differential attribution analysis of genes predicted to be upregulated in disease was performed, followed by TF- MoDISco clustering to identify motifs with high contributions (FIG. 58). In many of the diseases analyzed (specifically scleroderma, atopic dermatitis, Crohn’s disease, idiopathic pulmonary fibrosis, ulcerative colitis, and myocardial infarction), this procedure highlighted motifs corresponding to TF families known to regulate inflammatory pathways, such as JUN / FOS, STAT / IRF and C / EBP (see FIG.60I, providing examples of motifs found by TF- MoDISco based on the model’s differential attribution scores (disease vs. matched healthy pseudobulks) across multiple cell types and diseases, including all four cases shown in FIG. 60A – FIG. 60H), representing a shared inflammatory signature; FIG. 60J, providing an example of a dimeric motif identified by TF-MoDISco in attributions for fibroblasts in multiple fibrotic diseases, including both fibroblast examples shown in FIGS. 60A-B; FIG. 60K, providing an example of measured log fold change (disease vs. healthy) in TWIST1 expression in various diseases and cell types, separated by whether the TWIST1 dimeric motif is found by TF-MoDISco based on the disease vs. healthy differential attribution scores in that disease and cell type; and FIG.60L, providing an example of a NF-kB motif identified by TF-MoDISco in attributions for fibroblasts in Crohn’s Disease). In contrast, the contribution of these motifs was lower in chronic kidney disease, which is expected to have a weaker inflammatory contribution.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 82 of 126

[0364] Beyond this general inflammatory signature, a dimeric motif (FIG. 60L) was observed in fibroblasts - but not other cell types - particularly in fibrotic diseases, including Crohn’s disease, idiopathic pulmonary fibrosis, lung scleroderma, as well as COVID-19 in the lung parenchyma. This dimeric motif exhibits a very high-fidelity match to a ChIP-seq motif previously reported for the TF TWIST1 (TWST1.H12CORE.0.P.B) and represents the cooperative binding of TWIST1, together with a homeodomain factor (Kim et al. (2024)). A significant association between the presence of this motif in attributions and the upregulation of TWIST1 gene expression in fibroblasts in specific diseases was observed (Mann-Whitney U-test p-value=7 x 10-4; FIG.60K). It was previously reported that TWIST1, together with the homeodomain factor PRRX1 which may form the partner in the dimeric motif, controls fibroblast activation in wound healing and in pathological fibrotic conditions (Yeo et al. (2018)). Finally, an NFκB motif in fibroblasts specifically in Crohn’s disease was found (FIG. 60L), consistent with reports of NFκB activation in the gut of patients with inflammatory bowel diseases, particularly Crohn’s Disease (Schreiber et al. (1998); Han et al. (2017)).

[0365] Our analyses demonstrate that the model has learned known features of disease biology and can potentially be applied to interpret scRNA-seq data from disease conditions.

[0366] This example demonstrates that the disclosed model can be used to identify genomic DNA sequence features and motifs that drive cell type-specific disease responses. Example 7 – Design of Cell Type- and Disease-Biased Regulatory Element

[0367] Recent advances in genomic sequence modeling have demonstrated that trained oracle models can be used to generate sequences that drive tissue-specific expression (Taskiran et al. (2024); de Almeida et al. (2024); Gosai et al. (2023); Lal et al. (2024)). These model- designed sequences have been shown to drive specific and experimentally validated expression in human cell lines, as well as across various tissues in Drosophila. However, a major gap remains – current approaches lack the resolution to tune gene expression at the cell-type level, which is essential for applications such as Adeno-Associated Virus (AAV)-mediated gene therapy (D. Wang et al. (2019)). Delivering vectors that target specific cell types in diseased tissues while avoiding healthy cells could also open new therapeutic avenues.

[0368] The application of the model to evolve regulatory elements that are both cell-type and disease-enriched was explored. As an example, fibroblasts were targeted, which are implicated in the inflammatory response in Crohn’s disease. Fibroblasts contribute to tissueMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 83 of 126 remodeling, fibrosis, and the inflammatory response in Crohn’s disease, making them a relevant target for treatment. A synthetic construct composed of a random 200bp starting sequence paired with a cargo gene sequence was first generated to score expression (FIG.61, providing a schematic illustration showing promoter design through directed evolution with the Decima model). In order to predict the potency of the regulatory sequence with the model, this construct was inserted into a genomic safe harbor locus in silico. The starting sequence was predicted to have no expression in any cell type, providing a starting point for the design process (FIG. 62, providing an example of predicted expression of the cargo gene across healthy and diseased fibroblast and non-fibroblast cells over 100 rounds of directed evolution).

[0369] To achieve cell type-specific expression, a regulatory element was designed to drive expression specifically in fibroblasts relative to other gut associated cell types such as endothelial cells, macrophages, B cells, and T cells. 50 rounds of directed evolution was performed, maximizing expression in fibroblasts while minimizing expression in the outgroup cell types. After performing directed evolution, the resulting sequence was predicted to drive robust expression specific to fibroblasts and contained a fibroblast-implicated TWIST1 motif, demonstrating the ability of the approach to generate regulatory elements with cell type- specific expression within a complex tissue environment. This regulatory element was then further evolved to differentiate between fibroblasts from Crohn’s disease patients and those from healthy individuals. An additional 50 rounds of directed evolution was performed, optimizing expression in Crohn’s-derived fibroblasts relative to healthy fibroblasts. The final evolved sequence demonstrated a 6.84-fold higher predicted expression in disease-associated fibroblasts compared to other gut cell types and a 4.31-fold higher predicted expression compared to healthy fibroblasts (FIG. 63, providing an example of predicted specificity of cargo gene expression in fibroblasts and disease fibroblasts, which were improved in design in rounds 0-50 for cell-type specificity and in rounds 50-100 for disease-state specificity).

[0370] To understand the mechanism underlying disease specificity, in silico mutagenesis (ISM) was applied to the final evolved sequence, measuring the sensitivity of the predicted cell type- and disease-enriched expression to sequence perturbations. Several motifs were predicted to reduce expression of the cargo gene when perturbed, specifically in Crohn’s-derived fibroblasts (FIG.64, providing an example of how in silico mutagenesis (ISM) of the synthetic regulatory element reveals key sequence motifs whose perturbation is predicted to uniquelyMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 84 of 126 affect expression of the cargo gene in fibroblasts). Among these motifs, two TBP motifs (TATA-binding protein) were identified, which are essential in transcription initiation (FIG. 65, FIG. 66; FIG. 67, providing examples of how applying ISM with respect to disease fibroblast expression identifies key motifs generated in the design process for the fibroblast (FIG.65) and disease-fibroblast (FIG.66) evolved sequences, including a TWIST1, C / EBP, and IRF motif, which are implicated in fibroblast-specific and immune-specific regulation. These motifs match the HOCOMOCO v12 motifs (FIG. 67). This suggests that the model designed an effective and standalone promoter (Ponjavic et al. (2006)). The TWIST1 motif was consistent in both fibroblast-specific and disease-fibroblast-specific evolved sequences, aligning with previous findings implicating this motif in fibroblast-specific expression. As the sequence was evolve toward disease-fibroblasts, the completion or installation of C / EBP, STAT3, CREB3, and IRF motifs were observed. Several of these motifs were previously identified as associated with inflammatory response, indicating that the model is implanting motifs it has observed to drive inflammatory and cell-type specific expression.

[0371] This study demonstrates the disclosed model’s capability of designing highly specific regulatory elements that can distinguish not only between cell-types, but also between disease states.

[0372] Discussion: Sequence-to-function models have become pivotal in understanding how the genome sequence is interpreted and executed in various cellular contexts. These models are essential for generating mechanistic hypotheses for non-coding variants and designing synthetic regulatory sequences. Although recent models have expanded their biological coverage to include various experiments and tissues, their reliance on bulk samples from healthy individuals and cell lines limits the investigation of regulatory circuits specific to cell types or diseases. Single-cell RNA sequencing (scRNA-seq), for which substantial datasets are already publicly available, can overcome these limitations by providing extensive coverage of diverse cell states and diseases, thereby enriching our understanding of complex regulatory rules.

[0373] However, the field of single-cell genomics faces challenges in inferring regulatory activity in individual cell types and states when chromatin accessibility data is unavailable. This limitation also significantly restricts the generation of hypotheses regarding the regulatory elements driving differential expression between cell types and cellular states. The disclosedMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 85 of 126 model (“Decima”) bridges this gap by applying the utility of sequence-to-function models to single-cell datasets, enabling regulatory interpretation in various cell states and diseases, sc- eQTL effect predictions, and the design of regulatory sequences with cell type- and disease- enriched activity. The findings presented herein demonstrate that the model can learn regulatory rules related to cell type specificity, tissue residency, and disease-dependent cell states by replicating known regulatory syntax from the literature. The model can be used to add an extra layer of regulatory insight to the results generated by current scRNA-seq analyses, thereby enhancing the comprehension of biological phenomena.

[0374] One caveat of the described approach is that experimental batch effects might introduce spurious regulatory sequences in the downstream analyses. Modeling regulatory logic in cell populations that exist in a continuum and are difficult to categorize into distinct cell types also poses a challenge in the current formulation, where cells belonging to the same discrete group were collapsed into a single pseudobulk. Further, as the current focus is primarily on predicting expression differences between cell types rather than between individuals, cells from different individuals in the same study were also collapsed into the same pseudobulk.

[0375] Caveats of most existing sequence-to-expression models, including those using bulk expression data, also hold here: by defining instances as the sequence surrounding a gene and its expression across pseudobulks, the model captures cis-regulatory mechanisms. Models that incorporate broader cellular state information, for example defined by some appropriate representation of the full transcriptome, or by subset of appropriately selected genes, would be useful extensions that could capture some trans-regulation effect. Moreover, the output space of this model is strictly defined by the cell-type, tissue, disease combinations in the training set and thus the model is not designed to transfer to unseen cellular contexts.

[0376] Additionally, there is relatively low concordance between the disease versus healthy changes in expression measured in multiple studies of the same disease, presenting further complications in modeling the regulatory biology of diseases. The predictability of disease differential expression, arguably the most difficult task in our study, depends heavily on the model's performance for the specific genes and cell types involved in the disease, as well as the regulatory nature of the disease itself. Nonetheless, for diseases with clear disease-specific or enriched cell states, identifying regulatory drivers is not markedly different from identifyingMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 86 of 126 cell type-specific regulatory elements, highlighting the ability of the disclosed model to tackle regulatory challenges in disease biology.

[0377] As future work, large atlases of predicted functional variants will be compiled, including rare variants not captured in eQTL studies, in the context of hundreds of cell types. This will allow us to explore the colocalization of these variants with GWAS variants to link disease-associated variants to specific cell types and to identify transcription factors and other regulatory mechanisms and pathways that drive cell type-specific changes in activity in the context of disease. Further, the approach can be extended to design more complex regulatory elements that activate therapeutic expression in diseased cells, but turn off when the cells return to a healthy state, opening up a dynamically disease-conditioned approach to gene therapy.

[0378] This example demonstrates that the disclosed model can be used to design effective regulatory elements that are cell type-specific and / or disease state-specific. Furthermore, as described above, the model can be used to investigate and / or determine the mechanism underlying cell type- and / or disease-state specificity, e.g., by measuring the sensitivity of the predicted cell type-enriched and / or disease state-enriched gene expression to sequence perturbations. Example 8 – A Computational Framework for Predicting Regulatory Element Activity

[0379] Understanding how regulatory elements behave across human cell types is a central challenge in functional genomics. Experimental assays such as massively parallel reporter assays (MPRAs) and STARR-seq provide quantitative measurements of regulatory element activity but are generally limited to a small number of cell lines. ATAC-seq and related epigenomic profiling methods offer broader tissue coverage but do not directly measure the activity of individual sequences and remain constrained by sample availability. Computational models trained on these assays typically inherit the same limitations, performing well only for the at most two or three cell types that they were trained on, and failing to generalize to new contexts. Critically, no existing method enables direct prediction of how a given element will function across the full range of human cell types, and experimental validation in primary cells is prohibitively costly or infeasible at scale.

[0380] Disclosed herein is “Universal MPRA”, a computational framework for zero-shot prediction of regulatory element activity in any cell type, including those where individual element function has never been measured, such as primary tissues and rare cell types currentlyMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 87 of 126 inaccessible to reporter assays. As illustrated in FIG.68, libraries of candidate promoters and enhancers may be evaluated by designing in silico constructs and inserting them into genomic scoring loci without endogenous activity. Scalar (e.g., the “Decima” model described elsewhere herein) and / or profile (e.g., the “Borzoi” model) sequence-to-expression models can be used predict regulatory activity without additional training (i.e., zero-shot). Activity across the construct is scored to evaluate activity and specificity of the regulatory sequence across cell types.

[0381] FIG. 69 provides a non-limiting schematic illustration of the use of long-context DNA sequence models to predict functional genomics properties. Long-context DNA sequence models (e.g., the Borzoi or Decima models, as discussed elsewhere herein) may be configured to accept long DNA sequences (e.g., having lengths of at least 1 kb, 10 kb, 100 kb, 1,000 kb, 10,000 kb, 100,000 kb, 200,000 kb, 300,000 kb, 400,000 kb, 500,000 kb, 600,000 kb, 700,000 kb, 800,000 kb, 900,000 kb, or 106kb) as input, and output a prediction of, e.g., gene expression level (scalar valued) and / or a functional genomic property profile (e.g., a functional genomic property such as transcription factor binding, chromatin accessibility, gene expression, or transcription initiation, as obtained from sequencing-based functional genomics assays, as a function genomic locus).

[0382] FIG.70 provides a non-limiting schematic illustration of the use of a long-context DNA sequence model to process an input construct comprising, e.g., a candidate promoter sequence (of about 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, or 800 bp in length) and cargo gene sequence (the constructs generated in silico for one or more specified genomic insertion sites), and score the efficacy of the promoter sequence for gene expression based on one or more prediction functional genomics property profiles. In some instances, determining a proper metric for scoring candidate regulatory element (e.g., candidate promoter sequence) efficacy may be challenging.

[0383] FIG. 71 provides a non-limiting schematic illustration of a computational framework (“Element Scorer”) for using one or more long-context DNA models (e.g., 1, 2, 3, 4, 5, 6, 7, 8,9, or 10 long-context DNA models) to process a plurality of construct sequences (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, 100, or more than 100 construct sequences generated in silico for a plurality of different candidate genomic insertion sites (e.g., e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, 100, or more than 100 candidate genomic insertionMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 88 of 126 sites) comprising, e.g., a plurality of candidate promoter sequences and a cargo gene sequence inserted at the plurality of different candidate genomic insertion sites, and score the combinations of candidate promoter sequences, cargo gene sequence, and candidate genomic insertion sites for efficacy based on their predicted gene expression levels and / or one or more predicted genomic property profiles. Scoring of candidate regulatory element (e.g., candidate promoter sequence) activity and specificity can be performed by processing a corresponding predicted gene expression value (e.g., a scalar gene expression value) and / or the one or more predicted functional genomics property profiles over a specified genomic window corresponding to, e.g., the cargo gene sequence. Scoring can be performed using any of a variety of metrics, e.g., integration of predicted genomic functional property level over the specified genomic window, averaging of predicted functional genomic property level over the specified genomic window, determination of a maximum value of the predicted functional genomic property level over the specified genomic window, etc.

[0384] This approach (also referred to herein as “Universal MPRA”) builds on the observation that long-context DNA-to-expression models, trained across millions of endogenous loci, learn cell-type-specific regulatory grammar directly from sequence. Because these models are exposed to a wide range of genomic contexts during training, they implicitly capture logic governing cell-type specific enhancer, promoter, and chromatin interactions. Universal MPRA leverages this capacity to evaluate candidate elements through a flexible and systematic pipeline comprising: • Construction of in silico sequence constructs that combine regulatory elements with a cargo gene, • Insertion of the construct into scoring loci lacking endogenous activity, • Prediction of gene expression using long-context DNA-to-expression models, and • Scoring predicted activity across the construct to assess activity and specificity. This framework is modular and highly configurable, allowing users to substitute models, construct components, and design parameters.

[0385] Through careful benchmarking of scoring strategies, including scoring insertion locus and cargo gene selection, the use of normalization methods, and the use of ensemble approaches, long-context models can be demonstrated to accurately predict the behavior of individual regulatory elements.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 89 of 126

[0386] FIG.72 provides a non-limiting example of predicted gene expression data used to score a designed promoter and cargo gene sequence inserted at a specified genomic integration site during a directed evolution process, as described elsewhere herein. The iterative directed evolution process may comprise the use of in silico mutagenesis and one or more long-context DNA sequence prediction models (e.g., models configured to accept one or more input DNA construct sequences comprising a cargo gene sequence and flanking sequence regions that range in length from about 100 kb to about 1 Mb) and output a scalar value or profile prediction of, e.g., gene expression or a functional genomic property profile (e.g., ChIP-seq, ATAC-seq, RNA-seq, and / or CAGE-seq data) when the construct is inserted at one or more specified genomic insertion sites. The predicted functional genomic property profile data must then be scored to assess regulatory element (e.g., promoter sequence) activity. FIG. 72 shows an example of a plot of predicted RNA level (using the Borzoi model) as a function of genomic position for a construct comprising a SYN1 promoter sequence (2670 bp) and the EBFP cargo gene sequence inserted at a specified genomic locus (e.g., a safe-harbor or “gene desert” locus). The plot shows predicted RNA levels for both on-target expression and off-target expression. Scoring of regulatory element (e.g., promoter sequence) activity and specificity can be performed by processing the predicted RNA levels over a specified genomic window corresponding to, e.g., the cargo gene sequence. Scoring can be performed using any of a variety of metrics, e.g., integration of predicted RNA level over the specified genomic window, averaging of predicted RNA level over the specified genomic window, determination of a maximum value of predicted RNA level over the specified genomic window, etc.

[0387] FIG. 73 provides an example of data illustrating the distribution of designed promoter efficacy scores for an hSyn insertion into prescreened responsive safe-harbor sites versus insertion at other test sequence sites.

[0388] Universal MPRA achieves high concordance with MPRA measurements of promoters in HepG2 and K562 cells, and crucially, extends predictions to tissues and cell types that cannot be assayed directly (FIG.74, providing an example of predicted activities from the Decima model that show significant correlation with measured MPRA results in HepG2 (ρ = 0.79) cells, approaching experimental replicate concordance (ρ = 0.92); FIG.75, providing an example of predicted activities from the Decima model that show significant correlation with measured MPRA results in K562 (ρ = 0.78) cells, approaching experimental replicateMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 90 of 126 concordance (ρ = 0.76)). When paired with a profile-level model like Borzoi, neuron-specific promoter hSyn shows strong brain-specific activity across the construct, with predicted CAGE signal peaking at the promoter and RNA signal extending over the cargo gene (FIG. 76, providing an example of data for predicted activity of the brain-specific hSyn promoter (448bp) across RNA-seq, CAGE-seq, and ATAC-seq profiles; blue profiles show brain or neuronal activity, while grey denotes mean activity across all other tissues). In contrast, predicted ATAC-seq is less tissue-specific, highlighting its limited ability to resolve promoter activity. A single-cell RNA-seq pseudobulk model like Decima further reveals high expression exclusively in neurons but reduced activity in other brain cell types, such as glia and fibroblasts (FIG.77, providing an example of data for cell-type resolution hSyn predictions within brain tissue). Extending this analysis to a broader set of commonly used promoters reveals unexpected tissue biases: while promoters like hSyn and TNT show the predictable brain- and heart-specific activity, respectively, widely used "ubiquitous" promoters such as CAG and CMV exhibit substantial tissue preference (FIG. 78, providing an example of data for evaluation of constitutive promoters that reveals unexpected tissue biases; OPC = oligodendrocyte precursor; CT = corticothalamic; IT = intratelencephalic).

[0389] Together, these results demonstrate that genome-wide models can serve as general- purpose surrogates for cell-type-specific regulatory assays. By enabling flexible and model- agnostic analysis of regulatory constructs across primary tissues and rare cell types, Universal MPRA enables element selection, activity assessment, and de novo design, even in cell types where experimental measurement remains effectively impossible.

[0390] This example demonstrates that the disclosed methods and models can be used to predict regulatory element activity in any cell type. VIII. Example Computer System

[0391] FIG. 79 is a block diagram of a computer system 7900, in accordance with some embodiments. Computer system 7900 can be an example of one implementation for computing platform 102 described above in FIG.1.

[0392] FIG. 79 illustrates an example computing system 7900 that may be utilized to implement any of the methods and / or machine learning models described herein. In certain embodiments, the computing system 7900 may perform one or more steps of one or more methods described or illustrated herein. In certain embodiments, the computing system 7900MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 91 of 126 provide functionality described or illustrated herein. In certain embodiments, software running on the computing system 7900 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Certain embodiments include one or more portions of the computing systems 7900. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Moreover, reference to a computer system may encompass one or more computer systems, where appropriate.

[0393] This disclosure contemplates any suitable number of computing systems 7900. This disclosure contemplates computing system 7900 taking any suitable physical form. As example and not by way of limitation, computing system 7900 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computing system 7900 may include one or more computing systems 7900; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks.

[0394] Where appropriate, the computing system 7900 may perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example, and not by way of limitation, the computing system 7900 may perform in real time or in batch mode one or more steps of one or more methods described or illustrated herein. The computing system 7900 may perform at different times or at different locations one or more steps of one or more methods described or illustrated herein, where appropriate.

[0395] In certain embodiments, the computing system 7900 includes a processor 7902, memory 7904, storage 7906, an input / output (I / O) interface 7908, a communication interface 7910, and a bus 7912. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement. In certain embodiments, processor 7902 includes hardware for executing instructions, such as those making up a computer program. AsMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 92 of 126 an example, and not by way of limitation, to execute instructions, processor 7902 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 7904, or storage 7906; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 7904, or storage 7906. In certain embodiments, processor 7902 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 7902 including any suitable number of any suitable internal caches, where appropriate. As an example, and not by way of limitation, processor 7902 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches may be copies of instructions in memory 7904 or storage 7906, and the instruction caches may speed up retrieval of those instructions by processor 7902.

[0396] Data in the data caches may be copies of data in memory 7904 or storage 7906 for instructions executing at processor 7902 to operate on; the results of previous instructions executed at processor 7902 for access by subsequent instructions executing at processor 7902 or for writing to memory 7904 or storage 7906; or other suitable data. The data caches may speed up read or write operations by processor 7902. The TLBs may speed up virtual-address translation for processor 7902. In certain embodiments, processor 7902 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 7902 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 7902 may include one or more arithmetic logic units (ALUs); be a multi-core processor; or include one or more processors 7902. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0397] In certain embodiments, memory 7904 includes main memory for storing instructions for processor 7902 to execute or data for processor 7902 to operate on. As an example, and not by way of limitation, the computing system 7900 may load instructions from storage 7906 or another source (such as, for example, another computing system 7900) to memory 7904. Processor 7902 may then load the instructions from memory 7904 to an internal register or internal cache. To execute the instructions, processor 7902 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 7902 may write one or more results (which may beMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 93 of 126 intermediate or final results) to the internal register or internal cache. Processor 7902 may then write one or more of those results to memory 7904.

[0398] In certain embodiments, processor 7902 executes only instructions in one or more internal registers or internal caches or in memory 7904 (as opposed to storage 7906 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 7904 (as opposed to storage 7906 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple processor 7902 to memory 7904. Bus 7912 may include one or more memory buses, as described below. In certain embodiments, one or more memory management units (MMUs) reside between processor 7902 and memory 7904 and facilitate accesses to memory 7904 requested by processor 7902. In certain embodiments, memory 7904 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Moreover, where appropriate, this RAM may be single- ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 7904 may include one or more memory devices 7904, where appropriate. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.

[0399] In certain embodiments, storage 7906 includes mass storage for data or instructions. As an example, and not by way of limitation, storage 7906 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Storage 7906 may include removable or non-removable (or fixed) media, where appropriate. Storage 7906 may be internal or external to the computing system 7900, where appropriate. In certain embodiments, storage 7906 is non-volatile, solid-state memory. In certain embodiments, storage 7906 includes read-only memory (ROM). Where appropriate, this ROM may be mask- programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 7906 taking any suitable physical form. Storage 7906 may include one or more storage control units facilitating communication between processor 7902 and storage 7906, where appropriate. Where appropriate, storage 7906 may include one or more storages 7906. Although thisMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 94 of 126 disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.

[0400] In certain embodiments, I / O interface 7908 includes hardware, software, or both, providing one or more interfaces for communication between the computing system 7900 and one or more I / O devices. The computing system 7900 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and the computing system 7900. As an example, and not by way of limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device or a combination of two or more of these. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interfaces 7908 for them. Where appropriate, I / O interface 7908 may include one or more device or software drivers enabling processor 7902 to drive one or more of these I / O devices. I / O interface 7908 may include one or more I / O interfaces 7908, where appropriate. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.

[0401] In certain embodiments, communication interface 7910 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between the computing system 7900 and one or more other computer systems 7900 or one or more networks. As an example, and not by way of limitation, communication interface 7910 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 7910 for it.

[0402] As an example, and not by way of limitation, the computing system 7900 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, the computing system 7900 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTHMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 95 of 126 WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless network or a combination of two or more of these. The computing system 7900 may include any suitable communication interface 7910 for any of these networks, where appropriate. Communication interface 7910 may include one or more communication interfaces 7910, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0403] In certain embodiments, bus 7912 includes hardware, software, or both coupling components of the computing system 7900 to each other. As an example, and not by way of limitation, bus 7912 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Bus 7912 may include one or more buses 7912, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.

[0404] Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field- programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.

[0405] FIG. 80 illustrates a diagram 8000 of an example artificial intelligence (AI) architecture 8002 (which can be included as part of the one or more computing device(s) 7900MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 96 of 126 as discussed above with respect to FIG. 79) that can be utilized to determined one or more gene expression and / or functional genomics property predictions, in accordance with the disclosed embodiments. In certain embodiments, the AI architecture 8002 can be implemented utilizing, for example, one or more processing devices that may include hardware (e.g., a general purpose processor, a graphic processing unit (GPU), an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a microcontroller, a field-programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), and / or other processing device(s) that can be suitable for processing various molecular data and making one or more decisions based thereon), software (e.g., instructions running / executing on one or more processing devices), firmware (e.g., microcode), or some combination thereof.

[0406] In certain embodiments, as depicted by FIG. 80, the AI architecture 8002 may include machine learning (ML) algorithms and functions 8004, natural language processing (NLP) algorithms and functions 8006, expert systems 8008, computer-based vision algorithms and functions 80106910, speech recognition algorithms and functions 8012, planning algorithms and functions 8014, and robotics algorithms and functions 8016. In certain embodiments, the ML algorithms and functions 8004 may include any statistics-based algorithms that can be suitable for finding patterns across large amounts of data (e.g., “Big Data” such as genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data). For example, in certain embodiments, the ML algorithms and functions 8004 may include deep learning algorithms 8018, supervised learning algorithms 8020, and unsupervised learning algorithms 8022.

[0407] In certain embodiments, the deep learning algorithms 8018 may include any artificial neural networks (ANNs) that can be utilized to learn deep levels of representations and abstractions from large amounts of data. For example, the deep learning algorithms 8018 may include ANNs, such as a perceptron, a multilayer perceptron (MLP), an autoencoder (AE), a convolution neural network (CNN), a recurrent neural network (RNN), long short term memory (LSTM), a grated recurrent unit (GRU), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), aMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 97 of 126 generative adversarial network (GAN), and deep Q-networks, a neural autoregressive distribution estimation (NADE), an adversarial network (AN), attentional models (AM), a spiking neural network (SNN), deep reinforcement learning, and so forth.

[0408] In certain embodiments, the supervised learning algorithms 8020 may include any algorithms that can be utilized to apply, for example, what has been learned in the past to new data using labeled examples for predicting future events. For example, starting from the analysis of a known training data set, the supervised learning algorithms 8020 may produce an inferred function to make predictions about the output values. The supervised learning algorithms 8020 may also compare its output with the correct and intended output and find errors in order to modify the supervised learning algorithms 8020 accordingly. On the other hand, the unsupervised learning algorithms 8022 may include any algorithms that may applied, for example, when the data used to train the unsupervised learning algorithms 8022 are neither classified nor labeled. For example, the unsupervised learning algorithms 8022 may study and analyze how systems may infer a function to describe a hidden structure from unlabeled data.

[0409] In certain embodiments, the NLP algorithms and functions 8006 may include any algorithms or functions that can be suitable for automatically manipulating natural language, such as speech and / or text. For example, the NLP algorithms and functions 8006 may include content extraction algorithms or functions 8024, classification algorithms or functions 8026, machine translation algorithms or functions 8028, question answering (QA) algorithms or functions 8030, and text generation algorithms or functions 8032. In certain embodiments, the content extraction algorithms or functions 8024 may include a means for extracting text or images from electronic documents (e.g., webpages, text editor documents, and so forth) to be utilized, for example, in other applications.

[0410] In certain embodiments, the classification algorithms or functions 8026 may include any algorithms that may utilize a supervised learning model (e.g., logistic regression, naïve Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machine (SVM), and so forth) to learn from the data input to the supervised learning model and to make new observations or classifications based thereon. The machine translation algorithms or functions 8028 may include any algorithms or functions that can be suitable for automatically converting source text in one language, for example, into text in another language. The QA algorithms or functions 8030 may include any algorithms orMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 98 of 126 functions that can be suitable for automatically answering questions posed by humans in, for example, a natural language, such as that performed by voice-controlled personal assistant devices. The text generation algorithms or functions 8032 may include any algorithms or functions that can be suitable for automatically generating natural language texts.

[0411] In certain embodiments, the expert systems 8008 may include any algorithms or functions that can be suitable for simulating the judgment and behavior of a human or an organization that has expert knowledge and experience in a particular field (e.g., stock trading, medicine, sports statistics, and so forth). The computer-based vision algorithms and functions 8010 may include any algorithms or functions that can be suitable for automatically extracting information from images (e.g., photo images, video images). For example, the computer-based vision algorithms and functions 8010 may include image recognition algorithms 8034 and machine vision algorithms 8036. The image recognition algorithms 8034 may include any algorithms that can be suitable for automatically identifying and / or classifying objects, places, people, and so forth that can be included in, for example, one or more image frames or other displayed data. The machine vision algorithms 8036 may include any algorithms that can be suitable for allowing computers to “see”, or, for example, to rely on image sensors cameras with specialized optics to acquire images for processing, analyzing, and / or measuring various data characteristics for decision making purposes.

[0412] In certain embodiments, the speech recognition algorithms and functions 8012 may include any algorithms or functions that can be suitable for recognizing and translating spoken language into text, such as through automatic speech recognition (ASR), computer speech recognition, speech-to-text (STT) 8038, or text-to-speech (TTS) 8040 in order for the computing to communicate via speech with one or more users, for example. In certain embodiments, the planning algorithms and functions 8014 may include any algorithms or functions that can be suitable for generating a sequence of actions, in which each action may include its own set of preconditions to be satisfied before performing the action. Examples of AI planning may include classical planning, reduction to other problems, temporal planning, probabilistic planning, preference-based planning, conditional planning, and so forth. Lastly, the robotics algorithms and functions 8016 may include any algorithms, functions, or systems that may enable one or more devices to replicate human behavior through, for example, motions, gestures, performance tasks, decision-making, emotions, and so forth.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 99 of 126 IX. Example Descriptions of Terms

[0413] As used herein, “or” is inclusive and not exclusive, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A or B” means “A, B, or both,” unless expressly indicated otherwise or indicated otherwise by context. Moreover, “and” is both joint and several, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A and B” means “A and B, jointly or severally,” unless expressly indicated otherwise or indicated otherwise by context.

[0414] As used herein, “automatically” and its derivatives means “without human intervention,” unless expressly indicated otherwise or indicated otherwise by context.

[0415] As used herein, the term “nucleic acid”, “nucleic acid molecule”, or “oligonucleotide” are used interchangeably to refer to a polymer of nucleotide residues. The terms encompass oligonucleotides of any length.

[0416] As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably to refer to a polymer of amino acid residues. The terms encompass amino acid chains of any length, including full-length proteins with amino acid residues linked by covalent peptide bonds.

[0417] As used herein, a “sequence” refers to an oligonucleotide sequence or an amino acid sequence that includes an ordered set of nucleotides or amino acid identifiers, respectively.

[0418] As used herein, the terms “nucleic acid sequence” or “oligonucleotide sequence” are used interchangeably and refer to a sequence that identifies nucleotides of at least a portion of a nucleic acid molecule. In some cases, the nucleic acid sequence includes a variant sequence that includes a variant that is not observed in a corresponding reference sequence.

[0419] As used herein, a “peptide sequence” refers to a sequence that identifies amino acids of at least a portion of a peptide. In some cases, the peptide sequence includes a variant-coding sequence that includes a variant that is not observed in a corresponding reference sequence.

[0420] As used herein, a “reference sequence” may refer to a sequence that identifies nucleotides or amino acids within at least part of a non-mutant nucleic acid molecule or protein molecule, respectively, or to a wild-type nucleic acid sequence or protein sequence (e.g., a wild-type, parental sequence).

[0421] As used herein, a “representation” of a sequence or “sequence representation” can include a set of values that represent or identify amino acids in a protein or peptide sequenceMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 100 of 126 and / or a set of values that represent or identify nucleotides in a nucleic acid that encodes the protein or peptide sequence. For example, each nucleotide or amino acid can be represented by a binary string and / or vector of values that is distinct from every other binary string and / or vector of values representing every other nucleotide or amino acid. The sequence representation can be generated using, for example, one-hot encoding or using a BLOcks SUbstitution Matrix (BLOSUM) matrix. For example, a multi-dimensional (e.g., 20- or 21-dimensional) array be initialized (e.g., randomly or pseudo-randomly initialized). The initialized array may include, for each nucleotide or amino acid, a unique vector corresponding to that nucleotide or amino acid. The values can be fixed such that use of such a unique vector can be assumed to represent the corresponding nucleotide or amino acid. There can be multiple possible nucleic acid representations of a given protein or peptide sequence, given that any of multiple codons can encode a single amino acid.

[0422] As used herein, a “sample” can include tissue (e.g., a biopsy), single cell, multiple cells, fragments of cells, or an aliquot of body fluid. The sample can be obtained from a subject by means such as, for example, without limitation, venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, intervention, another type of sample collection means, or a combination thereof.

[0423] As used herein, a “subject” encompasses one or more cells, tissue, or an organism. The subject can be a human or non-human, whether in vivo, ex vivo, or in vitro, male or female. A subject can be a mammal, such as a human.

[0424] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims. X. Example Embodiments

[0425] Embodiments disclosed herein include:MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 101 of 126 1. A method for predicting gene expression levels from genomic sequence data, the method comprising: receiving, at one or more processors, genomic sequence data comprising a DNA sequence for a specified gene; inputting, using the one or more processors, the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific gene expression level; and outputting, using the one or more processors, a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene. 2. The method of embodiment 1, wherein the trained machine learning model is trained using a two-component loss function comprising a Poisson loss component and a multinomial negative log-likelihood component. 3. The method of embodiment 1, further comprising using the trained machine learning model to predict an impact of a gene sequence variant on gene expression level. 4. The method of any one of embodiments 1 to 3, further comprising using the trained machine learning model to predict gene expression patterns for a plurality of specified genes. 5. The method of any one of embodiments 1 to 4, further comprising using the trained machine learning model to predict differences in gene expression patterns for a plurality of specified genes in healthy versus diseased tissues. 6. The method of any one of embodiments 1 to 5, further comprising: (i) using the trained machine learning model to identify and compare gene sequences that are highly expressed in a given cell type, and (ii) clustering sequence features in the highly expressed gene sequences to identify transcription factor binding motifs.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 102 of 126 7. The method of any one of embodiments 1 to 6, further comprising using the trained machine learning model to design sequences for gene therapy, wherein the designed sequences are predicted to be actively expressed only in specific disease states. 8. The method of any one of embodiments 1 to 7, wherein the genomic sequence data comprises the DNA sequence for the specified gene and a flanking sequence for at least one of a 3’ end or a 5’ end of the specified gene sequence. 9. The method of any one of embodiments 1 to 8, wherein the genomic sequence data comprises a DNA sequence ranging in length from about 300,000 nucleotides to about 700,000 nucleotides. 10. The method of any one of embodiments 1 to 9, wherein the genomic sequence data comprises a DNA sequence of about 500,000 nucleotides in length. 11. The method of any one of embodiments 1 to 10, wherein the genomic sequence data is provided as input to the machine learning model as a one-hot encoded matrix. 12. The method of any one of embodiments 1 to 11, wherein the trained machine learning model comprises a deep neural network architecture. 13. The method of embodiment 12, wherein the deep neural network architecture comprises at least one convolution layer. 14. The method of embodiment 12 or embodiment 13, wherein the deep neural network architecture comprises at least one downsampling layer. 15. The method of any one of embodiments 12 to 14, wherein the deep neural network architecture comprises at least one self-attention layer. 16. The method of any one of embodiments 12 to 15, wherein the deep neural network architecture comprises at least one deconvolution layer with matched U-net connections. 17. The method of any one of embodiments 12 to 16, wherein the deep neural network architecture comprises at least one global mean pooling layer.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 103 of 126 18. The method of any one of embodiments 12 to 17, wherein the deep neural network architecture comprises a linear output layer. 19. The method of any one of embodiments 1 to 18, wherein the trained machine learning model is pre-trained on bulk gene expression profile data and / or epigenomic assay data. 20. The method of embodiment 19, wherein the bulk gene expression profile data comprises RNA-seq data and / or CAGE-seq data. 21. The method of embodiment 19 or embodiment 20, wherein the epigenomic assay data comprises DNase-seq data, ATAC-seq data, and / or ChIP-seq data. 22. The method of any one of embodiments 1 to 21, wherein the single cell gene expression data comprises single cell RNA sequencing data. 23. The method of any one of embodiments 1 to 22, wherein the plurality of cell types and cell states comprises at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 different types of cell. 24. The method of any one of embodiments 1 to 23, wherein the plurality of cell types and cell states comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130 different cell states. 25. The method of embodiment 24, wherein the different cell states comprise different development states, different metabolic states, different disease states, or any combination thereof. 26. The method of any one of embodiments 1 to 25, wherein the output comprises a prediction of at least 10, at least 25, at least 50, at least 75, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000 cell-type-specific and / or cell-state-specific gene expression levels for the specified gene. 27. A system comprising:MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 104 of 126 one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive genomic sequence data comprising a DNA sequence for a specified gene; input the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type- specific and / or a cell-state-specific gene expression level; and output a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene. 28. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive genomic sequence data comprising a DNA sequence for a specified gene; input the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific gene expression level; and output a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene. 29. A method for designing a DNA regulatory sequence efficacy, the method comprising: receiving, at one or more processors, at least one candidate DNA regulatory sequence;MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 105 of 126 receiving, at the one or more processors, a cargo gene sequence; generating, using the one or more processors, at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; inputting, using the one or more processors, the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level; determining, using the one or more processors, an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and selecting, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence. 30. The method of embodiment 29, wherein the at least one candidate DNA regulatory sequence comprises at least one promoter sequence. 31. The method of embodiment 29 or embodiment 30, wherein the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and wherein the method further comprises: for each of a plurality of iterations: generating a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; inputting the construct sequence into the trained machine learning model; determining an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence:MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 106 of 126 updating, using the one or more processors, the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or outputting, using the one or more processors, an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence. 32. The method of embodiment 31, further comprising repeating the method for each of a plurality of candidate insertion sites. 33. The method of any one of embodiments 31 to 32, further comprising repeating the method using at least one additional trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level. 34. The method of any one of embodiments 31 to 33, further comprising using the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence to develop a gene therapy or personalized anti-cancer vaccine. 35. The method of any one of embodiments 31 to 34, wherein the initial candidate DNA regulatory sequence comprises a random DNA sequence. 36. The method of any one of embodiments 31 to 35, wherein updating the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence comprises performing at least one single nucleotide substitution, single nucleotide insertion, or single nucleotide deletion. 37. The method of any one of embodiments 31 to 36, wherein updating initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence comprises performing at least one insertion, deletion, or substitution of a plurality of nucleotides. 38. The method of any one of embodiments 31 to 37, wherein updating the initial candidate DNA regulatory sequence or candidate DNA regulatory sequence comprises inserting nucleotides corresponding to at least one known regulatory sequence motif.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 107 of 126 39. The method of any one of embodiments 31 to 38, wherein the improved DNA regulatory sequence comprises a promoter sequence. 40. The method of any one of embodiments 31 to 39, wherein the initial candidate DNA regulatory sequence is between 100 and 1,000 nucleotides in length. 41. The method of any one of embodiments 31 to 40, wherein the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence is between 100 and 1,000 nucleotides in length. 42. The method of any one of embodiments 29 to 41, wherein the prediction of cell-type- specific and / or a cell-state-specific gene expression level is made at single nucleotide resolution. 43. The method of embodiment 42, wherein the efficacy score for the at least one candidate DNA regulatory sequence is based on a determination of a maximum predicted cell-type- specific and / or a cell-state-specific gene expression level, a mean predicted cell-type-specific and / or a cell-state-specific gene expression level, a median predicted cell-type-specific and / or a cell-state-specific gene expression level, or an integrated predicted cell-type-specific and / or a cell-state-specific gene expression level across the cargo gene sequence, or any combination thereof. 44. The method of embodiment 43, wherein the efficacy score for the at least one DNA regulatory sequence is further based on a determination of predicted cell-type-specific and / or a cell-state-specific off-target gene expression levels. 45. The method of any one of embodiments 29 to 44, wherein the at least one candidate insertion site comprises a safe harbor locus. 46. The method of any one of embodiments 29 to 45, wherein the genome is the human genome. 47. The method of any one of embodiments 30 to 46, wherein the trained machine learning model comprises a long-context DNA sequence prediction model.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 108 of 126 48. The method of any one of embodiments 29 to 47, wherein the trained machine learning model comprises a Borzoi model, an Enformer model, or a Decima model. 49. The method of any one of embodiments 29 to 48, wherein the at least one construct sequence comprising a candidate DNA regulatory sequence and the specified cargo gene is between 100 kilobases and 1 megabase in length. 50. The method of any one of embodiments 29 to 49, wherein sequence data for the at least on candidate DNA regulatory sequence is provided as input to the trained machine learning model as a one-hot encoded matrix. 51. The method of any one of embodiments 29 to 50, wherein the trained machine learning model comprises a deep neural network architecture. 52. The method of embodiment 51, wherein the deep neural network architecture comprises at least one convolution layer. 53. The method of embodiment 51 or embodiment 52, wherein the deep neural network architecture comprises at least one downsampling layer. 54. The method of any one of embodiments 51 to 53, wherein the deep neural network architecture comprises at least one self-attention layer. 55. The method of any one of embodiments 51 to 54, wherein the deep neural network architecture comprises at least one deconvolution layer with matched U-net connections. 56. The method of any one of embodiments 51 to 55, wherein the deep neural network architecture comprises at least one global mean pooling layer. 57. The method of any one of embodiments 51 to 56, wherein the deep neural network architecture comprises a linear output layer. 58. The method of any one of embodiments 52 to 57, wherein the machine learning model is pre-trained on bulk gene expression profile data and / or epigenomic assay data.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 109 of 126 59. The method of embodiment 58, wherein the bulk gene expression profile data comprises RNA-seq data and / or CAGE-seq data. 60. The method of embodiment 58 or embodiment 59, wherein the epigenomic assay data comprises DNase-seq data, ATAC-seq data, and / or ChIP-seq data. 61. The method of any one of embodiments 29 to 60, wherein the trained machine learning model is trained on single cell gene expression data for a plurality of cell types and cell states. 62. The method of embodiment 61, wherein the single cell gene expression data comprises single cell RNA sequencing data. 63. The method of embodiment 61 or embodiment 62, wherein the plurality of cell types and cell states comprises at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 different types of cell. 64. The method of any one of embodiments 61 to 63, wherein the plurality of cell types and cell states comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130 different cell states. 65. The method of embodiment 64, wherein the different cell states comprise different development states, different metabolic states, different disease states, or any combination thereof. 66. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive at least one candidate DNA regulatory sequence; receive a cargo gene sequence;MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 110 of 126 generate at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; input the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level; determine an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and select, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence. 67. The system of embodiment 66, wherein the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and further comprising instructions that, when executed by the one or more processors, cause the system to: for each of a plurality of iterations: generate a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; input the construct sequence into the trained machine learning model; determine an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: update the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; orMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 111 of 126 output an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence. 68. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive at least one candidate DNA regulatory sequence; receive a cargo gene sequence; generate at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; input the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell- state-specific gene expression level; determine an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and select, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence. 69. The non-transitory computer-readable storage medium of embodiment 68, wherein the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and further comprising instructions that, when executed by the one or more processors, cause the system to: for each of a plurality of iterations: generate a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; input the construct sequence into the trained machine learning model;MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 112 of 126 determine an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: update the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or output an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence. 70. A method for predicting mRNA properties from genomic sequence data, the method comprising: receiving, at one or more processors, genomic sequence data comprising a DNA sequence for a specified gene; inputting, using the one or more processors, the genomic sequence data to a machine learning model, wherein the machine learning model is configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific property of an mRNA molecule corresponding to a specified gene; and outputting, using the one or more processors, a prediction of a cell-type-specific and / or a cell-state-specific property of an mRNA molecule corresponding to the specified gene, wherein the predicted cell-type-specific and / or a cell-state-specific property comprises at least one of: (i) a predicted mRNA stability, or (ii) a predicted mRNA translation efficiency. 71. The method of embodiment 70, further comprising using the predicted mRNA stability and / or predicted mRNA translation efficiency to predict an expression level for a protein corresponding to the specified gene.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 113 of 126

[0426] The description provides preferred example embodiments only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the description of the preferred example embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It is understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.MOFO-358053825

Claims

ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 114 of 126 CLAIMS What is claimed is:

1. A method for predicting gene expression levels from genomic sequence data, the method comprising: receiving, at one or more processors, genomic sequence data comprising a DNA sequence for a specified gene; inputting, using the one or more processors, the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific gene expression level; and outputting, using the one or more processors, a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene.

2. The method of claim 1, wherein the trained machine learning model is trained using a two- component loss function comprising a Poisson loss component and a multinomial negative log- likelihood component.

3. The method of claim 1, further comprising using the trained machine learning model to predict an impact of a gene sequence variant on gene expression level.

4. The method of any one of claims 1 to 3, further comprising using the trained machine learning model to predict gene expression patterns for a plurality of specified genes.

5. The method of any one of claims 1 to 4, further comprising using the trained machine learning model to predict differences in gene expression patterns for a plurality of specified genes in healthy versus diseased tissues.

6. The method of any one of claims 1 to 5, further comprising: (i) using the trained machine learning model to identify and compare gene sequences that are highly expressed in a givenMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 115 of 126 cell type, and (ii) clustering sequence features in the highly expressed gene sequences to identify transcription factor binding motifs.

7. The method of any one of claims 1 to 6, further comprising using the trained machine learning model to design sequences for gene therapy, wherein the designed sequences are predicted to be actively expressed only in specific disease states.

8. The method of any one of claims 1 to 7, wherein the genomic sequence data comprises the DNA sequence for the specified gene and a flanking sequence for at least one of a 3’ end or a 5’ end of the specified gene sequence.

9. The method of any one of claims 1 to 8, wherein the genomic sequence data comprises a DNA sequence ranging in length from about 300,000 nucleotides to about 700,000 nucleotides.

10. The method of any one of claims 1 to 9, wherein the genomic sequence data comprises a DNA sequence of about 500,000 nucleotides in length.

11. The method of any one of claims 1 to 10, wherein the genomic sequence data is provided as input to the machine learning model as a one-hot encoded matrix.

12. The method of any one of claims 1 to 11, wherein the trained machine learning model comprises a deep neural network architecture.

13. The method of claim 12, wherein the deep neural network architecture comprises at least one convolution layer.

14. The method of claim 12 or claim 13, wherein the deep neural network architecture comprises at least one downsampling layer.

15. The method of any one of claims 12 to 14, wherein the deep neural network architecture comprises at least one self-attention layer.

16. The method of any one of claims 12 to 15, wherein the deep neural network architecture comprises at least one deconvolution layer with matched U-net connections.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 116 of 126 17. The method of any one of claims 12 to 16, wherein the deep neural network architecture comprises at least one global mean pooling layer.

18. The method of any one of claims 12 to 17, wherein the deep neural network architecture comprises a linear output layer.

19. The method of any one of claims 1 to 18, wherein the trained machine learning model is pre-trained on bulk gene expression profile data and / or epigenomic assay data.

20. The method of claim 19, wherein the bulk gene expression profile data comprises RNA- seq data and / or CAGE-seq data.

21. The method of claim 19 or claim 20, wherein the epigenomic assay data comprises DNase- seq data, ATAC-seq data, and / or ChIP-seq data.

22. The method of any one of claims 1 to 21, wherein the single cell gene expression data comprises single cell RNA sequencing data.

23. The method of any one of claims 1 to 22, wherein the plurality of cell types and cell states comprises at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 different types of cell.

24. The method of any one of claims 1 to 23, wherein the plurality of cell types and cell states comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130 different cell states.

25. The method of claim 24, wherein the different cell states comprise different development states, different metabolic states, different disease states, or any combination thereof.

26. The method of any one of claims 1 to 25, wherein the output comprises a prediction of at least 10, at least 25, at least 50, at least 75, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000 cell-type-specific and / or cell-state-specific gene expression levels for the specified gene.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 117 of 126 27. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive genomic sequence data comprising a DNA sequence for a specified gene; input the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type- specific and / or a cell-state-specific gene expression level; and output a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene.

28. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive genomic sequence data comprising a DNA sequence for a specified gene; input the genomic sequence data to a trained machine learning model, wherein the trained machine learning model is: trained on single cell gene expression data for a plurality of cell types and cell states, and configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific gene expression level; and output a prediction of a cell-type-specific and / or a cell-state-specific gene expression level for the specified gene.

29. A method for designing a DNA regulatory sequence efficacy, the method comprising:MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 118 of 126 receiving, at one or more processors, at least one candidate DNA regulatory sequence; receiving, at the one or more processors, a cargo gene sequence; generating, using the one or more processors, at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; inputting, using the one or more processors, the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level; determining, using the one or more processors, an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and selecting, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence.

30. The method of claim 29, wherein the at least one candidate DNA regulatory sequence comprises at least one promoter sequence.

31. The method of claim 29 or claim 30, wherein the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and wherein the method further comprises: for each of a plurality of iterations: generating a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; inputting the construct sequence into the trained machine learning model; determining an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; andMOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 119 of 126 based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: updating, using the one or more processors, the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or outputting, using the one or more processors, an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence.

32. The method of claim 31, further comprising repeating the method for each of a plurality of candidate insertion sites.

33. The method of any one of claims 31 to 32, further comprising repeating the method using at least one additional trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level.

34. The method of any one of claims 31 to 33, further comprising using the improved cell-type- specific and / or a cell-state-specific DNA regulatory sequence to develop a gene therapy or personalized anti-cancer vaccine.

35. The method of any one of claims 31 to 34, wherein the initial candidate DNA regulatory sequence comprises a random DNA sequence.

36. The method of any one of claims 31 to 35, wherein updating the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence comprises performing at least one single nucleotide substitution, single nucleotide insertion, or single nucleotide deletion.

37. The method of any one of claims 31 to 36, wherein updating initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence comprises performing at least one insertion, deletion, or substitution of a plurality of nucleotides.

38. The method of any one of claims 31 to 37, wherein updating the initial candidate DNA regulatory sequence or candidate DNA regulatory sequence comprises inserting nucleotides corresponding to at least one known regulatory sequence motif.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 120 of 126 39. The method of any one of claims 31 to 38, wherein the improved DNA regulatory sequence comprises a promoter sequence.

40. The method of any one of claims 31 to 39, wherein the initial candidate DNA regulatory sequence is between 100 and 1,000 nucleotides in length.

41. The method of any one of claims 31 to 40, wherein the improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence is between 100 and 1,000 nucleotides in length.

42. The method of any one of claims 29 to 41, wherein the prediction of cell-type-specific and / or a cell-state-specific gene expression level is made at single nucleotide resolution.

43. The method of claim 42, wherein the efficacy score for the at least one candidate DNA regulatory sequence is based on a determination of a maximum predicted cell-type-specific and / or a cell-state-specific gene expression level, a mean predicted cell-type-specific and / or a cell-state-specific gene expression level, a median predicted cell-type-specific and / or a cell- state-specific gene expression level, or an integrated predicted cell-type-specific and / or a cell- state-specific gene expression level across the cargo gene sequence, or any combination thereof.

44. The method of claim 43, wherein the efficacy score for the at least one DNA regulatory sequence is further based on a determination of predicted cell-type-specific and / or a cell-state- specific off-target gene expression levels.

45. The method of any one of claims 29 to 44, wherein the at least one candidate insertion site comprises a safe harbor locus.

46. The method of any one of claims 29 to 45, wherein the genome is the human genome.

47. The method of any one of claims 30 to 46, wherein the trained machine learning model comprises a long-context DNA sequence prediction model.

48. The method of any one of claims 29 to 47, wherein the trained machine learning model comprises a Borzoi model, an Enformer model, or a Decima model.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 121 of 126 49. The method of any one of claims 29 to 48, wherein the at least one construct sequence comprising a candidate DNA regulatory sequence and the specified cargo gene is between 100 kilobases and 1 megabase in length.

50. The method of any one of claims 29 to 49, wherein sequence data for the at least on candidate DNA regulatory sequence is provided as input to the trained machine learning model as a one-hot encoded matrix.

51. The method of any one of claims 29 to 50, wherein the trained machine learning model comprises a deep neural network architecture.

52. The method of claim 51, wherein the deep neural network architecture comprises at least one convolution layer.

53. The method of claim 51 or claim 52, wherein the deep neural network architecture comprises at least one downsampling layer.

54. The method of any one of claims 51 to 53, wherein the deep neural network architecture comprises at least one self-attention layer.

55. The method of any one of claims 51 to 54, wherein the deep neural network architecture comprises at least one deconvolution layer with matched U-net connections.

56. The method of any one of claims 51 to 55, wherein the deep neural network architecture comprises at least one global mean pooling layer.

57. The method of any one of claims 51 to 56, wherein the deep neural network architecture comprises a linear output layer.

58. The method of any one of claims 52 to 57, wherein the machine learning model is pre- trained on bulk gene expression profile data and / or epigenomic assay data.

59. The method of claim 58, wherein the bulk gene expression profile data comprises RNA- seq data and / or CAGE-seq data.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 122 of 126 60. The method of claim 58 or claim 59, wherein the epigenomic assay data comprises DNase- seq data, ATAC-seq data, and / or ChIP-seq data.

61. The method of any one of claims 29 to 60, wherein the trained machine learning model is trained on single cell gene expression data for a plurality of cell types and cell states.

62. The method of claim 61, wherein the single cell gene expression data comprises single cell RNA sequencing data.

63. The method of claim 61 or claim 62, wherein the plurality of cell types and cell states comprises at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, or at least 190 different types of cell.

64. The method of any one of claims 61 to 63, wherein the plurality of cell types and cell states comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130 different cell states.

65. The method of claim 64, wherein the different cell states comprise different development states, different metabolic states, different disease states, or any combination thereof.

66. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive at least one candidate DNA regulatory sequence; receive a cargo gene sequence; generate at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site;MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 123 of 126 input the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell-state-specific gene expression level; determine an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and select, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence.

67. The system of claim 66, wherein the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and further comprising instructions that, when executed by the one or more processors, cause the system to: for each of a plurality of iterations: generate a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; input the construct sequence into the trained machine learning model; determine an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: update the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or output an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence.MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 124 of 126 68. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive at least one candidate DNA regulatory sequence; receive a cargo gene sequence; generate at least one construct sequence by inserting the at least one candidate DNA regulatory sequence and the cargo gene sequence into a genome at at least one candidate insertion site; input the at least one construct sequence into a trained machine learning model configured to process an input construct sequence and predict a cell-type-specific and / or a cell- state-specific gene expression level; determine an efficacy score for the at least one candidate DNA regulatory sequence based on the prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the at least one construct sequence; and select, based at least in part on the efficacy score for the at least one candidate DNA regulatory sequence, a cell-type-specific and / or cell-state-specific combination of a DNA regulatory sequence and an insertion site for expression of the cargo gene sequence.

69. The non-transitory computer-readable storage medium of claim 68, wherein the at least one candidate DNA regulatory sequence comprises an initial candidate DNA regulatory sequence, and further comprising instructions that, when executed by the one or more processors, cause the system to: for each of a plurality of iterations: generate a construct sequence by inserting the initial candidate DNA regulatory sequence or a current candidate DNA regulatory sequence and the cargo gene sequence into a genome at a candidate insertion site; input the construct sequence into the trained machine learning model;MOFO-358053825ATTORNEY DOCKET PATENT APPLICATION 14639-20696.40 P39460-WO 125 of 126 determine an efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence based on a prediction of the cell-type-specific and / or the cell-state-specific gene expression level for the construct sequence; and based on the efficacy score for the initial candidate DNA regulatory sequence or the current candidate DNA regulatory sequence: update the initial candidate DNA regulatory sequence or the candidate DNA regulatory sequence and performing a next iteration; or output an improved cell-type-specific and / or a cell-state-specific DNA regulatory sequence.

70. A method for predicting mRNA properties from genomic sequence data, the method comprising: receiving, at one or more processors, genomic sequence data comprising a DNA sequence for a specified gene; inputting, using the one or more processors, the genomic sequence data to a machine learning model, wherein the machine learning model is configured to process genomic sequence data and predict a cell-type-specific and / or a cell-state-specific property of an mRNA molecule corresponding to a specified gene; and outputting, using the one or more processors, a prediction of a cell-type-specific and / or a cell-state-specific property of an mRNA molecule corresponding to the specified gene, wherein the predicted cell-type-specific and / or a cell-state-specific property comprises at least one of: (i) a predicted mRNA stability, or (ii) a predicted mRNA translation efficiency.

71. The method of claim 70, further comprising using the predicted mRNA stability and / or predicted mRNA translation efficiency to predict an expression level for a protein corresponding to the specified gene.MOFO-358053825