Unifying foundation model for functional genomics
A multi-modal, multi-masking-based deep learning model integrates DNA sequence and functional genomic data with adaptable masking strategies, addressing limitations of existing models by enhancing predictive and computational performance and enabling versatile genomic property predictions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2026-03-12
AI Technical Summary
Existing deep learning models for DNA sequence-based prediction in functional genomics are application-specific and cannot dynamically condition on arbitrary tasks or denoise sparsely sampled data, limiting their versatility and accuracy in predicting genomic properties.
A multi-modal, multi-masking-based deep learning model that integrates self-supervised learning on DNA sequences with supervised learning on functional genomic data, using a multiscale loss function and adaptable masking strategies to unify various prediction tasks into a single framework.
The model enhances predictive and computational performance by enabling flexible task reconfiguration, denoising sparsely sampled data, and improving accuracy across diverse genomic applications, facilitating efficient processing of multiple data modalities and reducing the need for multiple models.
Smart Images

Figure US2025044711_12032026_PF_FP_ABST
Abstract
Description
ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-11 of 94UNIFYING FOUNDATION MODEU FOR FUNCTIONAU GENOMICSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority benefit of United States Provisional Patent Application Serial No. 63 / 690,607, filed September 4, 2024, and of United States Provisional Patent Application Serial No. 63 / 781,561, filed April 1, 2025, the contents of each of which are incorporated herein by reference in their entirety.FIELD
[0002] This present disclosure generally relates to methods and systems for the analysis of genomic data, and more specifically to a multi-modal, multi-masking-based deep learning model configured to output sequence and / or functional genomics property predictions based on input data comprising DNA sequences and / or functional genomics data.BACKGROUND
[0003] Deep learning models for deoxyribonucleic acid (DNA) sequence-based prediction have achieved state-of-the-art performance on diverse tasks in regulatory genomics, and are used for, e.g., variant effect prediction and prioritization, understanding cis-regulatory logic, and designing regulatory elements. Bespoke variations of DNA sequence-based models that utilize different inputs and generate different outputs have been designed, e.g., models that predict gene expression from DNA sequence and epigenetic marks. Other variations omit DNA sequence data and instead map functional genomic profiles derived from one type of assay to those of others. Separately, methods have been developed for denoising functional genomic profiles derived from epigenetic assays.
[0004] Recently, masked-based, self-supervised modeling methods have been applied to applications in genomics. Within functional genomics, the standard approach is to first train a masked language model on DNA sequence data, and then fine-tune it on functional genomics data. However, this approach has yielded application-specific models that cannot perform DNA sequence-based prediction while also dynamically conditioning on an arbitrary set of tasks and / or known chromatin environments. These models cannot be used to directly design regulatory sequences given a fixed set of functional genomic property targets. Nor can they be used to denoise sparsely sampled functional genomic data. Thus, there is a need for improvedMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-12 of 94 methods and / or models configured to output refined predictions based on input data comprising DNA sequences and / or functional genomics data.SUMMARY
[0005] Disclosed herein are machine learning-based methods, systems, and programming for training and deploying a novel multimodal, multi-masking-based deep learning model configured to predict cell type-specific functional genomic traits. The model utilizes a new training strategy comprising: (i) self-supervised learning on DNA sequence, and (ii) supervised learning on functional genomic data, along with the use of a multiscale loss function. A key feature of the disclosed methods and systems is the ability of the core deep learning model to be reconfigured for different functional genomics prediction applications by changing the masking strategy used during training and / or inference.
[0006] The methods and systems described herein provide numerous technical advantages in terms of improved predictive and / or computational performance that derive from the model architecture and training approach. The disclosed model unifies the functionality of disparate prior art models into a single framework that is capable of handling a wide range of functional genomics prediction tasks - prediction tasks that previously required using separate, taskspecific models. As noted above, the core model can be reconfigured for different prediction tasks (such as sequence-to-function prediction, function-to-sequence prediction, denoising, and context-aware prediction, etc.} by changing the masking strategy used during training and / or inference. The model accepts multi-modal input comprising combinations of DNA sequence data and / or functional genomics data (e.g., DNasel hypersentitive sites sequencing (DNase- seq) data, Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) data, Chromatin Immunoprecipitation sequencing (ChlP-seq) data, Cap Analysis Gene Expression sequencing (CAGE-seq) data, etc.}, thereby enabling richer and more context-aware predictions. The model also supports variable-length input sequences, thereby allowing the model to process genomic regions of different sizes to capture both local and long-range genomic effects.
[0007] The disclosed methods and systems provide improved predictive and computational performance. For example, the loss function (e.g., a multi-scale Poisson loss function computed at multiple resolutions that match the resolutions of the deconvolution layers of a U-net model architecture) used during training improves the model’s ability to make accurate predictions atMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-13 of 94 both fine and coarse resolutions, as demonstrated by improved performance on benchmark tasks. The model can be configured to incorporate chromatin context and / or other functional data to improve prediction accuracy, particularly in genomic regions where sequence-only models underperform. The model can also be configured to denoise sparsely sampled or noisy functional genomics data, thereby improving the reliability of downstream analyses.
[0008] The versatility of the disclosed modelling methods supports a variety of prediction applications including, but not limited to, (i) function-to-sequence and sequence-to-function prediction, z.e., the model supports both the prediction of functional properties from sequence data, and the design of regulatory sequences to achieve desired functional property profiles; (ii) genotype prediction and de-identification, z.e., the model can infer individual genotypes from functional genomics data (e.g., ATAC-seq data), thereby enabling applications in privacy assessment and sample identification; and (iii) functional property perturbation and dependency analysis, i.e., the model can predict the effects of perturbing one functional genomic property on another, and can be used to study dependencies between different functional genomic property tracks.
[0009] The disclosed multi-modal, multi-masking-based deep learning model also provides several improvements over existing computer technologies in the functional genomics field, both at the level of computational efficiency and at the level of enabling new computational capabilities. Examples of these improvements include efficient utilization of computational resources through: (i) consolidation of multiple genomics prediction tasks (e.g., sequence-to- function, function-to-sequence, denoising, context-aware prediction, etc.) into a single core model, thereby eliminating the need to train, maintain, and deploy different models for different prediction tasks (and also reducing memory usage, storage requirements, and computational overhead); (ii) the multi-modal and multi-task model is designed to process multiple data modalities (e.g., DNA sequence and various functional genomics tracks) in parallel, leveraging shared computational pathways and reducing the need for duplicative data processing; and (iii) the use of a multi-scale loss function: the use of a multi-scale loss function enables the model to learn and make predictions at multiple resolutions in a single pass, thereby improving computational efficiency compared to running separate models or passes for different resolutions. The ability to reconfigure the disclosed multi-modal, multi-masking-based deep learning model for different prediction or design tasks eliminates the need to fully retrain orMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-14 of 94 reload new models. Parallel and scalable processing is enabled by optimizing the model’s architecture (e.g., comprising the use of convolutional and self-attention layers, parallel processing blocks, etc.) for implementation on modern computing hardware (such as graphical processing units (GPUs) and cloud computing platforms), thereby enabling efficient large-scale data processing and rapid inference. By consolidating the capability for performing multiple prediction tasks into a single model, the disclosed model can also reduce the need for developing, validating, and maintaining multiple specialized models, thereby streamlining research and development workflows in the functional genomics field.
[0010] The technical improvements described above result in faster, more accurate, and more resource-efficient computation for genomics prediction tasks, and enable new applications (such as context-aware prediction, desired functional property-based design of regulatory sequences, and the analysis of dependencies between functional genomics property tracks) that were not practical or possible with prior art computational systems, thereby expanding the functional capabilities of computer technology in the genomics field. These capabilities can, in turn, improve the efficacy of downstream development and manufacturing processes in the biotechnology and pharmaceutical industries (e.g., by facilitating the design and development of effective gene therapies (e.g., gene therapies directed to turning off expression of mutated genes, or to promoting expression of beneficial genes), Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-based drug delivery mechanisms (e.g., to improve the efficacy and precision of gene editing tools), and protein expression systems for biologies manufacturing (e.g., to maximize yields of expressed proteins), etc.).
[0011] Disclosed herein are methods for predicting a functional genomic property, comprising: obtaining an input genomic sequence derived from a test sample; obtaining at least one input functional genomic property dataset derived from the test sample; formulating an input data structure based on the input genomic sequence and the at least one input functional genomic property dataset; and predicting at least one output functional genomic property for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by: obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property datasetMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-15 of 94 derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the plurality of masked training input data structures using a multiscale loss function.
[0012] In some embodiments, the trained machine learning model is configured to receive a masked input data structure and predict the masked portion of the masked input data structure.
[0013] In some embodiments, the at least one predicted output functional genomic property for the test sample is different from the at least one input functional genomic property dataset derived from the test sample.
[0014] In some embodiments, the trained machine learning model comprises a multimodal machine learning model. In some embodiments, the multimodal machine learning model is configured to process genomic sequence data and at least 2, 3, 4, 5, 6, 7, 8 ,9, or 10 different functional genomic property datasets derived from the test sample as input.
[0015] In some embodiments, the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof. In some embodiments, the plurality of training samples comprise samples from at least one cell type and / or at least one species.
[0016] In some embodiments, the at least one input functional genomic property dataset comprises ATAC-Seq data, CAGE-Seq data, ChlP-Seq data, DNase-Seq data, H3K27ac-Seq data, or any combination thereof.
[0017] In some embodiments, the at least one output functional genomic property prediction comprises an ATAC-Seq activity, a CAGE-Seq activity, a ChlP-Seq activity, a DNase-Seq activity, an H3K27ac-Seq activity, or any combination thereof for at least one genomic locus.
[0018] In some embodiments, the prediction output by the trained machine learning model is used to predict a functional effect of a genomic variant. In some embodiments, the prediction output by the trained machine learning model is used to interpret a functional effect of a genomic variant. In some embodiments, the prediction output by the trained machine learning model is used to denoise a functional genomic property dataset for the test sample. In some embodiments, the prediction output by the trained machine learning model is used to identify dependencies between functional genomic properties in the test sample. In some embodiments,MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-16 of 94 the prediction output by the trained machine learning model is used to diagnose a disease. In some embodiments, the prediction output by the trained machine learning model is used to select a treatment for a disease. In some embodiments, the prediction output by the trained machine learning model is used to inform a disease prognosis.
[0019] Disclosed herein are methods for predicting a base at at least one specified genomic locus in a genomic sequence, comprising: obtaining the genomic sequence derived from a test sample, the genomic sequence comprising an unknown base at the at least one specified genomic locus; obtaining at least one functional genomic property dataset derived from the test sample; formulating an input data structure based on the genomic sequence and the at least one functional genomic property dataset; and predicting the base at the at least one specified genomic locus in the genomic sequence for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by: obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask at least one base in the genomic sequence of each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the masked training input data structures using a multiscale loss function to predict the at least one masked base.
[0020] In some embodiments, the pre-defined masking strategy comprises a variable masking strategy in which a different number of genomic loci and / or different genomic loci are masked in different training input data structures of the plurality of training input data structures.
[0021] In some embodiments, the method further comprises: identifying an individual associated with the test sample based on the predicted base at the at least one specified genomic locus for the test sample. In some embodiments, identifying the individual associated with the test sample comprises: obtaining genomic sequences derived from test samples from each of a plurality of individuals; obtaining at least one functional genomic property dataset derived from each of the plurality of test samples; formulating a plurality of input data structures based onMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-17 of 94 the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of test samples; inputting the plurality of input data structures into the trained machine learning model; outputting a predicted base at the at least one specified genomic locus in the genomic sequence for each of the plurality of test samples; and determining an identity of at least one individual based on the predicted base at the at least one specified genomic locus in the genomic sequence for each of the plurality of test samples.
[0022] In some embodiments, the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof. In some embodiments, the plurality of training samples comprise samples from at least one cell type and / or at least one species.
[0023] Also disclosed herein are methods for designing a genomic regulatory element comprising: obtaining at least one functional genomic property dataset generated based on a desired regulatory effect in a test sample; formulating an input data structure based on a candidate genomic sequence of specified length and the at least one functional genomic property dataset; predicting a base at one or more positions in the candidate genomic sequence by: inputting the input data structure into a trained machine learning model, the trained machine learning model configured to output a prediction of a base at one or more positions of a genomic sequence based on at least one input functional genomic property dataset; and substituting the predicted base at the one or more positions of the candidate genomic sequence for the previous base at the one or more positions of the candidate genomic sequence in the input data structure; and outputting the candidate genomic sequence as a designed genomic regulatory element for the test sample.
[0024] In some embodiments, the method further comprises repeating the steps of: (i) inputting the data structure into the trained machine learning model and (ii) substituting the predicted base at the one or more positions of the candidate genomic sequence for the previous base at the one or more positions of the candidate genomic sequence, for two or more iterations.
[0025] In some embodiments, the designed genomic regulatory element comprises an adeno-associated virus (AAV) vector.
[0026] In some embodiments, the trained machine learning model is trained on training data comprising genomic sequences and the at least one functional genomic property dataset derived from each of a plurality of training samples. In some embodiments, the training data isMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-18 of 94 formulated as a plurality of masked training input data structures generated from a plurality of training input data structures using a variable masking strategy in which a different number of genomic loci and / or different genomic loci are masked in different training input data structures of the plurality of training input data structures. In some embodiments, the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof. In some embodiments, the plurality of training samples comprise samples from at least one cell type and / or at least one species.
[0027] Disclosed herein are methods for predicting a genomic sequence and / or a functional genomic property, comprising: formulating an input data structure comprising a genomic sequence data component and / or at least one functional genomic property data component derived from a test sample, wherein at least a portion of the genomic sequence data component and / or the at least one functional genomic property data component is unknown; inputting the input data structure into a trained machine learning model configured to predict the unknown portion of the genomic sequence component and / or the at least one functional genomic property component, the machine learning model trained by: obtaining genomic sequences and / or at least one functional genomic property dataset derived from each of a plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and / or the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the genomic sequence and / or at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model to receive a masked training input data structure and predict the masked portion of the masked training input data structure; and outputting a prediction for the unknown portion of the genomic sequence data component and / or the at least one functional genomic property data component for the test sample.
[0028] In some embodiments, the trained machine learning model comprises a masked multi-modal machine learning model configured to process genomic sequence data and / or 2, 3, 4, 5, 6, 7, 8, 9, or 10 different functional genomic property datasets derived from the test sample. In some embodiments, the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof. In someMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-19 of 94 embodiments, the plurality of training samples comprise samples from at least one cell type and / or at least one species.
[0029] In some embodiments, the at least one functional genomic property prediction comprises an ATAC-Seq activity, a CAGE-Seq activity, a ChlP-Seq activity, a DNase-Seq activity, an H3K27ac-Seq activity, or any combination thereof for at least one genomic locus.
[0030] In some embodiments, the prediction output by the trained machine learning model is used to predict a functional effect of a genomic variant. In some embodiments, the prediction output by the trained machine learning model is used to interpret a functional effect of a genomic variant. In some embodiments, the prediction output by the trained machine learning model is used to denoise a functional genomic property dataset for the test sample. In some embodiments, the prediction output by the trained machine learning model is used to identify dependencies between functional genomic properties in the test sample. In some embodiments, the prediction output by the trained machine learning model is used to diagnose a disease. In some embodiments, the prediction output by the trained machine learning model is used to select a treatment for a disease. In some embodiments, the prediction output by the trained machine learning model is used to inform a disease prognosis.
[0031] In some embodiments, the input data structure comprises a matrix. In some embodiments, at least one row of the matrix comprises genomic sequence data. In some embodiments, the matrix comprises a first row comprising a first functional genomic property dataset. In some embodiments, the matrix comprises at least a second row comprising at least a second functional genomic property dataset.
[0032] In some embodiments, the trained machine learning model is trained using a multiscale loss function.
[0033] In some embodiments, the test sample comprises a cell sample, a blood sample, a biopsy sample, or a tissue sample.
[0034] In some embodiments, the at least one functional genomic property dataset comprises ATAC-Seq data, CAGE-Seq data, ChlP-Seq data, DNase-Seq data, H3K27ac-Seq data, or any combination thereof.
[0035] Disclosed herein are systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructionsMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-110 of 94 that, when executed by the one or more processors, cause the system to perform any of the methods described herein.
[0036] Also disclosed herein are non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to perform any of the methods described herein.
[0037] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.INCORPORATION BY REFERENCE
[0038] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0040] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-111 of 94
[0041] FIG. 1 provides a block diagram of an example prediction system, in accordance with some embodiments of the present disclosure.
[0042] FIG. 2 provides a non-limiting, schematic illustration of a machine learning model architecture, in accordance with some embodiments of the present disclosure.
[0043] FIG. 3 provides a non-limiting example of a process flowchart for training a multimodal, multi-masking-based deep learning model configured to process input sequence and / or functional genomic property data and output sequence and / or functional genomics property predictions, in accordance with some embodiments of the present disclosure.
[0044] FIG. 4 provides a non-limiting example of a process flowchart for predicting DNA sequences and / or cell type-specific functional genomics properties, in accordance with some embodiments of the present disclosure.
[0045] FIG. 5 provides a non-limiting example of a process flowchart for predicting a base at a specified genomic locus, in accordance with some embodiments of the present disclosure.
[0046] FIG. 6 provides a non-limiting example of a process flowchart for designing a genomic regulatory element, in accordance with some embodiments of the present disclosure.
[0047] FIG. 7 provides a non-limiting schematic illustration of a multi-modal, multimasking-based deep learning model (a functional genomics prediction model) configured to output sequence and / or functional genomics property predictions, in accordance with some embodiments of the present disclosure.
[0048] FIG. 8 provides a non-limiting schematic illustration of how incorporating different masks during inference (selected from a same distribution of masks used during model training) enables versatile use cases for the single core model that were hitherto addressed using separate models.
[0049] FIG. 9 provides a non-limiting example of correlation data for model performance on 12 tasks with and without using a multiscale loss function during model training.
[0050] FIG. 10 provides a non-limiting example of prediction accuracy data for model prediction of masked bases in language modeling mode.
[0051] FIG. 11 provides a non-limiting example of data for base-resolution prediction loss (Poisson NLL loss) for deployment of the model in sequence-to-functional genomic property prediction and functional genomic property prediction denoising modes.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-112 of 94
[0052] FIG. 12 provides a non-limiting example of data for Spearman counts correlation of the ChromBPNet model and the presently disclosed model on ATAC-seq data peaks from Fibroblasts and a GC -matched set of non-peak regions in a test set of genomic sequences from “test” chromosomes that were not included in the training data.
[0053] FIGS. 13A-B provide non-limiting examples of comparisons of observed and predicted functional genomic property profile data at a heterochromatin locus. FIG. 13A: data for model predictions when used in sequence to functional genomic property profile prediction mode (2 kb input sequence). FIG. 13B: data for model predictions when used in sequence + context to functional genomic property profile prediction mode (2 kb ± 65.5 kb input sequence).
[0054] FIG. 14 provides a non-limiting example of predicted versus measured (labeled) Fibroblast ATAC-seq data profiles using 131 kb input fibroblast sequences. The labeled profile has been inverted for ease of comparison. Inset: zoomed in view of the predicted and measured profile data in the indicated genomic region.
[0055] FIG. 15 provides a non-limiting example of a comparison of model prediction performance data for a prior art model (ChromBPNet) and different versions of the model disclosed herein.
[0056] FIG. 16 provides a non-limiting example of prediction accuracy data for a test set comprising DNase-seq and CAGE-seq data when the model was operated in sequence only (4 kb) versus sequence + context (196 kb) input modes.
[0057] FIG. 17 provides a non-limiting example of data for the correlation between model predictions when operated in sequence + context input mode versus model predictions when operated in sequence-only input mode.
[0058] FIG. 18A provides a non-limiting example of data that illustrates the distribution of ChromHMM states in differential genomic regions.
[0059] FIG. 18B provides a non-limiting example of data for comparison of sequence-only prediction of DNase-seq data in K562 cells versus experimentally-determined DNase-seq data for a genomic locus where including context in the model input improved the model’s prediction performance.
[0060] FIG. 19 provides a non-limiting example of use of the masking strategy used to predict gene expression (CAGE-seq ) profiles in K562 human chronic myelogenous leukemiaMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-113 of 94(CML) cells when a reporter sequence of about 2 kb was inserted at multiple genomic insertion sites.
[0061] FIG. 20 provides a non-limiting example of Spearman R correlation coefficient data for predicted versus measured Steensel TRIP-seq data in K562 cells.
[0062] FIG. 21 provides a non-limiting example of Spearman R correlation coefficient data for predicted versus measured Cohen TRIP-seq data in K562 cells.
[0063] FIG. 22A provides a non-limiting example of data comparing the functional genomic property prediction performance of the presently disclosed model to that of a prior art model.
[0064] FIG. 22B provides a non-limiting example of precision-recall data for model prediction performance for the model disclosed herein and for several prior art models (e.g., ChromBPNet, Enformer, and Borzoi) when deployed for dsQTL classification in lymphoblastoid cell lines (LCLs).
[0065] FIG. 22C provides a non-limiting example of data illustrating the effect of perturbing an epigenetic property (e.g., histone modification) on the prediction of another functional genomic property (e.g., CAGE-seq data) using the disclosed model.
[0066] FIG. 22D provides a non-limiting example of data showing the increase in prediction performance over sequence-only prediction mode when additional functional genomic property data tracks were include with genomic sequence data as part of the model input.
[0067] FIG. 23A provides a non-limiting example of ATAC signal as a function of position relative to the distance from a variant position.
[0068] FIG. 23B provides a masking strategy used for training and deployment of the disclosed model for genotype prediction, in accordance with some embodiments of the present disclosure.
[0069] FIG. 24A provides a non-limiting schematic illustration of the inference mask used with the disclosed model to predict bases in a genotype prediction task, in accordance with some embodiments of the present disclosure.
[0070] FIG. 24B provides a non-limiting example of base prediction accuracy data for deployment of the disclosed model for genotype prediction.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-114 of 94
[0071] FIG. 25 provides a non-limiting schematic illustration of use of the disclosed model for performing a “linking attack” to identify an individual based on a predicted ATAC-seq data-derived genotype.
[0072] FIG. 26 provides a non-limiting example of a histogram showing the number of individuals versus the log likelihood for a given individual exhibiting a predicted ATAC-seq data-derived genotype.
[0073] FIG. 27 provides a non-limiting examples of data for linking accuracy as a function of the number of sub-sampled sequence reads utilized in the ATAC-seq data-derived genotype prediction.
[0074] FIG. 28 provides a non-limiting schematic illustration of deployment of the disclosed model for design of tissue type-specific DNA regulatory sequences, in accordance with some embodiments of the present disclosure.
[0075] FIGS. 29A-29B provide non-limiting schematic illustration of the masking strategy used with the disclosed model for design of tissue type-specific DNA regulatory sequences, in accordance with some embodiments of the present disclosure. FIG. 29A: masking strategy used for model training. FIG. 29B: masking strategy used for inference.
[0076] FIG. 30 provides a non-limiting examples of data illustrating repeated predictions of DNase-seq data profiles for three different cell types for the indicated input data prompt, and the averaged output profiles for the three different call types.
[0077] FIG. 31 provides a non-limiting examples of data illustrating repeated predictions of DNase-seq data profiles for three different cell types for the indicated input data prompt, and the averaged output profiles for the three different call types.
[0078] FIG. 32 provides a non-limiting examples of data illustrating repeated predictions of DNase-seq data profiles for three different cell types for the indicated input data prompt, and the averaged output profiles for the three different call types.
[0079] FIG. 33 provides a non-limiting examples of data illustrating repeated predictions of DNase-seq data profiles for three different cell types for the indicated input data prompt, and the averaged output profiles for the three different call types.
[0080] FIG. 34 provides a non-limiting examples of data illustrating repeated predictions of DNase-seq data profiles for three different cell types for the indicated input data prompt, and the averaged output profiles for the three different call types.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-115 of 94
[0081] FIG. 35 provides a non-limiting example of data comparing cell type-specific sequence motifs identified in designed sequences predicted to generate the desired cell typespecific functional genomic property profiles.
[0082] FIG. 36 provides a non-limiting example of data providing a comparison of the predicted versus measured functional activity of generated DNA regulatory sequences based on stratification of the peaks in their corresponding functional genomics property profiles.
[0083] FIG. 37 provides a non-limiting example of data illustrating a comparison of functional genomics property profiles for actual DNA regulatory sequences and generated DNA regulatory sequences in three different cell types.
[0084] FIG. 38 provides a non-limiting example of a histogram showing the minimum Levenshtein distance (edit distance) between generated DNA regulatory sequences or test DNA regulatory sequences and all DNA sequences in the model training dataset.
[0085] FIG. 39 provides a non-limiting schematic illustration of a functional genomics prediction model trained to predict K562 DNase and CAGE-seq data from DNA sequence and H3K27ac ChlP-seq data in K562 cells.
[0086] FIG. 40A provides a non -limiting example of reference H3K27ac ChlP-seq data for K562 cells.
[0087] FIG. 40B provides a non-limiting example of H3K27ac ChlP-seq data for an in silico CRISPRi enhancer knockout generated by replacing the reference signal at the enhancer midpoint with background signal.
[0088] FIG. 41 provides non-limiting examples of input reference H3K27ac ChlP-seq data for K562 cells (upper panel), and predicted CAGE-seq data output by the model (lower panel).
[0089] FIG. 42 provides non-limiting examples of input H3K27ac ChlP-seq data for an in silico CRISPRi enhancer knockout in K562 cells (upper panel), and predicted CAGE-seq data output by the model (lower panel).
[0090] FIG. 43 provides a non-limiting example of a block diagram of a computer system, in accordance with some embodiments of the present disclosure.
[0091] FIG. 44 provides a non-limiting example of a block diagram of an example artificial intelligence (Al) architecture included as part of an example computing system, in accordance with some embodiments of the present disclosure.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-116 of 94
[0092] In the appended figures, similar components and / or features can have the same reference label. Further, various components of the same type can be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.DETAILED DESCRIPTION
[0093] Machine learning-based methods, systems, and programming for training and deploying a novel multimodal, multi-masking-based deep learning model configured to predict cell type-specific sequences and / or functional genomic traits are described. The disclosed model utilizes a new training strategy comprising: (i) self-supervised learning on DNA sequence, and (ii) supervised learning on functional genomic data, along with the use of a multiscale loss function. A key feature of the disclosed methods and systems is the ability of the core deep learning model to be reconfigured for different functional genomics prediction applications by changing the masking strategy used during training and / or inference.
[0094] The methods and systems described herein provide numerous technical advantages in terms of improved predictive and / or computational performance that derive from the model architecture and training approach used. For example, a single core model can be used to predict sequences and / or functional genomic properties (at single nucleotide resolution) from other functional genomic property data (with or without corresponding input DNA sequence data), and to denoise functional genomics measurements - applications that typically require separate machine learning models.
[0095] In sone instances, for example, the disclosed methods for predicting a functional genomic property can comprise: obtaining an input genomic sequence derived from a test sample; obtaining at least one input functional genomic property dataset derived from the test sample; formulating an input data structure based on the input genomic sequence and the at least one input functional genomic property dataset; and predicting at least one output functional genomic property for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by: obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating aMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-117 of 94 plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the plurality of masked training input data structures.
[0096] In some instances, the disclosed method for predicting a genomic sequence and / or a functional genomic property can comprise: formulating an input data structure comprising a genomic sequence component and / or at least one functional genomic property component, wherein at least a portion of the genomic sequence component and / or the at least one functional genomic property component is unknown; inputting the input data structure into a trained machine learning model configured to predict the unknown portion of the genomic sequence component and / or the at least one functional genomic property component, the machine learning model trained by: obtaining genomic sequences and / or the at least one functional genomic property dataset derived from each of a plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and / or the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the genomic sequence and / or at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model to receive a masked training input data structure and predict the masked portion of the masked training input data structure; and outputting the predicted portion of the genomic sequence and / or the at least one functional genomic property for the test sample.
[0097] In some instances, the disclosed methods for predicting a base at at least one specified genomic locus in a genomic sequence can comprise: obtaining the genomic sequence derived from a test sample, the genomic sequence comprising an unknown base at the specified genomic locus; obtaining at least one functional genomic property dataset derived from the test sample; formulating an input data structure based on the genomic sequence and the at least one functional genomic property dataset; and predicting the base at the at least one specified genomic locus in the genomic sequence for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by: obtainingMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-118 of 94 genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask at least one base in the genomic sequence of each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the masked training input data structures.
[0098] In some instances, the disclosed methods for designing a genomic regulatory element can comprise: obtaining at least one functional genomic property dataset generated based on a desired regulatory effect in a test sample; formulating an input data structure based on an unknown genomic sequence of specified length and the at least one functional genomic property dataset; predicting the unknown genomic sequence by: inputting the input data structure into a trained machine learning model configured to output a prediction of a base at one or more positions of an unknown genomic sequence; and adding the predicted base at the one or more positions of the unknown genomic sequence to the input data structure; outputting the predicted genomic sequence; and designing the genomic regulatory element for the test sample based on the predicted genomic sequence.
[0099] Due to improved predictive performance, the disclosed machine learning-based methods and model can be used in a variety of important biotechnology and pharmaceutical industry applications. The novel architectural and training features of the disclosed model enable novel application modes including, but not limited to, flexible input sequence length, in-context sequence prediction and design, functional perturbation prediction, and function-to- sequence prediction. These capabilities can, in turn, improve the efficacy of downstream development and manufacturing processes in the biotechnology and pharmaceutical industries (e.g. , by facilitating the design and development of effective gene therapies (e.g. , gene therapies directed to turning off expression of mutated genes, or to promoting expression of beneficial genes), CRISPR-based drug delivery mechanisms (e.g., to improve the efficacy and precision of gene editing tools), and protein expression systems for biologies manufacturing (e.g., to maximize yields of expressed proteins), etc.).MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-119 of 94
[0100] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention, and is provided in the context of a patent application and its requirements.I. Sequence-based Prediction Models and Sequence+Function-based Prediction Models
[0101] Existing state-of-the-art DNA sequence-based models for prediction of functional genomics properties comprise disparate models trained for specific applications, e.g., models that process DNA sequences to make predictions of functional genomics properties (e.g., regulatory activities) in different cell types. These models were initially configured to predict a single value (e.g., a binary or real -valued number) for a given regulatory activity at a given genomic locus. More recently, sequence-based prediction models for functional genomics have moved towards what may be known as profile models. Instead of predicting a single value at a single genomic locus, the models predict entire genomic profiles for functional genomic property data (e.g., plots that show experimentally-determined and / or predicted values for the number of sequence reads identified as a function of genomic locus for functional genomic assay data (e.g., ATAC-seq data, CAGE-seq data, Histone H3-Lysine 27 acetylation (H3K27ac-seq) data, etc. . These functional genomic property profiles can provide an illustration of how the DNA sequence is modified (or decorated) at different genomic loci to regulate gene expression, etc. These DNA sequence-based models can work on sequences as long as a few tens of thousands of base pairs (bp) in length.
[0102] Such models can then be used, for example, to predict the functional effects of genetic variants, e.g., by inserting the variant of interest into the input sequence and comparing the predicted effect on function for the variant to that for the wild-type or reference sequence. These models can also be used as “oracles” to design DNA sequences that have (or confer) a desired functional genomic profile in a given cell type. For example, if one wants to design a DNA sequence that has a functional genomic profile (e.g. , for transcription factor (TF) binding, or chromatin accessibility) comprising an activity of “one” at a first genomic locus and an activity of “zero” at another genomic locus one can propose a hypothetical DNA sequence and run different permutations of the proposed sequence through the model until the proposed sequence is predicted to convey the desired functional genomic properties.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-120 of 94
[0103] There are also variations of these sequence-based models that, instead of accepting genomic sequences as input, can also use a combination of sequence and other functional genomic data (e.g., functional genomic data profiles or “tracks”) as input and use that input to predict some other functional genomic property. That is, these models change the sequence-to- functional genomic property prediction paradigm into a sequence-plus-function-to-different functional genomic property prediction paradigm.
[0104] Finally, sequence-based masked language models may be similar to, e.g., ChatGPT, where the system can take a text-based sequence as input with some words masked and then predict the masked words from the text sequence. Similar masking may be performed using genomic sequences (e.g., DNA sequences), where certain bases of the DNA sequence are masked in the input, and the model is trained to predict the bases that were missing in the input. These models have been trained on input DNA sequences of different length, and can be trained on sequences from multiple species.II. Multimodal, Masked Model for Functional Genomics Applications
[0105] The disclosed functional genomics prediction model unifies the many disparate prior art models into a single framework that draws inspiration from so-called text-vision masked multimodal models in which, rather than masking within just one modality (e.g., either text or images), one can mask across multiple modalities (e.g., text and / or images).
[0106] The disclosed modelling method comprise a multimodal masked-base model for genomics sequence and / or functional genomic property prediction. The multimodal masked model concept can be extended to functional genomics applications by stacking different tracks of genomic sequence and / or functional genomics property datasets (i.e., different data modalities), masking some portion of the input data, and training the model to predict the missing portions. The human genome is very long (about 3 billion bases of DNA in total, split into different chromosomes with the longest chromosome being about 250 million bases in length). In some instances, the disclosed functional genomics prediction model can be trained using segments of genomic sequence ranging in length from about 10,000 bases up to about 100,000 bases.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-121 of 94III. Example Prediction System
[0107] FIG. 1 is a block diagram of an example prediction system, in accordance with some embodiments. Prediction system 100 can be used to determine a predicted DNA property, such as a predicted sequence, a predicted expression level, a predicted functional property, and / or predicted impact of a sequence variant on gene expression for a given input DNA sequence, e.g., a gene sequence, and / or given input functional genomics property profile(s). The prediction system 100 can include, e.g., computing platform 102, data store 104, and display system 106. Computing platform 102 may take any of a variety of forms. In some embodiments, for example, the computing platform 102 can include a single computer (or computer system), or multiple computers in communication with each other. In some embodiments, the computing platform 102 can be a cloud computing platform.
[0108] Data store 104 and display system 106 are each in communication with computing platform 102. In some examples, one or more of data store 104 and / or display system 106 can be considered part of, or otherwise integrated with, computing platform 102. Thus, in some examples, computing platform 102, data store 104, and display system 106 can be separate components in communication with each other, but in other examples, some combination of these components can be integrated together. Communication between the different components can be implemented using any number of wired communications links, wireless communications links, optical communications links, or a combination thereof.
[0109] The prediction system 100 can include a sequence analyzer 108, which can be implemented using hardware, software, firmware, or a combination thereof. In some embodiments, the sequence analyzer 108 is implemented as part of computing platform 102. The sequence analyzer 108 receives input data, e.g., sequence data 110 and / or functional data 111, for processing. For example, sequence data 110 can be provided as input into the sequence analyzer 108, can be retrieved from data store 104 or some other type of storage (e.g., cloud storage), can be accessed from cloud storage, or can be obtained in some other manner. In some cases, sequence data 110 can be retrieved from data store 104 in response to receiving user input entered by a user via an input device. Similarly, functional data 111 can be provided as input into the sequence analyzer 108, can be retrieved from data store 104 or some other type of storage (e.g., cloud storage), can be accessed from cloud storage, or can be obtained in someMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-122 of 94 other manner. In some cases, functional data 111 can be retrieved from data store 104 in response to receiving user input entered by a user via an input device.
[0110] In some embodiments, the sequence data 110 can be generated by processing a set of samples 112. The set of samples 112 may take the form of one or more biological samples from one or more subjects (e.g., a diseased sample, a healthy sample, or a combination thereof). In some embodiments, the set of samples 112 may include a sample obtained from a tumor of a subject. The tumor can be a manifestation of, for example, lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myelogenous leukemia, chronic myelogenous leukemia, chronic lymphocytic leukemia, T cell lymphocytic leukemia, non-small cell lung cancer, small-cell lung cancer, or a combination thereof. In some embodiments, the set of samples 112 may include a sample obtained from a subject diagnosed with a disease.[oni] In some embodiments, a sample in the set of samples 112 may include, for example, various nucleic acid molecules, e.g., various deoxyribonucleic acid (DNA) molecules, various ribonucleic acid (RNA) molecules (e.g., various messenger RNA (mRNA) molecules, various ribosomal RNA (rRNA) molecules, or various transfer RNA (tRNA) molecules), or a combination thereof. When the set of samples 112 includes a diseased sample, the nucleic acid molecules may include one or more mutant nucleic acid sequences (e.g., mutant gene sequences, mutant regulatory sequences, mutant promoter sequences, mutant mRNA sequences, etc.).
[0112] In some embodiments, the set of samples 112 may comprise biological samples that are processed to extract nucleic acid molecules and generate sequence data 110. In some embodiments, multiple samples in the set of samples 112 can be processed at the same or at different times. In some embodiments, the prediction system 100 includes a sample analyzer that is used in processing the set of samples 112 to generate the sequence data 110.
[0113] In some embodiments, a sample in the set of samples 112 may include synthetic nucleic acid molecules, e.g., synthetic DNA molecules, synthetic RNA molecules (e.g., synthetic mRNA molecules, synthetic rRNA molecules, or synthetic tRNA molecules), or a combination thereof.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-123 of 94
[0114] In some embodiments, a sample in the set of samples 112 may include designed or in silico nucleic acid sequences, e.g. , designed or in silico DNA sequences, designed or in silico RNA sequences (e.g., designed or in silico mRNA molecules, designed or in silico rRNA molecules, or designed or in silico tRNA molecules), or a combination thereof.
[0115] In some embodiments, a sample in the set of samples 112 may include, e.g., an extracted, synthetic, designed, or in silico DNA sequence 129. In some embodiments, the DNA sequence 129 may correspond to a gene or a sequence that comprises a 5’ end subsequence 120, a central subsequence 118, and / or a 3’ end subsequence 122. In some embodiments, the 5’ end subsequence 120 may correspond to a 5’ (upstream or start) sequence 128. In some embodiments, the central subsequence 118 may correspond to a gene coding sequence (CDS) 126. In some embodiments, the 3’ end subsequence may correspond to a 3’ (downstream or stop) sequence 130. One or more sub-sequences of, e.g., DNA sequence 129 (e.g., 5’ sequence 128, coding sequence 126, or 3’ sequence 130) can be processed separately or as a single sequence.
[0116] In some embodiments, DNA sequence 129 encodes for at least a portion of a corresponding protein. In some embodiments, DNA sequences (e.g., gene sequences), mRNA sequences, and / or protein sequences, or mutants thereof, can be identified by comparison to a corresponding reference sequence and / or by performing a search in an appropriate database (e.g., the RefSeq NCBI Reference Sequence Database, etc.) based on the DNA sequence, mRNA sequence, and / or protein sequence data obtained from the sample.
[0117] In some embodiments, the functional data 111 can be generated by processing a set of samples 112 or a set of samples that is different from the set of samples 112. As noted above, the set of samples (i.e., a set of samples that is that same as, or different from, the set of samples 112) may take the form of one or more biological samples from one or more subjects (e.g., a diseased sample, a healthy sample, or a combination thereof). In some embodiments, the set of samples 112 may include a sample obtained from a tumor of a subject.
[0118] In some embodiments, the functional data 11, may be derived by processing a set of samples and performing a functional genomics assay. Examples of functional genomics assays can include, but are not limited to, RNA sequencing (RNA-seq), chromatin immunoprecipitation (ChIP) followed by sequencing (ChlP-seq), and DNA accessibility assays (e.g., ATAC-seq, DNase-seq, or Formaldehyde-Assisted Isolation of Regulatory ElementsMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-124 of 94 sequencing (FAIRE-seq) assays), a CAGE-seq assay, an H3K27ac-seq assay, a methyl-seq assay, etc.
[0119] Sequence analyzer 108 receives the sequence data 110 and / or functional data 111 as input for processing. The sequence analyzer 108 can include one or more machine-learning model(s) 132 that process the sequence data 110 and / or functional data 111. In some embodiments, the sequence data 110 and / or functional data 111 is sent directly into the one or more machine-learning model(s) 132 for processing. In some embodiments, the sequence analyzer 108 preprocesses the sequence data 110 and / or functional data 111 prior to sending the sequence data 110 and / or functional data 111 to the one or more machine-learning model(s) 132 for processing. Pre-processing the sequence data 110 and / or functional data 111 may include, e.g., formulating an input data structure comprising DNA sequence data and / or one or more corresponding functional genomics property tracks.
[0120] The one or more machine-learning model(s) 132 can be implemented in any of a number of different ways. In some embodiments, a machine learning model of the one or more machine-learning model(s) 132 can be any type of model that uses a set of element-focused scores that represent properties of a set of nucleic acid sequence representations. The one or more machine-learning model(s) 132 can be trained a training mode or deployed in a prediction mode. In the training mode, the one or more machine-learning model(s) 132 are trained using training data 133. Examples of the training data 133 are described in more detail below. The machine-learning model 132 is trained such that it can be deployed and used in the prediction mode.
[0121] The one or more machine-learning model(s) 132 process the sequence data 110 (e.g., DNA sequence 129) and / or functional data 111 (e.g., DNase-seq data 122, ChlP-seq data 123, ATAC-seq data 124, and / or CAGE-seq data 125) via, a multi-modal, masked data processing engine, e.g., a multi-modal, masked DNA sequence + functional data processing engine 138. In some embodiments, separate DNA sequence processing engines for the coding sequence 126, the 5’ sequence 128, and / or the 3’ sequence 130 can enable improved predictive performance. Non-limiting examples of implementations for these different processing engines are described in greater detail below.
[0122] As used herein, the terms “processing engine” and “engine” identify at least one software component and / or a combination of at least one software component and at least oneMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-125 of 94 hardware component which are designed / programmed / configured to interact and / or communicate data to other software and / or hardware components including but not limited to other processing engines.
[0123] The one or more machine-learning model(s) 132 process the sequence data 110 and / or functional data 111 to generate an output that can be used, e.g., to generate a report 144. The report 144 may include the exact output of the one or more machine-learning model(s) 132, a transformed or filtered version of the output, or both. In some cases, the report 144 may include notifications, recommendations, alerts, or other information generated by the sequence analyzer 108 based on the output of the one or more machine-learning model(s) 132.
[0124] The report 144 can be an output that includes, for example, predicted nucleic acid sequence properties, e.g., a predicted DNA property, such as a predicted base, a predicted sequence, a predicted gene expression level, a predicted gene functional property, and / or a predicted impact of a sequence variant on gene expression for a given DNA sequence, e.g., a gene sequence, with respect to one or more input sequences.
[0125] In some embodiments, a report 144 can be displayed on a graphical user interface (GUI) 150 on the display system 106. A user may view and / or interact with the report 144 via the graphical user interface 150. In some embodiments, the user may use the report 144 to make decisions about, e.g., the development of a gene-based therapy, the development of a CRISPR- based delivery system, or the treatment of a subject from which at least one of the set of samples 112 was obtained (or collected).
[0126] In some embodiments, the prediction system 100 sends the report 144 to the remote system 152 (e.g., wirelessly). The remote system 152 can be a cloud computing platform, cloud storage, another computer system, a user device (e.g., a smartphone, a tablet, a laptop, etc.), or some other type of platform. In some embodiments, the remote system 152 can be a treatment manufacturing system (or machine) or a portion thereof.IV. Example Model Architecture
[0127] The machine-learning model 132 of the embodiments described herein may include multiple subsystems (or subnetworks). Each of the multiple subsystems can include, for example, an encoder, a transformer encoder, and / or one or more processing layers.
[0128] Machine-learning model 132 may include one or more encoders configured to, for example, transform an element of the input sequences (e.g, an amino acid sequence, a nucleicMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-126 of 94 acid sequence, a codon sequence, etc.) based on the other elements of the input sequences. An encoder can be a transformer encoder.
[0129] The machine-learning model 132 may include one or more processing layers such as self-attention layers or convolution layers, or a neural network such as a long-short term memory unit (LSTM), recurrent structure, or recurrent component. The machine-learning model 132 can implement, for example, one or more self-attention layers. Machine-learning model 132 can use a self-attention mechanism, global attention mechanism, soft attention mechanism, local attention mechanism, and / or hard attention mechanism. In some instances, the machine-learning model 132 does not include any convolutional layer, any recurrent structure, any LSTM unit, and / or any recurrent component. In some instances, the machine learning model 132 is not a recurrent machine-learning model and / or does not include a recurrent neural network. In some instances, the machine-learning model 132 includes a recurrent neural network and / or may use positional encoding to handle sliding windows of sequence elements across one or more sequences. In some instances, the machine-learning model 132 is not a convolutional machine-learning model and / or does not include a convolutional neural network.
[0130] The machine-learning model 132 may include processing blocks, such one or more first processing blocks used to process, e.g., one or more nucleic acid sequence representations independent from a second processing block used to process, e.g., a representation of a functional genomics assay data profile. In some embodiments, the second processing block may process part or all of, e.g., a functional genomics assay data profile. The independence of these processing blocks can facilitate parallel processing when using the machine-learning model 132. Further, the independence may improve the performance (e.g., accuracy of predictions) of the machine learning model 132.
[0131] The machine-learning model 132 can be configured such that an output value at any given layer depends not only on a corresponding input value but also on one or more (e.g., all) other input values. Thus, the machine-learning model 132, a loss function, and / or an optimization function can be configured to optimize an output corresponding to a single position representing, e.g., a level of gene expression for a given input DNA sequence for a specified gene. In some instances, the loss function can comprise supervised loss function such as binary cross entropy or unsupervised loss functions. In some instances, the unsupervisedMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-127 of 94 loss functions can include a contrastive loss or regularization losses (e.g., L1 / L2 losses). In some instances, the loss function can comprise auxiliary loss functions. In such instances, the auxiliary loss function can be used alongside a main loss function to train the machine learning model 132. Accordingly, in some instances, the auxiliary loss function can improve the learning process by adding additional information or constraints. In some instances, any of a plurality of outputs of transformer encoders may represent such an occurrence probability. The machine-learning model 132 can be trained accordingly. In some instances, an endpoint (e.g., surplus endpoint) may represent (in response to training) a gene expression level and / or a functional genomic property probability or likelihood. Aggregated outputs can be, for example, fed to another layer, subsystem, or processing block (e.g., that includes one or more of: a processing layer such as a self-attention layer, or an encoder such as a transformer encoder).
[0132] In some instances, one, two, or all dimensions of an output from another layer and / or another subsystem or processing block is the same size as the input fed to the other layer and / or other subsystem or processing block. In some instances, an input fed to this other layer and / or other subsystem or processing block has a length along one axis that is greater than or equal to a sum of one or more of, for example, a number of nucleotides in a gene sequence, a number of nucleotides in an 5’ end flanking sequence, or a number of nucleotides in a 3’ end flanking sequence. In some instances, the length of the input is one longer than the total number of nucleotides. The length of the input along the one axis may exceed the summed count of nucleotides when, for example, an additional feature vector is appended to the nucleotidespecific feature values. Another dimension of the input can include a number of features (e.g., determined via a hyperparameter). An output generated by the other layer and / or other subsystem or processing block may have the same size as the input.
[0133] A subset of values of the output generated by the other layer and / or other subnetwork can be processed by another neural network (e.g., a fully connected feedforward network). The subset of values may include a 1 -dimensional vector of values that may correspond to one set of feature values.
[0134] In some embodiments, a neural network within the machine-learning model 132 can be configured to output one or more results. The one or more results can include, for example, a numeric result, binary result, and / or categorical result. Each of the one or more results can predict whether and / or an extent to which a nucleic acid sequence (e.g., a gene sequence) isMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-128 of 94 likely to be expressed in a specified cell type and / or a specified cell state. The machine-learning model 132 may include one or more activation layers to produce an intermediate result (e.g., to transform a real-number interim value into a binary and / or categorical output). The machinelearning model 132 can be trained to generate multiple types of predictions (e.g., gene expression predictions and / or functional genomic property predictions). In some instances, a prediction can be binary or categorical. Other predictions can be non-binary or non-categorical. For example, a prediction can be scalar.
[0135] Machine-learning model 132 may include and / or can be included within an ensemble model. The ensemble model may include multiple (e.g., identical) sub-models that can be trained using different portions of the training dataset.
[0136] FIG. 2 provides a non-limiting, schematic illustration of a machine learning model architecture, in accordance with some embodiments of the present disclosure. The machine learning model architecture depicted in FIG. 2 can be, in some examples, based on a foundational model (e.g., the Borzoi model for prediction of RNA-seq coverage (Linder et al. (2023), “Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation”, bioRxiv preprint, http s: / doi.oi 1.0 .l 01 / 2023.08.30.555582) or the Enformer model (Avsec et al. (2021), “Effective gene expression prediction from sequence by integrating long-range interactions”, Nature Methods 18(10): 1196-1203). The Borzoi model is a sequence-to-function only model, and is based on a neural network architecture comprising a stack of convolution and downsampling layers, followed by a series of self-attention layers with relative positional encodings operating at 128 bp resolution. Repeated application of the convolution block achieves a 2-fold reduction of the sequence length and extracts local sequence patterns until each position in the sequence represents 128 bp. Then, repeated application of the self-attention (or transformer) block enables long-range interaction and exchange between every pair of sequence positions. The output is then up-sampled through a number of deconvolutional layers with matched U-net connections.
[0137] For the model disclosed herein, the model input is configured differently from that used in prior art models (e.g., the Borzoi model). Specifically, the model is configured to accept an input data structure comprising an input DNA sequence and / or one or more functional genomics property data tracks. The model also includes an additional input channel for each DNA or functional channel that represents the positions for that input sequence or functionalMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-129 of 94 genomic property track that have been masked. In some instances, the model can be configured to use separate convolution blocks for the DNA sequence and functional genomic property tracks, where the separate convolution blocks may be inserted before the transformer layer. Also, the output layer of the model is configured to output predicted DNA sequences and / or functional genomic property profiles at single base pair resolution.V. Example Model Training Process
[0138] The training of the disclosed model utilizes a new strategy comprising: (i) selfsupervised learning on DNA sequence data, and (ii) supervised learning on functional genomic data, along with the use of a multiscale loss function.
[0139] The weighting factors, bias values, and threshold values, or other computational parameters of the neural network (or other machine learning architecture), can be "taught" or "learned" in a training phase using one or more sets of training data (e.g., 1, 2, 3, 4, 5, or more than 5 sets of training data) and a specified training approach configured to solve, e.g., minimize, a loss function. A loss function can use an error term (e.g., mean squared error or median squared error) and / or an entropy term (e.g., cross entropy or binary cross entropy). The adjustable parameters for the neural network (e.g., deep learning model) may be determined based on input data from a training dataset using an iterative solver (such as a gradient-based method, e.g., backpropagation), so that the output value(s) that the neural network computes (e.g., a prediction of gene expression level in a given cell type) are consistent with the examples included in the training dataset. In some instances, multitask learning can be used, such that the model is simultaneously trained to predict each of two different types of results (e.g., gene expression level and mRNA stability). A static or non-static learning rate can be used. For example, learning rate annealing (e.g., using stepwise annealing or cosine annealing) can be used to reduce the learning rate over iterations. Validation-data assessment can be used to potentially terminate training early (e.g., upon determining that a performance target has been met).
[0140] FIG. 3 provides a non-limiting example of a flowchart for a process 300 for training a multi-modal, multi-masking-based deep learning model configured to process input sequence and / or functional genomic property data and output sequence and / or functional genomics property predictions. Process 300 can be performed, for example, using the prediction system 100 in FIG. 1 to train machine-learning model 132. In some instances, part or all of processMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-130 of 94300 can be performed at a remote computing system that is remote relative to a user device and / or laboratory. The remote computing system can be a cloud computing system.
[0141] At step 302 in FIG. 3, input training data structures comprising DNA sequence data and / or functional genomics property data are formulated.
[0142] In some instances, the training data comprises DNA sequences (e.g., gene sequences, or gene sequences comprising 5’- and / or 3’ end flanking sequence regions) at base pair resolution from a training data set comprising cell type-specific DNA sequences.
[0143] In some instances, the training data set can comprise from about 10,000 to about 1,000,000 DNA sequences (depending on the length of the sequences) and / or corresponding functional genomics property tracks. For example, in some instances the training data set can comprise at least 10,000, at least 50,000, at least 100,000, at least 150,000, at least 200,000, at least 300,000, at least 400,000, at least 500,000, at least 600,000, at least 700,000, at least 800,000, at least 900,000, or at least 1,000,000 DNA sequences and / or corresponding functional genomics property tracks. In some instances, the training data set can comprise at most 1,000,000, at most 900,000, at most 800,000, at most 700,000, at most 600,000, at most 500,000, at most 400,000, at most 300,000, at most 200,000, at most 150,000, at most 50,000, or at most 10,000 DNA sequences and / or corresponding functional genomics property tracks. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the training data set can comprise between about 10,000 and about 300,000 DNA sequences and / or corresponding functional genomics property tracks. Those of skill in the art will recognize that the training data set can comprise a number of DNA sequences having any value within this range, e.g., about 220,300 DNA sequences and / or corresponding functional genomics property tracks.
[0144] In some instances, the DNA sequences of the training data set may range in length from about 100 base pairs (bp) to about IM bp. In some instances, the DNA sequences of the training data set may have a length of at least 100 bp, at least 500 bp, at least 1,000 bp, at least 5,000 bp, at least 10,000 bp, at least 50,000 bp, at least 100,000 bp, at least 200,000 bp, at least 300,000 bp, at least 400,000 bp, at least 500,000 bp, at least 600,000 bp, at least 700,000 bp, at least 800,000 bp, at least 900,000 bp, or at least IM bp. In some instances, the DNA sequences of the training data set may have a length of at most IM bp, at most 900,000 bp, at most 800,000MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-131 of 94 bp, at most 700,000 bp, at most 600,000 bp, at most 500,000 bp, at most 400,000 bp, at most 300,000 bp, at most 200,000 bp, at most 100,000 bp, at most 50,000 bp, at most 10,000 bp, at most 5,000 bp, at most 1,000 bp, at most 500 bp, or at most 100 bp. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances, the DNA sequences of the training data set may have a length of between about 5,000 bp and about 200,000 bp. Those of skill in the art will recognize that the DNA sequences of the training data set can have a length of any value within this range, e.g., about 20,754 bases.
[0145] In some instances, the training data comprises one or more functional genomic property data profiles (or tracks) (e.g., 1, 2, 3, 4, 5, 6, 7, 8,9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000 or more than 10,000 tracks) corresponding to each of the DNA sequences in the training data set. In some instances, the one or more cell type-specific functional genomic property data profiles can comprise, e.g, RNA-seq data, ChlP-seq data, ATAC-seq data, CAGE-seq data, DNase-seq data, FAIRE-seq data, H3K27ac-Seq data, High-Throughput Chromosome Conformation Capture (Hi-C) data, Micrococcal Nuclease Chromosome Conformation Capture (Micro-C) data, Massively Parallel Reporter Assay (MPRA) data, methyl-seq data, etc., at base pair resolution. In some instances, the length of the functional genomics data tracks in the training data set may be equal to, longer, shorter, or similar in length to the length of the DNA sequences in the training data set.
[0146] In some instances, the training data can comprise DNA sequences and / or functional genomic property data profiles (or tracks) for at least 1 cell type, at least 2 cell types, at least 3 cell types, at least 4 cell types, at least 5 cell types, at least 6 cell types, at least 7 cell types, at least 8 cell types, at least 9 cell types, at least 10 cell types, at least 20 cell types, at least 30 cell types, at least 40 cell types, at least 50 cell types, at least 60 cell types, at least 70 cell types, at least 80 cell types, at least 90 cell types, at least 100 cell types, at least 125 cell types, at least 150 cell types, at least 175 cell types, or at least 200 cell types. In some instances, the training data can comprise DNA sequences and / or functional genomic property data profiles (or tracks) for any number of cell types, e.g, any number of cell types within the range of values described in this paragraph.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-132 of 94
[0147] In some instances, the input training data structure comprising DNA sequence data and / or functional genomics property data at base pair resolution may be formulated as a matrix, e.g., an L x (4+T) matrix, where L is the length of the DNA sequence (genomic context), T is the number of functional genomics property tracks, and the DNA sequence is one-hot encoded.
[0138] In some instances, the genomic sequence data can be provided as input to the machine learning model as a one-hot encoded matrix. One-hot encoding is a technique used in machine learning to convert categorical data (e.g., colors, or the bases in a DNA sequence) into a binary vector format that can be processed by a machine learning algorithms. Each category is represented as a binary vector where one bit has a value of “1” (hot) and the remaining bits have a value of “0” (cold).
[0148] In some instances, the training data set comprising DNA sequences and / or functional genomics property profile data may be split into training, validation, and test datasets (e.g., using an 80%, 10%, 10% split, respectively, of the initial training dataset, or any other split ratio known to those of skill in the art).
[0149] At step 304 in FIG. 3, the training data structures are masked for input into the model, where the configuration of the mask indicates the values to be predicted by the trained model (e.g., a predicted DNA base at a specified genomic locus, or a predicted functional genomics property). In some instances, the same masks may be used for training and inference. In some instances, different masks may be used for training and inference.
[0150] At step 306 in FIG. 3, iterative training of the model is performed, where the iterative training process includes steps of: (i) performing self-supervised learning on the DNA sequence training data, (ii) performing supervised learning on the functional genomic property training data, (iii) evaluating a loss function (e.g., a multiscale Poisson loss function) with respect to current predictions for all or a subset of the masked input data, (iv) updating model weights and biases, and (v) evaluating a termination condition. Self-supervised learning on the DNA sequence training data and supervised learning on the functional genomic property data can happen simultaneously. In the most general case, part of the DNA sequence may be masked in the input and some part of the functional tracks may also be masked. On the output side, both are predicted. The loss is computed on the output for both DNA and tracks at specified positions. In some instances, for example, a Poisson loss function can be used for the functionalMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-133 of 94 genomic property tracks and a cross entropy loss function can be used for the DNA sequence data.
[0151] During the training process, a Poisson loss function, for example, may be evaluated at multiple scales (or resolutions), as illustrated in FIG. 2. Intermediate prediction vectors, e.g., PL, 2, and pi, comprising predicted outputs at 2L-1bp, 21bp, and 1 bp resolution, respectively, are used to evaluate the Poisson loss function according to, e.g., log Pots (yL1 / )where L is the length of the input sequence training data, and yLis the input training data vector. In some instances, a Poisson loss function at single base pair resolution may be used instead of a multi-scale Poisson loss function.
[0152] Prior to beginning the iterative training process, model hyperparameters (z.e., settings such as the learning rate in neural networks, the number of hidden layers, or regularization parameters, that are set to control the model's structure, behavior, and learning process) and / or model weights and / or biases (ie., learned parameters) may be initialized. In some instances, model weights and / or biases can be initialized from scratch, e.g., using a random set of values. In some instance, model weights and / or biases can be initialized using the weights and / or biases from a trained foundational model. In some instances, all or a portion of the model (e.g., all or a portion of the layers in the model) may be trained and / or fine-tuned, (e.g., by training the model as a whole on a first set of training data, and then fine-tuning specified layers using a second set of training data). In some instances, model hyperparameters may be optimized by repeating several rounds of the iterative model training process, where the hyperparameters are adjusted in one or more rounds.
[0153] The iterative model training process can be repeated until, e.g., a specified termination condition is met. Examples of suitable termination conditions include, but are not limited to, an early stopping criterion (e.g., where the model’s performance on a validation dataset starts to degrade or ceases to improve significantly for a predefined number of consecutive iterations), reaching a fixed number of iterations, reaching a point where improvement in the loss function becomes insignificant between iterations, or the change in the weight vector (comprising values of each learned weight) between iterations falls below a predefined threshold.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-134 of 94
[0154] At step 308 in FIG. 3, optimal values of model weights and biases are output. In some instances, the model’s state (z.e., the set of weights and biases) at the point of optimal performance on a validation dataset (z.e., before decline or stagnation of performance sets in) can be restored and those values used as the final configuration of the model.
[0155] In some instances, the training dataset (e.g., paired sets of genomic sequence data and corresponding functional genomics properties) can be randomly parsed, shuffled, and / or divided to train various models within an ensemble of models.
[0156] In some instances, the disclosed models and methods may comprise retraining (or “fine-tuning”) a previously trained machine learning models (e.g., by iteratively retraining a previously trained model using one or more training datasets that differ from those used to train the model initially). In some instances, retraining the machine learning model may comprise using a continuous, e.g., online, machine learning model, i.e., where the model is periodically or continuously updated or retrained based on new training data. The new training data may be provided by, e.g., a single deployed local operational system, a plurality of deployed local operational systems, or a plurality of deployed, geographically-distributed operational systems. In some instances, the disclosed methods may employ, for example, pre-trained neural networks, and the pre-trained neural networks can be fine-tuned using additional dataset(s) to tune the pre-trained neural network for the same, or a different, prediction objective.
[0157] The training of the model (i.e., determination of the adjustable parameters of the model using an iterative solver) may or may not be performed using the same hardware as that used for deployment of the trained model.VI. Prediction of DNA Sequences and / or Functional Genomics Properties
[0158] FIG. 4 provides a non-limiting example of a flowchart for a process 400 for predicting DNA sequences and / or cell type-specific functional genomics properties, in accordance with some embodiments of the present disclosure. Process 400 can be implemented using, for example, the prediction system 100 described in FIG. 1. For example, process 400 can be implemented using the sequence analyzer 108 and the machine-learning model(s) 132 shown in FIG. 1.
[0159] At step 402 in FIG. 4, an input data structure is formulated, where the input data structure comprises: a genomic sequence data component and / or at least one functional genomic property data component derived from a test sample, and where at least a portion ofMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-135 of 94 the genomic sequence data component and / or the at least one functional genomic property data component is unknown.
[0160] In some instances, the genomic sequence data can comprise, e.g., the DNA sequence for a specified gene and, optionally, a flanking sequence for at least one of a 3’ end or a 5’ end of the specified gene sequence (e.g., any human gene, any non-human mammalian gene, etc. .
[0161] In some instances, the genomic sequence data (e.g., DNA sequence data) may comprise sequences that range in length from about 100 base pairs (bp) to about IM bp. In some instances, the DNA sequences of the training data set may have a length of at least 100 bp, at least 500 bp, at least 1,000 bp, at least 5,000 bp, at least 10,000 bp, at least 50,000 bp, at least 100,000 bp, at least 200,000 bp, at least 300,000 bp, at least 400,000 bp, at least 500,000 bp, at least 600,000 bp, at least 700,000 bp, at least 800,000 bp, at least 900,000 bp, or at least IM bp. In some instances, the DNA sequences of the training data set may have a length of at most IM bp, at most 900,000 bp, at most 800,000 bp, at most 700,000 bp, at most 600,000 bp, at most 500,000 bp, at most 400,000 bp, at most 300,000 bp, at most 200,000 bp, at most 100,000 bp, at most 50,000 bp, at most 10,000 bp, at most 5,000 bp, at most 1,000 bp, at most 500 bp, or at most 100 bp. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances, the DNA sequences of the training data set may have a length of between about 5,000 bp and about 200,000 bp. Those of skill in the art will recognize that the DNA sequences of the training data set can have a length of any value within this range, e.g., about 10,750 bases.
[0162] In some instances, the at least one functional genomic property data component can comprise, e.g., cell type-specific RNA-seq data, ChlP-seq data, ATAC-seq data, CAGE-seq data, DNase-seq data, FAIRE-seq data, H3K27ac-Seq data, Methyl-seq data, etc., for at least one genomic locus. In some instances, the at least one functional genomic property data component can comprise functional genomic property profile data at single base pair resolution. In some instances, the at least one functional genomic property data component can comprise 1, 2, 3, 4, 5, 6, 7, 8,9, 10, or more than 10 functional genomic property datasets (or tracks). In some instances, the length of the functional genomics data tracks in the training data set may be equal to, longer, shorter, or similar in length to the length of the input genomic sequences.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-136 of 94
[0163] In some instances, the input data structure comprises a matrix. In some instances, at least one row of the matrix (e.g., at least one, two, three, or four rows) comprises genomic sequence data. In some instances, the genomic sequence data is one-hot encoded. In some instances, the matrix comprises a first row comprising a first functional genomic property dataset. In some instances, the matrix comprises at least a second row (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10throw) comprising at least a second functional genomic property dataset (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10thfunctional genomic property dataset).
[0164] In some instances, the test sample comprises a cell sample, a blood sample, a biopsy sample, or a tissue sample. In some instances, the genomic sequence data comprises data for genomic DNA extracted from the test sample. In some instances, the functional genomic property data comprises functional genomic property data derived by performing an assay (e.g., an RNA-seq, ChlP-seq, ATAC-seq, DNase-seq, FAIRE-seq, CAGE-seq, H3K27ac-seq, or methyl-seq assay) on the test sample.
[0165] At step 404 in FIG. 4, the input data structure is input into a trained machine learning model, the machine learning model trained by: (i) obtaining genomic sequences and / or at least one functional genomic property dataset derived from each of a plurality of training samples, (ii) formulating a plurality of training input data structures based on the genomic sequences and / or the at least one functional genomic property dataset derived from each of the plurality of training samples, (iii) applying a pre-defined masking strategy to mask a portion of the genomic sequence and / or at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures, and (iv) training the machine learning model to receive a masked training input data structure and predict the masked portion of the masked training input data structure.
[0166] In some instances, the plurality of training samples can comprise a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.
[0167] In some instances, the plurality of training samples can comprise samples from at least one cell type (e.g., erythrocytes (red blood cells), platelets, bone marrow cells, endothelial cells, lymphocytes, hepatocytes, neurons, glial cells, epidermal cells, respiratory interstitial cells, adipocytes (fat cells), fibroblasts, muscle cells, etc.) and / or at least one species (e.g., human, ape, monkey, dog, cat, or rat, etc.).MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-137 of 94
[0168] In some instances, the plurality of training samples can comprise from about 1 to about 10,000 samples. For example, in some instances the training data set can comprise data for at least 1, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 500, at least 1,000, at least 1,500, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000 samples. In some instances, the training data set can comprise data for at most 10,000, at most 9,000, at most 8,000, at most 7,000, at most 6,000, at most 5,000, at most 4,000, at most 3,000, at most 2,000, at most 1,000, at most 500, at most 100, at most 90, at most 80, at most 70, at most 60, at most 50, at most 40, at most 30, at most 20, at most 10, or at most 1 DNA sample. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the training data set can comprise data for between about 1,000 and about 10,000 samples. Those of skill in the art will recognize that the training data set can comprise data for a number of samples having any value within this range, e.g., about 4,750 samples.
[0169] In some instances, the training input data structures can each comprise a matrix. In some instances, at least one row of the matrix (e.g., at least one, two, three, or four rows) comprises genomic sequence data. In some instances, the genomic sequence data is one-hot encoded. In some instances, the matrix comprises a first row comprising a first functional genomic property dataset. In some instances, the matrix comprises at least a second row (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th row) comprising at least a second functional genomic property dataset (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th functional genomic property dataset).
[0170] In some instances, the trained machine learning model comprises a masked, multimodal machine learning model configured to process genomic sequence data and / or 2, 3, 4, 5, 6, 7, 8, 9, or 10 different functional genomic property datasets derived from the test sample.
[0171] In some instances, the trained machine learning model can comprise a deep neural network architecture. In some instances, the deep neural network architecture can comprise at least one convolution layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 convolutional layers).
[0172] In some instances, the deep neural network architecture can comprise at least one downsampling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 downsampling layers).MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-138 of 94
[0173] In some instances, the deep neural network architecture can comprise at least one self-attention layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 self-attention layers).
[0174] In some instances, the deep neural network architecture can comprise at least one deconvolution layer with matched U-net connections (at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 deconvolution layers with matched U-net connections).
[0175] In some instances, the deep neural network architecture may not include a global mean pooling layer. In some instances, the deep neural network architecture can comprise a single global mean pooling layer. In some instances, the deep neural network architecture can comprise at least one global mean pooling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 global mean pooling layers).
[0176] In some instances, the deep neural network architecture can comprise a linear output layer.
[0177] As noted above, a key feature of the disclosed deep learning model is its ability to be reconfigured for different functional genomics prediction applications by changing the masking strategy used during training and / or inference. The model can be reconfigured to predict DNA sequences (e.g., one or more bases at a specified genomic locus) and / or functional genomics property profiles by varying the masking of the input data structures used during training and / or inference. Non-limiting examples of the use of masking to reconfigure the model’s output predictions are described in more detail below.
[0178] At step 406 in FIG. 4, a prediction for the unknown portion of the genomic sequence data component and / or the at least one functional genomic property data component of the input data structure for the test sample is output.
[0179] In some instances, the prediction output by the model can comprise a prediction of a single nucleotide base at a specified genomic locus, or a prediction of a DNA sequence over a specified genomic region or interval.
[0180] In some instances, the prediction output by the model can comprise at least one functional genomic property prediction . In some instances, the at least one functional genomic property prediction can comprise an ATAC-Seq activity, a CAGE-Seq activity, a ChlP-Seq activity, a DNase-Seq activity, an H3K27ac-Seq activity, or any combination thereof for at least one genomic locus.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-139 of 94
[0181] The machine learning model can be configured for any of a variety of functional genomics applications known to those of skill in the art. For example, in some instances the prediction output by the trained machine learning model can be used to predict a functional effect of a genomic variant.
[0182] In some instances, the prediction output by the trained machine learning model can be used to interpret a functional effect of a genomic variant.
[0183] In some instances, the prediction output by the trained machine learning model can be used to denoise a functional genomic property dataset for the test sample.
[0184] In some instances, the prediction output by the trained machine learning model can be used to identify dependencies between functional genomic properties in the test sample.
[0185] In some instances, the prediction output by the trained machine learning model can be used to diagnose a disease.
[0186] In some instances, the prediction output by the trained machine learning model can be used to select a treatment for a disease.
[0187] In some instances, the prediction output by the trained machine learning model can be used to inform a disease prognosis.
[0188] In some instances, for example, the trained machine learning model can be used for predicting a functional genomic property according to a method comprising: obtaining an input genomic sequence derived from a test sample; obtaining at least one input functional genomic property dataset derived from the test sample; formulating an input data structure based on the input genomic sequence and the at least one input functional genomic property dataset; and predicting at least one output functional genomic property for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by: obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the plurality of masked training input data structures.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-140 of 94
[0189] In some instances, the trained machine learning model can be configured to receive a masked input data structure and predict the masked portion of the masked input data structure.
[0190] In some instances, the at least one predicted output functional genomic property for the test sample can be different from the at least one input functional genomic property dataset derived from the test sample.
[0191] In some instances, the trained machine learning model can comprise a multimodal machine learning model. In some instances, the multimodal machine learning model is configured to process genomic sequence data and at least 2, 3, 4, 5, 6, 7, 8 ,9, or 10 different functional genomic property datasets derived from the test sample as input.
[0192] In some instances, the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof. In some instances, the plurality of training samples comprise samples from at least one cell type (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, or 200 cell types) and / or at least one species (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 species), where the species may include, e.g., human, ape, monkey, dog, cat, rat, or mouse, etc.).
[0193] In some instances, the at least one input functional genomic property dataset comprises ATAC-Seq data, CAGE-Seq data, ChlP-Seq data, DNase-Seq data, H3K27ac-Seq data, or any combination thereof.
[0194] In some instances, the at least one output functional genomic property prediction comprises an ATAC-Seq activity, a CAGE-Seq activity, a ChlP-Seq activity, a DNase-Seq activity, an H3K27ac-Seq activity, or any combination thereof for at least one genomic locus.
[0195] In some instances, the prediction output by the trained machine learning model is used to: (i) predict a functional effect of a genomic variant, (ii) interpret a functional effect of a genomic variant, (iii) denoise a functional genomic property dataset for the test sample, (iv) identify dependencies between functional genomic properties in the test sample, (v) diagnose a disease, (vi) select a treatment for a disease, and / or (vii) inform a disease prognosis.VII. Prediction of Bases at Specific Genomic Loci
[0196] FIG. 5 provides a non-limiting example of a flowchart for a process 500 for predicting a base at a specified genomic locus, in accordance with some embodiments of the present disclosure. Process 500 can be implemented using, for example, the prediction systemMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-141 of 94100 described in FIG. 1. For example, process 500 can be implemented using the sequence analyzer 108 and the machine-learning model(s) 132 shown in FIG. 1.
[0197] At step 502 in FIG. 5, a genomic sequence derived from a test sample is obtained, where the genomic sequence comprises an unknown base at at least one specified genomic locus.
[0198] As noted above with respect to the process depicted in FIG. 4, the genomic sequence can comprise, e.g., the DNA sequence for a specified gene and, optionally, a flanking sequence for at least one of a 3’ end or a 5’ end of the specified gene sequence (e.g., any human gene, any non-human mammalian gene, etc.).
[0199] In some instances, the genomic sequence data (e.g., DNA sequence data) may comprise sequences that range in length from about 100 base pairs (bp) to about IM bp. In some instances, the DNA sequences of the training data set may have a length of at least 100 bp, at least 500 bp, at least 1,000 bp, at least 5,000 bp, at least 10,000 bp, at least 50,000 bp, at least 100,000 bp, at least 200,000 bp, at least 300,000 bp, at least 400,000 bp, at least 500,000 bp, at least 600,000 bp, at least 700,000 bp, at least 800,000 bp, at least 900,000 bp, or at least IM bp. In some instances, the DNA sequences of the training data set may have a length of at most IM bp, at most 900,000 bp, at most 800,000 bp, at most 700,000 bp, at most 600,000 bp, at most 500,000 bp, at most 400,000 bp, at most 300,000 bp, at most 200,000 bp, at most 100,000 bp, at most 50,000 bp, at most 10,000 bp, at most 5,000 bp, at most 1,000 bp, at most 500 bp, or at most 100 bp. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances, the DNA sequences of the training data set may have a length of between about 5,000 bp and about 200,000 bp. Those of skill in the art will recognize that the DNA sequences of the training data set can have a length of any value within this range, e.g., about 5,650 bases.
[0200] In some instances, the test sample comprises a cell sample, a blood sample, a biopsy sample, or a tissue sample. In some instances, the genomic sequence data comprises data for genomic DNA extracted from the test sample. In some instances, the functional genomic property data comprises functional genomic property data derived by performing an assay (e.g., an RNA-seq, ChlP-seq, ATAC-seq, DNase-seq, FAIRE-seq, CAGE-seq, H3K27ac-seq, or methyl-seq assay) on the test sample.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-142 of 94
[0201] At step 504 in FIG. 5, at least one functional genomic property dataset derived from the test sample is obtained.
[0202] As noted above with respect to the process depicted in FIG. 4, in some instances the at least one functional genomic property data component can comprise, e.g., cell type-specific RNA-seq data, ChlP-seq data, ATAC-seq data, CAGE-seq data, DNase-seq data, FAIRE-seq data, H3K27ac-Seq data, Methyl-seq data, etc., for at least one genomic locus. In some instances, the at least one functional genomic property data component can comprise functional genomic property profile data at single base pair resolution. In some instances, the at least one functional genomic property data component can comprise 1, 2, 3, 4, 5, 6, 7, 8,9, 10, or more than 10 functional genomic property datasets (or tracks). In some instances, the length of the functional genomics data tracks in the training data set may be equal to, longer, shorter, or similar in length to the length of the input genomic sequences.
[0203] At step 506 in FIG. 5, an input data structure is formulated based on the genomic sequence and the at least one functional genomic property dataset.
[0204] As noted above with respect to the process depicted in FIG. 4, in some instances the input data structure comprises a matrix. In some instances, at least one row of the matrix (e.g., at least one, two, three, or four rows) comprises genomic sequence data. In some instances, the genomic sequence data is one-hot encoded. In some instances, the matrix comprises a first row comprising a first functional genomic property dataset. In some instances, the matrix comprises at least a second row (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th row) comprising at least a second functional genomic property dataset (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th functional genomic property dataset).
[0205] At step 508 in FIG. 5, the base at the at least one specified genomic locus is predicted by inputting the input data structure into a trained machine learning model, where the machine learning model trained by: obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask at least one base in the genomic sequence of each of the training input data structures to generateMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-143 of 94 a plurality of masked training input data structures, and training the machine learning model based on the masked training input data structures to predict the at least one masked base.
[0206] In some instances, the plurality of training samples can comprise a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.
[0207] In some instances, the plurality of training samples can comprise samples from at least one cell type (e.g., erythrocytes (red blood cells), platelets, bone marrow cells, endothelial cells, lymphocytes, hepatocytes, neurons, glial cells, epidermal cells, respiratory interstitial cells, adipocytes (fat cells), fibroblasts, muscle cells, etc.} and / or at least one species (e.g., human, ape, monkey, dog, cat, or rat, etc. .
[0208] In some instances, the plurality of training sample can comprise from about 1 to about 10,000 samples. For example, in some instances the training data set can comprise data for at least 1, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 500, at least 1,000, at least 1,500, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000 samples. In some instances, the training data set can comprise data for at most 10,000, at most 9,000, at most 8,000, at most 7,000, at most 6,000, at most 5,000, at most 4,000, at most 3,000, at most 2,000, at most 1,000, at most 500, at most 100, at most 90, at most 80, at most 70, at most 60, at most 50, at most 40, at most 30, at most 20, at most 10, or at most 1 DNA samples. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the training data set can comprise data for between about 500 and about 8,000 samples. Those of skill in the art will recognize that the training data set can comprise data for a number of samples having any value within this range, e.g., about 6,300 samples.
[0209] In some instances, the training input data structures can each comprise a matrix. In some instances, at least one row of the matrix (e.g., at least one, two, three, or four rows) comprises genomic sequence data. In some instances, the genomic sequence data is one-hot encoded. In some instances, the matrix comprises a first row comprising a first functional genomic property dataset. In some instances, the matrix comprises at least a second row (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th row) comprising at least a second functional genomic property dataset (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th functional genomic property dataset).MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-144 of 94
[0210] In some instances, the trained machine learning model comprises a masked, multimodal machine learning model configured to process genomic sequence data and / or 2, 3, 4, 5, 6, 7, 8, 9, or 10 different functional genomic property datasets derived from the test sample.
[0211] In some instances, the trained machine learning model can comprise a deep neural network architecture. In some instances, the deep neural network architecture can comprise at least one convolution layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 convolutional layers).
[0212] In some instances, the deep neural network architecture can comprise at least one downsampling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 downsampling layers).
[0213] In some instances, the deep neural network architecture can comprise at least one self-attention layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 self-attention layers).
[0214] In some instances, the deep neural network architecture can comprise at least one deconvolution layer with matched U-net connections (at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 deconvolution layers with matched U-net connections).
[0215] In some instances, the deep neural network architecture may not include a global mean pooling layer. In some instances, the deep neural network architecture can comprise a single global mean pooling layer. In some instances, the deep neural network architecture can comprise at least one global mean pooling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 global mean pooling layers).
[0216] In some instances, the deep neural network architecture can comprise a linear output layer.
[0217] As noted above, a key feature of the disclosed deep learning model is its ability to be reconfigured for different functional genomics prediction applications by changing the masking strategy used during training and / or inference. The model can be reconfigured to predict DNA sequences (e.g., one or more bases at a specified genomic locus) and / or functional genomics property profiles by varying the masking of the input data structures used during training and / or inference. For example, in some instances the pre-defined masking strategy comprises a variable masking strategy in which a different number of genomic loci and / or different genomic loci are masked in different training input data structures of the plurality of training input data structures. Any part of any functional genomic property track and / or DNA sequence can be masked.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-145 of 94
[0218] In some instances, the method further comprises identifying an individual associated with the test sample based on the predicted base at the at least one specified genomic locus for the test sample. In some instances, the identifying the individual associated with the test sample can comprise: obtaining genomic sequences derived from test samples from each of a plurality of individuals; obtaining at least one functional genomic property dataset derived from each of the plurality of test samples; formulating a plurality of input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of test samples; inputting the plurality of input data structures into the trained machine learning model; outputting a predicted base at the at least one specified genomic locus in the genomic sequence for each of the plurality of test samples; and determining an identity of at least one individual based on the predicted base at the at least one specified genomic locus in the genomic sequence for each of the plurality of test samples.VIII. Design of Genomic Regulatory Elements
[0219] FIG. 6 provides a non-limiting example of a flowchart for a process 600 for designing a genomic regulatory element, in accordance with some embodiments of the present disclosure. Process 600 can be implemented using, for example, the prediction system 100 described in FIG. 1. For example, process 600 can be implemented using the sequence analyzer 108 and the machine-learning model(s) 132 shown in FIG. 1.
[0220] At step 602 in FIG. 6, at least one functional genomic property dataset generated based on a desired regulatory effect is obtained for a test sample.
[0221] In some instances, the at least one functional genomic property dataset may comprise, for example, an ATAC-Seq dataset, CAGE-Seq dataset, ChlP-Seq dataset, DNase- Seq dataset, H3K27ac-Seq dataset, or any combination thereof.
[0222] In some instances, the at least one functional genomic property dataset may be generated manually, e.g., by drawing an activity curve as a function of genomic position that indicates the desired regulatory effect in the test sample (z.e., the target tissue type or cell type)
[0223] In some instances, the at least one functional genomic property dataset may be generated experimentally, e.g., by performing a sequencing-based assay on nucleic acid molecules extracted from a sample that is different from the test sample (z.e., the target tissue type or cell type), but that exhibits the desired regulatory effect.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-146 of 94
[0224] In some instances, the desired regulatory can be, e.g., increased expression for one or more genes (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 genes) in the test sample (e.g., a target tissue type or cell type), suppression of expression for one or more genes (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 genes) in the test sample (e.g., a target tissue type or cell type), or any combination thereof.
[0225] In some instances, the test sample may be a tissue sample or a cell sample (e.g., a target tissue type or cell type in which the designed regulatory element is intended to induce the desired regulatory effect). In some instances, the test sample may comprise a human tissue or cell sample. In some instances, the test sample may comprise a non-human tissue sample or cell sample (e.g., an ape, monkey, dog, cat, or rat sample, etc.).
[0226] At step 604 in FIG. 6, an input data structure is formulated based on a candidate genomic sequence of specified length and the at least one functional genomic property dataset.
[0227] The candidate genomic sequence can be, e.g., an initially random sequence or other candidate sequence (e.g., a partially known candidate sequence, or a candidate sequence based on a genomic sequence from a same or different tissue type or cell type). In some instances, the initial candidate sequence can be a fully empty sequence (e.g., a fully masked sequence) and the full sequence is generated iteratively, as illustrated in FIG. 29B.
[0228] In some instances, the candidate genomic sequence (e.g., candidate genomic regulatory element) can comprise a candidate promoter sequence, a candidate enhancer sequence, a candidate silencer sequence, a candidate cis-regulatory sequence, a candidate trans- regulatory sequence, or any combination thereof.
[0229] In some instances, the candidate genomic sequence can comprise a candidate DNA sequence ranging in length from about 100 nucleotides (or base pairs (bp)) to about 100,000 nucleotides (or base pairs (bp)), or longer. In some instances, the length of the candidate DNA sequence can be at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 3,500, at least 4,000, at least 4,500, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, at least 80,000, at least 90,000, or at least 100,000 nucleotides (or bp) in length. In some instances, the length of the candidate DNA sequence can be at most 100,000, at most 90,000, at most 80,000, at most 70,000, at mostMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-147 of 9460,000, at most 50,000, at most 40,000, at most 30,000, at most 20,000, at most 10,000, at most 9,000, at most 8,000, at most 7,000, at most 6,000, at most 5,000, at most 4,500, at most 4,000, at most 3,500, at most, 3,000, at most 2,500, at most 2,000, at most 1,500, at most 1,000, at most 900, at most 800, at most 700, at most 600, at most 500, at most 400, at most 300, at most 200, or at most 100 nucleotides (or bp) in length. Any of the lower and upper values described in this paragraph may be combined to form a range included within the present disclosure. For example, in some instances the length of the candidate DNA sequence can have a length ranging from about 200 to about 5,000 nucleotides. Those of skill in the art will recognize that the candidate DNA sequence can have a length of any value within this range, e.g., about 232 nucleotides.
[0230] In some instances the input data structure can comprise a matrix. In some instances, one row of the matrix comprises genomic sequence data. In some instances, the genomic sequence data is one-hot encoded. In some instances, the matrix comprises a first row comprising a first functional genomic property dataset. In some instances, the matrix comprises at least a second row (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th row) comprising at least a second functional genomic property dataset (e.g., a 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, or 10th functional genomic property dataset).
[0231] At step 606 in FIG. 6, the a base at one or more positions of the candidate genomic sequence is predicted by: (i) inputting the input data structure into a trained machine learning model, the trained machine learning model configured to output a prediction of a base at one or more positions of a genomic sequence based on at least one input functional genomic property dataset; and (ii) substituting the predicted base at the one or more positions of the candidate genomic sequence for the previous base at the one or more positions of the candidate genomic sequence in the input data structure.
[0232] In some instances, the trained machine learning model may comprise a deep neural network architecture. For example, in some instances the trained machine learning model may comprise a deep neural network architecture comprising a stack of convolution and downsampling layers. In some instances, the deep neural network architecture can comprise at least one convolution layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 convolutional layers), at least one downsampling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 downsampling layers), at least one self-attention layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 self-attention layers),MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-148 of 94 at least one deconvolution layer with matched U-net connections (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 deconvolution layer with matched U-net connections), at least one global mean pooling layer (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 global mean pooling layers), or any combination thereof. In some instances, the deep neural network architecture comprises a linear output layer.
[0233] In some instances, as described above, the trained machine learning model may comprise a neural network similar to that of, e.g., the Borzoi model, but where the model is configured to accept an input data structure comprising an input DNA sequence and / or one or more functional genomics property data tracks. In some instances, the output layer can be configured to output a predicted DNA base at one or more genomic loci.
[0234] In some instances, the trained machine learning model can be trained on training data comprising genomic sequences and at least one functional genomic property dataset derived from each of a plurality of training samples. In some instances, the plurality of training samples can comprise a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof. In some instances, the plurality of training samples can comprise samples from at least one cell type, cell state (e.g., disease state), and / or species (e.g., human, ape, monkey, dog, cat, or rat, etc.).
[0235] In some instances, the at least one cell type can comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, or at least 200 different types of cells.
[0236] In some instances, the at least one cell state (e.g., disease state) can comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, or at least 150 different cell states. In some instances, the different cell states can comprise different cellular development states, different cellular metabolic states, different disease states, or any combination thereof.
[0237] In some instances, the training data may be formulated as a plurality of masked training input data structures generated from a plurality of training input data structures using a variable masking strategy in which a different number of genomic loci and / or different genomic loci are masked in different training input data structures of the plurality of training input data structures. In some instances the training input data structure can comprise a matrix,MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-149 of 94 e.g., a matrix comprising one-hot encoded genomic sequence data and one or more functional genomics property data tracks, as described above for the input data structure.
[0238] In some instances, the trained machine learning model can be configured to output a prediction of a base at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 positions of a genomic sequence based on at least one input functional genomic property dataset.
[0239] At step 608 in FIG. 6, the candidate genomic sequence is output as a designed genomic regulatory element for the test sample.
[0240] In some instances, the method can further comprise repeating the steps of: (i) inputting the data structure into the trained machine learning model and (ii) substituting the predicted base at the one or more positions of the candidate genomic sequence for the previous base at the one or more positions of the candidate genomic sequence, for two or more iterations (e.g., for 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 230, 40, 50, 60, 70, 80, 90, 100, 200, 400, 600, 800, 1,000, or more than 1,000 iterations).
[0241] In some instances, the designed genomic regulatory sequence can comprise a promoter sequence, a enhancer sequence, a silencer sequence, a cis-regulatory sequence, a trans-regulatory sequence, or any combination thereof.
[0242] In some instances, the designed genomic regulatory element can comprise an adeno- associated virus (AAV) vector.IX. ExamplesExample 1 - A Multimodal, Multi-Masking Model for Prediction ofDNA Sequence and / or Functional Genomic Properties
[0243] This non-limiting example further describes the disclosed multimodal, multimasking-based deep learning model configured to predict cell type-specific functional genomic traits.
[0244] FIG. 7 provides a non-limiting schematic illustration of a multi-modal, multimasking-based deep learning model (e.g., a functional genomics prediction model, or “Nona”) configured to output sequence and / or functional genomics property predictions, in accordance with some embodiments of the present disclosure. The model can be configured to accept input data comprising a DNA sequence and one or more functional genomics property data sets (e.g., “tracks”), and output a prediction of masked bases in the input DNA sequence and / or maskedMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-150 of 94 portions of one or more functional genomics property tracks. Non-limiting examples of regulatory genomics tracks include, but are not limited to, DNase-seq data, ChlP-seq data, and CAGE-seq data over a specified genomic regions (e.g., a 131 kb region in this example).
[0245] In some instances, as illustrated in FIG. 7, the input data may be configured as a matrix comprising the DNA sequence data and the one or more functional genomics property data tracks. In some instances, the DNA sequence data may be one-hot encoded. In the schematic illustration of FIG. 7, the “tasks” correspond to single base-resolution predictions of sequence bases and / or sequence read counts in functional genomics experiments as a function of genomic position. The red blocks in the input matrix denote values that are masked.
[0246] FIG. 8 provides a non-limiting schematic illustration of how incorporating different masks during inference (selected from a same distribution of masks used during model training) enables a variety of different use cases for the model that were hitherto only addressable by training separate models. Non-limiting examples of masking strategies (and corresponding use cases) include the sequence-to-track masking mode (for variant effect prediction, sequence interpretation, sequence design, etc.), sequence-plus-context masking mode (for context-aware sequence design, etc.), positional masking mode (fortrack super-resolution (e.g., taking a track with low sequencing read-depth and enhancing it to generate a track corresponding to a higher sequencing read-depth), etc.), language modeling masking mode (for variant effect prediction, etc.), sub sequence-to-track masking mode (for design of short regulatory or gene insertion elements, e.g., adreno-associated virus (AAV) vectors, etc.), and track-to-track with sequence masking mode (for learning track-to-track dependencies, etc.).
[0247] Some entries of the input matrix are masked using a binary mask. The mask is concatenated with the inputs in order to distinguish real Os from masked Os. In one example, a U-Net architecture similar to that of the Borzoi model (Linder et al. (2023) was utilized. The architecture and training of the model using a multi-resolution loss function are described elsewhere herein (see FIG. 2 and the description thereof). The model outputs a matrix of shape L x (4+T) at single base resolution, and matrices of shape L / 21x T at coarser resolutions. During training, masks were sampled from a pre-specified distribution. During inference, incorporating different masks (from the same distribution) enabled versatile use cases that were hitherto only addressable by using separate models, as noted above.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-151 of 94
[0248] In this example, 2 kb sequence-to-task models (z.e., models accepting 2 kb long sequence and / or functional genomics property tracks as input) were trained for 12 prediction tasks (6 DNase-seq predictions, 3 ATAC-seq predictions, and 3 ChlP-seq prediction) in various cell types. Five models were trained using single base-resolution Poisson loss at the final output layer only, and five models were trained using the multi-scale loss function. Multi-scale training significantly improved performance on one task, z.e., NFE2 ChlP-seq predictions in K562 cells (FIG. 9, showing a non-limiting example of model prediction performance on 12 tasks with and without training using the multiscale loss function; the orange (larger) dot corresponds to predicted NFE2 ChlP-seq data in K562 cells). The multiscale loss function was subsequently used in further studies.
[0249] Next, a 131 kbp single base-resolution model was trained on the same 12 prediction tasks utilizing all 6 masks illustrated in FIG. 8. In masked language modeling mode, the disclosed functional genomics prediction model exhibited an accuracy of 49.6% (FIG. 10). Task-specific analyses were focused on the fibroblast ATAC-seq prediction task from Nair et al (2023), which is one of the 12 tasks. First, single base-resolution predictions in the sequence- to-task and task denoising masking modes were compared (FIG. 11, showing a non-limiting example of the single base-resolution prediction Poisson NLL loss (z.e., the negative loglikelihood loss with a Poisson distribution of the target) for sequence-to-task and task denoising modes). As expected, providing the model with the experimentally-determined bases at some positions improved the accuracy of base prediction at the masked positions, thus improving denoising ability.
[0250] Next, the sub sequence-to-task masking mode was employed to study how prediction in the central 2kb regions improves as the input sequence length is increased. The performance of the model in full 13 Ikb sequence-to-task masking mode was benchmarked against a 2kb ChromBPNet model (Nair et al (2023)), and achieved comparable performance (FIG. 12, showing a plot of Spearman correlation coefficient for total counts prediction in the central 2 kb region using ChromBPNet and the disclosed model on fibroblast peaks and a GC -matched set of non-peak regions in the test gene set). The performance improved when a genomic sequence + 129 kb of context (z.e., observations of 12 tasks for 64.5kb on either side of the central 2kb sequence region) was provided as input. The sequence-to-task masking mode overestimated accessibility primarily in heterochromatin in fibroblasts (FIG. 13A, showingMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-152 of 94 model predictions in sequence-to-task masking mode at a heterochromatin locus; the observed profile (black) is flipped along the x-axis), which was corrected in the sequence + context masking mode (FIG. 13B, showing model predictions in sequence + context masking mode for the heterochromatin locus). The sequence-to-task masking mode overestimated chromatin accessibility over the entire region. Use of the sequence + context masking mode with a context of 64.5kb on either side of the central 2kb region (highlighted in yellow) corrected the predicted counts in the central 2kb region.
[0251] New masking modes, larger scale models, and new applications for the study and design of regulatory sequences thereby enabled are currently under study. For example, novel applications enabled by the disclosed model include:
[0252] Context-aware predictions: In this mode, the model predicts epigenetic tracks in a local genomic window based on the DNA sequence and the observed epigenetic tracks in adjacent windows. This mode allows the model to integrate broader chromatin context information to improve predictions over sequence-only models. As discussed elsewhere herein, a 196kb functional genomics prediction model with a 4kb central mask trained on 7 DNase and 2 CAGE tracks from K562 cells showed improved performance on a test set of sequences and on prediction of expression of randomly integrated reporters across diverse genomic backgrounds measured by TRIP-seq experiments.
[0253] Sequence generation: In this mode, the model is trained as a masked language model conditioned on epigenetic profiles with variable masking rate. As discussed elsewhere herein, the model can then be used to generate regulatory elements that have desired epigenetic profiles by iteratively sampling bases. For a Ikb model trained with 3 DNase-seq profiles, the generated regulatory elements are predicted by Borzoi (Linder et al. (2023)) to have the desired regulatory profiles and contain relevant cell-type specific motifs.
[0254] Genotyping from ATAC-seq: Commonly used ATAC-seq fragment files, though devoid of variant level information, can inadvertently result in leakage of private patient information due to subtle changes to Tn5 transposon cut sites around the variant. A 128 bp conditional masked language model was trained using base-resolution ATAC-seq profiles from GM12878 cells. As described elsewhere herein, the model was then used on ATAC-seq from different individuals to predict their alleles at common variant loci. Given a set of -3200 individual genotypes from the AFGR project, a linking attack was performed that identified 83MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-153 of 94 sample donors from their ATAC-seq profiles with near perfect accuracy. The linking is robust to read subsampling, and is highly accurate at read depths as low as 5M.
[0255] Altogether, the disclosed model comprises a versatile training and inference paradigm that extends sequence-to-function and masked language modeling to novel applications in regulatory genomics. The disclosed functional genomics prediction model can also be used for studying epigenetic perturbations and extensions of language modeling, and that model predictions can shed light on context-specific dependencies between epigenetic tracks and regulatory elements.Example 2 - Sequence + Context Prediction Models for K562 Cells
[0256] This non-limiting example describes the training and use of a new machine learning model (“Nona”) to predict functional genomic properties for K562 human chronic myelogenous leukemia (CML) cells. The model was trained on data for human K562 and mouse myeloid cell lines, including data from 7 human DNase-seq assays, 2 human CAGE- seq assays, 3 mouse DNase-seq assays, and 3 mouse CAGE-seq assays. Previous studies have shown that training on data from human plus mouse leads to better predictive models that those trained on human data alone (see, e.g., Kelly (2020), “Cross-species regulatory sequence activity prediction”, PLOS Computational Biology, https: / / doi.org / 10.1371 / journal.pcbi.1008050; Avsec et al. (2021), “Effective gene expression prediction from sequence by integrating long-range interactions”, Nature Methods 18: 1196- 1203; Linder et al. (2025), “Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation”, Nature Genetics 57:949-961).
[0257] FIG. 14 provides a non-limiting example of predicted fibroblast ATAC-seq data versus measured (labeled) ATAC-seq data profiles using 131 kb input sequences. The labeled profile has been inverted for ease of comparison. The inset provides a zoomed in view of the predicted and measured profile data in the indicated genomic region.
[0258] FIG. 15 provides a non-limiting example of a comparison of model prediction performance data (Spearman R correlation coefficients) for a prior art model (ChromBPNet) and different versions of the model disclosed herein. Data for the model used in sequence-only prediction mode is shown in the left panel. Data for the model used in the sequence + context prediction mode (i.e., where the input comprises 129 kb of context data) is shown in the rightMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-154 of 94 panel. Operation in the sequence + context prediction mode resulted in significant improvements in model prediction accuracy.
[0259] FIG. 16 provides a non-limiting example of prediction accuracy data (Pearson correlation coefficients for comparison of predictions and experimentally determined values) for a test set comprising seven DNase-seq data tracks from K562 cells, and two CAGE-seq data tracks also from K562 cells, under different treatment conditions (the IDs are ENCODE accession) when the model was operated in sequence only (4 kb) versus sequence + context (196 kb) input modes at 128 bp resolution. Predictions were made for a test set of genomic regions used to evaluate model performance. As can be seen, providing context improved prediction performance. DNase-seq and CAGE-seq were for the K562 cell line. DNase-seq is an assay that measures regions of open chromatin (similar to ATAC-seq). CAGE-seq is an assay that measures gene expression by annotating transcription start sites.
[0260] FIG. 17 provides a non-limiting example of data for the correlation between model predictions when operated in sequence + context mode versus sequence-only mode for DNase- seq data for K562 cells treated with 1 pM vorinostat for 72 hours (ENCFF899YDP). The data indicate that the model predictions were altered for a subset of genomic loci when operated in sequence + context prediction mode (the sequence-only prediction mode tended to overpredict at these sites).
[0261] FIG. 18A provides a non-limiting example of data that illustrates the distribution of ChromHMM states in differential genomic regions. Genomic regions for which predictions were corrected by including context tended to be quiescent regions or heterochromatin regions.
[0262] FIG. 18B provides a non-limiting example of data for comparison of sequence-only prediction of DNase-seq data in K562 cells versus experimentally-determined DNase-seq data for a genomic locus where including context in the model input improved the model’s prediction performance. Operating the model in sequence + context prediction mode improved prediction performance (z.e., the predictions were closer to the actual (observed) values). In context aware mode, the model takes the low signal in the flanking regions into account and adjusts its prediction to be lower than that obtained using the sequence-only prediction mode. The data is for a test set genomic region corresponding roughly to chrl9:52436285-52553935 (hg38). The data tracks are DNase-seq data for K562 cells. For the ChromHMM states:MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-155 of 94 lilac indicates heterochromatin, light green indicates ZNF genes + repeats, white indicates quiescent states, and red, yellow, and dark green indicate transcription states.Example 3 -Evaluation of Model Performance on TRIP-seq Data
[0263] An experimental approach known as TRIP-seq (Akhtar et al. (2013)) was used as another means of evaluating the sequence-plus-context model in a more realistic setting. The key idea is that the same reporter gene (approximately 2KB in length) is randomly inserted in different parts of the genome in a population of cells, and the expression of the gene is then measured (e.g., using CAGE-seq data for K562 cells). FIG. 19 illustrates the masking strategy used to predict gene expression (CAGE-seq ) profiles in K562 human chronic myelogenous leukemia (CML) cells when a reporter sequence of about 2 kb was inserted at multiple genomic insertion sites. This provided a use case for the sequence-plus-context prediction model because information about the region of the genome where the reporter gene was integrated can be presented to the model along with the sequence data. The model was provided with the genomic sequence at the integration site and the other functional genomics property tracks around the region where the reporter gene was integrated. Benchmarking studies were performed with the sequence-only model and the sequence-plus-context model for a set of test genes, and the prediction results (as quantified by calculation of the Spearman correlation coefficient between the predicted gene expression and the measured gene expression) were also compared with those generated using another sequence-based model known as Enformer (FIG. 20, showing Spearman R correlation coefficient data for predicted versus measured Steensel TRIP-seq data in K562 cells). The performance of the sequence-plus-context model was about the same or better than that of the sequence-only model when making predictions for this specific functional genomic property track (CAGE-seq data for K562 cells), which is relevant for applications that involve integrating constructs into the genome. In general, the prediction performance of the sequence-plus-context model was better than that for the sequence-only model, and was as good or better than that for the prior art Enformer model. Similar improvements in prediction performance were observed in another TRIP-seq study (FIG. 21, showing Spearman R correlation coefficient data for predicted versus measured Cohen TRIP- seq data in K562 cells).MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-156 of 94Example 4 - dsQTL Classification in Lymphoblastoid Cell Lines (LCLs)
[0264] The mapping of expression quantitative trait loci (eQTLs) is an important tool for linking genetic variation to changes in gene regulation. DNase I sensitivity quantitative trait loci (dsQTLs) are variant single nucleotide polymorphism or insertion / deletion loci for which the DNase-seq read depth correlates significantly with genotype (Degner et al. (2012), “DNase I sensitivity QTLs are a major determinant of human expression variation”, Nature 482(7385): 390-394). It has been shown that genetic variants that modify chromatin accessibility and transcription factor binding are a major mechanism through which genetic variation leads to gene expression differences among humans.
[0265] This example illustrates application of the disclosed model to the task of classifying dsQTLs in lymphoblastoid cell lines (LCLs). A 662-track GM12878 and K562 model (196kb) was trained on all 6 masking modes described above. On sequence-to-track prediction performance (z.e., prediction of functional genomics properties based on an input genomic sequence), it approached the performance of the Enformer model on a test set of genes (FIG. 22A).
[0266] FIG. 22B provides a non-limiting example of precision-recall data for the disclosed model (z.e., the functional genomics prediction model) when operated in sequence-only prediction mode, and for several prior art models (e.g., ChromBPNet, Enformer, and Borzoi), when deployed for dsQTL classification in LCLs (AUPRC = area under the precision recall curve). As can be seen, the model’s performance was competitive with that of prior art models on dsQTL classification in LCLs. This example demonstrates that the disclosed model can be used to classify variants as being dsQTLs or not.
[0267] Adding individual tracks with sequence data enables new applications such as studying the effect of perturbing one epigenetic data track on the prediction for another (FIG. 22C, showing the effect of perturbing the epigenetic modification of histone H3 on predicted Cap Analysis Gene Expression sequencing (CAGE-seq) data in GM12878 cells), and for studying the dependence of functional genomic property tracks on each other (FIG. 22D, showing the increase in prediction performance over sequence-only prediction mode when additional functional genomic property data tracks were include with genomic sequence data as part of the model input for GM12878 and K562 cells). In this case, the track-track dependency primarily splits by cell line. This example demonstrates that the disclosed modelMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-157 of 94 can be used to predict the effect of perturbation of one functional genomic property on another, and also to evaluate the interdependencies of such functional genomic properties.Example 5 - Predicting Individual Genotypes
[0268] Another non-limiting use case example for the disclosed model is prediction of individual genotypes. A different masking strategy can be used to train the disclosed model to perform de-de-identification of ATAC-seq fragment files, as inspired by a recent paper (Walker et al. (2024) that showed that the genotypes of individual people (z.e., the actual DNA sequences of different people) can be derived from gene expression matrices. For example, measurement of gene expression for a set of different genes in cells derived from a blood sample can enable one to determine what DNA mutations are present in an individual even without detailed knowledge of the individual’s genomic DNA sequence, which in turn can allow one to link a given set of DNA mutations to the individual that it came from.
[0269] The disclosed model can perform something similar using ATAC-seq data (a sequencing-based assay technique used to assess genome-wide chromatin accessibility and determine which genomic DNA regions are most active), which can indicate whether or not an individual has a given mutation (e.g., aA mutation) at an indicated genomic locus. ATAC sequence read data is typically converted into a fragment file and de-identified by removing all sensitive information about mutations, etc., from the data set. Examples of sequence read data presented in the ATAC-seq fragment file format for various individuals can be found online at, for example, the 10X Genomics, Inc. website (www.1 Oxgenomi cs. com) . Although the identity of the specific individual and the mutations present in their genome are not included in this data, ATAC-seq data has the property that it’s extremely spiky. For example, if one looks at ATAC sequence read lengths and where they are positioned, the reads tend to accumulate at specific genomic loci with single base resolution due to the very specific sequence preference of the enzyme used to cut sequences when performing ATAC-seq studies. One tends to see a larger number of reads for sequences that include higher GC content compared to that for sequences that have lower GC.
[0270] To illustrate this, FIG. 23A provides a plot of normalized ATAC signal as a function of position relative to the distance from a variant position. The data plotted in this example represents ATAC seq signal summed up over 32 and 15 individuals respectively, where the 32 individuals have a T / T allele at the variant position (z.e., the genotypes at this position are bothMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-158 of 94T and T), and the 15 individuals have an A / A allele (the pattern for the 15 individuals has been reflected along the Y axis). As can be seen, while there’s an overall correspondence between genomic positions exhibiting high ATAC signal, there is a subtle difference in the shape of the profile around the variant position. The data indicates that for individuals who have a T / T allele in their genome (z.e., both parents’ genotypes had a T in this position) will have a profile shape as indicated, whereas individuals who have an A / A allele have a slightly different profile shape (z.e., different levels of ATAC signal at corresponding genomic loci). This further indicates that the presence of genetic variants in some individuals leads to differences in ATAC-seq data profiles and implies that one might be able to predict the genotype of an individual (z.e., the presence of a mutation at a specific genomic locus) from their ATAC-seq profile.
[0271] FIG. 23B illustrates a masking strategy that can be used to train a functional genomics prediction model that predicts genotype based on a human reference genome, z.e., a representative genome obtained by averaging across a pool of individuals rather than being derived from a specific individual’s genome. The training of the functional genomics prediction model may be similar to that used in masked language modeling, except that one track of functional genomic property data (e.g., ATAC-seq data) was included in the training data along with genomic sequence data. A variable masking rate may be used during training, z.e., a different set of base positions were masked in each training data element. For example, for a sequence that is seven bases in length, one might mask two bases in one training data element, and one base in a different training data element. A uniform masking probability may be used to ensure an equal chance of masking 1, 2, 3, 4, 5, 6, or all 7 base positions in this case, and then the functional genomics prediction model is trained to predict the missing base(s) from the bases that are presented in the training data. During inference, the functional genomics prediction model may be used to predict the genotype of an individual from whom adjacent sequence data and functional genomic property data were derived. The central base may be masked, z.e., the functional genomics prediction model is directed to predict the base at the central position of the 7 base sequence from the adjacent sequence data and functional genomic property data presented as input.
[0272] This represents a new use case for multimodal, masked machine learning models. The functional genomics prediction model has been shown to accurately predict the genotype of the person from whom the data set was derived. The mask used is illustrated in FIG. 24A,MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-159 of 94 and included 128 bp of ATAC-seq and genomic sequence data, with a central single nucleotide polymorphism (SNP) masked. A non-limiting example of data for base prediction accuracy is shown in FIG. 24B. The functional genomics prediction model had about a 90% prediction accuracy for being able to correctly identify an individual’s genotype based on their ATAC- seq data, even without training on the individual’s genome.
[0273] This implies that the functional genomics prediction model can be used to identify which individual a given sample came from (where the sample is from a collection of samples from different individuals for which ATAC-seq data is available) based on the ATAC-seq data profile. Assume that one has the ATAC-seq data track for an individual (e.g., Bob) that indicates the level of ATAC-seq signal at a number of known common variant positions. Based on these known variant positions (distributed throughout the genome), one can make predictions using the functional genomics prediction model described above. For example, at one variant position the functional genomics prediction model may predict that the individual likely has a cytosine (C) instead of adenine (A), guanine (G), or thymine (T) (FIG. 25, left panel). At another position, the functional genomics prediction model may predict that the individual likely has a different DNA base, e.g., an A instead of a C, G, or T (FIG. 25, left panel). Assume that one also has the genotypes of three individuals, e.g., Alice, Bob, and Charlie (FIG. 25, right panel). Given this information (and not knowing in advance that the ATAC-seq data set came from Bob), can one link this data set back to Bob? In this case, the functional genomics prediction model predicts (based on the ATAC-seq data) that there should be a C at the specified variant position. In accordance with the functional genomics prediction model’s prediction, Bob’s genotype has a C at the specified position in each gene allele. The functional genomics prediction model’s predictions may thus be used to identify an individual, or link a specific sample data set to a specific individual, based on genotype predictions. In another example, the functional genomics prediction model might predict that the individual has an A at the specified variant position. In this case, Bob has two A’s (one at the specified position in each of the two gene alleles), whereas both of the other individuals have an A / T pair, so again one can likely link the ATAC-seq data set to Bob.
[0274] FIG. 26 provides an example of a plot of the number of individuals observed versus the log likelihood that a given ATAC-seq derived genotype (for individual NA12878) was linked to that individual for a pool of about 3,200 individuals (data taken from the 1000MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-160 of 94Genomes Project). A genotype prediction was made for each individual using the functional genomics prediction model described above, and the predictions were then compared to the genotype predictions for the other individuals in the pool of about 3,200 individuals in total. As can be seen, the score for the individual (NA12878) from whom the ATAC-seq data actually came is much higher than that for any other individual. The model was able to very accurately predict which individual the ATAC-seq sample came from without having access to their sequence-level genotyping data.
[0275] As another example of the near-perfect model prediction performance, ATAC-seq data for 83 people (z.e., 83 AFGR ATAC-seq profiles) was compiled and the system was used to link the 83 ATAC-seq samples to the 3,200 individuals in the pool. In this case, the functional genomics prediction model was trained once on one individual, and then applied to every other individual in the pool. Based on the data for the original samples, the functional genomics prediction model was able to link all 83 samples with the correct individuals with 100% accuracy. Sub-sampling these experiments to make the experimental data increasingly noisy indicated that there was a read level (about 2 million reads) where the model’s performance to correctly link samples to individuals decreased (FIG. 27).Example 6 - Design of Regulatory Elements
[0276] As another non-limiting example of a use case, the disclosed functional genomics prediction model can be used to design regulatory elements. Recently, several groups have used DNA sequence-based deep learning models to design regulatory elements, z.e., elements (short nucleic acid sequences, typically 5 to 15 nucleotides long) such as promoters, that control the activity of different genes in a cell type-specific manner. For example, a given regulatory element might turn a gene on in retinal cells, but not in liver cells. The model disclosed herein can potentially be used to design longer cell type-specific regulatory sequences, e.g., enhancer sequences ranging in length from about 100 bp to about 2,000 bp, or even longer sequences of up to about 100 kb or longer that include multiple enhancer, promoter, and / or gene sequences.
[0277] FIG. 28 provides a non-limiting schematic illustration of deployment of the disclosed functional genomics prediction model for design or improvement of tissue typespecific DNA regulatory sequences based on high-level specifications, for example, the design of an enhancer sequence and / or promoter sequence that yields high expression of a given gene in one tissue (e.g., lung) which minimizing the expression of the gene in another tissue (e.g.,MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-161 of 94 brain). The design or improvement of a regulatory element can be based on an initial candidate sequence that is completely unknown, completely random, or based on a known regulatory element sequence.
[0278] The functional genomics prediction model may be trained in a manner similar to that used in the genotype prediction examples discussed above, but can then be used to generate novel DNA sequences. The masking strategy used for training the model is illustrated in FIG. 29A, while that used for inference is illustrated in FIG. 29B. In the conceptual masking strategy illustrated in FIG. 28A, there was one track of functional genomic property data provided as model input in addition to the genomic sequence data. In the experimental data examples that follow, up to three functional genomic property tracks have been used, but this can be extended to multiple tracks. During training, the functional genomics prediction model was presented with sequence and functional genomic property data where some bases have been masked (similar to a conditional mask model) but the functional genomic property data was left completely unmasked. A variable masking rate was used during training, z.e., three base positions might be masked for some training data elements, while only one base position might be masked in others, and so on. The trained model may then be used to generate a sequence that should induce the desired functional genomic property track.
[0279] During inference, the functional genomics prediction model may be presented with the desired functional genomic property track and all base positions in the input sequence (initially unknown) were masked. Model predictions were then used to fill in the base positions one or more bases at a time (e.g., 1, 2, 3, 4, or more than 4 bases at a time) through an iterative process. For example, in the first step of inference illustrated in FIG. 29B, the model made a set of predictions and, based on the set of predictions, the central base position in the sequence was determined be an A. In the next step of inference, the A was included in the input sequence, and the model may used to make another set of predictions. Based on the new set of predictions, the second base position in the input sequence was determined to be an A. The model can predict all masked bases simultaneously in each step. One can then choose to "accept" just one base, two bases, etc., or all bases at each step.
[0280] The process may then be repeated, one or more bases at a time, for a specified number of iterations (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more than 15 iterations) until an entirely new sequence had been generated based on the functional genomics propertyMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-162 of 94 track(s). The functional genomics prediction model can be used to iteratively sample positions one at a time in designed sequences that are conditioned on specific functional genomic properties.
[0281] As an example, a functional genomics prediction model was trained using three different DNase-seq tracks from different cell lines (z.e., DNase-seq data for GM12878, Hl- hESC, and fibroblast cells). These were publicly-available data sets from the Encode portal. The model was trained at two different functional property track lengths (1024 bp and 2048 bp). The results presented in the following figures are for the model trained on 1,024 bp functional property tracks. In this case, a uniform U(0,l) masking rate was used that was independent of base position. The masking is performed in two steps. First, a per-base masking rate parameter is determined by sampling from a Uniform distribution between 0 and 1. Then, the sampled rate is used for masking each base independently.
[0282] Examples of predicted DNase-seq data profiles corresponding to regulatory elements designed using a functional genomics prediction model prompted with three different cell type-specific functional properties are shown in FIG. 30, FIG. 32, FIG. 33, and FIG. 34. In these examples, the indicated DNase-seq data tracks were used as input to the model, and the model was used to design regulatory elements that would generate the input functional property track in the specified cell type. The DNase-seq data tracks for the designed sequence were then predicted using the Borzoi model (Linder et al. (2023)). In each of FIG. 30, FIG. 32, FIG. 33, and FIG. 34, the three cell type-specific DNase-seq data tracks used as input are shown at the left. Examples of the Borzoi -predicted DNase-seq data tracks for the resulting designed sequence are shown for the three cell types in the center of the figure. Averaged plots of the Borzoi -predicted DNase-seq data tracks for 128 of the resulting designed sequences are shown for the three cell types at the right.
[0283] In some of these examples (e.g., see FIG. 30 and FIG. 31), the model was used to generate a regulatory sequence that will have a functional property profile (e.g., a DNase-seq data profile) that is basically off in two of the cell lines and is high in the other cell line. The Borzoi model was used to predict the functional property profile for the designed sequence. Borzoi is a sequence-to-track only baseline model that has been trained to predict functional genomic property tracks from DNA sequence. As can be seen in these examples, the functional genomics prediction model disclosed herein does generate regulatory sequences that areMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-163 of 94 predicted to have high activity in one cell type, and low or no activity in the other cell types. On average, the regulatory sequences generated by the model have similar profiles to what were input into the functional genomics prediction model. In some examples, the model was prompted with functional property tracks (e.g., DNase-seq data profiles) with almost no signal (FIG. 32) or high signal (FIG. 33) in all three cell types, and again the model generated sequences that were predicted to have the correct corresponding functional property tracks.
[0284] In the example shown in FIG. 34, the system may go beyond generating short regulatory sequences that will invoke simple functional property tracks, to designing longer regulatory domains that can induce a more complex functional profile. In this example, the DNase-seq data profiles for two cells types (GM12878 and Fibroblasts) are similar, while that for the Hl-hESC cells exhibits a bimodal peak instead of a single. As can be seen, in some cases the model was able to generate sequences that also induce this complex profile.
[0285] FIG. 35 provides an example of motif analysis data that indicates that the generated sequences for different functional property inputs are enriched for known cell type-specific motifs. When the functional genomics prediction model was used to design regulatory sequences for a specific cell type, the generated sequences tended to have motifs that are important for that cell type or specific to that cell type. When the functional genomics prediction model was used to generate regulatory sequences that induce high functional property signal across all cell types, the generated sequences tended to include motifs like CTCF (a highly conserved transcription factor regulating the 3D structure of chromatin), which is known to have ubiquitous activity that is not cell-type dependent.
[0286] The functional genomics prediction model is also capable of generating regulatory sequences that induce quantitative functional activity (e.g., medium activity or high activity) for a specific cell type (FIG. 36, showing a comparison of the predicted versus measured functional activity of generated DNA regulatory sequences based on stratification of the peaks in their corresponding functional genomics property profiles). Based on a comparison of actual genomic sequences to the designed sequences, we have shown that the model is able to generate sequences that have a spectrum of predicted activity (FIG. 37, showing a comparison of functional genomics property profiles for actual DNA regulatory sequences and generated DNA regulatory sequences in the three different cell types), and that are distinct from the sequences that the model was trained on (FIG. 38, showing a histogram showing the minimumMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-164 of 94Levenshtein distance (edit distance) between generated DNA regulatory sequences or test DNA regulatory sequences and all DNA sequences in the model training dataset). For a model trained using all genomic sequence data except sequences from chromosome 7, the plot in FIG. 38 illustrates that a designed sequence generated by the model is as different from other training sequences as any sequence in chromosome 7 is from any of the training sequences, z.e., the model isn’t just copying sequences that have been included in the training data but rather is generating truly novel sequences.Example 7 - Predicting the Effects of Functional Perturbations on Gene Expression
[0287] This example illustrates the use of the disclosed functional genomics prediction model to predict the effect of “functional perturbations” (such as epigenetic modifications) on gene expression.
[0288] The disclosed model is capable of predicting the effects of variants in DNA sequence. However, in experiments such as CRISPR interference (CRISPRi) and CRISPR- mediated transcriptional activation (CRISPRa), perturbations are made not to the DNA sequence itself but instead to the epigenome, which can result in downstream effects on gene expression of nearby genes. These downstream effects can be captured by perturbations to functional genomic property profiles, such as CAGE-seq assay data, Histone ChlP-seq assay data (a technique used to map the genome-wide locations of histone modification), ATAC-seq assay data, and / or DNase-seq assay data.
[0289] It would be useful to be able to predict the effect of such “functional perturbations” on gene expression of nearby genes, as captured by, e.g., CAGE-seq data and other functional genomic property tracks.
[0290] FIG. 39 provides a schematic illustration of a model trained to predict K562 DNase and CAGE-seq data (masked in the input training data) from DNA sequence and H3K27ac ChlP-seq data (masked in the output data). The trained model was then used to predict the effect of perturbations to the input functional data track (H3K27ac ChlP-seq data) on the DNase-seq and Cage-seq data profile predictions without making any changes to the input DNA sequence.
[0291] FIG. 40A provides a non-limiting example of reference H3K27ac ChlP-seq data for K562 cells. FIG. 40B provides a non-limiting example of H3K27ac ChlP-seq data for an inMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-165 of 94 silico CRISPRi enhancer knockout generated by replacing the reference signal at the enhancer midpoint with background signal.
[0292] FIG. 41 shows the input reference H3K27ac ChlP-seq data for K562 cells (upper panel), and the predicted CAGE-seq data output by the model (lower panel). FIG. 42 shows the input H3K27ac ChlP-seq data for the in silico CRISPRi enhancer knockout in K562 cells (upper panel), and the predicted CAGE-seq data output by the model (lower panel). In this case, the model predicted a dramatic reduction in the predicted CAGE-seq signal after modification of the input H3K27ac ChlP-seq signal without any change to the input DNA sequence.
[0293] The disclosed model can thus be used to predict the effects of perturbations to functional genomics properties, such as epigenetic modifications generated using CRISPRi and CRISPRa techniques.X. Example Computer System
[0294] FIG. 43 is a block diagram of a computer system 4300, in accordance with some embodiments. Computer system 4300 can be an example of one implementation for computing platform 102 described above in FIG. 1.
[0295] FIG. 43 illustrates an example computing system 4300 that may be utilized to implement any of the methods and / or machine learning models described herein. In certain embodiments, the computing system 4300 may perform one or more steps of one or more methods described or illustrated herein. In certain embodiments, the computing system 4300 provide functionality described or illustrated herein. In certain embodiments, software running on the computing system 3900 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Certain embodiments include one or more portions of the computing systems 4300. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Moreover, reference to a computer system may encompass one or more computer systems, where appropriate.
[0296] This disclosure contemplates any suitable number of computing systems 4300. This disclosure contemplates computing system 4300 taking any suitable physical form. As example and not by way of limitation, computing system 4300 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-moduleMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-166 of 94(COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computing system 4300 may include one or more computing systems 4300; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks.
[0297] Where appropriate, the computing system 4300 may perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example, and not by way of limitation, the computing system 4300 may perform in real time or in batch mode one or more steps of one or more methods described or illustrated herein. The computing system 4300 may perform at different times or at different locations one or more steps of one or more methods described or illustrated herein, where appropriate.
[0298] In certain embodiments, the computing system 4300 includes a processor 4302, memory 4304, storage 4306, an input / output (I / O) interface 4308, a communication interface 4310, and a bus 4312. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement. In certain embodiments, processor 4302 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions, processor 4302 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 4304, or storage 4306; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 4304, or storage 4306. In certain embodiments, processor 4302 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 4302 including any suitable number of any suitable internal caches, where appropriate. As an example, and not by way of limitation, processor 4302 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches may be copies of instructions in memory 4304 or storage 4306, and the instruction caches may speed up retrieval of those instructions by processor 4302.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-167 of 94
[0299] Data in the data caches may be copies of data in memory 4304 or storage 4306 for instructions executing at processor 4302 to operate on; the results of previous instructions executed at processor 4302 for access by subsequent instructions executing at processor 4302 or for writing to memory 4304 or storage 4306; or other suitable data. The data caches may speed up read or write operations by processor 4302. The TLBs may speed up virtual-address translation for processor 4302. In certain embodiments, processor 4302 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 4302 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 4302 may include one or more arithmetic logic units (ALUs); be a multi-core processor; or include one or more processors 4302. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.
[0300] In certain embodiments, memory 4304 includes main memory for storing instructions for processor 4302 to execute or data for processor 4302 to operate on. As an example, and not by way of limitation, the computing system 4300 may load instructions from storage 4306 or another source (such as, for example, another computing system 4300) to memory 4304. Processor 4302 may then load the instructions from memory 4304 to an internal register or internal cache. To execute the instructions, processor 4302 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 4302 may write one or more results (which may be intermediate or final results) to the internal register or internal cache. Processor 4302 may then write one or more of those results to memory 4304.
[0301] In certain embodiments, processor 4302 executes only instructions in one or more internal registers or internal caches or in memory 4304 (as opposed to storage 4306 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 4304 (as opposed to storage 4306 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple processor 4302 to memory 4304. Bus 4312 may include one or more memory buses, as described below. In certain embodiments, one or more memory management units (MMUs) reside between processor 4302 and memory 4304 and facilitate accesses to memory 4304 requested by processor 4302. In certain embodiments, memory 4304 includes random access memory (RAM). This RAM may beMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-168 of 94 volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Moreover, where appropriate, this RAM may be singleported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 4304 may include one or more memory devices 4304, where appropriate. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.
[0302] In certain embodiments, storage 4306 includes mass storage for data or instructions. As an example, and not by way of limitation, storage 4306 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Storage 4306 may include removable or non-removable (or fixed) media, where appropriate. Storage 4306 may be internal or external to the computing system 4300, where appropriate. In certain embodiments, storage 4306 is non-volatile, solid-state memory. In certain embodiments, storage 4306 includes read-only memory (ROM). Where appropriate, this ROM may be maskprogrammed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 4306 taking any suitable physical form. Storage 4306 may include one or more storage control units facilitating communication between processor 4302 and storage 4306, where appropriate. Where appropriate, storage 4306 may include one or more storages 4306. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.
[0303] In certain embodiments, I / O interface 4308 includes hardware, software, or both, providing one or more interfaces for communication between the computing system 4300 and one or more EO devices. The computing system 4300 may include one or more of these EO devices, where appropriate. One or more of these EO devices may enable communication between a person and the computing system 4300. As an example, and not by way of limitation, an EO device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable EO device or a combination of two or more of these. An EO device may include one or more sensors. This disclosure contemplates any suitable EO devices and any suitable EO interfaces 3908 for them. Where appropriate, EO interface 4308 may include one or more device orMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-169 of 94 software drivers enabling processor 4302 to drive one or more of these EO devices. EO interface 4308 may include one or more EO interfaces 4308, where appropriate. Although this disclosure describes and illustrates a particular EO interface, this disclosure contemplates any suitable EO interface.
[0304] In certain embodiments, communication interface 4310 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between the computing system 4300 and one or more other computer systems 4300 or one or more networks. As an example, and not by way of limitation, communication interface 4310 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 4310 for it.
[0305] As an example, and not by way of limitation, the computing system 4300 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, the computing system 4300 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless network or a combination of two or more of these. The computing system 4300 may include any suitable communication interface 4310 for any of these networks, where appropriate. Communication interface 4310 may include one or more communication interfaces 4310, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.
[0306] In certain embodiments, bus 4312 includes hardware, software, or both coupling components of the computing system 4300 to each other. As an example, and not by way of limitation, bus 4312 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), aMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-170 of 94HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Bus 4312 may include one or more buses 4312, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0307] Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field- programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.
[0308] FIG. 44 illustrates a diagram 4400 of an example artificial intelligence (Al) architecture 4402 (which can be included as part of the one or more computing device(s) 4300 as discussed above with respect to FIG. 43) that can be utilized to determined one or more gene expression and / or functional genomics property predictions, in accordance with the disclosed embodiments. In certain embodiments, the Al architecture 4402 can be implemented utilizing, for example, one or more processing devices that may include hardware (e.g., a general purpose processor, a graphic processing unit (GPU), an application-specific integrated circuit (ASIC), a system-on-chip (SoC), a microcontroller, a field-programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), and / or other processing device(s) that can be suitable for processing various molecular data and making one or more decisions based thereon), software (e.g., instructionsMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-171 of 94 running / executing on one or more processing devices), firmware (e.g., microcode), or some combination thereof.
[0309] In certain embodiments, as depicted by FIG. 44, the Al architecture 4402 may include machine learning (ML) algorithms and functions 4404, natural language processing (NLP) algorithms and functions 4406, expert systems 4408, computer-based vision algorithms and functions 4410, speech recognition algorithms and functions 4412, planning algorithms and functions 4414, and robotics algorithms and functions 4416. In certain embodiments, the ML algorithms and functions 4404 may include any statistics-based algorithms that can be suitable for finding patterns across large amounts of data e.g., “Big Data” such as genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data). For example, in certain embodiments, the ML algorithms and functions 4404 may include deep learning algorithms 4418, supervised learning algorithms 4420, and unsupervised learning algorithms 4422.
[0310] In certain embodiments, the deep learning algorithms 4418 may include any artificial neural networks (ANNs) that can be utilized to learn deep levels of representations and abstractions from large amounts of data. For example, the deep learning algorithms 4418 may include ANNs, such as a perceptron, a multilayer perceptron (MLP), an autoencoder (AE), a convolution neural network (CNN), a recurrent neural network (RNN), long short term memory (LSTM), a grated recurrent unit (GRU), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and deep Q-networks, a neural autoregressive distribution estimation (NADE), an adversarial network (AN), attentional models (AM), a spiking neural network (SNN), deep reinforcement learning, and so forth.
[0311] In certain embodiments, the supervised learning algorithms 4420 may include any algorithms that can be utilized to apply, for example, what has been learned in the past to new data using labeled examples for predicting future events. For example, starting from the analysis of a known training dataset, the supervised learning algorithms 4420 may produce an inferred function to make predictions about the output values. The supervised learning algorithms 4420 may also compare its output with the correct and intended output and find errors in order to modify the supervised learning algorithms 4420 accordingly. On the other hand, the unsupervised learning algorithms 4422 may include any algorithms that may applied,MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-172 of 94 for example, when the data used to train the unsupervised learning algorithms 4422 are neither classified nor labeled. For example, the unsupervised learning algorithms 4422 may study and analyze how systems may infer a function to describe a hidden structure from unlabeled data.
[0312] In certain embodiments, the NLP algorithms and functions 4406 may include any algorithms or functions that can be suitable for automatically manipulating natural language, such as speech and / or text. For example, the NLP algorithms and functions 4406 may include content extraction algorithms or functions 4424, classification algorithms or functions 4426, machine translation algorithms or functions 4428, question answering (QA) algorithms or functions 4430, and text generation algorithms or functions 4432. In certain embodiments, the content extraction algorithms or functions 4424 may include a means for extracting text or images from electronic documents (e.g., webpages, text editor documents, and so forth) to be utilized, for example, in other applications.
[0313] In certain embodiments, the classification algorithms or functions 4426 may include any algorithms that may utilize a supervised learning model (e.g., logistic regression, naive Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machine (SVM), and so forth) to learn from the data input to the supervised learning model and to make new observations or classifications based thereon. The machine translation algorithms or functions 4428 may include any algorithms or functions that can be suitable for automatically converting source text in one language, for example, into text in another language. The QA algorithms or functions 4430 may include any algorithms or functions that can be suitable for automatically answering questions posed by humans in, for example, a natural language, such as that performed by voice-controlled personal assistant devices. The text generation algorithms or functions 4432 may include any algorithms or functions that can be suitable for automatically generating natural language texts.
[0314] In certain embodiments, the expert systems 4408 may include any algorithms or functions that can be suitable for simulating the judgment and behavior of a human or an organization that has expert knowledge and experience in a particular field (e.g., stock trading, medicine, sports statistics, and so forth). The computer-based vision algorithms and functions 4010 may include any algorithms or functions that can be suitable for automatically extracting information from images (e.g., photo images, video images). For example, the computer-based vision algorithms and functions 4410 may include image recognition algorithms 4434 andMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-173 of 94 machine vision algorithms 4436. The image recognition algorithms 4434 may include any algorithms that can be suitable for automatically identifying and / or classifying objects, places, people, and so forth that can be included in, for example, one or more image frames or other displayed data. The machine vision algorithms 4436 may include any algorithms that can be suitable for allowing computers to “see”, or, for example, to rely on image sensors cameras with specialized optics to acquire images for processing, analyzing, and / or measuring various data characteristics for decision making purposes.
[0315] In certain embodiments, the speech recognition algorithms and functions 4412 may include any algorithms or functions that can be suitable for recognizing and translating spoken language into text, such as through automatic speech recognition (ASR), computer speech recognition, speech -to-text (STT) 4438, or text-to-speech (TTS) 4440 in order for the computing to communicate via speech with one or more users, for example. In certain embodiments, the planning algorithms and functions 4414 may include any algorithms or functions that can be suitable for generating a sequence of actions, in which each action may include its own set of preconditions to be satisfied before performing the action. Examples of Al planning may include classical planning, reduction to other problems, temporal planning, probabilistic planning, preference-based planning, conditional planning, and so forth. Lastly, the robotics algorithms and functions 4416 may include any algorithms, functions, or systems that may enable one or more devices to replicate human behavior through, for example, motions, gestures, performance tasks, decision-making, emotions, and so forth.XI. Example Descriptions of Terms
[0316] As used herein, “or” is inclusive and not exclusive, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A orB” means “A, B, or both,” unless expressly indicated otherwise or indicated otherwise by context. Moreover, “and” is both joint and several, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A and B” means “A and B, jointly or severally,” unless expressly indicated otherwise or indicated otherwise by context.
[0317] As used herein, “automatically” and its derivatives means “without human intervention,” unless expressly indicated otherwise or indicated otherwise by context.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-174 of 94
[0318] As used herein, the term “nucleic acid”, “nucleic acid molecule”, or “oligonucleotide” are used interchangeably to refer to a polymer of nucleotide residues. The terms encompass oligonucleotides of any length.
[0319] As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably to refer to a polymer of amino acid residues. The terms encompass amino acid chains of any length, including full-length proteins with amino acid residues linked by covalent peptide bonds.
[0320] As used herein, a “sequence” refers to an oligonucleotide sequence or an amino acid sequence that includes an ordered set of nucleotides or amino acid identifiers, respectively.
[0321] As used herein, the terms “nucleic acid sequence” or “oligonucleotide sequence” are used interchangeably and refer to a sequence that identifies nucleotides of at least a portion of a nucleic acid molecule. In some cases, the nucleic acid sequence includes a variant sequence that includes a variant that is not observed in a corresponding reference sequence.
[0322] As used herein, a “peptide sequence” refers to a sequence that identifies amino acids of at least a portion of a peptide. In some cases, the peptide sequence includes a variant-coding sequence that includes a variant that is not observed in a corresponding reference sequence.
[0323] As used herein, a “reference sequence” may refer to a sequence that identifies nucleotides or amino acids within at least part of a non-mutant nucleic acid molecule or protein molecule, respectively, or to a wild-type nucleic acid sequence or protein sequence (e.g., a wild-type, parental sequence).
[0324] As used herein, a “representation” of a sequence or “sequence representation” can include a set of values that represent or identify amino acids in a protein or peptide sequence and / or a set of values that represent or identify nucleotides in a nucleic acid that encodes the protein or peptide sequence. For example, each nucleotide or amino acid can be represented by a binary string and / or vector of values that is distinct from every other binary string and / or vector of values representing every other nucleotide or amino acid. The sequence representation can be generated using, for example, one-hot encoding or using a BLOcks Substitution Matrix (BLOSUM) matrix. For example, a multi-dimensional (e.g., 20- or 21 -dimensional) array be initialized (e.g., randomly or pseudo-randomly initialized). The initialized array may include, for each nucleotide or amino acid, a unique vector corresponding to that nucleotide or amino acid. The values can be fixed such that use of such a unique vectorMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-175 of 94 can be assumed to represent the corresponding nucleotide or amino acid. There can be multiple possible nucleic acid representations of a given protein or peptide sequence, given that any of multiple codons can encode a single amino acid.
[0325] As used herein, a “sample” can include tissue (e.g., a biopsy), single cell, multiple cells, fragments of cells, or an aliquot of body fluid. The sample can be obtained from a subject by means such as, for example, without limitation, venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, intervention, another type of sample collection means, or a combination thereof.
[0326] As used herein, a “subject” encompasses one or more cells, tissue, or an organism. The subject can be a human or non-human, whether in vivo, ex vivo, or in vitro, male or female. A subject can be a mammal, such as a human.
[0327] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.XII. Example Embodiments
[0328] Embodiments disclosed herein include:1. A method for predicting a functional genomic property, comprising: obtaining an input genomic sequence derived from a test sample; obtaining at least one input functional genomic property dataset derived from the test sample; formulating an input data structure based on the input genomic sequence and the at least one input functional genomic property dataset; and predicting at least one output functional genomic property for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by:MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-176 of 94 obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the plurality of masked training input data structures using a multiscale loss function.2. The method of embodiment 1, wherein the trained machine learning model is configured to receive a masked input data structure and predict the masked portion of the masked input data structure.3. The method of embodiment 1 or embodiment 2, wherein the at least one predicted output functional genomic property for the test sample is different from the at least one input functional genomic property dataset derived from the test sample.4. The method of any one of embodiments 1 to 3, wherein the trained machine learning model comprises a multimodal machine learning model.5. The method of embodiment 4, wherein the multimodal machine learning model is configured to process genomic sequence data and at least 2, 3, 4, 5, 6, 7, 8 ,9, or 10 different functional genomic property datasets derived from the test sample as input.6. The method of any one of embodiments 1 to 5, wherein the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-177 of 947. The method of any one of embodiments 1 to 6, wherein the plurality of training samples comprise samples from at least one cell type and / or at least one species.8. The method of any one of embodiments 1 to 7, wherein the at least one input functional genomic property dataset comprises ATAC-Seq data, CAGE-Seq data, ChlP-Seq data, DNase-Seq data, H3K27ac-Seq data, or any combination thereof.9. The method of any one of embodiments 1 to 8, wherein the at least one output functional genomic property prediction comprises an ATAC-Seq activity, a CAGE-Seq activity, a ChlP- Seq activity, a DNase-Seq activity, an H3K27ac-Seq activity, or any combination thereof for at least one genomic locus.10. The method of any one of embodiments 1 to 9, wherein the prediction output by the trained machine learning model is used to predict a functional effect of a genomic variant.11. The method of any one of embodiments 1 to 10, wherein the prediction output by the trained machine learning model is used to interpret a functional effect of a genomic variant.12. The method of any one of embodiments 1 to 11, wherein the prediction output by the trained machine learning model is used to denoise a functional genomic property dataset for the test sample.13. The method of any one of embodiments 1 to 12, wherein the prediction output by the trained machine learning model is used to identify dependencies between functional genomic properties in the test sample.14. The method of any one of embodiments 1 to 13, wherein the prediction output by the trained machine learning model is used to diagnose a disease.15. The method of any one of embodiments 1 to 14, wherein the prediction output by the trained machine learning model is used to select a treatment for a disease.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-178 of 9416. The method of any one of embodiments 1 to 15, wherein the prediction output by the trained machine learning model is used to inform a disease prognosis.17. A method for predicting a base at at least one specified genomic locus in a genomic sequence, comprising: obtaining the genomic sequence derived from a test sample, the genomic sequence comprising an unknown base at the at least one specified genomic locus; obtaining at least one functional genomic property dataset derived from the test sample; formulating an input data structure based on the genomic sequence and the at least one functional genomic property dataset; and predicting the base at the at least one specified genomic locus in the genomic sequence for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by: obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask at least one base in the genomic sequence of each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the masked training input data structures using a multiscale loss function to predict the at least one masked base.18. The method of embodiment 17, wherein the pre-defined masking strategy comprises a variable masking strategy in which a different number of genomic loci and / or different genomic loci are masked in different training input data structures of the plurality of training input data structures.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-1 79 of 9419. The method of embodiment 17 or embodiment 18, further comprising: identifying an individual associated with the test sample based on the predicted base at the at least one specified genomic locus for the test sample.20. The method of embodiment 19, wherein identifying the individual associated with the test sample comprises: obtaining genomic sequences derived from test samples from each of a plurality of individuals; obtaining at least one functional genomic property dataset derived from each of the plurality of test samples; formulating a plurality of input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of test samples; inputting the plurality of input data structures into the trained machine learning model; outputting a predicted base at the at least one specified genomic locus in the genomic sequence for each of the plurality of test samples; and determining an identity of at least one individual based on the predicted base at the at least one specified genomic locus in the genomic sequence for each of the plurality of test samples.21. The method of any one of embodiments 17 to 20, wherein the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.22. The method of any one of embodiments 17 to 21, wherein the plurality of training samples comprise samples from at least one cell type and / or at least one species.23. A method for designing a genomic regulatory element comprising:MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-180 of 94 obtaining at least one functional genomic property dataset generated based on a desired regulatory effect in a test sample; formulating an input data structure based on a candidate genomic sequence of specified length and the at least one functional genomic property dataset; predicting a base at one or more positions in the candidate genomic sequence by: inputting the input data structure into a trained machine learning model, the trained machine learning model configured to output a prediction of a base at one or more positions of a genomic sequence based on at least one input functional genomic property dataset; and substituting the predicted base at the one or more positions of the candidate genomic sequence for the previous base at the one or more positions of the candidate genomic sequence in the input data structure; and outputting the candidate genomic sequence as a designed genomic regulatory element for the test sample.24. The method of embodiment 23, further comprising repeating the steps of: (i) inputting the data structure into the trained machine learning model and (ii) substituting the predicted base at the one or more positions of the candidate genomic sequence for the previous base at the one or more positions of the candidate genomic sequence, for two or more iterations.25. The method of embodiment 23 and embodiment 24, wherein the designed genomic regulatory element comprises an adeno-associated virus (AAV) vector.26. The method of any one of embodiments 23 to 25, wherein the trained machine learning model is trained on training data comprising genomic sequences and the at least one functional genomic property dataset derived from each of a plurality of training samples.27. The method of embodiment 26, wherein the training data is formulated as a plurality of masked training input data structures generated from a plurality of training input data structures using a variable masking strategy in which a different number of genomic lociMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-181 of 94 and / or different genomic loci are masked in different training input data structures of the plurality of training input data structures.28. The method of embodiment 26 or embodiment 27, wherein the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.29. The method of any one of embodiments 26 to 28, wherein the plurality of training samples comprise samples from at least one cell type and / or at least one species.30. A method for predicting a genomic sequence and / or a functional genomic property, comprising: formulating an input data structure comprising a genomic sequence data component and / or at least one functional genomic property data component derived from a test sample, wherein at least a portion of the genomic sequence data component and / or the at least one functional genomic property data component is unknown; inputting the input data structure into a trained machine learning model configured to predict the unknown portion of the genomic sequence component and / or the at least one functional genomic property component, the machine learning model trained by: obtaining genomic sequences and / or at least one functional genomic property dataset derived from each of a plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and / or the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the genomic sequence and / or at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model to receive a masked training input data structure and predict the masked portion of the masked training input data structure; andMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-182 of 94 outputting a prediction for the unknown portion of the genomic sequence data component and / or the at least one functional genomic property data component for the test sample.31. The method of embodiment 30, wherein the trained machine learning model comprises a masked multi-modal machine learning model configured to process genomic sequence data and / or 2, 3, 4, 5, 6, 7, 8, 9, or 10 different functional genomic property datasets derived from the test sample.32. The method of embodiment 30 or embodiment 31, wherein the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.33. The method of any one of embodiments 30 to 32, wherein the plurality of training samples comprise samples from at least one cell type and / or at least one species.34. The method of any one of embodiments 30 to 33, wherein the at least one functional genomic property prediction comprises an ATAC-Seq activity, a CAGE-Seq activity, a ChlP- Seq activity, a DNase-Seq activity, an H3K27ac-Seq activity, or any combination thereof for at least one genomic locus.35. The method of any one of embodiments 30 to 34, wherein the prediction output by the trained machine learning model is used to predict a functional effect of a genomic variant.36. The method of any one of embodiments 30 to 35, wherein the prediction output by the trained machine learning model is used to interpret a functional effect of a genomic variant.37. The method of any one of embodiments 30 to 36, wherein the prediction output by the trained machine learning model is used to denoise a functional genomic property dataset for the test sample.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-183 of 9438. The method of any one of embodiments 30 to 37, wherein the prediction output by the trained machine learning model is used to identify dependencies between functional genomic properties in the test sample.39. The method of any one of embodiments 30 to 38, wherein the prediction output by the trained machine learning model is used to diagnose a disease.40. The method of any one of embodiments 30 to 39, wherein the prediction output by the trained machine learning model is used to select a treatment for a disease.41. The method of any one of embodiments 30 to 40, wherein the prediction output by the trained machine learning model is used to inform a disease prognosis.42. The method of any one of embodiments 1 to 41, wherein the input data structure comprises a matrix.43. The method of embodiment 42, wherein at least one row of the matrix comprises genomic sequence data.44. The method of embodiment 42 or embodiment 43, wherein the matrix comprises a first row comprising a first functional genomic property dataset.45. The method of any one of embodiments 42 to 44, wherein the matrix comprises at least a second row comprising at least a second functional genomic property dataset.46. The method of any one of embodiments 20 to 45, wherein the trained machine learning model is trained using a multiscale loss function.47. The method of any one of embodiments 1 to 46, wherein the test sample comprises a cell sample, a blood sample, a biopsy sample, or a tissue sample.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-184 of 9448. The method of any one of embodiments 1 to 47, wherein the at least one functional genomic property dataset comprises ATAC-Seq data, CAGE-Seq data, ChlP-Seq data, DNase-Seq data, H3K27ac-Seq data, or any combination thereof.49. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method of any one of embodiments 1 to 48.50. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to perform the method of any one of embodiments 1 to 48.
[0329] The description provides preferred example embodiments only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the description of the preferred example embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It is understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.MOFO-358127518
Claims
ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-185 of 94CLAIMSWhat is claimed is:
1. A method for predicting a functional genomic property, comprising: obtaining an input genomic sequence derived from a test sample; obtaining at least one input functional genomic property dataset derived from the test sample; formulating an input data structure based on the input genomic sequence and the at least one input functional genomic property dataset; and predicting at least one output functional genomic property for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by: obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the plurality of masked training input data structures using a multiscale loss function.
2. The method of claim 1, wherein the trained machine learning model is configured to receive a masked input data structure and predict the masked portion of the masked input data structure.
3. The method of claim 1 or claim 2, wherein the at least one predicted output functional genomic property for the test sample is different from the at least one input functional genomic property dataset derived from the test sample.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-186 of 944. The method of any one of claims 1 to 3, wherein the trained machine learning model comprises a multimodal machine learning model.
5. The method of claim 4, wherein the multimodal machine learning model is configured to process genomic sequence data and at least 2, 3, 4, 5, 6, 7, 8 ,9, or 10 different functional genomic property datasets derived from the test sample as input.
6. The method of any one of claims 1 to 5, wherein the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.
7. The method of any one of claims 1 to 6, wherein the plurality of training samples comprise samples from at least one cell type and / or at least one species.
8. The method of any one of claims 1 to 7, wherein the at least one input functional genomic property dataset comprises ATAC-Seq data, CAGE-Seq data, ChlP-Seq data, DNase-Seq data, H3K27ac-Seq data, or any combination thereof.
9. The method of any one of claims 1 to 8, wherein the at least one output functional genomic property prediction comprises an ATAC-Seq activity, a CAGE-Seq activity, a ChlP-Seq activity, a DNase-Seq activity, an H3K27ac-Seq activity, or any combination thereof for at least one genomic locus.
10. The method of any one of claims 1 to 9, wherein the prediction output by the trained machine learning model is used to predict a functional effect of a genomic variant.
11. The method of any one of claims 1 to 10, wherein the prediction output by the trained machine learning model is used to interpret a functional effect of a genomic variant.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-187 of 9412. The method of any one of claims 1 to 11, wherein the prediction output by the trained machine learning model is used to denoise a functional genomic property dataset for the test sample.
13. The method of any one of claims 1 to 12, wherein the prediction output by the trained machine learning model is used to identify dependencies between functional genomic properties in the test sample.
14. The method of any one of claims 1 to 13, wherein the prediction output by the trained machine learning model is used to diagnose a disease.
15. The method of any one of claims 1 to 14, wherein the prediction output by the trained machine learning model is used to select a treatment for a disease.
16. The method of any one of claims 1 to 15, wherein the prediction output by the trained machine learning model is used to inform a disease prognosis.
17. A method for predicting a base at at least one specified genomic locus in a genomic sequence, comprising: obtaining the genomic sequence derived from a test sample, the genomic sequence comprising an unknown base at the at least one specified genomic locus; obtaining at least one functional genomic property dataset derived from the test sample; formulating an input data structure based on the genomic sequence and the at least one functional genomic property dataset; and predicting the base at the at least one specified genomic locus in the genomic sequence for the test sample by inputting the input data structure into a trained machine learning model, the machine learning model trained by: obtaining genomic sequences derived from a plurality of training samples; obtaining at least one functional genomic property dataset derived from each of the plurality of training samples;MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-188 of 94 formulating a plurality of training input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask at least one base in the genomic sequence of each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model based on the masked training input data structures using a multiscale loss function to predict the at least one masked base.
18. The method of claim 17, wherein the pre-defined masking strategy comprises a variable masking strategy in which a different number of genomic loci and / or different genomic loci are masked in different training input data structures of the plurality of training input data structures.
19. The method of claim 17 or claim 18, further comprising: identifying an individual associated with the test sample based on the predicted base at the at least one specified genomic locus for the test sample.
20. The method of claim 19, wherein identifying the individual associated with the test sample comprises: obtaining genomic sequences derived from test samples from each of a plurality of individuals; obtaining at least one functional genomic property dataset derived from each of the plurality of test samples; formulating a plurality of input data structures based on the genomic sequences and the at least one functional genomic property dataset derived from each of the plurality of test samples; inputting the plurality of input data structures into the trained machine learning model; outputting a predicted base at the at least one specified genomic locus in the genomic sequence for each of the plurality of test samples; andMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-189 of 94 determining an identity of at least one individual based on the predicted base at the at least one specified genomic locus in the genomic sequence for each of the plurality of test samples.
21. The method of any one of claims 17 to 20, wherein the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.
22. The method of any one of claims 17 to 21, wherein the plurality of training samples comprise samples from at least one cell type and / or at least one species.
23. A method for designing a genomic regulatory element comprising: obtaining at least one functional genomic property dataset generated based on a desired regulatory effect in a test sample; formulating an input data structure based on a candidate genomic sequence of specified length and the at least one functional genomic property dataset; predicting a base at one or more positions in the candidate genomic sequence by: inputting the input data structure into a trained machine learning model, the trained machine learning model configured to output a prediction of a base at one or more positions of a genomic sequence based on at least one input functional genomic property dataset; and substituting the predicted base at the one or more positions of the candidate genomic sequence for the previous base at the one or more positions of the candidate genomic sequence in the input data structure; and outputting the candidate genomic sequence as a designed genomic regulatory element for the test sample.
24. The method of claim 23, further comprising repeating the steps of: (i) inputting the data structure into the trained machine learning model and (ii) substituting the predicted base at the one or more positions of the candidate genomic sequence for the previous base at the one or more positions of the candidate genomic sequence, for two or more iterations.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-1 90 of 9425. The method of claim 23 and claim 24, wherein the designed genomic regulatory element comprises an adeno-associated virus (AAV) vector.
26. The method of any one of claims 23 to 25, wherein the trained machine learning model is trained on training data comprising genomic sequences and the at least one functional genomic property dataset derived from each of a plurality of training samples.
27. The method of claim 26, wherein the training data is formulated as a plurality of masked training input data structures generated from a plurality of training input data structures using a variable masking strategy in which a different number of genomic loci and / or different genomic loci are masked in different training input data structures of the plurality of training input data structures.
28. The method of claim 26 or claim 27, wherein the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.
29. The method of any one of claims 26 to 28, wherein the plurality of training samples comprise samples from at least one cell type and / or at least one species.
30. A method for predicting a genomic sequence and / or a functional genomic property, comprising: formulating an input data structure comprising a genomic sequence data component and / or at least one functional genomic property data component derived from a test sample, wherein at least a portion of the genomic sequence data component and / or the at least one functional genomic property data component is unknown; inputting the input data structure into a trained machine learning model configured to predict the unknown portion of the genomic sequence component and / or the at least one functional genomic property component, the machine learning model trained by:MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-191 of 94 obtaining genomic sequences and / or at least one functional genomic property dataset derived from each of a plurality of training samples; formulating a plurality of training input data structures based on the genomic sequences and / or the at least one functional genomic property dataset derived from each of the plurality of training samples; applying a pre-defined masking strategy to mask a portion of the genomic sequence and / or at least one functional genomic property dataset in each of the training input data structures to generate a plurality of masked training input data structures; and training the machine learning model to receive a masked training input data structure and predict the masked portion of the masked training input data structure; and outputting a prediction for the unknown portion of the genomic sequence data component and / or the at least one functional genomic property data component for the test sample.
31. The method of claim 30, wherein the trained machine learning model comprises a masked multi-modal machine learning model configured to process genomic sequence data and / or 2, 3, 4, 5, 6, 7, 8, 9, or 10 different functional genomic property datasets derived from the test sample.
32. The method of claim 30 or claim 31, wherein the plurality of training samples comprises a plurality of cell samples, blood samples, biopsy samples, tissue samples, or any combination thereof.
33. The method of any one of claims 30 to 32, wherein the plurality of training samples comprise samples from at least one cell type and / or at least one species.
34. The method of any one of claims 30 to 33, wherein the at least one functional genomic property prediction comprises an ATAC-Seq activity, a CAGE-Seq activity, a ChlP-SeqMOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-192 of 94 activity, a DNase-Seq activity, an H3K27ac-Seq activity, or any combination thereof for at least one genomic locus.
35. The method of any one of claims 30 to 34, wherein the prediction output by the trained machine learning model is used to predict a functional effect of a genomic variant.
36. The method of any one of claims 30 to 35, wherein the prediction output by the trained machine learning model is used to interpret a functional effect of a genomic variant.
37. The method of any one of claims 30 to 36, wherein the prediction output by the trained machine learning model is used to denoise a functional genomic property dataset for the test sample.
38. The method of any one of claims 30 to 37, wherein the prediction output by the trained machine learning model is used to identify dependencies between functional genomic properties in the test sample.
39. The method of any one of claims 30 to 38, wherein the prediction output by the trained machine learning model is used to diagnose a disease.
40. The method of any one of claims 30 to 39, wherein the prediction output by the trained machine learning model is used to select a treatment for a disease.
41. The method of any one of claims 30 to 40, wherein the prediction output by the trained machine learning model is used to inform a disease prognosis.
42. The method of any one of claims 1 to 41, wherein the input data structure comprises a matrix.
43. The method of claim 42, wherein at least one row of the matrix comprises genomic sequence data.MOFO-358127518ATTORNEY DOCKET PATENT APPLICATION146392069840-P39632-WO-193 of 9444. The method of claim 42 or claim 43, wherein the matrix comprises a first row comprising a first functional genomic property dataset.
45. The method of any one of claims 42 to 44, wherein the matrix comprises at least a second row comprising at least a second functional genomic property dataset.
46. The method of any one of claims 20 to 45, wherein the trained machine learning model is trained using a multiscale loss function.
47. The method of any one of claims 1 to 46, wherein the test sample comprises a cell sample, a blood sample, a biopsy sample, or a tissue sample.
48. The method of any one of claims 1 to 47, wherein the at least one functional genomic property dataset comprises ATAC-Seq data, CAGE-Seq data, ChlP-Seq data, DNase-Seq data, H3K27ac-Seq data, or any combination thereof.
49. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method of any one of claims 1 to 48.
50. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to perform the method of any one of claims 1 to 48.MOFO-358127518
Citation Information
Patent Citations
Methods and systems for sequence generation and prediction
US20230298698A1