A multi-omics search engine for integrated analysis of cancer genetic and clinical data

The multi-omics data index system addresses the challenges of data access and integration by ranking cancer-specific data indexes based on clinical relevance, enhancing cancer diagnosis and treatment through comprehensive tumor biology analysis.

JP7681817B2Active Publication Date: 2025-05-23HUMAN LONGEVITY INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021520420
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-10-12
Filing Date
2019-10-14
Publication Date
2025-05-23
Estimated Expiration
2039-10-14

AI Technical Summary

Technical Problem

Current systems face challenges in providing immediate access to cancer data, integrating multi-omics datasets for comprehensive tumor biology analysis, and effectively correlating prognostic, diagnostic, and therapeutic information with various data types for clinical insights and actionable biomarkers.

Method used

A multi-omics data index system that stores cancer-specific tokenized data, populates and indexes additional multi-omics data with annotations, and ranks relevant data indexes based on clinical actionability, pathogenicity, feature weights, or frequency in response to user queries.

Benefits of technology

Enables efficient and effective access to cancer data, integrates multi-omics datasets for a complete tumor biology picture, and provides clinically relevant insights and biomarkers, thereby improving cancer diagnosis and treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681817000003
    Figure 0007681817000003
  • Figure 0007681817000004
    Figure 0007681817000004
  • Figure 0007681817000005
    Figure 0007681817000005
Patent Text Reader

Abstract

A method for utilizing multi-omics data indexes for tumor profiling is provided, the method can include storing a plurality of multi-omics data indexes, each of the plurality of multi-omics data indexes including cancer-specific tokenized data, capturing additional multi-omics data and annotations associated with the additional multi-omics data, capturing the additional multi-omics data associated with the one or more indexes, indexing the captured additional multi-omics data and annotations while preserving gene names, gene variant names, and multi-omics mappings between different data streams for the same patient within a particular index to generate tokenized captured additional multi-omics data, receiving a user query, selecting one or more relevant multi-omics data indexes based on the user query, ranking the selected one or more multi-omics data indexes based on at least one of clinical actionability, pathogenicity, feature weight, or frequency, and returning the ranked one or more multi-omics data indexes to the user.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] As the importance of cancer gene sequencing grows, thousands of cancer genes, exomes, transcriptomes, proteomes, and other cancer data are sequenced by both private and public institutions (e.g., The Cancer Genome Atlas [TCGA], International Cancer Genome Consortium [ICGC]). Interpretation and analysis of tumor and normal sequencing data depends on the integrated analysis of both private and public genetic data and databases.

[0002] Industry, biopharmaceutical companies, research institutes, and international cancer consortia face hurdles, for example, (1) providing immediate access to any sample or subset of samples, (2) integrating multi-omics datasets to form a complete picture of tumor biology, and (3) effectively linking prognostic, diagnostic, and therapeutic information across all available data (e.g., genetic, transcriptional, proteomic, functional, medical, imaging, literature, etc.) to provide clinical insights and actionable insights into individual cancer patients as well as potential multi-omics prognostic, diagnostic, or therapeutic biomarkers.

[0003] Currently, publicly available data are scattered across publications, guidelines, and web-based resources. Ultimately, solutions that address the above three issues will enable widespread clinical use of cancer genetic analysis.

[0004] Data integration and harmonization pose particularly serious challenges in cancer sequencing: standardization and integration to enable users to incorporate multiple data sources and identify clinically and biologically relevant information. Furthermore, compared to germline sequence analysis, cancer genetic analysis requires extensive bioinformatics pipelines, generating multi-omics streams of data from the same sample. For example, for a typical cancer biopsy and normal blood sample, binary base calls (BCLs) of tumor DNA, normal DNA, tumor RNA, and possibly normal RNA must be converted to variant call format (VCF) via alignment to reference genes, deduplication, realignment, and variant recalibration. Furthermore, it is generally industry standard to run multiple somatic variant callers to derive a consensus set of somatic single nucleotide polymorphisms (SNVs) and small insertions and deletions (indels). Of even greater interest are pipelines for, for example, tumor copy number variation (CNV) detection, differential gene expression between tumor and normal RNA-Seq replicates, and data processing and gene fusion detection to confirm that mutations detected in somatic (tumor) DNA are also expressed in RNA. Of further interest will be the use of tools that call large structural variants and perform advanced bioinformatics to annotate cancer alterations and calculate relevant tumor characteristics (such as tumor mutation burden, gene mutation signature, microsatellite status, expressed neoantigens, HLA-normal gene typing) and identify clinically relevant tumor alterations.

[0005] Modern cancer profiling techniques can easily generate 25 gigabytes of multi-omics data per sample, meaning that researchers conducting medium-scale cancer biomarker discovery studies are easily faced with terabytes of raw data. Therefore, identifying relevant biomarkers is akin to "finding a needle in a haystack." Furthermore, once an analytical pipeline finishes running, there is virtually no way to interact with the results and generate new hypotheses.

[0006] The most common approach to addressing the issues of cancer data accessibility, multiplex integration, and utility is to design portals that display prefiltered data tables and analyses based on previously curated files and precomputed workflows. Examples of portals include Illumina BaseSpace Correlation Engine and Cohort Analyzer, WuXI nextCODE TCGA Portal, cBioPortal, IntOGen, Tumorscape, Tumorportal, Xena, ICGC Data Portal, St. JudePeCan, and QiagenOmicSoft. However, these portals typically limit the types of questions they can address and the additional analyses that can be performed. Furthermore, the data are typically inaccessible for exploration at many levels of the bioinformatics pipeline. Data within portals is often prefiltered, unintegrated, and typically not ranked. Furthermore, most portals do not host individual user data. Those that allow users to upload their own data typically do not provide a means to integrate their data with portal data, or allow users to derive advanced cancer analytics, making this data accessible and ranking it in terms of clinical actionability, pathogenicity, feature weight, or frequency.

[0007] Therefore, there is a need to provide systems and methods that effectively and efficiently provide immediate access to any sample or subset of samples. There is also a need to provide systems and methods that effectively and efficiently integrate multi-omic datasets to form a complete picture of tumor biology. Furthermore, there is a need to effectively and efficiently link prognostic, diagnostic, and therapeutic information to all available data (e.g., genetic, transcriptional, proteomic, functional, medical, imaging, literature) to stratify individual cancer patients and cohorts of patients regarding potential multi-omic prognostic or therapeutic biomarkers.

[0008] Summary of the Invention

[0009] Profiling. The method may include storing multiple multi-omics data indexes, each of the multiple multi-omics data indexes including cancer-specific tokenized data. The method may further include incorporating any additional multi-omics data and annotations associated with the additional multi-omics data, the additional multi-omics data associated with one or more indexes. The method may further include indexing the acquired additional multi-omics data and annotations to generate tokenized acquired additional multi-omics data while preserving gene names, gene variant names, and multi-omics mappings between different data streams for the same patient within a particular index. The method may further include receiving a user query. The method may further include selecting one or more relevant multi-omics data indexes based on the user query. The method may further include ranking the selected one or more multi-omics data indexes based on at least one of clinical action likelihood, pathogenicity, feature weight, or frequency. The method may further include returning the ranked one or more multi-omics data indexes to the user.

[0010] According to various embodiments, a non-transitory computer-readable medium is provided having stored thereon a program for causing a computer to execute a method for utilizing multi-omics data indexes for tumor profiling. The method may include storing a plurality of multi-omics data indexes, each of the plurality of multi-omics data indexes including cancer-specific tokenized data. The method may further include incorporating additional multi-omics data and annotations associated with the additional multi-omics data, the additional multi-omics data associated with one or more indexes. The method may further include indexing the acquired additional multi-omics data and annotations to generate tokenized acquired additional multi-omics data while preserving gene names, gene variant names, and multi-omics mappings between different data streams for the same patient within a particular index. The method may further include receiving a user query. The method may further include selecting one or more relevant multi-omics data indexes based on the user query. The method may further include ranking the selected one or more multi-omics data indexes based on at least one of clinical action likelihood, pathogenicity, feature weight, or frequency. The method may further include returning the ranked multi-omics data index or indexes to the user.

[0011] According to various embodiments, a system for utilizing a multi-omics data index for tumor profiling is provided. The system may include an indexing unit. The indexing unit may include a storage element configured to store multiple multi-omics data indexes, each of the multiple multi-omics data indexes including cancer-specific tokenized data. The indexing unit may further include an indexing engine. The indexing unit may be configured to retrieve additional multi-omics data and annotations associated with the additional multi-omics data, the additional multi-omics data associated with one or more indexes. The indexing unit may be further configured to index the retrieved additional multi-omics data and annotations while preserving gene names, gene variant names, and multi-omics mappings between different data streams for the same patient within a particular index, generating additional tokenized retrieved multi-omics data. The system may further include a user interface configured to receive a user query. The system may further include a query engine configured to select one or more relevant multi-omics data indexes from the indexing unit based on the user query. The system may further include a ranking engine configured to receive the selected one or more relevant multi-omics data indexes and rank the selected one or more multi-omics data indexes based on at least one of clinical action likelihood, pathogenicity, feature weight, or frequency. The ranking engine may be further configured to return the ranked one or more multi-omics data indexes to a user via a user interface.

[0012] According to various embodiments, a system for utilizing a multi-omics data index for tumor profiling is provided. The system may include an indexing unit. The indexing unit may include a storage element configured to store multiple multi-omics data indexes, each of the multiple multi-omics data indexes including cancer-specific tokenized data. The indexing unit may further include an indexing engine. The indexing unit may be configured to retrieve additional multi-omics data and annotations associated with the additional multi-omics data, and the additional multi-omics data associated with one or more indexes. The indexing unit may be further configured to index the retrieved additional multi-omics data and annotations while preserving gene names, gene variant names, and multi-omics mappings between different data streams for the same patient within a particular index, generating tokenized retrieved additional multi-omics data. The system may further include a user interface configured to receive a user query. The system may further include a query engine configured to select one or more relevant multi-omics data indexes from the indexing unit based on the user query. The query engine may be further configured to rank the selected one or more multi-omics data indexes based on at least one of clinical action likelihood, pathogenicity, feature weight, or frequency. The query engine may be further configured to return the ranked one or more multi-omics data indexes to a user via a user interface.

[0013] According to various embodiments, a multi-omic cancer search engine system is provided for tumor profiling, comprising: a storage element configured to store a plurality of integrated multi-omic indexes; an advanced cancer analysis software module; a multi-omic indexing pipeline; a ranking engine that reflects the clinical utility of multi-omic cancer alterations; a query engine that selects and combines relevant multi-omic indexes and returns ranked multi-omic alterations for individual samples and cohorts of samples; and a user interface configured to receive user queries and perform searches against cancer data.

[0014] Additional aspects will become apparent from the following detailed description, and the claims and drawings appended hereto. [Brief explanation of the drawings]

[0015] The foregoing illustrative examples of various aspects and implementations provide an overview or framework for understanding the nature and features of the claimed aspects and implementations.

[0016] FIG. 1 illustrates an example system architecture for a multi-omics cancer search engine according to various embodiments.

[0017] Figure 2a shows an example of multi-omics index organization according to various embodiments, and Figure 2b shows an example of hierarchical propagation of annotations and ranking of variants according to various embodiments.

[0018] FIG. 3 shows an example of a set of cancer analyses that are dynamically pre-computed and calculated for individual samples and cohorts, according to various embodiments.

[0019] Figure 4a shows an example of a wide and deep model for learning variant rankings, according to various embodiments. Figure 4b shows an example of a learn-to-rank engine that relies on a deep semantic similarity model (DSSM) for biomedical data, according to various embodiments.

[0020] 5a and 5b together illustrate an example workflow for operation of a query engine, according to various embodiments.

[0021] Figure 6 shows an example of a user interface according to various embodiments. For example, as shown, a single search box allows users to enter various queries and receive ranked results.

[0022] FIG. 7 illustrates example search results obtained for a particular syntax, according to various embodiments.

[0023] 8a and 8b show examples of search results obtained for a particular syntax, according to various embodiments.

[0024] FIG. 9 illustrates example search results returned from a user query, according to various embodiments.

[0025] FIG. 10 illustrates example search results returned from a user query, according to various embodiments.

[0026] FIG. 11 illustrates example search results returned from a user query, according to various embodiments.

[0027] FIG. 12 illustrates example search results returned from a user query, according to various embodiments.

[0028] FIG. 13 is a block diagram of a computer system according to various embodiments.

[0029] FIG. 14 shows a flowchart of a method for utilizing multi-omic data indexing for tumor profiling, according to various embodiments.

[0030] FIG. 15 illustrates a system for utilizing multi-omic data indexing for tumor profiling, according to various embodiments.

[0031] FIG. 16 illustrates a system for utilizing multi-omic data indexing for tumor profiling, according to various embodiments.

[0032] It should be understood that the drawings are not necessarily to scale, and that objects in the figures are not necessarily to scale relative to each other. The figures are representations intended to bring clarity and understanding to various embodiments of the devices, systems, and methods disclosed herein. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts. Furthermore, it should be understood that the drawings are not intended to limit the scope of the present teachings in any way. DETAILED DESCRIPTION

[0033] This specification describes various exemplary embodiments of a multi-omics search engine for integrated analysis of cancer genetic and clinical data, and related systems and methods, but the disclosure is not limited to these exemplary embodiments and applications or the manner in which the exemplary embodiments and applications operate or are described herein.

[0034] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments disclosed herein belong. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. References herein to "or" are intended to encompass "and / or" unless expressly stated otherwise.

[0035] The present disclosure describes systems and methods for operating a multi-omics search engine for the integrated analysis of cancer genetic and clinical data, and may be referred to herein by the abbreviation "Cancer Search" or "cancer search."

[0036] Unless otherwise defined, scientific and technical terms used in connection with the present teachings described herein shall have the meanings commonly understood by those skilled in the art. Furthermore, unless the context requires otherwise, the singular includes the plural, and the plural includes the singular. In general, the nomenclature utilized in connection with, and techniques of, cell and tissue culture, molecular biology, and protein and oligo- or polynucleotide chemistry and hybridization described herein are well known and commonly used in the art. Standard techniques are used, for example, for nucleic acid purification and preparation, chemical analysis, recombinant nucleic acid, and oligonucleotide synthesis. Enzymatic reactions and purification techniques are performed according to manufacturer's specifications, or as commonly accomplished in the art, or as described herein. The techniques and procedures described herein are generally performed according to conventional methods known in the art and as described in various general and more specific references cited and discussed throughout the specification. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbour Laboratory Press, Cold Spring Harbour, NY 2000). The nomenclature, and the laboratory procedures and techniques utilized in connection herein are those well known and commonly used in the art.

[0037] As used herein, "DNA" (deoxyribonucleic acid) refers to a chain of nucleotides consisting of four types of nucleotides: A (adenine), T (thymine), C (cytosine), and G (guanine), while its RNA (ribonucleic acid) is composed of four types of nucleotides: A, U (uracil), G, and C. Certain pairs of nucleotides specifically bind to each other in a complementary manner (called complementary base pairs). That is, adenine (A) pairs with thymine (T) (except in RNA, where adenine (A) pairs with uracil (U)), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand consisting of nucleotides complementary to those in the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequence data," "base sequence information," "base sequence," "gene sequence," "gene sequence," or "fragment sequence," or "nucleic acid sequence read" refers to any information or data that indicates the order of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine / uracil) in a DNA or RNA molecule (e.g., whole gene, whole transcriptome, exome, oligonucleotide, polynucleotide, fragment, etc.).

[0038] It should be understood that the present teachings contemplate sequence information obtained using all types of available techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, electronic signature-based systems, and the like. "Polynucleotide," "nucleic acid," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or their analogs) linked by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides usually range in size from a few monomeric units, from 3-4 to several hundred monomeric units. Whenever a polynucleotide, such as an oligonucleotide, is represented by a series of letters, such as "ATGCCTG," it is understood that the nucleotides are in 5'->3' order from left to right, and "A" represents deoxyadenosine. Unless otherwise specified, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents thymidine. The letters A, C, G, and T may be used to refer to the base itself, the nucleoside, or the nucleotide that constitutes the base, as is standard in the art.

[0039] The phrase "next-generation sequencing" (NGS) refers to sequencing technologies with increased throughput compared to traditional Sanger and capillary electrophoresis-based approaches, for example, with the ability to generate hundreds of thousands of relatively small sequence reads. Some examples of next-generation sequencing technologies include, but are not limited to, sequencing-by-synthesis, sequencing-by-ligation, and sequencing-by-hybridization. More specifically, Illumina's MISEQ, HISEQ, and NEXTSEQ systems and Life Technologies Corp.'s Personal Genome Machine (PGM) and SOLiD sequencing systems provide massively parallel sequencing of whole genes or targeted genes. For information regarding the SOLiD system and associated workflows, protocols, chemistries, etc., see PCT Publication No. WO2006 / 084132, filed February 1, 2006, entitled "Reagents, Methods, and Libraries for Bead-Based Sequencing," U.S. Patent Application No. 12 / 873,190, filed August 31, 2010, entitled "Low-Volume Sequencing System and Methods of Use," and U.S. Patent Application No. 12 / 873,132, filed August 31, 2010, entitled "High-Speed ​​Indexing Filter Wheels and Methods of Use," each of which is incorporated herein by reference in its entirety.

[0040] The phrase "sequencing run" refers to any step or portion of a sequencing experiment that is performed to determine some information associated with at least one biomolecule (e.g., a nucleic acid molecule).

[0041] As used herein, the phrase "genetic feature" refers to a genetic region having some annotated function (e.g., gene, protein-coding sequence, mRNA, tRNA, rRNA, repeat sequence, inverted repeat, miRNA, siRNA, etc.), or to a single or group of genes (DNA or RNA) that have undergone changes due to genetic / genetic variation (e.g., single nucleotide polymorphism / mutation, insertion / deletion sequence, copy number variation, inversion, etc.), mutation, recombination / crossover, or genetic drift for a particular species or subpopulation within a particular species.

[0042] As used herein, the term "biomarkers" refers to objectively measurable indicators of a biological state.

[0043] As used herein, the term "pathogenicity" refers to the property of a genetic change that increases an individual's susceptibility or predisposition to a particular disease or disorder. Also referred to as predisposing mutations, deleterious mutations, and disease-causing mutations.

[0044] As used herein, the term "germline" refers to tissue derived from a reproductive cell (egg or sperm) that becomes incorporated into the DNA of all cells in the body of the offspring. Germ cell mutations can be passed from parents to offspring.

[0045] As used herein, the term "somatic" refers to genetic changes acquired by cells during the process of cell division. Somatic mutations are distinct from germline mutations, which are genetic changes that occur in germ cells.

[0046] As used herein, the term "codon" refers to a trinucleotide sequence of DNA or RNA that corresponds to a specific amino acid.

[0047] As used herein, the term "UI (User Interface)" is an acronym for user interface.

[0048] As used herein, the term "query time" refers to the time at which a user submits a query.

[0049] As used herein, the terms "learning-to-rank" or "ranking engine" or "releavance-learning" refer to the application of machine learning, typically supervised, semi-supervised, or reinforcement learning, in building ranking models for information retrieval systems. The training data consists of lists of items with a specified partial order between the items in each list. This order is typically induced by giving each item a numeric or ordinal score or a binary judgment (such as "relevant" or "not relevant"). The goal of a ranking model is to rank, i.e., generate permutations of items in new, unseen lists, in a way that is in some sense "analogous" to the ranking of the training data.

[0050] As used herein, the term "latent space" or "hidden space" refers to the space in which features reside.

[0051] As used herein, the term "embedding" refers to the mapping of documents (e.g., text, images, structured data) into a low-dimensional latent space that preserves key properties of the objects.

[0052] As used herein, the term "deep-and-wide model" refers to a deep learning model that jointly trains a wide linear model (e.g., for memory) along with a deep neural network (e.g., for generalization).

[0053] As used herein, the term "language model" refers to a probability distribution over a sequence of words.

[0054] As used herein, the term "transformer model" refers to a deep learning model with the core idea of ​​self-attention, i.e., the ability to focus attention on different positions in an input sequence and compute a representation of that sequence.

[0055] As used herein, the term "BM25" refers to a broad family of statistical functions in information retrieval that consider the occurrence of each query term, i.e., term frequency (TF), in a document or set of documents and the corresponding inverse documents, and rank the set of documents based on the query terms that appear in each document, regardless of their proximity within the document.

[0056] As used herein, the term "RM3" refers to an information retrieval model useful for both relevance and pseudo-relevance feedback.

[0057] As used herein, the term "Deep Semantic Similarity Model (DSSM)" is an acronym that stands for Deep Semantic Similarity Model.

[0058] As used herein, the term "Siamese network" refers to an artificial neural network that uses the same weights while operating in tandem on two different input vectors to compute comparable output vectors.

[0059] As used herein, the term "FDA" is an acronym for the United States Food and Drug Administration.

[0060] As used herein, the term "NCCN (National Comprehensive Cancer Network)" is an acronym for National Comprehensive Cancer Network.

[0061] As used herein, the term "COSMIC (Catalogue of Somatic Mutations in Cancer)" is an acronym for Catalogue of Somatic Mutations in Cancer.

[0062] As used herein, the term "TCGA (The Cancer Genome Atlas)" is an acronym for The Cancer Gene Atlas.

[0063] As used herein, the term "CPRA (chromosome, position, reference, and alternative)" is an acronym for chromosome, position, reference, and alternative.

[0064] As used herein, the term "SNV (Single Nucleotide Variants)" is an acronym for single nucleotide polymorphism.

[0065] As used herein, the term "copy number variatns (CNV)" is an acronym for copy number variation.

[0066] As used herein, the term "BCL (Binary Base Call)" is an acronym for binary base call.

[0067] As used herein, the term "FAST()" refers to a text-based format for storing both biological sequences (typically nucleotide sequences) and their corresponding quality scores. Both the sequence characters and the quality scores are encoded with a single ASCII character each for simplicity.

[0068] As used herein, the term "BAM" refers to a binary format for storing sequence data.

[0069] As used herein, the term "VCF" is an acronym that stands for Variant Call Format and refers to a text file format used in bioinformatics to store variations in gene sequences.

[0070] As used herein, the term "Electronic Health Records (EHR)" is an acronym that stands for electronic health record.

[0071] As used herein, the term "ASCO (American Society of Clinical Oncology)" is an acronym that stands for the American Society of Clinical Oncology.

[0072] This disclosure describes various embodiments of a multi-omic search engine for integrated analysis of cancer genetic and clinical data, referred to herein for short as "Cancer Search." Cancer Search is an extension of the work presented in U.S. patent application Ser. No. 15 / 465,454, filed March 21, 2017, and entitled "Genomic Metabolic, and Microbiombic Search Engine," the contents of which are incorporated herein by reference in their entirety.

[0073] According to various embodiments, a general search engine architecture is provided that can be configured to adapt to specific needs for cancer multi-omic data. The general architecture may include various components, which are discussed in detail below with reference to FIG. 1. For example, the general architecture may include a web-based user interface, a query engine, an indexing pipeline that uses all annotations to index cancer multi-omic data, a cancer analysis software module, and a ranking engine. The query engine may be configured to respond to requests to search any combination of available multi-omic data streams for individual samples or cohorts. The cancer analysis (e.g., software module or engine) may be configured to pre-compute some features at query time and dynamically calculate others to derive important tumor characteristics. The ranking engine may be configured to pre-load default clinically actionable or pathogenicity-related rankings at indexing time and further enhance its rankings based on the detected query intent at query time. Details related to various data types, pipelines, engines, modules, and analyses are provided below.

[0074] The overall functionality of the user interface (UI) may be configured to provide a unified, responsive way to query and navigate multi-omics cancer search results. The UI may actively maintain state for the user search session. The UI may be configured to accept user queries, relay them to the query engine, render the resulting integrated multi-omics ranked results and their summary visualizations, and allow the user to interact with the search results. Users can interact with the search results in various ways through the UI. For example, they may provide relevance feedback, comment on the accuracy of the information presented by the search results (e.g., a particular annotation source / publication is outdated or inconsistent), promote / demoting / fix / delete type ratings of how well the results address the user's information needs, and mark specific results for inclusion in dynamic individual patient or cohort reports. More details related to the UI are discussed below.

[0075] Figure 1 illustrates a non-limiting example of the general architecture of a multi-omic cancer retrieval system 100. Samples (e.g., tumor and / or normal samples) may be added to the indexing pipeline or indexer 115 from the somatic workflow 120 or uploaded via a user interface 125. Non-limiting examples of upload formats may include FASTQ, BAM, tumor VCF, normal, and somatic. Non-limiting examples of upload formats may include FASTQ, BAM, tumor VCF, normal, somatic VCF, RNA-Seq variant confirmation VCF, tabular RNA-Seq differential gene expression, CNV VCF, structural variant VCF, fusion call VCF, or any combination thereof. Multi-omic data 110 may be cancer multi-omic data including BCL, FASTQ, BAM, VCF, tabular cancer data, text cancer data, and image cancer data. A set of annotation, literature, and phenotypic data 130 may be added to the indexer 115 via the annotation pipeline 135. The data may reside in a storage unit 170 (e.g., cloud storage, internal computer storage) or may be uploaded by a user via a dedicated search upload interface. Data added by the indexing pipeline 115 may be stored in one or more indexes 140. The system architecture may further include a cancer analysis engine or module 145 that can be configured to derive key characteristics of tumors during indexing and serving. The cancer analysis engine 145 can derive the key characteristics regardless of whether the analysis is for individual samples or cohorts. The user interface 125 may allow a user to enter queries and receive results provided by the query engine 150. The query engine 150 may be configured to accept user queries, select, pre-combine, aggregate, and summarize relevant multi-omics indexes, and return ranked multi-omics data or features.According to various embodiments, the system architecture may further include a load balancer 155 to accommodate bidirectional transfer of data between the UI 125 and the query engine 150 for multiple users. According to various embodiments, the system architecture may further include an authentication proxy 160 and may include an identity provider 175 (e.g., a third-party provider). Results retrieved from the indexer 115 may be ranked by a ranking engine 165 (e.g., a ranking learning engine), which may be configured to derive ranking models for, for example, variants, genes, pathways, phenotypes, text data, and images. Results obtained from the index are ranked by the ranking engine and displayed to the user in ranked order. As described in detail herein, the types of data that can be queried, analyzed, and ranked are vast, including genes, transcriptomes, epigenetics, chromatin accessibility data, microbiomes, proteomes, medical literature, phenotype data, text data, imaging data, annotation sources, cancer analysis, predictive models, and features that contribute to model accuracy. More details regarding various method and system implementations related to this example general architecture are provided below.

[0076] 14 , according to various embodiments, a method 1400 for utilizing multi-omics data indexes for tumor profiling is provided. The method can include, at step 1410, storing a plurality of multi-omics data indexes, each of the plurality of multi-omics data indexes including cancer-specific tokenized data. Further discussion related to, for example, features, multi-omics data indexes, and storing cancer-specific data is provided throughout this disclosure and is applicable to this specification and all embodiments discussed or contemplated herein.

[0077] The method may further include importing additional multi-omics data and annotations associated with the additional multi-omics data, the additional multi-omics data associated with one or more indexes, at step 1420. For example, further discussion related to annotation and import functionality is provided throughout this disclosure and is applicable to this and all embodiments discussed or contemplated herein.

[0078] The method may further include, at step 1430, indexing the acquired additional multi-omic data and annotations while preserving gene names, gene variant names, and multi-omic mappings between different data streams for the same patient within a particular index, generating tokenized and ingested additional multi-omic data. For example, further discussion related to indexing, gene names, gene variant names, and multi-omic mappings is provided throughout this disclosure and is applicable to this specification and all embodiments discussed or contemplated herein.

[0079] The method may further include receiving a user query at step 1440. For example, further discussion related to receiving functions and user queries is provided throughout this disclosure and is applicable to this and all implementations discussed or contemplated herein.

[0080] The method may further include selecting one or more relevant multi-omics data indexes based on the user query, at step 1450. Further discussion related to, for example, selection functions, pre-combining multi-omics indexes, and determining relevance is provided throughout this disclosure and is applicable to this specification and all embodiments discussed or contemplated herein.

[0081] The method may further include, in step 1460, ranking the selected one or more multi-omics data indexes based on at least one of clinical actionability, pathogenicity, feature weight, and frequency. Other ranking factors may also be included, such as factors related to query intent. Further discussion related to ranking is provided throughout this disclosure and is applicable to this and all embodiments discussed or contemplated herein.

[0082] The method may further include returning the ranked multi-omics data index or indexes to the user in step 1470. Further discussion related to, for example, return functionality, displays, and reports is provided throughout this disclosure and is applicable to this specification and all embodiments discussed or contemplated herein.

[0083] According to various embodiments, a non-transitory computer-readable medium is stored with a program for causing a computer to execute a method for utilizing multi-omics data indexing for tumor profiling, the steps of which may be similar to those described above or may be modified as desired.

[0084] The method may include storing a plurality of multi-omics data indexes, each of the plurality of multi-omics data indexes comprising cancer-specific tokenized data. Further discussion related to, for example, storing features, multi-omics data indexes, and cancer-specific data is provided throughout this disclosure and is applicable to this specification and all embodiments discussed or contemplated herein.

[0085] The method may further include incorporating additional multi-omic data and annotations associated with the additional multi-omic data, the additional multi-omic data associated with one or more indexes. For example, further discussion related to annotation and incorporating functions is provided throughout this disclosure and is applicable to this and all embodiments discussed or contemplated herein.

[0086] The method may further include indexing the acquired additional multi-omics data and annotations to generate tokenized acquired additional multi-omics data, while preserving gene names, gene variant names, and multi-omic mappings between different data streams for the same patient within a particular index. For example, further discussion related to indexing, gene names, gene variant names, and multi-omics mappings is provided throughout this disclosure and is applicable to this specification and all embodiments discussed or contemplated herein.

[0087] The method may further include receiving a user query. For example, further discussion related to functionality and receiving a user query is provided throughout this disclosure and is applicable to this and all embodiments discussed or contemplated herein.

[0088] The method may further include selecting one or more relevant multi-omics data indexes based on the user query. Further discussion related to, for example, selection functions and determining relevance is provided throughout this disclosure and is applicable to this and all embodiments discussed or contemplated herein.

[0089] The method may further include ranking the selected one or more multi-omics data indexes based on at least one of clinical actionability, pathogenicity, feature weight, or frequency. Note that the ranking can be further modified depending on the query objective (e.g., ranking in reverse order of frequency, ranking in order of feature contribution to a particular model prediction, ranking in reverse order of mutation signature contribution weight, etc.). Therefore, if no other ranking is requested and no other intent is readily inferred (or can be inferred), clinical actionability serves as the default ranking. For example, further discussion related to ranking functions and determinations is provided throughout this disclosure and is applicable to all embodiments discussed or contemplated herein.

[0090] The method may further include returning the ranked multi-omics data index or indexes to the user. For example, further discussion related to return functionality is provided throughout this disclosure and is applicable to this and all embodiments discussed or contemplated herein.

[0091] According to various embodiments, the multi-omics data may be selected from the group consisting of genetic, transcriptomic, epigenetic, chromatin accessibility data, microbiomics, proteomics, phenotypic, imaging, related literature, integrated multi-omics data, and combinations thereof. According to various embodiments, the plurality of multi-omics data indexes may further include tumor (somatic) genetic alterations, normal (germline) genetic alterations, and cancer annotation sources.

[0092] According to various embodiments, the methods discussed or contemplated herein may further include deriving a cancer analysis for one or more selected multi-omics data indexes. The cancer analysis may include tumor features selected from the group consisting of quality control, tumor mutation burden, gene mutation signatures, microsatellite instability status, neoantigens and their binding affinities, HLA allele typing, RNA confirmation variants, copy number variants, structural variants, non-coding regulatory variants, gene fusions, pathway enrichment, cancer driver identification, mutation summary, differential gene expression, immune signatures, and combinations thereof. According to various embodiments, the cancer analysis may be derived for individual samples or cohorts of samples. Additionally, the cancer analysis may include matching information regarding treatment outcomes of similar patients. According to various embodiments, the cancer analysis may include machine learning predictions and ranked features. According to various embodiments, the cancer analysis may include machine learning predictions and machine learning model features ranked in order of relevance to a particular prediction. The machine learning predictions may be selected from primary site classifiers, prediction of future metastatic site classifiers, prediction of microsatellite instability status, prediction of neo-antigen binding affinity, disease state stratification, cancer lineage determination, and combinations thereof. The cancer analysis may be dynamically calculated after receiving a user query. Deriving the cancer analysis may include utilizing deep neural networks or other machine learning methods (e.g., support vector classifiers, tree methods, ensemble methods). Deriving model feature importance may include gradient imputation or other feature importance methods.

[0093] According to various embodiments, the methods discussed or contemplated herein may further include propagation of annotations from higher levels of the gene hierarchy to lower levels of the gene hierarchy.

[0094] According to various embodiments, the methods discussed or contemplated herein may further include propagating the ranking of one or more selected multi-omics data indexes from higher levels of the genetic hierarchy to lower levels of the genetic hierarchy. The ranking may include clinical rankings of cancer variants and genes. The ranking may include the probability of enrichment of genes belonging to specific pathways. The ranking may include importance weights determined for features of machine learning models. The ranking may stratify the cohort by incorporating latent space representations of cancer data and subselecting representations to maximize disentanglement between responders and non-responders, short- and long-term progression-free survival, one and another cancer subtype, etc. The cohort may be stratified into responders and non-responders. The cohort can be stratified by long-progression-free survival time and short-progression-free survival time. The cohort can be stratified into various cancer subtypes. The latent space representation can be performed by neural networks or other dimensionality reduction methods (principal component analysis, individual component analysis, manifold learning, etc.). The neural networks may be selected from the group consisting of autoencoders, variational autoencoders, deep belief networks, restricted Boltzmann machines, feedforward, convolutional, recurrent, gated regression, long short-term memory, residual, and generative adversarial networks.

[0095] According to various embodiments, including methods discussed or contemplated herein, the ranking may further include a model for learning the ranking selected from the group consisting of a support vector machine, a boosted decision tree, a regression method, a neural network, and combinations thereof. The model for learning the ranking may include other machine learning models and deep neural networks. The ranking may further include a deep learning ranking. The ranking may further include similarities between query embeddings and indexed documents in a joint embedding space learned via deep learning techniques. The deep learning ranking may be derived from a deep learning model selected from the group consisting of a deep semantic similarity model, a deep and wide model, a deep language model, a learned deep learning text embedding, a learned named entity recognition, a Siamese neural network, and combinations thereof.

[0096] According to various embodiments, including methods discussed or contemplated herein, the multi-omics data may be selected from the group consisting of somatic (and germline) calling from whole-gene sequence data, somatic (and germline) calling from whole-exome sequence data, somatic (and germline) panel sequencing from fresh-frozen tissue, somatic (and germline) panel sequencing from formalin-fixed, paraffin-embedded tissue, somatic (and germline) panel sequencing from liquid biopsies, tumor and normal variant calls, data indexed as confirmed variants at the tumor / normal transcript RNA or gene expression level, epigenetic data, chromatin accessibility data, microbiome data, proteomics data, single-cell sequencing data, and combinations thereof. In various embodiments, the indexed multi-omics data may come from internal somatic calling and 16mmune pipelines, or may be provided or uploaded in real time from any external partner in FASTQ, BAM, VCF, and other tabular formats.

[0097] According to various embodiments, including methods discussed or contemplated herein, the multi-omics data index may further include extracted phenotypic data, which may be selected from the group consisting of electronic health records, clinical data, functional data, and combinations thereof.

[0098] According to various embodiments, including methods discussed or contemplated herein, the multi-omics data index may further include characterized / embedded imaging data. The characterized image data may be selected from the group consisting of histology slides, MRI images, x-rays, mammograms, ultrasounds, PET images, CT scans, and combinations thereof.

[0099] According to various embodiments, including methods discussed or contemplated herein, indexing the acquired additional multi-omics data and annotations may further include indexing derived data selected from the group consisting of cancer analysis, annotations, features extracted from image data, phenotypes, medical literature data, data embedding, and combinations thereof.

[0100] According to various embodiments, including methods discussed or contemplated herein, ranking may further include matching sample modifications with established drug target markers and available clinical trials. Ranking may stratify the cohort based on clinical variables of interest and / or statistical significance by detecting potential biomarkers, and returning one or more ranked multi-omic data indexes to the user may further include target identification of anti-cancer drugs in the cohort, including visualization of the stratification.

[0101] According to various embodiments, including methods discussed or contemplated herein, returning the ranked multi-omics data indexes to the user may further include dynamic creation of hyperlinked reports (e.g., including ranked changes where each entry is hyperlinked to a search query) for individual patients and / or cohorts, providing comprehensive profiling of the tumor or cancer. Returning the ranked multi-omics data indexes to the user may further include returning a summary visualization of the returned results along with the list of ranked results.

[0102] According to various embodiments, including methods discussed or contemplated herein, the user query may include user-uploaded data selected from the group consisting of a panel of variants, genes, pathways, disease states, and phenotypes of interest, where selecting includes querying individual sample or cohort data subselected by the uploaded data. The user query may be provided via a user interface and may include uploading data for indexing selected from the group consisting of genetic data, transcriptomic data, epigenetic data, chromatin accessibility data, microbiological data, proteomic data, phenotypic data, annotation data, and combinations thereof.

[0103] According to various embodiments, methods discussed herein or contemplated herein may include normalizing and / or expanding user queries, classifying the intent of the query, summarizing retrieved documents, and performing document retrieval based on the similarity between the query and documents in a latent space using deep learning methods.

[0104] According to various embodiments, including methods discussed or contemplated herein, at least one of indexing, selecting, and ranking includes utilizing a deep neural network.

[0105] According to various embodiments, the methods (and systems) discussed or contemplated herein may function to centralize vast amounts of cancer multi-omic data to provide a platform for oncologists, practitioners, research scientists, and other non-programmers to interrogate the cancer bioinformatics pipeline at a detailed level and gain clinical and biological insights into cancer biology and potential clinical treatments for cancer. Data types may include, for example, genetic (single nucleotide variations, tumor and normal indels, structural rearrangements, copy number variations, gene fusions, and tumor gene expression variations), transcriptional, epigenetic, chromatin accessibility, microbial, proteomic abundance and localization, medical literature data (publications, treatment guidelines, clinical trial inclusion / exclusion criteria), phenotypic data (functional, clinical, electronic medical records, histopathology and radiology reports), imaging data (histopathology slides, MRI scans, X-rays, mammograms, ultrasound, PET images, CT scans), cancer annotation sources (variants, genes, pathways, drugs), derived cancer analyses (tumor mutation burden, mutational signatures, microsatellite instability status, RNA-seq confirmed variants, differentially expressed genes, spatial omics lineage representation, neo-antigen binding affinity, MHC class I and class II molecules), etc.

[0106] As mentioned above and discussed in further detail below, various methods (and systems) described and contemplated herein include cancer analysis (e.g., as a step, function, engine, module, or software module) according to various embodiments. Cancer analysis provides users with access to important tumor characteristics, including, for example, tumor mutation burden, mutational signatures, spatial omics lineage representation, neo-antigen binding affinity to MHC class I and class II molecules, RNA sequence-confirmed mutations, differentially expressed genes, pathway enrichment, microsatellite instability status and microsatellite repeat loci, and feature imaging and clinical data extracted from the following: According to various embodiments, this data can be pre-computed for individual samples or dynamically calculated for cohort samples. According to various embodiments, cancer analysis can provide predictions from machine learning models and integration of those features ranked by their contribution to a specific classification. Specific classifications include, for example, prediction of primary site, future metastatic site, classification of variants as true or false positive, information on treatment outcomes of similar patients, anomaly detection of sequence quality, disease state prediction and actual representation using latent cohorts, etc. The advantage of returning features ranked by their contribution to a particular classification is that the model's predictions become more explainable to the user.

[0107] As mentioned above and discussed in further detail below, various methods (and systems) described and contemplated herein include multimodal ranking (e.g., as a step, function, engine, module, or software) according to various embodiments. Multimodal ranking provides an association learning engine to integrate multi-omics genetic data, annotation sources, literature data, clinical trial results, and significantly mutated genes into well-characterized cohorts to learn clinically actionable rankings of cancer data. In various embodiments, machine learning models may be used to weigh contributions from annotations of the multi-omics data. In various embodiments, deep learning and machine learning dimensionality reduction techniques may be used to derive latent space representations of cohorts of samples. In various embodiments, the learned embeddings may be used to rank genes, text, and image data.

[0108] As noted above and discussed in more detail below, various methods (and systems) and various embodiments described and contemplated herein may further include a mechanism (e.g., as a step, function, engine, module, or software module) for integrating and ranking multiple cancer annotation sources, including, for example, FDA labels, NCCN guidelines, clinical trials, CIViC, DocM, OncoKB, Mycancergenome, databases of genetic biomarkers for cancer therapeutics, TCGA, ICGC, COSMIC, NCI60, CCLE, Drugbank, ClinVar, HGMD, PGMD, PharmGKB, dbSNP, dbNSFP, 1000Genomes, EXAC, CPDB, KEGG, BioCarta, BioCyc, Reactome, GenMAPP, MsigDB, Brenda, CTD, HPRD, GXD, and BIND. In various embodiments, annotations and rankings can be propagated from higher levels of representation to lower levels (e.g., pathways from gene to variant, or pathways from gene to variant codon to full variant specification - chromosome, location, reference, alternative).

[0109] As discussed above and in further detail below, various methods (and systems) and various embodiments described and contemplated herein further include mechanisms (e.g., as steps, functions, engines, modules, or software modules) for integrating multiple deep learning models. The integration can function to provide neural data indexing (e.g., embedding multi-omics datasets individually or together, normalizing their respective latent spaces for DNA and RNA tumor alterations, embedding text data from electronic health records, clinical notes, literature, and annotations, embedding deep transformer models for named entity extraction and summarization, and embedding text and annotation data, and image data). The integration can also provide neural learning to rank models (e.g., deep semantic similarity models, convolutional deep semantic similarity models, iterative deep semantic similarity models, deep association matching models, interaction Siamese networks, lexical and semantic matching networks, DeepRank, etc.), which are used to address the feature engineering problem of learning to rank. Integration can provide neural query models (e.g., deep learning Transformer models for query normalization, synonym expansion, abbreviation expansion, term disambiguation, and alternative suggestion), and can serve to provide neural models for advanced cancer analysis (e.g., classification of site of origin, prediction of future metastatic sites, prediction of neoantigen binding affinity, classification of variants as true or false positives, drug and test matching, recommendation systems for treatments using information from indexed similar cases, loss, gain, and maintenance of allelic fractions, copy number variation, RNA expression at each location on serial biopsies, and models comparing deep learning autoencoder methods and other dimensionality reduction techniques for cohort analysis and stratification).

[0110] As noted above and explained in further detail below, the various methods (and systems) described and contemplated herein, and according to various embodiments, may further include (e.g., as steps, functions, engines, modules, or software modules) statistical, machine learning, and deep learning methods for identifying diagnostic, prognostic, or predictive biomarkers. When a user (e.g., an academic or industry researcher) enters a phenotypic query on a cohort of samples, various embodiments return ranked biomarkers that can stratify the cohort, their statistical significance, and a summary visualization of them. In various embodiments, validation queries may be suggested by a search engine to perform robust algorithmic and statistical validation. In various embodiments, the systems and methods can automatically suggest iterative hypothesis refinements through suggested query refinement, and according to various embodiments, statistical visualizations and analyses derived for cancer cohort queries include, for example, Kaplan-Meier survival analysis visualizations, log-rank test result visualizations, Cox proportional hazards regression analysis visualizations, tree-structured survival model visualizations, heat maps, scatter plots, box plots, and bar graphs providing statistical significance.

[0111] As noted above and explained in more detail below, various methods (and systems) described and contemplated herein, and according to various embodiments, may further include (e.g., as a step, function, engine, module, or software module) the interactive use and / or receipt of summary visualization and / or ranked variants, genes, pathways, derived cancer analysis, and output of the integrated machine learning model (e.g., cancer type classification, most likely site of recurrence). This may be provided via a query engine (described in more detail below). In various embodiments, the summary visualization may be dynamic, and every data point may be linked to a specific result returned.

[0112] As noted above and explained in more detail below, the various methods (and systems) and various embodiments described and contemplated herein can further provide interactive, fast access within 10,000, 5,000, 4,000, 3,000, 2,000, 1,000, 900, 800, 700, 500, 400, 300, 200, 100 milliseconds or less of access, or any range between the above values, to multi-omic cancer data ranked by clinical actionability, pathogenicity, feature weight, or frequency.

[0113] As noted above, the systems and methods described herein, and various embodiments thereof, can provide a universal search interface (as opposed to many different entry points). In various embodiments, all knowledge, e.g., multi-omics cancer data, samples, variants, genes, drugs, pathways, phenotypes, medical literature, imaging data, derived cancer analyses, tumor features and machine learning models to predict those features, uploads, and user data can be accessed from the same simple search interface.

[0114] As described above and in further detail below, the various methods (and systems) described and contemplated herein, and in various embodiments, can further (e.g., as a step, function, engine, module, or software module) provide the ability to compare serial biopsy samples, differences (increases, decreases, maintenance) between old and new cancer drivers, changes in mutant allele proportions, copy number changes, and RNA confirmation status changes of cancer alterations.

[0115] As noted above and discussed in further detail below, according to various embodiments, the various methods (and systems) described and contemplated herein can further provide (e.g., as steps, functions, engines, modules, or software modules) for a variety of comparison regimes. These regimes include, for example, (1) sample-to-sample comparison, any combination of multi-omics streams of data within the same patient, (2) sample-to-cohort comparison (e.g., comparing individual samples with TCGA subtypes of the same cancer), and (3) pairwise cohort comparison (e.g., comparing a cohort with a well-characterized TCGA cohort of the same cancer type).

[0116] According to various embodiments, the various methods (and systems) described and contemplated herein can be provided (e.g., as a step, function, engine, module, or software module) for dynamic uploading of a variant / gene drug discovery target panel (or a panel currently in use) from a user institution, with subsequent queries indicating use of the intersection of the uploaded panel and the stored multi-omics data for the sample.

[0117] As public domain data and as already discussed herein, a general genetic search is proposed to address the problem of immediate access to germline genetic data. This represents a significantly different problem from germline genetic profiling focused on Mendelian rare variants, GWAS hits, common disease burden testing and polygenic risk, and inherited risk. To effectively address all three major issues in comprehensive cancer characterization discussed above and herein, the systems and methods described herein, according to various embodiments provided and contemplated herein, include advanced cancer analysis of individual samples and cohorts, as well as a ranking engine (described in detail above and herein). The systems and methods described herein, according to various embodiments provided herein, augment all parts of existing general germline search systems to integrate multi-omics data during indexing and presentation, rank cancer alterations by their clinical relevance and pathogenicity, and apply the search engine paradigm to comprehensive cancer profiling of individual samples and cohorts. Additionally, the systems and methods described herein, according to various embodiments provided herein, may include cancer cohort stratification analysis built on top of a cancer search engine, which was completely missing from previous studies.

[0118] According to various embodiments, FIG. 15 illustrates a system 1500 provided for utilizing a multi-omics data index for tumor profiling. The system 1500 includes an indexing unit 1510. The indexing unit includes a storage element 1520 configured to store multiple multi-omics data indexes, each of the multiple multi-omics data indexes including cancer-specific tokenized data. The indexing unit 1510 may further include an indexing engine 1530. The indexing unit 1510 may be configured to ingest the additional multi-omics data and annotations associated with the additional multi-omics data via a data source 1540, the additional multi-omics data associated with one or more indexes. The indexing unit 1510 may be further configured to index the additional multi-omics data and annotations ingested from the data source 1540 while preserving gene names, gene variant names, and multi-omics mappings between different data streams for the same patient within a particular index, providing the tokenized ingested additional multi-omics data.

[0119] The system 1500 may further comprise a user interface 1550 configured to receive a user query 1560 .

[0120] The system 1500 may further include a query engine 1570 configured to select one or more relevant multi-omics data indexes from the indexing unit 1510 based on the user query 1560 .

[0121] The system 1500 may further comprise a ranking engine 1580 configured to receive the selected one or more relevant multi-omics data indexes (e.g., from the query engine 1570), rank the selected one or more multi-omics data indexes, and return the ranked one or more multi-omics data indexes to the user via the user interface 1550.

[0122] According to various embodiments, FIG. 16 illustrates a system 1600 provided for utilizing a multi-omics data index for tumor profiling. The system 1600 includes an indexing unit 1610. The indexing unit can include a storage element 1620 configured to store multiple multi-omics data indexes, each of the multiple multi-omics data indexes including cancer-specific tokenized data. The indexing unit 1610 further includes an indexing engine 1630. The indexing unit 1610 may be configured to ingest the additional multi-omics data and annotations associated with the additional multi-omics data via a data source 1640, the additional multi-omics data associated with one or more indexes. The indexing unit 1610 may be further configured to index the additional multi-omics data and annotations ingested from the data source 1640 while preserving gene names, gene variant names, and multi-omics mappings between different data streams for the same patient within a particular index, generating the tokenized ingested additional multi-omics data.

[0123] The system 1600 further comprises a user interface 1650 configured to receive a user query 1660 .

[0124] System 1600 further comprises a query engine 1670 configured to select one or more relevant multi-omics data indexes from indexing unit 1610 based on user query 1660. Query engine 1670 may be further configured to rank the selected one or more multi-omics data indexes based on clinical actionability, pathogenicity, feature weight, or frequency. Query engine 1670 may be further configured to return the ranked one or more multi-omics data indexes to the user via user interface 1650.

[0125] It should be noted that all of the preceding discussions regarding additional features according to various implementations, particularly with respect to the aforementioned methods and non-transitory computer-readable media, are applicable to features of the various system embodiments described and contemplated herein.

[0126] According to various embodiments, a computer-implemented system for utilizing multi-omic data indexes for tumor profiling is provided. The system includes computer storage, a digital processing device including at least one processor, an operating system configured to execute executable instructions, memory, and a computer program including instructions executable by the digital processing device to create a multi-omic cancer search engine application. The multi-omic cancer search engine application includes a plurality of integrated multi-omic indexes stored in the computer storage and a software module for providing advanced cancer analysis. The multi-omic cancer search engine application includes a software module providing a multi-omic indexing pipeline that ingests multi-omic cancer data, annotations, medical, and clinical data associated with the multi-omic gene and imaging data, tokenizes the data while preserving variant nomenclature, gene names, and drug names, and updates the index with the tokenized data. The multi-omic cancer search engine application further includes a software module responsible for ranking the integrated multi-omic data to reflect the clinical utility of cancer alterations. The multi-omics cancer search engine application may include a query engine that selects and combines relevant multi-omics indexes and returns ranked multi-omics changes for individual samples and cohorts of samples. The multi-omics cancer search engine application includes a software module that presents a user interface that allows a user to enter a user query and perform a faceted search against the multi-omics data.

[0127] According to various embodiments, a non-transitory computer-readable storage medium is provided that is encoded with a computer program including instructions executable by a processor to create a multi-omics cancer search engine application. The multi-omics cancer search engine application includes a plurality of integrated multi-omics indexes stored in computer storage and a software module that provides advanced cancer analysis. The multi-omics cancer search engine application includes a software module that provides a multi-omics indexing pipeline that ingests medical and clinical data related to multi-omics cancer data, annotations, multi-omics genetic and imaging data, and can update the index with gene and drug names and tokenized data while maintaining variant nomenclature. The multi-omics cancer search engine application further includes a software module responsible for ranking the integrated multi-omics data to reflect clinical utility, pathogenicity, frequency, and feature weights of returned results. The multi-omics cancer search engine application includes a query engine that selects and combines relevant multi-omics indexes and returns ranked multi-omics changes for individual samples and cohorts of samples. The multi-omic cancer search engine application includes a software module that presents a user interface that allows a user to enter a user query and perform a faceted search against the multi-omic data.

[0128] According to various embodiments, a computer-implemented method for providing a multi-omics cancer search engine application is provided. The multi-omics cancer search engine application includes a plurality of integrated multi-omics indexes stored in computer storage and a software module for providing advanced cancer analytics. The multi-omics cancer search engine application includes a software module for providing a multi-omics indexing pipeline that ingests medical and clinical data related to multi-omics cancer data, annotations, multi-omics genetic and imaging data, tokenizes the data while preserving variants, and updates the index with nomenclature, gene names, drug names, and tokenized data. The multi-omics cancer search engine application includes a software module responsible for ranking the integrated multi-omics data to reflect the clinical utility of cancer alterations, pathogenicity, frequency, and feature weights of returned results. The multi-omics cancer search engine application includes a query engine that selects and combines relevant multi-omics indexes and returns ranked multi-omics alterations for individual samples and cohorts of samples. The multi-omics cancer search engine application includes a software module that presents a user interface that allows a user to enter a user query and perform faceted searches on the multi-omics data. In various embodiments, the index is optimally formatted in a partially pre-bound configuration, and clinical rankings are pre-loaded to increase search speed and reduce lag time between search and results. In various embodiments, pre-binding of the multi-omics index occurs before the user enters a query.

[0129] It should be noted that all of the foregoing discussion and contemplated herein regarding additional functionality according to various embodiments, particularly with respect to the computer-implemented methods, computer-implemented systems, and non-transitory computer-readable media described above, is applicable to features of the various system embodiments described.

[0130] As mentioned above, in accordance with various embodiments described herein, the systems and methods are capable of unifying vast amounts of cancer multi-omic data. To date, this includes, for example, genetic (e.g., single nucleotide polymorphisms, tumor and normal indels, structural rearrangements, copy number variation, gene fusions, and tumor gene expression variants), transcriptomics (e.g., RNA-Seq variant confirmation and differential gene expression), epigenetic, chromatin accessibility, microbial, and proteomic abundance and localization, medical literature data (e.g., publications, treatment guidelines, clinical trial inclusion / exclusion criteria), phenotypic data (e.g., functional, clinical, EHR), imaging data (e.g., histology, MRI, X-ray, mammogram, ultrasound, PET images, CT scan), cancer annotation sources (e.g., variants, genes, pathways, drugs), derived cancer analyses (e.g., tumor mutation burden, mutational signatures, microsatellite instability status, spatial omics lineage representation, neoantigen binding affinity to MHC class I and class II molecules), and predictions from machine learning models and their features (e.g., primary site, microsatellite instability, likely sites of future metastasis, drug and study concordance). According to various embodiments, genetic data may be in the form of whole exome, whole gene, gene panel data, SNP arrays, etc. According to various embodiments, serial biopsy multi-omics data may be indexed for purposes of monitoring disease progression, development of drug resistance, and monitoring recurrence.

[0131] According to various embodiments, the indexed data can be in formats such as, but not limited to, variant call format (VCF), BAM, and FASTQ for both tumor and normal, or tumor only. According to various embodiments, the phenotypic data can be provided in tabular or raw formats (e.g., EHR, clinical notes, pdf reports).

[0132] As noted above, in accordance with various embodiments described herein, the systems and methods can include annotation sources, examples of which include, but are not limited to, FDA labels, NCCN guidelines, clinical trials, CIViC, DocM, OncoKB, Mycancergenome, databases of genetic biomarkers for cancer therapeutics, TCGA, ICGC, COSMIC, NCI60, CCLE, Drugbank, ClinVar, HGMD, PGMD, PharmGKB, dbSNP, dbNSFP, 1000Genomes, EXAC, CPDB, CADD, PolyPhen, dbNSFP, and many others.

[0133] According to various embodiments, the systems and methods described herein can further include drug target information that can be derived and integrated from multiple sources, including, for example, FDA labels, the NCCN Compendium of Drugs and Biologics, Thomson Micromedex DrugDex, the Elsevier Gold Standard Compendium of Clinical Pharmacology, the American Hospital Formulary Serving-Drug Information Compendium, ESMO guidelines, ASCO guidelines, and NCCN guidelines, as well as mutations annotated in other cancer knowledge databases, such as OncoKB, CIViC, DocM, and COSMIC. According to various embodiments, drug targets can be indexed at the variant, gene, and pathway level. According to various embodiments, drug indications, evidence, cancer type, reported side effects, and additional information can be stored in the search index.

[0134] As noted above, according to various embodiments described herein, the systems and methods include, or use of, a cancer analysis (or advanced cancer analysis) or a software module that provides the advanced cancer analysis. The software module provides both pre-computed (e.g., calculated at indexing time) and dynamic (e.g., calculated at query time) derived cancer analysis. According to various embodiments, the advanced analysis can also be visualized at query time. Figure 3 shows examples of pre-computed and dynamically calculated cancer analysis for individual samples and cohorts. The advanced analysis module integrates predictions from machine learning and deep learning models to predict key characteristics of tumor biology.

[0135] According to various embodiments, pre-calculated derived cancer analytics for individual samples may include, for example, but are not limited to, tumor mutational burden (an important biomarker for treatments such as immunotherapy), microsatellite instability status (an important cancer condition in which mismatch repair proteins are disabled), genetic mutation signatures (the potential etiologic and mechanistic basis of cancer), detected neoORFs (frameshift mutations that may lead to novel amino acid sequences that may be useful in cancer vaccines), detected neo-antigens, neo-antigen binding affinities for MHC class I and class II molecules, HLA allele typing (an important variable in cancer vaccine design), expressed immune genes (e.g., genes that play a role in response to immunotherapy treatment), RNA-seq validated variants, and differentially expressed genes.

[0136] According to various embodiments, dynamic advanced cancer analysis of individual samples includes, but is not limited to, pathway enrichment analysis of specific types of variants (e.g., non-silent variants based on a query), and spatial omics phylogenetic representation. According to various embodiments, dynamic advanced cancer analysis of cohorts of samples includes, but is not limited to, mutational signatures of the cohort, detection of significantly mutated genes and cancer drivers after collapsing recurrent somatic alterations within the same gene, correcting for the ratio of non-silent to silent variants, gene replication time, and other characteristics of cancer biology, disease state stratification, spatial omics phylogenetic representation, pathway enrichment analysis of subsets of variants (e.g., non-silent variants).

[0137] According to various embodiments, cancer analytics can include, for example, machine learning and deep learning models for predicting key features of tumor biology (e.g., tumor-only and tumor-normal classifiers for microsatellite instability, tumor-origin classification for metastatic tumors of unknown cause, deep learning and machine learning methods for patient-specific, tumor-only variant calling, neo-antigen binding prediction, machine learning models for inherited cancer risk prediction for various cancer types, machine learning models for immunotherapy outcome prediction, deep variant, gene, drug, and disease learning methods for classifying variants as true positive or false positive, processing literature, EHR, and clinical trial data, and Deep learning methods for identifying regions of interest and extracting features from unstructured histology and radiology slides and other image data; deep learning models for learning latent entities; multi-omics cancer pathogenesis; deep learning methods for drug and test matching; machine learning models for identifying similar patients; recommendation systems for cancer treatment based on outcomes from the treatment of similar patients; and machine learning and deep learning methods for cohort biomarker stratification and cohort disease state identification.

[0138] According to various embodiments described herein, the systems and methods can include, for example, deep learning embeddings of phenotypic data (e.g., learned from electronic health records, clinical and functional records), annotation sources, medical literature, or imaging data (e.g., histology slides, MRIs, X-rays, mammograms, ultrasounds, PET images, CT scans, etc.).

[0139] According to various embodiments described herein, the systems and methods include an advanced cancer analysis module that sets statistical thresholds for quality control and identifies outliers in indexed sequencing quality metrics. Some non-limiting examples of quality control metrics of interest may include tumor-normal match quality control (e.g., kinship and identity values), tumor and normal sequence metrics, such as the Freemix / Compair metric that reflects potential tumor / normal contamination, sequence metrics including average total coverage, percentage of reads adjusted, duplication rate, and Y / X ratio. Somatic sequencing quality control metrics include, but are not limited to, dbSNP variant count, dbSNP enrichment, dbSNP insertion / deletion rate, dbSNP transition / transversion ratio, and heterogeneous / homogeneous variant ratio (heterozygous / homozygous variant ratio).

[0140] In various embodiments, advanced cancer analysis (or related modules) can be provided, such as dynamic algorithms for mutation summarization, cancer driver identification, multiple biopsy comparison, and cohort stratification based on suspected (multi-omic) biomarkers in a cohort of samples. In various embodiments, sample-to-sample cohort comparisons, as well as comparisons of multiple cohorts, can be performed.

[0141] According to various embodiments described herein, the systems and methods involve indexing and centralizing vast amounts of cancer multi-omics data. As discussed in some detail above, data may include, for example, but are not limited to, genetic data (e.g., single nucleotide mutations, indels in tumor and normal, structural rearrangements, copy number variations, gene fusions, and expression variants of tumor genes), transcriptomic data, epigenetic data, chromatin accessibility data, microbiome data, proteomic abundance and localization data, medical literature data (e.g., publications, treatment guidelines, clinical trial inclusion / exclusion criteria), phenotypic data (e.g., functional, clinical, EHR), imaging data (e.g., histology slides, MRI, X-rays, mammograms, ultrasound, PET images, CT scans), cancer annotation sources (e.g., variants, genes, pathways, drugs), derived cancer analyses (e.g., tumor mutation burden, mutational signatures, differentially expressed genes, spatial omics lineage representation, primary site of origin, future metastatic sites, predictions and features from machine learning models of microarray instability status, neoantigen binding affinity for MHC class I and class II molecules).

[0142] Applicant has advantageously found that by indexing raw data along with derived analyses, predictions from machine learning and deep learning models and their (derived) features and embeddings may include improved interpretability of machine learning, iterative hypothesis generation, and refinement of continuous queries by users to better understand tumor biology.

[0143] According to various embodiments, and as described above, the systems and methods disclosed herein may include a software module for multi-omics indexing of cancer data, annotation, medical and clinical data related to genetic and imaging data, tokenizing the data as it is stored, and updating the index with variant nomenclature, gene names, drug names, and tokenized data. According to various embodiments, the multi-omics indexing step includes integrating and pre-combining multi-omics indexes at the variant, gene, pathway, cancer subtype, or sample level.

[0144] Specific to cancer annotation data, according to various embodiments, the systems and methods described herein include a software module that provides an indexing step (see above) or multi-omic indexing of cancer annotation data. Cancer annotation data includes, but is not limited to, FDA labels and NCCN guidelines, clinical trials, public cancer databases (CIViC, DocM, OncoKB, Mycancergenome, COSMIC, Database of Genetic Biomarkers for Cancer Therapeutics, ICGC, TCGA), public genetic databases (ClinVar, dbNSFP, dbSNP), and commercial data sources (HGMD, PGMD, PharmGKB, CPDB). In another aspect, the multi-omic indexing software module also indexes annotation sources that are not cancer-focused (ClinVar, dbNSFP, dbSNP, CPDB, HGMD, PGMD). According to various embodiments, the software module for multi-omic indexing is configured to integrate and pre-combine multi-omic annotation data at the level of variants, gene codon numbers, genes, pathways, cancer subtypes, or samples.

[0145] According to various embodiments, indexing further includes utilizing derived content embedding to index complex phenotypes, literature data, histopathology, MRI, X-ray, mammogram, ultrasound, PET images, and CT scan images.

[0146] The systems and methods described herein, according to various embodiments, further include indexing procedures in which multi-omics data integration during indexing is performed first at the sample level, and then at either the variant, gene codon number, gene, or pathway level, or any combination thereof, as shown in Figures 2a and 2b.

[0147] In a non-limiting example of multi-omics indexing integration shown in Figure 2a, the acquired multi-omics cancer data are selected from single nucleotide polymorphisms (SNVs) and small indels (represented as chromosome number, chromosomal location, reference, and alternative alleles—CPRA), copy number variations (CNVs), and RNA-confirmed variants. SNVs can be indexed from SNVs and small indels, including somatic VCFs. Copy number variations (CNVs) called in chromosomal regions (e.g., mapped at the gene level using the Advanced Cancer Analysis module) can be indexed from copy number-called VCFs (CNVs are also mapped at the gene level). RNA-Seq-confirmed variants can be obtained from RNA-Seq analysis (derived from the Advanced Cancer Analysis module). Multi-omics indexes can be combined to answer complex queries (e.g., to obtain SNVs and small indels that overlap CNV gains and losses, represented in the RNA of a group of samples). Differentially expressed genes can be derived, for example, from advanced analysis software modules. According to various embodiments, a combined multi-omics index can be generated via selected indexing methods, such as, but not limited to, KEYSxCPRA, KEYSxCNV, KEYSxCNV_RANGE, KEYSxCNV_GENE, KEYSxCPRA_RNA, and KEYSxGENE_RNA, to create indexes of copy number variants and confirmed RNA variants (see Figure 2a). Applicant also anticipates that cross-indexing of multiple streams of information can be performed, for example, querying any combination of multi-omics streams of data or the individual streams themselves, including at the variant, gene codon number, gene, pathway, and other levels.

[0148] Referring to the illustrated example of FIG. 2a, a first index table 210 describes single nucleotide polymorphisms and small indels in DNA with respect to their CPRA 212 (chromosome 214, location 216, reference 218, alternative allele 220) occurring in a sample with KEIS sample ID 222. A second index table 230 describes copy number variations (CNVs) with respect to their range 232 (chromosome 234, start 236, end 238) occurring in a sample with key sample ID 242. A third index table 250 describes variants in DNA(CPRA) 252 (see first index table 210) with respect to RNS-Seq occurring in a sample with key sample ID 262. A fourth index table 270 describes copy number variations CNVs 272 with their ranges versus single nucleotide polymorphisms and small indels in DNA(CPRA) 274.

[0149] Referring to the example shown in FIG. 2b, a CPRAxTEM ranking 300 is provided, consisting of a ranking of annotations (terms) aggregated at the CPRA level 310, the GENE_CODON level 312, and the GENE level 314. Equation 320 shows an example of how to calculate the rank at the GENE_CODON level of the CPRA. Equation 322 shows an example of how to calculate the rank at the GENE level of the CPRA. A fifth index table 330 provides an example of a CPRA with GENE_CODON mapping index table. A sixth index table 340 provides an example of an annotation index table at the GENE_CODON level. A seventh index table 350 provides an example of an annotation index table at the CPRA level.

[0150] As noted above, according to various embodiments described herein, the systems and methods provide a ranking of one or more selected multi-omics data indexes. In various embodiments, the ranking can occur without associated filtering of the available cancer multi-omics data. As noted above, accessible data can include, for example, variants, genes, pathways, RNA-seq-confirmed variants, differentially expressed genes, hyper- and hypomethylated regions, expressed proteins, copy number variants, structural variants, gene fusions, phenotypes, family history, annotations, drugs, clinical trial inclusion / exclusion criteria, derived analyses (e.g., mutational signature weights, microsatellite repeat loci, image data and features extracted from the images themselves, literature data and their embeddings), and machine learning model predictions and their features (e.g., microsatellite instability status and microsatellite instability loci, key origins and alterations identified as key features of the model in order of relative importance in the prediction, predicted metastatic sites and key features of the model, and predicted neo-antigen binding affinities for MHC class I and class II molecules). In various embodiments, any combination of different multi-omics streams or individual data streams may be returned based on a user query.

[0151] For example, Figure 2b shows an example of hierarchical propagation of annotations and ranking of variants (CPRA) accumulated by weighted ranking of variant-level CPRA x cpraTERM, codon-level CPRA x codonTERM, and gene-level CPRA x geneTERM annotations.

[0152] As noted above, in accordance with various embodiments described herein, systems and methods can provide integration and ranking of multiple cancer annotation sources, including, for example, FDA labels, NCCN guidelines, NCCN Compendium Biomarkers, clinical trials, CIViC, DocM, OncoKB, Mycancergenome, Database of Genetic Biomarkers for Cancer Drugs, TCGA, ICGC, COSMIC, NCI60, CCLE, DrugBank, ClinVar, HGMD, PGMD, PharmGKB, dbSNP, dbNSFP, 1000Genomes, EXAC, CPDB, KEGG, BioCarta, BioCyc, Reactome, GenMAPP, MSigDB, Brenda, CTD, HPRD, GXD, and BIND.

[0153] According to various embodiments, the multimodal ranking engine (or module) further includes an association learning engine that integrates, for example, annotation sources of well-characterized cohorts (such as TCGA), literature data, clinical trial results, and significantly mutated genes to learn clinically actionable rankings of multi-omic data in both individual patient and cohort query use case settings. In other embodiments, the learned rankings are based on predicted pathogenicity of variations with unknown clinical significance.

[0154] As described above, according to various embodiments described herein, the systems and methods can provide rankings of cancer genetic alterations in terms of clinical actionability, pathogenicity, feature weight, or frequency. According to various embodiments, the ranking model is derived by training a supervised learning model by learning to weigh features extracted for multi-omic cancer data. For variants (e.g., exact location and specific codon) or genes (e.g., mutation type is considered), this may include, for example, whether the genetic alteration / type change is related to FDA labeling, NCCN guidelines, NCCN Biomarker Compendium, ASCO guidelines, ESMO guidelines, or other leading cancer guidelines, and indications of indications / contraindications for specific drugs. Extraction of genetic alterations / type changes from other cancer annotation sources, such as clinical trials, OncoKB, Mycancergenome, CIViC, DocM, and databases of genetic biomarkers for cancer drugs, may also be performed. The resulting features include features extracted from other relevant annotation sources such as TCGA, TCGA Significantly Mutated Genes, COSMIC Cancer Gene Census, COSMIC, ICGC, Drugbank, Swissprot, dbNSFP, HGMD, PGMD, PharmGKB, ClinVar, population allele frequency data from HLI, HLI Cancer, TCGA, COSMIC, ICGC, 1000 Genes, EXAS, and Gnomad, relevant clinical trials, embeddings from PubMed, Medline, OMIM articles, and other medical literature, and named entity embeddings extracted from medical text.

[0155] According to various embodiments, the ranking is based on support vector regression, boosted trees, etc., and other machine learning models that weight information from annotation sources such as FDA, NCCN guidelines, NCCN Biomarker Compendium, curated cancer genes, COSMIC, TCGA significantly mutated genes, known hotspots, clinical trials, and in silico predicted loss / gain of function scores (CADD, FATHMM, SIFT, Polyphen, etc.).

[0156] According to various embodiments, three learning-to-rank methods are used to derive rankings: point-wise (e.g., logistic regression), pair-wise (e.g., RankSVM, RankBoost), and list-wise approaches (e.g., LambdaMart).

[0157] According to various embodiments, variant and gene rankings can be learned separately from rankings of other documents (such as medical literature), where a separate learn-to-rank model, including, for example, BM25, PageRank, RM3, and other text document ranking models, is trained to use the weighted transformed feature set.

[0158] According to various embodiments, variant and gene rankings are learned separately or as part of a deep and wide mode together with rankings of other document types. In some embodiments, ranking text documents utilizes deep learning language modeling (LM) to rank items by their probability of being a document given a query. According to various embodiments, the deep learning language model may be a Transformer model (e.g., BERT, RoBERTa, Xlnet, Albert) fine-tuned to relevant data. Such models may also be embeddings of large-scale, pre-trained language models. According to various embodiments, document relevance is generated using the textual and temporal portions of the document, for example, by deriving multiple classes of features, including: entity features and temporal features, both derived from a set of annotations for named entity recognition (NER) and temporal tagging.

[0159] According to various embodiments, to provide additional semantic understanding, deep learning methods (e.g., deep semantic similarity models, convolutional deep semantic similarity models, iterative deep semantic similarity models, deep relevance matching models, interaction Siamese networks, lexical and semantic matching networks, long short-term memory networks, Transformer networks, word embedding methods, DeepRank) are used to address the feature engineering task of learning to rank, primarily by using features automatically learned from the raw text of the query and documents. To this end, deep learning methods use various types of neural networks, whether convolutional or iterative.

[0160] According to various embodiments described herein, the systems and methods include rankings, including clinical rankings of cancer variants and genes, and deep learning rankings, where the deep learning rankings can be derived from a deep learning model selected from the group consisting of a deep semantic similarity model, a deep and wide model, a deep language model, trained deep learning text embeddings, trained named entity recognition, including a Siamese neural network, and combinations thereof.

[0161] Figure 4a shows an example of a wide and deep model for learning variant ranking. The wide part effectively memorizes sparse features and their interactions using cross-product feature transformations from various annotation sources, while the deep part can generalize to previously unseen feature interactions and literature embeddings.

[0162] Figure 4b shows an example of a learning-to-rank engine that relies on a deep semantic similarity model (see discussion above) for biomedical data. In the particular example shown in Figure 4, a Siamese network is used to identify the relationship between a query (Q) and related documents (D) by learning joint query and document embeddings. +) to learn the semantic similarity between the query and document embeddings R(Q, D). The relevance is estimated by the cosine similarity between the query and document embeddings R(Q, D). The network can minimize the cross-entropy loss for randomly sampled negative documents D:

number

[0163] After the ranking model is trained, document embeddings are pre-computed (e.g., as the centroid of all unit vectors of words in the document). At query time, embeddings of the query vector may be generated before evaluating the similarity of the query and document representations in the joint latent space. Figure 4b is illustrative only, and the specific queries and documents referenced are in no way limited to the type of query submitted and the documents analyzed.

[0164] According to various embodiments, the global ranking may be optimized for clinical actionability (or pathogenicity when clinical utility is unknown) and pre-loaded into the index so that the results (e.g., following a top-K algorithm) better meet specific information needs. According to various embodiments, re-ranking may include the use of weighted transformed features from language modeling or standard information retrieval models (e.g., PageRank, BM25, RM3).

[0165] According to various embodiments, ranking potential biomarkers in a cohort of samples is achieved by first learning a latent space representation of the multi-omic data stream (e.g., DNA and RNA, as discussed herein) and then clustering the representation. A set of features (e.g., biomarkers) responsible for the greatest disentanglement between sub-cohorts of interest is identified. According to various embodiments, a multi-omic unsupervised deep learning approach (e.g., a variational autoencoder) is constructed for this purpose. According to various embodiments, a deep generative adversarial network is constructed using periodic losses across multiple data streams. According to various embodiments, standard dimensionality reduction techniques (e.g., principal component analysis, individual component analysis, manifold learning) are used to transform sparse and extensive multi-omic data into a meaningful latent space. These approaches can advantageously improve the ability to detect multi-omic biomarkers.

[0166] As noted above, the systems and methods described herein may, according to various embodiments, propagate rankings learned from higher levels of the biological hierarchy to inform lower levels of the biological hierarchy. For example, gene-level rankings inform variant lever rankings where information about the occurrence of variants in various cancer annotations may not be available.

[0167] According to various embodiments, rankings of variants with missing annotations can be constructed as an ensemble of gene and mutation type rankings, e.g., an aggregation function is learned that takes these aspects into account to predict overall relevance, and then a traditional learning-to-rank algorithm is applied to learn the rankings.

[0168] According to various embodiments, clinically actionable and pathogenicity rankings may be pre-loaded into the index to speed up searches. According to various embodiments, ranking formulas learned for specific combinations of multi-omics streams may be applied at index search time.

[0169] As noted above, the systems and methods described herein, according to various embodiments, can include a ranking of the results returned for a particular user query, which may depend on the combination of multi-omics data streams queried and may vary depending on the user query, taking into account the clinical relevance of the individual multi-omics data streams and the combined multi-omics data stream according to user preferences.

[0170] According to various embodiments, the rank may be changed by the user (e.g., by promoting or demoting a returned result). According to various embodiments, the rank may be changed by indirect feedback from the user, such as, for example, click-through rate and dwell time for a particular returned result.

[0171] As noted above, the systems and methods described herein, according to various embodiments, provide for collecting user feedback via web interactivity to improve the multi-omic ranking of results. For example, variants, genes, pathways, and derivative analyses may be promoted or demoted in the list of returned results based on user feedback. According to various embodiments, additional curation information may be provided and stored in the index.

[0172] In various embodiments, the systems and methods described herein may provide an interface (or interaction with an interface) for collecting explicit user feedback regarding the relevance of the returned results (e.g., the user may find a result satisfactory / promote / save / save for reporting / pin / export a particular result, or the user may find a result unsatisfactory / demote / remove a result from the list of returned results).

[0173] In various embodiments, the systems and methods described herein facilitate the collection and analysis of implicit user feedback from search logs (e.g., analysis of clicks, dwell times, query sequences, number of results returned).

[0174] In various embodiments, a collaborative search user interface may be provided (or interacted with) to allow multiple users to collaboratively improve the quality of ranking multi-omic cancer alterations (e.g., in a virtual tumor board setting).

[0175] As noted above, the systems described herein, according to various embodiments, can include a query engine, which may be configured to accept at least one user-submitted query, select, aggregate, and summarize relevant multi-omics indexes, and return a ranked multi-omics index for individual samples and / or cohorts of cancer samples.

[0176] In various embodiments, the query engine may be a stateless server that accepts user queries (e.g., as HTTP POST requests) and may respond with a ranked list of results (e.g., as asynchronous JSON) based on a collection of pre-computed and pre-combined multi-omics index files. In various embodiments, the query engine may perform at least one of the following functions: (a) parsing the query and classifying the user's intent (e.g., does the user need variants, genes, pathways, samples, single sample data, cohort sample data, sample-to-cohort comparisons, cohort-to-cohort comparisons, publications, images); (b) providing query auto-correction (e.g., using log-tuned auto-correction deep learning models), providing selective synonym expansion and abbreviation expansion, generating alternative queries (e.g., using deep learning fine-tuned Transformer models), and (c) generating query queries based on the query intent (e.g., using log-tuned auto-correction deep learning models). (c) determining the appropriate multi-omics index combination to use; (e) ranking results by relevance to the predicted query intent (e.g., clinical relevance and pathogenicity—default ranking, frequency of some queries, mutual information for others, feature weights, etc.); (f) annotating and summarizing medical literature (e.g., using deep learning summarization techniques); and (g) processing interaction / feedback signals from the UI. In various embodiments, the query engine may enable sub-second latency for all queries and scalability to hundreds of thousands of concurrent users.

[0177] At least some of these functions are illustrated in the exemplary workflows of Figures 5a and 5b, which show a query engine workflow that functions to (1) generate synonym and abbreviation expansions, (2) generate alternative (similar) queries, (3) create content-based suggestions and provide query autocomplete and autocorrect functionality, (4) classify the intent of a user query (e.g., does the user need variants, genes, pathways, samples, single-sample data, cohort-sample data, sample-to-cohort comparisons, cohort-to-cohort comparisons, publications, or images), (5) perform neural information retrieval (e.g., based on joint embeddings of the query and indexed documents), and (6) provide document summaries (e.g., multi-source text summaries) that can be returned to the user via the system UI. According to various embodiments, topic-specific term embeddings may be used for query expansion, particularly in (2) above. According to various embodiments, for text data, the neural information retrieval model may consider both matches in the term space and matches in the latent space. Additionally, named entity extraction models for, for example, variants, genes, pathways, drugs, and cancer types may be integrated to improve recall. Note that the descriptions of the specific queries, data, and summaries referenced in Figures 5a and 5b are merely exemplary and are in no way limited to the types of queries submitted, documents analyzed, and summaries produced. For example, for the particular exemplary workflow shown throughout Figures 5a and 5b, given the specific parameters of that query, the query engine may conclude that TP53 loss-of-function events are highly common in cancer, but that the R248 variant not only results in loss of tumor suppression but also functions as a gain-of-function mutation that promotes tumorigenesis in mouse models (see annotation sources CIViC and Database of Genetic Biomarkers for Cancer Drugs [GDKB]).

[0178] As mentioned above, the systems and methods described herein, according to various embodiments, may facilitate the integration of query term expansion using deep learning models trained on biomedical literature and available medical ontologies (e.g., GO, UMLS, DO, MeSH, eVOC, HPO, MPO).

[0179] As noted above, the systems described herein, according to various embodiments, can facilitate the integration of neural information retrieval models to provide better semantic understanding for ranking documents, images, and annotations. In various embodiments, distributed word representations (such as those generated by word2vec) can be combined to generate embeddings for queries and documents, and average embeddings can be used to generate effective document similarity searches.

[0180] An effective method for query-specific ranking is to build a ranking schema for each query individually. However, training a model for each query suffers from a lack of labeled data for unseen queries. However, according to various embodiments, a cancer gene modification search engine may group query types and fine-tune the ranking of specific subsets of queries with high clinical importance (e.g., queries that return cancer alterations in order of clinical actionability and pathogenicity, queries that return genes in order of clinical actionability). A corpus of hand-labeled query and document pairs may be used to derive the clinical utility of variants and genes. In various embodiments, precision and recall of results are measured.

[0181] In various embodiments, the training corpus set may include a comprehensive set of cancer cases that have been manually reviewed by cancer analysts.

[0182] In various embodiments, the manual training corpus may be constructed, for example, by a cancer analyst / curator. The analyst / curator may examine, for example, (1) genetic alterations (>0.02p or q value from MutSigCV) that are significantly mutated within a well-characterized cohort of the same cancer type (e.g., TCGA, ICGC, internal cohort), (2) the rank of the significantly mutated genes, (3) if the detected mutation is of the same type (e.g., missense, indel, nonsense) as in the well-characterized cohort, (4) if the mutation is missense, whether it occurs in a hotspot, (5) the number of patients from the well-characterized cohort that have this mutation, and (6) if further testing is required for the mutation, its location, structure, and the type of cancer of patients who have the mutation.

[0183] As noted above, the systems and methods described herein, according to various embodiments, may provide a universal search interface (as opposed to many different entry points). In various embodiments, all knowledge, whether it be multi-omics cancer data, samples, variants, genes, drugs, pathways, phenotypes, medical literature, imaging data, derivative cancer analyses, tumor features and machine learning models to predict those features, user data uploads, etc., may be accessed through the same simple search interface.

[0184] As noted above, the systems and methods described herein, in various embodiments, may provide a checklist / terminal of key actionable and significant cancer changes, derived cancer analysis, and quality control metrics for clinicians or researchers working with either individual samples or cohorts of samples.

[0185] The systems and methods described herein, according to various embodiments, may provide important cancer and inherited cancer variants reported according to ACMG guidelines.

[0186] The systems and methods described herein, according to various embodiments, can provide dynamically hyperlinked individual patient and cohort reports, where cancer changes are ranked if at least some of the report entries are hyperlinked to multimodal cancer search queries. In various embodiments, the hyperlinked report content may be dynamically generated based on queries that users create and save for reporting purposes.

[0187] The systems and methods described herein, according to various embodiments, may include at least one of integrated multi-omics results, visualizations, images, medical literature, advanced cancer analytics, and data from all levels of the cancer bioinformatics pipeline (e.g., sequence coverage, percentage of base pair changes by type, dynamic reports generated by user queries saved for visualization reports of sequence reads, supporting individual variants).

[0188] The systems and methods described herein, according to various embodiments, may be implemented as a web service with a two-factor authentication and access control layer, where access is controlled by various entities to ensure that all clients can only access samples to which they are authorized and that analyses are not performed across independent data sets.

[0189] In various embodiments, a query can consist of natural language terms (which can be conceptually arbitrary) and may be combined with special operators. In various embodiments, a query can include a speech-to-text model. In various embodiments, the special operators may allow a user to explicitly refer to specific information (e.g., a specific client) or to impose specific constraints (e.g., providing only genes or pathways as results). In various embodiments, operators may include, for example, plus signs, minus signs, equal signs, ampersands, asterisks, quotation marks, parentheses, brackets, curly braces, backslashes, slashes, colons, semicolons, hash signs (#), at signs (@), tildes (~), equal signs (=), square brackets (>), less than signs (<), and the terms AND, OR, NOT, and EXCEPT. In various embodiments, a query consists of natural language terms combined with special operators. In various embodiments, the special operators may allow a user to explicitly refer to specific information.

[0190] Figure 6 shows an example user interface 600 with a single search box 610 that allows users to enter different queries and receive ranked results. Each variant may be displayed with a wealth of data, such as variant quality control, variant metrics, allele frequency compared to population databases, therapeutic drug annotations, comparisons to cancer databases and annotation sources, and the ability to view the variant and surrounding sequence reads using the Integrated Gene Variant Browser (IGV) and explore variants in the UCSC Gene Browser.

[0191] Using section 620 of UI 600, users can examine the location and quality of the variant call. Chromosomes, locations, and variants can be listed with the mutated base highlighted in a different color than the reference. Using the UCSC link, users can view the variant in a gene browser (allowing for detailed variant investigation). Actual sequence reads can be visualized using the IGV link, allowing users to, for example, determine the confidence of the variant call and see if the variant occurs in a cluttered region or if sequencing artifacts cause the call to be unreliable.

[0192] Section 630 of UI 600 provides gene-level information. Gene names are listed and can be clicked to go to more information about the variant, including a summary of the gene and the frequency of that variant in the TCGA data. This allows users to investigate whether the variant has been found, with the same frequency, and in other tumor types. Clinical trials and other relevant clinical information for that variant can be viewed. The HGVS tab displays variants at the protein level. The Ensembl tab displays the transcripts used to map the protein and lists the dbSNPrsID. Variants can be compared to the frequency found in healthy populations (see "HLI Healthy Allele Frequency" in Figure 6). The PubMed tab links to relevant papers about that variant in the scientific literature in PubMed.

[0193] Using section 640 of the UI600, users can perform quality control of variant calling. If RNA-Seq was also performed, the RNA-Seq allele fraction is displayed. The tumor vs. normal allele fraction and read depth allow users to judge the quality of the call and determine if there is evidence of the mutation in normal blood.

[0194] Box 650 of UI 600 provides clinical information, if available.

[0195] In various embodiments, the systems described herein may include an interface that allows a user to input or use a user query. In various embodiments, the methods described herein may provide for input of or use of a user query via an interface. As described above, in various embodiments, the user query may be voice-generated. In various embodiments, the user query includes, for example, a patient / individual ID number, a cohort name / ID number, a specific gene name or gene symbol, a specific annotation source, a variant, and / or a phenotype. In various embodiments, the input may be a checkbox or clickable button that limits or filters the output to sequences, for example, variants, genes, phenotype data, specific combinations of multi-omics data streams, and statistically significant mutations, genes, or pathways. In various embodiments, results may be sortable, designated as favorites where appropriate, exported to another program, or exported to a dynamically generated report. In various embodiments, individual search terms may be combined. In various embodiments, an individual (or user) may use additional user queries or filtering to search for additional information within a particular result set. Table 1 illustrates a non-exhaustive list of example information required, example user input, and example output. Table 1 is not an exclusive or exhaustive list of queries that a user can develop. [Table 1]

[0196] Note that all references to figures in Table 1 are for guidance only and are not intended to limit relative user input and output examples related to the type of information desired by the user. For example, Figure 7 shows example search results obtained for a particular syntax ("fda+nccn@PatientSeqID"), according to various embodiments.

[0197] Further, for example, Figures 8a and 8b show example search results obtained with a particular syntax ("@PatientSeqID afrac>0.05tmb") according to various embodiments. Figure 8b specifically shows a display of one of the non-silent mutations contributing to the overall tumor mutation burden for this particular example tumor. More specifically, Figures 8a and 8b show an example of search results obtained with the particular above-referenced syntax, where the user desires a tumor mutation burden value that counts only mutations with an allelic fraction greater than 5%. The tumor mutation burden may then be displayed against the backdrop of the Cancer Genome Atlas tumor mutation values ​​grouped by cohort. The number of non-silent mutation types found in the tumor sample can also be displayed in an illustrated pie chart (see Figure 8b). This display allows the user to quickly assess potential cancer subtypes, potential sequencing issues, and an overall assessment of what is behind the tumor mutation burden value. The central area of ​​the illustrated pie chart displays the total number of non-silent mutations. The total number of nonsilent mutations is further categorized by the type of nonsilent mutation identified, further referencing the area outside the central region of the pie chart (see the legend adjacent to the pie chart). In many cancers (as seen in this example), missense mutations are likely to be the most frequent. If microsatellite unstable frameshift mutations account for a large proportion of mutations, the pie chart display allows for quick exploration of that parameter. Various sequencing artifacts may also contribute to a high proportion of unusual mutation types in that cancer. The pie chart display can also be used to determine the clinical relevance of tumor mutation burden. Some immunotherapeutic agents work best in tumors composed primarily of frameshift mutations or other specific mutation types. Therefore, the pie chart display allows users to quickly evaluate these possibilities. Below the chart, the interface generates a ranked list of all nonsilent variants with allele fractions greater than 5% (Figure 8b displays a single hit due to lack of space).

[0198] Further, for example, FIG. 9 shows example search results returned from a user query according to various embodiments. In particular, FIG. 9 shows a non-limiting example of search results obtained with the specific syntax "@PatientSeqIDmutsig." A mutational signature is the overall pattern of base pair changes occurring in a tumor across all genes. A mutational signature may be derived by counting all base pair changes in context to arrive at an overall mutagenesis pattern. An easy-to-use definition of a mutational signature can be found at https: / / cancer.sanger.ac.uk / cosmic / signatures. Identifying mutational signatures can help guide treatment, explain the underlying cause of a tumor, and resolve mutations of unknown significance. Therefore, mutational signatures are important for analyzing the overall characteristics of a tumor.

[0199] Section A of Figure 9 displays an XY chart of the types of base pair substitution types (i.e., C>A, C>G, C>T, T>A, T>C, T>G) in the context of the base pairs surrounding the mutation (i.e., 3 bp, shown on the X-axis). The frequency of each mutation type is plotted on the Y-axis. In this example, the graph is compared to the COSMIC-identified signature to arrive at an overall mutational signature for the tumor.

[0200] Section B displays the overall mutation signature percentage found in the tumor on a pie chart. This display allows the user to determine the tumor's major signature along with any minor signatures identified. In this example, the dominant signature displayed from the melanoma tumor is S7, which is consistent with the literature. If the displayed mutation signature is unexpected for that cancer type, the user may wish to investigate further.

[0201] Mutational signatures can also help guide clinical decisions. For example, consider BRCA1 / 2 mutations in breast and ovarian cancer. PARP inhibitors can be used in BRCA1 / 2-mutated cases of breast and ovarian cancer. COSMIC signature 3 is characterized by defects in BRCA or pathway genes, so identifying a tumor with signature 3 indicates a BRCA mutation process even in the absence of an identified mutation. If a tumor contains a BRCA mutation of unknown significance, analyzing for the presence of signature 3 can help determine whether the mutation is functional. In either case, the potential benefit of PARP inhibitors may be explored.

[0202] Another function accessible here is the reconstruction weights of each of the 96 triplets (not shown).

[0203] Further, for example, according to various embodiments, Figure 10 shows example search results returned from a user query. In particular, Figure 10 shows a non-limiting example of search results obtained with the specific syntax "cohort:CohortID tmb." The focus in this case may be to identify the mutational burden of tumors in a cohort. The tumor mutation burden (TMB, mutations / mb) of each tumor within the cohort (circles with associated numeric TMB values) can be compared to the TMB of tumors of the same cancer type (in this case, pancreatic cancer—PAAD) by Cancer Genome Atlas (the remainder of the circles on the plot, and most of them, do not reference associated TMB values). TMB is represented on the Y-axis, allowing users to see whether the TMB identified for the cohort is consistent with prior knowledge about that cancer. The TCGA median for PAAD is displayed as a horizontal line in the center of the box. A boxplot representation allows users to see whether the cohort samples plot within the mean or outlier range found in TCGA.

[0204] Referring to FIG. 10, a cohort TMB chart 500 is provided displaying TMB 510 on the Y-axis 512. The tumor mutation burden (TMB, mutations / mb) of each tumor in the cohort is the first point 520, which has an associated numeric TMB value 522. These values ​​are compared to the TMB of tumors of the same cancer type (in this case pancreatic cancer - PAAD) by the Cancer Genome Atlas, represented by the second point 530, but with no associated TMB value, which in this example constitutes the majority of the captured points.

[0205] Further, for example, according to various embodiments, Figure 11 shows example search results returned from a user query. In particular, Figure 11 displays an integrated summary of multiple genetic alterations and clinical information examined in a cohort of samples in response to the user query "cohort:CohortID panel:cgc nonsilent," which asks for a summary of nonsilent mutations in the Cancer Gene Census panel for a specific cohort. Effectively, the query in this case is to identify whether samples from a specific cohort have the same number and type of mutations. Each tumor sample is displayed in a column, each gene in a row, and available clinical information can be added to the table. The plot can be stratified by any of the displayed clinical parameters. The plot can be sorted first by the most frequently mutated cancer gene in the cohort (see figure), and gene-level frequencies are displayed. The type of mutation (e.g., missense, nonsense, frameshift) can be identified by the type of variant using different box colors (see section B of Figure 11). In the illustrated example, the driver gene (NRAS) is a missense mutation, as expected. The total number of mutations for each sample can also be displayed, and users can use that information to sort the plot. The display feature allows users to perform detailed analysis of the cohort as well as identify specific alterations in individual samples. The plot allows for the co-occurrence or mutual exclusivity of mutations to be seen. Individual mutations are listed below the chart (not shown).

[0206] In the case shown in Figure 11, section A shows that the leftmost sample has the most mutations. The mutation types are fairly consistent across this cohort. In some cases, samples with a high number of frameshift mutations and very high mutation counts are observed. This observation may warrant further investigation to determine whether the sample is microsatellite unstable or has an artifact. Additionally, the third sample from the left does not have an NRAS mutation like the remaining samples. However, the number and type of mutations differ from the other cohorts. This observation may warrant more thorough investigation to determine whether this difference is artifactual or biological. Section C shows a mutation table plot that can be sorted using clinical data.

[0207] Further, for example, according to various embodiments, FIG. 12 shows example search results returned from a user query. In particular, FIG. 12 shows a non-limiting example of search results obtained with the specific syntax "cohort:responder cohort:non-responder egfr." Here, a user wishes to compare mutations in the EGFR gene across two subcohorts: responders and non-responders. Ranked individual mutations can be listed below (not shown). In this example, section A provides a schematic of the EGFR gene-level germline / somatic mutations in the two cohorts (cohort responders and cohort non-responders). Section B provides the 3D protein structure, highlighting the locations affected by hotspot mutations clustered near the drug (gefitinib) binding site in the two cohorts.

[0208] FIG. 13 is a block diagram illustrating a computer system 1000 in which some of the present teachings may be implemented, or in part of an embodiment. In various embodiments of the present teachings, the computer system 1000 may include a bus 1002 or other communication mechanism for communicating information. A processor 1004 coupled with the bus 1002 for processing information. In various embodiments, the computer system 1000 may also include memory 1006, which may be random access memory (RAM) or other dynamic storage device coupled to the bus 1002 for determining instructions to be executed by the processor 1004. The memory 1006 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by the processor 1004. In various embodiments, the computer system 1000 may further include a read-only memory (ROM) 1008 or other static storage device coupled to the bus 1002 for storing static information and instructions for the processor 1004. A storage device 1010, such as a magnetic disk or optical disk, may be provided and coupled to the bus 1002 for storing information and instructions.

[0209] In various embodiments, computer system 1000 may be coupled via bus 1002 to a display 1012, such as a cathode ray tube (CRT) or liquid crystal display (LCD), for displaying information to a computer user. Input devices 1014, including alphanumeric and other keys, may be coupled to bus 1002 for communicating information and command selections to processor 1004. Another type of user input device is a cursor control 1016, such as a mouse, trackball, cursor direction keys, etc., for conveying directional information and command selections to processor 1004 and controlling cursor movement on display 1012. This input device 1014 typically has two degrees of freedom in two axes, a first axis (i.e., x) and a second axis (i.e., y), which allows the device to specify a position in a plane. However, it should be understood that input devices 1014 that allow for three-dimensional (x, y, and z) cursor movement are also contemplated herein. Displays and input devices (or interfaces as used herein) are discussed in more detail herein with respect to functionality beyond the capabilities discussed herein.

[0210] Consistent with a particular implementation of the present teachings, results are provided by computer system 1000 in response to processor 1004 executing one or more sequences of one or more instructions contained in memory 1006. Such instructions may be read into memory 1006 from another computer-readable medium or a computer-readable storage medium, such as storage device 1010. Execution of the sequences of instructions contained in memory 1006 may cause processor 1004 to perform the processes described herein. Alternatively, hard-wired circuitry may be used in place of or in combination with software instructions to implement the present teachings. Thus, implementation of the present teachings is not limited to any specific combination of hardware circuitry and software.

[0211] The terms "computer-readable medium" (e.g., data store, data storage, etc.) or "computer-readable storage medium," as used herein, and as will be described in more detail below, refer to any medium that participates in providing instructions to processor 1004 for execution. Such media may take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Examples of non-volatile media may include, but are not limited to, optical, solid-state, and magnetic disks, such as storage device 1010. Examples of volatile media may include, but are not limited to, dynamic memory, such as memory 1006. Examples of transmission media may include, but are not limited to, coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 1002.

[0212] Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape or other magnetic media, CD-ROMs, other optical media, punch cards, paper tape, other physical media with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, other memory chips or cartridges, or other tangible media from which a computer can read. Further discussion of media is provided below.

[0213] In addition to computer-readable media, instructions or data may be provided as signals on a transmission medium included in a communication device or system to provide one or more sequences of instructions to the processor 1004 of the computer system 1000 for execution. For example, a communication device may include a transceiver having signals indicative of instructions and data. The instructions and data are configured to cause one or more processors to implement the functions outlined in this disclosure. Representative examples of data communication transmission connections include, but are not limited to, a telephone modem connection, a wide area network (WAN), a local area network (LAN), an infrared data connection, an NFC connection, etc. More information about data communications is provided below.

[0214] It should be understood that the methodologies described herein, including the flowcharts, diagrams, and accompanying disclosure, can be implemented using computer system 1000 as a standalone device or on a distributed network of shared computer processing resources, such as a cloud computing network.

[0215] It should be further appreciated that in certain embodiments, a machine-readable storage device is provided for storing non-transitory machine-readable instructions for performing or executing the methods described herein. The machine-readable instructions may control all aspects of the systems and methods described herein. Furthermore, the machine-readable instructions may be initially loaded into a memory module or accessed via the cloud or an API.

[0216] In various embodiments, the systems and methods described herein may include, or use of, a digital processing device. In various embodiments, the digital processing device may include one or more hardware central processing units (CPUs) or general-purpose graphics processing units (GPGPUs) that perform the device's functions. In various embodiments, the digital processing device further comprises an operating system configured to execute executable instructions. In various embodiments, the digital processing device may optionally be connected to a computer network. In various embodiments, the digital processing device may optionally be connected to the Internet to access the World Wide Web. In various embodiments, the digital processing device may optionally be connected to a cloud computing infrastructure. In various embodiments, the digital processing device may optionally be connected to an intranet. In various embodiments, the digital processing device may optionally be connected to a data storage device.

[0217] According to various embodiments, suitable digital processing devices include, by way of non-limiting example, server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, handheld computers, Internet appliances, mobile smartphones, tablet computers, and personal digital assistants. Those skilled in the art will recognize that many smartphones are suitable for use with the systems described herein. Those skilled in the art will also recognize that selected televisions, video players, and digital music players with optional computer network connectivity are suitable for use with the systems described herein. Suitable tablet computers are known to those skilled in the art and include those with flip, slate, and convertible configurations.

[0218] In various embodiments, the digital processing device includes an operating system configured to execute executable instructions. The operating system can be software, for example, that manages the device's hardware, provides services for running applications, and includes programs and data. Those skilled in the art will recognize that suitable server operating systems include, by way of example and not limitation, FreeBSD, OpenBSD, NetBSD, Linux, Apple® MacOSX Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those skilled in the art will recognize that suitable personal computer operating systems include, by way of non-limiting example, Microsoft® Windows®, Apple® MacOSX®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In various embodiments, the operating system is provided by cloud computing. Those skilled in the art will also recognize that suitable mobile phone operating systems include, by way of non-limiting example, Nokia® Symbian® OS, Apple® iOS®, ResearchInMotion® BlackBerry® OS, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.

[0219] In various embodiments, the device comprises a storage and / or memory device. A storage and / or memory device is one or more physical devices used to temporarily or permanently store data or programs. In various embodiments, the device is volatile memory, requiring power to maintain stored information. In various embodiments, the device is nonvolatile memory, retaining stored information when power is not applied to the digital processing device. In various embodiments, the nonvolatile memory comprises flash memory. In some embodiments, the nonvolatile memory comprises dynamic random access memory (DRAM). In various embodiments, the nonvolatile memory comprises ferroelectric random access memory (FRAM). In various embodiments, the nonvolatile memory comprises phase change random access memory (PRAM). In various embodiments, the device is a storage device, including, by way of non-limiting example, a CD-ROM, a DVD, a flash memory device, a magnetic disk drive, a magnetic tape drive, an optical disk drive, and cloud computing-based storage. In various embodiments, the storage and / or memory device is a combination of devices as disclosed herein.

[0220] In various embodiments, the digital processing device includes a display for transmitting visual information to a user. In various embodiments, the display is a cathode ray tube (CRT). In various embodiments, the display is a liquid crystal display (LCD). In various embodiments, the display is a thin film transistor liquid crystal display (TFT-LCD). In various embodiments, the display is an organic light emitting diode (OLED) display. In various embodiments, the OLED display is overlaid with a passive matrix OLED (PMOLED) or an active matrix OLED (AMOLED) display. In various embodiments, the display is a plasma display. In various embodiments, the display is a video projector. In various embodiments, the display is a combination of devices such as those disclosed herein.

[0221] In various embodiments, the digital processing device includes an input device for receiving information from a user. In various embodiments, the input device is a keyboard. In various embodiments, the input device is a pointing device, including, by way of non-limiting example, a mouse, trackball, trackpad, joystick, game controller, or stylus. In various embodiments, the input device is a touchscreen or multi-touchscreen. In various embodiments, the input device is a microphone for capturing voice or other audio input. In various embodiments, the input device is a video camera or other sensor for capturing motion or visual input. In various embodiments, the input device is a Kinect, Leap Motion, or the like. In various embodiments, the input device is a combination of devices, such as those disclosed herein.

[0222] In various embodiments, the systems disclosed herein may include one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of a networked digital processing device, and the methods herein may be performed. In various embodiments, the computer-readable storage medium is a tangible component of the digital processing device. In various implementations, the computer-readable storage medium is optionally removable from the digital processing device. In various embodiments, the computer-readable storage medium includes, by way of non-limiting example, CD-ROMs, DVDs, flash memory devices, solid-state memory, magnetic disk drives, magnetic tape drives, optical disk drives, cloud computing systems and services, and the like. In various embodiments, the programs and instructions are encoded on the media permanently, substantially permanently, semi-permanently, or non-transitoryly.

[0223] In various embodiments, the systems and methods disclosed herein may include or use at least one computer program. A computer program may include a set of instructions executable by a CPU of a digital processing device, written to perform specified tasks. Computer-readable instructions may be implemented as program modules, such as functions, objects, application programming interfaces (APs), data structures, etc., that perform particular tasks or implement particular abstract data types. Those skilled in the art will recognize that computer programs may be written in a variety of languages ​​and in various versions.

[0224] The functionality of the computer-readable instructions may be combined or distributed as desired in various environments. In various embodiments, a computer program includes one sequence of instructions. In various embodiments, a computer program includes multiple sequences of instructions. In various embodiments, a computer program is provided from one location. In various embodiments, a computer program is provided from multiple locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or any combination thereof.

[0225] In various embodiments, the computer program includes a web application. Those skilled in the art will recognize that web applications, in various embodiments, utilize one or more software frameworks and one or more database systems. In various embodiments, the web application is created on a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In various embodiments, the web application utilizes one or more database systems, including, by way of non-limiting example, relational, non-relational, object-oriented, associative, and XML database systems. In various embodiments, suitable relational database systems include, by way of non-limiting example, Microsoft® SQL Server, mySQL™, and Oracle®. Those skilled in the art will also recognize that, in various embodiments, web applications are written in one or more versions of one or more languages. Web applications can be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In various embodiments, the web application is written in part in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or Extensible Markup Language (XML). In various embodiments, the web application is written in part in a presentation definition language such as Cascading Style Sheets (CSS). In various embodiments, the web application is written in part in a client-side scripting language such as Asynchronous Javascript and XML (AJAX), Flash® Actionscript, Javascript, or Silverlight®.In various embodiments, the web application is written, in part, in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, Groovy, etc. In various embodiments, the web application is written, in part, in a database query language such as Structured Query Language (SQL). In various embodiments, the web application integrates with enterprise server products such as IBM® Lotus Domino®. In various embodiments, the web application includes a media player element. In various embodiments, the media player element utilizes one or more of a number of suitable multimedia technologies, including, by way of non-limiting example, Adobe® Flash®, HTML5, Apple® QuickTime®, Microsoft®, Siverty®, Java™, and Unity®.

[0226] In various embodiments, the computer program comprises a mobile application that is provided to the mobile digital processing device. In various embodiments, the mobile application is provided to the mobile digital processing device when it is manufactured. In various embodiments, the mobile application is provided to the mobile digital processing device via a computer network as described herein.

[0227] Mobile applications may be created using hardware, languages, and development environments known to those skilled in the art and with techniques known to those skilled in the art. Those skilled in the art will recognize that mobile applications can be written in a number of languages. Suitable programming languages ​​include, by way of non-limiting example, C, C++, C#, Objective-C, Java™, Javascript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, XHTML / HTML with or without CSS, or combinations thereof.

[0228] Suitable mobile application development environments are available from several sources. Commercially available development environments include, but are not limited to, Airplay SDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available free of charge, but are not limited to, Lazarus, MobiFlex, MoSync, and Phonegap. Additionally, mobile device manufacturers distribute software development kits, but are not limited to, the iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.

[0229] Those skilled in the art will recognize that several commercial forums are available for the distribution of mobile applications, including, by way of non-limiting example, the Apple® AppStore, Google® Play, Chrome WebStore, BlackBerry® AppWorld, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo DSiShop.

[0230] In various embodiments, the computer program includes a standalone application, which is a program that runs as an independent computer process rather than an add-on to an existing process, such as a plug-in. Those skilled in the art will recognize that standalone applications are often compiled. A compiler is a computer program that converts source code written in a programming language into binary object code, such as assembly language or machine language. Suitable compiled programming languages ​​include, by way of non-limiting example, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, VB.NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In various embodiments, the computer program includes one or more executable compliant applications.

[0231] In various embodiments, the computer program includes a web browser plug-in (e.g., an extension). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Software application manufacturers support plug-ins to allow third-party developers to create features that extend the application, support easy addition of new functionality, and reduce the application's size. Plug-ins allow customization of the software application's functionality. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display specific file types. Those skilled in the art will be familiar with several web browser plug-ins, including Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In various embodiments, the toolbar includes one or more web browser extensions, add-ins, or add-ons. In various embodiments, the toolbar includes one or more explorer bars, toolbars, or desk bands.

[0232] Those skilled in the art will recognize that several plug-in frameworks are available that allow for development of plug-ins in a variety of programming languages, including, by way of non-limiting example, C++, Delphi, Java™, PHP, Python™, VB .NET, or combinations thereof.

[0233] A web browser (also called an Internet browser) is a software application designed for use on network-connected digital processing devices to retrieve, display, and traverse information resources on the World Wide Web. Suitable web browsers include, by way of non-limiting example, Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, and KDE Konqueror. In various embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, minibrowsers, or wireless browsers) are designed for use on mobile digital processing devices such as, by way of non-limiting example, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, and personal digital assistants (PDAs). Suitable mobile web browsers include, by way of non-limiting example, Google® Android® browser, RIM BlackBerry® browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony PSP™ browser.

[0234] In various embodiments, the systems and methods disclosed herein include software, server, and / or database modules or incorporate their use in the methods according to various embodiments disclosed herein. Software modules may be created by techniques known to those skilled in the art using machines, software, and languages ​​known to those skilled in the art. The software modules disclosed herein are implemented in many ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, or a combination thereof. In further various embodiments, a software module comprises multiple files, multiple sections of code, multiple programming objects, multiple programming structures, or a combination thereof. In various embodiments, one or more software modules include, by way of non-limiting examples, a web application, a mobile application, and a standalone application. In various embodiments, a software module is within one computer program or application. In various embodiments, a software module is included in multiple computer programs or applications. In various embodiments, a software module is hosted on one machine. In various embodiments, a software module is hosted on multiple machines. In various embodiments, a software module is hosted on a cloud computing platform. In various embodiments, the software modules are hosted on one or more machines in one location. In various embodiments, the software modules are hosted on one or more machines in multiple locations.

[0235] In various embodiments, the systems and methods disclosed herein include one or more databases, or incorporate the use of the same in the methods according to various embodiments disclosed herein. Those skilled in the art will recognize that many databases are suitable for storing and retrieving user, query, token, and result information. In various embodiments, non-limiting examples include suitable databases such as relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. Other non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, and Sybase. In various embodiments, the database is internet-based. Further web, non-limiting examples include suitable web browsers such as Microsoft® Internet Explorer®, Mozilla® Firefox®, Google® Chrome, Apple® Safari®, Opera Software® Opera®, and KDE Konqueror. In various embodiments, the web browser is a mobile web browser. Mobile web browsers (also known as microbrowsers, minibrowsers, or wireless browsers) are designed for use on mobile digital processing devices, including, but not limited to, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, and personal digital assistants (PDAs).Suitable mobile web browsers include, by way of non-limiting example, Google® Android® browser, RIM BlackBerry® browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony PSP™ browser.

[0236] In various embodiments, the database is web-based. In various embodiments, the database is cloud computing-based. In other embodiments, the database is based on one or more local computer storage devices.

[0237] In various embodiments, the systems and methods disclosed herein include one or more features to prevent unauthorized access. Security measures, for example, protect a user's data. In various embodiments, the data is encrypted. In various embodiments, access to the system requires multi-factor authentication and an access control layer. In various embodiments, access to the system requires two-step authentication (e.g., a web-based interface). In various embodiments, two-step authentication requires the user to enter an access code sent to the user's email or mobile phone in addition to a username and password. In some cases, after failing to enter the proper username and password, the user is locked out of their account. In various embodiments, the systems and methods disclosed herein may include mechanisms to protect the anonymity of a user's genes and their searches across any genes.

[0238] The systems and methods described herein, in various embodiments, can assist oncologists in deriving clinical insights during case review or in collaborative settings during virtual tumor boards by enabling exploration of a patient or set of patient data at any level of the cancer bioinformatics pipeline, verifying which cancer alterations are real and do not represent sequencing artifacts, reporting quality control values, integrating multi-omics data streams with advanced analytics to provide a central dashboard or "must-see" checklist of cancer features and findings, and providing clinical, prognostic, diagnostic, and treatment information for each ranked result returned. In various embodiments, the multi-omics cancer searches described herein provide "augmented intelligence" to physicians to assist with clinical decisions.

[0239] Use of the systems and methods described herein, according to various embodiments, can include clinicians as users who can use the systems and methods described herein to perform comprehensive reporting of drug targets and key alterations in tumor (and normal) genes.

[0240] The systems and methods described herein can be used in virtual tumor boards, according to various embodiments. According to various embodiments, the systems and methods described herein can be used by individual clinicians as a checklist to ensure important tumor characteristics are not overlooked and to check available clinical trials within the oncologist's institution or worldwide. According to various embodiments, the systems and methods described herein can be used by oncologists during patient-oncologist visits. In various embodiments, multiple clinicians can use collaborative features to query, visualize, and re-rank clinically actionable and pathogenic cancer alterations, and navigate through available phenotypes, as well as image and literature data, in the virtual molecular tumor board to help determine the best diagnosis and treatment. Some non-limiting examples of questions that the systems and methods described herein can address include: What are the clinically relevant cancer variants? Are there potential treatments (FDA-approved, NCNN, clinical trials)? Are the mutations identified in the tumor real? Are they supported by high-quality sequence reads? Are there mutations in difficult-to-sequence regions? Is it present only in the tumor and not in normal? Is it expressed in RNA? Is this mutation functional? What are the global tumor characteristics, tumor mutation burden, or microsatellite instability? The system can display multiple metrics that can be used to determine both overall quality and the quality of single variants. Systems and methods according to various embodiments may provide for comparing a patient's mutations to those previously described in public datasets, such as the Cancer Gene Atlas (TCGA). Systems and methods according to various embodiments may provide for comparing multiple biopsies from the same patient.

[0241] In various embodiments, users of the systems and methods described herein can include biopharmaceutical or academic researchers, who can then perform, for example, cohort tumor profiling to characterize the genetic profiles of patients with good / poor prognosis, responders / non-responders, quality control checks, identification of drug discovery targets, stratification of cohorts for potential drug response biomarkers, and rapid, iterative hypothesis generation before conducting more extensive analyses on additional validation or test cohorts. In various embodiments, the system returns ranked biomarkers that can stratify cohorts, their statistical significance, and a summary visualization of them. In various embodiments, validation queries can be suggested by a search engine to perform robust algorithmic and statistical validation. In various embodiments, the system automatically suggests iterative hypothesis refinement through refinement of suggested queries.

[0242] In various embodiments, the systems and methods described herein can, for example, identify proteins, pathways, or mutational processes that correlate with survival, resistance, or response; dig deep into differences found in one group and compare them with other datasets; examine cohort quality control to ensure the cohort analysis is reliable and not skewed based on one of the quality control parameters; investigate anomalous results to ensure they are not due to systematic issues; drill down to individual samples, outliers, or anomalous results to confirm they are real results; investigate further and quickly obtain statistical significance of the analysis; perform multi-targeted data exploration; and search literature and annotation sources for potential therapeutics. Standard bioinformatics analysis generally lacks the ability to interactively query data and refine hypotheses using domain knowledge. Internal systems are typically based on database systems, which provide relevance ranking rather than search indexes (such as those described herein), can perform integration of multiple streams of information (e.g., genes, transcriptomics, annotation, literature), and include built-in machine learning models for relevance.

[0243] As noted above, the systems and methods described herein, in various embodiments, can be configured to provide dynamically hyperlinked individual patient and cohort variant reports. All items in the report are hyperlinked to the multimodal cancer search query. In various embodiments, the hyperlinked report content is dynamically generated based on the user's query and is highlighted and saved for reporting purposes.

[0244] As described above, the systems and methods described herein, and variant embodiments thereof, can be configured to have an expert review feature that allows a user to select which query results will be used to generate a hyperlinked live report.

[0245] In various embodiments, dynamic reports never go out of date and are updated based on newly indexed information. Additionally, users can be notified of new annotations, drugs, and clinical trials that are available.

[0246] In various embodiments, the systems and methods provided herein extend analysis beyond both static clinical reports and pre-computed cancer portal analyses, dynamically generating reports hyperlinked to individual patients or cohorts. Examples of such reports include, but are not limited to, tumor profiling, drug-to-test matching, immune reports for individual samples, cohort profiling reports for cohorts of samples, etc. Reports can be tailored based on user queries and, in various embodiments, include results returned by a multi-omic cancer search and pre-selected by the user.

[0247] Applicant has discovered that a dynamic reporting paradigm based on a multi-omics cancer search system is advantageous in terms of (1) user interaction with the data beyond the capabilities of standard static PDF reports that cannot be modified or updated after running an extensive bioinformatics pipeline; (2) ranking of all multi-omics cancer alterations in terms of clinical actionability, pathogenicity, feature weight, or frequency; (3) user query of the output of the pipeline at any level from BAM to VCF to output for more complex analysis; and (4) user view of not only the machine learning model predictions but also the list of ranked features that led to a particular prediction.

Claims

1. 1. A method for utilizing a multi-omics data index for tumor profiling, comprising: storing, in a computer, a plurality of multi-omics data indexes, wherein each of the plurality of multi-omics data indexes comprises cancer-specific tokenized data; populating the additional multi-omics data and annotations associated with the additional multi-omics data, the additional multi-omics data associated with one or more indexes; indexing the ingested additional multi-omics data and annotations, outputting the ingested additional multi-omics data into tokenized data while preserving gene names, gene variant names, and multi-omics mappings between different data streams for the same patient within a particular index; receiving a user query; selecting one or more relevant multi-omics data indexes based on the user query; ranking the selected one or more multi-omics data indexes based on at least one of clinical actionability, pathogenicity, feature weight, or frequency; returning the ranked one or more multi-omics data indexes to a user; and indexing the additional multi-omics data and annotations acquired, while simultaneously indexing the additional multi-omics data and propagating annotations from higher levels of the gene hierarchy to lower levels of the gene hierarchy. propagating the rankings of the selected one or more multi-omics data indexes from higher levels of a gene hierarchy to lower levels of a gene hierarchy; The method comprising:

2. 2. The method of claim 1, wherein the multi-omics data is selected from the group consisting of genetic, transcriptomic, epigenetic, chromatin accessibility, microbiomics, proteomic, phenotypic, imaging, related literature, integrated multi-omics data, and combinations thereof.

3. 2. The method of claim 1, wherein the plurality of multi-omics data indexes further comprises somatic genetic alterations, normal genetic alterations, and cancer annotation sources.

4. The method of claim 1, further comprising deriving a cancer analysis of the selected one or more multi-omics data indexes, wherein the cancer analysis comprises quality control, tumor mutation burden, gene mutation signatures, microsatellite instability status, neoantigens, HLA allele typing, RNA confirmation mutations, copy number variations, structural variations, non-coding regulatory variants, gene fusions, pathway enrichment, cancer driver identification, mutation summary, differential gene expression, immune signatures, matching information regarding treatment outcomes of similar patients, and combinations thereof.

5. 5. The method of claim 4, wherein the cancer analysis is derived for an individual sample or a cohort of samples.

6. 5. The method of claim 4, wherein the cancer analysis comprises machine learning predictions and ranked features.

7. 7. The method of claim 6, wherein the machine learning prediction is selected and determined from a group consisting of a primary site of a country of origin classifier, a prediction of future metastatic site classifier, a prediction of microsatellite instability status, a prediction of neo-antigen binding affinity, disease state stratification, cancer lineage, and combinations thereof.

8. 2. The method of claim 1, wherein the rankings include clinical and pathogenicity rankings of cancer variants and genes.

9. 2. The method of claim 1, wherein the ranking comprises stratifying a cohort by incorporating a latent space representation of cancer data.

10. 10. The method of claim 9, wherein the cohort is stratified into responders and non-responders.

11. 10. The method of claim 9, wherein the cohort is stratified into those with long progression-free survival and those with short progression-free survival.

12. 10. The method of claim 9, wherein the cohort is stratified into different subtypes of cancer.

13. 10. The method of claim 9, wherein the latent space representation is performed by a neural network.

14. 10. The method of claim 9, wherein the latent space representation is performed by a dimensionality reduction technique.

15. The method of claim 13, wherein the neural network is selected from the group consisting of autoencoders, variational autoencoders, deep belief networks, restricted Boltzmann machines, feedforward, convolutional, recurrent, gated regression, long short-term memory, residual, and generative adversarial networks.

16. 2. The method of claim 1 , wherein the ranking further comprises a model for learning the ranking selected from the group consisting of support vector machines, boosted decision trees, regression methods, neural networks, and combinations thereof.

17. The method of claim 1 , wherein the ranking further comprises a deep learning ranking.

18. 20. The method of claim 17, wherein the deep learning ranking is selected from the group of deep semantic similarity models, convolutional deep semantic similarity models, iterative deep semantic similarity models, deep relevance matching models, deep and wide models, deep language models, transformer networks, long short-term memory networks, trained deep learning text embeddings, trained named entity recognition, Siamese neural networks, interaction Siamese networks, lexical and semantic matching networks, and combinations thereof.

19. 2. The method of claim 1, wherein the multi-omics data is selected from the group consisting of somatic calls from whole genome sequence data, somatic calls from whole exome sequence data, somatic panel sequencing from fresh frozen tissue, somatic panel sequencing from formalin-fixed paraffin embedded tissue, somatic panel sequencing from liquid biopsies, tumor and normal variant calls, tumor / normal transcriptional data indexed as confirmed variants at the RNA or gene expression level, epigenetic data, chromatin accessibility data, microbiomics data, proteomics data, single cell sequencing data, and combinations thereof.

20. The method of claim 1 , wherein the multi-omics data index further comprises extracted phenotypic data.

21. 21. The method of claim 20, wherein the phenotypic data is selected from the group consisting of electronic health records, clinical data, functional data, and combinations thereof.

22. 2. The method of claim 1 , wherein the multi-omics data index further comprises characterized imaging data.

23. 23. The method of claim 22, wherein the characterized imaging data is selected from the group consisting of histology slides, MRI images, x-rays, mammograms, ultrasounds, PET images, CT scans, and combinations thereof.

24. 5. The method of claim 4, wherein the cancer analysis is dynamically calculated after receiving the user query.

25. 2. The method of claim 1, wherein indexing the acquired additional multi-omics data and annotations further comprises indexing derived data selected from the group consisting of cancer analysis, annotations, features extracted from image data, phenotypes, medical literature data and embeddings thereof, and combinations thereof.

26. 10. The method of claim 1, wherein the ranking further comprises matching sample alterations to established drug target labels and available clinical trials.

27. 2. The method of claim 1, wherein the ranking further comprises identifying anti-cancer drug targets in the cohort by detecting potential biomarkers that stratify cohorts based on clinical variables of interest and / or statistical significance, and wherein returning the ranked one or more multi-omics data indexes to the user comprises a stratification visualization.

28. 2. The method of claim 1, wherein returning the ranked one or more multi-omics data indexes to the user further comprises dynamic creation of hyperlinked reports for individual patients and / or cohorts providing comprehensive profiling of tumors.

29. 2. The method of claim 1, wherein the user query includes user uploaded data selected from the group consisting of a panel of variants, genes, pathways, disease states, and phenotypes of interest, and the selection is selected from individual sample or cohort data subselected by the uploaded data.

30. 2. The method of claim 1, wherein the user query can be provided via a user interface and includes uploading data for indexing selected from the group consisting of genomic data, transcriptomic data, epigenetic data, chromatin accessibility data, microbiomics data, proteomics data, phenotypic data, annotation data, and combinations thereof.

31. 13. The method of claim 1, further comprising: normalizing and / or expanding the user query, classifying the intent of the query, summarizing the retrieved documents, and performing document retrieval based on query-document similarity in a latent space using deep learning techniques.

32. 10. The method of claim 1 , wherein at least one of the indexing, selecting, and ranking comprises utilizing a deep neural network.

33. 5. The method of claim 4, wherein deriving the cancer analysis comprises utilizing a deep neural network.

34. 2. The method of claim 1 , wherein returning the ranked one or more multi-omics data indexes to a user further comprises returning a summary visualization of the returned results along with the list of ranked results.

35. A non-transitory computer-readable medium having stored thereon a program for causing a computer to execute a method for utilizing a multi-omics data index for tumor profiling, comprising: storing a plurality of multi-omics data indexes, each of the plurality of multi-omics data indexes comprising cancer-specific tokenized data; incorporating additional multi-omics data and annotations associated with the additional multi-omics data, said additional multi-omics data associated with one or more indexes; indexing the captured additional multi-omics data and annotations while preserving gene names, gene variant names, and multi-omic mappings between different data streams of the same patient in a particular index to generate the tokenized captured additional multi-omics data; receiving a user query; selecting one or more relevant multi-omics data indexes based on the user query; ranking the selected one or more multi-omics data indexes based on at least one of clinical actionability; returning the ranked one or more multi-omics data indexes to a user; and indexing the additional multi-omics data and annotations acquired, while simultaneously indexing the additional multi-omics data and propagating annotations from higher levels of the gene hierarchy to lower levels of the gene hierarchy. propagating the rankings of the selected one or more multi-omics data indexes from higher levels of a gene hierarchy to lower levels of a gene hierarchy; 16. A non-transitory computer-readable medium configured by a method comprising:

36. 36. The non-transitory computer readable medium of claim 35, wherein the multi-omics data is selected from the group consisting of genomic, transcriptomic, epigenetic, chromatin accessibility, microbiomics, proteomic, phenotypic, imaging, related literature, integrated multi-omics data, and combinations thereof.

37. 36. The non-transitory computer readable medium of claim 35, wherein the plurality of multi-omics data indexes further comprise somatic genomic alterations, normal genomic alterations, and cancer annotation sources.

38. The non-transitory computer-readable medium of claim 35, further comprising deriving a cancer analysis of the selected one or more multi-omics data indexes, wherein the cancer analysis includes quality control, tumor mutation burden, genomic mutation signatures, microsatellite instability status, neoantigens, HLA allele typing, RNA confirmatory mutations, copy number variations, structural variations, non-coding regulatory variants, gene fusions, pathway enrichment, cancer driver identification, mutation summary, differential gene expression, immune signatures, matching information regarding treatment outcomes of similar patients, and combinations thereof.

39. 40. The non-transitory computer readable medium of claim 38, wherein the cancer analysis is derived for an individual sample or a cohort of samples.

40. 40. The non-transitory computer readable medium of claim 38, wherein the cancer analysis comprises machine learning predictions and ranked features.

41. The non-transitory computer readable medium of claim 40, wherein the machine learning prediction is selected from the group consisting of: primary site of origin classifier, prediction of future metastatic site classifier, prediction of microsatellite instability status, prediction of neo-antigen binding affinity, disease state stratification, cancer lineage determination, and combinations thereof.

42. 36. The non-transitory computer readable medium of claim 35, wherein the rankings include clinical rankings of cancer variants and genes.

43. 36. The non-transitory computer readable medium of claim 35, wherein the ranking comprises stratifying cohorts by incorporating a latent space representation of cancer data.

44. 44. The non-transitory computer readable medium of claim 43, wherein the cohort is stratified into responders and non-responders.

45. 44. The non-transitory computer readable medium of claim 43, wherein the cohorts are stratified into those with long progression free survival and those with short progression free survival.

46. 44. The non-transitory computer-readable medium of claim 43, wherein the latent space representation is performed by a neural network.

47. 47. The non-transitory computer-readable medium of claim 46, wherein the neural network is selected from the group consisting of an autoencoder, a variational autoencoder, a deep belief network, a restricted Boltzmann machine, a feedforward network, a convolutional network, a recurrent network, a long short-term memory network, and a generative adversarial network.

48. 36. The non-transitory computer-readable medium of claim 35, wherein the ranking further comprises a model for learning the ranking selected from the group consisting of a support vector machine, a boosted decision tree, a regression model, a neural network, and combinations thereof.

49. 36. The non-transitory computer-readable medium of claim 35, wherein the ranking further comprises a deep learning ranking.

50. 50. The non-transitory computer-readable medium of claim 49, wherein the deep learning ranking is selected from the group consisting of a deep semantic similarity model, a deep wide model, a deep language model, trained deep learning text embeddings, trained named entity extraction, a Siamese neural network, and combinations thereof.

51. 36. The non-transitory computer readable medium of claim 35, wherein the multi-omics data is selected from the group consisting of somatic calling from whole genome sequence data, somatic calling from whole exome sequence data, somatic panel sequencing from fresh frozen tissue, somatic panel sequencing from formalin-fixed paraffin embedded tissue, somatic panel sequencing from liquid biopsies, tumor vs. normal variant calling, tumor / normal transcriptomics data indexed as variants confirmed at the RNA or gene expression level, epigenetic data, chromatin accessibility data, microbiological data, proteomic data, single cell sequencing data, and combinations thereof.

52. 36. The non-transitory computer readable medium of claim 35, wherein the multi-omics data index further comprises extracted phenotypic data.

53. 53. The non-transitory computer readable medium of claim 52, wherein the phenotypic data is selected from the group consisting of electronic health records, clinical data, functional data, and combinations thereof.

54. 36. The non-transitory computer readable medium of claim 35, wherein the multi-omics data index further comprises characterized imaging data.

55. 55. The non-transitory computer readable medium of claim 54, wherein the characterized imaging data is selected from the group consisting of histology slides, MRI images, x-rays, mammograms, ultrasounds, PET images, CT scans, and combinations thereof.

56. 40. The non-transitory computer readable medium of claim 38, wherein the cancer analysis is dynamically calculated after receiving the user query.

57. The non-transitory computer readable medium of claim 35, further comprising indexing of the acquired additional multi-omics data and annotations and indexing of derived data selected from the group consisting of cancer analysis, annotations, features extracted from image data, phenotypes, medical literature data and embeddings thereof, and combinations thereof.

58. 36. The non-transitory computer readable medium of claim 35, wherein the ranking further comprises matching sample alterations with established drug target labels and available clinical trials.

59. The non-transitory computer readable medium of claim 35, wherein the ranking further comprises identifying anti-cancer drug targets in the cohort by detecting potential biomarkers, stratifying the cohort based on clinical variables of interest and / or statistical significance, and returning the ranked multi-omics data index or indexes to the user comprises visualization of the stratification.

60. 36. The non-transitory computer-readable medium of claim 35, wherein returning the ranked one or more multi-omics data indexes to a user further comprises dynamic creation of hyperlinked reports for individual patients and / or cohorts providing comprehensive profiling of tumors.

61. 36. The non-transitory computer readable medium of claim 35, wherein the user query comprises user uploaded data selected from the group consisting of a panel of variants, genes, pathways, disease states, and phenotypes of interest, and the selection comprises selected from individual sample or cohort data subselected by the uploaded data.

62. 36. The non-transitory computer-readable medium of claim 35, wherein the user query can be provided via a user interface and includes uploading data for indexing selected from the group consisting of genomic data, transcriptomic data, epigenetic data, chromatin accessibility data, microbiomics data, proteomics data, phenotypic data, annotation data, and combinations thereof.

63. 36. The non-transitory computer-readable medium of claim 35, further comprising normalizing and / or expanding the query, classifying the intent of the query, summarizing the retrieved documents, and performing document retrieval based on similarity between the query and documents in a latent space using deep learning methods.

64. 36. The non-transitory computer-readable medium of claim 35, wherein at least one of the indexing, selecting, and ranking comprises utilizing a deep neural network.

65. 40. The non-transitory computer readable medium of claim 38, wherein deriving the cancer analysis comprises utilizing a deep neural network.

66. 36. The non-transitory computer-readable medium of claim 35, wherein returning the ranked one or more multi-omics data indexes to the user further comprises returning a summary visualization of the returned results along with the list of ranked results.

67. 1. A system for utilizing a multi-omics data index for tumor profiling, comprising: a storage element configured to store a plurality of multi-omics data indexes, each of the plurality of multi-omics data indexes including cancer-specific tokenized data; and an indexing unit, comprising an indexing engine configured to import additional multi-omics data and annotations associated with the additional multi-omics data, the additional multi-omics data associated with one or more indexes, and index the obtained additional multi-omics data and annotations while preserving gene names, gene variant names, and multi-omics mappings between different data streams of the same patient within the particular index, and generate tokenized obtained additional multi-omics data; a user interface configured to receive a user query; a query engine configured to select one or more relevant multi-omics data indexes from the indexing units based on the user query; and a ranking engine configured to receive the selected one or more relevant multi-omics data indexes and rank the selected one or more relevant multi-omics data indexes based on at least one of clinical actionability, pathogenicity, feature weight, or frequency. and indexing the acquired additional multi-omics data and annotations while propagating annotations from higher levels of gene hierarchy to lower levels of gene hierarchy if the indexing engine is configured to index the additional multi-omics data. the ranking engine is further configured to propagate rankings of the selected one or more multi-omics data indexes from higher levels to lower levels of a genomic hierarchy; A system comprising:

68. 68. The system of claim 67, wherein the multi-omics data is selected from the group consisting of genomic, transcriptomic, epigenetic, chromatin accessibility, microbiomics, proteomic, phenotypic, imaging, related literature, integrated multi-omics data, and combinations thereof.

69. 68. The system of claim 67, wherein the plurality of multi-omics data indexes further comprise somatic genetic alterations, normal genetic alterations, and cancer annotation sources.

70. 68. The system of claim 67, further comprising a cancer analysis engine configured to derive a cancer analysis of the selected one or more multi-omics data indexes, the cancer analysis including quality control, tumor mutational burden, genomic mutational signatures, microsatellite instability status, neoantigens, HLA allele typing, RNA confirmatory mutations, copy number variations, structural variations, non-coding regulatory variants, gene fusions, pathway enrichment, cancer driver identification, mutation summary, differential gene expression, immune signatures, matching information regarding treatment outcomes of similar patients, and combinations thereof.

71. The system of claim 70, wherein the cancer analysis is derived for an individual sample or a cohort of samples.

72. 71. The system of claim 70, wherein the cancer analysis includes machine learning predictions and ranked features.

73. 73. The system of claim 72, wherein the machine learning prediction is selected from the group consisting of a primary site classifier, a prediction of a future metastatic site classifier, a prediction of microsatellite instability status, a prediction of neo-antigen binding affinity, disease stratification, a determination of cancer lineage, and combinations thereof.

74. 68. The system of claim 67, wherein the ranks include clinical ranks of cancer variants and genes.

75. 68. The system of claim 67, wherein the ranking comprises stratifying cohorts by incorporating a latent space representation of cancer data.

76. 76. The system of claim 75, wherein the cohort is stratified into responders and non-responders.

77. 76. The system of claim 75, wherein the cohort is stratified into those with long progression-free survival and those with short progression-free survival.

78. 76. The system of claim 75, wherein the cohort is stratified into different cancer subtypes.

79. 76. The system of claim 75, wherein the latent space representation is performed by a neural network.

80. 80. The system of claim 79, wherein the neural network is selected from the group consisting of an autoencoder, a variational autoencoder, a deep belief network, a restricted Boltzmann machine, a feedforward, a convolutional, a recurrent, a gated regression, a long short-term memory, a residual, and a generative adversarial network.

81. 68. The system of claim 67, wherein the ranking engine further comprises a model for learning to rank selected from the group consisting of a support vector machine, a boosted decision tree, a regression model, a neural network, and combinations thereof.

82. 68. The system of claim 67, wherein the rank further comprises a deep learning rank.

83. 83. The system of claim 82, wherein the deep learning rank is created from a group consisting of deep learning models selected from a deep semantic similarity model, a deep wide model, a deep language model, trained deep learning text embeddings, trained named entity recognition, a Siamese neural network, and combinations thereof.

84. 68. The system of claim 67, wherein the multi-omics data is selected from the group consisting of somatic calls from whole genome sequence data, somatic calls from whole exome sequence data, somatic panel sequencing from fresh frozen tissue, somatic panel sequencing from formalin-fixed paraffin embedded tissue, somatic panel sequencing from liquid biopsies, tumor and normal variant calls, tumor / normal transcriptional data indexed as confirmed variants at the RNA or gene expression level, epigenetic data, chromatin accessibility data, microbiomics data, proteomics data, single cell sequencing data, and combinations thereof.

85. 68. The system of claim 67, wherein the multi-omics data index further comprises extracted phenotypic data.

86. 86. The system of claim 85, wherein the phenotypic data is selected from the group consisting of electronic health records, clinical data, functional data, and combinations thereof.

87. 68. The system of claim 67, wherein the multi-omics data index further comprises characterized imaging data.

88. 88. The system of claim 87, wherein the characterized image data is selected from the group consisting of histology slides, MRI images, x-rays, mammograms, ultrasounds, PET images, CT scans, and combinations thereof.

89. 71. The system of claim 70, wherein the cancer analysis is dynamically calculated after receiving the user query.

90. 68. The system of claim 67, wherein the indexing engine is further configured to index derived data selected from the group consisting of cancer analysis, annotations, features extracted from image data, phenotypes, medical literature data and embeddings thereof, and combinations thereof.

91. 68. The system of claim 67, wherein the ranking engine is further configured to match sample changes to established drug discovery target labels and available clinical trials.

92. The system of claim 67, wherein the ranking engine is further configured to identify anti-cancer drug targets within the cohort by detecting potential biomarkers that stratify the cohort based on clinical variables of interest and / or statistical significance, and further configured to return one or more ranked multi-omics data indexes to a user via visualization of the stratification.

93. The system of claim 67, wherein the ranking engine is configured to return one or more ranked multi-omics data indexes to a user via dynamic creation of hyperlinked reports of individual patients and / or cohorts providing comprehensive profiling of tumors.

94. 68. The system of claim 67, wherein the user query comprises user uploaded data selected from the group consisting of a panel of variants, genes, pathways, disease states, and phenotypes of interest, and wherein the selection comprises querying individual sample or cohort data subselected by the uploaded data.

95. 68. The system of claim 67, wherein the user interface is configured to receive a user query including data uploaded for indexing and selected from the group consisting of genomic data, transcriptomic data, epigenetic data, chromatin accessibility data, microbiomics data, proteomics data, phenotypic data, annotation data, and combinations thereof.

96. 68. The system of claim 67, wherein the query engine is further configured to normalize and / or expand user queries, classify query intent, summarize retrieved documents, and use deep learning techniques to perform document retrieval based on similarity between the query and documents in a latent space.

97. 68. The system of claim 67, wherein at least one of the indexing engine, the query engine, and the ranking engine is configured to utilize a deep neural network.

98. 71. The system of claim 70, wherein the cancer analysis engine is configured to derive the cancer analysis utilizing a deep neural network.

99. 68. The system of claim 67, wherein the ranking engine is further configured to return the ranked multi-omics data index or indexes to the user by returning a summary visualization of the returned results along with a list of ranked results.

Citation Information

Patent Citations

  • Method and apparatus for analyzing personalized multi-omics data

    US20140052380A1

  • Genomic, metabolomic, and microbiomic search engine

    US20170270212A1

  • Phenotype / disease specific gene ranking using curated, gene library and network based data structures

    US20180095969A1

  • Systems and methods for response prediction to chemotherapy in high grade bladder cancer

    WO2016118527A1