Systems and methods for early cancer detection and subtype classification.
The neural network-based cancer detection tool addresses inefficiencies in existing methods by using a VAE to adjust for batch effects and model nonlinear dependencies, resulting in improved cancer detection and subtype classification accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- エクサイ バイオ インコーポレイテッド
- Filing Date
- 2024-04-15
- Publication Date
- 2026-05-26
AI Technical Summary
Existing statistical methods for cancer detection using oncRNA data are inefficient in adjusting for batch effects and modeling dependencies between variables, leading to inaccurate predictions.
A neural network-based cancer detection and subtype classification tool using a variational autoencoder (VAE) that encodes input RNA count data into latent variables and decoder heads to generate predictive outputs, adjusting for batch effects through variational Bayesian estimation and semi-supervised training.
The AI-based tool achieves improved accuracy in cancer detection and subtype classification, outperforming conventional models by reducing batch effects and modeling complex nonlinear relationships, with enhanced performance on datasets from diverse sources.
Smart Images

Figure 2026516660000001_ABST
Abstract
Description
Technical Field
[0001] Inventors: Babak Alipanahi, Hani Goodarzi, Fereydoun Hormozdiari, and Mehran Karimzadeh (Cross - reference) This application is a non - provisional application of U.S. Provisional Application No. 63 / 496,344, filed on April 14, 2023, which is hereby expressly incorporated by reference in its entirety, and claims the benefit of its priority under 35 U.S.C. § 119.
[0002] (Technical Field) This embodiment generally relates to artificial intelligence (AI) - based diagnostic methods, and more particularly, to systems and methods for early cancer detection and subtype classification using orphan non - coding ribonucleic acid (oncRNA) count data in a biological sample of a subject based on AI.
Background Art
[0003] (Background) Recent medical research has shown that oncRNAs (a class of small RNAs (smRNAs) that are present in tumors and rarely present in healthy tissues) can be used in early cancer diagnostic methods. For example, detection of the presence, absence, and / or amount of oncRNAs or their functional fragments in a patient's sample can be used to diagnose and subtype classify cancer for that patient. Some existing statistical methods (e.g., ensemble logistic regression models, penalized logistic regression models, and linear regression models) have been trained to predict diagnostic outputs based on available oncRNA count inputs. However, linear models are often inefficient or even inadequate in adjusting for batch effects (e.g., data variations caused by technical and non - biological factors from data sources), or for modeling dependencies and interactions between independent variables.
Brief Description of the Drawings
[0004] [Figure 1] FIG. 1 is a schematic diagram showing a process of using an artificial intelligence (AI)-based diagnostic platform for cancer detection and subtype classification according to some embodiments described herein.
[0005] [Figure 2] FIG. 2 is a schematic diagram showing an exemplary architecture of the neural network model described in FIG. 1 according to embodiments described herein.
[0006] [Figure 3A] FIGS. 3A-3B provide a schematic block diagram showing an exemplary aspect of a two-stage training process of an AI-based cancer detection and subtype classification tool according to embodiments described herein. [Figure 3B] FIGS. 3A-3B provide a schematic block diagram showing an exemplary aspect of a two-stage training process of an AI-based cancer detection and subtype classification tool according to embodiments described herein.
[0007] [Figure 4] FIG. 4 provides a schematic block diagram showing a representative aspect of an estimation process / test process of an AI-based cancer detection and subtype classification tool according to embodiments described herein.
[0008] [Figure 5] FIG. 5 provides a schematic block diagram showing data augmentation of RNA data from a training dataset according to embodiments described herein.
[0009] [Figure 6]Figure 6 is a simplified diagram showing a computing device implementing an AI-based cancer detection and subtype classification module according to one embodiment described herein.
[0010] [Figure 7] Figure 7 is a simplified diagram showing a neural network structure that implements the cancer detection and subtype classification module described in Figure 6, according to some embodiments.
[0011] [Figure 8] Figure 8 is a simplified block diagram of a networked system suitable for implementing the cancer detection and subtype classification framework described in Figures 1 to 5, as well as other embodiments described herein.
[0012] [Figure 9A] Figures 9A to 9B provide illustrative logic flow diagrams illustrating a method for training a neural network-based model to generate cancer diagnosis predictions based on the framework shown in Figures 1 to 5, according to some embodiments described herein. [Figure 9B] Figures 9A to 9B provide illustrative logic flow diagrams illustrating a method for training a neural network-based model to generate cancer diagnosis predictions based on the framework shown in Figures 1 to 5, according to some embodiments described herein.
[0013] [Figure 10] Figure 10 provides an illustrative logic flow diagram illustrating a method for subtype classification of lung cancer samples using a neural network-based model based on the framework shown in Figures 1 to 5, according to some embodiments described herein.
[0014] [Figure 11]Figures 11 to 15B provide exemplary data plots illustrating the exemplary performance of the framework shown in Figures 1 to 5, according to some embodiments described herein. [Figure 12] Figures 11 to 15B provide exemplary data plots illustrating the exemplary performance of the framework shown in Figures 1 to 5, according to some embodiments described herein. [Figure 13A] Figures 11 to 15B provide exemplary data plots illustrating the exemplary performance of the framework shown in Figures 1 to 5, according to some embodiments described herein. [Figure 13B] Figures 11 to 15B provide exemplary data plots illustrating the exemplary performance of the framework shown in Figures 1 to 5, according to some embodiments described herein. [Figure 14A] Figures 11 to 15B provide exemplary data plots illustrating the exemplary performance of the framework shown in Figures 1 to 5, according to some embodiments described herein. [Figure 14B] Figures 11 to 15B provide exemplary data plots illustrating the exemplary performance of the framework shown in Figures 1 to 5, according to some embodiments described herein. [Figure 15] Figures 11 to 15B provide exemplary data plots illustrating the exemplary performance of the framework shown in Figures 1 to 5, according to some embodiments described herein. [Modes for carrying out the invention]
[0015] Embodiments of this disclosure and their advantages will be best understood by referring to the detailed description below. It should be understood that similar reference numerals are used to identify similar elements shown in one or more of the drawings, and the indications in the drawings are for the purpose of illustrating embodiments of this disclosure and not to limit this disclosure.
[0016] (Detailed explanation) As used herein, the term “Network” may include any hardware-based or software-based framework, including any artificial intelligence network or artificial intelligence system, neural network or neural system and / or any training or learning model implemented in or with them.
[0017] As used herein, the term “module” may include a hardware-based framework or a software-based framework that performs one or more functions. In some embodiments, such a module may be implemented in one or more neural networks.
[0018] As used herein, the term "small RNA" refers to RNA species that are approximately less than 200 nt (for example, in the range of 50 nt to 100 nt).
[0019] As used herein, the term "oncRNA" refers to a type of small RNA (smRNA) (typically small non-coding RNA (small ncRNA)) that is present in tumors but rarely in healthy tissues. As a non-limiting example, oncRNA may refer to small ncRNAs that (i) have a CPM of less than 0.09 in 95% of normal serum samples; (ii) have an adjusted p-value of <0.1 after association studies comparing tumor tissue with normal tissue, adjusted for known confounding factors (such as age and sex) using a generalized linear model; and (iii) are generally less than approximately 200 nt in length (e.g., in the range of 50 nt to 100 nt). Furthermore, although oncRNA species are non-coding RNA sequences, they may partially overlap with adjacent coding sequences. A representative embodiment of the detection and / or quantification of oncRNA molecules in the target sample can be found in PCT International Application Publication No. WO2022 / 040106, which is expressly incorporated herein by reference in its entirety.
[0020] Small RNAs can be secreted in extracellular vesicles (e.g., exosomes) derived from cells. Both mRNA and small non-coding RNA species have been found in extracellular vesicles. Therefore, extracellular vesicles can provide a stable source for reliable detection of RNA biomarkers by providing a medium for the movement of RNA contents and protection of RNA contents from degradation in the extracellular environment. Small ncRNA species can serve as "oncRNA" biomarkers when they are found to be present differently in biological samples from subjects with cancer compared to "normal" subjects (i.e., subjects without cancer). Small ncRNA species or sets of small ncRNA species are present differently between samples when the difference between the expression levels in cancer cells and normal cells is determined to be statistically significant. Common tests for statistical significance include, but are not limited to, the t-test, ANOVA, Kruskal-Wallis, Wilcoxon, Mann-Whitney, chi-squared, and Fisher's exact test. oncRNA biomarkers can be used alone or in combination to provide a measure of the relative likelihood of a subject having or not having cancer.
[0021] In one practice, small ncRNA biomarkers (i.e., oncRNAs) of cancer may be discovered by whole RNA sequencing and / or small RNA sequencing of multiple cancer types and subtypes derived from various primary tissues, as well as by identifying known or previously unknown small ncRNAs specifically expressed in cancer cells. More than 260,000 such RNAs have been identified in various cells / tissues and corresponding cancer types. These oncRNA biomarkers can be used to determine the type and status of cancer in a subject (e.g., a subject whose cancer status was previously unknown, or a subject suspected of having cancer). This can be achieved by determining the levels of one or more oncRNAs or combinations thereof in a biological sample derived from that subject. Differences in the levels of one or more of these oncRNA biomarkers compared to those in a biological sample derived from a normal subject indicate that the subject has the type and primary tissue cancer associated with those oncRNA biomarkers. The method may also be performed by determining the presence or absence of one or more of the identified oncRNAs, as well as the absolute number of RNA species detected (i.e., RNA species distinct from one another), where the absolute number of RNA species detected may be the absolute number of total RNA species, total small RNA species (i.e., less than a certain maximum length), and / or total small ncRNA species. Any two or more of these methods may be used to analyze the same biological sample; that is, the sample may be analyzed in terms of (1) the level of a particular oncRNA biomarker, (2) the presence or absence of such biomarker, and / or (3) the absolute number of RNA species detected in the sample.
[0022] Existing statistical methods (e.g., ensemble logistic regression models, penalized logistic regression models, and linear regression models) have been trained to predict diagnostic outputs based on RNA count inputs (generally including at least oncRNA count inputs). However, linear models are often inefficient, or even inadequate, in adjusting for batch effects (e.g., data variability caused by technical and abiotic factors from data sources) or modeling dependencies and interactions between independent variables. These shortcomings of linear models can become particularly apparent when these models are trained on datasets containing multiple smaller datasets of RNA cancer research data from diverse data sources (which often cause batch effects due to their diverse suppliers, data sources, and other sources of variation).
[0023] Furthermore, since typically only a fraction of oncRNAs may be present in the volume of blood sample, small RNA (smRNA) fingerprinting results in sparse patterns derived from thousands of individual oncRNA species. Given the zero-plus nature of oncRNA patterns, the underlying biological variability that distinguishes different cancer types or separates cancer from non-cancerous cells may be dominated by technical confounding factors (e.g., differences in sequencing depth, differences in RNA extraction, differences in sample processing, and other unknown sources of variability). Moreover, the sample collection process itself often contains known sources of variability that should be considered (such as biological differences between donors, age, sex, and BMI). Therefore, developing generalizable liquid biopsy assays may require considering the biological characteristics of the circulating biomarker of interest, as well as freeing them from technical and biological variability in sequencing data.
[0024] In light of the challenges described above in developing robust oncRNA-based diagnostic tools, embodiments described herein provide a neural network-based cancer detection and subtype classification tool for predicting the presence of a tumor, its primary tissue, and / or its subtype using RNA sequencing (smRNA-seq) data (e.g., oncRNA count data and total small RNA count data). More specifically, this AI-based cancer detection and subtype classification tool is based on a variational autoencoder (VAE) that encodes input RNA count data into latent variables and one or more decoder heads (e.g., classification heads) for generating predictive outputs (e.g., tumor presence, primary tissue, subtype classification, and / or their similarities).
[0025] In addition to its usefulness in cancer detection, subtype classification, and tissue identification, this method is useful in at least the following: detecting cancer stages (e.g., TNM stages or stage number); detecting changes in cancer pathways (e.g., changes in mitotic signaling pathways, metabolic pathways, or DNA repair); detecting genomic abnormalities; analyzing the cellular state of cancer; detecting and analyzing various aspects of the oncogenic process (e.g., germline pathogenic variants, copy number variations, and mutations (e.g., somatic driver mutations)); and detecting and analyzing other factors related to cancer detection or characterization.
[0026] Input RNA count data may be total RNA counts and / or counts of small RNAs, miRNAs, mRNAs, small ncRNAs, and small ncRNAs previously identified as oncRNAs. The term “input RNA count” is used herein to refer to any or all of the above. For example, in one embodiment, the input RNA count data includes oncRNA counts and (total) small ncRNA counts. In another representative example, the input RNA count data includes endogenous high-expression RNA biotype counts and oncRNA counts. In yet another representative example, the input RNA count data includes oncKmer data, where “oncKmer” or “onc Kmer” refers to an RNA sequence of size k that is significantly enriched in the sequencing reads of small RNA sequencing (smRNA-seq) reads from a tumor sample compared to a normal sample. The term “onc kmer” as used herein generally refers to a kmer of size k that is enriched (by a statistically significant difference) in the small RNA sequencing (smRNA-seq) reads from a cancer-derived sample compared to a non-cancer-derived sample. The length of the k-mer can vary, for example, being about 5 nucleotides, about 10 nucleotides, about 11 nucleotides, about 12 nucleotides, about 13 nucleotides, about 14 nucleotides, about 15 nucleotides, about 20 nucleotides, about 25 nucleotides, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 45 nucleotides, about 50 nucleotides, about 100 nucleotides, or about 200 nucleotides, and the above number of nucleotides in the k-mer may be the exact nucleotide value or an approximate nucleotide value.
[0027] In one embodiment, the predictive output may take the form of a probability distribution (e.g., P(tumor presence = yes | oncRNA count) and P(tumor presence = no | oncRNA count), and / or similar). An argmax operation may be performed on that probability distribution to generate the final classification output. As another example, a threshold may be applied for predicting the presence of a tumor, for example, if P(tumor presence = yes | oncRNA count) > Th, then the presence of cancer is determined. Predictive outputs for primary tissue or subtype may be obtained in a similar manner.
[0028] As another example, given an input oncRNA count, this AI-based cancer detection and subtype classification tool can generate predictions of cancer subtypes, thereby distinguishing between two or more subtypes of a particular type of cancer, each of which exhibits a distinct pathological and / or genetic signature. As an example, but not an exhaustive one, this method could be used to identify:
[0029] Breast cancer subtypes include estrogen receptor-positive (ER+), progesterone receptor-positive (PR+), human epidermal growth factor receptor-positive (HER2+), and triple-negative (TNBC), the latter characterized by the absence of expression of all of the above hormone receptors.
[0030] Colorectal cancer (CRC) subtypes (including those exhibiting chromosomal instability (CIN), microsatellite instability (MSI), consensus molecule subtype (CMS), or CpG island methylation phenotype (CIMP));
[0031] Hepatocellular carcinoma (HCC) subtypes (including steatohepatitis-like HCC, clear cell HCC, macrotrabecular-massive HCC, sclerosing HCC, chromophobic HCC, fibrous laminous HCC, neutril-rich HCC, and lymphocyte-rich HCC); and
[0032] Non-small cell lung cancer (NSCLC) subtypes (including adenocarcinoma, squamous cell carcinoma, and large cell carcinoma).
[0033] In one embodiment, given the large number of features in the RNA count input (e.g., hundreds of thousands) and the relatively small size of cancer datasets, the number of features can sometimes exceed the number of samples in a single dataset. Therefore, multiple datasets from various data sources are often used during training. This AI-based cancer detection and subtype classification tool uses variational Bayesian estimation and semi-supervised training to adjust for batch effects, to learn a low-dimensional distribution that explains the abiotic variability of the data, and to classify cancer status, primary tissue, cancer subtype, and / or other aspects of the detected cancer as described above. For example, its VAE can transform the oncRNA training data into a latent space so that batch effects resulting from training data from various suppliers, data sources, and / or other abiotic variability can be reduced or eliminated. Thus, this AI-based cancer detection and subtype classification tool can optimize its parameters by learning the statistical representation of the input dataset through its variational Bayesian objective function through a two-stage training process. Therefore, this VAE-based AI-based cancer detection and subtype classification tool models small RNA sequence read counts in serum while simultaneously eliminating batch effects resulting from the use of two or more suppliers, two or more data sources, and / or other known or unknown sources of variation.
[0034] In this way, this AI-based cancer detection and subtype classification tool can be trained on a large aggregated dataset composed of multiple smaller datasets that do not necessarily have to be identical. Therefore, this AI-based cancer detection and subtype classification tool can effectively combine and distill information from all sub-datasets, learn the complex nonlinear relationships between input features (such as oncRNA) and targets (tumor presence, subtype, etc.), and adjust for unknown / unobserved sources of variation in the data. All of this is achieved through a customized semi-supervised deep learning model that uses variational Bayesian principles to learn the statistical representation of the data, account for unknown sources of variation, model nonlinear dependencies between independent variables, and provide more accurate predictions. This makes it possible to combine data from multiple batches or even from multiple suppliers with various data acquisition and data processing protocols, and thus enable the construction of large AI models based on that combined dataset.
[0035] Because this AI model is trained on large, heterogeneous datasets, it can generalize better than conventional linear models or other simpler models. For example, with respect to a non-small cell lung cancer (NSCLC) dataset consisting of three separate batches from two suppliers, a conventional linear regression model achieves: Area under the curve (AUC): 0.85, Stage I sensitivity at 95% specificity: 36%, and subtype classification sensitivity comparing adenocarcinoma and squamous cell carcinoma for later stages (III / IV) at 70% specificity: 46%. The AI-based cancer detection and subtype classification tools described herein outperform conventional linear models, achieving: AUC: 0.98, Stage I sensitivity at 95% specificity: 85%, and subtype classification sensitivity comparing adenocarcinoma and squamous cell carcinoma for later stages (III / IV) at 70% specificity: 0.67.
[0036] Figure 1 is a simplified diagram illustrating process 100 using an AI-based diagnostic platform for cancer detection and subtype classification according to some embodiments described herein. Process 100 illustrates a liquid biopsy approach for cancer detection, using newly annotated oncRNA emerging from lung cancer and released from the tumor as a signature for cancer detection from blood.
[0037] In one embodiment, a biological sample 102 (e.g., NSCLC samples and tumor-adjacent normal samples derived from a public dataset (e.g., The Cancer Genome Atlas (TCGA) tissue dataset)) may be input to the oncRNA discovery module 110 for oncRNA selection. For example, to identify a set of oncRNAs, smRNA-sequencing data may be collected from 10,403 tumor samples and 679 adjacent normal tissue samples derived from TCGA, which span 32 distinct tissue types. Quality control is applied to the BAM files aligned to GRCh38, and reads that were <15 base pairs or considered low complexity based on a DUST score > 2 are removed. Furthermore, reads mapped to chrUn, chrMT, or other non-human transcripts are removed. Novel smRNA loci are identified by integrating all reads across the entire 11,082 TCGA samples after filtering, and by performing peak calling on the genomic coverage to identify a set of smRNA loci that were <200 base pairs long. This resulted in 74 million distinct candidate loci for feature discovery.
[0038] In another example, to discover lung tumor-specific oncRNAs, lung tumors (n=999) and all adjacent normal tissues (n=679) were considered, and candidate loci were filtered to those appearing in at least 1% of the samples, resulting in 1,293,892 smRNAs. A generalized linear regression model could identify smRNAs that were significantly more abundant in lung tumors compared to normal tissues. Such models can be adjusted for age, sex, and principal components to capture overall smRNA expression variability between tissues and batches. After multiple testing correction, smRNA features (FDR q<0.1) that were enriched in lung tumors (OR>1) and resulted in approximately 260k lung tumor-related oncRNAs, suggesting significant significance, were obtained for downstream use in serum.
[0039] In one embodiment, the above-described TCGA smRNA-seq database for identifying 255,393 NSCLC-specific oncRNAs through differential expression analysis between NSCLC tissue and non-cancerous tissue. The NSCLC oncRNA fingerprint 106 can be generated from TCGA NSCLC samples and tumor-adjacent normal samples 102, as well as one set of non-cancerous serum independent reference cohorts 104. The oncRNA fingerprint 106 can be input into an AI model 120 to identify rare smRNAs that are selectively expressed in lung tumors compared to normal lung tissue.
[0040] In one embodiment, serum mRNA data 115 may be generated from an intratissue dataset of patient serum 112. For example, patient serum 112 may be collected from 1,050 treatment-naïve individuals (419 with NSCLC and 631 without a history of cancer). These samples were supplied by two separate suppliers, each providing both cancer and control samples as shown in Table 1 below. Cell-free smRNA may be isolated from 0.5 mL of serum to quantify the expression of NSCLC-specific oncRNAs identified in the TCGA data. A total of 237,928 (93.15%) of the oncRNAs selected from the tissue samples were detected in at least one of those serum samples. Table 1: Sample statistical information. Sample size and key statistical aspects of the training and holdout validation sets. [Table 1]
[0041] Therefore, a serum mRNA profile 114 can be extracted from a set of patient serum 112. Such a serum mRNA profile 115 can then be input into an AI model 120 along with an oncRNA fingerprint 106 to train the AI model 120 to identify cancer-related oncRNA features. For example, given an input of an oncRNA profile (count data), the AI model 120 may generate an oncRNA-based prediction 116, which may indicate a cancer diagnosis (whether a cancer to be detected is present), primary tissue, cancer subtype, and / or one or more of these.
[0042] Figure 2 is a schematic diagram showing an exemplary architecture of the neural network model 120 described in FIG. 1 according to the embodiments described herein. The AI model 120 may include a neural network model of a customized and regularized multi-input semi-supervised variational autoencoder (VAE) 210 and a decoder 220.
[0043] In one embodiment, when given an oncRNA count matrix generated from training samples (e.g., from the NSCLC oncRNA fingerprint 106 and / or serum mRNA profile 115 in FIG. 1), x i ∈Z d (201) shows the counts for d oncRNAs for the i-th sample, and r i ∈Z m (202) shows the counts for m endogenous highly expressed smRNAs for the i-th sample, respectively. y i ∈{0,1} b ×R t contains b binary and t actual targets (cancer status), and v i ∈Z c contains c known confounding factors (sample source, processing batch, etc.), respectively.
[0044] In one embodiment, the oncRNA encoder 211 can encode the oncRNA count data x, which originally existed in a high-dimensional space, into a low-dimensional latent variable Z (231) in the latent space 230 using the mapping f z : X→Z (referred to as the oncRNA encoder 211). This oncRNA encoder 211 captures the characteristics of the variations in X (201).
[0045] In one embodiment, since the common source of variation in transcriptome data is derived from the sequenced total RNA, oncRNA may not be observed for one of two reasons: it is not secreted because it is not present; or it is actually present in the blood but is not detected in the experiment due to small volume blood sampling or limited sequencing. Therefore, an additional encoder (called library encoder 212) is used, with a normal distribution q as a proxy for the logarithm of the library size. l To calculate (l|r), one set of endogenous high-expression RNA r(202) can be encoded:f l :R→l. The encoded variable l∈R(232) may represent another unobserved random variable that takes into account the input RNA level and library sequencing depth. In other words, its library size is the mr in a given mini-batch. i This is a log-normal distribution with a prior derived from the logarithms of the mean and variance. As a result, l(232) shows a strong correlation with the total number of oncRNA reads, even though it does not originate from oncRNA.
[0046] For example, the oncRNA encoder 211 may include one hidden layer having 1,500 hidden units for encoding oncRNA, while the library encoder 212 may include one hidden layer having 1,500 units for encoding library size derived from endogenous RNA. The latent space 230 may include an embedding space for d=50 latent variables for learning the underlying Gaussian distribution of the oncRNA data, and an embedding space for s=1 latent variable for learning the library size distribution derived from endogenous RNA.
[0047] In one embodiment, the decoder 220 has a decoder output (235) x^=g(z)=g(f(x)) which is approximately the same as the input x(201) (for example, ||xx^||). 2A different mapping (g:Z→X) can be adopted, such that (where is small). In a variational autoencoder that replaces the deterministic mapping from x to z, x(201) is distributed q (usually Gaussian). z It is mapped to (z|x). When x is reconstructed, z(231) is its distribution q z (z|x) can be sampled, and using this sample, x^=p x A distribution can be generated for x, which has been reconstructed as (x|z).
[0048] In one embodiment, the decoder 220 may include an oncRNA dropout module 221 and an oncRNA abundance module 222. For example, similar to gene counts in cells in single-cell RNA-seq data, some oncRNAs may be observed in only a small number of samples, and their counts are mostly zero (also called zero-overload). Assuming that these non-zero counts follow a negative binomial distribution, the oncRNA counts can be represented as a conditionally zero-overload negative binomial (ZINB) distribution p(x|z,l) where z∈R k If k≪d is the latent embedding of x, then the oncRNA dropout module 221 has zero excess parameter φ. i It is possible to generate (f φ :Z→φ), the oncRNA abundance module 222 is f ρ :Z→ρ via the transcription scale parameter ρ i It can generate f ρ When the process includes a softmax step, it forces the expression of each oncRNA as only a fraction of all the oncRNAs being expressed.
[0049] In one embodiment, μ = ρ is the gamma-Poisson representation of the negative binomial distribution. i ×e l μ can provide the shape parameter of its gamma distribution, and the input-independent, learnable parameter θ represents its inverse dispersion. Thus, oncRNA distributions conditioned on parameters μ, θ, and φ can be generated.
[0050] During training, given training inputs x(201) and r(202), the oncRNA encoder 211 and library encoder 212 perform a low-dimensional Gaussian distribution q z (z|x) and q l (l|r) can be generated, and as a result, a zero-over-negative binomial distribution q x (x|z,l) has the ability to generate realistic in silico oncRNA profiles.
[0051] In detail, the first loss L KLZ The output latent variable distribution L of the oncRND encoder 211 KLZ =D KL (q z It can be calculated based on (z|x)||p(z)), D KL is the Kullback-Leibler divergence, and p(z)=N(0,I) is the prior distribution for z.
[0052] Second loss L KLL The output latent variable distribution L of the library encoder 212 is shown. KLL =D KL (q l It can be calculated based on (l|r)||p(l|r)), where p(l|r) is the prior log-normal distribution for l. Unlike z, the prior distribution for l is different for each batch, and its log-mean and log-standard deviation are calculated based on the value of r in each mini-batch B.
[0053] A third loss (which may be a reconstruction loss) can be calculated as the negative log-likelihood of a zero-over-negative binomial distribution describing the distribution of the input oncRNA data:
number
[0054] A fourth loss (which may be a control loss or a triple margin loss) can be calculated using a known confounding factor v with respect to z (derived from annotations in the training batch), as further described in Figure 3A:
number
[0055] The fifth loss is the cross-entropy loss L between the predicted sample label 241 and the original sample label (e.g., the contrast between cancer and control). CE It can be calculated as follows. For example, the cancer estimation module 240 predicts the label 241
number
[0056] In one embodiment, during training, one or more of these five losses may be minimized during backpropagation of the encoder 210 and / or decoder 220 in order to update their weights. For example, as further shown in relation to Figures 3A and 3B, these various losses may be used to update the encoder 210 and / or decoder 220 in one or more separate training phases. In one embodiment, a joint training loss may be calculated as a weighted sum of these five losses: L=λ1L KLZ +λ2L KLL +λ3L NLL +λ4L TML+ λ5L CE。 Subsequently, the encoder 210 and decoder 220 can be updated by backpropagation using a joint loss L. Further details regarding backpropagation of the neural network model for updating the weights of this neural network model can be discussed in Figure 10.
[0057] In this way, the semi-supervised training framework for AI model 120, which uses these five different types of loss, enables its representational learning to capture the desired biological signal (e.g., cancer detection) while simultaneously eliminating unwanted confounding factors (e.g., batch effects).
[0058] Figures 3A and 3B provide simplified block diagrams illustrating exemplary embodiments of a two-stage training process for an AI-based cancer detection and subtype classification tool according to embodiments described herein. For example, two types of RNA count data are used as training inputs X (201 in Figure 2, RNA biotype used for classification) and Q (202 in Figure 2, endogenous high-expression RNA biotype used solely to estimate the dataset library size as a latent variable).
[0059] For example, the training data could be a combination of smaller datasets of RNA biotypes originating from various data sources, vendors, or other suppliers. The training data may include RNA input counts, which are counts of specific RNA sequences previously established as oncRNAs. The training input sample X or S may take the form of oncRNA counts annotated with corresponding information (e.g., one or more labels such as: presence of tumor (yes / no), size, lymph node invasion and metastatic status (TNM), tumor subtype (e.g., adenocarcinoma or squamous cell carcinoma), primary tissue (e.g., lung), tumor gene expression profile, treatment plan and monitoring, predicted minimal residual disease (MRD), and / or similar).
[0060] For example, the data encoder 211 and / or library encoder 212 may have a hidden layer of size 1,500 and a z-axis with 50 dimensions. d The parameters of can be mapped to X, and z has one dimension. s Q can be mapped to the parameter. Decoder 220 may have one hidden layer for decoding oncRNA data from the latent distribution. A dropout rate (p=0.5) and L2 regularization (L2=2) may be employed. Classification layer 310 may have one hidden layer of size 25, and 50 normalized latent values can be mapped to generative predictions for each class.
[0061] Figure 3A shows a typical embodiment of the first training phase using triplet margin loss 242. Distance metric learning may be used to enable this AI-based cancer detection and subtype classification tool to learn the biometric representation of the data regardless of the supplier, dataset, and other technological variations. As described in Figure 2, training sample X(201) is encoded into latent variable 231 by data encoder 211, and training sample Q(202) is encoded into library latent variable 232 by library encoder 212. This data encoder 211 or library encoder 212 may be a VAE that projects the original RNA biotypes X(201) and Q(202) as latent variable Z and latent library size S in latent space.
[0062] In one embodiment, for each training sample indexed as i, the ω triplet is as follows, with respect to each confounding factor v c The following can be sampled: First, training samples i and j share the same classification label (y i =y j ) do not share the same confounding factors (
number
number
[0063] In one example, the sampled triplets of the training samples may be encoded into latent variables by the data encoder 211 and / or library encoder 212 in order to calculate the triplet margin loss 242:
number
[0064] In this way, the triplet margin loss 242 is calculated based on latent variables to train the VAE data encoder 211 and the VAE library encoder 212, indicating the positive anchors for which the model should minimize the distance of each sample in the embedding space, and the negative anchors for which the model should maximize their distance in the embedding space. In other words, if healthy control samples originate from sources A and B, and cancer samples originate from sources C and D, this VAE is trained to project input samples from data sources A and B close to each other but far from samples C and D in the embedding space. Similarly, a control loss may be calculated based on the latent representation to train this VAE.
[0065] During training, a cost function may be applied to move samples from various sources or processing batches that share the same label (e.g., cancer samples from various sources) closer together, while simultaneously moving samples with different labels (e.g., cancer samples from non-cancer samples) further apart.
[0066] Figure 3B shows a typical embodiment of a second training phase using cross-entropy loss 311 and KL information loss 312. In one embodiment, after a first training phase using triplet margin loss 242 described in relation to Figure 3A, in the second training phase, the trained VAE data encoder 211 and VAE library encoder 212 from the first training phase may be connected to one or more decoder heads 220, which then generate reconstructed RNA biotypes 235 used for classification X' from latent variables Z and S. A classification head 310 may be attached to predict the output distribution (e.g., tumor presence, primary tissue, tumor subtype, and / or their likenesses). Based on the predicted output distribution compared to the ground truth labels of tumor presence, primary tissue, and tumor subtype corresponding to input samples from that training dataset, a loss may be calculated. Such a loss may be cross-entropy loss 311 (for the classification head) or Kullback-Leibler (KL) information loss 312. Subsequently, these VAE encoders (e.g., data encoder 211 and library encoder 212) and the individual decoder heads 220 can be jointly updated based on their cross-entropy loss 311 and KL loss 312.
[0067] In one embodiment, the VAE encoder and the individual decoder head can be trained separately in two separate training phases, as shown in Figures 1A and 1B.
[0068] In another embodiment, the VAE encoder and individual decoder heads can be jointly trained based on a combined loss of cross-entropy loss, KL divergence, and triplet margin loss.
[0069] Figure 4 provides a simplified block diagram showing a typical embodiment of the estimation / test process of an AI-based cancer detection and subtype classification tool according to the embodiments described herein. In the estimation / test stage, the AI-based cancer detection and subtype classification tool, including a trained VAE and a cancer estimation neural network model (e.g., various classification heads for predicting / classifying various outputs), can generate predictive outputs in response to an input of oncRNA counts obtained from any biofluid or tissue sample. Depending on the type of classification head in the cancer estimation neural network model, the outputs may include, but are not limited to, detection of tumor presence 411, predicted size, lymph node invasion and metastatic status (TNM), predicted tumor subtype 413, predicted primary tissue 412, predicted tumor gene expression profile, predicted treatment plan and monitoring, predicted minimal residual disease (MRD), and / or similar.
[0070] For example, the cancer estimation neural network model 240 performs classification through a two-layer perceptron head. The input to the classification head is derived from a batch-normalized product (i.e., z × l) of oncRNA and library-size embeddings (e.g., 231, 232 in Figure 2). During training, to improve the robustness and sensitivity to noise of the model, q is calculated for each data point. z The latent variable is sampled from (z|x)η = 100 times. During estimation, the deterministic expectations of z and l are input into the classification head.
[0071] As shown in Figure 4, the cancer estimation neural network model 240 may include various classification heads to generate a probability distribution 411 of the presence of cancer (e.g., P(presence of cancer = yes | oncRNA count)). A threshold may be applied so that a prediction of the presence of cancer can be output if P(presence of cancer = yes | oncRNA count) > Th.
[0072] As another example, the cancer estimation neural network model 240 may generate probability distributions 412 of primary tissues (e.g., distributions in a predetermined set of possible primary tissues (e.g., {lung, not lung})). In one implementation, an argmax operation may be applied to these probability distributions to output one or more predicted primary tissues.
[0073] As another example, the cancer estimation neural network model 240 may generate a probability distribution 413 of cancer subtypes (e.g., a distribution in a predefined set of possible cancer subtypes (e.g., {adenocarcinoma, squamous cell carcinoma, null} (where the class "null" corresponds to the case where cancer does not exist))). In one implementation, an argmax operation may be applied to this probability distribution to output the predicted cancer subtype.
[0074] Similarly, the cancer estimation neural network model 240 may generate probability distributions for cancer TNM stages and / or treatment recommendations.
[0075] In one embodiment, the cancer estimation neural network model 240 may generate output predictions (e.g., {presence of cancer, primary tissue, subtype}) in a joint output vector. For example, the predicted presence of cancer, primary tissue, and / or subtype may be generated in parallel by one or more classification heads. Exemplary output vectors may take the form of {yes (cancer), lung, adenocarcinoma}, {yes (cancer), lung, squamous cell carcinoma}, {no (cancer), null, null}, {yes (cancer), not lung, null}, and / or similar forms.
[0076] In one embodiment, the cancer estimation neural network model 240 may generate output predictions in a progressive manner. For example, the cancer estimation neural network model may first determine the presence of cancer. If the presence of cancer is determined, the cancer estimation neural network model may use other classification heads to generate predicted cancer subtypes, primary tissues, and / or TNM stages of cancer.
[0077] The test stage / estimation stage described in Figure 4 may be applied in a cancer diagnostic method from RNA samples. For example, in one embodiment, the cancer diagnostic method may include a step of receiving and processing an RNA sample to identify small RNAs (e.g., oncRNAs). A representative embodiment of the detection and / or quantification of oncRNA molecules in a sample of interest can be found in PCT international application publication number WO2022 / 040106. This cancer diagnostic method further includes the step of supplying the detected small ncRNA data (e.g., oncRNA counts) to a data encoder shown in Figure 2, grouped as input data X (oncRNA counts for classification), and to a library encoder shown in Figure 2, grouped as input data Q (endogenous high-expression oncRNAs for library size estimation). The cancer estimation neural network model can then generate probability distributions indicating the likelihood of cancer, and / or probability distributions indicating the likelihood of a particular primary tissue or cancer subtype.
[0078] In another example, the test stage / predictive stage described in Figure 4 may be applied to a cancer diagnostic method from a biological sample (e.g., a blood sample) from a suspected cancerous subject. The cancer diagnostic method may include receiving and processing the biological sample from the subject in a laboratory setting. For example, the expression levels of one or more small non-coding RNA biomarkers may be determined in the biological sample from the subject. A sample from the subject is a sample originating from the subject. Such a sample may be further processed after it has been obtained from that subject. For example, RNA may be isolated from the sample. In this example, the RNA isolated from the sample is also a sample from the subject. A biological sample useful for determining the levels of one or more small non-coding RNA biomarkers may be obtained from essentially any source (including cells, tissues, and fluids from throughout the body).
[0079] In some embodiments, a biological sample useful for determining the level of one or more small noncoding RNA biomarkers is a sample containing circulating small ncRNA (e.g., extracellular small ncRNA). Extracellular small ncRNA circulates freely in a wide range of biomaterials (body fluids, e.g., fluids from the circulatory system (e.g., blood or lymph samples) or fluids from other body fluids (e.g., urine or saliva)). Therefore, in some embodiments, a biological sample useful for determining the level of one or more small ncRNA biomarkers is body fluid, e.g., blood, blood fraction, serum, plasma, urine, saliva, tears, sweat, semen, vaginal secretions, lymph, bronchial secretions, cerebrospinal fluid (CSF), whole blood, feces, interstitial fluid, synovial fluid, gastric acid, sebum, mucus, bile, etc. In some embodiments, the sample is a non-invasively obtained sample (e.g., a fecal sample). In some embodiments, the sample is a human serum sample.
[0080] The substance may be solid (e.g., biological tissue). The substance may include normal, healthy tissue. The tissue may relate to various types of organs. Non-limiting examples of organs include the brain, breast, liver, lungs, kidneys, prostate, ovaries, spleen, lymph nodes (including tonsils), thyroid gland, pancreas, heart, skeletal muscle, intestines, larynx, esophagus, stomach, or combinations thereof.
[0081] The substance may include tumors. Tumors may be benign (non-cancerous), precancerous, or malignant (cancerous), or any metastases thereof. The substance may include a mixture of normal healthy tissue or tumorous tissue. The tissue may be associated with various types of organs. Non-limiting examples of organs include the brain, breast, liver, lungs, kidneys, prostate, ovaries, spleen, lymph nodes (including tonsils), thyroid gland, pancreas, heart, skeletal muscle, intestines, larynx, esophagus, stomach, or combinations thereof.
[0082] In some embodiments, the substance may comprise various cells (including eukaryotic cells, prokaryotic cells, fungal cells, cardiac cells, lung cells, kidney cells, liver cells, pancreatic cells, germ cells, stem cells, induced pluripotent stem cells, gastrointestinal cells, blood cells, cancer cells, bacterial cells, bacterial cells isolated from human microbiome samples, and circulating cells in human blood). In some embodiments, the substance may comprise cellular contents (e.g., contents of a single cell or contents of multiple cells).
[0083] In some embodiments, any of the methods disclosed herein involves the use of small volume samples. In some embodiments, the disclosed method includes the steps of isolating total RNA or small RNA (e.g., small non-coding RNA) and / or amplifying total RNA or small RNA in a sample of about 20 microliters or less, about 40 microliters or less, about 80 microliters or less, about 100 microliters or less, about 200 microliters or less, about 300 microliters or less, about 400 microliters or less, about 500 microliters or less, about 600 microliters or less, about 700 microliters or less, about 800 microliters or less, about 900 microliters or less, about 1 milliliter or less, about 1.1 milliliters or less, about 1.2 milliliters or less, about 1.3 milliliters or less, about 1.4 milliliters or less, about 1.5 milliliters or less, about 1.6 milliliters or less, about 1.7 milliliters or less, about 1.8 milliliters or less, about 1.9 milliliters or less, and about 2.0 milliliters or less. In some embodiments, the sample size is a liquid sample of about 25 microliters to about 2 milliliters, in the form of the plasma, whole blood, or serum of the subject.
[0084] In some embodiments, the disclosed method includes the steps of isolating total RNA and / or amplifying non-coding RNA in a sample of serum of about 20 microliters or less, about 40 microliters or less, about 80 microliters or less, about 100 microliters or less, about 200 microliters or less, about 300 microliters or less, about 400 microliters or less, about 500 microliters or less, about 600 microliters or less, about 700 microliters or less, about 800 microliters or less, about 900 microliters or less, about 1 milliliter or less, about 1.1 milliliters or less, about 1.2 milliliters or less, about 1.3 milliliters or less, about 1.4 milliliters or less, about 1.5 milliliters or less, about 1.6 milliliters or less, about 1.7 milliliters or less, about 1.8 milliliters or less, about 1.9 milliliters or less, or about 2.0 milliliters.
[0085] Circulating small noncoding RNAs (SNLAs) include cellular SNLAs, extracellular SNLAs in microvesicles, extracellular SNLAs in exosomes, and extracellular SNLAs unrelated to cells or microvesicles (extracellular non-vesicular SNLAs). In some embodiments, a biological sample useful for determining the level of one or more SNLA biomarkers (e.g., a sample containing circulating SNLAs) may contain cells. In other embodiments, the biological sample may be cell-free or substantially cell-free (e.g., a serum sample). In some embodiments, a sample containing circulating SNLAs (e.g., extracellular SNLAs) is a blood-derived sample. Examples of blood-derived samples include plasma samples, serum samples, and blood samples. In other embodiments, a sample containing circulating SNLAs is a lymphatic fluid sample. Circulating small non-coding RNAs are also found in urine and saliva, and biological samples derived from these sources are similarly suitable for determining the levels of one or more small non-coding RNA biomarkers.
[0086] In some embodiments, any of the methods of this disclosure include operations for isolating total RNA or small RNA from a sample or cells or extracellular vesicles. Methods for isolating RNA for expression analysis from blood, plasma and / or serum (see, e.g., Tsui NB et al. (2002) Clin. Chem. 48, 1647-53 (the entire text is incorporated herein by reference)) and methods for isolating RNA for expression analysis from urine (see, e.g., Boom R et al. (1990) J Clin Microbiol. 28, 495-503 (the entire text is incorporated herein by reference)) are described.
[0087] In some embodiments, a biological sample may be subjected to one or more processing operations as part of the methods and systems described herein. Processing operations may be performed to isolate a substance from the biological sample, purify the biological sample, separate the biological sample into one or more fractions for further use or processing, quantify the amount of a substance, detect the presence or absence of one or more substances, transform or modify a substance for further downstream processing or analysis, or any combination thereof. For example, enzymes may be added to digest proteins to remove contaminants, or to inactivate nucleases that, if not inactivated, could degrade nucleic acids (RNA or DNA) during purification. One or more of these processing operations may follow immediately after sample collection, immediately before another processing or assay operation, or be performed concurrently with another processing or assay operation. These processing operations may also be performed on a sample or intermediate in its process, which has been appropriately stored for a specified amount of time. Any number of suitable processing operations may be performed on the biological sample or any part thereof.
[0088] The processing operations described herein may include, but are not limited to, immunoassays, enzyme-linked immunosorbent assays (ELISA), radioimmunoassays (RIA), ligand-binding assays, functional assays, enzyme assays, enzymatic processing (e.g., using kinases, phosphatases, ligases, transcriptases, and reverse transcriptases), enzymatic digestion (e.g., nucleases), spectroscopic assays (e.g., ultraviolet-visible spectroscopy, Fourier transform infrared spectroscopy, and circular dichroism), spectrophotometric assays (e.g., ultraviolet-visible spectroscopy), immunoprecipitation (IP), sequencing reactions, electrophoresis, chromatography, enrichment, pull-down methods, and mass spectrometry (MS). In some embodiments, the methods described herein may include omitting one or more processing operations.
[0089] One or more processing operations may be performed on a sample or a portion thereof. The methods and systems disclosed herein may include steps of performing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more processing operations on a sample or a portion thereof. One or more of these processing operations may be performed sequentially or simultaneously. One or more of these operations may be performed on the same sample or a portion thereof (e.g., aliquots, fractions), or they may be performed on separate samples.
[0090] In some embodiments, a biological sample containing or suspected of containing RNA is subjected to one or more processing steps to facilitate downstream processing steps (e.g., isolation). In some embodiments, a sample containing or suspected of containing RNA may be treated with one or more enzymes, cofactors, and / or other reagents to induce a end repair process. The sample containing or suspected of containing one or more RNAs is subjected to treatment with polynucleotide kinase (PNK) to ensure that terminally modified RNA species having a 3' phosphate group are not lost until further analysis. That is, the PNK enzyme dephosphorylates the 3' phosphate RNA species, thereby enabling their inclusion in subsequent polyadenylation and downstream processing steps.
[0091] In some embodiments, a biological sample containing or suspected to contain one or more RNAs is subjected to one or more processing operations to remove or modify RNA modifications that may inhibit further downstream processing (e.g., subsequent reverse transcription). Alternatively, chemical modifications may be removed or retained to determine the presence or absence of a relationship between the chemical modifications and a disease state or any pathological condition (e.g., cancer) in the subject or target population. Examples of chemically modified RNA bases include N 6 -Methyladenosine (m 6 A) Inosine (I), 5-methylcytosine (m 5C), Pseudouridine (Ψ), 5-Hydroxymethylcytosine (hm 5 C), N 1 -Methyladenosine (m 1 A), or 7-methylguanosine (m 7 G) could be cited, but is not limited to them. For example, methylated RNA (e.g., m 6 A,m 5 C, hm 5 Cm 1 A or m 7 A biological sample containing or suspected to contain G may be treated with one or more demethylases (e.g., AlkB) to remove alkyl groups that may interfere with downstream processing (e.g., by reverse transcriptase).
[0092] In some embodiments, the biological sample is subjected to one or more immunoprecipitation (IP) reactions (e.g., in vitro or in vivo crosslinking and IP reaction (CLIP)). The one or more IP reactions may enrich or pull down one or more target substances. In some embodiments, the IP processing operation includes a crosslinking operation to covalently bond two or more separate substances (e.g., to link a protein to DNA, or a protein to RNA). Immunoprecipitation of these crosslinked substances may provide an indicator of biomolecules associated with each other and / or may be used to enrich specific substances known or suspected to interact with each other. In one example, one or more target RNAs are crosslinked to one or more corresponding proteins. IP of those RNAs and crosslinked proteins allows for subsequent isolation and downstream processing of those target RNAs. In another example, the IP reaction involves targeting RNA modifications (e.g., adenosine modifications (e.g., m)). 6 A,m 1 A. This can be carried out using antibodies specific to selective polyadenylation (or RNA editing that converts adenosine to inosine), uridine modification (e.g., conversion to pseudouridine), or other RNA modifications mentioned above.
[0093] In some embodiments, the sample is subjected to one or more isolation operations. The isolation operations may target a general type of molecule (e.g., nucleic acids, e.g., RNA) or a specific molecule (e.g., a specific annotated RNA molecule).
[0094] In some embodiments, the processing operation may include a step of adding one or more substances (e.g., a spike-in step). These one or more spike-in substances may be for any appropriate purpose (including, but not limited to, quality control, enrichment of a target species, depletion of a non-target species, or any combination thereof). In some embodiments, the spike-in substances may include synthetic biomolecules corresponding to the target biomolecule (e.g., nucleic acids, e.g., RNA or modified RNA (i.e., RNA with the base modifications mentioned above)). In some embodiments, the spike-in substances may include endogenous or exogenous biomolecules. The spike-in molecules may be selected based on any appropriate properties (e.g., abundance or relative abundance, origin, sequence or part thereof, overall structure or local structure (e.g., secondary or tertiary structure), or any combination thereof). Quantification of spike-in substances after one or more downstream processing operations can be used for quality control and quantity control.
[0095] In some embodiments, the processing operation may include associating a target biomolecule or set of target biomolecules with one or more unique molecular identifiers (UMIs). UMIs may be used to associate biomolecules and indicate that they originate from the same sample or a portion thereof. In one example, a UMI (e.g., a nucleic acid barcode) may be assigned to or associated with an individual sample or a portion thereof. Alternatively, those UMIs may be assigned to or associated with an individual subject. UMIs associated with individual biological samples or subjects may allow for sample pooling during downstream processing.
[0096] In addition to associating biomolecules with specific samples or individuals, UMIs can also enable downstream process control and absolute quantification of biomolecules (e.g., nucleic acids, e.g., small non-coding RNAs). In such cases, a single UMI may correspond to one or substantially one target biomolecule. In one example, target biomolecules derived from a biological sample are tagged with individual UMIs corresponding to individual biomolecules. These UMIs enable quantification of their corresponding target biomolecules and allow for control over sequencing artifacts (e.g., PCR duplication).
[0097] In some embodiments, the UMIs include nucleic acid barcodes associated with nucleic acids derived from biological samples described herein. These nucleic acid barcodes are attached to or otherwise associated with sample-derived nucleic acids to produce a set of tagged nucleic acid constructs. In some embodiments, the target nucleic acids are associated with nucleic acid barcodes corresponding to individual samples. In some embodiments, the target nucleic acids are associated with nucleic acid barcodes corresponding to individual molecules.
[0098] One subset of processing operations may provide a different function than another subset of processing operations. For example, one processing operation may be performed to determine the presence of a substance in a sample, and a second assay may be performed to isolate that substance from the sample. The processing operations may be performed in any appropriate order. For example, first, the sample may be processed to determine the presence of a target substance. Then, a second processing operation may be used to isolate the target substance from the sample. If necessary, a third processing operation may be performed to purify the isolated substance. In another example, first, the sample may be processed to isolate a substance or a group of substances. Then, a second processing operation is performed on the isolate to determine the presence or absence of a target substance or group of target substances in the isolate. If necessary, a third processing operation is performed between the first and second processing operations to purify one or more of those target substances. In another example, a processing operation is performed to enrich a sample for the presence of one or more RNAs, for example, by enriching the DNA produced from those RNAs, which is typically carried out using a solid support (e.g., streptavidin beads) that specifically binds to the target DNA (which may be modified to have, for example, covalently bound biotin groups or bound to a biotinylated complementary nucleic acid sequence). Before or after the enrichment operation, another processing operation is performed as necessary to determine the presence of those one or more RNAs. Further combinations and permutations of processing operations for a sample are contemplated herein.
[0099] In some embodiments, the sample is processed or analyzed by a third party. For example, one party may perform a purification or enrichment operation on the sample. The purified or enriched sample may then be subjected to subsequent processing operations (e.g., quantification) by another party.
[0100] Following the step of processing a biological sample to detect small ncRNAs (e.g., oncRNA counts), the cancer diagnostic method further includes the step of supplying the detected small ncRNA data (e.g., oncRNA counts) to a data encoder shown in Figure 2, grouped as input data X (oncRNA counts for classification), and to a library encoder shown in Figure 2, grouped as input data Q (endogenous high-expression oncRNAs for library size estimation). The cancer estimation neural network model can then generate probability distributions indicating the likelihood of cancer, and / or probability distributions indicating the likelihood of primary tissue or cancer subtype, and / or predicted treatment.
[0101] In another embodiment, at the estimation stage shown in Figure 5, the input oncRNA count to the data encoder and / or library encoder is obtained as a count of specific RNA sequences previously established as oncRNAs.
[0102] Figure 5 provides a simplified block diagram illustrating data augmentation of RNA data from a training dataset according to an embodiment described herein. During training, the original RNA training data 501 (e.g., samples of oncRNA count data, RNA subtypes, and / or their equivalents) is augmented as part of the training. For example, specific RNAs (e.g., counts of specific RNA sequences) are enhanced (binarization 504). The model's dependence on RNA groups rather than individual RNAs is enforced through dropout to some RNAs (502). The stochasticity of this model is also reduced with respect to changes in the data distribution, for example, by converting oncRNA data points into latent variables (503). Thus, compared to conventional linear regression models that often cannot address batch effects, small RNA sequence read counts in serum can be accurately modeled while simultaneously eliminating batch effects depending on the supplier, data source, and other unknown sources of variation.
[0103] Figure 6 is a simplified diagram showing a computing device implementing an AI-based cancer detection and subtype classification module according to one embodiment described herein. As shown in Figure 6, the computing device 600 includes a processor 610 connected to memory 620. The operation of the computing device 600 is controlled by the processor 610. Although the computing device 600 is shown with only one processor 610, it is understood that the processor 610 may represent one or more central processing units, multicore processors, microprocessors, microcontrollers, digital signal processors, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), image processing units (GPUs), and / or similar entities in the computing device 600. The computing device 600 may be implemented as a standalone subsystem, as a board added to a computing device, and / or as a virtual machine.
[0104] Memory 620 may be used to store software executed by computing device 600 and / or one or more data structures used during the operation of computing device 600. Memory 620 may include one or more types of machine-readable media. Some common forms of machine-readable media include floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH®-EPROMs, any other memory chips or cartridges, and / or any other media to which a processor or computer is adapted for reading.
[0105] The processor 610 and / or memory 620 can be located in any suitable physical arrangement. In some embodiments, the processor 610 and / or memory 620 may be implemented on the same board, in the same package (e.g., system-in-package), on the same chip (e.g., system-on-chip), and / or similar arrangements. In some embodiments, the processor 610 and / or memory 620 may include distributed, virtualized, and / or containerized computing resources. Consistent with such embodiments, the processor 610 and / or memory 620 may be located in one or more data centers and / or cloud computing facilities.
[0106] In some examples, memory 620 may include non-temporary, tangible, machine-readable media containing executable code that, when executed by one or more processors (e.g., processor 610), can cause one or more processors to implement methods further described herein. For example, as shown, memory 620 includes instructions for a cancer detection and subtype classification module 630, which may be used to implement and / or emulate the System and Model and / or to implement any of the methods further described herein. The cancer detection and subtype classification module 630 may receive input 640 (e.g., input training data (e.g., training oncRNA data)) via a data interface 615 and generate output 650 which may be a predicted distribution relating to cancer detection, primary tissue and / or subtype, and / or predicted treatment.
[0107] The data interface 615 may include a communication interface, a user interface (e.g., a voice input interface, a graphical user interface, and / or similar). For example, the computing device 600 may receive input 640 from a networked database (e.g., a training dataset) via the communication interface. Alternatively, the computing device 600 may receive input 640 from a user (e.g., an oncRNA count) via the user interface.
[0108] In some embodiments, the cancer detection and subtype classification module 630 is configured to predict the presence of cancer, its primary tissue, and / or subtype in response to an input oncRNA count sample as described herein. The cancer detection and subtype classification module 630 may further include a data encoder submodule 631 (e.g., 211 in Figures 2, 3A-3B, and 4), a library encoder submodule 632 (e.g., 212 in Figures 2, 3A-3B, and 4), a decoder submodule 633 (e.g., 220 in Figures 2 and 4), and one or more classification heads 634 (e.g., 240 in Figures 2 and 4, or 310 in Figure 3B). Each of these submodules 631-634 may take the structure of a neural network as described in Figure 7.
[0109] In one embodiment, the cancer detection and subtype classification module 630 and its submodules 631-634 may be implemented by hardware, software, and / or a combination thereof.
[0110] In one embodiment, the cancer detection and subtype classification module 630 and one or more of its submodules 631-434 may be implemented by an artificial neural network. The neural network includes a computing system based on a collection of connected units or nodes (called neurons). Each neuron receives an input signal and then generates an output by a nonlinear transformation of that input signal. Neurons are often connected by edges, and adjustable weights are often associated with those edges. The neurons are often assembled into layers so that separate layers perform separate transformations on individual inputs and output the transformed input data to the next layer. Thus, the neural network may be stored in memory 620 as the structure of the layers of neurons (including parameters that describe the nonlinear transformations at each neuron and the weights associated with the edges connecting those neurons). Exemplary neural networks may be VAEs (e.g., employed by data encoder 631 and library encoder 632), multilayer perceptrons (MLPs) which may be employed by decoder submodule 633, and / or similar.
[0111] In one embodiment, the neural network-based cancer detection and subtype classification module 630 and one or more of its submodules 631-634 can be trained by updating the underlying parameters of this neural network based on the loss described in relation to Figures 1-5. For example, the loss (e.g., triplet margin loss, cross-entropy loss, KL divergence loss, and / or similar, described in relation to the training process in Figures 1A-1B) is a metric that evaluates how far the neural network model produces predicted output values from its target output value (also called the "ground truth" value). Given the loss, the negative gradient of the loss function is calculated individually for each weight of each layer. Such negative gradients are calculated iteratively in reverse, from the last layer to the input layers of this neural network, one layer at a time. The parameters of this neural network are updated iteratively in reverse, from the last layer to the input layers, based on the calculated negative gradients (backpropagation), in order to minimize the loss. This backpropagation from the final layer to the input layer can be performed over multiple training samples in multiple training epochs. In this way, the parameters of this neural network can be updated in a direction that results in a smaller or minimized loss, indicating that the neural network has been trained to produce predicted output values closer to the target output value.
[0112] Some examples of computing devices (e.g., computing device 600) may include a non-temporary, tangible, machine-readable medium containing executable code that, when executed by one or more processors (e.g., processor 610), can cause one or more processors to perform the method process described above. Some common forms of machine-readable medium that may contain such method processes include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH®-EPROMs, any other memory chips or cartridges, and / or any other media to which a processor or computer is adapted to read.
[0113] Figure 7 is a simplified diagram showing a neural network structure that implements the cancer detection and subtype classification module 630 described in Figure 6, according to some embodiments. In some embodiments, the cancer detection and subtype classification module 630, and / or one or more of its submodules 631-634, may be implemented at least partially through the artificial neural network structure shown in Figure 7B. The neural network includes a computing system based on a collection of connected units or nodes (called neurons) (e.g., 744, 745, 746). Neurons are often connected by edges, and adjustable weights (e.g., 751, 752) are often associated with those edges. The neurons are often assembled into layers so that separate layers can perform separate transformations on individual inputs and output the transformed input data to the next layer.
[0114] For example, the neural network architecture may include an input layer 741, one or more hidden layers 742, and an output layer 743. Each layer may contain multiple neurons, and the neurons between layers are interconnected according to a specific topology of the neural network topology. The input layer 741 receives input data (e.g., 640 in Figure 6) (e.g., oncRNA training data). The number of nodes (neurons) in the input layer 741 may be determined by the number of dimensions of the input data (e.g., the length of the vector of oncRNA count data). Each node in the input layer represents a feature or attribute of its input.
[0115] The hidden layer 742 is an intermediate layer between the input and output layers of the neural network. Two hidden layers 742 are shown in Figure 7B for illustrative purposes only, and it should be noted that any number of hidden layers can be used in the neural network structure. The hidden layer 742 can extract and transform input data through a series of weight calculations and activation functions.
[0116] For example, as discussed in Figure 6, the cancer detection and subtype classification module 630 receives an input 740 of oncRNA count data and transforms this input into an output 750 of a predicted cancer diagnosis. To perform this transformation, each neuron receives an input signal, performs a weighted sum of its inputs according to the weights assigned to each connection (e.g., 751, 752), and then applies an activation function associated with the individual neuron (e.g., 761, 762, etc.) to the result. The output of the activation function is passed to the next layer of neurons or acts as the final output of the network. The activation function may be the same or different across different layers. Exemplary activation functions include, but are not limited to, sigmoid, hyperbolic tangent, normalized linear unit (ReLU), leaky normalized linear unit (Leaky ReLU), softmax, and / or similar. In this way, after multiple hidden layers, the input data received at the input layer 741 is transformed into significantly different values that represent data features corresponding to the task that the neural network structure is designed to perform.
[0117] The output layer 743 is the final layer of the neural network structure. It generates the network's output or prediction based on the calculations performed in the preceding layers (e.g., 741, 742). The number of nodes in the output layer depends on the nature of the task being addressed. For example, in a binary classification problem, the output layer may consist of a single node representing the probability of belonging to one class. In a multi-class problem, the output layer may have multiple nodes, each representing the probability of belonging to a particular class.
[0118] Therefore, the cancer detection and subtype classification module 630, and / or one or more of its submodules 631-634, may include a transformative neural network structure with layers of neurons, weights, and activation functions describing nonlinear transformations at each neuron. Such a neural network structure is often implemented in one or more hardware processors 710 (e.g., graphics processing units (GPUs)). Exemplary neural networks may be VAEs and / or their derivatives.
[0119] In one embodiment, the cancer detection and subtype classification module 630 and its submodules 631-634 may be implemented by hardware, software, and / or a combination thereof. For example, the cancer detection and subtype classification module 630 and its submodules 631-634 may include a specific neural network structure implemented and executed on various hardware platforms 760 (e.g., CPUs (central processing units), GPUs (graphics processing units), FPGAs (field-programmable gate arrays), application-specific integrated circuits (ASICs), dedicated AI accelerators (such as TPUs (tensor processing units)), and specialized hardware accelerators specifically designed for the neural network computations described herein, and / or similar). Exemplary specific hardware for the neural network structure may include, but are not limited to, the Google Edge TPU, Deep Learning Accelerator (DLA), NVIDIA AI-focused GPUs, and / or similar. The 760 hardware used to implement the neural network structure is specially configured based on factors such as the complexity of the neural network, the scale of the task (e.g., the number of training iterations, the scale of the input data, the size of the training dataset, etc.), and the desired performance.
[0120] In one embodiment, the neural network-based cancer detection and subtype classification module 630, and one or more of its submodules 631-632, can be trained by iteratively updating the underlying parameters of the neural network (e.g., bias parameters and / or coefficients in activation functions 761, 762, such as weights 751, 752, etc.) based on the loss described in relation to Figures 2, 3A-3B. For example, during forward propagation, training data (e.g., oncRNA count data derived from training samples) is fed into the neural network. The data flows through the layers of the network (741, 742), each layer performing calculations based on its weights, biases, and activation functions until the output layer 743 produces the network output 750. In some embodiments, the output layer 743 produces an intermediate output, on which the network output 750 is based.
[0121] The output generated by the output layer 743 is compared to the expected output from the training data (e.g., "ground truth," e.g., the presence of correspondingly annotated cancers) to compute a loss function that measures the discrepancy between the predicted output and the expected output. For example, this loss function could be a triplet margin loss 242, a cross-entropy loss 311, a KL divergence loss 312, and / or similar. Given this loss, the negative gradient of the loss function is computed individually for each weight of each layer. Such negative gradients are computed iteratively in reverse, from the last layer 743 of the neural network to the input layer 741, one layer at a time. These gradients quantify the sensitivity of the network's output to changes in parameters. By propagating these gradients in reverse from the output layer 743 to the input layer 741, the chain rule of calculus is applied to efficiently compute these gradients.
[0122] The parameters of this neural network are updated in reverse (backpropagation) from the last layer to the input layer using an optimization algorithm to minimize its loss based on its calculated negative gradient. This backpropagation from the last layer 743 to the input layer 741 can be performed on multiple training samples over multiple iterative training epochs. In this way, the parameters of this neural network can be progressively updated in a direction that yields a smaller or minimized loss, indicating that the neural network has been trained to produce predicted output values closer to the target output value with improved predictive accuracy. Training can continue until a stopping criterion is met (e.g., reaching the maximum number of epochs or achieving satisfactory performance on validation data). At this point, the trained network can be used to make predictions about new, unseen data (e.g., generating predictions such as cancer diagnosis, primary tissue, and cancer subtype in response to an oncRNA count input).
[0123] Neural network parameters can be trained over multiple stages. For example, initial training (e.g., pre-training) may be performed on one set of training data, and then additional training stages (e.g., fine-tuning) may be performed on another set of training data. In some embodiments, all or some of the parameters of one or more neural network models used together may be frozen so that those "frozen" parameters are not updated during their training phase. This may allow, for example, a smaller subset of those parameters to be trained without the computational cost of updating all of them.
[0124] Therefore, this training process transforms the neural network into an "updated" trained neural network with updated parameters (e.g., weights, activation function, and bias). Thus, this trained neural network improves neural network techniques in cancer diagnostic methods.
[0125] Figure 8 is a simplified block diagram of a networked system 800 suitable for implementing the cancer detection and subtype classification framework described in Figures 1A to 4, as well as other embodiments described herein. In one embodiment, the system 800 includes a user device 810 that can be operated by a user 840, data vendor servers (845, 870, and 880), a server 830, and other types of devices, servers, and / or software components that operate to implement various methodologies according to the embodiments described herein. Exemplary devices and servers may include devices, standalone servers, and enterprise-class servers that run an OS (e.g., MICROSOFT® OS, UNIX® OS, LINUX® OS, or other suitable device-based OS and / or server-based OS), similar to the computing device 400 described in Figure 4. It can be understood that the devices and / or servers shown in Figure 8 may be arranged in other ways, and that the operations performed and / or services provided by such devices and / or servers may be combined or separated for a given embodiment, and may be performed by more or fewer devices and / or servers. One or more devices and / or servers may be operated and / or maintained by the same entity or separate entities.
[0126] The user device 810, the data vendor servers (845, 870, and 880), and the server 830 can communicate with each other via the network 860. The user device 810 can be used by users 840 (e.g., drivers, system administrators, etc.) to access various functions available to the user device 810 (which may include processes and / or applications associated with the server 830 to receive output data anomaly reports).
[0127] The user device 810, the data vendor server 845, and the server 830 may each include one or more processors, memory, and other suitable components for executing instructions (e.g., program code and / or data) stored in one or more computer-readable media to implement the various applications, data, and processes described herein. For example, such instructions may be stored in one or more computer-readable media (e.g., memory or data storage devices located inside and / or outside the various components of system 800, and / or accessible by network 860).
[0128] The user device 810 may be implemented as a communication device that can utilize appropriate hardware and software configured for wired and / or wireless communication with data vendor servers 845 and / or server 830. For example, in one embodiment, the user device 810 may be implemented as an autonomous vehicle, a personal computer (PC), a smartphone, a laptop computer / tablet computer, a wristwatch with appropriate computer hardware resources, eyeglasses with appropriate computer hardware (e.g., GOOGLE GLASS®), other types of wearable computing devices, portable communication devices, and / or other types of computing devices capable of transmitting and / or receiving data (e.g., an iPad® sold by APPLE®). Although only one communication device is shown, multiple communication devices may function similarly.
[0129] The user device 810 in Figure 8 includes a user interface (UI) application 812 and / or other applications 816 (which may correspond to processes, procedures, and / or applications executable on the associated hardware). For example, this user device 810 may receive messages from the server 830 in the form of a medical report (including the predicted presence, primary tissue, and subtype of cancer), and these messages may be displayed by the UI application 812. In other embodiments, the user device 810 may include additional or other modules having specialized hardware and / or software as needed.
[0130] In various embodiments, the user device 810 includes other applications 816 to provide functionality to the user device 810, as may be desired in a particular embodiment. For example, other applications 816 may include security applications for implementing client-side security features, programmatic client applications for interfacing with appropriate application programming interfaces (APIs) by the network 860, or other types of applications. Other applications 816 may also include communication applications (e.g., email applications, texting applications, voice applications, social networking applications, and IM applications that enable the user to send and receive emails, calls, texts, and other notifications through the network 860). For example, other applications 816 may be an email application or instant messaging application that receives prediction result messages from the server 830. Other applications 816 may include device interfaces and other display modules that can receive input and / or output information. For example, other applications 816 may include a processor-executable software program for asset management, including a graphical user interface (GUI) configured to provide the user 840 with an interface to view medical reports of diagnostic prediction results. User 840 could be a patient, a doctor, or an agent processing medical results.
[0131] The user device 810 may further include a database 818 stored in the temporary and / or non-temporary memory of the user device 810, which may store various applications and data and be available during the execution of various modules of the user device 810. The database 818 may store a user profile associated with user 840, predictions previously viewed or saved by user 840, historical data received from server 830, and / or other types of user-related information. In some embodiments, the database 818 may be local to the user device 810. However, in other embodiments, the database 818 may be external to the user device 810 and accessible by the user device 810 (such as a cloud storage system and / or database accessible by network 860).
[0132] The user device 810 includes at least one network interface component 817 adapted to communicate with the data vendor server 845 and / or server 830. In various embodiments, the network interface component 817 may be a DSL (e.g., digital subscriber line) modem, a PSTN (public switched telephone network) modem, an Ethernet® device, a broadband device, a satellite device, and / or various other types of wired or wireless network communication devices (including microwave communication devices, radio frequency communication devices, infrared communication devices, Bluetooth® communication devices, and short-range wireless communication devices).
[0133] The data vendor server 845 may correspond to a server that hosts the database 819 so that it provides training datasets (such as the oncRNA count dataset) to the server 830. The database 819 may be implemented by one or more relational databases, distributed databases, cloud databases, and / or similar entities.
[0134] The data vendor server 845 includes at least one network interface component 826 adapted to communicate with the user device 810 and / or the server 830. In various embodiments, the network interface component 826 may be a DSL (e.g., digital subscriber line) modem, a PSTN (public switched telephone network) modem, an Ethernet® device, a broadband device, a satellite device, and / or various other types of wired and / or wireless network communication devices (including microwave communication devices, radio frequency communication devices, infrared communication devices, Bluetooth® communication devices, and short-range wireless communication devices). For example, in one embodiment, the data vendor server 845 may transmit asset information from the database 819 to the server 830 via the network interface 826.
[0135] Server 830 may store the cancer detection and subtype classification module 630, as shown in Figure 6, along with its submodules. In some implementations, the cancer detection and subtype classification module 430 may receive data from database 819 via network 860 to data vendor server 845 and generate detection predictions. The generated results may also be transmitted via network 860 to user device 810 for review by user 840.
[0136] In one embodiment, the cancer detection and subtype classification module 430 may receive training datasets originating from multiple vendors 845, 870, and 880. This cancer detection and subtype classification module 430 may aggregate multiple datasets from various vendors into a larger training dataset for training. As described in relation to Figures 1A and 1B, this cancer detection and subtype classification module 430 uses variational Bayesian estimation and semi-supervised training to adjust for batch effects or other sources of variation (in this case, resulting from the use of various training data originating from three vendors 845, 870, and 880) to learn a low-dimensional distribution that explains the biological variability of the data and classify cancer states, primary tissues, and cancer subtypes.
[0137] The database 832 may be stored in the temporary and / or non-temporary memory of the server 830. In one embodiment, the database 832 may store data obtained from the data vendor server 845. In one embodiment, the database 832 may store parameters for the cancer detection and subtype classification module 430. In one embodiment, the database 832 may store previously generated prediction results and their corresponding input feature vectors. In another embodiment, the database 832 stores at least two of the above, combined, as necessary, with at least one additional type of information that may be valuable in this method.
[0138] In some embodiments, the database 832 may be local to the server 830. However, in other embodiments, the database 832 may be external to the server 830 and accessible by the server 830 (e.g., a cloud storage system and / or database accessible by the network 860).
[0139] Server 830 includes at least one network interface component 833 adapted to communicate with user devices 810 and / or data vendor servers (845, 870, or 880) via network 860. In various embodiments, the network interface component 833 may include DSL (e.g., digital subscriber line) modems, PSTN (public switched telephone network) modems, Ethernet® devices, broadband devices, satellite devices, and / or various other types of wired and / or wireless network communication devices (including microwave communication devices, radio frequency (RF) communication devices, and infrared (IR) communication devices).
[0140] Network 860 may be a single network or a combination of multiple networks. For example, in various embodiments, network 860 may include the Internet or one or more intranets, fixed telephone networks, wireless networks, and / or other suitable types of networks. Thus, network 860 may correspond to small-scale communication networks (e.g., private networks or local area networks) or large-scale networks (e.g., wide area networks or the Internet) accessible by various components of system 800.
[0141] (Example workflow) Figures 9A to 9B provide illustrative logic flow diagrams illustrating a method for training a neural network-based model to generate cancer diagnostic predictions based on the framework shown in Figures 1 to 5, according to some embodiments described herein. One or more of the processes of Method 900 can be implemented at least in part in the form of executable code stored in a non-temporary, tangible, machine-readable medium, which may cause one or more processors to perform one or more of those processes if they are executed by one or more processors. In some embodiments, Method 900 corresponds to the operation of a cancer detection and subtype classification module 630 (e.g., Figures 6 and 8) which is trained to generate cancer diagnostic predictions.
[0142] As shown, this method 900 includes a number of enumerated steps, but embodiments of this method 900 may include additional steps before, after, and between those enumerated steps. In some embodiments, one or more of those enumerated steps may be omitted or performed in a different order.
[0143] In step 901, a training sample of oncRNA count data (e.g., 201 in Figure 2) may be received by the communication interface during the training epoch.
[0144] In step 902, a positive sample having the same label as the training sample of the oncRNA count data and a negative sample having a different label from the training sample of the oncRNA count data can be sampled from the training sample batch, for example, as described in relation to Figure 3A.
[0145] In step 903, the first loss may be calculated based on the distance metrics between the training sample, the positive sample, and the negative sample in the latent space (e.g., the triplet margin loss 242 in Figure 3A).
[0146] In step 904, a second loss may be calculated based on the Kullback-Leibler divergence between the conditional distribution of the encoded latent variable of the training sample conditioned on that training sample and the prior distribution of the encoded latent variable (e.g., the KL divergence loss 312 in Figure 3B).
[0147] In step 905, a third loss may be calculated based on the reconstructed distribution of training samples from its encoded latent variable (e.g., 231 in Figure 2). For example, a decoder may generate the reconstructed distribution of training samples from the encoded latent variable.
[0148] In step 906, the decoder's classification head (e.g., 310 in Figure 3B) can generate the predicted classification of the training samples from its encoded latent variables.
[0149] In step 907, a fourth loss may be calculated as the cross-entropy between its predicted classification and the annotated labels of the training samples (e.g., the cross-entropy loss 311 in Figure 3B).
[0150] In step 908, the encoder and decoder may be jointly trained based on a joint loss as a weighted sum of their first, second, third, and fourth losses. Alternatively, the encoder may be trained at least partially based on a first loss (e.g., as shown in Figure 3A) in a first training stage, and the encoder and decoder may be trained at least partially based on a second or fourth loss (e.g., as shown in Figure 3B) in a second training stage following the first training stage.
[0151] Figure 10 provides an illustrative logic flow diagram illustrating a method for subtype classifying lung cancer samples by a neural network-based model based on the framework shown in Figures 1 to 5, according to some embodiments described herein. One or more of the processes of Method 900 can be implemented at least in part in the form of executable code stored in a non-temporary, tangible, machine-readable medium, which may cause one or more processors to perform one or more of those processes if they are executed by one or more processors. In some embodiments, Method 1000 corresponds to the operation of a cancer detection and subtype classification module 630 (e.g., Figures 6 and 8) which is trained to generate cancer diagnostic predictions.
[0152] As shown, Method 1000 includes a number of enumerated steps, but embodiments of Method 900 may include additional steps before, after, and between those enumerated steps. In some embodiments, one or more of those enumerated steps may be omitted or performed in a different order.
[0153] In step 1001, input oncRNA count data related to lung cancer samples obtained from the subject may be received via the communication interface.
[0154] In step 1002, the encoder (e.g., encoder 210 in Figure 2) can convert the input oncRNA count data into latent variables (e.g., 231 in Figure 2) in the latent space (e.g., 230 in Figure 2).
[0155] In step 1003, the decoder (e.g., 220 in Figure 2) may generate a cancer diagnostic prediction, including a first prediction about the presence of lung cancer (e.g., 411 in Figure 4), and a second prediction (e.g., 413 in Figure 4) about whether the subtype of the lung cancer sample is adenocarcinoma or squamous cell carcinoma, based on latent variables. [Examples]
[0156] (Examples) In this example, the embodiments described herein were applied to analyze oncRNA extracted from serum to investigate their usefulness for early detection and subtype classification (squamous cell carcinoma and adenocarcinoma) of non-small cell lung cancer (NSCLC). 887 serum samples were collected from Indivumed (Hamburg, Germany; 222 control samples from individuals with benign lung, breast, or colon disease, and 320 from individuals with NSCLC) and the MT Group (Los Angeles, CA; 345 control samples from individuals with no known history of cancer). These samples were collected retrospectively as part of four independent intrathecal studies. RNA isolated from 0.5 mL of serum from each individual was used to construct smRNA libraries, which were sequenced using 50 bp single-ended reads with a mean depth of 18.5 ± 6.5 × 1 million.
[0157] Using the Cancer Genome Atlas (TCGA) smRNA-seq database and a single cohort of tissue reference serum from non-cancer donors (for filtration of bona fide smRNAs), 255,953 distinct NSCLC-specific oncRNA species were identified. After processing serum samples for this study, 185,905 (72.6%) oncRNAs were detected in at least one sample.
[0158] To model samples from multiple suppliers and studies, the customized semi-supervised generative AI models described herein were used for statistical estimation, batch correction, and prediction of cancer presence and its subtypes. For comparison, a standard linear model with elastic network regularization was used as a baseline.
[0159] The clinical cohorts used for the training datasets for the training process described in Figures 3A and 3B were as follows: [Table 2] [Table 3]
[0160] (Workflow for smRNA-seq of RNA isolated from serum): (i) RNA was isolated from 1 mL of serum using the Zymo Research Quick-cfRNA serum and plasma kit (catalog number R105) according to the manufacturer's protocol; (ii) A library was prepared from the RNA isolated in step (i) using the Takara SMARTer smRNA-seq kit (catalog number 635031) obtained from Illumina; and (iii) The library prepared in step (ii) was sequenced using the Illumina NextSeq2000 instrument.
[0161] Using 10-fold cross-validation with the AI-based tools described herein, and using 10-fold cross-validation with a linear model, the AUCs were 0.98 (95% confidence interval: 0.97–0.99) and 0.85 (0.82–0.87), respectively. More importantly, the sensitivity for Stage I, with a 95% specificity, was 0.88 (0.82–0.93) for the AI model and 0.36 (0.28–0.45) for the linear model. The sensitivities for subsequent stages (II, III, and IV) were 0.93 (0.88–0.96) for the AI model and 0.42 (0.35–0.49) for the linear model, respectively. Regarding the detection of tumors smaller than 2 cm (T1a-b), this AI model achieved a sensitivity of 0.85 (0.73-0.94) with a 95% specificity, while the linear model had a sensitivity of 0.35 (0.22-0.49).
[0162] Furthermore, this AI-based tool was trained to distinguish between squamous cell carcinoma and adenocarcinoma in later-stage (III / IV) NSCLC using serum small RNA content. This achieved a sensitivity of 0.67 (0.53–0.8) at 70% specificity, while the linear model had a sensitivity of 0.46 (0.32–0.6) at 70% specificity.
[0163] Therefore, these results demonstrate that oncRNA profiling and the AI-based tools described herein can be applied to the accurate, sensitive, and early detection of NSCLC through sequencing of routine blood samples. Furthermore, this AI model may establish the role of oncRNA as a non-invasive biomarker for predicting patient outcomes by directly subtyping NSCLC from serum.
[0164] Figure 11 provides an illustrative diagram of the application of triplet margin loss to simulated data. The left panel shows label-independent embeddings, and the right panel shows embeddings with triplet margin loss constraints to minimize technical variability while preserving biological differences. For each sample, a positive anchor (same phenotype, different dataset) is sampled to minimize the embedding distance, and a negative anchor (different phenotype, arbitrary dataset) is sampled to maximize the embedding distance.
[0165] Figure 12 provides an exemplary loss convergence plot showing the convergence of the reconstruction loss, KL divergence loss based on the latent variable z(231), cross-entropy loss, triplet margin loss, KL divergence loss based on the latent variable l(232), and the classification accuracy during training, as described in relation to Figure 2.
[0166] Figures 13A–14B illustrate the exemplary performance of the neural network models for cancer detection and subtype classification described in Figures 1–10. To evaluate the performance of the neural network models for cancer detection and subtype classification, the dataset can be split into a 20% holdout and the remaining 80%. The neural network models for cancer detection and subtype classification can be trained on 80% of this data using a non-overlapping 10-fold cross-validation configuration. In each split, oncRNAs from one subset of TCGA within its training set were enriched in those cancer samples compared to control samples from each data source provider, resulting in a mean of 6,376 ± 60 (standard deviation) of oncRNAs per split. Five neural network models for cancer detection and subtype classification using different random seeds were trained for each split, and the scores for this test set were averaged.
[0167] As shown in Figure 13A, the neural network model for cancer detection and subtype classification (referred to as "Orion" in Figure 13A) achieved an area under the receiver operational characteristic curve (ROC) of 0.97 (95% confidence interval 0.96~0.98) and an overall sensitivity of 92% (88%~95%) at 90% specificity. A commonly used elastic network model, using the same set of oncRNAs for each training partition and identical configuration, had an ROC area of 0.84 (0.81~0.86) and an overall sensitivity of 49% (44%~55%). Other methods (e.g., XGBoost, k-nearest neighbor classifiers, and support vector machine classifiers) also performed worse than the neural network model for cancer detection and subtype classification. More importantly, the sensitivity (n=88) for Stage I, with 90% specificity, was 0.9 (0.83–0.94) for the neural network model for cancer detection and subtype classification, compared to 0.4 (0.31–0.49) for the elastic network model.
[0168] Figure 13B shows the performance metrics for binary classification on this holdout validation set. Calculating all threshold-dependent metrics (all except area under ROC) based on their cutoffs yields a 90% specificity on this 10-fold cross-validated training dataset. Bar heights represent point estimates of area under ROC, F1 score, Matthews correlation coefficient (MCC), sensitivity, and specificity. To assess the generalizability of the neural network model for cancer detection and subtype classification, a cutoff corresponding to 90% specificity was selected in this 10-fold cross-validation, and various classification metrics were measured on the holdout validation set. The neural network model for cancer detection and subtype classification showed high agreement in performance on this holdout validation set, while the performance of XGBoost and ElasticNet was at the lower limit of this 10-fold cross-validation measurement. In bootstrap analysis, the AUC of the neural network model for cancer detection and subtype classification was also significantly higher than that of ElasticNet (Δ AUC =0.13 (95% confidence interval: 0.11~0.16). The AUC of Orion and the AUC of XGBoost were relatively similar (Δ AUC =0.04 (0.03~0.05)), but the F1 score, the sensitivity of the neural network model for detecting and classifying this cancer at 90% specificity, and the generalizability to the validation set were also better for Orion compared to XGBoost (Δ F 1 = 0.07 (0.04 ~ 0.1), Δ 感度 = 0.12 (0.08~0.16).
[0169] As shown in Figure 14A, the sensitivity for later stages (II, III, and IV, n=243) was 0.97 (0.93~0.99) for the neural network model for cancer detection and subtype classification, and 0.55 (0.48~0.62) for the elastic network model. When detecting tumors smaller than 2 cm (T1a~b, n=52), the neural network model for cancer detection and subtype classification achieved a sensitivity of 0.87 (0.74~0.94) at 90% specificity. On the other hand, the elastic network model had a sensitivity of 0.31 (0.19~0.59) at 90% specificity.
[0170] As a measure of successful batch effect removal, these model scores for the control sample are expected to be similar and therefore indistinguishable from the sample supplier. The neural network model for cancer detection and subtype classification had an ROC area of 0.53 (0.47-0.58), suggesting that the model successfully removed the supplier's influence, while XGBoost had a larger ROC area of 0.62 (0.57-0.670) and ElasticNet had a larger ROC area of 0.59 (0.54-0.64).
[0171] Considering that the control sample in the cohort showed overrepresentation of individuals without a smoking history compared to the cancer sample (54% vs. 10%), we investigated the effect of smoking status on the model score. In the control sample, the neural network model validation set score for cancer detection and subtype classification had a ROC area of 0.6 (0.5-0.7) with respect to the presence of a smoking history, further confirming that there was little variation in the model score between individuals with and without a smoking history.
[0172] To identify the most important oncRNAs for the neural network model of cancer detection and subtype classification, Shapley Additive exPlanations (SHAP) (Lundberg et al., a unified approach to interpreting model predictions, Advances in Neural Information Processing Systems, volume 30, Curran Associates, Inc., 2017) is used to average values between model partitions. As shown in Figure 14B, among the high SHAP oncRNAs for this model, duplication or proximity of oncRNAs to some genes that are important in the pathogenesis and prognosis of lung cancer is observed.These included SOX2-OT (Dodangeh et al., Long non-coding RNA SOX2-OT enhances cancer biological traits via sponging to tumor suppressor mir-122-3p and mir-194-5p in non-small cell lung carcinoma, Scientific Reports, 13(1):12371, 2023), HSP90AA1 (Niu et al., Targeting hsp90 inhibits proliferation and induces apoptosis through akt1 / erk pathway in lung cancer, Frontiers in Pharmacology, 12:724192, 2022; Bhattacharyya et al., CDK1 and HSP90AA1 appear as the novel regulatory genes in non-small cell lung cancer: a bioinformatics approach, Journal of Personalized Medicine, 12(3):393, 2022), and FZD2 (Tuluhong et al., Fzd2 promotes tgf-β-induced epithelial-to-mesenchymal transition in breast cancer via activating notch signaling pathway, Cancer Cell International, 21:1-13, 2021).
[0173] Figures 15A and 15B provide illustrative performance charts showing the ability to classify tumor subtypes from circulating oncRNA. Understanding the histology of patients with NSCLC, in addition to early detection of cancer signals, has a significant impact on treatment selection and resistance mechanisms. Squamous cell carcinoma transformation of lung adenocarcinoma has been reported to occur after resistance to targeted therapy. Squamous cell carcinoma transformation has been reported as one of the mechanisms of acquired resistance to epidermal growth factor receptor (EGFR). Conventional methods for stratifying patients to evaluate squamous cell carcinoma transformation involve repeated biopsies of lung cancer patients, which can result in serious side effects (e.g., pneumothorax, hemorrhage, and air embolism).
[0174] Considering the tissue-specific nature of chromatin accessibility in various cancers, oncRNA expression patterns are specific to cancer type and subtype, allowing this model to non-invasively detect primary tissue from blood among various cancer types. We hypothesize that the biological differences between lung adenocarcinoma and squamous cell carcinoma are also reflected in serum oncRNA content, enabling this model to distinguish these major subtypes of NSCLC. While tumor tissue is highly different from normal tissue, the differences between given tumor subtypes are far less substantial. In NSCLC, for example, the pathologist agreement for various subtypes is close to 0.81. Consequently, subtype prediction in tumor histology is more difficult than cancer detection.
[0175] To evaluate this hypothesis, we will examine the potential to distinguish between two major NSCLC subtypes (adenocarcinoma and squamous cell carcinoma) using oncRNA in the blood. For this analysis, we will perform 20-fold cross-validation to adjust for reduced sample size, considering that this is an NSCLC-specific task.
[0176] Figure 15A shows Orion's ROC plot for distinguishing squamous cell carcinoma and adenocarcinoma in stage III / IV NSCLC samples. Figure 15B shows Orion's confusion matrix for subtype prediction at the 70% specificity cutoff. For later-stage tumors (stage III / IV), this cancer subtype classification model achieved an ROC area of 0.75 (95% confidence interval: 0.67~0.83) and a sensitivity of 0.71 (95% confidence interval: 0.56~0.84) at 70% specificity in distinguishing squamous cell carcinoma and adenocarcinoma samples in serum samples, as shown in Figures 15A and 15B.
[0177] This specification and accompanying drawings illustrating aspects, embodiments, practices, or applications of the present invention should not be construed as limitations. Various mechanical, compositional, structural, electronic, and operational modifications may be made without departing from the spirit and scope of this specification and the claims. In some cases, well-known circuits, structures, or techniques have not been illustrated or described in detail so as not to obscure the embodiments of this disclosure. Similar figures in two or more figures represent the same or similar elements.
[0178] This specification provides specific details describing certain embodiments consistent with the disclosure. Many specific details are provided to provide a complete understanding of those embodiments. However, it will be apparent to those skilled in the art that some embodiments can be carried out without using some or all of these specific details. The specific embodiments disclosed herein are intended to be illustrative, not limiting. Those skilled in the art will understand other elements that are not specifically described herein but are within the scope and spirit of the disclosure. Furthermore, to avoid unnecessary repetition, one or more features shown and described in relation to a particular embodiment may be incorporated into other embodiments without being specifically stated as not being incorporated, unless such one or more features would render the embodiment non-functional.
[0179] While exemplary embodiments have been shown and described, extensive modifications, alterations, and substitutions are intended in the above disclosure, and in some cases, some features of these embodiments may be used without corresponding use of other features. Those skilled in the art will recognize many variations, substitutions, and alterations. Therefore, the scope of the invention should be limited only by the appended claims, and it is appropriate that the claims be interpreted broadly in a manner consistent with the scope of the embodiments disclosed herein.
Claims
1. A method for generating cancer diagnosis predictions using a neural network-based model implemented on one or more hardware processors, The process involves receiving multiple samples of orphan noncoding ribonucleic acid (oncRNA) count data via a communication interface; The process of converting samples of oncRNA count data into latent variables in latent space using an encoder; and The process of generating a cancer diagnosis prediction based on the latent variables using a decoder. Methods that include...
2. The method according to claim 1, wherein the plurality of samples of oncRNA count data are received from one or more data sources.
3. The encoder, A first variational autoencoder (VAE) that encodes a first sample associated with a ribonucleic acid (RNA) subtype used for classification; and A second VAE encodes a second sample related to the endogenous high-expression RNA biotype used for library evaluation. The method according to claim 1, comprising one or more of the above.
4. The aforementioned cancer diagnosis prediction, The presence of cancer; Nuclear power plant organizations; and Cancer subtypes The method according to claim 1, comprising any of the above.
5. The process of receiving training samples of oncRNA count data during a training epoch; A step of sampling a positive sample having the same label as the training sample for oncRNA count data; A step of sampling negative samples having a different label from the training samples of oncRNA count data; and A step of calculating a first loss based on the distance metric between the training sample, the positive sample, and the negative sample in the latent space. The method according to claim 1, further comprising:
6. A step of calculating a second loss based on the Kullback-Leibler divergence between the conditional distribution of the encoded latent variable of the training sample conditioned by the training sample and the prior distribution of the encoded latent variable. The method according to claim 5, further comprising:
7. The decoder generates a reconstructed distribution of the training samples from the encoded latent variables; and A step of calculating a third loss based on the reconstructed distribution. The method according to claim 6, further comprising:
8. The process of generating the predicted classification of the training samples from the encoded latent variables using the classification head of the decoder; and A fourth loss is calculated as the cross-entropy between the predicted classification and the annotated labels of the training samples. The method according to claim 7, further comprising:
9. A step of training the encoder and the decoder based on a joint loss which is a weighted sum of the first loss, the second loss, the third loss, and the fourth loss. The method according to claim 8, further comprising:
10. A step of training the encoder in a first training stage based at least in part on the first loss; and A step of training the encoder and decoder in a second training stage following the first training stage, based at least in part on the second loss or the fourth loss. The method according to claim 8, further comprising:
11. The method according to claim 1, wherein the sample of oncRNA count data is associated with a lung cancer sample, and the cancer diagnosis prediction includes the prediction of the presence of lung cancer and the prediction of lung cancer subtypes, namely adenocarcinoma and squamous cell carcinoma.
12. A system that generates cancer diagnosis predictions using a neural network-based model, A communication interface for receiving multiple samples of orphan noncoding ribonucleic acid (oncRNA) count data; A memory for storing the neural network-based model and a plurality of processor-executable instructions; and One or more processors that execute the plurality of processor-executable instructions in order to carry out the operation, wherein the operation is The process of converting samples of oncRNA count data into latent variables in latent space using an encoder; and The process of generating a cancer diagnosis prediction based on the latent variables using a decoder. including one or more processors A system that includes this.
13. The system according to claim 12, wherein the plurality of samples of oncRNA count data are received from one or more data sources.
14. The encoder, A first variational autoencoder (VAE) that encodes a first sample associated with a ribonucleic acid (RNA) subtype used for classification; and A second VAE encodes a second sample related to the endogenous high-expression RNA biotype used for library evaluation. The system according to claim 12, comprising one or more of the above.
15. The aforementioned cancer diagnosis prediction, The presence of cancer; Nuclear power plant organizations; and Cancer subtypes The system according to claim 12, comprising any of the above.
16. The aforementioned operation, The process of receiving training samples of oncRNA count data during a training epoch; A step of sampling a positive sample having the same label as the training sample for oncRNA count data; A step of sampling negative samples having a different label from the training samples of oncRNA count data; and A step of calculating a first loss based on the distance metric between the training sample, the positive sample, and the negative sample in the latent space; A step of calculating a second loss based on the Kullback-Leibler divergence between the conditional distribution of the encoded latent variable of the training sample conditioned by the training sample and the prior distribution of the encoded latent variable. The decoder generates a reconstructed distribution of the training samples from the encoded latent variables; A step of calculating a third loss based on the reconstructed distribution. The process of generating the predicted classification of the training samples from the encoded latent variables using the classification head of the decoder; and A fourth loss is calculated as the cross-entropy between the predicted classification and the annotated labels of the training samples. The system according to claim 12, further comprising:
17. The aforementioned operation, A step of training the encoder and the decoder based on a joint loss which is a weighted sum of the first loss, the second loss, the third loss, and the fourth loss. The system according to claim 16, further comprising:
18. The system according to claim 12, wherein the sample of oncRNA count data is associated with a lung cancer sample, and the cancer diagnosis prediction includes predicting the presence of lung cancer and predicting lung cancer subtypes, namely adenocarcinoma and squamous cell carcinoma.
19. A method for subtype classification of lung cancer samples using a neural network-based model implemented on one or more hardware processors, A process of receiving input oncRNA count data related to lung cancer samples obtained from a target via a communication interface; A step of converting the input oncRNA count data into latent variables in the latent space using an encoder; and The process involves a decoder generating a cancer diagnosis prediction based on the latent variables, which includes a first prediction regarding the presence of lung cancer and a second prediction regarding whether the subtype of the lung cancer sample is adenocarcinoma or squamous cell carcinoma. Methods that include...
20. A method for diagnosing and predicting the treatment of cancer using a neural network-based model implemented on one or more hardware processors, A step of generating a cancer diagnosis prediction based on latent variables using a neural network-based model that converts oncRNA count data into latent variables; and The process of generating recommended treatment when the cancer diagnosis prediction indicates the presence of cancer. Methods that include...