Methods for developing cancer diagnostic models and their use in the development of cancer detection methods
A diagnostic model using miRNA expression profiles from multiple cancer types effectively detects a range of cancers with high sensitivity and specificity, addressing the limitations of current screening methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2026-04-03
AI Technical Summary
Current cancer screening methods are limited, with only four types covered by guidelines, leaving two-thirds of cancer diagnoses and 70% of cancer deaths uncovered, and there is a need for low-cost, non-invasive tests for early detection of multiple cancers.
A method involving constructing a training set with miRNA expression profiles from multiple cancer types and developing a diagnostic model using statistical models to rank miRNAs, calculating a diagnostic index, and classifying subjects based on this index.
The method achieves high sensitivity and specificity in detecting various cancers, including lung, colorectal, esophageal, gastric, liver, ovarian, pancreatic, biliary tract, bladder, glioma, and prostate cancers, with sensitivity ranging from 0.75 to 0.95 and specificity of 0.99 or higher.
Smart Images

Figure 2026510430000019 
Figure 2026510430000020 
Figure 2026510430000021
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims the interests of U.S. Provisional Patent Application No. 63 / 451,751, filed on 13 March 2023, and U.S. Provisional Patent Application No. 63 / 538,292, filed on 14 September 2023, the disclosures thereof incorporated herein by reference in their entirety.
[0002] Reference to electronically submitted sequence listings The contents of the sequence listing filed and submitted electronically with this specification, filed as Top300list.txt, size 172,425 bytes, and created on March 12, 2024, are incorporated herein by reference in their entirety.
[0003] The present invention generally relates to the technology of disease screening, detection and diagnosis, and more specifically to methods, kits, systems and non-temporary storage media for detecting one or more human cancers. [Background technology]
[0004] Cancer is a difficult and deadly disease for humankind. It is well recognized that early detection (or diagnosis) of cancer can significantly increase the survival rate of cancer patients. However, early-stage cancer patients are often asymptomatic, and in reality, these patients are often not diagnosed.
[0005] Currently, there are only four types of cancer for which recommended screening methods exist: breast cancer, colorectal cancer, lung cancer, and cervical cancer. It is estimated that two-thirds of cancer diagnoses and 70% of cancer deaths are not covered by existing cancer screening guidelines. Therefore, there is an urgent need to develop low-cost, ideally non-invasive or minimally invasive tests that enable the early detection of multiple cancers, i.e., multi-cancer early detection (MCED).
[0006] MicroRNAs (i.e., miRNAs) are small, single-stranded, non-coding RNA molecules averaging 22 nucleotides in length, encoded by corresponding genes in the human genome. These miRNAs are estimated to function by negatively regulating gene expression for over 50% of human genes. Abnormal miRNA expression is associated with many human cancers. Of particular note are miRNAs, which are extracellular circulating molecules released into circulation by tumor cells, are abundant and often stable in blood samples from cancer patients, and therefore miRNAs have potential as non-invasive biomarkers for cancer screening and diagnosis. Previously, a 4-miRNA diagnostic model was developed from a training set containing only lung cancer patients and matched non-cancer controls using publicly available serum miRNA microarray datasets. This model demonstrated approximately 80–100% sensitivity for 10 cancers and approximately 70% sensitivity for sarcomas and ovarian cancers, while maintaining approximately 99% specificity (Zhang A, et al. 2022). [Overview of the Initiative]
[0007] In a first aspect, the disclosure provides a method for developing a diagnostic model for detecting a cancer of interest (i.e., a method for developing a cancer diagnostic model).
[0008] The method includes (1) constructing a training set containing expression profiles of multiple miRNAs obtained from non-cancer subjects and cancer patients, and (2) developing a diagnostic model based on the training set. Here, cancer patients are configured to have at least two types of cancer, and step (2) further includes (a) a substep of performing differential expression analysis on the expression profiles of multiple miRNAs by a first statistical model so that the multiple miRNAs are ranked by adjusted p-values, and (b) a substep of constructing a diagnostic model based on a selected set of miRNA biomarkers from the multiple miRNAs. Here, the selected set of miRNA biomarkers contains at least one miRNA selected such that each has a ranking less than or equal to a predetermined cutoff m (m≧1).
[0009] Herein, according to some embodiments, the first statistical model can be selected from one of the following models, including but not limited to linear models (limma) of microarray data, logistic regression models, linear discriminant analysis (LDA) models, conditional logistic regression models, lasso regression models, ridge regression models, random forests, support vector machines, or probit regression models. According to some embodiments, a limma model is used as the first statistical model.
[0010] In a particular embodiment of the method for developing a cancer diagnostic model, substep (b) is: This involves calculating a diagnostic index based on the expression profile of a selected set of miRNA biomarkers, the diagnostic index being calculated using the formula:
number
[0011] Here, according to some embodiments, i t is a constant, and therefore the diagnostic index takes an unweighted approach. According to some other embodiments, the weight t i This is based on a second statistical model. Similar to the first statistical model, the second statistical model can also be selected from one of the following models, including but not limited to linear models (limma) of microarray data, logistic regression models, linear discriminant analysis (LDA) models, conditional logistic regression models, lasso regression models, ridge regression models, random forests, support vector machines, or probit regression models.
[0012] Here, the first and second statistical models may be different or substantially the same, at the discretion of the user. Furthermore, according to some embodiments, the limma model is used as both the first and second statistical models.
[0013] According to some embodiments, at least one miRNA in the selected miRNA biomarker set is a top n-position miRNA.
[0014] In the cancer diagnostic model development method, step (2) may optionally further include a substep (c) evaluating the performance of the diagnostic model through cross-validation on a training set. Here, cross-validation can be performed in at least two divisions (e.g., 2, 3, 5, 10 divisions). Optionally, substep (c) may further include at least one of the following: calculating the area under the curve (AUC) of the receiver operating characteristic (ROC) curve and evaluating the performance of the diagnostic model based thereon, or calculating the specificity and sensitivity of the diagnostic model and evaluating the performance of the diagnostic model based thereon. Note that, in addition to AUC and specificity / sensitivity, other parameters such as overall accuracy, positive likelihood ratio, negative likelihood ratio, positive predictor, negative predictor, and Euden's J may also be used to evaluate the performance of the diagnostic model in substep (c) of step (2).
[0015] In certain embodiments, the cancer diagnostic model development method may further include the step of (3) validating the diagnostic model based on a validation set, where the validation set includes expression profiles of a selected set of miRNA biomarkers obtained from non-cancer subjects and cancer patients with the cancer of interest. Optionally, step (3) further includes at least one of calculating the AUC of the ROC curve and evaluating the performance of the diagnostic model based thereon, or calculating the specificity and sensitivity of the diagnostic model and evaluating the performance of the diagnostic model based thereon. Note that, in addition to AUC and specificity / sensitivity, other parameters such as overall precision, positive likelihood ratio, negative likelihood ratio, positive predictor, negative predictor, and Euden's J may also be used to validate the diagnostic model in step (3) above.
[0016] A system for developing a diagnostic model for detecting a target cancer is further provided, the system comprising a processor and a non-temporary storage medium containing program instructions for execution by the processor. Here, the program instructions are configured to cause the processor to perform various steps of a cancer diagnostic model development method according to any embodiment described above in the first aspect.
[0017] A non-transitory storage medium configured to store computer-executable program instructions, wherein the program instructions cause a processor to execute, when executed by the processor, various steps of a cancer diagnosis model development method according to any of the embodiments as described above in a first aspect. A non-transitory storage medium is also provided.
[0018] In a second aspect, the present disclosure further provides a method for detecting a target cancer from a subject using a diagnosis model. Here, the diagnosis model is developed by a cancer diagnosis model development method according to any of the embodiments described above.
[0019] According to some embodiments, the cancer detection method includes the following three steps:
[0020] (A) Determining an expression profile of a selected miRNA biomarker set from a biological sample obtained from a subject;
[0021] (B) Calculating a diagnosis index of the biological sample based on the expression profile of the selected miRNA biomarker set, wherein the diagnosis index is calculated based on the formula:
Number
[0022] (C) Classifying the subject as having or not having the target cancer based on the calculated diagnosis index, wherein the subject is classified as having the target cancer if the calculated diagnosis index is greater than or equal to a predetermined threshold, and is classified as not having the target cancer otherwise. The step of classifying.
[0023] Here, each miRNA in the arbitrarily selected miRNA biomarker set is from the top 200 miRNAs listed in Table 1, and n is 200 or less.
[0024] According to some embodiments of cancer detection methods, the diagnostic index is calculated using weights from a limma model.
[0025] Here, according to some embodiments of the cancer detection method, at least one miRNA in the selected miRNA biomarker set is a top n-position miRNA.
[0026] According to some embodiments, n is between 4 and 200, and the classification can achieve an AUC of more than approximately 0.97 to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0027] According to several embodiments, n is between 4 and 200, and the classification can achieve a sensitivity of at least about 0.75 while maintaining a specificity of at least about 0.99 to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, according to several embodiments, the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity of at least about 0.99 to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, according to several embodiments, the classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity of at least about 0.99 to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, according to some embodiments, the classification can achieve a sensitivity of at least about 0.95 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, esophageal cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0028] According to several embodiments of the cancer detection method, the selected miRNA biomarker set consists of hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, and hsa-miR-663a. Furthermore, according to several embodiments, the classification can achieve a sensitivity of at least about 0.75 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, according to several embodiments, the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity value of about 0.99 for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, according to some embodiments, the classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity value of about 0.99 for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, according to some embodiments, the classification can achieve a sensitivity of at least about 0.95 while maintaining a specificity value of about 0.99 for detecting lung cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, according to some embodiments, the classification can achieve a sensitivity of at least about 0.99 while maintaining a specificity value of about 0.99 for detecting lung cancer, gastric cancer, biliary tract cancer, or bladder cancer.
[0029] According to several embodiments of the cancer detection method, the selected miRNA biomarker set consists of hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-663a, and hsa-miR-320a, and the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity value of about 0.99 for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0030] According to several embodiments of the cancer detection method, the selected set of miRNA biomarkers consists of the top 10 miRNAs in Table 1, and the classification can achieve a sensitivity of at least approximately 0.99 while maintaining a specificity value of approximately 0.99 for detecting gastric cancer, esophageal cancer, biliary tract cancer, or prostate cancer.
[0031] According to several embodiments of the cancer detection method, the selected set of miRNA biomarkers consists of the top 15 miRNAs in Table 1, and the classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity value of about 0.99 for detecting lung cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0032] In any embodiment of the cancer detection method described above, the expression profile of the selected miRNA biomarker set is obtained by at least one of Northern blotting, microarray analysis, RNA sequencing or RNA in situ hybridization, or a nucleic acid amplification procedure, the nucleic acid amplification procedure comprising at least one of reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR.
[0033] In any embodiment of the cancer detection method described above, the biological sample is a liquid biopsy sample selected from the group consisting of blood samples, serum samples, plasma samples, urine samples, saliva samples, and sputum samples.
[0034] A system for detecting a target cancer from a subject is also provided, the system comprising a processor and a non-temporary storage medium containing program instructions for execution by the processor, wherein the program instructions are configured to cause the processor to perform various steps of a cancer detection method according to any embodiment described above in a second aspect.
[0035] Also provided is a non-temporary storage medium configured to store computer executable program instructions, wherein, when executed by a processor, the program instructions cause the processor to perform various steps of a cancer detection method according to any embodiment described above in a second aspect.
[0036] Unless otherwise defined, terms used throughout this disclosure are defined as follows:
[0037] Generally, the term “subject” refers to mammals such as humans and primates including chimpanzees, pet animals including dogs and cats, domestic animals including cattle, horses, sheep and goats, and rodents including mice and rats. The term “healthy subject” refers to such mammals that do not have detectable cancer. Note that this entire disclosure relates more specifically to human subjects, but may be applied to other non-human mammals at the discretion of the individual.
[0038] Unless otherwise indicated or defined, terms or abbreviations such as “nucleic acid,” “nucleotide,” “polynucleotide,” “DNA,” “RNA,” and “miRNA” shall be used in accordance with common usage in the art.
[0039] As used herein, the term “polynucleotide” is interchangeable with “nucleic acid” and refers to nucleic acids including all RNA, DNA, and RNA / DNA (chimeras). DNA includes all cDNA, genomic DNA, and synthetic DNA. RNA includes all total RNA, mRNA, rRNA, miRNA, siRNA, snoRNA, snRNA, non-coding RNA, and synthetic RNA.
[0040] As used herein, the term “fragment” means a polynucleotide having a base sequence that comprises a contiguous portion of a polynucleotide, preferably having a length of 15 or more nucleotides, for example, 15, 16, 17, 18, or 19 nucleotides.
[0041] As used herein, the term “gene” is intended to include not only RNA and double-stranded DNA, but also each single-stranded DNA, such as the positive-sense strand (or sense strand) or complementary strand (or antisense strand) that constitutes a double helix. Genes are not particularly limited by their length. As used herein, “gene” includes, unless otherwise specified, double-stranded DNA including human genomic DNA, single-stranded DNA (positive-sense strand) including cDNA, single-stranded DNA (complementary strand) having a sequence complementary to the positive-sense strand, miRNA (miRNA) and its fragments, and all of their transcripts. “Genes” include not only “genes” represented by a specific base sequence (or sequence number), but also “nucleic acids” that have equivalent biological function to the RNA encoded by that gene, such as homologs (i.e., homologs or orthologues), variants (e.g., genetic polymorphs), and derivatives. Specific examples of "nucleic acids" encoding such congeners, variants, or derivatives include "nucleic acids" having a base sequence that hybridizes under the stringent conditions described later to a complementary sequence of a base sequence derived from any of the base sequences represented by sequence numbers 1 to 200, or a base sequence in which nucleotide "U" (or "u") is replaced with nucleotide "T" (or "t"). A "gene" is not particularly limited by its functional region and may include, for example, an expression regulatory region, a coding region, an exon, or an intron. A "gene" may be contained within a cell, released outside the cell and existing independently, or a "gene" may be encapsulated in a vesicle called an exosome.
[0042] Throughout this disclosure, the term “microRNA (miRNA)” is intended, unless otherwise specified, to mean a 15-25 nucleotide non-coding RNA that is transcribed as a hairpin-like RNA precursor, cleaved by a dsRNA-cleaving enzyme with RNase III cleavage activity, incorporated into a protein complex called RISC, and involved in the repression of mRNA translation. As used herein, the term “miRNA” includes not only “miRNA” represented by a specific base sequence (or sequence number), but also its precursors (pre-miRNA or pri-miRNA), and miRNAs with equivalent biological functions, such as homologs (i.e., orthologs), variants (e.g., genetic polymorphs), and derivatives. Such precursors, homologs, variants, or derivatives can be specifically identified using miRBase Release 20 (Kozomara and Griffiths-Jones, 2010). Examples of such precursors, congeners, variants, or derivatives include "miRNAs" having a nucleotide sequence that hybridizes to a complementary sequence of a specific nucleotide sequence represented by any of SEQ ID NOs: 1 to 200 under the stringent conditions described below. As used herein, the term "miRNA" may refer to the gene product of a miR gene. Such gene products include mature miRNAs (e.g., 15-25 nucleotide or 19-25 nucleotide non-coding RNAs involved in the repression of mRNA translation as described above) or miRNA precursors (e.g., the pre-miRNA or pri-miRNA described above).
[0043] As used herein, the term "probe" includes polynucleotides used to specifically detect RNA or polynucleotides derived from RNA resulting from gene expression, and / or polynucleotides complementary thereto.
[0044] As used herein, the terms “primer” or “amplification primer” include polynucleotides that specifically recognize and amplify RNA or polynucleotides derived from RNA resulting from gene expression, and / or polynucleotides complementary thereto.
[0045] In this context, a complementary polynucleotide (complementary or reverse chain) means a polynucleotide that has a complementary base relationship based on A:T(U) and G:C base pairs with a polynucleotide consisting of a full-length sequence or partial sequence thereof (where this full-length or partial sequence is conveniently referred to as the plus-chain) that is derived from a base sequence defined in any of Sequence IDs 1 to 200, or a base sequence in which nucleotide "U" (or "u") is replaced by nucleotide "T" (or "t"). However, such a complementary chain is not limited to a sequence that is completely complementary to the base sequence of the target plus-chain, but may have a complementary relationship to the extent that it enables hybridization to the target plus-chain under stringent conditions.
[0046] As used herein, the term “stringent conditions” refers to conditions under which a nucleic acid probe hybridizes to its target sequence to a greater extent than other sequences (e.g., a measurement of the mean of background measurements plus two times the standard deviation of background measurements). Stringent conditions are sequence-dependent and vary depending on the hybridization environment. Target sequences that are 100% complementary to a nucleic acid probe can be identified by controlling the stringency of the hybridization and / or washing conditions. Specific examples of “stringent conditions” are described below.
[0047] As used herein, the term "Tm value" means the temperature at which the double-stranded portion of a polynucleotide is denatured into single strands such that double-stranded and single-stranded portions exist in a 1:1 ratio.
[0048] As used herein, the term "variant" means, in the case of nucleic acids, a natural variant resulting from polymorphism, mutation, etc.; a variant having a deletion, substitution, addition, or insertion of one, two, or three or more nucleotides in a subsequence of a nucleotide sequence derived from any of the sequence codes 1 to 200, or a nucleotide sequence in which nucleotide "U" (or "u") is replaced by nucleotide "T" (or "t"); a variant containing a deletion, substitution, addition, or insertion of one or more nucleotides in a subsequence of a premature miRNA; a variant showing approximately 90% or more identity percentage, approximately 95% or more identity percentage, to each of these nucleotide sequences or subsequences; or a nucleic acid that hybridizes to a polynucleotide or oligonucleotide containing these nucleotide sequences or subsequences under the stringent conditions defined above. Variants can be prepared using well-known techniques such as site-directed mutagenesis or PCR-based mutagenesis.
[0049] The term "identity percentage (%)" can be determined by the presence or absence of introduced gaps using the BLAST or FASTA-based protein or gene search systems described above (Zhang et al., 2000; Altschul et al. 1990; Pearson et al. 1988).
[0050] The term "derivative" includes, but is not limited to, modified nucleic acids, such as derivatives labeled with fluorophores, modified nucleotides (for example, nucleotides containing groups such as halogens, alkyls such as methyl, alkoxys such as methoxy, thio or carboxymethyl, and nucleotides that have undergone rearrangement of bases, saturation of double bonds, deamination, substitution of oxygen molecules with sulfur atoms, etc.), PNA (peptide nucleic acid; Nielsen et al. 1991), and LNA (locked nucleic acid; Obika et al. 1998).
[0051] The "nucleic acid" that can specifically bind to the polynucleotide selected from the above miRNAs is a synthesized or prepared nucleic acid, specifically, a "nucleic acid probe" or "primer." The "nucleic acid" is used directly or indirectly to detect the presence or absence of cancer in a subject, to diagnose the severity, degree of improvement, or sensitivity of treatment of cancer, or to screen for candidate substances useful for the prevention, improvement, or treatment of cancer. The "nucleic acid" includes nucleotides, oligonucleotides, and polynucleotides that, in relation to the development of cancer, can specifically recognize and bind to the transcript represented by any of SEQ ID NOs: 1 to 200 or its synthetic cDNA nucleic acid in a living organism, particularly in samples such as bodily fluids (e.g., blood or urine). Based on the above properties, nucleotides, oligonucleotides, and polynucleotides can be effectively used as probes for detecting the above genes expressed in living organisms, tissues, or cells, or as primers for amplifying the above genes expressed in living organisms.
[0052] As used herein, the term “detection” is interchangeable with the terms “examination,” “measurement,” or “detection or decision support.” As used herein, the term “assessment” includes diagnostic or assessment support based on examination or measurement results.
[0053] When used within the scope of this disclosure, the terms “P-value,” “accuracy,” “AUC,” “sensitivity,” and “specificity” should be understood to have common definitions that are generally well understood by those skilled in the art, and specifically are defined as follows:
[0054] The term "p-value" or "p" refers to the probability that a statistical test will observe a more extreme statistic than what is actually calculated from the data under the null hypothesis. Therefore, a smaller "p" or "p-value" indicates a greater statistical significance between the subjects being compared.
[0055] The term "AUC" refers to the area under the receiver operating characteristic curve. The term "precision" refers to the value of (number of true positives + number of true negatives) / (total number of cases). Precision represents the proportion of correctly identified samples for all samples and serves as a primary indicator for evaluating detection performance.
[0056] As used herein, the term "sensitivity" refers to the value of (true positives) / (true positives + false negatives). High sensitivity enables the detection of cancer and leads to clinical therapeutic intervention.
[0057] As used herein, the term "specificity" refers to the value of (number of true negatives) / (number of true negatives + number of false positives). High specificity eliminates the need for unnecessary testing of healthy individuals who are misdiagnosed as cancer patients, leading to reduced patient burden and lower healthcare costs.
[0058] Unless otherwise specified, the following are the available techniques that may be used to determine the expression profiles of miRNA biomarker sets.
[0059] Furthermore, determining the expression profile of a miRNA biomarker set essentially involves determining the expression levels of all miRNAs included in that set. Preferably, the expression levels of all miRNAs included in the miRNA biomarker set can be determined simultaneously in a single, well-controlled experiment. Optionally, the expression levels of these miRNAs can also be determined in two or more experiments and using different experimental procedures.
[0060] As used herein, measuring or detecting the expression of any of the miRNAs included in a set of miRNA biomarkers includes measuring or detecting any nucleic acid transcript corresponding to the miRNA.
[0061] Typically, expression can be detected or measured based on miRNA or corresponding reverse-transcribed cDNA levels. Any quantitative or qualitative method can be used to measure RNA or cDNA levels. Suitable methods for detecting or measuring miRNA or cDNA levels include, for example, Northern blotting, microarray analysis, RNA sequencing, RNA in situ hybridization, or nucleic acid amplification procedures, such as real-time RT-PCR, also known as reverse transcription PCR (RT-PCR) or quantitative RT-PCR (qRT-PCR), or digital RT-PCR. Such methods are well known in the art (see, for example, Sambrook et al. 2012). Other techniques include digital multiplex analysis of gene expression, such as the nCounter® (NanoString Technologies, Seattle, Washington) gene expression assay, which is further described in US20100112710 and US20100047924.
[0062] The detection of a target nucleic acid generally involves hybridization between the target (e.g., miRNA or cDNA) and a probe. Sequences of miRNAs used in various oncogene expression profiles are known. Therefore, those skilled in the art can easily design hybridization probes for detecting these miRNAs (see, e.g., Sambrook et al. 2012). For example, polynucleotide probes that specifically bind to miRNA transcripts (or cDNA synthesized from miRNA transcripts) described herein can be prepared using the nucleic acid sequence of the miRNA or cDNA target itself by conventional techniques (e.g., PCR or synthesis). As used herein, the term “probe” means a portion or part of a polynucleotide sequence containing approximately 10 or more consecutive nucleotides, approximately 15 or more consecutive nucleotides, or approximately 20 or more consecutive nucleotides. In certain embodiments, a polynucleotide probe contains 10 or more nucleic acids, 15 or more nucleic acids, or 20 or more nucleic acids. To ensure sufficient specificity, when a probe is determined using, for example, the well-known Basic Local Alignment Search Tool (BLAST) algorithm (available from the National Center for Biotechnology Information (NCBI) in Bethesda, Maryland), it may have approximately 90% or more sequence identity with respect to the complementary strand of the target sequence, for example, approximately 95% or more (e.g., approximately 98% or more or approximately 99% or more).
[0063] Each probe may be substantially specific to its target to avoid any cross-hybridization and false positives. An alternative to using specific probes is to use specific reagents when deriving material from the transcript (e.g., when using target-specific primers during cDNA production or amplification). In either case, specificity may be achieved by hybridization to a subset of targets that is substantially unique within the group of miRNAs being analyzed; for example, hybridization to the polyA tail would not provide specificity. If the target has multiple splice variants, it is possible to design hybridization reagents that recognize a region common to each variant, and / or to use two or more reagents, each capable of recognizing one or more variants.
[0064] The stringency of a hybridization reaction is readily apparent to those skilled in the art and is generally determined by empirical calculations depending on the probe length, washing temperature, and salt concentration. Generally, longer probes may require higher temperatures for proper annealing, while shorter probes may require lower temperatures. Hybridization generally depends on the ability of the denatured nucleic acid sequence to re-anneal in the presence of a complementary strand in a sub-melting temperature environment. The higher the desired degree of homology between the probe and the hybridizable sequence, the higher the relative temperature that can be used. Consequently, higher relative temperatures tend to result in more stringent reaction conditions, while lower temperatures tend to have less so.
[0065] The terms “stringent conditions” or “high stringency conditions” as defined herein are not limited to, but include: (1) conditions in which washing is performed using low ionic strength and high temperature, e.g., 50°C with 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate; and (2) conditions in which a denaturing agent such as formamide is used during hybridization, e.g., 50% (v / v) formamide containing 0.1% bovine serum albumin, 0.1% Ficol, 0.1% polyvinylpyrrolidone, and 50 mM sodium phosphate buffer at pH 6.5 at 42°C, with 750 mM sodium chloride and 75 mM citrate. Conditions using sodium citrate together; or (3) conditions using 50% formamide, 5×SSC (0.75M NaCl, 0.075M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5× Denhardt's solution, sonicated salmon sperm DNA (50 μg / ml), 0.1% SDS, and 10% dextran sulfate at 42°C, followed by washing at 55°C in 0.2×SSC (sodium chloride / sodium citrate) and 50% formamide, and then a high-stringency wash with 0.1×SSC containing EDTA at 55°C. "Moderately stringent conditions" are, but are not limited to, those described by Sambrook et al. 1989 and include the use of less stringent washing solutions and hybridization conditions (e.g., temperature, ionic strength, and %SDS) than those described above. An example of moderately stringent conditions involves incubation overnight at 37°C in a solution containing 20% formamide, 5×SSC (150 mM NaCl, 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5× Denhardt's solution, 10% dextran sulfate, and 20 mg / mL of denatured shear salmon sperm DNA, followed by washing the filter in 1×SSC at approximately 37-50°C. Those skilled in the art are aware of methods for adjusting temperature, ionic strength, etc., as needed to accommodate factors such as probe length.
[0066] In certain embodiments, methods based on microarray analysis, Northern blotting, RNA in situ hybridization, or PCR are used. In this regard, measuring the expression of the aforementioned miRNAs in a biological sample may include, for example, contacting a sample containing or suspected of containing cancer cells with a polynucleotide probe specific to the miRNA of interest or a primer designed to amplify a portion of the miRNA of interest, and detecting the binding of the probe to a nucleic acid target or the amplification of the nucleic acid, respectively. Detailed protocols for designing PCR primers are known in the art (see, for example, Sambrook et al. 2012). In certain embodiments, the miRNA obtained from the sample may be subjected to qRT-PCR. Reverse transcription may be performed by any method known in the art, such as the use of the Omniscript RT Kit (Qiagen). The resulting cDNA may then be amplified by any amplification technique known in the art. Subsequently, miRNA expression may be analyzed by using a control sample, for example, as described below. As described herein, overexpression or underexpression of miRNAs compared to controls may be measured to reveal the miRNA expression profile for individual biological samples. Similarly, detailed protocols for preparing and using microarrays to analyze miRNA expression are known in the art and are described herein.
[0067] As used herein, RNA sequencing (RNA-seq), also known as whole transcriptome shotgun sequencing, refers to any of the various high-throughput sequencing techniques used to detect the presence and quantity of RNA transcripts in real time. See Wang, Z., M. Gerstein, and M. Snyder, RNA-Seq: a revolutionary tool for transcriptomics, NAT REV GENET, 2009.10(1):p.57-63. RNA-seq can be used to reveal a snapshot of genomic miRNAs in a sample at a given time point. In certain embodiments, miRNAs are converted to cDNA fragments via reverse transcription before sequencing, and in certain embodiments, miRNAs can be sequenced directly without conversion to cDNA. Adapters may be attached to the 5' and / or 3' ends of the miRNA, and the miRNA or cDNA may be optionally amplified, for example, by PCR. The fragments are then sequenced using high-throughput sequencing technologies, such as those available from Roche (e.g., the 454 platform), Illumina, Inc., and Applied Biosystem (e.g., the SOLiD system).
[0068] Please note the following.
[0069] The description of background art may contain information that may be useful in understanding the present invention. However, this does not constitute an endorsement that any information provided herein is prior art or relevant to the subject matter of this application, nor does it constitute an endorsement that any publication specifically or implicitly cited is prior art.
[0070] In interpreting both this specification and the claims, it should be further noted that all terms should be interpreted in the broadest possible way that is consistent with the context. In particular, the terms “comprises” and “comprising” should be interpreted as referring to elements, components, or steps in a non-exclusive manner, indicating that the referenced element, component, or step exists, is utilized, or can be combined with other elements, components, or steps not explicitly referenced. The meanings of “a,” “an,” and “the” include multiple references unless otherwise explicitly indicated by the context. Also, as used in this description, the meaning of “in” includes “in” and “on” unless otherwise explicitly indicated by the context. Numbers used to describe and claim specific embodiments of the present invention, such as quantities, concentrations, or other properties of components, or reaction conditions, should be understood to be modified in some cases by the term “approximately.” As used herein, terms such as “about,” “around,” and “approximately” mean, when referring to a specified measurable value (e.g., a parameter, quantity, duration, etc.), encompass the specified value and variations from that value, such as variations of + / - 20% or less, or + / - 10% or less, to the extent that such variations are appropriate to perform in the disclosed embodiments. Thus, the values themselves that the modifiers “about” or “approximately” refer to are also specifically disclosed. The enumeration of value ranges in this specification is intended merely as a convenient way to refer individually to each distinct value that falls within that range. The use of any examples or illustrative language (e.g., “etc.”) provided in reference to specific embodiments herein is intended merely to better illustrate the invention and not to limit the scope of the claimed invention. [Brief explanation of the drawing]
[0071] [Figure 1] A block diagram of a cancer diagnostic model development method provided by some embodiments of this disclosure is shown.
[0072] [Figure 2] This figure shows a computerized system according to some embodiments of the present disclosure.
[0073] [Figure 3A] This diagram illustrates the workflow for dataset and research design, showing the construction of training and validation datasets. [Figure 3B] This diagram shows the dataset and research design flow, illustrating the research design for model development and validation.
[0074] [Figure 4A] This diagram illustrates the development of a diagnostic model through cross-validation in a training set, showing the cross-validation process. [Figure 4B] This figure illustrates the development of diagnostic models through cross-validation in a training set, showing the performance of different diagnostic models using the top N miRNAs (N=1, 2, 3, ..., 40).
[0075] [Figure 5] This figure compares the diagnostic performance of different diagnostic models, which varies depending on the number of top miRNAs N (where N is between 4 and 200).
[0076] [Figure 6A] Figure 6B shows the diagnostic performance of the 4-miRNA model in a multi-cancer training set, with Figure 6B showing the ROC of the 4-miRNA model. [Figure 6B] Figure 6C shows the diagnostic performance of the 4-miRNA model in a multi-cancer training set, with Figure 6C being a scatter plot of the diagnostic index.
[0077] [Figure 7A] This figure shows the diagnostic performance of the 4-miRNA model in validation set 1 (lung cancer validation dataset), and displays the ROC of the 4-miRNA model. [Figure 7B]This figure shows the diagnostic performance of the 4-miRNA model in validation set 1 (lung cancer validation dataset), and displays a scatter plot of the diagnostic index. [Figure 7C] This figure shows the diagnostic performance of the 4-miRNA model in validation set 1 (lung cancer validation dataset), and displays a scatter plot of diagnostic indices from serum samples before and after surgery. [Figure 7D] This figure shows the diagnostic performance of the 4-miRNA model in validation set 1 (lung cancer validation dataset), and displays a scatter plot of diagnostic indices in clinical subsets. ADC: adenocarcinoma, SqCC: squamous cell carcinoma, LCC: large cell carcinoma, SCLC: small cell lung cancer.
[0078] [Figure 8A] This figure shows the diagnostic performance of the 4-miRNA model in validation sets 2 and 3, and the scatter plot of the diagnostic index in validation set 2 is shown. [Figure 8B] This figure shows the diagnostic performance of the 4-miRNA model in validation sets 2 and 3, and a scatter plot of the diagnostic index in validation set 3. [Modes for carrying out the invention]
[0079] Generally, when developing a diagnostic model to detect a specific disease (e.g., a specific type of cancer), it is considered necessary to use data specific to that particular disease (e.g., clinical data, molecular data, etc.) and to avoid using data from different diseases (e.g., different types of cancer). This is because data from different diseases (e.g., different types of cancer) introduces undesirable noise into the development of the diagnostic model, and therefore, the diagnostic model developed in this way does not function ideally.
[0080] To develop a cancer detection method, the inventors used serum miRNA microarray datasets obtained from multiple cancer types. Unexpectedly, the inventors found that including miRNA expression data from multiple cancer types led to the development of a diagnostic model that could sensitively and reliably detect a specific type of cancer. Briefly, in the inventors' study described in more detail in the example section below, the inventors' training set used miRNA expression datasets from a total of seven cancers (lung cancer, ovarian cancer, liver cancer, bladder cancer, esophageal cancer, gastric cancer, and prostate cancer), and surprisingly, the inventors' cross-validation using the training set and subsequent validation using different validation sets showed that the diagnostic model developed based on data from the seven cancers worked unexpectedly well not only for these seven cancers, but also, surprisingly, for most of the other cancers tested (biliary tract cancer, colorectal cancer, glioma, pancreatic cancer, and sarcoma).
[0081] Based on this research, a first aspect of this disclosure provides an approach to developing cancer diagnostic models.
[0082] This model development approach essentially involves developing a diagnostic model to detect a target cancer using data from multiple cancer types.
[0083] As shown in Figure 1, a method is provided for developing a diagnostic model for detecting a cancer of interest, provided according to some embodiments of the present disclosure, the method substantially comprising the following three main steps S100 to S300.
[0084] S100: Construct a training set containing miRNA expression profiles obtained from non-cancer subjects and patients with multiple cancers (two or more cancers).
[0085] S200: Develop a diagnostic model based on a training set, and
[0086] S300: Validate a diagnostic model based on a validation set that includes expression profiles of selected miRNA biomarker sets obtained from non-cancer subjects and cancer patients with the target cancer.
[0087] Step S100 involves constructing a training set to be used in Step S200 to build a diagnostic model. The training set includes multiple miRNA expression profiles (i.e., miRNA expression profiles) obtained from non-cancer subjects and multi-cancer patients, which will be used as controls and cases, respectively, when developing the cancer detection diagnostic model in Step S200.
[0088] Here, among patients with multiple cancers, there are at least two types of cancer. More specifically, among cancer patients, there are a total of N types of cancer (N≧2), that is, these patients with multiple cancers may include a first subset of patients having a first type of cancer, a second subset of patients having a second type of cancer, ..., an Nth subset of patients having an Nth type of cancer. According to some embodiments, at least two types of cancer may include the same type of cancer as the cancer of interest. According to some other embodiments, at least two types of cancer do not include the cancer of interest.
[0089] As used herein and throughout this disclosure, the term “cancer” means a specific “cancer type” as defined as a specific type of human disease characterized by an abnormal increase in the number of cells of a particular type (e.g., epithelial cells, connective tissue cells, blood cells, etc.) that may invade or spread to other parts of the body. Cancer types included herein may include, but are not limited to, carcinomas, sarcomas, lymphomas, germ cell tumors, blastomas, and may include both benign and malignant tumors. Non-limiting examples of cancer types include lung cancer, breast cancer, esophageal cancer, prostate cancer, stomach cancer, pancreatic cancer, liver cancer, ovarian cancer, colorectal cancer, biliary tract cancer, kidney cancer, bladder cancer, brain tumors (e.g., gliomas), sarcomas (e.g., osteosarcoma, chondrosarcoma, fibrosarcoma, etc.), leukemia, lymphoma, germ cell tumors (e.g., germ cell tumors, seminomas, etc.), hepatoblastoma, medulloblastoma, nephroblastoma, etc.
[0090] In an example for explanation, where the target cancer is pancreatic cancer, the cancer diagnostic model development method described above can be used to develop a diagnostic model that is specifically applicable to the detection of pancreatic cancer. Thus, miRNA expression profiles obtained from non-cancer subjects (i.e., controls) and multi-cancer patients (i.e., cases) can be utilized. Here, optionally, multi-cancer patients may include one subset of pancreatic cancer patients and at least one other subset of other cancer patients (e.g., lung cancer patients, glioma patients, etc.). Optionally, multi-cancer patients may include several subsets of cancer patients that are not pancreatic cancer patients (e.g., lung cancer patients, glioma patients, prostate cancer patients).
[0091] As used herein and throughout this disclosure, “expression profile” refers to a collection of expression information relating to a given set of biological molecules (e.g., miRNAs, proteins, DNA molecules, etc.), which may include expression data obtained for each member molecule in the given set of biological molecules. Various methods may exist for obtaining an expression profile. For example, a miRNA expression profile as described herein may optionally be obtained by Northern blotting, microarray analysis, RNA sequencing, or RNA in situ hybridization, or optionally by nucleic acid amplification procedures including reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR. A miRNA expression profile may be obtained directly from cancerous or tumor tissue from a patient with a specific type of cancer, or from some other tissue from the patient, such as blood, serum, plasma, saliva, or sweat. In the latter case, the miRNA thus assayed in the expression profile may exist as cell-free miRNA released into circulation from tumor cells. Regardless of the origin of the sample or the assay approach, the obtained miRNA expression profile should be interpreted as being included within this disclosure.
[0092] According to a particular embodiment, step S200, which develops a diagnostic model based on a training set, includes the following substeps:
[0093] S210: Differential expression analysis is performed on miRNA expression profiles in order to rank miRNAs by adjusted p-values, and
[0094] S220: Construct a diagnostic model based on a selected set of miRNA biomarkers.
[0095] Here, a substep S210 is performed to perform differential expression analysis on the miRNA expression profile, thereby ranking the miRNAs by adjusted p-values. Differential expression analysis is widely performed, for example, in the report by Zhang et al. 2014, which is incorporated in its entirety by reference, and can be performed by a first statistical model which can be selected from one of the following statistical models: a linear model (limma) model for microarray data, a logistic regression model, a linear discriminant analysis (LDA) model, a conditional logistic regression model, a lasso regression model, a ridge regression model, a random forest, a support vector machine, or a probit regression model. Note that other statistical models can also be used as the first statistical model.
[0096] Each of the terms “linear (limma) model of microarray data” (Ritchie et al. 2015), “logistic regression model” (Venable and Ripley 2002), “linear discriminant analysis (LDA) model” (Venable and Ripley 2002), “conditional logistic regression model” (Venable and Ripley 2002), “lasso regression model” (Tibshirani 1996), “ridge regression model” (Hoerl and Kennard 1970), “random forest” (Ripley 1996), “support vector machine” (Ripley 1996), and “probit regression model” (Venable and Ripley 2002) are essentially probabilistic statistical models that follow definitions generally understood by those skilled in the art, and their details can be found in the references listed in parentheses immediately following them.
[0097] As used herein, the term “adjusted p-value” means a p-value adjusted for multiple testing. The ranking of miRNAs (e.g., 1, 2, 3, ...) is based on adjusted p-values such that the miRNA with the smallest adjusted p-value has the highest rank (1), and miRNAs with smaller adjusted p-values have higher ranks.
[0098] Here, in substep S220, in which a diagnostic model is constructed based on the selected set of miRNA biomarkers, the selected set of miRNA biomarkers substantially contains at least one miRNA, and each miRNA in the selected set of miRNA biomarkers is selected such that its ranking is less than or equal to a predetermined cutoff m (i.e., not exceeding the cutoff m). Here, the "predetermined cutoff" can be any integer greater than or equal to 1 (i.e., m ≥ 1, for example, m = 1, 2, 3, 5, 10, 15, 20, 50, 100, 200, etc.).
[0099] According to some embodiments of the method for developing a cancer diagnosis model, the sub-step S220 of constructing a diagnosis model based on a selected miRNA biomarker set is to calculate a diagnosis index based on the expression profile of the selected miRNA biomarker set, and the diagnosis index is calculated by the formula:
Number
[0100] In the above formula (I), n is the total number of at least one miRNA in the selected miRNA biomarker set, miRNA i is the expression level of the i-th miRNA in the selected miRNA biomarker set, i is an integer greater than 0 and less than or equal to n (that is, 0 < i ≤ n), and t i is the weight for the i-th miRNA. Further, the total number of miRNAs in the selected miRNA biomarker set n is less than or equal to a preset cut-off m (that is, n ≤ m).
[0101] According to some embodiments, an unweighted approach is applied to calculate the diagnosis index, and thus the weight for each miRNA in the selected miRNA biomarker set is an equal number (for example, 1, 2, 0.5, etc.). In other words, for any miRNA in the selected miRNA biomarker set i for, t i = C (that is, a constant).
[0102] According to some other embodiments, a weighted approach is applied to calculate the diagnosis index, and thus the weight t i for the i-th miRNA is based on a second statistical model.
[0103] Here, depending on different embodiments, the second statistical model may be the same as or different from the first statistical model described above, and can be selected from limma models, logistic regression models, LDA models, conditional logistic regression models, lasso regression models, ridge regression models, random forests, support vector machines, or probit regression models.
[0104] According to some embodiments that use weighted approaches, such as the examples provided below, the limma model is selected as the second statistical model. Furthermore, according to some embodiments, both the first and second statistical models use the limma model. Thus, when differential expression analysis is performed on the miRNA expression profile in substep S210, coefficients can be generated for all miRNAs examined, and the coefficients corresponding to at least one miRNA in the selected set of miRNA biomarkers on which the diagnostic model was constructed can be used as weights in the diagnostic index calculation in substep S220.
[0105] According to several other embodiments employing a weighted approach, a second statistical model, based on obtaining a weight for each miRNA in a selected set of miRNA biomarkers, differs from the first statistical model, which is based on the ranking of miRNAs. For example, the limma model is used as the first model, and the LDA model is used as the second model.
[0106] According to some embodiments of the cancer diagnostic model development method described above, each of the n miRNAs in the selected miRNA biomarker set is a top n miRNA obtained in differential expression analysis in substep S210 as described above. In other words, the selected miRNA biomarker set contains n selected miRNAs, which together substantially represent top miRNAs with rankings 1, 2, 3, ..., n. In one example for illustrative purposes, the selected miRNA biomarker set may consist of four miRNAs with rankings 1, 2, 3, and 4, respectively, in substep S210.
[0107] It should be noted that the n(n) miRNAs in the selected miRNA biomarker set do not necessarily represent all of the top n miRNAs, and according to some other embodiments of this disclosure, they may represent only a subset of these top n miRNAs. In one example for illustrative purposes, the selected miRNA biomarker set may consist of five miRNAs having rankings 1, 2, 3, 4, and 6, respectively, in substep S210.
[0108] According to some embodiments of the cancer diagnostic model development method, after substep S220, step S200 may optionally further include the following substeps.
[0109] S230: Evaluate the performance of the diagnostic model through cross-validation in the training set.
[0110] As used herein, the term “cross-validation” refers to a technique for processing data within a given dataset, and in particular as used herein, it means dividing the training set into N equal parts (hence called N-fold cross-validation), with each part being used sequentially as the validation set and the remaining parts together as the training set. The diagnostic model is developed on the training set using the same approach described in S210 and S220, and then validated on the validation set. Here, the number of cross-validation divisions N can be any integer greater than or equal to 2 (i.e., N ≥ 2, e.g., 2, 3, 4, 5, ..., 10, 20, etc.).
[0111] The performance of the diagnostic model obtained in substep S220 can be evaluated in substep S230 based on different parameters (multiple parameters are possible).
[0112] According to some embodiments, the substep S230 for evaluating the performance of the diagnostic model includes calculating the area under the curve (AUC) of the receiver operating characteristic (ROC) curve and evaluating the performance of the diagnostic model based on it (i.e., based on the AUC). Here, a preset cutoff for the AUC (e.g., 0.80, 0.90, 0.95, 0.99, etc.) can be used to evaluate the performance of the diagnostic model based on the calculated AUC value, and if the AUC value calculated in this way is greater than or equal to the preset cutoff, the diagnostic model is determined to have good performance. However, if the AUC value calculated in this way is lower than the preset cutoff, the diagnostic model is determined to have not reached ideal performance or to have poor performance.
[0113] In some other embodiments, a substep S230 for evaluating the performance of a diagnostic model includes calculating the specificity and sensitivity of the diagnostic model and evaluating the performance of the diagnostic model based on these (based on the specificity and sensitivity). Here, a preset cutoff for sensitivity (e.g., 0.80, 0.90, 0.95, 0.99, etc.) to a given specificity (e.g., a preset value such as 0.90, 0.95, 0.98, or 0.99) can be used to evaluate the performance of the diagnostic model based on the calculated sensitivity and specificity values. If the sensitivity value for a given specificity calculated in this way is greater than or equal to the preset cutoff, the diagnostic model is determined to have good performance. However, if the sensitivity value for a given specificity calculated in this way is lower than the preset cutoff, the diagnostic model is determined to have not achieved ideal performance or to have poor performance.
[0114] In addition to the use of AUC in the ROC curve and the use of specificity and sensitivity, it should be noted that there are other parameters that can be used to evaluate the performance of the diagnostic model thus obtained in substep S220 of the cancer diagnostic model development method. Examples of these other parameters include, but are not limited to, overall accuracy, positive likelihood ratio, negative likelihood ratio, positive predictor, negative predictor, and Juden's J.
[0115] Furthermore, it should be noted that the substep S230, which evaluates the performance of the diagnostic model by cross-validation within the training set, is optional and may be skipped according to some embodiments of this disclosure.
[0116] In step S300, the diagnostic model thus obtained in step S200 (particularly substep S220) as described above is further validated using a different validation set than the training set, which does not contain overlapping subjects or patients. Here, the validation set includes expression profiles of a selected set of miRNA biomarkers obtained from non-cancer subjects (i.e., controls) and cancer patients with the cancer of interest (i.e., cases).
[0117] More specifically, step S300 can be performed by evaluating the same or different parameters as those used in the aforementioned substep S230, which evaluates the performance of the diagnostic model by cross-validation on a training set. According to some embodiments of the present disclosure, the AUC of the ROC curve is used for validation such that the diagnostic model is considered valid if the AUC is greater than or equal to a preset cutoff (e.g., 0.80, 0.90, 0.95, 0.99, etc.), or invalid if the AUC is lower than the preset cutoff. Here, the preset cutoff for AUC used in step S300 may be the same as or different from the preset cutoff for AUC used in substep S230 as described above. According to some other embodiments of this disclosure, sensitivity and specificity are used for validation such that if the sensitivity value for a given specificity calculated in this manner (e.g., 0.90, 0.95, 0.98, or 0.99) is greater than or equal to a preset cutoff (e.g., 0.80, 0.90, 0.95, 0.99), the diagnostic model is deemed valid, or if the sensitivity value for a given specificity calculated in this manner is lower than the preset cutoff, the diagnostic model is deemed invalid. There are also other parameters that can be used to validate the diagnostic model.
[0118] The disclosure further provides a computer-aided solution that is substantially useful for carrying out various steps and / or substeps of the cancer diagnostic model development method described above in a computerized and automated manner. Such a computer-aided solution can be applied in situations where the carrying out of various steps S100 to S300, and even various substeps S210 to S230 of step S200, of the cancer diagnostic model development method described above is automated by executing a software program containing program instructions within a computer, which brings advantages such as high efficiency and great convenience.
[0119] Specifically, such computer-based solutions may include a computerized system or computer system (or simply "the System"). The System includes a collection of hardware (e.g., a processor, memory, I / O interfaces, storage media, etc.) and software (i.e., computer programs, including operating system software and specific program software, etc.) configured to work together to collectively perform all or some of the steps of the cancer detection diagnostic model development method described above. According to some embodiments, the System comprises a processor (i.e., a controller) and a computer-readable non-temporary storage medium communicably coupled to the processor. The non-temporary storage medium is configured to contain software (i.e., program instructions) to be executed by the processor, and the program instructions are configured to cause the processor to perform various different steps and substeps in the cancer detection diagnostic model development method described above.
[0120] As used herein and throughout this disclosure, “processor” is interchangeable with “central controller” or “central processing unit (CPU)” and can be considered as a single-core or multi-core processor, or multiple processors for parallel processing. As used herein, the term “non-temporary” is intended to describe a tangible computer-readable storage medium other than propagating electromagnetic signals, but is not intended to limit the types of physical computer-readable storage devices included in this phrase. Examples may include any tangible or non-temporary storage or memory medium, such as electronic, magnetic, or optical media (e.g., disks or CD / DVD-ROMs), or non-volatile memory storage devices (e.g., “flash” memory).
[0121] As shown in Figure 2, in addition to the processor 10 and the computer-readable non-temporary storage medium 20, the system 100 may further include a bus 30, memory 40, I / O interface 50, and communication interface 60. The processor 10, storage medium 20, memory 40, I / O interface 50, and communication interface 60 are all coupled to communicate with each other via the bus 30.
[0122] The storage medium 20 stores computer executable program instructions, which, when executed by the processor 10, cause the processor 10 to perform steps (1) to (3) of the method described above. The memory 40 is configured to temporarily store program instructions obtained from the storage medium 20, and the processor 10 is configured to execute program instructions temporarily stored in the memory 40. The I / O interface 50 enables input and output between the system 100 and the user and provides control of the system 100. The communication interface 60 can enable the system 100 to communicate with another computing device for data exchange. Note that these computer hardware components can be located locally or remotely via a network such as an intranet, the internet, or the cloud.
[0123] In a second aspect, the disclosure further provides a method for detecting a cancer of interest from a subject (e.g., a human), wherein the cancer detection method is substantially based on a diagnostic model developed by a cancer detection diagnostic model according to any of the embodiments described above.
[0124] Specifically, the cancer detection method includes the following main steps:
[0125] S1000: A step to determine the expression profile of a selected set of miRNA biomarkers from a biological sample obtained from the subject.
[0126] S2000: A step of calculating a diagnostic index for a biological sample based on the expression profile of a selected set of miRNA biomarkers, wherein the diagnostic index is given by the formula:
number
[0127] S3000: A step of classifying subjects as having or not having the target cancer based on a calculated diagnostic index, wherein a subject is classified as having the target cancer if the calculated diagnostic index is equal to or greater than a predetermined threshold, and is classified as not having the target cancer otherwise.
[0128] Here, in step S1000 of the cancer detection method provided herein, the “biological sample” from the subject may be a liquid biopsy sample such as a blood sample, serum sample, plasma sample, urine sample, saliva sample, or sputum sample, but optionally may be a sample from tumor tissue. The “selected miRNA biomarker set” is substantially the set of miRNA biomarkers selected when developing a diagnostic model for detecting the cancer of interest by a cancer detection diagnostic model development method such as the one provided in the first aspect of this disclosure. Determination of the expression profile of the selected miRNA biomarker set from the biological sample obtained from the subject may be achieved by various probe-based approaches including Northern blotting, microarray analysis, RNA sequencing, or RNA in situ hybridization, or by various amplification-dependent approaches including reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR. When used herein, each of the above miRNA detection approaches should be understood within the scope of common definitions well understood by those skilled in the art.
[0129] In step S2000 of the cancer detection method provided herein, the diagnostic index can be calculated in the same manner as the cancer detection diagnostic model development method described above in the first embodiment of this disclosure.
[0130] In step S3000, the term “predetermined threshold” refers to a cut-point value of a diagnostic index that can be used to determine or decide whether a subject has the cancer of interest with a given specificity / sensitivity. The predetermined threshold is typically predetermined based on an existing dataset that includes a range of diagnostic index values obtained and calculated for an existing population of subjects known to have and / or not have the disease. International Patent Application No. WO2022261039A2 provides further details on how subjects are classified as having or not having the cancer of interest based on the calculated diagnostic index, and its disclosure is incorporated herein by reference in its entirety.
[0131] According to some embodiments, each miRNA in the selected miRNA biomarker set is from the top 200 miRNAs listed in Table 1 provided in Example 1 below, and the total number of miRNAs in the selected miRNA biomarker set (i.e., "n") is 200 or less. Note that each miRNA in the selected miRNA biomarker set may be from any number of top miRNAs ranked in substep S210 of step S200 of the cancer detection diagnostic model development method described above, which may be the top 100, top 50, or top 500, depending on the different embodiments.
[0132] According to some preferred embodiments, the diagnostic index in step S2000 of the cancer detection method is calculated using weights from a limma model. It should be noted that, according to some other embodiments, the diagnostic index can be calculated, alternatively and optionally, via weights from other statistical models such as logistic regression models, LDA models, conditional logistic regression models, lasso regression models, ridge regression models, random forests, support vector machines, or probit regression models. Alternatively or optionally, the diagnostic index can be calculated via an unweighted approach, in which case the weights for each miRNA in the selected miRNA biomarker set are equal numbers (e.g., 1, 2, 0.5, etc.). In other words, any miRNA in the miRNA biomarker set i Regarding t i = C (i.e., a constant).
[0133] According to some embodiments of the cancer detection method, each of the n miRNAs in the selected miRNA biomarker set is a top n miRNA obtained in differential expression analysis in substep S210 of step S200 of the cancer detection diagnostic model development method described above, in the first embodiment of the present disclosure. In other words, the selected miRNA biomarker set includes n selected miRNAs, which together substantially represent top miRNAs having rankings 1, 2, 3, ..., n. It should be noted that the n miRNAs in the selected miRNA biomarker set do not necessarily each represent all of the top n miRNAs, and according to some other embodiments of the present disclosure, they may represent only a subset of these top n miRNAs.
[0134] Herein, according to some embodiments of the cancer detection methods provided herein, n is between 4 and 200 (4 ≤ n ≤ 200), and the classification can achieve an AUC greater than approximately 0.97 to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0135] According to some embodiments, n is between 4 and 200 (4 ≤ n ≤ 200), and the classification can achieve a sensitivity of at least about 0.75 while maintaining a specificity of about 0.99 or higher to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0136] Furthermore, according to several embodiments, the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0137] Furthermore, according to several embodiments, the classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0138] Furthermore, according to several embodiments, the classification can achieve a sensitivity of at least about 0.95 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, esophageal cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0139] Furthermore, according to some embodiments, the classification can achieve a sensitivity of at least about 0.99 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, esophageal cancer, gastric cancer, biliary tract cancer, bladder cancer, or prostate cancer.
[0140] According to some embodiments of the cancer detection methods provided herein, the selected miRNA biomarker set consists of the top four miRNAs, including hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, and hsa-miR-663a. Thus, the classification can achieve a sensitivity of at least about 0.75 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Further optional, the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity value of about 0.99 for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, with optional selection, the classification can achieve a sensitivity of at least approximately 0.90 while maintaining a specificity value of approximately 0.99 to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, with optional selection, the classification can achieve a sensitivity of at least approximately 0.95 while maintaining a specificity value of approximately 0.99 to detect lung cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Furthermore, with optional selection, the classification can achieve a sensitivity of at least approximately 0.99 while maintaining a specificity value of approximately 0.99 to detect lung cancer, gastric cancer, biliary tract cancer, or bladder cancer.
[0141] According to some embodiments of the cancer detection methods provided herein, the selected set of miRNA biomarkers comprises the top five miRNAs, including hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-663a, and hsa-miR-320a, and the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity value of about 0.99 for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0142] According to some embodiments of the cancer detection methods provided herein, the selected set of miRNA biomarkers consists of the top 10 miRNAs in Table 1, and the classification can achieve a sensitivity of at least about 0.99 while maintaining a specificity value of about 0.99 for detecting gastric cancer, esophageal cancer, biliary tract cancer, or prostate cancer.
[0143] According to some embodiments of the cancer detection methods provided herein, the selected set of miRNA biomarkers consists of the top 15 miRNAs in Table 1, and the classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity value of about 0.99 for detecting lung cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
[0144] In any embodiment of the cancer detection method described above, the method optionally further includes the step of performing an evaluation of a subject, the evaluation including the diagnosis of cancer or the detection of cancer recurrence. Here, “diagnosis of cancer” means the detection of cancer in a subject that was previously not considered to have cancer, and “cancer recurrence” means the detection of cancer again in a subject that had cancer but was previously considered to have disappeared after cancer removal treatment.
[0145] In any embodiment of the methods described above, the method further optionally includes a step of performing a diagnostic procedure on a subject if the subject is classified as having cancer. The diagnostic procedure may optionally include a physical examination, pathological examination of a biopsy from the subject, immunohistochemical examination, or imaging tests such as X-ray, computed tomography (CT), ultrasound, and / or magnetic resonance imaging.
[0146] In any embodiment of the above method, the method optionally further includes the step of administering a treatment regimen to a subject if the subject is classified as having cancer. Herein, various known treatment regimens, including surgery, radiotherapy, chemotherapy, hormone therapy, targeted therapy, immunotherapy, or combinations thereof, can be administered in the method. These above treatment regimens are well established for each of the above various cancers.
[0147] The Disclosure further provides a computer-aided solution, including a computer system (including a processor and a non-temporary storage medium) for carrying out various steps and / or substeps of the cancer detection method described above. Details of such a computer-aided solution can be found in the computer-aided solution for cancer detection diagnostic model development method provided in a first aspect of the Disclosure.
[0148] This disclosure further provides a kit for detecting cancer from biological samples obtained from subjects, which is substantially used to carry out the cancer detection methods described above.
[0149] Where used herein and elsewhere in this disclosure, the term “Kit” refers to a collection of articles and / or instructions. Articles included in a Kit may be physical entities or their components. Examples of articles that may be included in a Kit as disclosed herein include one or more nucleic acids (e.g., polynucleotides) or one or more devices, apparatus or equipment (e.g., molecular arrays or microarrays containing one or more nucleic acids). Instructions included in a Kit may be instructions for specific steps to be performed (e.g., a manual) and may be printed on a physical medium (e.g., paper, card, etc.), a computer-readable storage medium (e.g., hard disk, compact disc or CD, flash drive, etc.), or even stored on the Internet (e.g., in an accessible cloud space).
[0150] The kit may include at least the following components (1) and (2) (i.e., articles and / or instructions):
[0151] Component (1): At least one nucleic acid that can obtain an expression profile of a selected miRNA biomarker set from a biological sample by specifically recognizing each miRNA in the selected miRNA biomarker set. Here, the selected miRNA biomarker set is substantially a set of miRNA biomarkers selected when developing a diagnostic model for detecting a cancer of interest by a cancer detection diagnostic model development method such as that provided in a first aspect of this disclosure.
[0152] Component (2): At least one instruction manual comprising a first instruction manual and a second instruction manual, wherein the first instruction manual substantially carries out step S2000 of calculating a diagnostic index of a biological sample based on the expression profile of a selected set of miRNA biomarkers. The second instruction manual substantially carries out step S3000 of classifying a subject as having or not having cancer, wherein the subject is classified as having cancer if the calculated diagnostic index is greater than or equal to a predetermined threshold, and as not having cancer otherwise.
[0153] Here, in component (1) of the kit, at least one nucleic acid may optionally include (a) a polynucleotide containing or consisting of the base sequence of each miRNA in the selected miRNA biomarker set, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof containing 15 or more consecutive nucleotides, or (b) a polynucleotide containing or consisting of a base sequence complementary to the base sequence of each miRNA in the selected miRNA biomarker set, a polynucleotide containing a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof containing 15 or more consecutive nucleotides, under stringent conditions.
[0154] According to a different embodiment, at least one instruction manual in component (2) of the kit may further include a third instruction manual for performing an assessment of the subject, wherein the assessment includes a diagnosis of cancer or detection of cancer recurrence. Alternatively, at least one instruction manual in component (2) of the kit may further include a fourth instruction manual for administering a treatment regimen to the subject if the subject is classified as having cancer.
[0155] According to some embodiments, at least one instruction manual in component (2) of the kit is a first additional instruction manual for obtaining an expression profile of a selected set of miRNA biomarkers, which further includes a first additional instruction manual including a procedure for performing Northern blotting, microarray analysis, RNA sequencing, or RNA in situ hybridization with at least one nucleic acid, where the at least one nucleic acid may optionally be placed on a molecular array.
[0156] According to some embodiments, the kit may further include at least one set of amplification primers, each set of which is capable of specifically amplifying each of at least one miRNA in a selected set of miRNA biomarkers from a biological sample. According to some embodiments, at least one instruction manual in component (2) of the kit may further include a second additional instruction manual for obtaining an expression profile of the selected set of miRNA biomarkers, which includes a procedure for performing reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR with at least one nucleic acid and at least one set of amplification primers.
[0157] Below, an example (i.e., Example 1) is provided to illustrate the present invention as described above in various embodiments of this disclosure.
[0158] Example 1
[0159] In this example, a large multi-cancer training set (hereinafter referred to as the "training set") was constructed, including multiple cancer species rather than a single cancer species, by using eight publicly available serum miRNA microarray datasets totaling 6,283 cancer patients and 5,130 non-cancer controls. Differentially expressed miRNAs were identified, and the optimal number of miRNAs for constructing a diagnostic model in the training set was determined using 10-fold cross-validation. The performance of the model was then evaluated on three validation sets. More detailed information is provided below.
[0160] method
[0161] Experimental Design and Construction of Training and Validation Datasets: Eight serum miRNA microarray datasets were identified from the Gene Expression Omnibus (GEO) (Zhang A et al., 2022; Zhang A and Hu H, 2022). All of these were constructed from the Japanese nationwide research project "Development and Diagnostic Techniques for the Detection of miRNAs in Body Fluids," which was designed to characterize serum miRNAs from over 50,000 participants across 13 cancers using a standardized microarray platform. These eight datasets were originally used to develop individual diagnostic models for cancers of the lung (GSE137140) (Asakura K, et al., 2020), ovary (GSE106817) (Yokoi A, et al., 2018), liver (GSE113740) (Yamamoto Y, et al., 2020), bladder (GSE113486) (Usuba W, et al., 2019), esophageal squamous epithelium (GSE122497) (Sudo K, et al., 2019), stomach (GSE164174) (Abe S, et al., 2021), prostate (GSE112264) (Urabe F, et al., 2019), and glioma (GSE139031) (Ohno M, et al., 2019), respectively. After removing duplicate cases, three large, independent datasets were constructed: a lung cancer dataset (n=3744) (Zhang A et al., 2022; Asakura K, et al., 2020), a combined dataset (n=3792) created by merging ovarian cancer, liver cancer, and bladder cancer datasets (Zhang A et al., 2022; Yokoi A, et al., 2018; Yamamoto Y, et al., 2020; Usuba W, et al., 2019), and a combined dataset (n=3877) created by merging esophageal squamous cell carcinoma, gastric cancer, prostate cancer, and glioma cancer datasets (Zhang A and Hu H, 2022; Sudo K, et al., 2019; Abe S, et al., 2021; Ohno M, et al., 2019; Urabe F, et al., 2019).Based on these three large datasets, a large training set was constructed to develop a diagnostic model for detecting multiple cancer types. This set included 1408 cancer patients from seven cancer types (208 lung cancer patients, and 200 each from ovarian, liver, bladder, esophageal, stomach, and prostate cancers) and 1408 age- and sex-matched non-cancer controls. Three separate, independent validation sets were formed for all remaining cases (Figures 3A and 3B).
[0162] Blood sample collection and miRNA microarray analysis: Serum sample collection and microarray expression analysis have been previously described (Asakura K et al., 2020). Briefly, total RNA was extracted from 300 μl of serum collected from cancer patients and non-cancer controls before surgery, labeled with the 3DGene® miRNA labeling kit, and hybridized to 3D-Gene® human miRNA oligochips (Toray Industries, Inc., Kanagawa Prefecture, Japan). The signal intensity for miRNA was determined after background removal and normalization using three pre-selected internal control miRNAs (miR-149-3p, miR-2861, and miR-4463).
[0163] Diagnostic Model Development: Identification of miRNA biomarkers and all model development work were performed only on a multi-cancer training set. Differential miRNA expression between cancer and non-cancer was evaluated using a linear model (limma) of microarray data (Ritchie ME, et al., 2015). miRNAs were then ranked based on the statistical significance of differential expression, and top-ranking miRNAs were used to construct diagnostic models for distinguishing cancer from non-cancer. A diagnostic index was calculated for each diagnostic model as a linear sum of the expression levels of selected miRNAs weighted by limma statistics. Ten-fold cross-validation was performed to determine the optimal number of miRNAs to be included in the final diagnostic model that yielded the highest area under the receiver operating characteristic (ROC) curve (AUC) for distinguishing cancer from non-cancer. The cutpoints for the diagnostic index were chosen to ensure at least 99% specificity (i.e., <1% false positive rate), as the models could be used as a screening tool for the general population at potential risk.
[0164] Diagnostic Model Validation: Three independent validation datasets contained mutually exclusive samples not used in model development, each offering different characteristics for validating the developed model. Validation Set 1 was not only a very large sample size set of lung cancer cases, but also included comprehensive patient-level clinicopathological data, in contrast to the other two validation datasets, allowing for evaluation of model performance against early-stage cancers and different histological subtypes. Validation Set 2 included samples from 12 other cancers, thus expanding the evaluation of model performance across multiple cancer types. Validation Set 3 included a large number of cases from four cancers, including two cancers with smaller sample sizes in Validation Set 2, enabling further individual validation of model performance.
[0165] Statistical Analysis: Diagnostic performance for detecting cancer versus non-cancer was measured using AUC, sensitivity, and specificity from ROC curve analysis. Sensitivity was defined as the proportion of cancer patients correctly identified as having cancer by the diagnostic model, and specificity was defined as the proportion of non-cancer participants correctly identified as not having cancer. Limma analysis was performed using the Bioconductor package limma (Ritchie ME, et al., 2015). All statistical analyses were performed using R version 4.2.1.
[0166] result
[0167] Participants and Datasets: Detailed demographic and clinical information for these cancer types with large sample sizes was provided in the original publication. Briefly, the lung cancer dataset (n=1566) had a mean age of 65 years, 57% were male, 62% were former or current smokers, 78% of tumors were adenocarcinoma, 14% were squamous cell carcinoma, and 87% were stage I or II. The bladder cancer dataset (n=392) included patients with a mean age of 68 years, 72% were male, 95% were non-metastatic, 88% were lymph node negative, 77% were T1, and 80% were high-grade. The ovarian cancer dataset (n=333) included patients with a mean age of 57 years, 35% were stage I or II, and 96% had epithelial carcinoma (55% serous, 19% clear cell, and 13% endometrial). The liver cancer dataset (n=348) consisted of patients with a mean age of 68 years, 78% male, and 70% at stage I or II. The esophageal cancer dataset (n=447) consisted of patients with a mean age of 67 years, 97% male, and 66% at stage I or II. The gastric cancer dataset (n=1267) consisted of patients with a mean age of 66 years, 77% male, and all patients at stage I or II. The glioma dataset (n=196) consisted of patients with a mean age of 56 years, and 57% male. Finally, the prostate cancer dataset (n=769) consisted of patients with a mean age of 68 years, 93% lymph node negative, and 92% non-metastatic.
[0168] Development of Cancer Diagnostic Models: All diagnostic model development work was performed on a multi-cancer training set containing 1408 cancer patients and 1408 age- and sex-matched non-cancer controls (Figure 3B). First, differential expression of miRNAs between cancer and non-cancer was evaluated using limma analysis. Then, miRNAs were ranked based on adjusted p-values. The top 200 differentially expressed miRNAs are listed in Table 1. Next, 10-fold cross-validation was performed using the multi-cancer training set (Figure 4A). The top four miRNAs (hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, and hsa-miR-663a) showed the highest AUC in ROC analysis and were therefore found to be included in the final diagnostic model (Figure 4B). Next, the diagnostic performance of different diagnostic models for detecting different cancers was compared using different validation sets. These different diagnostic models differed in the number of top miRNAs N (N is 4-200). The results are shown in Figure 5. The inventors calculated the diagnostic index by the weighted sum of the four miRNA expression levels and normalized it to a range of 0 to 10. This 4-miRNA model achieved an AUC value of 0.994 in the training set (Figure 6A). By selecting a cut point of 5.3, an overall specificity of >99% (i.e., <1% false positives) and an overall sensitivity of 94% were obtained across all non-cancer cases (Figure 6B). The AUC and sensitivity of the model for each of the seven cancers in the multi-cancer training set ranged from 0.985 and 84% for ovarian cancer to 0.998 and 100% for bladder cancer and gastric cancer, respectively (Figure 5). [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10] [Table 1-11]
[0169] Validation of the diagnostic model in an independent validation set 1: The performance of the 4-miRNA model was first evaluated in an independent validation set 1 (n=2859) including 1358 lung cancer patients and 1501 non-cancer controls. This model achieved an AUC of 1,000 (Figure 7A), with a specificity of 100% and a sensitivity of 99% (Figure 7B). Furthermore, analysis of paired serum samples (pre-operative vs. post-operative; n=180) confirmed that the diagnostic index in post-operative serum samples normalized to the level of non-cancer controls (Figure 7C). The performance of the 4-miRNA model was further evaluated across clinical subsets of validation set 1, as defined by clinical stage, TNM stage, and histological subtype. High sensitivity was observed for all clinical subsets. This model achieved at least 99% sensitivity for 22 of the 24 clinical subsets examined, except for stage IIB and T3 tumors (Figure 7D). In particular, this model showed a sensitivity of >99% for stage I lung cancer, as well as adenocarcinoma and squamous cell carcinoma.
[0170] Validation of the diagnostic model in independent validation sets 2 and 3: Independent validation set 2 included 1,438 patients and 1,623 non-cancer controls across 12 additional cancers. With the exception of breast cancer, the 4-miRNA model achieved at least 90% sensitivity for eight cancers (biliary tract, bladder, colorectal, esophageal, gastric, glioma, pancreatic, and prostate) and at least 75% sensitivity for the other three cancers (liver, ovarian, and sarcoma) (Figure 8A and Table 2). Notably, although the model had a reasonable AUC of 0.909 for breast cancer, the 1% sensitivity remained very low due to the high specificity requirement (Figure 8A and Table 2). An independent validation set 3 included 2079 patients from four cancers (esophageal, gastric, glioma, and prostate) and 598 non-cancer controls. The sample sizes for the four cancers were substantially larger than those in validation set 2 (247 vs. 124 for esophageal cancer, 1067 vs. 150 for gastric cancer, 196 vs. 40 for glioma, and 569 vs. 40 for prostate cancer). The 4-miRNA model achieved an AUC of >0.99 and sensitivity of >99% for all four cancers, similar to those observed in validation set 2 (Figure 8B and Table 2). The specificity of the model was slightly lower in validation set 3 than in validation set 2 (0.98 vs. 0.99) (Figure 8B). Therefore, for validation set 3, a sensitivity analysis with an adjusted diagnostic index cutpoint of 5.6 was considered to increase the specificity of the new model to 99%. With this new cutpoint, the model still achieved high sensitivity for all four types of cancer, including 99% for gastric cancer, 92% for glioma, 91% for prostate cancer, and 89% for esophageal cancer.
[0171] Consideration
[0172] Non-invasive screening tests for MCED by analyzing circulating cell-free nucleic acids and / or proteins in bodily fluids, particularly blood, have attracted considerable attention over the past decade. In this study, we report the development and validation of a serum 4-miRNA diagnostic model, demonstrating that the 4-miRNA model can simultaneously detect 12 cancers with high sensitivity (over 90% for 9 cancers and over 75% for 3 cancers) while still achieving a very high specificity of approximately 99% in three large, independent validation sets involving a total of 8,597 participants (4,875 cancer patients and 3,722 non-cancer individuals). Furthermore, the observation that the diagnostic index of postoperative serum samples decreased to normal levels suggests the potential usefulness of the model for monitoring treatment response and detecting recurrence.
[0173] Importantly, our model was able to detect early-stage cancer with high sensitivity. Specifically, in validation set 1 of lung cancer patients, the model detected stage I and II cancers with a sensitivity range of 98.4% to 99.6% (Figure 7D). In validation sets 2 and 3, individual patient-level stage information was not available, but overall stage information was obtained for 6 of the 12 cancers tested. Firstly, since all gastric cancer patients were stage I or II, the 100% sensitivity of our model was applied to early gastric cancer. Secondly, 88% and 93% of bladder cancer and prostate cancer patients had lymph node-negative disease, respectively. Therefore, given the 99% and 98% sensitivities for these two cancers, the sensitivity for stage I or II bladder and prostate cancer is also considered to be very high. Thirdly, 66% and 70% of esophageal cancer and liver cancer patients were stage I or II, respectively. It was reasonable to assume that the sensitivity for stage I or II of these two cancers did not deviate significantly from the 92% and 84% sensitivities reported for all stages included. In summary, based on the data currently available in three validation sets, we concluded that our 4-miRNA model achieves high sensitivity for stage I or II of six cancers (lung, stomach, bladder, prostate, esophagus, and liver).
[0174] Notably, simple four-parameter diagnostic models, such as those described herein, are not only significantly less expensive but can also be developed into in vitro diagnostic (IVD) trials using RT-qPCR, which allows for decentralized testing—an advantage over NGS-based trials typically performed as laboratory-developed tests (LDTs). These characteristics are important for promoting the adoption of MCED tests and improving their cost-effectiveness, as MCED tests are intended for high-risk or at-risk populations, particularly those in low-income communities.
[0175] In summary, our research provides proof-of-concept data for developing a blood screening test based on circulating cell-free miRNA expression profiles for 12 types of cancer, which account for an estimated 50% of new cancer cases and 63% of cancer deaths in the United States in 2022 (Siegel RL, et al., 2022). References JPEG2026510430000016.jpg244162
Claims
1. A method for developing a diagnostic model to detect a target cancer, (1) A step of constructing a training set that includes expression profiles of multiple miRNAs obtained from non-cancer subjects and cancer patients, wherein at least two types of cancer are present in the cancer patients, (2) A step of developing the diagnostic model based on the training set, (a) A substep in which differential expression analysis is performed on the expression profiles of the plurality of miRNAs by a first statistical model so that the plurality of miRNAs are ranked by adjusted p-values, (b) A substep of constructing the diagnostic model based on a selected set of miRNA biomarkers from the plurality of miRNAs, wherein the selected set of miRNA biomarkers each includes at least one miRNA having a rank less than or equal to a predetermined cutoff m, and m is an integer greater than 0. The steps include developing the diagnostic model and Methods that include...
2. The substep (b) in step (2) is, Calculating a diagnostic index based on the expression profile of the selected miRNA biomarker set, wherein the diagnostic index is given by the formula: [Math 1] The formula is calculated based on the following, where n is the total number of at least one miRNA in the selected miRNA biomarker set, n is an integer less than or equal to m, and miRNA i is the expression level of the i-th miRNA in the selected miRNA biomarker set, where i is an integer greater than 0 and less than or equal to n, and t i This involves calculating the diagnostic index, which is the weight for the i-th miRNA. The method according to claim 1, including the method described in claim 1.
3. t i The method according to claim 2, wherein is a constant.
4. t i The method according to claim 2, wherein the method is based on a second statistical model selected from one of the following: a linear model (limma) model of microarray data, a logistic regression model, a linear discriminant analysis (LDA) model, a conditional logistic regression model, a Lasso regression model, a Ridge regression model, a random forest, a support vector machine, or a probit regression model.
5. The method according to claim 4, wherein the first statistical model and the second statistical model are substantially the same.
6. The method according to claim 5, wherein a limma model is used as both the first statistical model and the second statistical model.
7. The method according to any one of claims 2 to 6, wherein at least one of the selected miRNAs in the miRNA biomarker set is a top n miRNA.
8. Step (2) is, (c) A substep to evaluate the performance of the diagnostic model through cross-validation in the training set. The method according to any one of claims 1 to 7, further comprising:
9. The method according to claim 8, wherein the cross-validation is divided into at least two parts.
10. The substep (c) which evaluates the performance of the diagnostic model by cross-validation in the training set, Calculate the area under the curve (AUC) of the receiver operating characteristic (ROC) curve and evaluate the performance of the diagnostic model based on it, or The specificity and sensitivity of the diagnostic model are calculated, and the performance of the diagnostic model is evaluated based on these. The method according to claim 8 or 9, comprising at least one of the following.
11. (3) Validating the diagnostic model based on a validation set, wherein the validation set includes expression profiles of the selected miRNA biomarker set obtained from non-cancer subjects and cancer patients with the target cancer. The method according to any one of claims 1 to 10, further comprising:
12. The above step (3) is, Calculate the AUC of the ROC curve and evaluate the performance of the diagnostic model based on it, or The specificity and sensitivity of the diagnostic model are calculated, and the performance of the diagnostic model is evaluated based on these. The method according to claim 11, comprising at least one of the following.
13. A system for developing diagnostic models to detect a target cancer, Processor and A non-temporary storage medium comprising program instructions for execution by the processor, wherein the program instructions cause the processor to perform the steps in the method according to any one of claims 1 to 12. A system that includes this.
14. A non-temporary storage medium for storing computer executable program instructions, wherein, when the program instructions are executed by a processor, the non-temporary storage medium causes the processor to execute the method according to any one of claims 1 to 12.
15. A method for detecting a target cancer from a subject using a diagnostic model, wherein the diagnostic model is developed by the method described in any one of claims 1 to 12.
16. To determine the expression profile of the selected miRNA biomarker set from the biological sample obtained from the subject, The diagnostic index of the biological sample is calculated based on the expression profile of the selected miRNA biomarker set, wherein the diagnostic index is given by the formula: [Math 2] The formula is calculated based on the following, where n is the total number of at least one miRNA in the selected miRNA biomarker set, and miRNA i is the expression level of the i-th miRNA in the selected miRNA biomarker set, where i is an integer greater than 0 and less than or equal to n, and t i This involves calculating the diagnostic index, which is the weight of the i-th miRNA, The process of classifying the subject as having or not having the target cancer based on the calculated diagnostic index, wherein the subject is classified as having the target cancer if the calculated diagnostic index is equal to or greater than a predetermined threshold, and as not having the target cancer otherwise. The method according to claim 15, including the method described in claim 15.
17. The method according to claim 16, wherein each miRNA in the selected miRNA biomarker set is from the top 200 miRNAs listed in Table 1, and n is 200 or less.
18. The method according to claim 16 or 17, wherein the diagnostic index is calculated via a model weighted using weights from the limma model.
19. The method according to claim 18, wherein at least one of the selected miRNAs in the miRNA biomarker set is a top n miRNA.
20. The method according to claim 19, wherein n is 4 or more and 200 or less, and the classification can achieve an AUC of more than about 0.97 to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
21. The method according to claim 19, wherein n is 4 or more and 200 or less, and the classification can achieve a sensitivity of at least about 0.75 while maintaining a specificity of about 0.99 or more for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
22. The method according to claim 21, wherein the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
23. The method according to claim 22, wherein the classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
24. The method according to claim 23, wherein the classification can achieve a sensitivity of at least about 0.95 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, esophageal cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
25. The method according to claim 24, wherein the classification can achieve a sensitivity of at least about 0.99 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, esophageal cancer, gastric cancer, biliary tract cancer, bladder cancer, or prostate cancer.
26. The method according to any one of claims 16 to 19, wherein the selected miRNA biomarker set consists of hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, and hsa-miR-663a.
27. The method according to claim 26, wherein the classification can achieve a sensitivity of at least about 0.75 while maintaining a specificity of about 0.99 or higher for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
28. The method according to claim 27, wherein the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity value of about 0.99 for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
29. The method according to claim 28, wherein the classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity value of about 0.99 for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
30. The method according to claim 29, wherein the classification can achieve a sensitivity of at least about 0.95 while maintaining a specificity value of about 0.99 for detecting lung cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
31. The method according to claim 30, wherein the classification can achieve a sensitivity of at least about 0.99 while maintaining a specificity value of about 0.99 for detecting lung cancer, gastric cancer, biliary tract cancer, or bladder cancer.
32. The method according to any one of claims 16 to 19, wherein the selected set of miRNA biomarkers comprises hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-663a, and hsa-miR-320a, and the classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity value of about 0.99 for detecting lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
33. The method according to any one of claims 16 to 19, wherein the selected set of miRNA biomarkers consists of the top 10 miRNAs in Table 1, and the classification can achieve a sensitivity of at least about 0.99 while maintaining a specificity value of about 0.99 for detecting gastric cancer, esophageal cancer, biliary tract cancer, or prostate cancer.
34. The method according to any one of claims 16 to 19, wherein the selected set of miRNA biomarkers consists of the top 15 miRNAs in Table 1, and the classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity value of about 0.99 for detecting lung cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.
35. The method according to any one of claims 16 to 34, wherein the expression profile of the selected miRNA biomarker set is obtained by at least one of Northern blotting, microarray analysis, RNA sequencing, RNA in situ hybridization, or a nucleic acid amplification procedure, the nucleic acid amplification procedure comprising at least one of reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR.
36. The method according to any one of claims 16 to 34, wherein the biological sample is a liquid biopsy sample selected from the group consisting of blood samples, serum samples, plasma samples, urine samples, saliva samples, and sputum samples.
37. A system for detecting a target cancer from a subject, Processor and A non-temporary storage medium comprising program instructions for execution by the processor, wherein the program instructions cause the processor to perform a step in the method according to any one of claims 16 to 36. A system that includes this.
38. A non-temporary storage medium for storing computer executable program instructions, wherein, when the program instructions are executed by a processor, the non-temporary storage medium causes the processor to execute the method according to any one of claims 16 to 36.