Cancer detection methods, kits and systems

JP2024523848A5Pending Publication Date: 2025-06-13MIRONCOL DIAGNOSTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023576034
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-09
Filing Date
2022-06-07
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Current cancer detection methods lack non-invasive tests that can simultaneously and accurately detect multiple cancer types with high specificity, particularly those requiring greater than 99% specificity for screening the general population.

Method used

A method utilizing a miRNA biomarker set, including hsa-miR-5100 and optionally other top-ranked miRNAs, to determine the expression profile from biological samples like blood, serum, or saliva, calculating a diagnostic index based on miRNA levels to classify the presence of cancers such as lung, biliary tract, bladder, colorectal, esophageal, gastric, glioma, liver, pancreatic, prostate, ovarian, and sarcoma, using statistical models for accuracy.

Benefits of technology

The method achieves diagnostic accuracy with AUC values ranging from 0.780 to 0.999, sensitivity of 68-99%, and specificity of 99-100%, effectively detecting multiple cancer types with high precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

SOLUTION: Provided herein are methods, kits and systems capable of detecting one or more human cancers with high accuracy. Based on a liquid biopsy sample from a subject, an expression profile of a miRNA biomarker set consisting of one or more miRNAs is determined, and then a diagnostic index is calculated based on which the subject is classified as having or not having cancer. The 4-miRNA biomarker model shows extremely high sensitivity of 99.0-100% for lung cancer and gastric cancer, 83.0-99.0% for biliary tract cancer, bladder cancer, colon cancer, esophageal cancer, glioma, liver cancer, pancreatic cancer, and prostate cancer, and 68.2-72.0% for ovarian cancer and sarcoma, while maintaining a specificity of 99.3%.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 208,506, filed June 9, 2021, the disclosure of which is incorporated herein by reference in its entirety.

[0002] The contents of the electronically submitted sequence listing (file name Top_miRNA_Seq.txt, size 15,063 bytes, created May 31, 2022) are incorporated by reference herein.

[0003] The present invention relates generally to the technical field of disease screening, detection and diagnosis, and more specifically to methods, kits, systems and non-transitory storage media for detecting cancer in one or more humans. [Background technology]

[0004] Despite the rapid development of diagnostic and therapeutic techniques in recent years, cancer remains a challenging and deadly disease for humans. It is well known that early detection of cancer is crucial to reduce cancer-related mortality, since early detection increases the chance of successful treatment. There is an urgent unmet need to develop tests that are ideally non-invasive and capable of detecting multiple cancer types early and simultaneously, such as blood tests that are the basis of the so-called multi-cancer early detection (MCED) paradigm. Such MCED tests often require very high specificity, preferably 99% or higher, to minimize false positives, in order to be able to screen the general population at risk.

[0005] miRNAs are small, single-stranded, non-coding RNA molecules with an average length of 22 nucleotides that are encoded by corresponding genes in the human genome. miRNAs function in the negative post-transcriptional regulation of gene expression, mainly by binding to complementary sequences in the 3' untranslated region (3'UTR) of mRNA molecules. miRNAs appear to regulate more than 50% of human genes, and aberrant expression of miRNAs is involved in many human cancers. Combined with their remarkable stability in blood and other body fluids, circulating extracellular miRNAs have the potential to serve as non-invasive biomarkers for cancer screening and diagnosis. Summary of the Invention [Means for solving the problem]

[0006] The present disclosure provides a multi-cancer detection approach (i.e., method, kit, and system) with a miRNA biomarker set consisting of at least one miRNA biomarker. The approach is substantially based on the expression profile of the miRNA biomarker set that can be determined from a biological sample obtained from a human subject. Such biological sample can be a liquid biopsy sample, including blood sample, serum sample, plasma sample, urine sample, saliva sample, or sputum sample, among others, thereby allowing non-invasive or minimally invasive detection of cancer. The approach can be employed to accurately and reliably detect whether a human subject is suffering from any of the following cancers, including lung cancer, biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma cancer, liver cancer, pancreatic cancer, prostate cancer, ovarian cancer, and sarcoma.

[0007] In one aspect, a method for detecting cancer in a biological sample obtained from a subject is provided. The method essentially comprises the following three steps (1) to (3):

[0008] Step (1): determining an expression profile of a miRNA biomarker set consisting of at least one miRNA from a biological sample, where the miRNA biomarker set consists of hsa-miR-5100.

[0009] Step (2): Calculate a diagnostic index for the biological sample based on the expression profile of the miRNA biomarker set. The diagnostic index is as follows:

number

[0010] Step (3): Classifying the subject as having or not having cancer based on the value of the calculated diagnostic index. If the calculated diagnostic index is equal to or greater than a predetermined threshold, the subject is classified as having cancer, and if not, the subject is classified as not having cancer.

[0011] Additionally, the method is configured to achieve a diagnostic accuracy having an AUC value of greater than about 0.780.

[0012] As used herein, an expression profile of a miRNA biomarker set is essentially a data set that includes expression level data determined for each miRNA member included in the miRNA biomarker set.

[0013] The term "predetermined threshold" refers to the cut-point value of the diagnostic index that can be used to determine whether a subject is affected by a cancer type with a given specificity / sensitivity. This is typically predetermined based on an existing data set consisting of a range of diagnostic index values ​​obtained and calculated for an existing population of subjects known to have disease and / or known not to have disease. For example, in Example 1 provided below, when the miRNA biomarker set consists of any one of the top 100 miRNAs (corresponding to SEQ ID NO: S: 1-100), the AUC can reach a level greater than 0.780 (i.e., for hsa-miR-1238-5p), and even reach about 0.999 (i.e., for the top four miRNAs: hsa-miR-5100, hsa-miR-1343-3p, hsa-miR-1290 and hsa-miR-4787-3p) (see Table 1).

[0014] According to some embodiments of the method, the miRNA biomarker set includes, in addition to hsa-miR-5100 (corresponding to SEQ ID NO: 1), one or more of the other 99 miRNAs listed in Table 1, namely, hsa-miR-1343-3p, hsa-miR-1290, hsa-miR-4787-3p, hsa-miR-6877-5p, hsa-miR-17-3p, hsa-miR-6765-5p, hsa-miR-1268b, hsa-miR-4258, hsa-miR-451a, hsa-miR-1228-5p ... miR-8073, hsa-miR-4454, hsa-miR-187-5p, hsa-miR-4286, hsa-miR-6746-5p, hsa-miR-663b, hsa-miR-6075, hsa-miR-5001-5p, hsa-miR-6789-5p, hsa-miR-4513, hsa-miR-3192-5p, hsa-miR-8060, hsa-miR-668-5p, hsa-miR-1268a, hsa-miR-1273g-3p, hsa-miR-4706, hsa-miR-124-3p, hsa-miR- 1260b, hsa-miR-4740-5p, hsa-miR-320b, hsa-miR-7977, hsa-miR-29b-3p, hsa-miR-4708-3p, hsa-miR-4525, hsa-miR-92b-3p, hsa-miR-4257, hsa -miR-4727-3p, hsa-miR-92a-3p, hsa-miR-663a, hsa-miR-6787-5p, hsa-miR-3131, hsa-miR-6802-5p, hsa-miR-654-5p, hsa-miR-6511b-5p, hsa-mi R-29b-1-5p, hsa-miR-4417, hsa-miR-4736, hsa-miR-6840-3p, hsa-miR-4710, hsa-miR-4635, hsa-miR-296-3p, hsa-miR-1199-5p, hsa-miR-7975, h sa-miR-4480, hsa-miR-3648, hsa-miR-371a-5p, hsa-miR-4771, hsa-miR-6717-5p, hsa-miR-1254, hsa-miR-1246, hsa-miR-23b-3p, hsa-miR-320a,hsa-miR-4687-5p, hsa-miR-191-5p, hsa-miR-320c, hsa-miR-6131, hsa-miR-4515, h sa-miR-342-5p, hsa-miR-4718, hsa-miR-23a-3p, hsa-miR-4455, hsa-miR-211-3p, h sa-miR-3122, hsa-miR-103a-3p, hsa-miR-4429, hsa-miR-920, hsa-miR-3194-3p, hs a-miR-4754, hsa-miR-1238-5p, hsa-miR-3191-3p, hsa-miR-4755-3p, hsa-miR-3688- 5p, hsa-miR-4529-5p, hsa-miR-6861-5p, hsa-miR-1469, hsa-miR-619-5p, hsa-miR-4448, hsa-miR-4658, hsa-miR-22-3p, hsa-miR-4776-5p, hsa-miR-320e, hsa-miR-1225-3p, hsa-miR-6875-5p, hsa-miR-4534, hsa-miR-4652-5p, hsa-miR-648, hsa-miR-4259, hsa-miR-107, and hsa-miR-650 are ranked based on adjusted P-values ​​and correspond to SEQ ID NOs: 2-100, respectively.

[0015] According to some other embodiments of the method, the miRNA biomarker set includes, in addition to hsa-miR-5100, one or more of the other top 50 miRNAs listed in Table 1, namely, hsa-miR-1343-3p, hsa-miR-1290, hsa-miR-4787-3p, hsa-miR-6877-5p, hsa-miR-17-3p, hsa-miR-6765-5p, hsa-miR-1268b, hsa-miR-4258, hsa-miR -451a, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-4454, hsa-miR-187-5p, hsa-miR-4286, hsa-miR-6746-5p, hsa-miR-663b, hsa-miR-6075, hsa-miR-5001-5p, hsa-miR-6789-5p, hsa-miR-4513, hsa-miR-3192-5p, hsa-miR-8060, hsa-miR-668-5p, hsa -miR-1268a, hsa-miR-1273g-3p, hsa-miR-4706, hsa-miR-124-3p, hsa-miR-1260b, hsa-miR-4740-5p, hsa-miR-320b, hsa-mi R-7977, hsa-miR-29b-3p, hsa-miR-4708-3p, hsa-miR-4525, hsa-miR-92b-3p, hsa-miR-4257, hsa-miR-4727-3p, hsa-miR-92 a-3p, hsa-miR-663a, hsa-miR-6787-5p, hsa-miR-3131, hsa-miR-6802-5p, hsa-miR-654-5p, hsa-miR-6511b-5p, hsa-miR-29b-1-5p, hsa-miR-4417, hsa-miR-4736, hsa-miR-6840-3p, and hsa-miR-4710 are ranked based on adjusted P-values ​​and correspond to SEQ ID NOs:2-50, respectively.

[0016] According to some other embodiments of the method, the miRNA biomarker set includes, in addition to hsa-miR-5100, one or more of the other top 20 miRNAs listed in Table 1, namely, hsa-miR-1343-3p, hsa-miR-1290, hsa-miR-4787-3p, hsa-miR-6877-5p, hsa-miR-17-3p, hsa-miR-6765-5p, hsa-miR-1268b, hsa-miR-4258, hsa-miR-451a, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-4454, hsa-miR-187-5p, hsa-miR-4286, hsa-miR-6746-5p, hsa-miR-663b, hsa-miR-6075, hsa-miR-5001-5p, and hsa-miR-6789-5p, ranked based on adjusted P-values ​​and corresponding to SEQ ID NOs: S:2-20, respectively. wherein, further optionally, the miRNA biomarker set consists of the top 20 miRNAs listed in Table 1 (corresponding to SEQ ID NOs: S:1-20, respectively).

[0017] According to some other embodiments of the method, the miRNA biomarker set further includes, in addition to hsa-miR-5100, one or more of the other top four miRNAs listed in Table 1, namely, hsa-miR-1343-3p, hsa-miR-1290, and hsa-miR-4787-3p, which are ranked based on adjusted P-values ​​and correspond to SEQ ID NOs: 2-4, respectively. Further optionally herein, the miRNA biomarker set consists of the top four miRNAs listed in Table 1, namely, hsa-miR-5100, hsa-miR-1343-3p, hsa-miR-1290, and hsa-miR-4787-3p, which correspond to SEQ ID NOs: 1-4, respectively.

[0018] The method can optionally be further configured to achieve diagnostic accuracy with higher AUC values.

[0019] According to some embodiments, the method is configured to achieve a diagnostic accuracy having an AUC value of greater than about 0.850, where optionally, the cancer that may be detected may be selected from the group consisting of lung cancer, biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma cancer, liver cancer, pancreatic cancer, prostate cancer, ovarian cancer, and sarcoma.

[0020] According to some embodiments, the method is configured to achieve a diagnostic accuracy having an AUC value of greater than about 0.950, where optionally, the cancers that may be detected are comprised of the group consisting of lung cancer, biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma cancer, liver cancer, ovarian cancer, pancreatic cancer, and prostate cancer.

[0021] According to some embodiments, the method is configured to achieve a diagnostic accuracy having an AUC value of greater than about 0.990, where optionally, the cancer that may be detected may be selected from the group consisting of lung cancer, biliary tract cancer, bladder cancer, esophageal cancer, gastric cancer, glioma cancer, and prostate cancer.

[0022] According to some embodiments, the method is configured to achieve a diagnostic accuracy having an AUC value of greater than about 0.999, wherein optionally, the cancer that may be detected may be lung cancer or gastric cancer.

[0023] According to different practical needs, the method can be optionally configured to achieve diagnostic accuracy with different sensitivity and specificity levels.

[0024] According to some embodiments, the method is configured to achieve a diagnostic accuracy having a sensitivity of greater than about 68.0% while having a specificity of greater than about 99.0%, where optionally, the cancers that may be detected are comprised of the group consisting of lung cancer, biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma cancer, liver cancer, pancreatic cancer, prostate cancer, ovarian cancer, and sarcoma.

[0025] According to some embodiments, the method is configured to achieve a diagnostic accuracy having a sensitivity greater than about 83.0% while having a specificity greater than about 99.0%, where optionally, the cancers that may be detected are comprised of the group consisting of lung cancer, biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma cancer, liver cancer, pancreatic cancer, and prostate cancer.

[0026] According to some embodiments, the method is configured to achieve a diagnostic accuracy having a sensitivity of greater than about 99.0% and a specificity of greater than about 99.0%, where optionally, the cancer that can be detected can be lung cancer or gastric cancer.

[0027] According to some embodiments of the method, in step (2) of calculating a diagnostic index for the biological sample based on the expression profile of the miRNA biomarker set, the diagnostic index is calculated via an unweighted model.

[0028] According to some other embodiments of the method, in step (2) of calculating a diagnostic index of the biological sample based on the expression profile of the miRNA biomarker set, the diagnostic index is calculated by a weighting model using weights from one selected from the group consisting of a logistic regression model, a linear discriminant analysis (LDA) model, a conditional logistic regression model, a lasso regression model, a ridge regression model, a random forest, a support vector machine, and a probit regression model, which is calculated via a weighting model using weights from one of the group consisting of Linear Models for Microarray Data (limma) models. Optionally, the diagnostic index is calculated by a weighting model using weights from the limma model.

[0029] As used herein, the terms "unweighted model" and "weighted model" shall be understood within the scope of the general definition well understood by those skilled in the art. The term "unweighted model" refers to a situation in which no weight is applied to each miRNA in the miRNA biomarker set when calculating the diagnostic index. Within the scope of this disclosure, with reference to formula (I), the phrase "the diagnostic index is calculated via an unweighted model" refers to a weighting factor that is equal for any miRNA in the miRNA biomarker set. i (For example, t i = 1). The term "weighting model" refers to a situation where a weight corresponding to each miRNA in the miRNA biomarker set is applied when calculating the diagnostic index. Within the scope of the present disclosure, with reference to formula (I), the phrase "the diagnostic index is calculated via a weighting model" refers to a weighting model that is applied to all t i It can be understood that the weights are not equal (i.e., there are at least two miRNAs with different weights).

[0030] Each of the terms, "Linear Models for Microarray Data (limma) model" (Ritchie et al. 2015), "logistic regression model" (Venable and Ripley 2002), "linear discriminant analysis (LDA) model" (Venable and Ripley 2002), "conditional logistic regression model" (Venable and Ripley 2002), "lasso regression model" (Tibshirani 1996), "ridge regression model" (Hoerl and Kennard 1970), "random forest" (Ripley 1996), "support vector machine" (Ripley 1996), and "probit regression model" (Venable and Ripley 2002), are essentially probability models and statistical models that follow definitions commonly appreciated by those skilled in the art, the details of which can be found in the references included immediately below.

[0031] For convenience, according to some embodiments, after step (2) and before step (3), the method may further include a normalization step of: obtaining a normalized diagnostic index based on the calculated diagnostic index. Correspondingly, step (3) is: classifying the subject as having cancer if the normalized diagnostic index is equal to or greater than the pre-set cut point; or classifying the subject as not having cancer if not.

[0032] Here, there can be different ways for the normalization step. According to some embodiments, the normalized diagnostic index is represented by formula (II):

number

[0033] More specifically, param location is a position parameter configured to shift the minimum value of the normalized diagnostic index to a first preset value, and param scale is a scale parameter configured to effectively scale the maximum value of the normalized diagnostic index to a second value, where the first and second preset values ​​are thus the minimum and maximum values, respectively, of a range of calculated normalized diagnostic index values ​​obtained from an existing population of subjects known to have cancer and subjects known not to have cancer, excluding outliers.

[0034] Optionally, multiple settings can be applied. For example, in the existing data set of EXAMPLE 1 below, where the diagnostic index values ​​have been determined to have a range from 600 to 1600 excluding outliers (reference), to shift the range between 0 (i.e. the first preset value) and 10 (i.e. the second current value), the param location and param scale can be set to 600 and 100, respectively. This normalization scheme is also adopted in Example 1 below.

[0035] Or, param location and param scale Alternatively, we can set param location and param scale Alternatively, we can set param location and param scale can be set to 350 and 250, respectively, and the final normalized diagnostic index can be set to 1 to 5.

[0036] In embodiments where the normalized diagnostic index is normalized to be between 0 and 10, the preset cut point can optionally be set as 5.1, thereby causing the method to have a specificity of greater than about 0.95, or can optionally be set as 6.0, thereby causing the method to have a specificity of greater than about 0.99.

[0037] In any embodiment of the above-described methods, the biological sample is a liquid biopsy sample from the group consisting of a blood sample, a serum sample, a plasma sample, a urine sample (Yun et al., 2012), a saliva sample (Park et al., 2009), and a sputum sample.

[0038] In any embodiment of the method as described above, in step (1) of determining an expression profile of a miRNA biomarker set consisting of at least one miRNA from a biological sample, the expression profile of the miRNA biomarker set can optionally be obtained by Northern blotting, by microarray analysis, RNA sequencing, or RNA in situ hybridization, or optionally by a nucleic acid amplification procedure consisting of reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR.

[0039] As used herein, each of the above miRNA detection approaches is to be understood within the general definition well understood by those skilled in the art. Details for carrying out these approaches to determine the expression profile of a miRNA biomarker set are provided below.

[0040] In any embodiment of the method as described above, the method optionally further comprises performing an evaluation of the subject, said evaluation comprising diagnosing cancer or detecting recurrence of cancer.

[0041] As used herein, a "diagnosis of cancer" refers to the discovery of cancer in a subject who was previously known to be cancer-free, and a "recurrence of cancer" refers to the re-discovery of cancer in a subject who previously underwent treatment to remove the cancer and is no longer cancer-free.

[0042] In any embodiment of the method as described above, the method optionally further comprises administering a therapeutic regimen to the subject if the subject is classified as having cancer. Here, the method can administer a variety of known therapeutic regimens, including surgery, radiation therapy, chemotherapy, hormonal therapy, targeted therapy, immunotherapy, or a combination thereof. These therapeutic regimens are well established for each of the different cancers listed above.

[0043] In any embodiment of the method as described above, the method optionally further comprises performing a diagnostic procedure on the subject if the subject is classified as having cancer, where the diagnostic procedure may optionally include a physical examination, pathological examination of a biopsy from the subject, immunohistochemistry, or an imaging test such as an X-ray, computed tomography (CT), ultrasound, and / or magnetic resonance imaging.

[0044] In a second aspect, the present disclosure further provides a kit for detecting cancer from a biological sample obtained from a subject, the kit being substantially adapted to perform the method according to the first aspect.

[0045] As used herein and elsewhere in this disclosure, the term "kit" refers to a collection of items and / or instructions. The items included in the kit may be physical entities or components thereof. Examples of items that may be included in the kit as disclosed herein may include one or more nucleic acids (e.g., polynucleotides), or one or more devices, apparatuses or instruments (e.g., molecular arrays or microarrays that include one or more nucleic acids). The instructions included in the kit may be descriptions (e.g., manuals) of specific steps to be performed, and may be printed on a physical medium (e.g., paper, card, etc.), a computer-readable storage medium (e.g., hard disk, compact disk or CD, flash drive, etc.), or may be stored on the Internet (e.g., accessible cloud space), etc.

[0046] The kit includes at least the following components (1) and (2) (i.e., articles and / or instructions):

[0047] Component (1): at least one nucleic acid capable of specifically recognizing each miRNA in a miRNA biomarker set, thereby obtaining an expression profile of the miRNA biomarker set from a biological sample, wherein the miRNA biomarker set comprises hsa-miR-5100 (SEQ ID NO: 1).

[0048] Component (2): At least one instruction including a first instruction and a second instruction, wherein the first instruction includes a first sub-instruction for calculating a diagnostic index for the biological sample based on an expression profile of the miRNA biomarker set, and the diagnostic index is represented by the formula:

number

[0049] Here, in component (1) of the kit, at least one nucleic acid may optionally comprise a polynucleotide capable of specifically hybridizing under stringent conditions to either: (a) SEQ ID NO:1, its derivative, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides, or (b) a polynucleotide comprising or consisting of a nucleotide sequence complementary to the nucleotide sequence of SEQ ID NO:1, its derivative, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides.

[0050] According to some embodiments of the kit, the miRNA biomarker set further comprises, in addition to hsa-miR-5100, one or more of the other 99 miRNAs listed in Table 1. Correspondingly, in the kit component (1), the at least one nucleic acid may optionally further comprise at least one polynucleotide capable of specifically hybridizing under stringent conditions to either (a) a polynucleotide comprising or consisting of any one of the nucleotide sequences of SEQ ID NOs: S.2-100, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides, or (b) a polynucleotide comprising or consisting of a nucleotide sequence complementary to any one of the nucleotide sequences of SEQ ID NOs: S.2-100, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides.

[0051] According to some embodiments of the kit, the miRNA biomarker set further comprises, in addition to hsa-miR-5100, one or more of the other top 50 miRNAs listed in Table 1. Correspondingly, in the component (1) of the kit, the at least one nucleic acid can optionally further comprise at least one polynucleotide capable of specifically hybridizing under stringent conditions to either (a) a polynucleotide comprising or consisting of any one of the nucleotide sequences of SEQ ID NOs: 2-50, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides, or (b) a polynucleotide comprising or consisting of a nucleotide sequence complementary to any one of the nucleotide sequences of SEQ ID NOs: 2-50, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides.

[0052] According to some embodiments of the kit, the miRNA biomarker set further comprises, in addition to hsa-miR-5100, one or more of the other top 20 miRNAs listed in Table 1. Correspondingly, in the kit component (1), the at least one nucleic acid can optionally further comprise at least one polynucleotide capable of specifically hybridizing under stringent conditions to either (a) a polynucleotide comprising or consisting of any one of the nucleotide sequences of SEQ ID NOs: 2-20, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides, or (b) a polynucleotide comprising or consisting of a nucleotide sequence complementary to any one of the nucleotide sequences of SEQ ID NOs: 2-20, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides.

[0053] Further optionally, the miRNA biomarker set herein consists of the top 20 miRNAs in Table 1, and correspondingly, in component (1) of the kit, at least one nucleic acid may further comprise a total of 20 polynucleotides capable of specifically hybridizing under stringent conditions to either a) a polynucleotide comprising or consisting of any one of the nucleotide sequences of SEQ ID NOs: 1 to 20, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides, or (b) a polynucleotide comprising or consisting of a nucleotide sequence complementary to any one of the nucleotide sequences of SEQ ID NOs: 1 to 20, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides.

[0054] According to some embodiments of the kit, the miRNA biomarker set further comprises, in addition to hsa-miR-5100, one or more of the other top four miRNAs listed in Table 1. Correspondingly, in the kit component (1), the at least one nucleic acid can optionally further comprise at least one polynucleotide capable of specifically hybridizing under stringent conditions to either (a) a polynucleotide comprising or consisting of any one of the nucleotide sequences of SEQ ID NOs: S: 2-4, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides, or (b) a polynucleotide comprising or consisting of a nucleotide sequence complementary to any one of the nucleotide sequences of SEQ ID NOs: S: 2-4, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides.

[0055] Here, optionally, the miRNA biomarker set consists of the top four miRNAs in Table 1, namely, hsa-miR-5100, hsa-miR-1343-3p, hsa-miR-1290, and hsa-miR-4787-3p, and correspondingly, in component (1) of the kit, at least one nucleic acid may further comprise a total of four polynucleotides capable of specifically hybridizing under stringent conditions to either (a) a polynucleotide comprising or consisting of a nucleotide sequence of SEQ ID NO:S:1 to 4, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof each comprising 15 or more consecutive nucleotides, or (b) a polynucleotide comprising or consisting of a nucleotide sequence complementary to a nucleotide sequence of SEQ ID NO:S:1 to 4, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof each comprising 15 or more consecutive nucleotides.

[0056] In the kit, in the first sub-instructions of the first instructions of component (2), the diagnostic index can be calculated via an unweighted model, or alternatively, via a weighted model using weights from one of the probability modeling statistical models provided above in the first aspect, where according to some embodiments of the kit, the diagnostic index is calculated via a weighted model using weights from the Limma model.

[0057] According to some embodiments of the kit, the predetermined threshold may be set as 1110 and the second instructions further indicate that classification using 1110 as the predetermined threshold has a specificity of greater than 0.95. According to some other embodiments of the kit, the predetermined threshold may be set as 1200 and the second instructions further indicate that such classification using 1200 as the predetermined threshold has a specificity of greater than 0.99.

[0058] According to some embodiments of the kit, the first instruction further includes a second sub-instruction for obtaining a normalized diagnostic index based on the diagnostic index calculated according to the first sub-instruction, and in the second instruction, the subject is classified as having cancer if the normalized diagnostic index is equal to or greater than a preset cut point, and is otherwise classified as not having cancer. This normalization process is substantially the same as the normalization process described in the first method aspect above, and will not be described here.

[0059] Optionally, the normalized diagnostic index is calculated via a weighted model using weights from the Limma model, with a first preset value being 0 and a second preset value being 10. Additionally, the preset cut points can be optionally set as 5.1 or 6.0, such that classification using the preset cut points has a specificity greater than 0.95 or 0.99, respectively.

[0060] According to different embodiments, the instructions for at least one of the components (2) of the kit can further include third instructions for performing an evaluation of a subject, said evaluation including diagnosing cancer or detecting recurrence of cancer; or can further include fourth instructions for administering a therapeutic regimen to the subject if the subject is classified as having cancer.

[0061] According to some embodiments, at least one of the instructions of component (2) in the kit may further comprise a first additional instruction for obtaining an expression profile of the miRNA biomarker set, which comprises a procedure for performing Northern blotting, microarray analysis, RNA sequencing, or RNA in situ hybridization with the at least one nucleic acid, where the at least one nucleic acid may optionally be disposed on a molecular array.

[0062] According to some embodiments, the kit may further comprise at least one set of amplification primers, each set capable of specifically amplifying each of at least one miRNA in the miRNA biomarker set from a biological sample. Thus, at least one instruction for component (2) in the kit may further comprise a second additional instruction for obtaining an expression profile of the miRNA biomarker set, which includes a procedure for performing reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR with at least one nucleic acid and at least one set of amplification primers.

[0063] In any embodiment of the kit as described above, the biological sample may be a liquid biopsy sample selected from the group consisting of a blood sample, a serum sample, a plasma sample, a urine sample, a saliva sample, and a sputum sample.

[0064] In a third aspect, the present disclosure further provides a system for detecting cancer in a subject, wherein the system is a computerized system consisting essentially of a collection of hardware (e.g., a processor, a memory, an I / O interface, a storage medium, etc.) and software (i.e., a computer program, including an operation system software and specific program software, etc.), which are configured to cooperate to collectively perform all or some steps of the method as described above in the first aspect. According to some embodiments, the system includes a processor and a non-transitory storage medium. The non-transitory storage medium is configured to include software (i.e., program instructions) for execution by the processor, and the program instructions are configured to cause the processor to perform various steps of the method according to various different embodiments of the method described above in the first aspect.

[0065] In a fourth aspect, the present disclosure further provides a non-transitory storage medium configured to store computer-executable program instructions which, when executed by a processor, cause the processor to perform methods according to various different embodiments of the method described above in the first aspect.

[0066] There may be various different embodiments of the above-mentioned system and non-transitory storage medium with respect to the following elements / features, including what miRNA components are included in the miRNA biomarker set, whether and how normalization is performed for the diagnostic index, how subjects are classified as having or not having cancer, what samples can be used for the biological sample, what detection accuracy level is achieved, etc. Specific details about these different embodiments may be referred to the various embodiments of the method as described in the first aspect, and are omitted herein for brevity.

[0067] Unless otherwise defined, definitions of terms used throughout this disclosure are as follows:

[0068] Generally, this refers to mammals such as humans, primates such as chimpanzees, pets such as dogs and cats, farm animals such as cows, horses, sheep and goats, and rodents such as mice and rats. Also, "healthy subjects" refers to mammals that do not have cancer to be detected. It should be noted that the entire disclosure, although more specifically directed to human subjects, may also be applied to other non-human mammals, if desired.

[0069] Terms or abbreviations such as "nucleic acid," "nucleotide," "polynucleotide," "DNA," "RNA," and "miRNA" follow common usage in the art unless otherwise indicated or defined.

[0070] As used herein, the term "polynucleotide" is interchangeable with "nucleic acid" and refers to nucleic acids including RNA, DNA, and RNA / DNA (chimeras). DNA includes cDNA, genomic DNA, and synthetic DNA. RNA includes total RNA, mRNA, rRNA, miRNA, siRNA, snoRNA, snRNA, non-coding RNA, and synthetic RNA.

[0071] As used herein, the term "fragment" refers to a polynucleotide having a nucleotide sequence that has a contiguous portion of a polynucleotide, desirably having a length of 15 or more nucleotides, for example 15, 16, 17, 18, 19, etc. nucleotides.

[0072] In this specification, the term "gene" is intended to include not only RNA and double-stranded DNA, but also each single-stranded DNA such as a plus strand (or sense strand) or a complementary strand (or antisense strand) that constitutes the double strand. The length of the gene is not particularly limited. In this specification, unless otherwise specified, "gene" includes double-stranded DNA including human genomic DNA, single-stranded DNA (plus strand) including cDNA, single-stranded DNA (complementary strand) having a base sequence complementary to the plus strand, miRNA (microRNA) and all of their fragments and transcription products. "Gene" includes not only "gene" represented by a specific base sequence (or sequence ID number), but also "nucleic acid" that codes for RNA having biological functions equivalent to the RNA encoded by the gene, such as congeners (i.e., homologs or orthologs), variants (e.g., genetic polymorphisms), and derivatives. Specific examples of "nucleic acids" encoding such congeners, variants, or derivatives include "nucleic acids" having a nucleotide sequence that hybridizes under stringent conditions described below with a complementary sequence of a nucleotide sequence represented by any of SEQ ID NOs: 1 to 100, or a nucleotide sequence derived from the nucleotide sequence by substituting the nucleotide "U" (or "u") with the nucleotide "T" (or "t"). The "gene" is not particularly limited by its functional region, and may include, for example, an expression control region, a coding region, an exon, or an intron. The "gene" may be contained within a cell, or may be released outside the cell and exist alone. Alternatively, the "gene" may be in a state encapsulated in a vesicle called an exosome.

[0073] Within the scope of the entire disclosure, the term "microRNA (miRNA)" is intended to mean a 15-25 nucleotide non-coding RNA that is transcribed as an RNA precursor having a hairpin-like structure, cleaved by a dsRNA cleaving enzyme having RNase III cleavage activity, incorporated into a protein complex called RISC, and involved in suppressing the translation of mRNA, but is not otherwise specified. The term "miRNA" as used herein includes not only "miRNA" represented by a specific nucleotide sequence (or sequence ID number), but also precursors of "miRNA" (pre-miRNA or pri-miRNA), and miRNAs having equivalent biological functions thereto, such as congeners (i.e., homologs or orthologs), variants (e.g., gene polymorphisms), and derivatives. Such precursors, conjugates, variants, or derivatives can be specifically identified using miRBase Release 20 (Kozomara and Griffiths-Jones, 2010), and examples thereof include "miRNAs" having a nucleotide sequence that hybridizes under stringent conditions described below with the complementary sequence of any specific nucleotide sequence represented by any of SEQ ID NOs: 1-100. As used herein, the term "miRNA" may refer to the gene product of a miRNA gene. Such gene products include mature miRNAs (e.g., 15-25 nucleotide or 19-25 nucleotide non-coding RNAs involved in translational repression of mRNAs as described above) or miRNA precursors (e.g., pre-miRNA or pri-miRNA).

[0074] As used herein, the term "probe" includes a polynucleotide used to specifically detect RNA resulting from expression of a gene, or a polynucleotide derived from an RNA, and / or a polynucleotide complementary thereto.

[0075] As used herein, the term "primer" or "amplification primer" includes a polynucleotide that specifically recognizes and amplifies RNA resulting from expression of a gene or RNA-derived polynucleotide, and / or a polynucleotide complementary thereto.

[0076] In this context, a complementary polynucleotide (complementary or reverse strand) consists of a nucleotide sequence defined in any of SEQ ID NOs. 1-100, or a nucleotide sequence derived from this nucleotide sequence by substituting nucleotides "U" (or "u") for nucleotides "T" (or "t"), or a subsequence thereof (herein, this full length or subsequence is referred to as a plus strand for convenience). However, such a complementary strand is not limited to a sequence that is completely complementary to the nucleotide sequence of the target plus strand, but may have a complementary relationship to the extent that hybridization to the target plus strand is possible under stringent conditions.

[0077] As used herein, "stringent conditions" refers to conditions under which a nucleic acid probe hybridizes to a target sequence to a greater extent than other sequences (for example, a measured value of at least the average value of background measurements plus twice the standard deviation of background measurements). Stringent conditions vary depending on the sequence and also on the environment in which hybridization is performed. A target sequence that is 100% complementary to a nucleic acid probe can be identified by controlling the stringent hybridization conditions and washing conditions. Specific examples of "stringent conditions" will be described later.

[0078] As used herein, the term "variant" refers, in the case of nucleic acids, to naturally occurring variants resulting from polymorphisms, mutations, etc.; A nucleotide sequence represented by any of SEQ ID NOs. 1 to 100, or a nucleotide sequence derived therefrom by substituting the nucleotide "U" (or "u") with the nucleotide "T" (or "t"), or a subsequence thereof; a variant comprising a deletion, substitution, addition or insertion of one or two or more nucleotides in the nucleotide sequence of a premature miRNA represented by any of SEQ ID NOs.; a variant comprising a deletion, substitution, addition or insertion of one or two or more nucleotides in the nucleotide sequence of a premature miRNA represented by any of SEQ ID NOs. Mutants containing insertions: the nucleotide sequence of an early onset miRNA of the sequence represented by 1 to 100, or a nucleotide sequence derived from said nucleotide sequence by substituting the nucleotide "U" (or "u") with the nucleotide "T" (or "t"), or a subsequence thereof; a variant exhibiting about 90% or more, about 95% or more, about 97% or more, about 98% or more, about 99% or more identity to each of these nucleotide sequences or their subsequences; or a nucleic acid that hybridizes under stringent conditions as defined above to a polynucleotide or oligonucleotide consisting of each of these nucleotide sequences or their subsequences. Mutants can be prepared using well-known techniques such as site-directed mutagenesis or PCR-based mutagenesis.

[0079] The term "percent (%) identity" can be determined with or without introduced gaps using the BLAST or FASTA based protein or gene search systems described above (Zhang et al., 2000; Altschul et al., 1990; Pearson et al., 1988).

[0080] The term "derivative" includes modified nucleic acids, e.g., derivatives labeled with a fluorophore or the like, derivatives containing modified nucleotides (e.g., nucleotides that have undergone, for example, halogen, alkyl such as methyl, methoxy, thio, alkoxy such as carboxymethyl, base rearrangements, double bond saturation, deamidation, replacement of an oxygen molecule with a sulfur atom, etc.), nucleotides including PNA (peptide nucleic acid; Nielsen et al., 1991), LNA (locked nucleic acid; Obika et al., 1998), etc.

[0081] The "nucleic acid" capable of specifically binding to a polynucleotide selected from the above miRNAs is a synthetic or prepared nucleic acid, specifically including a "nucleic acid probe" or a "primer." The "nucleic acid" is used directly or indirectly for detecting the presence or absence of cancer in a subject, diagnosing the severity, improvement, and treatment sensitivity of cancer, and screening of candidate substances useful for preventing, improving, and treating cancer. The "nucleic acid" includes nucleotides, oligonucleotides, and polynucleotides that can specifically recognize and bind to a transcription product represented by any of SEQ ID NOs: 1 to 100 or a synthetic cDNA nucleic acid thereof in a living body, particularly in a sample such as a body fluid (e.g., blood or urine), in relation to the occurrence of cancer. Based on the above-mentioned properties, the nucleotides, oligonucleotides, and polynucleotides of the present invention can be effectively used as a probe for detecting the gene expressed in a living body, tissue, cell, etc., or as a primer for amplifying the gene expressed in a living body.

[0082] As used herein, the term "detection" is interchangeable with the terms "testing," "measuring," or "detection or decision support." As used herein, the term "evaluation" is meant to include diagnosis or evaluation support based on test or measurement results.

[0083] As used within the scope of this disclosure, the terms "P-value", "accuracy", "AUC", "sensitivity", and "specificity" are generally understood to have common definitions that are well understood by those of skill in the art, and are specifically defined as follows:

[0084] The term "P value" or "P" can be considered interchangeable with "p value" or "p" and refers to the probability of observing a statistic in a statistical test that is more extreme than the statistic actually calculated from the data under the null hypothesis. Thus, the smaller the "P" or "P value", the more significant the difference between the comparison subjects.

[0085] "AUC" means the area under the Receiver Operating Characteristic curve. "Accuracy" means the value of (true positive number + true negative number) / (total number of cases). Accuracy indicates the ratio of correctly identified samples to all samples, and is the main index for evaluating detection performance.

[0086] As used herein, the term "sensitivity" refers to the value of (true positive number) / (true positive number+false negative number). High sensitivity allows cancer to be detected, leading to clinical therapeutic intervention.

[0087] In this specification, the term "specificity" means the value of (the number of true negatives) / (the number of true negatives + the number of false positives). High specificity can prevent unnecessary testing of healthy individuals who are mistakenly determined to be cancer patients, leading to a reduction in the burden on patients and medical costs.

[0088] Unless otherwise specified, available techniques that can be used to determine the expression profile of a miRNA biomarker set are summarized below.

[0089] It should be noted that determining the expression profile of miRNA biomarker set essentially includes determining the expression level of each and every miRNA contained in the miRNA biomarker set.Preferably, the expression levels of all miRNAs contained in the miRNA biomarker set can be determined simultaneously in one well-controlled experiment.However, optionally, the expression levels of these miRNAs can be determined in multiple experiments by different experimental procedures.

[0090] As used herein, measuring or detecting the expression of a miRNA included in a miRNA biomarker set includes measuring or detecting a nucleic acid transcript corresponding to the miRNA.

[0091] Typically, expression can be detected or measured based on miRNA or corresponding reverse transcription cDNA levels. Any quantitative or qualitative method for measuring RNA or cDNA levels can be used. Suitable methods for detecting or measuring miRNA or cDNA levels include, for example, Northern blotting, microarray analysis, RNA sequencing, RNA in situ hybridization, or nucleic acid amplification procedures such as reverse transcription PCR (RT-PCR) or real-time RT-PCR (also known as quantitative RT-PCR (qRT-PCR)), or digital RT-PCR. Such methods are well known in the art (e.g., Green and Sambrook et al.). Other techniques include digital multiplex analysis of gene expression, such as nCounter® (NanoString Technologies, Seattle, WA) gene expression assay, which is further described in US20100112710 and US20100047924.

[0092] Detection of a nucleic acid of interest generally requires hybridization of a target (e.g., miRNA or cDNA) with a probe. The sequences of miRNAs used in expression profiles of various cancer genes are known. Thus, a person skilled in the art can easily design hybridization probes for detecting those miRNAs (e.g., Green and Sambrook et al.). For example, polynucleotide probes that specifically bind to the miRNA transcripts (or cDNAs synthesized therefrom) described herein can be made by routine techniques (e.g., PCR or synthesis) using the nucleic acid sequence of the miRNA or cDNA target itself. As used herein, the term "probe" refers to a portion or part of a polynucleotide sequence consisting of about 10 or more consecutive nucleotides, about 15 or more consecutive nucleotides, about 20 or more consecutive nucleotides. In certain embodiments, the polynucleotide probe comprises 10 or more nucleic acids, 15 or more nucleic acids, or 20 or more nucleic acids. To confer sufficient specificity, the probes may have about 90% or greater sequence identity to the complement of the target sequence, such as about 95% or greater (e.g., about 98% or greater or about 99% or greater), as determined, for example, using the well-known Basic Local Alignment Search Tool (BLAST) algorithm (available through the National Center for Biotechnology Information (NCBI), Bethesda, Md.).

[0093] Each probe may be substantially specific for its target to avoid cross-hybridization and false positives. An alternative to using specific probes is to use specific reagents to obtain the material from the transcript (e.g., using target-specific primers during cDNA production or amplification). In both cases, specificity can be obtained by hybridization to a portion that is substantially unique within the group of miRNAs being analyzed. If a target has multiple splice variants, it is possible to design hybridization reagents that recognize a region common to each of the variants and / or to use multiple reagents, each recognizing one or more variants.

[0094] Stringent conditions for hybridization reactions can be readily determined by one of skill in the art and are generally an empirical calculation that depends on the length of the probe, the washing temperature, and the salt concentration. Longer probes generally require higher temperatures for proper annealing, whereas shorter probes require lower temperatures. Hybridization generally depends on the ability of denatured nucleic acid sequences to reanneal when complementary strands are present in an environment below their melting temperature. The higher the homology between the probe and the hybridizable sequence, the higher the relative temperature that can be used. As a result, higher relative temperatures result in more stringent reaction conditions, whereas lower relative temperatures result in more lenient reaction conditions.

[0095] As defined herein, "stringent conditions" or "even more stringent conditions" are specified by, but are not limited to, the following: (1) the use of low ionic strength and high temperature for washing, e.g., 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate at 50° C.; (2) the use of a denaturing agent, e.g., formamide, during hybridization, e.g., 50% (v / v) formamide with 0.1% bovine serum albumin / 0.1% Ficoll / 0.1% polyvinylpyrrolidone / 50 mM sodium phosphate buffer, pH 6.5, 750 mM sodium chloride; or (3) 50% formamide, 5x SSC (0.75M NaCl, 0.075M sodium citrate), 50mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5x Denhardt's solution, sonicated salmon sperm DNA (50μg / ml), 0.1% SDS, and 10% dextran sulfate at 42°C, 0.2x SSC (sodium chloride / sodium citrate) and 50% formamide at 42°C, 55°C, followed by more stringent washes in 0.1x SSC with EDTA at 55°C. "Moderately stringent conditions" include, but are not limited to, those described in Sambrook et al., 1989, and include the use of wash solutions and hybridization conditions (e.g., temperature, ionic strength, and % SDS) that are less stringent than those described above. An example of moderately stringent conditions is an overnight incubation at 37° C. in a solution consisting of 20% formamide, 5× SSC (150 mM NaCl, 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5× Denhardt's solution, 10% dextran sulfate, and 20 mg / mL denatured sheared salmon sperm DNA, followed by washing the filter in 1× SSC at approximately 37-50° C. One of skill in the art will know how to adjust temperature, ionic strength, etc., as necessary to accommodate factors such as probe length.

[0096] In certain embodiments, microarray analysis, Northern blot, RNA in situ hybridization, or PCR-based methods are used. In this regard, measuring the expression of the aforementioned miRNAs in a biological sample may, for example, consist of contacting a sample containing or suspected to contain cancer cells with a polynucleotide probe specific for the miRNA of interest, or a primer designed to amplify a portion of the miRNA of interest, and detecting the binding of the probe to the nucleic acid target or the amplification of the nucleic acid, respectively. Detailed protocols for designing PCR primers are known in the art (e.g., Green and Sambrook et al.). In certain embodiments, the miRNA obtained from the sample may be subjected to qRT-PCR. Reverse transcription may be performed by any method known in the art, such as using the Omniscript RT Kit (Qiagen). The obtained cDNA is amplified by any amplification technique known in the art. The miRNA expression is then analyzed using, for example, a control sample as described below. As described herein, the overexpression or underexpression of the miRNA relative to the control may be measured to determine the miRNA expression profile of an individual biological sample. Similarly, detailed protocols for preparing and using microarrays to analyze miRNA expression are known in the art and described herein.

[0097] As used herein, RNA-sequencing (RNA-seq), also known as Whole Transcriptome Shotgun Sequencing, refers to any of a variety of high-throughput sequencing techniques used to detect the presence and abundance of RNA transcripts in real time. See Wang, Z., M. Gerstein, and M. Snyder, RNA-Seq: a revolutionary tool for transcriptomics, NAT REV GENET, 2009. 10(1): p.57-63. RNA-seq can be used to reveal a snapshot of a sample's miRNAs from the genome at a single moment in time. In certain embodiments, miRNAs are converted to cDNA fragments via reverse transcription prior to sequencing, and in certain embodiments, miRNAs can be sequenced directly without conversion to cDNA. Adapters can be attached to the 5' and / or 3' ends of the miRNAs, and the miRNAs or cDNAs can be optionally amplified, for example, by PCR. The fragments are then sequenced using high throughput sequencing technology such as those available from Roche (eg, the 454 platform), Illumina, Applied Biosystem (eg, the SOLiD system). [Brief description of the drawings]

[0098] [Figure 1] Figures 1A-1C show case flow diagrams for the lung cancer dataset ( Figure 1A , split into discovery and validation sets) and the ovarian, liver, and bladder cancer datasets ( Figure 1B , merged into a single validation dataset after removing redundant samples), summarizing patient and tumor characteristics of lung, bladder, ovarian, and liver cancer patients, and demographic information of matched controls ( Figure 1C ); [Diagram 2]Figures 2A-2G show the development and validation of a 4-miRNA diagnostic model in a lung cancer dataset. Figure 2A shows the determination of the optimal number of miRNAs (dotted line) for the diagnostic model by 10-fold cross-validation in the discovery set. Figure 2B shows the ROC analysis in the discovery set. Figure 2C shows the distribution of normalized diagnostic indices in the discovery set. Figure 2D shows the ROC analysis in the validation set. Figure 2E shows the distribution of normalized diagnostic indices in the validation set. Figure 2F shows the comparison of normalized diagnostic indices in paired serum samples (pre- and post-surgery) of 180 lung cancer patients. Figure 2G shows the distribution of normalized diagnostic indices in the clinical subset of the validation set. The dotted horizontal line indicates the cut point of the normalized diagnostic indices of our model. The percentages shown in the graphs are the sensitivity in each cancer subgroup. [Diagram 3] Figures 3A and 3B show the performance of the 4-miRNA diagnostic model in the dataset of additional cancers, where Figure 3A shows the ROC analysis and Figure 3B shows the distribution of normalized diagnostic indices of the 4-miRNA model. The percentages shown in the graphs are the sensitivity of each cancer type and the specificity of non-cancer controls; [Figure 4] 4A and 4B show the ROC analysis and distribution of normalized diagnostic indices among age and sex groups in the lung cancer dataset. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0099] The present disclosure provides approaches, including methods, kits and computerized systems, that can accurately and reliably detect one or more human cancers in a subject based on the expression profile of a miRNA biomarker set consisting of at least one miRNA determined from a biological sample obtained from the subject.

[0100] In a first aspect of this section, a detection method is provided that can achieve a diagnostic accuracy having an AUC value of about 0.780 or greater, the method essentially comprising the following three steps:

[0101] Step (1): determining the expression profile of a miRNA biomarker set;

[0102] Step (2): Calculate a diagnostic index for the biological sample based on the expression profile of the miRNA biomarker set. The diagnostic index is calculated based on:

number

[0103] Step (3): Classifying the subject as having or not having cancer based on the value of the calculated diagnostic index. If the calculated diagnostic index is equal to or greater than a predetermined threshold, the subject is classified as having cancer, and if not, the subject is classified as not having cancer.

[0104] Here, the miRNA biomarker set includes hsa-miR-5100, and may optionally further include any one or combination of miRNAs listed in Table 1 (see Example 1). According to different embodiments, in addition to hsa-miR-5100, the miRNA biomarker set may further include miRNA(s) from the top 2-100 miRNAs of Table 1, or alternatively, may further include miRNA(s) from the top 2-50 miRNAs, or alternatively, may further include miRNA(s) from the top 2-20 miRNAs, or alternatively, may further include miRNA(s) from the top 2-4 miRNAs.

[0105] Preferably, the miRNA biomarker set consists of the top four miRNAs (i.e., hsa-miR-5100, hsa-miR-1343-3p, hsa-miR-1290, and hsa-miR-4787-3p). Here, according to different embodiments, there may be different AUC cutoff levels (e.g., 0.780, 0.850, 0.950, 0.990, and 0.999) or different sensitivity specificity levels (e.g., 68%-99%, 68%-99%, 83%-99%, and 99%-99%) at which the method can accurately detect at least a particular cancer type. For example, the method can accurately detect lung cancer and gastric cancer with AUC>0.999, and / or sensitivity>99.0%, specificity>99.0%.

[0106] There are various methods of calculating diagnostic index based on formula (I).Optionally, calculation can be based on non-weighted model or weighted model.In the latter situation, can optionally apply different models (such as Limma model, logistic regression model, etc.) to obtain the weight of miRNA in miRNA biomarker set.

[0107] Preferably, the diagnostic index is calculated through a weighting model using weights from the Limma model. Wherein, in step (3) of the method, the predetermined threshold can be set as 1110, so that the method has a specificity of greater than 0.95; or optionally, the predetermined threshold can be set as 1200, so that the method has a specificity of greater than 0.99.

[0108] Optionally, the diagnostic index calculated in step (2) can be further subjected to a normalization process, and step (3) can determine a cancer classification based on whether the normalized diagnostic index is below or above a pre-set cut point.

[0109] It should be noted that the selection of the normalization process is arbitrary. According to some embodiments, the normalization process may be based on a mathematical formula:

number

[0110] where, optionally, param location and param scale By selecting α,β,β as 600 and 1000, respectively, the normalized diagnostic index can be between 0 and 10. Under such normalization, the preset cut point can be set as 5.1 to give a specificity >0.95 or as 6.0 to give a specificity >0.99.

[0111] In this method, the biological sample can advantageously be a liquid biopsy sample, such as a blood sample, a serum sample, a plasma sample, a urine sample, a saliva sample, or a sputum sample. The determination of the expression profile of the miRNA biomarker set can be achieved by various probe-based approaches, including Northern blotting, microarray analysis, RNA sequencing, or RNA in situ hybridization, or various amplification-dependent approaches, including reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR.

[0112] Optionally, the method can further include performing an evaluation of the subject to determine whether the subject is diagnosed with cancer (if the subject did not previously have cancer) or whether the subject is suffering from a recurrence of cancer (if the subject has previously undergone treatment to remove the cancer or did not previously have cancer). For such purposes, the evaluation can further include a physical examination, pathological examination of a biopsy from the subject, immunohistochemistry, or imaging studies including x-rays, computed tomography (CT), ultrasound, magnetic resonance imaging, and the like.

[0113] Further optionally, the method can further include administering to the subject a therapeutic regimen, such as surgery, radiation therapy, chemotherapy, hormonal therapy, targeted therapy, immunotherapy, or a combination thereof, if the subject is classified as having cancer.

[0114] In the second aspect, there is further provided a kit which can be employed to specifically carry out the various steps of the methods according to different embodiments as described above in the first aspect of this section.

[0115] The kit essentially comprises certain items that can be used to determine the expression profile of a miRNA biomarker set (i.e., component (1) including one or more nucleic acids capable of specifically recognizing each miRNA in the miRNA biomarker set, and optionally one or more amplification primers), and specific instructions for calculating a diagnostic index and for cancer classification (i.e., component (2)).

[0116] Depending on the miRNA included in the miRNA biomarker set, each of the nucleic acids of component (1) consists of, or a polynucleotide consisting of, (a) a nucleotide sequence set forth in SEQ ID NO:S:1-100, 1-50, 1-20 or 1-4, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof comprising 15 or more consecutive nucleotides; or (b) a nucleotide sequence complementary to a nucleotide sequence set forth in SEQ ID NO:S:1-100, 1-50, 1-20 or 1-4, a derivative thereof, a variant thereof having at least 80% sequence identity, or a fragment thereof consisting of 15 or more consecutive nucleotides.

[0117] There may be various different embodiments of the kit with respect to the following elements / features, including what miRNA components are included in the miRNA biomarker set, whether and how normalization is performed for the diagnostic index, how to classify whether a subject has cancer or not, what samples can be used for the biological sample, what detection accuracy level is achieved, etc. Specific details about these different embodiments may be referred to the various embodiments of the method as described above, and are omitted herein for brevity.

[0118] In a third aspect of this section, there is further provided a computerized solution substantially useful for carrying out, in a computerized automated manner, various steps of the method as described above in the first aspect of this section.

[0119] Such a computer solution can be applied in situations where the implementation of the various steps (1) to (3) of the above-mentioned method is automated by running a software program consisting of program instructions on a computer, and brings the advantages of high efficiency and great convenience.

[0120] Specifically, such a computerized solution may include a computerized system or computer system comprising a processor (i.e., controller) and a computer-readable non-transitory storage medium communicatively coupled to the processor. The computer-readable non-transitory storage medium is configured to store program instructions executable by the processor, thereby causing the processor to perform various different steps in the method as described above, such as:

[0121] Step (1): determining the expression profile of a miRNA biomarker set;

[0122] Step (2): calculating a diagnostic index for the biological sample according to formula (I) based on the expression profile of the miRNA biomarker set; and

[0123] Step (3): The subject is classified as having or not having cancer based on the calculated value of the diagnostic index.

[0124] As used herein, a "processor" is construed as interchangeable with a "central control unit" or "central processing unit (CPU)" and may be considered a single-core or multi-core processor, or multiple processors for parallel processing. The term "non-transient" as used herein is intended to describe a tangible computer-readable storage medium that does not propagate electromagnetic signals, but is not intended to otherwise limit the type of physical computer-readable storage device encompassed by this phrase. Examples may include any tangible or non-transient storage or memory medium, such as electronic, magnetic, or optical media (e.g., disks or CD / DVD-ROMs), or non-volatile memory storage (e.g., "flash" memory).

[0125] 5, the system 100 may further include, in addition to the processor 10 and the computer-readable non-transitory storage medium 20, a bus 30, a memory 40, an I / O interface 50, and a communication interface 60. The processor 10, the storage medium 20, the memory 40, the I / O interface 50, and the communication interface 60 are all communicatively coupled to each other via the bus 30.

[0126] The storage medium 20 stores computer-executable program instructions that, when executed by the processor 10, cause the processor 10 to perform steps (1)-(3) of the above-mentioned method. The memory 40 is configured to temporarily store the program instructions retrieved from the storage medium 20, and the processor 10 is configured to execute the program instructions temporarily stored in the memory 40. The input / output interface 50 enables input / output between the system 100 and a user, and realizes control of the system 100. The communication interface 60 communicatively connects the system 100 to other computing devices, allowing data to be exchanged. It should be noted that these computer hardware components can be located locally or remotely over a network, such as an intranet, the Internet, or a cloud.

[0127] In the following, an example is provided to illustrate the invention as described above in various aspects of the present disclosure. EXAMPLES

[0128] Example 1 In this example, we develop and validate a circulating cell-free miRNA-based diagnostic signature for MCED by utilizing four large-scale miRNA microarray datasets based on a standardized microarray platform.

[0129] 2. Materials and Methods Study design Four microarray datasets with a total of 7536 unique participants, including 3604 cancer patients and 3932 non-cancer controls, were included in the current analysis, all of which were derived from studies originating from the Japanese national research project "Development and Diagnostic Technologies for miRNA Detection in Body Fluids," designed to characterize serum miRNAs in over 50,000 participants across 13 cancer types using a standardized microarray platform (Asakura et al. 2020; Yokoi et al. 2018; Usuba et al. 2019; Yamamoto et al. 2020). The four datasets were originally collected to develop diagnostic signatures for lung cancer (GSE137140), ovarian cancer (GSE106817), liver cancer (GSE113740), and bladder cancer (GSE113486), respectively.

[0130] The lung cancer dataset has the largest sample size in a single cancer type (n=1566) and non-cancer controls (n=2178). In the original lung cancer study, a 2-miRNA diagnostic model (referred to in this study as the "original 2-miRNA model") with high sensitivity and specificity for detecting lung cancer was established (Asakura et al.). The aim of this study was initially set to use this dataset to develop and validate a new diagnostic model that could potentially outperform the original 2-miRNA model in detecting lung cancer. As datasets for other cancer types were identified, the new model was evaluated for its performance in detecting other cancers.

[0131] 2.2. Participants and serum samples Collection of serum samples has been described in the original papers (Asakura et al., 2020; Yokoi et al., 2018; Usuba et al., 2019; Yamamoto et al., 2020). Briefly, serum samples were collected from cancer patients referred or admitted to the National Cancer Center Hospital (NCCH) between 2008 and 2016 before surgery and stored at 4°C for 1 week, then stored at −20°C until further use. Cancer patients who underwent preoperative chemotherapy and radiotherapy before serum collection were excluded. Serum samples from non-cancer controls with no history of cancer and no hospitalization in the past 3 months were collected along with routine blood tests at the outpatient clinics of three medical institutions: NCCH, National Center for Geriatrics and Gerontology (NCGG) Biobank, and Yokohama Minoru Clinic (YMC). Serum collected from NCCH was stored in the same manner as cancer patients, and serum collected from NCGG and YMC was stored at −80°C until use. This study was approved by the Institutional Review Board of the NCCH, the Ethics and Conflict of Interest Committee of the NCGG, and the Research Ethics Committee of Shintokai YMC Medical Corporation. Written informed consent was obtained from each participant.

[0132] 2.3.miRNA microarray expression analysis For details of the microarray analysis, see the original paper (Asakura et al.). Briefly, total RNA was extracted from 300 μL of serum and analyzed using 3DGene. (登録商標) The 3D-Gene was designed to probe 2588 miRNA sequences in miRBase Release 21 using the miRNA Labeling kit. (登録商標) The samples were hybridized to a Human miRNA Oligo Chip (Toray, Kanagawa, Japan). The following low-quality samples were excluded: negative control probe with coefficient of variation >0.15; and 3D-Gene (登録商標)Number of flagged probes identified as "uneven spot image" by Scanner >10. Presence of miRNA was determined if the signal intensity was greater than the mean + 2x standard deviation of the negative control signal, and the top and bottom 5% of ranked signal intensities were removed when using the negative control signal. Background subtraction was performed by subtracting the mean signal of the negative control signal (after removing the top and bottom 5% ranked by signal intensity) from the miRNA signal. Normalization between microarrays was achieved by calibration according to three preselected internal control miRNAs (miR-149-3p, miR-2861, miR-4463).

[0133] 2.4. Diagnostic model development The patients in the lung cancer dataset were divided into the same discovery and validation sets as in the original paper (Figure 1A) (Asakura et al. 2020). The reasons are: (1) the discovery set was selected by the original authors to be balanced between cancer and non-cancer in terms of age, sex, and smoking history; (2) 50% of the non-cancer patients in the discovery set were NCCH patients with the same serum storage conditions as the cancer patients, minimizing potential bias in the selection of miRNA candidates; (3) using the same discovery and validation sets allows for a direct performance comparison of the new diagnostic model with the original 2-miRNA model. Because the diagnostic model was developed from the lung cancer discovery set, after validation with the lung cancer validation set, its ability as a multi-cancer diagnostic model was further verified with a dataset combining other cancer types that were not used for model development.

[0134] A Linear Model for Microarray Data (limma) (Ritchie et al. 2015) was performed in the discovery set to assess the statistical significance of miRNA expression differences between lung cancer vs. non-cancer cases. A 10-fold cross-validation was performed in the discovery set to determine the optimal number of miRNAs for the optimal diagnostic model based on the area under the curve (AUC) of the Receiver's Operating Characteristics (ROC) curve analysis. The diagnostic index was calculated as a linear sum of miRNA expression levels weighted by the limma statistic. The cut point of the diagnostic index was selected to avoid misclassification of non-cancer controls in the discovery set to minimize false positives, as the diagnostic model may be used as a screening test for the general population at risk.

[0135] 2.5.Statistical analysis The diagnostic ability to discriminate between cancer and non-cancer was determined by the AUC, sensitivity, and specificity of the ROC curve analysis. Comparison of the AUC of two ROC curves was performed with the roc.test function using the bootstrap method in the pROC package. Comparison of the sensitivity of paired pre- and post-surgery samples in the lung cancer clinical subset was performed with the McNemar test. Limma analysis was performed using the Bioconductor package limma (The Bioconductor Open Source Software For Bioinformatics, accessed August 27, 2020). All statistical analyses were performed using R version 4.0.5 (The R Project for Statistical Computing, accessed July 15, 2020).

[0136] 3.Results Participants and Dataset The lung cancer dataset included 1566 lung cancer patients and 2178 non-cancer controls (Figure 1A) (Asakura et al.). The ovarian cancer dataset included 333 ovarian cancer patients and 2759 non-cancer controls, as well as patients with breast cancer, colon cancer, esophageal cancer, gastric cancer, liver cancer, lung cancer, pancreatic cancer, and sarcoma (Figure 1B) (Yokoi et al. 2018). The liver cancer and bladder cancer datasets included 345 liver cancer / 1033 non-cancer and 392 bladder cancer / 100 non-cancer patients, respectively, in addition to patients with biliary tract cancer, breast cancer, colon cancer, esophageal cancer, gastric cancer, glioma, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, and sarcoma (Figure 1B) (Usuba et al. 2019, Yamamoto et al. 2020). We left the lung cancer dataset intact and removed redundant samples in the other three datasets that had correlations greater than 0.99 between the datasets or with samples in the lung cancer dataset. We then merged the unique samples from the ovarian, liver, and bladder cancer datasets into a single non-lung cancer dataset containing a total of 3792 samples, including 2038 cancer patients and 1754 non-cancer controls across 12 cancer types (Figure 1B).

[0137] The lung cancer dataset was divided into a discovery set (n = 416) and a validation set (n = 3328) (Figure 1A), the same as in the original study. The discovery set included 208 lung cancer patients and 208 non-cancer controls matched for age, sex, and smoking status (Asakura et al. 2020). The validation set included 1358 lung cancer patients and 1970 non-cancer controls. The lung cancer patients were 57% male, 62% former or current smokers, 78% adenocarcinoma, 14% squamous cell carcinoma, 72% stage I, 15% stage II, and 13% stage III (Figure 1C).

[0138] The 392 bladder cancer patients had a mean age of 68 years, 72% were male, 5% had metastases, 12% had positive lymph nodes, 77% had T2 or lower, and 80% had high grade (Figure 1C). The 333 ovarian cancer patients had a mean age of 57 years, 25% had stage I, 10% had stage II, 55% had serous, 19% had clear cell, and 13% had endometriosis (Figure 1C). The 348 liver cancer patients had a mean age of 68 years, 78% were male, 37% had stage I, and 33% had stage II (Figure 1C). Detailed demographic and tumor characteristics for the other cancers were not provided by the original study. [Table 1] TIFF2024523848000008.tif205141TIFF2024523848000009.tif202141TIFF2024523848000010.tif205141TIFF2024523848000011.tif16141

[0139] 3.2.Development of diagnostic model The development of the diagnostic model was performed on the discovery set of the lung cancer dataset, which included 208 lung cancer patients and 208 non-cancer controls (Figure 1A). Limma analysis was used to assess the statistical significance of the miRNA expression differences between lung cancer patients and non-cancer controls. The top 100 differentially expressed miRNAs were listed in Table 1. 10-fold cross-validation showed that the diagnostic model using the top four miRNAs (hsa-miR-5100, hsa-miR-1343-3p, hsa-miR-1290, hsa-miR-4787-3p), ranked by adjusted p-value, yielded the best AUC in the ROC curve analysis (Figure 2A). The diagnostic index, calculated as a weighted sum of the four miRNA expression levels and normalized to a range of 0 to 10, showed a near-perfect AUC value of 0.999 (Figure 2B), which was numerically superior to the AUC value of 0.993 (p = 0.16) of the 2-miRNA model in the original paper (Asakura et al. 2020). The cut point of 6 was chosen to minimize false positives and to avoid misclassification of non-cancer controls in the discovery set, resulting in a sensitivity of 98% and a specificity of 100% (Figure 2C), compared to 99% for both the original 2-miRNA model (Asakura et al. 2020).

[0140] 3.3. Validation of diagnostic model on lung cancer validation set The performance of the 4-miRNA model was evaluated in a lung cancer validation set (n=3328) including 1358 lung cancer patients and 1970 non-cancer controls. The 4-miRNA model achieved an AUC of 0.999 (Figure 2D), which was significantly better (P=0.01) than the original 2-miRNA model (Asakura et al. 2020) AUC of 0.996. The original 2-miRNA model (Asakura et al. 2020) also had a sensitivity of 95% and a specificity of 99%, whereas the new model had a sensitivity and specificity of 99% (Figure 2E).

[0141] Furthermore, the performance of the 4-miRNA model was evaluated in clinical subsets of the validation set defined by clinical stage, T stage, N stage, M stage, and histology. In all clinical subsets, the 4-miRNA model showed a sensitivity of approximately 99% or higher (Figure 2G, Table 2), which was superior to the sensitivity of the original 2-miRNA model (Table 2). Especially for early lung cancer, for example, both patients with stage I lung cancer and patients with T1 tumors, the 4-miRNA model showed a sensitivity of 99% or higher (Figure 2G, Table 2), while the sensitivity of the 2-miRNA model was 95.4% and 95.9%, respectively (Table 2). Even in the common histological types of adenocarcinoma and squamous cell carcinoma, the 4-miRNA model showed superior performance compared to the original 2-miRNA model (Figure 2G, Table 2) (Table 2). [Table 2]

[0142] Paired serum sample (pre- and post-surgery) data were also available for 180 patients. The diagnostic index of the 4-miRNA model for post-surgery serum samples was reduced to normal levels below the diagnostic index cutpoint (Figure 2F).

[0143] 3.4. Application of the diagnostic model to additional cancer types The performance of the 4-miRNA model was further evaluated on a combined dataset of 3792 patients, including 2038 patients with 12 types of cancer and 1754 non-cancer controls. Bladder cancer, liver cancer, and ovarian cancer had the largest sample sizes, with more than 300 patients each. Except for breast cancer, where the 4-miRNA model did not work, the 4-miRNA model showed very strong performance, with AUC>0.95 in biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma, liver cancer, ovarian cancer, pancreatic cancer, and prostate cancer, and AUC0.876 in sarcoma (Figure 3A). Thus, the 4-miRNA model showed high sensitivity of 83.2-100% in biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma, liver cancer, pancreatic cancer, and prostate cancer, and reasonable sensitivity of 68.2% and 72.0% in ovarian cancer and sarcoma, respectively (Figure 3B). Furthermore, the 4-miRNA model maintained a high specificity of 99.3% against 1754 non-cancer controls independent of those included in the lung cancer dataset.

[0144] Further sensitivity analysis was performed using a different diagnostic index cut point of 5.1, which reduced the specificity to 95%, and the sensitivity increased for all 11 cancer types, with all 11 cancer types showing a sensitivity of 90% or higher, except for sarcoma, which had a sensitivity of 76.5% (Table 3).

[0145] [Table 3]

[0146] 4. Discussion In this example, we report the development and performance evaluation of a 4-miRNA diagnostic model for the early detection of multiple cancers. In a large independent validation set of 7120 individuals, including 3396 cancer patients and 3724 non-cancer patients, the 4-miRNA model was able to simultaneously detect 12 types of cancer (biliary, bladder, colorectal, esophageal, gastric, glioma, liver, lung, ovarian, pancreatic, prostate, and sarcoma) with high sensitivity (80-100% for 10 cancer types and ~70% for 2 cancer types), while maintaining a very high specificity of 99%, which is typically required for a screening test to be useful in the general population at risk. To our knowledge, this is the first MCED diagnostic model based on circulating free miRNAs. Interestingly, the diagnostic index in lung cancer patients decreased to the level of non-cancer controls after tumor resection.

[0147] Non-invasive screening tests that analyze circulating nucleic acids and / or proteins have been a driving force in the MCED campaign and have seen significant progress recently. Nearly all of the tests being developed for MCED are based on the evaluation of circulating tumor DNA, most of which utilize next-generation bisulfite sequencing technology to evaluate the methylation patterns of these tumor DNAs (Klein et al. 2021; Cohen et al. 2018; Chen et al. 2020; Cristiano et al. 2019). Two such tests, Galleri and PanSeer, have been developed as methylation-based epigenetic signatures (Klein et al. 2021; Chen 2020). In an analysis of a case-control study of the Circulating Cell-Free Genome Atlas (CCGA), Galleri surveyed over 100,000 methylated regions and found that sensitivity for 12 pre-specified cancers (anal, bladder, colon / rectum, esophagus, head and neck, liver / bile duct, lung, lymphoma, ovarian, pancreatic, plasma cell neoplasms, and stomach) was 67.6% in patients with stage I-III cancer (n=874) and increased to 76.3% (n=1346) when stage IV cancer was included, while specificity based on 1254 non-cancer controls reached 99.3% (Klein et al. 2021). On the other hand, the PanSeer assay, which targets only 477 methylated genomic regions, showed a high sensitivity of 95% in 98 individuals (prediagnostic samples) who were subsequently diagnosed with one of five cancers (gastric, esophageal, colorectal, lung, and liver) within 4 years of blood collection, but a low specificity of 96% in 207 healthy controls (Chen et al. 2020). However, what is puzzling about PanSeer is that when evaluated on 113 postdiagnostic plasma samples, the test showed a low sensitivity of only 88% (Chen et al. 2020). Another test called DELFI, based on genome-wide analysis of cell-free DNA fragmentation patterns by next-generation sequencing, achieved a sensitivity of 73% and a specificity of 98% (n = 215) in seven cancers (n = 208; breast, cholangiocarcinoma, colorectal, gastric, lung, ovarian, and pancreatic cancer) (Cristiano et al. 2019).Finally, CancerSEEK, a test combining the measurement of nine protein biomarkers with mutation detection of 16 genes in circulating cell-free DNA, showed 10-fold cross-validation and a median sensitivity of 70% (n=1005) and specificity of 99% (n=812) in eight cancers (n=1005; ovarian, liver, stomach, pancreas, esophagus, colon, lung, breast) (Cohen et al. 2018). In summary, MCED tests currently under development generally showed sensitivities in the 60–70% range when a high specificity of 99% was mandated. Compared to these tests, our diagnostic model is very simple, with only four miRNAs, yet showed substantially higher sensitivity, in the 80–100% range, for 10 of the 12 cancer types studied in a large cohort of over 7000 individuals. It is noteworthy that the simple diagnostic model is not only significantly lower in cost but can also be developed into an in vitro diagnostic (IVD) test using conventional technology platforms that allow decentralized testing, such as RT-PCR, which is an advantage over NGS-based tests that are typically performed as laboratory developed tests (LDTs). These attributes are critical in driving widespread adoption and compliance of MCED tests, as they are intended to target high-risk or at-risk general public.

[0148] Among the 13 cancer types examined in this study, only breast cancer was not successfully detected by the 4-miRNA diagnostic model. The reason for this underperformance is unclear, but it may indicate that breast cancer has a different miRNA expression profile or a different miRNA shedding pattern in the blood. Interestingly, Galleri and CancerSEEK also showed low sensitivity in breast cancer, 30.5% and 33%, respectively (Klein et al. 2021; Cohen et al. 2018), although their poor performance in breast cancer may not be clinically significant, since mammography screening is highly effective in detecting early breast cancer and reducing breast cancer mortality (Nelson et al. 2016).

[0149] The final diagnostic performance and clinical value of these MCED tests must be established in large-scale prospective screening trials in asymptomatic individuals. In the DETECT-A trial, which enrolled more than 10,000 asymptomatic women, 96 cancers from 10 cancer types were identified, with CancerSEEK's sensitivity of 27%, which rose to 52% when cancers detected by standard screening tests were added (Lennon et al.). Furthermore, when CancerSEEK was combined with PET-CT scans, the specificity was 99.6% and the positive predictive value (PPV) was 40.6%. Meanwhile, an interim analysis of 4,033 patients in the prospective study of the Galleri test, PATHFINDER, showed that 40 were positive, of which 18 were confirmed to have cancer, with a PPV of 45% (Beer et al., 2021). In our 4-miRNA diagnostic model, assuming a cancer prevalence of 1%, a conservative average sensitivity of 85, and a specificity of 99.3%, the PPV for screening asymptomatic individuals is 55%, which is significantly higher than the PPVs of the four single cancer screens recommended by the USPSTF (3.7–4.4%) (Lehman et al. 2017; US Food and Drug Administration Cologuard Summary of Safety and Effectiveness Data, 2014; and National Lung Screening Trial Research Team, 2013).

[0150] 5. Conclusion In summary, our study provided proof-of-concept data for a simple, affordable, blood-based diagnostic test to detect multiple cancers. The 12 cancer types detected in this study account for approximately 380,000 (~62%) of the estimated cancer deaths in the United States in 2021.

[0151] It should be noted that the above examples and data provided only cover the 12 types of cancers for which miRNA biomarker set, especially 4-miRNA biomarker set, shows excellent power in cancer detection with very high accuracy, but the types of cancers that miRNA biomarker set can be applied to are not limited.Therefore, the scope of the present disclosure is interpreted as covering other cancer types.The fact that the model provided in the present disclosure works in 12 cancer types out of the 13 cancer types studied strongly suggests that the present method is applicable to most, if not all, cancer types.

[0152] References Ritchie, ME;et al.(2015).limma powers differential expression analyzes for RNA-sequencing and microarray studies.Nucleic Acids Research 43(7),e47. Venables, WN and Ripley, BD (2002) Modern Applied Statistics with S.Fourth edition. Springer. Tibshirani, R(1996)."Regression Shrinkage and Selection via the lasso".Journal of the Royal Statistical Society.Series B (methodology).Wiley.58(1):267-88. Hoerl, AE and Kennard, RW (1970). "Ridge Regression: Biased Estimation for Nonorthogonal problems". Technometrics.12(1):55-67. Ripley,BD(1996)Pattern Recognition and Neural Networks.Cambridge University Press. Kozomara,A and Griffiths-Jones,S(2010)."MiRBase:integrating microRNA annotation and deep-sequencing data".Nucleic Acids Research.39(Database issue):D152-7. miRBase:the microRNA database:http: / / www.mirbase.org / The Bioconductor Open Source Software For Bioinformatics:http: / / www.bioconductor.org The R Project for Statistical Computing:https: / / www.r-project.org / Asakura,K;et al.(2020).A MiRNA-Based Diagnostic Model Predicts Resectable Lung Cancer in Humans with High Accuracy.Commun.Biol.3,134. Yokoi, A; et al.(2018).Integrated Extracellular MicroRNA Profiling for Ovarian Cancer Screening.Nat.Commun.9,4319. Usuba,W;et al.(2019).Circulating MiRNA Panels for Specific and Early Detection in Bladder Cancer.Cancer Sci.110,408-419. Yamamoto,Y;et al.(2020). Highly Sensitive Circulating MicroRNA Panel for Accurate Detection of Hepatocellular Carcinoma in Patients With Liver Disease.Hepatol.Commun.4,284-297. Klein,EA;et al.(2021).Clinical Validation of a Targeted Methylation-Based Multi-Cancer Early Detection Test Using an Independent Validation Set.Ann.Oncol.:Off.J.Eur.Soc.Med.Oncol.32,1167-1177. Cohen,JD;et al.(2018).Detection and Localization of Surgically Resectable Cancers with a Multi-Analyte Blood Test.Science.359,926-930. Chen,X;et al.(2020).Non-Invasive Early Detection of Cancer Four Years before Conventional Diagnosis Using a Blood Test.Nat.Commun.11,3475. Cristiano,S;et al.(2019).Genome-Wide Cell-Free DNA Fragmentation in Patients with Cancer.Nature.570,385-389. Nelson,HD;et al.(2016).Effectiveness of Breast Cancer Screening:Systematic Review and Meta-Analysis to Update the 2009 U.S.Preventive Services Task Force Recommendation.Ann.Intern.Med.164,244-255. Lennon,AM;et al.(2020).Feasibility of Blood Testing Combined with PET-CT to Screen for Cancer and Guide Intervention.Science.369,eabb9601. Beer,T;et al.(2021).Interim Results of PATHFINDER,a Clinical Use Study Using a Methylation-Based Multi-Cancer Early Detection Test.J.Clin.Oncol.39,3010. Lehman,CD;et al.(2017).National Performance Benchmarks for Modern Screening Digital Mammography:Update from the Breast Cancer Surveillance Consortium. Radiology.283,49-58. U.S.Food and Drug Administration Cologuard Summary of Safety and Effectiveness Data(Premarket Approval Application P130017);2014. National Lung Screening Trial Research Team;Church,TR;et al.(2013).Results of Initial Low-Dose Computed Tomographic Screening for Lung Cancer.New Engl.J.Med.2013,368,1980-1991. Nielsen,PE;et al.(1991).Sequence-selective recognition of DNA by strand displacement with a thymine-substituted polyamide.Science.254,p.1497-500. Obika,S;et al.(1998).Stability and structural features of the duplexes containing nucleoside analogues with a fixed N-type conformation,2'-O,4'-C-methyleneribonucleosides.Tetrahedron Lett..39,p.5401-5404. Green,MR and Sambrook,J.(2012).Molecular Cloning:A Laboratory Manual,4th Ed.,Cold Spring Harbor Press,Cold Spring Harbor,N.Y. Sambrook,J;et al.(1989).Molecular Cloning:A Laboratory Manual,New York:Cold Spring Harbor Press. Zhang,Z;et al.(2000).A greedy algorithm for aligning DNA sequences.J.Comput.Biol.7,p.203-214. Altschul,SF;et al.(1990).Basic local alignment search tool.Journal of Molecular Biology, Vol.215,p.403-410. Pearson, WR et al.(1988).Improved tools for biological sequence comparison.Proc.Natl.Acad.Sci.U.S.A.,Vol.85,p.2444-2448. Yun,SJ;et al.(2012).Cell-free microRNAs in urine as diagnostic and prognostic biomarkers of bladder cancer.Int J Oncol.2012 Nov;41(5):1871-8. Park,NJ;et al.(2009).Salivary microRNA: discovery,characterization,and clinical utility for oral cancer detection.Clin Cancer Res.2009 Sep 1;15(17):5473-7.

Claims

1. A system for detecting cancer in a subject, comprising: a processor; and a non-transitory memory medium containing program instructions for execution by the processor, the program instructions causing the processor to execute steps in a method comprising: the non-transitory memory medium, determining an expression profile of a miRNA biomarker set consisting of at least one miRNA from a biological sample from the subject, each of the at least one miRNA being as set forth in Table 1; said determining; calculating a diagnostic index for the biological sample based on the expression profile of the miRNA biomarker set, the diagnostic index being calculated based on the mathematical formula 【Number 1】 and Here, n is the total number of the at least one miRNA in the miRNA biomarker set, and the miRNA i is the expression level of the i th th miRNA in the miRNA biomarker set, i is an integer greater than 0 and less than or equal to n, and t i is the weight of the i th th miRNA, the calculating, and classifying whether the subject has cancer or not based on the calculated diagnostic index, classifying the subject as having cancer if the calculated diagnostic index is greater than or equal to a predetermined threshold, and classifying the subject as not having cancer otherwise; said classifying, wherein the method can achieve a diagnostic accuracy with an AUC value exceeding about 0.

780.

2. The system according to claim 1, wherein the miRNA biomarker set consists of hsa-miR-5100, hsa-miR-1343-3p, hsa-miR-1290, and hsa-miR-4787-3p.

3. The system according to claim 2, wherein the method can achieve a diagnostic accuracy with an AUC value greater than about 0.850, and the cancer is selected from the group consisting of lung cancer, biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma cancer, liver cancer, pancreatic cancer, prostate cancer, ovarian cancer, and sarcoma.

4. The system according to claim 2, wherein the method can achieve a diagnostic accuracy having a specificity greater than about 99.0% and a sensitivity greater than about 83.0%, and the cancer is selected from the group consisting of lung cancer, biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma cancer, liver cancer, pancreatic cancer, and prostate cancer.

5. The system according to any one of claims 1 to 4, wherein when calculating the diagnostic index of the biological sample based on the expression profile of the miRNA biomarker set, the diagnostic index is calculated through an unweighted model.

6. The system according to any one of claims 1 to 4, wherein when calculating the diagnostic index of the biological sample based on the expression profile of the miRNA biomarker set, the diagnostic index is calculated through a weighted model using weights selected from the group consisting of Linear Models for Microarray Data (limma) model, logistic regression model, linear discriminant analysis (LDA) model, conditional logistic regression model, lasso regression model, ridge regression model, random forest, support vector machine, and probit regression model.

7. The system according to claim 6, wherein the diagnostic index is calculated through a weighted model using weights from the limma model.

8. The system according to claim 1, after calculating the diagnostic index of the biological sample and before classifying whether the subject has the cancer, obtaining a normalized diagnostic index based on the calculated diagnostic index, classifying whether the subject has the cancer based on the calculated diagnostic index, if the normalized diagnostic index is equal to or greater than a preset cut-off point, classifying the subject as having the cancer, or otherwise, classifying the subject as not having the cancer, further including the obtaining.

9. In the system according to claim 8, when obtaining the normalized diagnostic index based on the calculated diagnostic index, the normalized diagnostic index is calculated based on the formula: 【Number 2】 calculated based on, Here, the param location and param scale are a location parameter and a scale parameter, respectively, configured such that the normalized diagnostic index falls within a range of not more than a first preset value and not more than a second preset value. System.

10. In the system according to claim 9, the miRNA biomarker set consists of hsa-miR-5100, hsa-miR-1343-3p, hsa-miR-1290, and hsa-miR-4787-3p.

11. The system according to claim 10, wherein the diagnostic index is calculated through a weighted model using weights from the LIMMA model, the first preset value is 0, and the second preset value is 10.

12. The system according to claim 11, wherein the preset cut-off point is 5.1, and the method can achieve a diagnostic accuracy having a specificity value greater than about 0.

95.

13. The system according to claim 11, wherein the preset cut-off point is 6.0, and the method can achieve a diagnostic accuracy having a specificity value greater than about 0.

99.

14. The system according to claim 1, wherein the biological sample is a liquid biopsy sample selected from the group consisting of a blood sample, a serum sample, a plasma sample, a urine sample, a saliva sample, and a saliva sample.

15. The system according to claim 1, wherein when determining the expression profile of the miRNA biomarker set, the expression profile of the miRNA biomarker set is obtained by at least one means of a Northern blot method, a microarray analysis method, an RNA sequencing method, an RNA in situ hybridization method, or a nucleic acid amplification procedure, and the nucleic acid amplification procedure includes at least one of reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR.