Method for developing cancer diagnostic model and use thereof in developing cancer detection method

By constructing a diagnostic model based on a set of miRNA biomarkers, the shortcomings of existing technologies in detecting multiple cancers have been addressed. This enables non-invasive detection of multiple cancers with high sensitivity and specificity, and is applicable to the early detection of multiple cancer types in blood samples.

CN120958128APending Publication Date: 2025-11-14MIRONCOL DIAGNOSTICS LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202480018819.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-14
Filing Date
2024-03-12
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Currently, only four types of cancer have recommended screening methods, while screening methods for other cancer types are insufficient. As a result, two-thirds of cancer diagnoses and 70% of cancer deaths are not covered, making it urgent to develop low-cost, non-invasive or minimally invasive early detection methods for multiple cancer types.

Method used

By constructing a training set and using statistical models such as linear and logistic regression models for microarray data, a variety of miRNA biomarker sets are screened out, a diagnostic model is established, and a diagnostic index is calculated for the detection of various cancer types. The detection is performed using a processor and non-transient storage media.

Benefits of technology

It achieves high sensitivity and specificity for the detection of multiple cancer types, with an AUC of over 0.97 and sensitivity and specificity of over 0.99, making it suitable for non-invasive detection of blood samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005593696190000031
    Figure BDA0005593696190000031
  • Figure BDA0005593696190000041
    Figure BDA0005593696190000041
  • Figure BDA0005593696190000191
    Figure BDA0005593696190000191
Patent Text Reader

Abstract

The invention provides a cancer diagnosis model development method. The cancer diagnosis model development method comprises the steps of constructing a training set and establishing a diagnosis model on the training set. The training set includes miRNA expression profiles from non-cancer subjects and cancer patients having two or more cancer types, and establishing the diagnostic model includes calculating a diagnostic index based on a selected miRNA biomarker set, the selected miRNA biomarker set being obtained from a miRNA ranking in a differential expression analysis of the miRNA expression profiles in the training set. Methods of detecting a target cancer by means of the diagnostic model so developed are also provided. A 4-miRNA-based diagnostic model shows high performance in a verification set, and can realize the sensitivity greater than or equal to 0.98 while maintaining the specificity of 0.99 when detecting various cancers including lung cancer, gastric cancer, biliary tract cancer, bladder cancer, prostate cancer and glioma.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 451,751, filed March 13, 2023, and U.S. Provisional Application No. 63 / 538,292, filed September 14, 2023, the disclosure of which is incorporated herein by reference in its entirety.

[0003] Reference to the electronically submitted sequence list

[0004] The sequence list content submitted electronically, filename: Top200list.xml, file size: 172,425 bytes, and creation date: March 12, 2024, has been submitted with this application, and its entire contents are incorporated herein by reference. Technical Field

[0005] This invention relates generally to the field of disease screening, detection and diagnosis technology, and more specifically to methods, reagent kits, systems and non-transient storage media for detecting one or more types of human cancer. Background Technology

[0006] Cancer is a highly challenging and often fatal disease for humans, and it is widely believed that early detection (or diagnosis) of cancer can significantly improve survival rates. However, early-stage cancer patients are often asymptomatic, making them difficult to diagnose in practice.

[0007] Currently, only four types of cancer (breast cancer, colorectal cancer, lung cancer, and cervical cancer) have recommended screening methods. It is estimated that two-thirds of cancer diagnoses and 70% of cancer deaths are not covered by existing cancer screening guidelines. Therefore, there is an urgent need to develop a low-cost, ideally non-invasive or minimally invasive, detection method capable of early detection of multiple cancer types—a method known as Multi-Cancer Early Detection (MCED).

[0008] MicroRNAs (miRNAs) are a class of small, single-stranded non-coding RNA molecules encoded by corresponding genes in the human genome, with an average length of 22 nucleotides. It is estimated that these miRNAs negatively regulate the expression of more than 50% of genes in humans. Aberrant miRNA expression has been shown to be associated with various human cancers. Of particular note is that miRNAs are typically abundant and stable in blood samples from cancer patients; they are extracellular circulating molecules released from tumor cells into the circulatory system, thus miRNAs possess the potential to serve as non-invasive biomarkers for cancer screening and diagnosis. Previously, a 4-miRNA diagnostic model was developed using publicly available serum miRNA microarray datasets from a training set consisting only of lung cancer patients and matched non-cancer controls; it demonstrated a detection sensitivity of ~80%-100% for 10 cancer types, and a detection sensitivity of ~70% for sarcoma and ovarian cancer, while maintaining ~99% specificity (Zhang A et al., 2022). Summary of the Invention

[0009] In a first aspect, this disclosure provides a method for developing a diagnostic model for detecting a target cancer (i.e., a cancer diagnostic model development method).

[0010] The method includes the following steps: (1) constructing a training set comprising expression profiles of multiple miRNAs obtained from non-cancer subjects and cancer patients; and (2) developing a diagnostic model based on the training set. Here, the cancer patients are configured such that at least two types of cancer are present in the cancer patients; and step (2) further comprises the following sub-steps: (a) performing differential expression analysis on the expression profiles of the multiple miRNAs using a first statistical model, such that the multiple miRNAs are ranked by an adjusted p-value; and (b) establishing a diagnostic model based on a set of miRNA biomarkers selected from the multiple miRNAs. Here, the set of selected miRNA biomarkers includes at least one miRNA, each selected such that its ranking is no higher than a preset cutoff value m (m≥1).

[0011] Here, the first statistical model may optionally be selected from one of the following models, including but not limited to: a linear model for microarray data (limma) model, a logistic regression model, a linear discriminant analysis (LDA) model, a conditional logistic regression model, a lasso regression model, a ridge regression model, a random forest, a support vector machine, or a probabilistic unit regression model. According to some implementations, the limma model is used as the first statistical model.

[0012] In some embodiments of the cancer diagnostic model development method, sub-step (b) includes: calculating a diagnostic index based on the expression profile of a selected set of miRNA biomarkers, wherein the diagnostic index is calculated based on the following formula:

[0013]

[0014] Where n is the total number of at least one miRNA in the selected miRNA biomarker set (n≤m), miRNA i The expression level of the i-th miRNA in the selected miRNA biomarker set (0 <i≤n),t i denoted as the weight of the i-th miRNA.

[0015] Here, according to some implementation methods, t i The value is a constant, and therefore the diagnostic index is calculated without weights. According to some other implementations, the weight t... i Based on the second statistical model. Similar to the first statistical model, the second statistical model may also be selected from one of the following models, including but not limited to: the microarray data linear model (limma) model, the logistic regression model, the linear discriminant analysis (LDA) model, the conditional logistic regression model, the lasso regression model, the ridge regression model, the random forest, the support vector machine, or the probabilistic unit regression model.

[0016] Optionally, the first statistical model and the second statistical model may be different or substantially the same. Further, according to some embodiments, the limma model may be used as both the first and second statistical models.

[0017] According to some implementation methods, at least one miRNA in the selected miRNA biomarker set is one of the top n miRNAs in the sorting.

[0018] In the cancer diagnostic model development method, step (2) may optionally further include the following sub-step: (c) evaluating the performance of the diagnostic model through cross-validation on the training set. Here, the cross-validation may be at least 2-fold (e.g., 2-fold, 3-fold, 5-fold, 10-fold, etc.). Further optionally, sub-step (c) includes at least one of the following: calculating the area under the curve (AUC) of the receiver operating characteristic (ROC) curve and evaluating the performance of the diagnostic model based on it; or calculating the specificity and sensitivity of the diagnostic model and evaluating the performance of the diagnostic model based on the specificity and sensitivity. It should be noted that, in addition to AUC and specificity / sensitivity, other parameters such as overall accuracy, positive likelihood ratio, negative likelihood ratio, positive predictive value, negative predictive value, Youden's J, etc., may also be used to evaluate the performance of the diagnostic model in sub-step (c) of step (2) above.

[0019] In some embodiments, the method for developing the cancer diagnosis model may further include the following steps: (3) validating the diagnosis model based on a validation set. Here, the validation set includes the expression profiles of a selected set of miRNA biomarkers obtained from non-cancer subjects and cancer patients suffering from the target cancer. Further optionally, step (3) includes at least one of the following: calculating the AUC of the ROC curve and evaluating the performance of the diagnosis model based thereon; or calculating the specificity and sensitivity of the diagnosis model and evaluating the performance of the diagnosis model based on the specificity and sensitivity. It should be noted that in addition to AUC and specificity / sensitivity, other parameters such as overall accuracy, positive likelihood ratio, negative likelihood ratio, positive predictive value, negative predictive value, Youden index, etc. can also be used to validate the diagnosis model in step (3) above.

[0020] There is further provided a system for developing a diagnosis model for detecting a target cancer, comprising: a processor; and a non-transitory storage medium containing program instructions executed by the processor. Here, the program instructions are configured to cause the processor to execute each of the steps of the cancer diagnosis model development method according to any of the above embodiments in the first aspect.

[0021] There is also provided a non-transitory storage medium configured to store computer-executable program instructions that, when executed by a processor, cause the processor to execute each of the steps of the cancer diagnosis model development method according to any of the above embodiments in the first aspect.

[0022] In a second aspect, the present disclosure further provides a method for detecting a target cancer from a subject by means of a diagnosis model. Here, the diagnosis model is developed by the cancer diagnosis model development method according to any of the above embodiments.

[0023] According to some embodiments, the cancer detection method includes the following steps:

[0024] (A) determining the expression profiles of a selected set of miRNA biomarkers from a biological sample obtained from the subject;

[0025] (B) calculating a diagnostic index of the biological sample based on the expression profiles of the selected set of miRNA biomarkers, wherein the diagnostic index is calculated based on the following formula:

[0026]

[0027] where n is the total number of at least one miRNA in the selected set of miRNA biomarkers, miRNA i is the expression level of the i-th miRNA in the selected set of miRNA biomarkers (0 < i ≤ n), and ti The weights of the i-th miRNA; and

[0028] (C) Based on the calculated diagnostic index, the subject is classified as having the target cancer or not having the target cancer, wherein if the calculated diagnostic index is greater than or equal to a predetermined threshold, the subject is classified as having the target cancer; otherwise, the subject is classified as not having the target cancer.

[0029] Here, optionally, each miRNA in the selected miRNA biomarker set is from the first 200 miRNAs listed in Table 1, and n is less than or equal to 200.

[0030] According to some implementations of the cancer detection method, weights from the limma model are used to calculate the diagnostic index.

[0031] Here, according to some embodiments of the cancer detection method, at least one miRNA in the selected miRNA biomarker set is one of the top n miRNAs in the sorting.

[0032] According to some implementations, n is not less than 4 and not greater than 200, and the classification can achieve an AUC greater than about 0.97 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, bile duct cancer, bladder cancer, glioma, or prostate cancer.

[0033] According to some embodiments, n is not less than 4 and not greater than 200, and the classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least about 0.75 while maintaining a specificity of at least about 0.99. Further according to some embodiments, the classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least about 0.80 while maintaining a specificity of at least about 0.99. Further according to some embodiments, the classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least about 0.90 while maintaining a specificity of at least about 0.99. Furthermore, according to some embodiments, this classification, when used to detect lung cancer, esophageal cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least about 0.95 while maintaining a specificity of not less than about 0.99. Furthermore, according to some embodiments, this classification, when used to detect lung cancer, esophageal cancer, gastric cancer, biliary tract cancer, bladder cancer, or prostate cancer, can achieve a sensitivity of at least about 0.99 while maintaining a specificity of not less than about 0.99.

[0034] According to some embodiments of the cancer detection method, the selected miRNA biomarker set consists of hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, and hsa-miR-663a. Further, according to some embodiments, this classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least approximately 0.75 while maintaining a specificity of not less than approximately 0.99. Further, according to some embodiments, this classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least approximately 0.80 while maintaining a specificity of approximately 0.99. Further, according to some embodiments, this classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least approximately 0.90 while maintaining a specificity of approximately 0.99. Further, according to some embodiments, this classification, when used to detect lung cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least approximately 0.95 while maintaining a specificity of approximately 0.99. Further, according to some embodiments, this classification, when used to detect lung cancer, gastric cancer, biliary tract cancer, or bladder cancer, can achieve a sensitivity of at least approximately 0.99 while maintaining a specificity of approximately 0.99.

[0035] According to some implementations of cancer detection methods, the selected set of miRNA biomarkers consists of hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-663a, and hsa-miR-320a; and this classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity of about 0.99 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.

[0036] According to some implementations of cancer detection methods, the selected set of miRNA biomarkers consists of the top 10 miRNAs from Table 1; and this classification can achieve a sensitivity of at least about 0.99 while maintaining a specificity of about 0.99 when used to detect gastric cancer, esophageal cancer, biliary tract cancer or prostate cancer.

[0037] According to some implementations of cancer detection methods, the selected set of miRNA biomarkers consists of the first 15 miRNAs from Table 1; and this classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity of about 0.99 when used to detect lung cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.

[0038] In any embodiment of the cancer detection method described above, the expression profile of the selected miRNA biomarker set is obtained by at least one of the following methods: RNA blotting, microarray analysis, RNA sequencing, or RNA in situ hybridization, or nucleic acid amplification program; wherein the nucleic acid amplification program includes at least one of reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR.

[0039] In any embodiment of the above-described cancer detection method, the biological sample is a liquid biopsy sample selected from the group consisting of blood samples, serum samples, plasma samples, urine samples, saliva samples, and sputum samples.

[0040] A system for detecting a target cancer in a subject is also provided, comprising: a processor; and a non-transient storage medium containing program instructions executable by the processor. Here, the program instructions are configured to cause the processor to perform various steps of the cancer detection method according to any of the embodiments described above in the second aspect.

[0041] A non-transient storage medium is also provided, which is configured to store computer-executable program instructions that, when executed by a processor, cause the processor to perform the steps of the cancer detection method according to any of the above embodiments of the second aspect.

[0042] Unless otherwise defined elsewhere, the terms used throughout this disclosure are defined as follows.

[0043] Generally, "subject" refers to mammals such as primates, including humans and chimpanzees; pets, including dogs and cats; livestock, including cattle, horses, sheep, and goats; and rodents, including mice and rats. The term "healthy subject" also refers to such mammals that do not have the cancer being tested. It should be noted that this disclosure as a whole focuses more specifically on human subjects, but may optionally be applied to other non-human mammals.

[0044] Unless otherwise stated or defined, terms or abbreviations such as “nucleic acid,” “nucleotide,” “polynucleotide,” “DNA,” “RNA,” and “miRNA” follow the common definitions used in the field.

[0045] As used herein, the terms "polynucleotide" and "nucleic acid" are used interchangeably and refer to all nucleic acids, including RNA, DNA, and RNA / DNA chimeras. DNA includes all of the following: cDNA, genomic DNA, and synthetic DNA. RNA includes all of the following: total RNA, mRNA, rRNA, miRNA, siRNA, snoRNA, snRNA, non-coding RNA, and synthetic RNA.

[0046] As used herein, the term "fragment" is a polynucleotide having a continuous nucleotide sequence with a polynucleotide portion, and preferably having a length of 15 or more nucleotides, such as 15, 16, 17, 18, 19 nucleotides, etc.

[0047] As used herein, the term "gene" is intended to include not only RNA and double-stranded DNA, but also the individual single-stranded DNA strands that make up the double strand, such as the positive strand (or sense strand) or the complementary strand (or antisense strand). There is no particular limitation on the length of a gene. Unless otherwise stated, as used herein, "gene" includes all of the following: double-stranded DNA, including human genomic DNA; single-stranded DNA (positive strand), including cDNA; single-stranded DNA (complementary strand) having a complement to the positive strand; miRNAs (miRNAs) and fragments thereof and their transcripts. "Gene" includes not only the "gene" represented by a specific nucleotide sequence (or SEQ ID NO), but also "nucleic acids" encoding RNA that have an equivalent biological function to the RNA encoded by that gene, such as congeners (i.e., homologs or orthologs), variants (e.g., genetic polymorphisms), and derivatives. Specific examples of such “nucleic acids” encoding homologues, variants, or derivatives may include: “nucleic acids” having a complementary sequence to the nucleotide sequence shown in any of SEQ ID NO:1 to 200 and hybridizing under the stringent conditions described below, or “nucleic acids” having a nucleotide sequence derived from the aforementioned nucleotide sequence by replacing the nucleotide “U” (or “u”) with the nucleotide “T” (or “t”). The functional regions of the “gene” are not particularly limited and may contain, for example, expression regulatory regions, coding regions, exons, or introns. A “gene” may be contained within a cell or may exist independently after being released outside the cell. Alternatively, a “gene” may also be in a state enclosed in vesicles called exosomes.

[0048] Within the entire scope of this disclosure, unless otherwise stated, the term “microRNA” or “miRNA” is intended to refer to a non-coding RNA of 15 to 25 nucleotides that is first transcribed into a hairpin-like RNA precursor, cleaved by a dsRNA-cleaving enzyme with RNase III activity, integrated into a protein complex called RISC, and involved in the repression of mRNA translation. As used herein, the term “miRNA” includes not only the “miRNA” represented by a specific nucleotide sequence (or SEQ ID NO), but also the precursor (pre-miRNA or pri-miRNA) of said “miRNA”, and miRNAs with equivalent biological functions, such as homologs (i.e., homologs or orthologs), variants (e.g., genetic polymorphisms), and derivatives. Such precursors, homologues, variants, or derivatives can be specifically identified using miRBase Release 20 (Kozomara and Griffiths-Jones, 2010), and examples may include “miRNAs” having a nucleotide sequence complementary to a specific nucleotide sequence shown by any of the sequences in SEQ ID NO:1 to 200 and hybridizing under the stringent conditions described below. As used herein, the term “miRNA” may also refer to the gene product of a miR gene. Such gene products include mature miRNAs (e.g., the 15- to 25-nucleotide or 19- to 25-nucleotide noncoding RNAs involved in mRNA translation repression described above) or miRNA precursors (e.g., the pre-miRNA or pri-miRNA described above).

[0049] As used herein, the term "probe" includes polynucleotides, or polynucleotides derived from said RNA, and / or their complementary polynucleotides, used for the specific detection of gene expression.

[0050] As used herein, the term "primer" or "amplification primer" includes a polynucleotide that specifically recognizes and amplifies the RNA produced by the expression of a gene, or a polynucleotide derived from said RNA, and / or its complementary polynucleotide.

[0051] In this context, a complementary polynucleotide (complementary strand or antisense strand) refers to the full-length sequence or a portion thereof (here, for convenience, the full-length or partial sequence is referred to as the positive strand) of a polynucleotide that forms a complementary base relationship based on A:T(U) and G:C base pairing with a nucleotide sequence defined by any of the sequences in SEQ ID NO:1 to 200, or a nucleotide sequence derived from said nucleotide sequence by replacing nucleotide "U" (or "u") with nucleotide "T" (or "t"), based on the nucleotide sequence defined by any of the sequences in SEQ ID NO:1 to 200. However, such complementary strands are not limited to sequences that are completely complementary to the target positive strand nucleotide sequence; the degree of complementarity is sufficient to allow hybridization with the target positive strand under stringent conditions.

[0052] As used herein, the term "stringent condition" refers to a condition under which the nucleic acid probe hybridizes more strongly with its target sequence than with other sequences (e.g., the measurement is equal to or greater than the background measurement mean plus the background measurement standard deviation × 2). Stringent conditions depend on the sequence itself and can vary depending on the environment in which hybridization takes place. By controlling the stringency of hybridization and / or washing conditions, target sequences that are 100% complementary to the nucleic acid probe can be identified. Specific examples of "stringent conditions" will be mentioned below.

[0053] As used herein, the term "Tm value" refers to the temperature at which the moiety of a polynucleotide unwinds into a single strand, resulting in a 1:1 ratio of double strands to single strands.

[0054] As used herein, the term "variant" in the context of nucleic acids means: a naturally occurring variant resulting from polymorphism, mutation, etc.; a variant containing one, two, or three or more nucleotide deletions, substitutions, additions, or insertions in, or in part of, the nucleotide sequence represented by any of the sequences in SEQ ID NO:1 to 200, or a nucleotide sequence derived from such nucleotide sequence by replacing nucleotide "U" (or "u") with nucleotide "T" (or "t"); a variant in, or in part of, the nucleotide sequence represented by any of the sequences in SEQ ID NO:1 to 200; a variant in, or in part of ... NO: A variant of a precursor miRNA containing one or more nucleotides of deletion, substitution, addition, or insertion in the nucleotide sequence of any of the sequences shown in NO:1 to 200, or a nucleotide sequence derived from such nucleotide sequence by replacing nucleotide "U" (or "u") with nucleotide "T" (or "t"); a variant having approximately 90% or higher, approximately 95% or higher, identity with each of these nucleotide sequences or their partial sequences; or a nucleic acid that hybridizes to a polynucleotide or oligonucleotide including each of these nucleotide sequences or their partial sequences under the stringent conditions defined above. Variants may be prepared using known techniques such as site-directed mutagenesis or PCR-based mutagenesis.

[0055] The term “percentage of identity (%)” can be determined using the aforementioned BLAST or FASTA-based protein or gene retrieval systems (Zhang et al., 2000; Altschul et al., 1990; Pearson et al., 1988) with or without vacancies.

[0056] The term “derivative” is intended to include, but is not limited to, modified nucleic acids, such as derivatives labeled with fluorophores, derivatives containing modified nucleotides (e.g., nucleotides containing groups such as halogens, alkyl groups such as methyl, alkoxy groups such as methoxy, thio, or carboxymethyl, and nucleotides modified by base rearrangement, double bond saturation, deamination, or replacement of oxygen molecules with sulfur atoms), PNAs (peptide nucleic acids; Nielsen et al., 1991), and LNAs (locked nucleic acids; Obika et al., 1998).

[0057] The "nucleic acid" capable of specifically binding to polynucleotides selected from the above-mentioned miRNAs is a synthetic or prepared nucleic acid, and particularly includes "nucleic acid probes" or "primers". The "nucleic acid" is used directly or indirectly to: detect whether a subject has cancer, diagnose the severity, degree of improvement, or treatment sensitivity of cancer, or screen candidate substances for cancer prevention, improvement, or treatment. Regarding cancer development, the "nucleic acid" includes nucleotides, oligonucleotides, and polynucleotides capable of specifically recognizing and binding to transcripts or synthetic cDNA nucleic acids represented by any of the sequences in SEQ ID NO:1 to 200 in vivo, especially in samples such as bodily fluids (e.g., blood or urine). Based on the above characteristics, nucleotides, oligonucleotides, and polynucleotides can be effectively used as probes to detect the above-mentioned genes expressed in vivo, in tissues, or in cells, or as primers to amplify the above-mentioned genes expressed in vivo.

[0058] As used herein, the term “detection” is used interchangeably with the terms “examination,” “measurement,” or “detection and decision support.” As used herein, the term “assessment” is intended to include diagnosis or assessment support based on examination or measurement results.

[0059] As used within the scope of this disclosure, each of the terms “p-value,” “accuracy,” “AUC,” “sensitivity,” and “specificity” is generally understood to conform to a common definition well known to those skilled in the art, and is specifically defined as follows:

[0060] As used in this article, the terms "p-value," "p-value," "p-value," "p-value," or "p" refer to the probability, in a statistical test, of observing a more extreme result than the statistic actually calculated from the data, assuming the null hypothesis is true. Therefore, the smaller the "p-value," the more significant the difference between the subjects being compared.

[0061] The term "AUC" refers to the area under the receiver operating characteristic (AUC) curve. The term "accuracy" refers to the value of (number of true positives + number of true negatives) / (total number of samples). Accuracy represents the proportion of correctly identified samples out of the total samples and is a core metric for evaluating test performance.

[0062] As used in this article, the term "sensitivity" refers to the value of (number of true positives) / (number of true positives + number of false negatives). High sensitivity allows cancer to be detected, thus enabling clinical treatment intervention.

[0063] As used in this article, the term "specificity" refers to the value of (number of true negatives) / (number of true negatives + number of false positives). High specificity can prevent healthy subjects from being misdiagnosed as cancer patients and undergoing unnecessary additional testing, thereby reducing the burden on patients and lowering medical costs.

[0064] Unless otherwise specified, the following outlines the available techniques that can be used to determine the expression profiles of miRNA biomarker sets.

[0065] It should be noted that determining the expression profile of a miRNA biomarker set essentially involves determining the expression level of each and every miRNA contained in the set. Preferably, the expression levels of all miRNAs in the set are determined simultaneously in a single, well-controlled experiment. Alternatively, the expression levels of these miRNAs can be determined in more than one experiment using different experimental procedures.

[0066] As used herein, measuring or detecting the expression of any miRNA contained in the miRNA biomarker set includes measuring or detecting any nucleic acid transcript corresponding to said miRNA.

[0067] Expression can typically be detected or measured based on miRNA or its corresponding reverse-transcribed cDNA levels. Any quantitative or qualitative method for measuring RNA or cDNA levels can be employed. Suitable methods for detecting or measuring miRNA or cDNA levels include, for example, RNA blotting, microarray analysis, RNA sequencing, RNA in situ hybridization, or nucleic acid amplification procedures such as reverse transcription PCR (RT-PCR) or real-time RT-PCR, also known as quantitative RT-PCR (qRT-PCR), and digital RT-PCR. These methods are well known in the art (see, for example, Sambrook et al., 2012). Other techniques include digital multiplex analysis of gene expression, such as… (NanoString Technologies, Seattle, WA) Gene expression assays, which are further described in US20100112710 and US20100047924.

[0068] Detection of target nucleic acids typically involves hybridization between a target (e.g., miRNA or cDNA) and a probe. The sequences of miRNAs used for various cancer gene expression profiles are known. Therefore, those skilled in the art can readily design hybridization probes for detecting these miRNAs (see, for example, Sambrook et al., 2012). For instance, polynucleotide probes that specifically bind to the miRNA transcripts (or cDNA synthesized therefrom) described herein can be prepared using conventional techniques (e.g., PCR or synthesis) based on the nucleic acid sequence of the miRNA or cDNA target itself. As used herein, the term "probe" refers to a portion or part of a polynucleotide sequence comprising about 10 or more consecutive nucleotides, about 15 or more consecutive nucleotides, or about 20 or more consecutive nucleotides. In some embodiments, the polynucleotide probe will comprise 10 or more nucleic acids, 15 or more nucleic acids, or 20 or more nucleic acids. To ensure sufficient specificity, the probe complement to the target sequence can have approximately 90% or more, such as approximately 95% or more (e.g., approximately 98% or more, or approximately 99% or more), as determined, for example, using the known Basic Local Alignment Search (BLAST) algorithm (available from the National Center for Biotechnology Information (NCBI) Bethesda, MD).

[0069] Each probe can be substantially specific to its target to avoid any cross-hybridization and false positives. An alternative to using specific probes is to employ specific reagents when obtaining material from the transcript (e.g., during cDNA production, or during amplification using target-specific primers). In both cases, specificity can be achieved by partial hybridization with a target that is substantially unique within the group of the miRNA being analyzed; for example, hybridization with a poly-A tail does not provide specificity. If the target has multiple splice variants, it is possible to design a hybridization reagent that recognizes regions common to each variant and / or use more than one reagent, where each reagent recognizes one or more variants.

[0070] The stringency of hybridization reactions can be readily determined by one skilled in the art and typically relies on empirical calculations of probe length, washing temperature, and salt concentration. Generally, longer probes may require higher temperatures for proper annealing, while shorter probes may require lower temperatures. When the complementary strand is present below its melting temperature, hybridization often depends on the re-annealing capability of the denatured nucleic acid sequence. The higher the desired degree of homology between the probe and the hybridizable sequence, the higher the usable relative temperature. As a result, higher relative temperatures tend to make reaction conditions more stringent, while lower temperatures are less stringent.

[0071] “Strict conditions” or “highly stringent conditions” as defined herein can be identified by, but are not limited to, the following conditions: (1) washing with low ionic strength and high temperature, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate at 50°C; (2) the use of a denaturing agent during hybridization, such as formamide, for example, 50% (v / v) formamide with 0.1% bovine serum albumin / 0.1% Ficoll / 0.1% polyvinylpyrrolidone / 50 mM sodium phosphate buffer at pH 6.5 with 750 mM sodium chloride and 75 mM sodium citrate at 42°C; or (3) the use of 50% formamide, 5x SSC (0.75 M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, and 5x SSC at 42°C. Denhardt solution, sonicated salmon sperm DNA (50 μg / ml), 0.1% SDS, and 10% dextran sulfate were used for washing with 0.2×SSC (sodium chloride / sodium citrate) at 42 °C and with 50% formamide at 55 °C, followed by highly stringent washing with 0.1×SSC containing EDTA at 55 °C. “Medium stringent conditions” are those described by Sambrook et al. (1989), but are not limited to them, and include the use of wash solutions and hybridization conditions (e.g., temperature, ionic strength, and SDS%) with lower stringency than those described above. An example of moderately stringent conditions involves overnight incubation at 37°C in a solution comprising: 20% formamide, 5×SSC (150 mM NaCl, 15 mM sodium citrate), 50 mM sodium phosphate (pH 7.6), 5×Denhardt solution, 10% dextran sulfate, and 20 mg / mL denaturing and cleaving salmon sperm DNA, followed by washing the filter with 1×SSC at approximately 37–50°C. Those skilled in the art will recognize how to adjust the temperature, ionic strength, etc., as necessary, depending on factors such as probe length.

[0072] In some embodiments, microarray analysis, RNA blotting, RNA in situ hybridization, or PCR-based methods are employed. In this regard, measuring the expression of the aforementioned miRNAs in a biological sample may include, for example, contacting a sample containing or suspected of containing cancer cells with a polynucleotide probe specific to the target miRNA, or with primers designed to amplify a portion of the target miRNA, and detecting the binding of the probe to the nucleic acid target, or the amplification of the nucleic acid, respectively. Detailed protocols for PCR primer design are known in the art (see, for example, Sambrook et al., 2012). In some embodiments, qRT-PCR can be performed on the miRNA obtained from the sample. Reverse transcription can be performed by any method known in the art, such as using the OmniscriptRT kit (Qiagen). Subsequently, the resulting cDNA can be amplified using any amplification technique known in the art. The expression of the miRNA can be analyzed, for example, using a control sample as described below. As described herein, overexpression or underexpression of the miRNA relative to a control can be measured to determine the miRNA expression profile of an individual biological sample. Similarly, detailed protocols for preparing and using microarrays to analyze miRNA expression are known in the art and are also described herein.

[0073] As used in this article, RNA sequencing (RNA-seq), also known as whole transcriptome shotgun sequencing, refers to any of the various high-throughput sequencing technologies used to detect the presence and content of RNA transcripts in real time. See Wang, Z., M. Gerstein, and M. Snyder, RNA-Seq: a revolutionary tool for transcriptomics, NATREV GENET, 2009.10(1):p.57-63. RNA-seq can be used to reveal the transient expression profile (snapshot) of miRNAs in a sample genome at a specific time point. In some embodiments, miRNAs are reverse transcribed into cDNA fragments before sequencing; in other embodiments, miRNAs can be sequenced directly without conversion to cDNA. Aptamers can be attached to the 5' and / or 3' ends of miRNAs, and miRNAs or cDNAs can optionally be amplified by, for example, PCR. These fragments were then sequenced using high-throughput sequencing technologies, such as those from Roche (e.g., the 454 platform), Illumina, Inc., and Applied Biosystems (e.g., the SOLiD system).

[0074] The following points should be noted.

[0075] Note that the background description includes information that may help in understanding the invention. However, it is not an admission that any information provided herein is prior art or related to the currently claimed subject matter, or that any publication explicitly or implicitly referenced is prior art.

[0076] It should also be noted that, in interpreting both the specification and the claims, all terms should be interpreted in the context and in the broadest possible sense. In particular, the terms "comprise" and "comprising" should be interpreted as referring to elements, components, or steps in a non-exclusive manner, indicating that the referenced element, component, or step may be present or utilized or combined with other elements, components, or steps not explicitly referenced. Unless the context clearly indicates otherwise, "a," "an," and "the" include plural references. Furthermore, as used herein, "in" includes both "in" and "on," unless the context clearly indicates otherwise. In describing and claiming certain embodiments of the invention, quantities used to express amounts of ingredients, properties such as concentration, reaction conditions, etc., should be understood to be, in some cases, modified by the term "about." As used herein, the terms “about,” “approximately,” “about,” or similar terms, when referring to a particular measurable value (such as a parameter, quantity, duration of time, etc.), are intended to encompass both the particular value itself and a range of deviations from that value—such as a deviation of + / -20% or less, or alternatively, a deviation of + / -10% or less, ..., provided that such deviations are reasonably practicable in the actual operation of the embodiments of this disclosure. Therefore, the values ​​modified by “about” or “about” are themselves also within the scope of this explicit disclosure. Ranges of numerical values ​​referenced herein are intended merely as a shorthand method to refer individually to each independent numerical value falling within that range. The use of any and all instances or exemplary language (e.g., “such as”) provided with respect to certain embodiments of this document is intended only to better illustrate the invention and does not constitute a limitation on the scope of the invention as otherwise claimed. Attached Figure Description

[0077] Figure 1 A block diagram illustrating a method for developing a cancer diagnostic model provided by some embodiments of the present disclosure is shown.

[0078] Figure 2 Computerized systems according to some embodiments of the present disclosure are illustrated.

[0079] Figures 3A-3B Together, they illustrate the workflow and research design of the dataset, in which Figure 3A The construction of the training and validation datasets is shown, where Figure 3B The research design for model development and validation is shown.

[0080] Figures 4A-4B This demonstrates the development of a diagnostic model using cross-validation on the training set, where Figure 4A The process of cross-validation is shown, and Figure 4B The performance of different diagnostic models constructed using the first N miRNAs (N = 1, 2, 3, ..., 40) is shown.

[0081] Figure 5 The diagnostic performance of different diagnostic models was compared, and the differences between these models lie in the number of the top N miRNAs (N ranges from 4 to 200).

[0082] Figures 6A-6B The diagnostic performance of the 4-miRNA model in a multi-cancer training set is shown, where... Figure 6B The ROC of the 4-miRNA model is shown, and Figure 6C shows a scatter plot of the diagnostic index.

[0083] Figures 7A-7D The diagnostic performance of the 4-miRNA model on validation set 1 (lung cancer validation dataset) is shown, where Figure 7A The ROC of the 4-miRNA model is shown. Figure 7B A scatter plot of the diagnostic index is shown. Figure 7C A scatter plot comparing the diagnostic indices of serum samples before and after surgery is shown. Figure 7D A scatter plot of diagnostic indices from the clinical subset is shown. ADC: adenocarcinoma; SqCC: squamous cell carcinoma; LCC: large cell carcinoma; SCLC: small cell lung cancer.

[0084] Figures 8A-8B The diagnostic performance of the 4-miRNA model on validation sets 2 and 3 is shown. Figure 8A A scatter plot of the diagnostic indices in validation set 2 is shown. Figure 8B A scatter plot of the diagnostic indices in validation set 3 is shown. Detailed Implementation

[0085] It is generally believed that if a diagnostic model is to be developed for the detection of a specific disease (such as a specific type of cancer), only data specific to that specific disease (such as clinical data, molecular data, etc.) should be used, and data from other diseases (such as different types of cancer) should be avoided. This is because data from different diseases (such as different types of cancer) can introduce interference signals into the development of the diagnostic model, thereby causing the developed diagnostic model to fail to achieve the desired performance.

[0086] In developing our cancer detection method, we used serum miRNA microarray datasets from more than one cancer type. Surprisingly, we found that incorporating miRNA expression data from more than one cancer type resulted in a diagnostic model that could sensitively and reliably detect specific cancer types. Briefly, as will be described in more detail in the Examples section below, our training set used miRNA expression datasets from a total of seven cancer types (lung cancer, ovarian cancer, liver cancer, bladder cancer, esophageal cancer, gastric cancer, and prostate cancer). Unexpectedly, our cross-validation using the training set and subsequent validation using different validation sets both demonstrated that the diagnostic model developed based on data from these seven cancer types not only performed excellently for these seven cancer types, but also surprisingly showed remarkable superiority for most other tested cancers (biliary tract cancer, colorectal cancer, glioma, pancreatic cancer, and sarcoma).

[0087] Based on the findings of this research, the first aspect of this disclosure provides a method for developing cancer diagnostic models.

[0088] This model development approach essentially involves using data from multiple cancer types to develop diagnostic models for detecting target cancers.

[0089] like Figure 1 An illustrative example is a method for developing a diagnostic model for detecting a target cancer, provided according to some embodiments of this disclosure, which essentially comprises the following three main steps S100-S300:

[0090] S100: Construct a training set including miRNA expression profiles obtained from non-cancer subjects and patients with multiple cancer types (≥2 types of cancer);

[0091] S200: Develop a diagnostic model based on the above training set; and

[0092] S300: Validate the diagnostic model based on a validation set, which includes expression profiles of a selected set of miRNA biomarkers obtained from non-cancer subjects and patients with the target cancer.

[0093] In step S100, the constructed training set will be used to establish the diagnostic model in step S200. This training set includes expression profiles of various miRNAs (i.e., miRNA expression profiles) obtained from non-cancer subjects and patients with multiple cancer types. When developing the cancer detection and diagnostic model in step S200, these expression profiles are essentially used as controls and cases, respectively.

[0094] Here, the multiple cancer patients include at least two cancer types. Specifically, these cancer patients include a total of N cancer types (N≥2), meaning the multiple cancer patients may include a first subset of patients with a first cancer type, a second subset of patients with a second cancer type, ..., an Nth subset of patients with an Nth cancer type. According to some embodiments, the at least two cancer types may include the same cancer type as the target cancer. According to some other embodiments, the at least two cancer types do not include the target cancer.

[0095] As used herein and throughout this disclosure, the term “cancer” refers to a specific “cancer type,” defined as a particular type of human disease involving an abnormally increased number of specific cell types (e.g., epithelial cells, connective tissue cells, blood cells, etc.) and the potential to invade or spread to other parts of the body. Cancer types covered herein include, but are not limited to, carcinomas, sarcomas, lymphomas, germ cell tumors, blastomas, etc., and may include both benign and malignant tumors. Non-limiting examples of cancer types may include lung cancer, breast cancer, esophageal cancer, prostate cancer, gastric cancer, pancreatic cancer, liver cancer, ovarian cancer, colorectal cancer, biliary tract cancer, kidney cancer, bladder cancer, brain tumors (e.g., gliomas), sarcomas (e.g., osteosarcoma, chondrosarcoma, fibrosarcoma, etc.), leukemia, lymphoma, germ cell tumors (e.g., germ cell tumors, seminoma, etc.), hepatoblastoma, medulloblastoma, etc.

[0096] In an illustrative example targeting pancreatic cancer, the above-described cancer diagnostic model development method can be used to develop a diagnostic model specifically designed for the detection of pancreatic cancer. This allows the use of miRNA expression profiles obtained from non-cancer subjects (i.e., controls) and patients with multiple cancer types (i.e., cases). Optionally, the multiple cancer types may include a subset of pancreatic cancer patients and at least one other subset of patients with other cancers (e.g., lung cancer patients, glioma patients, etc.). Optionally, the multiple cancer types may also include subsets of cancer patients who are not pancreatic cancer patients (e.g., lung cancer patients, glioma patients, prostate cancer patients).

[0097] As used herein and throughout this disclosure, “expression profile” refers to a collection of expression information about a given set of biomolecules (e.g., miRNAs, proteins, DNA molecules, etc.), which may include expression data obtained for each member molecule in that given set of biomolecules. Various methods may be used to obtain expression profiles. For example, the expression profiles of miRNAs described in this disclosure may optionally be obtained using RNA blotting, microarray analysis, RNA sequencing, or RNA in situ hybridization, or optionally using nucleic acid amplification procedures, including reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR. miRNA expression profiles may be obtained directly from cancerous or tumor tissue of a patient with a specific type of cancer, or from other tissues of the patient such as blood, serum, plasma, saliva, sweat, etc. In the latter case, the miRNAs measured in the expression profile may exist as cell-free miRNAs released into the circulatory system by tumor cells. Regardless of the sample source or assay method, the obtained miRNA expression profiles should be understood to be covered within the scope of this disclosure.

[0098] According to some implementations, step S200 of developing a diagnostic model based on a training set includes the following sub-steps:

[0099] S210: Perform differential expression analysis on the miRNA expression profile to rank the miRNAs using adjusted p-values; and

[0100] S220: Establish a diagnostic model based on a selected set of miRNA biomarkers.

[0101] Here, substep S210 performs differential expression analysis on the miRNA expression profile, thereby ranking the miRNAs using adjusted p-values. Differential expression analysis has been widely implemented, for example, as reported by Zhang et al. in 2014, the entire contents of which are incorporated herein by reference. It can be implemented using a first statistical model, which can be selected from one of the following statistical models: a microarray data linear model (limma) model, a logistic regression model, a linear discriminant analysis (LDA) model, a conditional logistic regression model, a lasso regression model, a ridge regression model, a random forest, a support vector machine, or a probabilistic unit regression model. It should be noted that other statistical models can also be used as the first statistical model in this paper.

[0102] Each of the terms "linear models for microarray data (limma) model" (Ritchie et al., 2015), "logistic regression model" (Venable and Ripley, 2002), "linear discriminant analysis (LDA) model" (Venable and Ripley, 2002), "conditional logistic regression model" (Venable and Ripley, 2002), "lasso regression model" (Tibshirani, 1996), "ridge regression model" (Hoerl and Kennard, 1970), "random forest" (Ripley, 1996), "support vector machine" (Ripley, 1996), and "probit regression model" (Venable and Ripley, 2002) is essentially a probability modeling statistical model that conforms to the definitions well-known to those skilled in the art, and the details thereof can be referred to the references cited immediately after each term.

[0103] As used herein, the term "adjusted p-value" refers to the p-value adjusted for multiple testing. The ranking of miRNAs (e.g., 1, 2, 3, …, etc.) is based on the adjusted p-value, such that the miRNA with the smallest adjusted p-value is ranked highest as 1, and the smaller the adjusted p-value, the higher the ranking of the miRNA.

[0104] Here, in sub-step S220 of establishing a diagnostic model based on a selected set of miRNA biomarkers, the selected set of miRNA biomarkers basically includes at least one miRNA, and each miRNA in the selected set of miRNA biomarkers is selected such that its ranking is less than or equal to (i.e., not greater than) a preset cut-off value m. Here, the "predetermined cut-off value" can be any integer greater than or equal to one (i.e., m ≥ 1, such as m = 1, 2, 3, 5, 10, 15, 20, 50, 100, 200, etc.).

[0105] According to some embodiments of the cancer diagnosis model development method, sub-step S220 of establishing a diagnostic model based on a selected set of miRNA biomarkers includes: calculating a diagnostic index based on the expression profiles of the selected set of miRNA biomarkers, wherein the diagnostic index is calculated based on the following formula:

[0106]

[0107] In the above formula (I), n is the total number of at least one miRNA in the selected set of miRNA biomarkers; miRNA i is the expression level of the i-th miRNA in the selected set of miRNA biomarkers; i is an integer greater than zero and less than or equal to n (i.e., 0 < i ≤ n); and, t iThe weight is assigned to the i-th miRNA. Furthermore, the total number n of miRNAs in the selected miRNA biomarker set must be less than or equal to a preset cutoff value m (i.e., n ≤ m).

[0108] According to some implementation methods, the diagnostic index is calculated using an unweighted method, and each miRNA in the selected miRNA biomarker set has an equal weight (e.g., 1, 2, 0.5, etc.). In other words, for any miRNA in the selected miRNA biomarker set... i , t i =C (i.e., a constant).

[0109] According to some other implementations, the diagnostic index is calculated using a weighted method, and thus the weight t of the i-th miRNA is... i Based on the second statistical model.

[0110] Depending on the implementation method, the second statistical model may be the same as or different from the first statistical model described above, and may be selected from limma model, logistic regression model, LDA model, conditional logistic regression model, lasso regression model, ridge regression model, random forest, support vector machine or probabilistic unit regression model, etc.

[0111] According to some implementations employing weighted methods, such as those shown in the examples below, the limma model is selected as the second statistical model. Further, according to some implementations, both the first and second statistical models employ the limma model. Thus, in sub-step S210, when performing differential expression analysis on the miRNA expression profile, coefficients are generated for all examined miRNAs, and the coefficients corresponding to at least one miRNA in the selected miRNA biomarker set upon which the diagnostic model is based can be used as weights when calculating the diagnostic index in sub-step S220.

[0112] In some other implementations employing weighting methods, a second statistical model based on obtaining the weights of each miRNA in a selected set of miRNA biomarkers differs from a first statistical model based on miRNA ranking. For example, the limma model is used as the first statistical model, while the LDA model is used as the second statistical model.

[0113] According to some implementation methods of the cancer diagnostic model development method described above, the n miRNAs in the selected miRNA biomarker set are the top n miRNAs obtained from the differential expression analysis in sub-step S210. In other words, the selected miRNA biomarker set includes n selected miRNAs, which collectively represent the top 1, 2, 3, ..., n miRNAs. In an exemplary instance, the selected miRNA biomarker set may consist of 4 miRNAs, which respectively have the order 1, 2, 3, and 4 in sub-step S210.

[0114] It should be noted that, according to some other embodiments of this disclosure, the n miRNAs in the selected miRNA biomarker set do not necessarily represent all of the top n miRNAs in the sequence, but may only represent a subset of the top n miRNAs in the sequence. In one exemplary instance, the selected miRNA biomarker set may consist of 5 miRNAs, which are respectively ordered as 1, 2, 3, 4 and 6 in sub-step S210.

[0115] According to some embodiments of the cancer diagnostic model development method, after sub-step S220, step S200 may optionally further include the following sub-steps:

[0116] S230: Evaluate the performance of the diagnostic model using cross-validation on the training set.

[0117] As used herein, the term "cross-validation" refers to a method of processing data in a given dataset; and specifically, as used herein, it refers to dividing the training set into N equal parts (hence the term N-fold cross-validation), where each part alternately serves as the validation set, and the remaining parts together serve as the training set. A diagnostic model is developed on this training set using the same method described in S210 and S220, and subsequently validated on the validation set. Here, the number of folds N in the cross-validation can be any integer greater than or equal to 2 (i.e., N≥2, for example, 2, 3, 4, 5, ..., 10, 20, etc.).

[0118] In sub-step S230, the performance of the diagnostic model obtained in sub-step S220 can be evaluated based on one or more different parameters.

[0119] According to some implementations, the sub-step S230 of evaluating the performance of the diagnostic model includes: calculating the area under the receiver operating characteristic (ROC) curve (AUC) and evaluating the performance of the diagnostic model based on it (i.e., based on the AUC). Here, to evaluate the performance of the diagnostic model based on the calculated AUC value, a preset cutoff value for the AUC (e.g., 0.80, 0.90, 0.95, 0.99, etc.) can be used: if the calculated AUC value is greater than or equal to the preset cutoff value, the diagnostic model is determined to have good performance; however, if the calculated AUC value is lower than the preset cutoff value, the diagnostic model is determined to have poor or suboptimal performance.

[0120] According to some other implementations, the sub-step S230 of evaluating the performance of the diagnostic model includes: calculating the specificity and sensitivity of the diagnostic model, and evaluating the performance of the diagnostic model based on these (specificity and sensitivity). Here, to evaluate the performance of the diagnostic model based on the calculated sensitivity and specificity values, a preset cutoff value for sensitivity (e.g., 0.80, 0.90, 0.95, 0.98, or 0.99, etc.) for a given specificity can be used: if the sensitivity value calculated for a given specificity is greater than or equal to the preset cutoff value, the diagnostic model is determined to have good performance; however, if the sensitivity value calculated for a given specificity is lower than the preset cutoff value, the diagnostic model is determined to have unsatisfactory or poor performance.

[0121] It should be noted that, in addition to using the AUC of the ROC curve and specificity and sensitivity, other parameters can be used to evaluate the performance of the diagnostic model obtained in sub-step S220 of the cancer diagnostic model development method. Examples of these other parameters include, but are not limited to: overall accuracy, positive likelihood ratio, negative likelihood ratio, positive predictive value, negative predictive value, Youden index, etc.

[0122] It should also be noted that the sub-step S230, which evaluates the performance of the diagnostic model through cross-validation in the training set, is only optional and may be omitted in some embodiments according to this disclosure.

[0123] Using step S300, the diagnostic model obtained in step S200 (and particularly in sub-step S220) is further validated using an independent validation set that does not overlap with the training set of subjects or patients. Here, the validation set includes expression profiles of a selected set of miRNA biomarkers obtained from non-cancer subjects (i.e., controls) and cancer patients with the target cancer (i.e., cases).

[0124] More specifically, step S300 can be implemented by evaluating one or more of the same or different parameters used in the aforementioned sub-step S230, which evaluates the performance of the diagnostic model through cross-validation on the training set. According to some embodiments of this disclosure, the validation uses the AUC of the ROC curve such that: when the AUC is greater than or equal to a preset cutoff value (e.g., 0.80, 0.90, 0.95, 0.99, etc.), the diagnostic model passes validation; conversely, when the AUC is lower than the preset cutoff value, it fails validation. Here, the preset AUC cutoff value used in step S300 may be the same as or different from the preset AUC cutoff value used in the aforementioned sub-step S230. According to some other embodiments of this disclosure, sensitivity and specificity are used for verification, such that: when the sensitivity value calculated for a given specificity (e.g., 0.90, 0.95, 0.98, or 0.99) is greater than or equal to a preset cutoff value (e.g., 0.80, 0.90, 0.95, 0.99), the diagnostic model passes verification; conversely, when the sensitivity value calculated for the given specificity is lower than the preset cutoff value, the verification fails. Furthermore, other parameters may also be used to verify the diagnostic model.

[0125] This disclosure further provides a computerized solution whose function is essentially to implement the various steps and / or sub-steps of the aforementioned cancer diagnostic model development method in a computerized and automated manner. This computerized solution is applicable to the following scenarios: automating the execution of each step S100-S300 of the aforementioned cancer diagnostic model development method, and each sub-step S210-S230 of step S200, by running a software program including program instructions in a computer; this computerized solution has advantages such as high efficiency and ease of operation.

[0126] More specifically, this computerized solution may include a computerized system or computer system (or simply "system"). This system comprises a collection of hardware (e.g., processor, memory, I / O interfaces, storage media, etc.) and software (i.e., computer programs, including operating system software and specific program software, etc.) configured to work together to jointly execute all or part of the steps of the cancer detection and diagnostic model development method described above. According to some embodiments, the system includes a processor (i.e., a controller) and a computer-readable non-transient storage medium communicatively coupled to the processor. The non-transient storage medium is configured to contain software (i.e., program instructions) executable by the processor, and the program instructions are configured to cause the processor to perform various steps and sub-steps of the cancer diagnostic model development method described above.

[0127] As used herein, and throughout this disclosure, the term "processor" may be used interchangeably with "central controller" or "central computing unit (CPU)," and may refer to a single-core processor or a multi-core processor, or multiple processors for parallel processing. As used herein, the term "non-transient" is intended to describe tangible computer-readable storage media, excluding propagating electromagnetic signals, but is not intended to limit the type of physical computer-readable storage device covered by the phrase. Examples may include any tangible or non-transient storage medium or memory medium, such as electronic media, magnetic media, or optical media (e.g., magnetic disks or CD / DVD-ROMs), or non-volatile memory storage (e.g., "flash memory"), etc.

[0128] like Figure 2 As illustrated, in addition to the processor 10 and the computer-readable non-transient storage medium 20, the system 100 may further include a bus 30, a memory 40, an I / O interface 50, and a communication interface 60. The processor 10, storage medium 20, memory 40, I / O interface 50, and communication interface 60 are all communicatively coupled to each other via the bus 30.

[0129] Storage medium 20 stores computer-executable program instructions, which, when executed by processor 10, cause processor 10 to perform steps (1)-(3) of the above method. Memory 40 is configured to temporarily store program instructions obtained from storage medium 20, and processor 10 is configured to execute program instructions temporarily stored in memory 40. I / O interface 50 allows input / output between system 100 and the user to control system 100. Communication interface 60 allows system 100 to communicatively connect with another computing device to exchange data. It should be noted that these computer hardware components can be deployed locally or remotely via networks such as intranets, the Internet, or cloud networks.

[0130] In a second aspect, this disclosure further provides a method for detecting a target cancer in a subject (e.g., a human). Here, the cancer detection method is implemented essentially by means of a diagnostic model developed according to the cancer detection diagnostic model described in any of the above embodiments.

[0131] Specifically, this cancer detection method includes the following main steps:

[0132] S1000: Determine the expression profile of a selected set of miRNA biomarkers from biological samples obtained from subjects;

[0133] S2000: Calculate the diagnostic index of the biological sample based on the expression profile of the selected miRNA biomarker set, wherein the diagnostic index is calculated based on the following formula:

[0134]

[0135] Here, n represents the total number of at least one miRNA in the selected miRNA biomarker set; miRNA i The expression level of the i-th miRNA in the selected miRNA biomarker set; i is an integer greater than zero and less than or equal to n; t i Let be the weight of the i-th miRNA; and

[0136] S3000: Based on the calculated diagnostic index, the subject is classified as having the target cancer or not having the target cancer. If the calculated diagnostic index is greater than or equal to a predetermined threshold, the subject is classified as having the target cancer; otherwise, the subject is classified as not having the target cancer.

[0137] Here, in step S1000 of the cancer detection method provided herein, the subject's "biological sample" may be a liquid biopsy sample such as a blood sample, serum sample, plasma sample, urine sample, saliva sample, or sputum sample, or optionally a sample from tumor tissue. The "selected set of miRNA biomarkers" is essentially the set of miRNA biomarkers selected when developing a diagnostic model for detecting a target cancer, using the cancer detection diagnostic model development method provided as in the first aspect of this disclosure. Determining the expression profile of the selected set of miRNA biomarkers from the biological sample obtained from the subject can be achieved using a variety of probe-based methods, including RNA imprinting, microarray analysis, RNA sequencing, or RNA in situ hybridization, or using a variety of amplification-dependent methods, including reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR. As used herein, the definition of each of the above miRNA detection methods is consistent with the general meaning well known to those skilled in the art.

[0138] In step S2000 of the cancer detection method provided herein, the diagnostic index can be calculated in the same way as in the cancer detection and diagnostic model development method described in the first aspect of this disclosure.

[0139] In step S3000, the term "predetermined threshold" refers to a cutoff value for a diagnostic index that can be used to determine whether a subject has a target cancer based on a given specificity / sensitivity. It is typically predetermined based on an existing dataset comprising a series of diagnostic index values ​​obtained and calculated for an existing population of subjects known to have the disease and / or known not to have the disease. International Patent Application No. WO 2022261039A2 details how subjects can be classified as having a target cancer based on a calculated diagnostic index; the entire contents of that patent application are incorporated herein by reference.

[0140] According to some implementation methods, each miRNA in the selected miRNA biomarker set is from the top 200 miRNAs listed in Table 1 provided in Example 1 below, and the total number of miRNAs in the selected miRNA biomarker set (i.e., "n") is less than or equal to 200. It should be noted that each miRNA in the selected miRNA biomarker set can be from any of the top few miRNAs after sorting in sub-step S210 of step S200 of the above-mentioned cancer detection and diagnostic model development method. Depending on different implementation methods, it can be the top 100, top 50, or top 500, etc.

[0141] According to some preferred embodiments, in step S2000 of the cancer detection method, the diagnostic index is calculated using weights from a limma model. It should be noted that, according to some other embodiments, the diagnostic index can also be alternatively and optionally calculated using weights obtained from other statistical models, such as logistic regression, LDA, conditional logistic regression, lasso regression, ridge regression, random forest, support vector machine, or probabilistic unit regression. It should also be noted that the diagnostic index can also be alternatively and optionally calculated using an unweighted method, whereby each miRNA in the selected miRNA biomarker set has an equal weight (e.g., 1, 2, 0.5, etc.). In other words, for any miRNA in the miRNA biomarker set... i , t i =C (i.e., a constant).

[0142] According to some embodiments of the cancer detection method, the n miRNAs in the selected miRNA biomarker set are the top n miRNAs obtained from differential expression analysis in sub-step S210 of step S200 of the cancer detection and diagnostic model development method described in the first aspect of this disclosure. In other words, the selected miRNA biomarker set includes n selected miRNAs, which collectively represent the top 1, 2, 3, ..., n miRNAs. It should be noted that, according to some other embodiments of this disclosure, the n miRNAs in the selected miRNA biomarker set do not necessarily represent all of the top n miRNAs, but may only represent a subset of the top n miRNAs.

[0143] Here, according to some implementations of the cancer detection method provided herein, n is not less than 4 and not greater than 200 (4≤n≤200), and this classification can achieve an AUC greater than approximately 0.97 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, bile duct cancer, bladder cancer, glioma, or prostate cancer.

[0144] According to some implementations, n is not less than 4 and not greater than 200 (4≤n≤200), and the classification can achieve a sensitivity of at least about 0.75 while maintaining a specificity of not less than about 0.99 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, bile duct cancer, bladder cancer, glioma, or prostate cancer.

[0145] Furthermore, according to some implementation methods, this classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity of not less than about 0.99 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.

[0146] Furthermore, according to some embodiments, this classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity of not less than about 0.99 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.

[0147] Furthermore, according to some implementation methods, this classification can achieve a sensitivity of at least about 0.95 while maintaining a specificity of not less than about 0.99 when used to detect lung cancer, esophageal cancer, gastric cancer, bile duct cancer, bladder cancer, glioma, or prostate cancer.

[0148] Furthermore, according to some implementation methods, this classification can achieve a sensitivity of at least about 0.99 while maintaining a specificity of not less than about 0.99 when used to detect lung cancer, esophageal cancer, gastric cancer, bile duct cancer, bladder cancer, or prostate cancer.

[0149] According to some embodiments of the cancer detection methods provided herein, the selected miRNA biomarker set consists of the first four miRNAs, including hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, and hsa-miR-663a. Thus, this classification achieves a sensitivity of at least approximately 0.75 while maintaining a specificity of not less than approximately 0.99 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Further optionally, this classification achieves a sensitivity of at least approximately 0.80 while maintaining a specificity of approximately 0.99 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer. Further optionally, this classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.90 while maintaining a specificity of approximately 0.99. Further optionally, this classification, when used to detect lung cancer, gastric cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.95 while maintaining a specificity of approximately 0.99. Further optionally, this classification, when used to detect lung cancer, gastric cancer, biliary tract cancer, or bladder cancer, achieves a sensitivity of at least approximately 0.99 while maintaining a specificity of approximately 0.99.

[0150] According to some implementations of the cancer detection methods provided herein, the selected miRNA biomarker set consists of the first 5 miRNAs, including hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-663a, and hsa-miR-320a; and this classification can achieve a sensitivity of at least about 0.80 while maintaining a specificity of about 0.99 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.

[0151] According to some embodiments of the cancer detection methods provided herein, the selected set of miRNA biomarkers consists of the top 10 miRNAs from Table 1; and this classification achieves a sensitivity of at least approximately 0.99 while maintaining a specificity of approximately 0.99 when used to detect gastric cancer, esophageal cancer, biliary tract cancer, or prostate cancer.

[0152] According to some embodiments of the cancer detection methods provided herein, the selected set of miRNA biomarkers consists of the first 15 miRNAs from Table 1; and this classification can achieve a sensitivity of at least about 0.90 while maintaining a specificity of about 0.99 when used to detect lung cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.

[0153] In any embodiment of the cancer detection method described above, the method optionally further includes a step of assessing the subject, wherein the assessment includes detection of cancer diagnosis or cancer recurrence. Here, the phrase "cancer diagnosis" refers to cancer detection in a subject who was previously known not to have cancer; while the phrase "cancer recurrence" refers to cancer detection again in a subject who previously had cancer, which was cleared by treatment and reached a cancer-free state.

[0154] In any embodiment of the above method, when a subject is classified as having cancer, the method may optionally further include the step of performing a diagnostic procedure on the subject. This diagnostic procedure may optionally include a physical examination, pathological examination of a biopsy sample from the subject, immunohistochemical examination, or imaging examinations such as X-ray, computed tomography (CT), ultrasound, and / or magnetic resonance imaging.

[0155] In any embodiment of the above method, when a subject is classified as having cancer, the method optionally further includes the step of administering a treatment regimen to the subject. Here, various known treatment regimens can be applied in the method, including surgical treatment, radiation therapy, chemotherapy, hormone therapy, targeted therapy, immunotherapy, or combinations of the above treatments. These treatment regimens have been well-developed for each of the different cancers mentioned above.

[0156] This disclosure further provides a computerized solution, including a computerized system (including a processor and non-transient storage media) for performing the various steps and / or sub-steps of the cancer detection method described above. Details of this computerized solution can be found in the computerized solution for the cancer detection and diagnostic model development method provided with respect to the first aspect of this disclosure.

[0157] This disclosure further provides a kit for detecting cancer in biological samples obtained from a subject, the kit being essentially used to perform the cancer detection method described above.

[0158] As used herein, and elsewhere in this disclosure, the term "kit" refers to a collection of articles and / or instructions. Articles included in a kit may be physical substances or components thereof. Examples of articles that may be included in a kit disclosed herein may include one or more nucleic acids (e.g., polynucleotides), or one or more devices, instruments, or apparatuses (e.g., molecular arrays or microarrays comprising one or more nucleic acids). Instructions included in a kit may be descriptions of specific steps to be performed (e.g., an instruction manual), which may be printed on a physical medium (e.g., paper, cards, etc.), stored on a computer-readable storage medium (e.g., hard disks, optical discs or CDs, flash drives, etc.), or even stored on the Internet (e.g., accessible cloud space).

[0159] The kit may include at least the following components (1) and (2) (i.e., the product and / or instructions):

[0160] Component (1): at least one nucleic acid, each nucleic acid capable of specifically recognizing each miRNA in the selected miRNA biomarker set, thereby allowing the expression profile of the selected miRNA biomarker set to be obtained from a biological sample. Here, the selected miRNA biomarker set is essentially the set of miRNA biomarkers selected when developing a diagnostic model for detecting a target cancer, using the cancer detection diagnostic model development method provided as in the first aspect of this disclosure.

[0161] Component (2): At least one instruction manual, including a first instruction manual and a second instruction manual. Here, the first instruction manual is configured to substantially implement step S2000 of calculating a diagnostic index of a biological sample based on the expression profile of a selected set of miRNA biomarkers; and the second instruction manual is configured to substantially implement step S3000 of classifying a subject as having or not having cancer, wherein if the calculated diagnostic index is greater than or equal to a predetermined threshold, the subject is classified as having cancer; otherwise, the subject is classified as not having cancer.

[0162] Herein, in component (1) of the kit, the at least one nucleic acid may optionally include a polynucleotide capable of hybridizing under stringent conditions with the following specificities: (a) a polynucleotide comprising or consisting of the nucleotide sequence of each miRNA in the selected miRNA biomarker set, its derivatives, variants with at least 80% sequence identity, or a fragment comprising 15 or more consecutive nucleotides; (b) a polynucleotide comprising or consisting of a nucleotide sequence complementary to the nucleotide sequence of each miRNA in the selected miRNA biomarker set, its derivatives, variants with at least 80% sequence identity, or a fragment comprising 15 or more consecutive nucleotides.

[0163] According to different implementations, at least one instruction manual in component (2) of the kit may further include a third instruction manual for evaluating a subject, wherein the evaluation includes detection of cancer diagnosis or cancer recurrence; or may further include a fourth instruction manual for administering a treatment regimen to a subject when the subject is classified as having cancer.

[0164] According to some embodiments, at least one instruction manual for component (2) of the kit may further include a first supplementary instruction manual for obtaining the expression profile of a selected set of miRNA biomarkers, including procedures for performing RNA imprinting, microarray analysis, RNA sequencing, or RNA in situ hybridization using at least one of the aforementioned nucleic acids. Here, the at least one nucleic acid may optionally be arranged on a molecular array.

[0165] According to some embodiments, the kit may further include at least one set of amplification primers, each set capable of specifically amplifying each of at least one miRNA from the selected miRNA biomarker set in a biological sample. Thus, at least one instruction manual in component (2) of the kit may further include a second supplementary instruction manual for obtaining the expression profile of the selected miRNA biomarker set, including procedures for performing reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR using at least one nucleic acid and at least one set of amplification primers.

[0166] The following provides an embodiment (i.e., Embodiment 1) to illustrate the inventive content described in various aspects of this disclosure.

[0167] Example 1

[0168] In this embodiment, a large multi-cancer training set (hereinafter referred to as the "training set") was constructed using eight publicly available serum miRNA microarray datasets, comprising a total of 6,283 cancer patients and 5,130 non-cancer controls. This training set included multiple cancer types rather than a single cancer type. Differentially expressed miRNAs were identified in the training set, and 10-fold cross-validation was used to determine the optimal number of miRNAs for building a diagnostic model within the training set. The model's performance was then evaluated in three validation sets. More detailed information is provided below.

[0169] method

[0170] Research Design and Construction of Training and Validation Datasets: Eight serum miRNA microarray datasets were identified from the Gene Expression Omnibus (GEO) (Zhang A et al., 2022; Zhang A and Hu H, 2022). These datasets all originated from the nationwide Japanese research project "Development and Diagnostic Technologies for the Detection of miRNAs in Body Fluids," which was designed to characterize serum miRNAs from over 50,000 participants across 13 cancer types using a standardized microarray platform. These eight datasets were initially used to develop individual diagnostic models for the following: lung cancer (GSE137140) (Asakura K et al., 2020), ovarian cancer (GSE106817) (Yokoi A et al., 2018), liver cancer (GSE113740) (Yamamoto Y et al., 2020), bladder cancer (GSE113486) (Usuba W et al., 2019), esophageal squamous cell carcinoma (GSE122497) (Sudo K et al., 2019), gastric cancer (GSE164174) (Abe S et al., 2021), prostate cancer (GSE112264) (Urabe F et al., 2019), and glioma (GSE139031) (Ohno M et al., 2019). After removing duplicate cases, three independent large datasets were obtained: a lung cancer dataset (n=3744) (Zhang A et al., 2022; Asakura K et al., 2020), a combined dataset by incorporating ovarian cancer, liver cancer, and bladder cancer datasets (n=3792) (Zhang A et al., 2022; Yokoi A et al., 2018; Yamamoto Y et al., 2020; Usuba W et al., 2019), and a combined dataset by incorporating esophageal squamous cell carcinoma, gastric cancer, prostate cancer, and glioma datasets (n=3877) (Zhang A and Hu H, 2022; Sudo K et al., 2019; Abe S et al., 2021; Ohno M et al., 2019; Urabe F et al., 2019). Based on these three large datasets, a large training set was constructed for developing diagnostic models to detect multiple cancer types. This set included 1408 cancer patients from seven cancer types (208 lung cancer patients, and 200 patients each from ovarian, liver, bladder, esophageal, gastric, and prostate cancer), and 1408 non-cancer controls matched for patient age and sex. All remaining cases formed three independent validation sets. Figure 3A and Figure 3B ).

[0171] Blood Sample Collection and miRNA Microarray Analysis: Serum sample collection and microarray expression analysis have been described previously (Asakura K et al., 2020). Briefly, total RNA was extracted from 300 μl of serum collected from cancer patients (preoperatively) and non-cancer controls; [The text then abruptly shifts to a different topic:] Blood sample collection and miRNA microarray analysis: Serum sample collection and microarray expression analysis have been described previously (Asakura K et al., 2020). [The text then abruptly shifts again:] Blood sample collection and miRNA microarray analysis: Serum sample collection and miRNA expression analysis have been described previously (Asakura K et al., 2020). The miRNA marker kit is used to label and will be associated with... Human miRNA oligonucleotide microarrays (Toray Industries, Kanagawa, Japan) were used for hybridization. The signal intensity of miRNAs was determined after background subtraction and normalization based on three pre-selected internal control miRNAs (miR-149-3p, miR-2861, and miR-4463).

[0172] Diagnostic Model Development: The identification of miRNA biomarkers and all model development work were completed solely within a multi-cancer training set. Differential miRNA expression between cancer and non-cancer groups was assessed using a microarray data linear model (limma) (Ritchie ME et al., 2015). Then, miRNAs were ranked based on the statistical significance of differential expression, and the top-ranked miRNAs were used to build diagnostic models to distinguish between cancer and non-cancer. A diagnostic index for each model was calculated by linearly summing the expression levels of the selected miRNAs with their corresponding limma statistics. Ten-fold cross-validation was performed to determine the optimal number of miRNAs to include in the final diagnostic model, which would have the highest area under the receiver operating characteristic (ROC) curve (AUC) when distinguishing between cancer and non-cancer. The cutoff point for the diagnostic index was chosen to ensure a model specificity of at least 99% (i.e., a false positive rate <1%), as the model may be used as a screening tool for high-risk general populations.

[0173] Diagnostic Model Validation: These three independent validation datasets contain mutually exclusive samples that were not used in model development, and each dataset has different characteristics, which can be used to validate the developed model. Validation set 1 not only contains a large number of lung cancer case samples, but also contains comprehensive patient-level clinicopathological data compared to the other two validation datasets, making it possible to evaluate the model's performance in early-stage cancer and different histological subtypes. Validation set 2 contains samples from 12 other cancer types, thus expanding the evaluation of the model's performance across multiple cancer types. Validation set 3 includes a large number of cases from four cancer types, including two cancer types with a smaller sample size in validation set 2, allowing for additional independent validation of the model's performance.

[0174] Statistical analysis: AUC analysis of ROC curves, sensitivity, and specificity were used to measure diagnostic performance for detecting cancer and non-cancer subjects. Sensitivity was defined as the proportion of cancer patients correctly identified as cancer by the diagnostic model, while specificity was defined as the proportion of non-cancer subjects correctly identified as non-cancer subjects. limma analysis was performed using the limma package from Bioconductor (Ritchie ME et al., 2015). All statistical analyses were performed using R version 4.2.1.

[0175] result

[0176] Participants and Datasets: Detailed demographic and clinical information for the cancer types with larger sample sizes has been described in the original publication. In brief, the lung cancer dataset (n=1566) included patients with a mean age of 65 years, 57% male, and 62% former or current smokers; adenocarcinoma accounted for 78% of the tumors, squamous cell carcinoma for 14%, and stage I or II for 87%. The bladder cancer dataset (n=392) included patients with a mean age of 68 years, 72% male; 95% were non-metastatic, 88% were lymph node negative, 77% were T1 stage, and 80% were high-grade. The ovarian cancer dataset (n=333) included patients with a mean age of 57 years, 35% stage I or II; epithelial ovarian cancer accounted for 96% (including serous, clear cell, and endometrioid carcinomas, which accounted for 55%, 19%, and 13% respectively histologically). The liver cancer dataset (n=348) included patients with a mean age of 68 years, 78% male, and stage I or II for 70%. The esophageal cancer dataset (n=447) consisted of patients with a mean age of 67 years, 97% being male, and 66% in stage I or II. The gastric cancer dataset (n=1267) included patients with a mean age of 66 years, 77% being male, and all in stage I or II. The glioma dataset (n=196) included patients with a mean age of 56 years and 57% being male. Finally, the prostate cancer dataset (n=769) included patients with a mean age of 68 years, 93% with negative lymph nodes, and 92% without metastasis.

[0177] Cancer diagnostic model development: All diagnostic model development work was carried out on a multi-cancer training set, which included 1408 cancer patients and 1408 age- and sex-matched non-cancer controls. Figure 3B First, limma analysis was used to assess differential miRNA expression between the cancer and non-cancer groups. miRNAs were ranked based on adjusted p-values. The top 200 differentially expressed miRNAs are listed in Table 1. Then, 10-fold cross-validation was performed using a multi-cancer training set. Figure 4AFurthermore, it was revealed that the top four miRNAs (hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, and hsa-miR-663a) had the highest AUC in the ROC analysis and were therefore included in the final diagnostic model. Figure 4B Then, the diagnostic performance of different diagnostic models using different validation sets for detecting different cancers was compared within their respective performance metrics. These different diagnostic models varied in the number of top-ranked miRNAs, N (N ranging from 4 to 200), and the relevant results are as follows: Figure 5 As shown. We calculated the diagnostic index by weighted summation of the expression levels of the four miRNAs and normalized it to a range of 0 to 10. This 4-miRNA model achieved an AUC of 0.994 on the training set. Figure 6A The cutoff value of 5.3 was chosen to demonstrate >99% specificity (i.e., false positive rate <1%) in non-cancer cases, while maintaining an overall sensitivity of 94%. Figure 6B The model's AUC and sensitivity for each of the seven cancer types in the multi-cancer training set ranged as follows: 0.985 and 84% for ovarian cancer, and 0.998 and 100% for bladder and stomach cancer, respectively. Figure 5 ).

[0178] Table 1. The top 200 miRNAs identified through differential expression analysis of multiple cancer training sets.

[0179]

[0180]

[0181]

[0182]

[0183]

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190]

[0191]

[0192]

[0193] Note: In the training set, differential miRNA expression was analyzed using limma in 1408 cancer patients across 7 cancer types and 1408 matched non-cancer controls. miRNAs were ranked by adjusted p-values ​​(calculated by limma), and the top 200 miRNAs were derived from this ranking list.

[0194] Validation of the diagnostic model on independent validation set 1: The performance of the 4-miRNA model was first evaluated on independent validation set 1 (n = 2859, including 1358 lung cancer patients and 1501 non-cancer controls). The model achieved an AUC of 1.000 (…). Figure 7A ), with a specificity of 100% and a sensitivity of 99%. Figure 7B Furthermore, analysis of paired serum samples (preoperative vs. postoperative; n=180) confirmed that the diagnostic indices in the postoperative serum samples had been normalized to the levels of non-cancer controls. Figure 7C Furthermore, the performance of the 4-miRNA model was evaluated in different clinical subsets of validation set 1 (divided by clinical stage, TNM stage, and histological subtype). High sensitivity was observed in all clinical subsets. Of the 24 clinical subsets examined, the model achieved at least 99% sensitivity for the remaining 22 subsets, excluding stage IIB and T3 tumors. Figure 7D In particular, the model demonstrated a sensitivity of >99% for stage I lung cancer, as well as for adenocarcinoma and squamous cell carcinoma.

[0195] Validation of the diagnostic model in independent validation sets 2 and 3: Independent validation set 2 included 1438 patients across 12 other cancer types and 1623 non-cancer controls. Except for breast cancer, the 4-miRNA model achieved at least 90% sensitivity for eight cancer types (biliary tract cancer, bladder cancer, colorectal cancer, esophageal cancer, gastric cancer, glioma, pancreatic cancer, and prostate cancer), and at least 75% sensitivity for three other cancer types (liver cancer, ovarian cancer, and sarcoma). Figure 8A (See Table 2). It is worth noting that although the model has a reasonable AUC of 0.909 for breast cancer, a sensitivity of 1% is still quite low due to the high specificity requirement. Figure 8A(See Table 2). The independent validation set 3 included 2079 patients and 598 non-cancer controls from four cancer types (esophageal cancer, gastric cancer, glioma, and prostate cancer), with sample sizes for all four cancer types significantly larger than their corresponding sample sizes in validation set 2 (esophageal cancer: 247 vs. 124; gastric cancer: 1067 vs. 150; glioma: 196 vs. 40; prostate cancer: 569 vs. 40). The 4-miRNA model achieved AUC >0.99 and sensitivity >99% for all four cancer types, similar to the results observed in validation set 2. Figure 8B (and Table 2). The model's specificity on validation set 3 was slightly lower than its specificity on validation set 2 (0.98 vs. 0.99). Figure 8B Therefore, for validation set 3, a sensitivity analysis was explored: the diagnostic index cutoff was adjusted to 5.6 to improve the specificity of the new model to 99%. With this new cutoff, the model still achieved high sensitivity for all four cancer types, including 99% for gastric cancer, 92% for glioma, 91% for prostate cancer, and 89% for esophageal cancer.

[0196] discuss

[0197] Over the past decade, non-invasive screening for cancer-related epidemiological disorders (MCEDs) by analyzing circulating cell-free nucleic acids and / or proteins in bodily fluids, particularly blood, has garnered significant attention. In this study, we report the development and validation of a serum 4-miRNA diagnostic model, demonstrating that in three large, independent validation sets comprising 8597 participants (4875 cancer patients across 13 cancer types and 3722 non-cancer individuals), this 4-miRNA model can simultaneously detect 12 cancer types with high sensitivity (sensitivity >90% for 9 cancer types and ≥75% for 3 cancer types) while maintaining extremely high specificity of ~99%. Furthermore, the observation that the diagnostic index of postoperative serum samples returned to normal levels suggests the model's potential application value in monitoring treatment response and detecting cancer recurrence.

[0198] Importantly, our model is capable of detecting early-stage cancer with high sensitivity. Specifically, in the validation set 1 of lung cancer patients, the model's detection sensitivity for stage I and II cancer ranged from 98.4% to 99.6%. Figure 7DIn validation sets 2 and 3, although individual patient-level staging information was unavailable, aggregate staging information was provided for 6 of the 12 cancer types examined. First, all gastric cancer patients were in stage I or II, therefore our model's 100% sensitivity applies to early gastric cancer detection. Second, 88% of bladder cancer patients and 93% of prostate cancer patients had lymph node-negative diseases. Therefore, given the model's sensitivities of 99% and 98% for these two cancers, respectively, its sensitivity for stage I or II bladder and prostate cancer should also be very high. Third, 66% of esophageal cancer patients and 70% of liver cancer patients were in stage I or II. It is reasonable to assume that the detection sensitivity for stage I or II of these two cancers should not differ significantly from the reported sensitivities of 92% and 84% including all stages. In summary, based on currently available data in the three validation sets, we conclude that our 4-miRNA model achieves high detection sensitivity for stage I or II disease in six cancer types (lung cancer, gastric cancer, bladder cancer, prostate cancer, esophageal cancer, and liver cancer).

[0199] It is worth noting that the simple four-parameter diagnostic model described in this article is not only significantly less expensive, but can also be developed into an in vitro diagnostic (IVD) test using RT-qPCR, enabling decentralized testing; this is an advantage over NGS-based tests, which are typically used as laboratory-developed tests (LDTs). These characteristics are crucial for promoting the widespread application and accessibility of MCED testing technology—because the target population for such technologies is the general population at high risk or at risk, especially low-income groups.

[0200] In summary, our study provides proof-of-concept data for the development of a blood screening assay based on circulating cell-free miRNA expression profiling that can detect 12 types of cancer; these 12 types are estimated to account for 50% of new cancer cases and 63% of cancer deaths in the United States in 2022 (Siegel RL et al., 2022).

[0201] References

[0202] Zhang C, et al. Analysis of differential gene expression and noveltranscript units of ovine muscle transcriptomes. PLoS One. 2014 Feb 26; 9(2): e89817.

[0203] Ritchie,ME;et al.(2015).limma powers differential expression analysesfor RNA-sequencing and microarray studies.Nucleic Acids Research 43(7),e47.

[0204] Venables,WN and Ripley,BD(2002)Modern Applied Statistics withS.Fourth edition.Springer.

[0205] Tibshirani,Robert(1996).“Regression Shrinkage and Selection via thelasso”.Journal of the Royal Statistical Society.Series B(methodological).Wiley.58(1):267-88.

[0206] Hoerl,Arthur E.;Kennard,Robert W.(1970).“Ridge Regression:BiasedEstimation for Nonorthogonal problems”.Technometrics.12(1):55-67.

[0207] Ripley,B.D.(1996)Pattern Recognition and Neural Networks.CambridgeUniversity Press.

[0208] Zhang A,et al.A Novel Blood-Based microRNA Diagnostic Model with HighAccuracy for Multi-Cancer Early Detection.Cancers 2022,14(6),1450.

[0209] Zhang A and Hu H.Independent validation of a novel noninvasive 4-microRNAdiagnostic model for multicancer early detection.Journal of ClinicalOncology.2022;40(16_suppl):3065-3065.

[0210] Asakura K,et al.A miRNA-based diagnostic model predicts resectablelung cancer in humans with high accuracy.Commun Biol.2020;3(1).

[0211] Yokoi A,et al.Integrated extracellular microRNA profiling for ovariancancer screening.Nat Commun.2018;9(1).

[0212] Yamamoto Y,et al.Highly Sensitive Circulating MicroRNA Panel forAccurate Detection of Hepatocellular Carcinoma in Patients With LiverDisease.Hepatol Commun.2020;4(2):284-297.

[0213] Usuba W,et al.Circulating miRNA panels for specific and earlydetection inbladder cancer.Cancer Sci.2019;110(1).

[0214] Sudo K,et al.Development and Validation of an Esophageal SquamousCellCarcinoma Detection Model by Large-Scale MicroRNA Profiling.JAMANetwOpen.2019;2(5):e194573.

[0215] Abe S,et al.A novel combination of serum microRNAs for the detectionofearly gastric cancer.Gastric Cancer.2021;24(4):835-843.

[0216] Ohno M,et al.Assessment of the Diagnostic Utility of SerumMicroRNAClassification in Patients With Diffuse Glioma.JAMA Netw Open.2019;2(12):e1916953.

[0217] Urabe F,et al.Large-scale Circulating microRNA Profiling for theLiquidBiopsy of Prostate Cancer.Clinical Cancer Research.2019;25(10):3016-3025.

[0218] Siegel RL,et al.Cancer statistics,2022.CA Cancer J Clin.2022;72(1):7-33。

Claims

1. A method for developing a diagnostic model for detecting a target cancer, comprising the following steps: (1) Construct a training set comprising expression profiles of multiple miRNAs obtained from non-cancer subjects and cancer patients; wherein at least two cancer types are present in the cancer patients; and (2) Developing the diagnostic model based on the training set includes the following sub-steps: (a) Differential expression analysis of the expression profiles of the various miRNAs was performed using a first statistical model, allowing the various miRNAs to be ranked by an adjusted p-value; and (b) Establish a diagnostic model based on a set of miRNA biomarkers selected from the plurality of miRNAs, wherein the set of selected miRNA biomarkers includes at least one miRNA, and the order of each miRNA is no higher than a preset cutoff value m, where m is an integer greater than zero.

2. The method according to claim 1, wherein, Sub-step (b) in step (2) includes: Based on the expression profiles of the selected miRNA biomarker set, a diagnostic index is calculated, wherein the diagnostic index is calculated based on the following formula: Where n is the total number of at least one miRNA in the selected miRNA biomarker set; n is an integer not greater than m, and miRNA i The expression level of the i-th miRNA in the selected miRNA biomarker set; i is an integer greater than zero and less than or equal to n; t i denoted as the weight of the i-th miRNA.

3. The method according to claim 2, wherein, t i It is a constant.

4. The method according to claim 2, wherein, t i Based on the second statistical model, which is selected from one of the following models: the microarray data linear model (limma model), the logistic regression model, the linear discriminant analysis (LDA) model, the conditional logistic regression model, the lasso regression model, the ridge regression model, the random forest, the support vector machine, or the probabilistic unit regression model.

5. The method according to claim 4, wherein, The first statistical model and the second statistical model are essentially the same.

6. The method according to claim 5, wherein, The limma model is used as both the first and second statistical models.

7. The method according to any one of claims 2-6, wherein, At least one miRNA in the selected miRNA biomarker set is one of the top n miRNAs in the sorting.

8. The method according to any one of claims 1-7, wherein, Step (2) further includes the following sub-steps: (c) Evaluate the performance of the diagnostic model using cross-validation on the training set.

9. The method according to claim 8, wherein, The cross-validation is at least 2-fold.

10. The method according to claim 8 or claim 9, wherein, Sub-step (c) of evaluating the performance of the diagnostic model through cross-validation on the training set includes at least one of the following: Calculate the area under the receiver operating characteristic (ROC) curve (AUC) and evaluate the performance of the diagnostic model based on it; or The specificity and sensitivity of the diagnostic model are calculated, and the performance of the diagnostic model is evaluated based on these.

11. The method according to any one of claims 1-10, further comprising: (3) Validate the diagnostic model based on a validation set, wherein the validation set comprises expression profiles of the selected set of miRNA biomarkers obtained from non-cancer subjects and cancer patients with the target cancer.

12. The method according to claim 11, wherein, Step (3) includes at least one of the following: Calculate the AUC of the ROC curve and evaluate the performance of the diagnostic model based on it; or The specificity and sensitivity of the diagnostic model are calculated, and the performance of the diagnostic model is evaluated based on these.

13. A system for developing a diagnostic model for detecting a target cancer, comprising: processor; and A non-transient storage medium containing program instructions executable by the processor, wherein the program instructions cause the processor to perform the steps of the method according to any one of claims 1-12.

14. A non-transient storage medium storing computer-executable program instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1-12.

15. A method for detecting a target cancer in a subject using a diagnostic model, wherein, The diagnostic model is developed using the method described in any one of claims 1-12.

16. The method of claim 15, comprising: Determine the expression profile of the selected set of miRNA biomarkers from biological samples obtained from the subject; Based on the expression profiles of the selected miRNA biomarker set, a diagnostic index for the biological sample is calculated, wherein the diagnostic index is calculated based on the following formula: Where n is the total number of at least one miRNA in the selected miRNA biomarker set; miRNA i The expression level of the i-th miRNA in the selected miRNA biomarker set; i is an integer greater than zero and less than or equal to n; t i The weights of the i-th miRNA; and Based on the calculated diagnostic index, the subject is classified as having the target cancer or not having the target cancer. If the calculated diagnostic index is greater than or equal to a predetermined threshold, the subject is classified as having the target cancer; otherwise, the subject is classified as not having the target cancer.

17. The method according to claim 16, wherein, Each miRNA in the selected miRNA biomarker set is from the first 200 miRNAs listed in Table 1, and n is less than or equal to 200.

18. The method according to claim 16 or claim 17, wherein, The diagnostic index is calculated using a weighted model derived from the weights of the limma model.

19. The method according to claim 18, wherein, At least one miRNA in the selected miRNA biomarker set is one of the top n miRNAs in the sorting.

20. The method according to claim 19, wherein, n is not less than 4 and not greater than 200, wherein the classification can achieve an AUC greater than about 0.97 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, bile duct cancer, bladder cancer, glioma, or prostate cancer.

21. The method according to claim 19, wherein, n is not less than 4 and not greater than 200, wherein the classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, bile duct cancer, bladder cancer, glioma, or prostate cancer, can achieve a sensitivity of at least about 0.75 while maintaining a specificity of not less than about 0.

99.

22. The method according to claim 21, wherein, The classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.80 while maintaining a specificity of not less than approximately 0.

99.

23. The method according to claim 22, wherein, The classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.90 while maintaining a specificity of not less than approximately 0.

99.

24. The method according to claim 23, wherein, The classification, when used to detect lung cancer, esophageal cancer, gastric cancer, bile duct cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.95 while maintaining a specificity of not less than approximately 0.

99.

25. The method according to claim 24, wherein, The classification, when used to detect lung cancer, esophageal cancer, gastric cancer, bile duct cancer, bladder cancer, or prostate cancer, achieves a sensitivity of at least approximately 0.99 while maintaining a specificity of not less than approximately 0.

99.

26. The method according to any one of claims 16-19, wherein, The selected set of miRNA biomarkers consists of hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, and hsa-miR-663a.

27. The method according to claim 26, wherein, The classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, sarcoma, bile duct cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.75 while maintaining a specificity of not less than approximately 0.

99.

28. The method according to claim 27, wherein, The classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.80 while maintaining a specificity of approximately 0.

99.

29. The method according to claim 28, wherein, The classification, when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.90 while maintaining a specificity of approximately 0.

99.

30. The method according to claim 29, wherein, The classification, when used to detect lung cancer, gastric cancer, bile duct cancer, bladder cancer, glioma, or prostate cancer, achieves a sensitivity of at least approximately 0.95 while maintaining a specificity of approximately 0.

99.

31. The method according to claim 30, wherein, The classification, when used to detect lung cancer, gastric cancer, bile duct cancer, or bladder cancer, achieves a sensitivity of at least approximately 0.99 while maintaining a specificity of approximately 0.

99.

32. The method according to any one of claims 16-19, wherein, The selected set of miRNA biomarkers consists of hsa-miR-5100, hsa-miR-1228-5p, hsa-miR-8073, hsa-miR-663a, and hsa-miR-320a; wherein the classification achieves a sensitivity of at least approximately 0.80 while maintaining a specificity of approximately 0.99 when used to detect lung cancer, colorectal cancer, esophageal cancer, gastric cancer, liver cancer, ovarian cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.

33. The method according to any one of claims 16-19, wherein, The selected set of miRNA biomarkers consists of the first 10 miRNAs from Table 1; wherein the classification achieves a sensitivity of at least approximately 0.99 while maintaining a specificity of approximately 0.99 when used to detect gastric cancer, esophageal cancer, biliary tract cancer, or prostate cancer.

34. The method according to any one of claims 16-19, wherein, The selected set of miRNA biomarkers consists of the first 15 miRNAs from Table 1; wherein the classification achieves a sensitivity of at least approximately 0.90 while maintaining a specificity of approximately 0.99 when used to detect lung cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, biliary tract cancer, bladder cancer, glioma, or prostate cancer.

35. The method according to any one of claims 16-34, wherein, The expression profiles of the selected miRNA biomarker set were obtained by at least one of the following methods: RNA blotting, microarray analysis, RNA sequencing, or RNA in situ hybridization, or nucleic acid amplification program, wherein the nucleic acid amplification program includes at least one of reverse transcription PCR (RT-PCR), quantitative RT-PCR (qRT-PCR), or digital RT-PCR.

36. The method according to any one of claims 16-34, wherein, The biological sample is a liquid biopsy sample selected from the group consisting of blood samples, serum samples, plasma samples, urine samples, saliva samples, and sputum samples.

37. A system for detecting a target cancer from a subject, comprising: processor; and A non-transient storage medium containing program instructions executable by the processor, wherein the program instructions cause the processor to perform the steps of the method according to any one of claims 16-36.

38. A non-transient storage medium storing computer-executable program instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 16-36.

Citation Information

Patent Citations

  • Stable nanoreporters

    US20100047924A1

  • Methods and computer systems for identifying target-specific sequences for use in nanoreporters

    US20100112710A1

  • Cancer detection method, kit, and system

    WO2022261039A2