Cancer diagnostic marker and reference marker comprising exon-junction
Cancer diagnostic markers using exon-junctions in RNA enhance sensitivity and accuracy by normalizing gene expression levels, addressing the limitations of current methods in detecting early-stage cancer.
Patent Information
- Application Number
- PCT/KR2025/006080
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-02
- Filing Date
- 2025-05-07
- Publication Date
- 2025-11-13
AI Technical Summary
Current cancer diagnostic methods lack accuracy in detecting early-stage cancer due to interference from genomic DNA and low-abundance transcripts, and there is a need for more sensitive markers to reflect subtle molecular changes.
Development of cancer diagnostic markers and reference markers utilizing exon-junctions in RNA, which include specific base sequences at exon-junction sites for normalizing gene expression levels and enhancing sensitivity in blood RNA profiling.
The proposed markers improve the accuracy of cancer detection by reducing interference from genomic DNA and enabling better quantification of low-abundance transcripts, making them effective for early-stage cancer diagnosis.
Smart Images

Figure KR2025006080_13112025_PF_FP_ABST
Abstract
Description
Cancer diagnostic markers and reference markers including exon-junctions
[0001] This application claims the benefit of Korean Patent Application No. 10-2024-0059941, filed May 7, 2024, and Korean Patent Application No. 10-2025-0042849, filed April 2, 2025, the entire contents of which are incorporated herein by reference.
[0002] The present invention relates to a cancer diagnostic marker and reference marker including an exon-junction, and more particularly, to a cancer diagnostic marker and reference marker including an exon-junction of RNA in blood.
[0003]
[0004] Cancer is causing a rise in deaths not only in Korea but also worldwide. Korea is home to a wide range of cancers, including stomach, breast, thyroid, lung, colon, and ovarian cancers. The causes of cancer can be categorized into congenital, genetic mutation, and acquired causes. Rather than being caused by mutations in specific genes, cancer is often caused by a complex interplay of multiple factors. Treatment options include surgical transplantation and removal, as well as chemotherapy and radiation therapy. While these methods have recently been shown to reduce cancer recurrence rates, ongoing research is ongoing to identify the underlying causes and predict prognosis.
[0005]
[0006] Meanwhile, platelets play a crucial role in the relationship between cancer and inflammation, emerging as valuable biomarkers. These cells engage in complex molecular interactions with tumor cells, and tumor-trained platelets undergo specific RNA splicing events in response to tumor signals. This interaction involves direct transfer of tumor-derived RNA via microvesicles, and platelets release specific factors to regulate tumor angiogenesis. These cancer-specific changes in platelet RNA profiles effectively reflect the body's systemic malignant response and provide useful diagnostic information.
[0007]
[0008] Recent advances in blood RNA analysis have shown remarkable potential for early-stage tumor detection. Previous studies have shown that blood RNA sequencing signatures have strong diagnostic potential for detecting ovarian cancer. Furthermore, blood RNA analysis offers the advantage of being a non-invasive method for diagnosing cancer.
[0009]
[0010] Therefore, there is a need to develop new diagnostic markers and / or diagnostic methods to more accurately diagnose cancer by selecting markers for cancer diagnosis from RNA isolated from blood and utilizing them.
[0011]
[0012] Accordingly, the present inventors developed a diagnostic method for accurate cancer detection utilizing blood RNA profiling, utilizing cancer diagnostic markers comprising exon junctions and reference markers for normalization. The markers of the present invention can help to better detect cancer-specific splicing events and reduce interference from contaminating genomic DNA. Furthermore, they enable more accurate quantification of low-abundance transcripts and better distinguish tumor-specific RNA signatures. Most importantly, they can enhance the sensitivity for detecting subtle molecular changes associated with early-stage disease, making them a powerful tool for early-stage cancer diagnosis.
[0013]
[0014] Accordingly, an object of the present invention is to provide a reference marker or a combination thereof for normalizing gene expression levels, which includes a base at position 1 and a base at position 2 in each of at least one selected from exon-junction numbers 1 to 6 in Table 1 below, and includes two or more bases consecutive in the 5' direction and / or the 3' direction of each chromosome.
[0015] Exon-Splice NumberGeneChromosomeStrandPosition1Position21SH3KBP1X-19595001196079372FYB15-39134350391348543PTP4A21-31911827319158944ITM2B13+48258237482587965ACTB7-552872055291606ACTB7-55278925528003
[0016]
[0017] Another object of the present invention is to provide a method for evaluating the expression level of a target gene in a sample, comprising a step of normalizing the expression level of the target gene in the sample with the reference marker or a combination thereof.
[0018]
[0019] Another object of the present invention is to provide a cancer diagnostic marker or a combination thereof, which comprises at least two bases consecutive in the 5' direction and / or the 3' direction of each chromosome, while including the bases at position 1 and position 2 in each of at least one selected from exon-junction numbers 7 to 28 of Table 2 below:
[0020] Exon-Splice NumberGeneChromosomeStrandPosition1Position27MAX14-65078037650937078MAX14-65076664650779129ACTN114-688772416887898810DAPP14+998661229986811611DAPP14+9986385699 86603312KIF2A5+623472256234804713TSPAN337+12916290812916447314TSPAN337+129 16756112916777215TSPAN337+12916690712916739816MTPN7-13595068313595151617MT PN7-13595163113597702818PTGS19+12237857412237877419IL1R22+1020262541020282 2520IL1R22+10200864310200956121DEFA18-6978647698001222DEFA38-7016100701667 523CD17719+433539944335420624FMO21+17120779117120879325NPRL316-11925613814 926TCN222+306230843062645927DEFA1B8-6997764699823328GCKR2+2749740027497561
[0021] Another object of the present invention is to provide a method for providing information for cancer diagnosis, comprising a step of measuring the expression level of the cancer diagnosis marker or a combination thereof in a sample obtained from an individual.
[0022]
[0023] Another object of the present invention is to provide a composition for diagnosing cancer, including a preparation capable of measuring the expression level of the marker or a combination thereof, and a cancer diagnosis kit including the same.
[0024]
[0025] Another object of the present invention is to provide a use of a preparation capable of measuring the expression level of the marker or a combination thereof for preparing a composition for diagnosing cancer.
[0026]
[0027] Another object of the present invention is to provide a method for diagnosing cancer comprising the following steps:
[0028] a) Step of extracting a sample;
[0029] b) a step of measuring the expression level of the cancer diagnostic marker or a combination thereof from the sample; and
[0030] c) A step of inputting the expression level of the cancer diagnosis marker or a combination thereof into an artificial intelligence cancer discrimination model pre-trained to diagnose cancer, comparing the output result value with a cut-off value, and diagnosing the presence or absence of cancer.
[0031]
[0032] In order to achieve the above-described object of the present invention, the present invention provides a reference marker or a combination thereof for normalizing gene expression level, which includes a base at position 1 and a base at position 2 in each of at least one selected from exon-junction numbers 1 to 6 of Table 1, and includes two or more bases consecutive in the 5' direction and / or the 3' direction of each chromosome.
[0033]
[0034] In order to achieve another object of the present invention, the present invention provides a method for evaluating the expression level of a target gene in a sample, comprising a step of normalizing the expression level of the target gene in the sample with the reference marker or a combination thereof.
[0035]
[0036] In order to achieve another object of the present invention, the present invention provides a cancer diagnostic marker or a combination thereof, which comprises two or more bases consecutive in the 5' direction and / or the 3' direction of each chromosome, while including bases at position 1 and position 2 in each of one or more selected from exon-junction numbers 7 to 28 of Table 2 below.
[0037]
[0038] In order to achieve another object of the present invention, the present invention provides a method for providing information for cancer diagnosis, comprising a step of measuring the expression level of the cancer diagnosis marker or a combination thereof in a sample obtained from an individual.
[0039]
[0040] In order to achieve another object of the present invention, the present invention provides a cancer diagnostic composition comprising a preparation capable of measuring the expression level of the marker or a combination thereof, and a cancer diagnostic kit comprising the same.
[0041]
[0042] In order to achieve another object of the present invention, the present invention provides a use of a preparation capable of measuring the expression level of the marker or a combination thereof for preparing a composition for diagnosing cancer.
[0043]
[0044] In order to achieve another object of the present invention, the present invention provides a cancer diagnosis method comprising the following steps:
[0045] a) Step of extracting a sample;
[0046] b) a step of measuring the expression level of the cancer diagnostic marker or a combination thereof from the sample; and
[0047] c) A step of inputting the expression level of the cancer diagnosis marker or a combination thereof into an artificial intelligence cancer discrimination model pre-trained to diagnose cancer, comparing the output result value with a cut-off value, and diagnosing the presence or absence of cancer.
[0048]
[0049] Hereinafter, the present invention will be described in detail.
[0050]
[0051] The present invention provides a reference marker or a combination thereof for normalizing gene expression levels, which includes a base at position 1 and a base at position 2 in each of at least one selected from exon-junction numbers 1 to 6 of Table 1, and includes two or more bases that are consecutive in the 5' direction and / or the 3' direction of each chromosome.
[0052]
[0053] In the present invention, the reference marker means a base sequence of a certain length that includes the base of an exon-junction site indicated by the gene described in Table 1 and the location information on the corresponding chromosome.
[0054]
[0055] In the above Table 1, genes and corresponding chromosomes for exon-junction numbers 1 to 6 are indicated, and the end base of the exon at the upper position where the exon-junction occurs (position 1) and the start base of the exon at the lower position (position 2) are indicated by the position numbers in the corresponding chromosomes. That is, in the present invention, the reference marker includes the junctions of positions 1 and 2 in each chromosome described in the above Table 1 (see Fig. 1).
[0056]
[0057] In one embodiment of the present invention, each of the single reference markers may be characterized by being composed of a base sequence including two or more bases that are consecutive in the 5' direction and / or the 3' direction, while including each base of position 1 and position 2 of any one of exon-junction numbers 1 to 6 in Table 1.
[0058]
[0059] In another aspect of the present invention, each of the single reference markers may be characterized by being composed of a base sequence including 2 to 300 bases consecutively in the 5' direction and / or the 3' direction, including each base at position 1 and position 2 in Table 1.
[0060]
[0061] In another embodiment of the present invention, each of the single reference markers comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, It may be characterized by being composed of a base sequence containing 200, 210, 220, 230, 240, 250, 260, 270, 280, 290 or 300 bases.
[0062]
[0063] The above reference marker is exemplarily explained through the sequences shown in Table 3 below. In Table 3 below, SEQ ID NOs 1 to 12 are defined as follows. In SEQ ID NOs 1 to 12, odd-numbered sequence numbers represent 150 base sequences in the 5' direction including position 1 of each exon-junction site specified in Table 1 above. For example, SEQ ID NO 1 represents a 150 base sequence in the 5' direction based on the base at position 1, including the base at position 1 (base 19595001 on chromosome X) in exon-junction number 1, and SEQ ID NO 11 represents a 150 base sequence in the 5' direction based on the base at position 1, including the base at position 1 of exon-junction number 6 (base 5527892 on chromosome 7). Next, even sequence numbers in SEQ ID NOs: 1 to 12 represent a sequence of 150 bases in the 3' direction including position 2 at each exon-junction site specified in Table 1 above. For example, SEQ ID NO: 2 represents a sequence of 150 bases in the 3' direction based on the base at position 2, including the base at position 2 of the exon-junction number 1 (base 19607937 on chromosome X), and SEQ ID NO: 12 represents a sequence of 150 bases in the 3' direction based on the base at position 2, including the base at position 2 of the exon-junction number 6 (base 5528003 on chromosome 7). That is, among the 150 bases included in each odd sequence number in Table 3 below, the 3'-terminal base is the base corresponding to position 1 in Table 1 above, and among the 150 bases included in each even sequence number, the 5'-terminal base is the base corresponding to position 2 in Table 1 above.
[0064]
[0065]
[0066] In one specific example of the present invention, the exon-junction reference marker may be composed of a base sequence that essentially includes the 3'-terminal base of the odd-numbered sequence number in Table 3 (i.e., the base corresponding to position 1 in Table 1) and the 5'-terminal base of the even-numbered sequence number (i.e., the base corresponding to position 2 in Table 1), and additionally includes at least one base that is continuous in the 5' direction of the odd-numbered sequence number based on position 1 and / or in the 3' direction of the even-numbered sequence number based on position 2.
[0067]
[0068] In one specific embodiment of the present invention, the exon-junction reference marker essentially includes a 3'-terminal base in an odd sequence number (i.e., a base corresponding to position 1 in Table 1) and a 5'-terminal base in an even sequence number (i.e., a base corresponding to position 2 in Table 1), and is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, It may be composed of a base sequence that additionally includes 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290 or 298 bases.
[0069]
[0070] The present invention also provides a combination of single reference markers for each exon-junction number in Table 1 above.
[0071]
[0072] The present invention also provides a method for evaluating the expression level of a target gene in a sample, comprising a step of normalizing the expression level of the target gene in the sample with the reference marker or a combination thereof.
[0073]
[0074] When analyzing gene expression levels, "normalization" refers to the process of compensating for various technical fluctuations that may occur during the experimental process, enabling fair comparison of expression levels between different samples. The reference markers of the present invention or a combination thereof can be utilized for normalizing gene expression levels.
[0075]
[0076] The above normalization can be performed by comparing the expression level of the target gene measured in each sample used in the experiment with the expression level of a reference marker. For example, even if the expression level of a specific gene is 100, this value may be exaggerated or underestimated depending on factors such as the amount of RNA extracted or the efficiency of the PCR reaction. In this case, measuring the expression level of the reference marker and setting it as a reference point allows for a more accurate interpretation of the expression level of the target gene. In other words, the step of normalizing the expression level of the target gene is essential for accurately comparing and evaluating the expression level of the target gene across each sample or experimental condition.
[0077]
[0078] For example, the normalization process in qPCR (quantitative PCR) is performed using the ΔΔCt method. The difference in the Ct values between the target gene and the reference marker within a sample is calculated to calculate ΔΔCt. The difference in ΔΔCt between the control and experimental groups is then calculated again to obtain ΔΔCt. Finally, this ΔΔCt value can be used to calculate the relative expression level in the form of 2^(-ΔΔ Ct). This calculated value allows for comparison of the degree of gene expression change between samples without technical deviation.
[0079]
[0080] In the method for evaluating the expression level of the gene of the present invention, the expression level of the gene may be analyzed by PCR (polymerase chain reaction), quantitative PCR (qPCR), RT-PCR (Reverse Transcription PCR), qRT-PCR, dPCR (digital PCR), or multiplex PCR.
[0081]
[0082] In the method for evaluating the gene expression level of the present invention, the sample may be selected from the group consisting of whole blood, plasma, serum, blood cells, anucleated cells in blood, exosomes in blood, and cfRNA (cell-free RNA) isolated from blood, but is not limited thereto.
[0083]
[0084] In the present invention, the “anucleated cell” refers to a cell without a nucleus, and includes red blood cells and platelets in blood.
[0085]
[0086] The present invention also provides a cancer diagnostic marker or a combination thereof, which comprises at least two bases consecutive in the 5' direction and / or the 3' direction of each chromosome, while including the bases at position 1 and position 2 in each of at least one selected from exon-junction numbers 7 to 28 of Table 2.
[0087]
[0088] In the present invention, the cancer diagnostic marker refers to a base sequence of a certain length that includes the bases of the exon-junction site indicated by the gene described in Table 2 and the location information on the corresponding chromosome.
[0089]
[0090] In the above Table 2, each gene and corresponding chromosome for exon-junction numbers 7 to 28 are indicated, and the end base of the exon at the upper position where the exon-junction occurs (position 1) and the start base of the exon at the lower position (position 2) are indicated by the position numbers in the corresponding chromosome. That is, the cancer diagnostic marker in the present invention includes the junctions at positions 1 and 2 in each chromosome described in the above Table 2 (see Fig. 1).
[0091]
[0092] In one embodiment of the present invention, each of the single cancer diagnostic markers may be characterized by being composed of a base sequence including two or more bases that are consecutive in the 5' direction and / or the 3' direction, while including each base of position 1 and position 2 of any one of exon-junction numbers 7 to 28 in Table 2.
[0093]
[0094] In another aspect of the present invention, each of the single cancer diagnostic markers may be characterized by being composed of a base sequence including 2 to 300 bases that are consecutive in the 5' direction and / or the 3' direction, while including each base at position 1 and position 2 in Table 2.
[0095]
[0096] In another embodiment of the present invention, each of the single cancer diagnostic markers comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, It may be characterized by being composed of a base sequence containing 210, 220, 230, 240, 250, 260, 270, 280, 290 or 300 bases.
[0097]
[0098] The above cancer diagnostic marker is exemplarily explained through the sequences shown in Table 4 below. In Table 4 below, SEQ ID NOs 13 to 56 are defined as follows. In SEQ ID NOs 13 to 56, odd-numbered sequence numbers represent 150 base sequences in the 5' direction including position 1 of each exon-junction site specified in Table 2 above. For example, SEQ ID NO: 13 represents a 150 base sequence in the 5' direction based on the base at position 1, including the base at position 1 of exon-junction number 7 (base 65078037 in chromosome 14), and SEQ ID NO: 55 represents a 150 base sequence in the 5' direction based on the base at position 1, including the base at position 1 of exon-junction number 28 (base 27497400 in chromosome 2). Next, in SEQ ID NOs: 13 to 56, even-numbered sequence numbers represent a 150-base sequence in the 3' direction including position 2 at each exon-junction site specified in Table 2 above. For example, SEQ ID NO: 14 represents a 150-base sequence in the 3' direction based on the base at position 2, including the base at position 2 of the exon-junction number 7 (base 65093707 in chromosome 14), and SEQ ID NO: 56 represents a 150-base sequence in the 3' direction based on the base at position 2, including the base at position 2 of the exon-junction number 28 (base 27497561 in chromosome 28). That is, among the 150 bases included in each odd sequence number in Table 4 below, the 3'-terminal base is the base corresponding to position 1 in Table 2 above, and among the 150 bases included in each even sequence number, the 5'-terminal base is the base corresponding to position 2 in Table 2 above.
[0099]
[0100]
[0101] In one specific example of the present invention, the cancer diagnostic marker may be composed of a base sequence that essentially includes the 3'-terminal base of the odd-numbered sequence number in Table 4 (i.e., the base corresponding to position 1 in Table 2) and the 5'-terminal base of the even-numbered sequence number (i.e., the base corresponding to position 2 in Table 2), and additionally includes at least one base that is continuous in the 5' direction of the odd-numbered sequence number based on position 1 and / or in the 3' direction of the even-numbered sequence number based on position 2.
[0102]
[0103] In one specific embodiment of the present invention, the cancer diagnostic marker essentially includes a 3'-terminal base in an odd sequence number (i.e., a base corresponding to position 1 in Table 2) and a 5'-terminal base in an even sequence number (i.e., a base corresponding to position 2 in Table 2), and is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, which are consecutive in the 5' direction of an odd sequence number based on position 1 and / or in the 3' direction of an even sequence number based on position 2. 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, It may be composed of a base sequence that additionally includes 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290 or 298 bases.
[0104]
[0105] The present invention also provides a combination of single cancer diagnostic markers of each exon-junction number in Table 2 above.
[0106]
[0107] In the present invention, the combination of cancer diagnostic markers may be interpreted to mean not only that the expression levels of two or more markers are used together as independent variables for cancer diagnosis, but also that the difference in expression levels of two or more markers, the sum, etc., are converted into a new variable and used for cancer diagnosis.
[0108]
[0109] In one embodiment of the present invention, the combination of markers may be a combination of markers including two or more consecutive bases in the 5' direction and / or the 3' direction of each chromosome, while including the bases at position 1 and position 2 in each of two or more selected from exon-junction numbers 7 to 18. In one specific example, the combination of markers may be a combination of exon-junction numbers 8 and 11, a combination of exon-junction numbers 8 and 10, a combination of exon-junction numbers 7 and 17, a combination of exon-junction numbers 17 and 18, a combination of exon-junction numbers 12 and 13, a combination of exon-junction numbers 12 and 15, a combination of exon-junction numbers 12 and 14, a combination of exon-junction numbers 15 and 16, and / or a combination of exon-junction numbers 9 and 16. In another specific example, a binary value can be calculated based on the relative magnitude of the expression levels of two markers included in each of the above marker combinations and utilized as a combination variable. For example, for a specific combination (x, y), if the expression level of x is greater than y, it is converted to 1, otherwise it is converted to 0, and the binary value obtained in this way can be utilized as a combination-specific discriminant variable.
[0110] The potential applications of marker combinations are not limited to this. The absolute value of the difference in expression levels between two markers can be calculated and compared to a predefined reference value, converting it into a quantitative or binary variable. This value can be used as input for cancer diagnostic algorithms based on combined characteristics rather than individual biomarkers. Combination markers can be utilized not only for simple expression level comparisons but also as discriminant variables for machine learning-based diagnostics.
[0111]
[0112] In one embodiment of the present invention, the combination of markers may be a combination of markers including two or more bases consecutive in the 5' direction and / or the 3' direction of each chromosome, while including bases at position 1 and position 2 in each of two or more selected from exon-junction numbers 19 to 28.
[0113]
[0114] In the present invention, the cancer may be selected from the group consisting of ovarian cancer, bladder cancer, bone cancer, blood cancer, breast cancer, melanoma, thyroid cancer, parathyroid cancer, bone marrow cancer, rectal cancer, throat cancer, laryngeal cancer, lung cancer, esophageal cancer, pancreatic cancer, colon cancer, stomach cancer, tongue cancer, skin cancer, brain tumor, uterine cancer, head or neck cancer, gallbladder cancer, oral cancer, colon cancer, anal cancer, central nervous system tumor, liver cancer, and colon cancer, but is not limited thereto.
[0115]
[0116] The present invention also provides a method for providing information for cancer diagnosis, comprising a step of measuring the expression level of the cancer diagnosis marker or a combination thereof in a sample obtained from an individual.
[0117]
[0118] In the present invention, the sample may be, for example, isolated from a known or suspected individual. The sample may be selected from the group consisting of whole blood, plasma, serum, blood cells, anucleated cells in blood, exosomes in blood, and cell-free RNA (cfRNA) isolated from blood, but is not limited thereto.
[0119]
[0120] The sample may be in its original form isolated from the subject, or may be further processed to remove or add components, such as cells, or to enrich one component compared to another. The sample may be isolated or obtained from the subject and transported to a sample analysis device. The sample may be stored and shipped at a desired temperature, such as room temperature, 4°C, -20°C, and / or -80°C.
[0121]
[0122] For example, a blood sample is collected from an individual for a liquid biopsy, and at this time, the collected blood can be used to determine whether to use it by checking the quality check (QC) indicators, thereby increasing the accuracy of the identification. Thereafter, one or more selected from the group consisting of anucleated cells such as platelets, exosomes, and cfRNA can be isolated from the collected blood sample. The separation method can be performed using a method known in the art, and preferably, the separation can be performed through centrifugation or the like. In the case of cfRNA, it can be used for cDNA synthesis directly from blood, plasma, serum, or fractions thereof.
[0123]
[0124] In the present invention, the subject may be a human, mammal, animal, pet, service animal, or companion animal. The subject may have cancer. The subject may have been treated with one or more therapies, such as surgery, treatment, medication, chemotherapy, antibodies, vaccines, or biological agents. The subject may or may not be in remission.
[0125]
[0126] In the present invention, the term "anucleate cell" refers to a cell without a nucleus and unable to produce daughter cells through cell division. The anucleate cell includes platelets, red blood cells, and any cell that lacks a nucleus due to incomplete cell division. Preferably, the anucleate cell is a platelet or red blood cell, and most preferably, a platelet.
[0127]
[0128] In the present invention, the 'exosome' refers to an extracellular vesicle having a vesicle structure with a nanometer size (e.g., 50-90 nm), and has a structure in which the inside and the outside of the exosome are separated by a lipid bilayer composed of cell membrane components of the cell from which it is derived, and contains cell membrane lipids, cell membrane proteins, nucleic acids, and cell components of the cell. The origin of the exosome in the present invention is not particularly limited, but may preferably be isolated from blood. Exosomes mediate the transport of mRNA, miRNA, DNA, and proteins between cells and play an important role in signal transmission and interaction inside and outside the cell. Exosomes can be isolated using any method known in the art without limitation, and for example, exosomes can be isolated using ultra-centrifugation isolation, size exclusion, immunoaffinity isolation, microfluidics chip technology, and polymeric method. Additionally, exosomes can be isolated using a commercially available exosome isolation kit (e.g., Exo2DTM EV isolation kit).
[0129]
[0130] The method of the present invention may include a step of isolating RNA from a sample. Isolation of RNA from a sample may be accomplished using various methods known in the art. For example, RNA isolation methods include, but are not limited to, the guanidine thiocyanate-cesium chloride ultracentrifugation method, the guanidine thiocyanate-hot phenol method, the guanidine hydrochloride method, and the acidic guanidine thiocyanate-phenol-chloroform method. In addition, commercially available RNA extraction reagents (e.g., RNA queous kit (Ambion Inc., Austin, TX), Micro-to-midi total RNA purification system (Invitrogen), NucleoSpin RNA II (BD Biosciences Clontech, Palo Alto, CA), RNeasy mini kit (Qiagen), GenElute mammalian total RNA kit (Sigma-Aldrich, and Trizol LS reagent (Invitrogen)) can be used according to the attached protocol. The RNA is isolated according to a method known in the art. The isolated RNA fraction can be further purified into only mRNA and used, if necessary. The purification method is not particularly limited as long as it is a known RNA purification method, but for example, mRNA can be purified by adsorbing mRNA to a biotinylated oligo (dT) probe, capturing mRNA using the binding of biotin / streptavidin to paramagnetic particles on which streptavidin is immobilized, washing the mRNA, and then eluting the mRNA. In addition, oligo (dT) A method of adsorbing mRNA onto a cellulose column and then eluting and purifying it may also be employed. However, for the method of the present invention, the mRNA purification process is not mandatory and may be performed optionally.
[0131]
[0132] A step of synthesizing complementary DNA (cDNA) for the isolated RNA may then be performed. Methods for synthesizing cDNA from RNA can be performed without limitation using methods known in the art. For example, reverse transcriptase and deoxyribonucleotides are added to RNA to copy the first DNA strand using the mRNA chain as a template. Then, mRNA is removed from the DNA-RNA hybrid double strands by treatment with an RNase (RNase H). Then, cDNA can be synthesized by treating the DNA strand produced by reverse transcription with a DNA polymerase to form a second DNA strand using the template, thereby completing the template.
[0133]
[0134] Thereafter, the expression level of the cancer diagnostic marker or a combination thereof may be measured using at least one method selected from the group consisting of reverse transcription polymerase chain reaction (RT-PCR), competitive RT-PCR, real-time RT-PCR, quantitative or semi-quantitative RT-PCR, quantitative or semi-quantitative real-time RT-PCR, in situ hybridization, fluorescence in situ hybridization (FISH), RNase protection assay (RPA), northern blotting, southern blotting, RNA sequencing, DNA chip, and RNA chip, preferably reverse transcription polymerase chain reaction (RT-PCR), competitive RT-PCR, real-time RT-PCR, quantitative or semi-quantitative RT-PCR, quantitative or semi-quantitative real-time RT-PCR, in situ hybridization, fluorescence in situ hybridization (FISH), RNase protection assay (RPA), northern blotting, southern blotting, RNA sequencing, DNA chip, and RNA chip. It can be measured using one or more methods selected from the group consisting of competitive RT-PCR, real-time RT-PCR, quantitative or semi-quantitative RT-PCR, quantitative or semi-quantitative real-time RT-PCR, and RNA sequencing.
[0135]
[0136] In one embodiment of the present invention, when the expression level of the marker is measured by reverse transcription polymerase chain reaction (RT-PCR), competitive RT-PCR, real-time RT-PCR, quantitative or semi-quantitative RT-PCR, or quantitative or semi-quantitative real-time RT-PCR, a process of normalizing the expression level with a reference marker according to any one or more of the exon-junction numbers 1 to 6 described above may be additionally performed.
[0137]
[0138] Next, whether or not a patient has cancer is determined based on the expression level of each marker above.
[0139]
[0140] In one embodiment of the present invention, the presence or absence of cancer can be determined by comparing the expression level of a marker measured in the specimen with a previously secured database of expression levels of each marker. For example, if the expression level of a specific marker identified as being upregulated in cancer patients in the previously secured database is increased in the individual compared to a normal control group, the individual can be determined to have cancer. This determination can be made using the expression levels of one or more markers.
[0141]
[0142] Preferably, the determination of whether the subject has cancer can be made by applying the expression level of a marker measured from a subject sample to a pre-learned cancer determination model.
[0143]
[0144] In the present invention, the determination of whether a subject has cancer may be determined by determining whether one or more types of cancer are present. Preferably, the determination of whether two or more types of cancer are present may be performed simultaneously or sequentially using information obtained from a single sample isolated from the subject.
[0145]
[0146] In one embodiment of the present invention, the discriminant model is trained using public data (e.g., GSE68086), and a model verified using this can be used. Typically, the entire set is divided into a training set and a validation set in a ratio of 6:4, and the cancer detection model is trained using the training set for the expression levels of each acquired marker, and the performance is verified using the validation set before use.
[0147]
[0148] In one embodiment of the present invention, by acquiring biomarker characteristics measured from the specimen and inputting them into a discriminant model, it is possible to determine whether a subject's sample is cancerous or normal. Furthermore, the discriminant model can output a discriminant score for cancer or normal as an output value. By comparing the output value with a predetermined cut-off value, it is possible to determine whether the subject has cancer.
[0149]
[0150] In the present invention, the machine learning algorithm that can be used when learning the cancer discrimination model should be interpreted as including all machine learning methods or types that are or will be disclosed in the art. For example, the machine learning algorithm may include (1) supervised learning, (2) unsupervised learning, (3) reinforcement learning, (4) semi-supervised learning, (5) neural networks, etc., and more specifically, may include, but is not limited to, Naive Bayes Classification, Logistic Regression, Decision tree, Random forest, Boosting (XGBoost / ensemble boosting / AdaBoost / Gradient Boost / LightGBM / CatBoost, etc.), Perceptron, Support Vector Machine, Quadratic classifiers, Clustering (K-means clustering, Bayesian network clustering, etc.), Deep Neural Network, etc.
[0151]
[0152] The present invention also provides a composition for diagnosing cancer, comprising a preparation capable of measuring the expression level of a marker or a combination thereof.
[0153]
[0154] In the present invention, the agent capable of detecting the marker may be a primer pair and / or a probe capable of amplifying the marker, preferably a primer pair or a probe that specifically binds to a sequence that includes two or more consecutive bases in the 5' direction and / or the 3' direction while including each base at position 1 and position 2 in each exon-junction number of Table 2.
[0155]
[0156] In the present invention, the “primer” refers to a short nucleic acid sequence having a short free 3’ hydroxyl group, which can form base pairs with a complementary template and functions as a starting point for copying the template strand. The primer can initiate DNA synthesis in the presence of a reagent for polymerization (DNA polymerase or reverse transcriptase) and four different dNTPs (deoxynucleoside triphospates) in an appropriate buffer solution and temperature. The primer may incorporate additional characteristics that do not change the basic properties of the primer that function as the starting point for DNA synthesis. In the present invention, the primers including the base sequences of SEQ ID NOs: 1 to 7 are concepts that include base sequences having a sequence homology of 95% or more, respectively. In the present invention, the primers can be chemically synthesized using a phosphoramidite solid support method or other well-known methods. Such nucleic acid sequences may also be modified using many means known in the art. Non-limiting examples of such modifications include methylation, "capping," substitution of one or more natural nucleotides with homologs, and modifications between nucleotides, such as modification with uncharged linkers (e.g., methyl phosphonate, phosphotriester, phosphoroamidate, carbamate, etc.) or charged linkers (e.g., phosphorothioate, phosphorodithioate, etc.). The nucleic acids may contain one or more additional covalently linked moieties, such as proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), intercalating agents (e.g., acridine, psoralen, etc.), chelating agents (e.g., metals, radioactive metals, iron, oxidizing metals, etc.), and alkylating agents. Additionally, in the present invention, the primer nucleic acid sequence may, if necessary, include a label detectable directly or indirectly by spectroscopic, photochemical, biochemical, immunochemical or chemical means.Examples of labels include enzymes (e.g., horseradish peroxidase, alkaline phosphatase), radioisotopes (e.g., 32P), fluorescent molecules, and chemical groups (e.g., biotin).
[0157]
[0158] In the present invention, the “probe” refers to a nucleic acid fragment such as RNA or DNA, which is short, a few bases, or a long, several hundred bases, and can specifically bind to mRNA, and is labeled so as to be able to confirm the presence or absence of a specific mRNA and the expression level. The probe can be produced in the form of an oligonucleotide probe, a single-stranded DNA probe, a double-stranded DNA probe, an RNA probe, etc. The selection of an appropriate probe and hybridization conditions can be appropriately selected according to techniques known in the art.
[0159]
[0160] The present invention also provides a cancer diagnostic kit comprising the composition.
[0161]
[0162] In the present invention, the diagnostic kit can be used to detect the marker or a combination thereof according to the present invention. The kit of the present invention can include primers and probes for detecting the marker or a combination thereof, as well as one or more other component compositions, solutions, or devices suitable for the analysis method.
[0163] As a specific example, the kit of the present invention may be a kit comprising a primer set specific for mRNA and / or complementary cDNA derived from a sample to be analyzed, an appropriate amount of DNA polymerase, a dNTP mixture, a PCR buffer solution, and water for performing PCR. The PCR buffer solution may contain KCl, Tris-HCl, and MgCl2. In addition, components necessary for performing electrophoresis to confirm whether a PCR product has been amplified may be additionally included in the kit of the present invention.
[0164] As another specific example, the kit of the present invention may be a kit comprising essential elements necessary for performing a DNA chip. The DNA chip kit may include a substrate to which cDNA corresponding to a gene or a fragment thereof is attached as a probe, reagents, preparations, enzymes, etc. for producing a fluorescently labeled probe, and the like. Furthermore, the substrate may additionally include cDNA corresponding to a quantitative control gene or a fragment thereof.
[0165] Meanwhile, the kit may include a stabilizer and / or a non-reactive dye for experimental convenience, stabilization, and improved reactivity.
[0166] The above non-reactive dye material should be selected from materials that do not affect the polymerase chain reaction, and is intended to be used for analysis or identification using the polymerase chain reaction product. Materials that satisfy these conditions may be water-soluble dyes such as rhodamine, Tamra, bleach, bromophenol blue, xylene cyanol, bromocresol red, and cresol red. The non-reactive dye material may be included in an amount of 0.0001 to 0.01 wt% based on the total weight of the composition, and is preferably included in an amount of 0.001 to 0.005 wt%. If added in an amount exceeding 0.01 wt% based on the total weight of the composition, there is a problem that a high concentration of water-soluble dye may act as a reaction inhibitor during the polymerase chain reaction.
[0167] In addition, the above polyhydric alcohols can be used as stabilizing substances to further stabilize the kit components of the present invention, and one or more substances selected from the group consisting of glucose, glycerol, mannitol, galactitol, glucitol, and sorbitol can be used.
[0168] The above kit components may be provided in liquid form, and are preferably dried to enhance stability, ease of storage, and long-term storability. The drying may be performed using any known drying method, such as room temperature drying, heat drying, freeze drying, or reduced pressure drying. Any drying method may be used as long as the components of the composition are not lost.
[0169]
[0170] The present invention also provides the use of a formulation capable of measuring the expression level of the marker or a combination thereof for preparing a composition for diagnosing cancer.
[0171]
[0172] The present invention also provides a method for diagnosing cancer comprising the following steps:
[0173] a) Step of extracting a sample;
[0174] b) a step of measuring the expression level of the cancer diagnostic marker or a combination thereof from the sample; and
[0175] c) A step of inputting the expression level of the cancer diagnosis marker or a combination thereof into an artificial intelligence cancer discrimination model pre-trained to diagnose cancer, comparing the output result value with a cut-off value, and diagnosing the presence or absence of cancer.
[0176]
[0177] Utilizing the cancer diagnostic markers of the present invention, including exon junctions, and reference markers for normalizing them, can better detect cancer-specific splicing events, reduce interference from contaminating genomic DNA, and more accurately quantify low-abundance transcripts, thereby better distinguishing tumor-specific RNA signatures. Furthermore, by increasing the sensitivity for detecting subtle molecular changes associated with early-stage disease, the markers can become powerful tools for early-stage cancer diagnosis.
[0178]
[0179] Figure 1 is a diagram illustrating the definition of an exon-junction and position 1 and position 2 in the present invention.
[0180]
[0181] Figures 2a and 2b are box plots showing the distribution of RNA sequencing results (log2CPM) and qPCR analysis results (Cq values) for five types of exon-junction reference markers, respectively.
[0182]
[0183] Figure 3 is a diagram showing the results of a qPCR dilution experiment to evaluate the detection limit of a reference marker, visualizing the change in Cq value according to serial dilution.
[0184]
[0185] Figure 4 is a diagram showing qPCR Cq values measured in four different samples for five types of exon-junction reference markers.
[0186]
[0187] Figures 5a to 5j are diagrams showing the results of marker analysis for the log2CPM values of the learning dataset.
[0188]
[0189] Figures 6a to 6j are diagrams showing marker correlation plots of scaled log2CPM values and reversed ΔCq values.
[0190]
[0191] Figures 7a to 7j are diagrams showing the results of marker analysis for the ΔCq value of the learning dataset.
[0192]
[0193] Figures 8a to 8c are diagrams showing box plots for the classification scores of the learning, test, and full datasets.
[0194]
[0195] Figure 9 is a diagram showing a receiver operating characteristic (ROC) curve according to a cancer prediction model algorithm using the marker of the present invention.
[0196]
[0197] Hereinafter, the present invention will be described in detail with reference to the following examples. However, the following examples are intended only to illustrate the present invention and the present invention is not limited thereto.
[0198]
[0199] Example 1: Selection of reference markers
[0200] Reference markers were selected using publicly available data and clinical sample learning sequencing data. Sequencing data were preprocessed to log2CPM (counts per million), where the average expression value for all variables ranged from 0 to 14.76. The overall average expression range was divided into five equally spaced intervals, and a marker representing each interval was selected. Markers were selected based on a low coefficient of variation (CV) and no difference in average expression values between normal and tumor samples.
[0201]
[0202] * Reference marker selection criteria
[0203] 1. CV<0.5 for normal samples among clinical samples, CV<0.5 for tumor samples
[0204] 2. CV<0.5 for normal samples and CV<0.5 for tumor samples among public data
[0205] 3. |log2(FoldChange)|<0.1 of clinical samples (FoldChange: mean of tumor samples / mean of normal samples)
[0206] 4. |log2(FoldChange)|<0.1 of public data
[0207] 5. log2CPM>1 in more than 90% of clinical samples
[0208] 6. log2CPM>1 in more than 90% of samples in the public data
[0209] 7. Intersection of the top 10 CVs of the entire clinical sample and the top 10 CVs of the public data
[0210] 8. If the representative marker is not determined after step 7, the clinical sample with the smallest CV is selected.
[0211]
[0212] Five exon-junction reference markers were selected based on the above criteria (Table 5).
[0213]
[0214] To confirm whether the five selected reference markers had stable Cq values in PCR experiments, a qPCR run was performed with four samples in duplicate, and the templates of the five reference markers were used after being diluted 1 / 10.
[0215]
[0216] As a result, it was confirmed that all five selected reference markers were very stable with a CV of less than 3% in the PCR experiment, and that the Cq values of the markers were located in each section (Table 5, Figs. 2a and 2b).
[0217]
[0218] Number Sequencing data average expression level interval GENEChromosome / junction Mean Sequencing CV(%) PCR CV(%)R1(0~2.88]SH3KBP1X / 19595001_196079372.6019.512.40R2(2.88~5 .75]FYB15 / 39134350_391348545.414.541.33R3(5.75~8.63]PTP4A21 / 319 11827_319158947.832.541.70R4(8.63~11.51]ITM2B13 / 48258237_482587 9610.212.410.70R5(11.51~14.38]ACTB7 / 5528720_552916013.171.962.45
[0219]
[0220] Example 2: Evaluation of the limit of detection (LOD) of the reference marker
[0221] An experiment was conducted to determine the minimum concentration or amount at which the target substance can be reliably detected in a sample using the given analysis method.
[0222]
[0223] Specifically, the synthesized cDNA was used as a non-template control (NTC), the original concentration, and 10 -1 From 10-7 After serial dilution, qPCR was performed in 5 replicates. LOD was evaluated by measuring the Cq value of the marker (ACTB / 7 / 5528720_5529160) with the highest expression among the 5 reference markers selected in Example 1 (Fig. 3).
[0224]
[0225] As a result, from the original concentration to 10 -3 We confirmed the change in Cq value that continuously increased up to the dilution stage, and 10 -5 In the dilution step, no Cq value was measured in 2 out of 5 replicates.
[0226] As a result of calculating the average Cq for each serial dilution interval, 10 -3 Until then, a Cq difference of a certain interval was confirmed and judged as a quantitative confidence interval (Fig. 3).
[0227] In the qPCR results (Fig. 4) using a total of five reference markers including the marker used in LOD measurement, the Cq values of all markers were 10 -3 It was confirmed that it was a stable reference marker as it was located at a lower level than the measured Cq value in the dilution section (Figs. 3 and 4).
[0228]
[0229] Example 3: Selection and Evaluation of Cancer Diagnostic Markers (1)
[0230] The cancer detection algorithm was developed using sequencing data. When utilizing sequencing data, exon-junction markers were combined in pairs (Table 6). The algorithm was trained and validated using 12 exon-junction markers (9 combinations) using 29 clinical sample training data sets (5 normal controls (HC), 6 ovarian cancers (OC), 8 endometrial cancers (EC), and 10 benign tumors) and 17 validation data sets (4 HC, 4 OC, 4 EC, and 5 benign).
[0231]
[0232] The marker expression values were normalized using the log2CPM (counts per million) data obtained from sequencing. Normalization was based on the expression values of five reference markers derived from the present invention, and the geometric mean of these reference markers was calculated and used as the normalization standard to compensate for differences in the overall expression levels of each sample and technical variability between analyses. In other words, normalization was performed by subtracting the average value of the reference marker for the corresponding sample from the log2CPM expression value of each marker.
[0233]
[0234] NumberGene NameChromosome / junctionCombination NumberC1MAX14 / 65078037_65093707Combination 3(1)C2MAX14 / 65076664_65077912Combination 1(1), Combination 2(1)C3ACTN114 / 68877241_68878988Combination 9(1)C4DAPP14 / 99866122_99868116Combination 2(2)C5DAPP14 / 99863856_99866033Combination 1(2)C6KIF2A5 / 62347225_62348047Combination 5(1), Combination 6(1) Combination 7 (1) C7TSPAN337 / 129162908_129164473 Combination 5 (2) C8TSPAN337 / 129167561_129167772 Combination 7 (2) C9TSPAN337 / 129166907_129167398 Combination 6 (2), Combination 8 (1) C10MTPN7 / 135950683_135951516 Combination 8 (2), Combination 9 (2) C11MTPN7 / 135951631_135977028 Combination 3 (2), Combination 4 (1) C12PTGS19 / 122378574_122378774 Combination 4 (2)
[0235]
[0236] In the data normalized in this way, the relative size of the expression values (Cq values) of the first marker (1) and the second marker (2) of each combination number in Table 6 was converted to binary values (1 or 0), and an algorithm (Random Forest) was trained based on the binary data.
[0237]
[0238] The performance of the cancer presence / absence determination algorithm was verified using 37 clinical sample evaluation data (HC 9, OC 8, EC 14, Benign 6), and the sensitivity for ovarian cancer was 0.875 (7 / 8), the sensitivity for endometrial cancer was 0.429 (6 / 14), and the specificity was 0.867 (13 / 15). The specificity for healthy samples without disease or symptoms in the algorithm was 0.778 (7 / 9), and the specificity for benign tumor samples was 1.000 (6 / 6) (Tables 7 and 8).
[0239]
[0240]
[0241]
[0242] * Acc: Accuracy
[0243] * Sen: Sensitivity
[0244] * Spe: Specificity
[0245] * AUC: Area Under the Curve
[0246] * BA(balanced accuracy) = (Sen+Spe) / 2
[0247]
[0248] Example 4: Selection and Evaluation of Cancer Diagnostic Markers (2)
[0249] 1. Experimental method
[0250] (1) Platelet RNA extraction
[0251] A total of 90 women were enrolled in the study from three different medical institutions between August 2022 and January 2025. They were divided into asymptomatic control, female benign tumor, and ovarian cancer (OC) groups. The OC group was further stratified by stage (I-IV) and included borderline ovarian tumors (BOTs).
[0252]
[0253] Blood samples from these participants were collected using 10 mL EDTA-coated, purple-capped BD Vacutainers (BD) and stored at 4°C for further processing. Platelets were isolated using a two-step centrifugation process within 48 hours according to a previously established protocol to maintain sample consistency. The extracted platelets were suspended in RNAlater (Thermo Scientific, Waltham, MA, USA), stored overnight at 4°C, and then stored long-term in a -80°C deep freezer.
[0254] Total RNA was extracted within 2 months using the mirVana RNA Isolation Kit from Thermo Scientific.
[0255]
[0256] (2) RNA sequencing
[0257] The quality of total RNA was assessed using BioAnalyzer 2100 (Agilent, Santa Clara, CA, USA), and samples with RNA Integrity Number (RIN) ≥ 6 or a distinct ribosome peak were considered high-quality and used for sequencing.
[0258] For RNA sequencing, 500 pg of platelet RNA was subjected to cDNA synthesis and amplification using the SMART-Seq v4 Ultra Low Input RNA Kit (Takara Bio, Mountain View, CA, USA), and quality control was performed using BioAnalyzer 2100.
[0259] The amplified cDNA was fragmented using Covaris sonication and then labeled with an Illumina sequencing index barcode using the Truseq Nano DNA Sample Prep Kit (Illumina, San Diego, CA, USA). This was followed by 8 cycles of PCR amplification and purification using AMPure XP beads.
[0260] The concentration and fragment size distribution of the final library were evaluated using TapeStation 4200 (Agilent, Santa Clara, CA, USA), and libraries within the 500–600 bp size range were pooled and subjected to 150 bp paired-end sequencing on the Illumina NovaSeq6000 platform (Illumina, San Diego, CA, USA).
[0261]
[0262] (3) Data preprocessing and quantification
[0263] To improve the quality of RNA sequencing data and ensure accurate expression analysis, several preprocessing steps were performed. First, adapters were removed and quality-based read trimming was performed using Trimmomatic (v. 0.39). The filtered sequencing reads were then aligned to the human GRCh38 reference genome using HISAT2 (v. 2.1.0).
[0264] The generated SAM files were converted to BAM format, and only the primary alignment was maintained using Samtools (v. 1.9). Using the generated BAM files, gene-level expression values were calculated in units of Fragments Per Kilobase of transcript per Million mapped reads (FPKM) and Transcripts Per Million (TPM), and exon-junction level expression values were quantified in units of Counts Per Million (CPM).
[0265]
[0266] [Calculation criteria for exon-splice level expression values (CPM)]
[0267] · A read containing 150 bp upstream and 150 bp downstream of a specific splice junction
[0268] · If the CIGAR string contains 'N' (splicing event exists)
[0269] · Splice leads aligned exactly with the annotation location of the given joint
[0270] This process yielded FPKM and TPM values for 60,624 genes and CPM values for 2,855,955 splice junctions from a clinical RNA-seq dataset. FPKM and TPM calculations were performed using StringTie (v. 2.1.7).
[0271]
[0272] (4) Quantitative real-time polymerase chain reaction (qRT-PCR)
[0273] PrimeScript for RT-PCR analysis TMcDNA was synthesized from platelet RNA using the RT reagent kit with gDNA Eraser (Takara Bio, Mountain View, CA, USA). Genomic DNA (gDNA) was removed by treatment with gDNA Eraser at 42°C for 2 minutes, followed by reverse transcription at 37°C for 15 minutes, and enzyme inactivation at 85°C for 5 seconds.
[0274] For qPCR analysis, the synthesized cDNA was diluted 1:100 and used as a template. The total 20 μL reaction solution contained 10 μL of FastFACT 2X qPCR master mix (BIOFACT, Daejeon, South Korea), 0.6 μL each of primers and probes, and 4.2 μL of RNase- / DNase-free water.
[0275] qPCR was performed in a CFX Opus 96 Real-Time PCR System (Bio-Rad, Hercules, CA, USA), and the thermal cycling conditions were as follows.
[0276] - Initial denaturation: 2 minutes at 50℃, 10 minutes at 95℃
[0277] - Pre-cycling: 20 seconds at 95℃, 1 minute at 65℃ (repeat 5 times)
[0278] - Amplification: 95°C for 20 seconds, 60°C for 1 minute (repeated 40 times)
[0279]
[0280] (5) Exon-junction RNA sequencing analysis and marker selection
[0281] RNA splicing mutations can be aberrantly activated in cancer cells, and changes in specific splice junctions can serve as potential biomarkers for differentiating cancerous tissues. This study aimed to identify ovarian cancer-specific splicing markers using exon-junction RNA sequencing data.
[0282]
[0283] The dataset for marker selection consisted of 33 asymptomatic controls, 16 patients with benign ovarian tumors, and 13 patients with ovarian cancer. To normalize the sequencing data, the CPM values at the junction level were log-transformed (log2CPM) for analysis.
[0284]
[0285] Marker candidates were selected based on the following criteria:
[0286] - 70-95% of asymptomatic and benign tumor samples will show log2CPM values below a predefined threshold ([0.2, 0.3, 0.4, 0.5]).
[0287] - Expression of a specific junction will be significantly elevated in at least one ovarian cancer sample.
[0288]
[0289] The exon-junction reference markers used in the experiment are shown in Table 9 below:
[0290]
[0291] NumberGene NameChromosome / junctionR6ACTB7 / 5527892_5528003
[0292]
[0293] Finally, candidate cancer diagnostic markers were validated through PCR experiments, and the correlation (R² value) between sequencing and PCR data was evaluated. Markers with no expression or an R² value of 0.4 or higher in the asymptomatic control group and benign tumor group were selected, ultimately resulting in 10 valid markers (Table 10).
[0294]
[0295] NumberGene NameChromosome / junctionC13IL1R22 / 102026254_102028225C14IL1R22 / 102008643_ 102009561C15DEFA18 / 6978647_6980012C16DEFA38 / 7016100_7016675C17CD17719 / 433 53994_43354206C18FMO21 / 171207791_171208793C19NPRL316 / 119256_138149C20TCN 222 / 30623084_30626459C21DEFA1B8 / 6997764_6998233C22GCKR2 / 27497400_27497561
[0296]
[0297] Algorithm cutoff and score calculation
[0298] In this study, we developed and validated a classification algorithm to distinguish between ovarian cancer and non-cancerous cases using PCR data.
[0299]
[0300] 1. Dataset composition
[0301] · Training dataset (used for marker selection and algorithm development)
[0302] o 16 patients with benign tumors
[0303] o 13 patients with ovarian cancer
[0304] · Test dataset (for independent validation)
[0305] o 21 patients with benign tumors
[0306] o 4 patients with ovarian cancer
[0307] o Asymptomatic Control: 34 subjects
[0308] o 2 patients with borderline ovarian tumor (BOT) (for further verification)
[0309] Asymptomatic control samples were used in the marker selection process, but were not included in model development and were only used in the validation process.
[0310]
[0311] 2. qPCR data processing and cutoff value setting
[0312] · Each marker was measured twice for each sample, and the average of the duplicate measurements was used.
[0313] · If the Cq value is not detected in qPCR, specify 41.
[0314] · Set cutoff values for each marker using benign tumor samples from the training dataset to remove non-specific signals.
[0315] · Cq values exceeding the set cutoff value are replaced with 41.
[0316] · Finally, normalize the qPCR Cq values using the △Cq method (Cq_target - Cq_ACTB).
[0317]
[0318] Through these cutoff and normalization processes, we developed an algorithm that effectively distinguishes between ovarian cancer and non-cancerous cases.
[0319]
[0320] Scoring cutoffs were established for seven individual markers and five marker composite variables (sum of markers) to achieve 99% specificity using the training dataset. Cutoff values were determined based on the average △Cq values of samples that fell on the borderline between benign tumors and ovarian cancer. Scores were then calculated by evaluating the difference between the observed PCR values and the marker-specific cutoff values.
[0321]
[0322] 2. Experimental results
[0323] We analyzed the sequencing data from the training dataset to determine whether candidate markers absent from the control group (benign tumors and asymptomatic controls) exhibited outlier expression levels in some ovarian cancer samples. With the exception of IL1R2_a, DEFA1_a, and CD177_a, the selected markers did not show significant differences between the ovarian cancer and benign tumor groups. However, at least two ovarian cancer samples showed significantly higher expression levels for each marker compared to most control samples. In particular, DEFA1B_a and CD177_a were completely absent from all control samples but were exclusively expressed in a subset of ovarian cancer cases (Figures 5A-5J).
[0324]
[0325] To verify that the selected candidate markers exhibited expression patterns consistent with those identified by sequencing, we performed linear regression analysis (R² values) to assess the correlation between sequencing data (scaled log2CPM) and qPCR data (inverted ΔCq). Some of the final 10 markers were absent in the control group, and the correlation between sequencing and qPCR data ranged from 0.44 to 0.98 (Figures 6A to 6J). Although the ΔCq values did not show a statistically significant difference between the ovarian cancer and benign tumor groups, some ovarian cancer samples exhibited lower ΔCq values than benign tumors, which was consistent with the sequencing results (Figures 7A to 7J). This suggests that these markers are not a product of sequencing noise but are indeed specifically expressed in ovarian cancer.
[0326]
[0327] Next, we developed an algorithm to distinguish ovarian cancer from benign tumors and asymptomatic controls using the ΔCq values of the selected markers. In the training dataset (n=29), the algorithm achieved 92.3% sensitivity, 100.0% specificity, and an area under the curve (AUC) of 0.938 (95% CI: 0.790–0.984) (Table 11, Figures 8a and 9). Furthermore, among the 34 asymptomatic controls included in the marker selection process but not in the algorithm development, 33 were accurately classified as controls, resulting in a specificity of 97.1%.
[0328]
[0329] ActualTotalOvarian cancerBenignPredictOvarian cancer12012Non-Ovarian cancer11617Total131629
[0330]
[0331] In the test dataset (n=25), the algorithm achieved 100.0% sensitivity, 85.7% specificity, and AUC 0.917 (95% CI: 0.746-0.976), successfully identifying all ovarian cancer samples despite the small sample size, even though this dataset was not used for marker selection or algorithm development (Table 12, Fig. 8b, Fig. 9).
[0332]
[0333] ActualTotalOvarian cancerBenignPredictOvarian cancer437Non-Ovarian cancer01818Total42125
[0334]
[0335] Additionally, one of the two BOT samples was classified as ovarian cancer. The overall performance, combining the training and test datasets, achieved a sensitivity of 94.1%, a specificity of 94.4%, and an area under the curve (AUC) of 0.933 (95% CI: 0.861–0.969) (Table 13, Figures 8c and 9). By implementing predefined cutoff values and score calculation methods for each marker, high sensitivity and specificity were achieved.
[0336]
[0337] ActualTotalOvarian cancerBenignAsymptomatic controlsPredictOvarian cancer163120Non-Ovarian cancer1343368Total17373488
[0338]
[0339] Utilizing the cancer diagnostic marker of the present invention, including exon junctions, and a reference marker for normalizing the marker, can better detect cancer-specific splicing events, reduce interference from contaminating genomic DNA, and more accurately quantify low-abundance transcripts, thereby better distinguishing tumor-specific RNA signatures. Furthermore, by increasing the sensitivity for detecting subtle molecular changes associated with early-stage disease, the marker can serve as a powerful tool for early-stage cancer diagnosis, demonstrating high potential for industrial application.
Claims
1. A reference marker for normalization of gene expression level, or a combination thereof, which includes a base at position 1 and a base at position 2 in each of at least one selected from exon-junction numbers 1 to 6 in [Table 1] below, and includes two or more bases consecutively in the 5' direction and / or the 3' direction of each chromosome: [Table 1] 2. A method for evaluating the expression level of a target gene in a sample, comprising a step of normalizing the expression level of the target gene in the sample with a reference marker according to paragraph 1 or a combination thereof.
3. A method according to claim 2, characterized in that the expression level of the gene is analyzed by PCR (polymerase chain reaction), quantitative PCR (qPCR), RT-PCR (Reverse Transcription PCR), qRT-PCR, dPCR (digital PCR), or multiplex PCR.
4. A method according to claim 2, characterized in that the sample is selected from the group consisting of whole blood, plasma, serum, blood cells, anucleated cells in blood, exosomes in blood, and cfRNA (cell-free RNA) isolated from blood.
5. A cancer diagnostic marker or a combination thereof, comprising two or more consecutive bases in the 5' direction and / or 3' direction of each chromosome, including the bases at position 1 and position 2 in each of at least one selected from exon-junction numbers 7 to 28 of Table 2 below: [Table 2] 6. A cancer diagnostic marker or a combination thereof, characterized in that the combination of markers in paragraph 5 is a combination of markers including two or more bases consecutive in the 5' direction and / or the 3' direction of each chromosome, while including the bases at position 1 and position 2 in each of two or more selected from exon-junction numbers 7 to 18.
7. A cancer diagnostic marker or a combination thereof, characterized in that the combination of markers in paragraph 5 is a combination of markers including two or more bases consecutive in the 5' direction and / or the 3' direction of each chromosome, while including the bases at position 1 and position 2 in each of two or more selected from exon-junction numbers 19 to 28.
8. A cancer diagnostic marker or a combination thereof, characterized in that the cancer in paragraph 5 is selected from the group consisting of ovarian cancer, bladder cancer, bone cancer, blood cancer, breast cancer, melanoma, thyroid cancer, parathyroid cancer, bone marrow cancer, rectal cancer, throat cancer, laryngeal cancer, lung cancer, esophageal cancer, pancreatic cancer, colon cancer, stomach cancer, tongue cancer, skin cancer, brain tumor, uterine cancer, head or neck cancer, gallbladder cancer, oral cancer, colon cancer, anal cancer, central nervous system tumor, liver cancer, and colon cancer.
9. A method for providing information for cancer diagnosis, comprising a step of measuring the expression level of a cancer diagnosis marker or a combination thereof according to any one of claims 5 to 7 in a sample obtained from an individual.
10. An information providing method according to claim 9, characterized in that the sample is selected from the group consisting of whole blood, plasma, serum, blood cells, anucleated cells in blood, exosomes in blood, and cfRNA (cell-free RNA) isolated from blood.
11. An information providing method according to claim 9, wherein the expression level of the cancer diagnostic marker or a combination thereof is measured using at least one method selected from the group consisting of reverse transcription polymerase chain reaction (RT-PCR), competitive RT-PCR, real-time RT-PCR, quantitative or semi-quantitative RT-PCR, quantitative or semi-quantitative real-time RT-PCR, in situ hybridization, fluorescence in situ hybridization (FISH), RNase protection assay (RPA), northern blotting, southern blotting, RNA sequencing, DNA chip, and RNA chip.
12. An information providing method characterized in that, in paragraph 9, it further comprises a step of normalizing the expression level of the cancer diagnostic marker or a combination thereof with a reference marker according to paragraph 1.
13. An information providing method characterized in that, in paragraph 9, the method further comprises a step of inputting the expression level of the cancer diagnosis marker or a combination thereof into an artificial intelligence cancer discrimination model pre-trained to diagnose cancer and comparing the output result value with a cut-off value to determine the presence or absence of cancer.
14. An information providing method according to claim 13, wherein the pre-learning is performed by any one machine learning algorithm selected from the group consisting of Naive Bayes Classification, Logistic Regression, Decision tree, Random forest, Boosting (XGBoost / ensemble boosting / AdaBoost / Gradient Boost / LightGBM / CatBoost, etc.), Perceptron, Support Vector Machine, Quadratic classifiers, Clustering (K-means clustering, Bayesian network clustering, etc.), and Deep Neural Network.
15. A composition for diagnosing cancer, comprising a preparation capable of measuring the expression level of a marker or a combination thereof according to Article 5.
16. A composition according to claim 15, wherein the preparation is a probe or primer set that specifically binds to each of the markers or a combination thereof.
17. A cancer diagnostic kit comprising the composition of Article 15.
18. Use of a preparation capable of measuring the expression level of a marker or a combination thereof according to claim 5 for preparing a composition for diagnosing cancer. 19.a) Step of extracting a sample; b) a step of measuring the expression level of a cancer diagnostic marker or a combination thereof according to any one of claims 5 to 7 from the sample; and c) A cancer diagnosis method comprising a step of inputting the expression level of the cancer diagnosis marker or a combination thereof into an artificial intelligence cancer determination model pre-trained to diagnose cancer and comparing the output result value with a cut-off value to diagnose the presence or absence of cancer.
Citation Information
Patent Citations
Remote management system for hoist
KR1020250042489A
Predictive maintenance system for oil filter
KR102757088B1
Apparatus for treating substrate and temperature control method
KR102889024B1
Methods and systems for evaluating tumor mutational burden
WO2017151524A1