Sample identification method and biomarker
By employing specific reference sequences to calculate ratios of short-chain nucleic acids, the method stabilizes test performance and achieves accurate identification of cancer samples, addressing variability issues in existing miRNA-based cancer detection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-03-13
AI Technical Summary
Existing methods for identifying cancer using miRNA as biomarkers are prone to variability and instability in test performance due to differences in specimen and dataset results, leading to inaccurate sample identification.
A method involving the use of specific reference sequences of 6 to 9 consecutive bases to calculate ratios of first and second short-chain nucleic acids in a sample, comparing these ratios with control samples to determine the presence or risk of a target disease, utilizing nucleic acid analysis and sequencing techniques.
The method provides accurate and reproducible identification of samples derived from affected individuals by stabilizing the test performance and enhancing discriminative power, with AUC values consistently above 0.7, indicating high diagnostic accuracy.
Smart Images

Figure 2026045714000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a sample identification method and biomarkers.
Background Art
[0002] A system for identifying cancer patients and healthy individuals by analyzing nucleic acids isolated from body fluids that are easily collected is known. For example, a widely verified system is an identification system using miRNA (microRNA). miRNA is a single-stranded nucleic acid of about 17 to 25 bases, and it has been clarified that it has a function of regulating gene expression, and it has been reported that its type and expression level change from the initial stage in various diseases. For example, in cancer patients, various miRNA amounts are used as cancer markers and are known to increase or decrease compared to healthy individuals. These findings propose quantitatively examining the target miRNA contained in a sample collected from a subject as a means for knowing whether the subject has cancer or not.
[0003] Tests using miRNA amount or miRNA concentration as an index are likely to show differences in the results and trends between specimens and datasets. As a result, the variation between data and results becomes large, and the test performance tends to become unstable.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The problem to be solved by the present invention is to provide a method for accurately identifying a sample derived from an affected patient.
Means for Solving the Problems
[0005] The method according to the embodiment is a method for identifying whether or not a sample originates from an infected subject. The method includes: setting one or more sets of a first reference sequence and a second reference sequence as a combination of two sequences whose relationship to the target disease is suggested based on predetermined criteria; obtaining the number of first short chain nucleic acids (ID-1) and the number of second short chain nucleic acids (ID-2) present in the short chain nucleic acid group contained in the sample derived from the subject and the sample derived from the control sample, respectively; calculating the ratio R1 = (number of ID-1 / number of ID-2) for the sample derived from the subject and the ratio R2 = (number of ID-1 / number of ID-2) for the sample derived from the control sample; and determining whether or not the sample derived from the subject originates from an subject that has or is at risk of developing the target disease by comparing the values of ratio R1 and ratio R2. The first reference sequence consists of a sequence of 6 to 9 consecutive bases. The second reference sequence consists of a sequence of 6 to 9 consecutive bases and is different from the base sequence of the first reference sequence. The first short nucleic acid contains a sequence at one of its positions that is in exact agreement with the first reference sequence. The second short nucleic acid contains a sequence at one of its positions that is in exact agreement with the second reference sequence. [Brief explanation of the drawing]
[0006] [Figure 1] A scheme diagram showing the first embodiment. [Figure 2] An example of a system for obtaining the number of first and second short-chain nucleic acids. [Figure 3] A figure showing the results of one example of the first embodiment. [Figure 4] A scheme diagram showing the second embodiment. [Figure 5] A box plot showing the results of Example 1. [Figure 6] A box plot showing the results of Example 1. [Figure 7] A box plot showing the results of Example 1. [Figure 8] A box plot showing the results of Example 1. [Figure 9] A box plot showing the results of Example 3. [Figure 10]A box plot showing the results of Example 4. [Figure 11] A box plot showing the results of Example 4. [Figure 12] A box plot showing the results of Example 4. [Figure 13] A box plot showing the results of Example 4. [Figure 14] A box plot showing the results of Example 4. [Modes for carrying out the invention]
[0007] The embodiments will be described below with reference to the attached drawings. In each embodiment, substantially identical components will be denoted by the same reference numerals, and their descriptions may be partially omitted. The drawings are schematic, and the relationship between the thickness of each part and its planar dimensions, the ratio of the thicknesses of each part, etc., may differ from those in reality.
[0008] (First embodiment) The first embodiment is a method for identifying whether or not a sample originates from a person suffering from a disease.
[0009] "Identification" of a sample can mean determining whether the source of the sample being tested is suffering from a disease, whether it is at risk of developing a disease, or whether it is at high risk of developing a disease.
[0010] The identification method according to this embodiment will be explained with reference to Figure 1. The sample identification method includes the following: setting a reference sequence set including a first reference sequence and a second reference sequence (S11): S12: Obtain the number of first short-chain nucleic acids (ID-1) and the number of second short-chain nucleic acids (ID-2) present in the short-chain nucleic acid group contained in the sample derived from the target and the sample derived from the control sample. Calculate the ratio R1 = (number of ID-1 / number of ID-2) for the sample derived from the target, and the ratio R2 = (number of ID-1 / number of ID-2) for the sample derived from the control sample. S13: By comparing the values of ratio R1 and ratio R2, determine whether the sample derived from the target originates from a subject that has or is at risk of developing the target disease.
[0011] First, in S11, a reference sequence set is established. The reference sequence set includes a first reference sequence and a second reference sequence. The first and second reference sequences are sequences that constitute a combination of two types of sequences whose relationship to the target disease is suggested based on predetermined criteria.
[0012] These reference sequences are sequences of 6 to 9 consecutive bases, for example, a sequence of 7 consecutive bases. The first and second reference sequences may be of the same length or different lengths. However, from the viewpoint of improving the efficiency of the sequence decoding process and equalizing the influence of the quality of sequence decoding results between reference sequence sets, it is preferable that the first and second reference sequences are of the same length. There may be one reference sequence set or multiple reference sequence sets.
[0013] The first and second reference sequences have different sequences. A sequence is considered different if at least one of the bases at the same position from the 5' end is different. For example, if the first and second reference sequences have the same base length, the sequence of bases from the 5' end to the 3' end does not need to be exactly the same. Alternatively, if the first and second reference sequences have different base lengths, the sequence of bases from the 5' end to the 3' end of the shorter reference sequence does not need to be exactly the same.
[0014] Here, the "disease" can be any disease to be detected, and can be any disease for which a certain relationship with the reference sequence is predicted or a tendency to be related is observed. For example, such a disease is cancer. Here, cancer includes all stages, for example, the state in which an event or tendency related to cancer has occurred in the target cell or tissue, the state in which cancer remains in the organ of origin, the state in which cancer has spread to surrounding tissues, the state in which cancer has metastasized to lymph nodes, and the state in which cancer has metastasized to distant organs, etc. For example, cancer can be at least one cancer selected from the group consisting of breast cancer, colorectal cancer, lung cancer, gastric cancer, pancreatic cancer, cervical cancer, uterine cancer, ovarian cancer, sarcoma, prostate cancer, bile duct cancer, bladder cancer, esophageal cancer, liver cancer, brain tumor, and kidney cancer.
[0015] The "predetermined determination criterion" can be, for example, AUC. For example, a specific AUC value, for example, an AUC value of 0.7 or more can be used as the determination criterion. Here, AUC is an index used to represent the accuracy of a test. For example, AUC can be obtained as follows using a graph. That is, with sensitivity on the vertical axis and (1 - specificity) on the horizontal axis ranging from 0 to 1, it is shown as the area under the ROC curve when a ROC curve is created by continuously changing the threshold. Here, the "specificity" is a value from 0 to 1 representing the ratio of test values below the threshold among the test values of the healthy group when it is determined to be positive when the test value exceeds a certain threshold. The closer the AUC value is to 1, that is, the closer the curve is to the upper left corner, the higher the diagnostic ability of the diagnostic method is shown, and the closer the AUC is to 0.5, that is, the closer the curve is to the diagonal line, the lower the diagnostic ability of the diagnostic method is shown.
[0016] Also here, the "healthy group" can be data obtained from healthy individuals. The "healthy individual" refers to a person who does not suffer from the disease to be identified, for example, a person who does not suffer from cancer, and for example, a person who does not suffer from a malignant neoplasm including cancer or malignant tumor, that is, a "non-cancer person" may be sufficient. The "predetermined determination criterion" can be derived using a sample whose origin, for example, whether it is from a non-cancer person or a cancer patient, is clear.
[0017] Next, in S12, the number of the first short-chain nucleic acids (the number of ID-1) and the number of the second short-chain nucleic acids (the number of ID-2) present in the short-chain nucleic acid groups contained in the sample derived from the subject and the sample derived from the control specimen are obtained.
[0018] The first short-chain nucleic acid (ID-1) is a short-chain nucleic acid containing a sequence that exactly matches the first reference sequence at any of its positions. Similarly, the second short-chain nucleic acid (ID-2) is also a short-chain nucleic acid containing a sequence that exactly matches the second reference sequence at any of its positions. Note that the length of the short-chain nucleic acid detected in the sample is not particularly limited, and it may have a base length longer than the reference sequence contained in the short-chain nucleic acid.
[0019] The method of the embodiment analyzes the short-chain nucleic acid group contained in the sample as described above. The "short-chain nucleic acid group" can be, for example, RNA such as mRNA, ncRNA (non-coding RNA), housekeeping ncRNA, tRNA, small ncRNA, miRNA, piRNA, tsRNA, IncRNA, etc. For example, ncRNA, housekeeping ncRNA, tRNA, small ncRNA, miRNA, piRNA, tsRNA, IncRNA, etc. are preferred. One preferred example is miRNA.
[0020] A "sample" may be cells, tissues and / or fluids collected from a subject, or mixtures thereof, or processed products obtained by appropriately processing them. A "sample" may be body fluids removed from the subject. If the sample is a body fluid, it may be, for example, serum or plasma, or other body fluids such as blood, interstitial fluid, urine, feces, sweat, saliva, oral mucosa, nasal mucosa, nasal mucus, pharyngeal mucosa, sputum, digestive fluid, gastric juice, lymph, cerebrospinal fluid, tears, breast milk, amniotic fluid, semen, vaginal fluid, or mixtures thereof. For example, the sample may be immediately after collection from the subject, cultured, preserved using a desired procedure, or the supernatant obtained after maintaining them in a desired liquid. For example, body fluids such as blood, serum, and plasma are preferred as samples because they are easy to collect.
[0021] The "subject" can be any animal or plant from which a sample should be taken. For example, the subject could be any mammal, including humans, or any mammal other than humans. For example, the subject could be any animal belonging to a group such as humans, primates such as monkeys, rodents such as mice, rats, or guinea pigs, companion animals such as dogs, cats, or rabbits, domesticated animals such as horses, cows, or pigs, or animals used for exhibitions.
[0022] A "control sample" is a sample whose origin and the state of the subject are known. In this embodiment, the term "control sample" includes not only a single sample but also multiple samples. Furthermore, "multiple" simply indicates the number of samples. In other words, a control sample may include multiple samples of the same type, or two or more different types of samples.
[0023] For example, the control sample is at least one selected from a group consisting of non-cancer samples, pancreatic cancer samples, breast cancer samples, and colorectal cancer samples. In this case, the control sample may include only the non-cancer sample group or one of the cancer sample groups, or it may include both the non-cancer sample group and the cancer sample group. In one example, the control sample includes both the non-cancer sample group and the colorectal cancer sample group.
[0024] A control sample is, for example, a healthy individual. A healthy individual is defined as an individual that is not suffering from the target disease. It is preferable that the healthy individual is a healthy individual without any disease or abnormality. However, as mentioned above, the control sample is not limited to healthy individuals as long as the condition of the subject from which it originates is known. The individual selected as the control sample may be a different individual from the subject analyzed by this method, but it is preferable that it be an individual belonging to the same species, i.e., a human if the subject is human. Furthermore, there are no particular limitations on the physical conditions of the control sample, such as age, sex, height and weight, or the number of control samples, but the physical conditions may be the same as or similar to those of the subject being tested by this analysis method, or they may be different.
[0025] Here, "obtaining the number of first short-chain nucleic acids (number of ID-1) and the number of second short-chain nucleic acids (number of ID-2)" can be in any form, as long as it enables the calculation of the ratios R1 = (number of ID-1 / number of ID-2) and R2 = (number of ID-1 / number of ID-2) in the following (S13). For example, the person performing the method may determine the number of ID-1 and ID-2 by manual experimentation, or they may be automatically derived by a reaction and / or analysis device, or a combination of both. Furthermore, these numbers may be actually output to an output unit such as a display, or they may be temporarily stored, for example, in the CPU or computer's calculation unit.
[0026] Figure 2 shows an example of a system that performs the process (S12) of obtaining the number of first short-chain nucleic acids (number of ID-1) and the number of second short-chain nucleic acids (number of ID-2).
[0027] The analysis unit has the function of analyzing nucleic acids according to a pre-stored program and includes a calculation unit and an output unit. The calculation unit, output unit, and detection unit are electrically connected to each other. For example, information obtained by the detection unit is sent to the calculation unit according to a program pre-stored in a storage unit (not shown) under the control of a computer. The calculation unit converts the sent information into the number of target short-chain nucleic acids. The output unit then outputs the number of the first short-chain nucleic acid (number of ID-1) and the number of the second short-chain nucleic acid (number of ID-2).
[0028] Next, in S13, the ratio R1 = (number of ID-1s / number of ID-2s) is calculated for the sample derived from the target sample, and the ratio R2 = (number of ID-1s / number of ID-2s) is calculated for the sample derived from the control sample.
[0029] If the control sample consists of multiple types of samples, the ratio R2 is calculated for each type. For example, if the control sample includes both non-cancer and breast cancer samples, two ratios R2 are calculated: one derived from the non-cancer samples and another derived from the breast cancer samples.
[0030] In S14, by comparing the value of ratio R1 with the value of ratio R2, it is determined whether or not the sample derived from the subject is derived from a subject that has or is at risk of developing the target disease.
[0031] Specifically, for example, the ratio R1 of the sample to be identified can be compared with the ratio R2 value derived from non-cancer samples and the ratio R2 value derived from breast cancer samples to determine which one it is closer to. If it is determined to be closer to the ratio R2 value derived from the disease, it can be determined that the sample originated from a subject who has the disease or is at risk of developing it.
[0032] Furthermore, for example, the ratio R1 of the sample to be identified can be compared with the ratio R2 value derived from non-cancer samples and the ratio R2 value derived from breast cancer samples, and it can be determined which one it is closer to. If it is determined to be closer to the ratio R2 value derived from non-cancer samples, it can be determined that the sample originates from a subject that has no or low risk of developing the disease.
[0033] Furthermore, to differentiate between samples from non-cancer individuals and samples from cancer patients, for example, it is possible to use a control sample with a clearly identifiable origin and a defined threshold. In this case, if the ratio R1 of the samples to be identified is greater than or less than the defined threshold, the sample can be determined to originate from an individual who has the disease or is at risk of developing it.
[0034] For example, the disease status of such a subject can be examined by using the method according to the embodiment. In other words, the method can also be used to diagnose the subject. The results obtained by the method may assist in a physician's diagnosis. Alternatively, the method can be used to assist in disease screening, such as in health examinations.
[0035] In this method, information regarding RNA in the sample may be obtained after the RNA has been extracted from the sample. The RNA extraction method may be a method known in itself, and for example, a commercially available kit may be used.
[0036] Table 1 shows an example of a reference sequence set.
[0037] [Table 1]
[0038] Table 1 shows an example of a reference sequence set when the target disease is pancreatic cancer. The first reference sequence is a continuous 7-nucleotide sequence represented as GGCAGTG from 5' to 3'. The second reference sequence is a continuous 7-nucleotide sequence represented as GTAGTGT from 5' to 3'.
[0039] Here, the first reference sequence and the second reference sequence can be swapped. Taking No. 1-1 as an example, the first reference sequence may be a sequence of 7 consecutive bases represented by GTAGTGT in the 5' to 3' direction, and the second reference sequence may be a sequence of 7 consecutive bases represented by GGCAGTG in the 5' to 3' direction.
[0040] In this embodiment, for example, the first short nucleic acid is a sequence that includes a sequence at any of its positions that perfectly matches the first reference sequence in Table 1, and the second short nucleic acid is a sequence that includes a sequence at any of its positions that perfectly matches the second reference sequence in Table 1.
[0041] Figure 3 shows the ratio R calculated for each sample group using the method according to this embodiment, with the sequence No. 1-1 as the reference sequence.
[0042] Figure 3 shows the ratio R on the vertical axis and the sample groups H(2021), H(2022), H(2023-1), H(2023-2) derived from four different non-cancer sample groups; BC(2021), BC(2022), BC(2023-1) derived from three different breast cancer sample groups; PC(2021), PC(2022), PC(2023-1), PC(2023-2) derived from four different pancreatic cancer sample groups; and CC(2022), CC(2023-1), CC(2023-2) derived from three different colorectal cancer sample groups on the horizontal axis. The numbers in parentheses indicate the analysis year for each sample group. Since 2023 was analyzed twice, it is listed as 2023-1 and 2023-2. Similarly, in Figures 5 to 14, which will be discussed later, the ratio R is plotted on the vertical axis and the sample groups are plotted on the horizontal axis.
[0043] By using this data as the ratio R2 for samples derived from control samples and comparing it with the ratio R1 for samples derived from the target subjects, it can be determined whether the samples derived from the target subjects originate from subjects who have or are at risk of developing the target disease. Using Figure 3 as an example, if the value of ratio R1 for samples derived from the target subjects is close to the value of ratio R2 for samples derived from the pancreatic cancer sample group, it can be determined that the sample originates from a subject who has or is at risk of developing pancreatic cancer. In the illustrated example, the ratio R for samples derived from each sample group corresponds to ratio R2.
[0044] According to the first embodiment, a method is provided for identifying a sample derived from a target.
[0045] (Second Embodiment) The second embodiment is a further method for identifying whether a sample originates from an infected subject or not. This method will be explained with reference to Figure 4. First, in S41, RNA molecules are extracted from the sample derived from the subject and from the sample derived from the control sample, and the sequences of each RNA molecule are decoded. Next, in S42, the number of first short-chain nucleic acids (ID-1) and second short-chain nucleic acids (ID-2) contained in the RNA molecules of the sample derived from the subject, which were decoded in S41, and the number of first short-chain nucleic acids (ID-1) and second short-chain nucleic acids (ID-2) contained in the RNA molecules of the sample derived from the control sample are obtained. Next, in S43, the ratio R1 = (number of ID-1 molecules / number of ID-2 molecules) is calculated for the sample derived from the subject, and the ratio R2 = (number of ID-1 molecules / number of ID-2 molecules) is calculated for the sample derived from the control sample. Subsequently, in S44, by comparing the value of ratio R1 with the value of ratio R2, it is determined whether or not the sample originates from an object that has or is at risk of developing the target disease.
[0046] Here, RNA sequencing can be performed using sequencing methods such as next-generation sequencers, nanopore sequencers, Sanger sequencing (electrophoresis, capillary sequencing), or Maxam-Gilbert sequencing. Alternatively, these methods may be used in combination.
[0047] According to the second embodiment, a method is provided that can identify samples originating from infected individuals with higher accuracy.
[0048] (Third embodiment) The third embodiment is a set of biomarkers for identifying samples derived from cancer patients. Here, biomarkers are also simply referred to as markers. Examples of candidate biomarker sequences are shown in Table 2.
[0049] [Table 2]
[0050] Nos. 1-1, 2-1, 3-1, and 4-1 are biomarker sets for identifying samples derived from pancreatic cancer patients. No. 5-1 is a biomarker set for identifying samples derived from breast cancer.
[0051] The marker sets shown in Table 2 can be used as sequence sets containing a first reference sequence and a second reference sequence. Each sequence is the full-length sequence of the first reference sequence and the second reference sequence, respectively. Sequences to be counted as the first short-chain nucleic acid (ID-1) are all sequences that contain a sequence that is exactly identical to the first reference sequence at any position within that sequence. The same applies to the second short-chain nucleic acid (ID-2).
[0052] (Example 1) Tables 3 and 4 show the AUC values obtained when samples were identified using the method according to this embodiment, with the five sequence sets listed in Table 2 serving as the first and second reference sequences. The samples identified were those classified by analysis year and those classified by disease, etc.
[0053] The sample groups, classified by analysis year, are: i) sample groups derived from non-cancerous, breast cancer, and pancreatic cancer samples analyzed in 2021; ii) sample groups derived from non-cancerous, breast cancer, colorectal cancer, and pancreatic cancer samples analyzed in 2022; iii) sample groups derived from non-cancerous, breast cancer, colorectal cancer, and pancreatic cancer samples analyzed in 2023; and iv) sample groups derived from non-cancerous, colorectal cancer, and pancreatic cancer samples analyzed in 2023.
[0054] The sample groups classified by disease are: non-cancerous, breast cancer, colorectal cancer, combinations of non-cancerous and breast cancer, combinations of non-cancerous and colorectal cancer, combinations of breast cancer and colorectal cancer, and combinations of non-cancerous, breast cancer, and colorectal cancer.
[0055] Each AUC value indicates the ability to distinguish samples originating from the target disease from the sample group described above. "Target Disease" lists the disease being identified, while "Comparison Group" lists the diseases included in the control group used to calculate the AUC value. The numbers in parentheses () represent the analysis year for each sample group. Since 2023 was analyzed twice, it is listed as 2023-1 and 2023-2. "ALL" refers to all years; for example, "Non-(ALL)" uses 96 samples from non-cancer individuals across all years as the comparison group. The same applies to Tables 9 and 10, described later.
[0056] Furthermore, Figures 3 and 5-8 show the ratio R calculated by the method according to this embodiment, using the sequences described in Tables 3 and 4 as reference sequences.
[0057] [Table 3]
[0058] [Table 4]
[0059] Tables 3 and 4 show that the AUC values for the sample groups classified by analysis year never fall below 0.7, and the small variation in AUC values across analysis years indicates that the identification method according to this embodiment possesses both high discriminative power and reproducibility. Furthermore, the AUC values for the sample groups classified by disease demonstrate that even when samples derived from diseases other than the target disease are included in addition to samples derived from non-cancer specimens, it is still possible to identify samples derived from the target disease.
[0060] Figures 3 and 5-8 clearly show that the ratios calculated for samples derived from the target cancer sample group and the ratios calculated for samples derived from non-cancer and non-target cancer sample groups show distinctly different trends. Furthermore, these figures clearly demonstrate that the variability of the ratio R between datasets is small when samples are derived from the same type of sample group, indicating that the method according to this embodiment has high reproducibility.
[0061] (Example 2) Table 5 shows the AUC values when the base lengths of the first and second reference sequences are varied. The AUC values are listed separately for each analysis year, calculated from sample groups including non-cancer samples, breast cancer samples, colorectal cancer samples, and pancreatic cancer samples.
[0062] The first and second reference sequences in Table 5 are variations of the reference sequence designated as No. 1-1, with altered base lengths. The first and second reference sequences, designated as No. 1-6 to 1-9, are sequences in which one or two bases are added to the 5' and 3' ends of the first and second reference sequences designated as No. 1-1, respectively. The first and second reference sequences, designated as No. 1-10, are sequences in which the 5' end of the first and second reference sequences designated as No. 1-1 is shortened by one base. In other words, the reference sequences designated as No. 1-6 to 1-9 are sequences that include the reference sequence designated as No. 1-1, and the reference sequences designated as No. 1-10 are sequences that are included in the reference sequence designated as No. 1-1.
[0063] As shown in Table 5, high AUC values are observed even when sequences longer or shorter than the reference sequence designated as No. 1-1 are used as the reference sequence. Therefore, a continuous sequence of 6 to 9 nucleotides in length can be used as the reference sequence in the method and biomarker set according to this embodiment.
[0064] [Table 5]
[0065] (Example 3) Table 6 and Figure 9 show a comparison between selecting short nucleic acids to be analyzed using the name of miRNA and selecting them using sequences containing a specific reference sequence.
[0066] The miRNAs used for comparison were three types: hsa-miR-34a-5p (SEQ ID NO: 1), which contains the sequence of seven consecutive bases GGCAGTG; hsa-miR-142-3p (SEQ ID NO: 2); and hsa-miR-486-5p (SEQ ID NO: 3), which contain the sequence of seven consecutive bases GTAGTGT. The sequences of these miRNAs are shown in Table 6.
[0067] Table 7 shows the AUC values for Nos. 6-1 to 6-3 when a sample was identified using the method according to this embodiment, with the first and second short-chain nucleic acids being the three types of miRNAs described above. For comparison, the data for No. 1-1 is also included.
[0068] [Table 6]
[0069] [Table 7]
[0070] From No. 6-1, when hsa-miR-34a-5p in the sample group was designated as the first short-chain nucleic acid and hsa-miR-142-3p as the second short-chain nucleic acid, the AUC values showed large variability between datasets. Furthermore, from No. 6-2, when hsa-miR-34a-5p was designated as the first short-chain nucleic acid and hsa-miR-486-5p as the second short-chain nucleic acid, the AUC values were low in all datasets. In addition, from No. 6-3, when hsa-miR-142-3p was designated as the first short-chain nucleic acid and hsa-miR-486-5p as the second short-chain nucleic acid, the AUC values showed large variability between datasets.
[0071] Figure 9 shows the ratio R when the miRNAs described as No. 6-1 to 6-3 are used as the first and second short-chain nucleic acids. Compared to when the reference sequence No. 1-1 was used, when the short-chain nucleic acids to be analyzed were selected by the name of miRNA, there was greater variability between datasets, and the difference between the ratio in samples derived from the target disease and the ratio in the comparison samples was small.
[0072] These results show that classifying the short nucleic acids to be analyzed by sequences containing specific reference sequences significantly improves discriminative power and reproducibility compared to classifying them by miRNA name.
[0073] (Example 4) The first and second reference sequences were established using the samples shown in Table 8.
[0074] [Table 8]
[0075] As shown in Table 8, the number of samples obtained was 24 each from non-cancer individuals, breast cancer patients, and pancreatic cancer patients analyzed in 2021; 24 each from non-cancer individuals, breast cancer patients, pancreatic cancer patients, and colorectal cancer patients analyzed in 2022; 24 each from non-cancer individuals, breast cancer patients, pancreatic cancer patients, and colorectal cancer patients analyzed in 2023; 24 from non-cancer individuals analyzed at different times within the same year; and 22 each from pancreatic cancer patients and colorectal cancer patients, for a total of 332 samples. The reference sequence was set as follows. 1) First, from the 2094 possible sequences of seven consecutive base pairs from the 2nd to 8th base pairs from the 5' end of miRNAs registered in the database, we arbitrarily selected multiple candidate sequences for the first reference sequence and the second reference sequence. 2) Next, the candidate sequences were randomly combined to form sequences that were different from each other, and combinations that showed a tendency to identify samples derived from the target disease using the method according to this embodiment were selected. As a result, five first reference sequence candidates and four or five second reference sequence candidates for each first reference sequence candidate were selected, resulting in a total of 23 combinations. 3) From 23 possible combinations, we obtained five combinations of reference sequences that can identify samples originating from the target disease with high AUC values and high reproducibility.
[0076] Tables 9 to 13 show the AUC values when a sample is identified using the method according to this embodiment, with the 23 sequence sets selected in 2) as the first and second reference sequences. Tables 9 to 12 are candidate sequence combinations for identifying samples derived from pancreatic cancer specimens. Table 13 is a candidate sequence combination for identifying samples derived from breast cancer specimens.
[0077] Furthermore, using the sequences No. 1-1 to No. 5-4 described in Tables 9 to 13 as reference sequences, the ratios R calculated for each sample group using the method according to this embodiment are shown in Figures 2 and 5 to 14.
[0078] [Table 9]
[0079] [Table 10]
[0080] [Table 11]
[0081] [Table 12]
[0082] [Table 13]
[0083] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents.
Claims
1. A method for identifying a sample originating from a subject; The relationship with the target disease is suggested by a combination of two sequences based on predetermined criteria. A first reference sequence consisting of a sequence of 6 to 9 consecutive bases, and A second reference sequence consisting of a sequence of 6 to 9 consecutive bases, and which is different from the base sequence of the first reference sequence. Set one or more sets: Obtain the number of first short-chain nucleic acids (number of ID-1) and the number of second short-chain nucleic acids (number of ID-2) present in the short-chain nucleic acid group contained in the sample derived from the subject and the sample derived from the control sample, respectively. Here, The first short nucleic acid contains a sequence at any of its positions that is in exact agreement with the first reference sequence. The second short nucleic acid contains a sequence at any of its positions that is in exact agreement with the second reference sequence; For the sample derived from the aforementioned subject, ratio R 1 = (Number of ID-1 / Number of ID-2), and For the sample derived from the aforementioned control sample, ratio R 2 = (Number of ID-1 / Number of ID-2) Calculate: The aforementioned ratio R 1 The value and the ratio R 2 By comparing it with the value of the subject, it is determined whether or not the sample derived from the subject has or is at risk of developing the disease of the subject; A method that includes this.
2. The method according to claim 1, wherein the length of the first reference array and the length of the second reference array are equal.
3. The method according to claim 1, wherein the short-chain nucleic acid is miRNA.
4. The method according to claim 1, wherein the first reference sequence and the second reference sequence are each a sequence of seven consecutive base pairs.
5. The method according to claim 1, wherein the first reference sequence and the second reference sequence are sequences selected from at least one combination of sequences shown in Table 2.
6. The method according to claim 1, wherein the disease is cancer.
7. The method according to claim 1, wherein the disease is pancreatic cancer or breast cancer, the control specimen is one or more types of specimens from which the condition of the subject from which it originates is known, and the control specimen is at least one selected from the group consisting of non-cancer specimens, pancreatic cancer specimens, breast cancer specimens and colorectal cancer specimens.
8. The method according to claim 1, wherein the sample derived from the subject is serum or plasma.
9. The method according to claim 1, wherein the method is performed using a sequence determination technique such as a next-generation sequencer, a nanopore sequencer, the Sanger method (electrophoresis, capillary sequencer), or the Maxam-Gilbert method.
10. A set of biomarkers for identifying whether a sample originates from an infected person, wherein at least one marker is selected from the sequence combinations shown in Table 2.
11. The marker set according to claim 10, wherein the disease the affected subject has is cancer.
12. The marker set according to claim 10, wherein the disease the subject has is pancreatic cancer.
13. The marker set according to claim 10, wherein the disease the subject has is breast cancer.