Classification method, classification device, classification system, classification program, and recording medium
By grouping the base sequences of nucleic acid molecules and setting representative sequences, and using similarity classification to determine the sequences, the problem of excessive computation is solved, achieving more efficient computation and disease identification.
Patent Information
- Application Number
- CN202480042612.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-30
- Filing Date
- 2024-05-22
- Publication Date
- 2026-01-20
AI Technical Summary
In comprehensive molecular analysis, especially in the classification and disease identification of millions of microRNAs in human blood, the excessive computational load leads to unresolved issues such as increased equipment costs and reduced throughput.
By grouping the base sequences of nucleic acid molecules into multiple groups and setting representative sequences, the determination sequences can be classified into each group based on similarity, thus reducing the amount of computation.
This reduces the amount of data processed when classifying the base sequences of nucleic acid molecules, improves computational efficiency, and reduces equipment costs and analysis time.
Smart Images

Figure CN121368801A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a classification method, a classification apparatus, a classification system, a classification program, and a recording medium. BACKGROUND
[0002] In recent years, a comprehensive quantitative analysis of nucleic acids represented by Next-Generation Sequencing (hereinafter sometimes referred to as NGS) is applied to the discrimination of diseases such as cancer.
[0003] With the development of comprehensive analysis technology, the number of analyzable molecules has increased exponentially. However, with the increase in the number of molecules to be analyzed, the amount of calculation required for analysis has increased exponentially, and as a result, serious problems such as an increase in equipment costs and a decrease in throughput due to the need for high-performance computing devices and long analysis times have often occurred.
[0004] In the past, in order to suppress such an increase in the amount of calculation, important molecules were selected by using statistical analysis and machine learning to reduce the number of molecules to be analyzed, thereby achieving a reduction in the amount of calculation (see Non-Patent Documents 1 and 2).
[0005] Non-Patent Document 1: Jin et al, 2017, Clinical Cancer Research, Evaluation of Tumor-Derived Exosomal miRNA as Potential Diagnostic Biomarkers for Early-Stage Non-Small Cell Lung Cancer Using Next-Generation Sequencing
[0006] Non-Patent Document 2: Asakura et al, 2020, Communications Biology, A miRNA-based diagnostic model predicts resectable lung cancer in humans with high accuracy SUMMARY
[0007] PROBLEMS TO BE SOLVED BY THE INVENTION
[0008] In the analysis process using comprehensive molecular species analysis, as a process with a large amount of processing (specifically, the amount of calculation), the following two steps can be cited.
[0009] Step 1: Classification of target nucleic acid molecules
[0010] Step 2: Disease discrimination prediction
[0011] Here, for example, there are about 2600 known microRNAs in microRNAs in human blood, and there are several million microRNAs in several μL of blood as a sample. Also, in the case of performing the above-described step 1 on several million microRNAs, for example, several million microRNAs are individually compared with the known about 2600 microRNAs and classified, and thus the processing amount (specifically, the calculation amount) is large.
[0012] In the method of screening important molecules using statistical analysis, machine learning (refer to Non-Patent Literatures 1 and 2), after detecting the molecular species comprehensively, the processing amount (specifically, the calculation amount) in the subsequent machine learning, statistical analysis is reduced by reducing the number of analyzed molecular species, and thus the processing amount in the above-described step 2 can be reduced.
[0013] However, in the method of Non-Patent Literatures 1 and 2, the reduction of the processing amount in the above-described step 1 cannot be achieved, and there are still problems of an increase in equipment cost and a reduction in throughput due to a high-performance computing device or a long time for analysis.
[0014] An object of the present disclosure is to provide a classification method, a classification device, a classification system, a classification program, and a recording medium that can reduce the processing amount when classifying base sequences of nucleic acid molecules.
[0015] Means for solving the problem
[0016] The classification method of one aspect of the present disclosure has: a measurement step of measuring a base sequence of a nucleic acid molecule in a measurement sample as a measurement sequence; and a classification step of classifying the measurement sequence into each of a plurality of groups obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule, based on a similarity between a representative sequence set as a base sequence representing each of the groups and the measurement sequence measured in the measurement step.
[0017] Effects of the Invention
[0018] According to the present disclosure, the processing amount when classifying base sequences of nucleic acid molecules can be reduced. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a flowchart showing an example of each step of the classification method of the present embodiment.
[0020] Figure 2 is a block diagram showing an example of a computer that functions as the classification device of the present embodiment.
[0021] Figure 3 is a block diagram showing an example of a functional configuration of the classification device of the present embodiment.
[0022] Figure 4 is a list showing 56 kinds of microRNAs as classification targets in the present embodiment.
[0023] Figure 5 is a list showing each group obtained by grouping the 56 kinds of microRNAs shown in Figure 4 is a list showing each group obtained by grouping the 56 kinds of microRNAs shown in DETAILED DESCRIPTION
[0024] Hereinafter, an example of an embodiment of the technology of the present disclosure will be described based on the drawings. Note that, in all the drawings, the same reference signs are attached to constituent elements and processing that assume the same function, action, and effect, and sometimes the repeated description is appropriately omitted. Each drawing is merely schematically shown to the extent that the technology of the present disclosure can be sufficiently understood. Therefore, the technology of the present disclosure is not limited to the illustrated examples. In addition, in the present embodiment, the description is sometimes omitted for the configurations and well-known configurations that are not directly related to the present disclosure.
[0025] <Classification method 10>
[0026] First, the classification method 10 of the present embodiment will be described. Figure 1 is a schematic view showing an example of each process of the classification method 10 of the present embodiment.
[0027] The classification method 10 is a method of measuring the base sequence of a nucleic acid molecule in a measurement sample and classifying the base sequence into each group. As the measurement sample, a body fluid (for example, blood, serum, urine, tears, saliva, sweat, semen, lymph, tissue fluid, a body cavity fluid (for example, pleural fluid, ascites, and the like), cerebrospinal fluid, amniotic fluid, vaginal fluid, nasal discharge, and the like), a tissue, and a cell, and the like can be given. As the collection target of the measurement sample, a human being or a non-human animal can be given. As the non-human animal, a non-human mammal (monkey, dog, cat, mouse, rat, rabbit, cow, horse, pig, and sheep, and the like), a bird (chicken, quail, and the like), and the like can be given.
[0028] As the nucleic acid molecule, for example, a microRNA can be given. Note that, as the nucleic acid molecule, a small RNA other than a microRNA, another RNA, DNA, and the like can be given. In addition, as the nucleic acid molecule, a nucleic acid molecule not composed of only ATGCU bases can be given. Specifically, for example, a nucleic acid subjected to modification such as DNA / RNA methylation and editing such as A-to-I RNA editing can be given. In this way, various nucleic acid molecules can become the applicable target of the present classification method.
[0029] Specifically, as shown inFigure 1 As shown, the classification method 10 has a preparation step 11, a measurement step 12, a classification step 13, a reclassification step 14, a totalization step 15, and a determination step 16. In the present embodiment, as an example, the preparation step 11, the measurement step 12, the classification step 13, the reclassification step 14, the totalization step 15, and the determination step 16 are sequentially executed.
[0030] Note that the classification method 10 is a method having the totalization step 15 and the determination step 16, and thus can also be referred to as a totalization method or a determination method. In addition, in the determination step 16, in a case where inspection or analysis, or the like is performed by determination, the classification method can also be referred to as an inspection method or an analysis method. Hereinafter, each step of the classification method 10 will be described.
[0031] <Preparation Step 11>
[0032] The preparation step 11 is a step executed as preparation for executing the classification step 13. In the present embodiment, the preparation step 11 is executed, for example, before the measurement step 12 and the classification step 13 are executed. Note that the preparation step 11 can be executed at least before the classification step 13 is executed, and can also be executed after the measurement step 12.
[0033] In the preparation step 11, a plurality of base sequences of a plurality of nucleic acid molecules are grouped into a plurality of groups by a specific rule, and a representative sequence is set as a base sequence representing each of the plurality of groups. Note that hereinafter, a base sequence other than the representative sequence among the base sequences constituting each of the plurality of groups will be sometimes referred to as a similar sequence.
[0034] Each of the plurality of groups grouped in the preparation step 11 can be constituted by the representative sequence and one or more similar sequences. That is, the representative sequence and one or more similar sequences belong to each of the plurality of groups. Note that a group can also be constituted only by the representative sequence. That is, as a group, a group having no similar sequence can also be included.
[0035] Here, for example, among microRNAs in human blood, there are about 2600 known microRNAs. In a case where the base sequences of the microRNAs in human blood are taken as classification targets, the base sequences of the about 2600 microRNAs are grouped.
[0036] The representative sequence can be set, for example, as a base sequence having the longest base sequence among the base sequences constituting a group. Note that as the representative sequence, it is not limited to a base sequence having the longest base sequence among the base sequences constituting a group, and for example, a base sequence having the highest expression amount among the base sequences constituting a group can also be set. A base sequence having the highest expression amount can be set, for example, by the following first method or second method.
[0037] Here, the base sequence of the nucleic acid molecule present in a large amount in the test sample is determined for each test sample. That is, for example, if the blood of a human being, the microRNA present in a large amount in the blood is determined. The first method is a method of predefining, as the representative sequence, the base sequence of the nucleic acid molecule present in a large amount in the test sample among the base sequences constituting the group, using this property. Specifically, in the first method, for example, based on the measurement data or statistical data, or the like, obtained by performing measurement on a plurality of test samples as the sample for setting, the base sequence of the nucleic acid molecule present in a large amount in the test sample among the base sequences constituting the group can be predefined as the representative sequence as the base sequence having the highest expression amount.
[0038] The second method is a method of setting, as the representative sequence, the base sequence having the highest expression amount among the base sequences constituting the group, based on the measurement results in the measurement step 12. In the second method, for example, in the measurement step 12, the base sequence having the largest amount measured among the base sequences constituting the group can be set as the representative sequence as the base sequence having the highest expression amount. Therefore, in the second method, the representative sequence is sometimes changed depending on the measurement results of the measurement step 12. In addition, in the second method, the preparation step 11 is performed after the measurement step 12.
[0039] In addition, as the representative sequence, it can not be a sequence present in nature, but a sequence artificially defined. As this sequence, for example, a sequence partially amplified by polymerase chain reaction (PCR) or the like can be cited.
[0040] The specific rule is a rule based on the similarity of the base sequences grouped into a plurality of groups (specifically, one or more similar sequences) with respect to the representative sequence, or the similarity of the base sequences grouped into a plurality of groups (specifically, the representative sequence and one or more similar sequences) with respect to each other.
[0041] For example, in the case where the similarity of one or more similar sequences with respect to the representative sequence is used as the specific rule, a difference of n bases or less (n = a natural number) of the similar sequence with respect to the representative sequence can be used as the specific rule. At this time, the similarity of the representative sequence with respect to the reverse complementary strand can also be included as the specific rule. That is, a difference of n bases or less (n = a natural number) with respect to the representative sequence or the reverse complementary strand thereof can be used as the specific rule.
[0042] In addition, in a case where the similarity of the base sequences grouped into a plurality of groups (specifically, the representative sequence and one or more similar sequences) to each other is specified as the specific rule, the difference of the base sequences to each other can be specified as the specific rule, which is within n bases (n = a natural number). At this time, the similarity of the base sequences to the reverse complementary strand can also be included as the specific rule. That is, the difference of the base sequences or the reverse complementary strand thereof can be specified as the specific rule, which is within n bases (n = a natural number).
[0043] The above difference includes, for example, substitution, more, less, addition, and deletion of bases. As the similarity, for example, a similarity score used in homology search of a base sequence can be used.
[0044] In judging the above similarity, the sequence on the 5' side and the 3' side in the base sequence (i.e., the sequence at the end) can be focused on. That is, for example, grouping can be performed by the similarity (specifically, the difference) of the sequence on the 5' side and the 3' side in the base sequence (i.e., the sequence at the end).
[0045] Note that the specific rule is not limited to the rule based on the similarity of one or more similar sequences to the representative sequence and the rule based on the similarity of the base sequences grouped into a plurality of groups (specifically, the representative sequence and one or more similar sequences) to each other.
[0046] For example, a base sequence having a specific motif (for example, a specific base sequence, a consensus sequence) can also be specified as the specific rule, and the base sequences can be grouped. Specifically, for example, a base sequence having a specific base sequence (refer to the following sequence) can be grouped as one group.
[0047] AT**A*C*A*************(* can be any of ATGCU)
[0048] In addition, for example, the base sequence of the precursor and the base sequence of the same type and the base sequence serving as the basis thereof can also be grouped as one group.
[0049] Further, the specific rules of the above grouping can be combined in a plurality to perform grouping.
[0050] In the present embodiment, the preparation step 11 can be performed, for example, when the classification method is initially executed, and the classification step 13 can be performed using the group grouped when the classification method is initially executed and the representative sequence set when the classification method is executed for the second time or more.
[0051] Therefore, the preparation process 11 need not be performed every time the classification method is executed. Note that in the case of changing the grouping of the groups and the setting of the representative sequences, and the like, the preparation process 11 can be performed once for every multiple of the classification method or every time the classification method is executed after the second time.
[0052] <Measurement process 12>
[0053] The measurement process 12 is a process of measuring the base sequence of the nucleic acid molecule in the measurement sample as a measurement sequence. Specifically, in the measurement process 12, using a next-generation sequencer (NGS) as a measurement device, the base sequence of the nucleic acid molecule in the measurement sample is measured as a measurement sequence.
[0054] <Classification process 13>
[0055] The classification process 13 is a process of classifying the measurement sequence measured in the measurement process 12 into each group on the basis of the similarity between the representative sequence set as the base sequence representing each of the plurality of groups and the measurement sequence. That is, in the classification process 13, the measurement sequence is classified into each group by collating the measurement sequence with the representative sequence on the basis of the similarity.
[0056] In the classification process 13, for example, the measurement sequence is classified into the group in which the similarity between the representative sequence and the measurement sequence is the highest. Specifically, in the classification process 13, for example, the group in which the similarity score used in the homology search of the base sequence shows the highest score is classified. In this case, the measurement sequence is classified into one group. In the classification process 13, all the measurement sequences measured in the measurement process 12 are classified into groups.
[0057] Note that in the classification process 13, the group in which the similarity between the representative sequence and the measurement sequence is equal to or higher than a threshold value can also be classified. In this case, the measurement sequence can also be classified into a plurality of groups. In this way, in the classification process 13, the measurement sequence can also be classified into a plurality of groups. That is, the measurement sequence need not be one-to-one with the groups.
[0058] The measurement sequence need not be the same sequence as any of the base sequences grouped. In addition, the measurement sequence classified by the group can be collated with other groups again and classified before being totaled.
[0059] <Re-classification process 14>
[0060] The re-classification process 14 is a process of re-classifying the measurement sequence classified into each of the plurality of groups into each of the sequences constituting the base sequence of the group after the classification process 13. That is, the measurement sequence classified into the group is further classified into any of the representative sequence belonging to the group and one or more similar sequences.
[0061] In the reclassification process 14, for example, the measurement sequence is classified into any one of the representative sequence and the sequence having the highest similarity to the measurement sequence among the one or more similar sequences. Specifically, in the reclassification process 14, for example, the classification is performed into the representative sequence and any one of the one or more similar sequences having the highest score of the similarity score used in the homology search of the base sequence.
[0062] Note that the reclassification process 14 is not necessarily performed, and the total process 15 can be performed without the reclassification process 14 after the classification process 13.
[0063] <Total process 15>
[0064] The total process 15 is a process of totaling the number of molecules of the measurement sequence classified into each group of the plurality of groups in the classification process 13. Specifically, in the total process 15, the number of molecules of the measurement sequence reclassified into the representative sequence and any one of the one or more similar sequences in the reclassification process 14 is totaled for each base sequence (i.e., the representative sequence and the one or more similar sequences).
[0065] Note that in the case where the total process 15 is performed without the reclassification process 14 after the classification process 13, the number of molecules of the measurement sequence classified into each group of the plurality of groups is totaled for each group in the total process 15.
[0066] <Determination process 16>
[0067] The determination process 16 is a process of performing a specific determination of the measurement sample based on the number of molecules totaled by the total process 15.
[0068] As the specific determination, a disease discrimination prediction of a provider of the sample can be given. The disease discrimination prediction can be performed, for example, by the distribution of the number of molecules of the measurement sequence (i.e., the expression amount of the nucleic acid) totaled for each base sequence (i.e., the representative sequence and the one or more similar sequences) in the total process 15.
[0069] Specifically, for example, the determination of the disease or the like using the expression amount of the microRNA is performed by an arbitrary determination formula using the number of molecules of each group as a variable. For example, the following determination formula can be used.
[0070] Determination formula = Y1 x expression amount of group 1 +... + Y n x expression amount of group n + X
[0071] (Y is a coefficient determined in advance for each group, and X is a constant determined in advance.)
[0072] It should be noted that machine learning can also be used to determine diseases based on the expression levels of microRNAs. Specifically, the expression levels of individual microRNAs and information on the presence or absence of diseases are used as teaching data to build a learning model. Then, the sequences and their expression levels measured in the totalization step 15 are used as input factors. Through this learning model, the presence or absence of diseases, which is the output factor, can be determined.
[0073] In addition, the determination process 16 is not limited to disease discrimination and prediction, but can also include basic biological analysis, regression prediction, anomaly detection, etc.
[0074] <Classification System 20>
[0075] Next, the classification system 20, which performs the above classification method, will be described. For example... Figure 2 As shown, the classification system 20 has a measuring device 21 and a classification device 30.
[0076] <Measuring Apparatus 21>
[0077] The measuring device 21 is an example of a measuring unit, and is the device that performs the above-described measuring procedure 12. That is, the measuring device 21 measures the base sequence of nucleic acid molecules in the sample as the measuring sequence. For example, NGS is used as the measuring device 21.
[0078] <Classification device 30>
[0079] The classification device 30, as an example of a classification unit, is the device that performs the classification step 13 described above. That is, the classification device 30 classifies the measured sequence into its respective group based on the similarity between a representative sequence (defined as a base sequence representing each group in a plurality of groups) and the measured sequence measured by the measuring device 21. These plurality of groups are obtained by grouping the base sequences of multiple nucleic acid molecules according to specific rules. Furthermore, the classification device 30 performs the reclassification step 14, the totalization step 15, and the determination step 16 described above.
[0080] The sorting device 30 has the function of a computer, such as Figure 2 As shown, it includes a CPU (Central Processing Unit) 31, a ROM (Read Only Memory) 32, a RAM (Random Access Memory) 33, a memory 34, an input unit 35, a display unit 36, and a communication interface (I / F) 37. All components are connected to each other via a bus 39 in a manner enabling communication.
[0081] CPU 31 is the central processing unit, executing various programs and controlling various components. Specifically, CPU 31 reads programs from ROM 32 or memory 34, and uses RAM 33 as its working area to execute the programs. CPU 31 performs control and various arithmetic operations according to the programs stored in ROM 32 or memory 34. It should be noted that CPU 31 is an example of a processor.
[0082] ROM32 records various programs and data. RAM33 serves as a working area to temporarily store programs or data. Memory 34 consists of an HDD (Hard Disk Drive) or SSD (Solid State Drive) and records various programs and data, including the operating system.
[0083] In this embodiment, for example, a classification program for performing the classification process described above is recorded in memory 34. The classification program can be a single program or a group of programs consisting of multiple programs or modules.
[0084] Additionally, the memory 34 stores group information for each group that has been grouped through the aforementioned preparation step 11, as well as sequence information of representative sequences set as base sequences representing each group. It should be noted that the classification procedure, group information, and sequence information can also be stored in the ROM 32. The ROM 32 or memory 34 functions as an example of a non-transitory recording medium.
[0085] The input unit 35 includes a pointing device such as a mouse and a keyboard for performing various inputs. Additionally, the input unit 35 receives information from the measurement sequence measured by the measuring device 21 as input.
[0086] The display unit 36 is, for example, an LCD screen, which displays various information. The display unit 36 may also be a touch screen, functioning as an input unit 35.
[0087] Communication interface 37 is an interface for communicating with other devices, such as using standards such as Ethernet (registered trademark), FDDI (Fiber Distributed Data Interface), and Wi-Fi (registered trademark).
[0088] like Figure 3 As shown, in the sorting device 30, the CPU 31 performs functions as the sorting function unit 150, the re-sorting unit 160, the totaling unit 170, and the determination unit 180 by executing the sorting program.
[0089] The classification function section 150 executes the above-described classification process 13. That is, the classification function section 150 classifies the measurement sequence into each of the groups based on the similarity between the representative sequence set based on the base sequence representing each of the plurality of groups obtained by grouping the base sequences of the plurality of nucleic acid molecules by the specific rule and the measurement sequence measured by the measurement device 21 (refer to the above-described classification process 13).
[0090] The reclassification section 160 executes the above-described reclassification process 14. That is, the measurement sequence classified into each of the plurality of groups by the classification function section 150 is reclassified into each of the base sequences constituting the group (refer to the above-described reclassification process 14).
[0091] The total section 170 executes the total process 15. That is, the total section 170 totals the number of molecules of the measurement sequence classified into each of the plurality of groups by the classification function section 150. Specifically, the total section 170 totals the number of molecules of the measurement sequence reclassified into any one of the representative sequence and one or more similar sequences by the reclassification section 160, per base sequence (that is, the representative sequence and one or more similar sequences) (refer to the above-described total process 15).
[0092] The determination section 180 executes the above-described determination process 16. That is, the determination section 180 performs the specific determination of the measurement sample based on the number of molecules totaled by the total section 170 (refer to the above-described determination process 16).
[0093] Note that, in the present embodiment, the classification system 20 has the measurement device 21 and the classification device 30, but as the classification system 20, one device can also be constituted. In this case, the one device functions as an example of the measurement section and the classification section.
[0094] In addition, the classification device 30 can also be constituted by a plurality of devices. For example, the classification device 30 can also be constituted by a plurality of (for example, four) devices that share the execution of the above-described classification process 13, reclassification process 14, total process 15, and determination process 16.
[0095] <Embodiment>
[0096] Next, as an embodiment, an example in which each process of the above-described classification method is executed using blood containing microRNA (an example of a nucleic acid molecule) as a measurement sample will be described. In the present embodiment, for convenience, each process of the above-described classification method is executed targeting 56 kinds of microRNA (refer to Table 1). Figure 4 ) as a measurement sample will be described. In the present embodiment, for convenience, each process of the above-described classification method is executed targeting 56 kinds of microRNA (refer to Table 1).
[0097] Note that the embodiment illustrates an example of the technology of the present disclosure, and the technology of the present disclosure is not limited to the content of the embodiment.
[0098] <Preparation step 11>
[0099] In the present embodiment, 56 kinds of microRNAs are grouped based on the above-described specific rule, and the representative sequence is set as shown in Table 1. The representative sequence is set as the longest base sequence in the group. Therefore, the similar sequence is a sequence having a difference within 3 bases in the same sequence compared with the longest representative sequence. The difference includes, for example, substitution, more, less, addition, and deletion of bases with respect to the representative sequence.
[0100] In the present embodiment, 56 kinds of microRNAs are grouped based on the above-described specific rule, and the representative sequence is set as shown in Table 1. The representative sequence is set as the longest base sequence in the group. Therefore, the similar sequence is a sequence having a difference within 3 bases in the same sequence compared with the longest representative sequence. The difference includes, for example, substitution, more, less, addition, and deletion of bases with respect to the representative sequence. Figure 5 Figure 4 In the present embodiment, 56 kinds of microRNAs are grouped based on the above-described specific rule, and the representative sequence is set as shown in Table 1. The representative sequence is set as the longest base sequence in the group. Therefore, the similar sequence is a sequence having a difference within 3 bases in the same sequence compared with the longest representative sequence. The difference includes, for example, substitution, more, less, addition, and deletion of bases with respect to the representative sequence.
[0101] The set representative sequences are described below.
[0102] CAAAAACCGCAAUUACUUUUGCA (SEQ ID NO: 1) (hsa-miR-548h-3p)
[0103] AAAAGCUGGGUUGAGAGGGCGA (SEQ ID NO: 2) (hsa-miR-320a-3p)
[0104] CAAAGUGCUUACAGUGCAGGUAG (SEQ ID NO: 3) (hsa-miR-17-5p)
[0105] UGAGGUAGUAGUUUGUGCUGUU (SEQ ID NO: 4) (hsa-let-7i-5p)
[0106] UAUUGCACUCGUCCCGGCCUCC (SEQ ID NO: 5) (hsa-miR-92b-3p)
[0107] CAUAUACAACUUACUACUUUCCC (SEQ ID NO: 6) (hsa-miR-98-3p)
[0108] UAGCAGCACGUAAAUAUUGGCG (SEQ ID NO: 7) (hsa-miR-16-5p)
[0109] <Measurement step 12>
[0110] In the present embodiment, the base sequence of the microRNA in the blood of the measurement sample is measured as a measurement sequence by NGS as a measurement device, and the number of measurement sequences is quantified.
[0111] <Classification step 13>
[0112] In the present embodiment, all of the measurement sequences measured in the measurement step 12 are classified into the group in which the similarity score with the respective representative sequences of the plurality of groups shows the highest score. Note that microRNAs other than the 56 kinds of microRNAs have been excluded from the classification targets in advance.
[0113] <Reclassification step 14>
[0114] In the present embodiment, after the classification step 13, all of the measurement sequences classified into each of the plurality of groups are reclassified into the representative sequence and each of the one or more similar sequences of the group. At this time, any one of the representative sequence and the one or more similar sequences in which the similarity score shows the highest score is reclassified.
[0115] <Total step 15>
[0116] In the present embodiment, the number of molecules of the measurement sequences reclassified into any one of the representative sequence and the one or more similar sequences in the reclassification step 14 is totaled for each base sequence (representative sequence and one or more similar sequences).
[0117] <Determination step 16>
[0118] In the present embodiment, disease discrimination prediction is performed by the distribution of the number of molecules of the measurement sequences (i.e., the expression amount of the nucleic acid) totaled for each base sequence (representative sequence and one or more similar sequences) in the total step 15.
[0119] Specifically, for example, disease or the like using the expression amount of microRNA is determined by the following determination formula with the number of molecules of each group as a variable.
[0120] Determination formula = Y1 x expression amount of group 1 +... + Y7 x expression amount of group 7 + X
[0121] (Y is a coefficient determined in advance for each group, and X is a constant determined in advance.)
[0122] Note that determination of disease or the like using the expression amount of microRNA can also be performed by machine learning. Specifically, the expression amount of each of the 56 kinds of microRNAs as an object and information on the presence or absence of disease are subjected to machine learning as teaching data, and a learning model is constructed. Then, the measurement sequences and the expression amounts thereof totaled in the total step 15 are input as factors, and the presence or absence of disease as an output factor can be determined by the learning model.
[0123] <Effects of the embodiment>
[0124] According to the present embodiment, as described later, in the classification step 13, it is possible to reduce the processing amount (specifically, the calculation amount) when classifying the base sequences of microRNAs.
[0125] For example, in a case where the number of measurement sequences is set to 10,000,000, the calculation amount required for collating 1 measurement sequence with 1 base sequence is set to A, and 56 kinds of microRNAs are classified into 56 kinds of measurement sequences, the calculation amount is as follows.
[0126] 10,000,000 (molecules) x 56 (kinds) = 560,000,000 times... (1)
[0127] According to (1), the total calculation amount = 560,000,000 x A... (2).
[0128] On the other hand, in the present embodiment, 56 kinds of microRNAs are grouped into 7 groups, and the measurement sequences are classified based on the similarity to the representative sequence, and thus the calculation amount is as follows.
[0129] 10,000,000 (molecules) x 7 (kinds) = 70,000,000 times... (3)
[0130] According to (3), the total calculation amount = 70,000,000 x A... (4).
[0131] According to (2), (4), 70,000,000 x A / 560,000,000 x A = 0.125... (5).
[0132] As is clear from (5), according to the present embodiment, the calculation amount can be reduced to 0.125 times as compared with the case where the measurement sequences are classified into 56 kinds of microRNAs.
[0133] In a case where the reclassification process 14 is performed, the following calculation amounts are required. Here, it is assumed that 2,000,000 (molecules) are classified into Group 1, 2, and 3, respectively, and 1,000,000 (molecules) are classified into Group 4, 5, 6, and 7, respectively.
[0134] Group 1: Total calculation amount = 2,000,000 (molecules) x 34 (kinds) x A = 68,000,000... (11)
[0135] Group 2: Total calculation amount = 2,000,000 (molecules) x 6 (kinds) x A = 12,000,000... (12)
[0136] Group 3: Total calculation amount = 2,000,000 (molecules) x 5 (kinds) x A = 10,000,000... (13)
[0137] Group 4: Total calculation amount = 1,000,000 (molecules) x 4 (kinds) x A = 4,000,000... (14)
[0138] Group 5: Total calculation amount = 1,000,000 (molecules) x 3 (kinds) x A = 3,000,000... (15)
[0139] Group 6: Total calculation amount = 1,000,000 (molecules) x 2 (kinds) x A = 2,000,000... (16)
[0140] Group 7: Total calculation amount = 1,000,000 (molecules) x 2 (kinds) x A = 2,000,000... (17)
[0141] The total calculation amount obtained by the total of (11) to (17) is as follows.
[0142] Total calculation amount = 111,000,000 x A... (18)
[0143] The sum of (4) and (18) (refer to (19) below) is the calculation amount in the classification process 13 and the reclassification process 14.
[0144] 70,000,000 x A + 111,000,000 x A = 181,000,000 x A... (19)
[0145] According to (2), (19), 181,000,000 x A / 560,000,000 x A = 0.323... (20).
[0146] As is clear from (20), according to the present embodiment, the calculation amount can be reduced to 0.323 times compared with the case where the present embodiment is not used.
[0147] As described above, the technical idea of the present disclosure is to, using the insight that similar base sequences exist in nucleic acids, collect and group similar base sequences, thereby reducing the processing amount (specifically, the calculation amount).
[0148] <Notes>
[0149] (Scheme 1)
[0150] A classification method has:
[0151] a measurement process of measuring, as a measurement sequence, a base sequence of a nucleic acid molecule in a measurement sample; and
[0152] a classification process of classifying, based on a similarity between a representative sequence set as a representative of each of a plurality of groups obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule and the measurement sequence measured in the measurement process, the measurement sequence to each of the groups.
[0153] (Scheme 2)
[0154] The classification method according to Scheme 1, wherein the totalizing process totals the number of molecules of the determination sequence classified into each of the plurality of groups in the above classification process.
[0155] (Scheme 3)
[0156] The classification method according to Scheme 2, wherein the classification process is followed by a reclassification process of reclassifying the determination sequence classified into each of the plurality of groups into each of the base sequences constituting the group,
[0157] The totalizing process totals the number of molecules of the determination sequence reclassified into each of the base sequences for each of the base sequences.
[0158] (Scheme 4)
[0159] The classification method according to Scheme 2 or Scheme 3, wherein the classification process is followed by a reclassification process of reclassifying the determination sequence classified into each of the plurality of groups into each of the base sequences constituting the group,
[0160] (Scheme 5)
[0161] The classification method according to any one of Schemes 1 to 4, wherein the specific rule is a rule based on similarity of the base sequence grouped into the plurality of groups to the representative sequence, or similarity of the base sequence grouped into the plurality of groups to each other.
[0162] (Scheme 6)
[0163] The classification method according to any one of Schemes 1 to 5, wherein the representative sequence is set to a base sequence having the longest length among the base sequences constituting the group.
[0164] (Scheme 7)
[0165] The classification method according to any one of Schemes 1 to 6, wherein the representative sequence is set to a base sequence having the highest expression amount among the base sequences constituting the group.
[0166] (Scheme 8)
[0167] The classification method according to any one of Schemes 1 to 7, wherein, in the classification process, the determination sequence is classified into a group in which similarity of the representative sequence to the determination sequence is above a threshold value, or a group in which similarity of the representative sequence to the determination sequence is the highest.
[0168] (Scheme 9)
[0169] The classification method according to any one of the schemes 1 to 8, wherein the measurement process uses a next-generation sequencer to measure the base sequence of the nucleic acid molecule in the measurement sample as the measurement sequence.
[0170] (Scheme 10)
[0171] A classification device including a processor,
[0172] The processor classifies the measurement sequence into each of the groups based on similarity between a representative sequence set as a base sequence representing each of a plurality of groups and the measurement sequence measured by the measurement unit, the plurality of groups being obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule.
[0173] (Scheme 11)
[0174] A classification system including:
[0175] a measurement unit that measures a base sequence of a nucleic acid molecule in a measurement sample as a measurement sequence; and
[0176] a classification unit that classifies the measurement sequence into each of the groups based on similarity between a representative sequence set as a base sequence representing each of a plurality of groups and the measurement sequence measured by the measurement unit, the plurality of groups being obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule.
[0177] (Scheme 12)
[0178] A classification program for causing a computer to execute a classification process of classifying a measurement sequence into each of a plurality of groups based on similarity between a representative sequence set as a base sequence representing each of the groups and the measurement sequence measured, the plurality of groups being obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule.
[0179] (Scheme 13)
[0180] A non-transitory recording medium recording a classification program for causing a computer to execute a classification process of classifying a measurement sequence into each of a plurality of groups based on similarity between a representative sequence set as a base sequence representing each of the groups and the measurement sequence measured, the plurality of groups being obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule.
[0181] The disclosure of Japanese Patent Application No. 2023-108941 filed on June 30, 2023 is incorporated herein by reference in its entirety. All of the documents, patent applications and technical standards cited in the present specification are incorporated by reference to the same extent as if such individual documents, patent applications and technical standards were specifically and individually indicated to be incorporated by reference.
Claims
1. A classification method, comprising: a measurement step of measuring a base sequence of a nucleic acid molecule in a measurement sample as a measurement sequence; and a classification step of classifying the measurement sequence into each of a plurality of groups obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule, based on similarity of a representative sequence set as a base sequence representing each of the groups to the measurement sequence measured in the measurement step.
2. The classification method of claim 1, wherein, Further comprising a totalizing step of totalizing the number of molecules of the measurement sequence classified into each of the plurality of groups in the classification step.
3. The classification method of claim 2, wherein, Further comprising a reclassification step of reclassifying the measurement sequence classified into each of the plurality of groups after the classification step into each of the base sequences constituting the group, the totalizing step totalizing the number of molecules of the measurement sequence reclassified into each of the base sequences for each of the base sequences.
4. The classification method of claim 2, wherein, Further comprising a determination step of performing a specific determination of the measurement sample based on the number of molecules totalized by the totalizing step.
5. The classification method of claim 1, wherein, The specific rule is a rule based on similarity of the base sequence grouped into the plurality of groups to the representative sequence or similarity of the base sequences grouped into the plurality of groups to each other.
6. The classification method of claim 1, wherein, The representative sequence is set as a base sequence longest among the base sequences constituting the group.
7. The classification method of claim 1, wherein, The representative sequence is set as a base sequence highest in expression amount among the base sequences constituting the group.
8. The classification method of claim 1, wherein, In the classification step, the measurement sequence is classified into a group in which similarity of the representative sequence to the measurement sequence is above a threshold value or a group in which similarity of the representative sequence to the measurement sequence is highest.
9. The classification method of claim 1, wherein, The measurement step uses a next-generation sequencer to measure a base sequence of a nucleic acid molecule in the measurement sample as a measurement sequence.
10. A classification apparatus comprising a processor, the processor classifying a measurement sequence measured into each of a plurality of groups obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule, based on similarity of a representative sequence set as a base sequence representing each of the groups to the measurement sequence.
11. A classification system, comprising: a measurement unit that measures a base sequence of a nucleic acid molecule in a measurement sample as a measurement sequence; and a classification unit that classifies the measurement sequence into each of a plurality of groups obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule, based on similarity of a representative sequence set as a base sequence representing each of the groups to the measurement sequence measured by the measurement unit.
12. A classification program for causing a computer to execute a classification process of: classifying a measurement sequence measured into each of a plurality of groups obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule, based on similarity of a representative sequence set as a base sequence representing each of the groups to the measurement sequence.
13. A non-transitory recording medium recording a classification program for causing a computer to execute a classification process of: classifying a measured sequence into each of a plurality of groups obtained by grouping base sequences of a plurality of nucleic acid molecules by a specific rule, based on similarity of a representative sequence set as a representative of each of the groups to the measured sequence.
2. The method of claim 1, wherein the step of determining the similarity of the measured sequence to the representative sequence comprises: determining the similarity of the measured sequence to the representative sequence based on a Hamming distance between the measured sequence and the representative sequence.
3. The method of claim 1, wherein the step of determining the similarity of the measured sequence to the representative sequence comprises: determining the similarity of the measured sequence to the representative sequence based on a number of mismatches between the measured sequence and the representative sequence.
4. The method of claim 1, wherein the step of determining the similarity of the measured sequence to the representative sequence comprises: determining the similarity of the measured sequence to the representative sequence
Citation Information
Patent Citations
Oil damper system
JP2023108941A