Property determination model generation method, property determination model, property determination method, property determination device, program, and recording medium
By selecting microRNAs with stable expression levels and excluding those affected by total microRNA concentration changes, the method improves the accuracy of disease detection models, addressing inaccuracies in existing methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ARKRAY INC
- Filing Date
- 2025-10-21
- Publication Date
- 2026-05-07
Smart Images

Figure JP2025037031_07052026_PF_FP_ABST
Abstract
Description
Method for generating property determination model, property determination model, property determination method, property determination apparatus, program, and recording medium
[0001] The present disclosure relates to a method for generating a property determination model, a property determination model, a property determination method, a property determination apparatus, a program, and a recording medium.
[0002] In the blood of patients suffering from cancer diseases such as pancreatic cancer, it is known that the expression levels of specific multiple microRNAs among multiple types of small RNAs such as multiple microRNAs (miRNAs) increase or decrease.
[0003] Therefore, a blood sample of a subject to be determined whether suffering from a cancer disease is collected, the expression level of the microRNA that significantly increases or decreases due to the cancer disease in the collected blood sample is measured, and it is determined whether the subject is suffering from the cancer disease (for example, see Non-Patent Documents 1 and 2 below).
[0004] Non-Patent Document 1 discloses obtaining sample data including quantitative values of multiple types of microRNAs in blood samples collected from cancer patients such as esophageal cancer, gastric cancer, colorectal cancer, hepatobiliary cancer, pancreatic cancer, lung cancer, breast cancer, prostate cancer, etc., and blood samples collected from healthy subjects, and using 534 types of microRNAs that are stably expressed in both the blood samples of healthy subjects and cancer patients, and performing machine learning using the quantitative value data of microRNAs in the blood samples of healthy subjects without cancer diseases and cancer patients with cancer diseases as training data to generate a learned model capable of determining the presence or absence of cancer diseases, and a determination method for determining the presence or absence of cancer diseases of a subject using this learned model.
[0005] Further, Non-Patent Document 2 discloses obtaining sample data including quantitative values of multiple types of microRNAs in blood samples of lung cancer patients and healthy subjects, and using 181 types of microRNAs that are stably expressed, and performing machine learning using the quantitative value data of microRNAs in the blood samples of healthy subjects without diseases and blood samples of lung cancer patients as training data to generate a learned model capable of determining the presence or absence of diseases, and a determination method for determining the presence or absence of lung cancer diseases of a subject using this learned model.
[0006] Non-patent documents 1 and 2 both generate disease detection models using machine learning with microRNAs that are highly expressed in healthy individuals and disease patients, respectively. Generally, when generating disease detection models using machine learning, the training data used as input often includes features that are considered to have less impact from data noise in order to improve detection accuracy. Since microRNAs with low expression levels in biological samples such as blood collected from subjects are considered to be highly affected by data noise, it is conceivable to select microRNAs with high expression levels when generating disease detection models. By selecting microRNAs with high expression levels in human blood, it is possible to narrow it down to 100 to 200 types, and it was thought that a disease detection model with high detection accuracy could be generated by inputting such highly expressed microRNAs as features.
[0007] Non-patent Literature 1: Kuno Suzuki1, Hideyoshi Igata, Motoki Abe, Yusuke Yamamoto, "Multiple cancer type classification by small RNA expression profiles with plasma samples from multiple facilities", Cancer Science (Wiley Online library), June 19, 2022, Vol. 113, No. 6, pp. 2144-2166 Non-patent Literature 2: Masayasu Inagaki, Makoto Uchiyama, Kanae Yoshikawa-Kawabe, Masafumi Ito, Hideki Murakami, Masaharu Gunji, Makoto Minoshima, Takashi Kohnoh, Ryota Ito, Yuta Kodama, Mari Tanaka-Sakai, Atsushi Nakase1, Nozomi Goto1, Yusuke Tsushima, Shoich Mori, Masahiro Kozuka, Ryo Otomo, Mitsuharu Hirai, Masahiko Fujino, Toshihiko Yokoyama, "Comprehensive circulating microRNA profile as a supersensitive biomarker for early-stage lung cancer screening", Journal of Cancer Research and Clinical Oncology, [online], published April 19, 2023, Internet<URL: https: / / doi.org / 10.1007 / s00432-023-04728-9>
[0008] To measure the expression level of microRNAs in blood samples, quantitative analysis of microRNAs is performed using NGS (Next-Generation Sequencing) technology. In this NGS analysis, microRNAs derived from tissues or bodily fluids are obtained and used to create a library. However, even when library creation is performed using the same protocol, differences in microRNA quantification can occur between samples (between individuals). Furthermore, the efficiency of microRNA extraction is not constant, and even with identical samples, variations in the total concentration of microRNA obtained and differences between reagent lots can occur. Moreover, even for microRNAs with high expression levels in blood, which are generally considered to have less data noise, there are differences in expression levels due to individual differences.
[0009] When interpreting measurement data obtained from samples with varying total microRNA concentrations, the obtained measurement data is normalized. However, even after normalizing the measurement data, it may not be possible to eliminate the influence of variations in the total microRNA concentration in the sample. Therefore, if a trained model is generated using microRNAs selected solely based on high expression levels, and the presence or absence of disease is determined using the generated disease detection model, there is a possibility that some subjects may not be correctly diagnosed as having the disease.
[0010] This disclosure provides a method for generating a property determination model, a property determination model, a property determination method, a property determination device, a program, and a recording medium that can generate a property determination model with higher determination accuracy when generating a property determination model that determines the presence or absence of a property of a subject by performing machine learning based on the quantitative values of multiple small RNAs in a biological sample taken from the subject.
[0011] A method for generating a property determination model according to one aspect of the present disclosure is a method for generating a property determination model that is generated by machine learning and determines the presence or absence of a property of a subject based on the quantitative values of multiple small RNAs in a biological sample taken from the subject, wherein multiple small RNAs are detected from biological samples taken from multiple learning subjects consisting of learning subjects having the property and learning subjects not having the property, and multiple small RNAs are selected as standard small RNAs to be used when generating the property determination model, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range, from among the multiple small RNAs detected from biological samples taken from multiple learning subjects consisting of learning subjects having the property and learning subjects not having the property, and multiple small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range, and these multiple small RNAs are selected as standard small RNAs to be used when generating the property determination model, and learning small RNAs to be used for machine learning of the property determination model are selected from the selected standard small RNAs based on certain conditions, and the property determination model is generated by performing machine learning using the presence or absence of the property in the multiple learning subjects and the respective quantitative values of the multiple small RNAs selected as learning small RNAs as training data.
[0012] According to this disclosure, by performing machine learning based on the quantitative values of multiple small RNAs in a biological sample collected from a subject, it becomes possible to generate a property determination model that determines whether or not a subject possesses certain properties, thereby generating a property determination model with higher determination accuracy.
[0013] This is a diagram showing the system configuration of a learning model generation system 20 according to one embodiment of the present disclosure. This is a block diagram showing the hardware configuration of a learning model generation device 30. This is a block diagram showing the functional configuration of a learning model generation device 30 realized by the execution of a generation program. This is a flowchart showing a method for generating a disease judgment model 130. This is a diagram showing the machine learning process of the disease judgment model 130. This is a block diagram showing the functional configuration of a disease judgment device 50 for performing disease judgment using the generated disease judgment model 130. This is a flowchart showing the process for determining whether or not a person has cancer using a disease judgment model 130 that has undergone machine learning. This is a diagram showing the process for determining whether or not a person has cancer. This is a diagram showing the comparison results of quantitative values when one sample #1 is library-formed in two ways: undiluted and 4-fold diluted. This is a diagram showing the absolute value of the difference in quantitative values when the top 50 miRNAs of sample #1 are library-formed in two ways: undiluted and 4-fold diluted. This is a diagram showing a list of the top 10 miRNAs among the top 100 miRNAs in quantitative value for which the difference due to dilution is reproducible (small p-value). Figure 11 is a graph showing the average difference in quantitative values for the 10 types of miRNAs shown. This figure lists the top 10 miRNAs with reproducible differences (small p-values) under dilution, as determined by the library creation protocol "Qiaseq". Figure 13 is a graph showing the average difference in quantitative values for the 10 types of miRNAs shown.
[0014] Next, embodiments of the present disclosure will be described in detail with reference to the drawings.
[0015] An example of an embodiment relating to the technology of this disclosure will be described below with reference to the drawings. Components and processes that perform the same operation, action, or function will be given the same reference numerals throughout the drawings, and redundant explanations may be omitted as appropriate. Each drawing is only a schematic representation to the extent that the technology of this disclosure can be fully understood. Therefore, the technology of this disclosure is not limited to the illustrated examples. Furthermore, in this embodiment, explanations of configurations not directly related to this disclosure or well-known configurations may be omitted.
[0016] Figure 1 shows the system configuration of a learning model generation system 20 according to one embodiment of the present disclosure.
[0017] In the learning model generation system 20 of this embodiment, a disease determination model is generated by the disease determination model generation method described below.
[0018] The disease detection model generation method of this embodiment generates a trained disease detection model that determines the presence or absence of a disease in a subject based on the quantitative values of multiple small RNAs. In this method, from among the multiple small RNAs detected in a biological sample such as the subject's blood, only small RNAs whose quantitative value changes when the total concentration of small RNAs in the biological sample changes exceed a predetermined threshold range are selected as standard small RNAs to be used when generating the disease detection model. The disease detection model is then generated by performing machine learning using the selected training small RNAs as training data.
[0019] As shown in Figure 1, the learning model generation system 20 of this embodiment consists of a next-generation sequencer 21 using NGS (Next-Generation Sequencing) technology and a learning model generation device 30.
[0020] In addition to the next-generation sequencer 21, quantitative PCR (Polymerase Chain Reaction) and flow cytometers can also be used as measurement devices, as long as they can measure the expression levels of multiple nucleic acid molecules.
[0021] The learning model generation system 20 of this embodiment generates a disease determination model capable of determining whether or not a subject has cancer, using blood 40 collected from the subject whose presence or absence of cancer is to be determined.
[0022] In this embodiment, the method for determining the presence or absence of cancer is described using blood collected from a subject. However, the presence or absence of cancer may also be determined using other biological samples, such as body fluids, cells, extracellular vesicles, or tissue fragments. Here, body fluids include, for example, serum, urine, tears, saliva, sweat, semen, lymph, tissue fluid, body cavity fluids (e.g., pleural fluid, ascites), cerebrospinal fluid, amniotic fluid, vaginal fluid, nasal mucus, etc. Cells include, for example, red blood cells, white blood cells, platelets, oral swabs, etc. Extracellular vesicles include, for example, exosomes, liposomes, etc. Tissue fragments include, for example, FFPE (Formalin Fixed Paraffin Embedded) specimens, biopsy specimens, frozen specimens, etc.
[0023] Furthermore, although this embodiment describes the case where the subject of the examination is a human, the subject of the examination is not limited to humans, and this disclosure is equally applicable when the subject of the examination is various animals other than humans, such as dogs and cats. In other words, the technology of this disclosure is also applicable when examining the presence or absence of cancer or other diseases using biological samples such as blood collected from various animals such as dogs and cats.
[0024] Furthermore, in this embodiment, the method for generating a disease determination model is described using the example of generating a disease determination model to determine whether or not a subject has cancer, but this disclosure is not limited to such cases. This disclosure is also applicable when generating a disease determination model to determine whether or not a subject has diseases other than cancer, or when generating a property determination model to determine specific properties of a subject. The properties of a subject to be determined by the property determination model are not limited to the subject's disease. The properties of a subject to be determined can be any properties that can be determined based on the expression levels of multiple small RNAs. Specific examples of properties to be determined include, for example, confirming the efficacy of a subject's medication, determining the likelihood of disease recurrence, confirming / determining lifestyle habits such as smoking and drinking history, and predicting biological age.
[0025] The next-generation sequencer 21 measures the expression levels of multiple small RNAs contained in the blood 40 collected from the subject. Specifically, the next-generation sequencer 21 amplifies the small RNAs extracted from the blood 40 and then measures the relative expression levels of each type of small RNA in relation to other small RNAs.
[0026] Here, small RNAs include, for example, microRNAs (hereinafter referred to as miRNAs). Other small RNAs besides miRNAs (e.g., piRNAs and tsRNAs) may also be used. The following explanation will use the case where the expression level of miRNAs is measured from a subject's blood 40 to determine whether or not they have cancer.
[0027] The next-generation sequencer 21 then outputs the measured miRNA expression level data as miRNA data to the learning model generation device 30.
[0028] The learning model generation device 30 uses miRNA data from the next-generation sequencer 21 and information on whether or not each subject has cancer to generate a disease determination model for determining whether or not a subject has cancer.
[0029] When quantitatively analyzing miRNA using NGS technology, biological samples derived from tissues or bodily fluids are obtained, and these biological samples are then compiled into a library. However, even when library compilation is performed using the same protocol, the total concentration of miRNA in biological samples will vary between blood samples (between individuals). Furthermore, the efficiency of miRNA extraction is not constant, and even with identical blood samples, variations in the total concentration of miRNA obtained, as well as differences between reagent lots, can occur. Consequently, the total concentration of miRNA used in the blood sample may affect the analysis results.
[0030] For example, an impact on sequence depth is anticipated. That is, when measurements are performed on blood samples with low total concentrations, a decrease in sequence depth (reading depth or volume) is expected. Therefore, when analyzing measurement data obtained from blood samples with varying total miRNA concentrations, normalization of the measurement data is necessary. Known normalization methods include those based on expression standards (e.g., housekeeping miRNAs or spike-in controls) and RPM (Read Per Million) correction based on the total number of reads per sample.
[0031] However, variability in the total miRNA concentration of blood samples can lead to problems that cannot be solved by normalization alone. This is the bias (influence) on library preparation. Using blood samples with different total miRNA concentrations can affect the efficiency of adapter ligation and reverse transcription during the library preparation process. Furthermore, this bias may differ depending on the type (sequence) of miRNA. In other words, even with the same miRNA, differences in total concentration can result in different measurement data.
[0032] Furthermore, variations in the total concentration of miRNAs in blood samples can lead to bias in library preparation, potentially causing the following problems: (1) The measurement data may contain noise due to the bias, degrading its quality and potentially canceling out differences in miRNA expression levels that should be detected. (2) When performing machine learning using the acquired measurement data, the presence of noise in the data may lead to incorrect learning. As a result, the accuracy of the resulting learning model may be reduced.
[0033] To solve the above problems, it is necessary to standardize the extraction and quality assessment of miRNA and use blood samples with the same total concentration as much as possible. However, there are the following problems with standardizing the total miRNA concentration among blood samples: (1) It is time-consuming and costly because it adds an extra step of standardizing the total miRNA concentration. There is also the possibility of human error. (2) When standardizing the total miRNA concentration, it is basically necessary to standardize the total miRNA concentration of other blood samples to that of the blood sample with the lowest total miRNA concentration. However, in biological samples derived from bodily fluids such as blood samples, the total miRNA concentration is inherently low, and in some cases the total concentration is close to the detection limit. In such cases, if the total miRNA concentration of all blood samples is standardized to that of the blood sample with the lowest total concentration, the overall quality of the measurement data will decrease. (3) In biological samples derived from bodily fluids such as blood samples, the miRNA concentration is inherently low, and in some cases it is not possible to measure the total concentration at all.
[0034] For the reasons stated above, in reality, many previous studies that performed NGS measurements of miRNA using blood samples such as serum and plasma have created libraries without standardizing the total concentration of miRNA.
[0035] Therefore, in the learning model generation system 20 of this embodiment, as will be explained in detail below, miRNAs whose quantitative values obtained by changes in the total concentration of miRNAs in a sample such as a blood sample fluctuate greatly are excluded from the learning miRNAs used to generate the learning model, thereby generating a disease diagnosis model with higher judgment accuracy.
[0036] Next, the hardware configuration of the learning model generation device 30 described above is shown in the block diagram in Figure 2.
[0037] The learning model generation device 30 has computer-like functionality and, as shown in Figure 2, includes a CPU (Central Processing Unit) 31, ROM (Read Only Memory) 32, RAM (Random Access Memory) 33, storage 34, input unit 35, display unit 36, and communication interface (I / F) 37. Each component is connected to the others via a bus 39 so as to be able to communicate with each other.
[0038] The CPU 31 (an example of a processor) is a central processing unit that executes various programs and controls various parts. Specifically, the CPU 31 reads a program from the ROM 32 or storage 34 and executes the program using the RAM 33 as a working area. The CPU 31 controls each of the above components and performs various calculations according to the program stored in the ROM 32 or storage 34.
[0039] ROM 32 stores various programs and data. RAM 33 temporarily stores programs or data as a working area. Storage 34 consists of an HDD (Hard Disk Drive) or SSD (Solid State Drive) and stores various programs, including the operating system, and various data.
[0040] In this embodiment, for example, a generation program for generating a disease diagnosis model is recorded in the storage 34. This generation program may be a single program or a group of programs consisting of multiple programs or modules. The generation program may also be recorded in the ROM 32. The ROM 32 and the storage 34 function as examples of non-temporary recording media.
[0041] An example of a processor is not limited to the general-purpose CPU mentioned above, but could also be a dedicated processor consisting of circuits specifically designed to perform a particular task. Furthermore, an example of a processor is not limited to a single unit, but could also be a system in which multiple units located in physically separate locations cooperate to perform the task.
[0042] The input unit 35 includes a pointing device such as a mouse and a keyboard, and is used to perform various inputs. Further, the input unit 35 receives, as an input, miRNA data representing the expression levels of a plurality of miRNAs measured by the next-generation sequencer 21.
[0043] The display unit 36 is, for example, a liquid crystal display, and displays various information. Further, the display unit 36 may adopt a touch panel method and function as an input unit 35 for inputting operations from the user.
[0044] The communication interface 37 is an interface for communicating with other devices, and performs data transmission and reception with an external device including the next-generation sequencer 21.
[0045] Next, the functional configuration of the learning model generation device 30 realized by executing the generation program described above is shown in the block diagram of FIG. 3.
[0046] As shown in FIG. 3, in the learning model generation device 30, the CPU 31 executes a determination program, thereby constituting an acquisition unit 110, a generation unit 120, and a disease determination model 130.
[0047] The acquisition unit 110 acquires miRNA data (hereinafter sometimes referred to as learning miRNA), which is the result of measuring the expression level of each of a plurality of miRNAs in the blood of a learning subject by measuring the blood of the learning subject.
[0048] Note that the expression levels of a plurality of miRNAs in the blood are absolute values derived from a living body, but since it is necessary to quantify the expression levels in the process of passing through a measuring device and reagent treatment, it is difficult to quantify the expression levels of miRNAs in the blood as absolute values. That is, the miRNA data relatively shows the expression level of miRNA by the next-generation sequencer 21 or data processing. However, if the expression level of miRNA in the blood can be quantified as an absolute value, the expression level of miRNA in the blood as an absolute value may be used as miRNA data.
[0049] The generation unit 120 generates a disease determination model 130 that determines whether or not a subject has cancer, based on the miRNA data (training miRNA) acquired by the acquisition unit 110 and information indicating whether or not each training subject has cancer, by performing machine learning. The disease determination model 130 generated by the generation unit 120 is then stored, for example, in a non-volatile storage device.
[0050] Next, the method for generating the disease determination model 130 described above is shown in the flowchart of Figure 4.
[0051] First, in step S101, the blood collection process involves collecting blood from the learning subjects, centrifuging it, and then dispensing the resulting serum into storage tubes in predetermined quantities, which are then stored in a deep freezer at -80°C. These learning subjects include both healthy individuals without cancer and patients with cancer.
[0052] Next, in step S102, the measurement process involves removing the frozen serum from the deep freezer, thawing it at room temperature, and then measuring the miRNA expression level using a next-generation sequencer 21.
[0053] Then, in the acquisition step S103, the learning model generation device 30 acquires the results of measuring the expression levels of multiple miRNAs in the serum of the learning subject, which were measured in the measurement step S102.
[0054] Next, in the selection step S104, the learning model generation device 30 selects several miRNAs from among several miRNAs detected in the blood collected from multiple learning subjects consisting of learning subjects with cancer and learning subjects without cancer, excluding miRNAs whose change in quantitative value when the total concentration of miRNAs in the blood sample changes exceeds a preset threshold range, to be used as standard miRNAs when generating the disease judgment model 130.
[0055] Here, miRNAs whose quantitative value changes when the total concentration of miRNAs in a blood sample changes, exceeding a predetermined threshold range, are, for example, miRNAs whose quantitative value difference in multiple blood samples obtained by diluting the same blood sample at different dilution ratios is greater than or equal to a predetermined threshold.
[0056] Furthermore, small RNAs whose quantitative value changes when the total concentration of miRNAs in a blood sample changes, exceeding a predetermined threshold range, are defined, for example, miRNAs that are within a predetermined rank when multiple blood samples, each diluted with different dilution ratios from the same blood sample, are ranked in descending order of the difference in quantitative values.
[0057] Furthermore, small RNAs whose quantitative value changes when the total concentration of miRNAs in a blood sample changes, exceeding a predetermined threshold range, are, for example, miRNAs whose rate of change in quantitative value exceeds a predetermined threshold in multiple blood samples obtained by diluting the same blood sample at different dilution ratios.
[0058] Then, in the selection step S105, the learning model generation device 30 selects from the selected standard miRNAs to be used for machine learning of the disease diagnosis model based on certain conditions. Specifically, the learning model generation device 30 selects from the selected standard miRNAs, within a predetermined rank from the top when sorted in descending order of quantitative values, for example, within the top 100, as the learning miRNAs to be used for machine learning of the disease diagnosis model.
[0059] Finally, in the learning process of step S106, the learning model generation device 30 uses the acquired miRNA data and information indicating whether or not each training subject has cancer as training data to perform machine learning on the disease judgment model 130. Figure 5 shows the machine learning process of the disease judgment model 130.
[0060] As can be seen by referring to Figure 5, in the disease determination system 20 of this embodiment, the machine learning of the disease determination model 130 is performed by linking the expression level data of learning miRNAs collected from the blood of healthy individuals among the learning subjects who do not have cancer, and the expression level data of learning miRNAs collected from the blood of patients with cancer among the learning subjects, with information on whether or not the learning subjects have cancer.
[0061] Thus, the generation unit 120 generates a disease judgment model 130 by performing machine learning using the presence or absence of cancer in multiple training subjects and the expression levels of multiple miRNAs selected as training miRNAs as training data. The disease judgment model 130 is a learning model generated by performing machine learning using the presence or absence of cancer in multiple training subjects and the expression levels of multiple miRNAs selected as training miRNAs as training data. A specific example of this learning model is described below.
[0062] <Examples of Learning Models> Various linear and nonlinear algorithms known as machine learning algorithms, or combinations of multiple algorithms, can be used. For example, the following algorithms can be used.
[0063] Random forest, Gradient boosting decision trees, Extreme gradient boosting decision trees, Light-gradient boosting machine, Neural networks, Regularized regression, Elastic-net regression, K-Nearest neighbors, Support vector machine, Generalized additive model
[0064] Next, we will explain the method for determining whether or not a subject under evaluation has cancer, using the disease determination model 130 generated by the generation method described above.
[0065] The functional configuration of the disease diagnosis device 50 for performing such disease diagnosis is shown in the block diagram in Figure 6.
[0066] As shown in Figure 6, the disease determination device 50 is composed of an acquisition unit 110, a disease determination model 130, and a determination unit 140. Here, the disease determination model 130 in Figure 6 is a trained disease determination model 130 generated by the learning model generation device 30 shown in Figure 3.
[0067] The acquisition unit 110 acquires miRNA data, which is the result of measuring the expression levels of multiple miRNAs in the blood of a subject being assessed for the determination of whether or not they have cancer, by measuring the blood of the subject being assessed.
[0068] The determination unit 140 then inputs miRNA data, which shows the results of measuring the expression levels of multiple miRNAs in the blood collected from the subject to be determined, into the disease determination model 130, thereby determining whether or not the subject has cancer.
[0069] Thus, the disease determination device 50 has a disease determination model 130, and by inputting miRNA data showing the results of measuring the expression levels of multiple miRNAs in the blood collected from the subject into this disease determination model 130, it functions as a property determination device that determines whether or not the subject has cancer.
[0070] The disease determination device 50 then outputs the determination result from the determination unit 140 to an external device or displays it on the display unit.
[0071] Next, Figure 7 shows a flowchart illustrating the process of determining whether or not a patient has cancer using the disease determination model 130, which has been machine-learned in this manner.
[0072] First, in step S201, the sampling process, blood is collected from the subject to determine whether or not they have cancer.
[0073] Next, in step S202, the measurement process involves measuring the expression level of miRNA in the blood sample collected from the subject using a next-generation sequencer 21.
[0074] Then, in the acquisition step S203, the disease determination device 50 acquires the results of measuring the expression levels of multiple miRNAs in the subject's blood sample, which were measured in the measurement step S202, as miRNA data.
[0075] Finally, in step S204, the disease presence / absence determination process, the disease determination device 50 inputs the acquired miRNA data into the trained disease determination model 130 to obtain a determination result indicating whether or not the subject has cancer. Figure 8 shows this process for determining the presence or absence of cancer.
[0076] As can be seen by referring to Figure 8, in the disease determination system of this embodiment, the expression level data of miRNAs taken from a blood sample of a subject to be determined to have cancer is input into a trained disease determination model 130 to obtain the determination result.
[0077] Furthermore, according to the disease determination model generation method of this embodiment, when generating a disease determination model that determines whether or not a subject has cancer by performing machine learning based on the quantitative values of multiple miRNAs in a blood sample taken from the subject, multiple miRNAs are selected as standard miRNAs to be used when generating the disease determination model, excluding miRNAs whose change in quantitative value when the total concentration of miRNAs in the blood sample changes exceeds a predetermined threshold range.
[0078] Therefore, miRNAs that are significantly affected by changes in the total concentration of miRNAs in the blood sample are excluded from the standard miRNAs. As a result, by selecting training miRNAs from these standard miRNAs and performing machine learning on the disease diagnosis model 130 using the selected training miRNAs as training data, it is possible to realize a disease diagnosis model with improved diagnosis accuracy.
[0079] Next, as an example, experimental data will show that generating a disease diagnosis model using the generation method in this embodiment results in a disease diagnosis model with higher accuracy than generating a disease diagnosis model using miRNAs that are highly expressed in the subject's blood.
[0080] (1) In this experiment, human serum-derived RNA was obtained from 20 individuals (samples #1 to #20). (2) For each RNA, a miRNA library was created using both the original solution and a 4-fold dilution (solvent: water). The library creation was performed according to "Eminaga et al., Quantification of microRNA Expression with Next-Generation Sequencing". (3) The generated libraries were then sequenced using the next-generation sequencer 21 or other equipment mentioned above to obtain measurement data showing the expression levels of multiple types of miRNA. (4) The obtained measurement data was subjected to RPM normalization before analysis.
[0081] First, Figure 9 shows a comparison of the quantitative values obtained when one sample, #1, was library-prepared in two ways: as a stock solution and as a 4-fold dilution. Note that in Figure 9, for the sake of clarity, the quantitative values of the top 20 miRNAs are shown.
[0082] Referring to Figure 9, it can be seen that even though it is the same miRNA, the measurement results by the NGS instrument change simply by diluting it fourfold with water. In other words, even a slight change in the efficiency of miRNA extraction can potentially change the measurement results by the NGS instrument.
[0083] Next, Figure 10 shows the absolute difference between the logarithmic RPM values (log2RPM) of the top 50 miRNAs in Sample #1 compared to the original solution and the 4-fold dilution. In Figure 10, the miRNAs are sorted in descending order of the absolute difference.
[0084] Referring to Figure 10, it can be seen that there are miRNAs that are greatly affected by the total concentration and miRNAs that are hardly affected by the total concentration.
[0085] Figures 9 and 10 show measurement results based on a single sample, #1, so it cannot be ruled out that the measurement data for the undiluted solution and the 4-fold dilution may simply be randomly scattered. Therefore, in the following section, we analyzed whether or not a difference in total concentration occurred in all 20 samples, #1 to #20.
[0086] Figure 11, described below, is a list of the top 10 miRNAs among the top 100 in quantitative value that show reproducible differences (small p-values) with respect to dilution. Referring to Figure 11, it can be seen that the p-values are sufficiently small, and that differences occur between the original solution and the 4-fold dilution in all samples #1 to #20.
[0087] Figure 12 shows a graph of the average difference in quantitative values for the 10 types of miRNAs shown in Figure 11.
[0088] The miRNAs shown in Figures 11 and 12 are particularly susceptible to differences in total concentration affecting the expression level measurements in the protocol of "Eminaga et al., Quantification of microRNA Expression with Next-Generation Sequencing," and can be said to be miRNAs that should be excluded from analysis when generating learning models or determining the presence or absence of disease.
[0089] Finally, Figures 13 and 14 show the measurement results when the same investigation as above was performed using the "Qiaseq miRNA Library Kit (QIAgen)" reagent to create a library. Figures 13 and 14 show the difference in quantitative values when libraries were created using the undiluted solution and a 4-fold dilution for 10 samples.
[0090] As can be seen by comparing Figures 11 and 12 with Figures 13 and 14, the types of miRNAs affected by differences in total concentration vary depending on the reagents used during library creation. This is thought to be due to differences in the types of ligation enzymes and reverse transcriptases used, which cause changes in bias. In other words, it is difficult to establish universal rules, and it is important to identify in advance the types of miRNAs that are greatly affected by total concentration for each method of acquiring measurement data, and to exclude them from the analysis when generating learning models or determining the presence or absence of disease.
[0091] The disclosure of Japanese Patent Application No. 2024-188788, filed on 28 October 2024, is incorporated herein by reference in its entirety. All documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0092] [Note] Preferred embodiments of this disclosure are noted below. (Embodiment 1) A method for generating a property determination model that determines the presence or absence of a subject's properties based on quantitative values of a plurality of small RNAs in a biological sample collected from a subject, which is generated by machine learning, comprising: selecting a plurality of small RNAs as standard small RNAs to be used when generating the property determination model, from a plurality of small RNAs detected in biological samples collected from a plurality of learning subjects consisting of learning subjects having the properties and learning subjects not having the properties, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range; selecting a learning small RNA to be used for machine learning of the property determination model from the selected standard small RNAs based on certain conditions; and generating the property determination model by performing machine learning using the presence or absence of the properties in the plurality of learning subjects and the respective quantitative values of the plurality of small RNAs selected as learning small RNAs as training data.
[0093] (Aspect 2) A method for generating a property determination model according to Aspect 1, wherein the small RNA whose quantitative value changes when the total concentration of small RNA in the biological sample changes, and which exceeds a predetermined threshold range, is the small RNA whose quantitative value difference in multiple biological samples obtained by diluting the same biological sample at different dilution ratios is greater than or equal to a predetermined threshold.
[0094] (Aspect 3) A method for generating a property determination model according to Aspect 1, wherein the small RNAs whose quantitative value changes when the total concentration of small RNA in the biological sample changes exceed a predetermined threshold range are small RNAs that are ranked in order of the greatest difference in quantitative values among multiple biological samples obtained by diluting the same biological sample at different dilution ratios, and are ranked from the top to the lowest.
[0095] (Aspect 4) A method for generating a property determination model according to Aspect 1, wherein the small RNA whose quantitative value changes when the total concentration of small RNA in the biological sample changes, and which exceeds a predetermined threshold range, is the small RNA whose rate of change in quantitative value is greater than or equal to a predetermined threshold in multiple biological samples obtained by diluting the same biological sample at different dilution ratios.
[0096] (Aspect 5) A method for generating a property determination model according to any one of aspects 1 to 4, wherein when selecting a training small RNA to be used for machine learning of the property determination model from among the selected standard small RNAs based on certain conditions, the small RNAs that are within a predetermined rank when sorted in descending order of quantitative value from among the selected standard small RNAs are selected as the training small RNAs to be used for machine learning of the property determination model.
[0097] (Aspect 6) A method for generating a property determination model according to any one of aspects 1 to 5, wherein the biological sample is blood and the property is the presence or absence of cancer.
[0098] (Aspect 7) A method for generating a property determination model according to any one of aspects 1 to 6, wherein the small RNA is a microRNA.
[0099] (Aspect 8) A property determination model that is generated by machine learning and determines the presence or absence of a subject's properties based on the quantitative values of multiple small RNAs in a biological sample taken from the subject, wherein multiple small RNAs are selected as standard small RNAs to be used when generating the property determination model, from among multiple small RNAs detected in biological samples taken from multiple learning subjects consisting of learning subjects having the properties and learning subjects not having the properties, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range, from among the multiple small RNAs selected as standard small RNAs, and learning small RNAs to be used for machine learning of the property determination model from among the standard small RNAs selected based on certain conditions, and the property determination model is generated by performing machine learning using the presence or absence of the properties in the multiple learning subjects and the respective quantitative values of the multiple small RNAs selected as learning small RNAs as training data.
[0100] (Aspect 9) A property determination method for determining whether a subject possesses the aforementioned property by inputting small RNA data, which shows the results of measuring the expression levels of multiple small RNAs in a biological sample taken from the subject, into the property determination model described in Aspect 8.
[0101] (Aspect 10) A property determination device having the property determination model described in Aspect 8, wherein the presence or absence of the said property of a subject is determined by inputting small RNA data, which shows the results of measuring the expression levels of multiple small RNAs in a biological sample taken from the subject, into the property determination model.
[0102] (Aspect 11) A program for generating a property determination model that determines the presence or absence of a subject's properties based on quantitative values of multiple small RNAs in a biological sample collected from a subject, which is generated by machine learning, comprising the steps of: selecting multiple small RNAs as standard small RNAs to be used when generating the property determination model, from among multiple small RNAs detected in biological samples collected from multiple learning subjects consisting of learning subjects having the properties and learning subjects not having the properties, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range; selecting learning small RNAs to be used for machine learning of the property determination model from among the selected standard small RNAs based on certain conditions; and generating the property determination model by performing machine learning using the presence or absence of the properties in the multiple learning subjects and the respective quantitative values of the multiple small RNAs selected as learning small RNAs as training data.
[0103] (Aspect 12) A non-temporary recording medium on which a program is recorded for generating a property determination model that is generated by machine learning and determines the presence or absence of a subject's properties based on the quantitative values of multiple small RNAs in a biological sample taken from a subject, the program comprising: selecting multiple small RNAs from multiple small RNAs detected in biological samples taken from multiple learning subjects consisting of learning subjects having the properties and learning subjects not having the properties, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a preset threshold range, as standard small RNAs to be used when generating the property determination model; selecting learning small RNAs from the selected standard small RNAs to be used for machine learning of the property determination model based on certain conditions; and generating the property determination model by performing machine learning using the presence or absence of the properties in the multiple learning subjects and the respective quantitative values of the multiple small RNAs selected as learning small RNAs as training data; and a program is recorded on which a computer is instructed to perform the following steps.
Claims
1. A method for generating a property determination model that is generated by machine learning and determines the presence or absence of a subject's properties based on the quantitative values of multiple small RNAs in a biological sample taken from the subject, comprising: selecting multiple small RNAs as standard small RNAs to be used when generating the property determination model, from among multiple small RNAs detected in biological samples taken from multiple learning subjects consisting of learning subjects having the properties and learning subjects not having the properties, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range; selecting learning small RNAs to be used for machine learning of the property determination model from among the selected standard small RNAs based on certain conditions; and generating the property determination model by performing machine learning using the presence or absence of the properties in the multiple learning subjects and the respective quantitative values of the multiple small RNAs selected as learning small RNAs as training data.
2. A method for generating a property determination model according to claim 1, wherein the small RNA whose quantitative value changes when the total concentration of small RNA in the biological sample changes, and which exceeds a predetermined threshold range, is the small RNA whose quantitative value difference in multiple biological samples obtained by diluting the same biological sample at different dilution ratios is greater than or equal to a predetermined threshold.
3. A method for generating a property determination model according to claim 1, wherein the small RNAs whose quantitative value changes when the total concentration of small RNAs in the biological sample changes, and which exceed a predetermined threshold range, are small RNAs that are ranked within a predetermined rank from the top when the differences in quantitative values are greatest among multiple biological samples obtained by diluting the same biological sample at different dilution ratios.
4. A method for generating a property determination model according to claim 1, wherein the small RNA whose quantitative value changes when the total concentration of small RNA in the biological sample changes, and whose quantitative value change rate is above a predetermined threshold range, is the small RNA whose quantitative value change rate is above a predetermined threshold in multiple biological samples obtained by diluting the same biological sample at different dilution ratios.
5. A method for generating a property determination model according to claim 1, wherein, when selecting small RNAs to be used for machine learning of the property determination model from among the selected standard small RNAs based on certain conditions, small RNAs that are within a predetermined rank when sorted in descending order of quantitative value from among the selected standard small RNAs are selected as small RNAs to be used for machine learning of the property determination model.
6. A method for generating a property determination model according to claim 1, wherein the biological sample is blood, and the property is the presence or absence of cancer.
7. The method for generating a property determination model according to claim 1, wherein the small RNA is a microRNA.
8. A property determination model that is generated by machine learning and determines the presence or absence of a subject's properties based on the quantitative values of multiple small RNAs in a biological sample taken from the subject, wherein multiple small RNAs are selected as standard small RNAs to be used when generating the property determination model, from among multiple small RNAs detected in biological samples taken from multiple learning subjects consisting of learning subjects having the properties and learning subjects not having the properties, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range, from among the multiple small RNAs selected as standard small RNAs, and learning small RNAs to be used for machine learning of the property determination model based on certain conditions, and the property determination model is generated by performing machine learning using the presence or absence of the properties in the multiple learning subjects and the respective quantitative values of the multiple small RNAs selected as learning small RNAs as training data.
9. A method for determining whether a subject possesses the aforementioned properties by inputting small RNA data, which shows the results of measuring the expression levels of multiple small RNAs in a biological sample taken from the subject, into the property determination model described in claim 8.
10. A property determination device having the property determination model described in claim 8, wherein the presence or absence of the said property of a subject is determined by inputting small RNA data, which shows the results of measuring the expression levels of multiple small RNAs in a biological sample taken from the subject, into the property determination model.
11. A program for generating a property determination model that determines the presence or absence of a subject's properties based on quantitative values of multiple small RNAs in a biological sample collected from a subject, which is generated by machine learning, comprising the steps of: selecting multiple small RNAs as standard small RNAs to be used when generating the property determination model, from among multiple small RNAs detected in biological samples collected from multiple learning subjects consisting of learning subjects having the properties and learning subjects not having the properties, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range; selecting learning small RNAs to be used for machine learning of the property determination model from among the selected standard small RNAs based on certain conditions; and generating the property determination model by performing machine learning using the presence or absence of the properties in the multiple learning subjects and the respective quantitative values of the multiple small RNAs selected as learning small RNAs as training data.
12. A non-temporary recording medium on which a program is recorded for generating a property determination model that is generated by machine learning and determines the presence or absence of a subject's properties based on the quantitative values of multiple small RNAs in a biological sample taken from a subject, the program comprising: selecting multiple small RNAs as standard small RNAs to be used when generating the property determination model, from among multiple small RNAs detected in biological samples taken from multiple learning subjects consisting of learning subjects having the properties and learning subjects not having the properties, excluding small RNAs whose change in quantitative value when the total concentration of small RNAs in the biological sample changes exceeds a predetermined threshold range; selecting learning small RNAs to be used for machine learning of the property determination model from among the selected standard small RNAs based on certain conditions; and generating the property determination model by performing machine learning using the presence or absence of the properties in the multiple learning subjects and the respective quantitative values of the multiple small RNAs selected as learning small RNAs as training data.
Citation Information
Patent Citations
Salivary mRNA profiling, biomarkers and related methods and kits
JP2007522819A
Ultrasound-sensitive detection of circulating tumor DNA by genome-wide integration
JP2021519607A
Machine learning implementation for multi-analyte assays of biological samples
JP2021521536A
Methods for analyzing cell-free RNA
JP2023534094A
Disease development determination device, disease development determination method, and disease development determination program
WO2018079840A1