Property determination model generation method, property determination model, property determination method, property determination device, program, and recording medium

By selecting microRNAs with normal distribution patterns or low variation for training, the disease determination model achieves improved accuracy in cancer diagnosis by minimizing individual expression level discrepancies.

WO2025253811A1PCT designated stage Publication Date: 2025-12-11ARKRAY INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/015947
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-06
Filing Date
2025-04-24
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing disease determination models using microRNAs for cancer diagnosis are prone to inaccuracies due to variations in microRNA expression levels among individuals, leading to incorrect diagnoses.

Method used

Selecting microRNAs with normal distribution patterns or low variation in expression levels across individuals for training data to generate a disease determination model through machine learning, excluding those with significant individual variations.

Benefits of technology

This approach enhances the accuracy of disease determination models by reducing the impact of individual variations in microRNA expression, resulting in more precise cancer diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025015947_11122025_PF_FP_ABST
    Figure JP2025015947_11122025_PF_FP_ABST
Patent Text Reader

Abstract

According to the present invention, from among a plurality of small RNAs detected from biological samples collected from a plurality of learning subjects having or not having a certain property, a plurality of small RNAs with which the shape of the distribution of the expression levels satisfies a predetermined specific condition or with which the variation in the distribution of the expression levels falls within a preset threshold range are discriminated, small RNAs for learning are selected from among the discriminated standard small RNAs, and machine learning is performed by using, as supervised learning data, the presence or absence of the property in the plurality of learning subjects and the individual expression levels of the plurality of small RNAs selected as the small RNAs for learning.
Need to check novelty before this filing date? Find Prior Art

Description

Property determination model generation method, property determination model, property determination method, property determination device, program, and recording medium

[0001] The present disclosure relates to a property determination model generation method, a property determination model, a property determination method, a property determination device, a program, and a recording medium.

[0002] It is known that the expression levels of specific microRNAs among multiple types of small RNAs such as microRNAs (miRNAs) increase or decrease in the blood of patients suffering from cancer diseases such as pancreatic cancer.

[0003] Therefore, blood samples are collected from subjects whose cancer status is to be determined, and the expression levels of microRNAs that are significantly increased or decreased by cancer in the collected blood samples are measured to determine whether the subject has cancer (see, for example, Non-Patent Documents 1 and 2 below).

[0004] Non-Patent Document 1 discloses a method of obtaining sample data including the expression levels of multiple types of microRNAs in blood samples collected from patients with cancer such as esophageal cancer, stomach cancer, colon cancer, hepatobiliary tract cancer, pancreatic cancer, lung cancer, breast cancer, and prostate cancer, as well as blood samples collected from healthy individuals; and using 534 types of microRNAs that are stably expressed in the blood samples of both healthy individuals and cancer patients, machine learning is performed using the expression level data of microRNAs in blood samples from healthy individuals without cancer and blood samples from cancer patients with cancer as training data to generate a trained model that can determine the presence or absence of cancer, and this trained model is used to determine the presence or absence of cancer in a subject.

[0005] Furthermore, Non-Patent Document 2 discloses a method of determining whether a subject has lung cancer by obtaining sample data containing the expression levels of multiple types of microRNAs in blood samples from lung cancer patients and healthy individuals, using 181 types of stably expressed microRNAs, and using the expression level data of microRNAs in blood samples from disease-free healthy individuals and blood samples from lung cancer patients as training data to generate a trained model that can determine the presence or absence of disease through machine learning, and then using this trained model to determine whether a subject has lung cancer.

[0006] In both Non-Patent Documents 1 and 2, a disease determination model is generated by machine learning using microRNAs that are highly expressed in healthy individuals and disease patients. Generally, when generating a disease determination model using machine learning, the training data used as input is often selected to have features that are thought to be less affected by data noise in order to improve determination accuracy. Since microRNAs with low expression levels in biological samples such as blood collected from a subject are thought to be more affected by data noise, it is thought that microRNAs with high expression levels should be selected when generating a disease determination model. Selecting microRNAs with high expression levels in human blood can narrow the number down to 100 to 200 types, and it has been thought that generating a disease determination model using such highly expressed microRNAs as features would enable the generation of a disease determination model with high determination accuracy.

[0007] Non-patent literature 1: Kuno Suzuki1, Hideyoshi Igata, Motoki Abe, Yusuke Yamamoto, "Multiple cancer type classification by small RNA expression profiles with plasma samples from multiple facilities", Cancer Science (Wiley Online library), June 19, 2022, Vol. 113, No. 6, pp. 2144-2166. Non-patent literature 2: Masayasu Inagaki, Makoto Uchiyama, Kanae Yoshikawa-Kawabe, Masafumi Ito, Hideki Murakami, Masaharu Gunji, Makoto Minoshima, Takashi Kohnoh, Ryota Ito, Yuta Kodama, Mari Tanaka-Sakai, Atsushi Nakase1, Nozomi Goto1, Yusuke Tsushima, Shoich Mori, Masahiro Kozuka, Ryo Otomo, Mitsuharu Hirai, Masahiko Fujino, Toshihiko Yokoyama, "Comprehensive circulating microRNA profile as a supersensitive biomarker for early-stage lung cancer screening", Journal of Cancer Research and Clinical Oncology [online], published April 19, 2023, Internet<URL: https: / / doi.org / 10.1007 / s00432-023-04728-9>

[0008] However, even among microRNAs that are highly expressed in blood and are thought to be less affected by data noise, it has been confirmed that there are microRNAs whose expression levels vary depending on the individual. Specifically, it has been confirmed that even if a microRNA is highly expressed in the blood of one subject, its expression level in the blood of another subject may be extremely low. Therefore, if a trained model is generated using microRNAs selected solely based on the criterion of high expression level and the generated disease determination model is used to determine the presence or absence of a disease, there is a possibility that some subjects will not be correctly diagnosed with the disease.

[0009] The present disclosure provides a property determination model generation method, a property determination model, a property determination method, a property determination device, a program, and a recording medium that can generate a property determination model with higher determination accuracy when generating a property determination model that determines the presence or absence of a property in a subject by performing machine learning based on the expression levels of multiple small RNAs in a biological sample collected from the subject.

[0010] A method for generating a property determination model according to one aspect of the present disclosure is a method for generating a property determination model by performing machine learning, which determines the presence or absence of a property in a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, the method comprising: selecting, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies predetermined specific conditions or whose expression level distribution variation is within a predetermined threshold range, as standard small RNAs to be used in generating the property determination model; selecting, from the selected standard small RNAs, training small RNAs to be used in the machine learning of the property determination model based on certain conditions; and generating the property determination model by performing machine learning using the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs as training data.

[0011] According to the present disclosure, by performing machine learning based on the expression levels of multiple small RNAs in a biological sample collected from a subject, it is possible to generate a property determination model with higher determination accuracy when generating a property determination model for determining whether or not the subject has a property.

[0012] 1 is a diagram showing a system configuration of a learning model generation system 20 according to an embodiment of the present disclosure. FIG. 1 is a block diagram showing the hardware configuration of a learning model generation device 30. FIG. 2 is a block diagram showing the functional configuration of the learning model generation device 30 realized by executing a generation program. FIG. 3 is a flowchart showing a method for generating a disease determination model 130. FIG. 4 is a diagram showing the state of machine learning processes of the disease determination model 130. FIG. 5 is a block diagram showing the functional configuration of a disease determination device 50 for performing disease determination using a generated disease determination model 130. FIG. 6 is a flowchart showing the process for determining the presence or absence of cancer using a disease determination model 130 that has undergone machine learning. FIG. 7 is a diagram showing the state of the process for determining the presence or absence of cancer. FIG. 8 is a Venn diagram showing the inclusion relationship of five types of training data when selecting miRNAs with a high normal distribution. FIG. 9 is a diagram showing evaluation results of determination accuracy in a total of 15 disease determination models when selecting miRNAs with a high normal distribution. FIG. 10 is a Venn diagram showing the inclusion relationship of five types of training data when selecting miRNAs with a CV of 100% or less. FIG. 11 is a diagram showing evaluation results of determination accuracy in a total of 15 disease determination models when selecting miRNAs with a CV of 100% or less.

[0013] Next, embodiments of the present disclosure will be described in detail with reference to the drawings.

[0014] An example of an embodiment of the technology of the present disclosure will be described below with reference to the drawings. Note that components and processes that perform the same operations, actions, and functions are given the same reference numerals throughout the drawings, and redundant explanations may be omitted as appropriate. Each drawing is merely a schematic illustration to allow a sufficient understanding of the technology of the present disclosure. Therefore, the technology of the present disclosure is not limited to the illustrated examples. Furthermore, in this embodiment, explanations of configurations that are not directly related to the present disclosure or well-known configurations may be omitted.

[0015] FIG. 1 is a diagram showing the system configuration of a learning model generation system 20 according to an embodiment of the present disclosure.

[0016] In the learning model generation system 20 of this embodiment, a disease determination model is generated by a disease determination model generation method described below.

[0017] In addition, the method for generating a disease determination model in this embodiment is a method for generating a trained disease determination model that determines the presence or absence of a disease in a subject based on the expression levels of multiple small RNAs.Of multiple small RNAs detected from a biological sample such as the subject's blood, only small RNAs whose expression level distribution meets certain conditions are selected as standard small RNAs to be used when generating the disease determination model, and the disease determination model is generated by performing machine learning using training small RNAs selected from the selected standard small RNAs as training data.

[0018] As shown in FIG. 1, the learning model generation system 20 of this embodiment is composed of a next-generation sequencer 21 using NGS (Next-Generation Sequencing) technology and a learning model generation device 30.

[0019] In addition to the next-generation sequencer 21, a quantitative PCR (Polymerase Chain Reaction) or a flow cytometer can also be used as a measuring device as long as it can measure the expression levels of multiple nucleic acid molecules.

[0020] The learning model generation system 20 of this embodiment uses blood 40 collected from a subject whose cancer status is to be determined to be present, to generate a disease determination model capable of determining whether or not the subject has cancer.

[0021] Although the present embodiment will be described using blood collected from a subject to determine the presence or absence of cancer, the presence or absence of cancer may also be determined using biological samples other than blood, such as body fluids, cells, extracellular vesicles, tissue fragments, etc. Here, body fluids include, for example, serum, urine, tears, saliva, sweat, semen, lymph, tissue fluid, body cavity fluid (e.g., pleural effusion, ascites, etc.), cerebrospinal fluid, amniotic fluid, vaginal fluid, nasal mucus, etc.; cells include, for example, red blood cells, white blood cells, platelets, oral swabs, etc.; extracellular vesicles include, for example, exosomes, liposomes, etc.; and tissue fragments include, for example, FFPE (Formalin Fixed Paraffin Embedded) specimens, biopsy specimens, frozen specimens, etc.

[0022] Furthermore, in this embodiment, the case where the subject is a human will be described, but the test subject is not limited to a human, and the present disclosure can be similarly applied to cases where the test subject is a variety of non-human animals such as dogs, cats, etc. In other words, the present disclosure can be applied to cases where a biological sample such as blood collected from a variety of animals such as dogs, cats, etc. is used to test for the presence or absence of cancer or other diseases.

[0023] Furthermore, in this embodiment, the method for generating a disease determination model is described using a case where a disease determination model for determining whether or not a subject has cancer is generated, but the present disclosure is not limited to such a case. The present disclosure can also be applied to cases where a disease determination model for determining whether or not a subject has a disease other than cancer is generated, or to cases where a property determination model for determining a specific property of a subject is generated. The property of a subject to be determined by the property determination model is not limited to the subject's disease. The property of a subject to be determined may be any property that can be determined based on the expression levels of multiple small RNAs. Specific properties to be determined include, for example, confirming the efficacy of a subject's drug, determining the possibility of disease recurrence, confirming / determining lifestyle habits such as checking the presence or absence of a smoking history or drinking history, and predicting physical age.

[0024] The next-generation sequencer 21 measures the expression levels of multiple small RNAs contained in blood 40 collected from a subject. Specifically, the next-generation sequencer 21 performs an amplification process to amplify small RNAs extracted from blood 40, and then measures the relative expression level of each type of small RNA relative to other small RNAs.

[0025] Here, examples of small RNA include microRNA (hereinafter referred to as miRNA). Small RNA may also be small RNA other than miRNA (e.g., piRNA and tsRNA). The following description will be given using a case where the expression level of miRNA is measured from the blood 40 of a subject to determine whether or not the subject is suffering from cancer.

[0026] Then, the next-generation sequencer 21 outputs the measured data on the expression levels of miRNA to the learning model generation device 30 as miRNA data.

[0027] The learning model generation device 30 uses the miRNA data from the next-generation sequencer 21 and information on whether each subject is suffering from a cancer disease to generate a disease determination model for determining whether or not the subject is suffering from a cancer disease.

[0028] Next, the hardware configuration of the learning model generation device 30 described above is shown in the block diagram of FIG.

[0029] 2, the learning model generation device 30 has a function as a computer, and includes a CPU (Central Processing Unit) 31, a ROM (Read Only Memory) 32, a RAM (Random Access Memory) 33, a storage 34, an input unit 35, a display unit 36, and a communication interface (I / F) 37. Each component is connected to each other via a bus 39 so as to be able to communicate with each other.

[0030] The CPU 31 (an example of a processor) is a central processing unit that executes various programs and controls each part. That is, the CPU 31 reads a program from the ROM 32 or the storage 34 and executes the program using the RAM 33 as a work area. The CPU 31 controls each of the above components and performs various arithmetic processing in accordance with the program stored in the ROM 32 or the storage 34.

[0031] The ROM 32 stores various programs and various data. The RAM 33 temporarily stores programs or data as a working area. The storage 34 is configured by an HDD (Hard Disk Drive) or an SSD (Solid State Drive) and stores various programs including the operating system and various data.

[0032] In this embodiment, for example, a generation program for generating a disease determination model is recorded in the storage 34. This generation program may be a single program, or a group of programs configured by multiple programs or modules. The generation program may be recorded in the ROM 32. The ROM 32 and the storage 34 function as an example of a non-transitory recording medium.

[0033] An example of a processor is not limited to the above-mentioned CPU, which is a general-purpose processor, but may be, for example, a dedicated processor configured with a circuit designed specifically for executing a specific process. Also, an example of a processor is not limited to a single processor, but may be a processor configured by multiple processors located in physically separate locations working together.

[0034] The input unit 35 includes a pointing device such as a mouse and a keyboard, and is used to input various information. The input unit 35 also receives miRNA data representing the expression levels of multiple miRNAs measured by the next-generation sequencer 21 as input.

[0035] The display unit 36 ​​is, for example, a liquid crystal display that displays various information. The display unit 36 ​​may also function as the input unit 35 that inputs operations from the user by employing a touch panel system.

[0036] The communication interface 37 is an interface for communicating with other devices, and transmits and receives data to and from external devices including the next-generation sequencer 21 .

[0037] Next, the functional configuration of the learning model generation device 30 realized by executing the generation program described above is shown in the block diagram of FIG.

[0038] As shown in FIG. 3, in the learning model generation device 30, an acquisition unit 110, a generation unit 120, and a disease determination model 130 are configured by a CPU 31 executing a determination program.

[0039] The acquisition unit 110 acquires miRNA data (hereinafter sometimes referred to as training miRNAs) that is the result of measuring the expression levels of each of multiple miRNAs in the blood of the training subject by measuring the blood of the training subject.

[0040] Although the expression levels of multiple miRNAs in blood are absolute values ​​derived from the living body, it is difficult to quantify the expression levels of miRNAs in blood as absolute values ​​because the expression levels must be quantified through processes such as measurement devices and reagent processing. In other words, miRNA data is a relative representation of the expression levels of miRNAs obtained by using a next-generation sequencer 21 or data processing. However, if the expression levels of miRNAs in blood can be quantified as absolute values, the absolute values ​​of the expression levels of miRNAs in blood may be used as miRNA data.

[0041] The generating unit 120 performs machine learning to generate a disease determination model 130 that determines whether or not a subject has cancer, based on the miRNA data (training miRNA) acquired by the acquiring unit 110 and information indicating whether or not each of the training subjects has cancer. The disease determination model 130 generated by the generating unit 120 is stored in, for example, a non-volatile storage device.

[0042] Next, a method for generating the disease determination model 130 described above is shown in the flowchart of FIG.

[0043] First, in the collection step of step S101, blood is collected from a study subject and centrifuged, and the resulting serum is dispensed into storage tubes in predetermined amounts and stored in a deep freezer at −80° C. The study subjects include both healthy subjects who are not affected by cancer and patients who are affected by cancer.

[0044] Next, in the measurement step of step S102, the frozen serum is removed from the deep freezer and thawed at room temperature, and then the expression level of miRNA is measured using the next-generation sequencer 21.

[0045] Then, in the acquisition step of step S103, the learning model generation device 30 acquires the results of measuring the expression levels of multiple miRNAs in the serum of the training subject measured in the measurement step of step S102.

[0046] Next, in the selection process of step S104, the learning model generation device 30 selects, from among the multiple miRNAs detected from blood collected from multiple training subjects consisting of training subjects suffering from cancer and training subjects not suffering from cancer, multiple miRNAs whose expression level distribution shape satisfies predetermined specific conditions or whose expression level distribution variation is within a predetermined threshold range, as standard miRNAs to be used when generating the disease determination model 130.

[0047] Here, miRNAs whose expression level distribution shape satisfies predetermined specific conditions are, for example, miRNAs whose index indicating the normal distribution of the expression level distribution is equal to or greater than a predetermined value. Whether the expression level distribution has normal distribution can be determined, for example, by a Shapiro-Wilk test. Specifically, the p-value in the Shapiro-Wilk test is used as an index indicating the normal distribution of the expression level distribution, and if this p-value is 0.05 or greater, it is determined that the result is not statistically significant, and the expression level distribution is determined to have normal distribution. In other words, miRNAs whose expression level distribution shows normality are selected as standard miRNAs.

[0048] Furthermore, miRNAs whose expression level distribution variability is within a predetermined threshold range are determined by using the coefficient of variation (CV) calculated by dividing the standard deviation of the expression level by the mean value as an index indicating the degree of variability in the expression level distribution, and miRNAs whose expression level distribution variability is equal to or less than a predetermined value are selected as standard miRNAs. Specifically, miRNAs whose expression level CV is 100% or less are determined to be miRNAs whose expression level varies little between individuals and are selected as standard miRNAs. Here, the present disclosure is not limited to miRNAs whose expression level distribution variability is within a predetermined threshold range, and miRNAs whose expression level distribution variability meets certain predetermined conditions may be selected as standard miRNAs to be used when generating the disease determination model 130.

[0049] Then, in the selection process of step S105, the learning model generation device 30 selects, from the selected standard miRNAs, training miRNAs to be used for machine learning of the disease determination model based on certain conditions. Specifically, the learning model generation device 30 selects, from the selected standard miRNAs, miRNAs that are within a predetermined ranking from the top when sorted in descending order of expression level, for example, within the top 100, as training miRNAs to be used for machine learning of the disease determination model.

[0050] Finally, in the learning process of step S106, the learning model generation device 30 uses the acquired miRNA data and information indicating whether each learning subject is affected by cancer as training data to perform machine learning of the disease determination model 130. The machine learning process of such a disease determination model 130 is shown in FIG.

[0051] As can be seen from Figure 5, in the disease determination system 20 of this embodiment, machine learning of the disease determination model 130 is performed by linking the expression level data of training miRNAs collected from the blood of healthy training subjects who are not suffering from cancer and the expression level data of training miRNAs collected from the blood of diseased training subjects who are suffering from cancer with information on whether the training subjects are suffering from cancer.

[0052] In this way, the generation unit 120 generates the disease determination model 130 by performing machine learning using as training data the presence or absence of cancer in multiple training subjects and the expression levels of each of the multiple miRNAs selected as training miRNAs. In this way, the disease determination model 130 is a learning model generated by performing machine learning using as training data the presence or absence of cancer in multiple training subjects and the expression levels of each of the multiple miRNAs selected as training miRNAs. Specific examples of this learning model are described below.

[0053] <Examples of learning models> Various linear and non-linear algorithms known as machine learning algorithms can be used, or multiple algorithms can be combined. For example, the following algorithms can be used:

[0054] Random forest Gradient boosting trees Extreme gradient boosting trees Light-gradient boosting machine (GBM) Neural networks Regularized regression Elastic-net regression K-Nearest neighbors Support vector machine Generalized additive model

[0055] Next, a method for determining whether or not a subject to be determined is suffering from cancer using the disease determination model 130 generated by the above-described generation method will be described.

[0056] The functional configuration of a disease assessment device 50 for performing such disease assessment is shown in the block diagram of FIG.

[0057] As shown in Fig. 6, the disease determination device 50 is configured with an acquisition unit 110, a disease determination model 130, and a determination unit 140. Here, the disease determination model 130 in Fig. 6 is the trained disease determination model 130 generated by the learning model generation device 30 shown in Fig. 3.

[0058] The acquisition unit 110 acquires miRNA data, which is the result of measuring the expression levels of each of multiple miRNAs in the blood of the subject to be assessed, by measuring the blood of the subject to be assessed for the presence or absence of cancer.

[0059] Then, the determination unit 140 determines whether or not the subject is suffering from cancer by inputting miRNA data showing the results of measuring the expression levels of multiple miRNAs in blood collected from the subject to be determined into the disease determination model 130.

[0060] In this way, the disease determination device 50 has a disease determination model 130, and by inputting miRNA data showing the results of measuring the expression levels of multiple miRNAs in blood collected from a subject into this disease determination model 130, it functions as a property determination device that determines whether or not the subject is suffering from a cancer disease.

[0061] In the disease diagnosis device 50, the diagnosis result in the diagnosis unit 140 is output to an external device or displayed on the display unit.

[0062] Next, the flow chart of FIG. 7 shows the process of determining the presence or absence of cancer using the disease determination model 130 that has undergone machine learning in this manner.

[0063] First, in the collection step of step S201, blood is collected from a subject to be determined whether or not the subject is affected by cancer.

[0064] Next, in the measurement step of step S202, the expression level of miRNA in the blood collected from the subject is measured by the next-generation sequencer 21.

[0065] Then, in the acquisition step of step S203, the disease assessment apparatus 50 acquires the results of measuring the expression levels of each of the multiple miRNAs in the blood of the subject measured in the measurement step of step S202 as miRNA data.

[0066] Finally, in the disease presence / absence determination process of step S204, the disease determination device 50 inputs the acquired miRNA data into the trained disease determination model 130 to obtain a determination result indicating whether or not the subject is suffering from cancer. Such a cancer presence / absence determination process is shown in Figure 8.

[0067] As can be seen from Figure 8, in the disease determination system of this embodiment, miRNA expression level data collected from the blood of a subject to be determined to be suffering from cancer is input into a trained disease determination model 130 to obtain a determination result.

[0068] According to the method for generating a disease determination model of this embodiment, by performing machine learning based on the expression levels of multiple miRNAs in blood collected from a subject, it is possible to generate a disease determination model with higher determination accuracy when generating a disease determination model for determining whether or not the subject is suffering from a cancer disease.

[0069] The reason why the method for generating a disease determination model according to this embodiment makes it possible to generate a disease determination model with higher determination accuracy will be explained below.

[0070] In the method for generating a disease determination model in this embodiment, instead of simply selecting miRNAs with high expression levels in the subject's blood as standard miRNAs, only miRNAs whose expression level distribution meets predetermined specific conditions are selected as standard miRNAs. Specifically, miRNAs with normal expression level distributions or miRNAs whose expression level distribution variability is within a predetermined threshold range are selected as standard miRNAs. Therefore, miRNAs whose expression levels vary significantly depending on individuals are excluded from the standard miRNAs. As a result, by selecting training miRNAs from these standard miRNAs and performing machine learning of the disease determination model 130 using the selected training miRNAs as training data, a disease determination model with improved determination accuracy can be realized.

[0071] Next, as an example, experimental data is used to show that a disease determination model with higher accuracy can be generated by generating a disease determination model using the generation method of this embodiment rather than generating a disease determination model using miRNAs that are highly expressed in the subject's blood.

[0072] Example 1: Selection of miRNAs whose expression levels have a normal distribution

[0073] In this experiment, blood was collected from a total of 360 subjects, including 204 healthy individuals and 156 lung cancer patients, and miRNA expression data was obtained from serum obtained by centrifuging the collected blood using NGS.

[0074] The acquired expression level data was then logarithmically transformed, and from the logarithmically transformed miRNA expression level data, 142 healthy subjects and 109 lung cancer patients were randomly extracted from each group of healthy subjects and lung cancer patients. The normal distribution of the expression levels of each miRNA (i.e., training miRNA) in the expression level data of the extracted subjects (i.e., training subjects) was confirmed using a Shapiro-Wilk test. Then, miRNAs with a p-value of 0.05 or greater (p-value ≧ 0.05) were determined to have a normal distribution (high normal distribution). For comparison, five learning models for determining the presence or absence of lung cancer were generated as disease determination models by performing machine learning using the following five types of training data. The inclusion relationship of these five types of training data is shown in the Venn diagram of Figure 9.

[0075] (1) Unsorted miRNA data: Training data (unsorted data) containing all 854 miRNA data detected from the subject's blood.

[0076] (2) miRNA data with a high normal distribution from healthy individuals. Training data (normal distribution data from healthy individuals) containing 137 types of miRNA data detected from the blood of healthy individuals, whose expression levels were determined to have a normal distribution.

[0077] (3) Highly normally distributed miRNA data from lung cancer patients. Training data (normally distributed lung cancer patient data) containing 122 types of miRNA data detected from the blood of lung cancer patients, whose expression levels were determined to have a normally distributed distribution.

[0078] (4) Union of miRNA data with high normal distribution in healthy subjects and lung cancer patients. Training data (union) including 183 types of miRNA data whose expression levels were determined to have a normal distribution in the blood of either healthy subjects or lung cancer patients.

[0079] (5) Intersection set of miRNA data with high normal distribution in healthy subjects and lung cancer patients. Training data (intersection set) including 76 types of miRNA data whose expression levels were determined to have a normal distribution in the blood of both healthy subjects and lung cancer patients.

[0080] Then, the miRNA expression level data of 62 healthy subjects and 47 lung cancer patients (i.e., subjects to be assessed for the presence or absence of a hypothetical cancer disease) was used as validation data, excluding the expression level data used to generate the disease assessment model from the blood miRNA data of all subjects, and the accuracy of the assessment of the generated disease assessment model was evaluated.

[0081] The evaluation results of such determination accuracy are shown in Figure 10. In Figure 10, for each of the five training data ((1) to (5)), machine learning was performed on learning models (A to C) using three of the multiple machine learning algorithms shown above to generate a total of 15 disease determination models, and the results of calculating the AUC (Area Under the Curve) value for each determination result of each disease determination model are shown. Here, AUC is an evaluation index for a two-class classification model. AUC takes a value ranging from 0 to 1, and a larger AUC means a higher classification accuracy of the two-class classification model.

[0082] Referring to Figure 10, compared to when a learning model was generated using training data containing 854 types of miRNA data without selecting miRNAs ((1): unselected data), the value of AUC, an evaluation index, was larger when machine learning was performed using training data from selected data of miRNAs with normally distributed expression levels, including miRNA data with a high normal distribution from healthy individuals (2), miRNA data with a high normal distribution from lung cancer patients (3), and the union of miRNA data with a high normal distribution from healthy individuals and lung cancer patients (4).

[0083] On the other hand, when machine learning was performed using only the intersection set (5) of miRNA data with high normal distributions from healthy individuals and lung cancer patients as training data, the AUC value was smaller than when machine learning was performed using unselected miRNA data.

[0084] To summarize the above evaluation results, the AUC values ​​for all learning models A to C are in the following order: intersection (5) < unsorted data (1) < normal distribution data of healthy subjects (2) < normal distribution data of lung cancer patients (3) < union (4).

[0085] From the above results, it can be seen that, basically, by performing machine learning using miRNAs whose expression levels have a normal distribution to generate a disease determination model, it is possible to generate a disease determination model with higher determination accuracy than by performing machine learning using unselected miRNAs. However, it can be seen that if a disease determination model is generated using only the intersection set (5) of miRNAs whose expression levels are normally distributed in the blood of both healthy individuals and lung cancer patients as training data, the determination accuracy may be lower than when unselected miRNAs are used.

[0086] Therefore, it is possible to select, as standard miRNAs to be used when generating a disease determination model, miRNAs included in the union (4) whose expression level distribution pattern in blood collected from either a training subject with lung cancer or a healthy training subject, or miRNAs included in lung cancer patient normal distribution data (3) whose expression level distribution pattern in blood collected from a training subject with lung cancer is normally distributed, or miRNAs included in healthy subject miRNA data (2) whose expression level has a high normal distribution pattern.

[0087] Example 2: Selection of miRNAs whose distribution of expression levels falls within a preset threshold range

[0088] In this experiment, blood was collected from a total of 360 subjects, including 204 healthy individuals and 156 lung cancer patients, and miRNA expression data was obtained from serum obtained by centrifuging the collected blood using NGS.

[0089] The acquired expression level data was then logarithmically transformed, and from the logarithmically transformed miRNA expression level data, 142 healthy subjects and 109 lung cancer patients were randomly selected from each group of healthy subjects and lung cancer patients. The coefficient of variation (CV) of the expression level of each miRNA (i.e., training miRNA) in the expression level data of the selected subjects (i.e., training subjects) was calculated. MiRNAs with a calculated CV of 100% or less were determined to have small differences in expression levels due to individual differences. For comparison, five types of training data were used to perform machine learning to generate five learning models for determining the presence or absence of lung cancer as disease determination models. The inclusion relationship of these five types of training data is shown in the Venn diagram in Figure 11.

[0090] (1) Unsorted miRNA data: Training data (unsorted data) containing all 854 miRNA data detected from the subject's blood.

[0091] (2) miRNA data with CV of 100% or less in healthy subjects. Training data including 384 types of miRNA data with expression levels of CV of 100% or less among miRNA data detected from the blood of healthy subjects (healthy subject CV of 100% or less data).

[0092] (3) miRNA data with CV of 100% or less from lung cancer patients. Training data including 399 types of miRNA data detected from the blood of lung cancer patients, whose expression level CV was determined to be 100% or less (lung cancer patient CV of 100% or less).

[0093] (4) Union of miRNA data with CV of 100% or less between healthy individuals and lung cancer patients. Training data (union) including 407 types of miRNA data whose expression levels were determined to have a CV of 100% or less in the blood of either healthy individuals or lung cancer patients.

[0094] (5) Intersection set of miRNA data with CV of 100% or less between healthy subjects and lung cancer patients. Training data (intersection set) including 376 types of miRNA data determined to have CV of expression levels of 100% or less in the blood of both healthy subjects and lung cancer patients.

[0095] Then, the miRNA expression level data of 62 healthy subjects and 47 lung cancer patients (i.e., subjects to be assessed for the presence or absence of a hypothetical cancer disease) was used as validation data, excluding the expression level data used to generate the disease assessment model from the blood miRNA data of all subjects, and the accuracy of the assessment of the generated disease assessment model was evaluated.

[0096] The evaluation results of such determination accuracy are shown in Figure 12. In Figure 12, machine learning of learning models (D to F) was performed using three of the multiple machine learning algorithms shown above for each of the five training data ((1) to (5)), generating a total of 15 disease determination models, and the AUC value was calculated for each determination result of each disease determination model.

[0097] Referring to Figure 12, for learning models E and F, compared to when a learning model was generated using training data ((1): unselected data) containing 854 types of miRNA data without selecting miRNAs, the value of AUC, an evaluation index, was larger when machine learning was performed using as training data miRNA data (2) with a CV of 100% or less from healthy subjects, miRNA data (3) with a CV of 100% or less from lung cancer patients, and the union (4) of miRNA data (4) with a CV of 100% or less from healthy subjects and lung cancer patients, from among data in which miRNAs with a CV of 100% or less were selected.

[0098] Furthermore, referring to Figure 12, for learning model D, compared to when a learning model was generated using training data ((1): unselected data) containing 854 types of miRNA data without selecting miRNAs, the value of AUC, an evaluation index, was larger when machine learning was performed using as training data miRNA data (3) from lung cancer patients with a CV of 100% or less and the union (4) of miRNA data from healthy individuals and lung cancer patients with a CV of 100% or less, among data in which miRNAs with a CV of 100% or less were selected.

[0099] On the other hand, in all of learning models D to F, when machine learning was performed using only the intersection set (5) of miRNA data for healthy individuals and lung cancer patients with a CV of 100% or less as training data, the AUC value was smaller than when machine learning was performed using unselected miRNA data.

[0100] To summarize the above evaluation results, the AUC values ​​for learning models E and F are in the following order: intersection (5) < unsorted data (1) < healthy subject CV 100% or less data (2) < lung cancer patient CV 100% or less data (3) < union (4)

[0101] In addition, in learning model D, the AUC values ​​are in the following order: intersection (5) < healthy subject CV 100% or less data (2) < unsorted data (1) < lung cancer patient CV 100% or less data (3) < union (4)

[0102] From the above results, it can be seen that, basically, by performing machine learning using miRNA data whose expression level CV is 100% or less to generate a disease determination model, it is possible to generate a disease determination model with higher determination accuracy than by performing machine learning using unselected miRNAs. However, it can be seen that when a disease determination model is generated using miRNAs whose expression level CV is 100% or less in the blood of both healthy individuals and lung cancer patients, the determination accuracy may be lower than when unselected miRNAs are used.

[0103] Furthermore, in the case of learning model D, by performing machine learning using miRNA data (2) in which the CV of healthy individuals is less than 100%, to generate a disease determination model, it can be seen that the determination accuracy may be lower than when unselected miRNAs are used.

[0104] From the above results, it can be seen that miRNAs included in the union (4) whose expression level CV is 100% or less (the variation in the distribution of expression levels is within a predetermined threshold range) in blood collected from either a training subject with lung cancer or a healthy training subject, or miRNAs included in the lung cancer patient CV 100% or less data (3) whose expression level CV is 100% or less (the variation in the distribution of expression levels is within a predetermined threshold range) in blood collected from a training subject with lung cancer, can be selected as standard miRNAs to be used when generating a disease determination model.

[0105] The disclosure of Japanese Patent Application No. 2024-92472, filed on June 6, 2024, is incorporated herein by reference in its entirety. All documents, patent applications, and technical standards mentioned herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.

[0106] [Additional Notes] Preferred aspects of the present disclosure are described below. (Aspect 1) A method for generating a property determination model generated by machine learning and for determining the presence or absence of a property in a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, the method comprising: selecting, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies a predetermined specific condition or whose expression level distribution variation is within a predetermined threshold range, as standard small RNAs to be used in generating the property determination model; selecting, from the selected standard small RNAs, training small RNAs to be used in the machine learning of the property determination model based on a certain condition; and generating the property determination model by performing machine learning using as training data the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs.

[0107] (Aspect 2) The method for generating a property determination model according to Aspect 1, wherein the small RNAs whose expression level distribution shape satisfies a predetermined specific condition are small RNAs whose expression level distribution has an index indicating normal distribution equal to or greater than a predetermined value.

[0108] (Aspect 3) A method for generating a property determination model according to Aspect 1, wherein a coefficient of variation calculated by dividing the standard deviation of the expression level by the mean value is used as an index indicating the degree of variation in the distribution of expression levels, and small RNAs for which the index is equal to or less than a predetermined value are selected as small RNAs whose variation in the distribution of expression levels is within a predetermined threshold range, thereby selecting standard small RNAs.

[0109] (Aspect 4) A method for generating a property determination model described in any one of Aspects 1 to 3, wherein, when selecting training small RNAs to be used for machine learning of the property determination model from the selected standard small RNAs based on certain conditions, small RNAs that are ranked within a predetermined order from the top when the selected standard small RNAs are sorted in order of increasing expression level are selected as training small RNAs to be used for machine learning of the property determination model.

[0110] (Aspect 5) A method for generating a property determination model according to any one of Aspects 1 to 4, wherein, in a biological sample collected from either a training subject having the property or a training subject not having the property, or in a biological sample collected from a training subject having the property, small RNAs whose expression level distribution shape satisfies a predetermined specific condition or whose expression level distribution variation is within a predetermined threshold range are selected as standard small RNAs to be used when generating the property determination model.

[0111] (Aspect 6) The method for generating a property determination model according to any one of Aspects 1 to 5, wherein the biological sample is blood, and the property is the presence or absence of cancer.

[0112] (Aspect 7) The method for generating a property determination model according to any one of Aspects 1 to 6, wherein the small RNA is a microRNA.

[0113] (Aspect 8) A property determination model generated by performing machine learning, for determining the presence or absence of a property in a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, wherein, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies a predetermined specific condition or whose expression level distribution variation is within a predetermined threshold range are selected as standard small RNAs to be used in the machine learning of the property determination model, from the selected standard small RNAs, training small RNAs to be used in the machine learning of the property determination model are selected based on certain conditions, and the property determination model is generated by performing machine learning using the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs as training data.

[0114] (Aspect 9) A property determination method for determining the presence or absence of the property in a subject by inputting small RNA data showing the results of measuring the expression levels of multiple small RNAs in a biological sample collected from the subject into the property determination model described in Aspect 8.

[0115] (Aspect 10) A property determination device having the property determination model described in Aspect 8, which determines whether or not a subject has the property by inputting small RNA data showing the results of measuring the expression levels of multiple small RNAs in a biological sample collected from the subject into the property determination model.

[0116] (Aspect 11) A program for generating a property determination model generated by performing machine learning and for determining the presence or absence of a property of a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, the program causing a computer to execute the following steps: selecting, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies a predetermined specific condition or whose expression level distribution variation is within a predetermined threshold range, as standard small RNAs to be used in generating the property determination model; selecting, from the selected standard small RNAs, training small RNAs to be used in machine learning of the property determination model based on a certain condition; and generating the property determination model by performing machine learning using as training data the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs.

[0117] (Aspect 12) A non-transitory recording medium having recorded thereon a program for generating a property determination model that is generated by performing machine learning and that determines the presence or absence of a property of a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, the non-transitory recording medium having recorded thereon a program for causing a computer to execute the following steps: selecting, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies a predetermined specific condition or whose expression level distribution variation is within a predetermined threshold range, as standard small RNAs to be used in generating the property determination model; selecting, from the selected standard small RNAs, training small RNAs to be used in machine learning of the property determination model based on a certain condition; and generating the property determination model by performing machine learning using as training data the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs.

Claims

1. A method for generating a property determination model generated by machine learning and used to determine the presence or absence of a property in a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, the method comprising: selecting, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies predetermined specific conditions or whose expression level distribution variation is within a predetermined threshold range, as standard small RNAs to be used in the machine learning of the property determination model; selecting, from the selected standard small RNAs, training small RNAs to be used in the machine learning of the property determination model based on certain conditions; and generating the property determination model by performing machine learning using as training data the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs.

2. The method for generating a property determination model according to claim 1, wherein a small RNA whose distribution shape of expression levels satisfies a predetermined specific condition is a small RNA whose index showing the normal distribution of expression levels is equal to or greater than a predetermined value.

3. A method for generating a property determination model as described in claim 1, wherein a coefficient of variation calculated by dividing the standard deviation of the expression level by the mean value is used as an index showing the degree of variation in the distribution of expression levels, and small RNAs for which the index is equal to or less than a predetermined value are selected as small RNAs whose variation in the distribution of expression levels is within a predetermined threshold range, and are therefore selected as standard small RNAs.

4. A method for generating a property determination model as described in claim 1, wherein, when selecting training small RNAs to be used for machine learning of the property determination model based on certain conditions from the selected standard small RNAs, small RNAs that are within a predetermined ranking from the top when sorted in order of highest expression level from the selected standard small RNAs are selected as training small RNAs to be used for machine learning of the property determination model.

5. A method for generating a property determination model as described in claim 1, wherein small RNAs whose expression level distribution shape satisfies predetermined specific conditions or whose expression level distribution variation is within a predetermined threshold range are selected as standard small RNAs to be used when generating the property determination model in a biological sample collected from either a training subject having the property or a training subject not having the property, or in a biological sample collected from a training subject having the property.

6. The method for generating a property determination model according to claim 1, wherein the biological sample is blood, and the property is the presence or absence of cancer.

7. The method for generating a property determination model according to claim 1, wherein the small RNA is a microRNA.

8. A property determination model generated by machine learning to determine the presence or absence of a property in a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, wherein, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies predetermined specific conditions or whose expression level distribution variation is within a predetermined threshold range are selected as standard small RNAs to be used in the machine learning of the property determination model; from the selected standard small RNAs, training small RNAs to be used in the machine learning of the property determination model are selected based on certain conditions; and the property determination model generated by performing machine learning using the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs as training data.

9. A method for determining whether or not a subject has the above-mentioned property by inputting small RNA data showing the results of measuring the expression levels of multiple small RNAs in a biological sample collected from the subject into the property determination model described in claim 8.

10. A property determination device having the property determination model according to claim 8, which determines whether or not a subject has the property by inputting small RNA data showing the results of measuring the expression levels of multiple small RNAs in a biological sample collected from the subject into the property determination model.

11. A program for generating a property determination model generated by machine learning and for determining the presence or absence of a property of a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, the program causing a computer to execute the following steps: selecting, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies predetermined specific conditions or whose expression level distribution variation is within a predetermined threshold range, as standard small RNAs to be used in generating the property determination model; selecting, from the selected standard small RNAs, training small RNAs to be used in machine learning of the property determination model based on certain conditions; and generating the property determination model by performing machine learning using as training data the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs.

12. A non-transitory recording medium having recorded thereon a program for generating a property determination model that is generated by performing machine learning and that determines the presence or absence of a property of a subject based on the expression levels of multiple small RNAs in a biological sample collected from the subject, the program causing a computer to execute the following steps: selecting, from multiple small RNAs detected in biological samples collected from multiple training subjects consisting of training subjects with the property and training subjects without the property, multiple small RNAs whose expression level distribution shape satisfies predetermined specific conditions or whose expression level distribution variation is within a predetermined threshold range, as standard small RNAs to be used in generating the property determination model; selecting, from the selected standard small RNAs, training small RNAs to be used in machine learning of the property determination model based on certain conditions; and generating the property determination model by performing machine learning using as training data the presence or absence of the property in the multiple training subjects and the expression levels of each of the multiple small RNAs selected as the training small RNAs.

Citation Information

Patent Citations

  • Machine learning system for pancreatic cancer diagnosis based on serum miRNA

    CN114550830A

  • DISEASE PRESENCE DETECTION DEVICE, DISEASE PRESENCE DETECTION METHOD, AND DISEASE PRESENCE DETECTION PROGRAM

    JP2022024092A

  • Methods and machine learning systems for predicting the likelihood or risk of having cancer

    US20180068083A1

  • Test method, test device, learning method, learning device, test program and learning program

    WO2021132547A1