Feature amount selection method, property determination model generation method, property determination model, property determination method, and property determination device
The method of selecting features from small RNA data by removing correlations and then applying significance testing addresses the challenge of controlling feature quantity, improving the machine learning model's performance by reducing overfitting and underfitting.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ARKRAY INC
- Filing Date
- 2025-10-27
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for selecting feature quantities from small RNA data for machine learning face challenges in controlling the number of features to avoid overfitting or underfitting, as multicollinearity affects the learned model's performance.
A method that selects features by removing correlations between features using multicollinearity removal followed by significance testing, ensuring appropriate feature quantity adjustment.
This approach allows for easier control of the number of remaining features, enhancing the robustness and performance of the machine learning model by reducing overfitting and underfitting risks.
Smart Images

Figure JP2025037656_07052026_PF_FP_ABST
Abstract
Description
Feature quantity selection method, method for generating property determination model, property determination model, property determination method, and property determination device
[0001] The present disclosure relates to a feature quantity selection method, a method for generating a property determination model, a property determination model, a property determination method, and a property determination device.
[0002] When performing machine learning on a learning model using small RNA data indicating a measurement result of measuring the expression level of small RNA in a biological sample-derived sample, it is conceivable to select feature quantities from the small RNA data.
[0003] Patent Document 1 discloses a method for removing multicollinearity as a feature quantity selection method for selecting feature quantities from data used for machine learning.
[0004] Non-Patent Document 1 discloses a method of combining a significance test and removal of multicollinearity as the feature quantity selection method. Patent Document 1: Japanese Unexamined Patent Application Publication No. 2023-83931 Non-Patent Document 1: Circulating plasma microRNAs in systemic sclerosis-associated pulmonary arterial hypertension
[0005] Here, if feature quantities having multicollinearity remain, it has an adverse effect on the learned model. Therefore, in the feature quantity selection method by removing multicollinearity, a threshold value is set within a range that does not have an adverse effect on the learned model. For this reason, it is difficult to control the number of remaining feature quantities in the feature quantity selection method by removing multicollinearity. And when selecting feature quantities in the order of significance test and removal of multicollinearity as in Non-Patent Document 1, it is difficult to control the number of remaining feature quantities, and the number of remaining feature quantities may extremely decrease. Thus, when the number of feature quantities is small, the AUC score indicating the determination performance of the learned model is likely to decrease. On the other hand, when the number of feature quantities is more than necessary, overfitting of the learned model is likely to occur. Therefore, it is desired to be able to easily select an appropriate number of appropriate feature quantities.
[0006] This disclosure aims to provide a way to easily adjust remaining features when selecting features to be used in machine learning for a learning model from small RNA data, which shows measurement results of the expression level of small RNAs in biological samples.
[0007] A feature selection method according to one aspect of this disclosure is a method for selecting features to be used in machine learning by a learning model from small RNA data showing measurement results of the expression level of small RNA in a biological sample, wherein features are selected from the small RNA data by removing correlations between features, and then features are further selected from those features by class differences.
[0008] According to this disclosure, when selecting features to be used for machine learning in a learning model from small RNA data, which shows measurement results of the expression level of small RNA in biological samples, the remaining features can be easily adjusted.
[0009] This is a diagram showing the system configuration of a learning model generation system 20 according to one embodiment of the present disclosure. This is a block diagram showing the hardware configuration of a learning model generation device 30. This is a block diagram showing the functional configuration of a learning model generation device 30 realized by the execution of a generation program. This is a flowchart showing the method for generating a disease judgment model 130. This is a flowchart showing the feature selection method. This is a diagram showing the procedure of the second selection step of the feature selection method. This is a diagram showing the machine learning process of the disease judgment model 130. This is a block diagram showing the functional configuration of a disease judgment device 50 for performing disease judgment using the generated disease judgment model 130. This is a flowchart showing the process for determining the presence or absence of cancer using a disease judgment model 130 that has undergone machine learning. This is a diagram showing the process for determining the presence or absence of cancer. This is graph 1 showing the performance of the disease judgment model under condition 1. This is graph 2 showing the performance of the disease judgment model under condition 2. This is graph 3 showing the performance of the disease judgment model under condition 3. This is graph 4 showing the performance of the disease judgment model under condition 4. This is a conceptual diagram of a model that performs multiple class classification from a single dataset. This is a conceptual diagram showing the computational cost when feature selection is performed in the order of multicollinearity removal and significance testing. This is a conceptual diagram illustrating the computational complexity when selecting features in the following order: significance testing, multicollinearity removal, and significance testing.
[0010] An example of an embodiment relating to the technology of this disclosure will be described below with reference to the drawings. However, the embodiments of this disclosure are not limited to the embodiments described below. In the embodiments described below, the components (including element steps, etc.) are not essential unless otherwise specified. The same applies to numerical values and their ranges, and do not limit the embodiments of this disclosure.
[0011] In this disclosure, the term "process" includes not only processes that are independent of other processes, but also processes that cannot be clearly distinguished from other processes, provided that the purpose of such process is achieved.
[0012] In this disclosure, the numerical range indicated using "~" includes the numbers before and after "~" as the minimum and maximum values, respectively.
[0013] In numerical ranges described in stages within this disclosure, the upper or lower limit of one numerical range may be replaced with the upper or lower limit of another numerical range described in stages. Furthermore, in numerical ranges described within this disclosure, the upper or lower limit of that range may be replaced with the values shown in the examples.
[0014] In embodiments of this disclosure, components and processes that perform the same function, operation, or function are given the same reference numerals throughout the drawings, and redundant descriptions may be omitted as appropriate.
[0015] When embodiments of this disclosure are described with reference to the drawings, each drawing is provided only schematically to the extent that the technology of this disclosure can be fully understood. Therefore, the technology of this disclosure is not limited to the illustrated examples. Furthermore, descriptions of configurations not directly related to this disclosure or well-known configurations may be omitted.
[0016] <Learning Model Generation System 20> First, the learning model generation system 20 according to this embodiment will be described. Figure 1 is a schematic diagram showing the learning model generation system 20 according to this embodiment.
[0017] The learning model generation system 20 generates a disease judgment model using a disease judgment model generation method. The disease judgment model generation method is a method of generating a disease judgment model by performing machine learning using features selected by a feature selection method. As shown in Figure 1, the learning model generation system 20 comprises a next-generation sequencer 21 using NGS (Next-Generation Sequencing) technology and a learning model generation device 30.
[0018] The following describes the disease determination model, the parts of the learning model generation system 20, the method for generating the disease determination model, the feature selection method, the disease determination method using the disease determination model, and the disease determination device 50 that executes the disease determination method.
[0019] <Disease Determination Model> The disease determination model is the model to be generated by the learning model generation system 20. In this embodiment, the disease determination model is a model that determines whether or not a subject has cancer based on the expression level of microRNA (hereinafter sometimes referred to as miRNA) in the blood 40 collected from the subject.
[0020] The subject is an example of a sample to be collected. In this embodiment, the case where the subject is human is described, but the sample to be collected is not limited to humans. The sample to be collected may be human or non-human animal. Examples of non-human animals include non-human mammals (monkeys, dogs, cats, mice, rats, rabbits, cattle, horses, pigs, and sheep, etc.) and birds (chickens, quail, etc.).
[0021] Blood 40 is an example of a biological sample. In this embodiment, the case in which blood is used will be described, but the biological sample is not limited to blood. Examples of biological samples that can be used include body fluids, cells, extracellular vesicles, tissue fragments, etc. Examples of body fluids include blood, serum, plasma, urine, tears, saliva, sweat, semen, lymph, tissue fluid, body cavity fluid (e.g., pleural fluid, ascites), cerebrospinal fluid, amniotic fluid, vaginal fluid, nasal mucus, etc. Examples of cells include red blood cells, white blood cells, platelets, and detached cells from mucous membranes such as oral cells contained in oral swabs. Examples of extracellular vesicles include exosomes and liposomes. Examples of tissue fragments include FFPE (Formalin Fixed Paraffin Embedded) specimens, biopsy specimens, and frozen specimens.
[0022] miRNA is an example of small RNA. In this embodiment, the use of miRNA will be described, but the small RNA is not limited to miRNA. Examples of small RNAs include piRNA and tsRNA.
[0023] Examples of cancers to be assessed include solid tumors and hematological cancers. Examples of solid tumors include head and neck cancer, thyroid cancer, breast cancer, lung cancer, pancreatic cancer, colorectal cancer, hepatocellular carcinoma, biliary tract cancer, stomach cancer, esophageal cancer, small intestine cancer, rectal cancer, colon cancer, kidney cancer, bladder cancer, testicular cancer, prostate cancer, ovarian cancer, uterine cancer, cervical cancer, skin cancer, brain tumors, malignant melanoma, and osteosarcoma. Examples of hematological cancers include leukemia (e.g., acute myeloid leukemia, chronic myeloid leukemia, acute lymphoblastic leukemia, and chronic lymphocytic leukemia), lymphoma (e.g., Hodgkin lymphoma and non-Hodgkin lymphoma), and multiple myeloma. The assessment may include one or more types of diseases.
[0024] A disease determination model is an example of a property determination model. In this embodiment, a disease determination model for determining whether or not a subject has cancer is described, but the property determination model is not limited to this disease determination model. The property determination model may also be used to determine whether or not a subject has a disease other than cancer (for example, diabetes), or to determine the specific properties of the sample being collected.
[0025] The properties to be determined include various properties that can be determined based on the expression level of small RNA. Examples of such properties include confirming the efficacy of the drug being collected, determining the likelihood of disease recurrence, determining lifestyle habits such as whether or not there is a history of smoking or drinking, and predicting biological age.
[0026] <Next-Generation Sequencer 21> The next-generation sequencer 21 (see Figure 1) measures the expression level of miRNAs contained in the blood 40 collected from the subject. The next-generation sequencer 21 amplifies the miRNAs extracted from the blood 40 and measures the relative expression level of each type of miRNA in relation to other miRNAs. The next-generation sequencer 21 then outputs the expression level data of the multiple miRNAs measured as miRNA data to the learning model generation device 30.
[0027] While miRNA expression levels are absolute values derived from living organisms, it is sometimes difficult to quantify the expression level in blood as an absolute value because it is necessary to quantify the expression level through measurement by measuring devices and reagent processing. Therefore, miRNA expression levels may be shown as relative values, and even such relative values still represent miRNA data. Thus, miRNA data may represent either absolute or relative values.
[0028] The next-generation sequencer 21 is an example of a measuring device. In this embodiment, the case in which the next-generation sequencer 21 is used will be described, but the measuring device is not limited to the next-generation sequencer 21. As the measuring device, a variety of measuring devices can be used, including known methods such as next-generation sequencers (NGS), DNA chips, quantitative PCR, and flow cytometry.
[0029] <Learning Model Generation Device 30> The learning model generation device 30 (see Figure 1) uses miRNA data from the next-generation sequencer 21 and information on whether or not each subject has cancer to generate a disease determination model for determining whether or not a subject has cancer.
[0030] Here, the hardware configuration of the learning model generation device 30 is shown in the block diagram in Figure 2. The learning model generation device 30 has the functionality of a computer and, as shown in Figure 2, includes a CPU (Central Processing Unit) 31, ROM (Read Only Memory) 32, RAM (Random Access Memory) 33, storage 34, input unit 35, display unit 36, and communication interface (I / F) 37. Each component is connected to the others via a bus 39 so as to be able to communicate with each other.
[0031] The CPU 31 (an example of a processor) is a central processing unit that executes various programs and controls various parts. Specifically, the CPU 31 reads a program from the ROM 32 or storage 34 and executes the program using the RAM 33 as a working area. The CPU 31 controls each of the above components and performs various calculations according to the program stored in the ROM 32 or storage 34.
[0032] ROM 32 stores various programs and data. RAM 33 temporarily stores programs or data as a working area. Storage 34 consists of an HDD (Hard Disk Drive) or SSD (Solid State Drive) and stores various programs, including the operating system, and various data.
[0033] In this embodiment, for example, a generation program for generating a disease diagnosis model is recorded in the storage 34. This generation program may be a single program or a group of programs consisting of multiple programs or modules. The generation program may also be recorded in the ROM 32. The ROM 32 and the storage 34 function as examples of non-temporary recording media.
[0034] An example of a processor is not limited to the general-purpose CPU mentioned above, but could also be a dedicated processor consisting of circuits specifically designed to perform a particular task. Furthermore, an example of a processor is not limited to a single unit, but could also be a system in which multiple units located in physically separate locations cooperate to perform the task.
[0035] The input unit 35 includes a pointing device such as a mouse and a keyboard, and is used for various types of input. The input unit 35 also accepts miRNA data as input, which represents the expression levels of multiple miRNAs measured by the next-generation sequencer 21.
[0036] The display unit 36 is, for example, a liquid crystal display, and displays various types of information. Further, the display unit 36 may function as an input unit 35 that adopts a touch panel method and inputs operations from the user.
[0037] The communication interface 37 is an interface for communicating with other devices, and transmits and receives data to and from an external device including the next-generation sequencer 21.
[0038] Next, the functional configuration of the learning model generation device 30 realized by executing the above-described generation program is shown in the block diagram of FIG. 3.
[0039] As shown in FIG. 3, in the learning model generation device 30, the CPU 31 functions as an acquisition unit 110 and a generation unit 120 by executing the generation program.
[0040] The acquisition unit 110 acquires miRNA data (hereinafter sometimes referred to as learning miRNA), which is the result of measuring the expression levels of a plurality of miRNAs in the blood of a learning subject by measuring the blood of the learning subject with the next-generation sequencer 21.
[0041] The generation unit 120 generates a disease determination model 130 for determining the presence or absence of a cancer disease in a subject based on the miRNA data (learning miRNA) acquired by the acquisition unit 110 and information indicating whether each of the learning subjects has cancer disease. Then, the disease determination model 130 generated by the generation unit 120 is stored, for example, in a non-volatile storage device.
[0042] <Method for Generating Disease Determination Model 130> Next, the method for generating the disease determination model 130 described above is shown in the flowchart of FIG. 4.
[0043] First, in the sampling step of step S101, blood is sampled from a learning subject, centrifuged, and the obtained serum is dispensed into storage tubes for each predetermined amount and stored at -80°C in a deep freezer. Note that this learning subject includes both healthy subjects who do not have cancer disease and patients who have cancer disease.
[0044] Next, in the measurement step of step S102, the serum frozen in the deep freezer is taken out, thawed at room temperature, and then the expression level of miRNA is measured by the next-generation sequencer 21.
[0045] Then, in the acquisition step of step S103, the learning model generation device 30 acquires learning miRNA data indicating the results of measuring the expression levels of a plurality of miRNAs in the blood of the learning subject measured in the measurement step of step S102.
[0046] Next, in the feature quantity selection step of step S104, the learning model generation device 30 selects, from the learning miRNA data, the feature quantities for machine learning of the learning model by a feature quantity selection method.
[0047] Here, the feature quantity selection method is shown in the flowchart of FIG. 5. As shown in FIG. 5, the feature quantity selection method includes a first selection step (step S401), a second selection step (step S402), and a third selection step (step S403). The first selection step, the second selection step, and the third selection step are executed in this order.
[0048] In the first selection step of step S401, the learning model generation device 30 selects, as the feature quantities used for machine learning of the disease determination model, the miRNAs within a preset rank from the top when the learning miRNA data is arranged in descending order of expression level. As the preset rank, for example, the 100th rank is used. Therefore, in the first selection step, for example, the top 100 types of miRNAs with high expression levels are selected from the learning miRNA data.
[0049] In the second selection step of step S402, the learning model generation device 30 selects feature quantities by removing multicollinearity from the feature quantities selected in the first selection step. Note that multicollinearity means that a certain feature quantity can be predicted by other feature quantities.
[0050] In the second selection step, for example, VIF (variance inflation factor) is used as an index of multicollinearity. Specifically, for each feature quantity selected in the first selection step, the VIF shown in the following formula is calculated.
[0051]
[0052] Ri 2 VIF is the coefficient of determination when feature i is subjected to multiple regression on all other features. Then, the feature that gives the maximum VIF is removed, and VIF is recalculated using the remaining features. This is repeated until the maximum value of VIF is less than 10.
[0053] For example, if, after calculating the VIF for each feature selected in the first selection step, the VIF for "miR-2" reaches the maximum value, as shown in Table 1 of Figure 6, then "miR-2" is deleted. If, after recalculating the VIF with the remaining features, the VIF for "miR-1" reaches the maximum value, as shown in Table 2 of Figure 6, then "miR-1" is deleted. This process is repeated until the maximum VIF value is less than 10. Finally, as shown in Table 3 of Figure 6, when the VIF is recalculated with the remaining features and the VIF for all features is less than 10, the process is terminated.
[0054] As a result, multicollinearity removal not only removes one data point from a pair of miRNAs whose correlation is already known, but also removes one data point from a pair of miRNAs whose correlation is unknown. For example, hsa-miR-493-5p and hsa-miR-432-5p are a pair of miRNAs whose correlation is unknown. Multicollinearity removal removes one data point from hsa-miR-493-5p and hsa-miR-432-5p from the miRNA data, and the other data point is selected as a feature.
[0055] Note that multicollinearity removal is one example of removing correlations between features, and while VIF was used as the multicollinearity metric, it is not limited to this. Other metrics besides VIF may also be used to remove correlations between features.
[0056] In the third selection step S403, the learning model generation device 30 further selects features from the features selected in the second selection step by performing a significance test. In this embodiment, the significance test is a significance test for the presence or absence of cancer.
[0057] In the third selection step, for example, the p-value is used as an indicator for significance testing. Specifically, features are selected using the following procedure.
[0058] 80% of the samples are randomly selected from both cancer and healthy control samples included in the training miRNA data (extraction process). Next, the p-values between the extracted cancer and healthy samples are calculated for the quantitative values of the top 100 miRNAs (calculation process). The above extraction process and calculation process are performed multiple times (for example, 6 times). The features are sorted in ascending order of the average p-values obtained multiple times. Then, from the features selected in the second selection step, for example, a threshold "p-value < 5 × 10" is set. -5 Select the features that meet the specified criteria. Note that this threshold can be set as appropriate. In addition, indicators other than the p-value may be used as indicators for significance testing.
[0059] Significance testing is one example of a test for class differences. In this embodiment, significance testing was used as an example of a class difference, but it is not limited to this. As a class difference, group difference indicators such as the F-score or Fisher score may also be used.
[0060] In the learning process of step S105, the learning model generation device 30 uses the acquired miRNA data and information indicating whether or not each training subject has cancer as training data to perform machine learning on the disease judgment model 130. Figure 7 shows the learning process of the disease judgment model 130.
[0061] As can be seen by referring to Figure 7, in the learning model generation system 20 of this embodiment, the machine learning of the disease judgment model 130 is performed by linking the expression level data of learning miRNAs collected from the blood of healthy subjects who are not suffering from cancer, and the expression level data of learning miRNAs collected from the blood of patients who are suffering from cancer, among the learning subjects, with information on whether or not the learning subjects have cancer.
[0062] Thus, the generation unit 120 generates a disease judgment model 130 by performing machine learning using the presence or absence of cancer in multiple training subjects and the expression levels of multiple miRNAs selected as training miRNAs as training data. The disease judgment model 130 is a learning model generated by performing machine learning using the presence or absence of cancer in multiple training subjects and the expression levels of multiple miRNAs selected as training miRNAs as training data. A specific example of this learning model is described below.
[0063] <Examples of Learning Models> Various linear and nonlinear algorithms known as machine learning algorithms, or combinations of multiple algorithms, can be used. For example, the following algorithms can be used.
[0064] Random forest, Gradient boosting decision trees, Extreme gradient boosting decision trees, Light-gradient boosting machine, Neural networks, Regularized regression, Elastic-net regression, K-Nearest neighbors, Support vector machine, Generalized additive model
[0065] <Method for Determining Cancer> Next, we will explain the method for determining whether or not a subject to be evaluated has cancer, using the disease determination model 130 generated by the generation method described above. The functional configuration of the disease determination device 50 for performing such disease determination is shown in the block diagram of Figure 8. The method for determining cancer is an example of a property determination method. The disease determination device 50 is an example of a property determination device.
[0066] The disease determination device 50, as shown in Figure 8, comprises an acquisition unit 110, a disease determination model 130, and a determination unit 140. Here, the disease determination model 130 in Figure 8 is a trained disease determination model 130 generated by the learning model generation device 30 shown in Figure 3.
[0067] The acquisition unit 110 acquires miRNA data, which is the result of measuring the expression levels of multiple miRNAs in the blood of a subject being assessed for the determination of whether or not they have cancer, by measuring the blood of the subject being assessed.
[0068] The determination unit 140 then inputs miRNA data, which shows the results of measuring the expression levels of multiple miRNAs in the blood collected from the subject to be determined, into the disease determination model 130, thereby determining whether or not the subject has cancer.
[0069] Thus, the disease determination device 50 has a disease determination model 130, and by inputting miRNA data showing the results of measuring the expression levels of multiple miRNAs in the blood collected from the subject into this disease determination model 130, it functions as a property determination device that determines whether or not the subject has cancer.
[0070] The disease determination device 50 then outputs the determination result from the determination unit 140 to an external device or displays it on the display unit.
[0071] Next, Figure 9 shows the flowchart of the process used to determine whether or not a patient has cancer using the disease determination model 130 that has undergone machine learning in this manner.
[0072] First, in step S201, the sampling process, blood is collected from the subject to be determined to be affected by cancer, centrifuged, and the resulting serum is dispensed into storage tubes in predetermined quantities and stored in a deep freezer at -80°C.
[0073] Next, in step S202, the measurement process involves removing the frozen serum from the deep freezer, thawing it at room temperature, and then measuring the miRNA expression level using a next-generation sequencer 21.
[0074] Then, in the acquisition process of step S203, the disease determination device 50 acquires the results of measuring the expression levels of multiple miRNAs in the subject's blood, which were measured in the measurement process of step S202, as miRNA data.
[0075] Finally, in step S204, the disease presence / absence determination process, the disease determination device 50 inputs the acquired miRNA data into the trained disease determination model 130 to obtain a determination result indicating whether or not the subject has cancer. Figure 10 shows this process for determining the presence or absence of cancer.
[0076] As can be seen by referring to Figure 10, in this embodiment, the disease determination device obtains a determination result by inputting miRNA expression data collected from the blood of a subject to be determined as having cancer into a pre-trained disease determination model 130.
[0077] <Effects of this embodiment> According to the feature selection method of this embodiment, as described above, after selecting features from miRNA data by removing multicollinearity, further features are selected from those features by significance testing.
[0078] Here, if features with multicollinearity remain, it will negatively affect the trained model. Therefore, in the feature selection method by removing multicollinearity, a threshold is set within a range that does not negatively affect the trained model. For this reason, it is difficult to control the number of remaining features in the feature selection method by removing multicollinearity.
[0079] On the other hand, even if features that do not show a significant difference remain, they are less likely to negatively impact the trained model. Therefore, the feature selection method using significance testing offers a high degree of flexibility in setting thresholds. For this reason, the feature selection method using significance testing makes it easier to control the number of remaining features.
[0080] As a result, when features are selected from miRNA data using significance testing, and then further selected from those features by removing multicollinearity, the number of remaining features may become extremely small.
[0081] In contrast, in this embodiment, feature selection is performed in the order of multicollinearity removal followed by significance testing, making it easier to control the number of remaining features.
[0082] <Performance Evaluation> The performance of the disease diagnosis model, which was machine-learned using the features selected by the feature selection method according to the embodiment described above, was confirmed.
[0083] In this evaluation, we assessed disease detection models that were machine-learned using features selected according to the following conditions 1 to 4. A logistic regression model with L2 regularization was generated as the disease detection model. As an indicator of overfitting, we plotted the difference in scores between identical healthy subjects measured at the time of model generation and two months later. In the evaluation results shown in Figures 11 to 14, identical healthy subjects are connected by lines.
[0084] Condition 1: Perform only the first selection step (selecting the top 100 miRNAs with the highest expression levels). Condition 2: Perform only the first and second selection steps (removing multicollinearity). Condition 3: Perform only the first and third selection steps (significance testing). Condition 4: Perform the first, second, and third selection steps.
[0085] Under condition 1, there were 100 features; under condition 2, there were 49 features; under condition 3, there were 22 features; and under condition 4, there were 12 features.
[0086] The evaluation results, as shown in Figure 11, show that under Condition 1, the score calculated from the measurement data two months later, including the difference between measurements, differs significantly from the score calculated from the measurement data at the time of model generation. Similar trends are observed under Conditions 2 (see Figure 12) and 3 (see Figure 13). In contrast, as shown in Figure 14, the score difference due to the difference between measurements decreases under Condition 4. Therefore, it can be seen that robustness was achieved as a result of selecting features using the feature selection method according to this embodiment.
[0087] <Modification> In the above embodiment, a model was generated to determine whether a dataset is cancerous or non-cancer from miRNA data, but the model is not limited to this.
[0088] For example, a model that performs multiple classifications from a single dataset may be generated. Specifically, as shown in Figure 15, it is possible to generate a model that performs classification 1 (cancer vs. severe disease), classification 2 (severe disease vs. mild disease), and classification 3 (mild disease vs. healthy individuals). In this case, as shown in Figure 16, the features selected by multicollinearity removal are subjected to feature selection by significance testing for classification 1 (called the first selection process), feature selection by significance testing for classification 2 (called the second selection process), and further feature selection by significance testing for classification 3 (called the third selection process). Therefore, the VIF calculation in the selection by multicollinearity removal is completed only once.
[0089] In contrast, when features are selected from miRNA data using significance testing, and then further selected from those features by removing multicollinearity, the computational load increases because VIF calculation is required for each classification, as shown in Figure 17.
[0090] As described above, the feature selection method of this embodiment, which selects features in the order of removing multicollinearity and then performing significance testing, can reduce the computational cost when generating a model that performs multiple classifications from a single dataset. The first selection process is an example of feature selection by significance testing for the first property. The second and third selection processes are examples of feature selection by significance testing for the second property.
[0091] The present invention is not limited to the embodiments described above, and various modifications, changes, and improvements are possible without departing from its spirit. The aforementioned modifications may be combined in any way appropriate.
[0092] <Notes> (Aspect 1) A feature selection method for selecting features to be machine-learned by a learning model from small RNA data showing measurement results of the expression level of small RNA in a biological sample, wherein after selecting features from the small RNA data by removing correlations between features, further features are selected from those features by inter-class differences. (Aspect 2) The feature selection method according to Aspect 1, wherein the small RNA data is data showing measurement results measured by a next-generation sequencer. (Aspect 3) The feature selection method according to Aspect 1 or Aspect 2, wherein the inter-class differences are inter-class differences regarding the presence or absence of cancer disease. (Aspect 4) The feature selection method according to any one of Aspects 1 to 3, wherein the small RNA is microRNA, and by removing correlations between features, one of the data of hsa-miR-493-5p and hsa-miR-432-5p is removed from the small RNA data, and the other data is selected as a feature. (Aspect 5) A method for generating a property determination model, which involves selecting features from small RNA data showing measurement results of the expression level of small RNA in a biological sample by removing correlations between features, further selecting features from those features based on interclass differences, and then performing machine learning using those features to generate a property determination model for determining the properties of the sample from which the biological sample was collected. (Aspect 6) The method for generating a property determination model according to Aspect 5, wherein the features selected by removing correlations between features are further selected based on interclass differences for at least two or more properties. (Aspect 6-1) The method for generating a property determination model according to Aspect 5, wherein the features selected by removing correlations between features are further selected based on interclass differences for a first property, and also for a second property. In Aspect 6-1, the selection of features based on interclass differences is not limited to only two properties, the first and the second property, but may be performed for three or more properties.(Aspect 7) A property determination model for determining the properties of a sample from which a biological sample was collected, which is generated by performing machine learning using features selected from small RNA data showing the results of measuring the expression level of small RNA in a biological sample, wherein the features are selected by removing correlations between features and then further selected from those features by interclass differences. (Aspect 8) A property determination method for determining the properties of a sample to be determined by inputting small RNA data showing the results of measuring the expression level of small RNA in a biological sample to be determined into the property determination model described in Aspect 7. (Aspect 9) A property determination device having the property determination model described in Aspect 7, which determines the properties of a sample to be determined by inputting small RNA data showing the results of measuring the expression level of small RNA in a biological sample to be determined into the property determination model.
[0093] The disclosure of Japanese Patent Application No. 2024-192437, filed on 31 October 2024, is incorporated herein by reference in its entirety. All documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
Claims
1. A method for selecting features to be used in machine learning for a learning model from small RNA data showing measurement results of the expression level of small RNA in biological samples, wherein the feature selection method involves selecting features from the small RNA data by removing correlations between features, and then further selecting features from those features based on inter-class differences.
2. The feature selection method according to claim 1, wherein the small RNA data is data showing measurement results measured by a next-generation sequencer.
3. The feature selection method according to claim 1, wherein the inter-class difference is the inter-class difference regarding the presence or absence of cancer.
4. The feature selection method according to claim 1, wherein the small RNA is a microRNA, and by removing the correlation between the features, one of the data, hsa-miR-493-5p and hsa-miR-432-5p, is removed from the small RNA data, and the other data is selected as a feature.
5. A method for generating a property determination model, which involves selecting features from small RNA data showing measurement results of the expression level of small RNA in a biological sample by removing correlations between features, further selecting features from those features based on class differences, and then performing machine learning using those features to generate a property determination model for determining the properties of the sample from which the biological sample was collected.
6. A method for generating a property determination model according to claim 5, wherein the selected features are determined by removing the correlation between the features, and the selection of features is made based on the inter-class differences for each of at least two of the properties.
7. A property determination model that is generated by performing machine learning using features selected from small RNA data showing measurement results of the expression level of small RNA in a biological sample, and which determines the properties of the subject from which the biological sample was collected, wherein the features are selected by removing correlations between features and then further selected from those features by class differences.
8. A property determination method for determining the properties of a biological sample to be determined by inputting small RNA data, which shows the result of measuring the expression level of small RNA in a biological sample to be determined, into the property determination model described in claim 7.
9. A property determination device having the property determination model described in claim 7, wherein the property of a target is determined by inputting small RNA data, which shows the result of measuring the expression level of small RNA in a biological sample to be determined, into the property determination model.