Determination method, specimen data imputation method, specimen data estimation method, and determination system

The method addresses accuracy issues in disease discrimination models by using a missing value imputation model to enhance prediction accuracy, even when specimen data is incomplete, by filling in missing values based on the relationship between small RNA expression data and other features.

WO2026094874A1PCT designated stage Publication Date: 2026-05-07ARKRAY INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ARKRAY INC
Filing Date
2025-10-27
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing disease discrimination models using machine learning face accuracy issues when specimen data contains missing values, particularly in features other than small RNA expression levels.

Method used

A method and system that complements missing values in specimen data using a missing value imputation model, learned from the relationship between small RNA expression data and other features, to enhance prediction accuracy in disease discrimination models.

Benefits of technology

Improves the accuracy of disease prediction by filling in missing data points, allowing the model to effectively utilize small RNA expression data even when other features are incomplete.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025037657_07052026_PF_FP_ABST
    Figure JP2025037657_07052026_PF_FP_ABST
Patent Text Reader

Abstract

This determination method comprises: determining whether specimen data including small RNA expression data as feature values contains a missing item with respect to feature values other than the small RNA expression data; if the specimen data includes a missing item, imputating the missing value of the specimen data on the basis of a missing value imputation model trained on the relationship between an index item and feature values included in the specimen data; and inputting the specimen data in which the missing value has been imputated to a determination model that determines a disease based on the small RNA expression data and feature values other than the small RNA expression data, thereby determining a disease of a subject associated with the specimen data.
Need to check novelty before this filing date? Find Prior Art

Description

Discrimination method, sample data supplementation method, sample data estimation method, and discrimination system

[0001] This disclosure relates to a discrimination method, a method for supplementing sample data, a method for estimating sample data, and a discrimination system.

[0002] For example, prior art documents include Munenori Kawai, Akihisa Fukuda, Ryo Otomo, Shunsuke Obata, Kosuke Minaga, Masanori Asada, Atsushi Umemura, Yoshito Uenoyama, Nobuhiro Hieda, Toshihiro Morita, Ryuki Minami, Saiko Marui, Yuki Yamauchi, Yoshitaka Nakai, Yutaka Takada, Kozo Ikuta, Takuto Yoshioka, Kenta Mizukoshi, Kosuke Iwane, Go Yamakawa, Mio Namikawa, Makoto Sono, Munemasa Nagao, Takahisa Maruno, Yuki Nakanishi, Mitsuharu Hirai, Naoki Kanda, Seiji Shio, Toshinao Itani, Shigehiko Fujii, Toshiyuki Kimura, Kazuyoshi Matsumura, Masaya Ohana, Shujiro Yazumi, Chiharu Kawanami, Yukitaka Yamashita, Hiroyuki Marusawa, Tomohiro Watanabe, Yoshito Ito, Masatoshi Kudo and Hiroshi Seno, "Early detection "of pancreatic cancer by comprehensive serum miRNA sequencing with automated machine learning," British Journal of Cancer, [online], published August 28, 2024, Internet.<URL: https: / / www.nature.com / articles / s41416-024-02794-5> There is.

[0003] In prior art documents, a discrimination model for discriminating pancreatic cancer has been created by performing machine learning using a plurality of miRNAs and CA19-9 for healthy individuals and pancreatic cancer patients. According to the created discrimination model, healthy individuals and pancreatic cancer patients are discriminated with high accuracy.

[0004] When discriminating the disease of a subject using a discrimination model created by performing machine learning for discriminating a disease using the small RNA expression level data and the feature amounts included in the specimen data such as a questionnaire, a survey form, and test values, there may be a case where the item of the feature amount of the subject that should be included in the specimen data is missing.

[0005] The present disclosure provides a discrimination method, a specimen data complement method, a specimen data estimation method, and a discrimination system for improving the prediction accuracy of a disease based on a discrimination model when there are missing values in the specimen data.

[0006] The discrimination method of the property determination model according to one aspect of the present disclosure includes determining whether there is an item missing for a feature amount other than the small RNA expression level data in the specimen data in which the small RNA expression level data is included in the feature amounts, and when the specimen data includes the missing item, complementing the missing value of the specimen data based on a missing value complement model learned about the relationship between the index item included in the specimen data and the feature amounts, and inputting the specimen data with the missing value complemented into a discrimination model for discriminating a disease based on the small RNA expression level data and the feature amounts other than the small RNA expression level data, and discriminating the disease of the subject related to the specimen data.

[0007] According to the present disclosure, there are provided a discrimination method, a specimen data complement method, a specimen data estimation method, and a discrimination system for improving the prediction accuracy of a disease based on a discrimination model when there are missing values in the specimen data.

[0008] This figure shows the system configuration of a disease discrimination system 20 according to one embodiment of the present disclosure. This is a block diagram showing the configuration of a learning model generation device 30. This is a flowchart showing the method for generating a disease discrimination model 52. This figure shows the machine learning process of the disease discrimination model 52. This is a flowchart showing the method for generating a missing value imputation model 54. This figure shows the machine learning process of the missing value imputation model 54. This is a flowchart showing the process for determining whether or not a person has cancer using the disease discrimination model 52 and the missing value imputation model 54 that have undergone machine learning.

[0009] Next, embodiments of the present disclosure will be described in detail with reference to the drawings.

[0010] An example of an embodiment relating to the technology of this disclosure will be described below with reference to the drawings. Components and processes that perform the same operation, action, or function will be given the same reference numerals throughout the drawings, and redundant explanations may be omitted as appropriate. Each drawing is only a schematic representation to the extent that the technology of this disclosure can be fully understood. Therefore, the technology of this disclosure is not limited to the illustrated examples. Furthermore, in this embodiment, explanations of configurations not directly related to this disclosure or well-known configurations may be omitted.

[0011] Figure 1 shows the system configuration of a disease discrimination system 20 according to one embodiment of the present disclosure. The disease discrimination system 20 of this embodiment is an example of a "discrimination system" in this disclosure. Furthermore, the disease discrimination system 20 of this embodiment generates a disease discrimination model 52 by the method for generating a disease discrimination model 52 described below.

[0012] The disease discrimination model 52 generation method of this embodiment is a method for generating a trained disease discrimination model 52 that determines the presence or absence of a disease in a subject based on the expression levels of multiple small RNAs. This is achieved by performing machine learning using multiple small RNAs detected from a biological sample such as the subject's blood as training data.

[0013] The disease discrimination system 20 of this embodiment consists of a next-generation sequencer 21 using NGS (Next-Generation Sequencing) technology and a learning model generation device 30, as shown in Figure 1.

[0014] In this embodiment, the disease discrimination system 20 may use a measuring device such as a quantitative PCR (Polymerase Chain Reaction) or a flow cytometer instead of a next-generation sequencer 21. In other words, the disease discrimination system 20 of this embodiment only needs to be equipped with a device that measures the expression levels of multiple nucleic acid molecules.

[0015] The disease discrimination system 20 of this embodiment generates a disease discrimination model 52 that can determine whether or not a subject has cancer, using blood 40 collected from the subject to be examined for the presence or absence of cancer.

[0016] This explanation uses the case of determining the presence or absence of cancer using blood collected from the subject, but the presence or absence of cancer may also be determined using other biological samples such as body fluids, cells, extracellular vesicles, tissue fragments, etc. Here, body fluids include, for example, serum, plasma, urine, tears, saliva, sweat, semen, lymph, interstitial fluid, body cavity fluids (e.g., pleural fluid, ascites, etc.), cerebrospinal fluid, amniotic fluid, vaginal fluid, nasal mucus, etc. Cells include, for example, red blood cells, white blood cells, platelets, oral swab fluid, etc. Extracellular vesicles include, for example, exosomes, liposomes, etc. Tissue fragments include, for example, FFPE (Formalin Fixed Paraffin Embedded) specimens, biopsy specimens, frozen specimens, etc.

[0017] Furthermore, although this embodiment uses the case where the subject is a human as an example, the subject of the examination is not limited to humans, and this disclosure is equally applicable when the subject of the examination is various animals other than humans, such as dogs and cats. In other words, this disclosure is also applicable when examining the presence or absence of cancer or other diseases using biological samples such as blood collected from various animals such as dogs and cats.

[0018] Furthermore, in this embodiment, the method for generating the disease discrimination model 52 is described using the case of generating a disease discrimination model 52 for determining whether or not a subject has cancer, but this disclosure is not limited to such a case. This disclosure is also applicable when generating a disease discrimination model 52 for determining whether or not a subject has diseases other than cancer, or when generating a property determination model for determining specific properties of a subject. The properties of a subject to be determined by the property determination model are not limited to the subject's disease. The properties of a subject to be determined can be any properties that can be determined based on the expression levels of multiple small RNAs. Specific properties to be determined include, for example, confirmation of the efficacy of a drug for a subject, determination of the likelihood of disease recurrence, confirmation / determination of lifestyle habits such as the presence or absence of smoking and drinking history, and prediction of biological age. It should be noted that "discrimination" or "determination" in this disclosure is not limited to binary classification such as the presence or absence of a disease, but also includes classifying the possibility of having a disease into three or more classes, or representing the possibility of having a disease as a continuous value such as a probability.

[0019] The next-generation sequencer 21 measures the expression levels of multiple small RNAs contained in the blood 40 collected from the subject. Specifically, the next-generation sequencer 21 amplifies the small RNAs extracted from the blood 40 and then measures the relative expression levels of each type of small RNA in relation to other small RNAs.

[0020] Here, small RNAs include, for example, microRNAs (hereinafter referred to as miRNAs). Note that small RNAs may also include small RNAs other than miRNAs (e.g., piRNAs (PIWI-interacting RNAs), tsRNAs (transfer RNA-derived small RNAs), siRNAs (small interfering RNAs), snRNAs (small nuclear RNAs), and snoRNAs (small nucleolar RNAs)), and fragmented RNAs (e.g., mRNAs (messenger RNAs), tRNAs (transfer RNAs), rRNAs (ribosomal RNAs)). The following explanation will use the case where the expression level of miRNAs is measured from the subject's blood 40 to determine whether or not they have cancer.

[0021] The next-generation sequencer 21 then outputs the measured small RNA expression data as sample data 12 to the learning model generation device 30.

[0022] The learning model generation device 30 uses sample data 12 from the next-generation sequencer 21 and information on whether or not each subject has cancer to generate a disease discrimination model 52 for determining whether or not a subject has cancer.

[0023] Figure 2 shows a block diagram of the hardware configuration of the learning model generation device 30 described above.

[0024] The learning model generation device 30 has computer-like functions and, as shown in Figure 2, includes a CPU (Central Processing Unit) 41, RAM (Random Access Memory) 42, flash memory 43, ROM (Read Only Memory) 44, input device 34, display unit 36, and communication unit 38. Each component is connected to the others via an input / output interface (I / O) 46 and a bus 45 so as to be able to communicate with each other.

[0025] The CPU 41 (an example of a processor) is a central processing unit that executes various programs and controls various parts. Specifically, the CPU 41 reads a program from the flash memory 43 or ROM 44 and executes the program using the RAM 42 as a working area. The CPU 41 controls each of the above components and performs various calculations according to the program stored in the flash memory 43 or ROM 44.

[0026] The flash memory 43 stores various programs and data. The flash memory 43 also stores the disease discrimination model 52 and the missing value imputation model 54, which will be described later. The RAM 42 temporarily stores programs or data as a working area. The ROM 44 is composed of an HDD (Hard Disk Drive), etc., and records various programs, including the operating system, and various data.

[0027] In this embodiment, for example, a generation program for generating a disease discrimination model 52 and a missing value imputation model 54 is recorded in ROM 44. This generation program may be a single program or a group of programs consisting of multiple programs or modules. The generation program may also be recorded in flash memory 43. Flash memory 43 and ROM 44 function as examples of non-temporary recording media.

[0028] An example of a processor is not limited to the general-purpose CPU mentioned above, but could also be a dedicated processor consisting of circuits specifically designed to perform a particular task. Furthermore, an example of a processor is not limited to a single unit, but could also be a system in which multiple units located in physically separate locations cooperate to perform the task.

[0029] The input device 34 includes a pointing device such as a mouse and a keyboard, and is used for various types of input.

[0030] The display unit 36 ​​is, for example, a liquid crystal display and displays various information. Alternatively, the display unit 36 ​​may employ a touch panel system and function as an input device 34 for receiving user input.

[0031] The communication unit 38 is an interface for communicating with other devices and transmits and receives data with external devices, including the next-generation sequencer 21. More specifically, the communication unit 38 receives sample data 12 as input, which represents the expression levels of multiple small RNAs measured by the next-generation sequencer 21.

[0032] In this embodiment, the learning model generation device 30 generates the disease discrimination model 52 by having the CPU 41 execute a generation program (not shown) recorded in the ROM 44.

[0033] One example of a disease that can be identified is cancer. Examples of cancers include oral cancer, pharyngeal cancer, lung cancer, esophageal cancer, stomach cancer, colon cancer, rectal cancer, colorectal cancer, liver cancer, gallbladder cancer, bile duct cancer, pancreatic cancer, laryngeal cancer, skin cancer, breast cancer, prostate cancer, kidney cancer, urinary tract cancer, bladder cancer, brain tumor, thyroid cancer, malignant lymphoma, multiple myeloma, leukemia, uterine cancer, cervical cancer, endometrial cancer, and ovarian cancer. Alternatively, other diseases may be used as examples of diseases that can be identified instead of cancer. Examples of other diseases include Alzheimer's disease and diabetes.

[0034] Furthermore, as shown in Figure 1, the input device 34 is capable of inputting the subject's medical history data 16 and clinical test data. More specifically, the results of each item in the medical history questionnaire (not shown) that the subject has answered in advance and the measured clinical test values ​​are input by the input device 34. The medical history questionnaire has fields for the subject to fill in regarding matters other than the small RNA expression level. More specifically, the medical history questionnaire has fields for the subject to fill in their gender, age, or active smoking status. The information entered in the medical history questionnaire is then input into the disease discrimination system 20 by the input device 34. The clinical test values ​​are the results of clinical tests that have been measured in advance or will be newly measured regarding matters other than the small RNA expression level in the subject. The results of the clinical tests are then input into the disease discrimination system 20 by the input device 34.

[0035] The questionnaire may also include information on the subject's active smoking status (e.g., pack-years, Brinkman index). Furthermore, the questionnaire may include information on the subject's passive smoking status, alcohol consumption (e.g., weekly alcohol consumption), obesity level (weight, height, BMI, waist circumference, etc.), exercise status, nutrition / dietary intake, history of viral infections, history of illness, family history of illness, and history of chemical exposure. These values ​​may be recorded as quantitative or qualitative data.

[0036] Furthermore, clinical laboratory values ​​may include tumor marker measurements such as CA19-9 and CEA for the subject. Clinical laboratory values ​​may also include the results of specimen tests, biopsies, and physiological function tests of the subject. These values ​​may be measured as quantitative or qualitative data.

[0037] (Procedure for generating a disease diagnosis model) Next, the procedure for generating the disease diagnosis model 52 will be explained with reference to the flowchart in Figure 3.

[0038] The learning model generation device 30 generates a disease discrimination model 52 that determines whether or not a subject has cancer, based on learning miRNA, which is training subject sample data 12 obtained from the next-generation sequencer 21, information indicating whether or not each training subject (training subject) has cancer, and the results of each training subject's answers to a questionnaire, by performing machine learning.

[0039] First, in step S102, the blood collection process involves collecting blood from the learning subjects, centrifuging it, and then dispensing the resulting serum into storage tubes in predetermined quantities, which are then stored in a deep freezer at -80°C. These learning subjects include both healthy individuals who do not have cancer and patients who do have cancer.

[0040] Next, in step S104, the measurement process, the frozen serum is removed from the deep freezer, thawed at room temperature, and then the expression level of small RNAs is measured using a next-generation sequencer 21.

[0041] Then, in the acquisition step S106, the learning model generation device 30 acquires the results of measuring the expression levels of multiple small RNAs in the serum of the training subject, which were measured in the measurement step S104. More specifically, in the acquisition step S106, the learning model generation device 30 acquires the expression levels of small RNAs as sample data 12 from the next-generation sequencer 21.

[0042] Furthermore, in the acquisition process of step S106, the learning model generation device 30 further acquires disease information 14, which is information indicating whether or not each learning subject has cancer. Also in the acquisition process of step S106, the learning model generation device 30 further acquires questionnaire data 16, which is the result of each learning subject's answers to the questionnaire.

[0043] Finally, in the learning process of step S108, the learning model generation device 30 performs machine learning of the disease discrimination model 52 using the acquired sample data 12, disease information 14, and medical interview data 16 as training data. Figure 4 shows the machine learning process of the disease discrimination model 52.

[0044] As can be seen by referring to FIG. 4, in the disease discrimination system 20 of the present embodiment, machine learning of the disease discrimination model 52 is executed by linking the expression level data of the learning small RNA collected from the blood of the learning subject, the information on whether or not the learning subject has cancer, and the response results of the questionnaire of the learning subject.

[0045] Thus, the learning model generation device 30 generates the disease discrimination model 52 by performing machine learning using the expression levels of the small RNAs of a plurality of learning subjects, the presence or absence of cancer in the learning subjects, and the response results of the questionnaires of the learning subjects as teacher data. Then, the disease discrimination model 52 generated by the learning model generation device 30 is stored in the flash memory 43.

[0046] Various linear and non-linear algorithms known as machine learning algorithms, or a combination of a plurality of algorithms can be used. Any algorithm may be used, examples of which include random forest, gradient boosting trees, extreme gradient boosting trees, light-gradient boosting machine, neural networks, regularized regression, elastic-net regression, K-nearest neighbors, support vector machine, and generalized additive model.

[0047] Then, the created disease discrimination model 52 is used to discriminate whether or not the subject has cancer from the small RNA contained in the subject's blood sample and the items described in the questionnaire. More specific discrimination procedures will be described later.

[0048] In addition, in creating the disease discrimination model 52, only one item of the questionnaire used may be used, or a plurality of items may be combined and used.

[0049] (Procedure for generating the missing value completion model) Next, the procedure for generating the missing value completion model 54 will be described while referring to the flowchart of FIG. 5.

[0050] The learning model generation device 30 generates a missing value completion model 54 that estimates the items described in the questionnaire of the subject based on the learning small RNA obtained from the next-generation sequencer 21 and the results of the respective learning subjects answering the questionnaire by performing machine learning.

[0051] First, the collection step of step S202 and the measurement step of step S204 are the same as the collection step of step S102 and the measurement step of step S104, respectively.

[0052] Then, in the acquisition step of step S206, the learning model generation device 30 acquires the results of measuring the expression levels of a plurality of small RNAs in the serum of the learning subject measured in the measurement step of step S204.

[0053] In addition, in the acquisition step of step S206, the learning model generation device 30 further acquires the questionnaire data 16, which is the result of each learning subject answering the questionnaire.

[0054] Finally, in the learning step of step S208, in the learning model generation device 30, machine learning of the missing value completion model 54 is executed using the acquired sample data 12 and questionnaire data 16 as teacher data.

[0055] As can be seen by referring to FIG. 6, in the disease discrimination system 20 of the present embodiment, machine learning of the missing value completion model 54 is executed in cooperation with the expression level data of the learning small RNA collected from the blood of the learning subject and the questionnaire answer results of the learning subject.

[0056] In this manner, the learning model generation device 30 generates a missing value imputation model 54 by performing machine learning using the expression levels of small RNAs from multiple training subjects and the answers to the training subjects' questionnaires as training data. The missing value imputation model 54 generated by the learning model generation device 30 is then stored in the flash memory 43. In other words, in this embodiment, the expression levels of small RNAs and the questionnaire data 16 are examples of "sample data". Furthermore, the disease discrimination model 52 determines whether a subject has cancer or not based on the small RNA expression level data included in the sample data 12 and the answers to the questionnaire included in the questionnaire data 16. In other words, in this embodiment, the small RNA expression level data is an example of an "indicator item".

[0057] Incidentally, subjects do not necessarily answer all of the items included in the questionnaire used for the machine learning described above. Reasons for not answering the questionnaire include human factors such as the subject's lack of awareness, as well as changes in the examination method, such as changes in the content of the questionnaire. In other words, there may be missing items in the features used for training to create the disease discrimination model 52.

[0058] Next, the operation procedure of the CPU 41 in this embodiment will be described with reference to Figure 7.

[0059] (Supplementary Procedure) First, in step S302, the CPU 41 acquires the sample data 12 and the medical history data 16. More specifically, the CPU 41 acquires the small RNA expression level and the information recorded on the medical history form from the subject's blood sample, similar to the acquisition process in step S106. In other words, the CPU 41 performs the role of an example of an "input unit" in this disclosure through the operation in step S302. Then, the CPU 41 proceeds to step S304.

[0060] Next, in step S304, the CPU 41 determines whether there are any missing items in the medical history data 16. More specifically, the CPU 41 determines in step S302 whether the information from the medical history form used in the machine learning of the disease discrimination model 52 has been obtained. That is, if multiple items from the medical history form are used in the machine learning of the disease discrimination model 52, and any one of those items is missing, the CPU 41 makes a positive determination in step S304.

[0061] If the CPU 41 determines that the result is positive in step S304, it proceeds to step S306. If the CPU 41 determines that the result is negative in step S304, it proceeds to step S308.

[0062] Next, in step S306, the CPU 41 fills in the missing items in the medical history data 16. More specifically, in step S306, the CPU 41 inputs the small RNA expression levels contained in the sample data 12 obtained in step S302 into the missing value completion model 54, thereby filling in the specific values ​​of the missing items in the medical history data 16. In other words, through the operation in step S306, the CPU 41 plays the role of an example of a "missing value completion unit" in this disclosure. The CPU 41 then temporarily stores the specific values ​​of the missing items in the RAM 42 and proceeds to step S308.

[0063] Next, in step S308, the CPU 41 inputs the values ​​of each item included in the medical history data 16 and the small RNA expression levels included in the sample data 12 into the disease discrimination model 52. Note that the medical history data 16 in step S308 also includes the values ​​of the missing medical history data 16 of the subject that were supplemented in step S306. The CPU 41 then determines whether or not the subject has a disease based on the disease discrimination model 52. In other words, the CPU 41 performs the role of an example of a "discrimination unit" in this disclosure through the operation in step S308. The CPU 41 then proceeds to step S310.

[0064] Next, in step S310, the CPU 41 outputs the discrimination result from the disease discrimination model 52 to the display unit 36. If the CPU 41 has imputed missing values ​​in step S306, it may include in the output to the display unit 36 ​​in step S310 that such imputation has been performed.

[0065] In this embodiment, the CPU 41 improves the accuracy of disease prediction based on the disease discrimination model 52 when there are missing values ​​in the sample data, based on the procedure described above.

[0066] Next, the operation and effects of the disease discrimination system 20 in this embodiment will be explained.

[0067] (Effects and Effects) In this embodiment of the discrimination method, sample data with missing items for features other than small RNA expression data is supplemented based on the missing value imputation model 54 and then input into the disease discrimination model 52. Therefore, according to this embodiment of the discrimination method, by predicting and imputing missing values ​​using the learned missing value imputation model 54, the accuracy of disease prediction based on the disease discrimination model 52 can be improved compared to inputting the sample data with missing items into the disease discrimination model 52 without imputation.

[0068] Furthermore, in the discrimination method according to this embodiment, features are calculated using small RNA expression data. Therefore, according to the discrimination method according to this embodiment, the accuracy of disease prediction can be improved without using data related to items other than small RNA expression data. In other words, according to the discrimination method according to this embodiment, the accuracy of prediction can be improved even when items other than small RNA expression data are missing from the sample data.

[0069] Furthermore, in the sample data completion method according to this embodiment, missing values ​​in sample data that have missing features other than small RNA expression data are completed based on a missing value completion model 54 that has learned the relationship between small RNA expression data and the features of those items. Therefore, according to the sample data completion method according to this embodiment, missing values ​​can be completed even if there are missing features in the sample data other than small RNA expression data.

[0070] Furthermore, in the sample data estimation method according to this embodiment, the features used to distinguish a specific disease are estimated based on the small RNA expression level data of the subject being tested for that specific disease. Therefore, according to the sample data estimation method according to this embodiment, even if there are missing items in the sample data for features other than small RNA expression level data, the features of the subject can be estimated.

[0071] Next, as an example, we will demonstrate with experimental data that when there are missing items in the sample data, the accuracy of the disease discrimination model 52's judgment results increases by imputing the missing values ​​using a missing value imputation model.

[0072] (Example 1) In this experiment, blood samples were first collected from 142 healthy subjects and 109 lung cancer patients to obtain training miRNAs. In addition, in this example, the smoking status of the subjects was added as a qualitative feature to generate a disease discrimination model 52. In this example, the "smoking status" item represents either a result of "smoking" or "not smoking" for the subject. A total of 107 types of miRNAs were used to generate the disease discrimination model 52.

[0073] The specific method for generating the disease discrimination model 52 is as described in steps S102 to S108 above.

[0074] Furthermore, in this embodiment, as qualitative data, a missing value completion model 54 was generated to supplement the subject's smoking status based on miRNA contained in the subject's blood data. The type of miRNA used to generate the missing value completion model 54 is the same as the type used to generate the disease discrimination model 52. In other words, the missing value completion model 54 is a learning model that takes miRNA as input data and outputs smoking status.

[0075] The specific method for generating the missing value imputation model 54 is as described in steps S202 to S208 above. In this embodiment, the AUC (Area Under the Curve) of the missing value imputation model 54 generated was 0.759.

[0076] Next, the generated disease discrimination model 52 and missing value imputation model 54 were validated by collecting blood samples from 62 healthy individuals and 47 lung cancer patients, which were used as sample data with missing features. It should be noted that the 62 healthy individuals and 47 lung cancer patients used for validation were all different from the training subjects used to generate the disease discrimination model 52 and missing value imputation model 54.

[0077] Table 1 shows the results of determining whether or not lung cancer was present, after supplementing the missing values ​​of the subjects' smoking status based on the generated missing value imputation model 54.

[0078]

[0079] Regarding the test examples in Table 1, these examples represent the results when the smoking status of the subjects was omitted and then supplemented, as described above. In other words, the test examples represent the verification results when a positive judgment was made in all cases at step S304 in Figure 7.

[0080] In Table 1, Comparative Example 1 involves inputting the subject's correct smoking status (i.e., no missing values). Comparative Example 2 involves inputting all smoking statuses as "smoker". Comparative Example 3 involves inputting all smoking statuses as "non-smoker". Comparative Example 4 involves inputting the subject's correct smoking status in reverse (i.e., "non-smoker" is entered when the subject smokes, and "smoker" is entered when the subject does not smoke).

[0081] Furthermore, each model in Table 1 refers to the generation algorithm for the lung cancer discrimination model. The numerical values ​​for the imputation methods indicate cases where the judgment result was incorrect. In other words, each numerical value listed in Table 1 represents the total number of false positives (i.e., cases where lung cancer was diagnosed when the patient did not have lung cancer) and false negatives (i.e., cases where lung cancer was diagnosed when the patient did have lung cancer) as a result of the lung cancer discrimination model.

[0082] Table 1 shows that the test example has more errors than Comparative Example 1. However, the test example has fewer errors than Comparative Examples 2 through 4. This suggests that even when there are qualitative data items with missing values, the accuracy of the disease discrimination model 52 can be improved by using the missing value imputation model 54.

[0083] (Example 2) In this experiment, blood samples were collected from 142 healthy subjects and 109 lung cancer patients to obtain training miRNAs. In this example, the smoking status of the subjects was also added as a quantitative feature to generate a disease discrimination model 52. In this example, the "smoking status" item is the pack-years value of the subjects. A total of 107 types of miRNAs were used to generate the disease discrimination model 52.

[0084] The specific method for generating the disease discrimination model 52 is as described in steps S102 to S108 above.

[0085] Furthermore, in this embodiment, as quantitative data, a missing value completion model 54 was generated to supplement the subject's smoking status based on miRNA contained in the subject's blood data. The type of miRNA used to generate the missing value completion model 54 is the same as the type used to generate the disease discrimination model 52. In other words, the missing value completion model 54 is a learning model that takes miRNA as input data and outputs smoking status.

[0086] The specific method for generating the missing value imputation model 54 is as described in steps S202 to S208 above. In this embodiment, the correlation coefficient of the missing value imputation model 54 generated was 0.487.

[0087] Next, the generated disease discrimination model 52 and missing value imputation model 54 were validated by collecting blood samples from 62 healthy individuals and 47 lung cancer patients, which were used as sample data with missing features. It should be noted that the 62 healthy individuals and 47 lung cancer patients used for validation were all different from the training subjects used to generate the disease discrimination model 52 and missing value imputation model 54.

[0088] Table 2 shows the results of determining the presence or absence of lung cancer after supplementing the missing values ​​of the subjects' smoking status based on the generated missing value imputation model 54.

[0089]

[0090] Regarding the test examples in Table 2, these examples represent the results when the smoking status of the subjects was omitted and then supplemented, as described above. In other words, the test examples represent the verification results when a positive judgment was made in all cases at step S304 in Figure 7.

[0091] In Table 2, Comparative Example 1 uses data where the subject's correct smoking status is entered (i.e., there are no missing values). Comparative Example 2 uses data where the smoking status is entered as the median of the subject's pack-years. Comparative Example 3 uses data where the smoking status is entered as the average of the subject's pack-years.

[0092] Furthermore, each model in Table 2 refers to the generation algorithm for the lung cancer discrimination model. The numerical values ​​for the imputation methods indicate cases where the judgment result was incorrect. In other words, each numerical value listed in Table 2 represents the total number of false positives (i.e., cases where lung cancer was diagnosed when the patient did not have lung cancer) and false negatives (i.e., cases where lung cancer was diagnosed when the patient did have lung cancer) as a result of the lung cancer discrimination model.

[0093] Table 2 shows that the test example has more errors than Comparative Example 1. However, the test example has fewer errors than Comparative Examples 2 and 3. This suggests that even when there are quantitative data items with missing values, the accuracy of the disease discrimination model 52 can be improved by using the missing value imputation model 54.

[0094] In the example described above, the case in which miRNA expression data in sample data 12 was used as an indicator item was explained, but the technology in this disclosure is not limited to this. That is, the disease discrimination model 52 may be machine-learned using subject sample data other than miRNA expression data. For example, multiple small RNAs other than miRNA may be used as indicator items in machine learning.

[0095] Furthermore, in the example described above, a missing value imputation model 54 was created using miRNA expression data as training data, but the technology in this disclosure is not limited to this. For example, a missing value imputation model 54 may be created using miRNA expression data and the response items from one or more questionnaire data 16 as training data, and estimating the response items that were not used as training data. Then, if the content of the response in the questionnaire data 16 is missing as a missing value, the miRNA expression data may be combined with the response items used for training from the content of the response in one or more questionnaire data 16 used for training, and the missing value may be imputed based on this combination. In this case, the combination of miRNA expression data and the response items used for training from the content of the response in the questionnaire data 16 is another example of an "indicator item" in this disclosure. Also, the response items used for training is an example of "other features" in this disclosure.

[0096] A more specific example is to combine miRNA expression data with gender responses in the questionnaire data 16 to supplement the missing values ​​in the questionnaire data 16, namely smoking status. In this case, in the learning process of step S208, the combination of miRNA expression data and gender responses in the questionnaire data 16 becomes the explanatory variable in the training data, and smoking status becomes the dependent variable in the training data. In this case, in step S306, the combination of miRNA expression data and gender responses in the questionnaire data 16 is used to supplement the smoking status.

[0097] Furthermore, while the above example described a case where the response content in the questionnaire data 16 was missing as a missing value, the technology in this disclosure is not limited to this. For example, the expression level of miRNA may be used as the missing value. That is, in this embodiment, if there is a missing item regarding the expression level of miRNA among the data included in the sample data 12, the expression level of the missing miRNA may be supplemented based on the expression levels of other miRNAs.

[0098] While embodiments of this disclosure have been described above with reference to the attached drawings, it is clear that any person with ordinary skill in the art to which this disclosure belongs could conceive of various modifications or applications within the scope of the technical idea described in the claims, and these too are naturally understood to fall within the technical scope of this disclosure.

[0099] Further preferred embodiments of this disclosure are shown below.

[0100] (Note 1) A discrimination method comprising: determining whether there are missing items in the features other than small RNA expression data in the sample data in which small RNA expression data is included as a feature; if the sample data includes the missing items, imputing the missing values ​​in the sample data based on a missing value imputation model that has been learned about the relationship between the index items included in the sample data and the features; and inputting the sample data with the imputed missing values ​​into a discrimination model that discriminates diseases based on small RNA expression data and features other than the small RNA expression data, and discriminating the disease of the subject related to the sample data.

[0101] (Note 2) The aforementioned indicator item is small RNA expression level data, as defined in Note 1, and the discrimination method is as described.

[0102] (Note 3) The discrimination method described in Note 2, wherein the indicator items further include other features other than the small RNA expression data.

[0103] (Note 4) The discrimination method described in Note 2 or Note 3, wherein the characteristics other than the small RNA expression data include the subject's sex, the subject's age, the subject's active smoking status, or the subject's tumor marker measurement value.

[0104] (Note 5) The discrimination model is a discrimination method described in any one of Notes 1 to 4, which estimates cancer as the disease.

[0105] (Note 6) A method for imputing sample data, which includes, when sample data in which small RNA expression data is included as a feature has missing items for features other than small RNA expression data, imputing the missing values ​​in the sample data based on a missing value imputation model that has been learned about the relationship between the small RNA expression data included in the sample data and the features in question.

[0106] (Note 7) A method for estimating specimen data, which includes estimating the characteristics of a subject related to the specimen data from the specimen data, based on an estimation model that has been learned about the relationship between small RNA expression data of a target for testing for a specific disease included in the specimen data and the characteristics used to distinguish the specific disease.

[0107] (Note 8) A discrimination system comprising: an input unit into which sample data in which small RNA expression data is included as a feature is input; a missing value completion unit that, when the sample data includes missing items for features other than small RNA expression data, completes the missing values ​​of the sample data based on a missing value completion model that has learned the relationship between the index items included in the sample data and the features; and a discrimination unit that inputs the sample data with the missing values ​​completed into a discrimination model that discriminates diseases based on small RNA expression data and features other than the small RNA expression data, and discriminates the disease of the subject related to the sample data.

[0108] The disclosure of Japanese Patent Application No. 2024-190994, filed on 30 October 2024, is incorporated herein by reference in its entirety. All documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as the individual documents, patent applications, and technical standards are incorporated herein by reference in the same manner as the individual documents, patent applications, and technical standards are incorporated herein by reference in the same manner as described herein.

[0109] 12 Sample data 14 Disease information 16 Medical interview data 20 Disease discrimination system 21 Next-generation sequencer 30 Learning model generation device 34 Input device 36 Display unit 38 Communication unit 40 Blood 41 CPU 42 RAM 43 Flash memory 44 ROM 45 Bus 46 Input / output interface 50 Disease discrimination device 52 Disease discrimination model 54 Missing value imputation model

Claims

1. A discrimination method comprising: determining whether there are missing items in the features other than small RNA expression data in the sample data in which small RNA expression data is included as a feature; if the sample data includes the missing items, imputing the missing values ​​in the sample data based on a missing value imputation model that has been learned about the relationship between the index items included in the sample data and the features; and inputting the sample data with the imputed missing values ​​into a discrimination model that discriminates diseases based on small RNA expression data and features other than small RNA expression data, and discriminating the disease of the subject related to the sample data.

2. The discrimination method according to claim 1, wherein the indicator item is small RNA expression level data.

3. The discrimination method according to claim 2, wherein the indicator item further includes other features other than the small RNA expression data.

4. The discrimination method according to claim 2, wherein the features other than the small RNA expression data include the subject's sex, the subject's age, the subject's active smoking status, or the subject's tumor marker measurement value.

5. The discrimination method according to any one of claims 1 to 4, wherein the discrimination model estimates cancer as the disease.

6. A method for imputing sample data, which includes, when sample data containing small RNA expression data as a feature has missing items for features other than small RNA expression data, imputing the missing values ​​in the sample data based on a missing value imputation model learned about the relationship between the small RNA expression data and the features contained in the sample data.

7. A method for estimating sample data, comprising: estimating the characteristics of a subject related to the sample data from the sample data, based on an estimation model learned about the relationship between small RNA expression data of a target for testing for a specific disease included in the sample data and the characteristics used to distinguish the specific disease.

8. A discrimination system comprising: an input unit into which sample data containing small RNA expression data as a feature is input; a missing value completion unit that, when the sample data contains missing items for features other than small RNA expression data, completes the missing values ​​of the sample data based on a missing value completion model that has learned the relationship between the index items included in the sample data and the features; and a discrimination unit that inputs the sample data with the missing values ​​completed into a discrimination model that discriminates diseases based on small RNA expression data and features other than the small RNA expression data, and discriminates the disease of the subject related to the sample data.

Citation Information

Patent Citations

  • Data complement program, data complement method, and data complement device

    JP2020154828A

  • Disease prevalence determination device, disease prevalence determination method, disease feature extraction device, and disease feature extraction method

    JP6280997B1

  • Morbidity determination assistance device, morbidity determination assistance method, and morbidity determination assistance program

    WO2020250995A1

  • Systems and methods for transforming and storing data from multiple studies

    WO2024026571A1