Information processing device, information processing method, and information processing program
The information processing device improves disease estimation accuracy by first estimating a general disease probability and then using that result to predict specific disease types, addressing the limitations of existing methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2026-03-24
AI Technical Summary
Existing techniques face challenges in accurately estimating specific diseases from multiple cancer types and suffer from decreased estimation performance when the probability of developing a specific type of cancer is low.
An information processing device and method that estimates a first probability of disease occurrence based on biomarker expression levels and then uses this probability to estimate a second probability for a specific disease type, employing machine learning models to improve accuracy.
Enhances disease estimation performance by dividing the classification problem into two stages, improving the accuracy of disease and specific disease type predictions.
Smart Images

Figure 2026052130000001_ABST
Abstract
Description
Technical Field
[0005]
[0001] Embodiments of the present invention relate to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Techniques for estimating diseases such as cancer that focus on the expression levels of miRNAs (microRNAs) in specimens such as blood have been proposed. For example, techniques for estimating the likelihood of developing cancer by comparing the expression levels of miRNAs obtained from a subject with those obtained from a healthy person have been disclosed. In addition, techniques for generating information regarding the presence or absence of multiple diseases by inputting information regarding multiple types of biomarkers into a single learned model have been disclosed.
[0003] However, the classification problem of directly estimating a specific disease from a specimen of a subject who may have any of multiple cancers is highly difficult, and the estimation performance may decrease. Also, when estimating the probability of developing a specific type of cancer, if the probability of the subject developing that type of cancer is low, the estimation performance of whether the subject has cancer itself may decrease. That is, in the prior art, the disease estimation performance may decrease.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0005] The problem that this invention aims to solve is to provide an information processing device, an information processing method, and an information processing program that can improve the performance of disease estimation. [Means for solving the problem]
[0006] The information processing device of this embodiment includes a processing unit that estimates a first probability of disease occurrence based on first information relating to the expression levels of one or more types of biomarkers in a sample, and estimates a second probability of disease occurrence for a type of disease based on the first information and the first probability of disease occurrence. [Brief explanation of the drawing]
[0007] [Figure 1] A schematic diagram of an information processing device. [Figure 2A] Diagram illustrating biomarker information. [Figure 2B] Diagram illustrating specimen-related information. [Figure 2C] A diagram illustrating the features. [Figure 3] A diagram illustrating the estimation of the first probability of illness. [Figure 4] Diagram illustrating the probability of first-stage illness. [Figure 5] A diagram illustrating the estimation of the second incidence probability. [Figure 6] Diagram illustrating the probability of second-degree illness. [Figure 7] A flowchart illustrating the flow of information processing. [Figure 8] Hardware configuration diagram. [Modes for carrying out the invention]
[0008] The information processing method, information processing apparatus, and information processing program of this embodiment will be described in detail below with reference to the attached drawings.
[0009] Figure 1 is a schematic diagram of an example of the information processing device 10 of this embodiment.
[0010] The information processing apparatus 10 is an information processing apparatus that executes a process of estimating the probability of suffering from each of a disease and the type of the disease of a subject based on biomarker information of a sample of the subject. The information processing apparatus 10 is constituted by one or more dedicated or general-purpose computers.
[0011] The information processing apparatus 10 includes a communication unit 12, a UI (user interface) unit 14, a storage unit 16, and a processing unit 20. The communication unit 12, the UI unit 14, the storage unit 16, and the processing unit 20 are communicably connected via a bus 18 or the like.
[0012] The communication unit 12 communicates with an external information processing apparatus or the like via a network or the like.
[0013] The UI unit 14 has a display function for displaying various information and an input function for receiving an operation input by a user. The display function is, for example, a display, a projection device, or the like. The input function is, for example, a pointing device such as a mouse and a touch pad, a keyboard, or the like. A touch panel in which the display function and the input function are integrally configured may be used.
[0014] The UI unit 14 may be configured to be communicably connected to the processing unit 20 by wire or wirelessly. The UI unit 14 may be provided outside the information processing apparatus 10, and the UI unit 14 and the processing unit 20 may be connected via a network or the like.
[0015] The storage unit 16 stores various data. The storage unit 16 may be provided outside the information processing apparatus 10. Further, at least one of the storage unit 16 and one or more functional units included in the processing unit 20 described later may be mounted on an external information processing apparatus communicably connected to the information processing apparatus
[0016] The processing unit 20 executes information processing in the information processing apparatus 10. The processing unit 20 includes an acquisition unit 20A, a feature quantity calculation unit 20B, a first probability of suffering estimation unit 20C, a second probability of suffering estimation unit 20D, and an output unit 20E.
[0017] The acquisition unit 20A, the feature amount calculation unit 20B, the first disease probability estimation unit 20C, the second disease probability estimation unit 20D, and the output unit 20E are realized by, for example, one or more processors. For example, each of the above units may be realized by causing a processor such as a CPU (Central Processing Unit) to execute a program, that is, by software. Each of the above units may be realized by a processor such as a dedicated IC or circuit, that is, by hardware. Each of the above units may be realized by using software and hardware in combination. When using a plurality of processors, each processor may realize one of the units or two or more of the units.
[0018] The acquisition unit 20A acquires one or more types of biomarker information and specimen-related information of the specimen.
[0019] A specimen is a sample derived from the living body of a subject. A specimen may be referred to as a biological sample or the like. The specimen is, for example, blood, serum, plasma, urine, saliva, gastric juice, etc., but is not limited thereto.
[0020] A subject is a living body for which the information processing apparatus 10 estimates the probability of suffering from a disease. The subject is, for example, a person, but may be a living organism other than a person. In the present embodiment, the case where the subject is a person will be assumed and described.
[0021] A disease is a pathological condition of the subject. The disease estimated by the information processing device 10 of this embodiment may be either a benign disease or a malignant disease. Benign diseases include benign tumors and various benign diseases of various organs and parts of the body. Examples of benign diseases include, but are not limited to, benign breast diseases, benign prostate diseases, benign pancreatic and biliary tract diseases, etc. Malignant diseases include malignant tumors, i.e., cancer. Cancer is classified into multiple types according to its classification criteria. When cancer is classified by site, the types of cancer include, but are not limited to, breast cancer, prostate cancer, pancreatic cancer, biliary tract cancer, colorectal cancer, stomach cancer, esophageal cancer, ovarian cancer, lung cancer, pancreatic cancer, bile duct cancer, uterine cancer, cervical cancer, liver cancer, leukemia, bladder cancer, malignant brain tumor, etc. Furthermore, the types of cancer may include types that are further subdivided from the above site-based classifications. For example, the types of cancer may include not only pancreatic cancer, which is cancer of the pancreas, but also more detailed classifications of pancreatic cancer, such as invasive ductal carcinoma and acinar cell carcinoma.
[0022] In this embodiment, we will explain assuming that the disease for which the probability of incidence is estimated by the information processing device 10 is cancer, and the type of disease for which the probability of incidence is estimated is the type of cancer.
[0023] Biomarker information refers to information that represents the expression levels of biomarkers (biological indicators) in a sample. Expression levels are expressed, for example, by concentration. Examples of biomarkers include, but are not limited to, miRNA (micro-RNA), plasma LDL (low-density lipoprotein), p53 gene, matrix metalloproteinase, and KRAS gene.
[0024] In this embodiment, the explanation will be based on the assumption that the biomarker is miRNA.
[0025] miRNAs are functional nucleic acids consisting of single-stranded RNA molecules with a length of 21 to 25 nucleotides. miRNAs have the function of suppressing the translation of various genes that have complementary target sites, and are known to regulate fundamental biological functions such as cell development, differentiation, proliferation, and cell death. More than 2,500 types of human miRNAs have been discovered.
[0026] As described above, this embodiment assumes that the disease for which the probability of incidence is estimated by the information processing device 10 is cancer, and that the type of disease for which the probability of incidence is estimated is the type of cancer. Furthermore, this embodiment assumes that the biomarker is miRNA (micro-RNA). However, the disease for which the probability of incidence is estimated by the information processing device 10 in this embodiment, the type of disease, and the biomarker used for their estimation are not limited to these.
[0027] Figure 2A is an explanatory diagram illustrating an example of biomarker information acquired by the acquisition unit 20A. Figure 2A shows the expression levels of three types of miRNAs 1-3 as an example. The numbers 1-3 following each miRNA are the identification information for the miRNA. These miRNAs 1-3 can be different types of miRNAs.
[0028] For example, the acquisition unit 20A acquires the measured expression levels of multiple types of miRNAs as biomarker information for multiple types of samples. Expression levels are expressed, for example, by concentration. Figure 2A shows an example in which the acquisition unit 20A acquires the concentrations of three types of miRNAs in a sample as biomarker information.
[0029] The types of miRNAs acquired by the acquisition unit 20A may be one or more, and are not limited to three types.
[0030] The method for measuring miRNA is not limited. For example, the acquisition unit 20A acquires miRNA expression levels as biomarker information, measured by at least one of the following methods: PCR (Polymerase Chain Reaction), LAMP (Loop-Mediated Isothermal Amplification), microarray, nanostring, and next-generation sequencing.
[0031] Specimen-related information refers to information about a specimen. For example, specimen-related information includes at least one of the following: information about biomarkers other than biomarkers obtained as biomarker information in the specimen; and information about the subject from whom the specimen was collected. Information about the subject includes information representing at least one of the following: the subject's age, height, weight, medical history, medication history, and various test results for the subject.
[0032] Figure 2B is an explanatory diagram illustrating an example of specimen-related information.
[0033] For example, the acquisition unit 20A acquires specimen-related information, including the subject's gender ("male") and age ("50 years old").
[0034] Returning to Figure 1, we continue the explanation.
[0035] The feature calculation unit 20B calculates feature quantities of biomarker expression levels represented by the biomarker information acquired by the acquisition unit 20A.
[0036] A feature is a relative index of the biomarker expression level relative to a reference index. The reference index is, for example, a reference variance, a statistical value, etc. In detail, the feature is calculated by correcting the biomarker expression level to a relative value relative to the reference expression level, and then determining the relative index relative to the reference index. The feature calculation unit 20B may use one of the multiple types of biomarker information acquired by the acquisition unit 20A as the reference expression level, or it may use a value that has been stored in advance.
[0037] The feature calculation unit 20B calculates features by performing one or more of the following in any order on the expression levels of multiple types of biomarkers represented by the multiple types of biomarker information acquired by the acquisition unit 20A: correction by calculation using a reference expression level, standardization, whitening, addition of new features by logarithmic calculation, and correction by calculation using sample-related information. For standardization, a standardization method based on the mean or variance can be used.
[0038] Figure 2C is an explanatory diagram of an example of a feature.
[0039] Figure 2C shows an example in which the feature calculation unit 20B calculates the feature quantities of two types of miRNAs shown in Figure 2C from the expression levels of three types of miRNAs represented by the three types of biomarker information shown in Figure 2A.
[0040] For example, the feature calculation unit 20B sets miRNA1 and miRNA2 as classification miRNAs and miRNA3 as the reference expression level miRNA. The feature calculation unit 20B then obtains a division result by dividing the expression levels of the classification miRNAs, miRNA1 and miRNA2, by the expression level of the reference expression level miRNA3. This division corresponds to a correction calculation using the reference expression level.
[0041] The feature calculation unit 20B then calculates the common logarithm of each of the division results obtained by dividing the expression levels of miRNA1 and miRNA2 by the concentration of miRNA3. Furthermore, the feature calculation unit 20B calculates the feature quantities of miRNA1 and miRNA2 by standardizing the calculated common logarithms of miRNA1 and miRNA2 by the respective mean and variance values.
[0042] Specifically, as shown in Figure 2A, we assume that the expression level of miRNA1 is "10,000 copies / μL" and the concentration of miRNA3 is "100,000 copies / μL". The division result when the expression level of miRNA1 is divided by the expression level of miRNA3, which is the baseline expression level, is "0.1", and the common logarithm of this division result "0.1" is "-1". The feature calculation unit 20B standardizes this common logarithm "-1" by setting the baseline mean value to "0" and the variance value to "100", and calculates "-0.1" as the feature of the expression level of miRNA1. The feature calculation unit 20B performs the same calculation for miRNA2 and calculates "-0.2" as the feature of the expression level of miRNA2.
[0043] The feature calculation unit 20B may calculate the feature quantities of each biomarker by performing multiple types of computational processing on the multiple types of biomarker information acquired by the acquisition unit 20A. Numerical calculations, statistical processing, and machine learning methods may be used for these computational processing, and the computational processing method may be adjusted in advance using information related to the other multiple types of biomarker information. Alternatively, each of the multiple types of biomarker information may be used directly as the feature quantity of each biomarker.
[0044] Returning to Figure 1, we continue the explanation.
[0045] The first incidence probability estimation unit 20C estimates the first incidence probability for a disease based on first information regarding the expression levels of one or more types of biomarkers in the sample.
[0046] The first incidence probability is information representing the probability of a subject contracting a disease. In other words, the first incidence probability is information representing the probability of having a disease. As described above, in this embodiment, we will explain assuming that the disease is cancer. In this case, the first incidence probability is information representing the probability that the subject has cancer.
[0047] The first information is information representing either the expression levels of one or more types of biomarkers represented by the biomarker information of one or more types acquired by the acquisition unit 20A, or the features calculated by the feature calculation unit 20B. In this embodiment, the first disease probability estimation unit 20C will be described as using the features calculated by the feature calculation unit 20B as the first information, as an example.
[0048] Furthermore, the first disease probability estimation unit 20C may estimate the first disease probability based on the first information and the specimen-related information. In this embodiment, one example of how the first disease probability estimation unit 20C estimates the first disease probability based on the feature quantities calculated by the feature quantity calculation unit 20B and the specimen-related information acquired by the acquisition unit 20A will be described.
[0049] Figure 3 is an explanatory diagram illustrating an example of the estimation of the first incidence probability by the first incidence probability estimation unit 20C.
[0050] The first disease probability estimation unit 20C inputs the feature quantities for each of the one or more types of biomarker information calculated by the feature quantity calculation unit 20B and the sample-related information acquired by the acquisition unit 20A into the first trained model M1, thereby obtaining the first disease probability as the output from the first trained model M1.
[0051] The first pre-trained model M1 is any machine learning model such as a pre-trained logistic regression model, linear regression model, generalized linear regression model, decision tree model, gradient boosting model, or neural network model. The first pre-trained model M1 only needs to be a pre-trained model that takes feature variables and sample-related information as input and outputs the first probability of disease.
[0052] Furthermore, when using biomarker information as the primary information, the first trained model M1 can be any pre-trained model that takes biomarker information and specimen-related information as input and outputs the first probability of disease.
[0053] Furthermore, if the first disease probability estimation unit 20C estimates the first disease probability from only the first information, the first trained model M1 only needs to be a trained model that has been pre-trained to take the first information as input and output the first disease probability.
[0054] Figure 4 is an explanatory diagram illustrating an example of the first probability of incidence.
[0055] For example, consider the case where the first incidence probability estimation unit 20C inputs the sample-related information shown in Figure 2B and the features shown in Figure 2C to the first trained model M1. In this case, the first incidence probability estimation unit 20C estimates the cancer incidence probability shown in Figure 4 as the first incidence probability, for example, as the output from the first trained model M1.
[0056] Returning to Figure 1, we continue the explanation.
[0057] The second incidence probability estimation unit 20D estimates the second incidence probability for each type of disease based on the first information and the first incidence probability.
[0058] The second incidence probability is information representing the probability of a subject contracting a particular type of disease. That is, the second incidence probability is information representing the probability of a subject contracting one or more types of diseases. As described above, in this embodiment, we will explain assuming that the disease is cancer. In this case, the second incidence probability is information representing the probability of contracting each type of cancer, such as breast cancer, prostate cancer, pancreatic cancer, biliary tract cancer, colorectal cancer, stomach cancer, esophageal cancer, ovarian cancer, lung cancer, pancreatic cancer, bile duct cancer, uterine cancer, cervical cancer, liver cancer, leukemia, bladder cancer, and malignant brain tumor. The second incidence probability estimation unit 20D may estimate the second incidence probability for a specific type of disease, or it may estimate the second incidence probability for each of multiple types of diseases.
[0059] As described above, the first information may be either the expression levels of one or more types of biomarkers represented by the biomarker information of one or more types acquired by the acquisition unit 20A, or the features calculated by the feature calculation unit 20B. In this embodiment, the case in which the second disease probability estimation unit 20D uses the features calculated by the feature calculation unit 20B as the first information will be described as an example.
[0060] Furthermore, the second incidence probability estimation unit 20D may estimate the second incidence probability based on the first information, the first incidence probability, and the specimen-related information. In this embodiment, one example of how the second incidence probability estimation unit 20D estimates the second incidence probability based on the first incidence probability estimated by the first incidence probability estimation unit 20C, the feature quantities calculated by the feature quantity calculation unit 20B, and the specimen-related information acquired by the acquisition unit 20A will be described.
[0061] Figure 5 is an explanatory diagram illustrating an example of the estimation of the second incidence probability by the second incidence probability estimation unit 20D.
[0062] The second disease probability estimation unit 20D inputs the first disease probability estimated by the first disease probability estimation unit 20C, the feature quantities for each expression level of one or more types of biomarkers calculated by the feature quantity calculation unit 20B, and the sample-related information acquired by the acquisition unit 20A to the second trained model M2, thereby obtaining the second disease probability as the output from the second trained model M2.
[0063] The second pre-trained model M2 is any machine learning model such as a pre-trained logistic regression model, linear regression model, generalized linear regression model, decision tree model, gradient boosting model, or neural network model. The second pre-trained model M2 only needs to be a pre-trained model that takes the first incidence probability, features, and sample-related information as input and outputs the second incidence probability.
[0064] Furthermore, when biomarker information is used as the first piece of information, the second pre-trained model M2 can be any pre-trained model that takes the first disease probability, biomarker information, and specimen-related information as input and outputs the second disease probability.
[0065] Furthermore, when the second disease probability estimation unit 20D estimates the second disease probability from the first disease probability and the first information, the second trained model M2 only needs to be a trained model that has been pre-trained to take the first disease probability and the first information as input and output the second disease probability.
[0066] Figure 6 is an explanatory diagram illustrating an example of the second probability of incidence.
[0067] For example, consider a case where the second incidence probability estimation unit 20D inputs the first incidence probability shown in Figure 4, the specimen-related information shown in Figure 2B, and the feature quantities shown in Figure 2C to the second trained model M2. In this case, for example, the second incidence probability estimation unit 20D estimates the second incidence probability shown in Figure 6 as the output from the second trained model M2. Figure 6 shows pancreatic cancer as the type of cancer, and the incidence probability of pancreatic cancer is shown as the second incidence probability. However, as mentioned above, the second incidence probability estimated by the second incidence probability estimation unit 20D can be any incidence probability for a given type of disease, and the type is not limited to pancreatic cancer.
[0068] Figure 6 shows an example of the second incidence probability estimation unit 20D estimating the second incidence probability for one type of cancer. However, as mentioned above, the second incidence probability estimation unit 20D may also estimate the second incidence probability for each of multiple types of cancer.
[0069] In this case, the second trained model M2 can be a trained model that takes the first probability, first information, and specimen-related information as input and outputs the second incidence probability for each of the multiple types of cancer.
[0070] Returning to Figure 1, we continue the explanation.
[0071] The output unit 20E outputs output information regarding the first incidence probability estimated by the first incidence probability estimation unit 20C and the second incidence probability estimated by the second incidence probability estimation unit 20D to the output device.
[0072] The output device is at least one of the UI unit 14, the storage unit 16, and an external information processing device connected to the information processing device 10 via the communication unit 12. For example, when the display function of the UI unit 14 is used as the output device, the output unit 20E displays the output information on the display of the UI unit 14. Also, for example, when the storage unit 16 is used as the output device, the output unit 20E outputs the output information by storing it in the storage unit 16. Furthermore, for example, when an external information processing device is used as the output device, the output unit 20E outputs the output information by transmitting it to the external information processing device.
[0073] The output information should be information relating to the first probability of illness estimated by the first probability of illness estimation unit 20C and the second probability of illness estimated by the second probability of illness estimation unit 20D.
[0074] For example, the output unit 20E outputs output information that includes the first probability of illness estimated by the first probability of illness estimation unit 20C and the second probability of illness estimated by the second probability of illness estimation unit 20D.
[0075] Specifically, we assume a case where the first incidence probability estimation unit 20C estimates the first incidence probability shown in Figure 4, and the second incidence probability estimation unit 20D estimates the second incidence probability shown in Figure 6. In this case, the output unit 20E outputs output information to the output device that includes "cancer incidence probability" and "0.8" as the first incidence probability, and "pancreatic cancer incidence probability" and "0.2" as the second incidence probability. In this case, the output unit 20E can output output information indicating that the probability of developing pancreatic cancer, a specific type of cancer, is low, but there is a high possibility of developing some other type of cancer.
[0076] Furthermore, the output unit 20E may output to an output device output information that includes at least one of a first determination result of whether or not a person has a disease based on a first probability of occurrence, and a second determination result of whether or not a person has a disease of a certain type based on a second probability of occurrence.
[0077] For example, the output unit 20E generates a first determination result indicating that the subject has cancer if the first probability of incidence estimated by the first probability of incidence estimated by the first probability of incidence estimated by the first probability of incidence estimated by the first probability of incidence estimated by the first probability of incidence estimated by the first probability of incidence estimated by the first threshold, and generates a first determination result indicating that the subject does not have cancer if the first probability of incidence estimated by the first probability of incidence estimated by the first probability of incidence estimated by the first threshold.
[0078] The first threshold is, for example, 0.5, but is not limited to this value. The first threshold can be predetermined. Furthermore, the first threshold may be changed as appropriate by user instructions for operation of the UI unit 14.
[0079] Furthermore, for example, if the second incidence probability estimated by the second incidence probability estimation unit 20D is greater than or equal to a predetermined second threshold, the output unit 20E generates a second determination result indicating that the subject has a specific type of cancer. Also, if the second incidence probability estimated by the second incidence probability estimation unit 20D is less than the second threshold, the output unit 20E generates a second determination result indicating that the subject does not have that type of cancer.
[0080] The second threshold is, for example, 0.5, but is not limited to this value. The second threshold can be predetermined. Furthermore, the second threshold may be changed as appropriate by user instructions for operation of the UI unit 14.
[0081] The output unit 20E then outputs output information including at least one of the first determination result and the second determination result to the output device.
[0082] The output format of the first judgment result indicating infection can be any of the following, and is not limited to: text information indicating "infection," symbols such as ○, animation, sound, etc. Similarly, the output format of the first judgment result indicating no infection can be any of the following, and is not limited to: text information indicating "no infection," symbols such as ×, animation, sound, etc.
[0083] Furthermore, the output format of the second judgment result indicating infection can be any of the following, and is not limited to: text information indicating "infection," symbols such as ○, animation, sound, etc. Similarly, the output format of the second judgment result indicating no infection can be any of the following, and is not limited to: text information indicating "no infection," symbols such as ×, animation, sound, etc.
[0084] Furthermore, if the output unit 20E determines that the first determination result based on the first incidence probability is "affected" and the second determination result based on the second incidence probability is "not affected", it may output to the output device output information that further includes information indicating the possibility of a disease of a type other than the type represented by the second incidence probability.
[0085] Furthermore, the output unit 20E may output to the output device, along with the first determination result, output information further including information indicating a higher risk level as the value of the first disease probability used to determine the first determination result increases. Similarly, the output unit 20E may output to the output device, along with the second determination result, output information further including information indicating a higher risk level as the value of the second disease probability used to determine the second determination result increases.
[0086] Furthermore, the output unit 20E may determine which of the multiple ranges to which the first incidence probability estimated by the first incidence probability estimation unit 20C belongs, and use information representing the degree of the first incidence probability corresponding to the range to which it belongs as the first determination result or the risk level described above.
[0087] Similarly, the output unit 20E may determine which of the following probability ranges the second incidence probability estimated by the second incidence probability estimation unit 20D belongs to, by classifying the entire range from a minimum of 0 to a maximum of 1.0 of the incidence probability into multiple ranges, and use information representing the degree of the second incidence probability corresponding to the range to which it belongs as the second determination result or the above risk level. For example, the output unit 20E may pre-set the range of pancreatic cancer incidence probability from 0.9 to 1.0 as risk level A, the range of 0.5 to 0.9 as risk level B, the range of 0.1 to 0.5 as risk level C, and the range of 0.0 to 0.1 as risk level D. Then, for example, if a pancreatic cancer incidence probability of "0.2" as shown in Figure 6 is estimated, the output unit 20E should output output information that includes risk level C, which corresponds to the range including 0.2, as the risk level of the second incidence probability.
[0088] Furthermore, the output unit 20E may output to the output device output information regarding the corrected first and second incidence probabilities, obtained by correcting at least one of the first and second incidence probabilities based on the first and second incidence probabilities.
[0089] For example, the first incidence probability, which is information representing the probability of having a disease estimated by the first incidence probability estimation unit 20C, may be less than the second incidence probability, which represents the probability of having a specific type of disease estimated by the second incidence probability estimation unit 20D. In such cases, the output unit 20E corrects at least one of the first incidence probability estimated by the first incidence probability estimation unit 20C and the second incidence probability estimated by the second incidence probability estimation unit 20D so that the relationship first incidence probability ≥ second incidence probability is satisfied.
[0090] For example, the output unit 20E corrects the first probability of illness estimated by the first probability of illness estimation unit 20C so that it is greater than or equal to the second probability of illness estimated by the second probability of illness estimation unit 20D, so that the relationship first probability of illness ≥ second probability of illness is satisfied.
[0091] Furthermore, the output unit 20E may correct the second incidence probability estimated by the second incidence probability estimation unit 20D so that it is less than the first incidence probability estimated by the first incidence probability estimation unit 20C, so as to satisfy the relationship first incidence probability ≥ second incidence probability.
[0092] Furthermore, the output unit 20E may correct both the first probability of illness estimated by the first probability of illness estimation unit 20C and the second probability of illness estimated by the second probability of illness estimation unit 20D so as to satisfy the relationship first probability of illness ≥ second probability of illness.
[0093] The output unit 20E then outputs output information regarding the corrected first probability of disease and the corrected second probability of disease to the output device.
[0094] Next, an example of the information processing flow performed by the information processing device 10 of this embodiment will be described.
[0095] Figure 7 is a flowchart showing an example of the information processing flow performed by the information processing device 10 of this embodiment.
[0096] The processing unit 20 executes steps S100 to S110 each time it acquires biomarker information and sample-related information for a sample collected from a subject.
[0097] In detail, the acquisition unit 20A of the processing unit 20 acquires biomarker information and specimen-related information (step S100). The acquisition unit 20A acquires biomarker information and specimen-related information from an external information processing device, for example, via the communication unit 12. Alternatively, the acquisition unit 20A may acquire biomarker information and specimen-related information stored in the storage unit 16. Furthermore, the acquisition unit 20A may acquire biomarker information and specimen-related information input by the user through operation instructions on the UI unit 14.
[0098] The feature calculation unit 20B calculates the feature of the biomarker expression level represented by the biomarker information obtained in step S100 (step S102). For example, if the biomarker information shown in Figure 2A and the sample-related information shown in Figure 2B are obtained in step 100, the feature calculation unit 20B calculates the feature shown in Figure 2, for example.
[0099] Next, the first incidence probability estimation unit 20C estimates the first incidence probability for cancer using the first information, which is the biomarker information obtained in step S100 or the feature quantities calculated in step S102, and the sample-related information obtained in step S100 (step S104). Through the processing in step S104, for example, the first incidence probability estimation unit 20C estimates the first incidence probability shown in Figure 4.
[0100] Next, the first incidence probability estimation unit 20C estimates the second incidence probability for each type of cancer using the first incidence probability estimated in step S104, the first information which is the biomarker information obtained in step S100 or the feature quantities calculated in step S102, and the specimen-related information obtained in step S100 (step S106). Through the processing in step S106, for example, the second incidence probability estimation unit 20D estimates the second incidence probability shown in Figure 6.
[0101] Next, the output unit 20E generates output information regarding the first probability of disease incidence estimated in step S104 and the second probability of disease incidence estimated in step S106 (step S108). Then, the output unit 20E outputs the output information generated in step S108 to the output device (step S110), and this routine ends.
[0102] As described above, the processing unit 20 of the information processing device 10 of this embodiment estimates a first probability of disease occurrence based on first information regarding the expression levels of one or more types of biomarkers in the sample, and estimates a second probability of disease occurrence based on the first information and the first probability of disease occurrence.
[0103] Thus, the processing unit 20 of the information processing device 10 in this embodiment does not directly estimate the probability of disease occurrence for a given disease type from the expression level of a biomarker, but rather first estimates a first probability of disease occurrence, and then uses the estimated first probability of disease occurrence to further estimate a second probability of disease occurrence for a given disease type. In other words, the processing unit 20 of the information processing device 10 in this embodiment estimates the probability of disease occurrence in steps.
[0104] Therefore, the processing unit 20 of the information processing device 10 in this embodiment can improve the estimation performance of the first incidence probability for a disease and the second incidence probability for a disease type compared to the case where the incidence probability is estimated in only one stage of processing. That is, the processing unit 20 estimates the incidence probability for a disease (first incidence probability) in the first stage, and in the second stage, uses the first incidence probability estimated in the first stage to estimate the incidence probability for a disease type, which is a more detailed classification (second incidence probability). Therefore, the processing unit 20 of the information processing device 10 in this embodiment can divide the classification problem through stepwise estimation and improve the disease estimation performance.
[0105] Therefore, the processing unit 20 of the information processing device 10 in this embodiment can improve the disease estimation performance.
[0106] Furthermore, the first information is the expression level of a biomarker, or a characteristic value of the biomarker expression level. Therefore, the processing unit 20 of this embodiment can further improve the disease estimation performance by using the expression level of a biomarker, or a characteristic value of the biomarker expression level, as the first information.
[0107] Furthermore, the processing unit 20 estimates a first disease probability based on the first information and specimen-related information concerning the specimen, and estimates a second disease probability based on the first information, the first disease probability, and the specimen-related information. In this way, the processing unit 20 of this embodiment can further improve its disease estimation performance by further estimating using specimen-related information.
[0108] Furthermore, the processing unit 20 outputs output information regarding the first and second probability of infection.
[0109] Therefore, the processing unit 20 can easily provide the estimation results to the user. Furthermore, for example, if the processing unit 20 outputs output information that includes a second incidence probability that is at least a predetermined probability lower than the first incidence probability, the processing unit 20 can provide information indicating that although the probability of contracting a specific type of disease represented by the second incidence probability is low, there is a high possibility that other types of diseases are occurring.
[0110] Furthermore, the processing unit 20 outputs output information regarding the corrected first and second incidence probabilities, which are obtained by correcting at least one of the first and second incidence probabilities based on the first and second incidence probabilities.
[0111] Therefore, in addition to the above effects, the processing unit 20 can output even more highly accurate output information regarding the probability of disease occurrence.
[0112] Next, an example of the hardware configuration of the information processing device 10 of the above embodiment will be described.
[0113] Figure 8 is a hardware configuration diagram of an example of the information processing device 10 of the above embodiment.
[0114] The information processing device 10 of the above embodiment includes a control device such as a CPU (Central Processing Unit) 90D, storage devices such as a ROM (Read Only Memory) 90E, RAM (Random Access Memory) 90F, and HDD (Hard Disk Drive) 90G, an I / F unit 90B which serves as an interface to various devices, an output unit 90A which outputs various information, an input unit 90C which accepts user operations, and a bus 90H which connects each unit, and has a hardware configuration that uses a normal computer.
[0115] In the information processing device 10 of the above embodiment, each of the above components is realized on the computer by the CPU 90D reading a program from the ROM 90E onto the RAM 90F and executing it.
[0116] The program for executing each of the above processes performed by the information processing device 10 in the above embodiment may be stored in the HDD 90G. Alternatively, the program for executing each of the above processes performed by the information processing device 10 in the above embodiment may be pre-installed and provided in the ROM 90E.
[0117] Furthermore, the program for executing the above-described process performed by the information processing device 10 of the above embodiment may be stored in an installable or executable file format on a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD (Digital Versatile Disc), or flexible disk (FD), and provided as a computer program product. Alternatively, the program for executing the above-described process performed by the information processing device 10 of the above embodiment may be stored on a computer connected to a network such as the Internet and provided by allowing download via the network. Alternatively, the program for executing the above-described process performed by the information processing device 10 of the above embodiment may be provided or distributed via a network such as the Internet.
[0118] Although this embodiment has been described above, it is presented as an example and is not intended to limit the scope of the invention. This novel embodiment can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. This embodiment and its variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of Symbols]
[0119] 10 Information Processing Devices 20 Processing Units 20B Feature Calculation Unit 20C First Incidence Probability Estimation Unit 20D Second Incidence Probability Estimation Unit 20E Output Section
Claims
1. Based on initial information regarding the expression levels of one or more biomarkers in the sample, the initial probability of disease incidence is estimated. Based on the first information and the first probability of incidence, a second probability of incidence for each type of disease is estimated. Processing unit, An information processing device equipped with the following features.
2. The aforementioned processing unit, Output information regarding the first and second incidence probabilities is output to the output device. The information processing apparatus according to claim 1.
3. The aforementioned processing unit, The first determination result of whether or not the disease is present based on the first probability of incidence, The second determination result of whether or not the person has the disease of the aforementioned type based on the second probability of incidence, Outputting the output information which includes at least one of the following, The information processing apparatus according to claim 2.
4. The aforementioned processing unit, Based on the first and second incidence probabilities, output information relating to the corrected first and second incidence probabilities is output, with at least one of the first and second incidence probabilities corrected. The information processing apparatus according to claim 2.
5. The first information is the expression level of the biomarker, or a characteristic quantity of the expression level of the biomarker. The information processing apparatus according to claim 1.
6. The aforementioned feature quantities are The expression level of the aforementioned biomarker is a relative index to a reference index, The information processing apparatus according to claim 5.
7. The aforementioned biomarker is a miRNA. The information processing apparatus according to claim 1.
8. The aforementioned disease is cancer. The aforementioned disease type is a type of cancer. The information processing apparatus according to claim 1.
9. The aforementioned processing unit, Based on the first information and the first incidence probability, the second incidence probability for each of the multiple types of the disease is estimated. The information processing apparatus according to claim 1.
10. The aforementioned processing unit, Based on the first information and the specimen-related information concerning the specimen, the first probability of disease is estimated. Based on the first information, the first probability of disease, and the specimen-related information, the second probability of disease is estimated. The information processing apparatus according to claim 1.
11. The aforementioned processing unit, Using the first trained model, the first probability of disease is estimated from the first information. Using the second trained model, the second incidence probability is estimated from the first information and the first incidence probability. The information processing apparatus according to claim 1.
12. An information processing method performed by an information processing device, Based on initial information regarding the expression levels of one or more biomarkers in the sample, the initial probability of disease incidence is estimated. Based on the first information and the first probability of incidence, a second probability of incidence for each type of disease is estimated. Information processing methods.
13. On the computer, Based on initial information regarding the expression levels of one or more biomarkers in the sample, the initial probability of disease incidence is estimated. Based on the first information and the first probability of incidence, a second probability of incidence for a particular type of disease is estimated. An information processing program for that purpose.
Citation Information
Patent Citations
Breast cancer detection kit or device and detection method
JP6804975B2
DISEASE PRESENCE DETECTION DEVICE, DISEASE PRESENCE DETECTION METHOD, AND DISEASE PRESENCE DETECTION PROGRAM
JP7021097B2
MicroRNA measurement method and kit
JP7299765B2
DISEASE INVESTIGATION DEVICE, DISEASE INVESTIGATION METHOD, AND DISEASE INVESTIGATION PROGRAM
JP7411619B2