A diagnostic support device for subjects suspected of having disease A or disease B, a trained model for diagnostic support, a diagnostic support method, a diagnostic support program, and a data structure for diagnostic support.

A machine learning-based diagnostic support system for primary aldosteronism accurately distinguishes between bilateral and unilateral forms using non-invasive biometric data, enhancing diagnostic efficiency and reducing the reliance on invasive procedures.

JP7894077B2Active Publication Date: 2026-07-23KANAZAWA UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KANAZAWA UNIV
Filing Date
2022-01-31
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing diagnostic methods for primary aldosteronism, particularly distinguishing between bilateral aldosteronism and unilateral primary aldosteronism, are complex and often fail to reach a timely diagnosis due to the need for multiple tests and invasive procedures like adrenal vein catheterization, leading to untreated cases.

Method used

A diagnostic support device and method using machine learning models trained on biometric data from thousands of cases, capable of predicting the presence of bilateral or unilateral primary aldosteronism without adrenal vein sampling or CT scans, incorporating features like serum potassium levels, potassium supplement dosage, and aldosterone-to-renin ratio.

Benefits of technology

The system achieves high detection sensitivity and specificity in diagnosing bilateral and unilateral primary aldosteronism, reducing the need for invasive tests and improving diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894077000007
    Figure 0007894077000007
  • Figure 0007894077000008
    Figure 0007894077000008
  • Figure 0007894077000009
    Figure 0007894077000009
Patent Text Reader

Abstract

To provide a diagnosis assistance device for diagnosing which of two diseases a subject is suffering from, a trained model for diagnosis assistance, a diagnosis assistance method, a diagnosis assistance program, and a diagnosis assistance data structure.SOLUTION: A diagnosis assistance device provided herein is configured to use a plurality of sets of biometric data of patients with bilateral aldosteronism (disease A) and a plurality of sets of biometric data of patients with unilateral primary aldosteronism (disease B) in a database as learning data to generate a plurality of explanatory variables for learning from the plurality of sets of biometric data for learning, and then predict (diagnosis assistance) or determine which of the disease A and the disease B a subject suspected of having the disease A or disease B is suffering from.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a diagnostic support device for a subject suspected of having disease A or disease B, a learned model for diagnostic support, a diagnostic support method, a diagnostic support program, and a data structure for diagnostic support for diseases A and B.

Background Art

[0002] Primary aldosteronism (PA) is a hypertensive disease with a higher incidence of cardiovascular diseases compared to essential hypertension, and an estimated 3 million patients in Japan. Approximately half of them are of the subtype aldosterone-producing adenoma (APA), and can be cured by removing the diseased adrenal gland. However, conventionally, until a disease type diagnosis is reached, it is necessary to perform screening tests, multiple confirmatory tests, CT scans, and adrenal vein catheterization tests in sequence. Due to the complexity of the diagnostic method, many cases could not reach a diagnosis and receive appropriate treatment. [[ID= ,15]]

[0003] Non-Patent Document 1 discloses an artificial intelligence system for predicting the subtype of primary aldosteronism. However, compared with the present system, 1) the markers used are different and 2) the results of adrenal CT scan examinations are required.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present invention aims to provide a diagnostic support device, a diagnostic support model, a diagnostic support method, a diagnostic support program, and a diagnostic support data structure for diagnosing which of two diseases (particularly bilateral aldosteronism or unilateral primary aldosteronism) a patient has. [Means for solving the problem]

[0006] To solve the above-mentioned problems, the present invention uses multiple biometric data from patients with bilateral aldosteronism (disease A) and multiple biometric data from patients with unilateral primary aldosteronism (disease B), respectively, from a database of nearly 5,000 adrenal disease cases, as training data. Multiple training explanatory variables are generated from these training biometric data, and a diagnostic support device, a trained model for diagnostic support, a diagnostic support method, a diagnostic support program, and a data structure for diagnostic support are constructed that can predict (assist in diagnosis) or determine whether a subject suspected of having disease A or disease B has either disease A or disease B. Furthermore, the present invention was completed by confirming that the diagnostic support device has high detection sensitivity and / or high specificity in assisting the specific diagnosis of bilateral aldosteronism or unilateral primary aldosteronism without using measurement data from adrenal vein sampling (AVS) and adrenal CT scans.

[0007] In other words, the present invention is as follows: 1. A diagnostic support device for disease A or disease B in a subject suspected of having disease A or disease B, A learning data storage unit that stores multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B as learning data, A model building unit constructs a learning model for predicting whether a patient has disease A or disease B from explanatory variables including the biometric data sets of patients with disease A and patients with disease B, using multiple biometric data sets of patients with disease A and multiple biometric data sets of patients with disease B for learning purposes. A predictive data receiving unit that receives one or more biometric measurement data from a subject suspected of having disease A or disease B as predictive data, and A prediction unit predicts or determines that a subject has either disease A or disease B by applying the biometric data of the prediction data received by the prediction data receiving unit to the learning model constructed by the model construction unit. A diagnostic support device characterized by being equipped with the following features. 2. The system further includes a learning missing value imputation unit that, when at least one biometric measurement data is missing from the group of learning biometric measurement data stored in the learning data storage unit, estimates and imputes the missing biometric measurement data from the other existing biometric measurement data. The diagnostic support device according to paragraph 1, wherein the model building unit is a model building unit that constructs a learning model using multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B, which have been supplemented by the learning missing value completion unit. 3. In the case where an imbalance occurs between multiple sets of biometric data for patients with disease A and multiple sets of biometric data for patients with disease B, the system further includes an imbalance data resolution unit, The diagnostic support device according to paragraph 1 or 2, wherein the model construction unit is a model construction unit that constructs the learning model using a group of biometric data obtained by resolving the imbalance between a group of biometric data obtained by a patient with disease A and a group of biometric data obtained by a patient with disease B, or a group of biometric data obtained by resolving the imbalance between a group of biometric data obtained by a patient with disease A and a group of biometric data obtained by a patient with disease B, which has been resolving the imbalance by the learning missing value imputation unit. 4. The diagnostic support device according to any one of items 1 to 3 above, wherein the model building unit is constructed from multiple learning models. 5. The diagnostic support device described in paragraph 4 above, wherein the learning model is constructed from one or more of the following: 1) Random Forest 2) K-Nearest Neighbor 3) Multi-Layer Perceptron 4) Support Vector Machine 5) Logistic Regression 6) Naive Bayes 7) LightGBM 6. A diagnostic support device which is one of items 1 to 5 above, wherein the combination of disease A and disease B is bilateral aldosteronism and unilateral primary aldosteronism. 7. The diagnostic support device according to claim 6, wherein the biological measurement data is any one of the following combinations. 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation 8. The diagnostic support device according to claim 6 or 7, wherein the biometric data does not include adrenal vein sampling (AVS) and / or adrenal CT scans. 9. A trained model for supporting the diagnosis of disease A or disease B in a subject suspected of having disease A or disease B, the trained model includes the following: 1) Multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B as training data, 2) A trained model for predicting or determining whether a patient has either disease A or disease B, based on explanatory variables including the biometric data sets, obtained by machine learning using multiple biometric data sets from patients with disease A and patients with disease B. 10. The combination of disease A and disease B is bilateral aldosteronism and unilateral primary aldosteronism, and the biometric data is one of the following combinations: The trained model described in item 9 above. 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation 11. A method for supporting the diagnosis of disease A or disease B in a subject suspected of having disease A or disease B, the method comprising the following steps: 1) A trained model generation process that generates a trained model capable of determining disease A or disease B by performing machine learning using multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B. 2) A step of predicting or determining whether a subject has either disease A or disease B by applying one or more biometric data of a subject suspected of having disease A or disease B to the trained model. 12. The combination of disease A and disease B is bilateral aldosteronism and unilateral primary aldosteronism, and the biometric data is one of the following combinations: The diagnostic support method described in item 11 above. 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation 13. A diagnostic support program for a subject suspected of having disease A or disease B, the program comprising the following steps: 1) A trained model generation process that generates a trained model capable of determining disease A or disease B by performing machine learning using multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B. 2) A step of predicting or determining whether a subject has either disease A or disease B by applying one or more biometric data of a subject suspected of having disease A or disease B to the trained model. 14. The combination of the disease A and the disease B is bilateral aldosteronism and unilateral primary aldosteronism, and the biological measurement data is any one of the following combinations: The diagnostic support program according to paragraph 13 above. 1) Serum K value before potassium preparation supplementation, serum K value, potassium preparation dosage, ARR, PAC 2) Serum K value before potassium preparation supplementation, serum K value, potassium preparation dosage, ARR 3) Serum K value before potassium preparation supplementation, serum K value, potassium preparation dosage 4) Serum K value before potassium preparation supplementation, serum K value 5) Serum K value before potassium preparation supplementation 15. A data structure for diagnosing the disease A or the disease B of a subject suspected of having the disease A or the disease B, the data structure including the following. 1) A plurality of biological measurement data groups of disease A patients and a plurality of biological measurement data groups of disease B patients as learning data, and 2) Data of a trained model for predicting or determining whether it is any one of the disease A or the disease B from explanatory variables including the biological measurement data groups using the plurality of biological measurement data groups of disease A patients and the plurality of biological measurement data groups of disease B patients for learning. 16. The combination of the disease A and the disease B is bilateral aldosteronism and unilateral primary aldosteronism, and the biological measurement data is any one of the following combinations: The data structure for diagnostic support according to paragraph 15 above. 1) Serum K value before potassium preparation supplementation, serum K value, potassium preparation dosage, ARR, PAC 2) Serum K value before potassium preparation supplementation, serum K value, potassium preparation dosage, ARR 3) Serum K value before potassium preparation supplementation, serum K value, potassium preparation dosage 4) Serum K value before potassium preparation supplementation, serum K value 5) Serum K value before potassium preparation supplementation

Advantages of the Invention

[0008] The present invention provides a diagnostic support device, a pre-trained model for diagnostic support, a diagnostic support method, a diagnostic support program, and a data structure for diagnostic support to assist in diagnosing which of two diseases (particularly bilateral aldosteronism or unilateral primary aldosteronism) has high detection sensitivity and / or specificity. [Brief explanation of the drawing]

[0009] [Figure 1] This is a block diagram showing an example of the functional configuration of a diagnostic support device according to the first embodiment. [Figure 2] This is a block diagram showing an example of the functional configuration of a diagnostic support device according to the second embodiment. [Figure 3] This is a block diagram showing an example of the functional configuration of a diagnostic support device according to the third embodiment. [Figure 4] This is a block diagram showing an example of the functional configuration of a diagnostic support device according to the fourth embodiment. [Figure 5] Selection of biometric data optimized for each learning model. (A) Selection of biometric data from the screening dataset, (B) Selection of biometric data from the confirmation test dataset. [Figure 6] Receiver operating characteristic curves for predictive diagnosis of aldosterone-producing adenomas. (A) Screening dataset, (B) Screening dataset with five most important variables, and (C) Confirmatory test dataset. [Modes for carrying out the invention]

[0010] (Target of this invention) The present invention relates to a diagnostic support device, a pre-trained model for diagnostic support, a diagnostic support method, a diagnostic support program, and a data structure for diagnostic support, for assisting in the diagnosis of which of two diseases (particularly bilateral aldosteronism or unilateral primary aldosteronism) a patient has.

[0011] (First Embodiment) Hereinafter, a first embodiment of the present invention will be described with reference to the drawings. Figure 1 is a block diagram showing an example of the functional configuration of the diagnostic support device 1 for disease A or disease B in a subject suspected of having disease A or disease B according to the first embodiment. As shown in Figure 1, the diagnostic support device 1A according to the first embodiment includes a learning data storage unit 2, a model construction unit 3, a trained model 4, a prediction data receiving unit 5, and a prediction unit 6 (functional blocks). For the sake of explanation, each component is shown separately, but each component may be combined, or other components may be included. The same applies to the other embodiments described below. The above configuration can be implemented using either hardware or software. For example, when implemented using software, each of the above functional blocks is actually configured with a computer's CPU, RAM, ROM, etc., and is realized by the operation of a program stored on a recording medium such as RAM, ROM, hard disk, or semiconductor memory. Furthermore, each configuration does not need to be installed on the same hardware; parts of the configuration may be installed on other media (e.g., the cloud). The same applies to the other embodiments described below.

[0012] 〇 Learning data storage unit 2 The learning data storage unit 2 receives multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B, respectively, as learning data. Furthermore, a portion of the training data (for example, about 20%) will be split into validation data. From now on, the training data will include the validation data. While there are no particular limitations on the partitioning method, methods such as holdout, cross-validation, k-fold cross-validation, and leave-one-out can be used. These data can be obtained, for example, from publicly known disease databases (e.g., the disease registry database of the Japan Research Association for Evidence-Based Disease Management (JRAS)).

[0013] ○ Biometric measurement data The biometric data refers to any biometric data obtained from patients with disease A and / or disease B. For example, the data in Table 1 below can be used as an example.

[0014] [Table 1]

[0015] The abbreviations in Table 1 above are as follows. The ATC / DDD index for antihypertensive drugs was calculated according to the Anatomical Therapeutic Chemical / Daily Defined Dose Index 2020. ARR: aldosterone-to-renin ratio, BMI: body mass index, BUN: blood urea nitrogen, CCT: captopril loading test, CKD: chronic kidney disease, DBP: diastolic blood pressure, eGFR: estimated glomerular filtration rate, FBS: fasting blood glucose, FUT: furosemide standing loading test, IHD: ischemic heart disease, NGSP: National Glycohemoglobin Standardization Program, PA: primary aldosteronism, PAC: plasma aldosterone level, PRA: plasma renin activity, SAS: sleep apnea syndrome, SBP: systolic blood pressure, s-Cl: serum chloride level, sK: serum potassium level, s-Na: serum sodium level, s-UA: serum uric acid level.

[0016] Model Construction Section 3 The model building unit 3 generates multiple training explanatory variables from multiple biometric data of patients with disease A and multiple biometric data of patients with disease B, received by the training data storage unit 2, as needed. That is, it synthesizes or transforms multiple biometric data to generate new explanatory variables. The method for creating new explanatory variables is, for example, some or all of those generated by feature engineering such as standardization, factor analysis, simple regression, residual calculation using multiple regression, principal component analysis, arithmetic operations, mean, variance, and standard deviation. The model building unit 3 uses multiple biometric data from patients with disease A and multiple biometric data from patients with disease B, received by the learning data storage unit 2, and, if necessary, the generated explanatory variables (hereinafter, the learning explanatory variables may also be referred to as learning data), to build a learning model for deriving diagnostic support for disease A or disease B from biometric data. Furthermore, the constructed learning model is stored in the pre-trained model 4.

[0017] Model building unit 3 constructs a learning model by applying known machine learning methods using the above-mentioned training data. For example, Random Forest, K-Nearest Neighbor, Multi-Layer Perceptron, Support Vector Machine, Logistic Regression, Naive Bayes, LightGBM, etc. can be used.

[0018] 〇 Pre-trained model 4 The trained model 4 saves and updates the trained model built by the model building unit 3. The trained model 4 also has the function of updating the trained model by repeatedly training it in the model building unit 3.

[0019] 〇 Prediction data receiving unit 5 The prediction data receiving unit 5 inputs the biometric measurement data of the subject (a subject suspected of having disease A or disease B) as prediction data. Preferably, the subject is one for whom biometric measurement data has been obtained, but for whom it is not yet diagnosed (confirmed diagnosis) whether it has disease A or disease B.

[0020] Prediction Unit 6 The prediction unit 6 generates multiple explanatory variables for prediction from the prediction data received by the prediction data receiving unit 5, if necessary. These explanatory variables can be generated using the same method as used in the model building unit 3. The prediction unit 6 applies the biometric measurement data and multiple predictive explanatory variables received by the prediction data receiving unit 5 to the trained model stored in the trained model 4, thereby predicting (assisting in diagnosis) whether the subject has disease A or disease B. Furthermore, the prediction unit 6 outputs the prediction results as follows: AUC (Area Under the Curve), detection sensitivity, specificity, PPV (Positive Predictive Value), NPV (Negative Predictive Value), F-score, and accuracy.

[0021] According to the first embodiment, a learning model constructed based on multiple biometric data from patients with disease A and multiple biometric data from patients with disease B can be used to predict (assist in diagnosis) whether a subject has disease A or disease B based on biometric data about that subject.

[0022] (Second embodiment) A second embodiment of the present invention will be described with reference to Figure 2. Figure 2 is a block diagram showing an example of the functional configuration of the diagnostic support device 1B for disease A or disease B in a subject suspected of having disease A or disease B, according to the second embodiment. In Figure 2, components with the same reference numerals as those shown in Figure 1 have the same function, so redundant explanations are omitted here.

[0023] 〇 Missing value imputation unit for learning 7 In the second embodiment, in addition to the first embodiment, a learning missing value imputation unit 7 is included. The learning missing value imputation unit 7, when it finds that at least one biometric measurement is missing from the learning biometric data received by the learning data storage unit 2, estimates and imputates the missing biometric measurement from the other existing biometric measurements. In other words, the learning missing value imputation unit 7, when there are missing values ​​in the learning biometric measurement data, imputes the missing values ​​by using other existing biometric measurements and estimating them through machine learning. For example, Random Forest (MissForest) can be used as the machine learning method applied.

[0024] 〇 Prediction missing value imputation unit 8 The system may also include a predictive missing value imputation unit 7 that, if at least one biometric value is missing from the biometric measurement data of a subject (a subject suspected of having disease A or disease B), estimates and imputes the missing biometric value from other existing biometric values.

[0025] According to the second embodiment, if there are missing values ​​in the training data, these can be imputed, and then explanatory variables can be generated, a learning model can be constructed, and the disease diagnosis of the subject can be assisted.

[0026] (Third embodiment) A third embodiment of the present invention will be described with reference to Figure 3. Figure 3 is a block diagram showing an example of the functional configuration of the diagnostic support device 1C for disease A or disease B in a subject suspected of having disease A or disease B, according to the third embodiment. In Figure 3, components with the same reference numerals as those shown in Figures 1 and 2 have the same function, so redundant explanations are omitted here.

[0027] 〇 Data Imbalance Resolution Unit 9 In the third embodiment, in addition to the first and second embodiments, an unbalanced data resolution unit 9 is included. The imbalance data resolution unit 9 resolves imbalances in the learning biometric measurement data received by the learning data storage unit 2 if imbalances occur. In other words, the unbalanced data resolution unit 9 resolves imbalances in the training biometric data by employing oversampling (e.g., SMOTE), undersampling, etc.

[0028] According to the third embodiment, if there is an imbalance in the training data, this can be resolved before generating explanatory variables, constructing a learning model, and assisting in the diagnosis of diseases in the subjects.

[0029] (Fourth embodiment) A fourth embodiment of the present invention will be described with reference to Figure 4. Figure 4 is a block diagram showing an example of the functional configuration of the diagnostic support device 1D for disease A or disease B in a subject suspected of having disease A or disease B, according to the fourth embodiment. In Figure 4, components with the same reference numerals as those shown in Figures 1, 2, and 3 have the same function, so redundant explanations are omitted here.

[0030] In the fourth embodiment, the model building unit 3 is constructed from multiple learning models (ensemble learning models). Examples of multiple learning models include 1) Random Forest, 2) K-Nearest Neighbor, 3) Multi-Layer Perceptron, 4) Support Vector Machine, 5) Logistic Regression, 6) Naive Bayes, and 7) LightGBM. The following are examples of methods for calculating the final prediction result by combining the results of each model. 1) The predicted probabilities are averaged, and the patient is classified as either disease A or disease B, with a threshold of 0.5. The results of each learning model are added together with the same weight and then averaged. 2) Divide the learning model into a primary learning model and a secondary learning model, assign individual weights to each learning model, and evaluate them based on an overall score. While it is preferable to use the same markers (biometric data) in each learning model, it is also acceptable to set markers that are optimal for each learning model. For example, multiple markers selected in a learning model using random forest can be adopted in other learning models as well.

[0031] According to the fourth embodiment, since the results of multiple learning models are combined to make predictions, it is possible to provide highly accurate assistance in diagnosing diseases in subjects.

[0032] (How to set the importance level of markers) The method for setting the importance of markers can be a publicly known method, but the following are some examples. When using a random forest architecture, the importance of markers can be determined by the Gini coefficient. A smaller Gini coefficient indicates that the data has been split more correctly. Specifically, by calculating how much the Gini coefficient decreases depending on which marker is used to split the data when constructing the decision tree, markers with a larger decrease in the Gini coefficient are considered to be more important.

[0033] (Hyperparameter tuning) In the learning model used in this embodiment, hyperparameter tuning can be performed using known methods such as grid search, random search, and Bayesian optimization. For example, tuning can be performed using Bayesian optimization as follows. 1) Random Forest: Number and depth of decision trees 2) K-Nearest Neighbor: The number of neighboring samples to consider when determining the label. 3) Multi-Layer Perceptron: Number of hidden layers and neurons, regularization parameters 4) Support Vector Machine: Regularization strength, RBF kernel coefficients 5) Logistic Regression: Strength of regularization 7) LightGBM: Regularization strength, number and depth of decision trees

[0034] (A combination of disease A and disease B) In this specification, the combination of disease A and disease B primarily refers to cases where there is suspicion that the subject has either disease A or disease B. In this example, the subject was diagnosed with primary aldosteronism, and examples of bilateral aldosteronism (disease A) and unilateral primary aldosteronism (disease B) are given, but the example is not particularly limited.

[0035] (Diagnostic assistance for specific diagnosis of bilateral aldosteronism or unilateral primary aldosteronism) The following embodiment confirms that specific diagnostic assistance for bilateral aldosteronism or unilateral primary aldosteronism can be performed using the following biometric data. 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation The serum potassium level before potassium supplementation (especially potassium supplements used to treat hypokalemia resulting from the treatment of primary aldosteronism) refers to, for example, the serum potassium level before treatment for primary aldosteronism. The potassium (K) preparation dosage refers to the amount of potassium (K) preparation used, for example, during the treatment of primary aldosteronism. If no potassium preparation was administered, the value will be 0. Serum potassium levels are not limited to any specific time period and may include measurements taken when a disease is suspected (for example, measurements taken when diagnosing aldosteronism), measurements taken at the time of biopsy (screening), or measurements taken after potassium supplementation. PAC (Plasma Aldosterone Level) is measured at any time, and may include measurements taken when a disease is suspected (for example, measurements taken when diagnosing aldosteronism), measurements taken at the time of biopsy (screening), or measurements taken after potassium supplementation. Preferably, the measurement taken when oral medications that affect renin or aldosterone are discontinued or changed is adopted, in accordance with the Japanese Society of Hypertension's guidelines for hypertension treatment. ARR (aldosterone concentration / renin activity ratio) is not limited to a specific time period and may be measured when a disease is suspected (for example, measured at the time of diagnosis of aldosteronism), measured at the time of biopsy (screening), or measured after potassium supplementation. Preferably, the measured value used is one taken when oral medications that affect renin or aldosterone are discontinued or changed, in accordance with the hypertension treatment guidelines of the Japanese Society of Hypertension. The above measurement data does not include measurement data from adrenal vein sampling and adrenal CT scans, which are conventionally used for definitive diagnosis.

[0036] (A pre-trained model for supporting the diagnosis of disease A or disease B in subjects suspected of having disease A or disease B) A trained model for supporting the diagnosis of disease A or disease B in a subject suspected of having disease A or disease B includes the following: 1) Multiple biometric data sets from patients with disease A and multiple biometric data sets from patients with disease B as training data. 2) A trained model for predicting whether a patient has disease A or disease B, based on explanatory variables including the biometric data sets, by machine learning using multiple biometric data sets from patients with disease A and patients with disease B. In a specific embodiment of the trained model, the combination of disease A and disease B is bilateral aldosteronism and unilateral primary aldosteronism, and the biometric data is one of the following combinations: 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation

[0037] (Methods for supporting the diagnosis of disease A or disease B in subjects suspected of having disease A or disease B) A method for supporting the diagnosis of disease A or disease B in a subject suspected of having disease A or disease B includes the following steps: 1) A trained model generation process that generates a trained model capable of determining disease A or disease B by performing machine learning using multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B. 2) A step of predicting whether a subject has either disease A or disease B by applying one or more biometric data from a subject suspected of having disease A or disease B to the trained model. As a specific embodiment of the diagnostic support method, the combination of disease A and disease B is bilateral aldosteronism and unilateral primary aldosteronism, and the biometric data is one of the following combinations: 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation

[0038] (Diagnostic support program for subjects suspected of having disease A or disease B) A diagnostic support program for subjects suspected of having disease A or disease B includes the following steps: 1) A trained model generation process that generates a trained model capable of determining disease A or disease B by performing machine learning using multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B. 2) A step of predicting whether a subject has either disease A or disease B by applying one or more biometric data of a subject suspected of having disease A or disease B to the trained model. In a specific embodiment of the diagnostic support program, the combination of disease A and disease B is bilateral aldosteronism and unilateral primary aldosteronism, and the biometric data is one of the following combinations: 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation

[0039] (Data structure for supporting the diagnosis of disease A or disease B in subjects suspected of having disease A or disease B) The data structure for supporting the diagnosis of disease A or disease B in subjects suspected of having disease A or disease B includes the following: 1) Multiple biometric data sets from patients with disease A and multiple biometric data sets from patients with disease B as training data. 2) Using multiple biometric data sets from patients with disease A and patients with disease B, data of a trained model used to predict whether a patient has disease A or disease B based on explanatory variables including the biometric data sets. As a specific embodiment of the data structure for diagnostic support, the combination of disease A and disease B is bilateral aldosteronism and unilateral primary aldosteronism, and the biometric data is one of the following combinations: 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation

[0040] The present invention will be described below with reference to examples, but the present invention is not limited in any way by these examples. The following examples have been approved by the Ethics Committee of the National Hospital Organization Kyoto Medical Center (leading institution) and the Ethics Committees of JRAS participating institutions. [Examples]

[0041] In this example, biometric data from patients definitively diagnosed with bilateral aldosteronism and patients definitively diagnosed with unilateral primary aldosteronism were used as training data to construct a high-precision ensemble learning model to assist in diagnosing which of the two diseases a patient has.

[0042] (Patient data used) Between January 2006 and December 2018, male or female patients aged 20-90 years diagnosed with primary aldosteronism (PA) by adrenal vein sampling (AVS) were registered in the JRAS (Japan Radar Aldosteronism Study). Clinical characteristics, biochemical findings, and confirmatory test results were collected electronically using an online registration system. System construction, data security, and registry data maintenance were entrusted to EPS Corporation. This study was conducted using a dataset validated in March 2020. It included patients with data including clinical features, biochemistry, and AVS. PA was diagnosed according to the guidelines of the Japan Endocrine Society and the Japanese Society of Hypertension.

[0043] (Training dataset used) The JRAS database includes PA diagnostic criteria: physical examination; medical history; electrolyte, renal function, glucose metabolism, and lipid metabolism parameters; aldosterone-to-renin ratio (ARR) at screening; and confirmatory test results.

[0044] (statistical analysis) Data were expressed as mean ± standard deviation (SD) or percentage. Differences between APA and BAH were analyzed. The Shapiro-Wilk test was used to test for normality. If normality was rejected, the Mann-Whitney U test was used. If normality was accepted, the Bartlett test was used to test for equal variances. If equal variances were found, Student's t-test was used; if equal variances were not found, Welch's t-test was used. P < 0.05 was considered statistically significant.

[0045] (Building a learning model) The algorithm development process included the general processes of 1) data creation and 2) model building and evaluation. A total of seven machine learning programs were used (Random Forest, K-Nearest Neighbor, Multi-Layer Perceptron, Support Vector Machine, Logistic Regression, Naive Bayes, and LightGBM). Six of the algorithms, excluding LightGBM, were developed by imputing missing values ​​in the database using MissForest. These seven machine learning programs were combined to form a single high-accuracy ensemble learning model.

[0046] (result) 1) Patient characteristics A total of 4057 PA patients were registered in the JRAS database. Based on inclusion and exclusion criteria, data from 545 patients with APA and 1352 patients with BAH were used in this example. Table 2 below shows the baseline characteristics and percentage of missing data for APA and BAH patients.

[0047] [Table 2]

[0048] The data in Table 2 above are shown as mean (SD) or percentage. The ATC / DDD index for antihypertensive drugs was calculated according to the Anatomical Therapeutic Chemical / Daily Defined Dose Index 2020. *, p<0.05 for BAH;**, p<0.01 for BAH;***, p<0.001 for BAH. APA: Aldosterone-producing adenoma, ARR: aldosterone-to-renin ratio (aldosterone concentration / renin activity ratio), BAH: Bilateral adrenal hyperplasia, BMI: Body Mass Index, BUN: Blood urea nitrogen, CCT: Captopril loading test, CKD: Chronic kidney disease, DBP: Diastolic blood pressure, eGFR: Estimated glomerular filtration rate, FBS: Fasting blood glucose, FUT: Furosemide standing loading test, IHD: Ischemic heart disease, NGSP: National Glycohemoglobin Standardization Program, PA: Primary aldosteronism, PAC: Plasma aldosterone level, PRA: Plasma renin activity, SAS: Sleep apnea syndrome, SBP: Systolic blood pressure, s-Cl: Serum chloride level, sK: Serum potassium level, s-Na: Serum sodium level, s-UA: Serum uric acid level.

[0049] 2) Feature importance analysis Table 3 below shows the functional importance ranking and importance scores in the Random Forest (RF) model for the screening and confirmatory test datasets. The five most important variables (explanatory variables) for the screening test dataset were serum potassium level at the initial visit, serum potassium level at supplementation, PAC, ARR, and potassium supplementation dose. The five most important variables (explanatory variables) for the confirmatory test dataset were serum potassium level at the initial visit, PAC at 60 and 90 minutes in the captopril loading test, PAC at 0 minutes in the furosemide standing loading test, and serum potassium level at supplementation. Figure 5 shows the relationship between the number of variables and mean AUC for the seven MLAs in the two datasets. In the screening dataset, RF, MLP, LightGBM, SVM, LR, KNN, and NB showed the best mean AUC using 19, 12, 13, 12, 12, 7, and 9 variables, respectively. In the confirmatory test dataset, RF, MLP, LightGBM, SVM, LR, KNN, and NB showed the best mean AUC using 23, 19, 25, 21, 21, 7, and 10 variables, respectively. Therefore, the inventors selected the number of variables obtained from the verification test dataset in order to develop the best-performing prediction model.

[0050] [Table 3]

[0051] 3) Results of ensemble learning We developed three ensemble learning models by combining seven MLAs and three datasets (screening data, data for five key variables, and validation data). Figure 6 shows the median receiver operating characteristic curves for the APA predictive model using the screening test dataset, the screening dataset with the five most important variables, and the confirmatory test dataset. The AUC for the screening test dataset was 0.898 ± 0.040, and the AUC for the five most important variables model was 0.886 ± 0.044 (see Table 4). The model constructed using the validation test dataset showed a good AUC value of 0.914±0.041 compared to the other two datasets. The predictive performance of only the seven MLAs is shown in Table 5. The highest average AUC was 0.914±0.042 for RF. Based on the above, we constructed an ensemble learning model with high accuracy and specificity to assist in the specific diagnosis of two diseases. In particular, in the specific diagnostic support for bilateral aldosteronism or unilateral primary aldosteronism, measurement data from adrenal vein sampling (AVS) and adrenal CT scans are not used. It was confirmed to have high detection sensitivity and / or high specificity.

[0052] [Table 4]

[0053] [Table 5] [Examples]

[0054] In this example, the following high-importance markers (measurement data) obtained in Example 1 were used for predictions using the high-precision ensemble learning model obtained in Example 1. 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum potassium level before potassium supplementation, serum potassium level, potassium supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum potassium levels before potassium supplementation

[0055] (result) The prediction results are shown in Table 6 below. The high-precision ensemble learning model obtained in this embodiment 1 can make predictions with high accuracy and high specificity even when using not only the five markers, but also a portion of the five markers.

[0056] [Table 6] [Industrial applicability]

[0057] The present invention makes it possible to provide a diagnostic support device, a diagnostic support device, a trained model for diagnostic support, a diagnostic support method, a diagnostic support program, and a data structure for diagnostic support, for assisting in the diagnosis of which of two diseases (particularly bilateral aldosteronism or unilateral primary aldosteronism) has high detection sensitivity and / or specificity.

Claims

1. A diagnostic support device for disease A or disease B in a subject suspected of having disease A or disease B, A learning data storage unit that stores multiple sets of biometric data from patients with disease A and multiple sets of biometric data from patients with disease B as learning data, A model building unit constructs a learning model for predicting whether a patient has disease A or disease B from explanatory variables including the biometric data sets, using multiple biometric data sets from patients with disease A and multiple biometric data sets from patients with disease B. A predictive data receiving unit that receives one or more biometric measurement data from a subject suspected of having disease A or disease B as predictive data, and A prediction unit predicts or determines that a subject has either disease A or disease B by applying the biometric data of the prediction data received by the prediction data receiving unit to the learning model constructed by the model construction unit. A diagnostic support device characterized by being equipped with, Here, the combination of disease A and disease B is bilateral aldosteronism and unilateral primary aldosteronism, and The biometric measurement data is a combination that includes any one of the following: 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum K level before potassium supplementation Diagnostic support device.

2. The learning data storage unit further includes a learning missing value imputation unit that, if at least one biometric measurement data is missing from the group of learning biometric measurement data stored in the learning data storage unit, estimates and imputes the missing biometric measurement data from the other existing biometric measurement data. The diagnostic support device according to claim 1, wherein the model construction unit is a model construction unit that constructs the learning model using a plurality of sets of biometric data from patients with disease A and a plurality of sets of biometric data from patients with disease B, which have been imputed by the learning missing value imputation unit.

3. If an imbalance occurs between multiple sets of biometric data for patients with disease A and multiple sets of biometric data for patients with disease B, in the set of biometric data for learning stored in the learning data storage unit, the unbalanced data resolution unit is further provided. The diagnostic support device according to claim 2, wherein the model construction unit is a model construction unit that constructs the learning model using a group of biometric data obtained by resolving the imbalance between a group of biometric data obtained by a patient with disease A and a group of biometric data obtained by a patient with disease B, or a group of biometric data obtained by resolving the imbalance between a group of biometric data obtained by a patient with disease A and a group of biometric data obtained by a patient with disease B, which has been resolving the imbalance that has been resolving by the learning missing value completion unit.

4. The diagnostic support device according to claim 1 or 2, wherein the model building unit is constructed from a plurality of learning models.

5. The diagnostic support device according to claim 4, wherein the learning model is constructed from one or more of the following: 1) Random Forest 2) K-Nearest Neighbor 3) Multi Layer Perceptron 4) Support Vector Machine 5) Logistic Regression 6) Naive Bayes 7) LightGBM

6. The diagnostic support device according to claim 1, wherein the biometric data does not include adrenal vein sampling (AVS) and / or adrenal CT scans.

7. The diagnostic support device according to claim 1, wherein the biometric data is a combination of pre-potassium supplement serum K level, serum K level, K preparation dosage, ARR, and PAC.

8. The diagnostic support device according to claim 1, wherein the biometric data is a combination of serum K level before potassium supplementation, serum K level, K preparation dosage, and ARR.

9. A diagnostic support program for bilateral aldosteronism or unilateral primary aldosteronism in a subject suspected of having bilateral aldosteronism or unilateral primary aldosteronism, Computers, A learning data storage unit that stores multiple sets of biometric data from patients with bilateral aldosteronism and multiple sets of biometric data from patients with unilateral primary aldosteronism as learning data. A model building unit constructs a learning model for predicting whether a patient has bilateral aldosteronism or unilateral primary aldosteronism, using multiple sets of biometric data from patients with bilateral aldosteronism and multiple sets of biometric data from patients with unilateral primary aldosteronism, based on explanatory variables including the biometric data sets. A predictive data receiving unit that receives one or more biometric data from a subject suspected of having bilateral aldosteronism or unilateral primary aldosteronism as predictive data, and A prediction unit predicts or determines whether the subject has bilateral aldosteronism or unilateral primary aldosteronism by applying the biometric data of the prediction data received by the prediction data receiving unit to the learning model constructed by the model construction unit. It is a diagnostic support program designed to function as such. The biometric measurement data is a combination that includes any one of the following: 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum K level before potassium supplementation Diagnostic support program.

10. A diagnostic support method performed by a diagnostic support device for bilateral aldosteronism or unilateral primary aldosteronism in a subject suspected of having bilateral aldosteronism or unilateral primary aldosteronism, the method comprising the following steps: 1) A diagnostic support device performs a trained model generation step in which it performs machine learning using multiple sets of biometric data from patients with bilateral aldosteronism and multiple sets of biometric data from patients with unilateral primary aldosteronism to generate a trained model capable of determining whether the patient has bilateral aldosteronism or unilateral primary aldosteronism. 2) The diagnostic support device predicts or determines that a subject has either bilateral aldosteronism or unilateral primary aldosteronism by applying one or more biometric data from a subject suspected of having bilateral aldosteronism or unilateral primary aldosteronism to the trained model, and Here, the biometric measurement data is a combination that includes any one of the following: 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum K level before potassium supplementation Diagnostic support methods.

11. A diagnostic support program for bilateral aldosteronism or unilateral primary aldosteronism in a subject suspected of having bilateral aldosteronism or unilateral primary aldosteronism, Here, a diagnostic support program is provided to cause the computer to perform the following steps: 1) A trained model generation process that generates a trained model capable of determining bilateral aldosteronism or unilateral primary aldosteronism by performing machine learning using multiple sets of biometric data from patients with bilateral aldosteronism and multiple sets of biometric data from patients with unilateral primary aldosteronism. 2) A step of applying one or more biometric data from a subject suspected of having bilateral aldosteronism or unilateral primary aldosteronism to the trained model to predict or determine that the subject has either bilateral aldosteronism or unilateral primary aldosteronism. Here, the biometric measurement data is a combination that includes any one of the following: 1) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR, PAC, 2) Serum K level before potassium supplementation, serum K level, K supplement dosage, ARR 3) Serum potassium level before potassium supplementation, serum potassium level, and potassium supplement dosage 4) Serum K level before potassium supplementation, serum K level 5) Serum K level before potassium supplementation Diagnostic support program.