Medical information processing apparatus, medical information processing system, medical information processing method, and medical information processing program

The medical information processing apparatus quantifies the influence of unobserved confounding factors using propensity scores and CDS models, addressing the challenge of unobserved confounding in causal inference and improving clinical decision-making reliability.

JP7717501B2Active Publication Date: 2025-08-04CANON MEDICAL SYST CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021099384
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-15
Publication Date
2025-08-04
Estimated Expiration
2041-06-15

AI Technical Summary

Technical Problem

Existing causal inference methods using machine learning struggle with accurately estimating causal effects due to the presence of unobserved confounding factors, lacking a means to quantify their influence, and relying on unrealistic assumptions.

Method used

A medical information processing apparatus and system that includes a first acquisition unit, a second acquisition unit, a first extraction unit, and a calculation unit to quantify the influence of unobserved confounding factors by learning from observed confounding factors and support information, using propensity scores and CDS models to improve causal inference accuracy.

Benefits of technology

Enables accurate estimation of causal effects by quantifying the influence of unobserved confounding factors, enhancing the reliability of clinical decision-making and causal inference processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717501000011
    Figure 0007717501000011
  • Figure 0007717501000012
    Figure 0007717501000012
  • Figure 0007717501000013
    Figure 0007717501000013
Patent Text Reader

Abstract

To properly perform causal inference.SOLUTION: A medical information processing apparatus according to an embodiment is provided with a first acquisition unit, a second acquisition unit, a first extraction unit, and a calculation unit. The first acquisition unit acquires a first numerical value corresponding to a result of determination made by a user based on an observed confounder. The second acquisition unit acquires a second numerical value corresponding to a result of determination made by the user based on the observed confounder and first support information for supporting the determination made by the user. The first extraction unit extracts a first difference between the first numerical value and the second numerical value. The calculation unit calculates the degree of influence of an unobserved confounder on the determination made by the user based on the first difference and the observed confounder.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed in this specification and the drawings relate to a medical information processing apparatus and a medical information processing system.

Background Art

[0002] Causal inference is a method for estimating the causal effect of an intervention or exposure on an outcome from data, and is used in a wide range of fields such as medicine, economics, politics, and marketing. In recent years, many methods for estimating individual causal effects from data using machine learning (for example, TARNet, Causal Forest, CMGP, GANITE, X-learner) have been proposed. In such causal inference using machine learning, in order to appropriately estimate the causal effect, it is necessary to identify all confounding factors that affect the causal relationship.

[0003] However, human expertise (domain knowledge) in the target field is theoretically indispensable for identifying confounding factors, and it is generally difficult to identify all confounding factors. Furthermore, since there is no means to strictly verify whether domain knowledge and the results of causal inference from data are correct, there remains room for unobserved confounding factors. As methods for estimating the causal effect in the presence of unobserved confounding factors, for example, randomized controlled trial (RCT), regression discontinuity design (RDD), instrumental variable (IV) method, and front-door criterion can be mentioned, but these are strict and not realistic. Also, many of the recently proposed causal inference methods using machine learning assume the absence of unobserved confounding factors, but the validity of this assumption is disregarded in actual analysis. Therefore, in order to appropriately estimate the causal effect in causal inference using machine learning, it is desirable to quantify the influence degree of unobserved confounding factors.

Prior Art Documents

Patent Documents

[0004] Patent Document 1 Japanese Unexamined Patent Application Publication No. 2020-168397 Summary of the Invention Problems to be Solved by the Invention

[0005] One of the problems to be solved by the embodiments disclosed in this specification and the drawings is to appropriately perform causal inference. However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problems. It is also possible to position, as other problems, the problems corresponding to the respective effects of the respective configurations shown in the embodiments described later. Means for Solving the Problems

[0006] The medical information processing apparatus according to the embodiment includes a first acquisition unit, a second acquisition unit, a first extraction unit, and a calculation unit. The first acquisition unit acquires a first numerical value corresponding to the result determined by the user based on the observed confounding factor. The second acquisition unit acquires a second numerical value corresponding to the result determined by the user based on the observed confounding factor and first support information for assisting the user's determination. The first extraction unit extracts a first difference between the first numerical value and the second numerical value. The calculation unit calculates the influence degree of the unobserved confounding factor on the user's determination based on the first difference and the observed confounding factor. Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

[0008] Hereinafter, a medical information processing apparatus and a medical information processing system according to an embodiment will be described with reference to the drawings. In the following embodiments, parts denoted by the same reference numerals perform the same operations, and redundant descriptions will be omitted as appropriate.

[0009] FIG. 1 is a configuration example of a medical information processing system 100 according to an embodiment. The medical information processing system 100 includes a medical information processing apparatus 1 and a medical record information database 2. In the medical information processing system 100, the medical information processing apparatus 1 and the medical record information database 2 are communicably connected to each other. Note that the medical information processing system 100 may be, for example, a hospital internal network (LAN) constructed within a specific medical institution, or a wide area network (WAN) constructed across a plurality of medical institutions via a network. That is, the medical information processing system 100 may be a network of any scale as long as the above communication path is constructed.

[0010] The medical information processing apparatus 1 is a computer that processes various types of medical-related information. Specifically, the medical information processing apparatus 1 acquires a data set 200 for causal inference (described later in FIG. 5) from the medical record information database 2 and performs various processes to quantify the influence degree of unobserved confounding factors. Note that the medical information processing apparatus 1 may be a workstation capable of executing high-speed processing.

[0011] The medical information database 2 stores various medical information for each patient. The medical information includes, for example, basic information (patient number, age, gender, date of birth, etc.), personal information (height, weight, blood type, medical history, presence or absence of chronic diseases, lifestyle habits (exercise, smoking, diet, drinking, stress, sleep), etc.), and disease information (disease name, stage, frailty score, treatment methods implemented (surgery or medication), prognosis after treatment, etc.). Further, the medical information includes medical images taken by various medical imaging diagnostic devices (CR (Computer Radiography) device, CT (Computed Tomography) device, MRI (Magnetic Resonance Imaging) device, UL (Ultrasound) device, RI (Radio Isotope) device, endoscope device, etc.). In the present embodiment, the medical information database 2 includes a data set 200 for causal inference. Note that the medical information database 2 may be stored in the medical information processing device 1.

[0012] Figure 2 is a configuration example of the medical information processing device 1 according to the embodiment. The medical information processing device 1 includes a processing circuit 11, a memory 12, a display 13, an input interface 14, and a communication interface 15. Each component is communicably connected to each other via a bus which is a common signal transmission path. Note that each component does not necessarily need to be realized by individual hardware. For example, at least two of the components may be realized by one piece of hardware.

[0013] The processing circuit 11 executes various operations by controlling the medical information processing device 1. The processing circuit 11 has processors such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), and a GPU (Graphics Processing Unit) as hardware. The processing circuit 11 realizes each function (for example, the acquisition function 111, the extraction function 112, the calculation function 113, the learning function 114, the update function 115, the estimation function 116, the output function 117) corresponding to each program by executing the program developed in the memory 12 via the processor. Note that each function may be realized by the processing circuit 11 that combines a plurality of processors.

[0014] The acquisition function 111 acquires a first numerical value corresponding to the result judged by the user based on the observed confounding factor. The acquisition function 111 also acquires a second numerical value corresponding to the result judged by the user based on the observed confounding factor and the first support information for assisting the user's judgment. The extraction function 112 extracts a first difference between the first numerical value and the second numerical value. The extraction function 112 also extracts a second difference between the first tendency score and the second tendency score. The first tendency score and the second tendency score are the predicted values of the first numerical value and the second numerical value, respectively. The calculation function 113 calculates the influence degree of the unobserved confounding factor on the user's judgment based on the first difference and the observed confounding factor. The learning function 114 learns the first parameter of the first function and the second parameter of the second function so as to minimize the prediction residual between the first difference and the second difference. The update function 115 updates the model that outputs the first support information using the influence degree of the unobserved confounding factor. The estimation function 116 estimates the causal effect of the user's judgment on the outcome based on the influence degree of the unobserved confounding factor. The output function 117 outputs second support information for assisting the user's judgment based on causal effects. Also, the output function 117 outputs the ratio of the influence degree of unobserved confounding factors in the second support information. Further, the output function 117 outputs candidates for unobserved confounding factors that affect the second support information.

[0015] The memory 12 stores information such as data and programs used by the processing circuit 11. The memory 12 has semiconductor memory elements such as RAM (Random Access Memory) as hardware. Note that the memory 12 may be a drive device that reads and writes information to and from external storage devices such as magnetic disks (floppy (registered trademark) disks, hard disks), magneto-optical disks (MO), optical disks (CD, DVD, Blu-ray (registered trademark)), flash memories (USB flash memories, memory cards, SSDs), and magnetic tapes. Note that the storage area of the memory 12 may be inside the medical information processing device 1 or in an external storage device. In the present embodiment, the memory 12 stores a first function that outputs a first tendency score, which is a predicted value of a first numerical value, with the observed confounding factor as an input, and a second function that outputs a second tendency score, which is a predicted value of a second numerical value, with the observed confounding factor as an input. Further, the memory 12 stores a CDS (Clinical Decision Support) model 3. The memory 12 is an example of a storage unit.

[0016] The CDS model 3 supports the clinical decision-making of users who use the medical information processing device 1. The users include, for example, medical staff such as doctors and nurses who treat patients. In the present embodiment, the CDS model 3 is assumed to output support information for assisting the judgment of a doctor who treats a patient, with a plurality of types of medical information regarding the patient as inputs. Without being limited to this, the CDS model 3 may output information (raw data, predictions, recommendations, etc.) that can change the doctor's judgment. The CDS model 3 is implemented by a machine learning model such as a neural network, for example.

[0017] The display 13 displays data generated by the processing circuit 11, data stored in the memory 12, data output by the CDS model 3, and the like. As the display 13, for example, any display including a cathode ray tube (CRT) display, a liquid crystal display (LCD), a plasma display, an organic electro-luminescence display (OELD), and a tablet terminal can be used.

[0018] The input interface 14 receives an input from a user who uses the medical information processing apparatus 1, converts the received input into an electrical signal, and outputs the electrical signal to the processing circuit 11. As the input interface 14, for example, any operation component including a mouse, a keyboard, a trackball, a switch, a button, a joystick, a touch pad, and a touch panel display can be used. Note that the input interface 14 may be a device that receives an input from an external input device that is separate from the medical information processing apparatus 1, converts the received input into an electrical signal, and outputs the electrical signal to the processing circuit 11.

[0019] The communication interface 15 communicates various data between the medical information processing apparatus 1 and the medical information database 2. As the communication standard, for example, DICOM (Digital Imaging and Communications in Medicine) can be used for communication related to medical image information, and HL7 (Health Level 7) can be used for communication related to medical character information.

[0020] Figure 3 shows an operation example of the medical information processing apparatus 1. In step S101, the medical information processing device 1 acquires a dataset 200 for causal inference by means of an acquisition function 111. Specifically, the medical information processing device 1 accesses a medical information database 2 via a communication interface 15 to acquire the dataset 200 for causal inference. The dataset 200 includes a first numerical value corresponding to the result judged by the user based on an observed confounding factor, and a second numerical value corresponding to the result judged by the user based on the observed confounding factor and first support information for assisting the user's judgment. Note that the dataset 200 may be stored in the medical information database 2 in advance, or may be newly collected by the medical information processing device 1 according to the method shown in FIG. 4.

[0021] In step S102, the medical information processing device 1 learns parameters of a prediction function of a propensity score by means of a learning function 114. Specifically, the medical information processing device 1 uses the acquired dataset 200 to learn a first parameter of a first function for predicting a first propensity score that is a predicted value of the first numerical value, and a second parameter of a second function for predicting a second propensity score that is a predicted value of the second numerical value. Details of parameter learning will be described later with reference to FIG. 6.

[0022] In step S103, the medical information processing device 1 calculates the influence degree of an unobserved confounding factor by means of a calculation function 113. Specifically, the medical information processing device 1 calculates, as the influence degree of the unobserved confounding factor, the difference between the first numerical value and the first propensity score predicted using the learned first parameter, or the difference between the second numerical value and the second propensity score predicted using the learned second parameter.

[0023] In step S104, the medical information processing device 1 estimates a causal effect by means of an estimation function 116. Specifically, the medical information processing device 1 estimates the causal effect of the user's judgment on the outcome based on the calculated influence degree of the unobserved confounding factor. Further, the medical information processing device 1 may update a model (CDS model 3) that outputs first support information for assisting the user's judgment by using the calculated influence degree of the unobserved confounding factor by means of an update function 115.

[0024] In step S105, the medical information processing apparatus 1 outputs support information by the output function 117. Specifically, the medical information processing apparatus 1 or the CDS model 3 outputs second support information for assisting the user's judgment based on the estimated causal effect.

[0025] In step S106, the medical information processing apparatus 1 outputs the influence degree of each confounding factor by the output function 117. Specifically, the medical information processing apparatus 1 outputs the ratio of the influence degree of the unobserved confounding factor in the second support information. Further, the medical information processing apparatus 1 may output candidates for unobserved confounding factors that affect the second support information by the output function 117.

[0026] FIG. 4 is an example of a method for collecting the dataset 200 for causal inference. Hereinafter, as an example of causal inference, attention is paid to the causal relationship between a doctor's judgment regarding a patient's treatment method (also referred to as treatment judgment) and the patient's survival period when the patient is treated based on the judgment. In this causal relationship, the doctor's judgment corresponds to the intervention T (Treatment), and the patient's survival period due to the intervention T corresponds to the outcome Y. At this time, it is considered that there are a plurality of confounding factors that distort the causal relationship between the intervention T and the outcome Y. The plurality of confounding factors are divided into objectively obvious and observable confounding factors (also referred to as observed confounding factors: W) for reasons such as data being obtained, and confounding factors for which data is not obtained, which are not objectively obvious and not observed, and factors for which data is obtained but which are not recognized as confounding factors (also referred to as unobserved confounding factors: U). These confounding factors affect the doctor's judgment T with different degrees of influence and also affect the patient's survival period Y. In the present embodiment, it is assumed that the doctor makes the judgment T while explicitly considering the observed confounding factor W and implicitly considering the unobserved confounding factor U. Note that the influence degree of each confounding factor on the doctor's judgment T is illustrated by arrows of different thicknesses.

[0027] To collect the dataset 200 for causal inference, in this method, the doctor determines the treatment method for the patient before and after the CDS model 3 presents the support information. Here, it is assumed that the influence degrees of the unobserved confounding factor U and the judgment error ε on the doctor's judgment are invariant or constant before and after the presentation of the support information. Conversely, the influence degree of the observed confounding factor W on the doctor's judgment changes before and after the presentation of the support information.

[0028] First, before the presentation of the support information (before CDS presentation), the doctor makes a judgment based on the observed confounding factor W and the unobserved confounding factor U. For example, assume that the observed confounding factor W is age W1 and stage W2, and the unobserved confounding factor U is frailty U1 and gender U2. The doctor makes a first judgment T regarding the treatment method for the patient considering the patient's age W1 and stage W2. Age W1 is a quantitative variable that can take any numerical value, and stage W2 is a qualitative variable with multiple categories. Specifically, the doctor makes the first judgment T while emphasizing the patient's age W1 more than the stage W2. At this time, it is assumed that the doctor implicitly makes the first judgment T while further considering the patient's frailty U1 and gender U2, which are unobserved confounding factors U. Specifically, the influence degree of frailty U1 is slightly higher than that of gender U2.

[0029] The first judgment T is a qualitative variable with multiple categories. In this embodiment, the first judgment T is a binary variable with two categories, "surgery" or "medication". Specifically, using a dummy variable, "surgery" is expressed as "T = 1", and "medication" is expressed as "T = 0". Of course, the first judgment T may be a multi-valued variable with three or more categories. That is, the first judgment T may be represented by an N-dimensional One-hot vector according to the number N (N is a natural number) of each category. The first judgment T is stored in the medical information database 2.

[0030] Subsequently, the medical information processing device 1 displays support information on the display 13 via the CDS model 3. Specifically, the medical information processing device 1 inputs the age W1 and the stage W2, which are the observed confounding factors W before CDS presentation, to the CDS model 3. Based on the input patient age W1 and stage W2, the CDS model 3 outputs support information to assist the doctor's judgment. For example, the CDS model 3 outputs, as support information, a treatment method recommended for the patient (also referred to as recommended treatment). Since the recommended treatment affects the doctor's judgment T' after CDS presentation but does not affect the patient's survival period Y, it is not included in the observed confounding factor W.

[0031] Not limited to this, the CDS model 3 may output support information that also affects the patient's survival period Y. For example, the CDS model 3 may take the patient's age W1 and stage W2 as inputs and output the frailty score W3 of the patient. Since the frailty score W3 affects the doctor's judgment T' after CDS presentation and also affects the patient's survival period Y, it is included in the observed confounding factor W. The doctor reconsiders the judgment of the treatment method for the patient by checking the support information displayed on the display 13. Note that the medical information processing device 1 may present the raw data of the observed confounding factor W to be referred to by the doctor for treatment judgment as support information. That is, the support information may be any factor that can change the doctor's treatment judgment.

[0032] Note that the support information may be a value or a calculated value composed of all or some of the observed confounding factors among the multiple observed confounding factors. As an example, when there are multiple observed confounding factors W1, W2, W3, and W4, the support information may be a value calculated from some of the observed confounding factors W1 and W2.

[0033] Finally, after presenting the support information (after presenting the CDS), the doctor makes a judgment based on the observed confounding factor W, the support information, and the unobserved confounding factor U. For example, the doctor makes a second judgment T' regarding the treatment method for the patient considering the patient's age W1, stage W2, and the recommended treatment presented by the CDS model 3. Here, the doctor makes the second judgment T' emphasizing the stage W2 more than the patient's age W1. As described above, assuming that the influence degrees of the unobserved confounding factor U and the error ε are invariant in the first judgment T and the second judgment T', the change in the doctor's judgment from the first judgment T to the second judgment T' can be regarded as being caused by the change in the influence degree of the observed confounding factor W.

[0034] The second judgment T' is a qualitative variable with multiple categories. In this embodiment, the second judgment T' is a binary variable with two categories of "surgery" or "medication". Specifically, using a dummy variable, "surgery" is expressed as "T' = 1", and "medication" is expressed as "T' = 0". Of course, the second judgment T' may be a multi-valued variable with three or more categories. That is, the second judgment T' may be expressed by an N-dimensional One-hot vector according to the number N (N is a natural number) of each category. In other words, the definitions of the first judgment T and the second judgment T' are the same. The second judgment T' is stored in the medical information database 2.

[0035] Also, the survival period Y of the patient, which is the result of implementing treatment on the patient based on the second judgment T', is stored in the medical information database 2. In this embodiment, the survival period Y is a quantitative variable that can take any numerical value. The survival period Y is the survival period Y (1) when the second judgment T' is "surgery" (T' = 1), and the survival period Y (0) when the second judgment T' is "medication" (T' = 0). (1) For one patient, either Y (0) or Y (1) is observed, but the other is not observed. Therefore, the unobserved outcome Y (0) or Y is also called a potential outcome.

[0036] Through the above series of judgment flowcharts, in the medical information database 2, for one patient, the observed confounding factors W1 and W2, the first judgment T, the second judgment T', and the outcome Y (1) or Y (0) are stored with their respective values associated. By repeating a similar flow for each of multiple patients, a dataset 200 for causal inference with the above respective values associated for each patient is collected. As described above, since an operation similar to having the user make judgments twice is performed in this method, it can be said that the dataset 200 is not pure observational data.

[0037] Figure 5 is an example of the dataset 200 for causal inference. In the dataset 200, for each of N (N is a natural number) patients, the observed confounding factors W1 and W2, the unobserved confounding factor U, the treatment judgments T and T', and the outcome Y (0) or Y (1) are stored with their respective values associated. For each patient, the values of the unobserved confounding factor U and the potential outcome Y (0) or Y (1) are unknown, so cells with unknown values are indicated by "?". Note that the unobserved confounding factors U1 and U2 are simply aggregated and shown as "U".

[0038] For example, for the patient represented by the patient number "1", the respective values are W1 = W1 1 , W2 = W2 1 , T = 1, T' = 1, Y (1) = Y (1) 1 That is, in other words, the patient's age W1 is W1 1 , the disease stage W2 is W2 1 . That is, according to the dataset 200, the doctor selects "surgery" as the treatment judgment T before CDS presentation for the patient, selects "surgery" as the treatment judgment T' after CDS presentation, and as a result of performing "surgery" on the patient based on the latter treatment judgment T', the patient is Y (1) 1It is possible to grasp a case where the patient survived only during a certain period. That is, it can be seen that the doctor's judgment did not change before and after the CDS presentation in this case.

[0039] Similarly, for the patient represented by patient number "2", each value is W1 = W1 2 , W2 = W2 2 , T = 0, T' = 1, Y (1) = Y (1) 2 That is. In other words, the patient's age W1 is W1 2 , the disease stage W2 is W2 2 . That is, according to the dataset 200, the doctor selected "medication" as the treatment judgment T for the patient before the CDS presentation, and "surgery" as the treatment judgment T' after the CDS presentation. As a result of performing "surgery" on the patient based on the latter treatment judgment T', it can be grasped that the patient survived only during a certain period Y (1) 2 . That is, it can be seen that the doctor's judgment changed before and after the CDS presentation in this case.

[0040] Next, the medical information processing device 1 learns based on the dataset 200 for causal inference, and thereby estimates the causal effect Y (1) - Y (0) of the doctor's treatment judgment T on the patient's survival period Y. Here, assume that the prediction formula of the outcome Y for estimating the causal effect Y (1) - Y (0) is represented by the following formula (1). Here, it is assumed that the outcome Y is predicted by a linear model, but the outcome Y may also be predicted by a non - linear model.

Equation

[0041] However, since the values of the unobserved confounding factor U in the dataset 200 are unknown, the partial regression coefficient β U representing the influence of the unobserved confounding factor U on the outcome Y cannot be calculated. Therefore, next, assume the following equation (2) obtained by eliminating the term "+β U U" in equation (1).

Equation

[0042] Therefore, in this embodiment, the medical information processing apparatus 1 estimates the causal effect by using the propensity score e, which is the probability that a patient is assigned to surgery (T = 1). The propensity score e is a function of the observed confounding factors W of 1 or more. Ideally, if the propensity score e is appropriately estimated using all the confounding factors W and U, the causal effect is also appropriately estimated. As shown in FIG. 4, assuming that the influence degree of the unobserved confounding factor U on the doctor's judgment is invariant before and after the CDS presentation, the change amount ΔT of the judgment from the first judgment T to the second judgment T' is predicted from the values of the observed confounding factors W in the dataset 200. The medical information processing apparatus 1 uses the first function f that predicts the first propensity score T ~ which is the predicted value of the first judgment T, and the second function g that predicts the second propensity score T' ~ which is the predicted value of the second judgment T' ~ to predict the change amount ΔT of the judgment. Here, the tilde (

[0043] FIG. 6 is an example of a method for learning the parameters of the prediction function of the propensity score. First, before the CDS presentation, the first function f takes the observed confounding factors W1 and W2 as inputs and outputs the first propensity score T ~ . The first function f is modeled as in the following equation (3) using the first parameters γ1 and γ2 that represent the influence degree of the observed confounding factors on the doctor's judgment before the CDS presentation. Here, it is assumed that the propensity score is predicted by a linear model, but the propensity score may be predicted by a non-linear model.

Equation

[0044] Similarly, after the CDS presentation, the second function g takes the observed confounding factors W1 and W2 as inputs and outputs the second propensity score T'. ~ The second function g is modeled as in the following equation (4) using the second parameters γ'1 and γ'2 that represent the influence degree of the observed confounding factors on the doctor's judgment after the CDS presentation.

Equation

[0045] As described above, the medical information processing apparatus 1 models the first function f and the second function g that respectively predict the true values T and T' of the treatment judgment before and after the CDS presentation. The true value ΔT of the judgment change from before the CDS presentation to after the CDS presentation can be predicted from the observed confounding factor W under the assumption that the influence degree of the unobserved confounding factor U is invariant. That is, the true value ΔT of the judgment change in the difference between before and after the CDS presentation can be predicted using the first function f and the second function g.

[0046] In the difference between before and after the CDS presentation, the third function h takes the observed confounding factors W1 and W2 as inputs and outputs the predicted value ΔT ~ of the judgment change. The third function h is modeled as in the following equation (5) using the first function f and the second function g.

Equation

[0047] Using the first prediction error, the second prediction error, and the third prediction error modeled as described above, the medical information processing apparatus 1 learns the parameters γ1, γ2, γ'1, γ'2. At this time, the loss function L for learning the parameters γ1, γ2, γ'1, γ'2 is expressed as the following formula (6).

Equation

[0048] The medical information processing apparatus 1 learns each of the parameters γ1, γ2, γ'1, γ'2 so as to minimize the value of the loss function L. The learning at this time is specifically expressed by the following formula (7).

Equation

[0049] As described above, the true value ΔT of the judgment change from before CDS presentation to after CDS presentation can be completely predicted from only the observed confounding factor W under the assumption that the influence degree of the unobserved confounding factor U is invariant. That is, in Equation (6), the third prediction residual becomes 0, and only the influence degree of the unobserved confounding factor U that cannot be explained by the observed confounding factor W in the first prediction residual and the second prediction residual remains as a residual. Therefore, the parameters γ1, γ2, γ´1, γ´2 calculated by minimizing the above residual in Equation (7) can be used to calculate the influence degree of the unobserved confounding factor U from Equation (6).

[0050] After the parameters γ1, γ2, γ´1, γ´2 are learned, the medical information processing apparatus 1 calculates the influence degree U´ of the unobserved confounding factor on the doctor's judgment T by the following Equation (8) or (9).

Equation

Equation

[0051] Here, assuming that there is a correlation between the influence degree U´ of the unobserved confounding factor on the doctor's judgment and the influence degree U of the unobserved confounding factor on the outcome, that is, the ratio of the breakdown of the unobserved confounding factor U is invariant, U is replaced by U´. In this way, the medical information processing apparatus 1 estimates the outcome Y using the following Equation (10).

Equation

[0052] Regarding the prediction of the outcome Y, an existing method that combines the propensity score and the prediction of the outcome (doubly robust estimation: Doubly Robust Estimation, X-learner, R-learner, DR-learner, etc.) may be used. Subsequently, the medical information processing apparatus 1 may calculate various causal effects (average treatment effect: ATE (Average Treatment Effect), conditional average treatment effect: CATE (Conditional Average Treatment Effect), individual treatment effect: ITE (Individual Treatment Effect), etc.) using the predicted outcome Y.

[0053] Also, the medical information processing apparatus 1 or the CDS model 3 may output support information based on the predicted causal effect. For example, when the sign of the predicted causal effect Y (1) -Y (0) is positive, the recommended treatment corresponding to the intervention T that causes the outcome Y (that is, T = 1) may be output as support information. Conversely, when the sign of the causal effect Y (1) -Y (1) -Y (0) is negative, when the sign of the causal effect Y (0)The recommended treatment corresponding to the intervention T that causes it (i.e., T = 0) may be output as support information. Further, the medical information processing apparatus 1 or the CDS model 3 may output the ratio of the influence degree of each confounding factor in the support information.

[0054] FIG. 7 is an example of the influence degree of each confounding factor on the support information. FIGS. 7(a) and 7(b) can be displayed on the display 13 of the medical information processing apparatus 1. In FIG. 7(a), the influence degree of each confounding factor in the support information presented by the medical information processing apparatus 1 for each patient (Patient A, Patient B, Patient C) is shown by a bar graph. Specifically, the influence degree of each confounding factor is such that the respective values obtained by normalizing each partial regression coefficient β1, β2, β' in Equation (10) U are equivalent to the ratios of the respective normalized values of β1, β2, β' U to the sum of the respective values. For example, the normalized value of β' U in the sum of the normalized partial regression coefficients β1, β2, β' U corresponds to the influence degree of the unobserved confounding factor U. Note that the influence degree of each original confounding factor before normalization remains unchanged.

[0055] For example, the influence degree of the observed confounding factor W on the support information presented to Patient A is "0.55", and the influence degree of the unobserved confounding factor U is "0.45". Similarly, the influence degree of the observed confounding factor W on the support information presented to Patient B is "0.70", and the influence degree of the unobserved confounding factor U is "0.30". The user who uses the medical information processing apparatus 1 can refer to FIG. 7(a) displayed on the display 13 to confirm the ratio of the influence degree of each confounding factor in the support information output in consideration of the influence degree of the unobserved confounding factor.

[0056] During the display of FIG. 7(a), the user who uses the medical information processing apparatus 1 can operate the input interface 14 to select a bar graph regarding a desired patient. For example, when the bar graph regarding Patient A is selected, the display screen shifts from FIG. 7(a) to FIG. 7(b).

[0057] In FIG. 7(b), the influence degree of the observed confounding factor W and the influence degree of the unobserved confounding factor U are both calculated, and the breakdown of the bar graph is displayed. Here, by analyzing predetermined data, the medical information processing apparatus 1 may display one or more candidates for the unobserved confounding factor U in the window 300. Specifically, in the window 300, “frailty score”, “gender”, “smoking status”,... are displayed as a plurality of candidates for the unobserved confounding factor. As a method for determining candidates for the unobserved confounding factor, for example, a user (data scientist or knowledge-providing physician) who executes and supports data analysis may manually select candidates. Alternatively, for example, the medical information processing apparatus 1 may determine, as candidates for the unobserved confounding factor U, confounding factors that were used in other data processing but were not selected as observed confounding factors in the processing result of the medical information processing apparatus 1.

[0058] To present candidates for the unobserved confounding factor U, for example, the medical information processing apparatus 1 puts one or more unobserved confounding factors U as part of the confounding factor W into the CDS model 3, and calculates the influence degree again in the same manner. If the influence degree of the unobserved confounding factor U decreases by a certain amount or more before and after the processing, the medical information processing apparatus 1 may present the factors put into the CDS model 3 as the above candidates. The above processing is premised on the existence of an unobserved confounding factor U that is obtained as data but not recognized as the observed confounding factor W.

[0059] As described above, the medical information processing apparatus 1 according to the embodiment has been described. The medical information processing apparatus 1 indirectly quantifies the influence degree of an unobserved confounding factor based on the influence degree of the observed confounding factor. According to the medical information processing apparatus 1, the influence degree of an unobserved confounding factor that affects a doctor's judgment can be quantified. As a result, the doctor can quantitatively evaluate the degree of reliability of causal inference. That is, the medical information processing apparatus 1 can improve the reliability of causal inference.

[0060] Here, assume a case where a doctor makes a judgment considering only the observed confounding factors. Similarly, in this case, the medical information processing device 1 acquires a first numerical value corresponding to the doctor's judgment before the presentation of the support information (CDS) and a second numerical value corresponding to the doctor's judgment after the presentation of the support information (CDS). Subsequently, the medical information processing device 1 calculates a first tendency score, which is a predicted value of the first numerical value, and a second tendency score, which is a predicted value of the second numerical value, based on the observed confounding factors. Finally, the medical information processing device 1 calculates the difference between the first numerical value and the first tendency score, or the difference between the second numerical value and the second tendency score, as the influence degree of the unobserved confounding factors. Therefore, when the doctor makes a judgment considering only the observed confounding factors, the influence degree of the unobserved confounding factors is calculated as "0". As a result, the user using the medical information processing device 1 can confirm that the doctor's judgment does not include the influence of unobserved confounding factors.

[0061] According to at least one of the embodiments described above, causal inference can be appropriately performed.

[0062] Although several embodiments have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, changes, and combinations of the embodiments can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope of the invention and the scope equivalent to the invention described in the claims.

Explanation of Reference Numerals

[0063] 1... Medical information processing device 2... Medical treatment information database 3... CDS model 11... Processing circuit 12... Memory 13... Display 14... Input interface 15... Communication interface 100... Medical information processing system 111… Acquisition function 112… Extraction function 113… Calculation function 114… Learning function 115… Update function 116… Estimation function 117… Output function 200… Dataset 300… Window

Claims

1. An acquisition unit that acquires a first numerical value corresponding to a result determined by a user based on an observed confounding factor and a second numerical value corresponding to a result determined by the user based on the observed confounding factor and first support information for assisting the user's determination; A storage unit that stores a first function that outputs a first tendency score, which is a predicted value of the first numerical value, using the observed confounding factor as an input, and a second function that outputs a second tendency score, which is a predicted value of the second numerical value, using the observed confounding factor as an input; An extraction unit that extracts a first difference between the first numerical value and the second numerical value and a second difference between the first tendency score output from the first function and the second tendency score output from the second function; A learning unit that learns a first parameter of the first function and a second parameter of the second function so as to minimize a prediction residual between the first difference and the second difference; A calculation unit that calculates, as an influence degree of an unobserved confounding factor on the user's determination, a difference between the first numerical value and the first tendency score predicted using the learned first parameter, or a difference between the second numerical value and the second tendency score predicted using the learned second parameter; A medical information processing apparatus comprising the above. [[ID= / / ID=7]]

2. An update unit that updates a model that outputs the first support information using the influence degree of the unobserved confounding factor, further comprising the medical information processing apparatus according to Claim 1. The medical information processing apparatus according to Claim 1.

3. An estimation unit that estimates a causal effect of the user's determination on an outcome based on the influence degree of the unobserved confounding factor, further comprising the medical information processing apparatus according to Claim 1 or Claim 2. The medical information processing apparatus according to Claim 1 or Claim 2.

4. A first output unit that outputs second support information for assisting the user's determination based on the causal effect, further comprising the medical information processing apparatus according to Claim 3. The medical information processing apparatus according to Claim 3.

5. A second output unit that outputs a ratio of the influence degree of the unobserved confounding factor in the second support information, further comprising the medical information processing apparatus according to Claim 4. The medical information processing apparatus according to Claim 4.

6. A third output unit that outputs candidates for unobserved confounding factors that affect the second support information, further comprising the medical information processing apparatus according to Claim 4 or Claim 5. The medical information processing apparatus according to Claim 4 or Claim 5.

7. A medical information processing system comprising a medical information database and a medical information processing apparatus, wherein the medical information database is Store a first numerical value corresponding to the result determined by the user based on the observed confounding factor and a second numerical value corresponding to the result determined by the user based on the observed confounding factor and first support information for assisting the user's determination. The medical information processing apparatus, An acquisition unit that acquires the first numerical value and the second numerical value; A storage unit that stores a first function that outputs a first tendency score, which is a predicted value of the first numerical value, using the observed confounding factor as an input, and a second function that outputs a second tendency score, which is a predicted value of the second numerical value, using the observed confounding factor as an input; An extraction unit that extracts a first difference between the first numerical value and the second numerical value and a second difference between the first tendency score output from the first function and the second tendency score output from the second function; A learning unit that learns a first parameter of the first function and a second parameter of the second function so as to minimize a prediction residual between the first difference and the second difference; A calculation unit that calculates, as an influence degree of an unobserved confounding factor on the user's determination, a difference between the first numerical value and the first tendency score predicted using the learned first parameter, or a difference between the second numerical value and the second tendency score predicted using the learned second parameter; A medical information processing system comprising the above.

8. A computer, Acquires a first numerical value corresponding to the result determined by the user based on the observed confounding factor and a second numerical value corresponding to the result determined by the user based on the observed confounding factor and first support information for assisting the user's determination, Stores a first function that outputs a first tendency score, which is a predicted value of the first numerical value, using the observed confounding factor as an input, and a second function that outputs a second tendency score, which is a predicted value of the second numerical value, using the observed confounding factor as an input, Extracts a first difference between the first numerical value and the second numerical value and a second difference between the first tendency score output from the first function and the second tendency score output from the second function, Learns a first parameter of the first function and a second parameter of the second function so as to minimize a prediction residual between the first difference and the second difference, Calculate the difference between the first numerical value and the first tendency score predicted using the learned first parameter, or the difference between the second numerical value and the second tendency score predicted using the learned second parameter as the influence degree of the unobserved confounding factor on the user's judgment. Medical information processing method.

9. On a computer, An acquisition function that acquires a first numerical value corresponding to the result judged by the user based on the observed confounding factor and a second numerical value corresponding to the result judged by the user based on the observed confounding factor and the first support information for assisting the user's judgment; A storage function that stores a first function that outputs a first tendency score, which is a predicted value of the first numerical value, with the observed confounding factor as an input, and a second function that outputs a second tendency score, which is a predicted value of the second numerical value, with the observed confounding factor as an input; An extraction function that extracts a first difference between the first numerical value and the second numerical value and a second difference between the first tendency score output from the first function and the second tendency score output from the second function; A learning function that learns the first parameter of the first function and the second parameter of the second function so as to minimize the prediction residual between the first difference and the second difference; A calculation function that calculates the difference between the first numerical value and the first tendency score predicted using the learned first parameter, or the difference between the second numerical value and the second tendency score predicted using the learned second numerical value as the influence degree of the unobserved confounding factor on the user's judgment; A medical information processing program for realizing the above.

Citation Information

Patent Citations

  • Medical treatment information analysis apparatus, method, and program

    JP2006163465A

  • Treatment selection support system and method

    JP2019095960A

  • Analyzing apparatus and analyzing method

    JP2019125240A

  • Causal relationship for machine learning system

    JP2019194849A

  • Acute care treatment systems dashboard

    JP2020168397A