Analytical device, analytical method, and analytical program
The analytical device enhances treatment effect estimation by stratifying patients based on predictive factors and updating treatment effect differences, addressing inaccuracies in existing systems and improving treatment selection accuracy.
Patent Information
- Application Number
- JP2022092187
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-07
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2042-06-07
AI Technical Summary
The existing comprehensive medical data analysis systems fail to accurately stratify patients based on treatment effects and treat predictive and prognostic factors separately, leading to inaccurate estimation of therapeutic effects.
An analytical device that stratifies patient populations by predictive factors, assigns weights to these factors, and uses a processor to execute a program for improved treatment effect estimation, employing a search process to identify branching conditions and divide patient groups based on predictor factors, updating differences in treatment effects to enhance accuracy.
Improves the accuracy of estimating therapeutic effects by effectively stratifying patients based on predictive factors, enhancing the precision of treatment selection.
Smart Images

Figure 0007766006000007 
Figure 0007766006000008 
Figure 0007766006000009
Abstract
Description
[Technical Field]
[0001] The present invention relates to an analysis device, an analysis method, and an analysis program for analyzing data. [Background technology]
[0002] While traditional medical care has promoted standardization and the creation of guidelines based on randomized controlled trials, it has become apparent that treatments are not effective for all patients and vary from person to person. Therefore, current medical care focuses on pursuing optimal treatment selection tailored to the characteristics of each individual patient. For example, a comprehensive medical data analysis system has been disclosed that subtypes (stratifies) patients based on their characteristics and analyzes the treatments and outcomes of similar patients (see Patent Document 1 below).
[0003] The comprehensive medical data analysis system includes a medical main server including an intelligent medical engine, which is communicatively coupled to a central database that is a confidential electronic medical record database, and is further communicatively coupled to hospitals, clinics, and other medical sources via a network. The intelligent medical engine receives a large number of medical records, potentially from different countries, regions, and continents. The electronic medical records are provided by the hospitals, clinics, and other medical sources and fed into the intelligent medical engine to enable large-scale analysis and correlation of patient medical records on a global scale. The analysis begins by grouping (classifying) the medical records into multiple levels of subgroups according to patient clinical parameters, disease templates, treatments, and outcomes. When a new patient is entered into the system, the patient's parameters and disease templates are matched to the closest subgroup to determine the likelihood of a favorable outcome. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Special Publication No. 2017-502439 [Non-patent literature]
[0005] [Non-Patent Document 1] Athey, Susan, et al, “Recursive partitioning for heterogeneous causal effects” Proceedings of the National Academy of Sciences 113.27 (2016): 7353-7360. Summary of the Invention [Problem to be solved by the invention]
[0006] However, the comprehensive medical data analysis system of Patent Document 1 does not divide the patients into subgroups based on the treatment effect. Furthermore, Non-Patent Document 1 handles factors related to the treatment (predictive factors) and factors unrelated to the treatment (prognostic factors) in the same way when estimating the treatment effect.
[0007] An object of the present invention is to improve the accuracy of estimating the therapeutic effect. [Means for solving the problem]
[0008] An analytical device according to one aspect of the invention disclosed in the present application is an analytical device having a processor that executes a program and a storage device that stores the program, wherein the storage device: Configure patient characteristics Among the factors Reflects sensitivity to treatment The weights for each predictor group are stored. The processor may include: The value of each factor in the above factor group for each patient and treatment selection values and outcome values corresponding to the factors. Acquire multiple patient data, including and an output process that outputs stratification results obtained by the stratification process. In the stratification process, the processor repeatedly executes a search process that searches for a branching condition for branching the patient group based on a specific predictor in the group of factors and its weight, and a division process that divides the patient group using the branching condition searched for by the search process. In the search process, the processor sets at least some of the patient data of the plurality of patient data as a search target group, selects predictor factors in the group of factors for the search target group and their weights, and divides the search target group into a first group and a second group based on the predictor factors. In the search process, the processor performs the following steps: a first treatment effect for the first group using the treatment selection value and the outcome value in the first group; and a second treatment effect for the second group using the treatment selection value and the outcome value in the second group. a difference calculation for calculating a first difference of a first loss function based on the first treatment effect before and after the division and a second difference of a second loss function based on the second treatment effect before and after the division; an update for updating the first difference before the difference calculation to the first difference after the difference calculation if the first difference after the difference calculation is larger than the first difference before the difference calculation, and updating the second difference before the difference calculation to the second difference after the difference calculation if the second difference after the difference calculation is larger than the second difference before the difference calculation; and a branching condition generation for generating a partial causal tree as the branching condition by setting the first group corresponding to the first difference after the update and the second group corresponding to the second difference after the update as the search target group, respectively, and It is characterized by: [Effects of the Invention]
[0009] According to the exemplary embodiment of the present invention, it is possible to improve the accuracy of estimating the therapeutic effect. Problems, configurations, and effects other than those described above will become apparent from the following description of the examples. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is an explanatory diagram showing an example of outcomes of prognostic and predictive factors. [Figure 2] FIG. 2 is an explanatory diagram showing an example of dividing a patient population by predictive factors in patient characteristics that are thought to have a significant effect on the treatment effect τ, and weighting the divided patient population during learning. [Figure 3] FIG. 3 is a block diagram illustrating an example of the hardware configuration of the analysis device. [Figure 4] FIG. 4 is a block diagram illustrating an example of the functional configuration of the analysis device. [Figure 5] FIG. 5 is an explanatory diagram illustrating an example of the weighting table shown in FIG. [Figure 6] FIG. 6 is an explanatory diagram illustrating an example of the healthcare DB shown in FIG. [Figure 7] FIG. 7 is an explanatory diagram showing an example of a patient data table. [Figure 8] FIG. 8 is an explanatory diagram showing an example of an input screen of the analyzer. [Figure 9] FIG. 9 is a flowchart showing an example of an analysis processing procedure performed by the analysis device. [Figure 10] FIG. 10 is an explanatory diagram showing an example of the stratification result. [Figure 11] FIG. 11 is an explanatory diagram showing another example of the stratification result. [Figure 12] FIG. 12 is a flowchart showing a detailed example of the processing procedure of the stratification process (step S902) shown in FIG. [Figure 13] FIG. 13 is a flowchart illustrating a detailed example of the procedure of the branch condition search process (step S1002) shown in FIG. [Figure 14] FIG. 14 is a box plot showing the prediction error improvement rate before and after division between the conventional method and Example 1. [Figure 15] FIG. 15 is a flowchart illustrating an example of a procedure for generating a weight table by the generating unit according to the second embodiment. [Figure 16]FIG. 16 is a histogram showing search results from a medical literature database. [Figure 17] FIG. 17 is a flowchart of an example of a procedure for generating a weight table according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Prognostic and predictive factors for outcome Figure 1 is an explanatory diagram showing an example of outcomes of prognostic and predictive factors. Outcomes are observed values, such as survival or death, time to progression, and tumor size, and are values that inherently include non-treatment-related effects and treatment effects. Non-treatment-related effects and treatment effects cannot be directly observed.
[0012] Graph 101 shows the pre- and post-treatment outcomes of patient groups A and B, where the patient population is grouped by the presence or absence of prognostic factors. Graph 102 shows the pre- and post-treatment outcomes of patient groups C and D, where the patient population is grouped by the presence or absence of predictive factors.
[0013] Prognostic factors and predictive factors are each a factor in a group of factors that constitute the characteristics of a patient (hereinafter referred to as patient characteristics), and are quantitative variables, i.e., covariates, that change depending on the outcome. Prognostic factors are factors that indicate an independent prognosis regardless of whether treatment is performed, such as patient age. Predictive factors are factors that reflect sensitivity to treatment, such as EGFR (Epidermal growth factor receptor), and are factors that show different treatment effects depending on the presence or absence of predictive factors.
[0014] In graph 101, patient group A is a group of patients with low values of prognostic factors indicating age (age low), and patient group B is a group of patients with higher values of prognostic factors indicating age than patient group A (age high). In graph 101, outcomes before and after treatment differ depending on patient groups A and B, but there is no difference in treatment effect τ (difference in outcomes before and after treatment) between patient groups A and B.
[0015] In graph 102, patient group C is a group of patients with high values of predictive factors indicating EGFR (EGFR+), and patient group D is a group of patients with lower values of predictive factors indicating EGFR than patient group C (EGFR-). In graph 102, outcomes before and after treatment vary depending on the patient groups C and D, and there is also a difference in treatment effect τ (difference in outcomes before and after treatment) between patient groups C and D. In graph 102, the treatment effect τ of patient group C is greater than the treatment effect τ of patient group D.
[0016] In this way, stratifying the patient population by predictive factors such as EGFR can support treatment selection through condition classification by treatment effect τ, but if stratification by predictive factors is not performed, the prediction accuracy of treatment effect τ decreases. For this reason, in the example shown below, predictive factors within patient characteristics that are thought to have a significant effect on treatment effect τ are identified in advance and weighted during learning, thereby improving the prediction accuracy of treatment effect τ.
[0017] FIG. 2 is an explanatory diagram showing an example of dividing a patient population by predictors in patient characteristics that are thought to have a significant effect on the treatment effect τ and weighting them during learning. A population 200 includes patients 201 belonging to a treatment group and patients 202 belonging to a non-treatment group. The treatment group is a group of patients who have received treatment for an injury or illness, and the non-treatment group is a group of patients who have not received treatment for an injury or illness. Furthermore, (+) indicates a response, and (-) indicates a non-response. Hereinafter, patients 201 and 202 who responded will be referred to as patients 201(+) and 202(+), and non-responding patients 201 and 202 will be referred to as patients 201(-) and 202(-).
[0018] That is, patient 201(+) is patient 201 whose injury or illness was cured by treatment, and patient 201(-) is patient 201 whose injury or illness was not cured even though he received treatment. Also, patient 202(+) is patient 202 whose injury or illness was cured despite not receiving treatment, and patient 202(-) is patient 202 whose injury or illness was not cured because he received no treatment. In Figure 2, for the sake of simplicity, the set of these six patients 201 and 202 is referred to as population 200.
[0019] Here, the analysis device divides a patient population 200 into two groups using a predictor x in the patient characteristics that is thought to have a significant effect on the treatment effect τ. One group is designated as subtype L, and the other group is designated as subtype R.
[0020] The estimated treatment effect τ(L) for subtype L is the difference between the outcome of patient 201(+) within subtype L and the outcomes of patients 202(+) and 202(-) within subtype L, and corresponds to the difference in treatment effect τ between patient groups C and D in Figure 1.
[0021] The estimated treatment effect τ(R) for subtype R is the difference between the outcomes of patients 201(+) and 201(-) within subtype R and the outcome of patient 202(+) within subtype R, and corresponds to the difference in treatment effect τ between patient groups C and D in Figure 1.
[0022] The analysis device weights the sum of squares of the estimated treatment effects τ(L) and τ(R) with a weight w(x) related to the predictor x that divides the population 200 into subtypes L and R, thereby learning the loss function f using the following equation (1) and predicting the treatment effect τ of the patient to be predicted using the loss function f.
[0023]
number
[0024] Here, l is an index indicating whether the therapeutic effect τ(l) is for subtype L or R. N(l) is the number of samples for subtype L. Details of the analyzer shown in Figs. 1 and 2 will be described below as Examples 1 to 3. [Example]
[0025] In the first embodiment, an analysis device in which the weight w(x) is specified in advance will be described. The present invention is not limited to the following embodiments.
[0026] <Example of hardware configuration of analytical equipment> FIG. 3 is a block diagram showing an example of the hardware configuration of an analysis device. The analysis device 300 includes a processor 301, a storage device 302, an input device 303, an output device 304, and a communication interface (communication IF) 305. The processor 301, the storage device 302, the input device 303, the output device 304, and the communication IF 305 are connected via a bus 306. The processor 301 controls the analysis device 300. The storage device 302 serves as a working area for the processor 301. The storage device 302 is a non-transitory or temporary recording medium that stores various programs and data. Examples of the storage device 302 include a read-only memory (ROM), a random access memory (RAM), a hard disk drive (HDD), and a flash memory. The input device 303 inputs data. Examples of the input device 303 include a keyboard, a mouse, a touch panel, a numeric keypad, a scanner, a microphone, and a sensor. The output device 304 outputs data. The output device 304 may be, for example, a display, a printer, or a speaker. The communication IF 305 connects to a network and transmits and receives data.
[0027] <Example of functional configuration of analytical equipment> 4 is a block diagram showing an example of the functional configuration of an analysis device. Analysis device 300 includes a generation unit 400, an acquisition unit 401, a stratification unit 402, an output unit 403, a healthcare DB 410, a patient data table 420, and a weighting table 430. Healthcare DB 410, patient data table 420, and weighting table 430 are specifically data structures stored in storage device 302 shown in FIG. 3, and are accessible by processor 301. Generation unit 400, acquisition unit 401, stratification unit 402, and output unit 403 are specifically functions realized by processor 301 executing programs stored in storage device 302 shown in FIG. 3.
[0028] The generation unit 400 generates a patient data table 420 by referring to a healthcare DB 410. The acquisition unit 401 acquires multiple pieces of patient data for identifying patients from the patient data table 420, and acquires weights from a weight table 430. The stratification unit 402 stratifies the patient groups acquired as patient data by the acquisition unit 401. The stratification unit 402 has a search unit 411 and an iteration unit 412. The search unit 411 searches for branching conditions for stratifying the patient groups. The iteration unit 412 repeatedly executes the search for branching conditions by the search unit 411 and the division of the patient groups using the branching conditions. The output unit 403 outputs the stratification results by the stratification unit 402.
[0029] Fig. 5 is an explanatory diagram showing an example of the weight table 430 shown in Fig. 4. The weight table 430 has fields of a predictor 501 and a weight 502. A combination of the value of the predictor 501 and the value of the weight 502 in the same row forms an entry that identifies one predictor 501.
[0030] As described above, the predictor 501 is a field that specifies factors that reflect sensitivity to treatment, and holds x1, x2, ..., xi, ..., xn (n is an integer greater than or equal to 1, and i is an integer that satisfies 1≦i≦n) as identification information that uniquely specifies the predictor. Hereinafter, the value of the predictor 501 may be referred to as the predictor xi. The weight 502 is an index value that indicates the significance of the treatment effect τ, and is input into the above formula (1). In this example, the larger the value of the weight 502, the higher the prediction accuracy of the treatment effect τ.
[0031] In the first embodiment, the weight table 430 is prepared in advance. The analysis device 300 can add, change, or delete entries in the weight table 430 and change the values of the weights 502 in response to user operations.
[0032] FIG. 6 is an explanatory diagram showing an example of the healthcare DB 410 shown in FIG. 4. The healthcare DB 410 has the following fields: patient ID 601, admission ID 602, treatment line 603, date 604, procedure 605, event 606, and patient characteristics 607. A combination of values in each field in the same row forms an entry that defines one piece of healthcare information. One or more entries exist for one patient. For example, if a patient is hospitalized three times, three entries exist for that patient. Note that FIG. 6 defines healthcare information for the injury or illness (e.g., cancer) to be analyzed.
[0033] The patient ID 601 is identification information that uniquely identifies a patient. The admission ID 602 is identification information that is assigned when the patient identified by the patient ID 601 is admitted. The treatment line 603 is a number that indicates the order of treatment.
[0034] The treatment line 603 is a number indicating the order of treatment by administering anticancer drugs for cancer. For example, when an anticancer drug is administered for the first time for a certain cancer, it is the first treatment, and the value of the treatment line 603 is "1", for the second treatment it is "2", for the third treatment it is "3", and so on.
[0035] The date 604 is the year, month, and day when treatment was administered by the treatment line 603. The procedure 605 is the content of the treatment administered by the treatment line 603. The event 606 is the result of administering the procedure 605 in the treatment line 603 (for example, progression, death, etc.).
[0036] The patient characteristics 607 are explanatory variables indicating a group of factors that are characteristic quantities of the patient identified by the patient ID 601 as of the date 604, and include covariates. Specifically, the patient characteristics 607 are clinical test values and the presence or absence of gene mutations, and include, for example, factors such as age 671, gender 672, blood pressure 673, and EGFR 674.
[0037] 7 is an explanatory diagram showing an example of a patient data table 420. The patient data table 420 is generated by the acquisition unit 401 with reference to the healthcare DB 410. Note that the patient data table 420 may be stored in the storage device 302 in advance.
[0038] The patient data table 420 is a table that summarizes the healthcare DB 410 on a patient-by-patient basis, and has fields such as a patient ID 601, survival time 701, outcome 702, treatment selection 703, and patient characteristics 607. A combination of values in each field in the same row forms an entry that defines the patient data for one patient.
[0039] If there are multiple entries for one patient in the healthcare DB 410, for example, the entry with the maximum treatment line 603 is used as the entry for the patient data table 420.
[0040] The survival time 701 is the number of days from the date 604 to the date of death of the patient identified by the patient ID 601, which is the value of the event 606. If the event 606 has no value, the number of days is the number of days until the current date.
[0041] The outcome 702 is an observed value, such as survival or death, progression-free time, or tumor size, and is a value that inherently includes both treatment-related and treatment-related effects. In the example of FIG. 7, the value of the outcome 702 is a numerical value that specifies survival or death. For example, "1" indicates survival, and "0" indicates death. The analysis device 300 refers to the event 606, and if the event 606 does not have a value, it stores "1," and if the event 606 contains a date of death, it stores "0."
[0042] Treatment selection 703 is a value indicating whether or not the patient identified by patient ID 601 selected a treatment, with "1" indicating that a treatment was selected and "0" indicating that a treatment was not selected. The analysis device 300 refers to treatment 605, and if treatment 605 has no value, stores "0", and if treatment 605 has a value, stores "1".
[0043] 8 is an explanatory diagram showing an example of an input screen of the analysis device 300. The input screen 800 is displayed on a display device, which is an example of the output device 304 of the analysis device 300, or on a display device of another computer that can communicate with the analysis device 300 via the communication IF 305. In addition, a user can input information into the input screen 800 by operating the input device 303 of the analysis device 300 or an input device of another computer.
[0044] The input screen 800 has a healthcare information setting item 801, a classification setting item 802, a treatment progress item 803, a response variable item 804, an explanatory variable item 805, a missing value processing item 806, a classification model item 807, a weight item 808, and an execute button 809.
[0045] The healthcare information setting item 801 is a user interface that allows a user to select a prediction target entry from the group of entries in the healthcare DB 410 shown in Fig. 6. The classification setting item 802 is a user interface that allows a user to select an item that classifies the group of entries in the healthcare information setting item 801 by classification information such as the patient's cancer stage or genes. This makes it possible to narrow down the group of entries in the healthcare information setting item 801. The treatment progress item 803 is a user interface that allows a user to select the patient's treatment line 603.
[0046] The objective variable item 804 is a user interface that allows the selection of objective variables output from the classification model f. As objective variables, for example, the event 606 or the procedure 605 of the patient to be predicted can be selected. The explanatory variable item 805 is a user interface that allows the selection of factors of the patient characteristics 607 that will be one or more explanatory variables of the patient to be predicted. In the example of FIG. 8, age 671, gender 672, and blood pressure 673 are selected by inputting a check mark.
[0047] The missing value handling item 806 is a user interface that allows the user to select how to handle missing values of explanatory variables. In the example of FIG. 8, "interpolation" is selected as the missing value handling. The classification model item 807 is a user interface that allows the user to select a classification model f. In the example of FIG. 8, a causal tree is selected as the classification model f.
[0048] The weight item 808 displays the weights 502 of the explanatory variables corresponding to the predictors 501 among the explanatory variables selected in the explanatory variable item 805. The user may refer to the weights 502 and deselect the explanatory variables in the explanatory variable item 805. For example, because the weight 502 of gender 672 is "1.0" and lower than the other weights 502, the user may exclude gender 672 from the explanatory variable item 805. The execute button 809 is a user interface that, when pressed, causes the analysis device 300 to execute the analysis process.
[0049] <Analysis processing> 9 is a flowchart showing an example of an analysis processing procedure by the analysis device 300. The analysis device 300 generates the patient data table 420 from the healthcare DB 410 using the acquisition unit 401 if the patient data table 420 has not yet been generated. Then, the analysis device 300 acquires the patient data, which is an entry, from the patient data table 420 using the acquisition unit 401 (step S901).
[0050] Next, the analysis device 300 executes stratification processing using the stratification unit 402 (step S902). The stratification processing (step S902) is processing for stratifying patients using patient data. After this, the analysis device 300 outputs the stratification results obtained by the stratification processing (step S902) using the output unit 403 (step S903), thereby completing the series of analysis processing. In step S903, the analysis device 300 may display the stratification results on a display, which is an example of the output device 304, may transmit the stratification results to another computer via the communication IF 305, or may store the stratification results in the storage device 302.
[0051] <Stratification results> FIG. 10 is an explanatory diagram showing an example of a stratification result. The stratification result shown in FIG. 10 is a causal tree 1000 with a tree structure. The causal tree 1000 is composed of nodes 1001 to 1005. At node 1001, the analysis target group, whose average treatment effect is "3," is divided into a patient group for which the predictor x1>0 and a patient group for which this is not the case. This predictor x1 and the division threshold value "0" for dividing the analysis target group are the branching conditions for node 1001. The patient group for which factor x1>0 is node 1002, which indicates patient group A, whose average treatment effect is "10," and the patient group for which factor x1>0 is not node 1003, whose average treatment effect is "1."
[0052] At node 1003, the patient group to be split, whose average treatment effect is "1," is split into a patient group whose predictor x2 > 0 and a patient group whose average treatment effect is not. The splitting threshold value "0" for splitting the group to be split is the branching condition of node 1003. The patient group whose predictor x2 > 0 becomes node 1004, which indicates patient group B whose average treatment effect is "0," and the patient group whose predictor x2 > 0 is not becomes node 1005, which indicates patient group C whose average treatment effect is "-5."
[0053] There are no branching conditions at nodes 1002, 1004, and 1005. Causal tree 1000 is formed by nodes 1001 to 1005, the connections between nodes 1001 to 1005, and the branching conditions at nodes 1001 and 1003.
[0054] The splitting threshold is, for example, a value of a predictor used to split the patient groups so that the number of patients in each group is equal. For example, it may be the minimum value of the predictor used in the splitting within a patient group with a large value of the predictor, the maximum value of the predictor used in the splitting within a patient group with a small value of the predictor, or the average value of the minimum and maximum values of the predictor.
[0055] Figure 11 is an explanatory diagram showing another example of the stratification result. The stratification result 1100 shown in Figure 11 is an example shown in a graph. The stratification result 1100 is a scatter plot that graphically depicts the relationship between covariates, Factor 1 and Factor 2, and the analysis target group is divided into patient groups A, B, and C. The covariates are not limited to the combination of Factor 1 and Factor 2, and other combinations can also be selected.
[0056] Furthermore, when the user operates the input device 303 to specify each of patient groups A, B, and C, the analysis device 300 may display characteristic information of the specified patient group. In Fig. 11, when patient group B is specified, characteristic information 1101 of patient group B is displayed.
[0057] <Stratification processing> 12 is a flowchart showing a detailed example of the processing procedure of the stratification process (step S902) shown in FIG. The analysis device 300 sets an analysis target group using the iterator 412 (step S1201). Specifically, for example, when step S1201 is executed for the first time, the analysis device 300 selects an analysis target group for the first execution from the patient data acquired in step S901. The analysis target group for the first execution may be the patient data or all entries in the patient data table 420, or may be a portion of the patient data that meets preset conditions, as long as it is one or more patient data.
[0058] Furthermore, the analysis device 300 sets an execution label [K, V] to the analysis target group when step S1201 is executed for the first time. For example, the execution label [K, V] is a combination of a key K and a value V. When step S1201 is executed for the first time, the key K is set to 1 and the value V is set to False. False indicates that the branch condition search process (step S1202) has not been executed, and once the branch condition search process (step S1202) is executed, the value V is updated to True, which indicates that the branch condition search process (step S1202) has been executed.
[0059] Next, the analysis device 300 executes a branching condition search process (step S1202) using the search unit 411. The branching condition search process (step S1202) is a process for searching for a condition (branching condition) for branching the analysis target group and generating a causal tree.
[0060] Next, the analysis device 300 causes the search unit 411 to update the value V=False of the execution label [K, V] of the analysis target group to the value V=True, which indicates that the branch condition search process (step S1202) has been executed (step S1203).
[0061] Next, the analysis device 300 determines, via the iterator 412, whether the treatment effect has changed before and after the division of the analysis target group (step S1204). Specifically, for example, the analysis device 300 tentatively divides the analysis target group to be divided based on the branching conditions of the causal tree, and generates two patient groups (hereinafter referred to as the first branch group and the second branch group; when no distinction is made, they will simply be referred to as branch groups). The analysis device 300 determines whether the treatment effect of either the first branch group or the second branch group has changed significantly compared to the treatment effect of the analysis target group to be divided.
[0062] For example, the analysis device 300 calculates a standard deviation that combines the difference in treatment effect between the first branch group and the analysis target group (hereinafter referred to as the first difference) and the difference in treatment effect between the second branch group and the analysis target group (hereinafter referred to as the second difference).The analysis device 300 then determines whether at least one of the first difference and the second difference is larger than the standard deviation.
[0063] The branch group, which is the comparison source for a difference larger than the standard deviation, is determined to have changed in treatment effect from the analysis group before splitting. If at least one of the first difference and the second difference is larger than the standard deviation, it is determined that the treatment effect has changed (step S1204: Yes), and the process proceeds to step S1205. If both the first difference and the second difference are equal to or smaller than the standard deviation, the process proceeds to step S1206.
[0064] Also , minutesIf the loss function does not improve in the branch condition search process (step S1202) (i.e., if None is returned as the branch condition search result), the analysis device 300 determines that there is no change in the treatment effect (step S1204: No) and proceeds to step S1206.
[0065] After step S1204: Yes, the analysis device 300 divides the analysis target group using the branching condition used in the provisional division in step S1204 (step S1205). Specifically, for example, the analysis device 300 divides the analysis target group at the parent node in the first step S1205, and when the loop is entered at step S1206: No, in the next step S1205, the analysis device 300 divides the analysis target group at the branched child node.
[0066] The analysis device 300 also assigns an execution label to each of the two groups split in step S1205, i.e., the first branch group and the second branch group. Specifically, for example, the analysis device 300 duplicates the execution label [K, V] of the analysis target group for each of the first branch group and the second branch group. The analysis device 300 then assigns a branch number "1" to the end of the key K in the execution label [K, V] of the first branch group and updates the value V from V=True to V=False. Similarly, the analysis device 300 assigns a branch number "2" to the end of the key K in the execution label [K, V] of the second branch group and updates the value V from V=True to V=False.
[0067] For example, if the execution label [K,V] of the analysis target group is [1,True], the execution label [K,V] of the first branch group will be [11,False], and the execution label [K,V] of the second branch group will be [12,False]. Then, proceed to step S1206.
[0068] The analysis device 300 determines whether a termination condition is satisfied (step S1206). The termination condition may be, for example, a preset number of executions of the group division (step S1205) (i.e., the branching depth) or a lower limit value for the number of samples in a group. Specifically, for example, if the number of executions of the group division (step S1205) is less than a predetermined number, the termination condition is deemed not to be satisfied (step S1206: No), and the process returns to step S1201. On the other hand, if the number of executions of the group division (step S1205) is greater than or equal to the predetermined number, the value V of each of the first and second branch groups is updated from V=False to V=True, the termination condition is deemed to be satisfied (step S1206: Yes), the stratification process (step S902) is terminated, and the process proceeds to step S903.
[0069] If the termination condition is the lower limit of the number of samples in a group, the analysis device 300 executes the group division (step S1205) and determines whether the number of samples in each of the first and second branch groups is below the lower limit of the number of samples in a group. If at least one of the first and second branch groups is below the lower limit of the number of samples in a group, the analysis device 300 determines that the termination condition is not met (step S1206: No) and returns to step S1201. On the other hand, if both the first and second branch groups are equal to or greater than the lower limit of the number of samples in a group, the analysis device 300 updates the value V of each of the first and second branch groups from V=False to V=True, determines that the termination condition is met (step S1206: Yes), terminates the stratification process (step S902), and proceeds to step S903.
[0070] If the treatment effect has not changed (step S1204: No), the analysis device 300 determines whether the number of samples in the analysis group is below the lower limit of the number of samples in a group. If the number of samples in the analysis group is below the lower limit of the number of samples in a group, the end condition is not met (step S1206: No), and the process returns to step S1201. On the other hand, if the number of samples in the analysis group is equal to or greater than the lower limit of the number of samples in a group, the value V of each of the first and second branch groups is updated from V=False to V=True, the end condition is met (step S1206: Yes), the stratification process (step S902) is terminated, and the process proceeds to step S903.
[0071] That is, if there is a group in which the value V of the execution label [K, V] is "False", it is determined that the termination condition is not satisfied (step S1206: No), and the process returns to step S1201.
[0072] If step S1206: No and the process returns to step S1201, the analysis device 300 sets the group whose execution label [K, V] has a value of "False" as the next analysis target group (step S1201), and similarly executes steps S1202 to S1206.
[0073] In the example of group division (step S1205) described above, the execution label [K, V] of the first branch group is [11, False], and the execution label [K, V] of the second branch group is [12, False]. Therefore, the first branch group and the second branch group are each set as an analysis target group (step S1201), and steps S1202 to S1206 are executed for each analysis target group.
[0074] Here, a specific description will be given using the causal tree 1000 shown in FIG. 10 as an example. First, at the time of initial execution, the analysis device 300 provisionally divides the analysis target group into a first branch group (x1>0: Yes) and a second branch group (x1>0: No) based on the branch condition (x1>0) of node 1001. Here, it is assumed that the treatment effect has changed for either the first branch group (x1>0: Yes) or the second branch group (x1>0: No) (step S1204: Yes). As a result, the analysis device 300 divides the analysis target group into the first branch group (x1>0: Yes) and the second branch group (x1>0: No) based on the branch condition (x1>0) of node 1001 (step S1205).
[0075] In addition, the analysis device 300 uses the execution label [1, True] of the analysis target group to generate the execution label [11, False] of the first branch group (x1>0: Yes) and the execution label [12, False] of the second branch group (x1>0: No).
[0076] The first branch group (x1>0: Yes) transitions to node 1002. Because node 1002 does not have a branch condition, the analysis device 300 terminates the search for the first branch group (x1>0: Yes) (step S1206: Yes) and updates its execution label [11, False] to [11, True].
[0077] The execution label of the second branch group (x1>0: No) is [12, False], and the value V is False. Therefore, the analysis device 300 sets the second branch group (x1>0: No) as the next analysis target group (step S1206: No → S1201).
[0078] The analysis device 300 identifies the node 1002 in the causal tree 1000 to which the analysis target group (x1>0: No) transitions, and updates its execution label [12, False] to [12, True].
[0079] Then, the analysis device 300 provisionally divides the analysis target group (x1>0: No) into a third branch group (x2>0: Yes) and a fourth branch group (x2>0: No) using the branching condition (x2>0). Here, it is assumed that the treatment effect has changed for either the third branch group (x2>0: Yes) or the fourth branch group (x2>0: No) (step S1204: Yes). The analysis device 300 divides the analysis target group (x1>0: No) into the third branch group (x2>0: Yes) and the fourth branch group (x2>0: No) using the branching condition (x2>0) (step S1205).
[0080] In addition, the analysis device 300 uses the execution label [12, True] of the analysis target group (x1>0: No) to generate the execution label [123, False] of the third branch group (x2>0: Yes) and the execution label [124, False] of the fourth branch group (x2>0: No).
[0081] The third branch group (x2>0: Yes) transitions to node 1004. Because node 1004 does not have a branch condition, the analysis device 300 terminates the search for the third branch group (x2>0: Yes) (step S1206: Yes) and updates its execution label [123, False] to [123, True].
[0082] Similarly, the fourth branch group (x2>0: No) transitions to node 1005. Because node 1005 does not have a branch condition, the analysis device 300 terminates the search for the fourth branch group (x2>0: No) (step S1206: Yes) and updates its execution label [124, False] to [124, True].
[0083] Then, the analysis device 300 outputs the run labels generated so far, the groups corresponding to the run labels, and the branching conditions used for the division as stratification results.
[0084] 9, the analysis device 300 outputs, as the stratification result, a causal tree, which is a tree structure from the initial analysis target group to the terminal branch group, via the output unit 403. At this time, the execution labels of each group in the stratification result may be reassigned in ascending order starting from 0, with the initial analysis target group as the starting position.
[0085] In this way, in the stratification process (step S902), a search is performed to maximize the therapeutic effect for each branch group generated by branching, and stratification that maximizes the therapeutic effect is realized.
[0086] <Branch Condition Search Process (Step S1002)> Fig. 13 is a flowchart showing a detailed example of the processing procedure of the branching condition search process (step S1002) shown in Fig. 10. The search unit 411 reads the weights 502 of the predictors 501 from the weight table 430 (step S1301).
[0087] Next, the searching unit 411 acquires a search target group from the analysis target group (step S1302). Specifically, for example, the searching unit 411 may use the analysis target group as the search target group as is, or may divide the analysis target group into training data and validation data. In this case, the training data becomes the search target group, and the validation data is used in treatment effect estimation (step S1306).
[0088] Next, the search unit 411 randomly selects factors that are covariates within the search target group, creates a list of the selected factors (factor list) (step S1303), and creates a list of values of the selected factors (factor value list) (step S1304). The factor list is a list of fields that indicate factors that are covariates, such as age 671, blood pressure 673, and EGFR 674. The factor group selected for the factor list is a group of factors whose number is smaller than all factors. A causal tree is created for each factor list.
[0089] The factor value list is a list including values (56 [years], 62 [years], ..., 90 [ml], 127 [ml], ...) of selected factors such as age 671, blood pressure 673, and EGFR 674.
[0090] Also, in step S1304, the search unit 411 identifies a preset predictor from the factor list, and extracts the value of the identified predictor (hereinafter, search target predictor) from the factor value list.
[0091] Through steps S1301, S1303, and S1304, the search unit 411 selects unselected predictors and their weights.
[0092] Next, the search unit 411 divides the search target group into two using the search target predictor (step S1305). This data division is a process of dividing the data into subtypes L and R according to the patient characteristics shown in FIG. 2. Each time the process returns from steps S1311 and S1312, a different predictor is selected as the search target predictor. Note that one of the divided groups is referred to as subtype L, as in FIG. 2, and the other group is referred to as subtype R.
[0093] Next, the searching unit 411 calculates the therapeutic effect τ for each of the subtypes L and R (step S1306). The therapeutic effect τ is calculated by the following formula (2).
[0094] τ(l)=E[Y|T=1]-E[Y|T=0]···(2)
[0095] If the subtype is L, then l = L, and if the subtype is R, then l = R. Y is the outcome (for example, event 606). T is a binary variable indicating treatment selection, with T = 1 indicating that treatment was selected (treatment 605 was performed), and T = 0 indicating that treatment was not selected (treatment 605 was not performed). Furthermore, E[] is the expected value calculation operator. E[] is, for example, the sum of outcome Y. The second treatment effects, treatment effects τ(L) and τ(R), are calculated using the above formula (2). When there is no need to distinguish between treatment effects τ(L) and τ(R), they are written as τ(l) (where l = L, R).
[0096] Next, the search unit 411 calculates the loss functions before and after division using the treatment effects τ(L) and τ(R) (step S1307). The loss function before division is LossPre, and the loss function after division is LossPost. First, the loss function before division, LossPre, is shown in the following formula (3).
[0097]
number
[0098] In the above formula (3), N on the right-hand side is the number of samples in the search target group. Also, τ on the right-hand side is the treatment effect before division, which is the first treatment effect. At the first execution, the treatment effect τ at the parent node is used. From the second loop onwards, the treatment effect τ(l) after the previous division becomes the treatment effect τ before division.
[0099] Also, x is the search target predictor identified in step S1305 among the predictors 501 (x1, x2, ..., xi, ..., xn), and W(x) is the weight 502 of the search target predictor.
[0100] Furthermore, when the analysis target group is divided into training data and validation data in step S1302, the loss function LossPre before division is expressed as the following formula (4) by adding a penalty term due to variance to the above formula (3).
[0101]
number
[0102] N on the right side of the above formula (4) train is the number of samples in the training data, i.e., the number of samples in the search target set, N. est is the number of samples in the validation data. T=1 is the variance of the sample belonging to treatment selection T=1 in the search target group, and S T=0 is the variance of samples in the search target group that belong to treatment selection T=0. Also, p is the proportion of samples in the search target group that belong to treatment selection T=1.
[0103] Furthermore, the entire right-hand sides of the above formulas (3) and (4) may be normalized by dividing them by the number of samples N in the search target group.
[0104] Next, the loss function LossPost after division is shown in the following formula (5): The loss function LossPost after division is a loss function that maximizes each estimated treatment effect τ(l).
[0105]
number
[0106] In the above formula (5), N(l) on the right side is the number of samples of subtype l. If the entire right sides of the above formulas (3) and (4) are normalized by dividing them by the number of samples N in the search target group, the entire right side of the above formula (5) may be normalized by dividing it by the number of samples in the search target group (total number of samples of subtypes L and R). Furthermore, val is a threshold value for dividing the range of factor x. W(x) may be used without using val.
[0107] Next, the search unit 411 calculates the differential Gain between the loss functions LossPre and LossPost before and after the division (step S1308). The differential Gain is an index indicating whether the loss function LossPost has improved as a result of the division.
[0108] Gain=LossPost-LossPre...(6)
[0109] Next, the search unit 411 determines whether the current differential gain is greater than the differential gain currently held (step S1309). The differential gain currently held is the differential gain held in step S1310 of the previous loop, and is used as the target value. However, at the time of the first execution, there is no differential gain currently held, so 0 is used as the initial value of the differential gain currently held.
[0110] If the current differential Gain is greater than the held differential Gain (step S1309: Yes), the search unit 411 updates the loss function before division applied this time, LossPre, with the loss function LossPost to obtain the new loss function before division, updates the held differential Gain with the current differential Gain, and obtains the branching condition when the two-division was performed in step S1305. In this way, the branching condition is searched for. Then, the process proceeds to step S1311.
[0111] On the other hand, if the current differential Gain is not greater than the held differential Gain (step S1309: No), the search unit 411 proceeds to step S1311 without updating the loss function LossPre before division and the held differential Gain.
[0112] Next, the searching unit 411 determines whether the division of the search target group into two (step S1305) satisfies a termination condition (step S1311). The termination condition is, for example, when there are no remaining predictors 501 that can be selected as search targets. If the division of the search target group into two (step S1305) does not satisfy the termination condition (step S1305: No), that is, when there are remaining predictors 501 that can be selected as search targets, the process returns to step S1304. In this case, the searching unit 411 sets each of the subtypes L and R that were determined to have a difference greater than the previous difference in step S1309 to be the next search target group.
[0113] On the other hand, if the termination condition is met (step S1311: Yes), that is, if there are no remaining predictors 501 that can be selected as search targets, one causal tree has been created, and the search unit 411 saves the created causal tree and proceeds to step S1312.
[0114] Next, the search unit 411 determines whether or not a termination condition for creating a causal tree has been satisfied (step S1312). The termination condition is, for example, a threshold value for the number of causal trees. If the termination condition has not been satisfied (step S1312: No) (if the number of created causal trees has not reached the threshold value), the process returns to step S1303, and the search unit 411 recreates the factor list.
[0115] On the other hand, if the termination condition is satisfied (step S1312: Yes), the search unit 411 outputs the created causal tree and proceeds to step S1203. As a result, the number of causal trees created is equal to the threshold value set in step S1312. Of the node groups constituting the causal tree, a node having a branch destination node includes the predictor and division threshold value used when the group was divided at that node.
[0116] <Simulation results> Next, the simulation results of the first embodiment will be explained with reference to FIG.
[0117] 14 is a box plot showing the prediction error improvement rate before and after division between the conventional method and Example 1. The conventional method is a method for calculating the prediction error improvement rate using an equation obtained by removing W(x) from the above equations (3) and (5).
[0118] Y j =η(x j )+T j τ(x j )···(7)
[0119] The above formula (7) is the formula for calculating the outcome. The subscript j is the patient ID 601. Y on the left side j is the outcome of the patient whose patient ID 601 value is j (hereinafter referred to as patient j). j) is the prognostic factor x of patient j j These effects are not related to treatment with T. j is the treatment choice T (= 0 or 1) for patient j. τ(x j ) is the predictor x j This is the therapeutic effect of
[0120] Here, η(x j ) is expressed by the following formula (8).
[0121]
number
[0122] Also, τ(x j ) is expressed by the following formula (9).
[0123]
number
[0124] The above formulas (8) and (9) are formulas showing a method of generating data by simulation, and table data similar to that shown in Figure 7 is created. The number of samples N for patient j is set to N=1000, and the treatment selection T j Here, among the factors x1 to x8, the factors x1 and x2 have a very large value of the weight 502 compared to the other factors x3 to x8.
[0125] In this simulation, the prediction error reduction rate before and after division was calculated using RMSE (Root Mean Square Error) to evaluate accuracy. In Example 1, since weighting is used, it can be confirmed that the prediction error improvement rate is improved and the coefficient of variation (CV) is significantly reduced. [Example]
[0126] Next, a second embodiment will be described. In the first embodiment, the explanation is given on the assumption that the weight table 430 exists, but in the second embodiment, the analysis device 300 generates the weight table 430. That is, in the second embodiment, the analysis device 300 generates the weight table 430 by using the generation unit 400, with reference to the patient data table 420. Note that in the second embodiment, differences from the first embodiment will be mainly explained, and therefore explanations of commonalities with the first embodiment will be omitted.
[0127] 15 is a flowchart illustrating an example of a procedure for generating a weight table 430 by the generating unit 400 according to the second embodiment. The generating unit 400 randomly samples entries defining patient data from the patient data table 420 (step S1501). The number of samples is set arbitrarily, for example, to 50%, 70%, or the like, of all samples in the patient data table 420. The generating unit 400 may also use samples that have not been sampled as verification data.
[0128] Next, the generating unit 400 outputs the group of samples sampled in step S1501 to the stratifying unit 402, and calls and executes the stratification process (step S902) shown in FIG. 12 from the stratifying unit 402 (step S902).
[0129] Next, the generating unit 400 acquires the value of the predictor 501 and the division threshold for each predictor 501 used in division from each branch group that is the stratification result of the stratification process (step S902) (step S1503).
[0130] Thereafter, the generation unit 400 determines whether or not a termination condition has been satisfied (step S1504). Specifically, the termination condition is, for example, when steps S1501 to S1503 have been executed a predetermined number of times. If the termination condition has not been satisfied (step S1504: No), that is, when steps S1501 to S1503 have not been executed a predetermined number of times, the generation unit 400 returns to step S1501. On the other hand, if the termination condition has been satisfied (step S1504: Yes), that is, when steps S1501 to S1503 have been executed a predetermined number of times, the generation unit 400 calculates weights 502 for each predictor 501 and stores them in the weight table 430 (step S1505).
[0131] Specifically, for example, the generating unit 400 calculates a statistic between the value of the predictor 501 and the division threshold for each predictor 501, and sets the calculated value as the weight 502. More specifically, for example, the weight 502 may be the difference between the maximum value of the predictor 501 and the division threshold, the difference between the median value of the predictor 501 and the division threshold, the difference between the most frequent value of the predictor 501 and the division threshold, or the difference between the average value of the predictor 501 and the division threshold. Alternatively, the weight 502 may be the number of occurrences of the value of the predictor 501.
[0132] In this way, analysis device 300 automatically learns the weights as medical knowledge. Therefore, the weight 502 can be increased for predictors that are used as branching conditions, thereby improving the accuracy of estimating the therapeutic effect.
[0133] 9, when the stratification process (step S902) is executed in FIG. 9, the generation unit 400 may use the stratification results to update the weight table 430. As a result, the more analyses are performed by the analysis device 300, the more reliable the weight table 430 becomes, and the more accurate the estimation of the therapeutic effect becomes.
[0134] Furthermore, in the first embodiment, an arbitrarily created weight table 430 is applied, but in the second embodiment, a computer having a generation unit 400 other than the analysis device 300 may generate the weight table 430 by the generation process according to the second embodiment, and the analysis device 300 may acquire the weight table 430 from the computer. [Example]
[0135] Next, a third embodiment will be described. In the first embodiment, the explanation is given on the assumption that the weight table 430 exists, but in the third embodiment, the analysis device 300 generates the weight table 430. That is, in the third embodiment, the analysis device 300 generates the weight table 430 by using the generation unit 400, referring to a medical literature database such as PubMed. Note that the third embodiment will mainly be explained with reference to differences from the first embodiment, and therefore explanations of commonalities with the first embodiment will be omitted.
[0136] Specifically, for example, the analysis device 300 causes the generation unit 400 to execute an abstract search on a medical literature database, statistically process the appearance rates of related terms, and set the results of the statistical processing as the weights 502 of the predictors 501. In this way, the analysis device 300 automatically learns medical knowledge.
[0137] FIG. 16 is a histogram showing search results from a medical literature database. The vertical axis of histogram 1600 is a column of factors contained in sentences searched by search keywords. The search keywords may be, for example, names of risk factors. The search keywords may also include conjunctions related to outcomes, such as "cause" and "relate."
[0138] 16 is factor weight 502. The generation unit 400 calculates the value of weight 502 so that the greater the number of times the search keyword appears in sentences searched for using the search keyword or the greater the number of sentences searched for using the search keyword, the higher the value of weight 502. However, if a sentence searched for using the search keyword contains a negative word such as "not," the generation unit 400 does not calculate a high value for weight 502, or calculates a low value for weight 502.
[0139] The generation unit 400 excludes factors whose weight 502 values are less than a predetermined threshold or are in the top k+1 or lower, and stores factors whose weight 502 values are greater than a predetermined threshold or are in the top k as predictive factors 501 together with the weight 502 in the weight table 430.
[0140] 17 is a flowchart illustrating an example of a process procedure for generating the weight table 430 according to the third embodiment. The generating unit 400 sets a search keyword by a user operation (step S1701). Next, the generating unit 400 transmits the search keyword to a medical literature database, searches the abstracts of each document in the medical literature database, and acquires the abstracts of documents corresponding to the search keyword from the medical literature database (step S1702).
[0141] Next, generation unit 400 searches the abstract acquired in step S1702 for factors included in the search keyword, and extracts sentences including the factors (step S1703).
[0142] Next, the generation unit 400 searches the sentences extracted in step S1703 for a conjunction related to outcome (for example, "cause" or "relate"), and increments the positive relation count Cpos for sentences containing the conjunction. The positive relation count Cpos is an evaluation value for sentences in which the relationship between the factor and the conjunction is positive, and the higher the count value, the greater the weight 502. On the other hand, if a sentence searched for using a conjunction related to outcome contains a negative word such as "not," the generation unit 400 increments the negative relation count Cneg.
[0143] Next, the generating unit 400 calculates the weight 502 for each factor (step S1705). The weight 502(w) is calculated, for example, by the following formula (10).
[0144] w=Cpos / C neg ···(10)
[0145] If the negation relation count Cneg in the denominator is never counted, Cneg=0 and calculation becomes impossible. Therefore, even if Cneg=0, equation (1) may be modified so that the denominator of equation (10) does not become 0.
[0146] Next, the generating unit 400 stores the calculated weights 502 in the weight table 430 (step S1706).
[0147] Thereafter, the generating unit 400 determines whether or not a termination condition has been met (step S1704). Specifically, the termination condition is, for example, when weights 502 have been calculated for all of the factors found in step S1703. If there are factors for which weights 502 have not been calculated (step S1707: No), the process returns to step S1703. On the other hand, if there are no factors for which weights 502 have not been calculated (step S1707: Yes), the generating unit 400 terminates this example of processing.
[0148] In this way, analysis device 300 automatically learns using medical knowledge as weights. Therefore, the weight 502 increases as factors are retrieved from a medical literature database, and when factors with medical evidence from medical literature are used as predictive factors, the accuracy of estimating treatment effects can be improved.
[0149] In the third embodiment, the abstracts of medical literature are searched, which enables the speed of the process of generating the weight table 430 to be increased compared to when the medical literature itself is searched. On the other hand, the generation unit 400 may search the medical literature itself. This improves the reliability of the weights 502 and the accuracy of estimating the therapeutic effect compared to when the abstracts of medical literature are searched.
[0150] Also, Examples 3 In the above, an arbitrarily created weight table 430 is applied, but in the first embodiment, a computer having a generation unit 400 other than the analysis device 300 may generate the weight table 430 by the generation process according to the third embodiment, and the analysis device 300 may acquire the weight table 430 from the computer.
[0151] As described above, the above-described analysis device 300 improves classification accuracy when stratifying patients by factors that contribute to treatment effects by weighting predictive factors inferred in advance from empirical knowledge and medical literature. This improves the accuracy of estimating treatment effects, enabling more accurate patient stratification.
[0152] In this way, the analysis device 300 can directly classify patients into subtypes based on the estimated therapeutic effect according to the patient's characteristics. Therefore, stratified patient groups are classified into subtypes with different therapeutic effects, which is expected to contribute to the selection of optimal treatments suited to the characteristics of individual patients. Therefore, it becomes possible to identify subtypes for which a certain drug is expected to be therapeutically effective.
[0153] The present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the spirit and scope of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to configurations including all of the described configurations. Furthermore, part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Furthermore, the configuration of another embodiment may be added to the configuration of one embodiment. Furthermore, part of the configuration of each embodiment may be added to, deleted from, or replaced with other configurations.
[0154] Furthermore, the aforementioned configurations, functions, processing units, processing means, etc. may be realized in part or in whole in hardware, for example by designing them as integrated circuits, or may be realized in software by a processor interpreting and executing a program that realizes each function.
[0155] Information such as programs, tables, and files that realize each function can be stored in storage devices such as memory, hard disks, and SSDs (Solid State Drives), or on recording media such as IC (Integrated Circuit) cards, SD cards, and DVDs (Digital Versatile Discs).
[0156] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines that are necessary for implementation. In reality, it can be considered that almost all components are interconnected. [Explanation of symbols]
[0157] 300 Analyzer 301 processor 302 Storage Devices 400 Generation part 401 Acquisition Department 402 Stratification Department 403 Output section 411 Search Department 412 Repeat Section 420 Patient Data Table 430 Weight Table
Claims
1. An analytical device having a processor that executes a program and a storage device that stores the program, the storage device stores a weight for each predictor group reflecting sensitivity to treatment among a group of factors constituting patient characteristics; The processor: an acquisition process for acquiring a plurality of patient data sets for each patient in the patient group, the patient data sets including values for each factor in the group of factors, and values of treatment selections and outcomes corresponding to the factors; a stratification process for stratifying the patient group using the plurality of patient data acquired by the acquisition process; an output process for outputting the stratification result obtained by the stratification process; In the stratification process, the processor repeatedly executes a search process for searching for a branching condition for branching the patient group based on a specific predictive factor in the group of factors and its weight, and a division process for dividing the patient group using the branching condition found by the search process; In the search process, the processor a division in which at least some patient data of the plurality of patient data is set as a search target group, predictive factors and their weights are selected from the factor group of the search target group, and the search target group is divided into a first group and a second group based on the predictive factors; a treatment effect calculation that calculates a first treatment effect for the first group using the treatment selection value and the outcome value for the first group, and calculates a second treatment effect for the second group using the treatment selection value and the outcome value for the second group; a difference calculation for calculating a first difference of a first loss function based on the first therapeutic effect before and after the division and a second difference of a second loss function based on the second therapeutic effect before and after the division; updating the first difference before the difference calculation to the first difference after the difference calculation if the first difference after the difference calculation is greater than the first difference before the difference calculation, and updating the second difference before the difference calculation to the second difference after the difference calculation if the second difference after the difference calculation is greater than the second difference before the difference calculation; and generating a branching condition by repeatedly executing this updating process while setting the first group corresponding to the first difference after the update and the second group corresponding to the second difference after the update as the search target groups, until no predictors remain; In the output process, the processor outputs, as the stratification result, a causal tree that combines the branching conditions from the start to the end of the repeated execution of the division process. An analytical device characterized by:
2. The analytical device according to claim 1 , The processor: a generation process of searching a medical literature database using search keywords including conjunctions related to the factors and outcomes, extracting sentences corresponding to the search keywords, calculating weights of the factors included in the search keywords, and storing the factors included in the search keywords in association with the weights in the storage device; An analytical device characterized by performing the above.
3. An analytical method performed by an analytical device having a processor that executes a program and a storage device that stores the program, the storage device stores a weight for each predictor group reflecting sensitivity to treatment among a group of factors constituting patient characteristics; The processor: an acquisition process for acquiring a plurality of patient data sets for each patient in the patient group, the patient data sets including values for each factor in the group of factors, and values of treatment selections and outcomes corresponding to the factors; a stratification process for stratifying the patient group using the plurality of patient data acquired by the acquisition process; an output process for outputting the stratification result obtained by the stratification process; In the stratification process, the processor repeatedly executes a search process for searching for a branching condition for branching the patient group based on a specific predictive factor in the group of factors and its weight, and a division process for dividing the patient group using the branching condition found by the search process; In the search process, the processor a division in which at least some patient data of the plurality of patient data is set as a search target group, predictive factors and their weights are selected from the factor group of the search target group, and the search target group is divided into a first group and a second group based on the predictive factors; a treatment effect calculation that calculates a first treatment effect for the first group using the treatment selection value and the outcome value for the first group, and calculates a second treatment effect for the second group using the treatment selection value and the outcome value for the second group; a difference calculation for calculating a first difference of a first loss function based on the first therapeutic effect before and after the division and a second difference of a second loss function based on the second therapeutic effect before and after the division; updating the first difference before the difference calculation to the first difference after the difference calculation if the first difference after the difference calculation is greater than the first difference before the difference calculation, and updating the second difference before the difference calculation to the second difference after the difference calculation if the second difference after the difference calculation is greater than the second difference before the difference calculation; and generating a branching condition by repeatedly executing this updating process while setting the first group corresponding to the first difference after the update and the second group corresponding to the second difference after the update as the search target groups, until no predictors remain; In the output process, the processor outputs, as the stratification result, a causal tree that combines the branching conditions from the start to the end of the repeated execution of the division process. An analytical method characterized by:
4. A processor that can access a storage device that stores weights for each predictor group that reflects sensitivity to treatment among factors that constitute patient characteristics, an acquisition process for acquiring a plurality of patient data sets for each patient in the patient group, the patient data sets including values for each factor in the group of factors, and values of treatment selections and outcomes corresponding to the factors; a stratification process for stratifying the patient group using the plurality of patient data acquired by the acquisition process; an output process for outputting a stratification result obtained by the stratification process; In the stratification process, the processor is caused to repeatedly execute a search process of searching for a branching condition for branching the patient group based on a specific predictive factor in the group of factors and its weight, and a division process of dividing the patient group using the branching condition found by the search process; In the search process, the processor a division in which at least some patient data of the plurality of patient data is set as a search target group, predictive factors and their weights are selected from the factor group of the search target group, and the search target group is divided into a first group and a second group based on the predictive factors; a treatment effect calculation that calculates a first treatment effect for the first group using the treatment selection value and the outcome value for the first group, and calculates a second treatment effect for the second group using the treatment selection value and the outcome value for the second group; a difference calculation for calculating a first difference of a first loss function based on the first therapeutic effect before and after the division and a second difference of a second loss function based on the second therapeutic effect before and after the division; updating the first difference before the difference calculation to the first difference after the difference calculation if the first difference after the difference calculation is greater than the first difference before the difference calculation, and updating the second difference before the difference calculation to the second difference after the difference calculation if the second difference after the difference calculation is greater than the second difference before the difference calculation; and generating a branching condition by repeatedly executing this updating until no predictors remain, while setting the first group corresponding to the first difference after the update and the second group corresponding to the second difference after the update as the search target groups, In the output process, the processor is caused to output, as the stratification result, a causal tree that combines the branching conditions from the start to the end of the repeated execution of the division process. An analysis program characterized by:
Citation Information
Patent Citations
Computerized medical planning method and system using mass medical analysis
JP2017502439A
Computational medical treatment plan method and system using mass medical analysis
JP2020149711A
Estimation of individual causal effects
US8688610B1