Information Processing Apparatus, Information Processing Method, and Program

By selecting parameters based on acquisition rate and frequency, the information processing apparatus generates high-quality training data, addressing the issue of insufficient and varying parameter frequencies in clinical data to enhance the accuracy of prognosis prediction models.

JP7717152B2Active Publication Date: 2025-08-01TERUMO KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023508986
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-23
Filing Date
2022-03-10
Publication Date
2025-08-01
Estimated Expiration
2042-03-10

AI Technical Summary

Technical Problem

The accuracy of prognosis prediction models using machine learning is compromised by varying parameter frequencies and combinations in clinical data, leading to insufficient training data and unsuitable parameter sets, which can result in inaccurate predictions.

Method used

An information processing apparatus and method that selects high-quality training data by calculating the acquisition rate and frequency of parameters from time-series clinical data, using thresholds to identify relevant parameters for training, and generating data formats that minimize data loss.

Benefits of technology

This approach ensures the generation of a required number of high-quality training data, reducing data loss and improving the accuracy of machine learning models for patient prognosis prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717152000001
    Figure 0007717152000001
  • Figure 0007717152000002
    Figure 0007717152000002
  • Figure 0007717152000003
    Figure 0007717152000003
Patent Text Reader

Abstract

This information processing device is used in a system using machine learning to predict a prognosis for a patient, and comprises: an input unit that receives the input of a plurality of sets of time series data corresponding to a plurality of patients and including a plurality of first parameters that relate to at least one of the status and the treatment of each patient; and a processing unit that calculates an acquisition rate and an acquisition frequency for each of the first parameters included in the plurality of sets of time series data, and uses at least one of the calculated acquisition rate and the calculated acquisition frequency to select, from the plurality of first parameters, a second parameter that is to be used in training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] In recent years, it has been proposed to utilize a system using machine learning to predict the prognosis of patients (see, for example, Patent Document 1). A prognosis prediction model for patients using machine learning is trained using data combining various parameters. For example, in the case of predicting the death of patients admitted to an intensive care unit (ICU), basic information such as age and gender, disease information, vital signs, drug administration information, etc. are used as training data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When training a prognosis prediction model using clinical data of a patient that has not been acquired for machine learning, the frequency and combination of parameters actually acquired vary depending on the patient's condition and disease. Therefore, even if a specific parameter is selected as a parameter for machine learning, there may be many parameters that have not been acquired. If the quality of the data is low due to many missing parameters in the parameters used for training data, there is a concern that the accuracy of prediction by machine learning will decrease. Therefore, if only the data of patients with a high sufficiency of parameters used for training data is to be used, the number of data that can be used as training data will decrease, and sufficient training may not be possible. Furthermore, if a parameter set suitable for machine learning cannot be obtained, the trained prediction model may not be applicable.

[0005] Therefore, the object of the present disclosure made by focusing on these points is to provide an information processing apparatus, an information processing method, and a program that can generate a required number of high-quality training data for machine learning from time-series data during the clinical practice of patients.

Means for Solving the Problems

[0006] An information processing apparatus according to an aspect of the present disclosure is an information processing apparatus used in a system for predicting the prognosis of a patient by machine learning, and includes an input unit that receives input of a plurality of time-series data corresponding to a plurality of patients, where the time-series data includes a plurality of first parameters related to at least one of the state and treatment of each patient, and a processing unit that calculates an acquisition rate and an acquisition frequency of each of the first parameters included in the plurality of time-series data, and selects a second parameter to be used for training data from the plurality of first parameters using at least one of the calculated acquisition rate and the acquisition frequency.

[0007] In one embodiment, the acquisition rate indicates a ratio of the plurality of time-series data that includes the first parameter.

[0008] In one embodiment, the acquisition frequency indicates a frequency at which the first parameter is included in data within a predetermined period of the time-series data.

[0009] In one embodiment, when at least one of the acquisition rate and the acquisition frequency of the first parameter exceeds a predetermined threshold, the processing unit selects the first parameter as the second parameter to be used for the training data.

[0010] In one embodiment, the threshold of the acquisition frequency varies among the plurality of first parameters.

[0011] In one embodiment, the threshold of the acquisition frequency is determined based on the number of the time-series data that includes the first parameter exceeding the threshold.

[0012] As one embodiment, the plurality of time series data includes a first time series data group and a second time series data group, and the processing unit executes a process of selecting the second parameter by combining the first time series data group and the second time series data group, and a process of separately selecting the second parameter from the first time series data group and the second time series data group.

[0013] As one embodiment, the input unit further receives input of additional information including at least one of initial symptoms, personal attributes, and diseases for each of the plurality of patients, and the processing unit groups the time series data into a plurality of groups based on the additional information, and executes a process of selecting the second parameter for each of the plurality of groups.

[0014] As one embodiment, the processing unit lengthens the predetermined period for calculating the acquisition frequency as time elapses.

[0015] As one embodiment, the processing unit generates training data using the selected second parameter.

[0016] As one embodiment, the processing unit generates the training data in a data format based on the acquisition frequency of the selected second parameter.

[0017] As one embodiment, the processing unit generates a learned model for predicting the prognosis of a patient using the training data.

[0018] As one embodiment, for each of a plurality of temporary thresholds, when at least one of the acquisition rate and the acquisition frequency of the first parameter exceeds the temporary threshold, the processing unit selects the first parameter as a temporary parameter to be used for the training data, generates the training data and test data using the temporary parameter, generates a learned model for predicting the prognosis of a patient using the training data, performs a process of determining the accuracy of the learned model using the test data, and selects the temporary parameter with the highest determined accuracy as the second parameter.

[0019] As one embodiment, the time series data includes at least one of information on drug administration information, vital signs, examination information, findings information, water intake information, water loss information, and treatment information.

[0020] As one embodiment, the drug administration information includes at least one of information on the type of administered drug, administration route, dose, and administration rate.

[0021] As one embodiment, the vital signs include at least one of information on body temperature, blood pressure, heart rate, respiratory rate, pulse rate, oxygen saturation, body weight value, central venous pressure, and inhaled oxygen concentration.

[0022] As one embodiment, the examination information includes at least one of information on blood test data, blood gas data, urine test, electrocardiogram, and imaging diagnosis results.

[0023] As one embodiment, the findings information includes at least one of information on congestion, cyanosis, and level of consciousness.

[0024] As one embodiment, the water intake information includes at least one of information on the amount of water drunk and the amount of infusion.

[0025] As one embodiment, the water loss information includes at least one of information on urine volume and blood loss volume.

[0026] As one embodiment, the treatment information includes at least any one of information on introduction of a dialysis device, removal of a dialysis device, and setting of a dialysis device, and information on introduction of a ventilator, removal of a ventilator, and setting of a ventilator.

[0027] An information processing method as one aspect of the present disclosure is an information processing method executed by an information processing device used in a system for predicting a patient's prognosis by machine learning, the method including: obtaining a plurality of time-series data corresponding to a plurality of patients, the time-series data including a plurality of first parameters related to at least one of the state and treatment of each patient; calculating an acquisition rate and an acquisition frequency of each of the first parameters included in the plurality of time-series data; and selecting, using at least one of the calculated acquisition rate and acquisition frequency, second parameters to be used for training data from the plurality of first parameters.

[0028] A program as one aspect of the present disclosure is a program for causing an information processing device used in a system for predicting a patient's prognosis by machine learning to execute information processing, the information processing including: obtaining a plurality of time-series data corresponding to a plurality of patients, the time-series data including a plurality of first parameters related to at least one of the state and treatment of each patient; calculating an acquisition rate and an acquisition frequency of each of the first parameters included in the plurality of time-series data; and selecting, using at least one of the calculated acquisition rate and acquisition frequency, second parameters to be used for training data from the plurality of first parameters.

Advantages of the Invention

[0029] According to the present disclosure, since the second parameters to be used for training data are selected using the acquisition rate and acquisition frequency of the first parameters included in the time-series data, it is possible to generate a required number of high-quality training data for performing machine learning from the time-series data at the time of a patient's clinical situation.

Brief Description of the Drawings

[0030]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0031] (Configuration of the Information Processing Apparatus) Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. An information processing apparatus 10 according to an embodiment is used for selecting parameters used in training data in a system for predicting a patient's prognosis by machine learning. The information processing apparatus 10 can use a computer such as a PC (Personal Computer) and a workstation. The information processing apparatus 10 may be arranged in a medical institution such as a hospital, or an information processing facility that aggregates information from a plurality of medical institutions. As shown in FIG. 1, the information processing apparatus 10 includes an input unit 11, a processing unit 12, an output unit 13, and a storage unit 14.

[0032] The input unit 11 is a part where the information processing apparatus 10 receives the input of time-series data. When the information processing apparatus 10 receives time-series data from another apparatus via a communication line, the input unit 11 includes a communication interface with the other apparatus. When the information processing apparatus 10 acquires time-series data stored in a storage medium such as a magnetic storage medium, a magneto-optical storage medium, or an optical storage medium, the input unit 11 may include a reading device for the storage medium.

[0033] The time-series data includes time-series information regarding at least one of the states and treatments of each patient acquired during clinical practice for a plurality of parameters. Let the parameters included in the time-series data be the first parameters. One time-series data may include information from before treatment to the end of treatment for one patient. The input unit 11 receives the input of a plurality of time-series data for use in machine learning. The plurality of time-series data related to a plurality of patients may be referred to as a time-series data group hereinafter. The time-series data group may include time-series data on the order of, for example, several hundreds, several thousands, or several tens of thousands.

[0034] The time-series data may include at least one of information on drug administration information, vital signs, examination information, findings information, water intake information (water IN information), water loss information (water OUT information), and treatment information for one patient.

[0035] In one embodiment, the drug administration information may include at least one of information on the type of administered drug, administration route, dose, and administration rate.

[0036] In one embodiment, the vital signs may include at least one of information on body temperature, blood pressure, heart rate, respiratory rate, pulse rate, oxygen saturation, body weight value, central venous pressure, and inhaled oxygen concentration.

[0037] In one embodiment, the examination information may include at least one of information on blood test data, blood gas data, urine test, electrocardiogram, and image diagnosis results.

[0038] In one embodiment, the finding information may include information on at least one of congestion, cyanosis, and level of consciousness.

[0039] In one embodiment, the water intake information may include information on at least one of the amount of water drunk and the amount of infusion fluid.

[0040] In one embodiment, the water loss information may include information on at least one of the urine volume and the amount of bleeding.

[0041] In one embodiment, the treatment information may include information on at least one of the introduction of a dialysis device, the removal of a dialysis device, and the settings of a dialysis device, as well as the introduction of a ventilator, the removal of a ventilator, and the settings of a ventilator.

[0042] Each parameter of the time-series data can be adopted as an explanatory variable for machine learning. A part of the time-series data can be a target variable for machine learning. For example, the introduction of a dialysis device and the introduction of a ventilator included in the treatment information can be an outcome (outcome) to be predicted, and thus can be a target variable.

[0043] The processing unit 12 executes various arithmetic processes. The processing unit 12 is configured to include one or more processors and a memory. The "processor" includes, but is not limited to, a general-purpose processor and a dedicated processor specialized for specific processing. A general-purpose processor can read a program stored in the memory and execute processing according to the program. The processing executed by the processing unit 12 will be described below.

[0044] The output unit 13 outputs the result of the processing by the processing unit 12 to the outside of the information processing apparatus 10. The output unit 13 may include a communication interface for another system. The output unit 13 may include a writing device for storing information in a storage medium such as a magnetic storage medium, a magneto-optical storage medium, or an optical storage medium. The information processing apparatus 10 may include a storage device inside.

[0045] The storage unit 14 can store information necessary for the processing performed by the processing unit 12, information generated by the processing unit 12, and programs executed by the processing unit 12. The storage unit 14 may be configured using any one or more of, for example, a semiconductor memory, a magnetic memory, and an optical memory. The semiconductor memory may include a volatile memory and a non-volatile memory. The magnetic memory may include, for example, a hard disk and a magnetic tape.

[0046] (Configuration of the processing unit) As shown in FIG. 2, the processing unit 12 includes a time-series data acquisition unit 21, a parameter acquisition rate calculation unit 22, a parameter acquisition frequency calculation unit 23, and a parameter selection unit 24. The processing unit 12 may further include a training data generation unit 25, a model generation unit 26, and a model evaluation unit 27. Each unit of the processing unit 12 may be a hardware module or a software module. The functions of each component of the processing unit 12 are executed by the processing unit 12.

[0047] The time-series data acquisition unit 21 is configured to acquire time-series data via the input unit 11. The time-series data includes information such as examinations and treatments of patients during clinical practice collected at medical institutions such as hospitals. A simplified example of the time-series data of one patient is shown in FIG. 3. "Date / Time" is information indicating the date and time when the data of each parameter included in the time-series data was acquired. "Parameter Name" includes the name of each parameter or information for identifying each parameter. "Value" is information indicating the value of the parameter specified by the parameter name. As an example, the "Date / Time", "Parameter Name", and "Value" include information such as "March 1, 2021, 10:10", "Blood Pressure", and "130", respectively. The format of the time-series data as shown in FIG. 3 is only an example. The time-series data can also be data grouped in time series for each parameter.

[0048] The parameter acquisition rate calculation unit 22 is configured to calculate the acquisition rate of each parameter included in a plurality of time series data. The acquisition rate indicates the ratio of each parameter included in the plurality of time series data. For example, if the number of time series data is 1000 and 900 of them contain a specific parameter, the acquisition rate of the parameter is 90%. The acquisition rate can be the ratio of the target parameter included in the entire time series data. The acquisition rate may be the ratio of the target parameter included in the time series data within a predetermined period.

[0049] The parameter acquisition frequency calculation unit 23 is configured to calculate the acquisition frequency of each parameter included in a plurality of time series data. The acquisition frequency indicates the frequency of the target parameter included in the data within a predetermined period of the time series data containing the target parameter. For example, if the predetermined period is one day and the number of times the target parameter is included in the time series data during that period is 4, the acquisition frequency is 4. The parameter acquisition frequency calculation unit 23 may calculate the acquisition frequency of each parameter for each time series data. The parameter acquisition frequency calculation unit 23 may calculate the average acquisition frequency, which is the average value of the acquisition frequencies of each parameter for all time series data. Further, the parameter acquisition frequency calculation unit 23 may calculate the acquisition frequency standard deviation, which is the standard deviation of the acquisition frequencies.

[0050] The parameter selection unit 24 is configured to select a plurality of parameters to be used for training data using at least one of the acquisition rate and the acquisition frequency calculated by the parameter acquisition rate calculation unit 22 and the parameter acquisition frequency calculation unit 23. The plurality of parameters to be used for training data are the second parameters. The parameter selection unit 24 selects the parameter as a parameter to be used for training data when at least one of the acquisition rate and the average acquisition frequency of the parameter exceeds a predetermined threshold. The parameters selected by the parameter selection unit 24 correspond to the explanatory variables in machine learning.

[0051] For example, assume a case where the acquisition rate and average acquisition frequency for each parameter are as shown in FIG. 4. The parameter selection unit 24, for example, sets the threshold for the acquisition rate at 80% and the threshold for the average acquisition frequency at 2, and assumes that it selects parameters that exceed both thresholds. In this case, since parameters 1, 3, and 4 in FIG. 4 exceed these thresholds, they are selected as the second parameters. Parameters 2 and 5 are not adopted as the second parameters.

[0052] The acquisition frequency standard deviation may be considered in parameter selection. Even when the average acquisition frequencies are the same, if the standard deviation is large, the acquisition frequency of the parameter has a greater variation compared to the case where the standard deviation is small. Therefore, even when the same threshold is set, the larger the standard deviation, the fewer the number of data exceeding the threshold may be. Thus, a parameter with a smaller acquisition frequency standard deviation may be preferentially selected over a parameter with a larger acquisition frequency standard deviation.

[0053] According to the characteristics of each parameter, the threshold for the average acquisition frequency may vary for each of the plurality of parameters. For example, it can be set according to the frequency that can be actually performed in clinical practice, such as once a day for collecting blood test data and three times a day for measuring blood pressure.

[0054] Also, for the training data for machine learning, only the time-series data in which the acquisition frequency of the selected parameter exceeds the acquisition frequency threshold is used. The acquisition frequency threshold may be determined in consideration of the number of time-series data whose acquisition frequency exceeds the threshold. If the threshold is set low, there may be many gaps in the parameter data and it may not be possible to obtain high-quality training data. However, if the threshold is set too high, since a sufficient number of time-series data are not included, it may not be possible to secure the number of training data required for machine learning. For example, the threshold may be set so that the number of time-series data exceeding the threshold is a predetermined number (for example, 10,000) or more. Also, for example, the threshold may be set so that a predetermined ratio (for example, 90%) or more of all the time-series data exceeds the threshold.

[0055] The parameter selection unit 24 may select some or all of the parameters from among a plurality of parameters in which at least one of the acquisition rate of the parameters and the average acquisition frequency exceeds a predetermined threshold, as parameters to be used for the training data. The parameters to be selected may be determined according to the prognosis (outcome) to be predicted. A plurality of parameters that may be selected according to the prognosis to be predicted may be stored in the storage unit 14.

[0056] In one embodiment, the processing unit 12 may output the plurality of parameters selected by the parameter selection unit 24 to another device via the output unit 13. The processing of the training data generation unit 25, the model generation unit 26, and the model evaluation unit 27 described below may be executed by another device.

[0057] The training data generation unit 25 generates a plurality of pieces of training data for machine learning using the parameters selected by the parameter selection unit 24 in the time series data. For example, as shown in FIG. 4, when parameters 2 and 5 among parameters 1 to 5 are not adopted, parameters 2 and 5 are not included in the training data. In this case, as exemplified in a simplified manner in FIG. 5, the training data includes data indicating the time series values of parameters 1, 3, and 4. One piece of training data may include data from before the start of treatment to the end of treatment of one patient for the selected parameters. The training data in FIG. 5 is an example. The training data can have various forms.

[0058] The training data generation unit 25 may generate training data in a data format based on the acquisition frequency of each selected parameter. For example, for parameters acquired every hour, the training data may be in a format where a total of 24 values per day are stored every hour. On the other hand, in the case of a parameter with an average acquisition frequency of three times per day, the training data generation unit 25 can set the training data in a format where three data are stored per day. By doing so, data loss in the training data can be reduced, so that machine learning can be performed with an algorithm that does not require correction processing or requires little correction processing.

[0059] The training data generation unit 25 may generate test data for verifying the accuracy of the learned model generated by machine learning in addition to the training data. For example, the training data generation unit 25 can use a predetermined ratio of the data generated from the time-series data of the parameters selected by the parameter selection unit 24 as training data and the rest as test data. The predetermined ratio can be, for example, 80% or the like.

[0060] The training data generation unit 25 may deliver the generated training data and test data to the model generation unit 26 and the model evaluation unit 27, respectively, in order to generate a learned model.

[0061] In one embodiment, the processing unit 12 may output the training data and test data generated by the training data generation unit 25 to another device via the output unit 13 in order to generate a learned model in another device. The processing of the model generation unit 26 and the model evaluation unit 27 described below may be executed in another device.

[0062] The model generation unit 26 generates a learned model for predicting the prognosis of a patient using the training data generated by the training data generation unit 25. The prognosis of a patient can be paraphrased as an outcome. The outcome includes information such as the life or death of the patient after treatment, whether hemodialysis has been introduced, whether a ventilator has been introduced, the length of stay in the ICU if the patient has entered the ICU, the severity score of the patient after treatment, the presence or absence of complications, and the blood pressure and heart rate of the patient after treatment. The outcome corresponds to the target variable in machine learning.

[0063] The model evaluation unit 27 is configured to evaluate the prediction accuracy of the learned model generated by the model generation unit 26 using the test data generated by the training data generation unit 25. For this purpose, first, the model evaluation unit 27 predicts the outcome using the learned model and the test data. Next, the model evaluation unit 27 calculates the prediction accuracy from the degree of agreement between the predicted outcome and the actual outcome.

[0064] In one embodiment, the prediction accuracy by the model evaluation unit 27 may be fed back to the setting of the threshold value in the parameter selection unit 24. The processing unit 12 may determine, as the threshold value in the parameter selection unit 24, the threshold value at which the best prediction accuracy can be obtained by machine learning.

[0065] For example, a plurality of provisional threshold values are prepared in advance in the processing unit 12. For each of the plurality of provisional threshold values, when at least one of the acquisition rate and acquisition frequency of each parameter exceeds the provisional threshold value, the parameter selection unit 24 selects the parameter as a provisional parameter to be used for the training data. The training data generation unit 25 generates training data and test data using the selected provisional parameters. The model generation unit 26 generates a learned model for predicting the prognosis of a patient using the training data. The model evaluation unit 27 performs a process of determining the accuracy of the learned model using the test data.

[0066] After performing the above processing for all provisional thresholds, the processing unit 12 selects the provisional parameter with the highest determined accuracy as the parameter for performing machine learning. Further, the processing unit 12 adopts the learned model corresponding to the selected parameter as the learned model for predicting the prognosis of the patient.

[0067] (Processing of Multiple Time-Series Data Groups) The information processing apparatus 10 may acquire a time-series data group including a plurality of time-series data from a plurality of medical institutions such as hospitals. The processing unit 12 may execute both the process of selecting the acquisition rate and parameters by combining the plurality of time-series data groups into one time-series data group and the process of selecting parameters for each time-series data group as an individual time-series data group. For example, assume that the processing unit 12 acquires a first time-series data group from a first medical institution and a second time-series data group from a second medical institution as time-series data. The processing unit 12 may execute the process of selecting parameters by combining the first time-series data group and the second time-series data group, and the process of selecting parameters individually from the first time-series data group and the second time-series data group.

[0068] The processing unit 12 may generate a learned model of the overall common part using the parameters selected for the entire time-series data. The processing unit 12 may generate a learned model of an individual medical institution using the parameters selected for the individual time-series data. The processing unit 12 can improve the prediction accuracy of the outcome in each medical institution by combining the learned model of the common part and the learned model of the individual medical institution. Combining the learned models includes, for example, taking a majority vote of a plurality of learned models and taking a weighted average.

[0069] (Grouping of Time-Series Data) In one embodiment, the information processing apparatus 10 may be further configured to receive, by the input unit 11, input of additional information including at least any one of a disease, initial symptoms, and personal attributes for each patient included in a plurality of patients. The disease may include, for example, names of diseases such as cerebral infarction and heart failure. The initial symptoms may include, for example, data of vital signs when the patient visits the emergency department and when the patient is admitted to the ICU. The initial symptoms may include, for example, information on the severity score. The severity score may include, for example, indices for severity assessment such as SAPS (2nd simplified acute physiology score) II and APACHE (Acute Physiology and Chronic Health Evaluation) II. The personal attributes may include, for example, gender, age, race, presence or absence of transportation by ambulance, and admission route.

[0070] The processing unit 12 may execute a process of grouping time-series data into a plurality of groups based on the additional information and selecting parameters for each of the plurality of groups. If the content of the disease, the initial symptoms, the severity, etc. are different, the parameters to be acquired and the content of the treatment are different. For example, for a patient with heart failure, a treatment model for heart failure is applied. Also, the treatment strategy for a patient with heart failure varies depending on the blood pressure value of the initial symptoms. Therefore, by grouping the time-series data according to the additional information, the processing unit 12 can collect the time-series data of patients having a common disease and similar symptoms, etc.

[0071] The processing unit 12 can execute processes such as calculating the acquisition rate and acquisition frequency of parameters, selecting parameters, and generating training data for the grouped time-series data. By grouping the time-series data of patients having common diseases and similar symptoms, etc., the ratio of information such as administration information of drugs and vital values specific to the diseases and symptoms can be collected more highly. Thereby, it is possible to select a combination of parameters highly correlated with the outcome (prognosis), and it can be expected that the data loss of the parameters used for the training data will be reduced. Further, by limiting the time-series data to be grouped to time-series data corresponding to a specific disease or symptom, it can be expected that the prediction accuracy by the learned model generated using the training data will be increased.

[0072] (Period for calculating the acquisition frequency) In one embodiment, the processing unit 12 can make the period for calculating the acquisition rate and acquisition frequency of parameters longer as time elapses. For example, since the vital values of patients entering and leaving the ICU may be frequently measured because the values are not stable immediately after admission, as time elapses and the values become stable, the measurement interval of the vital values becomes longer. Therefore, the period for calculating the acquisition rate and acquisition frequency of parameters can be made longer as time elapses, for example, immediately after entering the ICU, 1 hour, 3 hours, 1 day, and 3 days after admission.

[0073] For some time after entering the ICU, since the types of parameters clinically acquired are many and the acquisition frequency is also high, it can be expected that a highly accurate learned model can be generated even in a short period. On the other hand, when the period after entering the ICU becomes longer, since the types and numbers of parameters clinically acquired decrease, a combination of parameters different from immediately after entering the ICU may be selected.

[0074] (Example 1 of information processing method) Referring to FIG. 6, an example of an information processing method executed by the information processing apparatus 10 according to an embodiment will be described. FIG. 6 shows the flow of information processing executed by the processing unit 12 of the information processing apparatus 10. This processing can be executed by a processor included in the information processing apparatus 10 according to a program. Such a program can be stored in a non-transitory computer-readable medium. Non-transitory computer-readable media include, for example, but are not limited to, magnetic storage media, magneto-optical storage media, and semiconductor memories.

[0075] The flowchart of FIG. 6 assumes that the processing unit 12 of the information processing apparatus 10 has a time series data acquisition unit 21, a parameter acquisition rate calculation unit 22, a parameter acquisition frequency calculation unit 23, a parameter selection unit 24, and a training data generation unit 25. The processing unit 12 does not necessarily have a model generation unit 26 and a model evaluation unit 27.

[0076] First, the processing unit 12 acquires clinical time series data regarding a plurality of patients via the input unit 11 (step S101).

[0077] For each parameter (first parameter) included in the plurality of time series data acquired in step S101, the processing unit 12 calculates the acquisition rate for each parameter (step S102).

[0078] For each parameter (first parameter) included in the plurality of time series data acquired in step S101, the processing unit 12 calculates the acquisition frequency for each parameter (step S103).

[0079] The processing of step S102 and step S103 may be executed substantially simultaneously in parallel. Also, step S103 may be executed before step S102.

[0080] Based on the acquisition rate and acquisition frequency for each parameter, the processing unit 12 selects parameters (second parameters) to be adopted for training data from among the parameters included in the time series data (step S104).

[0081] The processing unit 12 generates training data and test data for machine learning using the data of the selected parameters in the plurality of time series data (step S105).

[0082] The processing unit 12 outputs the generated training data and test data to another device or a storage medium via the output unit 13 for use in machine learning (step S106).

[0083] The processing unit 12 may not execute step S105 and may only output the information of the parameters selected in step S104 to the outside in the next step S106. In that case, another device generates the training data and test data for machine learning.

[0084] (Example 2 of information processing method) With reference to FIG. 7, an example of an information processing method executed by the information processing apparatus 10 according to another embodiment will be described. The processes from step S201 to step S205 in the flowchart of FIG. 7 are the same as or similar to the processes from step S101 to S105 in FIG. 6, and thus the description of the common content will be omitted. The flowchart of FIG. 7 is premised on the processing unit 12 including the model generation unit 26 and the model evaluation unit 27.

[0085] In steps S201 to S205, the processing unit 12 executes a process of generating training data and test data for machine learning based on the time series data at the time of clinical acquisition from the input unit 11 in the same manner as in steps S101 to S105. However, in step S204, a plurality of thresholds are prepared, and one of the thresholds is selected for the parameter. The plurality of thresholds can be regarded as provisional thresholds.

[0086] After generating the training data and the test data in step S205, the processing unit 12 constructs a prognostic prediction model of a patient, which is a learned model by machine learning, using the generated training data (step S206).

[0087] The processing unit 12 estimates the prediction accuracy by the prognosis prediction model constructed in step S206 using the test data generated in step S205 (step S207). The processing unit 12 stores the prediction accuracy in the storage unit 14 in association with a provisional threshold value.

[0088] If the operations from step S204 to step S207 have not been completed for all of the plurality of provisional threshold values (step S208: No), the processing unit 12 changes the threshold value to a provisional threshold value for which the operation has not yet been performed (step S209).

[0089] After step S209, the processing unit 12 returns to step S204 and repeats the processing from step S204 to step S207.

[0090] If the operations from step S204 to step S207 have been completed for all of the plurality of provisional threshold values (step S208: Yes), the processing unit 12 adopts the prognosis prediction model with the highest prediction accuracy stored in the storage unit 14 (step S210) and ends the processing.

[0091] By doing so, the processing unit 12 can select threshold values for the acquisition rate and acquisition frequency that provide high prediction accuracy for parameter selection.

[0092] As described above, the information processing apparatus 10 calculates the acquisition rate and acquisition frequency of each parameter included in a plurality of time series data, and selects a parameter to be used for training data using at least one of the calculated acquisition rate and acquisition frequency. As a result, it is possible to easily generate a required number of high-quality training data for machine learning from the time series data at the time of a patient's clinical situation. Also, this makes it easy to generate training data for machine learning from the data at the time of clinical practice.

[0093] Also, in the above-described embodiment, a parameter in which at least one of the acquisition rate and the average acquisition frequency exceeds a threshold is selected, and training data is generated using time series data having data in which the selected parameter exceeds the threshold. As a result, training data with less data loss can be generated, and an improvement in the accuracy of machine learning can be expected. Further, in order to correct the data loss in the training data, it is not necessary to perform a process of estimating the missing portion of the data, or it can be reduced, so that the processing load can be reduced.

[0094] Although the above-described embodiments have been described as representative examples, it is obvious to those skilled in the art that many changes and substitutions are possible within the spirit and scope of the present disclosure. Therefore, the present disclosure should not be construed as being limited by the above-described embodiments, and various modifications and changes are possible without departing from the scope of the claims. For example, the functions included in each component or each step, etc., can be rearranged so as not to be logically contradictory, and a plurality of components or steps, etc., can be combined into one or divided.

Explanation of Reference Numerals

[0095] 10 Information processing apparatus 11 Input unit 12 Processing unit 13 Output unit 14 Storage unit 21 Time series data acquisition unit 22 Parameter acquisition rate calculation unit 23 Parameter acquisition frequency calculation unit 24 Parameter selection unit 25 Training data generation unit 26 Model generation unit 27 Model evaluation unit

Claims

1. An information processing apparatus used in a system for predicting a patient's prognosis by machine learning, an input unit that receives inputs of a plurality of time-series data corresponding to a plurality of patients, wherein the time-series data includes a plurality of first parameters related to at least one of the state and treatment of each patient, a processing unit that calculates an acquisition rate and an acquisition frequency of each of the first parameters included in the plurality of time-series data, and selects a second parameter to be used for training data from the plurality of first parameters using at least one of the calculated acquisition rate and acquisition frequency An information processing apparatus comprising.

2. The information processing apparatus according to claim 1, wherein the acquisition rate indicates a ratio of the first parameter included in the plurality of time-series data.

3. The information processing apparatus according to claim 1 or 2, wherein the acquisition frequency indicates a frequency of the first parameter included in data within a predetermined period of the time-series data.

4. The processing unit according to any one of claims 1 to 3, wherein when at least one of the acquisition rate and the acquisition frequency of the first parameter exceeds a predetermined threshold, the first parameter is selected as the second parameter to be used for the training data.

5. The information processing apparatus according to claim 4, wherein the threshold of the acquisition frequency is different among the plurality of first parameters.

6. The information processing apparatus according to claim 4 or 5, wherein the threshold of the acquisition frequency is determined based on the number of the time-series data including the first parameter exceeding the threshold.

7. The plurality of time-series data includes a first time-series data group and a second time-series data group, and the processing unit executes a process of selecting the second parameter by combining the first time-series data group and the second time-series data group, and a process of selecting the second parameter individually from the first time-series data group and the second time-series data group. The information processing apparatus according to any one of claims 1 to 6.

8. The input unit further receives input of at least one of initial symptoms, personal attributes, and additional information including a disease for each of the plurality of patients, and the processing unit groups the time-series data into a plurality of groups based on the additional information and executes a process of selecting the second parameter for each of the plurality of groups. The information processing apparatus according to any one of claims 1 to 7.

9. The information processing apparatus according to claim 3, wherein the processing unit makes the predetermined period for calculating the acquisition frequency longer as time elapses.

10. The information processing apparatus according to any one of claims 1 to 9, wherein the processing unit generates training data using the selected second parameter.

11. The information processing apparatus according to claim 10, wherein the processing unit generates the training data in a data format based on the acquisition frequency of the selected second parameter.

12. The information processing apparatus according to claim 10 or 11, wherein the processing unit generates a learned model for predicting the prognosis of a patient using the training data.

13. For each of a plurality of temporary thresholds, when at least one of the acquisition rate and the acquisition frequency of the first parameter exceeds the temporary threshold, the processing unit selects the first parameter as a temporary parameter to be used for the training data, generates the training data and test data using the temporary parameter, generates a learned model for predicting the prognosis of a patient using the training data, performs a process of determining the accuracy of the learned model using the test data, and selects the temporary parameter with the highest determined accuracy as the second parameter. The information processing apparatus according to claim 1.

14. The information processing apparatus according to any one of claims 1 to 13, wherein the time-series data includes at least one of information on drug administration, vital signs, examination information, findings information, water intake information, water loss information, and treatment information.

15. The information processing apparatus according to claim 14, wherein the information on drug administration includes at least one of information on the type of administered drug, administration route, dose, and administration rate.

16. The information processing apparatus according to claim 14, wherein the vital signs include at least one of information on body temperature, blood pressure, heart rate, respiratory rate, pulse rate, oxygen saturation, body weight value, central venous pressure, and inhaled oxygen concentration.

17. The information processing apparatus according to claim 14, wherein the inspection information includes at least any one of blood test data, blood gas data, urine test, electrocardiogram, and image diagnosis results.

18. The information processing apparatus according to claim 14, wherein the finding information includes at least any one of congestion, cyanosis, and consciousness level information.

19. The information processing apparatus according to claim 14, wherein the water intake information includes at least any one of the amount of water intake and the amount of infusion.

20. The information processing apparatus according to claim 14, wherein the water loss information includes at least any one of urine volume and blood loss volume.

21. The information processing apparatus according to claim 14, wherein the treatment information includes at least any one of the introduction of a dialysis device, the removal of a dialysis device, and the setting of a dialysis device, and the introduction of a ventilator, the removal of a ventilator, and the setting of a ventilator.

22. An information processing method executed by an information processing apparatus used in a system for predicting a patient's prognosis by machine learning, A step of acquiring a plurality of time series data corresponding to a plurality of patients, wherein the time series data includes a plurality of first parameters related to at least any one of the state and treatment of each patient, A step of calculating an acquisition rate and an acquisition frequency of each of the first parameters included in the plurality of time series data, A step of selecting a second parameter to be used for training data from the plurality of first parameters by using at least one of the calculated acquisition rate and the acquisition frequency An information processing method including.

23. A program for causing the information processing apparatus to execute information processing executed by the information processing apparatus used in a system for predicting a patient's prognosis by machine learning, The information processing includes A step of acquiring a plurality of time series data corresponding to a plurality of patients, wherein the time series data includes a plurality of first parameters related to at least any one of the state and treatment of each patient, A step of calculating an acquisition rate and an acquisition frequency of each of the first parameters included in the plurality of time series data, A step of selecting a second parameter to be used for training data from the plurality of first parameters by using at least one of the calculated acquisition rate and the acquisition frequency A program including.

Citation Information

Patent Citations

  • Diagnostic and prognostic methods for cardiovascular diseases and events

    JP2019507354A

  • Prognosis prediction system, prognosis prediction programming device, prognosis prediction device, prognosis prediction method and prognosis prediction program

    JP2020144471A

  • Data selection device, learning device, and program

    JP2021086558A