Clinical data classification method, clinical data classification device, and control program

The medical data classification method and device effectively cluster patients into phenotype groups using ambulance-to-treatment data, enhancing prediction of condition changes and enabling targeted medical interventions.

JP7798342B2Active Publication Date: 2026-01-14OSAKA UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022009695
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2026-01-14
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

Existing methods struggle to accurately classify patients with specific diseases or injuries into phenotype groups based on their medical data, making it difficult for medical professionals to predict future condition changes and provide appropriate interventions.

Method used

A medical data classification method and device that acquires and classifies medical data from patients from ambulance transport to treatment, using attribute, vital sign, and symptom information to cluster patients into phenotype groups, with a control program to facilitate this process.

Benefits of technology

Enables accurate classification of patients into phenotype groups, allowing for early identification of high-risk patients and informed medical interventions, aiding in treatment strategy development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798342000001
    Figure 0007798342000001
  • Figure 0007798342000002
    Figure 0007798342000002
  • Figure 0007798342000003
    Figure 0007798342000003
Patent Text Reader

Abstract

To classify a plurality of patients affected by a prescribed disease or having suffered an injury into a plurality of phenotype groups.SOLUTION: A medical care data classification method includes: a data acquisition step in which a computer acquires medical care data that pertains to each of a plurality of patients affected by a prescribed disease or having suffered an injury, the data having been created regarding the disease or injury of the patient until medical care for the disease or injury is completed in a medical institution into which the patient was emergency transported; and a clustering step in which the computer classifies the plurality of patients into a plurality of phenotype groups on the basis of the medical care data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for classifying clinical data, a device for classifying clinical data, and a control program. [Background technology]

[0002] For example, Patent Document 1 describes a prediction method for predicting the occurrence of fatal symptoms such as sepsis and cardiac arrest from vital signs. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2021-177429 Summary of the Invention [Problem to be solved by the invention]

[0004] While most patients transported to the hospital by ambulance recover smoothly after treatment, it is known that some patients' condition suddenly worsens after receiving treatment for disease or trauma. Thus, there may be various phenotypes in the prognosis of patients transported to the hospital by ambulance. One aspect of the present invention aims to appropriately classify multiple patients suffering from a predetermined disease or trauma into multiple phenotype groups. [Means for solving the problem]

[0005] In order to solve the above-mentioned problems, a medical data classification method according to one aspect of the present disclosure includes: a data acquisition step in which a computer acquires medical data relating to each of a plurality of patients who have suffered from a predetermined disease or an injury, the medical data being created regarding the patient's disease or the injury from the time the patient is rushed to a medical institution until the time the patient finishes receiving medical treatment for the disease or the injury at the medical institution; and a clustering step in which the computer classifies the plurality of patients into a plurality of phenotype groups based on the medical data, wherein the medical data includes at least attribute information, vital sign information, and symptom information of each of the plurality of patients, and the symptom information includes the location of the disease and the severity of the disease, or the location of the injury and the severity of symptoms at the location.

[0006] In order to solve the above-mentioned problems, a medical data classification device according to one aspect of the present disclosure includes a data acquisition unit that acquires medical data relating to each of a plurality of patients who have suffered from a specified disease or injury, the data being created from the time the patient is rushed to a medical institution for the disease or the injury until the time the patient finishes receiving medical treatment for the disease or the injury at the medical institution, and a clustering unit that classifies the plurality of patients into a plurality of phenotype groups based on the medical data, wherein the medical data includes at least attribute information, vital sign information, and symptom information of each of the plurality of patients, and the symptom information includes the location of the disease and the severity of the disease, or the location of the injury and the severity of symptoms at the location.

[0007] The medical data classification device according to each aspect of the present disclosure may be realized by a computer. In this case, the control program for medical data classification that causes the computer to operate as each part (software element) of the medical data classification device, thereby realizing the medical data classification device on the computer, and the computer-readable recording medium on which the program is recorded, also fall within the scope of the present invention. [Effects of the Invention]

[0008] According to one aspect of the present invention, a plurality of patients suffering from a given disease or trauma can be appropriately classified into a plurality of phenotype groups. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a functional block diagram showing an example of the configuration of a medical data classification system including a medical data classification device according to an aspect of the present invention. [Figure 2] 10 is a flowchart showing the flow of processing performed by the classification device. [Figure 3] FIG. 10 is a diagram illustrating an overview of classification processing. [Figure 4] 1 is a table showing the classification results of the analysis cohort by the classification device. [Figure 5] 1 is a table showing the classification results of the validation cohort by the classifier. [Figure 6] Heatmap showing survival rates and distribution of variables for each clinical phenotype. [Figure 7] FIG. 1 shows the results of hierarchical cluster analysis of the principal component scores of each phenotype. [Figure 8] FIG. 1 is a plot showing the results of a principal component analysis of the centroids of each phenotype. [Figure 9] FIG. 1 is a graph showing serum proteins differentially expressed between high mortality phenotypes and other phenotypes. [Figure 10] FIG. 1 shows the results of gene ontology (GO) enrichment analysis of proteins. [Figure 11] This is a diagram plotting each variable of serum proteomics analysis data and the standardized value of each variable. DETAILED DESCRIPTION OF THE INVENTION

[0010] [Embodiment 1] Hereinafter, one embodiment of the present invention will be described in detail.

[0011] (Technical idea of ​​the present disclosure) First, the technical concept of a method for classifying medical data according to one aspect of the present invention will be described below.

[0012] The treatment of a specific disease or injury is based on a combination of individual treatments for each subdivided injury site. One of the reasons that makes it difficult to develop new treatments for patients with a specific disease or injury is that the disease or injury is a highly heterogeneous and complex pathological condition that is related to multiple factors such as age, gender, type of injury, degree of injury, and biological response.

[0013] Among patients with a certain disease or injury, there are those whose condition rapidly worsens after consultation. In such cases, medical professionals who treat the patient when they are transported by ambulance or at their first visit to the hospital must determine how the patient's condition will change in the future based on the patient's medical data and their own experience, and provide necessary medical intervention. However, it can be difficult for medical professionals to accurately determine how the patient's condition will change in the future at the time of initial consultation.

[0014] Therefore, the inventors conducted intensive research to solve the above-mentioned problems, and as a result, discovered a method for classifying multiple patients into multiple phenotype groups based on medical data created from the time the patient is rushed to a medical institution until the time the patient finishes receiving treatment for the disease or injury at the medical institution.

[0015] Here, the medical data created from the time a patient is transported to a medical institution by ambulance until the time the patient has been treated for a disease or injury at the medical institution may be, for example, medical data when the patient is treated by a paramedic in the ambulance during the emergency transport. The medical data may also be medical data when the patient is first treated by a doctor after arriving at the medical institution. The medical data may include both medical data obtained during the emergency transport and medical data obtained after arriving at the medical institution. The medical data may also include multiple types of test results that may be performed during the medical treatment. The medical data may also include multiple test results, as well as the dates and times when the test results were obtained (chronological data). The contents of the medical data will be described in detail below.

[0016] (Configuration of medical data classification system 100) The configuration of the clinical data classification system 100 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the configuration of the clinical data classification system 100 including a clinical data classification device 1 that executes a clinical data classification method according to one embodiment of the present invention. Hereinafter, the clinical data classification device 1 will also be referred to as the classification device 1, and the clinical data classification system 100 will also be referred to as the classification system 100.

[0017] The classification system 100 includes a classification device 1, an external device 5 that transmits medical data to the classification device 1, and a presentation device 4 that acquires the classification results output from the classification device 1 and presents the classification results. Note that Fig. 1 shows an example in which the classification system 100 has been introduced into a medical institution H1.

[0018] The external device 5 may be, for example, a server device that stores and manages medical data, or a terminal device into which medical personnel input medical data.

[0019] 1 illustrates an example in which the classification device 1 acquires medical data from an external device 5 that is separate from the classification device 1, but is not limited to this. For example, the classification device 1 may be configured to be built into the external device 5.

[0020] The presentation device 4 may be a display and a speaker that can present information output from the classification device 1. In one example, the presentation device 4 may be a display provided in the classification device 1 or an external device 5. Alternatively, the presentation device 4 may be a computer and a tablet terminal used by a medical professional belonging to the medical institution H1.

[0021] The classification device 1 and the external device 5, and the classification device 1 and the presentation device 4 may be connected by wireless communication or by wired communication.

[0022] (Configuration of medical data classification device 1) The clinical data classification device 1 includes a control unit 2 and a storage unit 3. The storage unit 3 may store a classification model 31. The classification model 31 will be described later.

[0023] The storage unit 3 may store, in addition to the classification model 31, control programs for each unit executed by the control unit 2, OS programs, application programs, etc. The storage unit 3 may also store various data that is read when the control unit 2 executes these programs. The storage unit 3 is configured with a non-volatile storage device such as a hard disk or flash memory. In addition to the storage unit 3, the storage unit may also include a volatile storage device such as a RAM (Random Access Memory), which is used as a working area for temporarily storing data in the process of executing the various programs described above.

[0024] <Configuration of control unit 2> The control unit 2 may be configured by a control device such as a CPU (central processing unit) or a dedicated processor. Each part of the control unit 2 shown in Fig. 1 can be realized by a control device such as a CPU reading out a program stored in a storage unit 3 realized by a ROM (read only memory) or the like into a RAM (random access memory) or the like and executing the program.

[0025] The control unit 2 acquires target medical data, classifies patients, and outputs the classification results. The control unit 2 includes a data acquisition unit 21, a clustering unit 22, and an output unit 23.

[0026] [Data Acquisition Section] The data acquisition unit 21 acquires medical data from the external device 5. The data acquisition unit 21 may store the acquired medical data in the storage unit 3.

[0027] The medical data includes the following items for multiple patients: ·Attribute information Vital signs information Symptom information.

[0028] The attribute information may include, for example, the patient's age and gender. The attribute information may also include the Charlson Polypharmacy Scale (CPS). This allows the classification system 100 to classify the patient's medical data based on the patient's age, gender, and CPS.

[0029] The vital sign information may include, for example, at least one of respiratory rate, pulse rate, systolic blood pressure, body temperature, and level of consciousness (Glasgow Coma Scale (GCS) or Japan Coma Scale (JCS)). Thus, the classification system 100 can classify the medical data of patients based on various vital sign information.

[0030] For a patient suffering from a disease, the symptom information includes the location of the disease and the severity of the disease. Diseases may be distinguished from trauma, which will be described later. The specified disease may include at least one of myocardial infarction, cerebral infarction, heat stroke, cerebral hemorrhage, cardiac arrest, poisoning, disseminated intravascular coagulation (DIC), acute respiratory distress syndrome (ARDS), burns, hemorrhagic shock (including, for example, gastrointestinal bleeding, postpartum hemorrhage, etc.), and infectious diseases. Examples of infectious diseases include sepsis and COVID-19. The diseased organ may specifically be the myocardium, lungs, cerebrospinal cord, peripheral nerves, bones, liver, kidneys, spleen, gastrointestinal tract, bladder, uterus, ovaries, testes, skin, etc., and the severity of the disease may be the progression of the disease. The severity of the disease may also be indicated, for example, by one of multiple preset graded scales. This allows the classification system 100 to classify patient medical data based on the location of the disease and the severity of the disease.

[0031] For a patient who has sustained an injury, the symptom information may include the location of the injury and the severity of the injury at that location, which may include, for example, the head and neck, face, chest, abdomen, extremities, and body surface.

[0032] The symptom information may include at least one of information indicating the type of injury to the patient's body, the degree of injury, and biological reactions caused by the trauma. For example, the symptom information may be represented by an Abbreviated Injury Scale (AIS) or an Injury Severity Score (ISS). This allows the classification system 100 to classify the patient's medical data based on at least one of information indicating the type of injury to the patient's body, the degree of injury, and biological reactions caused by the trauma.

[0033] The type of injury is not particularly limited, and may be, for example, a type classified by a known score for assessing the severity of the injury. An example of a known score for assessing the severity of the injury is the Abbreviated Injury Scale (AIS) coding (https: / / www.jtcr-jatec.org / index_ais.html). Specific examples of AIS coding include AIS 2015, AIS 2005 Update 2008, and AIS90 Update 98. The latest AIS coding updates may be included at https: / / www.jtcr-jatec.org / index_ais.html. More specifically, the type of injury may be an abrasion, contusion, crush, contusion, incision, split, laceration, puncture, avulsion, burn, fracture, or the like. This allows the classification system 100 to classify patient medical data based on injuries classified in detail by AIS coding.

[0034] The data acquiring unit 21 may have a preprocessing function for performing preprocessing on the acquired medical data. For example, the data acquiring unit 21 may select and acquire only necessary items from among items included in the acquired medical data.

[0035] [Clustering Department] The clustering unit 22 inputs the medical data into a classification model described below, and outputs the classification results in which patients are classified into phenotype groups.

[0036] The number of phenotype groups classified by the clustering unit 22 is not particularly limited. Furthermore, the phenotype groups classified by the clustering unit 22 may include at least one phenotype group in which the proportion of patients with a poor prognosis within a predetermined period from the first medical examination is 45% or more. The first medical examination may be the first time a patient is examined by a medical professional, for example, when the patient is examined by a paramedic in an ambulance during emergency transport. Here, a patient who will have a poor prognosis within a predetermined period is, for example, a patient who will die within the predetermined period. Furthermore, the "predetermined period" in "from the first medical examination" is not particularly limited, and may be, for example, several hours, one day, or 30 days from the first medical examination.

[0037] The classification method performed by the clustering unit 22 is not particularly limited, and may be, for example, k-means, latent class analysis (LCA), silhouette analysis, consensus clustering, random forest, neural network, principal component analysis, hierarchical clustering, or the like. The clustering unit 22 may also perform processing to evaluate whether the number of classified phenotypes is appropriate. Specific evaluation methods include, for example, the elbow method and silhouette analysis. When the clustering unit 22 uses silhouette analysis as the evaluation method, processing may be performed to exclude patients whose silhouette coefficient in the analysis is evaluated to be a negative value from the classification targets.

[0038] The number of times that the clustering unit 22 performs classification processing on the medical data to be classified is not particularly limited, and may be performed multiple times.

[0039] For example, the clustering unit 22 may further input into a classification model only the medical data of patients classified into a phenotype group in which the proportion of patients who will develop a poor prognosis within a predetermined period from their first medical treatment is 45% or more, along with the classification results of the classification into phenotype groups, and output the classification results. That is, the clustering unit 22 may perform the classification process twice. The classification model used in the second classification may be a method different from the classification method used in the first classification. By the clustering unit 22 classifying patients twice, the calculation cost can be reduced.

[0040] The clustering unit 22 may also extract features of each classified phenotype group by calculating the mean, standard deviation, median, or interquartile range for each item of medical data included in each phenotype. It is preferable that the clustering unit 22 calculates the median and interquartile range (IQR) for each item of medical data.

[0041] [Output section] The output unit 23 outputs information indicating the classification results into a plurality of phenotype groups output from the clustering unit 22 to the presentation device 4. The output unit 23 may be configured to cause the presentation device 4 to present the medical data to be classified together with the information indicating the classification results.

[0042] This allows medical personnel who are presented with the classification results to use the classification results to determine the medical intervention required for the patient and whether or not medical intervention for the patient should be rushed.

[0043] Furthermore, the output unit 23 may be configured to cause the presentation device 4 to present information indicating that medical intervention is necessary, particularly for patients classified into a phenotype group in which the proportion of patients with a poor prognosis within a predetermined period after first medical treatment is 45% or more. With this configuration, the clinical data classification device 1 can present information indicating that medical intervention is necessary to medical personnel along with the classification results.

[0044] The method of presenting the classification results to the user may be any desired method. For example, as shown in Fig. 1, the classification results may be displayed on a presentation device 4, or may be output from a printer (not shown) and a speaker (not shown).

[0045] (Processing performed by the medical data classification device 1) Next, the processing executed by the medical data classification device 1 will be described with reference to Fig. 2. Fig. 2 is a flowchart showing an example of the flow of processing executed by the medical data classification device 1.

[0046] First, the data acquiring unit 21 acquires medical data created from the time a patient is transported by ambulance to the medical institution until the patient's first medical treatment (Step S1: data acquiring step) from the external device 5. Here, the data acquiring unit 21 may select and discard multiple items included in the medical data as necessary.

[0047] Next, the clustering unit 22 inputs the medical data into a classification model and outputs the classification results in which the patients are classified into phenotype groups (step S2: clustering step).

[0048] Next, the output unit 23 outputs the classification result obtained by the clustering unit 22 to the presentation device 4 (step S3), and the classification process executed by the clinical data classification device 1 ends.

[0049] Thus, a medical data classification method according to one aspect of the present invention includes the steps of acquiring medical data generated from the time a patient is transported to a medical institution by ambulance until the time the patient has completed treatment for a disease or injury at the medical institution, and classifying a plurality of patients into a plurality of phenotype groups. This allows a plurality of patients suffering from a predetermined disease or injury to be classified into a plurality of phenotype groups based on the medical data generated at an early stage, from the time the patient is transported to a medical institution by ambulance until the time the patient has completed treatment for a disease or injury at the medical institution.

[0050] The classification method disclosed herein can classify patients suffering from a certain disease or trauma at an early stage, which can identify new therapeutic target groups and aid in the development of treatment strategies. For example, by applying the classification method disclosed herein, clinicians can identify the likelihood of death in patients suffering from a certain disease or trauma at an early stage and implement appropriate medical interventions.

[0051] [Software implementation example] The functions of the medical data classification device 1 (hereinafter referred to as the "device") can be realized by a program that causes a computer to function as the device, and a program that causes a computer to function as each control block of the device (particularly each part included in the control unit 2).

[0052] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The control device and storage device execute the program, thereby realizing the functions described in each of the above embodiments.

[0053] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.

[0054] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.

[0055] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Example]

[0056] An embodiment of the present invention will be described below. Figure 3 is a diagram showing an outline of the classification process carried out in this embodiment.

[0057] We obtained medical data from 158,918 patients registered in the Japan Trauma Data Bank (2013-2017). All data were collected from the time a patient was transported to a medical institution by ambulance until the time they completed treatment for their illness or injury. Of the collected data, data from 12,565 patients were excluded from classification due to non-blunt trauma. Of the remaining 146,353 blunt trauma patients, 75,315 patients were excluded because they did not meet the inclusion criteria. Specifically, patients who did not meet the inclusion criteria were not transported (n = 36,267), had an ISS of 75 (n = 13,667), were in cardiopulmonary arrest at the time of admission (n = 6,555), were pregnant (n = 69), and had abnormal or missing data (n = 18,757).

[0058] <Classification of the analysis cohort (1)> Data from 71,038 patients were used as the analysis cohort for classification. The analysis cohort used for this classification is hereafter also referred to as analysis cohort D.

[0059] The variables used for classification were age, sex, CPS, respiratory rate (RR), pulse rate (HR), systolic blood pressure (SBR), Grass Cross-Coma Scale (GCS), head and neck AIS, facial AIS, thoracic AIS, abdominal AIS, extremity AIS, superficial AIS, trauma severity score (ISS), Revised Trauma Score (RTS), Trauma and Injury Severity Score (TRISS), and number of survivors.

[0060] RTS is a severity assessment based on physiological indicators. Specifically, RTS is calculated from the level of consciousness (GCS), systolic blood pressure, and respiratory rate. The most severe score is 0 points, and the best score is 7.84 points.

[0061] TRISS is a predicted survival rate calculated by adding physiological severity, anatomical severity, and age factors.

[0062] After averaging, each data set was clustered using the k-means method. The optimal number of clusters was evaluated using silhouette analysis, and data with negative silhouette coefficients were excluded from classification.

[0063] As a result of the classification, the analysis cohort was classified into eight phenotypes (D-1 to D-8). The classification results are shown in Figure 4. In Figure 4, [IQR] indicates the interquartile range. When comparing each phenotype with the survival rate of patients of each phenotype in Figure 4, it was found that the survival rate of one of the eight phenotypes (D-8) was extremely low. In other words, it was found that the mortality rate of one of the classified phenotypes was high.

[0064] Next, we further classified the data for the D-8 phenotype, which had a high mortality rate, using latent class analysis (LCA) (VarSelLCM package). As a result, the D-8 data were classified into four clusters (D-8α, D-8β, D-8γ, and D-8δ). Evaluation of the four phenotypes classified by LCA revealed that D-8α represented multiple trauma in young people (n=464), D-8β represented head trauma with hypothermia (n=178), D-8γ represented severe head injury in elderly people (n=957), and D-8δ represented multiple trauma with a predicted mortality rate higher than the actual mortality rate (n=579).

[0065] <Classification of validation cohort> Data from 42,780 patients (January 2013 to June 2015) registered in the Japan Trauma Data Bank were classified using the same method as for analysis cohort D. Data showing negative Silhouette coefficient values ​​were excluded from the cluster, and 38,097 patients were finally classified. The classification results are shown in Figure 5. As in the analysis cohort, the validation cohort was also classified into eight phenotypes, V-1 to V-8. Furthermore, it was found that patients classified into one of the phenotypes (V-8) had a higher mortality rate.

[0066] Next, we further classified the data for the V-8 phenotype, which had a high mortality rate, using the LCA method. As a result, the V-8 data were classified into four clusters (V-8α, V-8β, V-8γ, and V-8δ), similar to the D-8 data.

[0067] <Comparison between analysis cohort and validation cohort> Figure 6 shows heat maps showing the survival rates for each classified phenotype in the analysis and validation cohorts, as well as the distribution of variables for each phenotype. The upper panel shows the survival rates for each clinical phenotype as a bar graph, and the lower panel shows the heat maps with each variable listed, with each cell showing the median value for each variable for each phenotype.

[0068] Figure 7 shows the results of hierarchical clustering analysis of the principal component scores of the centroids of each phenotype. Principal component analysis was performed on each phenotype in the analysis cohort and validation cohort. Subsequently, the principal component scores of the centroids of each phenotype were calculated, and hierarchical cluster analysis was performed. Figure 7 confirms that clusters of high mortality phenotypes are paired together.

[0069] Figure 8 is a plot of the coordinates of the centers of gravity calculated in Figure 7. The horizontal axis is the first principal component axis, and the vertical axis is the second principal component axis. The size of each plot indicates the number of patients. Figure 8 visualizes pairs with similar phenotypes in the analysis cohort and the validation cohort.

[0070] <Classification of the analysis cohort (2) Evaluation of molecular pathogenesis> After biological invasion due to a certain disease or trauma, damage-associated molecular patterns (DMAPs) associated with tissue injury bind as ligands to pattern recognition receptors present on immune cells. Activated intracellular transcription factors bind to DNA in the nucleus, and proteins are translated from the transcribed messenger RNA (mRNA), leading to the progression of a systemic inflammatory response. Therefore, comprehensive evaluation of blood proteins using proteomics analysis makes it possible to understand the molecular pathology of trauma.

[0071] In this example, an analysis cohort different from the above-mentioned analysis cohort D was used, and after classification, each cluster was evaluated using serum proteomics analysis data.

[0072] Medical data from 90 patients admitted to the Advanced Emergency and Critical Care Center at Osaka University Hospital between February 2017 and March 2021 (Osaka University cohort in Figure 3) were collected and classified using the same method as in <Classification of Analysis Cohorts (1)>. Additionally, general research data from patients admitted to the hospital and serum proteomic analysis data analyzed within 72 hours of injury were used to evaluate the molecular pathology of each classified phenotype. This study was conducted in accordance with the Declaration of Helsinki and approved by the Osaka University Ethics Committee (IRB approval no. 16260, 885). The analysis cohort used for this classification is hereafter referred to as Analysis Cohort B.

[0073] As a result of the classification, analysis cohort B was classified into eight phenotypes (B-1 to B-8), similar to analysis cohort D. Each classified phenotype was evaluated using serum proteomics analysis data. In the serum proteomics analysis data, the contents of the components shown below were used as variables. Note that the concentrations of the components shown below may also be used as variables in serum proteomics analysis. Sodium Potassium Chloride Total bilirubin AST (aspartate aminotransferase) ALT (alanine aminotransferase) ·LDH (serum lactate dehydrogenase) BUN (urea nitrogen) Creatinine C-Reactive Protein ·WBC count (white blood cell count) Hemoglobin ·Platelet count PT-INR (prothrombin time) APTT (activated partial thromboplastin time) FDP (fibrinogen-fibrin degradation products) D-dimer ·Lactate The evaluation revealed that the B-8 phenotype exhibits excessive inflammation compared to other phenotypes, including an acute inflammatory response, dysregulated complement activation pathways, impaired coagulation, and platelet degranulation pathways.

[0074] In cohort B, we evaluated the differences in serum protein expression between the high mortality phenotype and other phenotypes. Specifically, differential expression analysis was performed using Welch's two-group t-test. The false discovery rate (FDR) was calculated using the Benjamin-Hochberg method (see Benjamini, Y. & Hochberg, Y. Journal of the Royal Statistical Society. Series B (Methodological) 57, 289-300 (1995)). Proteins with an FDR of less than 0.2 were considered to be significantly expressed.

[0075] Figure 9 shows a volcano plot of serum proteins whose expression was differentially determined between the high mortality phenotype and other phenotypes. Darker plots indicate proteins whose expression was upregulated, and lighter plots indicate proteins whose expression was downregulated. There were 11 upregulated proteins and 26 downregulated proteins.

[0076] Figure 10 shows the results of the Gene Ontology (GO) enrichment analysis of proteins. The top 12 significant GO terms are shown in Figure 10 along with their FDR-corrected p-values ​​calculated by the Benjamin-Hochberg method. The size of the plot indicates the number of differentially expressed proteins in the list of significant proteins associated with the GO term.

[0077] GO enrichment analysis revealed that phenotype B-8 (equivalent to phenotype D-8 / V-8) exhibited excessive inflammation, including enhanced acute inflammatory response, dysregulation of complement activation pathway, and coagulation disorders such as reduced coagulation and platelet granule pathway.

[0078] Figure 11 is a plot of each variable in the serum proteomics analysis data for phenotypes B-1 to B-7 and phenotype B-8 (high mortality rate) against the standardized values ​​of each variable. The variables were normalized so that the mean value of the variables was 0 and the standard deviation (SD) was 1. Phenotype B-8 was found to have high laboratory values ​​related to coagulation, fibrinolysis, and inflammation, and lower fibrinogen and platelet values ​​than the other phenotypes.

[0079] In this example, multiple patients were classified into multiple phenotype groups based on medical data collected from the time the patient was transported by ambulance due to their illness or injury until their first medical examination after their arrival at the medical institution. The analysis cohort was classified into 11 phenotype groups. The validation cohort was also classified into 11 phenotype groups, confirming the reproducibility of this classification method. [Explanation of symbols]

[0080] 1. Medical data classification device 21 Data Acquisition Section 22 Clustering Department 100 Medical Data Classification System

Claims

1. a data acquisition step in which a computer acquires medical data for each of a plurality of patients who have suffered from a predetermined disease or an injury, the medical data being created from the time the patient is transported to a medical institution by ambulance to the time the patient has finished receiving medical treatment for the disease or the injury at the medical institution; a clustering step in which the computer classifies the plurality of patients into a plurality of phenotype groups based on the clinical data; the medical data includes at least attribute information, vital sign information, and symptom information of each of the plurality of patients; The symptom information includes a site of a disease and a degree of the disease, or a site of an injury and a degree of the symptom at the site. Methods for classifying medical data.

2. The plurality of phenotype groups includes at least one phenotype group in which the proportion of patients with a poor prognosis is 45% or more within a predetermined period from the first medical treatment. The method for classifying medical data according to claim 1 .

3. The attribute information includes the age and sex of the patient. The method for classifying medical data according to claim 1 or 2.

4. If the patient is an injured patient, the symptom information includes at least one of information indicating a type of injury to the patient's body, a degree of injury, and a biological reaction caused by the injury. The method for classifying medical data according to any one of claims 1 to 3.

5. The type of trauma is a type classified by AIS (Abbreviated Injury Scale) coding, The method for classifying medical data according to any one of claims 1 to 4.

6. The predetermined disease is an acute disease including at least one of myocardial infarction, cerebral infarction, heat stroke, cerebral hemorrhage, cardiac arrest, poisoning, disseminated intravascular coagulation (DIC), acute respiratory distress syndrome (ARDS), burns, hemorrhagic shock, and infectious disease. The method for classifying medical data according to any one of claims 1 to 5.

7. The vital sign information includes at least one of the patient's respiratory rate, pulse rate, systolic blood pressure, level of consciousness (Glasgow Coma Scale (GCS) or Japan Coma Scale (JCS)), and body temperature; The method for classifying medical data according to any one of claims 1 to 6.

8. a data acquisition unit that acquires medical data for each of a plurality of patients who have suffered from a predetermined disease or injury, the medical data being created from the time the patient is transported to a medical institution by ambulance for the disease or injury until the time the patient has completed medical treatment for the disease or injury at the medical institution; a clustering unit that classifies the plurality of patients into a plurality of phenotype groups based on the clinical data, the medical data includes at least attribute information, vital sign information, and symptom information of each of the plurality of patients; The symptom information includes the location of a disease and the extent of the disease, or the location of an injury and the extent of the symptoms at the location. Medical data classification device.

9. A computer, a data acquisition step of acquiring medical data for each of a plurality of patients who have suffered from a predetermined disease or injury, the medical data being created from the time the patient is transported to a medical institution by ambulance regarding the disease or injury until the time the patient has completed medical treatment for the disease or injury at the medical institution; a clustering step of classifying the plurality of patients into a plurality of phenotype groups based on the clinical data, the medical data includes at least attribute information, vital sign information, and symptom information of each of the plurality of patients; The symptom information includes the location of a disease and the extent of the disease, or the location of an injury and the extent of the symptoms at the location. Control program.

Citation Information

Patent Citations

  • System, program and method for verifying diagnosis group classification

    JP2010157128A

  • Seriousness determination device, and seriousness determination method

    JP2013148996A

  • Data analysis support system and data analysis support method

    JP2018180993A

  • Emergency transfer support system

    JP2021111373A

  • Method for generating prediction result for early predicting occurrence of fatal symptoms of subject and device using same

    JP2021177429A