Cohort extraction method, cohort extraction device, and cohort extraction program that implement the same

The method of generating a history table with incremental updates and bit strings for condition satisfaction addresses the inefficiencies of conventional cohort extraction methods, enabling faster and more efficient cohort extraction and re-extraction.

JP7692064B2Active Publication Date: 2025-06-12KAKAO HEALTHCARE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023576223
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-07
Filing Date
2022-05-11
Publication Date
2025-06-12
Estimated Expiration
2042-05-11

AI Technical Summary

Technical Problem

Conventional cohort extraction methods are inefficient as they require repeated queries to a Clinical Data Warehouse (CDW) when conditions change, leading to unnecessary work and prolonged extraction times.

Method used

A method involving the generation of a history table that tracks events of each patient step-by-step, with bit strings indicating condition satisfaction at each stage, allowing for incremental updates and efficient re-extraction of cohorts based on changing conditions.

Benefits of technology

This approach significantly reduces the time and resources required for cohort extraction by allowing quick calculation of patient and event numbers at each stage and enabling efficient re-extraction of cohorts when conditions change.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692064000005
    Figure 0007692064000005
  • Figure 0007692064000006
    Figure 0007692064000006
  • Figure 0007692064000007
    Figure 0007692064000007
Patent Text Reader

Abstract

A method for operating a cohort extraction device, the method including the steps of receiving a cohort generation condition and extracting events corresponding to the cohort generation condition from a clinical data warehouse; generating a first history table including an event identifier, a patient identifier, and a bit string indicating satisfaction of a first stage condition for each extracted event; receiving a condition for a current stage and identifying a patient of the current stage having an event corresponding to the current stage condition from among patients included in a history table of a previous stage, updating the bit string for each event of the patient of the current stage included in the history table of the previous stage, and generating a history table of the current stage by adding a new event extracted in the current stage; and generating a cohort table using the history table of the final stage after sequentially generating history tables by stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to patient cohort extraction.

Background Art

[0002] Cohort extraction is very important because researchers conduct medical research using cohorts extracted from a Clinical Data Warehouse (CDW). Therefore, researchers try to determine whether a cohort satisfying various conditions is appropriate and extract a cohort with an appropriate number of patients while changing the conditions.

[0003] However, conventional cohort extraction devices receive conditions and output a group of patients who satisfy all the conditions from the CDW, but the number of patients extracted according to the conditions is variable. Therefore, researchers have to repeat the cohort extraction work from the huge CDW while changing the conditions, which takes a considerable amount of time until they obtain a satisfactory cohort. Also, as the number of conditions increases, the amount of queries increases, and unnecessary work is repeated because patients with unchanged conditions have to be extracted again.

Summary of the Invention

Problems to be Solved by the Invention

[0004] This disclosure provides a method for extracting a cohort step by step, a cohort extraction device and a cohort extraction program that implement this method.

[0005] Specifically, this disclosure provides a method for extracting a cohort by generating a history table including events of each patient for each step and updating a bit string indicating whether each event satisfies the condition in the history table.

Means for Solving the Problems

[0006] A method for operating a cohort extraction device according to an embodiment, comprising: receiving cohort generation conditions, and extracting events corresponding to the cohort generation conditions from a clinical data warehouse; generating a first history table including event identifiers, patient identifiers, and a bit string indicating satisfaction of conditions in the first stage for each of the extracted events; receiving conditions for the current stage, identifying patients in the current stage having events corresponding to the conditions for the current stage among the patients included in the immediately previous stage's history table, updating the bit string for each event of the patients in the current stage included in the immediately previous stage's history table, and adding newly extracted events in the current stage to generate a history table for the current stage; and after sequentially generating stage-by-stage history tables, generating a cohort table using the history table for the final stage.

[0007] Each history table generated stage by stage includes events satisfying the conditions for that stage, and an event identifier, a patient identifier, and a bit string indicating whether the conditions have been satisfied up to that stage for each event are described. The bit string is specified with digits indicating whether the conditions for each stage have been satisfied with 1 or 0.

[0008] The step of generating the history table for the current stage can confirm the events of the patients in the current stage from the history table of the immediately previous stage, update the bit string of the confirmed events to a value indicating satisfaction of the conditions for the current stage, and record it in the history table for the current stage.

[0009] The step of generating the history table for the current stage can record the identifier of the new event, the patient identifier, and the bit string indicating satisfaction of the conditions for the current stage in the history table for the current stage if a new event is extracted in the current stage. The bit string of the new event is described such that the value of the digit specified for the current stage is 1 and the values of the digits specified for other stages are 0.

[0010] The step of generating the history table of the current stage identifies patients at the previous stage among the patients included in the history table of the immediately previous stage who do not have an event corresponding to the conditions of the current stage, and does not record the events of the patients at the previous stage in the history table of the current stage.

[0011] If the number of events or patients extracted at a specific stage is requested, the operation method may further include a step of calculating the number of events or the number of patients using the history table of the specific stage.

[0012] The operation method further includes a step of receiving a change condition for a specific stage, a step of bringing in the history table of the immediately previous stage generated at the immediately previous stage of the specific stage, identifying patients at the specific stage among the patients included in the history table of the immediately previous stage who have an event corresponding to the change condition of the specific stage, updating a bit string for each event of the patients at the specific stage included in the history table of the immediately previous stage, and adding newly extracted events at the specific stage to regenerate the history table of the specific stage.

[0013] The operation method may further include a step of sequentially regenerating the history table of the stage after the specific stage using the regenerated history table of the specific stage.

[0014] A method for operating a cohort extraction device according to another embodiment, comprising the steps of: receiving a condition; identifying, based on the clinical data of patients included in the first history table generated in the immediately previous step, the patients at the current stage who satisfy the condition among the patients included in the first history table; recording the event identifiers, patient identifiers, and updated bit strings of all events of the patients at the current stage included in the first history table in a second history table; when a new event corresponding to the condition is extracted, recording the event identifier, patient identifier, and bit string indicating the event extracted at the current stage of the new event in the second history table; and storing the second history table as the history table at the current stage.

[0015] For all events of the patients at the current stage included in the first history table, a bit string in which the value of the digit designated for the current stage is updated to 1 in the bit string recorded in the first history table is recorded in the second history table.

[0016] For the new event, a bit string in which the value of the digit designated for the current stage is 1 and the values of the digits designated for other stages are 0 is recorded in the second history table.

[0017] Among the events included in the first history table, the events of patients at the previous stage that do not have an event corresponding to the condition are not recorded in the second history table.

[0018] A computer program stored on a computer-readable storage medium according to still other embodiments, the computer program including instruction words executed by at least one processor, the method comprising: receiving cohort generation conditions; extracting events corresponding to the cohort generation conditions from a clinical data warehouse; generating a first history table including event identifiers, patient identifiers, and a bit string indicating satisfaction of conditions in a first stage for each of the extracted events; receiving current stage conditions, identifying current stage patients having events corresponding to the current stage conditions among patients included in the immediately preceding stage history table, updating the bit string for each event of the current stage patients included in the immediately preceding stage history table, adding newly extracted events in the current stage, and generating a current stage history table; and generating a cohort table using the history table of a final stage after sequentially generating stage-by-stage history tables.

[0019] Each history table generated step by step includes events that satisfy the conditions of that step, and an event identifier, a patient identifier, and a bit string indicating whether the conditions have been satisfied up to that step for each event are described. The bit string has digits specified to indicate whether the conditions of each step are satisfied with 1 or 0.

[0020] The step of generating the current stage history table checks the events of the current stage patients from the immediately preceding stage history table, updates the bit string of the confirmed events to a value indicating satisfaction of the current stage conditions, records it in the current stage history table, and if a new event is extracted in the current stage, the identifier of the new event, the patient identifier, and the bit string indicating satisfaction of the current stage conditions can be recorded in the current stage history table.

Advantages of the Invention

[0021] According to the embodiment, in order to manage, in a history table, the events of each patient extracted step by step and the bit strings indicating whether the stage-by-stage conditions of each event are satisfied, a plurality of history tables are used to quickly calculate the number of patients and the number of events at each stage. As a result, researchers can quickly determine the appropriateness of the cohort.

[0022] According to the embodiment, based on the bit strings indicating whether the stage-by-stage conditions of each event are satisfied, it is possible to quickly confirm the stage at which the event was extracted and the stage at which the event satisfies the conditions.

[0023] According to the embodiment, after completing the event extraction up to the final stage, if the conditions for a specific stage need to be changed, a new history table including the events that satisfy the changed conditions can be generated using the history table generated at the immediately previous stage.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0025] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present invention pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein. And, in order to clearly explain the present invention in the drawings, parts that are unnecessary for the explanation are omitted, and similar parts are denoted with similar reference numerals throughout the specification.

[0026] Throughout the specification, when a certain part "includes" a certain component, this means that, unless otherwise stated to the contrary, it does not exclude other components but can further include other components. Also, terms such as "... part", "... device", "module", etc. described in the specification mean a unit that processes at least one function or operation, and this is realized by hardware, software, or a combination of hardware and software.

[0027] FIG. 1 and FIG. 2 are diagrams for explaining a conventional cohort extraction method.

[0028] Referring to FIG. 1, a conventional cohort extraction device 10 receives cohort conditions (criteria) (condition 1, condition 2,..., condition n) from a researcher and extracts K patients who satisfy all the conditions from a Clinical Data Warehouse (CDW) 20 that stores various patient data. The conventional cohort extraction device 10 outputs a cohort table including the data of K patients.

[0029] If a researcher wants to change or delete Condition 1, they can input the changed conditions into the conventional cohort extraction device 10 and obtain a cohort composed of M patients who satisfy all the conditions. However, in the conventional cohort extraction device 10, if even one of the input conditions is changed, the cohort extraction operation has to be carried out again. Therefore, the cohort extraction operation has to be repeated, and patients with unchanged conditions have to be extracted again, resulting in unnecessary work being repeated. Also, if the number of conditions increases, the query volume increases, and the extraction time may be quite long.

[0030] Referring to FIG. 2, the conventional cohort extraction device 10 can receive cohort conditions (Condition 1, Condition 2,..., Condition n) from a researcher step by step and extract K patients while gradually reducing the number of patients. That is, the conventional cohort extraction device 10 can extract a first patient group that satisfies Condition 1, extract a second patient group that satisfies Condition 2 from the first patient group, and extract a third patient group that satisfies Condition 3 from the second patient group while extracting a K-patient group.

[0031] Since the patient groups extracted at each stage are patients who satisfy all the conditions up to that stage, the researcher can obtain patients who satisfy all the conditions set from the first stage to the current stage. In this way, the conventional cohort extraction device 10 focuses on extracting patients and only identifies patients who satisfy all the conditions up to the current stage (for example, hypertension diagnosis, 50s, male, prescription of Drug A, prescription of Drug B). Therefore, the researcher can only know that the extracted patients correspond to all the conditions up to the current stage (for example, hypertension diagnosis, 50s, male, prescription of Drug A, prescription of Drug B), but it is difficult to know whether Drugs A and B were prescribed together, separately, whether Drug A was prescribed at the time of hypertension diagnosis, or at the time of other diagnoses. If the researcher wants to obtain a cohort in which Drugs A and B are prescribed together, they have to analyze the patient data and reselect the patients.

[0032] On the one hand, if there is only one attribute for which a search is desired, such as keyword search, the search device only needs to extract the desired target from the one-dimensional data. However, even when extracting a cohort from the clinical data of a certain patient, data suitable for the conditions must be retrieved from tables for each attribute such as age, gender, main diagnosis name, secondary diagnosis name, diagnosis date, name of the drug administered, prescription date, etc. Therefore, although the cohort extraction operation becomes geometrically slower in terms of search speed according to the amount of tables, characteristics of attributes, and search conditions, if such operations are repeated every time the conditions are changed, time and resources will be wasted.

[0033] Hereinafter, a cohort extraction method that improves such a conventional method will be described in detail.

[0034] FIG. 3 is a diagram for explaining a cohort extraction device.

[0035] Referring to FIG. 3, the cohort extraction device 100 is a computing device operated by at least one processor. By the processor of the cohort extraction device 100 executing the instruction words included in the computer program, the operations of the present disclosure are performed. The computer program includes instruction words (instructions) described so that the processor executes the operations of the present disclosure and is stored in a non-transitory computer readable storage medium. The computer program is downloaded via a network or sold in a product form and is provided in computing devices at various sites such as research institutes and hospitals.

[0036] The cohort extraction device 100 extracts a cohort from a clinical data warehouse (CDW) 20 that stores various patient data. Although the types of patient data extracted from the clinical data warehouse (CDW) 20 are diverse, for convenience, they are generically referred to as clinical data. Also, although the cohort extraction device 100 can extract patient data from various storages, for convenience, it will be described as extracting from the clinical data warehouse.

[0037] The cohort extraction device 100 receives conditions step by step, extracts events corresponding to the conditions for each step, arranges the events for each patient, and generates a history table containing the events for each patient. Here, an event is information that can be confirmed from the clinical data warehouse (CDW) 20, and means information that classifies events, actions, etc. that occurred to a patient at a certain point in time. For example, an event is defined as a diagnosis event (e.g., the breakdown of a diabetes diagnosis that received a disease code of E10 - E14), a drug prescription event (e.g., the breakdown of Aspirin being prescribed), a test event (e.g., the breakdown of receiving a low-density lipoprotein (LDL) cholesterol test), a hospitalization event (e.g., the breakdown of an emergency room visit), etc. Here, the conditions can include cohort generation (entry) conditions (e.g., a person who has been diagnosed with hypertension at least once), and detailed conditions to be extracted (e.g., drugs, age, etc.). The detailed conditions are defined as including or excluding the relevant items, and may also be defined within a range.

[0038] After first generating the history table 1 for the cohort generation (entry) conditions, the cohort extraction device 100 separately generates history tables 2,..., history table n using the conditions (criteria) input step by step.

[0039] The history table includes a bit string that indicates the presence or absence of condition satisfaction up to each current step for each event with 0 and 1. Steps are assigned to each digit of the bit string. If the value of the bit is 1, it indicates that the condition for that step is satisfied, and if the value of the bit is 0, it can indicate that the condition for that step is not satisfied. For example, when the bit string is 10 bits long, "0000000001" indicates an event that satisfies the condition for step 1, "0000000011" indicates an event that satisfies the conditions for step 1 and step 2, and "0000000010" indicates an event that satisfies the condition for step 2.

[0040] The cohort extraction device 100 identifies patients at the current stage who have events corresponding to the conditions at the current stage among the patients included in the previous-stage history table. Then, the cohort extraction device 100 generates a current-stage history table composed of events that satisfy the conditions at the current stage.

[0041] At this time, if there is an event of a patient at the current stage existing in the previous-stage history table, the cohort extraction device 100 updates the bit string of the event (for example, updates from "0000000001" to "0000000011"), adds the event extracted at the current stage to the new event, and generates a current-stage history table. The new event describes a bit string (for example, "0000000010") in which the bit assigned to the current stage is "1".

[0042] The cohort extraction device 100 identifies patients at the previous stage who do not have events corresponding to the conditions at the current stage among the patients included in the previous-stage history table. Then, the cohort extraction device 100 does not bring the events of the patients at the previous stage from the previous-stage history table to the current-stage history table.

[0043] The history table is generated by stage, described in units of events, and although multiple events may be described for some patients, the events of patients having at least one event corresponding to the conditions at that stage are described. The schema of such a history table is defined in various ways. For example, as shown in Table 1, events are described for each row, event information is described in the columns, and they are aligned for each patient. The event information can include a patient identifier (person_ID), a visit identifier (visit_ID), an event start date (start_date), an event end date (end_date), an event type (event_type), and a detailed condition type (criteria_type). Here, the visit identifier (visit_ID), the event start date (start_date), and the event end date (end_date) can be used as event identifiers for distinguishing events.

[0044]

Table 1

[0045] In Table 1, the patient identifier (person_ID) is an identifier that distinguishes patients who meet the conditions. The visit identifier (visit_ID) is an identifier that distinguishes the visit during which the event occurred. The event start date (start_date) and the event end date (end_date) indicate the start date and the end date of the event. The event type (event_type) is stage information of the event, which is represented by a bit string indicating the presence or absence of condition satisfaction up to each current stage with 0 and 1, and is updated according to the stage. The detailed condition type (criteria_type) is information indicating the detailed conditions from which the event was extracted, and describes the detailed conditions when the event was first extracted.

[0046] The cohort extraction device 100 can calculate and output the number of patients and the number of events from the history tables of each stage. Therefore, the researcher can easily judge the appropriateness of the extracted cohort by looking at the number of patients and the number of events.

[0047] The cohort extraction device 100 can quickly extract only the events having a specific event type from the history table. For example, if the cohort extraction device 100 extracts the events with the event type described as "********11" from the history table, it can calculate the number of events generated by the patients who meet condition 2 among the events that meet condition 1, and calculate the number of patients having the events that meet condition 1 and condition 2 based on the patient identifiers of the events described as "********11". Therefore, the cohort extraction device 100 does not need to newly generate an SQL query for calculating the number of events and the number of patients and extract them from the CDW, and only needs to perform a bit operation on the column of the event type in the history table, so that the number of events and the number of patients can be quickly calculated.

[0048] The cohort extraction device 100 can generate a cohort table from the history table at the final stage or a specific stage and output it. The cohort table includes various clinical data of the patients included in the history table.

[0049] On the other hand, after the researcher has completed the event extraction up to the final stage, there may be a case where the researcher wants to change the conditions at a specific stage. In this case, the researcher inputs the specific stage to be changed and the change conditions into the cohort extraction device 100. Then, the cohort extraction device 100 brings in the history table generated immediately before the specific stage among the stored history tables, and can generate a new history table for the specific stage that includes events satisfying the change conditions using this history table.

[0050] Hereinafter, a method for the cohort extraction device 100 to generate a history table step by step will be described in detail.

[0051] FIGS. 4 to 6 are diagrams for explaining the cohort extraction method as an example.

[0052] With reference to FIGS. 4 to 6, a method for the cohort extraction device 100 to generate a history table step by step will be described as an example.

[0053] The condition for stage 1, which is the first stage, is a cohort generation (entry) condition. For example, it may be a person who has been diagnosed with hypertension at least once. The condition for stage 2 is assumed to be a drug. An event in which a drug or a specific drug is prescribed is extracted at stage 2. The condition for stage 3 is assumed to be age. Patients belonging to a specific age group are extracted at stage 3.

[0054] First, referring to FIG. 4, the cohort extraction device 100 receives the conditions of stage 1 and extracts hypertension diagnosis events corresponding to the conditions of stage 1 from the clinical data warehouse (CDW). For example, 9 events, event1, event2, ..., event9 are extracted. Assume that event1 and event2 are the hypertension diagnosis events of patient A, event3 is the hypertension diagnosis event of patient B, event4 and event5 are the hypertension diagnosis events of patient C, event6 is the hypertension diagnosis event of patient D, event7 and event8 are the hypertension diagnosis events of patient E, and event9 is the hypertension diagnosis event of patient F.

[0055] The cohort extraction device 100 stores the events extracted according to the conditions of stage 1 in the history table 1, and can record a bit string indicating whether the conditions are satisfied up to the current stage together with the patient identifier and the event identifier (visit identifier, event start date, event end date) for each event type. The cohort extraction device 100 can generate the history table of stage 1 as shown in Table 2. For convenience, the values of the event start date (start_date) and the event end date (end_date) in the history table are omitted.

[0056]

Table 2

[0057] Referring to Table 2, since the events are extracted in stage 1, "0000000001" with the most significant digit assigned to stage 1 being 1 is described in the event type. Since stage 1 is a cohort generation condition, the detailed condition type indicating the detailed conditions under which the event was extracted has a null value.

[0058] When the number of events extracted in stage 1 is requested, the cohort extraction device 100 can calculate the number of rows in the history table of stage 1 where the event type (event_type) is "0000000001" and output the number of events 9.

[0059] When the number of patients extracted in stage 1 is requested, the cohort extraction device 100 can calculate the number classified as the patient identifier (person_ID) from the history table of stage 1 and output the number of patients 6.

[0060] Referring to FIG. 5, the cohort extraction device 100 receives the condition (drug) of stage 2 and generates a history table 2 including events that satisfy the condition of stage 2 from the history table 1.

[0061] The cohort extraction device 100 refers to the clinical data warehouse (CDW) to identify the current-stage patients among the patients included in the history table 1 of stage 1 who have events corresponding to the condition (drug) of stage 2. Then, the cohort extraction device 100 updates the bit string of all events of the current-stage patients recorded in the history table 1 of stage 1 (for example, updates from "0000000001" to "0000000011"), adds the events extracted in stage 2 as new events, and generates the history table 2 of stage 2. At this time, the cohort extraction device 100 identifies the patients (previous-stage patients) among the patients included in the history table 1 who do not have any events corresponding to the condition (drug) of stage 2, and excludes the events of the previous-stage patients from coming to the history table of stage 2.

[0062] For example, assume that among the patients included in the history table 1, patient B does not have any events corresponding to the condition (drug) of stage 2. Assume that event10 - event14 are newly extracted in stage 2. Then, the cohort extraction device 100 can generate the history table 2 as shown in Table 3. The number of patients recorded in the history table 2 is 5, and the number of events is 13.

[0063]

Table 3

[0064] Referring to Table 3, since Patient B does not have an event corresponding to the condition (drug) in Stage 2, event3, which is the hypertension diagnosis event of Patient B, is not recorded in the history table 2.

[0065] The event types "0000000001" of event1, event2, event4 - event9 included in the history table 1 are events of the current stage patients who have events corresponding to the condition (drug) in Stage 2, so they are updated to "0000000011" where the second digit assigned to Stage 2 is 1.

[0066] event10 - event14 newly extracted in Stage 2 are added to the history table 2, and the event types are described as "0000000010" where the second digit from the end assigned to Stage 2 is 1. Also, since event10 - event14 are first extracted under the conditions of Stage 2, "drug" is described in the detailed condition type (criteria_type).

[0067] When the number of events extracted in Stage 2 is requested, the cohort extraction device 100 can calculate the number of rows in the history table 2 where the event type (event_type) is "0000000010" and output the number of events 5.

[0068] Referring to FIG. 6, the cohort extraction device 100 receives the conditions of Stage 3 and generates a history table 3 including the events that satisfy the conditions of Stage 3 from the history table 2.

[0069] The cohort extraction device 100 refers to the Clinical Data Warehouse (CDW) to identify the patients at the current stage who have events corresponding to the stage 3 conditions among the patients included in the history table 2 at stage 2. Then, the cohort extraction device 100 updates the bit string of the events of the patients at the current stage recorded in the history table 2 at stage 2 (for example, updates from "0000000011" to "0000000111"). And the cohort extraction device 100 can add the new events extracted at stage 3 to the history table 3 at stage 3.

[0070] If there are patients at the previous stage among the patients included in the history table 2 who do not meet the stage 3 conditions, the cohort extraction device 100 deletes the events of these patients.

[0071] On the other hand, when the condition is age / gender, the calculation conditions for age / gender can be the earliest event, the latest event, and each event of the patient.

[0072] For example, among the patients included in the history table 2, assume that patient D does not meet the stage 3 condition (age), and the remaining patients are patients at the current stage who meet the stage 3 conditions. Then, as shown in Table 4, the cohort extraction device 100 can generate a history table 3 that does not include event 6 and event 12 of patient D, who is a patient at the previous stage. The cohort extraction device 100 updates the bit string of the events of the patients at the current stage recorded in the history table 2 at stage 2. The third digit assigned to stage 3 in the bit string is updated to 1.

[0073] In addition, the cohort extraction device 100 adds the newly extracted events in step 3 to the history table 3 of step 3. However, when the age calculation condition is the earliest event of the patient, as shown in Table 4, new event 15, new event 16, new event 17, and new event 18, which have the same event identifiers as event 1, event 4, event 7, and event 9, the earliest events of patients A, C, E, and F respectively, can be added to the history table 3. Then, the cohort extraction device 100 describes age in the detailed condition type (criteria_type) of new event 15, new event 16, new event 17, and new event 18.

[0074]

Table 4

[0075] On the other hand, since event 15, event 16, event 17, and event 18 extracted under the age / gender condition have the same event identifiers (visit identifier, event start date, event end date) as event 1, event 4, event 7, and event 9, when calculating the number of events, the events extracted under the age / gender condition are excluded from the number of events. Therefore, the number of patients recorded in the history table 3 is 4, and the number of events is calculated as 11. The cohort extraction device 100 can identify events from each history table whose detailed condition type is age / gender (criteria_type = "age", criteria_type = "gender") and exclude them from the total number of events.

[0076] In this way, the cohort extraction device 100 generates a history table including events of each patient for each stage, and updates a bit string indicating whether a condition is satisfied for each event in the history table. Therefore, each time the cohort extraction device 100 searches for the number of patients satisfying the condition, it is not necessary to create an SQL query, and the number of patients and the number of events at each stage can be quickly calculated using a plurality of history tables. In particular, based on the bit string displayed for the event type, it is possible to quickly confirm the stage at which the event was extracted and the stage at which the event satisfies the condition.

[0077] FIG. 7 is a diagram for explaining a cohort re-extraction method using a history table.

[0078] Referring to FIG. 7, after the cohort extraction device 100 first generates a history table 1 for cohort generation (entry) conditions, it separately generates history tables 2, ..., history table n using the conditions input for each stage.

[0079] Thereafter, when the researcher changes the conditions at stage k (for example, stage 3), the cohort extraction device 100 can generate a new history table 3 corresponding to the changed conditions at stage 3 using the history table 2 at the immediately previous stage, stage 2. The cohort extraction device 100 can sequentially regenerate the history tables for the stages after stage 3 using the newly generated history table 3 regenerated in this way.

[0080] In this way, even if the researcher changes the conditions, the cohort extraction speed can be improved because the history table before the change can be used as it is and only the events for the changed conditions need to be extracted.

[0081] FIG. 8 is a flowchart of the cohort extraction method.

[0082] Referring to FIG. 8, the cohort extraction device 100 receives cohort generation conditions at the first stage and extracts events corresponding to the cohort generation conditions from the clinical data warehouse (CDW) (S110).

[0083] The cohort extraction device 100 generates a first history table including the event identifier (visit identifier, event start date, event end date) of the extracted event, the patient identifier, and a bit string indicating satisfaction of the first stage conditions (S120).

[0084] Thereafter, the cohort extraction device 100 receives the conditions of the current stage and extracts events corresponding to the conditions of the current stage from the clinical data of the patients included in the history table of the immediately previous stage (S130).

[0085] The cohort extraction device 100 identifies the patients of the current stage for which events corresponding to the conditions of the current stage have been extracted among the patients included in the history table of the immediately previous stage, updates the bit string of the events of the patients of the current stage included in the history table of the immediately previous stage, adds the newly extracted event at the current stage for the first time, and generates the history table of the current stage (S140). The cohort extraction device 100 identifies the patients of the previous stage that do not have events corresponding to the conditions of the current stage among the patients included in the history table of the immediately previous stage, and does not store the events of the patients of the previous stage stored in the history table of the immediately previous stage in the history table of the current stage.

[0086] The cohort extraction device 100 determines whether the current stage is the final stage (S150). If the current stage is not the final stage, the cohort extraction device 100 waits in a state where it can receive the conditions of the next extraction stage. If the cohort extraction device 100 is requested to end or generate a cohort table, it can determine that the current stage is the final stage.

[0087] If the current stage is the final stage, the cohort extraction device 100 generates a cohort table using the history table of the final stage (S160).

[0088] In this way, after the cohort extraction device 100 sequentially generates the step-by-step history table, it generates a cohort table using the history table of the final step.

[0089] FIG. 9 is a hardware configuration diagram of a computing device according to an embodiment.

[0090] Referring to FIG. 9, the cohort extraction device 100 can be realized by a computing device operated by at least one processor.

[0091] The cohort extraction device 100 can include one or more processors 110, a memory 130 for loading a computer program executed by the processor 110, a storage device 150 for storing the computer program and various data, and a communication interface 170. In addition, the cohort extraction device 100 can further include various components.

[0092] The processor 110 is a device that controls the operation of the cohort extraction device 100, and can be a processor in various forms that processes instruction words included in a computer program. For example, it can be configured to include at least one of a CPU (Central Processing Unit), an MPU (Micro Processor Unit), an MCU (Micro Controller Unit), a GPU (Graphic Processing Unit), or any form of processor well-known in the technical field of the present disclosure.

[0093] The memory 130 stores various data, instructions, and / or information. The memory 130 can load the computer program from the storage device 150 so that the instruction words described to execute the operation of the present disclosure are processed by the processor 110. The memory 130 can be, for example, a ROM (read only memory), a RAM (random access memory), or the like.

[0094] The storage device 150 can non-temporarily store computer programs and various types of data. The storage device 150 is configured to include a non-volatile memory such as a ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), flash memory, a hard disk, a removable disk, or any form of computer-readable recording medium well-known in the technical field to which the present disclosure pertains.

[0095] The communication interface 170 may be a wired / wireless communication module that supports wired / wireless communication. The communication interface 170 can be connected to the Clinical Data Warehouse (CDW) 20.

[0096] The computer program includes instructions to be executed by the processor 110, is stored in a non-transitory computer-readable storage medium, and the instructions cause the processor 110 to execute the operations of the present disclosure. The computer program can be downloaded via a network or sold in a product form.

[0097] A computer program can include instruction words that receive cohort generation conditions, extract events corresponding to the cohort generation conditions from a clinical data warehouse (CDW), and generate a first history table including event information of the extracted events, patient identifiers, and a bit string indicating whether the conditions up to the current stage are satisfied. Then, the computer program receives the conditions of the current stage, identifies the patients at the current stage who have events corresponding to the conditions of the current stage among the patients included in the history table of the immediately previous stage, updates the bit strings of the events of the patients at the current stage included in the history table of the immediately previous stage, adds the events extracted at the current stage to new events, and can include instruction words for generating a history table of the current stage. The program can determine whether the current stage is the final stage, and if the current stage is the final stage, can include instruction words for generating a cohort table using the history table of the final stage. If the current stage is not the final stage, the computer program can include instruction words for waiting in a state where the conditions of the next extraction stage can be received.

[0098] The embodiments of the present disclosure described above are not only realized by devices and methods, but may also be realized by a program that realizes functions corresponding to the configurations of the embodiments of the present disclosure or a recording medium on which the program is recorded.

[0099] Although the embodiments of the present disclosure have been described in detail above, the scope of rights of the present disclosure is not limited thereto, and various modifications and improvements made by those skilled in the art using the basic concepts of the present disclosure defined in the following claims also belong to the scope of rights of the present disclosure.

Claims

1. A method for operating a cohort extraction device, comprising: receiving cohort generation conditions and extracting events corresponding to the cohort generation conditions from a clinical data warehouse; generating a first history table including event identifiers, patient identifiers, and bit strings indicating satisfaction of conditions in the first stage for each of the extracted events; receiving conditions for the current stage, identifying patients in the current stage having events corresponding to the conditions for the current stage among the patients included in the immediately previous stage's history table, updating the bit strings for each event of the patients in the current stage included in the immediately previous stage's history table, and adding newly extracted events in the current stage to generate a history table for the current stage; generating a cohort table using the history table of the final stage after sequentially generating stage-by-stage history tables; and the method of operation including the above.

2. Each history table generated stage by stage includes events that satisfy the conditions for that stage, and event identifiers, patient identifiers, and bit strings indicating whether the conditions up to that stage are satisfied are described. The bit string is a digit that indicates whether the conditions for each stage are satisfied with 1 or 0, according to the method of operation described in Claim 1.

3. The stage of generating the history table for the current stage confirms the events of the patients in the current stage from the history table of the immediately previous stage, updates the bit string of the confirmed events to a value indicating satisfaction of the conditions for the current stage, and records it in the history table for the current stage, according to the method of operation described in Claim 1.

4. The stage of generating the history table for the current stage if a new event is extracted in the current stage, records the identifier of the new event, the patient identifier, and the bit string indicating satisfaction of the conditions for the current stage in the history table for the current stage. The bit string of the new event is described such that the value of the digit specified for the current stage is 1 and the values of the digits specified for other stages are 0, according to the method of operation described in Claim 1.

5. The stage of generating the history table for the current stage identifies patients in the previous stage who do not have events corresponding to the conditions for the current stage among the patients included in the history table of the immediately previous stage, and does not record the events of the patients in the previous stage in the history table for the current stage, according to the method of operation described in Claim 1.

6. If the number of events or patients extracted at a specific stage is requested, calculating the number of events or the number of patients using the history table of the specific stage The operation method according to claim 1, further comprising:

7. Receiving a change condition for a specific stage; Bringing in the history table of the immediately preceding stage generated at the stage immediately preceding the specific stage; Identifying patients at the specific stage having events corresponding to the change condition of the specific stage among the patients included in the history table of the immediately preceding stage, updating a bit string for each event of the patients at the specific stage included in the history table of the immediately preceding stage, and adding newly extracted events at the specific stage to regenerate the history table of the specific stage The operation method according to claim 1, further comprising:

8. Using the regenerated history table of the specific stage to sequentially regenerate the history tables of the stages after the specific stage The operation method according to claim 7, further comprising:

9. A computer program stored in a computer-readable storage medium and including instruction words executed by at least one processor, Receiving a cohort generation condition and extracting events corresponding to the cohort generation condition from a clinical data warehouse; Generating a first history table including an event identifier, a patient identifier, and a bit string indicating satisfaction of a condition at the first stage for each extracted event; Receiving a condition for the current stage, identifying patients at the current stage having events corresponding to the condition for the current stage among the patients included in the history table of the immediately preceding stage, updating a bit string for each event of the patients at the current stage included in the history table of the immediately preceding stage, and adding newly extracted events at the current stage to generate a history table for the current stage; After sequentially generating stage-by-stage history tables, generating a cohort table using the history table of the final stage A computer program including instruction words described to perform the above.

10. Each history table generated stage by stage Includes events that satisfy the conditions of the stage, and an event identifier, a patient identifier, and a bit string indicating whether the conditions have been satisfied up to the stage for each event are described. The computer program according to claim 9, wherein the bit string is specified with a digit indicating the presence or absence of satisfaction of the conditions at each stage by 1 or 0.

11. The step of generating the history table for the current stage is checking the event of the patient at the current stage from the history table of the immediately previous stage, updating the bit string of the checked event to a value indicating satisfaction of the conditions at the current stage, and recording it in the history table of the current stage, If a new event is extracted at the current stage, recording the identifier of the new event, the patient identifier, and the bit string indicating satisfaction of the conditions at the current stage in the history table of the current stage. The computer program according to claim 9.

Citation Information

Patent Citations

  • Incidence evaluating device, and program

    JP2006259809A

  • Information processing method, device, and program

    JP2015076031A

  • Systems and methods for model-assisted cohort selection

    JP2020516997A

  • Automatic Generation of User and Non-User Cohorts for The Estimation of Drug's Effect from Longitudinal Observational Data

    US20200005907A1