Program, information processing system, and information processing method
The program unifies medical data into standard terms using sequential conversion rules, addressing format inconsistencies and enhancing predictive model training and performance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MEDICU INC
- Filing Date
- 2024-11-06
- Publication Date
- 2026-05-19
AI Technical Summary
Differences in electronic medical record formats and description rules across medical institutions lead to varied representations of the same medical data, hindering effective data unification and utilization in machine learning applications.
A program and system that unifies character strings into standard medical terms by applying sequential conversion rules, including expansion, partial matching, and string conversion, to ensure consistent data representation for training predictive models.
Enables the unification of medical data into standardized terms, facilitating accurate training of predictive models for patient condition prediction and sudden change detection, thereby improving data consistency and model performance.
Smart Images

Figure 2026081810000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a program, an information processing system, and an information processing method.
Background Art
[0002] Conventionally, learning models obtained by machine learning have been used in various scenarios, and are also used in the medical field. Patent Document 1 discloses a prediction determination model that can accurately predict CRT non-responders. This prediction determination model is obtained by machine learning using data of patients who may become non-responders as teacher data.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] [[ID=3)5]]There are cases where it is desired to use data obtained in medicine, such as learning data in machine learning. In such cases, it is desirable that the data formats are unified to some extent. However, due to differences in the formats of electronic medical records used by each medical institution and differences in the description rules for each doctor, there is a problem that even data with the same meaning is generated in various ways.
[0005] The present invention has been made in consideration of such points, and aims to unify character strings indicating the same content into predetermined terms.
Means for Solving the Problems
[0006] The program of the present invention is a program for causing a computer to function as an acquisition unit and a unification unit, The acquisition unit acquires a string, The aforementioned unification section is, Determine whether the aforementioned string matches a pre-registered standard medical term. If the string does not match the standard medical term, Multiple string conversion rules, whose application order is predetermined, are applied sequentially to the string. The method is characterized by determining whether the converted string obtained by applying the conversion rule matches the standard medical term, and if it matches the standard medical term, unifying the string to the standard medical term.
[0007] In the program of the present invention, The aforementioned unification section is, If the string is modified by applying the conversion rule, the conversion rule may be applied to the modified string in the order of application.
[0008] In the program of the present invention, The aforementioned program further enables the computer to function as a learning unit, The learning unit may use the training data, which includes the unified string, to train a predictive model that predicts the patient's condition.
[0009] In the program of the present invention, The aforementioned program further enables the computer to function as a learning unit, The learning unit may use the training data, which includes the unified string, to train a predictive model that predicts sudden changes in a patient's condition.
[0010] In the program of the present invention, The aforementioned multiple conversion rules may include expansion rules that expand a collective notation of multiple disease names into multiple notations corresponding to those disease names.
[0011] In the program of the present invention, The aforementioned collective notation of multiple disease names may include numerical notation.
[0012] In the program of the present invention, The plurality of conversion rules may include a string conversion rule for converting a predetermined string into the standard medical term.
[0013] In the program of the present invention, The plurality of conversion rules may include a partial match rule for deleting a predetermined string.
[0014] In the program of the present invention, The plurality of conversion rules include an expansion rule for expanding a notation combining a plurality of disease names into a plurality of disease names, and a partial match rule for comparing a part of the string with the standard medical term, and the partial match rule may be applied prior to the expansion rule.
[0015] In the program of the present invention, The plurality of conversion rules include a string conversion rule for converting a predetermined string into the standard medical term, and a partial match rule for comparing a part of the string with the standard medical term, and the string conversion rule may be applied prior to the partial match rule.
[0016] In the program of the present invention, The plurality of conversion rules include an expansion rule for expanding a notation combining a plurality of disease names into a plurality of disease names, and a string conversion rule for converting a predetermined string into the standard medical term, and the string conversion rule may be applied prior to the expansion rule.
[0017] In the program of the present invention, The above multiple conversion rules are: The system may include a conversion rule that converts a first medical term into a second medical term that encompasses the first medical term.
[0018] In the information processing system of the present invention, An information processing system comprising an acquisition unit and a unification unit, The acquisition unit acquires a string, The aforementioned unification section is, Determine whether the aforementioned string matches a pre-registered standard medical term. If the string does not match the standard medical term, Multiple string conversion rules, whose application order is predetermined, are applied sequentially to the string. The method is characterized by determining whether the converted string obtained by applying the conversion rule matches the standard medical term, and if it matches the standard medical term, unifying the string to the standard medical term.
[0019] In the information processing method of the present invention, An information processing method performed by a computer having a control unit, The control unit performs the steps of obtaining a string, The control unit determines whether the string matches a pre-registered standard medical term, and if the string does not match a standard medical term, it applies a plurality of string conversion rules, whose application order is predetermined, to the string in order, determines whether the converted string obtained by applying the conversion rules matches a standard medical term, and if it matches a standard medical term, it unifies the string to the standard medical term. It is characterized by including. [Effects of the Invention]
[0020] According to the program, information processing system, and information processing method of the present invention, strings representing the same content can be unified into predetermined terms. [Brief explanation of the drawing]
[0021] [Figure 1] This is an overall diagram of the information processing system. [Figure 2] This is a flowchart of the model management process. [Figure 3] This figure shows an example of the data structure of medical data. [Figure 4] This flowchart shows the detailed process for standardizing medical terminology. [Figure 5] This figure shows an example of how to write a string. [Figure 6] This figure shows an example of the data structure for treatment information related to intravenous infusion at Hospital A. [Figure 7] This figure shows an example of the data structure for treatment information related to intravenous infusion at Hospital B. [Figure 8] This figure shows an example of data structure for treatment information related to intravenous administration, recorded in a time-series format. [Figure 9] This figure shows an example of the data structure for treatment information regarding the use of ventilators at Hospital A. [Figure 10] This figure shows an example of the data structure for treatment information regarding the use of ventilators at Hospital B. [Figure 11] This figure shows an example of data structure for treatment information regarding the use of ventilators, recorded in a time-series format. [Modes for carrying out the invention]
[0022] The embodiment will now be described with reference to the drawings. Figure 1 is an overall configuration diagram of the information processing system 1 according to this embodiment. The information processing system 1 according to this embodiment acquires medical data generated at each of multiple medical institutions and generates a sudden change prediction model that predicts sudden changes in a patient's condition using the medical data as training data. Furthermore, the information processing system 1 according to this embodiment performs sudden changes in a patient's condition by using the sudden change prediction model generated in this manner.
[0023] "Sudden deterioration" refers to a predetermined symptom that necessitates new treatment. Examples of "sudden deterioration" include sudden deterioration requiring admission to the intensive care unit, sudden deterioration in arterial oxygen saturation requiring oxygen administration or mechanical ventilation, sudden deterioration in blood pressure or pulse requiring vasopressors, massive fluid resuscitation or blood transfusion, the development of disseminated intravascular coagulation requiring treatment, and impending or reversible cardiac arrest requiring vasopressors, massive fluid resuscitation, blood transfusion or aortic clamping. Furthermore, it includes the initiation of cardiopulmonary resuscitation or death in hospital as a result of any of these. Typically, "sudden deterioration" is a sudden deterioration in arterial oxygen saturation requiring oxygen administration or mechanical ventilation, sudden deterioration in blood pressure or pulse requiring vasopressors, massive fluid resuscitation or blood transfusion, the development of disseminated intravascular coagulation requiring treatment, and impending or reversible cardiac arrest requiring vasopressors, massive fluid resuscitation, blood transfusion or aortic clamping.
[0024] The information processing system 1 generates a sudden change prediction model and uses it to predict sudden changes. However, as an alternative, the information processing system 1 may predict not only sudden changes but also any predetermined state of the patient. Specifically, the information processing system 1 may generate a prediction model to predict the patient's state and use this prediction model to predict the state of the target patient. Examples of such states include decreased blood pressure and decreased oxygen saturation.
[0025] The information processing system 1 comprises a model management device 10, a user terminal 20, and multiple hospital servers 30. The model management device 10 generates a rapid change prediction model and performs rapid change prediction using the rapid change prediction model. The user terminal 20 is an information processing device used by users of the model management device 10. The two hospital servers 30 are located in different hospitals (Hospital A and Hospital B). In this embodiment, for the sake of explanation, only two hospital servers 30 are shown, but the information processing system 1 may have three or more hospital servers 30. In this case, the three or more hospital servers 30 are assumed to be located in different hospitals.
[0026] The model management device 10, user terminals 20, and multiple hospital servers 30 are connected via network N.
[0027] The model management device 10 is configured, for example, by a computer, and mainly comprises a control unit 100, a storage unit 110, and a communication unit 120.
[0028] The control unit 100 includes a processor such as a CPU (Central Processing Unit) and controls the operation of the model management device 10. The communication unit 120 includes a communication interface that communicates with external devices wirelessly or via wired connection. The control unit 100 transmits and receives data between the user terminal 20 and the hospital server 30 via the communication unit 120.
[0029] The storage unit 110 includes, for example, an HDD (Hard Disk Drive), RAM (Random Access Memory), ROM (Read Only Memory), and SSD (Solid State Drive). Furthermore, the storage unit 110 is not limited to one built into the model management device 10, but may be a storage medium that can be detachably attached to the model management device 10 (for example, a USB memory stick). The storage unit 110 stores programs executed by the control unit 100, sudden change prediction models, and various other data.
[0030] The storage unit 110 of this embodiment includes a learning data DB 111 and a standard medical term DB 112. The learning data DB 111 stores learning data corresponding to medical data acquired from the hospital server 30. The standard medical term DB 112 stores medical terms that have been pre-set as standard medical terms. Medical terms are terms used in the medical field, such as disease names, treatment names, and equipment names. For example, for disease names, Japanese disease names corresponding to ICD-10 are stored as standard medical terms (standard disease names). Note that the standard medical term DB 112 only needs to contain pre-registered medical terms and is not limited to Japanese disease names corresponding to ICD-10.
[0031] The user terminal 20 is composed of a computer and mainly comprises a communication unit 200, a control unit 210, a display unit 220, and an operation unit 230. The communication unit 200 is similar to the communication unit 120 of the model management device 10 and includes a communication interface for communicating with external devices wirelessly or via wired connection.
[0032] The control unit 210 is similar to the control unit 100 of the model management device 10, includes a processor, and controls the operation of the user terminal 20. The display unit 220 is, for example, a monitor, and displays various screens. The display unit 220 displays, for example, medical data, the results of sudden change predictions, etc. The operation unit 230 is, for example, a keyboard, and can give various commands to the control unit 210.
[0033] The hospital server 30 is composed of computers and the like. The configuration of the hospital server 30 is the same as that of the model management device 10, and mainly consists of a control unit, a storage unit, and a communication unit.
[0034] Next, the configuration of the control unit 100 of the model management device 10 will be described. The control unit 100 functions as an acquisition unit 101, a unification unit 102, a learning unit 103, and a prediction unit 104 by executing a program stored in the storage unit 110. In the following, the processes described as being executed by the acquisition unit 101, the unification unit 102, the learning unit 103, and the prediction unit 104 are processes performed by the control unit 100 by executing a program.
[0035] The acquisition unit 101 acquires medical data from each of the multiple hospital servers 30 via the communication unit 120. The unification unit 102 unifies the terminology included in the medical data to the standard medical terminology registered in the standard medical terminology DB 112. The unification unit 102 also unifies the format of each data included in the medical data. The learning unit 103 uses the medical data, which has been unified in terms and format by the unification unit 102, as training data to train the sudden change prediction model. The prediction unit 104 performs sudden change prediction by inputting the patient's medical data into the sudden change prediction model obtained through training by the learning unit 103. The processing of each functional unit will be described in detail later.
[0036] Figure 2 is a flowchart showing the model management process performed by the model management device 10. In the model management process, first, the acquisition unit 101 acquires medical data from each of the multiple hospital servers 30 (step S100). The medical data includes, for example, strings entered by user operations in the electronic medical record.
[0037] Figure 3 shows an example of the data structure of medical data. Medical data includes patient information, treatment information, and measurement results. Patient information includes basic patient information such as name, age, and blood type. Treatment information includes information about treatment, such as the use of a ventilator, oral administration, injection administration, and intravenous administration. Measurement results include measurement results, i.e., the patient's biological information, such as electrocardiogram waveforms, arterial pressure waveforms, and arterial blood oxygen saturation waveforms. Furthermore, the measurement results also include information on whether or not there was a sudden change in the patient's condition.
[0038] In Figure 2, after processing in step S100, the unification unit 102 unifies the medical terms included in the medical data to the standard medical terms registered in the standard medical terminology DB 112 (step S102). The medical data includes medical terms used in medicine, such as disease names, drug names, procedure names, and equipment names. However, for example, the term "atlas fracture" may be written as "atlas fracture" or as "C1 fracture." Similarly, for "radial fracture" and "ulnar fracture," both may be written, or they may be combined into a single term such as "radial-ulnar fracture."
[0039] For medical data to be used as training data for a sudden change prediction model, it is preferable that medical terms with the same meaning be registered under the same term. Therefore, in this embodiment, the unification unit 102 performs a process to unify each medical term included in the medical data into a standard medical term.
[0040] For example, "atlas fracture" is registered as a standard medical term in the Standard Medical Terminology Database 112. Therefore, "atlas fracture" included in medical data is stored as "atlas fracture" in the training data database 111. On the other hand, "C1 fracture" included in medical data is converted (unified) to its corresponding standard medical term, "atlas fracture."
[0041] After the processing in step S102, the unification unit 102 unifies the format of the numerical data included in the medical data to a predetermined standard format (step S104). This process will be described in detail later. Next, the unification unit 102 stores the medical data, which has been unified in terms of medical terminology and standard format through the above process, as training data in the training data DB 111 (step S106).
[0042] Next, the learning unit 103 generates a sudden change prediction model by using the learning data stored in the learning data DB 111 (step S108). Specifically, the learning unit 103 generates a sudden change prediction model by performing supervised machine learning using the learning data stored in the learning data DB 111. Next, the prediction unit 104 performs a sudden change prediction using the sudden change prediction model based on the medical data of the patient to be predicted (step S110). Thus, the model management process is completed.
[0043] Note that the execution order of the medical term unification process in step S102 and the medical data format unification process in step S104 is not limited to the embodiment. As another example, step S102 and step S104 may be performed in parallel, or the process of step S102 may be performed after the process of step S104.
[0044] FIG. 4 is a flowchart showing detailed processing in the medical term unification process (step S102). In the medical term unification process, a process of unifying medical terms is performed. Hereinafter, the medical term unification process will be described by taking a disease name as an example.
[0045] In the medical term unification process (step S102), the unification unit 102 first obtains the character string included in the medical data obtained in step S100 as a target character string to be processed, and as a preprocess, unifies the notations of Chinese characters and symbols (step S200). In the preprocess, for example, the notation of "颈" is unified to "頚".
[0046] Next, the unification unit 102 checks whether the target character string after the preprocess exactly matches the standard medical terms stored in the standard medical term DB 112 (step S202). If the target character string exactly matches the standard medical terms (Y in step S202), the unification unit 102 recognizes the character string as a standard medical term (step S204) and completes the process.
[0047] On the other hand, if the target string does not exactly match a standard medical term (N in step S202), the unification unit 102 performs a line splitting process (step S206). In the line splitting process (step S206), the unification unit 102 splits the target string into multiple target strings according to the line splitting rule. The line splitting rule splits a target string written across multiple lines into multiple target strings. According to the line splitting rule, for example, as shown in Figure 5, the unification unit 102 splits a string displayed across two lines into two strings, each of which becomes a target string.
[0048] In the example in Figure 5, the first line contains the text "radial fracture" and the second line contains the text "ulnar fracture". In this case, "radial fracture ulnar fracture" is the target string, but in the line splitting process (step S206), the target string is split into two: "radial fracture" and "ulnar fracture". Similarly, if the target string has three or more lines, it is split into three or more target strings. If this process results in a change, i.e., if the target string is split into two or more target strings (Y in step S208), the unification unit 102 proceeds to step S202 and performs the processing from step S202 onwards for each target string.
[0049] Furthermore, if there are no changes (N in step S208), the unification unit 102 proceeds to step S210. In step S210, the unification unit 102 performs numerical expansion processing according to the numerical expansion rules. The numerical expansion rules rewrite notations that include multiple contents, such as notations indicating a numerical range, into notations for each individual content. For example, an atlas fracture may be written as a C1 fracture. In addition, there may be notations that group multiple bones together, such as "C1~3 fracture". In response to this, the unification unit 102 expands the notation grouped as "C1~3 fracture" into multiple notations corresponding to multiple disease names, such as "C1,2,3 fracture," according to the numerical expansion rules. Similarly, "TH2~7 fracture" is expanded to "TH2,3,4,5,6,7 fracture."
[0050] If a change occurs during the numerical expansion process, i.e., if numerical expansion is performed (Y in step S212), the unification unit 102 proceeds to step S202 and performs the subsequent processing using the numerically expanded string as the target string. If there is no change (N in step S212), the unification unit 102 proceeds to step S214. In step S214, the unification unit 102 performs string conversion processing according to the string conversion rules.
[0051] The string conversion rules convert predetermined strings into standard medical terms that are pre-associated with these strings. For example, if any of the characters "vertebral arch," "spinous process," "vertebral body," or "fracture" are detected after "C7," the string conversion rules will convert them to "7th cervical vertebral arch," "7th cervical vertebral spinous process," "7th cervical vertebral body," and "7th cervical vertebral fracture," respectively. In addition, the string conversion process removes predetermined characters from within parentheses. Furthermore, expressions that modify a disease name after it, such as "2nd degree burn," are changed to modify it from the beginning, such as "2nd degree burn." If a change occurs in the string conversion process (step S214), i.e., if a string conversion is performed (Y in step S216), the unification unit 102 proceeds to step S202 and performs the subsequent processing using the converted string as the target string.
[0052] Furthermore, if there are no changes (N in step S216), the unification unit 102 proceeds to step S218. In step S218, the unification unit 102 performs partial matching processing according to the partial matching rule. The partial matching rule detects whether there is a partial match by comparing a part of the target string with the standard medical terms in the standard medical term DB 112. In the partial matching processing (step S218), the unification unit 102 further considers the target string to be a partial match if it matches a standard medical term when a word such as "disease," "injury," "symptom," or "syndrome" is added to the target string according to the partial matching rule.
[0053] Then, when there is a partial match (Y in step S218), the unification unit 102 determines that the matched character string is recognized as a standard medical term (step S220) and completes the process. For example, when "depressive disorder" is simply described as "depression", it is converted to "depressive disorder" and recognized as a standard medical term. Further, when it is described as "depression", in the previous preprocessing (step S200), "depression" is converted to "depression", and in step S220, it is converted to "depressive disorder" and then recognized as a standard medical term.
[0054] Also, in the partial match process (step S218), when the unification unit 102 contains predetermined characters such as "right", "left", "both sides", "multiple", "worsening", "acute", etc. at the beginning of the target character string, these are deleted. If the character string after deletion matches the standard medical term, the target character string is regarded as a partial match. For example, when "right femoral fracture" is the target character string, "right" is deleted and it becomes "femoral fracture". Since "femoral fracture" matches the standard medical term, it is recognized as a standard medical term. Also, when "bilateral rib fractures" is the target character string, "both sides" is deleted and it becomes "rib fractures". Since "rib fractures" matches the standard medical term, it is recognized as a standard medical term. Also, for example, when "acute exacerbation of pulmonary edema" is the target character string, "acute" and "exacerbation" are deleted and it becomes "pulmonary edema". And since "pulmonary edema" matches the standard medical term, it is recognized as a standard medical term.
[0055] On the other hand, there may be cases where a character string containing the character "multiple", such as "multiple myeloma", is registered as a standard medical term. For such character strings, in the process of confirming whether there is an exact match (step S202) immediately after the preprocessing (step S200), they are recognized as matching the standard medical term. Therefore, it is possible to prevent the necessary "multiple" from being deleted by the partial match process.
[0056] Furthermore, partial matching is performed after string conversion. For example, in string conversion, "burn 2nd degree" is converted to "2nd degree burn" (prefixing of post-modifier). If partial matching were performed before this, "burn" would match the standard medical term, and a partial match would be determined for "burn," resulting in the loss of the "2nd degree" information. In contrast, by performing string conversion before partial matching, such loss of information can be prevented.
[0057] Furthermore, in partial matching processing, a word is replaced with a word representing a broader concept, and then the presence or absence of a partial match is checked again. For example, "ascending colon" is a part of "colon," and "colon" is a part of "large intestine." That is, "colon" encompasses "ascending colon," and "large intestine" encompasses "colon," indicating an inclusion relationship. On the other hand, if the target string is "ascending colon cancer," but "ascending colon cancer" is not registered as a standard medical term, it will be aggregated into "large intestine cancer." In such cases, partial matching processing is performed on the target string "ascending colon cancer," and since there is no match, "ascending colon" is converted to the broader concept of "colon," which encompasses "ascending colon," and partial matching processing is performed again on "colon cancer." In this case as well, if there is no match, partial matching processing is performed again on "large intestine cancer." Finally, a match is determined.
[0058] For example, if the string conversion process uniformly converts "ascending colon" to "large intestine," then all standard medical terms that include "ascending colon" will also be converted to "large intestine," which is undesirable. Therefore, in this embodiment, the partial matching process is configured to sequentially convert to the medical terms that are included. This ensures accurate conversion to standard medical terms.
[0059] If no partial match is found during the partial match processing (N in step S218), the unification unit 102 proceeds to step S222. In step S222, the unification unit 102 performs the expansion process. In the expansion process (step S222), the unification unit 102 divides the combined disease names into individual disease names.
[0060] As a result, for example, the description "radial and ulnar fracture" is expanded to "radial fracture" and "ulnar fracture." Similarly, the description "distal radius shaft fracture" is expanded to "distal radius fracture" and "radial shaft fracture." Furthermore, the description "C1, C2 fracture" is expanded to "C1 fracture" and "C2 fracture." Note that "C1 fracture" is converted to "atlas fracture" in the string conversion process (step S214) that is executed after the processing of steps S202 to S228 is repeated, as described below. Similarly, "C2 fracture" is converted to "axis fracture" in the string conversion process (step S214).
[0061] Then, if there is a change during the expansion process (step S222), that is, if expansion has been performed (Y in step S224), the unification unit 102 proceeds to step S202 and performs the processing from step S202 onward using the expanded string as the target string. In this way, the string conversion process is applied after the expansion process. This prevents, for example, "C1,2 fracture" from being converted to only "atlas fracture" or only "axis fracture".
[0062] The expansion process is performed after the partial matching process. For example, suppose the target string is "back first-degree burn". This string is registered as a standard medical term. On the other hand, its decomposition, "back burn" and "first-degree burn", are also registered as standard medical terms. Therefore, if the expansion process is performed before the partial matching process, "back first-degree burn" will be decomposed into "back burn" and "first-degree burn" against the author's intent. In contrast, in this embodiment, the expansion process is performed after the partial matching process, so it is possible to prevent such an expansion from occurring against the author's intent.
[0063] Furthermore, the expansion process is performed after the string conversion process. For example, suppose the target string is "Second-degree burn (TBSA 48%)". In this case, during the string conversion process, "TBSA 48%" in the target string is converted to "of 40-49% of body surface area", and the target string is converted to "Second-degree burn of 40-49% of body surface area". Then, during the expansion process, it is converted to "Second-degree burn, burn of 40-49% of body surface area". On the other hand, if the expansion process is performed before the string conversion process, it will be split into "Second-degree burn" and "TBSA 48%", resulting in an interpretation contrary to the author's intent.
[0064] If there are no changes in step S224 (N in step S224), the unification unit 102 proceeds to step S226. In step S226, the unification unit 102 performs extraction processing. In the extraction processing (step S226), the unification unit 102 searches for parentheses, etc., in the target string and extracts the text if the characters inside the parentheses match a disease name. For example, for the entry "Other (subarachnoid hemorrhage)", the characters inside the parentheses are removed so that it becomes "Subarachnoid hemorrhage Other". In this case, the "Subarachnoid hemorrhage" part is determined to be a partial match in the subsequent partial match processing (step S218).
[0065] Then, if there is a change in the extraction process (step S226), that is, if extraction has been performed (Y in step S228), the unification unit 102 proceeds to step S202 and performs the processing from step S202 onward using the extracted string as the target string.
[0066] Furthermore, if there are no changes (N in step S228), the unification unit 102 proceeds to step S230. In step S230, the unification unit 102 performs a combination search. In the combination search process (step S230), the unification unit 102 attempts a partial match by sequentially delimiting the target string character by character.
[0067] For example, if the target string is "utsu kyūkyū dakinoku" (depression acute drug poisoning), the unification unit 102 splits it into "u" and "tsu kyūkyū dakinoku". The unification unit 102 treats each of the two resulting strings as target strings and determines whether each matches a standard medical term. If neither string matches a standard medical term, the unification unit 102 changes the delimiter position. This splits it into "utsu" and "kyūkyū dakinoku". Then, the unification unit 102 treats each string as a target string and determines whether each matches a standard medical term. In this way, the unification unit 102 repeats the process, changing the delimiter position one character at a time, until it matches a standard medical term. If, in the combination search, some of the strings match a standard medical term (Y in step S232), the unification unit 102 repeats the process from step S202 for the remaining strings. If there are no matches in the combination search, the unification unit 102 completes the medical term unification process.
[0068] As described above, in the medical terminology standardization process, the standardization unit 102 first determines whether the target string is an exact match to a standard term. Then, the standardization unit 102 applies multiple conversion rules, such as line splitting, numerical expansion, word conversion, partial matching, expansion, extraction, and combination search, in a predetermined order of application. In this way, the standardization unit 102 converts the target string appropriately according to each conversion rule and determines again each time whether it matches a standard medical term. This ensures that the target string is correctly converted to the corresponding standard medical term.
[0069] In the above, the transformation rules for combined searches are applied after the rules for partial matches. Therefore, it is possible to prevent unnecessary combined searches from being applied to medical terms that partially match.
[0070] Next, the standard format unification process (step S104), which was explained with reference to Figure 2, will be described. The standard format unification process is a process that unifies the data format of each data contained in the medical data into a predetermined standard format. The unification unit 102 unifies each data contained in the medical data into a predetermined standard format.
[0071] The reference format includes a static format that does not include time, and a time-series format that includes time. Patient information is static information that does not include time, and is recorded in the training data DB111 as training data in static format. Treatment information and measurement results are information that corresponds to time, i.e., dynamic information that includes time, and are recorded in the training data DB111 as training data in time-series format.
[0072] Furthermore, time series formats are classified into point time series formats and interval time series formats. Point time series formats are data formats for evaluation that are managed based on predetermined points in time. Information in point time series formats includes information on oral medication administered by patients, information on the administration of injectable drugs, etc. Here, information on oral medication refers to information that a certain drug was administered orally at a certain time. Information on the administration of injectable drugs refers to information that a certain drug was administered in injection form at a certain time. Other examples of information in point time series formats include vital signs (pulse, blood pressure, respiratory rate, etc.), blood tests (white blood cell count, creatinine, urea nitrogen, etc.), blood gas analysis (pH, arterial oxygen partial pressure, lactate level, etc.), etc.
[0073] Furthermore, with regard to the administration of injectable drugs, the time required for administration may or may not be controlled. For example, with anesthetics, the time may be controlled to the extent of slow administration, while with antibiotics, the administration time may be precisely controlled, such as administering them over one hour at 6:00, 14:00, and 22:00 each day. In this case, for anesthetics, a point time series format is preferable, and for antibiotics, a segment time series format is preferable. Thus, with regard to the administration of injectable drugs, the time series format, whether a point time series format or a segment time series format, is predetermined depending on the type of drug administered and the method of administration.
[0074] Interval time series format is a data format for evaluation that is managed based on a predetermined interval. Information in interval time series format includes information on administration by intravenous infusion, information on procedures for which a period has been specified by healthcare professionals, etc. Information on administration by intravenous infusion indicates that administration by intravenous infusion was continuously performed over a certain interval (the period from the first time point to the second time point). Information on procedures for which a period has been specified indicates that a ventilator was used during a certain interval. Procedures include the use of medical devices such as ventilators. Examples of such medical devices include IABP (Intra Aortic Balloon Pumping) and ECMO (Extracorporeal Memberane Oxygenation). Other examples of information in interval time-series format include intravenous fluid information (physiological saline, norepinephrine, aminoglycosides, etc.), ventilator settings (inspired oxygen concentration, pressure support, positive end-expiratory pressure, etc.), and non-invasive positive pressure ventilation (inspired oxygen concentration, pressure support, positive end-expiratory pressure, etc.). Thus, the evaluation target information that indicates each evaluation target related to the patient included in the medical data is standardized to a predetermined standard format for each type of evaluation target information.
[0075] In the standard format unification process (S104), the unification unit 102, for example, if numerical information for a ventilator is given in point time series format, rewrites it into interval time series format. On the other hand, the unification unit 102 manages information regarding oral administration and injection in point time series format.
[0076] For example, regarding drug administration, one medical institution's information system records the amount of drug administered every hour, while another medical institution's information processing system records the start and end times of drug administration and the rate of administration. However, in drug administration, it is preferable to record information that allows healthcare professionals to reproduce similar procedures, and such information is data for a specific period, such as what drugs are administered and over what period. Therefore, this information regarding drug administration is managed in a time-series format. On the other hand, information such as oral administration, injection administration, blood pressure measurements, and vital signs requires data at a specific point in time, and therefore, this information is managed in a point-time series format.
[0077] For example, the data sets shown in Figures 6 and 7 are treatment information related to intravenous drug administration, included in medical data transmitted from Hospital A and Hospital B, respectively. The treatment information from Hospital A shown in Figure 6 includes patient ID, time, order ID, bottle ID, status, drug, and dosage (ml), with the dosage recorded on a flow rate basis. Patient ID is the patient's identification information. Order ID is the order's identification information. Bottle ID is the drug bottle's identification information. Status is the status of the treatment.
[0078] On the other hand, the treatment information for Hospital B shown in Figure 7 includes patient ID, patient name, order ID, time, administration type, bottle ID, order type, drug ID, drug, solution volume, drug amount, and number of units. Only the administration amount is recorded, and data on flow rate is not included. Administration type is the type of administration. Order type is the type of procedure. Drug ID is the identification information for the drug. The two data sets also differ in the data they include, such as whether or not the patient's name and administration type are included. Furthermore, the treatment information for Hospitals A and B is associated with time, and the drug amount and other information at that time are recorded, resulting in point-time series information. Therefore, the unification unit 102 unifies all of this treatment information into interval time series format.
[0079] Figure 8 shows an example of the data structure of treatment information related to intravenous drug administration, recorded in interval time series format. In this embodiment, the unification unit 102 unifies the numerical data included in the medical data acquired by the acquisition unit 101, specifically the treatment information related to intravenous drug administration, into interval time series format.
[0080] In the interval time-series data structure, treatment information regarding intravenous drug administration includes administration ID, patient ID, start time, end time, administration rate, whether it is a bolus administration or not, drug name, and drug ratio. The administration ID is the identification information for the administration. The start time and end time are the start and end times of the procedure, respectively. Thus, in the interval time-series data structure, information such as the administration rate is managed for intravenous drug administration within the interval from the start time to the end time. In this way, in the interval time-series format, information regarding the evaluation target, intravenous drug administration, is recorded on an interval basis. The unification unit 102 unifies the treatment information regarding intravenous drug administration into the interval time-series data structure shown in Figure 8.
[0081] The data sets shown in Figures 9 and 10 represent treatment information regarding the use of ventilators, included in medical data transmitted from Hospital A and Hospital B, respectively. The treatment information from Hospital A in Figure 9 includes patient ID, time, instruction ID, status, item, and value. The instruction ID is the identifying information for the treatment instruction. The status indicates the device being used. The item indicates the setting item, and the value indicates the value for that setting item.
[0082] The treatment information for Hospital B shown in Figure 10 includes patient ID, time, item name, and value. In all treatment information, information such as ventilation mode and oxygen concentration is managed for the interval from start time to end time.
[0083] Figure 11 shows an example of the data structure of treatment information related to ventilator use, recorded in interval time series format. In this embodiment, the unification unit 102 unifies the numerical data included in the medical data acquired by the acquisition unit 101, specifically the treatment information related to ventilator use, into interval time series format.
[0084] In the interval time-series data structure, treatment information related to ventilator use includes ventilator ID, patient ID, start time, end time, ventilation mode, oxygen concentration, positive end-expiratory pressure, inspiratory pressure, tidal volume, inspiratory time, respiratory rate, and inspiratory support pressure. Thus, the unification unit 102 unifies the treatment information related to ventilator use into the interval time-series data structure shown in Figure 12.
[0085] As described above, in the information processing system 1 according to this embodiment, it is possible to unify strings to standard medical terms by determining whether a string matches a pre-registered standard medical term, and if it does not match, by applying conversion rules sequentially to the string.
[0086] The embodiments described above are merely examples for carrying out the present invention, and various other embodiments can be adopted. For example, various modifications and changes are possible within the scope of the gist of the present invention as described in the claims, such as applying one modification to another. For example, some of the components of the above embodiments may be omitted, or the order of processing may be changed or omitted.
[0087] As a variation of this embodiment, the processing performed by the model management device 10 may be implemented by multiple devices. That is, some functions of the model management device 10 may be implemented by a first device, and other functions of the model management device 10 may be implemented by a second device.
[0088] For example, the first device may perform processing up to the generation of a sudden change prediction model, and the second device may acquire the sudden change prediction model and perform sudden change prediction using the sudden change prediction model. Another example is that the model management device 10 may be provided integrally with a single hospital server 30. Furthermore, in this case, the medical data managed by the hospital server 30 may be used as training data to generate the sudden change prediction model.
[0089] According to the program, information processing system, and information processing method of this embodiment, which has the above configuration, the acquisition unit 101 acquires a string, the unification unit 102 determines whether the string matches a pre-registered standard medical term, and if the string does not match a standard medical term, it applies a plurality of string conversion rules, whose application order is predetermined, to the string in order, determines whether the converted string obtained by applying the conversion rules matches a standard medical term, and if it matches a standard medical term, unifies the string to a standard medical term. This makes it possible to unify strings that indicate the same content into a predetermined term.
[0090] Furthermore, according to the program, information processing system, and information processing method of this embodiment, if a string is changed by applying a conversion rule, the unification unit may apply the conversion rule to the changed string in the order of application. This makes it possible to unify terminology accurately.
[0091] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the learning unit 103 may learn a predictive model to predict the patient's condition using the training data containing the unified string.
[0092] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the multiple conversion rules may include expansion rules that expand a combined notation of multiple disease names into multiple notations corresponding to those disease names. This makes it possible to convert multiple disease names into individual disease names.
[0093] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the collective notation of multiple disease names may include numerical notation. This makes it possible to convert the collective notation including numerical notation into individual disease names.
[0094] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the multiple conversion rules may include string conversion rules that convert predetermined strings into standard medical terms. This makes it possible to unify strings into predetermined terms.
[0095] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the multiple conversion rules may include a partial matching rule that deletes a predetermined string. This makes it possible to unify the terminology into a predetermined term that does not include unnecessary terms.
[0096] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the multiple conversion rules include an expansion rule that expands a combined notation of multiple disease names into multiple disease names, and a partial matching rule that compares a part of a string with standard medical terms, and the partial matching rule may be applied before the expansion rule. This makes it possible to unify strings more accurately.
[0097] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the multiple conversion rules include string conversion rules that convert a predetermined string into a standard medical term, and partial matching rules that compare a part of the string with a standard medical term, and the string conversion rules may be applied before the partial matching rules. This makes it possible to unify strings more accurately.
[0098] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the multiple conversion rules include an expansion rule that expands a combined notation of multiple disease names into multiple disease names, and a string conversion rule that converts a predetermined string into the standard medical term, and the string conversion rule may be applied before the expansion rule. This makes it possible to unify strings more accurately.
[0099] Furthermore, according to the program, information processing system, and information processing method of this embodiment, the multiple conversion rules may include a conversion rule that converts a first medical term into a second medical term that encompasses the first medical term. This allows for more accurate string unification.
[0100] 1. Information Processing System 10 Model Management Devices 20 User Terminals 30 Hospital Servers 100 Control Unit 101 Acquisition Department 102 Ministry of Unification 103 Learning Department 104 Prediction Section 110 Storage section 111 Training Data Database 112 Standard Medical Terminology Database 120 Communications Department 200 Communications Department 210 Control Unit 220 Display section
Claims
1. A program for causing a computer to function as an acquisition unit and a unification unit, The acquisition unit acquires a string, The aforementioned unification section is, Determine whether the aforementioned string matches a pre-registered standard medical term. If the string does not match the standard medical term, Multiple string conversion rules, whose application order is predetermined, are applied sequentially to the string. A program that, each time the conversion rule is applied, determines whether the resulting converted string matches the standard medical term, and if it matches the standard medical term, unifies the string to the standard medical term.
2. The aforementioned unification section is, The program according to claim 1, wherein if the string is changed by applying the conversion rule, the conversion rule is applied to the changed string in the order of application.
3. The aforementioned program further enables the computer to function as a learning unit, The program according to claim 1, wherein the learning unit learns a predictive model to predict the patient's condition using learning data that includes the unified string.
4. The aforementioned program further enables the computer to function as a learning unit, The program according to claim 1, wherein the learning unit learns a predictive model to predict sudden changes in a patient's condition using learning data that includes the unified string.
5. The program according to claim 1, wherein the plurality of conversion rules include an expansion rule that expands a notation that combines multiple disease names into a plurality of notations corresponding to the plurality of disease names.
6. The program according to claim 5, wherein the notation for grouping the aforementioned multiple disease names includes numerical notation.
7. The program according to claim 1, wherein the plurality of conversion rules include string conversion rules that convert predetermined strings into standard medical terms.
8. The program according to claim 1, wherein the plurality of conversion rules include a partial match rule that deletes a predetermined string.
9. The above multiple conversion rules are: A rule for expanding a collective notation of multiple disease names into multiple disease names, A partial matching rule that compares a portion of the aforementioned string with the aforementioned standard medical term, Includes, The program according to claim 1, wherein the partial matching rule is applied before the expansion rule.
10. The above multiple conversion rules are: A string conversion rule that converts a predetermined string into the aforementioned standard medical term, A partial matching rule that compares a portion of the aforementioned string with the aforementioned standard medical term, Includes, The program according to claim 1, wherein the string conversion rule is applied before the partial matching rule.
11. The above multiple conversion rules are: A rule for expanding a collective notation of multiple disease names into multiple disease names, A string conversion rule that converts a predetermined string into the aforementioned standard medical term, Includes, The program according to claim 1, wherein the string conversion rule is applied before the expansion rule.
12. The above multiple conversion rules are: The program according to claim 1, comprising a conversion rule that converts a first medical term into a second medical term that encompasses the first medical term.
13. An information processing system comprising an acquisition unit and a unification unit, The acquisition unit acquires a string, The aforementioned unification section is, Determine whether the aforementioned string matches a pre-registered standard medical term. If the string does not match the standard medical term, Multiple string conversion rules, whose application order is predetermined, are applied sequentially to the string. An information processing system that determines whether the converted string obtained by applying the conversion rule matches the standard medical term, and if it matches the standard medical term, unifies the string to the standard medical term.
14. An information processing method performed by a computer having a control unit, The control unit performs the steps of obtaining a string, The control unit determines whether the string matches a pre-registered standard medical term, and if the string does not match a standard medical term, it applies a plurality of string conversion rules, whose application order is predetermined, to the string in order, determines whether the converted string obtained by applying the conversion rules matches a standard medical term, and if it matches a standard medical term, it unifies the string to the standard medical term. Information processing methods, including those mentioned above.