Text data analysis system, text data analysis method, and computer program
By designing a text data analysis system, using technical means such as data collection, string extraction and multiple searches, the problem of difficulty in extracting medical procedures, drugs and disease names from text data in the prior art is solved, and efficient and accurate extraction effects are achieved without adding main database keywords.
Patent Information
- Application Number
- JP2021117846
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-16
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-07-16
AI Technical Summary
The prior art is difficult to extract strings representing medical procedures, drugs, and disease names from textual data without the need to add more keywords to the main database, resulting in increased data capacity and inefficient extraction.
A text data analysis system was designed, which can extract strings representing medical procedures, drugs and disease names from text data through steps such as data collection, string extraction, first search, prefix and suffix processing, secondary search and similar string extraction without adding more keywords to the main database.
The function of extracting strings representing medical procedures, drugs and disease names from text data is realized, which improves extraction efficiency, reduces the increased demand for the main database keywords, and improves the accuracy and efficiency of data processing.
Smart Images

Figure 0007674176000001 
Figure 0007674176000002 
Figure 0007674176000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a text data analysis system, a text data analysis method, and a computer program. [Background technology]
[0002] Conventionally, a document reading device that reads a document using an optical character reader is known (Patent Document 1). This technology has an error recognition database that stores an error recognition character string including an error recognition character and a correction character string for correcting the error recognition character in correspondence with each other, and searches the error recognition database for a character string in the read data obtained by reading a document using the optical character reader, and creates correction data by converting an error recognition character string into a corresponding correction character string in the case of an error recognition character string. In addition, the success rate of error recognition is increased by adding characters that were not correctly corrected to the error recognition database. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 3349699 Summary of the Invention [Problem to be solved by the invention]
[0004] Incidentally, when searching for a string that matches a specified keyword in a master that stores strings representing medical procedures, medicines, and names of injuries and illnesses, if the string listed in the medical bill or other document is not identical to the string stored in the master, or if the string listed in the medical bill or other document cannot be accurately converted into text data, it is not possible to extract the string representing the medical procedure or other document.
[0005] In response to this, there has been a technique in the past that, when a character string that matches a keyword is not stored in the master, increases the success rate of a search by adding the keyword to the master. However, this technique has problems in that it is not possible to extract a character string before a new keyword is added to the master, and the data capacity of the master increases as new keywords are added to the master.
[0006] The present invention has been made in consideration of the above background, and aims to provide a text data analysis system, a text data analysis method, and a computer program that can extract character strings representing medical procedures, medicines, and names of injuries and illnesses from text data without adding more keywords than necessary to the master. [Means for solving the problem]
[0007] A text data analysis system for achieving the above object includes a data acquisition means for acquiring text data, a character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means, a first search means for searching a medical practice / medicine master storing character strings representing medical practices or medicines, using the character string extracted by the character string extraction means as a first keyword, to see if a character string matching the first keyword is found, and if a character string matching the first keyword is found as a result of the search, a first search means for outputting information on the medical practice or medicine corresponding to the hit character string, a first character string generation means for generating a character string by removing the prefix from the first keyword, when a character string matching the first keyword is not found as a result of the search by the first search means, and a second search means for searching the medical practice / medicine master data set for determining whether a character string matching the second keyword is found, using a character string generated by the search means as a second keyword, and for outputting information on the medical practice or medicine corresponding to the hit character string when a character string matching the second keyword is found as a result of the search; a second character string generation means for generating a character string obtained by removing at least parentheses and the character string enclosed by the parentheses from the second keyword when a character string matching the second keyword is not found as a result of the search by the second search means; a similar character string extraction means for extracting a character string including the third keyword from the medical practice / medicine master data set, using the character string generated by the second character string generation means as a third keyword; and an output means for outputting information on at least one medical practice or medicine corresponding to the character string extracted by the similar character string extraction means.
[0008] According to such a system, character strings representing medical procedures or medicines can be extracted from text data without adding more keywords than necessary to the master.
[0009] In addition, the text data analysis system may further include a first similarity calculation means for calculating a similarity between each of the character strings extracted by the similar character string extraction means and the second keyword when there are a plurality of character strings extracted by the similar character string extraction means, and the output means may be configured to output information on a medical procedure or medicine corresponding to a character string among the character strings extracted by the similar character string extraction means, the similarity calculated by the first similarity calculation means being equal to or greater than a predetermined value.
[0010] This allows for the output of narrowed-down information on medical procedures and medicines.
[0011] In addition, the medical procedure / medicine master stores character strings representing one or more medical procedures or medicines in correspondence with index strings which are character strings included in the character strings representing the one or more medical procedures or medicines, and the text data analysis system may further include a second similarity calculation means for calculating a similarity between the third keyword and the index string if the similar string extraction means is unable to extract a string including the third keyword from the medical procedure / medicine master, and the similar string extraction means may be configured to extract a string corresponding to the index string from the medical procedure / medicine master if there is an index string whose similarity calculated by the second similarity calculation means is equal to or greater than a predetermined value.
[0012] This makes it possible to more reliably extract character strings representing medical procedures or medicines from the text data.
[0013] Furthermore, the text data analysis system may further include an element string extraction means for, when there is no index string for which the similarity calculated by the second similarity calculation means is equal to or higher than a predetermined value, extracting at least one element string from the second keyword by referring to an element string master that stores element strings that are specific strings included in strings representing medical procedures or medicines, and the similar string extraction means may be configured to extract a string including the fourth keyword from the medical procedure / drug master by using the element string extracted by the element string extraction means as a fourth keyword.
[0014] This makes it possible to more reliably extract character strings representing medical procedures or medicines from the text data.
[0015] In addition, the text data analysis system may further include a third similarity calculation means for calculating, when there are a plurality of element strings extracted by the element string extraction means, a similarity between the element string located near the center of the second keyword and the string extracted by the similar string extraction means, using the element string located near the center of the second keyword as a fourth keyword, before calculating a similarity between the element string and the second keyword for other element strings, and the output means may be configured to output information on a medical procedure or medicine corresponding to the string, when a string having a similarity equal to or higher than a predetermined value is included among the strings for which the third similarity calculation means has previously calculated a similarity.
[0016] This makes it possible to reduce the amount of processing required to output information on medical procedures or medicines, thereby increasing the processing speed.
[0017] The second character string generating means may be configured to generate a character string by further removing at least one of the following character strings (1) to (5) from the second keyword: (1) Leading or trailing spaces (2) Any spaces in the middle and the characters following the spaces (3) Bullet (4) Commas (5) A number and a character string immediately following the number that indicates the unit
[0018] This makes it easier to narrow down information on medical procedures and medicines.
[0019] A text data analysis system for achieving the above object includes a data acquisition means for acquiring text data, a character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means, a first search means for searching a disease name master storing character strings representing names of injuries or illnesses, using the character string extracted by the character string extraction means as a first keyword, to see if a character string matching the first keyword is found, and if a character string matching the first keyword is found as a result of the search, a first search means for outputting information on the disease name corresponding to the found character string, a first character string generation means for generating a character string by removing the suffix from the first keyword, by referring to a suffix master storing suffixes which are specific character strings added to the end of a character string, if a character string matching the first keyword is not found as a result of the search by the first search means, and a second search means for searching the injury / illness name master as a key word to see whether a character string matching the second keyword is found, and when a character string matching the second keyword is found as a result of the search, outputting information on the injury / illness name corresponding to the hit character string; a second character string generation means for generating a character string from the second keyword by removing the prefix, by referring to a prefix master which stores a prefix which is a specific character string added to the beginning of a character string, when a character string matching the second keyword is not found as a result of the search by the second search means; and a third search means for searching the injury / illness name master as a third keyword to see whether a character string matching the third keyword is found, and when a character string matching the third keyword is found as a result of the search, outputting information on the injury / illness name corresponding to the hit character string.
[0020] According to such a system, it is possible to extract character strings representing names of illnesses and injuries from text data without adding more keywords than necessary to the master.
[0021] Furthermore, the text data analysis system can be further configured to include: element string extraction means for, if a search by the third search means does not hit a string matching the third keyword, extracting at least one element string from the third keyword by referring to an element string master that stores element strings that are specific strings contained in an injury or illness name; similar string extraction means for extracting strings including the fourth keyword from the injury or illness name master by using the element string extracted by the element string extraction means as a fourth keyword; and output means for outputting information on at least one injury or illness name corresponding to the string extracted by the similar string extraction means.
[0022] This makes it possible to more reliably extract character strings representing the names of injuries and illnesses from the text data.
[0023] In addition, the text data analysis system may further include a similarity calculation means for calculating a similarity between a character string extracted by the similar character string extraction means and the third keyword, and the output means may be configured to output information on an injury or illness name corresponding to a character string among the character strings extracted by the similar character string extraction means, the similarity calculated by the similarity calculation means being equal to or greater than a predetermined value.
[0024] This allows the output of information on the names of injuries and illnesses to be narrowed down.
[0025] Furthermore, when there are a plurality of element strings extracted by the element string extraction means, the similarity calculation means calculates the similarity between the element string located close to the center of the third keyword and the third keyword for the strings extracted by the similar string extraction means as a fourth keyword before calculating the similarity between the element string located close to the center of the third keyword and the third keyword for the strings extracted by the similar string extraction means, and the output means can be configured to output information on the name of the injury or illness corresponding to the string if the string includes a string whose similarity is equal to or higher than a predetermined value among the strings whose similarity has been calculated earlier.
[0026] This makes it possible to reduce the amount of processing required before the information on the name of the injury or illness is output, thereby increasing the processing speed.
[0027] Further, a text data analysis method for achieving the above-mentioned object includes a means provided in a computer, the means including a data acquisition step of acquiring text data, a string extraction step of extracting a group of strings representing one item from the text data acquired in the data acquisition step, a first search step of using the string extracted in the string extraction step as a first keyword to search a medical practice / medicine master storing strings representing medical practices or medicines for a string matching the first keyword, and if a string matching the first keyword is found as a result of the search, a first search step of outputting information on the medical practice or medicine corresponding to the hit string, and a first string generation step of, if a string matching the first keyword is not found as a result of the search in the first search step, generating a string by removing the prefix from the first keyword by referring to a prefix master storing a prefix which is a specific string added to the beginning of a string, and The method includes executing a second search step of searching the medical practice / medicine master data set for a character string generated in the first string generation step as a second keyword to see if a character string matching the second keyword is found, and outputting information on the medical practice or medicine corresponding to the hit character string when a character string matching the second keyword is found as a result of the search; a second string generation step of generating a character string obtained by removing at least parentheses and the character string enclosed by the parentheses from the second keyword when a character string matching the second keyword is not found as a result of the search in the second search step; a similar string extraction step of extracting a character string including the third keyword from the medical practice / medicine master data set for the character string generated in the second string generation step as a third keyword; and an output step of outputting information on at least one medical practice or medicine corresponding to the character string extracted in the similar string extraction step.
[0028] According to this method, character strings representing medical procedures or medicines can be extracted from text data without adding more keywords than necessary to the master.
[0029] Further, a text data analysis method for achieving the above-mentioned object includes a means provided in a computer, the means including a data acquisition step of acquiring text data, a string extraction step of extracting a group of strings representing one item from the text data acquired in the data acquisition step, a first search step of searching a disease name master storing strings representing names of injuries and illnesses for a string matching the first keyword using the string extracted in the string extraction step as a first keyword, and outputting information on the disease name corresponding to the hit string when a string matching the first keyword is found as a result of the search in the first search step, a first string generation step of generating a string by removing the suffix from the first keyword by referring to a suffix master storing a suffix that is a specific string added to the end of a string, and a second string generation step of generating a string by removing the suffix from the first keyword when a string matching the first keyword is not found as a result of the search in the first search step, and a second search step of searching the injury / illness name master using a character string as a second keyword to see if a character string matching the second keyword is found, and if a character string matching the second keyword is found as a result of the search, outputting information on the injury / illness name corresponding to the hit character string; a second string generation step of generating a character string from the second keyword by removing the prefix, by referring to a prefix master that stores a prefix, which is a specific character string attached to the beginning of a character string, if a character string matching the second keyword is not found as a result of the search in the second search step; and a third search step of searching the injury / illness name master using the character string generated in the second string generation step as a third keyword to see if a character string matching the third keyword is found, and if a character string matching the third keyword is found as a result of the search, outputting information on the injury / illness name corresponding to the hit character string.
[0030] According to this method, it is possible to extract character strings representing names of illnesses and injuries from text data without adding more keywords than necessary to the master.
[0031] Further, a computer program for achieving the above object includes a computer including: data acquisition means for acquiring text data; character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; first search means for searching a medical practice / medicine master storing character strings representing medical practices or medicines using the character string extracted by the character string extraction means as a first keyword to see if a character string matching the first keyword is found, and when a character string matching the first keyword is found as a result of the search, first search means outputs information on the medical practice or medicine corresponding to the hit character string; first character string generation means for generating a character string by removing the prefix from the first keyword, when a character string matching the first keyword is not found as a result of the search by the first search means, by referring to a prefix master storing a prefix which is a specific character string added to the beginning of a character string; The medical practice / medicine master database includes a second search means for searching the medical practice / medicine master database for a string matching the second keyword using the string generated by the string generation means as a second keyword, and for outputting information on the medical practice or medicine corresponding to the hit string when the search results in a string matching the second keyword being hit; a second string generation means for generating a string from the second keyword by removing at least parentheses and the string enclosed by the parentheses when the search results in no string matching the second keyword being hit; a similar string extraction means for extracting a string including the third keyword from the medical practice / medicine master database using the string generated by the second string generation means as a third keyword; and an output means for outputting information on at least one medical practice or medicine corresponding to the string extracted by the similar string extraction means.
[0032] According to such a program, character strings representing medical procedures or medicines can be extracted from text data without adding more keywords than necessary to the master.
[0033] Further, a computer program for achieving the above object includes a computer including: data acquisition means for acquiring text data; character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; a first search means for searching, using the character string extracted by the character string extraction means as a first keyword, a disease name master storing character strings representing names of injuries and illnesses, whether a character string matching the first keyword is found, and, if a character string matching the first keyword is found as a result of the search, a first search means for outputting information on the disease name corresponding to the found character string; a first character string generation means for, if a character string matching the first keyword is not found as a result of the search by the first search means, referring to a suffix master storing suffixes which are specific character strings added to the end of a character string, generating a character string by removing the suffix from the first keyword; and a second search means for generating the character string generated by the first character string generation means. The present invention is characterized in that the second search means searches the injury / illness name master as a second keyword to see if a character string matching the second keyword is found, and if a character string matching the second keyword is found as a result of the search, outputs information about the injury / illness name corresponding to the hit character string; and if a character string matching the second keyword is not found as a result of the search by the second search means, refers to a prefix master which stores a prefix, which is a specific character string attached to the beginning of a character string, and generates a character string from the second keyword by removing the prefix; and a third search means uses the character string generated by the second character string generation means as a third keyword to search the injury / illness name master to see if a character string matching the third keyword is found, and if a character string matching the third keyword is found as a result of the search, functions as the third search means which outputs information about the injury / illness name corresponding to the hit character string.
[0034] According to such a program, it is possible to extract character strings representing names of illnesses and injuries from text data without adding more keywords than necessary to the master. Effect of the Invention
[0035] According to the present invention, character strings representing medical procedures, medicines, and names of injuries and illnesses can be extracted from text data without adding more keywords than necessary to the master. [Brief description of the drawings]
[0036] [Figure 1] FIG. 1 is a block diagram of a text data analysis system according to a first embodiment. [Diagram 2] FIG. 1 is a diagram explaining the first medical procedure / medicine master. [Diagram 3] FIG. 2 is a diagram explaining the second medical procedure / medicine master. [Figure 4] FIG. 1A is a diagram for explaining a prefix master, and FIG. [Diagram 5] FIG. 2 is a diagram illustrating a first example of a process in the text data analysis system of the first embodiment. [Figure 6] FIG. 11 is a diagram illustrating a second example of the processing in the text data analysis system of the first embodiment. [Figure 7] FIG. 11 is a diagram illustrating a third example of the processing in the text data analysis system of the first embodiment. [Figure 8] FIG. 8 is a diagram continuing from FIG. 7 for explaining a third example of processing in the text data analysis system of the first embodiment. [Figure 9] 4 is a flowchart illustrating the operation of the text data analysis system of the first embodiment. [Figure 10] 10 is a flowchart continuing from FIG. 9, illustrating the operation of the text data analysis system according to the first embodiment. [Figure 11] FIG. 11 is a block diagram of a text data analysis system according to a second embodiment. [Figure 12]FIG. 13 is a diagram for explaining the injury / illness name master. [Figure 13] FIG. 1A is a diagram for explaining a suffix master, FIG. 1B is a diagram for explaining a prefix master, and FIG. [Figure 14] FIG. 11 is a diagram illustrating a first example of a process in the text data analysis system of the second embodiment. [Figure 15] FIG. 11 is a diagram illustrating a second example of the processing in the text data analysis system of the second embodiment. [Figure 16] FIG. 15 is a diagram illustrating a second example of processing in the text data analysis system of the second embodiment. [Figure 17] 10 is a flowchart illustrating the operation of the text data analysis system according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0037] Next, a first embodiment will be described. As shown in Fig. 1, the text data analysis system 1 according to the first embodiment is a system that converts and extracts items related to medical procedures and items related to pharmaceuticals from text data created based on items described in, for example, a medical certificate, a medical treatment statement, a prescription statement, etc., into a format included in a basic master defined by the Ministry of Health, Labor and Welfare. The text data analysis system 1 includes a data acquisition means 11, a character string extraction means 12, a first search means 13, a first character string generation means 14, a second search means 15, a second character string generation means 16, a similar character string extraction means 17, a first similarity calculation means 18, an output means 19, a second similarity calculation means 20, an element character string extraction means 21, and a third similarity calculation means 22.
[0038] The text data analysis system 1 is made up of a computer including a CPU, RAM, ROM, etc. (not shown), and a storage device 90. The text data analysis system 1 realizes each means by loading a computer program stored in the ROM or storage device 90 into the RAM and executing it. In other words, the computer program causes the computer constituting the text data analysis system 1 to function as a data acquisition means 11, a character string extraction means 12, a first search means 13, a first character string generation means 14, a second search means 15, a second character string generation means 16, a similar character string extraction means 17, a first similarity calculation means 18, an output means 19, a second similarity calculation means 20, a component character string extraction means 21, and a third similarity calculation means 22.
[0039] The storage device 90 stores a medical practice / medicine master, a prefix master, and an element character string master. The medical procedure / drug master records character strings that represent medical procedures or drugs. Here, the medical procedures include at least one of the medical procedures listed in the medical procedure master, the medical procedures listed in the dental procedure master, and the dispensing procedures listed in the dispensing procedure master. Also, the drugs include at least one of the drugs listed in the drug master and the specific equipment listed in the specific equipment master. The medical procedure master, dental procedure master, dispensing procedure master, drug master, and specific equipment master are basic masters established by the Ministry of Health, Labor and Welfare.
[0040] In the first embodiment, a medical procedure listed in the medical procedure master is exemplified as a medical procedure, and a medicine listed in the medicine master is exemplified as a medicine.
[0041] The medical practice / drug master data includes a first medical practice / drug master data and a second medical practice / drug master data.
[0042] As shown in Figure 2, the first medical procedure / drug master is configured as a table that stores character strings representing medical procedures (abbreviated kanji names and basic kanji names) listed in the medical procedure master or the medical fee schedule, in correspondence with fee schedule category numbers, and also stores character strings representing drugs (kanji names and basic kanji names) listed in the drug master, in correspondence with drug price standards codes.
[0043] As shown in Fig. 3, the second medical practice / drug master is data prepared to facilitate the similar string extraction means 17 to extract strings including a third keyword described later, and stores strings representing one or more medical practices or drugs in association with index strings that are strings included in the strings representing the one or more medical practices or drugs. In more detail, the second medical practice / drug master is configured as a table that stores index strings prepared to facilitate the extraction of strings including the third keyword, such as "esophageal resection reconstruction", "endoscopic colonic polyp mucosal resection", and "lens reconstruction", in association with at least one fee schedule classification number or drug price code that corresponds to a string representing a medical practice or drug including the index string.
[0044] As shown in FIG. 4(a), the prefix master is configured as a table that stores prefixes, which are specific character strings attached to the beginning of character strings. A prefix is a character string that may be attached to the beginning of a medical procedure or medicine described in a medical treatment statement, etc., and indicates, for example, the name of a body part or a position such as "lower", "upper", "left", "left side", "right", or "right side". In the present invention, a character string includes a single character. A prefix can be determined, for example, by referring to modifiers used for prefixes listed in the modifier master of the basic master defined by the Ministry of Health, Labor and Welfare.
[0045] As shown in Fig. 4(b), the element string master is configured as a table that stores element strings. An element string is a specific string included in a string representing a medical procedure or medicine, such as "lens" or "insertion." As the element string, a string that does not generate too many candidates when narrowing down the string representing a medical procedure or medicine, for example, a string that can narrow down the candidates to about 500 or less, is adopted.
[0046] Returning to FIG. 1, the data acquisition means 11 acquires text data. The data acquisition means 11 may be configured to read and acquire text data that is previously created and stored in the storage device 90 or a storage medium, or may be configured to acquire text data inputted by an input device such as a keyboard connected to the text data analysis system 1. The previously created text data may be, for example, data generated by reading a medical care statement or the like with a scanner and performing optical character recognition (OCR). The data acquisition means 11 may also be configured to acquire text data generated by reading a medical care statement or the like with a scanner connected to the text data analysis system 1 and performing OCR.
[0047] The character string extraction means 12 extracts a group of character strings representing one item from the text data acquired by the data acquisition means 11. When a group of character strings representing one item exists across multiple lines as shown in Fig. 5(a), the character string extraction means 12 extracts the character strings as a single line of character strings as shown in Fig. 5(b).
[0048] The first search means 13 uses the character string extracted by the character string extraction means 12 as a first keyword and searches the medical practice / medicine master for a character string that matches the first keyword. In detail, the first search means 13 refers to the first medical practice / medicine master and searches for a character string that matches the first keyword. If the search results in a character string that matches the first keyword, the first search means 13 outputs information on the medical practice or medicine corresponding to the hit character string.
[0049] Here, in the first embodiment, the first search means 13, the second search means 15, and the output means 19 output at least one of the following as information on medical procedures or pharmaceuticals: the name of a medical procedure and the fee table classification number, and the name of a pharmaceutical product and the drug price standard code. The name of a medical procedure is at least one of the abbreviated kanji name and the basic kanji name listed in the medical procedure master, and the name of a pharmaceutical product is at least one of the kanji name and the basic kanji name listed in the pharmaceutical master.
[0050] The output method is arbitrary. For example, the data may be displayed on a display connected to the text data analysis system 1, may be printed by a printer connected to the text data analysis system 1, or may be transmitted to a user terminal connected to the text data analysis system 1 via a network such as the Internet.
[0051] When the search by the first search means 13 does not find a character string matching the first keyword, the first character string generating means 14 refers to the prefix master and generates a character string from the first keyword by removing the prefix. Also, when the first keyword does not have a prefix stored in the prefix master, the first character string generating means 14 sets the generated character string as the first keyword.
[0052] The second search means 15 uses the character string generated by the first character string generating means 14 as a second keyword and searches the medical practice / medicine master for a character string that matches the second keyword. In detail, the second search means 15 refers to the first medical practice / medicine master and searches for a character string that matches the second keyword. If the search results in a character string that matches the second keyword, the second search means 15 outputs information on the medical practice or medicine corresponding to the hit character string.
[0053] If the search by the second search means 15 does not find a string matching the second keyword, the second character string generating means 16 generates a string by removing parentheses and the character string enclosed by the parentheses from the second keyword. The second character string generating means 16 also generates a string by further removing the following character strings (1) to (5) from the second keyword. (1) Leading or trailing spaces (2) Any spaces in the middle and the characters following the spaces (3) Bullet (4) Commas (5) A number and a character string immediately following the number that indicates the unit
[0054] As an example, the second character string generating means 16 first performs a process of removing any spaces at the beginning or end of the second keyword, and then performs a process of removing any spaces in the middle and any character strings following the spaces. Next, the second character string generating means 16 performs a process of removing parentheses such as parentheses "(", ")" and brackets "", """ and character strings enclosed by the parentheses, and then performs a process of removing a dot "·" and a comma ",". Furthermore, the second character string generating means 16 performs a process of removing numbers and character strings representing units immediately following the numbers, such as "10mg" and "2%".
[0055] In addition, when the first keyword and the second keyword are the same character string, the second search means 15 may not search for a character string that matches the second keyword, and the second character string generation means 16 may generate a character string by removing parentheses and the character string enclosed by the parentheses from the second keyword.
[0056] Furthermore, if the second keyword does not include parentheses, a character string enclosed by the parentheses, and the above character strings (1) to (5), the second character string generating means 16 regards the generated character string as the second keyword.
[0057] The similar string extraction means 17 extracts a string including the third keyword from the medical practice / drug master using the string generated by the second string generation means 16 as the third keyword. In detail, the similar string extraction means 17 refers to the second medical practice / drug master and the first medical practice / drug master to extract a string including the third keyword (hereinafter also referred to as a "first similar string"). More specifically, the similar string extraction means 17 extracts a fee table category number or a drug price standard code associated with the first similar string from the third keyword and the second medical practice / drug master, and extracts a first similar string associated with the extracted fee table category number or drug price standard code from the first medical practice / drug master.
[0058] When the similar string extracting means 17 extracts a plurality of first similar strings, the first similarity calculating means 18 calculates the similarity between each of the first similar strings extracted by the similar string extracting means 17 and the second keyword.
[0059] Here, in the first embodiment, the first similarity calculation means 18, the second similarity calculation means 20 and the third similarity calculation means 22 calculate the similarity between two character strings based on at least one of the Levenshtein distance and the Jaro-Winkler distance, as an example.
[0060] The output means 19 outputs information on at least one medical practice or medicine corresponding to the first similar string extracted by the similar string extraction means 17. In particular, when the similar string extraction means 17 extracts one first similar string, the output means 19 outputs information on the medical practice or medicine corresponding to the first similar string. Furthermore, when the similar string extraction means 17 extracts a plurality of first similar strings, the output means 19 outputs information on the medical practice or medicine corresponding to the first similar string among the plurality of first similar strings, the similarity calculated by the first similarity calculation means 18 being equal to or greater than a predetermined value.
[0061] In the first embodiment, the similarity is calculated as a value between 0 and 1, with a larger value indicating a higher similarity between the two character strings, and a value of 1 indicating that the two character strings match. The predetermined value of the similarity is, for example, 0.8. Note that the text data analysis system 1 may be configured such that the user can arbitrarily set the predetermined value of the similarity.
[0062] If the similar string extraction means 17 is unable to extract a string (first similar string) containing the third keyword from the medical procedure / medicine master, the second similarity calculation means 20 calculates the similarity between the third keyword generated by the second string generation means 16 and an index string stored in the second medical procedure / medicine master.
[0063] When there is an index string whose similarity calculated by the second similarity calculation means 20 is equal to or greater than a predetermined value, the similar string extraction means 17 extracts a string corresponding to the index string from the medical practice / drug master. In detail, when there is an index string whose similarity to the third keyword is equal to or greater than a predetermined value, the similar string extraction means 17 extracts a point table category number or drug price code corresponding to the similar string from the second medical practice / drug master, and extracts a string (hereinafter also referred to as a "corresponding string") corresponding to the extracted point table category number or drug price code from the first medical practice / drug master.
[0064] When there are multiple corresponding strings extracted by the similar string extraction means 17, the first similarity calculation means 18 calculates the similarity between each of the corresponding strings extracted by the similar string extraction means 17 and the second keyword, and the output means 19 outputs information on the medical procedure or medicine corresponding to the corresponding strings whose similarity calculated by the first similarity calculation means 18 is equal to or higher than a predetermined value.
[0065] If there is no index string whose similarity calculated by the second similarity calculation means 20 is equal to or greater than a predetermined value, the element string extraction means 21 refers to the element string master and extracts at least one element string from the second keyword.
[0066] The similar string extraction means 17 extracts strings including the fourth keyword from the medical practice / medicine master, using the element string extracted by the element string extraction means 21 as the fourth keyword. In more detail, the similar string extraction means 17 extracts strings including the fourth keyword (hereinafter also referred to as "second similar strings") from the first medical practice / medicine master. In the first embodiment, when there are multiple element strings extracted by the element string extraction means 21, the similar string extraction means 17 first extracts the second similar string from the medical practice / medicine master, using the element string (hereinafter also referred to as "priority element string" in the first embodiment) located close to the center of the second keyword as the fourth keyword.
[0067] The third similarity calculation means 22 calculates the similarity between the second keyword and the second similar string extracted by the similar string extraction means 17. In detail, the third similarity calculation means 22 calculates the similarity between the second keyword and the second similar string extracted by the similar string extraction means 17 with the priority element string as the fourth keyword before calculating the similarity between the second keyword and other element strings.
[0068] The output means 19 outputs information on the medical practice or medicine corresponding to the second similar string when there is a string whose similarity is equal to or greater than a predetermined level among the second similar strings whose similarity has been previously calculated by the third similarity calculation means 22. When the output means 19 has output the information on the medical practice or medicine corresponding to the second similar string whose similarity has been previously calculated, thereafter the similar string extraction means 17 does not extract second similar strings for other element strings, and the third similarity calculation means 22 does not calculate similarities.
[0069] If there is no string with a similarity equal to or greater than a predetermined level among the second similar strings whose similarity has been calculated earlier, similar string extraction means 17 next extracts a second similar string from the first medical practice / drug master using an element string located close to the center of the second keyword as a fourth keyword, and third similarity calculation means 22 calculates the similarity between the extracted second similar string and the second keyword. Then, if there is a string with a similarity equal to or greater than a predetermined level among the second similar strings whose similarity has been calculated, output means 19 outputs information on the medical practice or drug corresponding to the second similar string.
[0070] When the element string extraction means 21 extracts a plurality of element strings, and the positions of the second keywords of the two element strings from the center are the same, the similar string extraction means 17 extracts second similar strings from the medical practice / medicine master with each element string as a fourth keyword, and calculates the similarity between the extracted second similar strings and the second keyword. Then, the output means 19 outputs information on the medical practice or medicine corresponding to the second similar string if there is a string with a similarity equal to or greater than a predetermined value among the second similar strings for which the similarity has been calculated. In this case, the similarity may be calculated for the string with a smaller number of extracted second similar strings before the string with a larger number, and if there is a string with a similarity equal to or greater than a predetermined value among the second similar strings for which the similarity has been calculated first, the information on the medical practice or medicine corresponding to the second similar string may be output and the process may end.
[0071] Here, the processing in the text data analysis system 1 of the first embodiment will be described with reference to a specific example. As a first example, as shown in FIG. 5(c), the first search means 13 uses the character string “Left lens reconstruction surgery (in case of inserting an intraocular lens, etc.)” acquired by the data acquisition means 11 and extracted by the character string extraction means 12 as a first keyword, and searches the first medical procedure / drug master for a character string that matches the first keyword.
[0072] If the search by the first search means 13 does not result in a string matching the first keyword, the first string generation means 14 generates the string "Lens reconstruction surgery (in the case of inserting an intraocular lens, etc.)" by removing the prefix "left" from the first keyword, by referring to the prefix master, as shown in Figure 5 (d).
[0073] The second search means 15 uses the character string generated by the first character string generation means 14, “Lens reconstruction surgery (intraocular lens insertion, etc.)” as a second keyword, and searches the first medical procedure / drug master for a character string that matches the second keyword.
[0074] If the search by the second search means 15 does not result in a string matching the second keyword, the second string generation means 16 generates the string "lens reconstruction surgery" by removing the parentheses and the string "(in case of inserting an intraocular lens, etc.)" enclosed in the parentheses from the second keyword, as shown in Figure 5 (e).
[0075] The similar string extraction means 17 extracts, from the second medical practice / drug master, the point table section numbers "K2821I", "K2821B", "K2822" and "K2823" associated with the string including the third keyword (first similar string) using the string "Lens reconstruction surgery" generated by the second string generation means 16 as the third keyword. After that, the similar string extraction means 17 extracts, from the first medical practice / drug master, four first similar strings shown in Fig. 5(f) associated with the extracted point table section numbers.
[0076] When there are a plurality of first similar strings extracted by the similar string extraction means 17, the first similarity calculation means 18 calculates the similarity between each of the first similar strings and the second keyword “lens reconstruction surgery (insertion of an intraocular lens, etc.)”, as shown in FIG. 5(g).
[0077] The output means 19 outputs information on medical procedures and medicines corresponding to the first similar strings having a similarity calculated by the first similarity calculation means 18 of a predetermined value of 0.8 or more among the first similar strings extracted by the similar string extraction means 17. Specifically, the output means 19 outputs information on the medical procedure name "Lens reconstruction surgery (when an intraocular lens is inserted) (other)" and the point table category number "K2821-ro", and the medical procedure name "Lens reconstruction surgery (when an intraocular lens is not inserted)" and the point table category number "K2822". The output means 19 may be configured to further output information on the similarity calculated by the first similarity calculation means 18.
[0078] Next, we will explain the second example. In the second example, for the same character string as in the first example, "Left lens reconstruction surgery (in case of inserting an intraocular lens, etc.)", the part "lens reconstruction surgery" is erroneously read as "water crystal restructuring surgery" by OCR.
[0079] In the second example, as shown in FIG. 6(a), the first search means 13 searches the first medical procedure / drug master for a character string that matches the first keyword, using the character string "left vitreous body resection (when inserting an intraocular lens, etc.)" as the first keyword, and since no match is found, as shown in FIG. 6(b), the first character string generation means 14 refers to the prefix master and generates the character string "left vitreous body resection (when inserting an intraocular lens, etc.)" by removing the prefix "left" from the first keyword.
[0080] The second search means 15 searches the first medical procedure / drug master for a character string that matches the second keyword, using the character string "left hydrated body resection (when inserting an intraocular lens / other)" as the second keyword, and since no match is found, the second character string generation means 16 generates the character string "left hydrated body resection" by removing the parentheses and the character string "(when inserting an intraocular lens / other)" enclosed by the parentheses from the second keyword, as shown in Figure 6(c).
[0081] The similar string extraction means 17 attempts to extract a string containing the third keyword from the second medical procedure and pharmaceutical master, with the string "aqueous humor reconstruction technique" as the third keyword. However, in the second example, a string containing the third keyword cannot be extracted. Therefore, as shown in FIG. 6(d), the second similarity calculation means 20 calculates the similarity between the third keyword "aqueous humor reconstruction technique" and each index string stored in the second medical procedure and pharmaceutical master, respectively.
[0082] When there is an index string for which the similarity calculated by the second similarity calculation means 20 is equal to or greater than a predetermined value of 0.8, the similar string extraction means 17 extracts the string (corresponding string) corresponding to the index string from the first medical procedure and pharmaceutical master. Specifically, since there is an index string "intraocular lens implantation for cataract reconstruction" for which the similarity is equal to or greater than the predetermined value of 0.8, the similar string extraction means 17 extracts the four corresponding strings shown in FIG. 5(e) from the first medical procedure and pharmaceutical master.
[0083] When there are a plurality of corresponding strings extracted by the similar string extraction means 17, as shown in FIG. 6(f), the first similarity calculation means 18 calculates the similarity between each corresponding string and the second keyword "aqueous humor reconstruction technique (when inserting an intraocular lens and other cases)", respectively. The output means 19 outputs information on the medical procedure and pharmaceutical corresponding to the corresponding string for which the similarity calculated by the first similarity calculation means 18 is equal to or greater than a predetermined value of 0.8 among the corresponding strings extracted by the similar string extraction means 17.
[0084] Next, a third example will be described. The third example is a case where the item "short hand 3 (intraocular lens implantation for cataract reconstruction and other cases (one side))" described in the medical record, etc. is misread as "hedge hand 3 (finger for cataract reconstruction, lens insertion in the inkstone, and other cases (one example))" by OCR.
[0085] In the third example, as shown in FIG. 7(a), the first search means 13 searches the first medical practice / drug master for a character string that matches the first keyword, using the character string "Kakete 3 (Lens reconstruction finger / Intra-capsule lens insertion / Other (One example))" as the first keyword, and since there is no match, the first character string generation means 14 attempts to generate a character string from the first keyword by referring to the prefix master and removing the prefix. In the third example, since the first keyword does not have a prefix stored in the prefix master, as shown in FIG. 7(b), the first character string generation means 14 generates the same character string as the first keyword, "Kakete 3 (Lens reconstruction finger / Intra-capsule lens insertion / Other (One example))."
[0086] The second search means 15 searches the first medical procedure / drug master for a string that matches the second keyword, using the string "Kakite 3 (Lens reconstruction finger, intra-lens insertion, and other items (one example))" as the second keyword, and since no string is found, the second string generation means 16 generates the string "Kakite 3" by removing the parentheses and the string "(Lens reconstruction finger, intra-lens insertion, and other items (one example))" enclosed by the parentheses from the second keyword, as shown in Figure 7(c).
[0087] The similar string extraction means 17 attempts to extract a string including the third keyword from the second medical practice / medicine master using the string "KANJI TE 3" as the third keyword, but in the third example, it is not possible to extract a string including the third keyword. Therefore, the second similarity calculation means 20 calculates the similarity between the third keyword "KANJI TE 3" and each index string stored in the second medical practice / medicine master.
[0088] In the third example, the similarities calculated by the second similarity calculation means 20 are all less than the predetermined value of 0.8, and as shown in Figure 7(d), the element string extraction means 21 refers to the element string master and extracts the element strings “lens” and “insertion” from the second keyword (see Figure 7(b)).
[0089] Since there are multiple element strings extracted by the element string extraction means 21, the similar string extraction means 17 first selects the element string “lens” located near the center of the second keyword as the fourth keyword, and extracts strings (second similar strings) including the fourth keyword “lens” from the first medical practice / drug master, as shown in Figure 8(a).
[0090] As shown in FIG. 8(b), the third similarity calculation means 22 calculates the similarity between each of the extracted second similar strings and the second keyword “Kakete 3 (cataract reconstruction finger, intraocular lens insertion, and others (one example))”.
[0091] The output means 19 outputs information on medical procedures or medicines corresponding to similar strings having a similarity calculated by the third similarity calculation means 22 of the second similar strings extracted by the similar string extraction means 17 that is equal to or greater than a predetermined value of 0.8. Specifically, the output means 19 outputs information on the medical procedure name "Tan 3 (Lens reconstruction surgery, intraocular lens insertion, other, one side)" and the point table division number "A4003ho", the medical procedure name "Tan 3 (Lens reconstruction surgery, intraocular lens insertion, other, both sides)" and the point table division number "A4003he", and the medical procedure name "Tan 3 (Lens reconstruction surgery, intraocular lens insertion, other, one side) (lifestyle care)" and the point table division number "A4003ho". The output means 19 may be configured to further output information on the similarity calculated by the third similarity calculation means 22.
[0092] Next, an example of the operation of the text data analysis system 1 of the first embodiment (a text data analysis method executed by the means included in the computer that constitutes the text data analysis system 1) will be described with reference to a flowchart.
[0093] 9, the text data analysis system 1 first acquires text data (S110) (data acquisition step). Next, the text data analysis system 1 extracts a group of character strings representing one item from the text data acquired in the data acquisition step (S120) (character string extraction step).
[0094] Next, the text data analysis system 1 uses the character string extracted in the character string extraction step as a first keyword and searches the first medical practice / medicine master for a character string matching the first keyword (S131) (first search step). If the search results in a character string matching the first keyword (S132, Yes), the system proceeds to step S183.
[0095] On the other hand, if the search in the first search step does not result in a string that matches the first keyword (S132, No), the text data analysis system 1 refers to the prefix master and generates a string from the first keyword by removing the prefix (S140) (first string generation step).
[0096] Next, the text data analysis system 1 searches the first medical practice / drug master for a character string that matches the second keyword using the character string generated in the first character string generation step as a second keyword (S151) (second search step). If the search results in a character string that matches the second keyword (S152, Yes), the system proceeds to step S183.
[0097] On the other hand, if the search in the second search step does not result in a string matching the second keyword (S152, No), the text data analysis system 1 generates a string from the second keyword by removing parentheses and the string enclosed by the parentheses (S160) (second string generation step).
[0098] Next, the text data analysis system 1 extracts a string (first similar string) including the third keyword from the second medical practice / drug master and the first medical practice / drug master using the string generated in the second string generation step as the third keyword (S171) (similar string extraction step).Then, the text data analysis system 1 determines whether the first similar string has been extracted (S172).
[0099] If a first similar string is extracted (S172, Yes), the text data analysis system 1 determines whether a plurality of first similar strings are extracted (S173). If a single first similar string is extracted (S173, No), the process proceeds to step S183.
[0100] On the other hand, if there are multiple strings extracted in the similar string extraction step (S173, Yes), the text data analysis system 1 calculates the similarity between each of the first similar strings extracted in the similar string extraction step and the second keyword (S181) (first similarity calculation step).Then, the text data analysis system 1 extracts first similar strings whose similarity calculated in the first similarity calculation step is equal to or greater than a predetermined value from among the first similar strings extracted in the similar string extraction step (S182), and proceeds to step S183.
[0101] Here, for example, a medical care statement, a prescription statement, etc., may contain multiple items related to medical procedures such as the name of surgery, items related to medicines, or items other than the items, so that the text data may contain multiple groups of character strings representing one item. Therefore, in step S183, the text data analysis system 1 determines whether there is a next group of character strings representing one item. If there is a next character string (S183, Yes), the text data analysis system 1 returns to step S120 and executes the subsequent processes, and if there is no next character string (S183, No), the text data analysis system 1 proceeds to step S191.
[0102] In step S191, the text data analysis system 1 extracts information on medical practices and medicines corresponding to the character strings found in the first search step or the second search step, and outputs the extracted medical practice and medicine information (S192). Alternatively, the text data analysis system 1 extracts information on at least one medical practice and medicine corresponding to the first similar character string extracted in the similar character string extraction step (S191), and outputs the extracted medical practice and medicine information (S192) (output step).
[0103] In step S172, if a string containing the third keyword (first similar string) cannot be extracted from the second medical practice / drug master and the first medical practice / drug master (No), as shown in FIG. 10, the text data analysis system 1 calculates the similarity between the third keyword and each index string stored in the second medical practice / drug master (S201) (second similarity calculation step).
[0104] Then, the text data analysis system 1 determines whether there is an index string whose similarity calculated in the second similarity calculation step is equal to or greater than a predetermined value (S202). If there is an index string whose similarity is equal to or greater than a predetermined value (Yes in S202), the text data analysis system 1 extracts a string corresponding to the index string (corresponding string) from the second medical practice / drug master and the first medical practice / drug master (S203).
[0105] Then, the text data analysis system 1 proceeds to step S173 in FIG. 9, determines whether there are multiple extracted corresponding character strings, and if there is one (S173, No), proceeds to step S183, and if there are multiple extracted corresponding character strings (S173, Yes), calculates the similarity between each of the extracted corresponding character strings and the second keyword (S181), and executes the subsequent processing.
[0106] Returning to FIG. 10, in step S202, if there is no index string whose similarity calculated in the second similarity calculation step (S201) is equal to or greater than a predetermined value (S202, No), the text data analysis system 1 refers to the element string master and extracts element strings from the second keyword (S211) (element string extraction step).
[0107] Next, the text data analysis system 1 determines whether there are multiple extracted element strings (S212). If there are multiple extracted element strings (S212, Yes), the text data analysis system 1 first determines an element string from which the second similar string is to be extracted, specifically, an element string located near the center of the second keyword (S213), and extracts a string (second similar string) containing the fourth keyword from the first medical practice / medicine master using the determined element string as the fourth keyword (S214). If there is only one extracted element string (S212, No), the text data analysis system 1 extracts the second similar string from the first medical practice / medicine master using the element string as the fourth keyword (S214).
[0108] Next, the text data analysis system 1 calculates the similarity between the extracted second similar string and the second keyword (S221) (third similarity calculation step). Next, the text data analysis system 1 determines whether there is a second similar string whose calculated similarity is equal to or greater than a predetermined level (S222). If there is no second similar string whose similarity is equal to or greater than a predetermined level (S222, No), the process returns to step S213 to determine the next element string from which the second similar string is to be extracted first, and executes the subsequent processes. On the other hand, if there is a second similar string whose similarity is equal to or greater than a predetermined level in step S222 (Yes), the text data analysis system 1 proceeds to step 182 in FIG. 9, extracts a second similar string whose similarity is equal to or greater than a predetermined level (S182), and executes the subsequent processes.
[0109] According to the first embodiment described above, it is possible to extract character strings representing medical procedures or medicines from text data without adding more keywords than necessary to the master. Also, character strings representing medical procedures or medicines can be extracted from text data by converting them into a format included in the basic master defined by the Ministry of Health, Labor and Welfare.
[0110] Furthermore, by further providing the first similarity calculation means 18, it is possible to narrow down and output information on medical procedures and medicines.
[0111] Furthermore, by further providing the second similarity calculation means 20, character strings representing medical procedures or medicines can be extracted more reliably from the text data.
[0112] Furthermore, by further providing the element character string extraction means 21, character strings representing medical procedures or medicines can be extracted more reliably from the text data.
[0113] In addition, a third similarity calculation means 22 is further provided, and when there is a second similar string whose similarity is equal to or higher than a predetermined level among the second similar strings whose similarity has been previously calculated, information on the name of the injury or illness corresponding to the second similar string is output and processing is terminated, thereby reducing the amount of processing required until information on the medical procedure or medicine is output and increasing the processing speed.
[0114] In addition, the second character string generating means 16 not only removes parentheses and character strings enclosed by those parentheses from the second keywords, but also further removes the above character strings (1) to (5), making it easier to narrow down information on medical procedures and pharmaceuticals.
[0115] In the first embodiment, the second character string generating means 16 removes the above character strings (1) to (5) from the second keyword in the second character string generating step, but it is sufficient that the second character string generating means 16 is configured to remove at least one of the above character strings (1) to (5) from the second keyword. Also, for example, the second character string generating means 16 may be configured to remove only parentheses and character strings enclosed by the parentheses from the second keyword in the second character string generating step.
[0116] In the first embodiment, the similar string extraction means 17 extracts the first similar string by referring to the second medical practice / drug master and the first medical practice / drug master, but it may be configured to extract the first similar string by referring to only the first medical practice / drug master. That is, the text data analysis system may not be configured to include the second medical practice / drug master. Also, the first medical practice / drug master may store index strings in association with strings representing medical practices / drugs.
[0117] In addition, in the first embodiment, the medical procedure / drug master stored fee schedule category numbers and drug price standard codes, but it may also store other codes, such as medical procedure codes from the medical procedure master and drug codes from the drug master.
[0118] Next, a second embodiment will be described. In the following, differences from the first embodiment will be described in detail, and the same elements will be denoted by the same reference numerals, and the description thereof will be omitted as appropriate.
[0119] 11, the text data analysis system 1 according to the second embodiment is a system that converts items related to the names of injuries and illnesses from text data created based on items written on a medical certificate or the like into a format included in an ICD10-compatible standard disease name master and extracts the items. The text data analysis system 1 includes a data acquisition means 11, a character string extraction means 12, a first search means 23, a first character string generation means 24, a second search means 25, a second character string generation means 26, a third search means 27, an element character string extraction means 28, a similar character string extraction means 29, a similarity calculation means 30, an output means 31, and a storage device 90.
[0120] In the second embodiment, the computer program stored in the ROM or storage device 90 causes the computer constituting the text data analysis system 1 to function as data acquisition means 11, character string extraction means 12, first search means 23, first character string generation means 24, second search means 25, second character string generation means 26, third search means 27, element string extraction means 28, similar character string extraction means 29, similarity calculation means 30, and output means 31.
[0121] The storage device 90 stores an injury / disease name master, a suffix master, a prefix master, and an element character string master. As shown in Figure 12, the injury / illness name master is configured as a table that stores character strings representing injury / illness names. The injury / illness name master stores character strings representing injury / illness names included in the ICD10-compatible standard illness name master in association with ICD codes.
[0122] As shown in Fig. 13(a), the suffix master is configured as a table that stores suffixes, which are specific character strings that are added to the end of character strings. Suffixes are character strings that may be added to the end of the name of an injury or illness written on a medical certificate, for example, "suspected of," "post-surgery of," "pre-surgery of," "aggravation of," "post-treatment of," "secondary infection of," etc. Suffixes can be determined, for example, by referring to modifiers used for suffixes listed in the modifier master.
[0123] As shown in Fig. 13(b), the prefix master is configured as a table that stores prefixes, which are specific character strings that are added to the beginning of a character string. Prefixes are character strings such as "lower", "acute", "upper", "left", "left side", "right", and "right side" that are sometimes added to the beginning of the name of an injury or illness written on a medical certificate, etc. Prefixes can be determined, for example, by referring to modifiers used for prefixes listed in the modifier master.
[0124] As shown in Fig. 13(c), the element string master is configured as a table that stores element strings. An element string is a specific string included in the name of an injury or illness, such as "pearl," "tendon," "room," "ear," "shrink," "crush," and "middle ear." As the element string, a string that does not generate too many candidates when narrowing down the injury or illness name, for example, a string that can narrow down the candidates to about 500 or less, is adopted.
[0125] Returning to Fig. 11, the first search means 23 searches the injury / illness name master for a character string that matches the first keyword, using the character string extracted by the character string extraction means 12 as the first keyword. If a character string that matches the first keyword is found as a result of the search, the first search means 23 outputs information on the injury / illness name corresponding to the hit character string.
[0126] Here, in the second embodiment, the first search means 23, the second search means 25, the third search means 27 and the output means 31 output information on the injury or illness name and the ICD code as information on the injury or illness name.
[0127] When the search by the first search means 23 does not find a character string matching the first keyword, the first character string generating means 24 refers to the suffix master and generates a character string from the first keyword by removing the suffix. Also, when the first keyword does not have a suffix stored in the suffix master, the first character string generating means 24 sets the generated character string as the first keyword.
[0128] The second search means 25 searches the injury / illness name master for a character string that matches the second keyword, using the character string generated by the first character string generation means 24 as a second keyword. If a character string that matches the second keyword is found as a result of the search, the second search means 25 outputs information on the injury / illness name corresponding to the hit character string.
[0129] If the search by the second search means 25 does not find a character string matching the second keyword, the second character string generating means 26 refers to the prefix master and generates a character string from the second keyword by removing the prefix.
[0130] In addition, when the first keyword and the second keyword are the same character string, the second search means 25 may not search for a character string that matches the second keyword, and the second character string generation means 26 may generate a character string by removing the suffix from the second keyword.
[0131] Furthermore, when the secondary keyword does not have a prefix stored in the prefix master, the second character string generating means 26 sets the generated character string as the secondary keyword.
[0132] The third search means 27 searches the injury / illness name master for a character string that matches the third keyword, using the character string generated by the second character string generation means 26 as a third keyword. When a character string that matches the third keyword is found as a result of the search, the third search means 27 outputs information on the injury / illness name corresponding to the hit character string.
[0133] When the search by the third search means 27 does not find a string matching the third keyword, the element string extraction means 28 refers to the element string master and extracts at least one element string from the third keyword.
[0134] In addition, when the second keyword and the third keyword are the same character string, the third search means 27 may not search for a character string that matches the third keyword, and the element character string extraction means 28 may extract the element character string from the third keyword.
[0135] The similar string extraction means 29 extracts strings (hereinafter also referred to as "similar strings") including the fourth keyword from the injury / disease name master, using the element strings extracted by the element string extraction means 28 as the fourth keyword.
[0136] The similarity calculation means 30 calculates the similarity between the similar string extracted by the similar string extraction means 29 and the third keyword. As an example, the similarity calculation means 30 calculates the similarity between the similar string and the third keyword based on at least one of the Levenshtein distance and the Jaro-Winkler distance. In the second embodiment, the similarity is calculated as a numerical value between 0 and 1.
[0137] The output means 31 outputs information on at least one injury or illness name corresponding to the similar string extracted by the similar string extraction means 29. In detail, the output means 31 outputs information on the injury or illness name corresponding to a string having a similarity calculated by the similarity calculation means 30 equal to or greater than a predetermined value, among the similar strings extracted by the similar string extraction means 29.
[0138] In the second embodiment, when there are multiple element strings extracted by the element string extraction means 28, the similar string extraction means 29 first extracts similar strings from the injury / illness name master using the element string (hereinafter also referred to as the "priority element string" in the second embodiment) that is located close to the center of the third keyword as the fourth keyword.
[0139] The similarity calculation means 30 calculates the similarity between the priority element string as the fourth keyword and the third keyword for the similar string extracted by the similar string extraction means 29 before calculating the similarity between the priority element string as the fourth keyword and the third keyword ... for the similar string extracted by the similar string extraction means 29.
[0140] If there is a string whose similarity is equal to or greater than a predetermined value among the similar strings whose similarity has been calculated previously, the output means 31 outputs information on the injury or illness name corresponding to the similar string. After the output means 31 outputs information on the injury or illness name corresponding to the similar string whose similarity has been calculated previously, the similar string extraction means 29 does not extract similar strings from other element strings, and the similarity calculation means 30 does not calculate similarities.
[0141] If there is no string with a similarity equal to or higher than a predetermined level among the similar strings whose similarity has been calculated earlier, the similar string extraction means 29 next extracts similar strings from the injury / disease name master using an element string located near the center of the third keyword as the fourth keyword, and the similarity calculation means 30 calculates the similarity between the extracted similar string and the third keyword. Then, if there is a string with a similarity equal to or higher than a predetermined level among the similar strings whose similarity has been calculated, the output means 31 outputs information on the injury / disease name corresponding to the similar string.
[0142] When the element string extraction means 28 extracts a plurality of element strings, and the positions of the third keyword from the center of two element strings are the same, the similar string extraction means 29 extracts similar strings from the illness / injury name master with each element string as the fourth keyword, and calculates the similarity between the extracted similar strings and the third keyword. Then, when there is a string whose similarity is equal to or higher than a predetermined value among the similar strings whose similarity is calculated, the output means 31 outputs information on the illness / injury name corresponding to the similar string. In this case, the similarity may be calculated for the string with fewer extracted similar strings before the string with more extracted similar strings, and when there is a string whose similarity is equal to or higher than a predetermined value among the similar strings whose similarity is calculated first, information on the illness / injury name corresponding to the similar string may be output and the process may end.
[0143] Here, the processing in the text data analysis system 1 of the second embodiment will be described with reference to a specific example. As a first example, as shown in FIG. 14(a), the first search means 23 uses the character string “suspected acute influenza pneumonia” acquired by the data acquisition means 11 and extracted by the character string extraction means 12 as a first keyword, and searches the injury / illness name master for a character string that matches the first keyword.
[0144] If the search by the first search means 23 does not result in a string matching the first keyword, the first string generation means 24 refers to the suffix master and generates the string "acute influenza pneumonia" by removing the suffix "suspected of" from the first keyword, as shown in Figure 14(b).
[0145] The second search means 25 uses the character string "acute influenza pneumonia" generated by the first character string generating means 24 as a second keyword and searches the injury / disease name master for a character string that matches the second keyword.
[0146] If the search by the second search means 25 does not result in a string matching the second keyword, the second string generation means 26 refers to the prefix master and generates the string "influenza pneumonia" by removing the prefix "acute" from the second keyword, as shown in Figure 14(c).
[0147] The third search means 27 searches the illness name master for a character string that matches the third keyword, using the character string "influenza pneumonia" generated by the second character string generation means 26 as the third keyword. If the search results in a character string that matches the third keyword, the third search means 27 outputs information on the illness name corresponding to the hit character string. Specifically, the third search means 27 outputs information on the illness name "influenza pneumonia" and the ICD code "J110", as shown in Fig. 14(d).
[0148] As a second example, as shown in FIG. 15(a), the first search means 23 uses the character string “Postoperative period for right pearlescent otitis media” acquired by the data acquisition means 11 and extracted by the character string extraction means 12 as a first keyword, and searches the injury / illness name master for a character string that matches the first keyword.
[0149] If the search by the first search means 23 does not result in a string matching the first keyword, the first string generation means 24 refers to the suffix master and generates the string "right pearlescent otitis media" by removing the suffix "post-operative" from the first keyword, as shown in Figure 15(b).
[0150] The second search means 25 uses the character string "right pearlescent otitis media" generated by the first character string generating means 24 as a second keyword and searches the injury / disease name master for a character string that matches the second keyword.
[0151] If the search by the second search means 25 does not result in a string matching the second keyword, the second string generation means 26 refers to the prefix master and generates the string "pearlogenic otitis media" by removing the prefix "right" from the second keyword, as shown in Figure 15(c).
[0152] The third search means 27 searches the injury / disease name master for a character string that matches the third keyword, using the character string "pearlogenic otitis media" generated by the second character string generation means 26 as a third keyword.
[0153] If the search by the third search means 27 does not result in a string matching the third keyword, the element string extraction means 28 refers to the element string master and extracts the element strings “pearl” and “middle ear” from the third keyword, as shown in FIG. 15(d).
[0154] Since there are multiple element strings extracted by the element string extraction means 28, the similar string extraction means 29 first takes the element string “middle ear”, which is located near the center of the third keyword, as the fourth keyword, and extracts strings (similar strings) containing the fourth keyword “middle ear” from the injury / illness name master, as shown in Figure 16(a).
[0155] As shown in FIG. 16(b), the similarity calculation means 30 calculates the similarity between each of the extracted similar character strings and the third keyword "pearlogenic otitis media."
[0156] The output means 31 outputs information on the disease name corresponding to the similar character strings having a similarity calculated by the similarity calculation means 30 of a predetermined value, for example, 0.8 or more, among the similar character strings extracted by the similar character string extraction means 29. Specifically, the output means 31 outputs information on the disease name "cholesteatomatous otitis media" and the ICD code "H71", the disease name "acute otitis media" and the ICD code "H669", and the disease name "chronic otitis media" and the ICD code "H669". The output means 31 may be configured to further output information on the similarity calculated by the similarity calculation means 30.
[0157] Next, an example of the operation of the text data analysis system 1 of the second embodiment will be described with reference to a flowchart.
[0158] 17, the text data analysis system 1 searches the injury / disease name master for a character string that matches the first keyword (S231) (first search step) using the character string extracted in the character string extraction step (S120) as a first keyword. If the search results in a character string that matches the first keyword (S232, Yes), the system proceeds to step S304.
[0159] On the other hand, if the search in the first search step does not result in a string matching the first keyword (S232, No), the text data analysis system 1 refers to the suffix master and generates a string from the first keyword by removing the suffix (S240) (first string generation step).
[0160] Next, the text data analysis system 1 searches the injury / disease name master for a character string that matches the second keyword using the character string generated in the first character string generation step as a second keyword (S251) (second search step). If the search results in a character string that matches the second keyword (S252, Yes), the system proceeds to step S304.
[0161] On the other hand, if the search in the second search step does not result in a string that matches the second keyword (S252, No), the text data analysis system 1 refers to the prefix master and generates a string from the second keyword by removing the prefix (S260) (second string generation step).
[0162] Next, the text data analysis system 1 searches the injury / disease name master for a character string that matches the third keyword, using the character string generated in the second character string generation step as a third keyword (S271) (third search step). If the search results in a character string that matches the third keyword (S272, Yes), the system proceeds to step S304.
[0163] On the other hand, if the search in the third search step does not result in a string matching the third keyword (S272, No), the text data analysis system 1 refers to the element string master and extracts an element string from the third keyword (S280) (element string extraction step).
[0164] Next, the text data analysis system 1 determines whether there are multiple extracted element strings (S291). If there are multiple extracted element strings (S291, Yes), the text data analysis system 1 first determines an element string from which a similar string is to be extracted, specifically, an element string located near the center of the third keyword (S292), and extracts a string (similar string) including the fourth keyword from the injury / illness name master using the determined element string as the fourth keyword (S293). If there is only one extracted element string (S291, No), the text data analysis system 1 extracts a similar string from the injury / illness name master using the element string as the fourth keyword (S293) (similar string extraction step).
[0165] Next, the text data analysis system 1 calculates the similarity between the extracted similar string and the third keyword (S301) (similarity calculation step). Next, the text data analysis system 1 determines whether there is a similar string whose calculated similarity is equal to or greater than a predetermined level (S302). If there is no similar string whose similarity is equal to or greater than a predetermined level (S302, No), the process returns to step S292 to determine the next element string from which a similar string will be extracted first, and executes the subsequent processes. On the other hand, if there is a similar string whose similarity is equal to or greater than a predetermined level in step S302 (Yes), the text data analysis system 1 extracts a similar string whose similarity is equal to or greater than a predetermined level (S303) and proceeds to step S304.
[0166] In step S304, the text data analysis system 1 determines whether there is a group of character strings representing the next item. If there is a next character string (S304, Yes), the system returns to step S120 and executes the subsequent processes. If there is no next character string (S304, No), the system proceeds to step S311.
[0167] In step S311, the text data analysis system 1 extracts information on the injury or illness name corresponding to the character string hit in the first search step, the second search step, or the third search step, and outputs the extracted information on the injury or illness name (S312). Alternatively, the text data analysis system 1 extracts information on at least one injury or illness name corresponding to the similar character string extracted in the similar character string extraction step (S311), and outputs the extracted information on the injury or illness name (S312) (output step).
[0168] According to the second embodiment described above, it is possible to extract character strings representing names of illnesses from text data without adding more keywords than necessary to the master. Also, it is possible to convert character strings representing names of illnesses from text data into a format included in the ICD10-compliant standard disease name master and extract the same.
[0169] Furthermore, by further providing the element string extraction means 28, the similar string extraction means 29, and the output means 31, it is possible to more reliably extract strings representing the names of injuries and illnesses from the text data.
[0170] Furthermore, by further providing a similarity calculation means 30, it is possible to narrow down and output information on names of injuries and illnesses.
[0171] Furthermore, if there is a similar string whose similarity is equal to or higher than a predetermined level among the similar strings whose similarity has been previously calculated, information on the name of the injury or illness corresponding to the similar string is output and processing is terminated, thereby reducing the amount of processing required to output the information on the name of the injury or illness and increasing processing speed.
[0172] In the second embodiment, the injury / disease name master stores ICD codes, but may store other codes, for example, disease name management numbers of the ICD10-compatible standard disease name master. In the second embodiment, the injury / disease name master is created based on the ICD10-compatible standard disease name master, but may be created based on the injury / disease name master of the basic master.
[0173] Although the embodiment has been described above, the present invention is not limited to the above embodiment, and can be modified as appropriate as exemplified below.
[0174] For example, in the above embodiment, when there are multiple extracted element strings, the similarity of the element string located near the center of the specified keyword is calculated before the other element strings, but this is not limited to this. For example, when there are multiple extracted element strings, similar strings are extracted for each extracted element string, and the similarity of the extracted element string with fewer similar strings is calculated first, and if there is a string with a similarity equal to or higher than a predetermined level among the similar strings whose similarity is calculated first, information on the medical treatment or the like corresponding to the similar string is output and the process is terminated. Also, the element string master may be stored in association with a number or the like indicating the priority when calculating the similarity.
[0175] In the above embodiment, the similarity calculation means calculates the similarity between the similar strings extracted by the similar string extraction means and a predetermined keyword, and the output means outputs information corresponding to the similar strings whose similarity calculated by the similarity calculation means is equal to or greater than a predetermined value, but this is not limiting. For example, the output means may be configured to output information corresponding to the similar strings extracted by the similar string extraction means without calculating the similarity if the number of similar strings extracted by the similar string extraction means is equal to or less than a predetermined value.
[0176] In addition, the elements described in the above-mentioned embodiment and the modified examples may be implemented in any combination. For example, when the first embodiment and the second embodiment are implemented in combination, the prefix master of the first embodiment and the prefix master of the second embodiment may be a common master, and the element string master of the first embodiment and the element string master of the second embodiment may be a common master. [Explanation of symbols]
[0177] 1. Text data analysis system 11 Data Acquisition Methods 12 String extraction means 13 First search method 14 First string generation means 15 Second search method 16 Second string generation means 17 Similar character string extraction means 18 First similarity calculation means 19 Output Method 20 Second similarity calculation means 21 Element String Calculation Method 22 Third similarity calculation means 23 First search method 24 First string generation means 25 Second search method 26 Second string generation means 27 Third Search Method 28 Element string extraction method 29 Similar character string extraction means 30 Similarity calculation method 31 means of contribution
Claims
1. A data acquisition means for acquiring text data; a character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; a first search means for searching a medical practice / drug master storing character strings representing medical practices or drugs, using the character string extracted by the character string extraction means as a first keyword, to see if a character string matching the first keyword is found, and for outputting information on the medical practice or drug corresponding to the character string that is found as a result of the search; a first character string generating means for generating a character string obtained by removing the prefix from the first keyword when a character string matching the first keyword is not found as a result of a search by the first search means, by referring to a prefix master storing a prefix which is a specific character string added to the beginning of a character string; a second search means for searching the medical practice / drug master for a character string that matches the second keyword, using the character string generated by the first character string generation means as a second keyword, and for outputting information on the medical practice or drug corresponding to the character string that matches the second keyword when the search results in a hit; a second character string generating means for generating a character string by removing at least parentheses and the character string enclosed by the parentheses from the second keyword when a character string matching the second keyword is not found as a result of the search by the second search means; a similar character string extracting means for extracting a character string including the third keyword from the medical practice / drug master using the character string generated by the second character string generating means as a third keyword; and an output unit that outputs information on at least one medical practice or drug corresponding to the character string extracted by the similar character string extraction unit.
2. a first similarity calculation means for calculating a similarity between each of the character strings extracted by the similar character string extraction means and the second keyword when the similar character string extraction means extracts a plurality of character strings, The text data analysis system according to claim 1, characterized in that the output means outputs information on medical procedures or medicines corresponding to character strings among the character strings extracted by the similar character string extraction means, the character strings having a similarity calculated by the first similarity calculation means that is equal to or greater than a predetermined value.
3. The medical practice / drug master stores character strings representing one or more medical practices or drugs in association with index character strings that are character strings included in the character strings representing the one or more medical practices or drugs, The text data analysis system further includes a second similarity calculation means for calculating a similarity between the third keyword and the index string when the similar string extraction means is unable to extract a string including the third keyword from the medical practice / drug master, 3. The text data analysis system according to claim 1, wherein when there is an index string whose similarity calculated by the second similarity calculation means is equal to or greater than a predetermined value, the similar string extraction means extracts a string corresponding to the index string from the medical procedure / drug master.
4. and an element string extracting means for extracting at least one element string from the second keyword when there is no index string having a similarity equal to or greater than a predetermined value calculated by the second similarity calculating means, by referring to an element string master storing element strings that are specific character strings included in character strings representing medical procedures or medicines; The text data analysis system according to claim 3, characterized in that the similar string extraction means extracts strings including the fourth keyword from the medical procedure / drug master, using the element string extracted by the element string extraction means as a fourth keyword.
5. a third similarity calculation means for calculating a similarity between the second keyword and a character string extracted by the similar character string extraction means, the third similarity calculation means being configured to calculate a similarity between the second keyword and a character string extracted by the similar character string extraction means, the fourth keyword being a fourth keyword, before calculating a similarity between the second keyword and other character strings, when a plurality of element character strings are extracted by the element character string extraction means; The text data analysis system according to claim 4, characterized in that, when a character string whose similarity is equal to or greater than a predetermined value is included among the character strings whose similarity has been previously calculated by the third similarity calculation means, the output means outputs information about a medical procedure or medicine corresponding to the character string.
6. The text data analysis system according to any one of claims 1 to 5, characterized in that the second character string generation means generates a character string by further removing at least one of the following character strings (1) to (5) from the second keyword. (1) Leading or trailing spaces (2) Any spaces in the middle and the characters following the spaces (3) Bullet (4) Commas (5) A number and a character string immediately following the number that indicates the unit
7. The computer includes: A data acquisition step of acquiring text data; a character string extraction step of extracting a group of character strings representing one item from the text data acquired in the data acquisition step; a first search step of searching a medical practice / drug master storing character strings representing medical practices or drugs, using the character string extracted in the character string extraction step as a first keyword, to see if a character string matching the first keyword is found, and outputting information on the medical practice or drug corresponding to the hit character string when a character string matching the first keyword is found as a result of the search; a first character string generating step of generating a character string obtained by removing the prefix from the first keyword by referring to a prefix master that stores a prefix, which is a specific character string added to the beginning of a character string, when a character string matching the first keyword is not found as a result of the search in the first search step; a second search step of searching the medical practice / drug master for a character string that matches the second keyword using the character string generated in the first character string generation step as a second keyword, and outputting information on the medical practice or drug corresponding to the hit character string when the character string matches the second keyword as a result of the search; a second character string generating step of generating a character string by removing at least parentheses and the character string enclosed by the parentheses from the second keyword when a character string matching the second keyword is not found as a result of the search in the second search step; a similar string extraction step of extracting a string including the third keyword from the medical practice / drug master using the string generated in the second string generation step as a third keyword; and an output step of outputting information on at least one medical practice or drug corresponding to the character string extracted in the similar character string extraction step.
8. Computer, A data acquisition means for acquiring text data; a character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; a first search means for searching a medical practice / drug master storing character strings representing medical practices or drugs, using the character string extracted by the character string extraction means as a first keyword, to see if a character string matching the first keyword is found, and for outputting information on the medical practice or drug corresponding to the character string that is found as a result of the search; a first character string generating means for generating a character string obtained by removing the prefix from the first keyword when a character string matching the first keyword is not found as a result of a search by the first search means, by referring to a prefix master storing a prefix which is a specific character string added to the beginning of a character string; a second search means for searching the medical practice / drug master for a character string that matches the second keyword, using the character string generated by the first character string generation means as a second keyword, and for outputting information on the medical practice or drug corresponding to the character string that matches the second keyword when the search results in a hit; a second character string generating means for generating a character string by removing at least parentheses and the character string enclosed by the parentheses from the second keyword when a character string matching the second keyword is not found as a result of the search by the second search means; a similar character string extracting means for extracting a character string including the third keyword from the medical practice / drug master using the character string generated by the second character string generating means as a third keyword; A computer program that functions as an output means for outputting information on at least one medical practice or drug corresponding to the character string extracted by the similar character string extraction means.
Citation Information
Patent Citations
Receipt examination system
JP2003196383A
Character recognition processing method, and method and device for recognition processing of content of medical fee bill
JP2005275510A
Disease name encoding method and disease name encoding program
JP2006309585A
Bill medical computer processing system code acquisition device
JP2008077589A
Document reading device and document reading processing program
JP3349699B2