Text data analysis system, text data analysis method, and computer program
The text data analysis system efficiently extracts medical, pharmaceutical, and disease-related character strings from text data by iteratively refining keywords and utilizing similarity calculations, addressing the challenge of increasing data capacity in existing systems.
Patent Information
- Application Number
- JP2025047496
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-12
AI Technical Summary
Existing text data analysis systems struggle to extract character strings representing medical acts, pharmaceuticals, and disease names from text data without increasing the data capacity of the master by adding unnecessary keywords.
A text data analysis system that includes data acquisition, character string extraction, and multiple search mechanisms to iteratively refine keywords by removing prefixes, suffixes, and parentheses, and utilizing similarity calculations to extract relevant information from masters storing medical acts, pharmaceuticals, and disease names.
Enables efficient extraction of character strings representing medical acts, pharmaceuticals, and disease names from text data without significantly increasing the master's data capacity, thereby improving processing speed and reducing unnecessary keyword additions.
Smart Images

Figure 2025089415000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a text data analysis system, a text data analysis method, and a computer program.
Background Art
[0002] Conventionally, a document reading device that reads a document with an optical character reader is known (Patent Document 1). In this technology, an error recognition database that stores an error recognition character string including misrecognized characters and a correction character string for correcting the misrecognized characters in correspondence is provided. The character string of the read data obtained by reading a document with an optical character reader is searched in the error recognition database, and correction data is created by converting it into the corresponding correction character string in the case of an error recognition character string. In addition, for those that are not correctly corrected, the error recognition success rate is increased by adding them to the error recognition database.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, when searching whether a character string that matches a keyword hits in a master that stores character strings representing medical acts, pharmaceuticals, and disease names using a predetermined character string as a keyword, if the character string described in a medical record or the like is not the same as the character string stored in the master, or if the character string described in a medical record or the like cannot be accurately converted into text data, it is impossible to extract a character string representing a medical act or the like as it is.
[0005] On the other hand, conventionally, when a character string matching a keyword is not stored in the master, there is a technique of increasing the success rate of search by adding the keyword to the master. However, the character string cannot be extracted before adding a new keyword to the master, and there is a problem that the data capacity of the master increases as new keywords are added to the master.
[0006] The present invention has been made in view of the above background, and an object thereof is to provide a text data analysis system, a text data analysis method, and a computer program that can extract character strings representing medical acts, pharmaceuticals, and disease names from text data without adding keywords to the master more than necessary.
Means for Solving the Problems
[0007] The text data analysis system for achieving the above-described object includes: a data acquisition means for acquiring text data; a character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; a first search means for searching, as a first keyword, the character string extracted by the character string extraction means, in a medical act or pharmaceutical master that stores character strings representing medical acts or pharmaceuticals, to determine whether a character string matching the first keyword is found. If a character string matching the first keyword is found as a result of the search, the first search means outputs information on the medical act or pharmaceutical corresponding to the found character string; if a character string matching the first keyword is not found as a result of the search by the first search means, a first character string generation means refers to a prefix master that stores a specific character string (prefix) attached to the beginning of a character string, and generates a character string obtained by removing the prefix from the first keyword; a second search means for searching, as a second keyword, the character string generated by the first character string generation means, in the medical act or pharmaceutical master, to determine whether a character string matching the second keyword is found. If a character string matching the second keyword is found as a result of the search, the second search means outputs information on the medical act or pharmaceutical corresponding to the found character string; if a character string matching the second keyword is not found as a result of the search by the second search means, a second character string generation means generates a character string obtained by removing at least parentheses and the character string enclosed by the parentheses from the second keyword; a similar character string extraction means for extracting, as a third keyword, a character string including the third keyword generated by the second character string generation means, from the medical act or pharmaceutical master; and an output means for outputting information on at least one medical act or pharmaceutical corresponding to the character string extracted by the similar character string extraction means.
[0008] According to such a system, it is possible to extract a character string representing a medical act or a pharmaceutical from text data without adding keywords to the master more than necessary.
[0009] In addition, when there are a plurality of character strings extracted by the similar character string extraction means, the text data analysis system further includes a first similarity calculation means for calculating the similarity between the character strings extracted by the similar character string extraction means and the second keyword respectively. The output means can be configured to output information on medical acts or pharmaceuticals corresponding to the character strings among the character strings extracted by the similar character string extraction means, for which the similarity calculated by the first similarity calculation means is a predetermined value or more.
[0010] According to this, it is possible to narrow down and output information on medical acts and pharmaceuticals.
[0011] In addition, the medical act / pharmaceutical master stores in association a character string representing one or more medical acts or pharmaceuticals and an index character string that is a character string included in the character string representing the one or more medical acts or pharmaceuticals. When the similar character string extraction means cannot extract a character string including the third keyword from the medical act / pharmaceutical master, the text data analysis system further includes a second similarity calculation means for calculating the similarity between the third keyword and the index character string. When there is an index character string for which the similarity calculated by the second similarity calculation means is a predetermined value or more, the similar character string extraction means can be configured to extract a character string corresponding to the index character string from the medical act / pharmaceutical master.
[0012] According to this, it is possible to more reliably extract a character string representing a medical act or a pharmaceutical from text data.
[0013] In addition, when there is no index string with a similarity calculated by the second similarity calculation means that is equal to or greater than a predetermined value, the text data analysis system refers to an element string master that stores an element string, which is a specific string included in a string representing a medical act or a pharmaceutical product, and extracts at least one element string from the second keyword. The similar string extraction means can be configured to extract a string including the fourth keyword from the medical act / pharmaceutical product master using the element string extracted by the element string extraction means as the fourth keyword.
[0014] According to this, a string representing a medical act or a pharmaceutical product can be more reliably extracted from text data.
[0015] In addition, when there are a plurality of element strings extracted by the element string extraction means, the text data analysis system further includes a third similarity calculation means for calculating the similarity with the second keyword earlier than other element strings for the string extracted by the similar string extraction means using the element string located closer to the center of the second keyword as the fourth keyword. The output means can be configured to output information on a medical act or a pharmaceutical product corresponding to the string when there is a string with a similarity equal to or greater than a predetermined value among the strings for which the third similarity calculation means has calculated the similarity earlier.
[0016] According to this, the processing amount until the information on a medical act or a pharmaceutical product is output can be reduced, and the processing speed can be increased.
[0017] In addition, the second string generation means can be configured to generate a string obtained by further removing at least one of the following strings (1) to (5) from the second keyword. (1) Blanks at the beginning or end (2) Blanks in the middle and the string after the blank (3) Filled-in circles (4) Commas (5) Numbers and strings representing units immediately following the numbers
[0018] According to this, it is possible to easily narrow down medical procedures and pharmaceutical information.
[0019] Further, a text data analysis system for achieving the above object includes: a data acquisition means for acquiring text data; a character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; a first search means for searching, using the character string extracted by the character string extraction means as a first keyword, whether a character string matching the first keyword hits in a disease name master that stores character strings representing disease names, and when a character string matching the first keyword hits as a result of the search, a first search means for outputting information on the disease name corresponding to the hit character string; when a character string matching the first keyword does not hit as a result of the search by the first search means, a first character string generation means for generating a character string obtained by removing a suffix, which is a specific character string appended to the end of the character string, by referring to a suffix master that stores the suffix, using the character string obtained by removing the suffix from the first keyword as a second keyword, a second search means for searching whether a character string matching the second keyword hits in the disease name master, and when a character string matching the second keyword hits as a result of the search, a second search means for outputting information on the disease name corresponding to the hit character string; when a character string matching the second keyword does not hit as a result of the search by the second search means, a second character string generation means for generating a character string obtained by removing a prefix, which is a specific character string appended to the beginning of the character string, by referring to a prefix master that stores the prefix, using the character string obtained by removing the prefix from the second keyword as a third keyword, a third search means for searching whether a character string matching the third keyword hits in the disease name master, and when a character string matching the third keyword hits as a result of the search, a third search means for outputting information on the disease name corresponding to the hit character string.
[0020] According to such a system, it is possible to extract a character string representing a disease or injury name from text data without adding keywords to the master more than necessary.
[0021] Further, when the text data analysis system does not find a character string that matches the third keyword as a result of the search by the third search means, it refers to an element character string master that stores an element character string, which is a specific character string included in the disease or injury name, and extracts at least one element character string from the third keyword. The text data analysis system may further include an element character string extraction means, a similar character string extraction means for extracting a character string including the fourth keyword from the disease or injury name master using the element character string extracted by the element character string extraction means as the fourth keyword, and an output means for outputting information on at least one disease or injury name corresponding to the character string extracted by the similar character string extraction means.
[0022] According to this, it is possible to more reliably extract a character string representing a disease or injury name from text data.
[0023] Further, the text data analysis system may further include a similarity calculation means for calculating a similarity between the character string extracted by the similar character string extraction means and the third keyword, and the output means may be configured to output information on a disease or injury name corresponding to a character string among the character strings extracted by the similar character string extraction means, for which the similarity calculated by the similarity calculation means is equal to or greater than a predetermined value.
[0024] According to this, it is possible to narrow down and output information on a disease or injury name.
[0025] Further, when there are a plurality of element character strings extracted by the element character string extraction means, the similarity calculation means calculates the similarity between the third keyword and the character string extracted by the similar character string extraction means using the element character string located closer to the center of the third keyword as the fourth keyword, earlier than other element character strings. The output means may be configured to output information on a disease or injury name corresponding to a character string for which the similarity is equal to or greater than a predetermined value among the character strings for which the similarity has been calculated earlier.
[0026] According to this, it is possible to reduce the processing amount until the information of the disease name is output and increase the processing speed.
[0027] Further, the text data analysis method for achieving the above object is such that the means provided in the computer includes a data acquisition step of acquiring text data, a character string extraction step of extracting a group of character strings representing one item from the text data acquired in the data acquisition step, and using the character string extracted in the character string extraction step as a first keyword, a first search step of searching whether a character string matching the first keyword hits in a medical act or pharmaceutical master that stores character strings representing medical acts or pharmaceuticals. When a character string matching the first keyword hits as a result of the search, a first search step of outputting information on the medical act or pharmaceutical corresponding to the hit character string; when a character string matching the first keyword does not hit as a result of the search in the first search step, referring to a prefix master that stores a specific character string attached to the head of the character string, a first character string generation step of generating a character string obtained by removing the prefix from the first keyword; using the character string generated in the first character string generation step as a second keyword, a second search step of searching whether a character string matching the second keyword hits in the medical act or pharmaceutical master. When a character string matching the second keyword hits as a result of the search, a second search step of outputting information on the medical act or pharmaceutical corresponding to the hit character string; when a character string matching the second keyword does not hit as a result of the search in the second search step, a second character string generation step of generating a character string obtained by removing at least parentheses and the character string surrounded by the parentheses from the second keyword; using the character string generated in the second character string generation step as a third keyword, a similar character string extraction step of extracting a character string including the third keyword from the medical act or pharmaceutical master; and an output step of outputting information on at least one medical act or pharmaceutical corresponding to the character string extracted in the similar character string extraction step.
[0028] According to such a method, it is possible to extract a character string representing a medical act or a pharmaceutical product from text data without adding keywords to the master more than necessary.
[0029] In addition, the text data analysis method for achieving the above object includes steps in which means provided in a computer perform a data acquisition step of acquiring text data, a character string extraction step of extracting a group of character strings representing one item from the text data acquired in the data acquisition step, a first search step of searching, as a first keyword, for a character string that matches the character string extracted in the character string extraction step in a disease name master that stores a character string representing a disease name, and when a character string that matches the first keyword is found as a result of the search, outputting information on the disease name corresponding to the found character string in the first search step; when a character string that matches the first keyword is not found as a result of the search in the first search step, referring to a suffix master that stores a suffix which is a specific character string appended to the end of a character string, and generating a character string obtained by removing the suffix from the first keyword in a first character string generation step; a second search step of searching, as a second keyword, for a character string that matches the character string generated in the first character string generation step in the disease name master, and when a character string that matches the second keyword is found as a result of the search, outputting information on the disease name corresponding to the found character string in the second search step; when a character string that matches the second keyword is not found as a result of the search in the second search step, referring to a prefix master that stores a prefix which is a specific character string appended to the beginning of a character string, and generating a character string obtained by removing the prefix from the second keyword in a second character string generation step; and a third search step of searching, as a third keyword, for a character string that matches the character string generated in the second character string generation step in the disease name master, and when a character string that matches the third keyword is found as a result of the search, outputting information on the disease name corresponding to the found character string in the third search step.
[0030] According to such a method, it is possible to extract a character string representing a disease name from text data without adding keywords to the master more than necessary.
[0031] In addition, a computer program for achieving the above object causes a computer to include: data acquisition means for acquiring text data; character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; first search means for searching, as a first keyword, the character string extracted by the character string extraction means, in a medical act or pharmaceutical master that stores a character string representing a medical act or a pharmaceutical, to determine whether a character string matching the first keyword is found. If, as a result of the search, a character string matching the first keyword is found, the first search means outputs information on the medical act or pharmaceutical corresponding to the found character string; if, as a result of the search by the first search means, a character string matching the first keyword is not found, first character string generation means for referring to a prefix master that stores a prefix which is a specific character string attached to the head of the character string, and generating a character string obtained by removing the prefix from the first keyword; second search means for searching, as a second keyword, the character string generated by the first character string generation means, in the medical act or pharmaceutical master, to determine whether a character string matching the second keyword is found. If, as a result of the search, a character string matching the second keyword is found, the second search means outputs information on the medical act or pharmaceutical corresponding to the found character string; if, as a result of the search by the second search means, a character string matching the second keyword is not found, second character string generation means for generating a character string obtained by removing at least parentheses and the character string enclosed by the parentheses from the second keyword; similar character string extraction means for extracting, as a third keyword, a character string including the character string generated by the second character string generation means, from the medical act or pharmaceutical master; and output means for outputting information on at least one medical act or pharmaceutical corresponding to the character string extracted by the similar character string extraction means.
[0032] According to such a program, it is possible to extract a string representing a medical act or a pharmaceutical from text data without adding keywords to the master more than necessary.
[0033] Also, a computer program for achieving the above object causes a computer to include: data acquisition means for acquiring text data; string extraction means for extracting a group of strings representing one item from the text data acquired by the data acquisition means; first search means for searching, as a first keyword, the string extracted by the string extraction means in a disease name master that stores a string representing a disease name to determine whether a string matching the first keyword hits, and when a string matching the first keyword hits as a result of the search, outputting information on the disease name corresponding to the hit string; when a string matching the first keyword does not hit as a result of the search by the first search means, first string generation means for generating a string obtained by removing a suffix, which is a specific string appended to the end of the string, by referring to a suffix master that stores the suffix; second search means for searching, as a second keyword, the string generated by the first string generation means in the disease name master to determine whether a string matching the second keyword hits, and when a string matching the second keyword hits as a result of the search, outputting information on the disease name corresponding to the hit string; when a string matching the second keyword does not hit as a result of the search by the second search means, second string generation means for generating a string obtained by removing a prefix, which is a specific string appended to the beginning of the string, by referring to a prefix master that stores the prefix; third search means for searching, as a third keyword, the string generated by the second string generation means in the disease name master to determine whether a string matching the third keyword hits, and when a string matching the third keyword hits as a result of the search, functioning as third search means for outputting information on the disease name corresponding to the hit string.
[0034] According to such a program, it is possible to extract a character string representing a disease name from text data without adding keywords to the master more than necessary.
Advantages of the Invention
[0035] According to the present invention, it is possible to extract character strings representing medical procedures, pharmaceuticals, and disease names from text data without adding keywords to the master more than necessary.
Brief Description of the Drawings
[0036]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiments for Carrying Out the Invention
[0037] Next, the first embodiment will be described. As shown in FIG. 1, the text data analysis system 1 according to the first embodiment is a system that converts and extracts items related to medical practices and items related to pharmaceuticals from text data created based on items described in, for example, medical certificates, medical treatment details, dispensing details, etc. into the formats listed in the basic masters defined by the Ministry of Health, Labour and Welfare. The text data analysis system 1 includes a data acquisition means 11, a character string extraction means 12, a first search means 13, a first character string generation means 14, a second search means 15, a second character string generation means 16, a similar character string extraction means 17, a first similarity calculation means 18, an output means 19, a second similarity calculation means 20, an element character string extraction means 21, and a third similarity calculation means 22.
[0038] The text data analysis system 1 consists of a computer including a CPU, RAM, ROM, etc. (not shown) and a storage device 90. The text data analysis system 1 realizes each means by reading a computer program stored in the ROM or the storage device 90 into the RAM and executing it. In other words, the computer program causes the computer constituting the text data analysis system 1 to function as a data acquisition means 11, a character string extraction means 12, a first search means 13, a first character string generation means 14, a second search means 15, a second character string generation means 16, a similar character string extraction means 17, a first similarity calculation means 18, an output means 19, a second similarity calculation means 20, an element character string extraction means 21, and a third similarity calculation means 22.
[0039] The storage device 90 stores a medical practice and pharmaceutical master, a prefix master, and an element character string master. The medical practice and pharmaceutical master stores character strings representing medical practices or pharmaceuticals. Here, medical practices include at least one of the medical practices listed in the medical examination master, the dental examination master, and the dispensing practice master. Also, pharmaceuticals include at least one of the pharmaceuticals listed in the pharmaceutical master and the specific equipment listed in the specific equipment master. The medical examination master, the dental examination master, the dispensing practice master, the pharmaceutical master, and the specific equipment master are basic masters defined by the Ministry of Health, Labour and Welfare.
[0040] In the first embodiment, as an example of a medical practice, a medical practice listed in the medical examination master is exemplified, and as an example of a pharmaceutical, a pharmaceutical listed in the pharmaceutical master is exemplified.
[0041] The medical practice and pharmaceutical master includes a first medical practice and pharmaceutical master and a second medical practice and pharmaceutical master.
[0042] As shown in FIG. 2, the first medical procedure and pharmaceutical master is configured as a table that stores, in association with each other, a character string representing a medical procedure (abbreviated Chinese name and basic Chinese name) listed in the medical treatment master or the medical fee point table and a point table classification number, and a character string representing a pharmaceutical product (Chinese name and basic Chinese name) listed in the pharmaceutical master and a drug price standard code.
[0043] As shown in FIG. 3, the second medical procedure and pharmaceutical master is data prepared to facilitate the extraction of a character string including a third keyword to be described later by the similar character string extraction means 17, and stores, in association with each other, a character string representing one or more medical procedures or pharmaceutical products and an index character string that is a character string included in the character string representing the one or more medical procedures or pharmaceutical products. Specifically, the second medical procedure and pharmaceutical master is configured as a table that stores, in association with each other, an index character string prepared to facilitate the extraction of a character string including a third keyword, for example, index character strings such as "esophagectomy and reconstruction", "endoscopic mucosal resection of large intestine polyps", "lens reconstruction", etc., and at least one point table classification number or drug price standard code corresponding to the character string representing a medical procedure or pharmaceutical product including the index character string.
[0044] As shown in FIG. 4(a), the prefix master is configured as a table that stores prefixes, which are specific character strings attached to the beginning of a character string. Prefixes may be attached to the beginning of medical procedures or pharmaceutical products described in medical certificates, etc., and are, for example, strings representing part names or positions such as "lower", "upper", "left", "left side", "right", "right side", etc. In the present invention, a character string includes the case of a single character. The prefix can be determined with reference to, for example, the modifiers used for the prefixes listed in the modifier master of the basic master defined by the Ministry of Health, Labour and Welfare.
[0045] As shown in FIG. 4(b), the element string master is configured as a table for storing element strings. The element string is a specific string included in a string representing a medical act or a pharmaceutical product, for example, strings such as "lens" and "insertion". As the element string, a string is adopted such that the number of candidates that hit when narrowing down the string representing the medical act or the pharmaceutical product does not become too large, for example, a string that can narrow down the candidates to about 500 cases or less.
[0046] Returning to FIG. 1, the data acquisition means 11 acquires text data. The data acquisition means 11 may be configured to read and acquire, for example, pre-created text data stored in a storage device 90, a storage medium, or the like, or may be configured to acquire text data input by an input device such as a keyboard connected to the text data analysis system 1. The text data created in advance may be, for example, data generated by reading a medical record or the like with a scanner and performing optical character recognition (OCR). Further, the data acquisition means 11 may be configured to read a medical record or the like with a scanner connected to the text data analysis system 1 and acquire text data generated by OCR.
[0047] The string extraction means 12 extracts a group of strings representing one item from the text data acquired by the data acquisition means 11. As shown in FIG. 5(a), when a group of strings representing one item exists across multiple lines, the string extraction means 12 extracts the string as a single-line string as shown in FIG. 5(b).
[0048] The first search means 13 searches the medical act / pharmaceutical master to see if a string that matches the first keyword, which is the string extracted by the string extraction means 12, hits. Specifically, the first search means 13 refers to the first medical act / pharmaceutical master and searches to see if a string that matches the first keyword hits. If the first search means 13 finds that a string that matches the first keyword hits as a result of the search, it outputs information on the medical act or pharmaceutical product corresponding to the hit string.
[0049] Here, in the first embodiment, the first search means 13, the second search means 15, and the output means 19 output at least one of the information of the medical act or pharmaceutical product, that is, the information of the medical act name and the score table classification number, and the information of the pharmaceutical product name and the drug price standard code. Note that the medical act name is at least one of the abbreviated Chinese character name and the basic Chinese character name listed in the medical treatment master, and the pharmaceutical product name is at least one of the Chinese character name and the basic Chinese character name listed in the pharmaceutical product master.
[0050] Also, the output method is arbitrary. For example, it may be output by a method of displaying on a display connected to the text data analysis system 1, or it may be output by a method of printing by a printer connected to the text data analysis system 1, or it may be output by a method of transmitting information to a terminal of a user connected to the text data analysis system 1 through a network such as the Internet.
[0051] When the first character string generation means 14 does not hit a character string that matches the first keyword as a result of the search by the first search means 13, the first character string generation means 14 refers to the prefix master and generates a character string obtained by removing the prefix from the first keyword. Also, when there is no prefix stored in the prefix master in the first keyword, the first character string generation means 14 sets the generated character string as the first keyword.
[0052] The second search means 15 searches whether a character string that matches the second keyword hits from the medical act / pharmaceutical product master using the character string generated by the first character string generation means 14 as the second keyword. Specifically, the second search means 15 refers to the first medical act / pharmaceutical product master and searches whether a character string that matches the second keyword hits. When the second search means 15 hits a character string that matches the second keyword as a result of the search, the second search means 15 outputs the information of the medical act or pharmaceutical product corresponding to the hit character string.
[0053] When, as a result of the search by the second search means 15, no string that matches the second keyword is found, the second string generation means 16 generates a string obtained by removing parentheses and the string enclosed by the parentheses from the second keyword. Further, the second string generation means 16 generates a string obtained by further removing the following strings (1) to (5) from the second keyword. (1) Whitespace at the beginning or end (2) Whitespace in the middle and the string after the whitespace (3) Filled circle (4) Comma (5) Numbers and the string representing the unit immediately following the numbers
[0054] As an example, the second string generation means 16 first executes a process of removing the whitespace at the beginning or end from the second keyword, and then executes a process of removing the whitespace in the middle and the string after the whitespace. Next, the second string generation means 16 executes a process of removing parentheses such as round parentheses "(", ")" and square brackets "「", "」" and the string enclosed by the parentheses, and then executes a process of removing the filled circle "·" and the comma "、". Further, the second string generation means 16 executes a process of removing numbers and the string representing the unit immediately following the numbers, for example, "10mg", "2%", etc.
[0055] Note that when the first keyword and the second keyword are the same string, the second search means 15 may search for whether a string that matches the second keyword is found, or the second string generation means 16 may generate a string obtained by removing parentheses and the string enclosed by the parentheses, etc. from the second keyword.
[0056] Further, when there are no parentheses, the string enclosed by the parentheses, and the above strings (1) to (5) in the second keyword, the second string generation means 16 uses the second keyword as the string to be generated.
[0057] The similar string extraction means 17 extracts a string containing the third keyword from the medical act and pharmaceutical master, using the string generated by the second string generation means 16 as the third keyword. Specifically, the similar string extraction means 17 refers to the second medical act and pharmaceutical master and the first medical act and pharmaceutical master to extract a string containing the third keyword (hereinafter also referred to as the "first similar string"). More specifically, the similar string extraction means 17 extracts the score table classification number or the drug price standard code associated with the first similar string from the third keyword and the second medical act and pharmaceutical master, and extracts the first similar string associated with the extracted score table classification number or drug price standard code from the first medical act and pharmaceutical master.
[0058] When there are multiple first similar strings extracted by the similar string extraction means 17, the first similarity calculation means 18 calculates the similarity between each of the first similar strings extracted by the similar string extraction means 17 and the second keyword.
[0059] Here, in the first embodiment, the first similarity calculation means 18, the second similarity calculation means 20, and the third similarity calculation means 22 calculate the similarity between two strings based on at least one of the Levenshtein distance and the Jaro-Winkler distance, as an example.
[0060] The output means 19 outputs information on at least one medical act or pharmaceutical corresponding to the first similar string extracted by the similar string extraction means 17. Specifically, when there is one first similar string extracted by the similar string extraction means 17, the output means 19 outputs information on one medical act or pharmaceutical corresponding to the first similar string. Also, when there are multiple first similar strings extracted by the similar string extraction means 17, the output means 19 outputs information on the medical act or pharmaceutical corresponding to the first similar string among the multiple first similar strings, for which the similarity calculated by the first similarity calculation means 18 is equal to or greater than a predetermined value.
[0061] In the first embodiment, the similarity is calculated as a numerical value from 0 to 1, where a larger numerical value indicates a higher similarity between the two character strings, and in the case of 1, it indicates that the two character strings match. The predetermined value of the similarity is, for example, 0.8. Note that the text data analysis system 1 may be configured such that the user can arbitrarily set the predetermined value of the similarity.
[0062] When the similar string extraction means 17 cannot extract a character string (first similar string) including the third keyword from the medical act / drug master, the second similarity calculation means 20 calculates the similarity between the third keyword generated by the second character string generation means 16 and the index string stored in the second medical act / drug master.
[0063] When there is an index string for which the similarity calculated by the second similarity calculation means 20 is equal to or greater than a predetermined value, the similar string extraction means 17 extracts the character string corresponding to the index string from the medical act / drug master. Specifically, when there is an index string for which the similarity with the third keyword is equal to or greater than a predetermined value, the similar string extraction means 17 extracts the score table classification number or the drug price standard code associated with the similar string from the second medical act / drug master, and extracts the character string (hereinafter also referred to as "corresponding string") associated with the extracted score table classification number or drug price standard code from the first medical act / drug master.
[0064] When there are a plurality of corresponding strings extracted by the similar string extraction means 17, the first similarity calculation means 18 calculates the similarity between each of the corresponding strings extracted by the similar string extraction means 17 and the second keyword, and the output means 19 outputs the information on the medical act or drug corresponding to the corresponding string for which the similarity calculated by the first similarity calculation means 18 is equal to or greater than a predetermined value.
[0065] When there is no index string for which the similarity calculated by the second similarity calculation means 20 is equal to or greater than a predetermined value, the element string extraction means 21 refers to the element string master and extracts at least one element string from the second keyword.
[0066] The similar string extraction means 17 extracts a string containing the fourth keyword from the medical procedure and drug master, using the element string extracted by the element string extraction means 21 as the fourth keyword. Specifically, the similar string extraction means 17 extracts a string containing the fourth keyword (hereinafter also referred to as the "second similar string") from the first medical procedure and drug master. In the first embodiment, when there are a plurality of element strings extracted by the element string extraction means 21, the similar string extraction means 17 first extracts the second similar string from the medical procedure and drug master, using the element string closest to the center of the second keyword (hereinafter also referred to as the "priority element string" in the first embodiment) as the fourth keyword.
[0067] The third similarity calculation means 22 calculates the similarity between the second similar string extracted by the similar string extraction means 17 and the second keyword. Specifically, the third similarity calculation means 22 calculates the similarity between the second similar string extracted by the similar string extraction means 17 using the priority element string as the fourth keyword and the second keyword earlier than for other element strings.
[0068] When there is a string with a similarity equal to or greater than a predetermined value among the second similar strings for which the third similarity calculation means 22 has calculated the similarity earlier, the output means 19 outputs information on the medical procedure or drug corresponding to the second similar string. When the output means 19 outputs information on the medical procedure or drug corresponding to the second similar string for which the similarity has been calculated earlier, thereafter, the similar string extraction means 17 does not extract the second similar string for other element strings, and the third similarity calculation means 22 does not calculate the similarity.
[0069] When there is no string with a similarity degree equal to or higher than a predetermined value among the second similar strings whose similarity degrees have been calculated previously, the similar string extraction means 17 uses, as the fourth keyword, the element string located closer to the center of the second keyword, and extracts the second similar string from the first medical act / drug master. The third similarity degree calculation means 22 calculates the similarity degree between the extracted second similar string and the second keyword. Then, when there is a string with a similarity degree equal to or higher than a predetermined value among the second similar strings for which the similarity degree has been calculated, the output means 19 outputs the information on the medical act or drug corresponding to the second similar string.
[0070] When there are a plurality of element strings extracted by the element string extraction means 21 and the positions of the two element strings from the center of the second keyword are the same, the similar string extraction means 17 extracts the second similar string from the medical act / drug master using each element string as the fourth keyword, and calculates the similarity degree between the extracted second similar string and the second keyword respectively. Then, when there is a string with a similarity degree equal to or higher than a predetermined value among the second similar strings for which the similarity degree has been calculated, the output means 19 outputs the information on the medical act or drug corresponding to the second similar string. In this case, for the second similar strings with the smaller number of extractions, calculate the similarity degree earlier than those with the larger number. When there is a string with a similarity degree equal to or higher than a predetermined value among the second similar strings for which the similarity degree has been calculated earlier, it is also possible to output the information on the medical act or drug corresponding to the second similar string and end the process.
[0071] Here, the processing in the text data analysis system 1 of the first embodiment will be described while showing a specific example. As a first example, as shown in FIG. 5(c), the first search means 13 uses the string “left lens reconstruction surgery (when inserting an intraocular lens · others)” obtained by the data acquisition means 11 and extracted by the string extraction means 12 as the first keyword, and searches the first medical act / drug master to see if a string matching the first keyword hits.
[0072] When the first string generation means 14 does not find a string that matches the first keyword as a result of the search by the first search means 13, as shown in FIG. 5(d), the string "intraocular lens reconstruction (when inserting an intraocular lens and others)" is generated by removing the prefix "left" from the first keyword with reference to the prefix master.
[0073] The second search means 15 searches the first medical act / drug master to see if a string that matches the second keyword, which is the string "intraocular lens reconstruction (when inserting an intraocular lens and others)" generated by the first string generation means 14, hits.
[0074] When the second string generation means 16 does not find a string that matches the second keyword as a result of the search by the second search means 15, as shown in FIG. 5(e), the string "intraocular lens reconstruction" is generated by removing the parentheses and the string "(when inserting an intraocular lens and others)" enclosed by the parentheses from the second keyword.
[0075] The similar string extraction means 17 extracts the score table classification numbers "K2821 A", "K2821 B", "K2822", and "K2823" associated with the strings (first similar strings) containing the third keyword from the second medical act / drug master, with the string "intraocular lens reconstruction" generated by the second string generation means 16 as the third keyword. Then, the similar string extraction means 17 extracts the four first similar strings shown in FIG. 5(f) associated with the extracted score table classification numbers from the first medical act / drug master.
[0076] When there are a plurality of first similar strings extracted by the similar string extraction means 17, as shown in FIG. 5(g), the first similarity calculation means 18 calculates the similarity between each first similar string and the second keyword "intraocular lens reconstruction (when inserting an intraocular lens and others)".
[0077] The output means 19 outputs information on medical acts and pharmaceuticals corresponding to the first similar character strings extracted by the similar character string extraction means 17, among which the similarity calculated by the first similarity calculation means 18 is 0.8 or more. Specifically, the output means 19 outputs information on the medical act name "Intraocular lens implantation (when inserting an intraocular lens) (others)" and the score table classification number "K2821ロ", and the medical act name "Intraocular lens implantation (when not inserting an intraocular lens)" and the score table classification number "K2822". Note that the output means 19 may be configured to further output information on the similarity calculated by the first similarity calculation means 18.
[0078] Next, a second example will be described. In the second example, for the same character string "Left intraocular lens implantation (when inserting an intraocular lens · others)" as in the first example, the part of "intraocular lens implantation" is misread as "intraocular lens reconstruction" by OCR.
[0079] In the case of the second example, as shown in Fig. 6(a), the first search means 13 uses the character string "Left intraocular lens reconstruction (when inserting an intraocular lens · others)" as the first keyword to search the first medical act and pharmaceutical master to see if a character string matching the first keyword hits. Since it does not hit, as shown in Fig. 6(b), the first character string generation means 14 generates a character string "Intraocular lens reconstruction (when inserting an intraocular lens · others)" by removing the prefix "Left" from the first keyword with reference to the prefix master.
[0080] The second search means 15 uses the character string "Intraocular lens reconstruction (when inserting an intraocular lens · others)" as the second keyword to search the first medical act and pharmaceutical master to see if a character string matching the second keyword hits. Since it does not hit, as shown in Fig. 6(c), the second character string generation means 16 generates a character string "Left intraocular lens reconstruction" by removing the parentheses and the character string "(when inserting an intraocular lens · others)" surrounded by the parentheses from the second keyword.
[0081] The similar string extraction means 17 attempts to extract a string containing the third keyword from the second medical procedure and pharmaceutical master using the string "aqueous humor reconstruction technique" as the third keyword. However, in the second example, a string containing the third keyword cannot be extracted. Therefore, as shown in FIG. 6(d), the second similarity calculation means 20 calculates the similarity between the third keyword "aqueous humor reconstruction technique" and each index string stored in the second medical procedure and pharmaceutical master, respectively.
[0082] When there is an index string for which the similarity calculated by the second similarity calculation means 20 is equal to or greater than a predetermined value of 0.8, the similar string extraction means 17 extracts the string (corresponding string) corresponding to the index string from the first medical procedure and pharmaceutical master. Specifically, since there is an index string "intraocular lens implantation" with a similarity equal to or greater than the predetermined value of 0.8, the similar string extraction means 17 extracts the four corresponding strings shown in FIG. 5(e) from the first medical procedure and pharmaceutical master.
[0083] When there are a plurality of corresponding strings extracted by the similar string extraction means 17, as shown in FIG. 6(f), the first similarity calculation means 18 calculates the similarity between each corresponding string and the second keyword "aqueous humor reconstruction technique (when inserting an intraocular lens and others)". The output means 19 outputs the information on the medical procedure and pharmaceutical corresponding to the corresponding string for which the similarity calculated by the first similarity calculation means 18 is equal to or greater than a predetermined value of 0.8 among the corresponding strings extracted by the similar string extraction means 17.
[0084] Next, a third example will be described. The third example is a case where the item "Short hand 3 (intraocular lens implantation, etc. (one side))" described in the medical record, etc. is misread as "Hedge hand 3 (intraocular lens reconstruction finger, intraocular lens insertion, etc. (one example))" by OCR.
[0085] In the case of the third example, as shown in FIG. 7(a), the first search means 13 uses the character string "Kaki 3 (intraocular lens reconstruction finger, insertion of lens in the inkstone, etc. (one example))" as the first keyword to search the first medical act / drug master to see if a character string that matches the first keyword hits. Since it does not hit, the first character string generation means 14 attempts to generate a character string obtained by removing the prefix from the first keyword with reference to the prefix master. In the third example, since there is no prefix stored in the prefix master for the first keyword, as shown in FIG. 7(b), the first character string generation means 14 generates the same character string as the first keyword, "Kaki 3 (intraocular lens reconstruction finger, insertion of lens in the inkstone, etc. (one example))".
[0086] The second search means 15 uses the character string "Kaki 3 (intraocular lens reconstruction finger, insertion of lens in the inkstone, etc. (one example))" as the second keyword to search the first medical act / drug master to see if a character string that matches the second keyword hits. Since it does not hit, as shown in FIG. 7(c), the second character string generation means 16 generates the character string "Kaki 3" obtained by removing the parentheses and the character string "(intraocular lens reconstruction finger, insertion of lens in the inkstone, etc. (one example))" enclosed by the parentheses from the second keyword.
[0087] The similar character string extraction means 17 attempts to extract a character string containing the third keyword from the second medical act / drug master using the character string "Kaki 3" as the third keyword. However, in the third example, a character string containing the third keyword cannot be extracted. Therefore, the second similarity calculation means 20 calculates the similarity between the third keyword "Kaki 3" and each index character string stored in the second medical act / drug master respectively.
[0088] In the third example, the similarities calculated by the second similarity calculation means 20 are all less than the predetermined value of 0.8. As shown in FIG. 7(d), the element character string extraction means 21 refers to the element character string master and extracts the element character strings "lens" and "insertion" from the second keyword (see FIG. 7(b)).
[0089] Since there are a plurality of element character strings extracted by the element character string extraction means 21, the similarity string extraction means 17 first extracts, as the fourth keyword, the element character string "lens" that is located close to the center of the second keyword, and extracts, from the first medical act and pharmaceutical master, a character string (second similar string) including the fourth keyword "lens" as shown in FIG. 8(a).
[0090] As shown in FIG. 8(b), the third similarity calculation means 22 calculates the similarity of each of the extracted second similar character strings with the second keyword "Kaki 3 (lens reconstruction finger, intraocular lens insertion, etc. (single example))".
[0091] The output means 19 outputs information on a medical act or a pharmaceutical product corresponding to a similar character string in which the similarity calculated by the third similarity calculation means 22 is equal to or greater than a predetermined value of 0.8 among the second similar character strings extracted by the similar character string extraction means 17. Specifically, the output means 19 outputs information on the medical act name "Short hand 3 (lens reconstruction surgery, intraocular lens insertion, etc., single side)" and the score table classification number "A4003 Ho", the medical act name "Short hand 3 (lens reconstruction surgery, intraocular lens insertion, etc., both sides)" and the score table classification number "A4003 He", and the medical act name "Short hand 3 (lens reconstruction surgery, intraocular lens insertion, etc., single side) (life care and recuperation)" and the score table classification number "A4003 Ho". Note that the output means 19 may be configured to further output information on the similarity calculated by the third similarity calculation means 22.
[0092] Next, an example of the operation of the text data analysis system 1 of the first embodiment (a text data analysis method executed by means included in a computer constituting the text data analysis system 1) will be described with reference to a flowchart.
[0093] As shown in FIG. 9, the text data analysis system 1 first acquires text data (S110) (data acquisition step). Next, the text data analysis system 1 extracts a group of character strings representing one item from the text data acquired in the data acquisition step (S120) (character string extraction step).
[0094] Next, the text data analysis system 1 searches the first medical act and drug master to check if a string that matches the string extracted in the string extraction step (the first keyword) hits (S131) (the first search step). If, as a result of the search, a string that matches the first keyword hits (S132, Yes), the process proceeds to step S183.
[0095] On the other hand, if, as a result of the search in the first search step, a string that matches the first keyword does not hit (S132, No), the text data analysis system 1 refers to the prefix master and generates a string obtained by removing the prefix from the first keyword (S140) (the first string generation step).
[0096] Next, the text data analysis system 1 searches the first medical act and drug master to check if a string that matches the string generated in the first string generation step (the second keyword) hits (S151) (the second search step). If, as a result of the search, a string that matches the second keyword hits (S152, Yes), the process proceeds to step S183.
[0097] On the other hand, if, as a result of the search in the second search step, a string that matches the second keyword does not hit (S152, No), the text data analysis system 1 generates a string obtained by removing parentheses and the string enclosed by the parentheses, etc. from the second keyword (S160) (the second string generation step).
[0098] Next, the text data analysis system 1 extracts a string that includes the third keyword (the first similar string) from the second medical act and drug master and the first medical act and drug master, using the string generated in the second string generation step as the third keyword (S171) (the similar string extraction step). Then, the text data analysis system 1 determines whether the first similar string could be extracted (S172).
[0099] When the first similar character string can be extracted (S172, Yes), the text data analysis system 1 determines whether there are multiple extracted first similar character strings (S173). If there is only one extracted first similar character string (S173, No), the process proceeds to step S183.
[0100] On the other hand, when there are multiple character strings extracted in the similar character string extraction step (S173, Yes), the text data analysis system 1 calculates the similarity between each of the first similar character strings extracted in the similar character string extraction step and the second keyword (S181) (first similarity calculation step). Then, the text data analysis system 1 extracts the first similar character strings among the first similar character strings extracted in the similar character string extraction step, where the similarity calculated in the first similarity calculation step is equal to or greater than a predetermined value (S182), and proceeds to step S183.
[0101] Here, for example, in a medical record or a prescription, there may be multiple items related to medical procedures such as the name of the surgery and items related to pharmaceuticals, or there may be other items other than these items. Therefore, in the text data, there may be multiple groups of character strings representing a single item. Thus, in step S183, the text data analysis system 1 determines whether there is the following group of character strings representing a single item. If there is the following character string (S183, Yes), the process returns to step S120 and subsequent processing is executed. If there is no such character string (S183, No), the process proceeds to step S191.
[0102] In step S191, the text data analysis system 1 extracts the information on the medical procedure and pharmaceuticals corresponding to the character string hit in the first search step or the second search step, and outputs the extracted information on the medical procedure and pharmaceuticals (S192). Alternatively, the text data analysis system 1 extracts at least one piece of information on the medical procedure and pharmaceuticals corresponding to the first similar character string extracted in the similar character string extraction step (S191), and outputs the extracted information on the medical procedure and pharmaceuticals (S192) (output step).
[0103] In step S172, when a character string (first similar character string) including the third keyword cannot be extracted from the second medical act / drug master and the first medical act / drug master (No), as shown in FIG. 10, the text data analysis system 1 calculates the similarity between the third keyword and each index character string stored in the second medical act / drug master (S201) (second similarity calculation step).
[0104] Then, the text data analysis system 1 determines whether there is an index character string whose similarity calculated in the second similarity calculation step is equal to or greater than a predetermined value (S202). When there is an index character string whose similarity is equal to or greater than the predetermined value (S202, Yes), the text data analysis system 1 extracts a character string (corresponding character string) corresponding to the index character string from the second medical act / drug master and the first medical act / drug master (S203).
[0105] Thereafter, the text data analysis system 1 proceeds to step S173 in FIG. 9 and determines whether there are a plurality of extracted corresponding character strings. If there is one (S173, No), it proceeds to step S183. If there are a plurality (S173, Yes), it calculates the similarity between each of the extracted corresponding character strings and the second keyword (S181) and executes the subsequent processing.
[0106] Returning to FIG. 10, in step S202, when there is no index character string whose similarity calculated in the second similarity calculation step (S201) is equal to or greater than a predetermined value (S202, No), the text data analysis system 1 refers to the element character string master and extracts an element character string from the second keyword (S211) (element character string extraction step).
[0107] Next, the text data analysis system 1 determines whether there are multiple extracted element strings (S212). If there are multiple extracted element strings (S212, Yes), the text data analysis system 1 determines the element string for extracting the second similar string first, specifically, the element string located closer to the center of the second keyword (S213). Using the determined element string as the fourth keyword, the text data analysis system 1 extracts a string (second similar string) including the fourth keyword from the first medical act / drug master (S214). If there is only one extracted element string (S212, No), the text data analysis system 1 extracts the second similar string from the first medical act / drug master using the element string as the fourth keyword (S214).
[0108] Next, the text data analysis system 1 calculates the similarity between the extracted second similar string and the second keyword (S221) (third similarity calculation step). Next, the text data analysis system 1 determines whether there is a second similar string whose calculated similarity is equal to or greater than a predetermined value (S222). If there is no second similar string whose similarity is equal to or greater than the predetermined value (S222, No), the process returns to step S213 to determine the next element string for extracting the second similar string first, and the subsequent processing is executed. On the other hand, in step S222, if there is a second similar string whose similarity is equal to or greater than the predetermined value (Yes), the text data analysis system 1 proceeds to step 182 in FIG. 9, extracts the second similar string whose similarity is equal to or greater than the predetermined value (S182), and executes the subsequent processing.
[0109] According to the above first embodiment, it is possible to extract a string representing a medical act or a drug from text data without adding more keywords to the master than necessary. In addition, it is possible to convert and extract a string representing a medical act or a drug from text data into a format included in the basic master defined by the Ministry of Health, Labour and Welfare.
[0110] In addition, by further providing the first similarity calculation means 18, it is possible to narrow down and output information on medical acts and drugs.
[0111] Moreover, by further providing the second similarity calculation means 20, a character string representing a medical act or a pharmaceutical product can be more reliably extracted from the text data.
[0112] Moreover, by further providing the element character string extraction means 21, a character string representing a medical act or a pharmaceutical product can be more reliably extracted from the text data.
[0113] Moreover, by further providing the third similarity calculation means 22, when there is a second similar character string having a similarity equal to or greater than a predetermined value among the second similar character strings for which the similarity has been calculated previously, information on the disease name corresponding to the second similar character string is output and the process is terminated. Therefore, the processing amount until the information on the medical act or the pharmaceutical product is output can be reduced and the processing speed can be increased.
[0114] Moreover, since the second character string generation means 16 not only removes parentheses and the character string surrounded by the parentheses from the second keyword, but also further removes the above-mentioned character strings (1) to (5), it is possible to easily narrow down the information on medical acts and pharmaceutical products.
[0115] In the first embodiment, the second character string generation means 16 removes the above-mentioned character strings (1) to (5) from the second keyword in the second character string generation step. However, any configuration may be used as long as at least one of the above-mentioned character strings (1) to (5) is removed from the second keyword. Further, for example, the second character string generation means 16 may be configured to remove only parentheses and the character string surrounded by the parentheses from the second keyword in the second character string generation step.
[0116] In the first embodiment, the similar string extraction means 17 extracts the first similar string with reference to the second medical procedure and drug master and the first medical procedure and drug master. However, for example, it may be configured to extract the first similar string with reference to only the first medical procedure and drug master. That is, the text data analysis system may be configured not to include the second medical procedure and drug master. Further, an index string may be stored in the first medical procedure and drug master in correspondence with a string representing a medical procedure and drug or the like.
[0117] In the first embodiment, the medical procedure and drug master stores the score table classification number and the drug price standard code. However, it may store other codes, for example, the medical procedure code of the medical examination procedure master and the drug code of the drug master.
[0118] Next, the second embodiment will be described. In the following, differences from the first embodiment will be described in detail, and the same points will be denoted by the same reference numerals for the same elements, and the description will be omitted as appropriate.
[0119] As shown in FIG. 11, the text data analysis system 1 according to the second embodiment is a system that extracts by converting an item related to the name of a disease or injury into a format included in the ICD10 corresponding standard disease name master from text data created based on items described in a medical certificate or the like. The text data analysis system 1 includes a data acquisition means 11, a character string extraction means 12, a first search means 23, a first character string generation means 24, a second search means 25, a second character string generation means 26, a third search means 27, an element character string extraction means 28, a similar character string extraction means 29, a similarity calculation means 30, an output means 31, and a storage device 90.
[0120] In the second embodiment, the computer program stored in the ROM or the storage device 90 causes the computer constituting the text data analysis system 1 to function as data acquisition means 11, character string extraction means 12, first search means 23, first character string generation means 24, second search means 25, second character string generation means 26, third search means 27, element character string extraction means 28, similar character string extraction means 29, similarity calculation means 30, and output means 31.
[0121] The storage device 90 stores a disease name master, a suffix master, a prefix master, and an element character string master. As shown in FIG. 12, the disease name master is configured as a table that stores character strings representing disease names. The disease name master stores, in association with each other, a character string representing a disease name listed in the ICD10-compatible standard disease name master and an ICD code.
[0122] As shown in FIG. 13(a), the suffix master is configured as a table that stores suffixes, which are specific character strings appended to the end of a character string. Suffixes may be appended to the end of a disease name described in a medical certificate or the like. Examples of suffixes include character strings such as "suspected of", "post-operative", "pre-operative", "deterioration of", "after treatment of", and "secondary infection of". The suffixes can be determined, for example, with reference to the modifiers used for the suffixes listed in the modifier master.
[0123] As shown in FIG. 13(b), the prefix master is configured as a table that stores prefixes, which are specific character strings appended to the beginning of a character string. Prefixes may be appended to the beginning of a disease name described in a medical certificate or the like. Examples of prefixes include character strings such as "lower", "acute", "upper", "left", "left side", "right", and "right side". The prefixes can be determined, for example, with reference to the modifiers used for the prefixes listed in the modifier master.
[0124] As shown in FIG. 13(c), the element string master is configured as a table for storing element strings. The element string is a specific string included in the disease name, for example, strings such as "pearl", "tendon", "chamber", "ear", "contraction", "ulcer", "middle ear", etc. As the element string, a string that does not cause too many candidates to be hit when narrowing down the disease name is adopted, for example, a string that can narrow down the candidates to about 500 cases or less.
[0125] Returning to FIG. 11, the first search means 23 searches the disease name master to see if a string matching the first keyword, which is the string extracted by the string extraction means 12, is hit. If, as a result of the search, the first search means 23 hits a string that matches the first keyword, it outputs the information of the disease name corresponding to the hit string.
[0126] Here, in the second embodiment, the first search means 23, the second search means 25, the third search means 27, and the output means 31 output the disease name and the ICD code information as the disease name information.
[0127] If, as a result of the search by the first search means 23, a string that matches the first keyword is not hit, the first string generation means 24 refers to the suffix master and generates a string obtained by removing the suffix from the first keyword. Also, when there is no suffix stored in the suffix master in the first keyword, the first string generation means 24 sets the generated string as the first keyword.
[0128] The second search means 25 searches the disease name master to see if a string that matches the second keyword, which is the string generated by the first string generation means 24, is hit. If, as a result of the search, the second search means 25 hits a string that matches the second keyword, it outputs the information of the disease name corresponding to the hit string.
[0129] When, as a result of the search by the second search means 25, no string matching the second keyword is found, the second string generation means 26 refers to the prefix master and generates a string obtained by removing the prefix from the second keyword.
[0130] When the first keyword and the second keyword are the same string, the second search means 25 may not search for whether a string matching the second keyword is found, and the second string generation means 26 may generate a string obtained by removing the suffix from the second keyword.
[0131] Also, when there is no prefix stored in the prefix master in the second keyword, the second string generation means 26 sets the string to be generated as the second keyword.
[0132] The third search means 27 searches the disease name master to find out whether a string matching the third keyword is found, using the string generated by the second string generation means 26 as the third keyword. When, as a result of the search, a string matching the third keyword is found, the third search means 27 outputs information on the disease name corresponding to the found string.
[0133] When, as a result of the search by the third search means 27, no string matching the third keyword is found, the element string extraction means 28 refers to the element string master and extracts at least one element string from the third keyword.
[0134] When the second keyword and the third keyword are the same string, the third search means 27 may not search for whether a string matching the third keyword is found, and the element string extraction means 28 may extract an element string from the third keyword.
[0135] The similar string extraction means 29 extracts a string (hereinafter also referred to as a "similar string") including the fourth keyword from the disease name master, using the element string extracted by the element string extraction means 28 as the fourth keyword.
[0136] The similarity calculation means 30 calculates the similarity between the similar character strings extracted by the similar character string extraction means 29 and the third keyword. As an example, the similarity calculation means 30 calculates the similarity between the similar character strings and the third keyword based on at least one of the Levenshtein distance and the Jaro-Winkler distance. Also in the second embodiment, the similarity is calculated as a numerical value from 0 to 1.
[0137] The output means 31 outputs information on at least one disease name corresponding to the similar character strings extracted by the similar character string extraction means 29. Specifically, the output means 31 outputs information on the disease name corresponding to the character strings among the similar character strings extracted by the similar character string extraction means 29 for which the similarity calculated by the similarity calculation means 30 is equal to or greater than a predetermined value.
[0138] In the second embodiment, when there are a plurality of element character strings extracted by the element character string extraction means 28, the similar character string extraction means 29 first extracts, as the fourth keyword, the element character string located closer to the center of the third keyword (hereinafter also referred to as the "priority element character string" in the second embodiment) from the disease name master as a similar character string.
[0139] The similarity calculation means 30 calculates the similarity between the similar character strings extracted by the similar character string extraction means 29 with the priority element character string as the fourth keyword and the third keyword earlier than for other element character strings.
[0140] If there is a character string with a similarity equal to or greater than a predetermined value among the similar character strings for which the similarity has been calculated earlier, the output means 31 outputs information on the disease name corresponding to the similar character string. When the output means 31 outputs information on the disease name corresponding to the similar character string for which the similarity has been calculated earlier, thereafter, the similar character string extraction means 29 does not extract similar character strings for other element character strings, and the similarity calculation means 30 does not calculate the similarity.
[0141] If there is no string with a similarity degree equal to or higher than a predetermined value among the similar strings for which the similarity degree has been calculated previously, the similar string extraction means 29 extracts, from the disease name master, a similar string using, as the fourth keyword, an element string located closer to the center of the third keyword. Then, the similarity degree calculation means 30 calculates the similarity degree between the extracted similar string and the third keyword. When there is a string with a similarity degree equal to or higher than a predetermined value among the similar strings for which the similarity degree has been calculated, the output means 31 outputs information on the disease name corresponding to the similar string.
[0142] When there are a plurality of element strings extracted by the element string extraction means 28 and the positions of the two element strings from the center of the third keyword are the same, the similar string extraction means 29 extracts, from the disease name master, similar strings using each element string as the fourth keyword, and calculates the similarity degree between each of the extracted similar strings and the third keyword. When there is a string with a similarity degree equal to or higher than a predetermined value among the similar strings for which the similarity degree has been calculated, the output means 31 outputs information on the disease name corresponding to the similar string. In this case, for the smaller number of the extracted similar strings, the similarity degree may be calculated earlier than that of the larger number of similar strings. When there is a string with a similarity degree equal to or higher than a predetermined value among the similar strings for which the similarity degree has been calculated earlier, the information on the disease name corresponding to the similar string may be output and the process may be terminated.
[0143] Here, the processing in the text data analysis system 1 according to the second embodiment will be described while showing a specific example. As a first example, as shown in FIG. 14(a), the first search means 23 searches the disease name master to see if a string matching the first keyword, which is the string "suspected acute influenza pneumonia" obtained by the data acquisition means 11 and extracted by the string extraction means 12, is hit.
[0144] When the search by the first search means 23 fails to hit a string matching the first keyword, the first string generation means 24 generates, as shown in FIG. 14(b), a string "acute influenza pneumonia" obtained by removing the suffix "suspected" from the first keyword with reference to the suffix master.
[0145] The second search means 25 searches the disease name master to check whether a character string matching the second keyword, which is the character string "acute influenza pneumonia" generated by the first character string generation means 24, is found.
[0146] When the search by the second search means 25 does not find a character string matching the second keyword, the second character string generation means 26 refers to the prefix master and generates a character string "influenza pneumonia" obtained by removing the prefix "acute" from the second keyword, as shown in Fig. 14(c).
[0147] The third search means 27 searches the disease name master to check whether a character string matching the third keyword, which is the character string "influenza pneumonia" generated by the second character string generation means 26, is found. When the search by the third search means 27 finds a character string matching the third keyword, the third search means 27 outputs information on the disease name corresponding to the found character string. Specifically, as shown in Fig. 14(d), the third search means 27 outputs information on the disease name "influenza pneumonia" and the ICD code "J110".
[0148] As a second example, as shown in Fig. 15(a), the first search means 23 searches the disease name master to check whether a character string matching the first keyword, which is the character string "postoperative right otitis media with cholesteatoma" obtained by the data acquisition means 11 and extracted by the character string extraction means 12, is found.
[0149] When the search by the first search means 23 does not find a character string matching the first keyword, the first character string generation means 24 refers to the suffix master and generates a character string "right otitis media with cholesteatoma" obtained by removing the suffix "postoperative" from the first keyword, as shown in Fig. 15(b).
[0150] The second search means 25 searches the disease name master to check whether a character string matching the second keyword, which is the character string "right otitis media with cholesteatoma" generated by the first character string generation means 24, is found.
[0151] When the second string generation means 26 does not find a string that matches the second keyword as a result of the search by the second search means 25, as shown in FIG. 15(c), it refers to the prefix master and generates a string "otitis media with pearl" obtained by removing the prefix "right" from the second keyword.
[0152] The third search means 27 searches the disease name master to find out whether a string that matches the third keyword hits, using the string "otitis media with pearl" generated by the second string generation means 26 as the third keyword.
[0153] When the element string extraction means 28 does not find a string that matches the third keyword as a result of the search by the third search means 27, as shown in FIG. 15(d), it refers to the element string master and extracts the element strings "pearl" and "middle ear" from the third keyword.
[0154] Since there are a plurality of element strings extracted by the element string extraction means 28, the similarity string extraction means 29 first uses the element string "middle ear" that is located closer to the center of the third keyword as the fourth keyword, and extracts from the disease name master a string (similarity string) that includes the fourth keyword "middle ear" as shown in FIG. 16(a).
[0155] As shown in FIG. 16(b), the similarity calculation means 30 calculates the similarity of each of the extracted similarity strings with the third keyword "otitis media with pearl".
[0156] The output means 31 outputs information on the disease name corresponding to the similarity string among the similarity strings extracted by the similarity string extraction means 29, for which the similarity calculated by the similarity calculation means 30 is a predetermined value, for example, 0.8 or more. Specifically, the output means 31 outputs information on the disease name "cholesteatomatous otitis media" and the ICD code "H71", the disease name "acute otitis media" and the ICD code "H669", and the disease name "chronic otitis media" and the ICD code "H669". Note that the output means 31 may be configured to further output information on the similarity calculated by the similarity calculation means 30.
[0157] Next, an example of the operation of the text data analysis system 1 according to the second embodiment will be described with reference to a flowchart.
[0158] As shown in FIG. 17, the text data analysis system 1 searches the disease name master to check whether a character string matching the first keyword (the character string extracted in the character string extraction step (S120)) is hit (S231) (first search step). If, as a result of the search, a character string matching the first keyword is hit (S232, Yes), the process proceeds to step S304.
[0159] On the other hand, if, as a result of the search in the first search step, a character string matching the first keyword is not hit (S232, No), the text data analysis system 1 refers to the suffix master and generates a character string obtained by removing the suffix from the first keyword (S240) (first character string generation step).
[0160] Next, the text data analysis system 1 searches the disease name master to check whether a character string matching the second keyword (the character string generated in the first character string generation step) is hit (S251) (second search step). If, as a result of the search, a character string matching the second keyword is hit (S252, Yes), the process proceeds to step S304.
[0161] On the other hand, if, as a result of the search in the second search step, a character string matching the second keyword is not hit (S252, No), the text data analysis system 1 refers to the prefix master and generates a character string obtained by removing the prefix from the second keyword (S260) (second character string generation step).
[0162] Next, the text data analysis system 1 searches the disease name master to check if a string that matches the string generated in the second string generation step hits as the third keyword (S271) (third search step). And if, as a result of the search, a string that matches the third keyword hits (S272, Yes), the process proceeds to step S304.
[0163] On the other hand, if, as a result of the search in the third search step, a string that matches the third keyword does not hit (S272, No), the text data analysis system 1 refers to the element string master and extracts element strings from the third keyword (S280) (element string extraction step).
[0164] Next, the text data analysis system 1 determines if there are multiple extracted element strings (S291). If there are multiple extracted element strings (S291, Yes), the text data analysis system 1 determines the element string that extracts similar strings first, specifically, the element string located closer to the center of the third keyword (S292), and uses the determined element string as the fourth keyword to extract a string (similar string) including the fourth keyword from the disease name master (S293). Also, if there is one extracted element string (S291, No), the text data analysis system 1 uses the said element string as the fourth keyword to extract a similar string from the disease name master (S293) (similar string extraction step).
[0165] Next, the text data analysis system 1 calculates the similarity between the extracted similar character strings and the third keyword (S301) (similarity calculation step). Next, the text data analysis system 1 determines whether there is a similar character string whose calculated similarity is equal to or greater than a predetermined value (S302). If there is no similar character string whose similarity is equal to or greater than the predetermined value (S302, No), the process returns to step S292 to determine the next element string for extracting the similar character string, and the subsequent processing is executed. On the other hand, in step S302, if there is a similar character string whose similarity is equal to or greater than the predetermined value (Yes), the text data analysis system 1 extracts the similar character string whose similarity is equal to or greater than the predetermined value (S303) and proceeds to step S304.
[0166] In step S304, the text data analysis system 1 determines whether there is a group of character strings representing one item next. And if there is the following character string (S304, Yes), the process returns to step S120 to execute the subsequent processing, and if there is no following character string (S304, No), the process proceeds to step S311.
[0167] In step S311, the text data analysis system 1 extracts the information of the disease name corresponding to the character string hit in the first search step, the second search system, or the third search step, and outputs the extracted disease name information (S312). Or, the text data analysis system 1 extracts the information of at least one disease name corresponding to the similar character string extracted in the similar character string extraction step (S311), and outputs the extracted disease name information (S312) (output step).
[0168] According to the above second embodiment, it is possible to extract the character string representing the disease name from the text data without adding keywords to the master more than necessary. In addition, the character string representing the disease name can be extracted by converting it into the format described in the ICD10-compatible standard disease name master.
[0169] In addition, by further providing the element string extraction means 28, the similar string extraction means 29, and the output means 31, it is possible to more reliably extract the string representing the disease or injury name from the text data.
[0170] In addition, by further providing the similarity calculation means 30, it is possible to narrow down and output the information of the disease or injury name.
[0171] Also, when there is a similar string with a similarity of a predetermined value or more among the previously calculated similar strings, the information of the disease or injury name corresponding to the similar string is output and the process is terminated. Therefore, the processing amount until the information of the disease or injury name is output can be reduced and the processing speed can be increased.
[0172] Note that in the second embodiment, the disease or injury name master stores the ICD code, but it may store other codes, for example, the disease name management number of the ICD10-compatible standard disease name master. Also, in the second embodiment, the disease or injury name master is created based on the ICD10-compatible standard disease name master, but it may be created based on, for example, the disease or injury name master of the basic master.
[0173] As described above, the embodiments have been explained, but the present invention is not limited to the above embodiments and can be appropriately modified and implemented as exemplified below.
[0174] For example, in the above embodiment, when there are a plurality of extracted element strings, for the element string located closer to the center of the predetermined keyword, the similarity is calculated earlier than other element strings, but it is not limited to this. For example, when there are a plurality of extracted element strings, for each of the extracted element strings, similar strings are extracted, and for the one with a smaller number of extracted similar strings than others, the similarity is calculated earlier. When there is a string with a similarity of a predetermined value or more among the similar strings for which the similarity has been calculated earlier, the information such as the medical act corresponding to the similar string may be output and the process may be terminated. Also, the element string master may store the element string in correspondence with a number or the like representing the priority order when calculating the similarity.
[0175] Also, in the above-described embodiment, the similarity calculation means calculates the similarity between the similar character strings extracted by the similar character string extraction means and a predetermined keyword, and the output means outputs information corresponding to the similar character strings for which the similarity calculated by the similarity calculation means is equal to or greater than a predetermined value. However, the present invention is not limited to this. For example, the output means may be configured to output information corresponding to the similar character strings extracted by the similar character string extraction means without calculating the similarity if the number of similar character strings extracted by the similar character string extraction means is equal to or less than a predetermined value.
[0176] In addition, each of the elements described in the above-described embodiment and modification examples may be implemented in any combination. Further, for example, when implementing a combination of the first embodiment and the second embodiment, the prefix master of the first embodiment and the prefix master of the second embodiment may be a common master, and the element character string master of the first embodiment and the element character string master of the second embodiment may be a common master.
Explanation of Reference Numerals
[0177] 1 Text data analysis system 11 Data acquisition means 12 Character string extraction means 13 First search means 14 First character string generation means 15 Second search means 16 Second character string generation means 17 Similar character string extraction means 18 First similarity calculation means 19 Output means 20 Second similarity calculation means 21 Element character string calculation means 22 Third similarity calculation means 23 First search means 24 First character string generation means 25 Second search means 26 Second character string generation means 27 Third search means 28 Element character string extraction means 29 Similar character string extraction means 30 Similarity calculation means 31 Output means
Claims
1. A data acquisition means for acquiring text data; a character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; a first search means for searching, using the character string extracted by the character string extraction means as a first keyword, whether a character string matching the first keyword is found in an injury / illness name master that stores character strings representing injury / illness names, and for outputting information on the injury / illness name corresponding to the character string that is found as a result of the search; a first character string generating means for generating a character string obtained by removing the suffix from the first keyword by referring to a suffix master storing a suffix that is a specific character string added to the end of a character string when a character string matching the first keyword is not found as a result of a search by the first search means; a second search means for searching the injury / illness name master for a character string that matches the second keyword using the character string generated by the first character string generation means as a second keyword, and for outputting information on the injury / illness name corresponding to the character string that matches the second keyword when the search results in a character string that matches the second keyword; a second character string generating means for generating a character string obtained by removing the prefix from the second keyword when a character string matching the second keyword is not found as a result of the search by the second search means, by referring to a prefix master storing a prefix which is a specific character string added to the beginning of a character string; a third search means for searching the illness / injury name master for a character string that matches the third keyword, using the character string generated by the second character string generation means as a third keyword, and for outputting information on the illness / injury name corresponding to the matched character string when a character string that matches the third keyword is found as a result of the search.
2. an element character string extracting means for extracting at least one element character string from the third keyword by referring to an element character string master that stores an element character string that is a specific character string included in the name of an injury or illness, when a character string matching the third keyword is not found as a result of the search by the third search means; a similar character string extracting means for extracting a character string including the fourth keyword from the injury / disease name master using the element character string extracted by the element character string extracting means as a fourth keyword; 2. The text data analysis system according to claim 1, further comprising: an output unit that outputs information on at least one injury or illness name corresponding to the character string extracted by the similar character string extraction unit.
3. a similarity calculation unit that calculates a similarity between the character string extracted by the similar character string extraction unit and the third keyword, The text data analysis system according to claim 2, characterized in that the output means outputs information on names of illnesses or injuries corresponding to character strings having a similarity calculated by the similarity calculation means that is equal to or greater than a predetermined value, among the character strings extracted by the similar character string extraction means.
4. the similarity calculation means, when there are a plurality of element strings extracted by the element string extraction means, calculates a similarity between the element string located near the center of the third keyword and the third keyword for the string extracted by the similar string extraction means as a fourth keyword before calculating a similarity between the element string and the third keyword for the element strings extracted by the similar string extraction means, The text data analysis system according to claim 3, characterized in that, when a character string having a similarity equal to or higher than a predetermined value is included in the character strings whose similarity has been previously calculated, the output means outputs information on the name of an injury or illness corresponding to the character string.
5. The computer includes: A data acquisition step of acquiring text data; a character string extraction step of extracting a group of character strings representing one item from the text data acquired in the data acquisition step; a first search step of searching for a character string that matches the first keyword from an injury or illness name master that stores character strings representing injury or illness names, using the character string extracted in the character string extraction step as a first keyword, and outputting information on the injury or illness name corresponding to the hit character string when a character string that matches the first keyword is found as a result of the search; a first character string generating step of, when a character string matching the first keyword is not found as a result of the search in the first search step, generating a character string by removing the suffix from the first keyword by referring to a suffix master that stores a suffix, which is a specific character string added to the end of a character string; a second search step of searching the injury / disease name master for a character string that matches the second keyword using the character string generated in the first character string generation step as a second keyword, and outputting information on the injury / disease name corresponding to the character string that matches the second keyword when the search result shows that the character string matches the second keyword; a second character string generating step of generating a character string obtained by removing the prefix from the second keyword by referring to a prefix master that stores a prefix, which is a specific character string added to the beginning of a character string, when a character string matching the second keyword is not found as a result of the search in the second search step; a third search step of using the character string generated in the second character string generation step as a third keyword to search the injury / illness name master for a character string matching the third keyword, and if a character string matching the third keyword is found as a result of the search, outputting information on the injury / illness name corresponding to the hit character string.
6. Computer, A data acquisition means for acquiring text data; a character string extraction means for extracting a group of character strings representing one item from the text data acquired by the data acquisition means; a first search means for searching, using the character string extracted by the character string extraction means as a first keyword, whether a character string matching the first keyword is found in an injury / illness name master that stores character strings representing injury / illness names, and for outputting information on the injury / illness name corresponding to the character string that is found as a result of the search; a first character string generating means for generating a character string obtained by removing the suffix from the first keyword by referring to a suffix master storing a suffix that is a specific character string added to the end of a character string when a character string matching the first keyword is not found as a result of a search by the first search means; a second search means for searching the injury / illness name master for a character string that matches the second keyword using the character string generated by the first character string generation means as a second keyword, and for outputting information on the injury / illness name corresponding to the character string that matches the second keyword when the search results in a character string that matches the second keyword; a second character string generating means for generating a character string obtained by removing the prefix from the second keyword when a character string matching the second keyword is not found as a result of the search by the second search means, by referring to a prefix master storing a prefix which is a specific character string added to the beginning of a character string; A computer program characterized in that the computer program functions as a third search means for searching the illness / injury name master for a character string that matches the third keyword, using the character string generated by the second character string generation means as a third keyword, and for outputting information on the illness / injury name corresponding to the hit character string when a character string that matches the third keyword is found as a result of the search.
Citation Information
Patent Citations
Document information retrieving device
JP1991286371A
Information retrieving device
JP1992133173A
Disease name identifying device
WO2006121115A1
Document reading device and document reading processing program
JP3349699B2