Information processing device, information processing method, and program
By accessing multiple insurance databases to calculate patient ratios and generate training data that matches these ratios, the method addresses biases in the Japanese medical insurance system, improving the predictive accuracy of machine learning models.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- JMDC CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-05-13
AI Technical Summary
The Japanese medical insurance system's complexity, with multiple insurance types and varying demographics, leads to biased training data when using a single insurer database, reducing the predictive accuracy of machine learning models.
An information processing device accesses multiple health insurance databases to calculate patient ratios for different insurance types, generating training data by sampling to match these ratios, thereby reducing bias and improving predictive accuracy.
This approach allows for the creation of high-quality training data that reduces bias, enhancing the accuracy of machine learning models in predicting health outcomes.
Smart Images

Figure 0007858124000001_ABST
Abstract
Description
Technical Field
[0004] , , , ,
[0005] , , ,
[0003] , ,
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Conventionally, there has been known a technique of acquiring, from a user's communication device, description information of a document received by a user who is an insured person from a medical institution and having no description of a disease name, and predicting a corresponding disease from information on medical treatment acts or pharmaceuticals received by the user at the medical institution based on the description information (Patent Document 1). In this technique, a machine learning model based on data associating previously acquired information on medical treatment acts or pharmaceuticals with disease information is used to predict a disease corresponding to the user's medical treatment act or pharmaceutical information.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] This invention has been made in view of the above problems, and its objective is to realize a technology that can generate training data with reduced or removed bias, thereby improving the prediction accuracy of machine learning models. [Means for solving the problem]
[0007] To solve this problem, for example, the information processing device of the present invention is an information processing device that can access a database storing multiple health insurance information based on different insurance types, A determination means that, for each type of insurance, obtains the number of insured persons and a predetermined number of patients from the stored health insurance information, and determines the ratio of the predetermined number of patients to the number of insured persons for that type of insurance, A calculation means for calculating the estimated number of patients by accumulating the aforementioned ratios for each type of insurance onto a predetermined base, The system includes a sampling means for generating training data for machine learning by sampling from the database in a manner that corresponds to the results of comparing the estimated number of patients across different insurance types. [Effects of the Invention]
[0008] According to the present invention, it is possible to generate training data with reduced or removed bias, thereby improving the predictive accuracy of machine learning models. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an example of a support system according to an embodiment of the present invention. [Figure 2] Block diagram showing an example of the functional configuration of the information processing device according to this embodiment. [Figure 3] This figure illustrates an example of a comparison table according to this embodiment. [Figure 4] Block diagram showing an example of the functional configuration of the communication device according to this embodiment. [Figure 5] A diagram illustrating an example of a training data creation support process. [Figure 6] Flowchart illustrating the operation of the learning data creation support process according to this embodiment. [Figure 7A] This diagram illustrates an example of a health insurance database managed by a health insurance association according to this embodiment. [Figure 7B] This figure illustrates an example of a national health insurance database according to this embodiment. [Figure 8A] This diagram illustrates an example of a sample source database extracted from a health insurance database managed by a health insurance association, using a specific drug as the key. [Figure 8B] This diagram illustrates an example of a sample source database extracted from the National Health Insurance Database using a specific drug as a key. [Modes for carrying out the invention]
[0010] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims, and not all combinations of features described in the embodiments are essential to the invention. Two or more of the features described in the embodiments may be combined in any way. Furthermore, identical or similar configurations will be given the same reference numeral, and redundant descriptions will be omitted.
[0011] <Configuration of the support system> First, with reference to Figure 1, the configuration of the support system 10 according to an embodiment of the present invention will be described. The support system 10 is a system that assists in the creation of training data used for machine learning, and includes, for example, an information processing device 100 and a communication device 110 used by a user 20. The user 20 includes, for example, a user who checks and analyzes the number of patients with illnesses or injuries or the number of prescriptions for medicines at a pharmaceutical company, medical institution, or insurance institution. The user 20 checks the number of patients nationwide, or the number of patients under specific conditions (specific age or gender), or checks the trend of the number of patients over time. The information processing device 100 is a server managed by a company (also called a service company) that operates a service that provides medical data such as information on the number of patients with illnesses or injuries or information on the number of prescriptions for medicines. The service company generates and provides medical data that is valuable to its customers, i.e., users 20, by analyzing and appropriately processing raw data collected from insurers and medical institutions.
[0012] Service companies receive data from insurers in the medical insurance system, medical institutions such as hospitals, and pharmacies such as dispensing pharmacies. The data received from insurers is health insurance information related to the medical care of members, and includes, for example, medical claims, dispensing claims, health checkup results, and member ledgers. Health insurance information may include information that can be extracted from claims, such as information showing when, at which medical institution, what illness or injury a member was diagnosed with, what treatment they received, and what medications they were prescribed. The data received from medical institutions includes, for example, medical claims, DPC (Diagnosis Procedure Combination) survey data, and electronic medical record data. The data received from pharmacies includes, for example, dispensing claims and prescriptions.
[0013] In this embodiment, the information processing apparatus 100 has a database that stores a plurality of health insurance information based on a plurality of different insurance types in the medical insurance system, and is configured to be accessible thereto. In other embodiments, the information processing apparatus 100 may not have a database but may be configured to be accessible to the database. In this case, the information processing apparatus 100 may appropriately access an external database constructed by collecting the above information from an insurer of social insurance, a local government, or the like by an enterprise different from the service enterprise.
[0014] The database possessed by the information processing apparatus 100 holds health insurance information for each insurance type. The database may be a set of databases provided for each insurance type, or may be configured such that each record of a single database has a field for the insurance type. Hereinafter, the former case will be described.
[0015] The information processing apparatus 100 acquires health insurance information such as receipt data from the insurer (health insurance union) 130 of the union-administered health insurance, the insurer (local government) 140 of the national health insurance, and the insurer (local government) 150 of the late-stage elderly medical system, and stores it in the corresponding database of itself. For the purpose of only illustration, in this embodiment, it is assumed that the information processing apparatus 100 does not acquire health insurance information from the insurer (mutual aid association) 160 of the mutual aid association insurance nor from the insurer (National Health Insurance Association) 170 of the association health insurance. In such an example, the number of patients and the like regarding the health insurance information that has not been acquired can be calculated based on the prescription rate and the like of the union-administered insurance health insurance by the processing described later. On the other hand, health insurance information may also be acquired from the insurer (mutual aid association) 160 of the mutual aid association insurance and the insurer (National Health Insurance Association) 170 of the association health insurance. In this embodiment, the insurance type refers to the insurance type in the current medical insurance system in Japan, and takes any value of union-administered health insurance, national health insurance, late-stage elderly medical system, mutual aid association insurance, and association health insurance. However, it is obvious to those skilled in the art mentioned in this specification that the technical idea of this embodiment can also be applied even when there are future changes or new establishments in the insurance type.
[0016] Service companies possess multiple databases with different coverages and characteristics. By referring to these multiple databases, service companies can provide analysis results with higher resolution than when relying on a single database. For example, according to a specified key, data can be extracted and integrated from multiple databases of different insurance types, and one large learning data for machine learning can be generated. It is expected that a machine learning model trained with this learning data will have a wider application range than a machine learning model trained with learning data generated from a single database, and the probability of improving both the accuracy and precision of prediction increases.
[0017] The information processing device 100 receives the specification of a key corresponding to the purpose of the machine learning model, calculates the ratio of the number of patients corresponding to the key for each insurance type based on the database it owns, and calculates the estimated number of patients from the ratio and external statistical data. The information processing device 100 determines the number of samplings from the databases of each insurance type so as to correspond to the ratio of the estimated number of patients, and generates learning data by performing sampling with the determined number of samplings. Thereby, it is possible to generate learning data with reduced bias (bias), as if randomly sampled from all over Japan. In other words, it is possible to generate high-quality learning data like a "miniature Japan" that corrects the bias of the database.
[0018] The information processing device 100 can communicate via a network 120 such as the Internet with a communication device 110. In this embodiment, the case where the information processing device 100 is a server on the cloud will be described as an example. However, the information processing device 100 is not limited to this, and may be realized by a virtual machine, an edge server, a computer placed on a small-scale network, or the like. Also, in this embodiment, the case where the communication device 110 is, for example, a personal computer will be described as an example. However, the communication device 110 is not limited to this, and may be, for example, another terminal capable of using a service for estimating the number of patients, such as a smartphone or a tablet-type terminal, using a browser or a dedicated application.
[0019] <Configuration of the information processing device> Next, an example of the functional configuration of the information processing device 100 will be described with reference to Figure 2. Note that each of the functional blocks described may be integrated or separated, and the functions described may be implemented in other blocks. Furthermore, what is described as hardware may be implemented in software, and vice versa.
[0020] The communication unit 201 includes a communication circuit that communicates with various devices via a network. The communication unit 201 receives information processed by the control unit 204 from the communication partner device (e.g., communication device 110) and transmits information processed by the control unit 204 to the communication partner device. The power supply unit 202 is a power supply that provides the power necessary for the operation of the information processing device 100.
[0021] The storage unit 203 includes, for example, a non-volatile storage medium such as a hard disk or semiconductor memory, and includes various programs executed by the control unit 204 of the information processing device 100, as well as various data and databases (DBs) used by the control unit 204. The various programs include a program for executing the learning data creation support process according to this embodiment, as well as an operating system, framework, libraries, etc. The various data includes, for example, the setting values of the information processing device 100. The storage unit 203 also stores the union-managed health insurance DB 231, the national health insurance DB 232, the late-stage elderly health insurance DB 233, statistical data 234, learning data 235, a comparison table 236, and a machine learning model 237.
[0022] The union-managed health insurance DB231 includes health insurance information based on the union-managed health insurance described above. The national health insurance DB232 includes health insurance information based on the national health insurance described above. The late-stage elderly health insurance DB233 includes health insurance information based on the late-stage elderly medical care system described above. The statistical data 234 is generated based on statistical data published by a third party other than the service company, such as an insurer or public institution, and holds data such as the total number of members for each type of insurance, and further broken down by member attributes, such as gender and / or age group. In this embodiment, the example will be described using gender and age group as member attributes, but member attributes may also include residential area. Residential area may be classified, for example, by prefecture or by municipality. The learning data 235 is generated from the health insurance information held in the union-managed health insurance DB231, national health insurance DB232, and late-stage elderly health insurance DB233 by the learning data creation support process according to this embodiment. The machine learning model 237 is a supervised machine learning model and is trained using the training data 235. The machine learning model 237 may be implemented using known supervised machine learning techniques such as GPT (Generative Pre-trained Transformer)-3, GPT-3.5, GPT-4 provided by openAI, LLaMA (Large Language Model Meta AI), and Bloom (BigScience Language Open-science Open-access Multilingual) provided by Meta.
[0023] The control unit 204 includes a central processing unit (CPU) 210 and RAM 211. The control unit 204 controls the operation of various parts within the control unit 204 and the operation of various parts of the information processing device 100 by loading and executing programs stored in the memory unit 203 into the RAM 211. The control unit 204 also performs learning data creation support processing using the various parts within the control unit 204.
[0024] RAM211 includes a volatile storage medium such as DRAM, and temporarily stores parameters and processing results for the control unit 204 to execute the program.
[0025] The key reception unit 220 receives a key specification from the user that corresponds to the purpose of the machine learning model 237. The key reception unit 220 may also receive the key specification by displaying a screen (not shown) prompting the user to input a key and obtaining the key entered by the user on that screen. The key may be a specific value of a field included in the health insurance information. For example, if the machine learning model 237 is a model that makes predictions about a particular drug, that particular drug may be specified as the key. Alternatively, if the machine learning model 237 is a model that makes predictions about a particular illness or injury, that particular illness or injury may be specified as the key. The following describes the case in which a particular drug is specified as the key. However, it will be apparent to those skilled in the art who have read this specification that the technical idea of this embodiment is also applicable when a particular illness or injury or other field value is specified as the key. Furthermore, the specific drug may be arbitrary, and may be one, multiple, or all of them. Moreover, in this embodiment, the case in which the key specification is received from the user will be described as an example. However, the key specification may be predetermined to be a specific drug or the like, and processing may be performed without receiving a key specification from the user. Alternatively, if no key is specified, all drugs may be processed, or other processing may be performed. Specifying a key increases the number of patients associated with that specific key, improving accuracy compared to creating a dataset without a key. On the other hand, not specifying a key results in a dataset that is not limited to a specific population, thus reducing costs.
[0026] The decision unit 221 retrieves the number of subscribers in the database and the number of patients corresponding to a specified specific medicine from the corresponding database for each insurance type for which there is a database corresponding to the information processing device 100, and determines the prescription rate, which is the ratio of the number of patients corresponding to the specific medicine to the number of subscribers for that insurance type. The number of patients corresponding to a specific medicine may be the number of patients who were prescribed the specific medicine (hereinafter referred to as the number of prescribed patients), or it may be the number of times the specific medicine was prescribed, but the former case will be explained below. The decision unit 221 refers to the union-managed health insurance DB 231, the national health insurance DB 232, and the late-stage elderly health insurance DB 233, and counts the number of prescribed patients for each attribute (gender, age group, and / or residential area). The decision unit 221 registers the count result as the number of prescribed patients in the company's DB in the comparison table 236. The decision unit 221 refers to the union-managed health insurance DB 231, the national health insurance DB 232, and the late-stage elderly health insurance DB 233, and counts the number of subscribers for each gender and age group. The decision unit 221 calculates the prescription rate for each insurance type, gender, and age group by dividing the number of patients prescribed by the company's database by the number of subscribers, and registers it in the comparison table 236. The prescription rates corresponding to insurance types and subscriber attributes will be various values, and these differences indicate differences in the characteristics of insurance types and subscriber attributes for a particular drug.
[0027] Figure 3 illustrates an example of a comparison table 236 according to this embodiment. The comparison table 236 includes, as an example, insurance type 901, gender 902, age group 903, number of patients prescribed by the company's database 904, prescription rate 905, estimated number of patients prescribed nationwide 906, relative ratio of the number of patients prescribed nationwide 907, and number of patients prescribed after sampling 908.
[0028] Returning to Figure 2, the calculation unit 222 calculates the estimated number of prescription patients nationwide by multiplying the prescription rate by a predetermined denominator for each insurance type, gender, and age group, and registers it in the comparison table 236. The predetermined denominator is the number of subscribers corresponding to each insurance type, gender, and age group, obtained from the statistical data 234. For example, the calculation unit 222 obtains the number of male subscribers aged 5-9 under health insurance managed by a health insurance association from the statistical data 234, and calculates the estimated number of prescription patients nationwide aged 5-9 under health insurance managed by a health insurance association by multiplying this number of subscribers by the prescription rate for male subscribers aged 5-9 under health insurance managed by a health insurance association.
[0029] The calculation unit 222 calculates the estimated number of prescription patients nationwide for insurance types for which health insurance information is not held in the database of the information processing device 100 by multiplying a predetermined denominator for that insurance type by the prescription rate of one of the insurance types for which a database corresponding to the information processing device 100 exists, and registers this calculation in the comparison table 236. For example, for the Japan Health Insurance Association (Kyokai Kenpo), for which the information processing device 100 does not have information, the prescription rate of a union-managed health insurance, which is for the same type of worker, is used. In this case, the calculation unit 222 obtains the number of members of a predetermined gender and age group for Kyokai Kenpo from statistical data 234, and calculates the estimated number of prescription patients nationwide for a predetermined gender and age group for Kyokai Kenpo by multiplying this number of members by the prescription rate of a predetermined gender and age group for a union-managed health insurance.
[0030] The sampling unit 223 generates training data 235 for machine learning by sampling from databases 231, 232, and 233 to correspond to the results of comparing the estimated number of prescription patients nationwide between different insurance types and between different subscriber attributes. The sampling unit 223 identifies the minimum non-zero (0) value among the number of prescription patients in the company's database corresponding to each insurance type, gender, and age group pair, and sets the identified minimum value as the criterion for determining the sampling size. The sampling unit 223 refers to the comparison table 236 and calculates the ratio between the estimated number of prescription patients nationwide corresponding to all insurance type, gender, and age group pairs as the relative ratio of prescription patients nationwide. In this case, the value of the relative ratio of prescription patients nationwide for the insurance type, gender, and age group pair that gives the minimum value set as the criterion for determining the sampling size is set to 1. For each insurance type, gender, and age group pair, the sampling unit 223 calculates the number of prescription patients after sampling by multiplying the set minimum value by the relative ratio of prescription patients nationwide and registers it in the comparison table 236. In this way, the number of prescription patients after sampling is determined according to the relative ratio of prescription patients nationwide.
[0031] By setting the minimum number of patients prescribed by the company's database as the criterion for determining the sampling size, it is possible to prevent or suppress the number of patients prescribed after sampling from exceeding the number of patients prescribed by the company's database.
[0032] In the example in Figure 3, the minimum number of patients prescribed in the company's database with insurance type "National Health Insurance", gender "Male", and age group "0-4 years" is identified as "100 people", and its relative ratio to the total number of prescribed patients nationwide is set to 1. The sampling unit 223 calculates the number of prescribed patients after sampling for insurance type "National Health Insurance", gender "Male", and age group "0-4 years" as 100 people × 1 = 100 people, and calculates the number of prescribed patients after sampling for insurance type "Union-managed Health Insurance", gender "Male", and age group "5-9 years" as 100 people × 2 = 200 people.
[0033] The number of prescription patients after sampling only needs to correspond to the estimated number of prescription patients nationwide. In addition to calculating it by preserving the ratio as in the example above, the number of prescription patients after sampling can also be calculated by, for example, multiplying the relative ratio of prescription patients nationwide by a predetermined weight and then multiplying by the minimum value.
[0034] The sampling unit 223 extracts health insurance information for patients prescribed specific medications from the union-managed health insurance DB 231, the national health insurance DB 232, and the late-stage elderly health insurance DB 233, respectively, and registers it in the respective source databases. For each combination of insurance type, gender, and age group, the sampling unit 223 refers to the source database for that insurance type, randomly extracts health insurance information for the number of patients prescribed after sampling from the health insurance information of patients belonging to that gender and age group, and registers it in the training data 235. For health insurance associations whose health insurance information is not held in the database of the information processing device 100, the sampling unit 223 randomly extracts health insurance information for the number of patients prescribed after sampling from the source database of union-managed health insurance, and registers it in the training data 235.
[0035] In the example in Figure 3, the sampling unit 223 refers to the sampling source database for health insurance under the category of "Health Insurance under Union Management," gender "Male," and age group "0-4 years old." It randomly extracts health insurance information for 100 patients (the number of patients prescribed after sampling) from the health insurance information of 1000 patients belonging to the gender "Male" and age group "0-4 years old," and registers it in the training data 235. For the category of "National Health Insurance," gender "Male," and age group "0-4 years old," the number of patients prescribed after sampling is also 100. Since the National Health Insurance sampling source database also holds health insurance information for 100 patients of the gender "Male" and age group "0-4 years old," all of this information is extracted and registered in the training data 235.
[0036] The training data 235 generated by the sampling unit 223 is used to train a machine learning model 237 corresponding to a specific drug designated as a key. In this embodiment, the case where the information processing device 100 has the machine learning model 237 is described, but it is not limited to this, and the machine learning model 237 may reside in another device other than the information processing device 100, or it may be held by a third party other than the service company, such as the user 20.
[0037] <Communication device configuration> Furthermore, an example configuration of the communication device 110 will be described with reference to Figure 4. Figure 4 shows an example of the functional configuration of the communication device 110 in this embodiment. Note that each of the functional blocks described may be integrated or separated, and the functions described may be implemented in other blocks. Also, what is described as hardware may be implemented in software, and vice versa.
[0038] The communication unit 301 includes, for example, a communication circuit, and communicates with the information processing device 100 via mobile communication such as 5G or LTE, or via wireless communication such as WiFi, to send and receive necessary data.
[0039] The operation unit 303 includes a keyboard or touch panel and accepts operations on various operation screens displayed on the display unit 305. The display unit 305 includes a display panel such as an LCD or OLED and displays a GUI for various operations. For example, the display unit 305 receives information generated by the information processing device 100 and displays it on the display unit 305 via, for example, a web browser or a dedicated application. The display unit 305 may display the information received from the information processing device 100 as is, or it may select a part of the received information or reconfigure the received information to configure a display screen.
[0040] The recording unit 306 includes, for example, a non-volatile memory such as a semiconductor memory, and stores set user information, programs executed by the control unit 302, and the like.
[0041] The control unit 302 includes a CPU 310 and RAM 311. For example, the CPU 310 executes a program recorded in the recording unit 306 to control the operation of each functional block within the control unit 302 and each part within the communication device 110.
[0042] <Explanation of the training data creation support process> Next, with reference to Figure 5, a specific example of the learning data creation support process will be explained. In this example, patients are not divided by gender or age group, and the estimated number of prescription patients nationwide is calculated for each type of insurance. The process for calculating the estimated number of prescription patients nationwide for health insurance managed by a health insurance association will be explained. The calculation unit 222 obtains the number of members of health insurance managed by a health insurance association, 401, from the statistical data 234 in the storage unit 203.
[0043] Furthermore, the determination unit 221 obtains the number of members and the number of prescribed patients from the union-managed health insurance DB 231 stored in the memory unit 203, and determines the prescription rate 405 (for a specific drug) in the union-managed health insurance.
[0044] The calculation unit 222 calculates the estimated number of prescription patients nationwide for a specific drug under health insurance managed by health insurance managed by associations, 402, by multiplying the prescription rate 405 in health insurance managed by associations by the number of members 401 in those associations.
[0045] This section explains the process for calculating the estimated number of prescription patients nationwide for the Japan Health Insurance Association (Kyokai Kenpo). The calculation unit 222 obtains the number of Kyokai Kenpo members, 424, from the statistical data 234 in the storage unit 203.
[0046] The calculation unit 222 calculates the estimated number of patients nationwide prescribed for a specific drug by the Japan Health Insurance Association (Kyokai Kenpo) by multiplying the prescription rate of 405 in health insurance managed by the association by the number of members of the Japan Health Insurance Association (Kyokai Kenpo) (424).
[0047] This section explains the process for calculating the estimated number of prescription patients nationwide under the National Health Insurance. The calculation unit 222 obtains the number of National Health Insurance subscribers, 407, from the statistical data 234 in the storage unit 203.
[0048] Furthermore, the determination unit 221 obtains the number of subscribers and the number of prescribed patients from the National Health Insurance DB 232 stored in the memory unit 203, and determines the prescription rate 411 (for a specific drug) in the National Health Insurance.
[0049] The calculation unit 222 calculates the estimated number of patients nationwide prescribed a specific drug under the National Health Insurance system (403) by adding the prescription rate (411) under the National Health Insurance system to the number of National Health Insurance subscribers (407).
[0050] This section explains the process for calculating the estimated number of prescription patients nationwide under the medical care system for the elderly. The calculation unit 222 obtains the number of subscribers to the medical care system for the elderly, 421, from the statistical data 234 in the memory unit 203.
[0051] Furthermore, the decision unit 221 obtains the number of subscribers and the number of prescribed patients from the late-stage elderly insurance DB 233 stored in the memory unit 203, and determines the prescription rate 422 (for a specific drug) in the late-stage elderly medical care system.
[0052] The calculation unit 222 calculates the estimated number of patients nationwide prescribed a specific drug under the medical care system for the elderly (423) by multiplying the prescription rate (422) under the medical care system for the elderly (421) by the number of members under the medical care system for the elderly (423).
[0053] Next, a series of operations for the learning data creation support process in the information processing device will be explained with reference to Figure 6. Each operation described in this series of operations is realized by the control unit 204 deploying the program stored in the storage unit 203 to the RAM 211 and executing it, thereby enabling the functions of each part of the control unit 204 described above.
[0054] In S601, the key reception unit 220 accepts the designation of a pharmaceutical product. In S602, the decision unit 221 determines the prescription rate for each member attribute based on the number of members in the health insurance union DB 231 and the number of patients prescribed the designated pharmaceutical product. The counting of the number of prescribed patients in the health insurance union DB 231 will be explained with reference to Figure 7A. Figure 7A is a diagram illustrating an example of the health insurance union DB 231. The health insurance information held in the health insurance union DB 231 includes, as an example, health insurance information identification information 701, medical institution number 702, consultation date 703, insured person identification information 704, information indicating diagnosed illness or injury 705, information indicating prescribed pharmaceutical product 706, and information indicating the insurer 707. For example, when the decision unit 221 obtains the number of patients prescribed a designated pharmaceutical product, it counts the number of patients in the information indicating prescribed pharmaceutical product 706 for which the designated pharmaceutical product is listed.
[0055] Returning to Figure 6, in S603, the calculation unit 222 calculates the estimated number of prescription patients nationwide by multiplying the prescription rate by the number of members of the health insurance association, obtained from public statistics, for each attribute of the member.
[0056] In S604, the decision unit 221 determines the prescription rate for each member attribute based on the number of members in the National Health Insurance DB 232 and the number of patients prescribed the specified medicine. The counting of the number of prescribed patients in the National Health Insurance DB 232 will be explained with reference to Figure 7B. Figure 7B is a diagram illustrating an example of the National Health Insurance DB 232. The health insurance information held in the National Health Insurance DB 232 includes, as an example, health insurance information identification information 711, medical institution number 712, consultation date 713, insured person identification information 714, information indicating diagnosed illness or injury 715, information indicating prescribed medicine 716, and information indicating the insurer 717. For example, when the decision unit 221 obtains the number of patients prescribed the specified medicine, it counts the number of patients in the information indicating prescribed medicine 716 for which the specified medicine is listed.
[0057] Returning to Figure 6, in S605, the calculation unit 222 calculates the estimated number of prescription patients nationwide by multiplying the prescription rate by the number of National Health Insurance members obtained from official statistics, for each attribute of the member.
[0058] In S606, the determination unit 221 determines the prescription rate for each member attribute of the medical system for the elderly based on the number of members in the elderly insurance DB 233 and the number of patients prescribed the designated medicine. In S607, the calculation unit 222 calculates the estimated number of prescribed patients nationwide for the medical system for the elderly by multiplying the prescription rate by the number of members of the medical system for the elderly obtained from public statistics for each member attribute.
[0059] In S608, the calculation unit 222 calculates the estimated number of prescription patients nationwide for the Japan Health Insurance Association (Kyokai Kenpo) by multiplying the prescription rate of health insurance managed by the association by the number of Kyokai Kenpo members obtained from public statistics, for each attribute of the member.
[0060] In S609, the sampling unit 223 determines the ratio of the estimated number of prescription patients nationwide calculated for each of the following: health insurance managed by a health insurance association, national health insurance, medical care system for the elderly, and health insurance association, as the relative ratio of prescription patients nationwide. In S610, the sampling unit 223 determines the number of prescription patients after sampling for each insurance type and for each attribute of the insured, according to the determined relative ratio of prescription patients nationwide. In S611, the sampling unit 223 performs sampling from each DB (the source database of sampling) according to the determined number of prescription patients after sampling, and in S612, the sampling unit 223 generates learning data 235 by integrating the health insurance information (records) obtained by sampling.
[0061] Figure 8A illustrates an example of a record from the source database extracted from the union-managed health insurance DB231 using a specific drug as the key, that matches the age of 40 and the gender of male. Figure 8B illustrates an example of a record from the source database extracted from the national health insurance DB232 using a specific drug as the key, that matches the age of 40 and the gender of male.
[0062] Health insurance managed by health insurance associations primarily covers employees of companies and their dependents who are under 75 years old. On the other hand, national health insurance primarily covers self-employed individuals, pensioners, the unemployed, and non-regular employees who are under 75 years old. Because the characteristics of the insured differ in this way, the characteristics of the health insurance information held in the respective databases 231 and 232, such as the prescription rate of specific drugs or the prevalence rate of specific illnesses and injuries, may also differ. For example, if a person develops a severe mental illness such as schizophrenia, it may become difficult for them to work, and they may be forced to leave the company they were employed by. As a result, the patient leaves the health insurance association they belonged to and switches to national health insurance. Therefore, if you only look at the data from health insurance associations, the number of schizophrenia patients may appear lower than it actually is, and conversely, the data from national health insurance may include a relatively large number of schizophrenia patients, creating a bias.
[0063] If we simply add up all the records shown in Figure 8A and all the records shown in Figure 8B to create training data, the probability of the disease "schizophrenia" existing in that training data is 3 / 9 = 33%. In contrast, if the processing according to this embodiment determines that the relative ratio of prescription patients nationwide for insurance type "Union-managed health insurance", gender "male", and age group "40 years old" is "1", and the relative ratio of prescription patients nationwide for insurance type "National health insurance", gender "male", and age group "40 years old" is "1", then training data is generated that includes all the records (2) in Figure 8B and 2 records randomly selected from the 7 records in Figure 8A. An example of the training data generated in this way is shown below. Age Gender Injury / Illness 40-year-old male with schizophrenia (from Figure 8B) 40-year-old male with schizophrenia (from Figure 8B) 40-year-old male suffering from depression (Figure 8A, second row) 40-year-old male with dementia (Figure 8A, 5th row) The probability of the disease "schizophrenia" being present in the training data is 2 / 4 = 50%. If the first record is selected by random sampling from Figure 8A, the probability of the disease "schizophrenia" being present in the training data becomes 3 / 4 = 75%. If the value of 33% in the case of simple summation is based on the characteristics of service companies, such as being able to collect a relatively large amount of data from health insurance managed by unions, and the characteristics of health insurance managed by unions, such as the fact that there are few people diagnosed with schizophrenia in the health insurance DB231, then the method according to this embodiment can reduce or eliminate the influence of these characteristics, and high-quality training data, such as that obtained by random sampling from patients nationwide, can be generated.
[0064] As described above, the information processing device 100 according to this embodiment accepts the specification of a key corresponding to the purpose of the machine learning model 237, calculates the ratio of the number of patients corresponding to the key for each insurance type based on its own databases 231, 232, and 233, and calculates the estimated number of patients from that ratio and external statistical data 234. The information processing device 100 determines the number of samples to be taken from each insurance type database 231, 232, and 233 to correspond to the ratio of the estimated number of patients, and generates training data 235 by performing sampling with the determined number of samples. This makes it possible to generate training data with reduced bias, as if randomly sampled from all of Japan.
[0065] Furthermore, in the information processing device 100 according to this embodiment, the estimated number of patients is calculated for each subscriber attribute, and the corresponding sampling number is determined. Therefore, the bias in the generated training data can be further reduced.
[0066] Furthermore, in the information processing device 100 according to this embodiment, for the National Health Insurance Association (Kyokai Kenpo), for which the information processing device 100 does not have information, the estimated number of patients is calculated using the prescription rate of union-managed health insurance, which is for the same type of worker. The corresponding number of samples is then determined, and sampling is performed from the union-managed health insurance DB 231. This allows sampling to be performed while taking into account insurance types for which the service company does not have information, thereby further reducing bias in the generated training data.
[0067] The training data 235 generated by the information processing device 100 according to this embodiment may be used to train a machine learning model that predicts indications from prescription receipt information, a machine learning model that predicts prescriptions from disease names, a machine learning model that recommends pharmaceuticals from disease names, or a machine learning model that predicts the demand for pharmaceuticals. Alternatively, the generated training data 235 itself may be provided to customers as a product, or database research may be conducted using the generated training data 235. Alternatively, a potential patient estimation model is a machine learning model that offers accuracy advantages by using the training data 235 generated by the information processing device 100 according to this embodiment. The potential patient estimation model identifies patients who actually have disease A but are not recorded as disease A on the medical claim form (= potential patients) from the disease, medication, and medical treatment data on the medical claim form. In conventional methods that generate training data from a single database, for example, the number of potential patients nationwide is calculated from the ratio of potential patients in the health insurance union DB 231 to the health insurance population and the Japanese population, but the calculation result will inherit the bias specific to health insurance unions. In contrast, by using the method according to this embodiment, potential patients are identified in MiniJapan, and the number of potential patients nationwide is estimated from the ratio of the MiniJapan population to the Japanese population, so the accuracy of the output of the potential patient estimation model trained on such training data may be improved. Also, compared to the case where it is implemented only with health insurance union DB 231, using MiniJapan allows implementation with a single model for multiple DBs, which is superior in terms of operation and maintenance.
[0068] Alternatively, by using the training data 235 generated by the information processing device 100 according to this embodiment, there are two machine learning models that are superior in terms of operation and maintenance: the health age model and the severe illness model. "Health age" (registered trademark) is an index for easily understanding one's health status (https: / / kenko-nenrei.jp / guide.html). The health age model predicts future medical expenses from health checkup results and replaces future medical expenses with an index called health age. When using conventional methods that generate training data from a single database, for example, a health age model trained only on the union-managed health insurance DB 231 has difficulty handling health checkup data from local governments. However, by using a model trained on training data that reflects MiniJapan generated by the method according to this embodiment, a single model can handle both health checkup data from union-managed health insurance and health checkup data from local governments, and since it is a self-contained single model, it is superior in terms of operation and maintenance. The severe illness model is a model that predicts the occurrence of future severe diseases (cardiovascular disease, cerebrovascular disease, renal failure requiring dialysis) from health checkups and medical claims. This model also excels in terms of operation and maintenance for the same reasons. Thus, the generated, less biased, and standardized training data 235 has diverse potential applications.
[0069] (Other embodiments) The above example described the case where a specific drug is used as the key, but a specific disease or injury may also be used as the key. In this case, for example, the decision unit 221 first calculates the prevalence rate (the ratio of the number of patients with the specific disease or injury to the number of insured persons) for each type of insurance from the data in each insurer database 231, 232, and 233. Then, by multiplying this prevalence rate by the number of insured persons in the national statistical data, the estimated number of patients with the specific disease or injury throughout Japan is calculated for each gender, age, and type of insurance. This allows the "ideal patient composition ratio" for all of Japan to be determined.
[0070] The calculation unit 222 samples data from each insurer database 231, 232, and 233 to match the calculated "ideal patient composition ratio". For example, if the composition ratio calculation results in a high proportion of schizophrenia patients belonging to the National Health Insurance, more data on those patients is extracted from the National Health Insurance database. The data extracted from each database in this way is integrated to generate training data 235. The machine learning model 237 trained using the training data 235 generated by this process achieves higher prediction accuracy compared to conventional models.
[0071] Furthermore, although the above embodiment also described the case where the key reception unit 220 does not accept a key specification, the following processing can also be performed in this case. For example, the number of members of union-managed health insurance, national health insurance, medical care system for the elderly, and Japan Health Insurance Association can be obtained from statistical data 234 for each attribute (gender, age group, and / or residential area), and the ratio of the number of members of each insurer for each attribute can be calculated. Then, members can be sampled from each DB, such as union-managed health insurance DB 231, to create a dataset so that it is equal to the calculated ratio of the number of members of each insurer for each attribute. As mentioned above, for Japan Health Insurance Association, data from union-managed health insurance DB 231, which is for the same workers, may be used. In this way, it becomes possible to generate datasets according to the composition ratio of Japan for each attribute and insurance type.
[0072] The invention is not limited to the embodiments described above, and various modifications and changes are possible within the scope of the gist of the invention. [Explanation of Symbols]
[0073] 100...Information processing device, 110...Communication device, 220...Key reception unit, 221...Determination unit, 222...Calculation unit, 223...Sampling unit
Claims
1. An information processing device that can access a database storing multiple health insurance information based on different insurance types, A determination means that, for each type of insurance, obtains the number of insured persons and a predetermined number of patients from the stored health insurance information, and determines the ratio of the predetermined number of patients to the number of insured persons for that type of insurance, A calculation means for calculating the estimated number of patients by accumulating the aforementioned ratios for each type of insurance onto a predetermined base, An information processing device comprising: sampling means for generating training data for machine learning by sampling from the database in a manner that corresponds to the results of comparing the estimated number of patients between different insurance types.
2. Furthermore, it includes a means for receiving key specifications from the user, The information processing apparatus according to claim 1, wherein the determination means obtains the number of subscribers and the number of patients corresponding to a specified key from the stored health insurance information for each type of insurance, and determines the ratio of the number of patients corresponding to the specified key to the number of subscribers for that type of insurance.
3. The aforementioned reception means accepts the designation of a specific pharmaceutical product as the key designation. The aforementioned determination means obtains the number of insured persons and the number of patients corresponding to a specified drug from the stored health insurance information for each type of insurance, and determines the prescription rate, which is the ratio of the number of patients corresponding to the specified drug to the number of insured persons for that type of insurance. The calculation means calculates the estimated number of prescribed patients by accumulating the prescription rate for each type of insurance onto a predetermined base number. The information processing apparatus according to claim 2, wherein the sampling means generates training data for machine learning by sampling from the database in accordance with the results of comparing the estimated number of prescribed patients between different insurance types.
4. The determination means determines the ratio for each type of insurance and for each attribute of the policyholder, The calculation means calculates the estimated number of patients by multiplying the ratio corresponding to the insurance type and the attribute of the insured, determined by the determination means, by a predetermined denominator corresponding to the insurance type and the attribute of the insured, for each insurance type and for each attribute of the insured. The information processing apparatus according to claim 1, wherein the sampling means performs sampling from the database to correspond to the results of comparing the estimated number of patients between different insurance types and between different subscriber attributes.
5. The information processing device according to claim 4, wherein the subscriber's attributes include gender, age group, and / or residential area.
6. The information processing apparatus according to claim 3, wherein the training data generated by the sampling means is used to train a machine learning model corresponding to the specific pharmaceutical product.
7. The information processing device according to claim 1, wherein the calculation means calculates an estimated number of patients by multiplying a predetermined denominator for an insurance type for which health insurance information is not held in the database by the ratio of one of the different insurance types.
8. The information processing apparatus according to claim 1, wherein the calculation means calculates an estimated number of patients by adding the ratio to the number of subscribers for each type of insurance obtained from statistical data provided by a public institution.
9. An information processing method performed in an information processing device that can access a database storing multiple health insurance information based on different insurance types, A determination step is to obtain the number of subscribers and a predetermined number of patients from the stored health insurance information for each type of insurance, and to determine the ratio of the predetermined number of patients to the number of subscribers for that type of insurance. A calculation process for calculating the estimated number of patients by accumulating the aforementioned ratios for each type of insurance onto a predetermined base, An information processing method comprising: a sampling step of generating training data for machine learning by sampling from the database in a manner that corresponds to the results of comparing the estimated number of patients between different insurance types.
10. A program for causing a computer to function as one of the means of an information processing apparatus according to any one of claims 1 to 8.