Data model configuration method for learning data, and learning data generation device

By using a filter to specify abstraction or detail levels of data items, the method addresses the challenge of setting detailed objective functions and explanatory variables, enabling the generation of learning data with tailored granularity for machine learning models, thus enhancing data integration and model suitability.

JP7710969B2Active Publication Date: 2025-07-22HITACHI LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021189669
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-22
Publication Date
2025-07-22
Estimated Expiration
2041-11-22

AI Technical Summary

Technical Problem

Existing methods for generating training data for machine learning models lack direction for partial abstraction or subdivision, making it difficult to set detailed objective functions and explanatory variables according to the type of inference required.

Method used

A method and device for generating learning data that utilize a filter to specify abstraction or detail levels of data items, allowing for the classification of data into target and explanatory variables, and enable partial abstraction or subdivision based on a hierarchical structure of databases.

Benefits of technology

Enables the generation of learning data with arbitrarily changed data granularity, suitable for the specific use and purpose of the machine learning model, facilitating efficient and effective data integration from multiple fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710969000001
    Figure 0007710969000001
  • Figure 0007710969000002
    Figure 0007710969000002
  • Figure 0007710969000003
    Figure 0007710969000003
Patent Text Reader

Abstract

To enable partial abstraction or fragmentation when learning data is generated from existing data.SOLUTION: Disclosed is a data model configuration method for configuring a data model for learning data for machine learning, in which, when a data item of a database based on the learning data has a hierarchical structure of an abstraction degree or a detail degree, the specification of the abstraction degree or the detail degree of the data item is made possible for each data item, and a filter for distributing the data items to an objective variable and an explanatory variable is used to configure a data model for extracting the data item to be used for the learning data from the database.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the generation of training data used in machine learning, and particularly to a technique for generating training data by converting data in a predetermined format into data in a desired format.

Background Art

[0002] In recent years, inferences using machine learning models have been put into practical use. A machine learning model is learned by training data and functions as a function approximator that obtains a predetermined output (answer) for a predetermined input (problem). Configurations such as Deep Neural Network (DNN) that make up a machine learning model and machine learning techniques for learning it are known.

[0003] Regarding inferences using machine learning models, various applications such as image analysis, speech recognition, and data analysis are known. However, in order to perform inferences for desired applications accurately, it is important to obtain appropriate training data.

[0004] As training data for performing supervised learning, it is necessary to prepare a set of problems (explanatory variables) and correct answers (objective variables). Also, it is desirable that the training data is sufficient in quality and quantity.

[0005] The cost of creating such training data has been a practical issue. At this time, it is expected to efficiently prepare sufficient-quality and -quantity training data by extracting and using sets of explanatory variables and objective variables from various existing databases.

[0006] Patent Document 1 showed that by selecting features step by step, it is possible to gradually narrow down the features that have a large impact on the output result of the learning model.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

SUMMARY OF THE INVENTION

PROBLEMS TO BE SOLVED BY THE INVENTION

[0008] In Patent Document 1, the feature amount is selected step by step from a large classification to a medium classification and then to a small classification, but there is no direction of abstraction or specification of the range for selecting abstraction.

[0009] That is, depending on the type of inference to be executed on the machine learning model, adjustments such as abstracting some parts of the learning data while not abstracting other parts are required. However, conventionally, it has been difficult to set a detailed objective function partially or to set detailed explanatory variables partially.

[0010] Therefore, an object of the present invention is to enable partial abstraction or subdivision when generating learning data from existing data.

MEANS FOR SOLVING THE PROBLEMS

[0011] A preferred aspect of the present invention is a method for constructing a data model for learning data for machine learning. When the data items in the database that are the basis of the learning data have a hierarchical structure of abstraction levels or detail levels, an information processing apparatus can specify the abstraction level or detail level of the data items for each data item, and uses a filter that distributes the data items into target variables and explanatory variables to extract the data items to be used for learning data from the database, thereby constructing a data model for learning data.

[0012] Another preferred aspect of the present invention is a learning data generation device for generating learning data for machine learning, comprising a learning data generation unit. When the data items in the database that is the basis of the learning data have a hierarchical structure of abstraction level or detail level, the learning data generation unit enables the specification of the abstraction level or detail level of the data items for each data item, and uses a filter for classifying data items into target variables and explanatory variables to extract data to be used as target variables or explanatory variables for learning data from the database.

[0013] Another preferred aspect of the present invention is a machine learning method in which an information processing device learns a machine learning model using the target variable and explanatory variable obtained above.

Advantages of the Invention

[0014] When generating learning data from existing data, partial abstraction or subdivision can be enabled.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17

Mode for Carrying Out the Invention

[0016] Embodiments of the present invention will be described below with reference to the drawings. Note that the present invention is not limited by the following description.

[0017] In the configurations of the examples described below, the same reference numerals are commonly used for the same parts or parts having similar functions among different drawings, and duplicate descriptions may be omitted.

[0018] When there are a plurality of elements having the same or similar functions, they may be described with different subscripts attached to the same reference numeral. However, when it is not necessary to distinguish between the plurality of elements, the subscripts may be omitted in the description.

[0019] In this specification and the like, notations such as "first", "second", "third", etc. are attached to identify components, and do not necessarily limit the number, order, or content thereof. Also, numbers for identifying components are used for each context, and the numbers used in one context do not necessarily indicate the same configuration in other contexts. Further, it does not prevent a component identified by a certain number from also having the functions of a component identified by another number.

[0020] In the drawings and the like, the positions, sizes, shapes, ranges, etc. of the respective components shown may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings and the like.

[0021] The publications, patents, and patent applications cited in this specification form part of the description of this specification as they are.

[0022] In this specification, components represented in the singular form shall include the plural form unless otherwise clearly indicated in the context.

[0023] In the embodiments described below, when providing a data analysis environment service that multiplies data in multiple fields, an appropriate data model is provided. In this embodiment, the data model has a function of defining at least the elements of data that become the target variable and the elements of data that become the explanatory variables. Also, the data model may include, as additional detailed information, a definition of the relationship between data elements. In this case, the data model is defined as "a data model in which the elements of data that become the target variable, the elements of data that become the explanatory variables, and the relationship between the data elements are defined".

[0024] Conventionally, it has been difficult to adjust the abstraction level between fields and the level of detail within a field. That is, there has been no direction for data abstraction or specification of the range where abstraction is stopped, and it has been difficult to set a partially detailed objective function or partially detailed explanatory variables.

[0025] In the following embodiment, an integrated filter having three types of filter functions, namely, a target variable / explanatory variable allocation filter, an abstraction avoidance filter, and an abstraction filter, is applied to the most detailed data layer. This makes it possible to provide an appropriate data model when providing a data analysis environment service that combines data from multiple fields.

[0026] Furthermore, the optimal integrated filter can be calculated by automatic parameter tuning of the integrated filter. In other words, it is possible to calculate the optimal integrated filter and achieve the optimal balance between fields.

[0027] According to such an embodiment, a data analysis is performed that includes a proposal for an integrated data model that fits the training data. Ta This will enable the service and make it possible to obtain learning data that integrates data from multiple fields.

[0028] 1 is a conceptual diagram showing the concept of a method for generating a data model of learning data described in the embodiment. A filter 100 is applied to data items of existing databases DB1, DB2, and DB3 to generate a data model 200.

[0029] Existing databases generally have a hierarchical structure defined by the database creator, and are structured in stages from higher-level data items (classifications) to lower-level data items (individual items), such as major classifications, medium classifications, minor classifications, and individual items. Various existing databases can be used as databases DB1, DB2, and DB3. Databases in one or more fields can be used.

[0030] In the filter 100, filter conditions are set by experts or the like having knowledge in the relevant field according to the use or purpose of the machine learning model to be created, and the filter conditions are stored as filter data. The filter 100 acts on data items in the database DB.

[0031] Filter 100 includes an abstraction filter that groups data items in a database into higher-level data items, an abstraction avoidance filter that does not apply the abstraction filter to predetermined data items, a purpose-explanation factor distribution filter that distributes data items in a database into target variables and explanatory variables, and the like. Filter 100 also specifies integration conditions when integrating a plurality of databases.

[0032] Data model 200 is a data model for generating (extracting) learning data from data in a database. Data items that become one or more target variables and data items that become one or more explanatory variables are specified. When data is extracted from database DB according to the definition of data model 200, learning data can be generated.

Example

[0033] An example of generating learning data from an existing database will be described. In this example, learning data is generated when machine learning whether a person who has a certain specific disease has other diseases. In this example, an appropriate data model for generating learning data is provided.

[0034] FIG. 2 is a table diagram showing an example of data items 300 in a database related to cardiovascular diseases. The alphabetic and numeric codes in the table are the codes of ICD10, an international disease classification, and represent disease names.

[0035] In the example of FIG. 2, the major classification is the whole of cardiovascular diseases, the middle classification is divided into two classifications, namely, "ischemic heart disease" and "cerebrovascular disease", which are further classified into the heart and the brain, the minor classification is divided into four classifications, and eight specific disease names are defined as individual items. Data items that mean the classification of database data often adopt a hierarchical structure in this way. The seven-digit code corresponding to the ICD10 code is the code determined by the Ministry of Health, Labour and Welfare.

[0036] In an actual database, data is stored for each patient ID and event, for example, according to the individual items shown in FIG. 2.

[0037] Figure 3 shows an example of the structure of the data model 200 to which the learning data should conform. Learning data generally consists of a set of inputs (explanatory variables) and expected outputs (target variables) of a machine learning model. If actual data is used in the database, the correct target variable can be obtained for the explanatory variable in question.

[0038] When generating learning data from a database having the data items of Figure 2 according to the data model of Figure 3, the machine learning model can learn, for example, to estimate the risk that a patient having symptoms of "ischemic heart disease" (excluding "acute myocardial infarction" and "myocardial infarction") or "cerebrovascular disease" as an explanatory variable will develop symptoms of "acute myocardial infarction" or "myocardial infarction" as the target variable. Alternatively, conversely, the explanatory variable may be estimated from the target variable.

[0039] In this embodiment, in order to automatically generate learning data from the database, a data model is generated and processed using the concept of a filter.

[0040] Figure 4 is an example of a filter for extracting learning data according to the data model 200 of Figure 3 from the database defined by the data item 300 of Figure 2.

[0041] In the filter 100, the abstraction filter specifies the level of abstraction for medium classification. This specifies that among the data items 300 of Figure 2, "ischemic heart disease" and "cerebrovascular disease", which are medium classifications, are to be used as data items. That is, the large classification, small classification, and individual items are ignored, and the medium classification corresponding to the individual items is extracted as learning data.

[0042] In the abstraction avoidance filter, the abstraction filter indicates that it is not applied to the individual items of "acute myocardial infarction" and "myocardial infarction". For this reason, the data of these individual items is extracted as it is in the learning data.

[0043] In the purpose - explanatory factor allocation filter, for the extracted data, "acute myocardial infarction" and "myocardial infarction" are specified as target variables, and the others are specified as explanatory variables.

[0044] When the filter 100 of the conditions in FIG. 4 is applied to the data item 300 in FIG. 2, the data model 200 in FIG. 3 can be generated, and learning data can be generated by extracting data from the database according to the data model.

[0045] According to this embodiment, based on the existing database, learning data with arbitrarily changed data granularity (data abstraction level) can be generated, and learning suitable for the use and purpose of the machine learning model can be executed. In the above example, the machine learning model can be configured to estimate the risks focusing particularly on acute myocardial infarction and myocardial infarction of individual items based on the middle - classification diseases.

Example

[0046] In Example 2, an example of generating learning data by integrating databases in multiple fields will be described. Here, an example of integrating the database in the disease field and the database in the dispensing field is shown. Combining and integrating data from multiple fields in this way is important in the field of machine learning. However, simply combining both data will include data with low importance and the data volume will become extremely large, increasing the load of the learning process. Therefore, data selection during integration is important.

[0047] In this example, an example of generating learning data by integrating the database in the disease field based on the data item 300 shown in FIG. 2 and the database in the dispensing field based on the data item 300 - 2 shown in FIG. 5 will be described. This is a case of machine - learning what explanatory variables a person who has been prescribed a certain drug and has a certain disease had. In this embodiment, adjustment of the abstraction (detail) level between fields and adjustment of the abstraction (detail) level within a field are enabled.

[0048] Figure 5 shows data item 300-2 of the database in the dispensing field for pharmaceuticals for the nervous system and sensory organs. The configuration is the same as that of the disease classification database in Figure 2, and the drug code is described in the individual items. Note that in the data structures of Figure 2 and Figure 5, the classification has three levels, but it may also have one level or four or more levels. That is, the number of classification levels is at the discretion of the database designer.

[0049] Figure 6 is a diagram showing the concept of the integration filter 100U that integrates data item 300 of the disease field database and data item 300-2 of the dispensing field database. The filter includes an abstraction filter, an abstraction avoidance filter, and a purpose / description factor sorting filter, similar to Example 1.

[0050] As shown in Figure 6, the filter conditions for the purpose / description factor sorting filter and the abstraction avoidance filter are set for each individual item. In the examples shown in Figure 2 and Figure 5, both the data item 300 in the disease field and the data item 300-2 in the dispensing field have eight individual items. In Figure 6, the filter conditions for each individual item are illustrated in the order of the data items shown in Figure 2 and Figure 5. Note that the number of items is an example, and the number is at the discretion of the filter designer.

[0051] Also, the abstraction filter is set for each database. In this example, the data items in the disease field are abstracted to medium classification, and the data items in the dispensing field are abstracted to small classification.

[0052] The filter conditions for the data item 300 in the disease field are the same as those in Example 1.

[0053] For the filter conditions for data item 300-2 in the dispensing field, in the purpose / description factor distribution filter, "Buscopan tablets 10 mg" and "Gavapentin tablets 5 mg" are specified as the target variables, "Myslee tablets 5 mg" and "Phenobarbital powder 10%" are not used, and the others are specified as explanatory variables. In the abstraction filter, the subcategory is specified as the abstraction level. In the abstraction avoidance filter, "Akineton tablets 1 mg", "Pramipexole hydrochloride tablets", "Buscopan tablets 10 mg", and "Gavapentin tablets 5 mg" are used as individual items.

[0054] Figure 7 is a diagram showing an example of the integrated data model 200U integrated by the integrated filter 100U.

[0055] As target variables, "acute myocardial infarction" and "myocardial infarction" are extracted from the individual items in the disease field. Also, as target variables, "Buscopan tablets 10 mg" and "Gavapentin tablets 5 mg" are extracted from the individual items in the dispensing field.

[0056] As explanatory variables, from the data items in the disease field, the medium category "ischemic heart disease (excluding the two individual items used as target variables)" and "cerebrovascular disease" are extracted. Also, as explanatory variables, from the data items in the dispensing field, the subcategory "hypnotics and sedatives (excluding the two individual items)", "anti-Parkinson agents (excluding the two individual items which are abstraction-avoided)", "autonomic nerve agents", and "anticonvulsants (excluding the two individual items used as target variables which are abstraction-avoided)" are extracted.

[0057] In this way, the level of abstraction of the data used for variables can be freely set to either be set for each classification or to use individual items as they are. For example, detailed explanatory variable settings such as using individual items for the data items of interest and grouping less important items by classification become possible. Note that in the above example, both target variables and explanatory variables are extracted from each database, but it is also possible to extract only target variables or only explanatory variables.

[0058] Figure 8 schematically shows the detailed conditions of the data model for integrating the variables obtained from the database in the disease field and the variables obtained from the database in the dispensing field. The integrated target variable has the target variable of "acute myocardial infarction" or "myocardial infarction", and is assumed to have the target variable of " Pa Busco Tablet 100 mg" or "Gabaron Tablet 5 mg". The integrated explanatory variable is assumed to have any of the explanatory variables illustrated in Figure 8.

[0059] The learning data obtained by this data model is suitable for learning which of the various medical histories or medications that have the symptoms of "acute myocardial infarction" or "myocardial infarction" and have a history of being prescribed " Pa Busco Tablet 100 mg" or "Gabaron Tablet 5 mg" are closely related to the target variable.

[0060] The above example is just an example. Depending on the content of the estimation to be performed on the machine learning model, the target variable and the explanatory variable obtained from multiple databases under the desired conditions can be combined by well-known logical operations to create an integrated target variable and an integrated explanatory variable.

[0061] Figure 9A is a conceptual diagram showing a configuration example of the data files of the database DB1 in the disease field and the database DB2 in the dispensing field. The two databases DB1 and DB2 can be cross-referenced by the personal ID. The personal database DBP stores the personal ID and other bibliographic items. Similarly hereinafter, the content of the bibliographic items is arbitrary.

[0062] Figure 9B is a diagram showing an example of the disease receipt file 901, which is the content of one data file of the database DB1 in the disease field. Associated with the personal ID, bibliographic items such as the medical receipt number and the diagnosis date, and individual items such as the disease name, disease name code, and ICD10 code are recorded. The individual items are hierarchically classified according to the classification shown in Figure 2.

[0063] FIG. 9C is a diagram showing an example of a dispensing receipt file 902, which is the content of one data file of the database DB2 in the dispensing field. Bibliographic items such as a dispensing receipt number and a prescription date are recorded in association with an individual ID, and individual items such as a dispensing name and a drug efficacy classification code are also recorded. The individual items are hierarchically classified according to the classification shown in FIG. 5.

[0064] A large number of data files such as those shown in FIGS. 9B and 9C are stored in the database DB1 in the disease field and the database DB2 in the dispensing field. The data files can be specified by a medical receipt number or a dispensing receipt number. Data is extracted from these data based on the data model shown in FIG. 7.

[0065] Note that although the above example of the data file is an independent data file for each medical treatment or dispensing event for each individual ID, it may be integrated data for each individual ID in advance.

[0066] FIG. 10 shows an example in which the integrated filter 100U of FIG. 6 is applied to the database DB1 in the disease field and the database DB2 in the dispensing field, and the data items of the integrated data model 200U of FIG. 7 are extracted. For each individual item or abstracted item shown in FIG. 7, whether it corresponds or not is indicated by 1 and 0. Then, based on the detailed conditions shown in FIG. 8, when the data file is extracted, the desired learning data can be obtained.

[0067] The example of FIG. 10 shows an example in which the content of the data file extracted by the integrated filter 100U is integrated as a big table 1000. The first row from the top corresponds to the content of the disease receipt file 901 of FIG. 9B, and the second row from the top corresponds to the content of the dispensing receipt file 902 of FIG. 9C.

[0068] The disease field database DB1 and the dispensing field database DB2 each contain multiple files, and usually multiple files are associated with a single personal ID. In the example of Figure 10, 10 data files associated with the individual with the personal ID "F20011" are shown. As already explained, the extracted items are classified into target variables and explanatory variables by the target - explanatory factor sorting filter.

[0069] According to the logical formula of the detailed conditions of the data model in Figure 8, from this big table 1000, ("acute myocardial infarction" or "myocardial infarction") and ("Buscopan tablets 100mg" or "Gavalon tablets 5mg") are set as the integrated target variable. Also, those having any of the explanatory variables are used as the integrated explanatory variables. Pa The data with the personal ID "F20011" has both "acute myocardial infarction" and "Buscopan tablets 10mg" as target variables and has the integrated target variable, so it can be used as learning data. Using this learning data, learn what integrated explanatory variables the person having the said integrated target variable had and the relationship between each item of the integrated explanatory variables.

[0070] As described above, the extracted and integrated data becomes teacher data including integrated explanatory variables (problems) and integrated target variables (answers), so it can be used as learning data for a machine learning model.

[0071]

Example

[0072] In the explanations so far, the filter 100 was assumed to include an abstraction filter that aggregates the data items in the database into higher - level data items, an abstraction avoidance filter that does not apply the abstraction filter to predetermined data items, and a target - explanatory factor sorting filter that sorts the data items in the database into target variables and explanatory variables.

[0073] ​The above example designs the filter from the perspective of whether to abstract small-granularity data (individual items) into large-granularity data (classification). However, conversely, it is also possible to design the filter from the perspective of whether to refine (concretize) large-granularity data into small-granularity data.

[0074] Figure 11 is a diagram showing the concept of an integrated filter 100U-2, which is an alternative to the integrated filter 100U in Figure 6. Instead of the abstraction avoidance filter in Figure 6, it is equipped with a refinement avoidance filter.

[0075] In the integrated filter 100U of Figure 6, in principle, individual items are abstracted into large classifications to small classifications, and individual items that are not abstracted by the abstraction avoidance filter are specified. Specifically, the first filter determines whether each individual item is a target variable, an explanatory variable, or not used, the second filter determines the abstraction of each individual item, and the third filter determines whether to avoid the abstraction of each individual item.

[0076] In the integrated filter 100U-2 of Figure 11, the first filter determines whether each classification is a target variable, an explanatory variable, or not used, the second filter determines the refinement of each classification, and the third filter determines whether to avoid the refinement of each classification.

[0077] In the integrated filter 100U-2 of Figure 11, in principle, classifications are refined into individual items, and individual items that are not refined (i.e., classifications are used as items) by the refinement avoidance filter are specified. The integrated filter 100U and the integrated filter 100U-2 have exactly the same function as a result.

Example

[0078] When constructing an integrated data model, it may be considered necessary to adjust the composition ratio between the fields of the target variable and the explanatory variable according to the inferences to be made by the target machine learning model and the characteristics of the underlying database. In that case, it is desirable to visualize the characteristics of the integrated data model 300U.

[0079] FIG. 12 is a diagram showing an example of a GUI (Graphical User Interface) that visualizes the inter-field composition ratio of the integrated data model 200U shown in FIGS. 7 and 8. For each of the integrated target variable and the integrated explanatory variable, it shows the data granularity (abstraction level), the number of adopted items, and the ratio.

[0080] For example, for the integrated target variable, since two individual items are adopted from the disease field database and the dispensing field database respectively, the ratio is 50% each.

[0081] For the integrated explanatory variable, two medium classifications ("ischemic heart disease" and "cerebrovascular disease") are adopted from the disease field database, three small classifications ("anti-Parkinson agent", "autonomic nerve agent", and "anticonvulsant") and two individual items ("Akineton tablets 1 mg" and "Pramipexole hydrochloride tablets") are adopted from the dispensing field database, for a total of five items.

[0082] In the example of FIG. 12, the ratio shows the ratio of the total number without distinguishing classifications or individual items, but it can also be shown by the granularity of the items. Alternatively, weighting according to the granularity of the items may be performed.

[0083] In a preferred embodiment, the integrated data model 200U and the integrated filter 100U therefor are designed by an expert with knowledge in the application field of the machine learning model and recorded in advance as data. At that time, it is desirable to create and store in advance a plurality of types having different characteristics so that they can be selected later.

[0084] FIG. 13 is a diagram showing an example of a GUI that compares and visualizes the characteristics of three integrated data models having different characteristics. The left table is in the same format as the table in FIG. 12 and shows the characteristics of the integrated filters A, B, and C for different integrated data models. The right diagram shows, for each integrated data model, the ratio of data adopted from the two databases. For example, for integrated filter A, 50% is adopted from the disease field and the dispensing field respectively, and for integrated filter B, 33% is adopted from the disease field and 67% is adopted from the dispensing field.

[0085] A user who wants to create learning data can select an integrated data model with desired characteristics by referring to these GUIs.

[0086] FIG. 14 is a diagram showing an example of a GUI for a user to specify the characteristics of a data model and select a data model having characteristics close to the specified characteristics.

[0087] The user specifies a desired inter-domain composition ratio in region 1401. On the system side, an integration filter for generating a data model having the same or closest characteristics to the specified characteristics is displayed in region 1402.

[0088] In this way, a data model having desired characteristics can be used.

Example

[0089] A specific system example for realizing the above example and an example of a processing flow will be described.

[0090] FIG. 15 is a block diagram of a learning data generation system for applying an integration filter to databases in a plurality of fields to generate learning data based on a desired data model.

[0091] The learning data generation system 1500 can be configured by an information processing device such as a general server. Similar to a general server, it includes a processing device CPU, a memory MEM, an input device IN, an output device OUT, and a bus (not shown) connecting each part. The program executed in the learning data generation system 1500 shall be stored in the memory MEM in advance.

[0092] In this embodiment, functions such as calculation and control are realized by a program stored in the memory MEM being executed by the processing device CPU, and by collaborating with other hardware to perform defined processing. The program executed by the processing device CPU, its function, or the means for realizing its function may be referred to as "function", "means", "section", "unit", "module", etc.

[0093] In this embodiment, the memory MEM stores a learning data generation unit 1501 and a machine learning unit 1502 as software for executing the processing described later. The memory MEM can be configured by, for example, a semiconductor memory device.

[0094] Further, the learning data generation system 1500 can access the storage device 1510 and utilize the data stored in the storage device 1510. Also, the learning data generation system 1500 can record data in the storage device 1510. The storage device 1510 can be configured by, for example, a magnetic storage device or the like.

[0095] In this embodiment, it is assumed that a database DB1 in the first field and a database DB2 in the second field are stored in the storage device 1510 in advance. The database DB1 in the first field is, for example, a database in the medical field, and the database DB1 in the second field is, for example, a database in the dispensing field (see FIGS. 9A to 9C). In this example, the number of databases is two, but the number is arbitrary.

[0096] In this embodiment, it is assumed that filter data FT is stored in the storage device 1510 in advance. A specific example of the filter data FT is, for example, the integrated filter 100U shown in FIG. 6. It is assumed that a plurality of types of integrated filters 100U with different characteristics are stored in the filter data FT in advance.

[0097] In addition, the learning data generation unit 1501 generates a big table TB and learning data TD from the database DB1 in the first field and the database DB2 in the second field, and records them in the storage device 1510. A specific example of the big table TB is the big table 1000 (see FIG. 10). In the embodiment, the data stored in the storage device 1510 is described in a table-form data structure, but it may be expressed in a data structure such as a list or a queue.

[0098] Based on the big table TB, for each personal ID, the target variable and the explanatory variable are aggregated according to the conditions of the logical formula shown in FIG. 8, for example. Then, for example, learning data indicating what diseases a person who has a disease such as acute myocardial infarction or myocardial infarction and has been prescribed buscopan tablets 10 mg or gabalon tablets 5 mg has suffered from or what drugs have been prescribed can be obtained. Since the existing database contains information on a large number of people, learning data for a large number of people can be obtained in the same way.

[0099] The machine learning unit 1502 performs machine learning using the obtained learning data TD. In the embodiment, the learning data generation system 1500 includes the machine learning unit 1502, but it may be a completely independent and separate configuration. If learning is performed by providing the learning data TD in an arbitrary method, the effects of this embodiment can be obtained. Since the machine learning method itself may be a known method, the details are omitted.

[0100] In the description of the embodiment, the "program" may be used as the subject for explanation. However, since the program is executed by the processing device CPU to perform the defined processing while using the memory MEM, the input device IN, and the output device OUT, it may be described with the processing device CPU as the subject. Also, the processing disclosed with the program as the subject may be the processing performed by the learning data generation system 1500. Further, part or all of the program may be realized by dedicated hardware.

[0101] In addition, as examples of the input device IN, a keyboard and a pointer device can be considered, but other known devices may also be used. As examples of the output device OUT, a display and a printer can be considered, but other known devices may also be used. Further, as the input device IN and the output device OUT, an interface for communicating with other external devices may be included.

[0102] The above configuration may be composed of a single computer, or any part of the processing device CPU, the memory MEM, the input device IN, and the output device OUT may be composed of other computers connected by a network. Further, the storage device 1510 may be a part of the learning data generation system 1500, or may be connected to a system separate from the learning data generation system 1500 via a network.

[0103] In this embodiment, functions equivalent to those configured by software can also be realized by hardware such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0104] FIG. 16A is a flowchart showing the flow of the learning data generation process executed by the learning data generation unit 1501.

[0105] The learning data generation unit 1501 accesses the filter data FT in the storage device 1510, reads out the file of one or more integrated filters 100U specified by the user, and acquires the filter conditions (S1601). Hereinafter, an example of integrating two databases into one integrated filter will be described. As already described, the number of databases to be integrated is arbitrary depending on the specification of the integrated filter. When a plurality of integrated filters are read out, the same processing as below may be repeated for the number of filters.

[0106] The learning data generation unit 1501 accesses the database DB1 in the first field of the storage device 1510 and acquires data (S1602-1). In this example, it is assumed that the data is hierarchically classified into major classification, medium classification, minor classification, and individual items (see FIG. 2).

[0107] The learning data generation unit 1501 classifies the individual items of the acquired data into individual items, explanatory variables, or unused according to the filter conditions (see FIG. 6) (S1603-1).

[0108] For the target variable and explanatory variable of the individual items in the first field, the learning data generation unit 1501 selects whether to avoid or not avoid abstraction according to the filter conditions (see FIG. 6) (S1604-1).

[0109] The learning data generation unit 1501 determines the abstraction level of the individual items in the first field according to the filter conditions (see FIG. 6) (S1605-1). In this example, there are four types of abstraction levels: major classification, medium classification, minor classification, and no abstraction.

[0110] Through the above processing, it is possible to abstract the individual items in the database of the first field by specifying a range. In the above example, abstraction is performed after classifying the individual items into variables, but variable classification may also be performed after abstraction. Also, although the abstraction level is determined at the end, it may be determined at the beginning. That is, the order of the flow is not limited to the example in FIG. 16.

[0111] The learning data generation unit 1501 accesses the database DB2 in the second field of the storage device 1510 and performs the processing of S1602-2 to S1605-2 in the same manner as above. The same applies when the number of databases to be integrated is three or more.

[0112] As a result of the processing shown in FIG. 16A, variables in the first field and variables in the second field that are abstracted by specifying a range and classified into a target variable and an explanatory variable, as shown in the big table 1000 in FIG. 10, for example, can be obtained.

[0113] FIG. 16B is a flowchart showing the flow of the learning data generation process executed by the learning data generation unit 1501 following FIG. 16A.

[0114] The learning data generation unit 1501 obtains the detailed conditions of the integrated target variable from the file of the integrated filter 100U acquired in S1601 (S1606). This is generally obtained in the form of logical expressions such as AND, OR, NOT, NOR, etc.

[0115] The learning data generation unit 1501 generates an integrated target variable based on the obtained logical expression by combining the target variables of the first field and the target variables of the second field (S1607).

[0116] The learning data generation unit 1501 obtains the detailed conditions of the integrated explanatory variable from the file of the integrated filter 100U acquired in S1601 (S1608). This is generally obtained in the form of logical expressions such as AND, OR, NOT, NOR, etc.

[0117] The learning data generation unit 1501 generates an integrated explanatory variable set by combining the explanatory variables of the first field and the explanatory variables of the second field based on the obtained logical expression (S1609).

[0118] The learning data generation unit 1501 calculates the inter-field composition ratio of the integrated target variable (S1610). This can be easily obtained from the integrated data model 200U.

[0119] The learning data generation unit 1501 calculates the inter-field composition ratio of the integrated explanatory variable (S1611). This can be easily obtained from the integrated data model 200U.

[0120] The processing regarding one set of integrated filter conditions ends here. The data including the integrated target variable and the integrated explanatory variable set can be used as learning data.

Example

[0121] In Example 6, taking Example 5 as the basic configuration, an example of generating learning data by automatically tuning the parameters of the integrated filter and training a machine learning model with the learning data to generate a prediction model will be described.

[0122] FIG. 17 is a flowchart for explaining an example of automatically tuning the parameters of the integrated filter executed by the learning data generation unit 1501.

[0123] The learning data generation unit 1501 sets ideal values for the constituent ratios between fields of the target variable and the explanatory variable (S1701). Specifically, the learning data generation unit 1501 displays a GUI shown in FIG. 14 on a display device (a specific example of the output device OUT in FIG. 15), allows the user to operate the scale of area 1401, and enables the user to input the constituent ratio between fields of the target variable and the explanatory variable that the user desires.

[0124] The learning data generation unit 1501 sets N integrated filter condition files (S1702). The N integrated filter condition files are files that the learning data generation unit 1501 reads desired files from the filter data FT in the storage device 1510. The N files may be selected by the user or automatically selected according to predetermined rules.

[0125] For each of the N integrated filter conditions, the learning data generation unit 1501 calculates the constituent ratio between fields of the target variable and the explanatory variable in the same manner as the processes of S1610 and S1611 in FIG. 16B (S1703).

[0126] The learning data generation unit 1501 selects an integrated filter condition in which the constituent ratio between fields of the target variable and the explanatory variable is close to the ideal value set in process S1701 (S1704).

[0127] The learning data generation unit 1501 constructs a big table according to the selected integrated filter condition (S1705). Specifically, each item of the big table 1000 in FIG. 10 is determined.

[0128] For each data file, the learning data generation unit 1501 inputs numerical values into the big table 1000 for each component in the file. This process is performed for all the files in the database (S1706).

[0129] As described above, for example, the big table 1000 in FIG. 10 can be obtained. Further, detailed conditions of the data model (see, for example, FIG. 8) are applied to the big table, and the corresponding data is used as learning data. As described above, this data is example data indicating what explanatory variables a person has (or does not have) when having a predetermined target variable. Therefore, by using it as learning data for a machine learning model, the relationship between the target variable and the explanatory variables can be learned.

[0130] Although not shown in the configuration, the generation of the prediction model by machine learning (S1707) can be performed by the machine learning unit 1502 using known hardware and software.

[0131] According to the above embodiment, efficient machine learning can be realized, so that energy consumption is reduced, carbon emissions are reduced, global warming can be prevented, and a sustainable society can be realized.

Explanation of Signs

[0132] Database DB1, DB2, DB3, filter 100, data model 200, big table 1000, learning data generation system 1500, learning data generation unit 1501, storage device 1510

Claims

1. A method for constructing a data model for learning data for machine learning, comprising: When data items that represent the classification of the data in the database that is the basis of the learning data have a hierarchical structure of abstraction levels or levels of detail, Preparing filter data that specifies at least one of the hierarchies and assigns which part of the data item is the target variable and which part is the explanatory variable; An information processing apparatus acts on the data item based on the filter data, enables specification of the abstraction level or level of detail of the data item for each data item, and uses a filter that assigns the data item to a target variable and an explanatory variable, To extract data items to be used for learning data from the database, a data model is constructed that defines data items to be target variables and data items to be explanatory variables. A method for constructing a data model for learning data.

2. When the data item has a hierarchical structure of classification and individual items, The filter has the functions of a first filter, a second filter, and a third filter, The first filter determines whether each individual item is a target variable, an explanatory variable, or not used, The second filter determines the abstraction of each individual item, The third filter determines whether to avoid the abstraction of each individual item. The method for constructing a data model for learning data according to Claim 1.

3. When the data item has a hierarchical structure of classification and individual items, The filter has the functions of a first filter, a second filter, and a third filter, The first filter determines whether each classification is a target variable, an explanatory variable, or not used, The second filter determines the refinement of each classification, The third filter determines whether to avoid the refinement of each classification. The method for constructing a data model for learning data according to Claim 1.

4. When a plurality of databases that are the basis of the learning data are used and the data items of each database have a hierarchical structure of abstraction levels or levels of detail, Applying the filter to each of the plurality of databases and functioning as an integration filter that extracts and integrates data items to be used for learning data from the plurality of databases. The method for constructing a data model for learning data according to Claim 1.

5. The filters applied to the plurality of databases each have different characteristics. The method for constructing a data model for learning data according to Claim 4.

6. Calculating the ratio of the target variable to the explanatory variable, which the integrated filter extracts from each of the plurality of said databases. The method for constructing a data model for learning data according to claim 5.

7. Preparing a plurality of candidates for the integrated filter. Each candidate of the integrated filter calculates the ratio of the target variable to the explanatory variable extracted from each database. Selecting an integrated filter that realizes the ratio of the target variable to the explanatory variable closest to the input value. The method for constructing a data model for learning data according to claim 6.

8. A learning data generation device for generating learning data for machine learning, comprising a learning data generation unit. The learning data generation unit When the data item that represents the classification of the data in the database that is the basis of the learning data has a hierarchical structure of abstraction or detail Comprising filter data for designating at least one of the hierarchies and allocating which part of the data item is the target variable and which part is the explanatory variable. Acting on the data item based on the filter data, making it possible to specify the abstraction or detail level for each data item, and using a filter that allocates the data item as the target variable and the explanatory variable to construct a data model that defines the data item to be the target variable and the data item to be the explanatory variable. Using the data model to extract data to be used as the target variable or the explanatory variable for learning data from the database. Learning data generation device.

9. When using a plurality of databases that are the basis of the learning data and the data items of each database have a hierarchical structure of abstraction or detail Applying the filter to each of the plurality of databases, and functioning as an integrated filter that extracts and integrates data to be used for learning data from each of the plurality of databases. The learning data generation device according to claim 8.

10. The filters applied to the plurality of databases each have different characteristics. The learning data generation device according to claim 9.

11. The filter further determines whether to exclude the data item of the database. The learning data generation device according to claim 9.

12. The filter has a function of generating an integrated target variable and an integrated explanatory variable by performing a logical operation on at least one of the target variable and the explanatory variable extracted from the plurality of databases. The learning data generation device according to claim 9.

13. The learning data generation unit has a function of selecting the ratio of the target variable and the explanatory variable extracted from each of the plurality of databases. The learning data generation device according to claim 9.

14. The learning data generation unit calculates the ratio of the target variable and the explanatory variable that each of the plurality of types of integrated filters extracts from each database, and selects an integrated filter that realizes the ratio of the target variable and the explanatory variable closest to the input value. The learning data generation device according to claim 9.

Citation Information

Patent Citations

  • Data division program, recording medium with the program recorded thereon, data distribution device and data distribution method

    JP2008299382A

  • Data analysis device

    JP2020135068A

  • Information processing device, information processing method, and program

    JP2020184212A

  • Information process system, and program

    WO2009025022A1

  • Information processing system, feature value explanation method and feature value explanation program

    WO2018180970A1