AI learning data creation support system, AI learning data creation support method, and AI learning data creation support program
The AI learning data creation support system addresses data scarcity issues for AI models by calculating required data and generating supplementary queries, efficiently collecting data for healthcare applications while ensuring privacy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HITACHI LTD
- Filing Date
- 2022-01-14
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies face challenges in efficiently collecting learning data for AI models, particularly for healthcare applications requiring high accuracy, especially when dealing with rare conditions like lung cancer risk, where data scarcity is a significant issue.
An AI learning data creation support system that extracts and collects learning data from databases by receiving user inputs, calculating the required data quantity, and generating supplementary queries to ensure sufficient data is gathered for training, using a processor and input device to manage the process.
The system efficiently collects training data for AI models, ensuring adequate data is obtained while protecting privacy, particularly for healthcare applications, by determining the necessary data quantity and generating supplementary queries when needed.
Smart Images

Figure 0007868979000001 
Figure 0007868979000002 
Figure 0007868979000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an AI learning data creation support system, an AI learning data creation support method, and an AI learning data creation support program that extract and collect learning data for training an AI model from at least one learning database.
Background Art
[0002] Techniques for obtaining desired information from a vast amount of information that can be acquired via the Internet have been disclosed. For example, in the technique described in Patent Document 1, a sub-web including a list of paths of websites on the Internet weighted based on the relevance to topics of interest to the user or the user's characteristics is created. Then, by using the sub-web for Internet site search, a search engine can easily execute a focused search of Internet sites. Therefore, when using the technique described in Patent Document 1, information on Internet sites related to the user's interests and characteristics can be collected by searching using a search engine.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, even if information on Internet sites related to the user's characteristics can be collected using the technique described in Patent Document 1, it may not be easy to extract and collect learning data for an AI model that includes information on specific multiple data items from a database.
[0005] In particular, AI models for healthcare, used to analyze and predict the health status of individuals and groups, are expected to perform important analyses related to human health. However, depending on the content of the analysis performed by the AI model, it may not be easy to collect training data. For example, if the analysis is about the lung cancer risk (likelihood of developing the disease) in patients with rare disease A, it is difficult to collect training data because very few people have had rare disease A in the past and subsequently developed lung cancer. Furthermore, if a high degree of accuracy is required for the analysis results of an AI model for healthcare, it may be difficult to collect training data.
[0006] The objective of this invention is to provide an AI learning data creation support system, an AI learning data creation support method, and an AI learning data creation support program that can efficiently collect learning data for training an AI model. [Means for solving the problem]
[0007] An AI learning data creation support system, which is one aspect of the invention disclosed in this application, is an AI learning data creation support system that extracts and collects learning data for training an AI model from at least one learning database, and comprises a storage device that stores at least one program, a processor that executes the program stored in the storage device, and an input device that receives input from a user, wherein the processor executes the program and receives input of a learning profile which consists of item values corresponding to each of a plurality of data items and includes information on data to be analyzed by the AI model and information on the type of the AI model, obtains a first query to be used for extracting the learning data, calculates the number of first learning data to be extracted from the learning database by the first query using the learning database, calculates the required number of learning data necessary for training the AI model using the information on the type of the AI model included in the learning profile, determines whether the number of first learning data is equal to or greater than the required number, and if it is determined that the number of first learning data is less than the required number, generates a supplementary query to be used for extracting the learning data based on the learning profile. [Effects of the Invention]
[0008] According to the present invention, training data for training an AI model can be collected efficiently. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 shows an example of a functional block diagram of the AI learning data creation support system in Example 1. [Figure 2] Figure 2 shows an example of a hardware configuration diagram for the AI learning data creation support system in Example 1. [Figure 3] Figure 3 shows an example of a personal profile and the first query. [Figure 4] Figure 4 shows a configuration conditions database and an example of a configuration conditions table stored in the configuration conditions database. [Figure 5] Figure 5 shows an example of a search criteria database. [Figure 6] Figure 6 shows an example of an algorithm requirement table. [Figure 7] Figure 7 shows an example of a table showing the required number of analysis items. [Figure 8] Figure 8 is an explanatory diagram showing an example of a personal profile input screen displayed on the client device for the user to enter their personal profile and the first query. [Figure 9] Figure 9 is an explanatory diagram showing an example of a query input screen displayed on the client device for the user to enter the first query. [Figure 10] Figure 10 is an explanatory diagram showing an example of a query input screen displayed on the client device for the user to enter the first query. [Figure 11] Figure 11 is a flowchart showing an example of the training data acquisition process in Example 1. [Figure 12] Figure 12 is a flowchart showing an example of the processing of the supplement query generation subroutine in Example 1. [Figure 13] Figure 13 illustrates how to generate the second supplementary query. [Figure 14] Figure 14 is an explanatory diagram showing an example of a supplement query display screen, which is displayed on the client device's screen to show the user the supplement queries registered in the supplement query list and the number of supplement data items. [Figure 15] Figure 15 is a flowchart showing an example of the processing of the supplement query generation subroutine in Example 2. [Figure 16] Figure 16 is a flowchart showing an example of the training data acquisition process in Example 3. [Modes for carrying out the invention]
[0010] Hereinafter, embodiments will be described with reference to the drawings. The examples are illustrative for explaining the present invention, and for the sake of clarity of explanation, appropriate omissions and simplifications have been made. The present invention is not limited to the examples, and the technical scope of the present invention includes all application examples that conform to the idea of the present invention.
[0011] Also, in the drawings and the following description, the same reference numerals may be given to the same parts or parts having the same or similar functions, or different subscripts may be attached to the same reference numeral for explanation, or the subscripts may be omitted for explanation. Also, unless otherwise specified, each component may be plural or singular.
[0012] The positions, sizes, shapes, ranges, etc. of the components shown in the drawings may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings.
[0013] Also, in the following description, various information may be described using expressions such as "table", "table", "list", "queue", etc., but the various information may be represented by other data structures. Also, in order to indicate that the various information does not depend on the data structure, "table" etc. can be called "management information". Identification information may be described using expressions such as "identification information", "identifier", "name", "ID", "number", etc., but these can be replaced with each other.
[0014] In addition, when describing a process in a sentence with "program" or "functional unit" as the subject, the program or functional unit is executed by a processor that is a processing unit or an arithmetic unit, for example, an MP (Micro Processor), a CPU (Central Processing Unit), or a GPU (Graphics Processing Unit), and performs a defined process. The processor performs processing while using storage resources (e.g., memory) and a communication interface device (e.g., a communication port). Therefore, the subject of a sentence with "program" or "functional unit" as the subject may be replaced with a processor, a processing unit, or an arithmetic unit. Also, the entity performing the process executed by the program may be a processor, an arithmetic unit, or a processing unit, or a controller, a device, a system, a computer, or a node having a processor, or a dedicated circuit performing a specific process. Here, the dedicated circuit is, for example, an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), a CPLD (Complex Programmable Logic Device), etc.
[0015] The program may be installed in the computer from a program source. The program source may be, for example, a program distribution server or a storage medium readable by the computer. When the program source is a program distribution server, the program distribution server includes a storage resource for storing the processor and the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. Also, two or more programs may be realized as one program, or one program may be realized as two or more programs.
Example
[0016] The AI training data creation support system 1 extracts and collects training data for training an AI model from at least one training database. The trained AI model then analyzes the data to be analyzed. The AI model to be trained could be, for example, a traffic AI model in transportation (such as a model for predicting optimal routes), an industrial AI model related to product manufacturing (such as a model for estimating equipment failure diagnosis), or a healthcare AI model related to medicine.
[0017] In the following example, the AI model to be trained is a healthcare AI model used for analyzing and predicting the health status of individuals and groups, and the data to be analyzed is personal information, including information on an individual's health status. This makes it easy for the AI training data creation support system 1 to collect training data, allowing many people to collect training data without having to refer to personal information and consider how to collect it. Therefore, by collecting training data for the healthcare AI model, the AI training data creation support system 1 can collect training data while protecting the privacy of the people being analyzed. Personal information may also include information such as diagnostic history contained in medical records and genetic information. Furthermore, the training data to be collected will be appropriately changed depending on the AI model to be trained. For example, if the AI model to be trained is a failure diagnosis estimation model related to product manufacturing, the training data to be collected would be, for example, data that associates information on the characteristics of manufacturing equipment with the circumstances of failures.
[0018] <System Configuration> Figure 1 shows an example of a functional block diagram of the AI learning data creation support system 1 in Example 1. As shown in Figure 1, the AI learning data creation support system 1 is connected to the client device 2 and the external learning database server 3 via a network NW.
[0019] Client device 2 can transmit personal information (data to be analyzed) to be analyzed by the AI model, as well as a first query for extracting training data from the training database, which is input by the user of client device 2, to the AI training data creation support system 1. Furthermore, client device 2 is equipped with a display device, such as a screen, to display information to the user.
[0020] The external learning database server 3 has an external learning database, which is a type of learning database that stores learning data for training an AI model. The AI learning data creation support system 1 can extract learning data from the external learning database server 3 using queries.
[0021] The network NW can be a wired network or a wireless network. The communication network NW can be a global network like the internet or a local area network (LAN).
[0022] As shown in Figure 1, the AI learning data creation support system 1 comprises a learning data acquisition unit 11 and a supplementary query generation unit 12. The AI learning data creation support system 1 also stores a first learning database 21, a setting condition database 22, a search condition database 23, an algorithm requirement table 24, and an analysis content requirement table 25.
[0023] The learning data acquisition unit 11 accepts input of a personal profile (learning profile) from the user, as will be explained in detail later using the flowchart in Figure 11. The personal profile, as will be explained in detail using Figure 3, consists of item values corresponding to multiple data items and includes personal information (data to be analyzed) to be analyzed by the AI model to be trained, and information on the type of AI model (AI model algorithm, analysis content).
[0024] Furthermore, the learning data acquisition unit 11 acquires a first query (see Figure 3) to be used for extracting learning data. The learning data acquisition unit 11 calculates the number of first learning data to be extracted from the learning database by the first query using the learning database. The learning data acquisition unit 11 calculates the required number of learning data necessary for training the AI model using the information of the type of AI model included in the learning profile. The learning data acquisition unit 11 determines whether the number of first learning data is equal to or greater than the required number. If the learning data acquisition unit 11 determines that the number of first learning data is equal to or greater than the required number, it extracts the first learning data from the learning database using the first query and outputs it. If the learning data acquisition unit 11 determines that the number of first learning data is less than the required number, it causes the supplementary query generation unit 12 to generate supplementary queries based on the learning profile, and receives the supplementary queries generated by the supplementary query generation unit 12 and receives the received supplementary queries ERI extracts supplementary data from the training database and outputs it, and the first query Extract the first training data from the training database and output it.
[0025] The supplemental query generation unit 12 generates supplemental queries to supplement the training data, as will be explained in detail later using the flowcharts in Figures 12 and 15.
[0026] The first learning database 21 is a database that stores learning data and a statistical information file 21a. The statistical information file 21a includes statistical information such as information representing the number of records, information regarding the maximum and minimum values of data for each column, and a histogram representing the distribution of data for each column. Typically, databases have statistical information files similar to the statistical information file 21a. The AI learning data creation support system 1 can also access learning databases other than the first learning database 21 (for example, the external learning database of the external learning database server 3) and extract learning data.
[0027] The setting conditions database 22, as will be described in detail later using Figure 4, is a database that includes a range table, a statistical coefficient table, and domain item information. The range table stores, in association with at least one data item of the data to be analyzed for the learning profile, and for each of those at least one data item, a range of multiple item values. The statistical coefficient table stores, in association with one or more data items of the first learning data, and for each of those one or more data items, a range of statistical values and statistical coefficients. The domain item information stores, in association with domain items related to the personal profile (learning profile), and for each of the domain items, a range of domain items.
[0028] The search criteria database 23, as will be explained in detail later using Figure 5, is a database that stores multiple search criteria records that associate previously created data (personal information) with past queries used to extract training data related to that data.
[0029] The algorithm requirement table 24, as will be explained in detail later using Figure 6, stores the algorithms of the AI models and the algorithm requirement, which represents the number of training data points required for training the AI model of that algorithm.
[0030] The table 25, which details the analysis content and the number of training data required for training the AI model, will be described later using Figure 7, in detail. It stores the analysis content of the AI model and the number of training data required for training the AI model for that analysis content.
[0031] Figure 2 shows an example of the hardware configuration diagram of the AI learning data creation support system 1 in Embodiment 1. As shown in Figure 2, the AI learning data creation support system 1 has a processor 31, main memory 32, sub-memory 33, input device 34, output device 35, network I / F 36, and a bus 37 connecting them. The AI learning data creation support system 1 can be implemented using a general information processing device such as a PC or server computer.
[0032] The processor 31 reads data and programs stored in the secondary memory 33 into the main memory 32 and executes the processing defined by the program.
[0033] The main memory 32 has volatile elements such as RAM and stores programs executed by the processor 31 and data.
[0034] The secondary storage device 33 is a device that stores programs, data, etc., and has a non-volatile memory element such as an HDD (Hard Disk Drive) or SSD (Solid State Drive). The secondary storage device 33 stores the first learning database 21, the setting condition database 22, the search condition database 23, the algorithm requirement table 24, and the analysis content requirement table 25, as described above.
[0035] Furthermore, the learning data acquisition program 11a and the supplementary query generation program 12a are installed in the sub-memory 33. The learning data acquisition unit 11 and the supplementary query generation unit 12 described above using Figure 1 are realized when the processor 31 reads the learning data acquisition program 11a and the supplementary query generation program 12a stored in the sub-memory 33 into the main memory 32 and executes them.
[0036] The input device 34 is a device that accepts user input such as a keyboard or mouse, and acquires information entered through user input. The output device 35 is a device that outputs information such as a display, and presents information to the user, for example, by displaying it on a screen.
[0037] The network interface 36 is an interface for sending and receiving data via the network NW to devices such as client device 2 and external learning database server 3. The AI learning data creation support system 1 can send and receive data with devices such as client device 2 and external learning database server 3 that are connected to the network NW using the network interface 36. The network interface 36 can receive information input from the user of client device 2, and thus the network interface 36 also functions as an input device. Furthermore, the network interface 36 can send data to client device 2 via the network NW and display the data on the display of client device 2, and thus the network interface 36 also functions as an output device.
[0038] The client device 2 and the external learning database server 3 can be configured using the same hardware resources as the AI learning data creation support system 1.
[0039] <Various Data Structures> Figure 3 shows an example of a personal profile and the first query. The personal profile (learning profile) 302 has item values corresponding to each of the multiple data items 301, and includes personal information (data to be analyzed) to be analyzed by the AI model and information about the type of AI model. The data items 301 include multiple data items related to personal information (data to be analyzed) to be analyzed by the AI model, and multiple data items related to information about the type of AI model (AI model algorithm and analysis content).
[0040] The data items related to personal information (data to be analyzed) consist of diagnostic items and other items. Diagnostic items are the items that correspond to the analysis results that the AI model analyzes, and are, so to speak, the dependent variables. The data items other than diagnostic items are, so to speak, the dependent variables. The training data (first training data, first supplementary data, second supplementary data) is created so that the trained AI model can analyze the item values of the diagnostic items using the item values of the data items other than the diagnostic items.
[0041] In the personal profile 302 shown in Figure 3, the diagnostic item is "UA" as an example. After training, the AI model analyzes the personal information (data to be analyzed) in the personal profile (training profile) and outputs a value for "UA" as the analysis result. The diagnostic items can be arbitrarily set, for example, the amount of medication prescribed or the method of treatment on the human body.
[0042] Figure 3 shows an example of the search range (search conditions) 303 of the first query included in the first query. The first query is used to extract the first training data from the first training database 21 (training database). When the AI model is trained using supervised learning, the first training data can be used as training data. In the first training data, the item values for the diagnostic items represent correct and incorrect answers. Therefore, the first query and supplementary queries (first supplementary query and second supplementary query) are set up so that data containing item values corresponding to the diagnostic items can be extracted from the training database.
[0043] Figure 4 shows an example of the setting condition database 22 and the setting condition table 22a stored in the setting condition database 22. The setting condition database 22 has setting condition tables for each of the multiple diagnostic items (dependent variables). In the example in Figure 4, in addition to setting condition table 22a, setting condition tables 22b and 22c are shown as examples in the setting condition database 22, and the illustration of other setting condition tables is omitted.
[0044] The setting conditions table 22a includes a range table (data item 401, first range 403 to third range 405), a statistical coefficient table (data item 401, statistical value type 408 to second statistical coefficient 412), and domain item information (domain item 406, domain item range 407).
[0045] The range table stores, in association with at least one data item 401 of the personal information (data to be analyzed) in the personal profile (learning profile), and for each of the at least one data item 401, a range of multiple item values (e.g., first range 403 to third range 405).
[0046] Data item 401 is a data item corresponding to the personal profile. Importance 402 is the importance of the personal information item values in the personal profile. In Figure 4, Importance 402 is shown as three numbers, 1 to 3, as an example. Also, the smaller the number, the higher the importance. The first range 403 to the third range 405 are ranges of values used to set the search range to be included in the supplementary query (second supplementary query) when creating a supplementary query from the personal profile. Although ranges other than the first range 403 to the third range 405 are omitted from the illustration, the setting condition table 22a has the first range 403 to the nth range set. The first range to the nth range is set considering importance.
[0047] The statistical coefficient table stores, in association with one or more data items 401 of the first training data, the type of statistical value 408, the range of the statistical value (first statistical range 409, second statistical range 411, etc.), and the statistical coefficient (first statistical coefficient 410, second statistical coefficient 412, etc.) for each of the one or more data items 401.
[0048] Domain item information stores domain item 406 related to the personal profile (learning profile) and domain item range 407 associated with domain item 406. Domain item 406 is an item that is considered to have significant meaning (a large influence) with respect to the diagnostic items (dependent variables) of the personal profile (learning profile). Furthermore, domain item 406 may or may not be included in the data items of the personal profile. Domain item range 407 is the range of values considered to be reasonable for domain item 406.
[0049] Statistical value 408 is the type of statistical value (e.g., skewness) calculated for the first training data extracted from the training database by the first query. The statistical value of the first training data is calculated for data items for which the statistical value type is set in statistical value 408 of the setting conditions table 22a. As will be described in detail later, the first statistical range 409 is the statistical... Value 4 This is the range of statistical values for 08, and the first statistical coefficient 410 is the statistical coefficient corresponding to the first statistical range 409. Similarly, the second statistical range 411 is also statistical Value 4 This is the range of statistical values for 08, and the second statistical coefficient 412 is the statistical coefficient corresponding to the second statistical range 411. The setting condition table 22a stores multiple such combinations of statistical ranges and statistical coefficients.
[0050] Figure 5 shows an example of a search criteria database. The search criteria database 23 stores multiple search criteria records that associate previously created data to be analyzed (personal information) with past queries used to extract learning data related to the previously analyzed data. ID 501 is an ID that identifies the search criteria record. Past query 502 is the past query for each search criteria record. IF 503 is the interface when using past query 502. Search target 504 is the name of the database that the search criteria record is searching.
[0051] Personal Profile 505 includes previously created data (personal information) that has been analyzed. Modifiable Items 506 are data items from the previously analyzed data in Personal Profile 505 that are considered to have a small correlation with the analysis results of the AI model, and for which the search range may be expanded to any range. Creation Date 507 is the date and time the record was created.
[0052] Figure 6 shows an example of the Algorithm Requirement Table 24. The Algorithm Requirement Table 24 stores the algorithms of the AI model and the number of training data required by that algorithm. ID 601 is the ID that identifies the algorithm. Algorithm 602 is the algorithm of the AI to be trained. Characteristics 603 is the characteristics column for algorithm 602, and the number of algorithms required corresponding to algorithm 602 of the AI model is listed in the "Data size (Samples)" column.
[0053] Figure 6 shows examples of algorithm 602, including Logistic Regression, DNN (Deep Neural Network), and SVM (Support Vector Machine). Feature 603 also includes the number of algorithms required for the AI model's algorithm 602 (in the "Data size (Samples)" column). Feature 603 includes, as an example, "preparation_time," an estimate of the time required for training with the training data; "fairness," an example of a desirable statistical value for the training data; "AUC," an estimated value of AUC (Area Under Curve), an example of the accuracy of the analysis results of the trained AI model; and the number of algorithms required for the AI model's algorithm 602 (in the "Data size (Samples)" column).
[0054] Figure 7 shows an example of the Analysis Content Required Quantity Table 25. The Analysis Content Required Quantity Table 25 stores the analysis content 702 of the AI model and the Analysis Content Required Quantity 703, which represents the number of training data required for the AI model to learn the analysis content. In the Analysis Content Required Quantity Table 25 in Figure 7, ID 701 is the ID that identifies the analysis content of the AI model. Analysis content 702 is the analysis content of the AI to be trained, and is sometimes referred to as the "problem". Analysis Content Required Quantity 703 is the number of training data required for the AI model to learn the analysis content 702. As an example, Figure 7 shows classification and regression as analysis content 702, and examples of the corresponding Analysis Content Required Quantity 703 are shown.
[0055] <Processing Procedure> In Example 1, the user inputs their personal profile and a first query into client device 2. Next, client device 2 sends the personal profile and the first query to AI learning data creation support system 1. When AI learning data creation support system 1 receives the personal profile and the first query from client device 2, it starts the learning data acquisition process. Alternatively, the user may directly input their personal profile and the first query into AI learning data creation support system 1, and upon receiving this input, AI learning data creation support system 1 may start the learning data acquisition process.
[0056] Figure 8 is an explanatory diagram showing an example of a personal profile input screen displayed to the client device 2 for the user to enter their personal profile and first query. The personal profile input screen 800 shown in Figure 8 includes an input field 801 for entering the personal profile, a query input button 802, and a submit button 803.
[0057] Input field 801 is where the user enters their personal profile. For example, "UA" is entered in the "subject" field as the diagnostic item to be analyzed by the trained AI model, and "Male" is entered in the "sex" field as the gender. In addition, "DNN" is entered in the "AI" field as the algorithm of the AI model to be trained, "classification" is entered in the "problem" field as the analysis content of the AI model to be trained, and "50" is entered in the "required_auc" field as AUC (Area Under Curve), which is an example of the accuracy of the analysis results of the trained AI model, and "50" representing the target value of "50%".
[0058] When the user presses the query input button 802, a query input screen for entering the first query is displayed on the client device 2. When the user presses the send execution button 803, the personal profile and the information of the first query entered by the user are sent from the client device 2 to the AI learning data creation support system 1 via the network NW.
[0059] Figures 9 and 10 are explanatory diagrams showing examples of query input screens displayed to the client device 2 for a user to enter a first query. The query input screen 900a shown in Figure 9 has a field 901a where the user enters the first query. The query input screen 900b shown in Figure 10 has a list selection button 901b for the user to select a data item table for entering the contents of the first query, and a data item table 902b. In the example in Figure 10, the user selects "Patient_basic_table" with the list selection button 901b, and "Patient_basic_table" is displayed in the data item table 902b. When the user clicks the checkbox in the data item table 902b to set the search conditions to be included in the first query, the client device 2 converts the data item table 902b into the first query.
[0060] Next, using Figure 11, we will explain the learning data acquisition process executed by the learning data acquisition unit 11 of the AI learning data creation support system 1. Figure 11 is a flowchart showing an example of the learning data acquisition process of the AI learning data creation support system 1. As mentioned above, when the AI learning data creation support system 1 receives the personal profile and the first query from the client device 2, it starts the learning data acquisition process shown in the flowchart in Figure 11.
[0061] The AI learning data creation support system 1 (processor 31) stores the profile and the first query received from the client device 2 (step S101).
[0062] Next, the AI learning data creation support system 1 extracts and saves a setting condition table 22a related to the diagnostic items of the personal profile from the setting condition database 22 (see Figure 4) (step S102).
[0063] Next, the AI learning data creation support system 1 calculates and stores the number and statistics of the first learning data extracted from the first learning database 21 by the first query using the statistics information file 21a of the first learning database 21 (step S103). Here, the AI learning data creation support system 1 estimates the number of the first learning data using the statistics information file 21a of the first learning database 21 in a known method as shown below. Furthermore, for all data items for which a statistical value type is set in statistical value 408 (see Figure 4) of the setting condition table 22a, the AI learning data creation support system 1 calculates the statistical value of the type set in statistical value 408 for the first learning data using the statistics information file 21a in a known method and sets it as the statistical value. Here, the AI learning data creation support system 1 calculates the number and statistical values of the first learning data using the statistical information file 21a. This makes it easier for the AI learning data creation support system 1 to calculate the number and statistical values of the first learning data compared to when the AI learning data creation support system 1 extracts the first learning data from the first learning database 21 and calculates the number and statistical values of the first learning data.
[0064] Databases typically have a statistics file. This statistics file contains statistical information such as the number of records, the maximum and minimum values of data for each column, and histograms showing the distribution of data for each column. For example, it is possible to estimate the number of records Ra in which the value of data item A is recorded. Furthermore, the number Raa in which the value of data item A falls within range A can be estimated from the histogram information. This allows us to estimate the proportion Rpa (Rpa = Raa / Ra) of records in which the value of data item A falls within range A. Similarly, it is possible to estimate the number Rb in which the value of data item B is recorded. And it is possible to estimate the proportion Rpb in which the value of data item B falls within range B. Therefore, the number of records AB in which the value of data item A is within range A and the value of data item B is within range B can be estimated as the product of the number of records Ra in which the value of data item A is recorded, the percentage Rpa of records in which the value of data item A is within range A, and the percentage Rpb of records in which the value of data item B is within range B (number of records AB = number of records Ra × percentage Rpa × percentage Rpb). In this way, the number of first training data is calculated by calculating the product of the number of records in which the data items are recorded and the percentage of records. Furthermore, the statistical values "skewness" and "kurtosis" of the data item values in the first training data can be estimated from the histogram of the data items, etc.
[0065] Also, for example, as shown in Figure 4 Setting Conditions Table 22a In this example, the statistical value 408 has the BMI term set to "skewness," and the AI learning data creation support system 1 is statistical Using information file 21a, the "skewness" value of the BMI in the first training data is calculated and used as the BMI statistic. Similarly, in the example in Figure 4, statistical values such as "skewness" are calculated for the LDL-C and γGT terms, and used as the statistical values for each term. Note that "skewness" is an example of a statistical value that represents the variability of the first training data, and other statistical values may be used instead of "skewness". For example, "kurtosis" may be used as the statistical value, or both "skewness" and "kurtosis" may be used.
[0066] Next, the AI learning data creation support system 1 associates the personal profile with the first query and saves it in the search condition database (see Figure 5) (step S104). Here, the data item with the lowest importance, which is assigned importance level 3 in the setting condition table 22a, may be set as the modifiable item in the search condition record (see Figure 5) (a data item for which the search range may be expanded to any range, see Figure 5).
[0067] Next, the AI learning data creation support system 1 calculates the upper limit of the required number, and based on the upper limit of the required number, the AI model algorithm (type of AI model), the setting condition table 22a, and the statistical values of the first learning data, calculates the number of data required for training the AI model as the required number and saves it (step S105). Here, the upper limit of the required number is an approximate value of the number of first learning data that the AI learning data creation support system 1 can acquire from the first learning database 21 within a first acceptable time interval (e.g., 6 hours) that is considered sufficiently short. The first acceptable time interval is set in advance. If the number of first learning data is less than or equal to the upper limit of the required number (number of first learning data ≤ upper limit of the required number), it can be determined that the time required to acquire the first learning data is sufficiently short. On the other hand, if the number of first learning data is greater than the upper limit of the required number (number of first learning data > upper limit of the required number), it can be determined that the time required to acquire the first learning data is too long.
[0068] The required upper limit is, for example, the product of the first training data acquisition speed and the first allowable time interval. The first training data acquisition speed represents the number of first training data that can be acquired from the first training database 21 per unit time. The AI training data creation support system 1 calculates the first training data acquisition speed based, for example, on the specifications of the processor 31 such as the number of cores and clock speed of the processor 31, the estimated utilization rate (operating status) of the processor 31 that can be allocated to acquire the first supplementary data, and the read / write speed of the main memory 32. The AI training data creation support system 1 may also measure the first training data acquisition speed by executing a predetermined program. The AI training data creation support system 1 then calculates the product of the first training data acquisition speed and the first allowable time interval and sets this as the required upper limit.
[0069] To calculate the required number, the following are used: the upper limit of the required number, the required number of algorithms table 24, the required number of analysis content tables 25, the statistical values calculated in step S103, and the setting conditions table 22a. As mentioned above, information on the algorithm and analysis content of the AI model to be trained is included in the personal profile. For example, in the personal profile shown in Figure 3, the algorithm is "Deep Neural Network (DNN)" and the analysis content is "Classification".
[0070] In calculating the required number, first, the number of algorithms required for the AI model is extracted from the algorithm requirement table 24, shown as an example in Figure 6, and the number of analysis content required for the AI model is extracted from the analysis content requirement table 25, shown as an example in Figure 7. The larger of the algorithm requirement and the analysis content requirement is taken as the model requirement M.
[0071] For example, in the algorithm requirement table 24 shown in Figure 6, the number of algorithms required for the AI model algorithm "DNN" is 100,000. Similarly, in the example analysis content requirement table 25 shown in Figure 7, the number of algorithms required for the AI model analysis content "classification" is 10,000. The larger of these two data counts, 100,000, becomes the model requirement M (model requirement M = 100,000). Note that while the algorithm requirement table 24 and analysis content requirement table 25 were used above, they can be modified as appropriate as follows. For example, a database could be created in advance that combines the algorithm requirement table 24 and analysis content requirement table 25, storing pairs of algorithms and analysis content in association with the model requirement M. Alternatively, the model requirement M could be calculated using only the algorithm requirement table 24. Or, the model requirement M could be calculated using only the analysis content requirement table 25. Furthermore, the required number of models M may be calculated by considering factors other than the AI model's algorithm and analysis content.
[0072] Furthermore, for each data item from which statistical values are calculated, statistical coefficients are calculated as follows, and the largest of the calculated statistical coefficients is defined as the maximum statistical coefficient C. The product of the model required number M and the maximum statistical coefficient C is defined as the required number D (required number D = model required number M × maximum statistical coefficient C). Moreover, if the required number D is greater than the upper limit of the required number (required number D > upper limit of the required number), the required number D is set as the upper limit of the required number. The statistical coefficients are those corresponding to the range containing the statistical value from the first statistical range to the nth statistical range (any of the first to nth statistical coefficients).
[0073] In the example of setting condition table 22a in Figure 4, if the statistical value of BMI is 0.4, the statistical value (0.4) falls into the second statistical range 411, and the second statistical coefficient 412 corresponding to the second statistical range 411, which is 10, is used as the statistical coefficient for the data item BMI (statistical coefficient = 10). Similarly, if the statistical value of the data item LDL-C is 0.1, the statistical value falls into the first statistical range 409, and the value of the first statistical coefficient 410, which is 1, becomes the statistical coefficient for the data item LDL-C (statistical coefficient = 1). Furthermore, if the maximum value of the statistical coefficients for all data items is 10, the maximum statistical coefficient C becomes 10. As described above, if the required number of models M is 100,000, the required number D will be 1,000,000 (= required number of models M 100,000 × maximum statistical coefficient 10).
[0074] Furthermore, if the required number D (required number D = product of the required model number M and the maximum statistical coefficient C) is greater than the upper limit of the required number (required number D > upper limit of the required number), the time required to acquire the required number D of the first training data is considered to be too long, so the required number D is set to the upper limit of the required number (required number D = upper limit of the required number). This allows the AI training data creation support system 1 to generate (extract) the first training data, as well as the first and second supplementary data described later, more reliably. Note that in step S105, the AI training data creation support system 1 does not need to calculate the upper limit of the required number, and furthermore, if the required number D is greater than the upper limit of the required number (required number D > upper limit of the required number), it does not need to set the required number D to the upper limit of the required number.
[0075] Furthermore, the required number D may be calculated by considering the AI model's learning method. For example, similar to the statistical coefficients mentioned above, statistical coefficients related to the learning method may be created to calculate the required number D. Examples of learning methods include the leave-one-out method, which involves extracting one training data set from the entire training data set as test data and performing cross-validation using the remaining training data as training data, as well as the hold-out method and the cross-validation method.
[0076] Next, returning to Figure 11, the AI learning data creation support system 1 determines whether the number of first learning data calculated in step S103 is greater than or equal to the required number calculated in step S105 (required number ≤ number of first learning data) (step S106). If it is determined that the number of first learning data is greater than or equal to the required number (required number ≤ number of first learning data) (step S106: YES), the system proceeds to step S107. If it is determined that the number of first learning data is less than the required number (required number > number of first learning data) (step S106: NO), the system proceeds to step S108.
[0077] Next, the AI learning data creation support system 1 uses the first query to extract the first learning data from the first learning database, outputs the extracted first learning data, and terminates processing (step S107). Here, the output of the first learning data may be one of the following: For example, the first learning data is sent to the client device 2. A file containing the first learning data is sent to the client device 2. The file containing the first learning data is stored in the sub-memory 33. The first learning data is output to the output device 35 and presented to the user of the AI learning data creation support system 1. The first learning data is sent to the client device 2, and the client device 2 presents the first learning data to the user. Here, the presentation to the user by the client device 2 may be output to the display of the client device 2. For example, the standard output displayed on the display of the client device 2 may be used. Standard output is the data output destination that a device (such as the device's operating system) uses as a standard when a program running on the computer is not specifically designated.
[0078] Next, the AI learning data creation support system 1 calculates the difference between the required number and the number of first learning data, and saves the difference as the target number of supplements (target number of supplements = required number - number of first learning data) (step S108).
[0079] Next, the AI learning data creation support system 1 calls a supplement query generation subroutine (step S109). The supplement query generation subroutine is a process executed by the supplement query generation unit 12 of the AI learning data creation support system 1, and generates supplement queries in order to supplement the learning data.
[0080] Next, the AI learning data creation support system 1 extracts the first learning data from the first learning database using the first query, extracts supplemental data from the database using the supplemental query, outputs the first learning data and supplemental data, and terminates processing (step S110). Here, the output of the first learning data and supplemental data may be the following, similar to step S107 described above. For example, the first learning data and supplemental data are sent to the client device 2. A file containing the first learning data and supplemental data is sent to the client device 2. The file containing the first learning data and supplemental data is stored in the sub-memory 33. The first learning data and supplemental data are sent to the client device 2, and the client device 2 presents the first learning data and supplemental data to the user. Here, the presentation to the user by the client device 2 may be output to the display of the client device 2. For example, the standard output displayed on the display of the client device 2 may be used.
[0081] Next, referring to Figure 12, we will explain the processing of the supplemental query generation subroutine executed by the supplemental query generation unit 12 of the AI learning data creation support system 1 using Figures 13 and 14. Figure 12 is a flowchart showing an example of the processing of the supplemental query generation subroutine.
[0082] The AI learning data creation support system 1 extracts at least one search condition record from the search condition database that contains past data to be analyzed whose similarity to the personal information (data to be analyzed) of the personal profile (learning profile) is greater than a predetermined similarity threshold, and saves the past queries of the extracted at least one search condition record as first supplement query candidates (step S201). As described above using Figure 3, the personal information of the personal profile includes item values of various data items.
[0083] Similarity is, for example, the ratio of the number of data items in personal information in a personal profile (excluding the number of name and ID data items) to the total number of data items (excluding the number of name and ID data items) included in both the personal information in the personal profile and the historical data (personal information) analyzed in the search criteria records. In other words, "Similarity = Number of data items included in both / Number of data items in personal information". Furthermore, the more data items included in both the personal information in the personal profile and the historical data (personal information) analyzed in the search criteria records, the higher the similarity. Name and ID are considered to be information that is less related to an individual's characteristics, while other data items are considered to be more related to an individual's characteristics. In calculating similarity, by excluding the number of name and ID data items from the total number of data items, the similarity becomes a similarity related to an individual's characteristics. This results in a suitable similarity.
[0084] For example, suppose the personal information data items in a personal profile are "ID, diagnostic item, name, age, height, BMI, LDL-C," and the data items in the historical data to be analyzed in the search criteria record are "diagnostic item, name, age, height." The number of data items related to an individual's characteristics included in the personal profile is 5, excluding the data items "ID" and "name." The number of data items included in both the personal information in the personal profile and the historical data to be analyzed (personal information) in the search criteria record is 3, the same as the data item "diagnostic item, age, height." The similarity (= number of data items included in both / number of data items in personal information) is 3 / 5 = 0.6.
[0085] The similarity threshold is a pre-set threshold for similarity, for example, 0.5.
[0086] In step S201, the domain field range (see Figure 4) is added as a search condition to the past query of the search condition record containing the past analysis target data, where the similarity to the personal information of the individual profile is greater than the similarity threshold, and this is used as the first candidate for supplementary query. For example, in the example of the domain field range shown in Figure 4, domain field range 407 is "4.2≦HbA1c≦6.2". In the example in Figure 4, the AI learning data creation support system 1 first extracts the past analysis target data from the setting condition database 22, where the similarity to the personal information of the individual profile is greater than the similarity threshold. Then, the query with the domain field range "4.2≦HbA1c≦6.2" added as a search condition to the past query of the search condition record containing the extracted past analysis target data is used as the first candidate for supplementary query.
[0087] As described above using Figure 4, the domain items relate to the individual profile (learning profile). Furthermore, the domain items are considered to be of significant importance (have a large influence) with respect to the diagnostic items (dependent variables) of the individual profile (learning profile). The domain item range is the range of values considered to be reasonable for the domain items. The first supplementary data, which is the learning data, is generated (extracted) based on the first supplementary query selected from the first supplementary query candidates. Therefore, the AI learning data creation support system 1 generates the first supplementary query that includes the domain item range as a search condition by adding the domain item range as a search condition to the first supplementary query candidates. This makes the first supplementary data (learning data) more suitable data with a higher correlation to the diagnostic items (dependent variables).
[0088] Note that there are 506 modifiable fields in the search criteria record (Figure 5Alternatively, the search range (see reference) can be expanded appropriately (for example, by 10%) from past queries to generate a query, and the generated query with search conditions based on domain field ranges can be used as the first candidate for supplementary query. Alternatively, the personal profiles of the search condition records in the search condition database 23 can be extracted using the first query, and the past queries related to the extracted personal profiles can be used with search conditions based on domain field ranges to create the first candidate for supplementary query.
[0089] Next, the AI learning data creation support system 1 estimates the number of first supplement candidate data extracted from the learning database for each first supplement query candidate using the statistical information file of the learning database, calculates an upper limit on the number of data, and selects first supplement query candidates whose number of first supplement candidate data is less than or equal to the upper limit on the number of data as the first supplement query, and saves the first supplement query in association with the number of first supplement queries (step S202). Here, as shown in the search target 504 of Figure 5, depending on the first supplement query candidate, the corresponding learning database may be a learning database other than the first learning database 21 owned by the AI learning data creation support system 1. If the learning database corresponding to the first supplement query candidate is the first learning database 21, the number of first supplement query candidates is the number of data (number of data records) obtained by removing data that overlaps between the first supplement candidate data and the first data from the first supplement candidate data. The number of duplicate data entries (data count) is the number of data entries extracted from the first training database 21 by a query that adds the search conditions of the first supplementary query candidate to the search conditions of the first query. The number of first supplementary candidate data entries is the number of data entries extracted by the first supplementary query candidate minus this number of duplicate data entries. The AI training data creation support system 1 calculates the number of data entries extracted by the first supplementary query candidate and the number of duplicate data entries using the first training database 21, and then calculates the number of first supplementary candidate data entries by taking the difference between the number of data entries extracted by the first supplementary query candidate and the number of duplicate data entries.
[0090] The training database typically contains a statistical information file. In step S202, the AI training data creation support system 1 estimates the number of first supplemental data to be extracted by the first supplemental query candidate, using the statistical information file contained in the training database designated as the first supplemental query candidate, in the same manner as in step S103 of the training data acquisition process in Figure 11.
[0091] The upper limit of the number of data points is an estimated number of first supplement candidate data points that the AI learning data creation support system 1 can acquire from the learning database within a second allowable time interval (e.g., 6 hours) that is considered sufficiently short. The second allowable time interval is set in advance. The AI learning data creation support system 1 calculates the upper limit of the number of data points to acquire by, for example, the product of the first supplement data acquisition speed and the second (predetermined) allowable time interval. The first supplement data acquisition speed represents the number of first supplement candidate data points that can be acquired from the learning database per unit time. The AI learning data creation support system 1 calculates the first supplement data acquisition speed based on, for example, the specifications of the processor 31, such as the number of cores and clock speed of the processor 31, the estimated utilization rate (operating status) of the processor 31 that can be allocated to acquire the first supplement candidate data, the read / write speed of the main memory 32, and the transmission / reception speed with the network. Furthermore, the AI learning data creation support system 1 may execute a predetermined program to measure the speed of acquiring the first supplementary data.
[0092] If the number of first-order supplementary data is less than or equal to the upper limit of the number of data (number of first-order supplementary data ≤ upper limit of the number of data), then the time required to acquire the first-order supplementary data can be judged as sufficiently short. On the other hand, if the number of first-order supplementary data is greater than the upper limit of the number of data (number of first-order supplementary data > upper limit of the number of data), then the time required to acquire the first-order supplementary data can be judged as too long.
[0093] The AI learning data creation support system 1 selects the first supplemental query candidates whose number of first supplemental candidate data is less than or equal to the upper limit of the number of data (number of first supplemental candidate data ≤ upper limit of the number of data) as the first supplemental query. The AI learning data creation support system 1 also stores the first supplemental queries in association with the number of first supplemental queries (number of first supplemental candidate data). This allows the AI learning data creation support system 1 to more reliably generate (extract) the first supplemental data using the first supplemental queries. In step S202, the AI learning data creation support system 1 may not calculate the upper limit of the number of data, and may also select all first supplemental query candidates as the first supplemental query regardless of the upper limit of the number of data.
[0094] Let's assume that m (or more) of the first supplementary queries were extracted. Let's label the extracted queries 1 through m as the first supplementary queries.
[0095] Next, the AI learning data creation support system 1 generates and saves the second supplemental queries 1 to n based on the personal profile and the range table (setting condition table 22a) (step S203).
[0096] Figure 13 illustrates how to generate the second supplemental query. Figure 13 includes data item 401, personal information 1301, first range 403, column 1302 of second supplemental query 1, second range 404, column 1303 of second supplemental query 2, third range 405, and column 1304 of second supplemental query 3. Here, data item 401, first range 403, second range 404, and third range 405 are the same as the range table in setting condition table 22a shown in Figure 4. Second supplemental query 1, shown in column 1302 of second supplemental query 1, is a query that includes a search range that expands the item values of personal information 1301 to the first range 403. For example, in a row where data item 401 is "Diagnostic Item", personal information 1301 is UA, the first range is ±5, and due to the nature of UA, the minimum value of UA is 0, so the search range for "Diagnostic Item" in the second supplemental query 1 is 0 to 10. Similarly, in a row where data item 401 is "Age", personal information 1301 is 68, the first range is ±3, so the search range for the second supplemental query 1 is 65 to 71. As explained above, the second supplemental query 2 shown in column 1303 of the second supplemental query 2 and the second supplemental query 3 shown in column 1304 of the second supplemental query 3 are generated, and furthermore, the second supplemental queries 4 to n (not shown) are generated corresponding to the fourth range to the nth range.
[0097] Next, the AI learning data creation support system 1 estimates the number of second supplemental data extracted by each second supplemental query 1 to n, and stores it in association with the second supplemental queries 1 to n (step S204).
[0098] Here, the AI learning data creation support system 1 estimates the number of second supplemental data 1 to n extracted from the first learning database 21 by the second supplemental queries 1 to n using the statistical information file 21a of the first learning database 21, in the same manner as in step S202 described above. That is, the number of second supplemental data 1 to n is the number of data (data count) obtained by subtracting the data that overlaps between the second supplemental data 1 to n and the first data from the second supplemental data 1 to n. The number of overlapping data (data count) is the number of data (data count) extracted from the first learning database 21 by a query that adds the search conditions of the second supplemental queries 1 to n to the search conditions of the first query. The number of second supplemental data 1 to n is the number obtained by subtracting this number of overlapping data (data count) from the number of data (data count) extracted by the second supplemental queries 1 to n. The AI learning data creation support system 1 calculates the number of data extracted by the second supplementary queries 1 to n and the number of duplicate data using the first learning database 21. Furthermore, it calculates the number of second supplementary data 1 to n by taking the difference between the number of data extracted by the second supplementary queries 1 to n and the number of duplicate data.
[0099] Furthermore, the training database used to extract the second supplementary data 1 to n in the second supplementary queries 1 to n may be a training database other than the first training database 21 (for example, the external training database of the external training database server 3). Also, second supplementary queries in which the number of second supplementary data is greater than the required upper limit (number of second supplementary data > required upper limit) may be excluded from the second supplementary queries 1 to n. This allows the AI training data creation support system 1 to generate (extract) the first supplementary data more reliably.
[0100] Next, the AI learning data creation support system 1 adds the top 1 to 5 (a predetermined number) queries from the first supplementary queries 1 to m, in order of priority, to the supplementary query list (not shown) (step S205). Here, priority is, for example, the number of first supplementary data. That is, the first supplementary query with a large number of first supplementary data is given priority and added to the supplementary query list. The supplementary query list is a list in which queries selected from the first supplementary queries 1 to m and the second supplementary queries 1 to n to supplement the first queries are registered in order of the number of supplementary data.
[0101] Next, the AI learning data creation support system 1 selects the top-ranked second supplementary query from among the second supplementary queries 1 to n, and then... 2 The number of supplemental data points is associated with the number of supplemental queries, and they are added to the supplemental query list (step S206). Here, "higher ranking" means that queries that are closer to the second supplemental query 1 are ranked higher (second supplemental query 1 > second supplemental query 2 > ... > second supplemental query n).
[0102] Furthermore, the number of second supplemental queries and their corresponding supplemental data that are currently registered in the supplemental query list will be replaced with the number of second supplemental queries and their corresponding supplemental data that are ranked number one and have not been previously registered in the supplemental query list. This means that the second supplemental queries registered in the supplemental query list will be modified to broaden the search range for at least one data item, the number of second supplemental data for the modified second supplemental query will be calculated, and the number of second supplemental data registered in the supplemental query list will be replaced with the calculated number of second supplemental data.
[0103] Next, the AI learning data creation support system 1 checks if the sum of the number of first supplementary data and the number of second supplementary data registered in the supplementary query list is equal to or greater than the target number of supplementary data (Σ number of supplementary data in the supplementary query list). ≧ Determine whether the target number of replenishments has been met (Step S207). The sum of the number of first replenishment data and the number of second replenishment data registered in the replenishment query list is greater than or equal to the target number of replenishments (Σ number of replenishment data in the replenishment query list).≧ If it is determined that the target number of replenishments is not met (Step S207: YES), proceed to Step S208, where the sum of the number of first replenishment data and the number of second replenishment data registered in the replenishment query list is less than the target number of replenishments (Σ number of replenishment data in the replenishment query list). < If it is determined that the target replenishment quantity is not met (Step S207: NO), return to Step S205.
[0104] If, at this point, the sum of the number of first supplementary data points and the number of second supplementary data points registered in the supplementary query list is determined to be greater than or equal to the target number of supplementary data points (target number of supplementary data points = required number - number of first training data points) (target number of supplementary data points = required number - number of first training data points ≤ Σ number of supplementary data points in the supplementary query list) (step S207: YES), then we can consider the following: That is, the total number of data points extracted by the queries registered in the supplementary query list plus the number of first training data points extracted by the first query will be greater than or equal to the required number of data points needed to train the AI model (required number ≤ number of first training data points + Σ number of supplementary data points in the supplementary query list). As a result, a sufficient number of training data points can be collected using the queries registered in the supplementary query list and the first query.
[0105] Next, the AI learning data creation support system 1 presents the supplemental queries (first supplemental query and second supplemental query) and the number of supplemental data entries registered in the supplemental query list to the user in order of priority (step S208). That is, it presents the supplemental queries to the user using an output device so that the user can select which supplemental queries to use from the first supplemental query and the second supplemental query. Here, the presentation to the user is such that when the AI learning data creation support system 1 sends the supplemental query list to the client device 2, the client device 2 displays the supplemental queries and the number of supplemental data entries registered in the supplemental query list on its display in order of priority, based on the supplemental query list. Furthermore, the user of the client device 2 selects the supplemental query to be used to supplement the first query from the displayed supplemental queries.
[0106] Note: Client device 2 no de Instead of displaying it on the display, the data may be output to the output device 35 of the AI learning data creation support system 1 and presented to the user of the AI learning data creation support system 1, allowing the user to select a supplementary query.
[0107] Figure 14 is an explanatory diagram showing an example of a supplement query display screen, which is displayed on the client device 2's screen to show the user the number of supplement queries and supplement data registered in the supplement query list.
[0108] In the supplement query display screen 1400 shown in Figure 14, supplement queries are displayed from top to bottom in order of priority. Here, priority is, for example, the number of supplement data items. The supplement query display screen 1400 includes a submit button 1401 and a target number of supplement items 1402. The supplement query display screen 1400 also includes a checkbox 1411 and the number of supplement data items extracted by supplement query 1410 1412, related to supplement query 1410 with priority 1. The supplement query display screen 1400 also includes a checkbox 1421 and the number of supplement data items extracted by supplement query 1420 1422, related to supplement query 1420 with priority 2. The supplement query display screen 1400 also includes a checkbox 1431 and the number of supplement data items extracted by supplement query 1430 1432, related to supplement query 1430 with priority 3.
[0109] The user of client device 2 can select supplementary queries to be used to supplement the first query by pressing checkboxes 1411, 1421, and 1431. Once the user has finished selecting the supplementary queries, they press the send button 1401. This causes client device 2 to send the supplementary queries selected by the user to the AI learning data creation support system 1.
[0110] In the supplement query display screen 1400 in Figure 14, the supplement queries 1410 and 1420, which have priority 1 and priority 2 respectively and are checked in the checkboxes 1411 and 1421, are selected as supplement queries, while the supplement query 1430, which has priority 3 and is unchecked in the checkbox 1431, is not selected.
[0111] Next, returning to Figure 12, the AI learning data creation support system 1 accepts the input of the supplementary query selected by the user, saves it as a supplementary query, and terminates processing (step S209). After terminating processing, the AI learning data creation support system 1 performs the processing of step S110 of the learning data acquisition process in Figure 11. In step S110, the AI learning data creation support system 1 extracts the first learning data from the learning database using the first query, and extracts supplementary data (first supplementary data, second supplementary data) from the learning database using the supplementary query selected by the user in step 209. Then, the AI learning data creation support system 1 outputs the first learning data and supplementary data to the output device. 3 Output is generated using either I / F 5 or network I / F 36.
[0112] Thus, in Example 1, the AI learning data creation support system 1 generates supplementary queries that can be used to acquire supplementary data to supplement the first learning data. This allows for efficient collection of learning data for training the AI model.
[0113] Furthermore, the AI learning data creation support system 1 can easily collect learning data for training an AI model by outputting first learning data and supplementary data.
[0114] Furthermore, the AI training data creation support system 1 calculates the required number based on the algorithm and analysis content of the AI model to be trained. Therefore, the required number is set more appropriately, and a more reasonable number of training data can be collected.
[0115] Furthermore, the AI learning data creation support system 1 calculates the required number based on the statistical values of one or more data items in the first learning data. Therefore, the required number is set more appropriately, and a more reasonable number of learning data can be collected.
[0116] Furthermore, the AI learning data creation support system 1 generates a first supplementary query from past queries created in the search condition database 23. This allows for the efficient collection of learning data for training the AI model.
[0117] Furthermore, the AI learning data creation support system 1 generates a second supplementary query using personal information (data to be analyzed) from the individual profile (learning profile). This allows for the efficient collection of learning data for training the AI model.
[0118] Furthermore, the system accepts input for a first and second supplementary query selected by the user, and uses either the first or second supplementary query selected by the user to create supplementary data. This allows the training data collected using the supplementary queries to be made into more appropriate training data. [Examples]
[0119] In Example 1, in the processing of the supplement query generation subroutine shown in the flowchart in Figure 12, the user selects a supplement query from the first supplement query and the second supplement query registered in the supplement query list (steps S208 to S209 in Figure 12). The difference between Example 2 and Example 1 is that the AI learning data creation support system 1 generates the supplement query without the user selecting a supplement query. In Example 2, parts and configurations that have the same functions as those in Example 1 are given the same reference numerals and their descriptions are omitted.
[0120] Figure 15 is a flowchart showing an example of the processing of the supplement query generation subroutine in Example 2. The processing in steps S301 to S307 of the flowchart in Figure 15 is the same as the processing in steps S201 to S207 of the flowchart of the supplement query generation subroutine in Example 1 shown in Figure 12, so the explanation is omitted.
[0121] In step S308, the AI learning data creation support system 1 saves the supplemental queries registered in the supplemental query list as supplemental queries and terminates the process.
[0122] Thus, in Example 2, supplementary queries are automatically generated without the user having to select them, allowing for efficient collection of training data. [Examples]
[0123] In Example 1, the first query generated by the user of client device 2 is used in the training data acquisition process. In Example 3, unlike Example 1, the first query is generated by the AI training data creation support system 1. In Example 3, parts and configurations that have the same functions as those in the AI training data creation support system 1 of Example 1 are given the same reference numerals and their descriptions are omitted.
[0124] In Example 3, the AI learning data creation support system 1, upon receiving a personal profile from the client device 2, starts the learning data acquisition process shown in the flowchart in Figure 16.
[0125] Figure 16 is a flowchart showing an example of the training data acquisition process in Example 3.
[0126] The AI learning data creation support system 1 stores the personal profile received from the client device 2 (step S401).
[0127] Next, the AI learning data creation support system 1 reads and saves the setting conditions table 22a related to the personal profile from the setting conditions database 22 (step S402). Note that the process in step S402 is the same as the process in step S102 of the flowchart of the learning data acquisition process in Example 1 shown in Figure 11. Also, as described above with reference to Figure 4, the setting conditions table 22a includes a range table.
[0128] Next, the AI learning data creation support system 1 generates and saves a first query based on the range table (setting condition table 22a) and the personal profile (step S403). Here, the first query is the second supplementary query 1 of Example 1, which was explained using Figure 13.
[0129] Accordingly, in the processing of the supplement query generation subroutine in Example 3 (see Figure 12), in the process of generating the second supplement query 1 to the second supplement query n, which corresponds to step S203 in the flowchart of Figure 12, the second supplement query 2 to the second supplement query n of Example 1 are generated, and these are designated as the second supplement query 1 to the second supplement query n-1 of Example 3. That is, the second supplement query 2 to the second supplement query n-1 of Example 1 n Therefore, we move it up one step and designate it as the second supplementary query 1 to the second supplementary query n-1 of Example 3.
[0130] The processes in steps S404 to S411 of the flowchart shown in Figure 16 are the same as the processes in steps S103 to S110 of the flowchart for the learning data acquisition process in Example 1 shown in Figure 11, so their explanation is omitted.
[0131] Thus, in Example 3, the AI training data creation support system 1 generates the first query, eliminating the need for the user to create the first query. This allows for efficient collection of training data for training the AI model.
[0132] It should be noted that the present invention is not limited to the embodiments described above, but includes various modifications and equivalent configurations within the spirit of the attached claims. For example, the embodiments described above are explained in detail to make the present invention easier to understand, and the present invention is not necessarily limited to having all the configurations described. Furthermore, some of the configurations of one embodiment may be replaced with those of another embodiment. Furthermore, some of the configurations of one embodiment may be added to those of another embodiment. Furthermore, some of the configurations of each embodiment may be added, deleted, or replaced with other configurations. [Explanation of Symbols]
[0133] 1: Training data creation support system 2: Client device 3: External learning database server 11: Training Data Acquisition Unit 11a: Training data acquisition program 12: Supplemental Query Generation Unit 12a: Supplemental query generation program 21: First learning database 21a: Statistical information file 22: Configuration Condition Database 22a: Setting conditions table 23: Search Criteria Database 24: Algorithm Required Number Table 25: Table of required number of analysis items 31: Processor 32: Main memory 33: Secondary storage device 34: Input device 35: Output device 36: Network Interface 37: Bus
Claims
1. An AI training data creation support system that extracts and collects training data for training an AI model from at least one training database, The system comprises a storage device for storing at least one program, a processor for executing the program stored in the storage device, and an input device for receiving input from a user. The aforementioned storage device is An algorithm requirement table stores the algorithm of the AI model and the number of training data required for training the AI model using the algorithm, in association with the algorithm requirement table. An analysis content required number table stores the analysis content of the AI model and the number of training data required for training the AI model based on the analysis content, in association with each other. A search condition database that stores multiple search condition records that associate previously created data to be analyzed with past queries used to extract the learning data related to the said data to be analyzed, The system includes a statistical coefficient table that stores one or more data items in association with the range of statistical values and statistical coefficients for each of those one or more data items, The processor executes the program, The system accepts input of a first query configured to extract from the learning database a learning profile consisting of item values corresponding to each of multiple data items, including the data to be analyzed by the AI model, the algorithm and analysis content of the AI model, and learning data including item values corresponding to the diagnostic items included in the data to be analyzed in the learning profile. The number of first training data extracted from the training database by the first query is calculated using the training database. With respect to the first training data, statistical values are calculated for each data item. The statistical coefficients associated with the range to which the statistic value belongs are extracted from the statistical coefficient table. Referring to the table of required algorithms and the table of required analysis content, the number of algorithms and analysis content required for training the AI model is calculated using the information on the AI model's algorithms and analysis content included in the training profile. The larger of the number of algorithms required and the number of analysis contents required is taken as the number of models required M, and the product of the largest statistical coefficient among the calculated statistical coefficients, which is the maximum statistical coefficient, and the number of models required M is taken as the required number D. Then, it is determined whether the number of the first training data is equal to or greater than the required number D. If it is determined that the required number D is less than the specified value, at least one search condition record containing the past data to be analyzed is extracted from the search condition database in which the similarity of the learning profile to the data to be analyzed is greater than a predetermined similarity threshold. The past query within the extracted at least one search condition record is used as a supplementary query for extracting the learning data, and the first learning data extracted using the first query and the supplementary data extracted using the supplementary query are output. AI training data creation support system.
2. An AI learning data creation support system according to claim 1, The aforementioned AI learning data creation support system is Furthermore, it includes an output device that outputs the aforementioned learning data, The aforementioned processor, If it is determined that the number of the first training data is equal to or greater than the required number D, the first training data is extracted from the training database using the first query and output from the output device. If it is determined that the number of the first training data is less than the required number D, the first query extracts the first training data from the training database and outputs it from the output device, and the supplementary query extracts supplementary data from the training database and outputs it from the output device. AI training data creation support system.
3. An AI learning data creation support system according to claim 1, The learning profile includes domain item information that associates domain items that affect the diagnostic items of the learning profile with domain item ranges, which are ranges of values for those domain items. The processor generates a first supplementary query that includes the domain item range as a search condition. AI training data creation support system.
4. An AI learning data creation support system according to claim 1, The processor extracts from the search condition database at least one search condition record that includes the past data to be analyzed in which the similarity of the learning profile to the data to be analyzed is greater than a predetermined similarity threshold, and uses the past queries of the extracted at least one search condition record as at least one first supplement query candidate. The number of first supplement candidate data extracted from the training database by the first supplement query candidate is estimated using the training database. The product of a first supplementary data acquisition rate, which represents the number of first supplementary candidate data that can be obtained from the learning database per unit time, and a predetermined allowable time interval is calculated as the upper limit of the number of data. The first supplemental query candidate for which the number of first supplemental candidate data is less than or equal to the upper limit of the number of data is defined as the first supplemental query. AI training data creation support system.
5. An AI learning data creation support system according to claim 1, The aforementioned AI learning data creation support system is Furthermore, it includes an output device that outputs the aforementioned learning data, The aforementioned processor, The output device is used to present to the user a supplementary query to be used from the at least one supplementary query, The system accepts input for the supplementary query selected by the user. The first query extracts the first training data from the training database and outputs it using the output device. The system extracts supplemental data from the learning database using the supplemental query selected by the user and outputs it using the output device. AI training data creation support system.
6. An AI learning data creation support system according to claim 1, The aforementioned AI learning data creation support system is Furthermore, an output device that outputs the aforementioned training data, It comprises a supplement query list for registering a first supplement query to be used as a supplement query, The aforementioned processor, The target number of supplements is calculated by subtracting the number of the first learning data from the required number D. The number of first supplementary data extracted from the training database by the at least one first supplementary query is calculated using the training database. A predetermined number of the first supplementary queries, in order of priority, are registered in the supplementary query list, paired with the number of the first supplementary data items. From among the first supplemental queries not registered in the supplemental query list, a predetermined number of the first supplemental queries, in order of predetermined priority, are added to the supplemental query list along with the number of first supplemental data, and this process is repeated until the sum of the number of first supplemental data is greater than the target number of supplemental data. The first query extracts the first training data from the training database and outputs it from the output device, and the first supplementary query registered in the supplementary query list extracts supplementary data from the training database. Output from the aforementioned output device, AI training data creation support system.
7. An AI learning data creation support system according to claim 1, The aforementioned AI model is an AI model for healthcare, and the data to be analyzed includes personal information, as part of an AI learning data creation support system.
8. An AI training data creation support system provides a method for creating AI training data, which involves extracting and collecting training data for training an AI model from at least one training database. The aforementioned AI learning data creation support system is The system comprises a storage device for storing at least one program, a processor for executing the program stored in the storage device, and an input device for receiving input from a user. The aforementioned storage device is An algorithm requirement table stores the algorithm of the AI model and the number of training data required for training the AI model using the algorithm, in association with the algorithm requirement table. An analysis content required number table stores the analysis content of the AI model and the number of training data required for training the AI model based on the analysis content, in association with each other. A search condition database that stores multiple search condition records that associate previously created data to be analyzed with past queries used to extract the learning data related to the said data to be analyzed, The system includes a statistical coefficient table that stores one or more data items in association with the range of statistical values and statistical coefficients for each of those one or more data items, The processor executes the program, The system accepts input of a first query configured to extract from the learning database a learning profile consisting of item values corresponding to each of multiple data items, including the data to be analyzed by the AI model, the algorithm and analysis content of the AI model, and learning data including item values corresponding to the diagnostic items included in the data to be analyzed in the learning profile. The number of first training data extracted from the training database by the first query is calculated using the training database. With respect to the first training data, statistical values are calculated for each data item. The statistical coefficients associated with the range to which the statistic value belongs are extracted from the statistical coefficient table. Referring to the table of required algorithms and the table of required analysis content, the number of algorithms and analysis content required for training the AI model is calculated using the information on the AI model's algorithms and analysis content included in the training profile. The larger of the number of algorithms required and the number of analysis contents required is taken as the number of models required M, and the product of the largest statistical coefficient among the calculated statistical coefficients, which is the maximum statistical coefficient, and the number of models required M is taken as the required number D. Then, it is determined whether the number of the first training data is equal to or greater than the required number D. If it is determined that the required number D is less than the specified value, at least one search condition record containing the past data to be analyzed is extracted from the search condition database in which the similarity of the learning profile to the data to be analyzed is greater than a predetermined similarity threshold. The past query within the extracted at least one search condition record is used as a supplementary query for extracting the learning data, and the first learning data extracted using the first query and the supplementary data extracted using the supplementary query are output. A method for supporting the creation of AI training data.
9. An AI training data creation support program executed on the processor of an AI training data creation support system, which extracts and collects training data for training an AI model from at least one training database, The aforementioned AI learning data creation support system is The system comprises a storage device for storing at least one program, a processor for executing the program stored in the storage device, and an input device for receiving input from a user. The aforementioned storage device is An algorithm requirement table stores the algorithm of the AI model and the number of training data required for training the AI model using the algorithm, in association with the algorithm requirement table. An analysis content required number table stores the analysis content of the AI model and the number of training data required for training the AI model based on the analysis content, in association with each other. A search condition database that stores multiple search condition records that associate previously created data to be analyzed with past queries used to extract the learning data related to the said data to be analyzed, The system includes a statistical coefficient table that stores one or more data items in association with the range of statistical values and statistical coefficients for each of those one or more data items, The aforementioned processor, The system accepts input of a first query configured to extract from the learning database a learning profile consisting of item values corresponding to each of multiple data items, the data to be analyzed by the AI model, the algorithm and analysis content of the AI model, and learning data containing item values corresponding to diagnostic items included in the data to be analyzed in the learning profile. The number of first training data extracted from the training database by the first query is calculated using the training database. With respect to the first training data, calculate the statistical value for each data item. The statistical coefficients corresponding to the range to which the statistical value belongs are extracted from the statistical coefficient table. By referring to the table of required algorithms and the table of required analysis content, the number of algorithms and analysis content required for training the AI model is calculated using the information on the AI model's algorithms and analysis content included in the training profile. The larger of the number of algorithms required and the number of analysis contents required is defined as the number of models required M. The product of the largest statistical coefficient among the calculated statistical coefficients, which is the maximum statistical coefficient, and the number of models required M is defined as the required number D. Then, the system determines whether the number of the first training data is greater than or equal to the required number D. If it is determined that the required number D is less than the specified value, at least one search condition record containing the past data to be analyzed in which the similarity of the learning profile to the data to be analyzed is greater than a predetermined similarity threshold is extracted from the search condition database, the past query in the extracted at least one search condition record is used as a supplementary query for extracting the learning data, and the first learning data extracted using the first query and the supplementary data extracted using the supplementary query are output. AI training data creation support program.