Data analysis device, data analysis method, and data analysis program

The data analysis apparatus and method address the challenge of intuitively understanding data set tendencies by calculating prediction performance indices and embedding data sets into a specific space, thereby simplifying the analysis of data trends for machine learning models.

WO2025109751A1PCT designated stage expired Publication Date: 2025-05-30NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/042175
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It is difficult to intuitively grasp the tendency of data sets used for learning in machine learning models, particularly in relation to their prediction performance.

Method used

A data analysis apparatus and method that calculates an index of prediction performance for a machine learning model using a data set with explanatory and objective variables, embeds this data set into a predetermined space using the calculated index, and outputs information representing the embedded space.

Benefits of technology

Facilitates easy grasping of the tendency of data sets used for learning machine learning models, enhancing the understanding and utilization of data trends for future operations and model updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023042175_30052025_PF_FP_ABST
    Figure JP2023042175_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A data analysis device according to the present invention can support a user in decision making by comprising: a calculation unit that calculates an index for prediction performance of a machine learning model by using a data set including an explanatory variable and an objective variable and a prediction result obtained by inputting the explanatory variable included in the data set to the machine learning model; an embedding unit that embeds the data set in a predetermined space by using the index calculated by the calculation unit; and an output unit that outputs information representing the space in which the data set has been embedded.
Need to check novelty before this filing date? Find Prior Art

Description

Data analysis device, data analysis method, and data analysis program

[0001] The present disclosure relates to a data analysis device, a data analysis method, and a data analysis program.

[0002] Techniques for evaluating the performance of machine learning models are known. One example of a technique for evaluating the performance of a machine learning model is the technique described in Patent Literature 1.

[0003] Japan Special Table No. 2021-532488

[0004] In the actual operation of machine learning models, large amounts of data sets are used, and the models are updated periodically. For effective operation, it is important to understand the trends in the data sets used for learning and use this information for future operation and model updates. However, there is a problem in that it is difficult to understand the trends in the data sets. In particular, it is difficult to intuitively understand the trends in the data sets related to the predictive performance of the machine learning model. This is also true for the technology described in Patent Document 1.

[0005] The present disclosure has been made in consideration of the above-mentioned problems, and one exemplary purpose thereof is to provide a technology that makes it easier to understand trends in datasets used to train machine learning models, trends that are related to the predictive performance of the machine learning models.

[0006] A data analysis apparatus according to an exemplary aspect of the present disclosure includes a calculation means for calculating an index of the predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model, an embedding means for embedding the dataset into a predetermined space using the index calculated by the calculation means, and an output means for outputting information representing the space into which the dataset is embedded.

[0007] A data analysis method according to an exemplary aspect of the present disclosure includes: a calculation process in which at least one processor calculates an index of predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model; an embedding process in which the at least one processor embeds the dataset into a predetermined space using the index calculated in the calculation process; and an output process in which the at least one processor outputs information representing the space into which the dataset is embedded.

[0008] A data analysis program according to an exemplary aspect of the present disclosure is a program that causes a computer to function as a data analysis device, and causes the computer to function as: a calculation means that calculates an index of the predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model; an embedding means that embeds the dataset into a predetermined space using the index calculated by the calculation means; and an output means that outputs information representing the space into which the dataset is embedded.

[0009] According to one exemplary aspect of the present disclosure, an exemplary effect is achieved in that a technology can be provided that makes it easier to understand trends in a dataset used to train a machine learning model, trends that are related to the predictive performance of the machine learning model.

[0010] Fig. 1 is a block diagram showing the configuration of a data analysis device according to the present disclosure; Fig. 2 is a flow diagram showing the flow of a data analysis method according to the present disclosure; Fig. 3 is a block diagram showing the configuration of an information processing device according to the present disclosure; Fig. 4 is a diagram showing an example of a performance index according to the present disclosure; Fig. 5 is a diagram showing an example of a distribution map according to the present disclosure; Fig. 6 is a block diagram showing the configuration of a computer functioning as each device according to the present disclosure.

[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0012] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0013] (Configuration of Data Analysis Apparatus) The configuration of the data analysis apparatus 1 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the data analysis apparatus 1. As shown in Fig. 1, the data analysis apparatus 1 includes a calculation unit 11, an embedding unit 12, and an output unit 13.

[0014] The calculation unit 11 calculates an index of the predictive performance of the machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model. The embedding unit 12 embeds the dataset into a predetermined space using the index calculated by the calculation unit 11. The output unit 13 outputs information representing the space into which the dataset is embedded.

[0015] (Effects of the Data Analysis Device) As described above, the data analysis device 1 employs a configuration including a calculation unit 11 that calculates an index of the predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model, an embedding unit 12 that embeds the dataset into a predetermined space using the index calculated by the calculation unit 11, and an output unit 13 that outputs information representing the space into which the dataset is embedded. Therefore, the data analysis device 1 provides the effect of making it easier to grasp trends in the dataset used to train the machine learning model that are related to the predictive performance of the machine learning model.

[0016] (Flow of Data Analysis Method) The flow of data analysis method S1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of data analysis method S1. As shown in Fig. 2, data analysis method S1 includes calculation processing S11, embedding processing S12, and output processing S13.

[0017] In a calculation process S11, at least one processor calculates an index of the predictive performance of the machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model. In an embedding process S12, at least one processor embeds the dataset into a predetermined space using the index calculated in the calculation process S11. In an output process S13, at least one processor outputs information representing the space into which the dataset is embedded.

[0018] (Effects of Data Analysis Method) As described above, the data analysis method S1 employs a configuration including: a calculation process in which at least one processor calculates an index of the predictive performance of the machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model; an embedding process in which the at least one processor embeds the dataset into a predetermined space using the index calculated in the calculation process; and an output process in which the at least one processor outputs information representing the space into which the dataset is embedded. Thus, the data analysis method S1 has the effect of making it easier to grasp trends in the dataset used to train the machine learning model, trends related to the predictive performance of the machine learning model.

[0019] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0020] (Overview of Information Processing Device) An information processing device 1A according to the present disclosure is a device for analyzing a dataset. The dataset includes an explanatory variable and a response variable. Examples of datasets include, but are not limited to, datasets in the medical and healthcare fields. For example, if the dataset to be analyzed is a dataset related to predicting hospital bed occupancy rates, the dataset includes, as an example, an explanatory variable indicating at least one of the month, day, season, weather, temperature, humidity, and whether or not there is an epidemic of infectious disease. Furthermore, as an example, the dataset includes a response variable indicating the hospital bed occupancy rate.

[0021] The information processing device 1A also analyzes a dataset using an evaluation model, which is generated and updated by machine learning and is capable of predicting a value corresponding to a target variable of a dataset when an explanatory variable of the dataset to be analyzed is input.

[0022] The dataset may be, for example, a dataset used to generate or update a predictive model generated by machine learning. For example, when a predictive model is updated periodically, the information processing device 1A analyzes changes in trends in the dataset used to update the predictive model and outputs the analysis results as information to be used in the operation of the predictive model. By checking the output information, a user of the information processing device 1A can understand changes in trends in the dataset and use the information to operate the predictive model. Note that a predictive model used in actual operation may be used as the evaluation model.

[0023] Furthermore, the dataset to be analyzed is not limited to the above-mentioned examples. The dataset according to the present disclosure may be, for example, a dataset indicating purchase histories collected at each of a plurality of stores.

[0024] (Configuration of information processing device) The configuration of the information processing device 1A will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of the information processing device 1A. The information processing device 1A is an example of a data analysis device according to the present disclosure. The information processing device 1A includes a control unit 10A, a storage unit 20A, a communication unit 30A, an input unit 40A, and an output unit 50A.

[0025] (Communication Unit) The communication unit 30A communicates with devices external to the information processing device 1A via a communication line. While the specific configuration of the communication line does not limit the present exemplary embodiment, examples of the communication line include a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination thereof. The communication unit 30A transmits data supplied from the control unit 10A to other devices, and supplies data received from other devices to the control unit 10A.

[0026] (Input Unit) The input unit 40A is configured to receive input to the information processing device 1A, and includes, for example, input devices such as a keyboard, a mouse, a touch panel, a camera, a microphone, etc. The input unit 40A may also be configured to receive data from the input devices via an interface such as a USB (Universal Serial Bus).

[0027] (Output Unit) The output unit 50A is a component for performing output from the information processing device 1A, and includes, for example, output devices such as a display, a printer, a touch panel, a speaker, etc. The output unit 50A may be configured to include, for example, an interface such as a USB, and to output data to the output device via the interface.

[0028] (Storage Unit) The storage unit 20A stores various types of information referenced by the control unit 10A. An example of such information is a data set D i (i = 1, 2, ..., n; n is an integer equal to or greater than 1), evaluation model h k (i=1, 2, ..., l; l is an integer equal to or greater than 1), and a performance index P j (j=1, 2, . . . m; m is an integer of 1 or more).

[0029] (Dataset) Dataset D i As described above, the data sets Di include a plurality of pairs of explanatory variables and response variables. The plurality of data sets Di may have a time sequence relationship with each other.

[0030] (Evaluation model) Evaluation model h k is an example of a machine learning model according to the present disclosure. Here, the evaluation model h k is stored in the storage unit 20A, the evaluation model h k This means that the parameters that define the above are stored in the storage unit 20A.

[0031] (Performance index) Performance index P j is the evaluation model h calculated by the index vector calculation unit 12A described later. k For example, the evaluation model h k is a classification model, the performance index Pj For example, accuracy, precision, recall, F-measure, cross-entropy loss, 01 loss, etc. may be used as the evaluation model h k is a regression model, the performance index P j For example, the coefficient of determination, the mean square error, the mean absolute error, etc. may be used as the performance index P j For example, the parameter includes at least one of accuracy, precision, recall, F-measure, cross-entropy loss, 01 loss, coefficient of determination, mean squared error, and mean absolute error.

[0032] (Control Unit) The control unit 10A includes a data set acquisition unit 11A, an index vector calculation unit 12A, an embedding vector calculation unit 13A, and an output control unit 14A. The index vector calculation unit 12A is an example of a calculation means according to the present disclosure. The embedding vector calculation unit 13A is an example of an embedding means according to the present disclosure. The output control unit 14A is an example of an output means according to the present disclosure.

[0033] (Data Set Acquisition Unit) The data set acquisition unit 11A acquires data set D i As an example, the data set acquisition unit 11A acquires the data set D from a storage destination (which may be a storage device within the information processing device 1A or a storage device outside the information processing device 1A) designated by the user of the information processing device 1A. i By reading out the data set D i The data set acquisition unit 11A also acquires the data set D from another device via the communication unit 30A. i By receiving the data set D i The data set acquisition unit 11A may acquire the data set D input to the input unit 40A. i may be obtained.

[0034] (Index Vector Calculation Unit) The index vector calculation unit 12A calculates the data set D i and dataset D i The explanatory variables included in the evaluation model h k The prediction results obtained by inputting k The performance index P is an index of the predictive performance of jHere, the evaluation model h k When there are a plurality of data sets D i The explanatory variables included in multiple evaluation models h k Using the multiple prediction results obtained by inputting into k For each of the performance indexes P j At this time, the index vector calculation unit 12A calculates the data set D i and dataset D i Evaluation model h generated by machine learning using k Dataset D i The prediction results obtained by inputting the explanatory variables included in j In addition, the performance index P j The evaluation model h used to calculate k is one or more datasets D i The model may be generated using the method described above, or may be generated by other methods.

[0035] FIG. 4 shows the performance index P calculated by the index vector calculation unit 12A. j In the example of FIG. 4, the index vector calculation unit 12A calculates the data set D 1 Regarding (i) the evaluation model h 1 Dataset D 1 (ii) the accuracy of the evaluation model h 2 Dataset D 1 (iii) the accuracy of the evaluation model h 3 Dataset D 1 (iv) the accuracy of the evaluation model h 1 Dataset D 1 F-measure for (v) evaluation model h 2 Dataset D 1 F-measure for (vi) evaluation model h 3 Dataset D 1 The index vector calculation unit 12A calculates six performance indices, namely, the F-measure for the data set D 2 , D 3 Similarly, six performance indices are calculated for each of the above.

[0036] In the example of FIG. 4, the index vector calculation unit 12A calculates one evaluation model h k Regarding multiple performance indicators P j However, the index vector calculation unit 12A calculates one evaluation model h k One performance index P j In addition, the performance index P calculated by the index vector calculation unit 12A may be calculated as j is the evaluation model h j The index vector calculation unit 12A may vary depending on two or more evaluation models h j Common performance index P j may be calculated.

[0037] The index vector calculation unit 12A calculates the performance index P j Using the data set D i The feature vector v i Generate a feature vector v i is expressed by the following formula, for example: i = [P 1 (h 1 , D i ), ..., P m (h k , D i ) ] where P j (h k , D i ) is the evaluation model h k Dataset D i Performance index P for j In the example of FIG. 4, the feature vector v i is a vector with six components.

[0038] (Embedding Vector Calculation Unit) The embedding vector calculation unit 13A embeds a plurality of data sets into a predetermined space using the indices calculated by the index vector calculation unit 12A. For example, the embedding vector calculation unit 13A generates a feature vector v representing the indices calculated by the index vector calculation unit 12A. i is an embedding vector e iThe embedding method includes, but is not limited to, a predetermined feature vector conversion, PCA, T-SNE, U-MAP, AutoEncoder, etc. In the example of FIG. 4, the embedding vector calculation unit 13A converts the feature vector v i is dimensionally compressed to obtain a two-dimensional embedding vector e i = [0.72, 0.68].

[0039] (Output Control Unit) The output control unit 14A outputs information representing the space in which the dataset is embedded. As an example, the output control unit 14A outputs the information representing the space to a display. More specifically, the output control unit 14A displays, for example, a distribution map representing the space in which the dataset is embedded on a display. Furthermore, as an example, the output control unit 14A may output the information by writing the information to a storage destination (which may be a storage device within the information processing device 1A or a storage device external to the information processing device 1A) designated by the user of the information processing device 1A. Furthermore, the output control unit 14A may output the information by transmitting the information via the communication unit 30A, or may output the information to an output device other than a display.

[0040] An example of the information representing the space is an image representing the space, and more specifically, a distribution diagram showing the space. Fig. 5 is a diagram showing an example of a distribution diagram that the output control unit 14A displays on the display. In the distribution diagrams 61 and 62 in Fig. 5, a plurality of data sets D i are embedded in a two-dimensional space. In the distribution maps 61 and 62, each of the multiple dots represents a data set D i It corresponds to multiple datasets D i are connected by lines showing the time series.

[0041] The user of the information processing device 1A can read the embedded data set D i For example, the spatial location of dataset D i If the figure that appears by the lines connecting the data sets D is a periodic figure, it can be seen that the correspondence between the objective variable and the explanatory variable changes periodically. iIf the figure formed by the lines connecting the data points does not have periodicity, it indicates that the data trend is changing and is not reproducible. Furthermore, as shown in Figure 5, users can intuitively grasp the speed of data change from the intervals between the embedding points of the data set. This allows users to make operational decisions.

[0042] The output control unit 14A also stores a plurality of data sets D i A first data set D selected from α A second data set D that is closer to the second data set D in the space than the other data sets. β It may also output information recommending the following.

[0043] (Effects of the Information Processing Apparatus) As described above, according to the information processing apparatus 1A, a plurality of data sets D i The evaluation model h k The data is embedded in a low-dimensional space based on the predictive performance of the model and presented to the user. By checking the presented information, the user can easily grasp the trend of the data set and evaluate the analysis results. k In particular, the information processing device 1A can be used to generate a plurality of evaluation models h k The performance index P obtained using j Using the dataset D i By embedding the feature vectors, it is possible to reduce the influence of noise and redundant features that do not affect prediction, and to achieve embedding and visualization that are useful for operations.

[0044] In addition, in actual operation, it is particularly important whether the prediction performance deteriorates (the correspondence between the objective variable and the explanatory variable). i By displaying this, it is possible to effectively visualize changes in the correspondence between the objective variable and the explanatory variables.

[0045] In the information processing device 1A, the embedding vector calculation unit 13A is configured to embed a plurality of data sets in the space. A user of the information processing device 1A can, for example, select a plurality of data sets D embedded in the space.i By checking multiple data sets D i In this way, according to the information processing device 1A, it is possible to compare the trends of the multiple data sets D i This has the effect of making it easier to grasp the trends.

[0046] In addition, in the information processing device 1A, the plurality of data sets D i In this way, the information processing device 1A employs a configuration in which the data sets D i This has the effect of making it easier to grasp changes in trends.

[0047] In the information processing device 1A, the index vector calculation unit 12A calculates one data set D i The explanatory variables included in multiple evaluation models h k Using the multiple prediction results obtained by inputting into k A configuration is adopted in which an index of prediction performance is calculated for each of the data sets D i Multiple evaluation models h k By using the above method for evaluation, it is possible to obtain the effect of analyzing the trends of the data set with greater precision.

[0048] In the information processing device 1A, the index vector calculation unit 12A calculates a data set D including explanatory variables and a target variable. i and dataset D i Evaluation model h generated by machine learning using k Dataset D i The prediction results obtained by inputting the explanatory variables included in j The evaluation model h is calculated as follows. k The dataset D used for learning i The performance index of the evaluation model h k By calculating using i This has the effect of enabling more accurate analysis of trends.

[0049] In addition, in the information processing device 1A, the performance index P jThe performance index P includes at least one of accuracy, precision, recall, F-measure, cross-entropy loss, 01 loss, coefficient of determination, mean square error, and mean absolute error. j By embedding i This has the effect of enabling more accurate analysis of trends.

[0050] In the information processing device 1A, the embedding vector calculation unit 13A calculates the feature vector v representing the index calculated by the index vector calculation unit 12A. i The feature vector v i The embedding vector e of a dimension lower than the dimension of i The feature vector v i is expressed as a lower dimensional embedding vector e i By converting the dataset D i In particular, the effect of making it easier to intuitively grasp the trend of data set D i By embedding into two dimensions, it becomes easier to draw the graph of the embedding space.

[0051] In the information processing device 1A, the information representing the space is a distribution map showing the space, and in the distribution map, a plurality of data sets D i are connected by line segments that represent time series. i are connected by a line segment representing the time series, and the data set D i This has the effect of making it easier to grasp the time-series changes in the trend.

[0052] In the information processing device 1A, the output control unit 14A outputs the data set D i The configuration is adopted in which a distribution map representing the space in which the data set D is embedded is displayed on a display. i This has the effect of making the trends easier to understand by visualizing them.

[0053] In the information processing device 1A, the output control unit 14A outputs a plurality of data sets D i A first data set D selected from α A second data set D that is closer to the second data set D in the space than the other data sets. β In this way, the information processing device 1A outputs information recommending the selected first data set D α It is possible to notify the user of data sets that have similar trends.

[0054] [Example of implementation by software] Some or all of the functions of the data analysis device 1 and the information processing device 1A (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.

[0055] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 6. Figure 6 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.

[0056] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.

[0057] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0058] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0059] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0060] [Appendix 1] The present disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims. [Appendix A] The present disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.

[0061] (Appendix A1) A data analysis device comprising: a calculation means that calculates an index of predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model; an embedding means that embeds the dataset into a predetermined space using the index calculated by the calculation means; and an output means that outputs information representing the space into which the dataset is embedded.

[0062] (Appendix A2) The data analysis apparatus according to Appendix A1, wherein the embedding means embeds a plurality of data sets into the space.

[0063] (Supplementary Note A3) The data analysis device according to Supplementary Note A2, wherein the plurality of data sets have a time sequence relationship with each other.

[0064] (Appendix A4) The data analysis device according to any one of Appendices A1 to A3, wherein the calculation means calculates an index of predictive performance for each of a plurality of machine learning models using a plurality of prediction results obtained by inputting explanatory variables included in one dataset into the plurality of machine learning models.

[0065] (Appendix A5) The data analysis device according to any one of Appendices A1 to A4, wherein the calculation means calculates the index using a dataset including explanatory variables and a target variable, and a prediction result obtained by inputting the explanatory variables included in the dataset into a machine learning model generated by machine learning using the dataset.

[0066] (Appendix A6) The data analysis device according to any one of Appendices A1 to A5, wherein the indicators include at least one of accuracy, precision, recall, F-measure, cross-entropy loss, 01 loss, coefficient of determination, mean squared error, and mean absolute error.

[0067] (Supplementary Note A7) The data analysis apparatus according to any one of Supplementary Notes A1 to A6, wherein the embedding means converts a feature vector representing the index calculated by the calculation means into a vector having a dimension lower than that of the feature vector.

[0068] (Supplementary Note A8) The data analysis device according to Supplementary Note A7, wherein the information representing the space is a distribution map representing the space, and in the distribution map, the plurality of data sets are connected by line segments representing time series.

[0069] (Appendix A9) The data analysis device according to any one of appendices A1 to A8, wherein the output means displays, on a display, a distribution map representing a space in which the dataset is embedded.

[0070] (Supplementary Note A10) The data analysis device according to Supplementary Note A2, wherein the output means outputs information recommending a second dataset that is closer in distance in the space to the first dataset selected from the plurality of datasets than other datasets.

[0071] [Appendix B] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0072] (Appendix B1) A data analysis method comprising: a calculation process in which at least one processor calculates an index of predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model; an embedding process in which the at least one processor embeds the dataset into a predetermined space using the index calculated in the calculation process; and an output process in which the at least one processor outputs information representing the space in which the dataset is embedded.

[0073] (Supplementary Note B2) The data analysis method according to Supplementary Note B1, wherein in the embedding process, the at least one processor embeds a plurality of data sets into the space.

[0074] (Supplementary Note B3) The data analysis method according to Supplementary Note B2, wherein the plurality of data sets have a time sequence relationship with each other.

[0075] (Appendix B4) The data analysis method according to any one of Appendices B1 to B3, wherein in the calculation process, the at least one processor calculates an index of predictive performance for each of a plurality of machine learning models using a plurality of prediction results obtained by inputting explanatory variables included in one dataset into the plurality of machine learning models.

[0076] (Appendix B5) The data analysis method according to any one of Appendices B1 to B4, wherein in the calculation process, the at least one processor calculates the index using a dataset including explanatory variables and a target variable, and a prediction result obtained by inputting the explanatory variables included in the dataset into a machine learning model generated by machine learning using the dataset.

[0077] (Appendix B6) The data analysis method according to any one of Appendices B1 to B5, wherein the index includes at least one of accuracy, precision, recall, F-measure, cross-entropy loss, 01 loss, coefficient of determination, mean squared error, and mean absolute error.

[0078] (Supplementary Note B7) The data analysis method according to any one of Supplementary Notes B1 to B6, wherein in the embedding process, the at least one processor converts a feature vector representing the index calculated in the calculation process into a vector having a dimension lower than that of the feature vector.

[0079] (Appendix B8) The data analysis method according to Appendix B7, wherein the information representing the space is a distribution map showing the space, and in the distribution map, the plurality of data sets are connected by line segments showing time series.

[0080] (Supplementary Note B9) The data analysis method according to any one of Supplementary Notes B1 to B8, wherein in the output process, the at least one processor displays, on a display, a distribution map representing a space in which the dataset is embedded.

[0081] (Appendix B10) The data analysis method according to Appendix B2, wherein in the output process, the at least one processor outputs information recommending a second dataset that is closer in distance in the space to a first dataset selected from the plurality of datasets than other datasets.

[0082] [Appendix C] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0083] (Appendix C1) A program that causes a computer to function as a data analysis device, the data analysis program causing the computer to function as: a calculation means that calculates an index of the predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model; an embedding means that embeds the dataset into a predetermined space using the index calculated by the calculation means; and an output means that outputs information representing the space into which the dataset is embedded.

[0084] (Supplementary Note C2) The data analysis apparatus according to Supplementary Note C1, wherein the embedding means embeds a plurality of data sets into the space.

[0085] (Supplementary Note C3) The data analysis device according to Supplementary Note C2, wherein the plurality of data sets have a time sequence relationship with each other.

[0086] (Appendix C4) The data analysis program according to any one of Appendices C1 to C3, wherein the calculation means calculates an index of predictive performance for each of a plurality of machine learning models using a plurality of prediction results obtained by inputting explanatory variables included in one dataset into the plurality of machine learning models.

[0087] (Appendix C5) The data analysis program according to any one of Appendices C1 to C4, wherein the calculation means calculates the index using a dataset including explanatory variables and a target variable, and a prediction result obtained by inputting the explanatory variables included in the dataset into a machine learning model generated by machine learning using the dataset.

[0088] (Appendix C6) The data analysis program according to any one of Appendices C1 to C5, wherein the indicators include at least one of accuracy, precision, recall, F-measure, cross-entropy loss, 01 loss, coefficient of determination, mean squared error, and mean absolute error.

[0089] (Supplementary Note C7) The data analysis program according to any one of Supplementary Note C1 to C6, wherein the embedding means converts a feature vector representing the index calculated by the calculation means into a vector having a dimension lower than that of the feature vector.

[0090] (Appendix C8) The data analysis program according to Appendix C7, wherein the information representing the space is a distribution map showing the space, and in the distribution map, the plurality of data sets are connected by line segments showing time series.

[0091] (Appendix C9) The data analysis program according to any one of appendices C1 to C8, wherein the output means displays, on a display, a distribution map representing a space in which the dataset is embedded.

[0092] (Supplementary Note C10) The data analysis program according to Supplementary Note C2, wherein the output means outputs information recommending a second dataset that is closer in distance in the space to the first dataset selected from the plurality of datasets than other datasets.

[0093] [Appendix D] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0094] (Appendix D1) A data analysis device including at least one processor, the at least one processor executing: a calculation process that calculates an index of predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model; an embedding process that embeds the dataset into a predetermined space using the index calculated in the calculation process; and an output process that outputs information representing the space into which the dataset is embedded.

[0095] The data analysis device may further include a memory, and the memory may store a program for causing the at least one processor to execute each of the processes.

[0096] (Appendix D2) The data analysis apparatus according to Appendix D1, wherein in the embedding process, the at least one processor embeds a plurality of data sets into the space.

[0097] (Appendix D3) The data analysis device according to appendix D2, wherein the plurality of data sets have a time sequence relationship with each other.

[0098] (Appendix D4) The data analysis device according to any one of Appendices D1 to D3, wherein in the calculation process, the at least one processor calculates an index of predictive performance for each of a plurality of machine learning models using a plurality of prediction results obtained by inputting explanatory variables included in one dataset into the plurality of machine learning models.

[0099] (Appendix D5) The data analysis device according to any one of Appendices D1 to D4, wherein in the calculation process, the at least one processor calculates the index using a dataset including explanatory variables and a target variable, and a prediction result obtained by inputting the explanatory variables included in the dataset into a machine learning model generated by machine learning using the dataset.

[0100] (Appendix D6) The data analysis device according to any one of Appendices D1 to D5, wherein the indicators include at least one of accuracy, precision, recall, F-measure, cross-entropy loss, 01 loss, coefficient of determination, mean squared error, and mean absolute error.

[0101] (Appendix D7) The data analysis device according to any one of Appendices D1 to D6, wherein in the embedding process, the at least one processor converts a feature vector representing an index calculated in the calculation process into a vector having a dimension lower than that of the feature vector.

[0102] (Appendix D8) The data analysis device according to appendix D7, wherein the information representing the space is a distribution map representing the space, and in the distribution map, the plurality of data sets are connected by line segments representing time series.

[0103] (Appendix D9) The data analysis apparatus according to any one of appendices D1 to D8, wherein in the output process, the at least one processor displays, on a display, a distribution map representing a space in which the dataset is embedded.

[0104] (Appendix D10) The data analysis device according to Appendix D2, wherein in the output process, the at least one processor outputs information recommending a second dataset that is closer in distance in the space to a first dataset selected from the plurality of datasets than other datasets.

[0105] [Appendix E] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0106] (Appendix E1) A non-transitory recording medium having recorded thereon a program that causes a computer to function as a data analysis device, the data analysis program causing the computer to execute the following: a calculation process that calculates an index of the predictive performance of a machine learning model using a dataset including explanatory variables and a target variable and a prediction result obtained by inputting the explanatory variables included in the dataset into the machine learning model; an embedding process that embeds the dataset into a predetermined space using the index calculated in the calculation process; and an output process that outputs information representing the space into which the dataset is embedded.

[0107] REFERENCE SIGNS LIST 1 Data analysis device 1A Information processing device 11 Calculation unit 11A Data set acquisition unit 12 Embedding unit 12A Index vector calculation unit 13, 50A Output unit 13A Embedding vector calculation unit 14A Output control unit

Claims

1. A data analysis apparatus comprising: a calculation means for calculating an index of prediction performance of a machine learning model using a data set including an explanatory variable and an objective variable, and a prediction result obtained by inputting the explanatory variable included in the data set into the machine learning model; an embedding means for embedding the data set into a predetermined space using the index calculated by the calculation means; and an output means for outputting information representing the space into which the data set is embedded.

2. The data analysis apparatus according to claim 1, wherein the embedding means embeds a plurality of data sets into the space.

3. The data analysis apparatus according to claim 2, wherein the plurality of data sets have a temporal relationship with each other.

4. The data analysis apparatus according to any one of claims 1 to 3, wherein the calculation means calculates an index of prediction performance for each of the plurality of machine learning models using a plurality of prediction results obtained by inputting the explanatory variable included in one data set into the plurality of machine learning models.

5. The data analysis apparatus according to any one of claims 1 to 4, wherein the calculation means calculates the index using a data set including an explanatory variable and an objective variable, and a prediction result obtained by inputting the explanatory variable included in the data set into a machine learning model generated by machine learning using the data set.

6. The data analysis apparatus according to any one of claims 1 to 5, wherein the index includes at least one of accuracy, precision, recall, F-value, cross-entropy loss, 01 loss, coefficient of determination, mean squared error, and mean absolute error.

7. The data analysis apparatus according to any one of claims 1 to 6, wherein the embedding means converts a feature vector representing the index calculated by the calculation means into a vector having a dimension lower than the dimension of the feature vector.

8. The information representing the space is a distribution map showing the space, and in the distribution map, the plurality of data sets are connected by line segments indicating a time series. The data analysis apparatus according to claim 7.

9. The data analysis apparatus according to any one of claims 1 to 8, wherein the output means displays a distribution map representing the space into which the data set is embedded on a display.

10. The data analysis apparatus according to claim 2, wherein the output means outputs information for recommending a second data set whose distance in the space from the first data set selected from the plurality of data sets is closer than that of other data sets.

11. A data analysis method comprising: a calculation process in which at least one processor calculates an index of prediction performance of a machine learning model using a data set including an explanatory variable and an objective variable, and a prediction result obtained by inputting the explanatory variable included in the data set into the machine learning model; an embedding process in which the at least one processor embeds the data set in a predetermined space using the index calculated in the calculation process; and an output process in which the at least one processor outputs information representing the space in which the data set is embedded.

12. A program for causing a computer to function as a data analysis apparatus, the program causing the computer to function as: a calculation means for calculating an index of prediction performance of a machine learning model using a data set including an explanatory variable and an objective variable, and a prediction result obtained by inputting the explanatory variable included in the data set into the machine learning model; an embedding means for embedding the data set in a predetermined space using the index calculated by the calculation means; and an output means for outputting information representing the space in which the data set is embedded; a data analysis program.

Citation Information

Patent Citations

  • Learning method by means of neural network

    JP1993189398A

  • Accuracy-estimating-model generating system and accuracy estimating system

    WO2016152053A1

  • Model analysis device, model analysis method, and recording medium

    WO2023181322A1