Analysis device

The analytical device addresses data loss issues in materials development by visually tracking data retention and missing data, facilitating efficient experimental planning.

JP2026014675APending Publication Date: 2026-01-29TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116041
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

In materials development using materials informatics, data loss during the cycle of experimental data addition and analysis makes it difficult to plan efficient experiments, and the impact of missing data on the experimental process is unclear.

Method used

An analytical device that generates a table indicating data retention and missing locations within a dataset, allowing users to visualize and understand changes in data retention states through different display modes for retained and missing data.

Benefits of technology

Enables users to easily track and understand data retention changes, guiding efficient experimental planning by highlighting important data locations for addition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014675000001_ABST
    Figure 2026014675000001_ABST
Patent Text Reader

Abstract

To enable a user to easily grasp a change in a data holding state.SOLUTION: The analysis device includes a storage unit that stores a data set including data representing a composition, a structure, or performance of a material, a generation unit that generates a table indicating a holding state of data in the data set stored in the storage unit, and an output unit that outputs the table generated by the generation unit. The storage unit stores a data set in association with a cycle indicating a process in which a state of the data set transitions due to addition of data. The generation unit generates a table for each cycle in such a manner that, among storage locations of data constituting the table, a loss location that is a storage location where data is lost and a holding location that is a storage location where data is held are displayed in different display modes. The output unit outputs the table generated by the generation unit for each cycle. The analysis device may include a prediction unit configured to predict an objective variable from an explanatory variable using a machine-learned regression model, and the data set may include the objective variable and a data item designated as the explanatory variable.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an analysis device. [Background technology]

[0002] Patent Document 1 discloses a support device for supporting optimization analysis in accordance with experimental design, which includes a selection means for selecting, from a list of multiple candidates displayed on a display screen, an error factor or a signal factor to be assigned to a level value table used to prepare data necessary for optimization analysis, and a display control means for identifying a candidate that has already been selected as a control factor from the candidate list and changing the display attribute so that it can be distinguished from other candidates. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2002-259464 Summary of the Invention [Problem to be solved by the invention]

[0004] In materials development using materials informatics, it is common to increase the amount of data contained in a dataset by repeating the cycle of adding experimental data, analyzing the data, planning the next experimental point, and then adding more experimental data. When repeating this cycle, data loss may occur, meaning that some data has not been acquired. Unintentional data loss makes it difficult to plan efficient experiments and, ultimately, to carry out efficient materials development, so it is important to understand the retention status of the data contained in the dataset.

[0005] Even if an experiment is performed using the experimental design method described in Patent Document 1, missing data may occur, and it is difficult to grasp how the missing data is compensated for and how much missing data remains.

[0006] The present invention has been made in view of the above circumstances, and has as its object to allow a user to easily grasp changes in the state of data retention. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems, the analytical device of the present invention is an analytical device that analyzes data representing the composition, structure, or performance of a material, and comprises: a memory unit that stores a dataset including the data; a generation unit that generates a table indicating the retention state of the data in the dataset stored in the memory unit; and an output unit that outputs the table generated by the generation unit, wherein the memory unit stores the dataset in correspondence with a cycle that indicates a process in which the state of the dataset changes due to the addition of the data; the generation unit generates the table for each cycle by displaying, in different display modes, missing locations, which are storage locations where the data is missing, and retained locations, which are storage locations where the data is retained, among the storage locations of the data that constitute the table; and the output unit outputs the table generated by the generation unit for each cycle. [Effects of the Invention]

[0008] According to the present invention, it is possible to allow the user to easily understand changes in the state of data retention. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram showing the configuration of an information processing system including an analysis device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating a storage device and a generating unit shown in FIG. [Figure 3] FIG. 2 is a diagram for explaining a generating unit shown in FIG. 1; [Figure 4] FIG. 2 is a diagram illustrating a specific portion shown in FIG. 1. [Figure 5] FIG. 2 is a diagram for explaining a calculation unit and a determination unit shown in FIG. [Figure 6]1. FIG. 4 is a diagram illustrating another function of the generation unit shown in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Components with the same reference numerals in each embodiment have similar components in each embodiment unless otherwise specified, and description thereof will be omitted.

[0011] Fig. 1 is a diagram showing the configuration of an information processing system 1 including an analysis device 4 of this embodiment. Fig. 2 is a diagram explaining a storage device 50 and a generation unit 42 shown in Fig. 1. Fig. 3 is a diagram explaining the generation unit 42 shown in Fig. 1.

[0012] The information processing system 1 is a system that analyzes data input by a user using an analysis device 4 and presents the analysis results to the user. The information processing system 1 includes a user terminal 2 and a server 3 that includes the analysis device 4. The user terminal 2 and the server 3 are connected to each other via a network N so that they can communicate with each other.

[0013] The user terminal 2 includes a processing device 21 and a display device 22. The processing device 21 includes a processor and a memory, and the processor executes a program to realize the functions of the user terminal 2. The display device 22 includes a display and displays the processing results of the processing device 21.

[0014] The user terminal 2 accepts data input by the user and transmits it to the analysis device 4 of the server 3. The user terminal 2 receives the analysis results of the analysis device 4 from the server 3, displays them, and presents them to the user.

[0015] The analysis device 4 analyzes data input by a user. The data input to the analysis device 4 is, for example, data representing the composition or structure of a material obtained from various measurement data related to material analysis, such as XRD (X-ray diffraction) measurement data or SEM (Scanning Electron Microscope) measurement data. The material is, for example, a metal material, a resin material, or a coating material used in a vehicle. As data analysis, the analysis device 4 predicts, for example, the performance of a material from the data representing the composition or structure of the material. The performance of the material may be, for example, battery performance, magnetic performance, rigidity, thermoplasticity, tensile performance, or mechanical durability. Furthermore, as data analysis, the analysis device 4 may predict candidate compositions or structures of the material from the data representing the performance of the material.

[0016] The analysis device 4 includes a processing device 40 and a storage device 50. The storage device 50 includes a storage for storing data received from the user terminal 2 and analysis results of the analysis device 4. For example, the storage device 50 stores data representing the composition or structure of a material as data received from the user terminal 2, and stores data representing the performance of the material as analysis results of the analysis device 4. That is, the storage device 50 stores a data set 60 including data representing the composition, structure, or performance of a material, as shown in the upper diagram of FIG. 2 .

[0017] In the example shown in FIG. 2, the dataset 60 is a collection of multiple sample data 61. Each sample data 61 consists of multiple data divided into multiple data items representing the composition, structure, or performance of a material, and a sample name (or sample ID). The dataset 60 includes data items designated by the user as the objective variables and explanatory variables of the prediction unit 41, which will be described later. In the example shown in FIG. 2, the explanatory variables are data items representing the composition of the material, and the objective variables are data items representing the performance of the material, but the analysis device 4 of this embodiment is not limited to this. The explanatory variables may be data items representing the structure of the material, and the objective variables may be data items representing the performance of the material. The explanatory variables may be data items representing the performance of the material, and the objective variables may be data items representing candidates for the composition or structure of the material.

[0018] The storage device 50 stores the dataset 60 in association with a cycle that indicates a process in which the state of the dataset 60 changes due to the addition of data. That is, a cycle is a process from a first point in time when the dataset 60 is in a first state to a second point in time when data is added to the dataset 60 and the dataset 60 changes to a second state.

[0019] In the context of materials informatics, a cycle consists of at least one of the following processes: (1) Integration of additional data into existing data (60 existing datasets) (2) Preprocessing (the process of converting data into a format suitable for modeling) (3) Feature extraction (the process of extracting descriptors from data that can be used for modeling using domain knowledge or statistical methods) (4) Model construction (the process of using the extracted descriptors to build and train a machine learning model to predict the properties or performance of the material). (5) Model evaluation (the process of evaluating the performance of the constructed machine learning model using validation data. In this process, the model's predictive accuracy or generalization ability is evaluated.) (6) The process of proposing experimental point plans (the process of using the constructed machine learning model to predict the composition or structure of materials with desirable properties or performance, and identifying candidate experimental conditions to be carried out next). (7) Data collection process (the process of conducting experiments under the conditions specified in the previous process (6) and obtaining data)

[0020] For example, the storage device 50 stores each piece of data constituting the data set 60 at the time point (second time point) when the data set 60 makes a state transition, the timestamp of the time point when the data set 60 makes a state transition, and the number of times the data set 60 makes a state transition (number of transitions in a cycle) in association with each other. This allows the storage device 50 to store the data set 60 in association with the cycle.

[0021] The processing device 40 includes a processor and a memory, and the processor executes a program to realize various functions. The processing device 40 includes a prediction unit 41, a generation unit 42, an identification unit 43, a calculation unit 44, a determination unit 45, and an output unit 46 as the various functions.

[0022] The prediction unit 41 predicts a dependent variable from explanatory variables included in a dataset 60. The prediction unit 41 predicts a dependent variable from explanatory variables using a regression model 41a that has been trained in advance by machine learning. The regression model 41a is constructed using a known model. For example, the regression model 41a is constructed using a linear model such as Lasso or Ridge, or a decision tree model such as Random Forest. The prediction unit 41 stores the predicted result in a dataset 60 in the storage device 50.

[0023] The generation unit 42 generates a table 70 indicating the retention state of data in the dataset 60 stored in the storage device 50, and outputs the table 70 to the output unit 46. The output unit 46 outputs the table 70 generated by the generation unit 42 to the communication device of the server 3. The communication device of the server 3 transmits the table 70 output from the output unit 46 to the user terminal 2. The table 70 transmitted to the user terminal 2 is displayed on the display device 22.

[0024] Table 70 is made up of multiple storage locations (data storage areas) in which data is stored, as shown in the lower diagram of Figure 2. These storage locations are divided into storage locations 71 where data is stored and missing locations 72 where data is missing.

[0025] The generation unit 42 generates the table 70 by displaying the retained portion 71 and the missing portion 72 in different display modes. The display modes of the retained portion 71 and the missing portion 72 are not particularly limited as long as the colors, patterns, shapes, etc. displayed in the retained portion 71 and the missing portion 72 are different from each other. In this embodiment, as shown in the lower diagram of FIG. 2, the retained portion 71 is shown in white, and the missing portion 72 is shown in dark gray. Note that the value of the data corresponding to the retained portion 71 may be displayed in the retained portion 71.

[0026] The generation unit 42 generates the table 70 for each cycle. For example, as shown in FIG. 3 , the generation unit 42 generates the table 70 for each cycle specified by the user. The output unit 46 outputs the table 70 generated by the generation unit 42 for each cycle. The user terminal 2 displays a slider bar 75 together with the table 70. The slider bar 75 is a user interface for switching the table 70 displayed on the user terminal 2 for each cycle. The slider bar 75 accepts operation of the slider 76 by the user. Upon accepting operation of the slider 76, the user terminal 2 requests the server 3 to generate and transmit the table 70 corresponding to the cycle corresponding to the position at which the slider 76 was operated. The generation unit 42 generates the table 70 in response to the request and transmits it to the user terminal 2 via the output unit 46.

[0027] This allows the table 70 displayed on the user terminal 2 to be switched for each cycle specified by the user, so that the analysis device 4 can easily allow the user to understand which data was added in that cycle. Thus, the analysis device 4 allows the user to easily understand changes in the data retention state.

[0028] As shown in FIG. 3, the generation unit 42 generates the form of the table 70 of the previous cycle in accordance with the form (the number and shape of rows and columns) of the table 70 of the last cycle to be switched. At this time, all storage locations of unacquired data in the table 70 of the previous cycle are generated as missing locations 72. In the example shown in FIG. 3, the table 70 of cycle 1 does not contain data of samples "#008" and "#009", but the table 70 of cycle 2 has added data of samples "#008" and "#009". In this case, rows of samples "#008" and "#009" are generated in the table 70 of cycle 1, and all storage locations included in the rows are generated as missing locations 72.

[0029] As a result, the form of table 70 does not change between cycles, and analysis device 4 can make it easier for the user to compare the data retention state between cycles. Therefore, analysis device 4 allows the user to easily understand changes in the data retention state.

[0030] Fig. 4 is a diagram illustrating the identification unit 43 shown in Fig. 1. In Fig. 4, the Nth cycle (N is a natural number) is represented as "cycle N".

[0031] The identification unit 43 identifies a storage location of data that is highlighted in the table 70 generated by the generation unit 42. Specifically, the identification unit 43 identifies a storage location that has changed from a missing location 72 to a retention location 71 in response to a transition between cycles as a difference location between cycles. In the example shown in FIG. 4, when a transition occurs from cycle N-2 to cycle N-1, a missing location 72a in cycle N-2 changes to a retention location 71 in cycle N-1. In this case, the identification unit 43 identifies the storage location that has changed from the missing location 72a to the retention location 71 as a difference location 73a. The generation unit 42 generates a table 70 in which the difference location 73a identified by the identification unit 43 is highlighted, as the table 70 for cycle N-1. Similarly, in the example shown in FIG. 4, when a transition occurs from cycle N-1 to cycle N, a missing location 72b in cycle N-1 changes to a retention location 71 in cycle N. In this case, the identification unit 43 identifies the storage location that has changed from the missing location 72b to the holding location 71 as the difference location 73b. The generation unit 42 generates the table 70 in which the difference location 73b identified by the identification unit 43 is highlighted, as the table 70 for cycle N.

[0032] This allows the analysis device 4 to make it easier for the user to understand the difference in the data retention state between cycles, for example, the difference in the data retention state between the current cycle and the immediately preceding cycle. Thus, the analysis device 4 allows the user to easily understand the change in the data retention state.

[0033] Furthermore, the identification unit 43 identifies a missing portion 72 that exists in the same storage location over multiple M cycles (M is an integer equal to or greater than 2 that can be arbitrarily set by the user). In the example shown in FIG. 4, the missing portion 72c in cycle N-2 remains as the missing portion 72 even after transition to cycle N, resulting in a state in which data is missing over two cycles. In this case, the identification unit 43 identifies the missing portion 72c that exists in the same storage location over M=2 cycles. The generation unit 42 generates a table 70 in which the missing portion 72c identified by the identification unit 43 is highlighted, as the table 70 for cycle N. At this time, the generation unit 42 may generate information (e.g., a message 81) indicating that the same storage location is missing over multiple cycles. The output unit 46 may output the information generated by the generation unit 42 together with the table 70 to a communication device of the server 3, and display the information together with the table 70 on the user terminal 2.

[0034] This allows the analysis device 4 to make it easier for the user to understand storage locations where data is missing over multiple cycles, and to alert the user. Therefore, the analysis device 4 can make the user easily understand changes in the data retention state, and can also urge the user to add data.

[0035] The highlighting mode is not particularly limited. For example, the highlighting mode may be such that the corresponding storage location is displayed in a distinctive color different from the normal display color, or such that the corresponding storage location is displayed by blinking, or such that the corresponding storage location is displayed surrounded by a distinctive frame.

[0036] FIG. 5 is a diagram illustrating the calculation unit 44 and the determination unit 45 shown in FIG.

[0037] The calculation unit 44 calculates feature importance, which is an index for evaluating the degree of influence of explanatory variables on the prediction of the dependent variable. Specifically, the calculation unit 44 calculates feature importance of explanatory variables input to the regression model 41a. The calculation unit 44 can calculate the feature importance using a known method. For example, if the regression model 41a is a linear model, the calculation unit 44 can calculate the feature importance based on regression coefficients. For example, if the regression model 41a is a decision tree model, the calculation unit 44 can calculate the feature importance based on Gini impurity, cross entropy, or the like. The calculation unit 44 may represent the calculated feature importance in a format that is easy for the user to understand, such as a graph 90 shown in FIG. 5, output it to the output unit 46, and display it on the user terminal 2.

[0038] The identification unit 43 identifies storage locations corresponding to explanatory variables whose feature importance calculated by the calculation unit 44 is equal to or greater than a threshold. The generation unit 42 generates a table 70 in which the storage locations identified by the identification unit 43 are highlighted, and also generates information (e.g., a message 82) urging the user to preferentially add data that compensates for missing locations 72 included in the storage locations identified by the identification unit 43. The output unit 46 outputs the table 70 generated by the generation unit 42 and the information together to a communication device of the server 3, and causes the table 70 and the information to be displayed together on the user terminal 2.

[0039] In the example shown in FIG. 5, the feature importance of the explanatory variable "composition A" is equal to or greater than a threshold. The identification unit 43 identifies the storage location of the entire column of "composition A." The generation unit 42 creates a table 70 in which the entire column of "composition A" is highlighted by surrounding it with a thick border. Furthermore, the generation unit 42 generates a message 82 that prompts the user to preferentially add data that compensates for missing portions 72 contained in the storage location of the entire column of "composition A." The output unit 46 outputs the table 70 and the message 82 generated by the generation unit 42 together to the communication device of the server 3, and displays the table 70 and the message 82 together on the user terminal 2.

[0040] This allows the analysis device 4 to make it easier for the user to understand the storage location of data indicating explanatory variables with high feature importance, and can alert the user to prioritize adding data to missing portion 72 included in the storage location. Thus, the analysis device 4 can make the user easily understand changes in the data retention state, and can urge the user to add data indicating important explanatory variables.

[0041] The judgment unit 45 determines the priority between a first plan that adds data to compensate for the missing portion 72 and a second plan that adds new sample data 61 based on the degree of influence that the missing portion 72 has on the behavior of the regression model 41a.

[0042] Specifically, first, the determination unit 45 estimates the degree of influence that the missing portion 72 has on the behavior of the regression model 41a. The degree of influence that the missing portion 72 has on the behavior of the regression model 41a refers to the degree to which the prediction result (objective variable) of the regression model 41a changes when data is added to the missing portion 72.

[0043] The method for estimating the degree of influence of the missing portion 72 on the behavior of the regression model 41a is not particularly limited, but the determination unit 45 can estimate the degree of influence using the following method. For example, the determination unit 45 constructs multiple regression models 41a by substituting multiple pieces of data for the missing portion 72, and inputs test data into each of the constructed regression models 41a to obtain a prediction result for each regression model 41a. The determination unit 45 then calculates the variance (prediction accuracy) of the prediction result for each acquired regression model 41a, and estimates the calculated variance as the degree of influence. The determination unit 45 estimates the degree of influence for all missing portions 72 included in the table 70.

[0044] The plurality of data to be substituted for the missing portion 72 may be prepared, for example, by sampling from a normal distribution whose variables are the mean value and standard deviation of the column that includes the missing portion 72. Alternatively, the plurality of data to be substituted for the missing portion 72 may be prepared by sampling from a uniform distribution whose variables are the maximum and minimum values ​​of the column that includes the missing portion 72.

[0045] Next, the determination unit 45 determines whether to prioritize the first plan or the second plan based on the estimated degree of influence. Specifically, if the estimated degree of influence is equal to or greater than a reference value, the determination unit 45 determines that the degree of influence is high and therefore determines that the first plan should be prioritized over the second plan. If there are multiple missing portions 72 whose degrees of influence are equal to or greater than the reference value, the determination unit 45 may determine that the first plan should be executed preferentially in descending order of the missing portion 72 with the highest degree of influence. On the other hand, if the estimated degree of influence is less than the reference value, the determination unit 45 determines that the degree of influence is low and therefore determines that the second plan should be prioritized over the first plan.

[0046] 5, if the degree of influence of the missing portion 72d on the behavior of the regression model 41a is equal to or greater than a reference value, the determination unit 45 determines that the first plan, which adds data to compensate for the missing portion 72d, is to be prioritized over the second plan. In the example shown in FIG. 5, if the degree of influence of all missing portions 72 is less than a reference value, the determination unit 45 determines that the second plan, which adds new sample data 61, is to be prioritized over the first plan.

[0047] The identification unit 43 identifies a storage location corresponding to the first plan or the second plan according to the determination result of the determination unit 45. The generation unit 42 generates a table 70 in which the storage location identified by the identification unit 43 is highlighted, and generates information (e.g., a message 83 or a message 84) that prompts the execution of the first plan or the second plan according to the determination result of the determination unit 45. The output unit 46 outputs the table 70 generated by the generation unit 42 and the information together to the communication device of the server 3, and causes the table 70 and the information to be displayed together on the user terminal 2.

[0048] 5 , if the determination unit 45 determines that the first plan for adding data to compensate for the missing portion 72d is to be prioritized, the identification unit 43 identifies the missing portion 72d as the storage location corresponding to the first plan. The generation unit 42 creates a table 70 in which the missing portion 72d is highlighted by surrounding it with a thick border. Furthermore, the generation unit 42 generates a message 83 urging the user to prioritize the execution of the first plan for adding data to compensate for the missing portion 72d. The output unit 46 outputs the table 70 and the message 83 generated by the generation unit 42 together to the communication device of the server 3, and displays the table 70 and the message 83 together on the user terminal 2.

[0049] 5, if the determination unit 45 determines that the second plan is to be prioritized, the identification unit 43 identifies the storage location of the entire row "#008" in which the new sample data 61 is stored as the storage location corresponding to the second plan. The generation unit 42 generates the storage location of the entire row "#008" as a missing portion 72e, and creates a table 70 in which the entire row "#008" is highlighted by surrounding it with a thick border. Furthermore, the generation unit 42 generates a message 84 that prompts the user to prioritize the execution of the second plan in which the new sample data 61 is added. The output unit 46 outputs the table 70 and the message 84 generated by the generation unit 42 together to the communication device of the server 3, and displays the table 70 and the message 84 together on the user terminal 2.

[0050] This allows the analysis device 4 to make it easier for the user to understand the missing portion 72 that has a high impact on the behavior of the regression model 41a, and to call the user's attention to prioritize adding data to the missing portion 72. Furthermore, the analysis device 4 can make it easier for the user to understand the priority between a first plan for adding data to compensate for the missing portion 72 and a second plan for adding new sample data 61, and can provide the user with guidelines for experimental planning. Thus, the analysis device 4 can make it easier for the user to understand changes in the data retention state, and can prompt the user to add data to compensate for the missing portion 72 that has a high impact, and can provide the user with guidelines for experimental planning.

[0051] FIG. 6 is a diagram illustrating another function of the generating unit 42 shown in FIG.

[0052] As shown in the upper diagram of FIG. 6 , the dataset 60 stored in the storage device 50 may include not only data representing the composition, structure, or performance of a material, but also information indicating the retention status of various measurement data, such as XRD measurement data or SEM measurement data. The generation unit 42 may generate a table 70 indicating not only the retention status of data representing the composition, structure, or performance of a material, but also the retention status of the measurement data. For example, as shown in the lower diagram of FIG. 6 , the generation unit 42 may generate the table 70 by associating the information indicating the retention of the measurement data with each sample by adding a storage location of an icon indicating the retention of the measurement data as a new column to the table 70. The generation unit 42 may generate the table 70 by displaying a retention location 71 in which the information indicating the retention of the measurement data is stored differently from a missing location 72f in which the information indicating the retention of the measurement data is not stored.

[0053] This allows the analysis device 4 to allow the user to easily grasp changes in the state of data retention, and also allows the user to easily grasp which samples have been measured.

[0054] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to these embodiments and various modifications can be made without departing from the spirit of the present invention. In the present invention, elements of one embodiment can be added to elements of another embodiment, elements of one embodiment can be replaced with elements of another embodiment, or some of the elements of one embodiment can be deleted. [Explanation of symbols]

[0055] 1...information processing system, 2...user terminal, 21...processing device, 22...display device, 3...server, 4...analysis device, 40...processing device, 41...prediction unit, 41a...regression model, 42...generation unit, 43...identification unit, 44...calculation unit, 45...determination unit, 46...output unit, 50...storage device (storage unit), 60...dataset, 61...sample data, 70...table, 71...storage location, 72, 72a to 72f...missing location, 73a, 73b...difference location, 75...slider bar, 76...slider, 81 to 84...message (information), 90...graph, N...network

Claims

1. An analytical device for analyzing data representing a composition, structure, or performance of a material, comprising: a storage unit that stores a dataset including the data; a generating unit that generates a table indicating a retention state of the data in the data set stored in the storage unit; an output unit that outputs the table generated by the generation unit, the storage unit stores the data set in association with a cycle indicating a process in which a state of the data set transitions due to the addition of the data; the generation unit generates the table for each cycle by displaying, in different display modes, a missing portion, which is the storage portion where the data is missing, and a retained portion, which is the storage portion where the data is retained, among the storage portions of the data that constitute the table; The output unit outputs the table generated by the generation unit for each cycle. An analytical device characterized by:

2. an identifying unit for identifying the storage location to be highlighted in the table; the identifying unit identifies the defective portion that exists in the same storage location over a plurality of cycles, The generating unit generates the table in which the missing portion identified by the identifying unit is highlighted.

2. The analysis device according to claim 1 .

3. an identifying unit for identifying the storage location to be highlighted in the table; the identifying unit identifies the storage location that has changed from the missing location to the retained location in response to the transition of the cycle as a difference location between the cycles; The generating unit generates the table in which the difference portion identified by the identifying unit is highlighted.

2. The analysis device according to claim 1 .

4. the dataset includes data items designated in advance as a response variable and an explanatory variable, The analysis device a prediction unit that predicts the dependent variable from the explanatory variables; a calculation unit that calculates a feature importance that is an index for evaluating the degree of influence of the explanatory variables on the prediction of the objective variable; an identification unit that identifies the storage location that is highlighted in the table, the identifying unit identifies the storage location corresponding to the explanatory variable for which the feature importance calculated by the calculating unit is equal to or greater than a threshold; the generating unit generates the table in which the storage location identified by the identifying unit is highlighted, and generates information that prompts the user to preferentially add the data that compensates for the missing location included in the storage location identified by the identifying unit; The output unit outputs the table and the information generated by the generation unit.

2. The analysis device according to claim 1 .

5. The dataset is a set of a plurality of sample data, each of which includes data items designated in advance as a response variable and an explanatory variable, and each of which is made up of a plurality of data; The analysis device a prediction unit that predicts the dependent variable from the explanatory variables using a regression model; a determination unit that determines the priority of a first plan for adding the data that compensates for the missing portion and a second plan for adding new sample data; an identification unit that identifies the storage location that is highlighted in the table, the determination unit estimates a degree of influence of the missing portion on a behavior of the regression model, and determines that the first plan is prioritized over the second plan when the estimated degree of influence is equal to or greater than a reference value, and determines that the second plan is prioritized over the first plan when the estimated degree of influence is less than the reference value; The identification unit identifies the storage location corresponding to the first plan or the second plan according to a determination result of the determination unit; the generation unit generates the table in which the storage location identified by the identification unit is highlighted, and generates information to prompt execution of the first plan or the second plan according to a determination result of the determination unit; The output unit outputs the table and the information generated by the generation unit.

2. The analysis device according to claim 1 .

Citation Information

Patent Citations

  • Device and method for supporting experimental design, and program therefor

    JP2002259464A