Analysis device

The analysis device addresses the challenge of tracking data distribution changes by generating cycle-specific maps, enhancing user understanding of data evolution and material development trends.

JP2026017020APending Publication Date: 2026-02-04TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024117633
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

In materials development using materials informatics, it is challenging to understand the order in which data distribution is interpolated for each cycle, making it difficult to grasp changes in data distribution states.

Method used

An analysis device that generates maps showing data distribution by setting different display modes for data added in each cycle, allowing visualization of how data distribution changes over time.

Benefits of technology

Enables users to easily understand changes in data distribution states, facilitating better comprehension of experimental trends and material development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017020000001_ABST
    Figure 2026017020000001_ABST
Patent Text Reader

Abstract

To enable a user to easily grasp a change in a distribution state of data.SOLUTION: The analysis device includes a storage unit that stores a data set including data representing a composition, a structure, or performance of a material, a generation unit that generates a map indicating a distribution of data included in the data set stored in the storage unit, and an output unit that outputs the map generated by the generation unit. The storage unit stores a data set in association with a cycle indicating a process in which a state of the data set transitions due to addition of data. The generation unit generates the map for each cycle such that a display mode of data added according to the transition of the cycle among the data used for generating the map is different from a display mode of other data. The analysis device may include a prediction unit configured to predict an objective variable from an explanatory variable using a machine-learned regression model, and the data set may include the objective variable and a data item designated as the explanatory variable.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an analysis device. [Background technology]

[0002] Patent Document 1 discloses a support device for supporting optimization analysis in accordance with experimental design, which includes a selection means for selecting, from a list of multiple candidates displayed on a display screen, an error factor or a signal factor to be assigned to a level value table used to prepare data necessary for optimization analysis, and a display control means for identifying a candidate that has already been selected as a control factor from the candidate list and changing the display attribute so that it can be distinguished from other candidates. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2002-259464 Summary of the Invention [Problem to be solved by the invention]

[0004] In materials development using materials informatics, it is common to increase the amount of data included in a dataset by repeating the cycle of adding experimental data, analyzing the data, planning the next experimental point, and then adding more experimental data. When repeating this cycle, sample data may be acquired for multiple samples while determining priorities. However, once sample data is acquired, it can sometimes be difficult to understand the order in which the data distribution was complemented for each cycle.

[0005] Even if an experiment is performed using the experimental design method disclosed in Patent Document 1, it is difficult to understand the order in which the data distribution is interpolated for each cycle.

[0006] The present invention has been made in view of the above circumstances, and has as its object to allow a user to easily grasp changes in the distribution state of data. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems, the analysis device of the present invention is an analysis device that analyzes data representing the composition, structure, or performance of a material, and is equipped with a memory unit that stores a dataset including the data, a generation unit that generates a map showing the distribution of the data included in the dataset stored in the memory unit, and an output unit that outputs the map generated by the generation unit, wherein the memory unit stores the dataset in association with a cycle that shows a process in which the state of the dataset changes with the addition of the data, and the generation unit generates the map for each cycle by setting a display mode for the data added in accordance with the transition of the cycle among the data used to generate the map to a display mode that is different from that for the other data. [Effects of the Invention]

[0008] According to the present invention, it is possible for a user to easily grasp changes in the distribution state of data. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram showing the configuration of an information processing system including an analysis device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram for explaining a data set stored in the storage device shown in FIG. 1. [Figure 3] FIG. 2 is a diagram for explaining a generating unit and a selecting unit shown in FIG. [Figure 4] FIG. 2 is a diagram for explaining a calculation unit and a generation unit shown in FIG. [Figure 5] 5 is a diagram for explaining another example of the representative points shown in FIG. 4. [Figure 6] 5 is a diagram for explaining another example of the representative points shown in FIG. 4. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Components with the same reference numerals in each embodiment have similar components in each embodiment unless otherwise specified, and description thereof will be omitted.

[0011] FIG. 1 is a diagram showing the configuration of an information processing system 1 including an analysis device 4 of this embodiment. FIG. 2 is a diagram explaining a dataset 60 stored in a storage device 50 shown in FIG. 1. FIG. 3 is a diagram explaining a generation unit 42 and a selection unit 43 shown in FIG. 1. Note that in FIGS. 2 and 3, the Nth cycle (N is a natural number) is represented as "cycle N." A map 70 shown in FIG. 3 shows the distribution of data included in the dataset 60 shown in FIG. 2.

[0012] The information processing system 1 is a system that analyzes data input by a user using an analysis device 4 and presents the analysis results to the user. The information processing system 1 includes a user terminal 2 and a server 3 that includes the analysis device 4. The user terminal 2 and the server 3 are connected to each other via a network N so that they can communicate with each other.

[0013] The user terminal 2 includes a processing device 21 and a display device 22. The processing device 21 includes a processor and a memory, and the processor executes a program to realize the functions of the user terminal 2. The display device 22 includes a display and displays the processing results of the processing device 21.

[0014] The user terminal 2 accepts data input by the user and transmits it to the analysis device 4 of the server 3. The user terminal 2 receives the analysis results of the analysis device 4 from the server 3, displays them, and presents them to the user.

[0015] The analysis device 4 analyzes data input by a user. The data input to the analysis device 4 is, for example, data representing the composition or structure of a material obtained from various measurement data related to material analysis, such as XRD (X-ray diffraction) measurement data or SEM (Scanning Electron Microscope) measurement data. The material is, for example, a metal material, a resin material, or a coating material used in a vehicle. As data analysis, the analysis device 4 predicts, for example, the performance of a material from the data representing the composition or structure of the material. The performance of the material may be, for example, battery performance, magnetic performance, rigidity, thermoplasticity, tensile performance, or mechanical durability. Furthermore, as data analysis, the analysis device 4 may predict candidate compositions or structures of the material from the data representing the performance of the material.

[0016] The analysis device 4 includes a processing device 40 and a storage device 50. The storage device 50 includes a storage for storing data received from the user terminal 2 and analysis results of the analysis device 4. For example, the storage device 50 stores data representing the composition or structure of a material as data received from the user terminal 2, and stores data representing the performance of the material as analysis results of the analysis device 4. That is, as shown in FIG. 2, the storage device 50 stores a data set 60 including data representing the composition, structure, or performance of a material.

[0017] In the example shown in FIG. 2, the dataset 60 is a collection of multiple sample data 61. Each sample data 61 consists of multiple data divided into multiple data items representing the composition, structure, or performance of a material, and a sample name (or sample ID). The dataset 60 includes data items designated by the user as the objective variables and explanatory variables of the prediction unit 41, which will be described later. In the example shown in FIG. 2, the explanatory variables are data items representing the composition of the material, and the objective variables are data items representing the performance of the material, but the analysis device 4 of this embodiment is not limited to this. The explanatory variables may be data items representing the structure of the material, and the objective variables may be data items representing the performance of the material. The explanatory variables may be data items representing the performance of the material, and the objective variables may be data items representing candidates for the composition or structure of the material.

[0018] The storage device 50 stores the dataset 60 in association with a cycle that indicates a process in which the state of the dataset 60 changes due to the addition of data. That is, a cycle is a process from a first point in time when the dataset 60 is in a first state to a second point in time when data is added to the dataset 60 and the dataset 60 changes to a second state.

[0019] In the context of materials informatics, a cycle consists of at least one of the following processes: (1) Integration of additional data into existing data (60 existing datasets) (2) Preprocessing (the process of converting data into a format suitable for modeling) (3) Feature extraction (the process of extracting descriptors from data that can be used for modeling using domain knowledge or statistical methods) (4) Model construction (the process of using the extracted descriptors to build and train a machine learning model to predict the properties or performance of the material). (5) Model evaluation (the process of evaluating the performance of the constructed machine learning model using validation data. In this process, the model's predictive accuracy or generalization ability is evaluated.) (6) The process of proposing experimental point plans (the process of using the constructed machine learning model to predict the composition or structure of materials with desirable properties or performance, and identifying candidate experimental conditions to be carried out next). (7) Data collection process (the process of conducting experiments under the conditions specified in the previous process (6) and obtaining data)

[0020] For example, the storage device 50 stores each piece of data constituting the data set 60 at the time point (second time point) when the data set 60 makes a state transition, the timestamp of the time point when the data set 60 makes a state transition, and the number of times the data set 60 makes a state transition (number of transitions in a cycle) in association with each other. This allows the storage device 50 to store the data set 60 in association with the cycle.

[0021] The processing device 40 includes a processor and a memory, and the processor executes a program to realize various functions. The processing device 40 includes a prediction unit 41, a generation unit 42, a selection unit 43, a calculation unit 44, and an output unit 45 as the various functions.

[0022] The prediction unit 41 predicts a dependent variable from explanatory variables included in a dataset 60. The prediction unit 41 predicts a dependent variable from explanatory variables using a regression model 41a that has been trained in advance by machine learning. The regression model 41a is constructed using a known model. For example, the regression model 41a is constructed using a linear model such as Lasso or Ridge, or a decision tree model such as Random Forest. The prediction unit 41 stores the predicted result in a dataset 60 in the storage device 50.

[0023] The generation unit 42 generates a map 70 that indicates the distribution of data included in the dataset 60 stored in the storage device 50, and outputs the map 70 to the output unit 45. The output unit 45 outputs the map 70 generated by the generation unit 42 to the communication device of the server 3. The communication device of the server 3 transmits the map 70 output from the output unit 45 to the user terminal 2. The map 70 transmitted to the user terminal 2 is displayed on the display device 22.

[0024] The generation unit 42 generates, for example, a scatter plot as the map 70. The generation unit 42 generates the map 70 by changing the display mode of the map 70 for each cycle. Specifically, the generation unit 42 generates the map 70 for each cycle by setting the display mode of data used to generate the map 70, which is added in accordance with the transition of the cycle, to a display mode that is different from that of the other data.

[0025] 2, in cycle N, data set 60 has missing data 62. When transitioning from cycle N to cycle N+1, data 63 is added to make up for part of missing data 62, and in cycle N+1, part of missing data 62 is resolved, but missing data 64 remains. When transitioning from cycle N+1 to cycle N+2, data 65 is added to make up for missing data 64, and in cycle N+2, missing data 64 is resolved.

[0026] In the example shown in Fig. 2, the generation unit 42 generates the map 70 by setting the display mode of the data 63 added when transitioning from cycle N to cycle N+1 to a different display mode from that of the data already held in cycle N, as shown in Fig. 3. In the example shown in Fig. 3, the generation unit 42 generates the map 70 by using hatched circles as the legend for the data 63 added in cycle N+1 and white circles as the legend for the data already held in cycle N.

[0027] Similarly, in the example shown in Fig. 2, the generation unit 42 generates the map 70 in such a way that the display mode of the data 65 added when transitioning from cycle N+1 to cycle N+2 is different from the display mode of the data already held in cycle N+1, as shown in Fig. 3. In the example shown in Fig. 3, the generation unit 42 generates the map 70 in such a way that the legend for the data 65 added in cycle N+2 is a black circle, and the legend for the data already held in cycle N+1 is a white circle and a hatched circle.

[0028] This allows the analysis device 4 to visualize how the data distribution changes for each cycle, allowing the user to easily understand. Therefore, the analysis device 4 allows the user to easily understand the changes in the data distribution state.

[0029] The selection unit 43 selects data items to be displayed as coordinate axes constituting the map 70 from among multiple data items constituting the dataset 60. The number of data items selected by the selection unit 43 is any number equal to or greater than one. Specifically, the selection unit 43 selects data items specified by the user as data items to be displayed as the coordinate axes. The user terminal 2 displays an input form 80 together with the map 70. The input form 80 is a user interface for specifying data items to be selected as variables of the coordinate axes constituting the map 70 displayed on the user terminal 2. The input form 80 has check boxes 81 corresponding to the data items and accepts user operation of the check boxes 81. Upon accepting operation of the check boxes 81, the user terminal 2 requests the server 3 to generate and transmit a map 70 showing the distribution of data belonging to the data items corresponding to the checked check boxes 81. In response to the request, the selection unit 43 selects data items to be displayed as coordinate axes constituting the map 70 and outputs the selected data items to the generation unit 42. The generating unit 42 generates a map 70 indicating the distribution of data belonging to the data item selected by the selecting unit 43 , and transmits the map 70 to the user terminal 2 via the output unit 45 .

[0030] This allows the analysis device 4 to generate a map 70 according to the data item specified by the user and display it on the user terminal 2, so that the analysis device 4 can easily allow the user to understand how the data belonging to the data item is distributed. Thus, the analysis device 4 allows the user to easily understand changes in the distribution of the data.

[0031] In particular, the multiple data items constituting the dataset 60 include data items specified as the response variable and explanatory variables of the prediction unit 41. The selection unit 43 selects at least one of the response variable and explanatory variables as data items to be displayed as coordinate axes constituting the map 70. The generation unit 42 generates the map 70 showing the distribution of data belonging to at least one of the response variable and explanatory variables selected by the selection unit 43.

[0032] This allows the user to easily understand the relationship between the explanatory variables and the response variable and how the distribution changes for each cycle. Thus, the analysis device 4 allows the user to easily understand changes in the distribution state of the data.

[0033] 3, three check boxes 81, "composition A," "composition B," and "performance value 2," are checked, and the generation unit 42 generates a two-dimensional scatter plot as the map 70, so three maps 71 to 73 are generated. However, the present invention is not limited to this, and the generation unit 42 may generate a three-dimensional scatter plot as the map 70, or may generate a map 70 other than a scatter plot.

[0034] FIG. 4 is a diagram illustrating the calculation unit 44 and the generation unit 42 shown in FIG.

[0035] The calculation unit 44 calculates, for each cycle, a representative point 75 in the distribution of data used to generate the map 70. The representative point 75 is, for example, a center of gravity in the distribution of data used to generate the map 70. The generation unit 42 generates the map 70 by superimposing the representative point 75 for each cycle calculated by the calculation unit 44 on the map 70.

[0036] In the upper diagram of FIG. 4, a map 74 with the variable P and the variable Q as coordinate axes is shown. Assume that a data set 60 containing data used for generating the map 74 is stored in the storage device 50 for each cycle, such as "Cycle 1", "Cycle 2", ···, "Cycle N-1", "Cycle N". When calculating the representative point 75 to be superimposed on the map 74, the calculation unit 44 first selects the cycle for which the representative point 75 is to be calculated. The number of selected cycles is any number of 2 or more. For example, assume that the calculation unit 44 selects "Cycle X1", "Cycle X2", and "Cycle X3" (where X1 < X2 < X3). In this case, the calculation unit 44 specifies the data used for generating the map 74 from the data set 60 at the time of "Cycle X1", and calculates the representative point 75 using the specified data. Similarly, the calculation unit 44 specifies the data used for generating the map 74 from the data set 60 at the time of "Cycle X2", and calculates the representative point 75 using the specified data. Similarly, the calculation unit 44 specifies the data used for generating the map 74 from the data set 60 at the time of "Cycle X3", and calculates the representative point 75 using the specified data. As shown in the lower diagram of FIG. 4, the generation unit 42 generates a map 70 by superimposing the representative points 75 of "Cycle X1", "Cycle X2", and "Cycle X3" calculated by the calculation unit 44 on the map 74.

[0037] Thereby, the analysis device 4 can visualize how the overall distribution of the data changes for each cycle, and can easily let the user grasp the trend of the experimental plan and thus the trend of material development. Therefore, the analysis device 4 can easily let the user grasp the change in the distribution state of the data.

[0038] FIG. 5 is a diagram for explaining another example of the representative point 75 shown in FIG. 4. FIG. 6 is a diagram for explaining another example of the representative point 75 shown in FIG. 4.

[0039] While the representative point 75 shown in FIG. 4 is the center of gravity of the data distribution, the calculation unit 44 may calculate the maximum or minimum value in the data distribution for each cycle as a representative point 76 for each cycle, as shown in FIG. 5. In FIG. 5, the maximum value in the data distribution for each cycle is shown as the representative point 76. In the example shown in FIG. 5, the dashed line, the one-dot-dash line, and the two-dot-dash line perpendicular to the coordinate axis of the variable P indicate the maximum value of the variable P in "Cycle X1," "Cycle X2," and "Cycle X3," respectively. In the example shown in FIG. 5, the dashed line, the one-dot-dash line, and the two-dot-dash line perpendicular to the coordinate axis of the variable Q indicate the maximum value of the variable Q in "Cycle X1," "Cycle X2," and "Cycle X3," respectively.

[0040] Furthermore, the calculation unit 44 may calculate a Pareto set in the data distribution for each cycle as a representative point 77 for each cycle, as shown in FIG. 6. The Pareto set is a set of Pareto-optimal solutions when the data distribution is considered as a multi-objective optimization problem. In other words, the Pareto set refers to a set of all solutions that are not dominated by other points. The user can arbitrarily specify whether a larger or smaller value is to be dominant for each variable. In the example shown in FIG. 6, the larger value of both variable P and variable Q is the dominant direction.

[0041] In this way, the calculation unit 44 calculates the center of gravity in the distribution of data for each cycle, the maximum or minimum value in the distribution of data for each cycle, or the Pareto set in the distribution of data for each cycle as the representative points 75 to 77 for each cycle.

[0042] This allows the analysis device 4 to easily visualize how the overall data distribution changes for each cycle, allowing the user to more easily grasp the trend of the experimental design and, in turn, the trend of material development. Thus, the analysis device 4 allows the user to easily grasp the change in the data distribution state.

[0043] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to these embodiments and various modifications can be made without departing from the spirit of the present invention. In the present invention, elements of one embodiment can be added to elements of another embodiment, elements of one embodiment can be replaced with elements of another embodiment, or some of the elements of one embodiment can be deleted. [Explanation of symbols]

[0044] 1...information processing system, 2...user terminal, 21...processing device, 22...display device, 3...server, 4...analysis device, 40...processing device, 41...prediction unit, 41a...regression model, 42...generation unit, 43...selection unit, 44...calculation unit, 45...output unit, 50...storage device (storage unit), 60...data set, 61...sample data, 62...missing data, 63...added data, 64...missing data, 65...added data, 70-74...map, 75-77...representative point, 80...input form, 81...check box, N...network

Claims

1. An analytical device for analyzing data representing a composition, structure, or performance of a material, comprising: a storage unit that stores a dataset including the data; a generation unit that generates a map showing a distribution of the data included in the dataset stored in the storage unit; an output unit that outputs the map generated by the generation unit, the storage unit stores the data set in association with a cycle indicating a process in which a state of the data set transitions due to the addition of the data; The generation unit generates the map for each cycle by setting a display mode of the data added in accordance with the transition of the cycle among the data used for generating the map to a display mode different from that of the other data. An analytical device characterized by:

2. The data set is a collection of a plurality of sample data, the sample data consists of a plurality of the data divided into a plurality of data items, the analysis device further includes a selection unit that selects the data items to be displayed as coordinate axes that form the map; The generating unit generates the map showing the distribution of the data belonging to the data item selected by the selecting unit.

2. The analysis device according to claim 1 .

3. A prediction unit that predicts a response variable from an explanatory variable using a regression model is further provided, the plurality of data items include data items designated in advance as the objective variable and the explanatory variable, The selection unit selects at least one of the objective variable and the explanatory variable as the data item to be displayed as the coordinate axis.

3. The analysis device according to claim 2.

4. a calculation unit that calculates, for each cycle, a representative point in the distribution of the data used to generate the map; The generating unit generates the map by superimposing the representative point for each cycle on the map using the calculating unit.

2. The analysis device according to claim 1 .

5. The calculation unit calculates, as the representative point for each cycle, a center of gravity in the distribution for each cycle, a maximum value or a minimum value in the distribution for each cycle, or a Pareto set in the distribution for each cycle.

5. The analysis device according to claim 4.

Citation Information

Patent Citations

  • Device and method for supporting experimental design, and program therefor

    JP2002259464A