Analysis method, analyzer, analysis program and computer-readable storage medium storing analysis program
The analysis method addresses the limitation of existing methods by calculating scores based on the number of levels in directed graph structures for each acquisition condition, enabling systematic analysis and comprehensive understanding of Bayesian networks in applications involving different vehicle models.
Patent Information
- Application Number
- JP2023193185
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-23
AI Technical Summary
Existing methods for analyzing Bayesian networks are limited in their ability to systematically analyze changes in graph structures when acquisition conditions for variables change, particularly in applications involving different vehicle models.
An analysis method that generates a graph based on multiple variables and calculates a score characterizing the graph, by determining directed graph structures for each acquisition condition and calculating scores that increase or decrease based on the number of levels in the graph structure.
This method allows for systematic analysis of the association between hierarchical directed graph structures and acquisition conditions, providing a quantitative measure of changes in graph structures and promoting comprehensive understanding and intellectual discovery.
Smart Images

Figure 2025080143000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an analysis method, an analysis device, an analysis program, and a computer-readable storage medium storing the analysis program. [Background technology]
[0002] A so-called Bayesian network is known as an example of a directed graph that is composed of nodes corresponding to a plurality of variables and visualizes the relationships between the variables.
[0003] A Bayesian network is a method for modeling dependencies between variables by graphing the dependencies. Using a Bayesian network, it is possible to gain insights that are difficult to reach using conventional rules of thumb and classical statistical analysis alone, and ultimately to support knowledge discovery.
[0004] In recent years, the application of Bayesian networks to industry has been progressing. As one example, Patent Document 1 below discloses a method for visualizing a Bayesian network and an example of the application of the method to an engineering phenomenon.
[0005] According to Patent Document 1, a hierarchical directed acyclic graph structure can be constructed to reflect the relationships between variables, and the hierarchical structure can be visualized. Based on the visualized graph structure, it becomes possible to verify previously known hypotheses and to encourage the creation of hypotheses themselves.
[0006] For example, the above-mentioned Patent Document 1 gives an example of an application to an engineering phenomenon in which multiple measurement points are set on the center line of a model car roof in a situation where wind flows along the roof. According to the above-mentioned Patent Document 1, a graph structure that is consistent with the wind flow direction is confirmed. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] JP 2021-111063 A Summary of the Invention [Problem to be solved by the invention]
[0008] However, the method described in Patent Document 1 is based on the premise of using variables acquired under specific conditions, and there is still room for exploration into what changes will occur in the graph structure when the acquisition conditions for each variable, such as dependencies between variables, change.
[0009] For example, the application example described in Patent Document 1 merely obtains a graph structure limited to a specific vehicle model. A method for systematically analyzing the model structure obtained for each vehicle model when the vehicle model as an acquisition condition is changed has not been known until now.
[0010] The present disclosure has been made in consideration of these points, and its purpose is to enable systematic analysis of the association between a hierarchical directed graph structure that reflects the relationships between variables and the acquisition conditions for each variable. [Means for solving the problem]
[0011] A first aspect of the present disclosure relates to an analysis method for generating a graph based on a plurality of variables and calculating a score that characterizes the graph by using a computer including a storage unit that stores a program and a calculation unit that executes the program stored in the storage unit. The analysis method includes the steps of: determining, for each acquisition condition of the plurality of variables, a directed graph structure that is composed of nodes corresponding to the plurality of variables and is hierarchical so as to reflect a relationship between the variables, by the calculation unit; determining, with one of the plurality of variables as a target variable and the remaining as explanatory variables, a node that constitutes the directed graph structure and corresponds to the target variable as a leaf node, a node that is connected to the leaf node via one or more edges in the directed graph structure and corresponds to the explanatory variable as a parent node, and a number of edges from the parent node to the leaf node as a number of levels, by the calculation unit; and calculating, for each acquisition condition and the explanatory variable, a score that is set to increase or decrease according to the number of levels.
[0012] According to the first aspect, a score that increases or decreases according to the number of layers is calculated for each acquisition condition and explanatory variable. By calculating such a score, it is possible to quantify the change in the graph structure that occurs when the acquisition condition is changed for each explanatory variable. This quantification makes it possible to systematically analyze the relationship between the graph structure and the acquisition condition.
[0013] Furthermore, quantification such as a score is suitable for computation, which allows for more appropriate analysis without the bias of the analyst.
[0014] In addition, simply visualizing the directed graph structures determined for each acquisition condition individually can make it difficult to make a comprehensive comparison between the structures. In contrast, eliminating preconceptions as described above promotes comprehensive understanding by the analyst, which can lead to intellectual discovery and support for idea generation.
[0015] Furthermore, according to a second aspect of the present disclosure, the calculation unit may cross-tabulate the scores obtained for the acquisition conditions and the explanatory variables so as to construct a partitioned table with one of the acquisition conditions and the explanatory variables as the side of the table and the other as the head of the table.
[0016] According to the second aspect, by cross-tabulating the scores, the scores can be compiled by acquisition conditions and explanatory variables, enabling a more systematic analysis that is free from the analyst's preconceptions.
[0017] Furthermore, according to a third aspect of the present disclosure, the calculation unit performs correspondence analysis on the cross-tabulated scores to map the relationship between the acquisition conditions and the explanatory variables onto a two-dimensional plane.
[0018] According to the third aspect, the relationship between the acquisition conditions and the explanatory variables is visualized as a relative distance on a two-dimensional plane, whereby, when a response variable is given, it is possible to visualize the acquisition conditions that strongly affect the relationship between the response variable and the explanatory variables.
[0019] In addition, by mapping the relationship between the acquisition conditions and explanatory variables on a two-dimensional plane, when multiple acquisition conditions are set for multiple explanatory variables, the relationships can be visualized at a bird's-eye view. The directed graph structure and the relationship between each variable and the acquisition conditions can be analyzed at a bird's-eye view.
[0020] According to a fourth aspect of the present disclosure, the calculation unit may determine, as the directed graph structure, a directed acyclic graph structure indicating a Bayesian network.
[0021] By using a Bayesian network in a directed graph structure, it is possible to visualize the dependency of explanatory variables on the objective variable. It is possible to systematically analyze the effect of changes in acquisition conditions on such dependency. This is useful for the analysis of various phenomena, including engineering phenomena.
[0022] According to a fifth aspect of the present disclosure, the plurality of variables are a plurality of time series data generated by obtaining a predetermined parameter for each of a predetermined period, and the analysis method may include a step in which the calculation unit generates a category data set for each of the plurality of time series data by discretizing a value at each time of each of the plurality of time series data into a multi-level system, and a step in which, when determining the directed graph structure, the calculation unit determines a graph structure that maximizes a conditional probability that the plurality of category data sets are realized when the element g is given, where G is a set of directed acyclic graph structures representing a Bayesian network in which each of the plurality of category data sets is a node, and g is a graph structure that is an element of set G.
[0023] Here, "discretization to a multi-level system" refers to a process of discretizing the time series data values into integer levels. In addition, the multiple variables may be a common variable measured simultaneously at multiple locations, or multiple types of variables measured simultaneously at a common location.
[0024] The fifth aspect shows a method for constructing a Bayesian network with each of a plurality of category data sets as a node. Here, by discretizing each time series data into a multi-level system in advance, it becomes possible to construct a Bayesian network even for time series data that changes continuously over time.
[0025] Further, instead of discretizing each time-series data with respect to time, by discretizing it with respect to its own value, a categorical data set that includes temporal relationships can be generated. As a result, a graph structure that includes temporal dependencies such as causal relationships can be generated. By generating such a graph structure, even when the temporal dependency between categorical data sets is not explicitly shown, a graph structure that suggests the dependency can be obtained. As a result, from the visualized graph structure, the temporal dependency can be traced backward, and thus, it becomes possible to verify conventionally known hypotheses or to newly create hypotheses themselves.
[0026] Further, according to the sixth aspect of the present disclosure, after the directed graph structure is determined, the operation unit may sequentially extract combinations of the leaf nodes and the parent nodes in ascending order of the number of layers.
[0027] According to the sixth aspect, by extracting the dependency between nodes for each number of layers based on the number of layers defined as described above, a graph structure that reaches the target variable as a leaf node can be extracted without omission and without excess or deficiency. At the same time, when performing the extraction, since the number of layers corresponding to each combination is naturally given, the score can be efficiently calculated while minimizing the amount of calculation. As a result, the efficiency of computer operation can be improved.
[0028] Further, according to the seventh aspect of the present disclosure, each of the plurality of variables may indicate a variable that characterizes the running of a vehicle, and the acquisition condition may indicate a condition set to distinguish at least one of the vehicle type, driving proficiency, and running speed.
[0029] According to the seventh aspect, the analysis method can systematically analyze the influence of vehicle type, running speed, etc. on running. This contributes to, for example, the improvement of automobile models and automobile control.
[0030] Furthermore, according to an eighth aspect of the present disclosure, the calculation unit may be configured to set the score corresponding to a node that is separated from the leaf node among the nodes that constitute the directed graph structure as a lower limit value, and to calculate the score so that it increases relative to the lower limit value as the number of hierarchical levels becomes smaller.
[0031] Formulating the score as in the eighth embodiment lends itself to systematic analysis.
[0032] A ninth aspect of the present disclosure relates to an analysis device configured by a computer including a storage unit that stores a program and a calculation unit that executes the program stored in the storage unit, generating a graph based on a plurality of variables and calculating a score that characterizes the graph. The analysis device includes: a graph structure determination means for determining, for each acquisition condition of the plurality of variables, a directed graph structure configured by nodes corresponding to the plurality of variables and hierarchically arranged to reflect a relationship between the variables; a hierarchical number acquisition means for acquiring the hierarchical number for each acquisition condition and the explanatory variable when one of the plurality of variables is set as a target variable and the remaining variables as explanatory variables, a node that constitutes the directed graph structure and corresponds to the target variable is set as a leaf node, a node that is connected to the leaf node in the directed graph structure via one or more edges and corresponds to the explanatory variable is set as a parent node, and the number of edges passing from the parent node to the leaf node is set as a hierarchical number; and a score calculation means for calculating, for each acquisition condition and the explanatory variable, a score set to increase or decrease according to the hierarchical number.
[0033] According to the ninth aspect, it becomes possible to systematically analyze the association between a hierarchical directed graph structure that reflects the relationships between variables and the acquisition conditions of each variable.
[0034] A tenth aspect of the present disclosure relates to an analysis program for generating a graph based on a plurality of variables and calculating a score that characterizes the graph by causing a computer including a storage unit that stores a program and a calculation unit that executes the program stored in the storage unit to execute the analysis program. The analysis program causes the computer to execute the following steps: determining, for each acquisition condition of the plurality of variables, a directed graph structure that is composed of nodes corresponding to the plurality of variables and is hierarchical so as to reflect the relationship between the variables, with one of the plurality of variables being a target variable and the remaining being explanatory variables, a node that constitutes the directed graph structure and corresponds to the target variable being a leaf node, a node that is connected to the leaf node in the directed graph structure via one or more edges and corresponds to the explanatory variable being a parent node, and the number of edges passing from the parent node to the leaf node being the number of levels, with the calculation unit acquiring the number of levels for each of the acquisition conditions and the explanatory variables, and calculating a score that is set to increase or decrease according to the number of levels for each of the acquisition conditions and the explanatory variables.
[0035] According to the tenth aspect, it becomes possible to systematically analyze the association between a hierarchical directed graph structure that reflects the relationships between variables and the acquisition conditions of each variable.
[0036] An eleventh aspect of the present disclosure relates to a computer-readable storage medium. The storage medium stores the analysis program. Effect of the Invention
[0037] As described above, according to the present disclosure, it becomes possible to systematically analyze the association between a hierarchical directed graph structure that reflects the relationships between variables and the acquisition conditions for each variable. [Brief description of the drawings]
[0038] [Figure 1]FIG. 1 is a diagram illustrating an example of a hardware configuration of an analysis device. [Diagram 2] FIG. 2 is a diagram illustrating an example of the software configuration of the analysis device. [Diagram 3] FIG. 3 is a flow chart illustrating the steps of the analysis method. [Figure 4] FIG. 4 is a flowchart illustrating the procedure of the graph structuring analysis method. [Diagram 5] FIG. 5 is a flow chart illustrating the procedure of the pre-processing process. [Figure 6] FIG. 6 is a flowchart illustrating the procedure of the main processing. [Figure 7] FIG. 7 is a flow chart illustrating the steps of the post-processing process. [Figure 8] FIG. 8 is a diagram showing a specific example of the post-processing process. [Figure 9] FIG. 9 is a diagram showing specific examples of a plurality of variables obtained according to acquisition conditions. [Figure 10] FIG. 10 is a diagram showing a specific example of a plurality of acquisition conditions. [Figure 11A] FIG. 11A is a diagram showing a specific example of a hierarchical directed graph structure. [Figure 11B] FIG. 11B is a diagram showing a specific example of a hierarchical directed graph structure. [Figure 11C] FIG. 11C is a diagram showing a specific example of a hierarchical directed graph structure. [Figure 11D] FIG. 11D is a diagram showing a specific example of a hierarchical directed graph structure. [Figure 12] FIG. 12 is a diagram showing a specific example of cross tabulation. [Figure 13] FIG. 13 is a diagram showing a specific example of correspondence analysis. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0039] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. Note that the following description is merely an example.
[0040] <1.Device configuration> FIG. 1 is a diagram illustrating an example of the hardware configuration of a graph structuring analysis device according to the present disclosure (specifically, a computer 1 constituting the analysis device), and FIG. 2 is a diagram illustrating an example of the software configuration.
[0041] 1, the computer 1 includes a central processing unit (CPU) 3 that controls the entire computer 1, a read only memory (ROM) 5 that stores a boot program and the like, a random access memory (RAM) 7 that functions as a main memory, and a solid state drive (SSD) 9 as a secondary storage device. Note that instead of the SSD 9, a hard disk drive (HDD) or the like can also be used as the secondary storage device.
[0042] Among these elements, the CPU 3 executes various programs. The CPU 3 functions as a calculation unit in this embodiment. The RAM 7 and the SSD 9 temporarily or continuously store the programs executed by the CPU 3. The RAM 7 and the SSD 9 each function as a storage unit in this embodiment.
[0043] The computer 1 also includes a display 11, a graphics memory (Video RAM: VRAM) 13 that stores image data to be displayed on the display 11, and a keyboard 15 and a mouse 17 as man-machine interfaces. The display 11 functions as a display unit that displays a screen based on the results of calculations by the CPU 3. The computer 1 according to this embodiment can also transmit and receive data to and from external devices via a communication interface 21.
[0044] As illustrated in FIG. 2, the program memory of SSD9 stores an operating system (OS) 19, a pre-processing program 29A, a main processing program 29B, a post-processing program 29C, a hierarchical number acquisition program 292, a score calculation program 293, a correspondence analysis program 294, an application program 39, etc.
[0045] Among these programs, the pre-processing program 29A, the main processing program 29B, and the post-processing program 29C are programs for executing a graph structured analysis (hereinafter referred to as "GSA") described later, and constitute a graph structured analysis program 291. Hereinafter, this will be referred to as the GSA program 291.
[0046] GSA is a big data analysis method that combines probability theory (Bayesian estimation) and graph theory, proposed by the inventor of the present application. In this embodiment, GSA is used to determine the graph structure. Note that it is not essential to use GSA to determine the graph structure.
[0047] The GSA program 291, the hierarchical number acquisition program 292, the score calculation program 293, and the correspondence analysis program 294 constitute the analysis program 29 in this embodiment. The hierarchical number acquisition program 292 may be constituted by a post-processing program 29C of the GSA program 291.
[0048] Here, the analysis program 29 is a program for executing an analysis method described below, and is configured to cause the computer 1 to execute each step constituting the analysis method. The analysis program 29 is pre-stored in a computer-readable storage medium 18. This storage medium 18 is a tangible storage medium constituted by a disk medium or the like.
[0049] In the program memory of the SSD 9, each program constituting the analysis program 29 is started in response to an instruction input from the keyboard 15, the mouse 17, etc. At that time, each program is loaded from the SSD 9 to the RAM 7 and executed by the CPU 3.
[0050] Meanwhile, a plurality of variables 49 to be analyzed are stored in the data memory of the SSD 9. Each of the plurality of variables 49 may be sequence data. The sequence data constituting each of the variables 49 may be time-series data. Each of the variables 49 as time-series data may be generated by acquiring a predetermined parameter over a predetermined period of time. Here, the word "acquire" may be read as "measure" or "record" depending on the type of parameter.
[0051] The values of the multiple variables 49 to be analyzed may vary depending on the acquisition conditions. For example, when parameters that characterize the behavior of a vehicle are used as the multiple variables 49, the acquired value of each variable 49 may vary depending on the type of vehicle as an acquisition condition.
[0052] Therefore, the data memory of the SSD 9 stores values of the multiple variables 49 according to the acquisition conditions of the multiple variables 49. For example, if n is a natural number between 1 and N, then the reference symbol 49n in Fig. 2 indicates multiple variables 49 obtained under the nth acquisition condition (simply referred to as "nth condition" in Fig. 2).
[0053] In the analysis method according to the present embodiment, one of the multiple variables 49 is set as a dependent variable, and the rest are set as explanatory variables. The variable 49 selected as the dependent variable is the same under all acquisition conditions.
[0054] The data memory of the SSD 9 also stores a first data set 59 indicating a graph structure generated by the GSA program 291. The first data set 59 is generated for each acquisition condition. The graph structure corresponding to the first data set 59 is configured to arrive at the same objective variable under all acquisition conditions.
[0055] The data memory of the SSD 9 also stores a second data set 69 indicating the hierarchical level s acquired by the hierarchical level acquisition program 292. The second data set 69 is generated for each acquisition condition. The hierarchical level s corresponding to the second data set 69 is acquired for each acquisition condition and for each explanatory variable.
[0056] The data memory of the SSD 9 also stores a third data set 79 indicating the scores calculated by the score calculation program 293. The third data set 79 is generated for each acquisition condition. The scores corresponding to the third data set 79 are calculated for each acquisition condition and for each explanatory variable.
[0057] In addition, various data generated by each program constituting the analysis program 29 and the execution results of the application program 39 are stored in the data memory of the SSD 9 or in the RAM 7 as the main memory, as necessary.
[0058] <2. Overview of analysis method> Fig. 3 is a flowchart illustrating the procedure of an analysis method. The analysis method illustrated in Fig. 3 generates a graph based on a plurality of variables 49, and calculates and visualizes a score that characterizes the graph.
[0059] As shown in FIG. 3, the analysis method is carried out by sequentially executing a graph structure determination process (step S1), a hierarchy number acquisition process (step S2), a score calculation process (step S3), and a correspondence analysis process (step S4).
[0060] The analysis program 29 is configured to cause the computer 1 to execute these processes. That is, of these processes, the graph structure determination process is performed by the CPU 3 executing the above-mentioned GSA program 291. Similarly, the hierarchy number acquisition process is performed by the CPU 3 executing the hierarchy number acquisition program 292, the score calculation process is performed by the CPU 3 executing the score calculation program 293, and the correspondence analysis process is performed by the CPU 3 executing the correspondence analysis program 294.
[0061] When the CPU 3 executes the GSA program 291 etc., an analysis device is configured by the computer 1. That is, the computer 1 functions as an analysis device including a graph structure determination means for executing a graph structure determination process (step S1), a level number acquisition means for executing a level number acquisition process (step S2), a score calculation means for executing a score calculation process (step S3), and a correspondence analysis means for executing a correspondence analysis process (step S4).
[0062] For example, in the graph structure determination process, the CPU 3 determines a directed graph structure constituted by nodes corresponding to each of the multiple variables 49. This directed graph structure is a hierarchical graph structure that reflects the relationships between the variables 49, and is determined for each acquisition condition of the multiple variables 49.
[0063] Before describing the details of the analysis method, the GSA used in the graph structure determination process will be specifically described below. Note that it is not essential to use the GSA in the graph structure determination process. Any method can be used as long as it can determine a hierarchical directed graph structure.
[0064] <3. Details of GSA> Figure 4 is a flow chart illustrating the procedure of GSA. The method illustrated in Figure 4 uses a computer 1 to generate a Bayesian network based on multiple variables (also referred to as "multiple time-series data" in this chapter) 49 and visualize the Bayesian network.
[0065] As shown in FIG. 4, the GSA is implemented by sequentially executing a pre-processing process (step S11), a main processing process (step S12), and a post-processing process (step S13).
[0066] Of these processes, the pre-processing process is performed by the CPU 3 executing the above-mentioned pre-processing program 29A. Similarly, the main processing process is performed by the CPU 3 executing the main processing program 29B, and the post-processing process is performed by the CPU 3 executing the post-processing program 29C.
[0067] By the CPU 3 executing the pre-processing program 29A etc., the computer 1 functions as a graph structure analysis device equipped with a pre-processing means for executing the pre-processing process, a main processing means for executing the main processing process, and a post-processing means for executing the post-processing process.
[0068] Below, we will explain each process that makes up GSA in order.
[0069] (3-1. Pre-processing) Fig. 5 is a flowchart illustrating the procedure of the pre-processing process. The flowchart illustrated in Fig. 5 shows the process performed in step S11 in Fig. 4. That is, when the control process proceeds to step S11 in Fig. 4, the CPU 3 sequentially executes steps S111 to S113 in Fig. 5.
[0070] Specifically, in step S111 of FIG. 5, the CPU 3 reads the time series data 49 (particularly, the time series data 49 corresponding to one acquisition condition).
[0071] Here, assume that a plurality of time series data 49 is composed of p (p is an integer of 2 or more) time series data 49. Also, as variables corresponding to the p time series data 49, p variables x 1 ,…,x p are considered. Assume that each of the p variables is individually set at time t 1 <…<t N .
[0072] In this case, as the time series data 49,
[0073]
Equation
Equation
[0074] Subsequently, in step S112 of FIG. 5, the CPU 3 as an arithmetic unit discretizes the value of each time series data 49 at each time into a multi-level system, thereby generating category data corresponding to each time series data 49.
[0075] Specifically, in this step S112, the data set (the i-th time series data 49) S i is converted into a category data set i consisting of r
Mathematics
Mathematics
[0076] That is, at the stage of Equation (2), the data set S i is composed of N elements classified by time. On the other hand, at the stage of Equation (3), the category data set C i is composed of r i (<N) elements classified by other parameters.
[0077] Preferably, the value of the data set S i at a predetermined time is represented by the variable x, the average value related to the time of the data set S i is μ i , the standard deviation related to the time of the data set S i is σ i , and the category data set corresponding to the data set S i is C i . Then, the mapping φ i is
Mathematics
[0078] More preferably, if the largest integer less than or equal to x is [x], the mapping φ i is
Mathematics
[0079] In particular, as shown in (6), the map φ i depends on the absolute value of the variable x itself, not on the time t at which the variable x was measured. Therefore, the categorical data set C generated through Eq. (6) i appears to have no dependency on time t.
[0080] In addition, if the time series data 49 is judged to have strong non-stationarity over time, the data set S i Divide the data set S i A process may be added in which each variable x constituting the above is divided into present and past variables in advance.
[0081] 5, the CPU 3 stores the categorized time-series data 49 in the RAM 7 or the SSD 9. The stored data is read as necessary in the main processing or other processes. When step S113 is completed, the control process returns from the flow illustrated in FIG. 5 and proceeds to step S12 in FIG.
[0082] For the sake of brevity, we use X, which shows the time series data 49 before categorization, and X, which shows the time series data 49 after categorization. c In addition, a discrete variable that has the i-th column vector component of a data string X consisting of p columns as categorical data is simply denoted as x i We will handle it as such.
[0083] (3-2. Main processing process) Fig. 6 is a flowchart illustrating the procedure of the main processing. The flowchart illustrated in Fig. 6 shows the processing performed in step S12 in Fig. 4. That is, when the control process proceeds to step S12 in Fig. 4, the CPU 3 sequentially executes steps S121 to S123 in Fig. 6.
[0084] In the main processing, the CPU 3 calculates each discrete variable (the i-th column vector component of the data string X) x i A Bayesian network is constructed with the nodes as follows.
[0085] Here, the construction of a Bayesian network is performed by searching for a directed acyclic graph (DAG) structure that represents the Bayesian network. This DAG structure is searched for as a graph structure that maximizes the conditional probability of a data string X with the graph structure as a condition.
[0086] 6, the CPU 3 reads category data (specifically, the categorized data string X). Then, in step S122, the CPU 3 sets a network score based on the category data read in step S121.
[0087] In this embodiment, we define all sets of DAG structures of Bayesian networks that can be expressed by p nodes as G p Let p discrete variables x 1 ,…,x p Using a data sequence X consisting of p discrete variables, we can find the optimal graph structure g∈G p Learn.
[0088] The procedure for setting the network score will be described below.
[0089] Graph structure g∈G p Given, for each i∈{1,…,p}, x i Let Π be the set of parent nodes directly connected to i ⊂{x 1 ,…,x p}, and Π i The number of patterns that can be taken as states is q i Then, q i teeth,
number
number
[0090] The union of all i, j, and k of
number
number
number
number
[0091] The posterior distribution when using equation (11) as the likelihood function and equation (12) as the prior distribution is organized using Bayes' theorem. Specifically, the posterior distribution organized through Bayes' theorem is
number
number
[0092] In this embodiment, the category data set C i p discrete variables x corresponding to i Each of these is regarded as a node constituting the Bayesian network. In other words, the Bayesian network in this embodiment is a Bayesian network consisting of p discrete variables x i It is constructed by connecting them with a single arrow (edge).
[0093] And the discrete variable x i Let G(G as mentioned above) be a set of directed acyclic graph structures that represent the Bayesian network constructed by interconnecting p (corresponding to ), and if the graph structure that is an element of set G is g, then we define a categorical data set C under the condition that element g is given. i We set an overall probability distribution, which is equal to the probability that the data sequence X is realized when the element g is taken as a condition, and is equal to the network score shown in equation (14).
[0094] Then, in step S123 following step S122, the CPU 3 determines a graph structure that maximizes the network score. The graph structure is determined by determining g∈G p This is done through a metaheuristic search algorithm called tabu search.
[0095] For details of tabu search, see, for example, “Bouckaert, R., Bayesian belief networks: from construction to inference, Ph.D. Thesis, University of Utrecht, 1995” and “Acid, S., and de Campos, LM, Searching for Bayesian network structures in the space of restricted acyclic partially directed graphs, Journal of Artificial Intelligence Research 18, pp. 445-490, 2003.”
[0096] In addition, the hyperparameter α ijkIn determining the ratio of prior and data, we adopt the network score “BDeu (Bayesian Dirichlet equivalence uniform)” recommended in “Ueno, M., Learning networks determined by the ratio of prior and data, In Proc. of 26th Conf. on Uncertainty in Artificial Intelligence, pp. 598-605, 2010” and “Ueno, M., Robust learning of Bayesian networks for prior belief, In Proc. of 27th Conf. on Uncertainty in Artificial Intelligence, pp. 698-707, 2011.”. In other words, we adopt the constraint that corresponds to a special case of the sufficient condition to satisfy “likelihood equivalence” described in “Heckerman, D., Geiger, D., and Chickering, DM, Learning Bayesian networks: The combination of knowledge and statistical data, Machine learning, 20, pp. 197-243, 1995.”
number
[0097] Thereafter, in step S124 following step S123, the CPU 3 stores the graph structure determined in step S123 in the RAM 7 or SSD 9. When step S124 is completed, the control process returns from the flow illustrated in Figure 6 and proceeds to step S13 in Figure 4.
[0098] (3-3. Post-processing) Fig. 7 is a flowchart illustrating the procedure of the post-processing process. The flowchart illustrated in Fig. 7 shows the process performed in step S13 in Fig. 4. That is, when the control process proceeds to step S13 in Fig. 4, the CPU 3 sequentially executes steps S131-S136 in Fig. 7.
[0099] Below, multiple variables x i One of the variables is the dependent variable, and the rest are the explanatory variables. To distinguish the dependent variable from the other explanatory variables, we use the term "x i " is sometimes written as "y" instead of ".
[0100] In the post-processing process, the graph structure g ∈ G obtained in the main processing process is p The system extracts a predetermined substructure from the graph and executes a process to visualize at least a part of the graph structure.
[0101] In particular, the post-processing process is a process in which a node constituting a graph structure g has a node corresponding to a target variable y as a child node, and a node connected to the child node via one or more edges in the graph structure g has an explanatory variable x j The node corresponding to the node is regarded as the parent node, and the number of edges when connecting the parent node and the child node in the shortest distance is regarded as the number of layers. In this case, the CPU 3 extracts combinations of child nodes and parent nodes for each number of layers (specifically, in order from the smallest number of layers).
[0102] This creates a set of parent nodes corresponding to the child nodes as the objective variable y (a substructure of g that reaches y and has variables other than the objective variable y, xj It is possible to extract the complete set of parent nodes (hierarchical structure composed of parent nodes).
[0103] In this embodiment, a specific objective variable y is designated as a child node of the graph structure g, and the end (so-called "leaf node") of the graph structure g is configured by the objective variable y. In the post-processing process, the CPU 3 sequentially extracts combinations of leaf nodes and parent nodes in ascending order of the number of layers. In other words, in the post-processing process, the CPU 3 extracts combinations of child nodes and parent nodes so that a graph structure g with the objective variable y as its end is constructed.
[0104] Note that the so-called "child node" can refer to a node located at the end of the graph structure g, or a node located downstream of a parent node and directly connected to the parent node, but in this specification, the "child node" refers to the latter node. Among the child nodes, those that fall into the former category will be called "leaf nodes" as mentioned above.
[0105] The above-mentioned extraction process will be described in detail below with reference to FIG.
[0106] First, in step S131 of Fig. 7, the CPU 3 reads the graph structure g. The graph structure g read in this step is equal to the graph structure determined by the main processing process.
[0107] Next, in step S132, based on the settings stored in advance in the SSD 9 or the like or the settings manually input by the user, the CPU 3 extracts a leaf node x from the extraction target (objective variable y). i Specify the leaf node x i The objective variable y as a function of the end (terminal) of the graph structure g as described above.
[0108] In the subsequent step S133, the CPU 3 iFor each i ∈ {1,…,p}, list the first-level parent nodes for
[0109] In this embodiment, the “sth hierarchical parent node” refers to the x i When x is a leaf node, it refers to the parent node that can be reached via s edges in the shortest time. i Let H be the set of parent nodes in the sth hierarchical level for i (s). For convenience, x i The set of 0th hierarchical parent nodes for is
number
number
number
number
[0110] Next, in step S134, the CPU 3 i Set H of sth hierarchical parent nodes for i Extract (s) recursively for each s.
[0111] Specifically, the set of general sth hierarchical parent nodes H i (s) is H i Among the first-level parent nodes for each element of (s-1), H i (1),…,H i The set of all sets that do not belong to any of the sets in (s-1), i.e., sequentially for s ≥ 2
number
number
number
number
[0112] By repeatedly calculating equation (20) for each s, H i (s) can be calculated recursively. The calculated H i (s) is stored in the RAM 7 or the SSD 9.
[0113] Next, in step S135, the CPU 3 i Each element x that composes (s) m From x i A set of edges L that are passed through to reach i Extract (s) recursively for each s.
[0114] Specifically, all edges in the graph structure g, i.e., the child nodes x m and the child node x m The first-level parent node x corresponding to l The pair (x l ,x m ), then (xl ,x m ) is the set of all
number
number
number
[0115] By repeatedly calculating equations (24) to (26) for each s, L i (s) can be calculated recursively. i (s) is stored in the RAM 7 or the SSD 9.
[0116] Thus, x i The maximum subgraph g that shows the entire substructure of a graph structure g when (i) But x i Includes all parent nodes up to the s(i)th level of
number
number
[0117] Finally, in step S136, the CPU 3 calculates the maximum subgraph g given by equations (27) to (28).(i) , that is, x set to the objective variable y i and the graph hierarchized by s indicating the number of layers is stored in the RAM 7 or SSD 9. When step S136 is completed, the control process returns from the flow illustrated in Fig. 7, and the graph structuring analysis method illustrated in Fig. 4 is terminated.
[0118] -Example of post-processing process- FIG. 8 is a diagram showing a specific example of the post-processing process. Here, the case of p=6, i.e., six-dimensional time series data 49, will be described. By the above-mentioned pre-processing process, six discrete variables x 1 ~x 6 As shown in graph G11 of FIG. 8(a), the six discrete variables x 1 ~x 6 A DAG structure g∈G 6 is assumed to have been obtained.
[0119] In this example, a DAG structure g∈G 6 x 6 The maximum subgraph g with (6) In this process, first, as shown in step S133 of FIG. 7, all nodes x 1 ~x 6 List the first-level parent nodes for .
[0120] Specifically, as can be seen from graph G11 in FIG. 8(a), in graph structure g, node x 1 ,x 2 ,x 3 ,x 4 ,x 5 ,x 6 The set of first-level parent nodes directly connected to is
number
number
[0121] Similarly, the set of edges for the entire graph structure g is
number
[0122] Therefore, as illustrated in step S135 of FIG. 7, for each level s ∈ {1, 2, 3, 4, 5, 6}, the set H 6 Each element x that composes (s) m x 6 A set of edges L that are passed through when 6 Recursively extracting (s),
number
[0123] Thus, the maximum subgraph g (6) teeth,
[0124]
number
[0125] Finally, the maximum subgraph g (6) When the above is displayed on the display 11, the content corresponding to the graph G17 in FIG. 8(c) is displayed. At that time, the CPU 3 calculates six discrete variables x 1 ~x 6 Among them, each variable x 1 ~x 5 For x as the objective variable y 6 The value of the number of layers s for is stored in the RAM 7 or SSD 9.
[0126] In graph G17, x 2 and x 3 is the x as the objective variable y 6 is the first-level parent node for x 4 is the x as the objective variable y 6 In other words, the objective variable y = x 6 In this case, x 2 and x 3 The number of levels is 1 (s=1), and x 4 The number of layers is 2 (s=2). Others, x 1 and x 5 is the x as the objective variable y 6 can be considered as a variable that has no dependency on
[0127] As described above, the graph structure analysis method is configured such that the number of levels corresponding to each variable x can be naturally obtained in its post-processing. The analysis method shown in FIG. 3 utilizes such a number of levels.
[0128] Hereinafter, returning to the flow of FIG. 3, each process constituting the same flow will be sequentially described based on the above description of the GSA.
[0129] <4. Details of the analysis method> (4-1. Graph structure determination process) First, in step S1, the CPU 3 executes a graph structure determination process based on the plurality of variables 49 shown in FIG. 2. By executing the graph structure process, the CPU 3 determines a directed graph structure composed of nodes corresponding to each of the plurality of variables 49 and hierarchically arranged to reflect the relationships between the variables 49. At this time, the directed graph structure is determined for each acquisition condition of the plurality of variables 49.
[0130] Specifically, the CPU 3 according to the present embodiment executes the GSA illustrated in FIGS. 4 to 8 when determining the graph structure. By executing the GSA, a directed acyclic graph structure indicating a Bayesian network is determined as the directed graph structure.
[0131] More specifically, the CPU 3 executes each process constituting the GSA for each of the plurality of variables 491 to 49N having different acquisition conditions. That is, the CPU 3 individually executes a preprocessing process, a main processing process, and a post-processing process for each of the N acquisition conditions.
[0132] As a result, in the preprocessing process, the CPU 3 generates, for each acquisition condition, for example, N categories of data based on the time-series data 49 acquired for each predetermined condition. In the subsequent main processing process, the CPU 3 determines a directed acyclic graph structure as the directed graph structure for each acquisition condition.
[0133] Then, in the post-processing process that is performed after the directed graph structure is determined, the CPU 3 sequentially extracts combinations of leaf nodes and parent nodes in ascending order of the number of layers.
[0134] Specifically, the CPU 3 determines the graph structure g and the number of levels s for each acquisition condition with the same selection of the objective variable y. The shape of the graph structure g and the size of the number of levels s may vary for each acquisition condition. The graph structure g determined by the CPU 3 is stored in the RAM 7 or SSD 9 as the first data set 59 shown in Fig. 2, while the number of levels s determined by the CPU 3 is stored in the RAM 7 or SSD 9 as the second data set 69 illustrated in the same figure.
[0135] In order to illustrate the changes for each acquisition condition, specific examples of the multiple variables 49 and multiple acquisition conditions are presented. Fig. 9 is a diagram showing specific examples of the multiple variables 49 obtained for each acquisition condition.
[0136] 9, multiple variables 49 each indicate a variable (particularly, time series data of each variable) that characterizes the traveling of the vehicle. Specifically, in this example, a total of 29 types of variables (i.e., p=29) are used, and the steering angle of the vehicle is used as the objective variable.
[0137] The remaining 28 explanatory variables are the vehicle altitude, the vehicle's distance traveled, the vehicle's vertical relative speed with respect to the ground (vertical speed_ground), the vehicle's longitudinal relative speed with respect to the ground (longitudinal speed_ground), the vehicle's lateral relative speed with respect to the ground (lateral speed_ground), the vehicle's longitudinal acceleration (longitudinal acceleration_vehicle), the vehicle's lateral acceleration (lateral acceleration_vehicle), the vehicle's yaw angle (yaw angle_vehicle), the vehicle's pitch angle (pitch angle_vehicle), the vehicle's roll angle (roll angle_vehicle), and the vehicle's sideslip angle (sideslip angle_vehicle). , roll rate of the vehicle body (Roll Rate_Vehicle Body), pitch rate of the vehicle body (Pitch Rate_Vehicle Body), yaw rate of the vehicle body (Yaw Rate_Vehicle Body), steering torque, X coordinate of the vehicle body relative to the centerline of the road (Centerline X Coordinate), Y coordinate of the vehicle body relative to the centerline of the road (Centerline Y Coordinate), distance of the vehicle body from the centerline (Centerline Distance), inclination angle of the vehicle path relative to the centerline (Centerline Path Angle), and curvature of the centerline on the road (Centerline Road Curvature).
[0138] In the following explanation, all 28 explanatory variables are distinguished by the variable numbers shown in the second column of Fig. 9. For example, "variable N1" indicates the variable with "variable number = 1", that is, the vehicle altitude.
[0139] On the other hand, Fig. 10 is a diagram showing a specific example of a plurality of acquisition conditions. The acquisition conditions may be conditions set to distinguish at least one of the vehicle type, the driving speed, and the driving proficiency. In the illustrated example, a total of 27 types of acquisition conditions are used, which are set to distinguish each of the vehicle type, the driving speed, and the driving proficiency (see the second column in Fig. 10).
[0140] For example, the vehicle types shown in the third column of Fig. 10 are vehicle types A, B, and C, and the vehicle running speeds shown in the fourth column of the same figure are composed of a first speed on the low side and a second speed on the high side. Here, the specific values of the first speed and the second speed are common to each vehicle type in this example for simplicity. More generally, it is not essential that the first speed and the second speed are common to each vehicle type.
[0141] Moreover, the fifth column in Figure 10 shows the driving proficiency (proficiency level) for the corresponding vehicle type and driving speed. The lower the proficiency level value, the earlier the data was measured (i.e., the driver is not accustomed to driving that vehicle type and has low driving proficiency). Moreover, the "integration" in the fifth column shows the integration of the explanatory variables of all proficiency levels for the corresponding vehicle type and driving speed in the time direction.
[0142] In the following explanation, all 27 types of acquisition conditions are represented by the condition IDs shown in the sixth column of Fig. 10. For example, "Condition B11" indicates the acquisition condition in the ninth row, that is, the explanatory variable at low proficiency level (proficiency level 1) acquired by driving vehicle type B at the first speed.
[0143] Figures 11A, 11B, 11C, and 11D are diagrams showing specific examples of hierarchical directed graph structures. Figure 11A shows a graph G21 obtained under condition A11, and Figure 11B shows a graph G22 obtained under condition A12. Similarly, Figure 11C shows a graph G23 obtained under condition A13, and Figure 11D shows a graph G24 obtained under condition A14.
[0144] In these graphs, edges connecting nodes with the same number of hierarchies, edges extending across multiple hierarchies, etc. are omitted in order to focus on the dependency with the variable N0 as the objective variable y.
[0145] Nodes surrounded by thick frames in Figures 11A to 11D correspond to variables related to the center line of the road, as shown by variables N24 to N28 in Figure 9. As the driving proficiency level increases, the dependency of these nodes on the steering angle shifts from a more direct dependency (low number of layers) to a more indirect dependency (high number of layers). This can be interpreted as suggesting that various factors related to the center line gradually mature from conscious to unconscious in determining the steering angle, that is, it becomes possible to deal with them unconsciously without being strongly conscious of the center line.
[0146] (4-2. Hierarchy number acquisition process) In the next step S2, the CPU 3 executes a hierarchical number acquisition process based on the first data set 59 shown in Fig. 2. By executing the hierarchical number acquisition process, the CPU 3 acquires the number of hierarchical levels for each acquisition condition and explanatory variable, assuming that one of the multiple variables 49 is a target variable and the rest are explanatory variables, and the number of hierarchical levels is the number of edges from a parent node corresponding to the explanatory variable to a leaf node corresponding to the target variable.
[0147] Here, when GSA is used in step S1, the number of hierarchical levels is necessarily calculated in the post-processing process when constructing the graph G21, as explained with reference to FIG.
[0148] Therefore, in the case of this embodiment, the post-processing process constituting the GSA essentially functions as a layer number acquisition process.
[0149] In this case, the CPU 3 reads in this step S2 the third data set 79 that has been configured by storing the calculation results of the post-processing in the RAM 7 or the SDD in step S1.
[0150] For example, in the case of FIG. 11A, the number of levels of the variable N28 is 1 (s=1). In addition, in the case of FIG. 11B, the number of levels of the same variable N28 is 2 (s=2), and in the case of FIG. 11C, the number of levels of the variable N28 is 3. On the other hand, in the case of FIG. 11D, the variable N28 is not linked to the variable N0 as the objective variable. In this case, the variable N28 can be regarded as a variable that does not have a dependent relationship with the objective variable.
[0151] (4-3. Score calculation process) In the next step S3, the CPU 3 executes a score calculation process based on the second data set 69 shown in Fig. 2. By executing the score calculation process, the CPU 3 calculates a score that is set to increase or decrease according to the number of layers for each acquisition condition and explanatory variable. When calculating the score, the CPU 3 sets the score corresponding to a node that is separated from the objective variable as a leaf node among the nodes that constitute the directed graph structure as a lower limit value, and calculates the score so that it increases relative to the lower limit value as the number of layers decreases.
[0152] Here, a node separated from the objective variable as a leaf node refers to a variable that does not have a direct or indirect dependency on the objective variable, such as the variable N28 in Fig. 11D. For such explanatory variables, the CPU 3 sets the score to a lower limit value (for example, zero).
[0153] Then, the CPU 3 increases the score as the number of layers becomes smaller, that is, as the dependency with the objective variable becomes more direct. The score can be calculated using any function that has the number of layers as an argument and satisfies such an increasing tendency.
[0154] For example, in this embodiment, the number of hierarchical levels acquired for each acquisition condition and explanatory variable is denoted by s as described above, and the maximum number of hierarchical levels when all acquisition conditions are considered is denoted by s. max Let s and s max If the score calculated based on is denoted as Sc, the CPU 3 can calculate the score based on the following formula (34).
[0155]
number
[0156] For example, for an explanatory variable (s=1) that is directly dependent on the objective variable, the score is necessarily 100. max = 5, and for explanatory variables separated from the objective variable as described above, Sc = 0, making it possible to calculate scores for all acquisition conditions and all explanatory variables.
[0157] For example, in the case of Fig. 11A, the score of the variable N28 is the upper limit value "Sc = 100". In addition, in the case of Fig. 11B, the score of the same variable N28 is a value "Sc = 80" that is slightly smaller than the upper limit, and in the case of Fig. 11C, the score of the variable N28 is an even smaller value "Sc = 60". On the other hand, in the case of Fig. 11D, the variable N28 is not connected to the variable N0 as the objective variable, so its score is the lower limit value "Sc = 0".
[0158] In this way, the changes in the graph structure brought about by changes in the acquisition conditions can be quantified by the score. By visualizing the progress of the scores for each explanatory variable, it becomes possible to discover new knowledge about the relationship between the acquisition conditions and each explanatory variable.
[0159] The scores calculated by the CPU 3 are stored in the RAM 7 or the SSD 9 as the third data set 79 shown in FIG. 2 for each acquisition condition and explanatory variable.
[0160] In addition, a correlation coefficient between explanatory variables 49 may be calculated, and the score value corresponding to each explanatory variable 49 may be adjusted according to the calculation result.
[0161] (4-4. Correspondence analysis process) In the following step S4, the CPU 3 executes a correspondence analysis process based on the third data set 79 shown in Fig. 2. By executing the correspondence analysis process, the CPU 3 performs cross-tabulation of scores and correspondence analysis on the cross-tabulated scores.
[0162] First, the CPU 3 cross-tabulates the scores obtained for the acquisition conditions and explanatory variables so as to construct a contingency table with one of the acquisition conditions and explanatory variables as the table side and the other as the table head. A specific example of cross-tabulation is as shown in FIG. 12. In this example, a contingency table is constructed with the acquisition conditions as the table side and the explanatory variables as the table head. Each score calculated based on the above formula (34) is displayed in each cell of this contingency table.
[0163] Next, CPU3 performs correspondence analysis on the cross-tabulated scores to visualize the relationship between the acquisition conditions and the explanatory variables by mapping them on a two-dimensional plane. CPU3 executes correspondence analysis by computing a constrained optimization problem configured to maximize the correlation coefficient between the items on the table side (row categories) and the table top (column categories). A specific example of mapping is shown in FIG. 13. In FIG. 13, the plots shown as squares correspond to the acquisition conditions, and the plots shown as diamonds correspond to the explanatory variables.
[0164] Focusing on the acquisition conditions in Fig. 13, there are many plots on the +X side corresponding to the first speed on the low side, and many plots on the -X side corresponding to the second speed on the high side. Therefore, tentatively, it can be considered that "the X-axis distinguishes between high and low vehicle speeds under the acquisition conditions."
[0165] Similarly, there are many plots on the +Y side that correspond to vehicle type C, and many plots on the -Y side that correspond to vehicle types A and B. Therefore, a tentative interpretation can be that "the Y axis distinguishes between vehicle types under the acquisition conditions."
[0166] On the other hand, looking at the explanatory variables in Figure 13, the variable N5 corresponding to "lateral speed_ground" is located on the X-axis. Therefore, as a tentative interpretation, it can be considered that "the X-axis is related to the magnitude of the lateral movement." Similarly, the Y-axis can be interpreted as "the +Y side is related to the path angle information, and the -Y side is related to the lateral position information."
[0167] Furthermore, in the first quadrant, since the acquisition conditions corresponding to vehicle type A and vehicle type C at the first speed in particular are close to the explanatory variables related to the roll angle and yaw angle, it is possible to give the interpretation that "when driving vehicle type A and vehicle type C at the first speed in particular, the roll angle and yaw angle affect the steering angle."
[0168] In addition, in the second quadrant, it can be interpreted that "at the low level of proficiency, while all three vehicle models are conscious of the distance and angle from the center line, lateral movement affects the steering angle."
[0169] In addition, in the third quadrant, it can be interpreted that "at the high level of proficiency, for both vehicle types A and B, when the lateral speed is relatively high, awareness of the center line affects the steering angle."
[0170] In the fourth quadrant, the plot corresponding to vehicle type B in particular is far from the plot corresponding to "lateral speed_ground" (variable N5). On the other hand, the plot corresponding to vehicle type B in particular is relatively close to the plot corresponding to "longitudinal speed_vehicle body" (variable N4). Therefore, it can be interpreted that "in the case of vehicle type B, longitudinal movement has a stronger effect on steering than the correlation with lateral movement."
[0171] Also, when focusing on the first and fourth quadrants, it can be seen that the plot corresponding to vehicle type B is farther away from the plot corresponding to "roll angle_body" (variable N11) than the plots corresponding to vehicle types A and C. This suggests that "when driving vehicle type B, it is possible to drive without being aware of the roll angle."
[0172] Also, it can be seen that the plot corresponding to vehicle type B is farther away from the plot corresponding to "lateral speed_ground" (variable N5) than the plots corresponding to vehicle types A and C. This suggests that "when driving vehicle type B, it is possible to drive without being strongly conscious of lateral motion."
[0173] Furthermore, the plot corresponding to "longitudinal acceleration_vehicle body" (variable N6) is far from the plots corresponding to any of the acquisition conditions. This tendency suggests the interpretation that "the longitudinal acceleration of the vehicle body does not affect the determination of the steering angle as compared to other explanatory variables." In fact, in any of Figures 11A to 11D, variable N6 does not have a dependency relationship with the objective variable (variable N0). In other words, the structures of graphs G21 to G24 support the interpretation of two-dimensional plots.
[0174] Taking these points into account, we can further interpret Vehicle type C: Mainly angular deviation, and precise position feedback is not possible for lateral deviation Vehicle type B: Mainly lateral deviation, but position feedback is also possible for angular deviation Model A: Between model C and model B, feedback is not working well This is a possibility.
[0175] <5. Significance of the analysis method> According to the embodiment, as illustrated in Fig. 12, a score that increases or decreases according to the number of hierarchical levels is calculated for each acquisition condition (see the top of the table in Fig. 12) and explanatory variable (see the bottom of the table in Fig. 12). By calculating such a score, it is possible to quantify, for each explanatory variable, the change in the graph structure that occurs when the acquisition condition is changed. This quantification makes it possible to systematically analyze the association between the graph structure and the acquisition condition.
[0176] Furthermore, quantification such as a score is suitable for calculation by the computer 1. This makes it possible to realize a more appropriate analysis that is free from the analyst's preconceptions.
[0177] In addition, simply visualizing the directed graph structures determined for each acquisition condition individually can make it difficult to make a comprehensive comparison between the structures. In contrast, eliminating preconceptions as described above promotes comprehensive understanding by the analyst, which can lead to intellectual discovery and support for idea generation.
[0178] Furthermore, by cross-tabulating the scores as shown in the example in Figure 12, the scores can be compiled by acquisition conditions and explanatory variables, enabling a more systematic analysis that is free from the analyst's preconceptions.
[0179] In addition, the relationship between the acquisition conditions and the explanatory variables is visualized as a relative distance on a two-dimensional plane, as shown in Fig. 13. This makes it possible to visualize acquisition conditions that strongly affect the relationship between an objective variable (for example, a variable N0 corresponding to a steering angle) when the objective variable is given.
[0180] Furthermore, by mapping the relationship between the acquisition conditions and explanatory variables on a two-dimensional plane, when multiple acquisition conditions are set for multiple explanatory variables as shown in Figures 9 and 10, the relationships can be visualized from a bird's-eye view. The directed graph structure and the relationship between the acquisition conditions of each variable can be analyzed from a bird's-eye view.
[0181] As described with reference to Fig. 6 and Fig. 8, the use of a Bayesian network in a directed graph structure makes it possible to visualize the dependency relationship between explanatory variables and a target variable. As shown in Figs. 11A to 11D, the influence of changes in acquisition conditions on such dependency relationships can be systematically analyzed. This contributes to the analysis of various phenomena, such as engineering phenomena.
[0182] Furthermore, as shown in equations (3) and (4), by discretizing each time series data 49 into a multi-level system in advance, it is possible to construct a Bayesian network even for time series data that changes continuously over time.
[0183] Moreover, by discretizing each time series data 49 not with respect to time but with respect to its own value, a categorical data set that includes a temporal relationship can be generated. This makes it possible to generate a graph structure that includes temporal dependencies such as causal relationships. By generating such a graph structure, even if the temporal dependency between categorical data sets is not explicitly stated, a graph structure that suggests the dependency can be obtained. As a result, the temporal dependency can be traced back from the visualized graph structure, which in turn makes it possible to verify previously known hypotheses or create new hypotheses.
[0184] In addition, as explained with reference to Fig. 8, by extracting dependencies between nodes for each number of layers, it is possible to extract a graph structure that leads to a target variable as a leaf node without omission or excess or deficiency. At the same time, since the number of layers corresponding to each combination is naturally given during the extraction, it is possible to efficiently calculate the score while minimizing the amount of calculation. This allows the efficiency of calculations by the computer 1 to be improved.
[0185] 13, the analysis method according to the embodiment can systematically analyze the influence of the vehicle type, driving speed, etc. on driving. This contributes to the improvement of automobile models and automobile controls, for example.
[0186] In addition, formulating the score, for example as in equation (34), is useful for systematic analysis.
[0187] <6. Other application examples> In the embodiment, the explanatory variables and objective variables shown in FIG. 9 and the acquisition conditions shown in FIG. 10 are given as specific examples, but application examples of the analysis method are not limited to such specific examples.
[0188] (First variant: aerodynamic performance of the vehicle) The present disclosure can be applied to the verification of aerodynamic performance. In this case, an index for distinguishing vehicle models can be used as an acquisition condition, and pressure time series data acquired at the same measurement position for each vehicle model can be used as an explanatory variable. Then, by using the air resistance coefficient (Cd value) as the objective variable, it is possible to assist in the search for vehicle models with relatively superior aerodynamic performance.
[0189] (Second variant: Marketing field) This disclosure can be applied to the marketing field. In this case, an index that distinguishes brands can be used as an acquisition condition, and time-dependent changes in factors such as purchasing behavior and preferences by brand can be used as explanatory variables. Then, by using the conversion achievement rate of customers as the objective variable, it is possible to support the discovery of hypotheses about factors that are strengths for each brand.
[0190] (Third variant: Healthcare field) The present disclosure can be applied to the field of healthcare. In this case, an index for distinguishing subjects or scenes can be used as an acquisition condition, and sensing data for each body part can be used as an explanatory variable. Then, by using an index indicating the subject's health condition as the objective variable, it is possible to assist in the search for which body part is strongly associated with the health condition for each subject or scene.
[0191] (Fourth variant: connected field) The present disclosure can be applied to the connected field. In this case, an index for distinguishing vehicles can be used as an acquisition condition, and data transmitted and received via CAN and GPS can be used as an explanatory variable. Then, by using a variable indicating the state of the vehicle or the driver as the objective variable, it is possible to assist in searching for which data is strongly linked to the state of the vehicle or the driver for each vehicle, and further, which vehicle is in an abnormal state.
[0192] (Fifth Variation: Personal Characteristics) The present disclosure can be applied to individual characteristics. In this case, an index that distinguishes the scene in which the subject is placed can be used as the acquisition condition, and the physiological data of the subject can be used as the explanatory variable. Then, by using an index that indicates the behavior or emotion of the subject as the objective variable, it is possible to support the verification of the connection between each scene and the physiological data, and ultimately which scene is most suitable for the subject.
[0193] <7. Other embodiments> In the above embodiment, a configuration implemented by one computer 1 is exemplified, but the present disclosure is not limited to this example. The analysis method and analysis program 29 according to the present disclosure may be executed using multiple computers 1, such as by having a first computer execute the process related to GSA, while having a second computer execute the process related to cross tabulation and correspondence analysis. In addition, the computer 1 in the present disclosure also includes parallel computers such as supercomputers and PC clusters.
[0194] <Industrial Applicability> As described above, the present disclosure is useful for analyzing dependencies in various fields such as automotive engineering, marketing, and healthcare, as well as for the calculations of computers themselves, and has industrial applicability. [Explanation of symbols]
[0195] 1. Computer 3 CPU (arithmetic unit) 7 RAM (memory section) 9 SSD (storage unit) 18 Storage medium 29 Analysis Program 291 Graph Structure Analysis Program 49 Variables S1 Graph structure decision process S2 Layer number acquisition process S3 score calculation process S4 Correspondence Analysis Process
Claims
1. 1. An analysis method for generating a graph based on a plurality of variables and calculating a score that characterizes the graph by using a computer including a storage unit that stores a program and a calculation unit that executes the program stored in the storage unit, comprising: determining, by the calculation unit, a directed graph structure configured by nodes corresponding to each of the plurality of variables and hierarchically arranged so as to reflect relationships between the variables, for each acquisition condition of the plurality of variables; a step of the calculation unit acquiring the number of levels for each of the acquisition conditions and the explanatory variables, when one of the plurality of variables is designated as an objective variable and the remaining variables are designated as explanatory variables, a node that constitutes the directed graph structure and corresponds to the objective variable is designated as a leaf node, a node that is connected to the leaf node via one or more edges in the directed graph structure and corresponds to the explanatory variable is designated as a parent node, and the number of edges passing from the parent node to the leaf node is designated as a number of levels; The calculation unit calculates a score set to increase or decrease according to the number of hierarchical levels for each of the acquisition conditions and the explanatory variables.
13. An analytical method comprising:
2. 2. The method of claim 1, The calculation unit cross-tabulates the scores obtained for the acquisition conditions and the explanatory variables so as to construct a contingency table in which one of the acquisition conditions and the explanatory variables is the table side and the other is the table head.
13. An analytical method comprising:
3. The method of claim 2, The calculation unit performs a correspondence analysis on the cross-tabulated scores to map the relationship between the acquisition conditions and the explanatory variables onto a two-dimensional plane.
13. An analytical method comprising:
4. 2. The method of claim 1, The calculation unit determines a directed acyclic graph structure representing a Bayesian network as the directed graph structure.
13. An analytical method comprising:
5. The method of analysis according to claim 4, The plurality of variables are a plurality of time series data generated by acquiring predetermined parameters over a predetermined period of time, The calculation unit generates a category data set for each of the plurality of time series data by discretizing a value at each time of each of the plurality of time series data into a multi-level system; When determining the directed graph structure, a set of directed acyclic graph structures representing a Bayesian network in which each of the plurality of category data sets is a node is defined as G, and a graph structure forming an element of the set G is defined as g, the calculation unit determines a graph structure that maximizes a conditional probability that the plurality of category data sets are realized when the element g is given.
13. An analytical method comprising:
6. 2. The method of claim 1, After the directed graph structure is determined, the calculation unit sequentially extracts combinations of the leaf nodes and the parent nodes in ascending order of the number of layers.
13. An analytical method comprising:
7. 2. The method of claim 1, each of the plurality of variables indicates a variable that characterizes a running of the vehicle; The acquisition conditions indicate conditions set to distinguish at least one of the vehicle type, driving skill level, and driving speed.
13. An analytical method comprising:
8. 2. The method of claim 1, The calculation unit sets the score corresponding to a node that is separated from the leaf node among the nodes that configure the directed graph structure as a lower limit value, and calculates the score so that the score increases with respect to the lower limit value as the number of layers decreases.
13. An analytical method comprising:
9. An analysis device that is configured by a computer having a storage unit that stores a program and a calculation unit that executes the program stored in the storage unit, generates a graph based on a plurality of variables, and calculates a score that characterizes the graph, a graph structure determination means for determining a directed graph structure, which is constituted by nodes corresponding to the plurality of variables, and which is hierarchically arranged so as to reflect relationships between the variables, for each acquisition condition of the plurality of variables; a hierarchical level acquisition means for acquiring the number of hierarchical levels for each of the acquisition conditions and the explanatory variables, when one of the plurality of variables is a target variable and the remaining variables are explanatory variables, a node that constitutes the directed graph structure and corresponds to the target variable is a leaf node, a node that is connected to the leaf node in the directed graph structure via one or more edges and corresponds to the explanatory variable is a parent node, and the number of edges passing from the parent node to the leaf node is the number of hierarchical levels; a score calculation means for calculating a score set to increase or decrease according to the number of hierarchical levels for each of the acquisition conditions and the explanatory variables; 1. An analytical device comprising:
10. An analysis program for generating a graph based on a plurality of variables and calculating a score that characterizes the graph by being executed by a computer including a storage unit that stores a program and a calculation unit that executes the program stored in the storage unit, the analysis program comprising: The computer includes: determining, by the calculation unit, a directed graph structure configured by nodes corresponding to each of the plurality of variables and hierarchically arranged so as to reflect relationships between the variables, for each acquisition condition of the plurality of variables; a step of the calculation unit acquiring the number of levels for each of the acquisition conditions and the explanatory variables, when one of the plurality of variables is designated as an objective variable and the remaining variables are designated as explanatory variables, a node that constitutes the directed graph structure and corresponds to the objective variable is designated as a leaf node, a node that is connected to the leaf node in the directed graph structure via one or more edges and corresponds to the explanatory variable is designated as a parent node, and the number of edges passing from the parent node to the leaf node is designated as a number of levels; the calculation unit calculates a score for each of the acquisition conditions and the explanatory variables, the score being set to increase or decrease according to the number of hierarchical levels. An analysis program characterized by:
11. The analysis program according to claim 10 is stored. A computer-readable storage medium comprising:
Citation Information
Patent Citations
Graph-structured analysis method, graph-structured analysis program, and computer-readable storage medium that stores the graph-structured analysis program
JP2021111063A