Structural equation modeling support device, structural equation modeling support method, and structural equation modeling support program
The structural equation modeling support device aids in constructing path diagrams by grouping nodes based on predefined rules, enhancing the accuracy and validity of structural equation models through Bayesian networks and LiNGAM methods, addressing the challenge of latent variable determination and model construction.
Patent Information
- Application Number
- JP2023220926
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-07-09
AI Technical Summary
In structural equation modeling, users face challenges in determining which factors to set as latent variables and constructing an appropriate path diagram, as existing methods do not provide adequate support for creating such models.
A structural equation modeling support device that includes a rule information storage unit, a group processing unit, and an output unit to assist in creating a path diagram by grouping nodes based on predefined rules, such as similarity of node texts and edge connections, and generating graphs using Bayesian networks or LiNGAM methods.
Facilitates the creation of path diagrams by grouping nodes effectively, allowing users to construct more accurate and valid structural equation models through the use of predefined rules and multiple graph generation methods, including Bayesian networks and LiNGAM, even with continuous or discrete data values.
Smart Images

Figure 2025103497000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a structural equation modeling support device, a structural equation modeling support method, and a structural equation modeling support program for assisting in creating a path diagram in structural equation modeling.
Background Art
[0002] Structural Equation Modeling (SEM), also called Covariance Structure Analysis (CSA), is a method of modeling and analyzing the relationships between a large number of variables set as hypotheses using a linear combination equation that can be represented by a so-called path diagram of a directed graph. This structural equation modeling is a useful method for verifying the validity of hypotheses, and has advantages such as being able to analyze including factors that cannot be directly observed, called latent variables, and being able to set multiple dependent variables in one analysis. Such structural equation modeling is disclosed, for example, in Patent Document 1.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, in structural equation modeling, although latent variables can be incorporated into the model, it is difficult for the user (analyst) who generates the model (path diagram, structural equation) to determine what factors should be set as latent variables and what model should be set to generate the path diagram (structural equation) of the directed graph representing the model.
[0005] The present invention is an invention made in view of the above circumstances, and an object thereof is to provide a structural equation modeling support device, a structural equation modeling support method, and a structural equation modeling support program that can assist in creating a path diagram in structural equation modeling.
Means for Solving the Problems
[0006] As a result of various studies, the present inventor has found that the above object is achieved by the following present invention. That is, a structural equation modeling support device according to an aspect of the present invention is a device that supports the creation of a path diagram in structural equation modeling, and includes a rule information storage unit that stores a grouping rule, which is a rule for grouping a plurality of nodes in a graph into one group, a group processing unit that extracts nodes to be grouped from a predetermined graph based on the grouping rule stored in the rule information storage unit, and an output unit that outputs the extraction result extracted by the group processing unit. Preferably, in the above-described structural equation modeling support device, the group processing unit randomly selects one node as a starting node from among the unprocessed nodes among the plurality of nodes in the graph, and repeatedly executes a grouping process of extracting nodes to be grouped based on the grouping rule stored in the rule information storage unit for the selected starting node until there are no unprocessed nodes.
[0007] Such a structural equation modeling support device outputs a plurality of nodes that are grouped into one as an extraction result. Therefore, a user (analyst) who creates a model (path diagram, structural equation) can refer to the extraction result and assign one new node to the plurality of nodes grouped into one to create a new graph that is the basis of the path diagram from a predetermined graph, thereby assisting in creating a path diagram in structural equation modeling.
[0008] In another aspect, in the above-described structural equation modeling support device, the grouping rule includes a rule for grouping, into one group, nodes up to a predetermined number that are sequentially connected via edges from a starting node.
[0009] Such a structural equation modeling support device outputs, as an extraction result, nodes up to a predetermined number that are sequentially connected via edges from a starting node. Thus, by referring to the extraction result, the user can create a new graph in which nodes up to a predetermined number that are sequentially connected via edges from a starting node in a predetermined graph are grouped into one group.
[0010] In another aspect, in the above-described structural equation modeling support device, the grouping rule further includes a rule for selecting a predetermined number of edges from among the plurality of edges when one node has a plurality of edges.
[0011] Such a structural equation modeling support device further outputs, as an extraction result, nodes connected to a predetermined number of edges selected from among the plurality of edges when one node has a plurality of edges. Thus, by referring to the extraction result, the user can create a new graph in which nodes connected to a predetermined number of edges selected from among the plurality of edges when one node has a plurality of edges in a predetermined graph are grouped into one group.
[0012] In another aspect, in the above-described structural equation modeling support device, each of the plurality of nodes in the graph has a sentence, and the grouping rule further includes a rule for grouping, into one group, two nodes connected to both ends of one edge when the sentences provided in the two nodes are similar to each other.
[0013] Such a structural equation modeling support device further outputs the two nodes as extraction results when each of the texts provided for each of the two nodes connected to both ends of one edge is similar to each other. Therefore, by referring to the extraction results, the user can create a new graph from a predetermined graph, taking into account the similarity of the texts provided for each of the two nodes.
[0014] In another aspect, in the above-described structural equation modeling support device, each of the plurality of nodes in the graph has a text, and the grouping rule is that when each of the texts provided for the starting node and each of the destination nodes connected to the starting node via an edge is similar to each other, the starting node and the destination nodes are grouped together into one, and the similar grouping process is repeatedly executed until grouping is no longer possible, with the destination node as a new starting node.
[0015] Such a structural equation modeling support device outputs the starting node and the destination nodes connected to the starting node via an edge as extraction results when each of the texts provided for each of the starting node and the destination nodes is similar to each other. Therefore, by referring to the extraction results, the user can create a new graph from a predetermined graph, taking into account the similarity of the texts provided for each of the two nodes.
[0016] In another aspect, in the above-described structural equation modeling support device, each of the plurality of nodes in the graph has a text, and the grouping rule includes a rule for extracting a node having a text similar to the text provided for the starting node from among the plurality of nodes in the graph, and performing a similar grouping process of grouping the starting node and the extracted node together into one.
[0017] Such a structural equation modeling support device outputs, as an extraction result, a plurality of nodes each having a sentence similar to another. Therefore, by referring to the extraction result, the user can create a new graph from a predetermined graph, taking into account the similarity of the sentences included in the nodes.
[0018] In another aspect, the above-described structural equation modeling support device further includes a fitness index processing unit that obtains a fitness index of a path diagram created from the graph based on the extraction result extracted by the group processing unit, and the output unit further outputs the fitness index of the path diagram obtained by the fitness index processing unit.
[0019] Such a structural equation modeling support device obtains and outputs a fitness index of a path diagram. Therefore, the user can recognize the fitness of the path diagram created by referring to the extraction result, and can recreate the path diagram or cause the structural equation modeling support device to re-extract according to the fitness of the path diagram.
[0020] In another aspect, the above-described structural equation modeling support device further includes a graph generation unit that generates the graph by a plurality of generation methods, and an input unit that receives any one of the plurality of generation methods, and the graph generation unit generates the graph by the generation method received by the input unit among the plurality of generation methods.
[0021] Such a structural equation modeling support device can change the graph generation method according to the fitness of the path diagram.
[0022] In another aspect, the above-described structural equation modeling support device further includes a graph generation unit that generates the graph by a Bayesian network generation method.
[0023] According to this, a structural equation modeling support device can be provided that further includes a graph generation unit that generates the graph by a Bayesian network generation method.
[0024] In another aspect, in the above-described structural equation modeling support device, for each of a plurality of different variables, from first data having a plurality of data values, for each of the plurality of data values, by discretizing the data value into a multi-level system, a discretization processing unit that generates second data having a plurality of the discretized data values for each of the plurality of variables is further provided, the graph generation unit generates the graph from the second data generated by the discretization processing unit, and the output unit further outputs the second data generated by the discretization processing unit.
[0025] Since such a structural equation modeling support device discretizes data values into a multi-level system, even if the data values of the first data are continuous values, a graph (Bayesian network) can be generated by the Bayesian network generation method.
[0026] In another aspect, in the above-described structural equation modeling support device, for each of a plurality of different variables, from data having a plurality of data values, a graph generation unit that generates the graph, and a normal distribution determination unit that determines whether the data is a normal distribution are further provided, and when the determination result of the normal distribution determination unit is a normal distribution, the graph generation unit generates the graph by the Bayesian network generation method, and when the determination result of the normal distribution determination unit is not a normal distribution, the graph generation unit generates the graph by LiNGAM.
[0027] Such a structural equation modeling support device can select a suitable graph generation method according to whether it is a normal distribution, and suitable graph generation can be expected.
[0028] In another aspect, in the above-described structural equation modeling support device, for each of a plurality of different variables, from data having a plurality of data values, a graph generation unit that generates the graph is further provided, and the data values are discrete values or continuous values.
[0029] According to this, a structural equation modeling support device in which the data values are discrete values or continuous values can be provided.
[0030] A structural equation modeling support method according to another aspect of the present invention is a method for supporting the creation of a path diagram in structural equation modeling, and is based on a grouping rule, which is a rule stored in a rule information storage unit for grouping a plurality of nodes in a graph into one group, and extracts nodes to be grouped from a predetermined graph. A group processing step, and an output step of outputting the extraction result extracted in the group processing step.
[0031] Such a structural equation modeling support method outputs a plurality of nodes that are grouped into one as an extraction result. Therefore, the user who generates the model can refer to the extraction result and assign one new node to the plurality of nodes that are grouped into one, thereby creating a new graph that is the basis of the path diagram from a predetermined graph. Therefore, it is possible to support the creation of a path diagram in structural equation modeling.
[0032] A structural equation modeling support program according to another aspect of the present invention is a program for causing a computer to function as any one of the above-described structural equation modeling support devices.
[0033] According to this, a structural equation modeling support program can be provided, and this structural equation modeling support program has the same effects as the above-described structural equation modeling support devices.
Effects of the Invention
[0034] The structural equation modeling support device, the structural equation modeling support method, and the structural equation modeling support program according to the present invention can support the creation of a path diagram in structural equation modeling.
Brief Description of the Drawings
[0035]
Figure 1
Figure 2
Figure 3
Figure 4
Embodiments for Carrying Out the Invention
[0036] Hereinafter, one or more embodiments of the present invention will be described with reference to the drawings. However, the scope of the invention is not limited to the disclosed embodiments. In each figure, components with the same reference numerals indicate the same components, and the description thereof will be omitted as appropriate. In this specification, when referring to components generically, reference numerals without subscripts are used, and when referring to individual components, reference numerals with subscripts are used.
[0037] The structural equation modeling support device in the embodiment is a device that supports the creation of a path diagram in structural equation modeling (covariance structure analysis). This structural equation modeling support device includes a rule information storage unit that stores a grouping rule, which is a rule for grouping a plurality of nodes in a graph into one group, a group processing unit that extracts nodes to be grouped from a predetermined graph based on the grouping rule stored in the rule information storage unit, and an output unit that outputs the extraction result extracted by the group processing unit. Hereinafter, such a structural equation modeling support device, as well as the structural equation modeling support method and structural equation modeling support program implemented thereon, will be described more specifically.
[0038] FIG. 1 is a block diagram showing the configuration of the structural equation modeling support apparatus according to the embodiment. FIG. 2 is a diagram for explaining group processing in the structural equation modeling support apparatus. In FIG. 2, ○ represents a node, a directed line segment represents an edge, and the integer value within ○ is a node number (serial number), which is an identifier (ID) for specifying and identifying the node. FIG. 3 is a diagram showing a path diagram as an example. In FIG. 3, ○ represents a node, a directed line segment represents an edge, the integer value within ○ is the node number of the node, and the character within ○ is the variable of the node.
[0039] The structural equation modeling support apparatus 1000 according to the embodiment includes, for example, as shown in FIG. 1, a control processing unit 1, an input unit 2, an output unit 3, an interface unit (IF unit) 4, and a storage unit 5.
[0040] The input unit 2 is connected to the control processing unit 1 and is a device that inputs various data necessary for operating the structural equation modeling support apparatus 1000, such as various commands, for example, a command for instructing the start of processing, and data for generating a graph and grouping rules, etc., and is, for example, a keyboard, a mouse, and a plurality of input switches assigned with predetermined functions. The output unit 3 is connected to the control processing unit 1 and is a device that outputs commands and data input from the input unit 2, and processing results, etc., according to the control of the control processing unit 1, and is, for example, a display device such as a CRT display, an LCD (liquid crystal display device), and an organic EL display, and a printing device such as a printer.
[0041] Note that the input unit 2 and the output unit 3 may be composed of a touch panel. When configuring this touch panel, the input unit 2 is a position input device that detects and inputs an operation position, such as a resistive film method or a capacitance method, and the output unit 3 is a display device. In this touch panel, the position input device is provided on the display surface of the display device, and one or a plurality of input content candidates that can be input to the display device are displayed. When the user touches the display position where the input content to be input is displayed, the position is detected by the position input device, and the display content displayed at the detected position is input to the structural equation modeling support device 1000 as the user's operation input content. In such a touch panel, since the user can easily understand the input operation intuitively, a structural equation modeling support device 1000 that is easy for the user to handle is provided.
[0042] The IF unit 4 is connected to the control processing unit 1 and is a circuit that inputs and outputs data to and from external devices, for example, according to the control of the control processing unit 1. Examples include an RS-232C interface circuit using a serial communication method, an interface circuit using the Bluetooth (registered trademark) standard, and an interface circuit using the USB standard. Further, the IF unit 4 may be a communication interface circuit that transmits and receives communication signals to and from external devices, such as a data communication card or a communication interface circuit conforming to the IEEE802.11 standard.
[0043] The storage unit 5 is a circuit connected to the control processing unit 1 and stores various predetermined programs and various predetermined data according to the control of the control processing unit 1. The various predetermined programs include, for example, a control processing program, and the control processing program includes, for example, a control program, a discretization processing program, a graph creation program, a group processing program, and a fitness index processing program, etc. The control program is a program that controls each part 2-5 of the structural equation modeling support device 1000 according to the functions of the respective parts. The discretization processing program is a program that, for each of a plurality of M variables that are different from each other, from an M×N first data having a plurality of N data values, for each of the plurality of data values, by discretizing the data value into a multi-level system, generates second data having a plurality of the discretized data values for each of the plurality of variables. The graph creation program is a program that generates a graph from predetermined data. The group processing program is a program that extracts nodes that form a group from a predetermined graph based on the grouping rules stored in the storage unit 5. The fitness index processing program is a program that obtains the fitness index of a path diagram created based on the extraction result extracted by the group processing program from the graph. The various predetermined data includes, for example, data for generating a graph and data necessary for executing these programs, such as grouping rules.
[0044] Such a storage unit 5 includes, for example, a ROM (Read Only Memory) which is a non-volatile storage element, an EEPROM (Electrically Erasable Programmable Read Only Memory) which is a rewritable non-volatile storage element, etc. And the storage unit 5 includes a RAM (Random Access Memory) etc. which serves as a working memory of the so-called control processing unit 1 for storing data etc. generated during the execution of the predetermined program. Also, the storage unit 5 may be configured to include a hard disk device or a solid state drive (SSD) having a relatively large storage capacity.
[0045] The memory unit 5 functionally includes a rule information storage unit 51. The rule information storage unit 51 stores grouping rules. The grouping rules are rules for grouping a plurality of nodes in a graph into one. The grouping rules include, for example, a first rule for grouping nodes up to a predetermined number (first number, connection number) sequentially connected via edges from a starting node into one. The connection number is, for example, appropriately set in advance and stored in the rule information storage unit 51. Alternatively, for example, at the time of initialization, the connection number is randomly generated as one or more and stored in the rule information storage unit 51. The grouping rules include, for example, a second rule for selecting a predetermined number (second number, branching number) of edges from among the plurality of edges when one node has a plurality of edges, in addition to the first rule. The branching number is, for example, appropriately set in advance and stored in the rule information storage unit 51. Alternatively, for example, at the time of initialization described later, the branching number is randomly generated as one or more and stored in the rule information storage unit 51. When each of a plurality of nodes in a graph has a sentence, the grouping rules include, for example, a third rule for grouping two nodes connected to both ends of one edge into one when the sentences provided in each of the two nodes are similar to each other, with respect to at least one of the first and second rules. Therefore, when the grouping rules include the first and third rules, if the third rule is not satisfied before reaching a predetermined number from the starting node according to the first rule, the grouping is terminated up to the node at that time. Note that the second and third rules may be independent, respectively, but these are equivalent to the case where the connection number in the first rule is one in combination with the first rule.Alternatively, for example, when each of a plurality of nodes in a graph includes a sentence, the grouping rule includes a fourth rule of repeatedly executing similar grouping processing (first similar grouping processing) of grouping the starting node and the destination nodes connected to the starting node via edges into one group until grouping is no longer possible, where each sentence included in the starting node and each of the destination nodes is similar to each other. Alternatively, for example, when each of a plurality of nodes in a graph includes a sentence, the grouping rule includes a fifth rule of extracting a node including a sentence similar to the sentence included in the starting node from among the plurality of nodes in the graph and executing similar grouping processing (first similar grouping processing) of grouping the starting node and the extracted node into one group.
[0046] The control processing unit 1 is a circuit for assisting in creating a path diagram in structural equation modeling by grouping a plurality of nodes in a graph into one group and controlling each of the parts 2 to 5 of the structural equation modeling support device according to the functions of the respective parts. The control processing unit 1 is configured to include, for example, a CPU (Central Processing Unit) and its peripheral circuits. When the control processing program is executed, a control unit 11, a discretization processing unit 12, a graph generation unit 13, a group processing unit 14, and a goodness-of-fit index processing unit 15 are functionally configured in the control processing unit 1.
[0047] The control unit 11 controls each of the parts 2 to 5 of the structural equation modeling support device 1000 according to the functions of the respective parts and is in charge of the overall control of the structural equation modeling support device 1000.
[0048] The discretization processing unit 12 has a plurality of different variables x1, x2, ···, x M each having a plurality of N data values, an M×N first data x1(1), x1(2), ···, x1(N), x2(1), x2(2), ···, x2(N), x M(1), x M (2), ···, x M (N), for each of the plurality of data values, by discretizing the data value into a multi-level system, for each of the plurality of variables, second data including a plurality of the discretized data values is generated. The data values of the first data may be discrete values such as answers to a questionnaire expressed on a Likert scale, or may be continuous values such as measurement data, for example. For the plurality M of variables x1, x2, ···, x M the average value of the i-th variable x i in (1 ≤ i ≤ M) is taken as μ i and the standard deviation of this variable x i is taken as σ i and the discretization function for discretizing this variable x i is taken as ψ i and when the largest integer less than or equal to a is taken as [a], for the discretization, for example, the formula; ψ i (x) = [(x - μ i ) / σ i is used. That is, each data value x i of the variable x i (1), x i (2), ···, x i (N) is normalized (standardized) and integerized to be discretized into a multi-level system. By standardization, second data that absorbs outliers in the first data can be generated, and deterioration in the accuracy of structural equation modeling can be suppressed.
[0049] The graph generation unit 13 generates a graph from predetermined data. The graph generation unit 13 may generate the graph, for example, by a Bayesian network generation method of one generation method. However, in the present embodiment, the graph generation unit 13 generates the graph by a plurality of generation methods. More specifically, the input unit 2 receives any one of the plurality of generation methods, and the graph generation unit 13 generates the graph by the generation method received by the input unit 2 among the plurality of generation methods. The plurality of generation methods includes, for example, the Bayesian network generation method and LiNGAM (Linear Non-Gaussian Acyclic Model). The Bayesian network generation method is a known method for causal analysis suitable for data of a normal distribution and is disclosed in, for example, Japanese Patent Application Laid-Open No. 2021-111063 (Patent No. 7375556). In the Bayesian network generation method disclosed in Japanese Patent Application Laid-Open No. 2021-111063, the graph generation unit 13 may generate the graph from the first data before discretization by the discretization processing unit 12. However, in the present embodiment, the graph is generated from the second data discretized by the discretization processing unit 12. In this Bayesian network generation method, the graph generation unit 13, for example, for a plurality of M variables x1, x2,..., x in the second data ψ1(x1(1)), ψ1(x1(2)),..., ψ1(x1(N)), ψ2(x2(1)), ψ2(x2(2)),..., ψ2(x2(N)),..., ψ M (x M (1)), ψ M (x M (2)),..., ψ M (x M (N)) MLet \(G\) be a set of directed acyclic graph structures that represent Bayesian networks with each as a node, and let \(g\) be a graph structure that is an element of the set \(G\). When the given element is \(g\), determine a graph structure that maximizes the conditional probability that the plurality of category data sets are realized. In the graph structure, any node in the graph structure is a child node, and a node connected to the child node via one or more edges in the graph structure is a parent node. When the number of edges passed through from the child node to the parent node is defined as the number of hierarchical levels, by extracting the combination of the child node and the parent node for each hierarchical level, from the plurality \(M\) of variables \(x_1, x_2, \cdots, x\) M From this, a directed acyclic graph (a model representing a causal relationship) is generated. LiNGAM is a known method for causal analysis suitable for data with non-normal distributions. In this LiNGAM, for example, the graph generation unit 13 selects, from among a plurality \(M\) of variables \(x_1, x_2, \cdots, x\) M so-called exogenous variables, and repeats a process of removing other variables affected by the extracted exogenous variables and then removing the extracted exogenous variables until no exogenous variables can be extracted, from the plurality \(M\) of variables \(x_1, x_2, \cdots, x\) M to generate a directed acyclic graph (a model representing a causal relationship). For LiNGAM, for example, Causal Analysis manufactured by NEC is used.
[0050] The group processing unit 14 extracts nodes that form a group from a predetermined graph based on the grouping rules stored in the rule information storage unit 51. More specifically, for example, among a plurality of nodes in a predetermined graph, the group processing unit 14 randomly selects one node as a starting node from among the unprocessed nodes, and performs a grouping process of extracting nodes that form a group based on the grouping rules stored in the rule information storage unit 51 for the selected starting node, and repeats this until there are no unprocessed nodes.
[0051] For example, when the grouping rule includes the first and second rules, and the number of connections in the first rule is 2 and the number of branches in the second rule is 2, in FIG. 2, among the 11 nodes, the node with node number "2" (the second node) is selected as the starting node by random selection. Among the three edges connected to the second node according to the second rule, the edge (the 12th edge) connected to the node with node number "1" (the first node) and the edge (the 23rd edge) connected to the node with node number "3" (the third node) are selected by random selection. First, according to the first and second rules, for the second node which is the starting node, the first and third nodes are extracted. Since there is no node connected to the first node, the extraction according to the first rule for the first node ends. On the other hand, according to the first rule for the third node, the third node is used to extract the third node with node number "4" (the fourth node) from the second node. The second node, the first node, the third node, and the fourth node are grouped together into one group. From the 11 nodes shown in FIG. 2, the same processing as described above is executed for the remaining 7 nodes excluding the first to fourth nodes grouped into this one group, and such processing is executed until there are no unprocessed nodes. Note that for the starting node in the processing after the second time of extracting the nodes of the group, it may be randomly selected, or the next node connected to the node finally selected by the first rule may be selected.
[0052] Alternatively, for example, when the grouping rule includes the first and third rules and the number of connections in the first rule is two, in FIG. 2, the second node is selected as the starting node by random selection from among the 11 nodes. The sentence included in the second node (second sentence) is similar to the sentence included in the third node (third sentence) and the sentence included in the node with node number "6" (sixth node), while the second sentence is not similar to the sentence included in the first node (first sentence). The sixth sentence is similar to the sentence included in the node with node number "9" (ninth node), while the third sentence is not similar to the sentence included in the fourth node (fourth sentence). In this case, according to the first and third rules, the third node, the sixth node, and the ninth node are extracted for the second node which is the starting node, and the second node, the third node, the sixth node, and the ninth node are grouped together into one group. From the 11 nodes shown in FIG. 2, the same processing as described above is executed for the remaining 7 nodes excluding the second, third, sixth, and ninth nodes grouped into this one group, and such processing is executed until there are no unprocessed nodes.
[0053] Alternatively, for example, when the grouping rule includes the fourth rule, in FIG. 2, the sixth node is selected as the starting node by random selection from among the 11 nodes. The sixth sentence and the ninth sentence are similar, and the ninth sentence and the sentence (tenth sentence) included in the node with node number "10" (tenth node) are similar. On the other hand, the ninth sentence is not similar to the sentence (eighth sentence) included in the node with node number "8" (eighth node) and the sentence (eleventh sentence) included in the node with node number "11" (eleventh node). In this case, according to the fourth rule, the ninth node and the tenth node are extracted for the sixth node which is the starting node, and the sixth node, the ninth node, and the tenth node are grouped together into one group. From the 11 nodes shown in FIG. 2, the same processing as described above is executed for the remaining 8 nodes excluding the sixth, ninth, and tenth nodes grouped into this one group, and such processing is executed until there are no unprocessed nodes. Note that for the starting node in the processing after the second time of extracting the group nodes, it may be randomly selected, or the next node connected to the node finally selected by the fourth rule may be selected.
[0054] Alternatively, for example, when the grouping rule includes the fifth rule, in FIG. 2, the third node is selected as the starting node by random selection from among the 11 nodes. The third sentence is similar to each of the eighth and eleventh sentences, while the third sentence is not similar to each sentence included in the remaining nodes. In this case, according to the fifth rule, the eighth node and the eleventh node are extracted for the third node which is the starting node, and the third node, the eighth node, and the eleventh node are grouped together into one group. From the 11 nodes shown in FIG. 2, the same processing as described above is executed for the remaining 8 nodes excluding the third, eighth, and eleventh nodes grouped into this one group, and such processing is executed until there are no unprocessed nodes.
[0055] Note that in the above cases, it is also possible for one node to form one group.
[0056] In the determination of the similarity of articles, for example, the group processing unit 14 performs morphological analysis to decompose each of the two articles to be determined for similarity into the minimum units where language has meaning, and converts each result of the morphological analysis into a vector based on the number of occurrences of words by means of a method such as Bag-of-Words or TF-IDF. Then, it calculates the cosine similarity of each of the converted vectors, compares the value of the calculated cosine similarity with a predetermined threshold value (similarity determination threshold value) set in advance. As a result of the comparison, when the value of the calculated cosine similarity is greater than or equal to the similarity determination threshold value, it determines that the two articles are similar to each other; when the value of the calculated cosine similarity is less than the similarity determination threshold value, it determines that the two articles are not similar.
[0057] In addition to the above similarity determination method, for the determination of the similarity between articles, a method may be used in which the articles are vectorized using the entire article or words included in the article by using techniques based on distributed representations of words and articles such as One-Hot vectorization, Doc2Vec, word2vec, fastText, ELMo (Embedding from Language), or methods based on Transformer such as BERT, and the similarity between the articles is determined by calculating the similarity between the two vectors. Also, the similarity between the vectors may be calculated by a method of calculating distances such as Euclidean distance, Manhattan distance, Minkowski distance, etc. in addition to the cosine similarity.
[0058] Alternatively, for example, the group processing unit 14 performs morphological analysis on each of the two sentences to be determined for similarity, determines whether a preset co-occurrence rule holds between the results of the morphological analysis, and determines that the two sentences are similar to each other when the co-occurrence rule holds, and determines that the two sentences are not similar when the co-occurrence rule does not hold. The co-occurrence means that when a certain word appears in a certain sentence, another limited word frequently appears in that sentence. Therefore, by combining the certain word and the other limited word as a rule (co-occurrence rule) and determining whether each of the two sentences includes all the words in the co-occurrence rule, the determination of similarity becomes possible. In this case, a preset ontology representing the relationship between words may be used. For example, if it is set in the ontology that word A and word B are mutually replaceable, and word A is included in the co-occurrence rule, the group processing unit 14 first determines the validity of the co-occurrence rule with word A, and then determines the validity of the co-occurrence rule with word B. When the co-occurrence rule holds for at least one of them, it is determined as similar. By using the ontology, the similarity range can be adjusted.
[0059] By using any one of the third to fifth rules as described above, for example, group nodes can be extracted from a graph of a questionnaire survey in which the questionnaire question items are nodes and the question sentences are the sentences. In particular, the structural equation modeling support device 1000 in the embodiment is useful when modeling a questionnaire survey with a large number of question items by structural equation modeling.
[0060] Note that the graph generation unit 13 of the present embodiment generates a directed acyclic graph, but the group processing unit 14 may extract group nodes not only from a directed graph but also from an undirected graph.
[0061] The control unit 11 outputs the second data obtained by the discretization processing unit and the extraction result extracted by the group processing unit 14 to the output unit 3.
[0062] The fitness index processing unit 15 obtains the fitness index of the path diagram created based on the extraction result extracted by the group processing unit 14 from the graph. First, the user refers to the extraction result of the group processing unit 14 output to the output unit 3, creates a path diagram (structural equation) of the directed graph representing the model from the graph, and inputs the created path diagram to the input unit 2. At this time, the edges may be determined with reference to the respective edges of each node grouped into a group. For example, when the path diagram shown in FIG. 3 is created from the graph, the structural equation is as follows. Equation: x3 = b 31 x1 + b 32 x2 + e3, x4 = b 43 x3 + e4, x5 = b 54 x4 + e5, where, b mn is the coefficient between the variable (data) x m of the m-th m-node and the variable (data) x n of the n-th n-node, and e m is the disturbance (noise) at the m-th node. When receiving the input of this path diagram, the fitness index processing unit 15 obtains the fitness index of the input path diagram. As the fitness index serving as an index of the validity of the model, for example, CFI (Comparative Fit Index) and RMSEA (Root Mean Square Error of Approximation) are used.
[0063] The CFI is obtained by the following Equation 1, and generally, the closer the value is to 1, the better the model is judged. The RMSEA is obtained by the following Equation 2, and generally, if the value is 0.05 or less, the model is judged to be a good model.
[0064]
Equation
[0065]
Equation
[0066] The goodness-of-fit index is obtained after generating a model by obtaining the coefficients in the structural equation of the path diagram using data. For this reason, the goodness-of-fit index processing unit 15 is configured by, for example, SPSS Amos, which is SEM software manufactured by IBM. In this SPSS Amos, when the structural equation of the path diagram and data are input, the coefficients in the structural equation of the path diagram are obtained, a model is generated, and goodness-of-fit indices such as the above-mentioned CFI and RMSEA are obtained.
[0067] The control unit 11 outputs the goodness-of-fit index of the path diagram obtained by the goodness-of-fit index processing unit 15 to the output unit 3.
[0068] These control processing unit 1, input unit 2, output unit 3, IF unit 4, and storage unit 5 can be configured by, for example, a computer such as a desktop type or a notebook type.
[0069] Next, the operation of the present embodiment will be described. FIG. 4 is a flowchart showing the operation of the structural equation modeling support device.
[0070] When the power of the structural equation modeling support device 1000 having such a configuration is turned on, it executes initialization of each necessary part and starts its operation. In the control processing unit 1, a control unit 11, a discretization processing unit 12, a graph generation unit 13, a group processing unit 14, and a goodness-of-fit index processing unit 15 are functionally configured by executing its control processing program.
[0071] The user (operator) inputs, for example, M×N first data for generating a graph, which has a plurality of data values N for each of a plurality of different variables M, into the structural equation modeling support apparatus 1000 from, for example, the input unit 2 or via the IF unit 4 from a storage medium storing the first data. In FIG. 4, when the structural equation modeling support apparatus 1000 receives the input of the first data (S1), the discretization processing unit 12 of the control processing unit 1 generates second data by discretizing each of the plurality of data values from the first data into a multi-level system (S2).
[0072] Subsequently, the structural equation modeling support apparatus 1000 receives, for example, the input of a Bayesian network generation method, and the graph generation unit 13 of the control processing unit 1 generates a graph from the second data (S3).
[0073] Subsequently, the structural equation modeling support apparatus 1000 extracts nodes that form groups from the generated graph based on the grouping rules stored in the rule information storage unit 51 by the group processing unit 14 of the control processing unit 1 (S4).
[0074] Subsequently, the structural equation modeling support apparatus 1000 outputs the obtained second data and the extracted extraction result (group) to the output unit 3 by the control unit 11 of the control processing unit 1 (S5).
[0075] The user refers to the extraction result of the group processing unit 14 output to the output unit 3, creates a path diagram (structural equation) of the directed graph from the graph, and inputs the created path diagram into the input unit 2.
[0076] Subsequently, when the structural equation modeling support apparatus 1000 receives the input of the path diagram (S6), the fitness index processing unit 15 of the control processing unit 1 generates a model by structural equation modeling (S7) and obtains fitness indexes such as CFI and RMSEA (S8).
[0077] Then, the structural equation modeling support device 1000 outputs the obtained fitness index to the output unit 3 by the control unit 11 of the control processing unit 1 and ends this process (S9). Note that the control unit 11 may output the extraction result, the fitness index, etc. to an external device via the IF unit 4 as necessary.
[0078] The user refers to the output fitness index, adopts the created path diagram, discards the created path diagram and creates a new path diagram to execute from process S6 again, or executes again from process S1 (or process S3) without changing the generation method or by changing the generation method. Thus, a suitable path diagram can be created.
[0079] As described above, the structural equation modeling support device 1000 in the embodiment, the structural equation modeling support method and the structural equation modeling support program implemented thereon output, as an extraction result, a plurality of nodes that can be grouped into one in a predetermined graph. Therefore, a user (analyst) who creates a model (path diagram, structural equation) can create a new graph that is the basis of the path diagram from the predetermined graph by referring to the extraction result and assigning one new node to the plurality of nodes grouped into one. Thus, it is possible to support the creation of a path diagram in structural equation modeling.
[0080] In the first rule of the grouping rule, the structural equation modeling support device 1000, the structural equation modeling support method, and the structural equation modeling support program output, as an extraction result, nodes up to a predetermined number sequentially connected via edges from the starting node. Therefore, the user can create a new graph in which nodes up to a predetermined number sequentially connected via edges from the starting node are grouped into one by referring to the extraction result.
[0081] The above structural equation modeling support device 1000, structural equation modeling support method, and structural equation modeling support program, in the second rule of the grouping rules, further output, as an extraction result, a node connected to a predetermined number of edges selected from the plurality of edges when one node has a plurality of edges. Therefore, by referring to the extraction result, the user can create a new graph in which, from the predetermined graph, when one node has a plurality of edges, the nodes connected to a predetermined number of edges selected from the plurality of edges are grouped into one.
[0082] The above structural equation modeling support device 1000, structural equation modeling support method, and structural equation modeling support program, in the third rule of the grouping rules, further output, as an extraction result, the two nodes when each sentence provided in each of the two nodes connected to both ends of one edge is similar to each other. Therefore, by referring to the extraction result, the user can create a new graph in which, from the predetermined graph, the similarity of each sentence provided in each of the two nodes is taken into consideration.
[0083] The above structural equation modeling support device 1000, structural equation modeling support method, and structural equation modeling support program, in the fourth rule of the grouping rules, further output, as an extraction result, the starting node and the destination node when each sentence provided in each of the starting node and the destination node connected to the starting node via an edge is similar to each other. Therefore, by referring to the extraction result, the user can create a new graph in which, from the predetermined graph, the similarity of each sentence provided in each of the two nodes is taken into consideration.
[0084] The above structural equation modeling support device 1000, structural equation modeling support method, and structural equation modeling support program output, as extraction results, a plurality of nodes each having sentences similar to each other in the fifth rule of the grouping rules. Therefore, by referring to the extraction results, the user can create a new graph from a predetermined graph, taking into account the similarity of the sentences provided in the nodes.
[0085] The above structural equation modeling support device 1000, structural equation modeling support method, and structural equation modeling support program obtain and output a goodness-of-fit index of a path diagram. Therefore, the user can recognize the goodness-of-fit of the path diagram created by referring to the extraction results, and can recreate the path diagram or cause the structural equation modeling support device 1000 to re-extract according to the goodness-of-fit of the path diagram.
[0086] Since the above structural equation modeling support device 1000, structural equation modeling support method, and structural equation modeling support program can generate a graph by a plurality of generation methods, the generation method of the graph can be changed according to the goodness-of-fit of the path diagram.
[0087] Since the above structural equation modeling support device 1000, structural equation modeling support method, and structural equation modeling support program discretize data values into a multi-level system, a graph (Bayesian network) can be generated by the Bayesian network generation method even if the data values of the first data are continuous values.
[0088] According to an embodiment, a structural equation modeling support device 1000 including a graph generation unit 13 that generates a graph by a Bayesian network generation method can be provided. According to an embodiment, a structural equation modeling support device 1000 in which the data values are discrete values or continuous values can be provided.
[0089] In the above-described embodiment, the graph generation unit 13 generates the graph from M×N data having a plurality of N data values for each of a plurality of M variables that are different from each other. The structural equation modeling support apparatus 1000 further functionally includes, in the control processing unit 1, a normal distribution determination unit 16 that determines whether or not the data is a normal distribution, as indicated by a broken line in FIG. 1. When the determination result of the normal distribution determination unit 16 is a normal distribution, the graph generation unit 13 generates the graph by a Bayesian network generation method. When the determination result of the normal distribution determination unit 16 is not a normal distribution, the graph generation unit 13 may generate the graph by LiNGAM. More specifically, the normal distribution determination unit 16 determines, for each of the plurality of variables (nodes), whether or not a plurality of data values in the variable are a normal distribution. When the ratio of the number of variables determined to be a normal distribution to the total number of variables is equal to or greater than a predetermined threshold (normal distribution determination ratio threshold (e.g., 80 [%], 90 [%], etc.)), it is determined that the data is finally a normal distribution. When the ratio is less than the normal distribution determination region position, it is determined that the data is not finally a normal distribution. Such a structural equation modeling support apparatus 1000 can select a suitable graph generation method according to whether or not it is a normal distribution, and suitable graph generation can be expected.
[0090] In order to represent the present invention, the embodiments have been appropriately and fully described above with reference to the drawings. However, it should be recognized that those skilled in the art can easily make changes and / or improvements to the above-described embodiments. Therefore, as long as the modified forms or improved forms implemented by those skilled in the art do not depart from the scope of the claims described in the claims, the modified forms or the improved forms are interpreted as being included in the scope of the claims.
Explanation of Signs
[0091] 1000 Structural equation modeling support apparatus 1 Control processing unit 2 Input unit 3 Output unit 4 Interface unit (IF unit) 5 Storage unit 11 Control Unit 12 Discretization Processing Unit 13 Graph Generation Unit 14 Group Processing Unit 15 Fitness Index Processing Unit 16 Normal Distribution Judgment Unit
Claims
1. A structural equation modeling support device for assisting in creating a path diagram in structural equation modeling, comprising: a rule information storage unit that stores a grouping rule, which is a rule for grouping a plurality of nodes in a graph into one group; a group processing unit that extracts nodes to be grouped from a predetermined graph based on the grouping rule stored in the rule information storage unit; and an output unit that outputs the extraction result extracted by the group processing unit. The structural equation modeling support device.
2. The grouping rule includes a rule for grouping nodes up to a predetermined number sequentially connected via edges from a starting node into one group. The structural equation modeling support device according to claim 1.
3. The grouping rule further includes a rule for selecting a predetermined number of edges from among the plurality of edges when one node has a plurality of edges. The structural equation modeling support device according to claim 2.
4. Each of the plurality of nodes in the graph has a sentence. The grouping rule further includes a rule for grouping two nodes connected to both ends of one edge into one group when the sentences provided in each of the two nodes are similar to each other. The structural equation modeling support device according to claim 2.
5. Each of the plurality of nodes in the graph has a sentence. The grouping rule includes a rule for repeatedly executing a similar grouping process of grouping the starting node and the destination node into one group when the sentences provided in the starting node and the destination node connected to the starting node via an edge are similar to each other, using the destination node as a new starting node until grouping is no longer possible. The structural equation modeling support device according to claim 1.
6. Each of the plurality of nodes in the graph has a sentence. The grouping rule includes a rule for executing a similar grouping process of extracting nodes having sentences similar to the sentence provided in the starting node from among the plurality of nodes in the graph and grouping the starting node and the extracted nodes into one group. The structural equation modeling support device according to claim 1.
7. Further comprising a fitness index processing unit that obtains a fitness index of a path diagram created from the graph based on the extraction result extracted by the group processing unit. The output unit further outputs the fitness index of the path diagram obtained by the fitness index processing unit. The structural equation modeling support device according to claim 1.
8. Further comprising a graph generation unit that generates the graph by a plurality of generation methods, And an input unit that accepts any one of the plurality of generation methods, The graph generation unit generates the graph by the generation method accepted by the input unit among the plurality of generation methods. The structural equation modeling support device according to claim 7.
9. Further comprising a graph generation unit that generates the graph by a Bayesian network generation method. The structural equation modeling support device according to claim 1.
10. Further comprising a discretization processing unit that generates second data including a plurality of discretized data values for each of the plurality of variables by discretizing each of the plurality of data values for each of the plurality of different variables from first data including a plurality of data values into a multi-level system. The graph generation unit generates the graph from the second data generated by the discretization processing unit. The output unit further outputs the second data generated by the discretization processing unit. The structural equation modeling support device according to claim 9.
11. Further comprising a graph generation unit that generates the graph from data including a plurality of data values for each of a plurality of different variables, And a normal distribution determination unit that determines whether the data is a normal distribution. When the determination result of the normal distribution determination unit is a normal distribution, the graph generation unit generates the graph by a Bayesian network generation method, and when the determination result of the normal distribution determination unit is not a normal distribution, the graph generation unit generates the graph by LiNGAM. The structural equation modeling support device according to claim 1.
12. Further comprising a graph generation unit that generates the graph from data including a plurality of data values for each of a plurality of different variables. The data values are discrete values or continuous values. The structural equation modeling support device according to claim 1.
13. A structural equation modeling support method for assisting in creating a path diagram in structural equation modeling, A grouping process step of extracting nodes to be grouped from a predetermined graph based on a grouping rule, which is a rule for grouping a plurality of nodes in a graph into one and storing it in a rule information storage unit, An output step of outputting an extraction result extracted in the grouping process step, A structural equation modeling support method.
14. A structural equation modeling support program for causing a computer to function as the structural equation modeling support device according to any one of Claims 1 to 12.
Citation Information
Patent Citations
Method, apparatus, and system for estimating causal relationships between observed variables
JP6743934B2