Model Generation Device, Inference Device, Model Generation Method, and Storage Medium
Through the combination of feature braiding network and speculator, the serious problem of feature loss in the prior art is solved, and high-precision solution of complex graph structure tasks is realized, and computing resource consumption is reduced.
Patent Information
- Application Number
- CN202180081830.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-29
- Filing Date
- 2021-09-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-09-14
AI Technical Summary
When using neural networks to solve complex graph structure tasks, the prior art has serious feature loss and the network level cannot be deepened, resulting in insufficient inferred accuracy.
The machine learning model of feature braiding network and speculator is used to extract the feature information of the graph through multiple feature braiding layers, and the encoder is used to derive relative feature quantities to avoid excessive smoothing and support network hierarchy deepening.
The calculation accuracy of complex graph structure tasks is improved, high-precision speculation results are achieved, and computing resource consumption is reduced.
Smart Images

Figure CN116635877B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a model generation device, an inference device, a model generation method, and a computer-readable storage medium. Background Art
[0002] In various inference tasks such as solving matching and path search, graphs are sometimes used to represent various information. A graph basically includes vertices (nodes) and branches (edges). Depending on the task, graphs with various structures such as directed graphs (digraphs) and bipartite graphs can be used.
[0003] As an example, in solving the stable matching problem, a bipartite graph can be used. The stable matching problem is an allocation problem in a two-sided market such as the matching between job seekers and employers. A bipartite graph is a graph configured such that a vertex set is divided into two subsets and there are no branches between vertices within each subset (that is, the multiple vertices constituting the bipartite graph are divided into belonging to either of the two subsets and there are no branches within each subset).
[0004] As an example of a method for expressing the stable matching problem, each subset is associated with a party to which the objects to be matched belong, and the vertices within each subset are associated with the objects to be matched (e.g., job seekers / employers). Each object (vertex) has an expected degree (e.g., expected rank) for an object belonging to the other party. The matching is expressed by a set of branches including branches, where the branches represent combinations of vertices belonging to one subset and vertices belonging to the other subset and are allocated in a non-overlapping vertex manner.
[0005] When there is a combination of vertices with a higher expected degree than the combination of branches belonging to the set of branches constituting the matching, the branch representing this combination blocks the matching (i.e., it means that there is another more preferable matching than the matching represented by this set of branches). In the stable matching problem, a combination of vertices is searched in such a way that there is no such blocking pair and a certain criterion (e.g., fairness) is satisfied.
[0006] Conventionally, to solve such tasks represented by graphs, various algorithms have been developed. The GS algorithm proposed in Non-Patent Document 1 is known as an example of a classical method for solving the stable matching problem. However, such a method using a specific algorithm lacks generality because it becomes unusable even with a slight change in the content of the task, and moreover, since it includes many manual operations, the generation cost of the solver is high.
[0007] Therefore, in recent years, as a general method for solving tasks expressed by graphs, research on methods using machine learning models such as neural networks has been advanced. For example, Non-Patent Document 2 proposes to solve the stable matching problem using a five-layer neural network. According to the method using such a machine learning model, as long as learning data corresponding to the task is prepared, a learned model (solver) capable of obtaining the ability to solve the task can be generated. Therefore, the generality of the generated model can be improved, and the generation cost can be reduced.
[0008] Prior Art Documents
[0009] Non-Patent Documents
[0010] Non-Patent Document 1: David Gale, Lloyd S. Shapley, "College admissions and the stability of marriage", "The American Mathematical Monthly", 69(1): 9-15, 1962. (David Gale, Lloyd S Shapley, "College admissions and the stability of marriage", The American Mathematical Monthly, 69(1): 9-15, 1962.)
[0011] Non-Patent Document 2: Shira Li, "Deep Learning for Two-Sided Matching Markets", PhD thesis, Harvard University, 2019. (Shira Li, "Deep Learning for Two-Sided Matching Markets", PhD thesis, Harvard University, 2019.) Summary of the Invention
[0012] Problems to be Solved by the Invention
[0013] The inventors of the present application have found the following problems in the network models (such as Graph Convolutional Networks) proposed in Non-Patent Document 2 and the like for solving inference tasks for data having a graph structure. That is, in order to solve complex tasks with high accuracy, it is desirable to deepen the layers of the network. However, in the conventional method, convolution operation for the features of each vertex is performed by calculating the weighted sum of the features of adjacent vertices for each type of edge. Due to the calculation of the weighted sum, excessive smoothing occurs, so that the more layers are stacked, the more features are lost. Therefore, the layers of the network cannot be deepened (for example, five layers in the method proposed in Non-Patent Document 2), and due to this, it is difficult to improve the accuracy of inferring the solution of the task.
[0014] The present invention has been accomplished in view of such a situation, and an object thereof is to provide a technique for improving the accuracy of inferring a solution to a task expressed by a diagram.
[0015] Technical means for solving the problem
[0016] To solve the above problems, the present invention adopts the following configuration.
[0017] That is, the model generation device according to one aspect of the present invention includes: an acquisition unit configured to acquire a plurality of training graphs; and a learning processing unit configured to perform machine learning of a speculation module using the acquired plurality of training graphs. The speculation module includes a feature weaving network and a speculator. The feature weaving network is configured to receive an input of an input graph and output feature information related to the input graph by including a plurality of feature weaving layers. The speculator is configured to speculate a solution to a task of the input graph based on the feature information output from the feature weaving network. Each feature weaving layer is configured to receive inputs of a three-layer first input tensor and a second input tensor. The first input tensor is configured to have: a first axis arranged with first vertices belonging to a first set of the input graph as elements, a second axis arranged with second vertices belonging to a second set of the input graph as elements, and a third axis arranged with feature amounts related to branches flowing from the first vertices to the second vertices as elements. The second input tensor is configured to have: a first axis arranged with the second vertices as elements, a second axis arranged with the first vertices as elements, and a third axis arranged with feature amounts related to branches flowing from the second vertices to the first vertices as elements. Each feature weaving layer includes an encoder. And each feature weaving layer is configured to: fix the third axis of the second input tensor and cycle through the other axes of the second input tensor one by one, and connect the feature amounts of each element of the first input tensor and the cycled second input tensor to generate a three-layer first connection tensor, fix the third axis of the first input tensor and cycle through the other axes of the first input tensor one by one, and connect the feature amounts of each element of the second input tensor and the cycled first input tensor to generate a three-layer second connection tensor, divide the generated first connection tensor into each element of the first axis and input it to the encoder, perform the operation of the encoder to generate a three-layer first output tensor corresponding to the first input tensor, and divide the generated second connection tensor into each element of the first axis and input it to the encoder, perform the operation of the encoder to generate a three-layer second output tensor corresponding to the second input tensor. The encoder is configured to derive relative feature amounts of each element based on the feature amounts of all the input elements. The first feature weaving layer among the plurality of feature weaving layers is configured to receive the first input tensor and the second input tensor from the input graph. The feature information includes the first output tensor and the second output tensor output from the last feature weaving layer among the plurality of feature weaving layers. The machine learning is configured in such a way that each training graph is input to the feature weaving network as the input graph, thereby training the speculation module so that the speculation result obtained by the speculator conforms to the correct solution of the task for each training graph.
[0018] In the said structure, the features of the input graph are extracted by a feature weaving network. In the feature weaving network, the data of the input graph is expressed by input tensors. Each input tensor is configured to represent the feature quantity of the out-branch from the vertex belonging to each set to the vertex belonging to another set. That is, the feature weaving network is configured to refer to the features of the input graph not for each vertex but for each branch.
[0019] In each feature weaving layer, the axes of the input tensors are aligned and connected to generate a connection tensor. Each connection tensor is configured to represent the feature quantities of the out-branch and the in-branch based on the vertices belonging to each set. Also, in each feature weaving layer, each connection tensor is divided into each element of the first axis and input to the encoder to perform the operations of the encoder. Thus, it is possible to reflect the feature quantities of other branches in the feature quantities of each branch based on each vertex. The more the feature weaving layers are stacked, the more the feature quantities of the branches related to the vertices departing from each vertex can be reflected in the feature quantities of the branches related to each vertex.
[0020] In addition, the encoder is configured to derive the relative feature quantities of each branch not by calculating the weighted sum of the features of the adjacent branches for each vertex but based on the feature quantities of all the input branches. According to the structure of the said encoder, it is possible to prevent excessive smoothing even when the feature weaving layers are stacked. Therefore, it is possible to deepen the hierarchy of the feature weaving network according to the difficulty of the task. As a result, in the feature weaving network, it is possible to expect the extraction of appropriate features of the input graph for the solution task. Thus, in the predictor, it is possible to expect high-precision prediction based on the obtained feature quantities. Therefore, according to the said structure, even for complex tasks, by deepening the hierarchy of the feature weaving network, it is possible to improve the accuracy of predicting the solution of the task expressed by the graph for the generated trained prediction module.
[0021] In the model generation device of the said aspect, the feature weaving network can be configured to have a residual structure. The so-called residual structure refers to a structure having a shortcut as described below, that is, the output of at least one of the multiple feature weaving layers is input not only to the next adjacent feature weaving layer but also to other feature weaving layers (feature weaving layers separated by two or more layers) after the next adjacent feature weaving layer. According to the said structure, with the residual structure, it is possible to easily deepen the hierarchy of the feature weaving network. Thus, it is possible to improve the prediction accuracy of the generated trained prediction module.
[0022] In the model generation device of the above aspect, the encoder of each feature weaving layer may include a single encoder that is commonly used among elements of the first axis of the first connection tensor and the second connection tensor. According to the above structure, by using a common encoder in the operation of each element of each connection tensor in each feature weaving layer, the number of parameters of the inference module can be reduced. Thereby, the computational amount of the inference module (that is, the consumption of computational resources used in the operation of the inference module) can be suppressed, and the efficiency of machine learning can be achieved.
[0023] In the model generation device of the above aspect, each feature weaving layer may also include a normalizer corresponding to each connection tensor, and the normalizer is configured to normalize the operation result of the encoder. In the case where the obtained first input tensor and the second input tensor belong to different distributions, there may be a situation where there are two domains of the input to the encoder. Whenever the encoder (feature weaving layer) is stacked, the bias caused by the difference in the distributions is amplified. As a result, a single encoder may fall into a state of solving two different tasks instead of a common task. As a result, the accuracy of the encoder may deteriorate, leading to the deterioration of the overall accuracy of the inference module. To solve this problem, according to the above structure, the operation result of the encoder is normalized by the normalizer, thereby suppressing the amplification of the bias, and thus preventing the deterioration of the accuracy of the inference module.
[0024] In the model generation device of the above aspect, each feature weaving layer may also be configured to assign a conditional code to the third axis of each of the first input tensor and the second input tensor before performing the operation of the encoder. In the above structure, when the encoder encodes the feature quantity, it can identify the source of the data based on the conditional code. Therefore, even in the case where the first input tensor and the second input tensor belong to different distributions, during the process of machine learning, the encoder can be trained by means of the conditional code to eliminate the bias caused by the difference in the distributions. Therefore, according to the above structure, the deterioration of the accuracy of the inference module caused by different distributions to which each input tensor belongs can be prevented.
[0025] In the model generation device of the above aspect, each training graph may be a directed bipartite graph, and the multiple vertices constituting the directed bipartite graph may be divided into any one of two subsets. The first set of the input graph may correspond to one of the two subsets of the directed bipartite graph, and the second set of the input graph may correspond to the other of the two subsets of the directed bipartite graph. According to the above structure, in the scenario where the given task is expressed by a directed bipartite graph, for the generated trained inference module, the accuracy of inferring the solution of the task can be improved.
[0026] In the model generation device of the above aspect, the task may also be a matching task between two groups, and it is a task of determining the best pairing of objects belonging to each group. Two subsets of the directed bipartite graph may correspond to the two groups in the matching task, and the vertices belonging to each subset of the directed bipartite graph may also correspond to the objects belonging to each group. Moreover, the feature quantity related to the branch flowing from the first vertex to the second vertex may correspond to the degree of expectation of an object belonging to one of the two groups for an object belonging to the other group, and the feature quantity related to the branch flowing from the second vertex to the first vertex may also correspond to the degree of expectation of an object belonging to the other of the two groups for an object belonging to one of them. According to the above structure, for the generated and trained inference module, it is possible to improve the accuracy of inferring the solution of the two-sided matching expressed by the directed bipartite graph.
[0027] In the model generation device of the above aspect, each training graph may be an undirected bipartite graph, and the multiple vertices constituting the undirected bipartite graph may be divided into any one of the two subsets. The first set of the input graph may correspond to one of the two subsets of the undirected bipartite graph, and the second set of the input graph may correspond to the other of the two subsets of the undirected bipartite graph. According to the above structure, when the given task is expressed by an undirected bipartite graph, for the generated and trained inference module, it is possible to improve the accuracy of inferring the solution of the task.
[0028] In the model generation device of the above aspect, the task may also be a matching task between two groups, and it is a task of determining the best pairing of objects belonging to each group. Two subsets of the undirected bipartite graph may correspond to the two groups in the matching task, and the vertices belonging to each subset of the undirected bipartite graph may also correspond to the objects belonging to each group. The feature quantity related to the branch flowing from the first vertex to the second vertex and the feature quantity related to the branch flowing from the second vertex to the first vertex may both correspond to the cost or reward of pairing an object belonging to one of the two groups with an object belonging to the other group. According to the above structure, for the generated and trained inference module, it is possible to improve the accuracy of inferring the solution of the one-sided matching expressed by the undirected bipartite graph.
[0029] In the model generation device of the above aspect, each training graph may be a directed graph. The first vertex belonging to the first set of the input graph may correspond to the starting point of the directed branch constituting the directed graph, and the second vertex belonging to the second set may also correspond to the ending point of the directed branch. Moreover, the branch flowing out from the first vertex to the second vertex may correspond to the directed branch flowing out from the starting point to the ending point, and the branch flowing out from the second vertex to the first vertex may also correspond to the directed branch flowing into the ending point from the starting point. According to the above structure, when the given task is expressed by a general directed graph, for the generated and trained inference module, it is possible to improve the accuracy of inferring the solution of the task.
[0030] In the model generation device of the above aspect, the directed graph may be configured as a network, and the task may be to infer the situation occurring on the network. The feature quantity related to the branch flowing out from the first vertex to the second vertex and the feature quantity related to the branch flowing out from the second vertex to the first vertex may both correspond to the connection attribute between the respective vertices constituting the network. According to the above structure, for the generated and trained inference module, it is possible to improve the accuracy of inferring the situation occurring on the network expressed by a general directed graph.
[0031] In the model generation device of the above aspect, each training graph may be an undirected graph. The first vertex belonging to the first set of the input graph and the second vertex belonging to the second set may respectively correspond to the respective vertices constituting the undirected graph. The first input tensor and the second input tensor input to each feature weaving layer may be the same as each other. The first output tensor and the second output tensor generated by each feature weaving layer may be the same as each other, and the process of generating the first output tensor and the process of generating the second output tensor may be executed jointly. According to the above structure, when the given task is expressed by a general undirected graph, for the generated and trained inference module, it is possible to improve the accuracy of inferring the solution of the task.
[0032] In the model generation device of the above aspect, each training graph may be a vertex feature graph including a plurality of vertices, and is a vertex feature graph in which each vertex has an attribute. The first vertex belonging to the first set of the input graph and the second vertex belonging to the second set may respectively correspond to the respective vertices constituting the vertex feature graph. The feature quantity related to the branch flowing out from the first vertex to the second vertex may also correspond to the attribute of the vertex of the vertex feature graph corresponding to the first vertex, and the feature quantity related to the branch flowing out from the second vertex to the first vertex may also correspond to the attribute of the vertex of the vertex feature graph corresponding to the second vertex. According to the above structure, when the given task is expressed by a vertex feature graph (for example, a vertex weighted graph) in which the vertex has a feature (attribute), for the generated and trained inference module, it is possible to improve the accuracy of inferring the solution of the task.
[0033] In the model generation device of the above aspect, for each element of the third axis of the first input tensor and the second input tensor from the input graph, information representing the relationship between corresponding vertices of the vertex feature map can be connected or multiplied, and the task can be a task of inferring a situation derived from the vertex feature map. According to the above structure, for the generated and trained inference module, it is possible to improve the accuracy of inferring a situation related to an object expressed by a vertex feature map.
[0034] In the model generation device of the above aspect, the task can be a task of inferring the relationship between the vertices constituting the vertex feature map. According to the above structure, for the generated and trained inference module, it is possible to improve the accuracy of inferring the relationship between the objects expressed by the vertices of the vertex feature map.
[0035] In the model generation device of each of the above aspects, a branch is configured to represent a combination of two vertices (i.e., represent the relationship between two vertices). However, the structure of the branch is not limited to such an example. In another example, a branch may be configured to represent a combination of three or more vertices. That is, the graph may be a hypergraph. For example, a model generation device according to an aspect of the present invention includes: an acquisition unit configured to acquire a plurality of training graphs; and a learning processing unit configured to perform machine learning of a prediction module using the acquired plurality of training graphs. Each training graph is a K-part hypergraph. K is 3 or more. The plurality of vertices constituting the hypergraph are divided into any one of K subsets. The prediction module includes a feature weaving network and a predictor. The feature weaving network is configured to receive an input of an input graph by including a plurality of feature weaving layers and output feature information related to the input graph. The predictor is configured to predict a solution to a task of the input graph based on the feature information output from the feature weaving network. Each feature weaving layer receives an input of K (K + 1)-layer input tensors. The i-th input tensor among the K input tensors is configured such that in the first axis, the i-th vertex belonging to the i-th subset of the input graph is arranged as an element, and in the j-th axis from the second axis to the K-th axis, the vertices belonging to the (j - 1)-th subset are arranged as elements in a cyclic manner starting from the i-th one, and in the (K + 1)-axis, the feature amounts related to the branches flowing out from the i-th vertex belonging to the i-th subset to the combinations of the vertices belonging to (K - 1) other subsets are arranged as elements. Each feature weaving layer includes an encoder. Each feature weaving layer is configured to: for the i-th input tensor from the first to the K-th, fix the (K + 1)-axis of the other (K - 1) input tensors other than the i-th one, and cyclically shift the other axes other than the (K + 1)-axis of the other input tensors in a manner consistent with each axis of the i-th input tensor, and connect the feature amounts of the elements of the i-th input tensor and the cycled other input tensors, thereby generating K (K + 1)-layer connection tensors, and dividing each of the generated K connection tensors into elements of the first axis and inputting them to the encoder, and performing the operation of the encoder, thereby generating K (K + 1)-layer output tensors respectively corresponding to the K input tensors. The encoder is configured to derive the relative feature amounts of the elements based on the feature amounts of all the input elements. The first feature weaving layer among the plurality of feature weaving layers is configured to receive K input tensors from the input graph. The feature information includes the K output tensors output from the last feature weaving layer among the plurality of feature weaving layers. And the machine learning is configured such that each training graph is input as an input graph to the feature weaving network, thereby training the prediction module so that the prediction result obtained by the predictor conforms to the correct answer of the task for each training graph. According to the above structure, when the given task is expressed by a hypergraph, for the generated trained prediction module, it is possible to improve the accuracy of predicting the solution to the task.
[0036] Moreover, the form of the present invention is not limited to the form of the model generation device. One aspect of the present invention may also be a speculation device configured to use the trained speculation module generated by the model generation device of any of the above forms to speculate on the solution to the task of the graph. For example, the speculation device according to one aspect of the present invention includes: an acquisition unit configured to acquire an object graph; a speculation unit configured to use the speculation module trained by machine learning to speculate on the solution to the task of the acquired object graph; and an output unit configured to output information related to the result of speculating on the solution to the task. The speculation module is configured in the same manner as described above. The process of speculating on the solution to the task of the object graph using the trained speculation module is configured by inputting the object graph as an input graph into the feature weaving network and obtaining the result of speculating on the solution to the task from the speculator. In addition, the speculation device can be rewritten as, for example, a matching device, a prediction device, etc. according to the type of task in the applicable scenario.
[0037] Moreover, as other forms of the model generation device and the speculation device of each of the above forms, one aspect of the present invention may also be an information processing method for implementing all or a part of the above structures, may also be a program, and may also be a computer-readable storage medium such as a computer or other device or machine storing such a program. Here, the so-called computer-readable storage medium refers to a medium that stores information such as a program through electrical, magnetic, optical, mechanical, or chemical actions. Moreover, one aspect of the present invention may also be a speculation system including the model generation device and the speculation device of any of the above forms.
[0038] For example, the model generation method according to one aspect of the present invention is an information processing method, and a computer executes the following steps: acquiring a plurality of training graphs; and performing machine learning of the speculation module using the acquired plurality of training graphs. Moreover, for example, the model generation program according to one aspect of the present invention is a program for causing a computer to execute the following steps, that is: acquiring a plurality of training graphs; and performing machine learning of the speculation module using the acquired plurality of training graphs. The speculation module is configured in the same manner as described above. The machine learning is configured by inputting each training graph as an input graph into the feature weaving network, thereby training the speculation module so that the speculation result obtained by the speculator conforms to the correct answer to the task of each training graph.
[0039] Effects of the Invention
[0040] According to the present invention, it is possible to improve the accuracy of speculating on the solution to the task expressed by the graph. Brief Description of the Drawings
[0041] Figure 1 Schematically illustrate an example of the scenario to which the present invention is applied.
[0042] Figure 2A An example of the structure of the speculation module that schematically illustrates an embodiment.
[0043] Figure 2B An example of the structure of the feature weaving layer that schematically illustrates an embodiment.
[0044] Figure 3 An example of the hardware structure of the model generation device that schematically illustrates an embodiment.
[0045] Figure 4 An example of the hardware structure of the speculation device that schematically illustrates an embodiment.
[0046] Figure 5 An example of the software structure of the model generation device that schematically illustrates an embodiment.
[0047] Figure 6A An example of the structure of the encoder disposed in the feature weaving layer that schematically illustrates an embodiment.
[0048] Figure 6B An example of the structure of the encoder disposed in the feature weaving layer that schematically illustrates an embodiment.
[0049] Figure 7 An example of the software structure of the speculation device that schematically illustrates an embodiment.
[0050] Figure 8 A flowchart showing an example of the processing flow of the model generation device according to an embodiment.
[0051] Figure 9 A flowchart showing an example of the processing flow of the speculation device according to an embodiment.
[0052] Figure 10 An example that schematically illustrates the case of passing through two feature weaving layers according to an embodiment.
[0053] Figure 11 An example that schematically illustrates another scenario to which the present invention is applied (a scenario in which a task is expressed by a directed bipartite graph).
[0054] Figure 12 An example that schematically illustrates another scenario to which the present invention is applied (a scenario in which a task is expressed by an undirected bipartite graph).
[0055] Figure 13 An example that schematically illustrates another scenario to which the present invention is applied (a scenario in which a task is expressed by a general directed graph).
[0056] Figure 14Illustrate an example of another scenario to which the present invention is applicable (a scenario expressing a task by a general undirected graph) schematically.
[0057] Figure 15 Illustrate an example of the arithmetic processing of a feature weaving layer in a scenario expressing a task by a general undirected graph schematically.
[0058] Figure 16 Illustrate an example of another scenario to which the present invention is applicable (a scenario expressing a task by a vertex feature graph with relationality between vertices) schematically.
[0059] Figure 17 Illustrate an example of another scenario to which the present invention is applicable (a scenario in which a task of inferring the relationality between objects is expressed by a vertex feature graph) schematically.
[0060] Figure 18 Illustrate an example of another scenario to which the present invention is applicable (a scenario expressing a task by a hypergraph) schematically.
[0061] Figure 19 Illustrate an example of the arithmetic processing of a feature weaving layer in a scenario expressing a task by a hypergraph schematically.
[0062] Figure 20 Illustrate an example of the processing procedure of machine learning of a modification example schematically.
[0063] Figure 21 Illustrate an example of the input form of a speculation module of a modification example schematically.
[0064] Figure 22 Illustrate an example of the structure of a speculation module of a modification example schematically.
[0065] Figure 23 Illustrate an example of the structure of a feature weaving layer of a modification example schematically.
[0066] Figure 24A Show the results of calculating the success rate of matching for verification samples of each size for the examples.
[0067] Figure 24B Show the results of calculating the success rate of matching for verification samples of each size for the comparative examples.
[0068] Figure 25A Show the results of calculating the success rate of fair matching for verification samples of each size for the examples.
[0069] Figure 25B Show the results of calculating the success rate of fair matching for verification samples of each size for the comparative examples.
[0070] Figure 25CResults showing the average fairness cost of all matching results calculated for the second and third embodiments.
[0071] Figure 26A Results showing the success rate of balanced matching for verification samples of each size calculated for the embodiments.
[0072] Figure 26B Results showing the success rate of balanced matching for verification samples of each size calculated for the comparative example.
[0073] Figure 26C Results showing the average balanced cost of all matching results calculated for the second and fourth embodiments.
[0074] Explanation of symbols
[0075] 1: Model generation device
[0076] 11: Control unit
[0077] 12: Storage unit
[0078] 13: Communication interface
[0079] 14: External interface
[0080] 15: Input device
[0081] 16: Output device
[0082] 17: Driver
[0083] 81: Model generation program
[0084] 91: Storage medium
[0085] 111: Acquisition unit
[0086] 112: Learning processing unit
[0087] 113: Saving processing unit
[0088] 125: Learning result data
[0089] 2: Deduction device
[0090] 21: Control unit
[0091] 22: Storage unit
[0092] 23: Communication interface
[0093] 24: External interface
[0094] 25: Input device
[0095] 26: Output device
[0096] 27: Driver
[0097] 82: Speculation program
[0098] 92: Storage medium
[0099] 211: Acquisition unit
[0100] 212: Speculation unit
[0101] 213: Output unit
[0102] 221: Object diagram
[0103] 30: Training diagram
[0104] 5: Speculation module
[0105] 50: Feature weaving network
[0106] 500: Feature weaving layer
[0107] 505: Encoder
[0108] 55: Speculator Detailed implementation manners
[0109] Hereinafter, an implementation manner of one aspect of the present invention (hereinafter also referred to as "this implementation manner") will be described based on the accompanying drawings. However, the implementation manner described below is merely an illustration of the present invention in all aspects. Of course, various improvements or deformations can be made without departing from the scope of the present invention. That is, when implementing the present invention, specific structures corresponding to the implementation manner can also be appropriately adopted. In addition, the data that appears in this implementation manner is described in natural language, but more specifically, it is specified in a computer-recognizable pseudo language, command, parameter, machine language, etc.
[0110] §1 Application example
[0111] Figure 1 An example of a scenario to which the present invention is applied is schematically illustrated. As Figure 1 shown, the speculation system 100 of this implementation manner includes a model generation device 1 and a speculation device 2.
[0112] (Model generation device)
[0113] The model generation device 1 of the present embodiment is a computer configured to generate a trained inference module 5 by implementing machine learning. Specifically, the model generation device 1 acquires a plurality of training graphs 30. The type of each training graph 30 can be appropriately selected according to the ability of the inference task obtained by the inference module 5. Each training graph 30 may include, for example, a directed bipartite graph, an undirected bipartite graph, a general directed graph, a general undirected graph, a vertex feature graph, a hypergraph, etc. The model generation device 1 uses the acquired plurality of training graphs 30 to implement machine learning of the inference module 5. Thereby, a trained inference module 5 is generated.
[0114] (Inference device)
[0115] On the other hand, the inference device 2 of the present embodiment is a computer configured to use the trained inference module 5 to infer the solution to the task of the graph. Specifically, the inference device 2 acquires an object graph 221. The object graph 221 is a graph that is the object of inferring the solution to the task, and its type is the same as that of the training graph 30 and can be appropriately selected according to the ability obtained by the trained inference module 5. The inference device 2 uses the inference module 5 trained by machine learning to infer the solution to the task of the acquired object graph 221. And, the inference device 2 outputs information related to the result of inferring the solution to the task.
[0116] (Inference module)
[0117] Further use Figure 2A to illustrate an example of the structure of the inference module 5. Figure 2A An example of the structure of the inference module 5 of the present embodiment is schematically illustrated. The inference module 5 includes a feature weaving network 50 and an inferencer 55. The feature weaving network 50 includes a plurality of feature weaving layers 500 arranged in series. Thus, the feature weaving network 50 is configured to receive the input of the input graph and output feature information related to the input graph. The inferencer 55 is configured to infer the solution to the task of the input graph based on the feature information output from the feature weaving network.
[0118] Further use Figure 2B to illustrate an example of the structure of the feature weaving network 50. Figure 2B An example of the structure of the feature weaving layer 500 of the present embodiment is schematically illustrated. Each feature weaving layer 500 is configured to receive the input of a three-layer first input tensor (Z A l ) and a second input tensor (Z B l ). The suffix "l" corresponds to the layer number of the feature weaving layer 500. That is, each input tensor (Z A l , Z B l)Corresponds to the input to the (l + 1)-th feature weaving layer 500. The suffix "l" of the input tensor substitutes integers from 0 to L - 1. L represents the number of feature weaving layers 500. As long as the value of L is 2 or more, it can be appropriately set according to the embodiment.
[0119] The first input tensor (Z A l ) is configured to have: a first axis arranged with the first vertices belonging to the first set (A) of the input graph as elements, a second axis arranged with the second vertices belonging to the second set (B) of the input graph as elements, and a feature quantity (D l dimensions) as an element and arranged on the third axis related to the branches flowing from the first vertex to the second vertex. In this embodiment, for the sake of explanation, it is assumed that the number of first vertices is N, the number of second vertices is M, and the dimension of the feature quantity is D l . The values of N, M, and D l can be appropriately set according to the embodiment. In addition, for convenience, it is assumed that N is 4 and M is 3 in the figure. The second input tensor (Z B l ) is configured to have: a first axis arranged with the second vertices belonging to the second set (B) as elements, a second axis arranged with the first vertices belonging to the first set (A) as elements, and a feature quantity (D l dimensions) as an element and arranged on the third axis related to the branches flowing from the second vertex to the first vertex.
[0120] Each feature weaving layer 500 includes an encoder 505 and is configured to perform the following operations. That is, each feature weaving layer 500 fixes the third axis of the second input tensor and makes the other axes except the third axis of the second input tensor cycle one by one In this embodiment, since the number of axes of the second input tensor (Z B l ) is three, making the other axes except the third axis cycle one by one is equivalent to swapping the first axis and the second axis. When understood from the elements of the first axis, each element arranged on the third axis of the second input tensor whose axis is cycled corresponds to the feature quantity related to the branches flowing from the second vertex into the first vertex. In addition to this, each feature weaving layer 500 concatenates (cat operation) the feature quantities of each element of the first input tensor (Z A l ) and the second input tensor whose axis is cycled . An example of the cat operation is: combining at the end of the third axis of the first input tensor (Z A l ) with the second input tensor whose axis is cycled Through these operations, each characteristic braided layer 500 generates the first connection tensor of the three layers.
[0121] Moreover, each characteristic braided layer 500 fixes the third axis of the first input tensor and makes the other axes other than the third axis of the first input tensor circulate one by one. Similar to the operation of the second input tensor, looping through the axes other than the third axis of the first input tensor is equivalent to swapping the first axis and the second axis of the first input tensor. Each element arranged on the third axis of corresponds to a feature quantity associated with a branch flowing from the first vertex to the second vertex. In addition, each feature braided layer 500 converts the second input tensor (Z B l ) and the first input tensor to loop over The feature quantities of each element of are concatenated (cat operation). An example of cat operation is: B l )'s third axis is combined with the first input tensor to loop over the axis Through these operations, each characteristic braided layer 500 generates a second connection tensor of three layers.
[0122] Then, each characteristic braided layer 500 divides the generated first connection tensor into each element of the first axis, and obtains each element (z ai l ∈R 1×M×2Dl ) is input to the encoder 505, and the encoder 505 performs the operation. Each feature braided layer 500 obtains the operation result of the encoder 505 for each element, thereby generating a tensor (Z A l ) corresponds to the first output tensor (Z A l+1 ∈R N×M×Dl+1 ). In addition, each characteristic braided layer 500 divides the generated second connection tensor into each element of the first axis, and obtains each element (z bj l ∈R 1×N×2Dl ) is input to the encoder 505, and the encoder 505 performs the operation. Each feature braided layer 500 obtains the operation result of the encoder 505 for each element, thereby generating a second input tensor (Z B l ) corresponds to the second output tensor (Z B l+1 ∈RM×N×Dl+1 )。The dimension (D l+1 ) of the elements on the third axis of each output tensor, i.e., the feature quantity, can be appropriately determined according to the implementation manner, and can be the same as or different from the dimension (D l ) before encoding. Each feature weaving layer 500 is configured to perform each of the above operations.
[0123] The encoder 505 is configured to derive the relative feature quantity of each element from the feature quantities of all the input elements. That is, each element on the first axis of the object connection tensor corresponds to each vertex of the object set (the first vertex a i , the second vertex b j ), and the object connection tensor represents a state in which the feature quantities of the outflow branches and inflow branches with respect to each vertex are connected. The encoder 505 is configured to derive the relative feature quantity of the object outflow branch from all the feature quantities related to the outflow branches and inflow branches with respect to the object vertex. The structure of the encoder 505 is not particularly limited as long as it can perform such an operation, and can be appropriately selected according to the implementation manner. Specific examples will be described later.
[0124] As Figure 2B shown, in an example of the present embodiment, the encoder 505 of each feature weaving layer 500 may include a single encoder commonly used among the elements (z ai l , z bj l ) on the first axis of the first connection tensor and the second connection tensor. That is, the feature weaving network 50 can be configured to include a single encoder for each feature weaving layer 500 (including L encoders for the L-layer feature weaving layer 500). Thereby, the number of parameters of the inference module 5 can be reduced. As a result, the calculation amount of the inference module 5 can be suppressed, and the efficiency of machine learning can be achieved.
[0125] However, the number of encoders 505 constituting the feature weaving layer 500 is not limited to such an example. In another example, a common encoder can be used among the elements (z ai l ) on the first axis of the first connection tensor, and among the elements (z bj l) Another common encoder is used during that time. That is, the encoder 505 may include two encoders respectively used by each connection tensor. Thus, it is possible to encode the feature amounts of the branches seen from the first vertex and the feature amounts of the branches seen from the second vertex respectively, and thus it is possible to reflect the respective individual features through inference. As a result, an improvement in the accuracy of the speculation module can be expected. When the distributions of the feature amounts related to the first set and the second set are different, the encoder 505 can be set separately with respect to the first connection tensor and the second connection tensor. Furthermore, in another example, the encoder 505 can be set separately corresponding to each element (z ai l 、z bj l ) of the first connection tensor and the second connection tensor.
[0126] As Figure 2A shown, the first feature weaving layer (that is, the feature weaving layer arranged closest to the input side) 5001 among the multiple feature weaving layers 500 is configured to receive a first input tensor 601 (Z A 0) and a second input tensor 602 (Z B 0) from the input graph. In one example, the input graph is given in a general graph format, and each input tensor (601, 602) can be directly obtained from the given input graph. In another example, each input tensor (601, 602) can be obtained from the result of a prescribed arithmetic process performed on the input graph. The prescribed arithmetic process can be performed by an arithmetic unit with an arbitrary structure such as a convolutional layer, for example. That is, an arbitrary arithmetic unit can also be provided before the feature weaving network 50. Furthermore, in another example, the input graph can also be directly obtained in the format of each input tensor (601, 602).
[0127] In one example, the feature weaving layers 500 can be continuously arranged in series. At this time, other feature weaving layers (that is, the second and subsequent feature weaving layers) 500 other than the first feature weaving layer 5001 are configured to receive each output tensor generated by the feature weaving layer 500 arranged immediately before it as each input tensor. That is, the feature weaving layer 500 arranged at the (l + 1)-th position is configured to receive each output tensor (Z A l+1 、Z B l+1 ) generated by the l-th feature weaving layer as each input tensor (where l is an integer from 1 to L - 1).
[0128] However, the configuration of each feature weaving layer 500 is not limited to such an example. For example, other types of layers (such as a convolutional layer, a fully connected layer, etc.) may be inserted between adjacent feature weaving layers 500. At this time, the inserted other layers are configured to perform arithmetic processing on each output tensor obtained from the feature weaving layer 500 configured immediately before it. The feature weaving layer 500 configured after the inserted other layers is configured to obtain each input tensor from the arithmetic result of the other layer. The number of the inserted other layers is not particularly limited and can be appropriately selected according to the embodiment. The structure of the inserted other layers can be appropriately determined according to the embodiment. In the case where other types of layers are inserted at multiple positions, the structure of each layer can be the same or different depending on the inserted positions.
[0129] The feature information includes a first output tensor 651(Z A L ) and a second output tensor 652(Z B L ) output from the last feature weaving layer (i.e., the feature weaving layer configured closest to the output side) 500L among the multiple feature weaving layers 500. The speculator 55 is appropriately configured to derive a result for speculating on the solution to the task from each output tensor (651, 652). The structure of the speculator 55 is not particularly limited as long as it can perform such arithmetic, and can be appropriately determined according to the embodiment (such as the structure of each output tensor, the derivation format of the speculation result, etc.).
[0130] The speculator 55 may include, for example, a data table, a functional formula, rules, etc. In an example of the case where a functional formula is included, the speculator 55 may include, for example, a neural network having an arbitrary structure such as a single convolutional layer and a single fully connected layer. At this time, the speculator 55 may include multiple nodes (neurons), and the weights of the connections between the nodes and the thresholds of the nodes are an example of the parameters of the speculator 55. The threshold of each node may also include an arbitrary activation function. Moreover, the output format of the speculator 55 is not particularly limited and can be appropriately selected according to the embodiment. In one example, the output of the speculator 55 may be configured to directly represent the speculation result. In another example, the output of the speculator 55 may be configured to indirectly represent the speculation result. At this time, by performing a prescribed arithmetic process such as a threshold determination on the output of the speculator 55, the speculation result can be derived from the output of the speculator 55.
[0131] In addition, as Figure 2A shown, the feature weaving network 50 may be configured to have a residual structure. The so-called residual structure refers to a structure having a shortcut as described below, that is, the output of at least one of the multiple feature weaving layers is input not only to the next adjacent feature weaving layer but also to other feature weaving layers (feature weaving layers separated by two or more layers) after the next adjacent feature weaving layer.
[0132] Figure 2A In the example of Figure 2A , the feature weaving network 50 has a shortcut 501 that omits two layers of the feature weaving layer 500. Thus, the output of the feature weaving layer 500 immediately before the shortcut 501 is input to the feature weaving layer 500 three layers ahead. Figure 2A In the example of Figure 2A , for example, a shortcut 501 is provided on the output side of the first feature weaving layer 5001. Therefore, the fourth feature weaving layer 500 is configured to obtain an input tensor from the output tensor of the third feature weaving layer 500 and the output tensor of the first feature weaving layer 5001.
[0133] However, the number of feature weaving layers 500 omitted by the shortcut 501 is not limited to such an example. The number of feature weaving layers 500 omitted by the shortcut 501 can be one layer, or can also be three or more layers. Moreover, the interval for setting the shortcut 501 is not particularly limited and can be appropriately determined according to the implementation. The shortcut 501 can be set, for example, at a specified interval such as every three layers.
[0134] By having such a residual structure, it is possible to easily deepen the hierarchy of the feature weaving network 50 (that is, to easily increase the number of feature weaving layers 500). By deepening the hierarchy of the feature weaving network 50, it is possible to improve the accuracy of feature extraction performed by the feature weaving network 50. As a result, it is possible to achieve an improvement in the prediction accuracy of the prediction module 5. In addition, the structure of the feature weaving network 50 is not limited to such an example. The residual structure can also be omitted.
[0135] As Figure 1 and Figure 2A shown, in the model generation device 1, the machine learning is configured in such a way that each training graph 30 is input as an input graph to the feature weaving network 50, thereby training the prediction module 5 so that the prediction result obtained by the predictor 55 conforms to the correct answer (true value) of the task for each training graph 30. On the other hand, in the prediction device 2, the process of predicting the solution to the task for the object graph 221 using the trained prediction module 5 is configured in such a way that the object graph 221 is input as an input graph to the feature weaving network 50, and the result of predicting the solution to the task is obtained from the predictor 55.
[0136] (Feature)
[0137] As described above, in the feature weaving network 50 of the inference module 5 of the present embodiment, instead of referring to the features of the input graph for each vertex, the features of the input graph are referred to for each branch. In each feature weaving layer 500, through the encoder 505, the feature amounts related to other branches are reflected in the feature amounts related to each branch based on each vertex. The more the feature weaving layers 500 are stacked, the more the feature amounts of the branches related to the vertices departing from each vertex can be reflected in the feature amounts of the branches related to each vertex. In addition, the encoder 505 is configured to derive the relative feature amounts of each branch from the feature amounts of all the input branches for the target vertex, instead of calculating the weighted sum of the features of the adjacent branches for each vertex. According to the structure of the encoder 505, it is possible to prevent excessive smoothing even when the feature weaving layers 500 are stacked. Therefore, the hierarchy of the feature weaving network 50 can be deepened according to the difficulty of the task. As a result, in the feature weaving network 50, the features of the input graph appropriate for the solution task can be extracted. That is, an improvement in the accuracy of feature extraction performed by the feature weaving network 50 can be expected. Thus, in the inference unit 55, high-precision inference can be expected based on the obtained feature amounts. Therefore, according to the present embodiment, even for a complex task, the hierarchy of the feature weaving network 50 can be deepened, and accordingly, for the inference module 5, an improvement in the accuracy of inferring the solution to the task expressed by a graph can be achieved. In the model generation device 1, generation of a trained inference module 5 capable of accurately inferring the solution to the task of the input graph can be expected. In the inference device 2, by using such a trained inference module 5, high-precision inference of the solution to the task of the target graph 221 can be expected.
[0138] In addition, Figure 1 in an example, the model generation device 1 and the inference device 2 are connected to each other via a network. The type of the network can be appropriately selected from, for example, the Internet, a wireless communication network, a mobile communication network, a telephone network, a private network, and the like. However, the method of exchanging data between the model generation device 1 and the inference device 2 is not limited to such an example and can be appropriately selected according to the embodiment. For example, between the model generation device 1 and the inference device 2, data can be exchanged using a storage medium.
[0139] Moreover, Figure 1 in an example, the model generation device 1 and the inference device 2 each include separate computers. However, the structure of the inference system 100 of the present embodiment is not limited to such an example and can be appropriately determined according to the embodiment. For example, the model generation device 1 and the inference device 2 may also be an integrated computer. Further, for example, at least one of the model generation device 1 and the inference device 2 may include multiple computers.
[0140] §2 Structural Example
[0141] [Hardware Structure]
[0142] <Model Generation Device>
[0143] Figure 3 Schematically illustrate an example of the hardware structure of the model generation device 1 of the present embodiment. As Figure 3 shown, the model generation device 1 of the present embodiment is a computer electrically connected by a control unit 11, a storage unit 12, a communication interface 13, an external interface 14, an input device 15, an output device 16, and a driver 17. In addition, Figure 3 in this case, the communication interface and the external interface are referred to as "communication I / F" and "external I / F".
[0144] The control unit 11 includes a central processing unit (CPU) as a hardware processor, a random access memory (RAM), a read only memory (ROM), etc., and is configured to perform information processing based on programs and various data. The storage unit 12 is an example of a memory, and includes, for example, a hard disk drive, a solid state drive, etc. In the present embodiment, the storage unit 12 stores various information such as a model generation program 81, a plurality of training diagrams 30, and learning result data 125.
[0145] The model generation program 81 is a program for causing the model generation device 1 to execute the information processing ( Figure 8 ) of machine learning described later, and the machine learning generates a trained inference module 5. The model generation program 81 includes a series of commands for the information processing. The plurality of training diagrams 30 are used for the machine learning of the inference module 5. The learning result data 125 represents information related to the trained inference module 5 generated by the implementation of machine learning. In the present embodiment, the learning result data 125 is generated as a result of executing the model generation program 81. Details will be described later.
[0146] The communication interface 13 is, for example, a wired Local Area Network (LAN) module, a wireless LAN module, etc., and is an interface for performing wired or wireless communication via a network. The model generation device 1 can use the communication interface 13 to perform data communication via the network with other information processing devices. The external interface 14 is, for example, a Universal Serial Bus (USB) port, a dedicated port, etc., and is an interface for connecting to an external device. The type and number of the external interfaces 14 can be arbitrarily selected. The training diagram 30 can also be obtained, for example, by a sensor such as a camera. Alternatively, the training diagram 30 can also be generated by another computer. At this time, the model generation device 1 can be connected to the sensor or another computer via at least one of the communication interface 13 and the external interface 14.
[0147] The input device 15 is, for example, a device for input such as a mouse, a keyboard, etc. Moreover, the output device 16 is, for example, a device for output such as a display, a speaker, etc. An operator such as a user can operate the model generation device 1 by using the input device 15 and the output device 16. The training diagram 30 can also be obtained by input via the input device 15.
[0148] The drive 17 is, for example, a Compact Disc (CD) drive, a Digital Versatile Disc (DVD) drive, etc., and is a drive device for reading various information such as programs stored in the storage medium 91. The storage medium 91 is a medium that stores the programs and other information by electrical, magnetic, optical, mechanical, or chemical action in a manner that can be read by a computer, other devices, machines, etc. At least any one of the model generation program 81 and the plurality of training diagrams 30 can also be stored in the storage medium 91. The model generation device 1 can also obtain at least any one of the model generation program 81 and the plurality of training diagrams 30 from the storage medium 91. Additionally, Figure 3 In this case, as an example of the storage medium 91, disc-shaped storage media such as CDs and DVDs are illustrated. However, the type of the storage medium 91 may not be limited to disc-shaped, and may also be other than disc-shaped. As a storage medium other than disc-shaped, for example, a semiconductor memory such as a flash memory can be cited. The type of the drive 17 can be arbitrarily selected according to the type of the storage medium 91.
[0149] In addition, regarding the specific hardware structure of the model generation device 1, components can be appropriately omitted, replaced, or added according to the implementation. For example, the control unit 11 may also include multiple hardware processors. The hardware processors may include a microprocessor, a Field-Programmable Gate Array (FPGA), a Digital Signal Processor (DSP), etc. The storage unit 12 may also include the RAM and ROM included in the control unit 11. At least any one of the communication interface 13, the external interface 14, the input device 15, the output device 16, and the driver 17 may be omitted. The model generation device 1 may also include multiple computers. In this case, the hardware structures of the respective computers may be the same or different. Moreover, the model generation device 1 may be a general server device, a Personal Computer (PC), etc., in addition to the information processing device designed specifically for the provided service.
[0150] <Inference Device>
[0151] Figure 4 An example of the hardware structure of the inference device 2 of the present embodiment is schematically illustrated. As Figure 4 shown, the inference device 2 of the present embodiment is a computer formed by electrically connecting a control unit 21, a storage unit 22, a communication interface 23, an external interface 24, an input device 25, an output device 26, and a driver 27.
[0152] The control unit 21 to the driver 27 of the inference device 2 and the storage medium 92 can be respectively configured in the same manner as the control unit 11 to the driver 17 and the storage medium 91 of the model generation device 1. The control unit 21 includes a CPU, a RAM, a ROM, etc. as hardware processors, and is configured to execute various information processes based on programs and data. The storage unit 22 includes, for example, a hard disk drive, a solid state drive, etc. In the present embodiment, the storage unit 22 stores various information such as an inference program 82 and learning result data 125.
[0153] The inference program 82 is a program for causing the inference device 2 to execute an information process ( Figure 9 ) described later for performing an inference task using the trained inference module 5. The inference program 82 includes a series of commands for the information process. At least any one of the inference program 82 and the learning result data 125 may be stored in the storage medium 92. Moreover, the inference device 2 may obtain at least any one of the inference program 82 and the learning result data 125 from the storage medium 92.
[0154] Similar to the training diagram 30, the object diagram 221 can also be obtained by a sensor such as a camera, for example. Alternatively, the object diagram 221 can also be generated by another computer. At this time, the speculation device 2 can be connected to the sensor or another computer via at least one of the communication interface 23 and the external interface 24. Alternatively, the object diagram 221 can also be obtained by input via the input device 25.
[0155] In addition, regarding the specific hardware structure of the speculation device 2, components can be appropriately omitted, replaced, or added according to the embodiment. For example, the control unit 21 can also include multiple hardware processors. The hardware processors can include microprocessors, FPGAs, DSPs, etc. The storage unit 22 can also include the RAM and ROM included in the control unit 21. At least any one of the communication interface 23, the external interface 24, the input device 25, the output device 26, and the driver 27 can also be omitted. The speculation device 2 can also include multiple computers. At this time, the hardware structures of each computer can be the same or different. Moreover, the speculation device 2 can be a general server device, a general PC, a tablet PC, a terminal device, etc., in addition to the information processing device designed specifically for the provided service.
[0156] [Software Structure]
[0157] <Model Generation Device>
[0158] Figure 5 Schematically illustrate an example of the software structure of the model generation device 1 of this embodiment. The control unit 11 of the model generation device 1 expands the model generation program 81 stored in the storage unit 12 into the RAM. And the control unit 11 interprets and executes the commands included in the model generation program 81 expanded in the RAM through the CPU to control each component. Thus, as Figure 5 shown, the model generation device 1 of this embodiment operates as a computer including an acquisition unit 111, a learning processing unit 112, and a saving processing unit 113 as software modules. That is, in this embodiment, each software module of the model generation device 1 is implemented by the control unit 11 (CPU).
[0159] The acquisition unit 111 is configured to acquire a plurality of training graphs 30. The learning processing unit 112 is configured to perform machine learning of the inference module 5 using the acquired plurality of training graphs 30. The inference module 5 includes a feature weaving network 50 and an inferencer 55. The machine learning is configured such that each training graph 30 is input as an input graph to the feature weaving network 50, thereby training the inference module 5 so that the inference result obtained by the inferencer 55 conforms to the correct answer (true value) of the task for each training graph 30. The correct answer (true value) can be given, for example, by blocking pairs in a case without a stable matching problem, satisfying a specified rule such as a specified scale, or the like. Alternatively, for each acquired training graph 30, a correct answer label (teaching signal) can be associated. Each correct answer label can be configured to represent the correct answer (true value) of the corresponding training graph 30.
[0160] The saving processing unit 113 is configured to generate information related to the trained inference module 5 generated by machine learning as learning result data 125, and save the generated learning result data 125 to a specified storage area. The learning result data 125 can be appropriately configured to include information for reproducing the trained inference module 5.
[0161] (Encoder)
[0162] As Figure 2A and Figure 2B shown, the feature weaving network 50 includes a plurality of feature weaving layers 500, and each feature weaving layer 500 includes an encoder 505. The structure of the encoder 505 can be appropriately selected according to the embodiment as long as the relative feature amounts of the respective elements can be derived from the feature amounts of all the input elements. In order to cope with the case where the number of elements connecting tensors varies due to differences such as the number of vertices of the input graph and the presence or absence of branches, it is preferable that the encoder 505 is configured to accept variable-length inputs. In the present embodiment, as the structure of the encoder 505, any one of the following two structural examples can be adopted.
[0163] (1) First structural example
[0164] Figure 6A An example of the structure of the encoder 5051 of the present embodiment is schematically illustrated. The encoder 5051 is an example of the encoder 505. In the first structural example, the encoder 5051 includes a bidirectional long short-term memory.
[0165] The bidirectional LSTM includes a plurality of nodes (neurons). The weights of the connections between the respective nodes and the thresholds of the respective nodes are examples of the parameters of the encoder 5051. The threshold of each node may also include any activation function. An example of the specific structure of the bidirectional LSTM is proposed, for example, in the literature: Alex Graves and Jurgen Schmidhuber, "Framewise phoneme classification with bidirectional LSTM and other neural network architectures", Neural Networks, 2005.
[0166] Thus, the encoder 5051 is configured to receive, for each element corresponding to the second axis, each feature amount arranged on the third axis as a variable-length sequence, bidirectionally depend on the previously input feature amount and the subsequently input feature amount, and derive, for each element corresponding to the second axis, each feature amount arranged on the third axis of the output tensor. According to the encoder 5051, it is possible to receive an input of a variable-length concatenated tensor derived from a graph having an arbitrary structure, and appropriately derive the relative feature amount of the object outflow branch from all the feature amounts related to the outflow branches and inflow branches with respect to the object vertex.
[0167] (2) Second structural example
[0168] Figure 6B An example of the structure of the encoder 5052 of the present embodiment is schematically illustrated. The encoder 5052 is an example of the encoder 505. In the second structural example, the encoder 5052 includes a first combiner 5054, an integrator 5055, and a second combiner 5056.
[0169] The first combiner 5054 is configured to combine the feature quantities 6051 arranged on the third axis for each element corresponding to the second axis of the connection tensor (i.e., combine the data from each input tensor). The integrator 5055 is configured to integrate the combination results 6054 of the first combiner 5054 for each element of the second axis. In one example, the integrator 5055 may be configured to calculate statistical information such as the maximum value, minimum value, average value, etc. for each element of the third axis (feature quantity). The second combiner 5056 is configured to derive, for each element corresponding to the second axis, the respective feature quantities 6056 arranged on the third axis of the output tensor from the feature quantities 6051 arranged on the third axis for each element corresponding to the second axis and the integration result 6055 obtained by the integrator 5055. In one example, each combiner (5054, 5056) may be provided respectively for each element corresponding to the second axis. In another example, a single one of each combiner (5054, 5056) may also be shared among the elements of the second axis.
[0170] As long as the operations can be executed, the structures of the first combiner 5054, the integrator 5055, and the second combiner 5056 are not particularly limited and can be appropriately selected according to the implementation manner. Each combiner (5054, 5056) may include, for example, any functional form, etc. In one example, the parameters constituting the functional form correspond to the parameters of each combiner (5054, 5056). The parameters of the first combiner 5054 and the second combiner 5056 are an example of the parameters of the encoder 5052.
[0171] According to the encoder 5052, by combining the integration result 6055 obtained by the integrator 5055 with the feature quantity 6051 of each element of the second axis through the second combiner 5056, it is possible to appropriately derive the relative feature quantity of the object outflow branch from all the feature quantities related to the outflow branches and inflow branches with respect to the object vertex. Moreover, each combiner (5054, 5056) is configured to be able to execute operations independently for each element, and the operation of the integrator 5055 is configured to be applicable to a variable-length element sequence. Thus, the encoder 5052 can be configured to receive a variable-length input.
[0172] (Machine learning)
[0173] As Figure 5 shown, by inputting each training graph 30 as an input graph into the inference module 5 and performing the forward arithmetic processing of the feature weaving network 50 and the inferencer 55, the inference result for each training graph 30 can be obtained. The learning processing unit 112 is configured to adjust the parameters of the inference module 5 (such as the parameters of the encoder 505 and the inferencer 55 of each of the feature weaving layers 500) in machine learning so that the error between the inference result obtained for each training graph 30 and the correct answer becomes smaller.
[0174] The saving processing unit 113 is configured to generate learning result data 125 for reproducing the trained inference module 5 generated by the machine learning. As long as the trained inference module 5 can be reproduced, the structure of the learning result data 125 is not particularly limited and can be appropriately determined according to the embodiment. As an example, the learning result data 125 may include information indicating the values of the respective parameters obtained by the adjustment of the machine learning. Depending on the situation, the learning result data 125 may include information indicating the structure of the inference module 5. The structure can be determined, for example, according to the number of layers, the type of each layer, the number of nodes included in each layer, the connection relationship between the nodes of adjacent layers, and the like.
[0175] <Inference Device>
[0176] Figure 7 An example of the software structure of the inference device 2 of the present embodiment is schematically illustrated. The control unit 21 of the inference device 2 expands the inference program 82 stored in the storage unit 22 into the RAM. And the control unit 21 interprets and executes the commands included in the inference program 82 expanded in the RAM by the CPU to control each component. Thus, as Figure 7 shown, the inference device 2 of the present embodiment operates as a computer including an acquisition unit 211, an inference unit 212, and an output unit 213 as software modules. That is, in the present embodiment, each software module of the inference device 2 is also implemented by the control unit 21 (CPU) in the same manner as the model generation device 1.
[0177] The acquisition unit 211 is configured to acquire the object graph 221. The inference unit 212 includes the trained inference module 5 by holding the learning result data 125. The inference unit 212 is configured to use the trained inference module 5 to infer the solution to the task of the acquired object graph 221. The process of using the trained inference module 5 to infer the solution to the task of the object graph 221 is configured as follows: the object graph 221 is input as an input graph to the feature weaving network 50, and the result of inferring the solution to the task is obtained from the inferencer 55. The output unit 213 is configured to output information related to the result of inferring the solution to the task.
[0178] <Others>
[0179] Regarding each software module of the model generation device 1 and the inference device 2, it will be described in detail using the operation examples described later. In addition, in the present embodiment, an example in which each software module of the model generation device 1 and the inference device 2 is implemented by a general-purpose CPU has been described. However, part or all of the software modules may also be implemented by one or more dedicated processors (e.g., a graphics processing unit). Each module may also be implemented as a hardware module. Moreover, regarding the software structure of the model generation device 1 and the inference device 2 respectively, the omission, replacement, and addition of software modules may be appropriately performed according to the embodiment.
[0180] §3 Operation Example
[0181] [Model Generation Device]
[0182] Figure 8 It is a flowchart showing an example of the processing flow related to the machine learning performed by the model generation device 1 of the present embodiment. The processing flow of the model generation device 1 described below is an example of the model generation method. However, the processing flow of the model generation device 1 described below is merely an example, and each step can be changed as much as possible. Moreover, for the following processing flow, the omission, replacement, and addition of steps can be appropriately performed according to the embodiment.
[0183] (Step S101)
[0184] In step S101, the control unit 11 operates as the acquisition unit 111 and acquires a plurality of training graphs 30.
[0185] Each training graph 30 can be appropriately generated according to the task capabilities obtained by the inference module 5. The conditions for the object of the solving task can be appropriately given, and each training graph 30 can be generated according to the given conditions. Moreover, each training graph 30 can be obtained according to existing data representing an existing transportation network, an existing communication network, etc. In addition to this, each training graph 30 can be obtained according to image data. The image data can be obtained either by a camera or can be appropriately generated by a computer.
[0186] The data format of each training graph 30 is not particularly limited and can be appropriately selected according to the embodiment. In one example, each training graph 30 may include an adjacency list or an adjacency matrix. In another example, each training graph 30 may include other data formats (e.g., image data, etc.) other than the adjacency list and the adjacency matrix. As an example of other data formats, each training graph 30 may include image data. At this time, each pixel may correspond to a vertex, and the relationship between pixels may correspond to an edge.
[0187] The correct answer (true value) for each training figure 30 can be appropriately given. In one example, the correct answer (true value) can be given according to a specified rule. In another example, the correct answer (true value) can be represented by a correct answer label (teaching signal). At this time, correct answer labels corresponding to each training figure 30 can be appropriately generated, and the generated correct answer labels can be associated with each training figure 30. Thus, each training figure 30 can also be generated in the format of a data set associated with a correct answer label.
[0188] Each training figure 30 can be automatically generated by the operation of a computer, or can also be manually generated by at least partially including the operation of an operator. Moreover, the generation of each training figure 30 can be performed by the model generation device 1, or can also be performed by other computers other than the model generation device 1. That is, the control unit 11 can automatically or manually generate each training figure 30 through the operation of an operator via the input device 15. Alternatively, the control unit 11 can, for example, obtain each training figure 30 generated by other computers via a network, a storage medium 91, etc. A part of the plurality of training figures 30 can be generated by the model generation device 1, and the others can be generated by one or more other computers. At least any one of the plurality of training figures 30 can also be generated, for example, by a generation model including a machine learning model (for example, the generation model included in an adversarial generation network).
[0189] The number of the obtained training figures 30 is not particularly limited and can be appropriately determined according to the implementation manner in a manner capable of implementing machine learning. When a plurality of training figures 30 are obtained, the control unit 11 advances the process to the next step S102.
[0190] (Step S102)
[0191] In step S102, the control unit 11 operates as the learning processing unit 112 and performs machine learning of the inference module 5 using the obtained plurality of training figures 30.
[0192] As an example of the machine learning process, first, the control unit 11 performs an initial setting of the inference module 5 that is the object of the machine learning process. The initial values of the structure and parameters of the inference module 5 can be given by a template, or can also be determined by the input of an operator. Moreover, in the case of performing additional learning or re-learning, the control unit 11 can also perform the initial setting of the inference module 5 based on the learning result data obtained through machine learning in the past.
[0193] Next, the control unit 11 trains the inference module 5 through machine learning (that is, adjusts the values of the parameters of the inference module 5) so that the result of inferring the solution to the task for each training figure 30 conforms to the correct answer (true value). For the training process, a probability gradient descent method, a mini-batch gradient descent method, etc. can be used.
[0194] As an example of the training process, first, the control unit 11 inputs each training graph 30 to the inference module 5 and performs forward arithmetic processing. In one example, the control unit 11 obtains each input tensor (601, 602) from each training graph 30 and inputs the obtained input tensors (601, 602) to the outermost feature knitting layer 5001 of the feature knitting network 50 configured on the input side. Next, the control unit 11 performs the arithmetic processing of each feature knitting layer 500 forward. In the arithmetic processing of each feature knitting layer 500, the control unit 11 generates each connection tensor from each input tensor. Subsequently, the control unit 11 uses the encoder 505 to generate each output tensor from each connection tensor. In the present embodiment, either of the two encoders (5051, 5052) can be used as the encoder 505. As a result of performing the arithmetic processing of each feature knitting layer 500 forward, the control unit 11 obtains each output tensor (651, 652) from the last feature knitting layer 500L. And the control unit 11 inputs the obtained output tensors (651, 652) to the predictor 55 and performs the arithmetic processing of the predictor 55. Through the above series of forward arithmetic processing, the control unit 11 obtains the inference result of the solution to the task of each training graph 30 from the predictor 55.
[0195] Next, the control unit 11 calculates the error between the obtained inference result and the corresponding correct solution. As described above, the correct solution can be given according to a specified rule or can be given through the corresponding correct solution label. For calculating the error, a loss function can be used. The loss function can be appropriately set according to, for example, the task, the format of the correct solution, etc. Subsequently, the control unit 11 calculates the gradient of the calculated error. The control unit 11 uses the calculated gradient of the error and calculates, in order from the output side, the error of the parameter values of the inference module 5 including the predictor 55 and the encoder 505 of each feature knitting layer 500 by the error backpropagation method. And the control unit 11 updates the parameter values of the inference module 5 based on the calculated errors. The degree of updating the parameter values can be adjusted by the learning rate. The learning rate can be given either by the designation of an operator or as a set value within the program.
[0196] The control unit 11 adjusts the values of the respective parameters through the series of update processes so that the sum of the errors calculated for each training graph 30 becomes smaller. For example, the control unit 11 may repeat the adjustment of the values of the respective parameters performed through the series of update processes until a specified condition is satisfied, such as the sum of the calculated errors becoming below a threshold after a specified number of executions. As a result of the machine learning process, the control unit 11 may generate a trained inference module 5 that has acquired the ability to perform a desired inference task (i.e., obtain a solution to a given graph inference task) corresponding to the training graph 30 used. When the machine learning process is completed, the control unit 11 advances the process to the next step S103.
[0197] (Step S103)
[0198] In step S103, the control unit 11 operates as a save processing unit 113 and generates information related to the trained inference module 5 generated by machine learning as learning result data 125. Then, the control unit 11 saves the generated learning result data 125 to a specified storage area.
[0199] The specified storage area may be, for example, a RAM in the control unit 11, the storage unit 12, an external storage device, a storage medium, or a combination thereof. The storage medium may be, for example, a CD, a DVD, etc., and the control unit 11 may also save the learning result data 125 to the storage medium via a drive 17. The external storage device may be, for example, a data server such as a Network Attached Storage (NAS). In this case, the control unit 11 may also use the communication interface 13 to save the learning result data 125 to the data server via a network. Moreover, the external storage device may also be an external storage device connected to the model generation device 1 via an external interface 14.
[0200] When the saving of the learning result data 125 is completed, the control unit 11 ends the processing flow of the model generation device 1 related to this operation example.
[0201] In addition, the generated learning result data 125 may be provided to the inference device 2 at an arbitrary timing. For example, the control unit 11 may also forward the learning result data 125 to the inference device 2 as a process of step S103 or independently of the process of step S103. The inference device 2 may also obtain the learning result data 125 by receiving the forwarding. Moreover, for example, the inference device 2 may use the communication interface 23 to access the model generation device 1 or the data server via a network to obtain the learning result data 125. Moreover, for example, the inference device 2 may obtain the learning result data 125 via a storage medium 92. Moreover, for example, the learning result data 125 may be pre-loaded into the inference device 2.
[0202] Furthermore, the control unit 11 can also update or newly generate the learning result data 125 by periodically or aperiodically repeating the processes of the steps S101 to S103. During the repetition, at least a part of the training graph 30 for machine learning can be appropriately changed, corrected, added, deleted, etc. Also, the control unit 11 can use any method to provide the updated or newly generated learning result data 125 to the inference device 2, thereby updating the learning result data 125 held by the inference device 2.
[0203] [Inference device]
[0204] Figure 9 FIG. 9 is a flowchart showing an example of a processing flow related to the execution of an inference task by the inference device 2 of the present embodiment. The processing flow of the inference device 2 described below is an example of an inference method. However, the processing flow of the inference device 2 described below is merely an example, and each step can be changed as much as possible. Also, for the following processing flow, steps can be appropriately omitted, replaced, and added according to the embodiment.
[0205] (Step S201)
[0206] In step S201, the control unit 21 operates as an acquisition unit 211 and acquires an object graph 221 that is the object of the inference task. The structure of the object graph 221 is the same as that of the training graph 30. The data format of the object graph 221 can be appropriately selected according to the embodiment. The object graph 221 can be generated, for example, according to conditions input via the input device 25 or the like. Alternatively, the object graph 221 can be obtained from existing data. In addition, the object graph 221 can be obtained from image data. The control unit 21 can also directly acquire the object graph 221. Alternatively, the control unit 21 can indirectly acquire the object graph 221 via a network, a sensor, another computer, a storage medium 92, etc., for example. When the object graph 221 is acquired, the control unit 21 advances the processing to the next step S202.
[0207] (Step S202)
[0208] In step S202, the control unit 21 operates as the estimation unit 212, and sets the trained estimation module 5 with reference to the learning result data 125. Then, the control unit 21 uses the trained estimation module 5 to estimate the solution to the task of the acquired object graph 221. The arithmetic processing of the estimation can be the same as the forward arithmetic processing in the training process of the machine learning. The control unit 21 inputs the object graph 221 to the trained estimation module 5 and executes the forward arithmetic processing of the trained estimation module 5. As a result of executing the arithmetic processing, the control unit 21 can obtain, from the estimator 55, the result of estimating the solution to the task of the object graph 221. When the estimation result is obtained, the control unit 21 advances the process to the next step S203.
[0209] (Step S203)
[0210] In step S203, the control unit 21 operates as the output unit 213 and outputs information related to the estimation result.
[0211] The output target and the content of the output information can be appropriately determined according to the embodiment. For example, the control unit 21 may directly output the estimation result obtained through step S202 to the output device 26 or the output device of another computer. Moreover, the control unit 21 may perform certain information processing based on the obtained estimation result. And the control unit 21 may output the result of performing the information processing as the information related to the estimation result. In the output of the result of performing the information processing, it may include controlling the operation of the controlled device according to the estimation result. The output target may be, for example, the output device 26, the output device of another computer, the controlled device, etc.
[0212] When the output of the information related to the estimation result is completed, the control unit 21 ends the processing flow of the estimation device 2 related to this operation example. In addition, the control unit 21 may continuously repeat the series of information processing of steps S201 to S203. The timing of repetition can be appropriately determined according to the embodiment. Thus, the estimation device 2 can be configured to continuously repeat the estimation task.
[0213] [Features]
[0214] As described above, in the feature weaving network 50 of the estimation module 5 according to this embodiment, it is configured to refer to the features of the input graph not for each vertex but for each branch. In each feature weaving layer 500 of the feature weaving network 50, through the encoder 505, the feature amounts related to other branches are reflected in the feature amounts related to each branch based on each vertex.
[0215] Figure 10Schematically illustrate an example of the case where two feature weaving layers 500 of the present embodiment are used to reflect the feature amounts related to other branches. Before being input to the feature weaving layer 500, the first input tensor (Z A la ) represents the feature amounts related to the branch flowing out from the first vertex to the second vertex based on the first vertex. The second input tensor (Z B la ) represents the feature amounts related to the branch flowing out from the second vertex to the first vertex based on the second vertex. The first connection tensor represents the state of combining the feature amounts related to the branch flowing out from the first vertex to the second vertex and the branch flowing into the second vertex.
[0216] When encoding the first connection tensor in the encoder 505(1) of the first feature weaving layer 500, the feature amounts of the branch flowing out to the second vertex as the object and the branch flowing into the second vertex as the object are woven, and the relationship with other second vertices (that is, the feature amounts of the branch flowing out to other second vertices and the branch flowing into other second vertices) is reflected in the weaving of the feature amounts. Therefore, each element of the first output tensor (Z A la+1 ) obtained through the encoder 505(1) represents the following feature amounts, that is: regarding the first vertex as the object, on the basis of reflecting the relationship with other second vertices, the feature amounts of the branch flowing out to the second vertex as the object and the branch flowing into the second vertex as the object are woven, and the feature amounts derived therefrom. Similarly, each element of the second output tensor (Z B la+1 ) represents the following feature amounts, that is: regarding the second vertex as the object, on the basis of reflecting the relationship with other first vertices, the feature amounts of the branch flowing out to the first vertex as the object and the branch flowing into the first vertex as the object are woven, and the feature amounts derived therefrom.
[0217] Before passing through the second feature weaving layer 500, each input tensor is derived from each output tensor (Z A la+1 , Z B la+1 ), and after the axes are cycled, the input tensors are connected. Thus, each connection tensor is generated. The first connection tensor at this stage represents the state of combining the feature amounts related to the branch flowing out from the first vertex to the second vertex and the branch flowing into the second vertex, and the relationship with other second vertices based on the first vertex and the relationship with other first vertices based on the second vertex are reflected in each feature amount. Therefore, the first output tensor (Z Ala+2 ) Each element represents a feature quantity as described below, that is: regarding the first vertex as the object, on the basis of reflecting the relationship with other second vertices and the relationship with other first vertices based on each second vertex, further weaving the feature quantity of the branch flowing out to the second vertex as the object and the branch flowing into the second vertex as the object, and thus deriving the feature quantity.
[0218] That is, in the feature quantity represented by each element of the first output tensor (Z A la+2 ) and the feature quantity related to the branch between the first vertex as the object and the second vertex as the object, it reflects the feature quantity related to the branch within a range of size 2 from the first vertex as the object (that is, the branch between the first vertex as the object and other second vertices, and the branch between each second vertex and other first vertices). As Figure 10 shown, the same is true in the process of backpropagation of the machine learning. In the process of backpropagation, the object element (diagonal element) of the first output tensor obtained through two feature weaving layers 500 is shown as the diagonal part of the figure. Every time a feature weaving layer 500 (encoder 505) is traced back, it expands to a range of size 1.
[0219] Therefore, the more the feature weaving layers 500 are stacked, the more the feature quantity of the branch related to the vertex away from each vertex can be reflected in the feature quantity of the branch related to each vertex. In addition, as described above, according to the number of the feature weaving layers 500, the range of other branches to be referred to expands. Therefore, preferably, the number of the feature weaving layers 500 is equal to or greater than the size of the input graph. For the size of the input graph, any two vertices can be selected, and it is calculated according to the maximum value of the branches passed on the shortest path from one of the selected two vertices to the other. By setting the number of the feature weaving layers 500 to be equal to or greater than the size of the input graph, the feature quantity related to all other branches (especially the feature quantity related to the branches far from the object branch in the graph) can be reflected in the feature quantity related to each branch. Thus, the features of each branch can be extracted from the whole graph. As an example, when the input graph corresponds to an image, the features of each region can be extracted on the basis of reflecting the relationship between distant regions.
[0220] In addition, the encoder 505 is configured to derive the relative feature amount of each branch from the feature amounts of all the input branches with respect to the target vertex, rather than calculating the weighted sum of the features of the adjacent branches for each vertex. According to the structure of the encoder 505, it is possible to prevent excessive smoothing even when the feature knitting layer 500 is superimposed. Therefore, the hierarchy of the feature knitting network 50 can be deepened according to the difficulty of the task. As a result, in the feature knitting network 50, it is possible to extract the features of the input graph appropriate for the solution task. That is, it is possible to expect an improvement in the accuracy of feature extraction performed by the feature knitting network 50. Thereby, in the estimator 55, it is possible to expect high-precision estimation based on the obtained feature amount.
[0221] Therefore, according to the present embodiment, even for a complex task, the hierarchy of the feature knitting network 50 can be deepened. Correspondingly, for the inference module 5, it is possible to improve the accuracy of inferring the solution to the task expressed by a graph. In the model generation device 1, through the processes of step S101 and step S102, it is possible to expect the generation of a trained inference module 5 that can accurately infer the solution to the task of the input graph. In the inference device 2, through the processes of step S201 to step S203, by using such a trained inference module 5, it is possible to expect high-precision inference of the solution to the task of the target graph 221.
[0222] §4 Variations
[0223] The embodiments of the present invention have been described in detail above, but the foregoing description is illustrative of the present invention in all aspects. Of course, various improvements or modifications can be made without departing from the scope of the present invention. For example, the following changes can be made. In addition, hereinafter, for the same components as those in the above embodiments, the same reference numerals are used, and the description of the same points as those in the above embodiments is appropriately omitted. The following variations can be combined as appropriate.
[0224] <4.1>
[0225] The inference system 100 of the above embodiment can be applied to scenarios for solving various tasks that can be expressed by a graph. The types of graphs (training graph 30, target graph 221) can be appropriately selected according to the task. For example, the graph can be a directed bipartite graph, an undirected bipartite graph, a general directed graph, a general undirected graph, a vertex feature graph, a hypergraph, etc. Tasks can be, for example, bipartite matching, unilateral matching, situation inference (including prediction), graph feature inference, in-graph search, vertex categorization (graph segmentation), inference of the relationship between vertices, etc. The following shows specific examples of limiting the applicable scenarios.
[0226] (A) Scenario using a directed bipartite graph
[0227] Figure 11 An example of an application scenario of the inference system 100A illustrating the first specific example schematically. The first specific example is an example in which the described embodiment is applied in a scenario where a directed bipartite graph is used as a graph expressing given conditions (objects of a solving task). The inference system 100A of the first specific example includes a model generation device 1 and an inference device 2A. The inference device 2A is an example of the inference device 2.
[0228] In the first specific example, each training graph 30A and object graph 221A (input graph) is a directed bipartite graph 70A. The multiple vertices constituting the directed bipartite graph 70A are divided into either of two subsets (71A, 72A) such that there are no branches between the vertices within each subset (71A, 72A). The first set of the input graph corresponds to one of the two subsets (71A, 72A) of the directed bipartite graph 70A, and the second set of the input graph corresponds to the other of the two subsets (71A, 72A) of the directed bipartite graph 70A. Each training graph 30A and object graph 221A is an example of each training graph 30 and object graph 221 in the described embodiment. In addition, the number of vertices and the presence or absence of directed branches between the vertices in the directed bipartite graph 70A can be appropriately determined to appropriately express the object of the solving task.
[0229] As long as the task can be expressed by a directed bipartite graph, its content is not particularly limited and can be appropriately selected according to the embodiment. In one example, the task can be a bipartite matching. Specifically, the task can be a matching task between two groups and a task of determining the best pairings between the objects belonging to each group.
[0230] At this time, the two subsets (71A, 72A) of the directed bipartite graph 70A correspond to the two groups in the matching task. The vertices belonging to each subset (71A, 72A) of the directed bipartite graph 70A correspond to the objects belonging to each group. The branch flowing out from the first vertex to the second vertex corresponds to a directed branch starting from the first vertex and ending at the second vertex. The branch flowing out from the second vertex to the first vertex corresponds to a directed branch starting from the second vertex and ending at the first vertex.
[0231] The feature quantity related to the branch flowing out from the first vertex to the second vertex corresponds to the degree of expectation of the object (first vertex) belonging to one of the two groups for the object (second vertex) belonging to the other. The feature quantity related to the branch flowing out from the second vertex to the first vertex corresponds to the degree of expectation of the object (second vertex) belonging to the other of the two groups for the object (first vertex) belonging to one of them.
[0232] As long as it can represent the degree of desired grouping (matching), the form of the degree of expectation is not particularly limited and can be appropriately selected according to the implementation. The degree of expectation can be, for example, the desired rank, score, etc. The score represents the degree of desired grouping numerically. The score can be given appropriately. In one example, after a list of desired ranks is given from each object, the desired ranks of each object represented in the list can be numerically converted through a specified operation to obtain the score, and the obtained score can be used as the degree of expectation.
[0233] The matching objects can be appropriately selected according to the implementation. The matching objects (hereinafter, the matching objects are also referred to as "agents") can be, for example, male / female, roommate, job seeker / employer, employee / allocation target, patient / doctor, energy provider / recipient, etc. In addition, corresponding to the task being matching, the speculation device 2A can be read as a matching device.
[0234] The speculator 55 can be appropriately configured to speculate the matching result based on the feature information. As an example, the speculator 55 can include a single convolutional layer. Thus, the speculator 55 can be configured to input the first output tensor (Z A L ) into the convolutional layer and perform the arithmetic processing of the convolutional layer. Thus, the first output tensor (Z A L ) is converted into the speculation result (q A ∈R N×M ) of the matching based on the object (the first vertex) belonging to one of the groups. Moreover, the speculator 55 can be configured to input the second output tensor (Z B L ) into the convolutional layer and perform the arithmetic processing of the convolutional layer. Thus, the second output tensor (Z B L ) is converted into the speculation result (q B ∈R M×N ) of the matching based on the object (the second vertex) belonging to the other group. Furthermore, the speculator 55 can be configured to average the respective speculation results (q A 、q B ) obtained for each group (agent), thereby calculating the speculation result (q) of the matching (Equation 1 below).
[0235] [Equation 1]
[0236]
[0237] T represents transpose. The speculation result (q) can be expressed as an N×M matrix. In addition, in the speculator 55, either with respect to each output tensor (Z A L 、Z BL ) and commonly use a single convolutional layer, or alternatively, respective convolutional layers may be provided for each output tensor (Z A L 、Z B L ).
[0238] The correct solution in machine learning can be appropriately given in a manner that obtains a matching result that meets the desired criteria. In one example, in the machine learning of the inference module 5, the loss function defined below can be used. That is, first, the inference result (q) can be preprocessed for each group using the softmax function according to the following equations 2 to 4.
[0239] [Equation 2]
[0240] q eA = softmax(q)…(Equation 2)
[0241] [Equation 3]
[0242] q eB = softmax(q T )…(Equation 3)
[0243] [Equation 4]
[0244]
[0245] q e represents the preprocessed inference result. The suffix (i, j) of q e ij represents the element in the i-th row and j-th column. The same applies to the suffix of q eA and q eB . A represents one of the groups (the first set), and B represents the other group (the second set). The min function selects the minimum value among the given independent variables. The loss function λ e used to train in a manner that does not have blocking pairs using the preprocessed inference result (q s can be defined by Equations 5 to 7.
[0246] [Equation 5]
[0247]
[0248] [Equation 6]
[0249]
[0250] [Equation 7]
[0251]
[0252] v A represents the degree of expectation of an object belonging to one group for an object belonging to another group. v B represents the degree of expectation of an object belonging to another group for an object belonging to one of the groups. In one example, the degree of expectation can be obtained by numerically converting the order of expectation so that the higher the order of expectation, the higher the value of the degree of expectation, and the value is greater than 0 and less than or equal to 1 corresponding to the order of expectation. a i represents the (i-th) object belonging to one of the groups, b j represents the (j-th) object belonging to another group. v A ij represents the degree of expectation of the i-th object belonging to one of the groups for the j-th object belonging to another group. v B ji represents the degree of expectation of the j-th object belonging to another group for the i-th object belonging to one of the groups. The max function selects the maximum value among the given independent variables.
[0253] g(a i ; b j , q e ) outputs a value greater than 0 when, for any object b c (c ≠ j) belonging to another group, a i prefers b c more than b j (that is, the degree of expectation for b j is higher than the degree of expectation for b c ). Similarly, g(b j ; a i , q e ) outputs a value greater than 0 when, for any object a c (c ≠ i) belonging to one of the groups, b j prefers a c more than a i (that is, the degree of expectation for a i is higher than the degree of expectation for a c ). When both g(a i ; b j , q e ) and g(b j ; a i , q e ) output values greater than 0, {a i , b j} becomes a blocking pair. According to the loss function λ s , the matching that generates such a blocking pair can be suppressed.
[0254] Moreover, each element of the inference result (q) can be configured to represent the probability of matching between the corresponding objects. At this time, the inference result (q) is preferably a symmetric doubly stochastic matrix. The loss function λ for converging to such an inference result c can be defined by the following equations 8 to 9.
[0255] [Equation 8]
[0256]
[0257] [Equation 9]
[0258]
[0259] * represents an arbitrary element. For example, q eA i* represents the vector formed by arranging all the elements of the i-th row of the matrix q eA , and q eB *i represents the vector formed by arranging all the elements of the i-th column of the matrix q eB . The C function averages the correlation of the inference result (assignment vector) for each group. By using the correlation-based loss, it is possible to calculate the loss only with respect to the symmetry of the matrix without being affected by the difference in the norms of q eA and q eB .
[0260] Furthermore, in order to obtain a stable match with respect to specified criteria such as satisfaction and fairness, a loss function for converging the inference result to satisfy the specified criteria can be added. In one example, at least any one of the three loss functions (λ f , λ e , λ b ) defined by the following equations 10 to 15 can also be added.
[0261] [Equation 10]
[0262]
[0263] [Equation 11]
[0264]
[0265] [Equation 12]
[0266]
[0267] [Equation 13]
[0268]
[0269] [Equation 14]
[0270]
[0271] The S function calculates the total expected degree in the matching result. Therefore, according to the loss function λ f , it is possible to suppress the gap between the total expected degree of one group and the total expected degree of another group in the matching result, that is, it is possible to improve the fairness of the speculation result. Moreover, according to the loss function λ e , it is possible to increase the total expected degree in the matching result, that is, it is possible to improve the satisfaction of the speculation result. The loss function λ b is the same as the loss function λ f and the loss function λ e . Therefore, according to the loss function λ b , it is possible to improve the fairness and satisfaction of the speculation result in a balanced state.
[0272] [Equation 15]
[0273] λ = w s λ s + w c λ c + w f λ f + w b λ b …(Equation 15)
[0274] Therefore, in the machine learning of the speculation module 5, the loss function λ defined by the above Equation 15 can be used. w s , w c , w f , w e and w b are weights that specify the priorities of the respective loss functions (λ s , λ c , λ f , λ e , λ b ). The values of w s , w c , w f , w e and w b can be specified appropriately. In machine learning, by substituting the speculation result into the loss function λ, the error is calculated, and the parameter values of the speculation module 5 are adjusted in the direction of reducing the value of the loss function λ (for example, approaching the lower limit, simply becoming smaller, etc.) by backpropagating the gradient of the calculated error.
[0275] In the first specific example, during the machine learning of the model generation device 1 and during the inference process of the inference device 2A, the directed bipartite graph 70A is processed as the input graph. The initial feature weaving layer 5001 of the feature weaving network 50 is configured to receive a first input tensor 601A and a second input tensor 602A from the directed bipartite graph 70A. The first input tensor 601A represents a feature quantity related to a directed branch extending from a vertex belonging to one of the two subsets (71A, 72A) to a vertex belonging to the other subset. The second input tensor 602A represents a feature quantity extending from a vertex belonging to the other subset to a vertex belonging to one of the subsets. When the feature quantity is one-dimensional, each input tensor (601A, 602A) can be expressed as a matrix (N×M, M×N). Except for these points, the structure of the first specific example can be the same as that of the above-described embodiment.
[0276] (Model generation device)
[0277] In the first specific example, the model generation device 1 can generate the trained inference module 5 through the same processing flow as that of the above-described embodiment, and the trained inference module 5 has the ability to infer the solution to the task of the directed bipartite graph 70A.
[0278] That is, in step S101, the control unit 11 acquires a plurality of training graphs 30A. Each training graph 30A includes a directed bipartite graph 70A. The vertices and branches in each training graph 30A can be appropriately given in a manner that can appropriately represent the conditions of the training object. In step S102, the control unit 11 performs machine learning of the inference module 5 using the acquired plurality of training graphs 30A. Through the machine learning, a trained inference module 5 can be generated, and the trained inference module 5 has the ability to infer the solution to the task of the directed bipartite graph 70A. In addition, when the inference module 5 is made to learn the ability to perform the matching task as the task, in the machine learning, the loss function λ can be used. In the machine learning, the control unit 11 can calculate an error by substituting the inference result into the loss function λ, and adjust the value of the parameter of the inference module 5 in the direction of reducing the value of the loss function λ by backpropagating the gradient of the calculated error. In step S103, the control unit 11 generates learning result data 125 representing the generated trained inference module 5, and stores the generated learning result data 125 in a specified storage area. The learning result data 125 can be provided to the inference device 2A at any time.
[0279] (Inference device)
[0280] The hardware structure and software structure of the inference device 2A may be the same as those of the inference device 2 of the above-described embodiment. In the first specific example, the inference device 2A can infer the solution of the task for the directed bipartite graph 70A through the same processing flow as that of the inference device 2 described above.
[0281] That is, in step S201, the control unit of the inference device 2A operates as an acquisition unit to acquire the object graph 221A. The object graph 221A includes a directed bipartite graph 70A. The settings of the vertices and branches in the object graph 221A can be appropriately given in a manner that appropriately expresses the conditions of the inference object. In step S202, the control unit operates as an inference unit to infer the solution of the task for the acquired object graph 221A using the trained inference module 5. Specifically, the control unit inputs the object graph 221A into the trained inference module 5 and performs the forward operation processing of the trained inference module 5. As a result of performing the operation processing, the control unit can obtain the result of the solution to the inference task for the object graph 221A from the inference unit 55. When the trained inference module 5 obtains the ability to perform the matching task, the control unit can obtain the matching inference result. At this time, the control unit can also perform the operation of equations 1 to 4 in the processing of the inference unit 55, and then perform the matching inference result on the obtained q. e ij An argument of the maximum (argmax) operation is performed to obtain a matching inference result.
[0282] In step S203, the control unit operates as an output unit and outputs information related to the inference result. As an example, when the control unit obtains the inference result of the match, the obtained inference result can also be directly output to the output device. Thus, the inference device 2A can also urge the operator whether to adopt the inference result. As another example, when the inference result of the match is obtained, the control unit can also make the matching of at least part of the inference result established (determined). At this time, the matching object (vertex) can be added at any time, and the inference device 2A can also repeatedly perform the matching of free objects (vertices) with each other. The object (vertex) that has been matched by the execution of the matching task can be excluded from the objects of the matching task after the next time. Moreover, the established match can be released at any time, and the object (vertex) whose match has been released can be added as the object of the subsequent matching task. Thus, the inference device 2A can be configured to perform matching online and in real time.
[0283] (feature)
[0284] According to the first specific example, in a scenario where a given task (such as the matching task) is expressed by a directed bipartite graph 70A, an improvement in the inference accuracy of the inference module 5 can be achieved. In the model generation device 1, the generation of a trained inference module 5 that can expect to accurately infer the solution of a task expressed by a directed bipartite graph 70A is anticipated. In the inference device 2A, by using such a trained inference module 5, an accurate inference of the solution of a task expressed by a directed bipartite graph 70A can be expected.
[0285] (B) Scenario using an undirected bipartite graph
[0286] Figure 12 An example of an application scenario of the inference system 100B according to the second specific example is schematically illustrated. The second specific example is an example in which the above-described embodiment is applied in a scenario where an undirected bipartite graph is used as a graph expressing given conditions (objects of a solving task). The inference system 100B according to the second specific example includes a model generation device 1 and an inference device 2B. The inference device 2B is an example of the inference device 2.
[0287] In the second specific example, each training graph 30B and object graph 221B (input graph) are undirected bipartite graphs 70B. The multiple vertices constituting the undirected bipartite graph 70B are divided into either of two subsets (71B, 72B) such that there are no branches between the vertices within each subset (71B, 72B). The first set of the input graph corresponds to one of the two subsets (71B, 72B) of the undirected bipartite graph 70B, and the second set of the input graph corresponds to the other of the two subsets (71B, 72B) of the undirected bipartite graph 70B. Each training graph 30B and object graph 221B are examples of each training graph 30 and object graph 221 in the above-described embodiment. In addition, the number of vertices and the presence or absence of branches between the vertices in the undirected bipartite graph 70B can be appropriately determined in a manner that appropriately expresses the object of the solving task.
[0288] As long as the task can be expressed by an undirected bipartite graph, its content is not particularly limited and can be appropriately selected according to the embodiment. In one example, the task can be one-sided matching. Specifically, the task can be a matching task between two groups and a task of determining the best pairing between the objects belonging to each group. At this time, the two subsets (71B, 72B) of the undirected bipartite graph 70B correspond to the two groups in the matching task. The vertices belonging to each subset (71B, 72B) of the undirected bipartite graph 70B correspond to the objects belonging to each group. The feature quantity related to the branch flowing from the first vertex to the second vertex and the feature quantity related to the branch flowing from the second vertex to the first vertex both correspond to the cost or reward of pairing the objects belonging to one of the two groups with the objects belonging to the other group.
[0289] The matching object can be appropriately selected according to the embodiment. The matching object can be, for example, a transport robot / luggage, a room / guest (automatic check-in at an accommodation facility), a seat / guest (e.g., seat allocation in public transportation, amusement facilities, etc.). In another example, the matching can be performed to track multiple objects among multiple images (e.g., consecutive frames in a moving image). At this time, the first image can correspond to one group, the second image can correspond to another group, and the matching object (vertex) can be an object (or object region) detected in each image. By matching the same object between each image, the tracking of the object can be performed.
[0290] The cost represents the degree of hindrance to matching. The reward represents the degree of promotion of matching. As an example, in the scenario of matching a transport robot and luggage, the distance from the target transport robot to the target luggage can be set as the cost. As another example, in the scenario of matching a seat and a guest, the preference of the target guest (e.g., a preferred position such as the window side, aisle side, etc.) can be set as the reward for the target seat. The numerical expression of the cost or reward can be appropriately determined according to the embodiment. In addition, corresponding to the task being matching, the speculation device 2B can be read as a matching device.
[0291] In another example, the feature map obtained in a convolutional neural network can be understood as a three-layer tensor of height × width × feature amount. Correspondingly, the feature map can be understood as an undirected bipartite graph having h vertices belonging to one of the subsets and w vertices belonging to the other subset. Therefore, the undirected bipartite graph 70B can include such a feature map, and the task can be a specified inference of an image. The inference of an image can be, for example, segmenting a region, detecting an object, etc. These specified inferences can be performed in various scenarios such as a scenario of extracting an object reflected in a captured image obtained by a vehicle camera, a scenario of detecting a person reflected in a captured image obtained by a camera arranged on the street, etc. In addition, the structure of the convolutional neural network is not particularly limited and can be appropriately determined according to the embodiment. Moreover, the feature map can be obtained either as an output of the convolutional neural network or as an intermediate operation result of the convolutional neural network.
[0292] The predictor 55 may be appropriately configured to derive the speculation result of the solution to the task of the undirected bipartite graph 70B, such as the matching and the specified inference of the image, from the feature information. The correct answer in machine learning may be appropriately given according to the content of the task of the undirected bipartite graph 70B. In the case where the task is the matching task, similar to the first specific example, the correct answer in machine learning may be appropriately given in such a way as to obtain a matching result that satisfies the desired criterion. In the case where the task is the specified inference of the image, for example, the correct answer label indicating the ground truth of the region segmentation, the ground truth of the object detection result, etc., of the inference correct answer may be associated with each training graph 30B. In machine learning, the error may be calculated using the correct answer label.
[0293] In the second specific example, during the machine learning of the model generation device 1 and during the speculation process of the speculation device 2B, the undirected bipartite graph 70B is processed as the input graph. The first feature weaving layer 5001 of the feature weaving network 50 is configured to receive the first input tensor 601B and the second input tensor 602B from the undirected bipartite graph 70B. In the undirected bipartite graph 70B, each branch has no directionality, so the feature quantity related to the branch flowing from the first vertex to the second vertex and the feature quantity related to the branch flowing from the second vertex to the first vertex are the same as each other. Therefore, the second input tensor 602B from the input graph received by the first feature weaving layer 5001 can be generated by transposing the first axis and the second axis of the first input tensor 601B from the input graph. In the case where the feature quantity is one-dimensional, each input tensor (601B, 602B) can be expressed as a matrix (N×M, M×N). Except for these points, the structure of the second specific example may be the same as that of the above-described embodiment.
[0294] (Model generation device)
[0295] In the second specific example, the model generation device 1 may generate the trained speculation module 5 through the same processing flow as that of the above-described embodiment, and the trained speculation module 5 has the ability to speculate on the solution to the task of the undirected bipartite graph 70B.
[0296] That is, in step S101, the control unit 11 acquires a plurality of training graphs 30B. Each training graph 30B includes an undirected bipartite graph 70B. The vertices and branches in each training graph 30B can be appropriately given in a manner that can appropriately represent the conditions of the training object. In step S102, the control unit 11 performs machine learning of the inference module 5 using the acquired plurality of training graphs 30B. Through the machine learning, a trained inference module 5 can be generated, and the trained inference module 5 has obtained the ability to solve the inference task of the undirected bipartite graph 70B. In step S103, the control unit 11 generates learning result data 125 representing the generated trained inference module 5 and stores the generated learning result data 125 in a specified storage area. The learning result data 125 can be provided to the inference device 2B at any time.
[0297] (Inference device)
[0298] The hardware structure and software structure of the inference device 2B can be the same as those of the inference device 2 in the above-described embodiment. In the second specific example, the inference device 2B can infer the solution to the task of the undirected bipartite graph 70B through the same processing flow as that of the inference device 2.
[0299] That is, in step S201, the control unit of the inference device 2B operates as an acquisition unit and acquires an object graph 221B. The object graph 221B includes an undirected bipartite graph 70B. The vertices and branches in the object graph 221B can be appropriately given in a manner that can appropriately represent the conditions of the inference object. When the trained inference module 5 has obtained the ability to perform a specified inference on an image, the control unit can perform the operation of the convolutional neural network as preprocessing, thereby acquiring the object graph 221B (feature map).
[0300] In step S202, the control unit operates as an inference unit and uses the trained inference module 5 to infer the solution to the acquired object graph 221B. Specifically, the control unit inputs the object graph 221B into the trained inference module 5 and performs the forward arithmetic processing of the trained inference module 5. As a result of performing the arithmetic processing, the control unit can obtain the result of the solution to the inference task of the object graph 221B from the inferencer 55. When the trained inference module 5 has obtained the ability to perform the matching task, the control unit can obtain a matching inference result. When the trained inference module 5 has obtained the ability to perform a specified inference on an image, the control unit can obtain a specified inference result for the image from which the object graph 221B is acquired.
[0301] In step S203, the control unit operates as an output unit and outputs information related to the speculation result. As an example, the control unit may directly output the obtained speculation result to the output device. As another example, when the matching speculation result is obtained, similar to the first specific example, the control unit may also make at least a part of the speculation result match. At this time, further similar to the first specific example, the speculation device 2B may be configured to perform matching online and in real time. For example, in the case of the automatic registration / seat allocation, the speculation device 2B may append the corresponding vertices to each subset as the guests arrive and the rooms / seats become available, and allocate vacant rooms / vacant seats through matching according to the guests' requests. And the speculation device 2B may delete the vertices corresponding to the allocated guests and rooms / seats from each subset. Further, as another example, in the case of the object tracking, the control unit may track one or more objects in each image based on the speculation result. At this time, the speculation device 2B may be configured to append the target object as a matching object when it is detected in the image, and exclude the target object from the matching objects when it is not detected in the image (for example, runs out of the shooting range, is blocked by the shadow of an object, etc.). Thus, the speculation device 2B may be configured to perform object tracking in real time.
[0302] (Feature)
[0303] According to the second specific example, in a scenario where the given task is expressed by an undirected bipartite graph 70B, an improvement in the speculation accuracy of the speculation module 5 can be achieved. In the model generation device 1, it is possible to expect the generation of a trained speculation module 5 that can accurately speculate on the solution of the task expressed by the undirected bipartite graph 70B. In the speculation device 2B, by using such a trained speculation module 5, it is possible to expect a high-precision speculation on the solution of the task expressed by the undirected bipartite graph 70B.
[0304] (C) Case of adopting a general directed graph
[0305] Figure 13 An example of the applicable scenario of the speculation system 100C of the third specific example is schematically illustrated. The third specific example is an example in which the above-described embodiment is applied in a scenario where a general directed graph (which may also be simply referred to as a directed graph) is used as the graph expressing the given conditions (the object of the solution task). The speculation system 100C of the third specific example includes a model generation device 1 and a speculation device 2C. The speculation device 2C is an example of the speculation device 2.
[0306] In the third specific example, each training graph 30C and object graph 221C (input graph) is a general directed graph 70C. Circular marks, square marks, and star marks in the graph represent vertices. The first vertex belonging to the first set of the input graph corresponds to the starting point of the directed branch that constitutes the general directed graph 70C, and the second vertex belonging to the second set corresponds to the ending point of the directed branch. Therefore, as Figure 13 shown, the number of vertices in the first set and the second set is the same, and the vertices belonging to each set are common (i.e., M = N). Moreover, the branch flowing from the first vertex to the second vertex corresponds to the directed branch flowing from the starting point to the ending point, and the branch flowing from the second vertex to the first vertex corresponds to the directed branch flowing from the starting point into the ending point. Each training graph 30C and object graph 221C is an example of each training graph 30 and object graph 221 in the above-described embodiment. In addition, the number of vertices in the general directed graph 70C and the presence or absence of directed branches between the vertices can be appropriately determined according to a manner that appropriately represents the object of the solving task.
[0307] As long as the task can be expressed by a general directed graph, its content is not particularly limited and can be appropriately selected according to the embodiment. In one example, the general directed graph 70C can be configured to represent a network. The network can correspond to, for example, a transportation network such as a road network, a railway network, an air route network, or a shipping route network. Alternatively, the network can correspond to a communication network. Correspondingly, the task can be a task of inferring (including predicting) a situation occurring on the network. The inferred situation can be, for example, estimating the time taken for movement between two vertices, predicting the occurrence of an abnormality, searching for the best path, etc.
[0308] At this time, the feature quantity related to the branch flowing from the first vertex to the second vertex and the feature quantity related to the branch flowing from the second vertex to the first vertex can both correspond to the connection attributes between the respective vertices constituting the network. When the network corresponds to a transportation network, each vertex can correspond to a transportation base (such as an intersection, a main location, a station, a port, an airport, etc.), and the connection attribute can be, for example, the type of road (such as a general road, a highway, etc.), the traffic allowance of the road, the time taken for movement, the distance, etc. When the network corresponds to a communication network, each vertex can correspond to a communication base (such as a server, a router, a switch, etc.), and the connection attribute can be, for example, the type of communication, the communication allowance, the communication speed, etc. The feature quantity can include one or more attribute values. In addition, corresponding to the prediction of the situation, the inference device 2C can be read as a prediction device.
[0309] In another example, the task may be to partition a general directed graph (i.e., classify the vertices that make up the general directed graph). As a specific example of the task, the general directed graph 70 may be a weighted graph representing the similarity of movement directions obtained based on the result of associating feature points between two consecutive images. At this time, the partitioning of the graph may be performed for the purpose of detecting dense subgraphs to determine the accuracy of the association. Alternatively, the general directed graph 70 may be a graph representing feature points within an image. At this time, the partitioning of the graph may be performed to classify feature points for each object. In various cases, the feature quantity may be appropriately set in a manner that represents the relationship between feature points.
[0310] In machine learning, the ability to partition the graph can be trained to achieve purposes such as determining the accuracy of the association and classifying for each object. Furthermore, a prescribed arithmetic processing (e.g., certain approximation algorithms) can be applied to the subgraphs obtained by the partitioning. At this time, in machine learning, the ability to partition the graph can be trained to be optimal for a prescribed operation (e.g., the approximation accuracy of the approximation algorithm becomes higher).
[0311] Furthermore, in another example, the task may be to infer the features exhibited in a general directed graph. As a specific example of the task, the general directed graph 70C may be, for example, a graph representing the quantitative change curve of an object (e.g., a person, an object, etc.) such as a knowledge graph. At this time, the task may be to infer the features of the object derived from the quantitative change curve represented by the graph.
[0312] The inferencer 55 may be appropriately configured to derive, from the feature information, a speculation result of the solution to tasks such as the state of affairs speculation, graph partitioning, and feature speculation for the general directed graph 70C. As an example, the output of the inferencer 55 may include a one - hot vector for each vertex representing the speculation result. The correct answer in machine learning may be appropriately given according to the content of the task for the general directed graph 70C. For example, a correct answer label representing the true value of the state of affairs, the true value of the partitioning, the true value of the feature, etc., of the speculation may be associated with each training graph 30C. In machine learning, the error may be calculated using the correct answer label. In addition to this, the correct answer in machine learning may be appropriately given, for example, in a manner that obtains a speculation result that satisfies a desired criterion, such as maximizing the traffic volume, minimizing the movement time of each moving body, maximizing the communication volume of the entire network, etc.
[0313] In the third specific example, during the machine learning of the model generation device 1 and during the inference process of the inference device 2C, the general directed graph 70C is processed as the input graph. The initial feature weaving layer 5001 of the feature weaving network 50 is configured to receive a first input tensor 601C and a second input tensor 602C from the general directed graph 70C. Here, as described above, the first vertex belonging to the first set of the input graph corresponds to the starting point of the directed branch, the second vertex belonging to the second set corresponds to the ending point of the directed branch, the number of vertices in the first set and the second set is the same, and the vertices belonging to each set are common. Therefore, the second input tensor 602C from the input graph received by the initial feature weaving layer 5001 can be generated by transposing the first axis and the second axis of the first input tensor 601C from the input graph. When the feature amount is one-dimensional, each input tensor (601C, 602C) can be expressed as a matrix (N×N). Except for these points, the structure of the third specific example can be the same as that of the above-described embodiment.
[0314] (Model generation device)
[0315] In the third specific example, the model generation device 1 can generate a trained inference module 5 through the same processing flow as in the above-described embodiment, and the trained inference module 5 has the ability to obtain the solution to the task of the general directed graph 70C.
[0316] That is, in step S101, the control unit 11 acquires a plurality of training graphs 30C. Each training graph 30C includes the general directed graph 70C. The setting of the vertices and directed branches in each training graph 30C can be appropriately given in a manner that can appropriately represent the conditions of the training object. In step S102, the control unit 11 performs machine learning of the inference module 5 using the acquired plurality of training graphs 30C. Through the machine learning, a trained inference module 5 that has the ability to obtain the solution to the inference task of the general directed graph 70C can be generated. In step S103, the control unit 11 generates learning result data 125 representing the generated trained inference module 5 and stores the generated learning result data 125 in a specified storage area. The learning result data 125 can be provided to the inference device 2C at any time.
[0317] (Inference device)
[0318] The hardware structure and software structure of the inference device 2C can be the same as those of the inference device 2 in the above-described embodiment. In the third specific example, the inference device 2C can infer the solution to the task of the general directed graph 70C through the same processing flow as that of the inference device 2.
[0319] That is, in step S201, the control unit of the speculation device 2C operates as an acquisition unit to acquire the object graph 221C. The object graph 221C includes the general directed graph 70C. The vertices and directed branches in the object graph 221C can be set in a manner that appropriately expresses the conditions of the speculation object.
[0320] In step S202, the control unit operates as a speculation unit and uses the trained speculation module 5 to speculate on the solution to the task of the acquired object graph 221C. Specifically, the control unit inputs the object graph 221C into the trained speculation module 5 and executes the forward arithmetic processing of the trained speculation module 5. As a result of executing the arithmetic processing, the control unit can obtain the result of speculating on the solution to the task of the object graph 221C from the speculator 55.
[0321] In step S203, the control unit operates as an output unit and outputs information related to the speculation result. As an example, the control unit can also directly output the obtained speculation result to the output device. As another example, when the speculation result of the situation occurring on the network is obtained, the control unit can also output an alarm informing of this situation when the possibility of an abnormality is above the threshold. Furthermore, as another example, when the speculation result of the optimal path in the communication network is obtained, the control unit can also send data packets according to the speculated optimal path.
[0322] (Feature)
[0323] According to the third specific example, in the scenario where the given task is expressed by the general directed graph 70C, the speculation accuracy of the speculation module 5 can be improved. In the model generation device 1, the generation of the trained speculation module 5 that can accurately speculate on the solution to the task expressed by the general directed graph 70C can be expected. In the speculation device 2C, by using such a trained speculation module 5, an accurate speculation on the solution to the task expressed by the general directed graph 70C can be expected.
[0324] (D) Case of using a general undirected graph
[0325] Figure 14 An example of the application scenario of the speculation system 100D of the fourth specific example is schematically illustrated. The fourth specific example is an example in which the above-described embodiment is applied in a scenario where a general undirected graph (also simply referred to as an undirected graph) is used as the graph expressing the given conditions (the object of the solution task). The speculation system 100D of the fourth specific example includes a model generation device 1 and a speculation device 2D. The speculation device 2D is an example of the speculation device 2.
[0326] In the fourth specific example, each training graph 30D and the object graph 221D (input graph) are general undirected graphs 70D. The first vertex belonging to the first set and the second vertex belonging to the second set of the input graph respectively correspond to the respective vertices constituting the general undirected graph 70D. Similar to the third specific example, the number of vertices in the first set and the second set is the same, and the vertices belonging to each set are common (i.e., M = N). The branches flowing from the first vertex to the second vertex and the branches flowing from the second vertex to the first vertex correspond to the undirected branches between the corresponding two vertices. Each training graph 30D and the object graph 221D are examples of the respective training graphs 30 and the object graph 221 in the above-described embodiment. In addition, the number of vertices in the general undirected graph 70D and the presence or absence of the branches between the vertices can be appropriately determined according to the manner of appropriately expressing the object of the solution task.
[0327] As long as the task can be expressed by a general undirected graph, its content is not particularly limited and can be appropriately selected according to the embodiment. Except that the branches do not have directionality, the task in the fourth specific example can be the same as that in the third specific example. That is, in one example, the general undirected graph 70D can be configured to represent a network. Correspondingly, the task can be: inferring (including predicting) the situation occurring on the network. At this time, the feature quantity related to the branch flowing from the first vertex to the second vertex and the feature quantity related to the branch flowing from the second vertex to the first vertex can both correspond to the connection attributes between the respective vertices constituting the network. In addition, corresponding to the predicted situation, the inference device 2D can be read as a prediction device. In another example, the task can be to segment the general undirected graph. Further, in another example, the task can be to infer the features exhibited in the general undirected graph.
[0328] The inferencer 55 can be appropriately configured to derive from the feature information the inference results of the solutions to the tasks for the general undirected graph 70D, such as the situation inference, graph segmentation, and feature inference. In one example, the output of the inferencer 55 can include a one-hot vector representing each vertex of the inference result. The correct answer in machine learning can be appropriately given according to the content of the task for the general undirected graph 70D. For example, the correct answer label representing the true value of the situation, the true value of the segmentation, the true value of the feature, etc., of the inference can be associated with each training graph 30D. In machine learning, the error can be calculated using the correct answer label. In addition to this, the correct answer in machine learning can be appropriately given in such a way as to obtain an inference result that satisfies the desired criterion.
[0329] In the fourth specific example, during the machine learning of the model generation device 1 and during the inference process of the inference device 2D, the general undirected graph 70D is processed as the input graph. The first feature weaving layer 5001 of the feature weaving network 50 is configured to receive a first input tensor 601D and a second input tensor 602D from the general undirected graph 70D. When the feature amount is one-dimensional, each input tensor (601D, 602D) can be represented by a matrix (N×N).
[0330] Figure 15 Schematically illustrate an example of the arithmetic processing of each feature weaving layer 500 in the fourth specific example. In the fourth specific example, since the vertices belonging to the first set and the vertices belonging to the second set are the same, as in the third specific example, the second input tensor 602D from the input graph received by the first feature weaving layer 5001 can be generated by transposing the first axis and the second axis of the first input tensor 601D from the input graph. However, in the general undirected graph 70D, since the branches do not have directionality, the second input tensor 602D generated by transposition is the same as the first input tensor 601D. Therefore, the first input tensor (Z A l ) and the second input tensor (Z B l ) input to each feature weaving layer 500 are the same as each other, and the first output tensor (Z A l+1 ) and the second output tensor (Z B l+1 ) generated by each feature weaving layer 500 are also the same as each other.
[0331] Therefore, as Figure 15 shown, in each feature weaving layer 500, the process of generating the first output tensor and the process of generating the second output tensor can be executed jointly. That is, the series of processes of generating the first output tensor from the first input tensor and the series of processes of generating the second output tensor from the second input tensor can be constituted singly, and either of them can be omitted. In other words, the series of processes of generating the first output tensor and the series of processes of generating the second output tensor can be regarded as executing two processes by executing any one of these processes. Figure 15 Illustrate a scenario where the series of processes of generating the second input tensor is omitted. Except for these points, the structure of the fourth specific example can be the same as that of the above-described embodiment.
[0332] (Model Generation Device)
[0333] Return Figure 14, in the fourth specific example, the model generation device 1 can generate the trained inference module 5 through the same processing flow as in the above-described embodiment. The trained inference module 5 has obtained the ability to infer the solution to the task for the general undirected graph 70D.
[0334] That is, in step S101, the control unit 11 acquires a plurality of training graphs 30D. Each training graph 30D includes a general undirected graph 70D. The vertices and branches in each training graph 30D can be appropriately given in a manner that can appropriately represent the conditions of the training object. In step S102, the control unit 11 performs machine learning on the inference module 5 using the plurality of acquired training graphs 30D. Through the machine learning, a trained inference module 5 that has obtained the ability to infer the solution to the task for the general undirected graph 70D can be generated. In step S103, the control unit 11 generates learning result data 125 representing the generated trained inference module 5 and stores the generated learning result data 125 in a specified storage area. The learning result data 125 can be provided to the inference device 2D at any time.
[0335] (Inference device)
[0336] The hardware structure and software structure of the inference device 2D can be the same as those of the inference device 2 in the above-described embodiment. In the fourth specific example, the inference device 2D can infer the solution to the task for the general undirected graph 70D through the same processing flow as that of the inference device 2.
[0337] That is, in step S201, the control unit of the inference device 2D operates as an acquisition unit and acquires the object graph 221D. The object graph 221D includes a general undirected graph 70D. The vertices and branches in the object graph 221D can be appropriately given in a manner that can appropriately represent the conditions of the inference object.
[0338] In step S202, the control unit operates as an inference unit and uses the trained inference module 5 to infer the solution to the task for the acquired object graph 221D. Specifically, the control unit inputs the object graph 221D to the trained inference module 5 and executes the forward arithmetic processing of the trained inference module 5. As a result of executing the arithmetic processing, the control unit can obtain the result of the solution to the inference task for the object graph 221D from the inference unit 55.
[0339] In step S203, the control unit operates as an output unit and outputs information related to the speculation result. Similar to the third specific example, as an example, the control unit may also directly output the obtained speculation result to the output device. As another example, when the speculation result of the situation generated on the network is obtained, the control unit may output an alarm indicating this situation when the possibility of an abnormality occurrence is equal to or higher than a threshold value. Further, as another example, when the speculation result of the optimal path in the communication network is obtained, the control unit may send data packets based on the speculated optimal path.
[0340] (Feature)
[0341] According to the fourth specific example, in a scenario where the given task is expressed by a general undirected graph 70D, an improvement in the speculation accuracy of the speculation module 5 can be achieved. In the model generation device 1, the generation of a trained speculation module 5 that can speculate with high accuracy on the solution of the task expressed by the general undirected graph 70D can be expected. In the speculation device 2D, by using such a trained speculation module 5, a high-accuracy speculation on the solution of the task expressed by the general undirected graph 70D can be expected.
[0342] (E) Case of adopting a vertex feature graph given the relationship between vertices
[0343] Figure 16 An example of an application scenario of the speculation system 100E of the fifth - 1st specific example is schematically illustrated. The fifth - 1st specific example is an example in which the above - described embodiment is applied in a scenario where a vertex feature graph is adopted as the graph expressing the given conditions (objects of the solution - solving task), and the relationship between vertices constituting the adopted vertex feature graph is known. The speculation system 100E of the fifth - 1st specific example includes a model generation device 1 and a speculation device 2E. The speculation device 2E is an example of the speculation device 2.
[0344] In the fifth - 1st specific example, each training graph 30E and object graph 221E (input graph) is a vertex feature graph 70E including a plurality of vertices, and is a vertex feature graph 70E in which each vertex has an attribute (i.e., vertex feature). The type and number of attributes possessed by each vertex can be appropriately determined according to the embodiment. Each vertex may have one or more attribute values. Weight is an example of an attribute, and a vertex - weighted graph is an example of the vertex feature graph 70E.
[0345] The first vertex belonging to the first set and the second vertex belonging to the second set of the input graph respectively correspond to the vertices constituting the vertex feature graph 70E. Therefore, the number of vertices in the first set and the second set is the same, and the vertices belonging to each set are common. The feature quantity related to the branch flowing from the first vertex to the second vertex corresponds to the attribute of the vertex of the vertex feature graph 70E corresponding to the first vertex, and the feature quantity related to the branch flowing from the second vertex to the first vertex corresponds to the attribute of the vertex of the vertex feature graph 70E corresponding to the second vertex. That is, in the 5-1 specific example, the attribute value of each vertex is understood as the feature quantity of each branch. Thus, the vertex feature graph 70E is processed as a branch feature graph in which each branch has an attribute, similar to the first to fourth specific examples. The branch weighted graph that assigns weights to each branch is an example of the branch feature graph. The feature quantity may include more than one type of attribute value.
[0346] Each training graph 30E and the object graph 221E are examples of the respective training graphs 30 and the object graph 221 in the above embodiment. In addition, the number of vertices, the attributes, and the relationality of the branches between the vertices in the vertex feature graph 70E can be appropriately determined in a manner that appropriately represents the object of the solving task. Hereinafter, for convenience of explanation, it is assumed that the number of vertices is N.
[0347] As long as the task can be expressed by the vertex feature graph, its content is not particularly limited and can be appropriately selected according to the embodiment. In one example, the task may be to infer a state of affairs derived from the vertex feature graph 70E.
[0348] As a specific example of the task, the vertex feature graph 70E can be configured to represent a chemical formula or a physical crystal structure, and the chemical formula represents a component. At this time, each vertex may correspond to an element, and the attribute of each vertex may correspond to an attribute of the element such as the type of the element. The relationality between the vertices may correspond to the connection relationship between the elements. The inferred state of affairs may be: inferring features related to the component such as the properties possessed by the component. In addition to this, the inferred state of affairs may be: inferring the process of the substance that generates the component. For the expression of the process, for example, natural language, symbol strings, etc. can be used. In the case of using symbol strings, the correspondence relationship between the symbol strings and the process can be predefined by rules. Furthermore, the symbol strings may include commands for controlling the actions of the production apparatus that constitutes the substance for generating the component. The production apparatus may be, for example, a computer, a controller, a robot apparatus, etc. The commands may be, for example, control instructions, computer programs, etc.
[0349] As another specific example of the task, the vertex feature map 70E can be configured to represent the features of multiple objects (such as people, objects, etc.). At this time, each vertex can correspond to each object, and the attributes of each vertex can correspond to the features of each object. The features of each object can be extracted from the graph expressing the variable quantity curve through the third specific example or the fourth specific example. The speculated situation can be: to perform an optimal match between objects. The matching objects can be, for example, guest / advertisement, patient / doctor, person / job, etc.
[0350] As yet another specific example of the task, the vertex feature map 70E can be configured to represent a power supply network. At this time, each vertex can correspond to a power production node or a power consumption node. The attributes of the vertex can be configured to represent, for example, the production quantity, consumption quantity, etc. Each branch can correspond to a wire. The speculated situation can be: to speculate on the optimal combination of production nodes and consumption nodes.
[0351] The speculator 55 can be appropriately configured to derive from the feature information a speculation result of the solution to the task of the vertex feature map 70E such as the situation speculation. The correct answer in machine learning can be appropriately given according to the content of the task of the vertex feature map 70E. For example, a correct answer label representing the truth value, etc. of the speculated situation can be associated with each training graph 30E. In machine learning, the error can be calculated using the correct answer label. In addition to this, the correct answer in machine learning can be appropriately given in a manner that obtains a speculation result that satisfies the desired criteria.
[0352] In the 5-1st specific example, during the machine learning of the model generation device 1 and during the speculation process of the speculation device 2E, the vertex feature map 70E is processed as an input graph. The first feature weaving layer 5001 of the feature weaving network 50 is configured to receive a first input tensor 601E and a second input tensor 602E from the vertex feature map 70E. A tensor (∈R N×1×D0 ) obtained by arranging vertices on the first axis and arranging the feature quantities representing the attributes of the vertices on the third axis, and replicating each element along the second axis N times ( Figure 16 “repeated × N” of
[0353] ), whereby each input tensor (601E, 602E) can be obtained. However, in this state, the relationship between vertices is not reflected in each input tensor (601E, 602E). T)。The relationship between vertices can be defined, for example, using attributes of branches such as the presence or absence of branches, the types of branches, and characteristic quantities related to the branches. The expression form of the relationship between vertices is not particularly limited and can be appropriately determined according to the implementation. In one example, when indicating the presence or absence of branches, the relationship between vertices can be expressed using the two values {0, 1}. In another example, when reflecting attributes such as the types of branches, the relationship between vertices can be expressed using continuous values.
[0354] By arranging each vertex on the first axis and the second axis and arranging the numerical information representing the relationship between vertices (the attributes of the branches between vertices) on the third axis, it is possible to generate information 75E applicable to the first input tensor 601E in the form of a tensor (E ∈ R N×N×DE )). DE represents the dimension of the numerical information that represents the relationship between vertices. By transposing the first axis and the second axis of the information 75E(E), it is possible to generate information 75E(E T ) applicable to the second input tensor 602E. In one example, DE can be one-dimensional, and the information 75E representing the relationship between vertices can be expressed as a matrix.
[0355] The connection can be: combining the information 75E representing the relationship between vertices to the third axis (cat operation) of each input tensor (601E, 602E). In addition to this, the information 75E representing the relationship between vertices can also be combined to other dimensions (such as a new fourth axis) of each input tensor (601E, 602E). Moreover, when the relationship between vertices is defined only by the presence or absence of branches, the relationship between vertices can also be reflected in each input tensor (601E, 602E) by omitting the part where there is no relationship between vertices (no branches). Except for these points, the structure of the 5-1 specific example can be the same as that of the above-mentioned implementation.
[0356] (Model generation device)
[0357] In the 5-1 specific example, the model generation device 1 can generate a trained inference module 5 through the same processing flow as the above-mentioned implementation, and the trained inference module 5 has the ability to obtain the solution to the task of inferring the vertex feature map 70E.
[0358] That is, in step S101, the control unit 11 acquires a plurality of training diagrams 30E. Each training diagram 30E includes a vertex feature map 70E. The vertices, attributes, and branches in each training diagram 30E can be appropriately given in a manner that can appropriately represent the conditions of the training object. In step S102, the control unit 11 performs machine learning on the plurality of acquired training diagrams 30E for the inference module 5. Through the machine learning, it is possible to generate a trained inference module 5 that has the ability to obtain a solution to the inference task for the vertex feature map 70E. In step S103, the control unit 11 generates learning result data 125 representing the generated trained inference module 5 and stores the generated learning result data 125 in a specified storage area. The learning result data 125 can be provided to the inference device 2E at any time.
[0359] (Inference device)
[0360] The hardware structure and software structure of the inference device 2E can be the same as those of the inference device 2 in the above-described embodiment. In the 5-1 specific example, the inference device 2E can infer the solution to the task for the vertex feature map 70E through the same processing flow as that of the inference device 2.
[0361] That is, in step S201, the control unit of the inference device 2E operates as an acquisition unit and acquires an object diagram 221E. The object diagram 221E includes a vertex feature map 70E. The vertices, attributes, and branches in the object diagram 221E can be appropriately given in a manner that can appropriately represent the conditions of the inference object.
[0362] In step S202, the control unit operates as an inference unit and uses the trained inference module 5 to infer the solution to the task for the acquired object diagram 221E. Specifically, the control unit inputs the object diagram 221E to the trained inference module 5 and executes the forward arithmetic processing of the trained inference module 5. As a result of executing the arithmetic processing, the control unit can obtain the result of inferring the solution to the task for the object diagram 221E from the inferencer 55.
[0363] In step S203, the control unit operates as an output unit and outputs information related to the inference result. As an example, the control unit can also directly output the obtained inference result to an output device. As another example, when the object diagram 221E is configured to represent an object component (chemical formula or physical crystal structure) and the process of the substance that generates the object component is inferred, the control unit can also control the operation of the production device according to the process indicated by the obtained inference result. As another example, when the matching between object objects or the combination of production nodes and consumption nodes is inferred, the control unit can also make the matching or combination valid (determined).
[0364] (Feature)
[0365] According to the 5-1 specific example, in the scenario where a task is expressed by the vertex feature map 70E with known relationships between vertices, an improvement in the inference accuracy of the inference module 5 can be achieved. In the model generation device 1, it is expected to generate a trained inference module 5 that can accurately infer the solution to the task expressed by the vertex feature map 70E. In the inference device 2E, by using such a trained inference module 5, it is expected to accurately infer the solution to the task expressed by the vertex feature map 70E.
[0366] (F) Case of using a vertex feature map with unknown relationships between vertices
[0367] Figure 17 An example of the application scenario of the inference system 100F of the 5-2 specific example is schematically illustrated. The 5-2 specific example is an example in which the present embodiment is applied in a scenario where a vertex feature map is used as a graph expressing given conditions (objects of the solution task), and the relationships between the vertices constituting the used vertex feature map are unknown. The inference system 100F of the 5-2 specific example includes a model generation device 1 and an inference device 2F. The inference device 2F is an example of the inference device 2.
[0368] Basically, the 5-2 specific example is the same as the 5-1 specific example except that the relationships between vertices (i.e., branch information) are not given. That is, each training graph 30F and object graph 221F (input graph) is a vertex feature map 70F including a plurality of vertices, and is a vertex feature map 70F in which each vertex has an attribute. The first vertex belonging to the first set and the second vertex belonging to the second set of the input graph respectively correspond to the vertices constituting the vertex feature map 70F. The feature quantity related to the branch flowing from the first vertex to the second vertex corresponds to the attribute of the vertex corresponding to the first vertex in the vertex feature map 70F, and the feature quantity related to the branch flowing from the second vertex to the first vertex corresponds to the attribute of the vertex corresponding to the second vertex in the vertex feature map 70F. Each training graph 30F and object graph 221F is an example of each training graph 30 and object graph 221 in the present embodiment. In addition, the number and attributes of the vertices in the vertex feature map 70F can be appropriately determined in a manner that appropriately expresses the object of the solution task.
[0369] As long as the task can be expressed by a vertex feature map with unknown relationships between vertices, its content is not particularly limited and can be appropriately selected according to the embodiment. In one example, the task may be to infer the relationships between the vertices constituting the vertex feature map 70F.
[0370] As a specific example of the task, the vertex feature map 70F can be configured to correspond to an image (e.g., a scene graph) of a detected object. At this time, each vertex can also correspond to the detected object, and the attributes of each vertex can also correspond to the attributes of the detected object. Inferring the relationship between each vertex can be inferring the features presented in the image (image recognition). The specific example can be adopted in various scenarios of image recognition.
[0371] As another specific example of the task, the vertex feature map 70F can be configured to correspond to an image that maps one or more persons and detects the feature points of the persons (e.g., a scene graph, a skeleton model). At this time, each vertex can also correspond to the detected feature points of the persons (e.g., joints, etc.), and the attributes of each vertex can also correspond to the attributes of the detected feature points (e.g., the type, position, inclination, etc. of the joints). Inferring the relationship between each vertex can be inferring the type of the action of the person (action analysis). The specific example can be adopted, for example, in various scenarios of action analysis such as the dynamic image analysis of sports (especially team competitions).
[0372] As yet another specific example of the task, the vertex feature map 70F can be configured to represent the properties of the constituent. The properties of the constituent can be defined, for example, by the type of influence on the human body (pharmacological effect, side effect, persistence, reaction when in contact with the skin, etc.), the degree of various effects (efficiency of the photovoltaic effect, efficiency of the thermoelectric effect, infrared absorption / reflection efficiency, coefficient of friction, etc.). At this time, each vertex can correspond to a specific property, and the attributes of each vertex can be configured to represent the degree of the specific property. Inferring the relationship between each vertex can be: inferring (generating) the structure of the constituent that satisfies the properties shown by the vertex feature map 70F (e.g., chemical formula, physical crystal structure).
[0373] The inferencer 55 can be appropriately configured to derive from the feature information the inference result of the solution to the tasks of the vertex feature map 70F such as the above-mentioned image recognition, action analysis, and structure generation. The correct answer in machine learning can be appropriately given according to the content of the task of the vertex feature map 70F. For example, the correct answer label representing the true value of image recognition, the true value of action analysis, the true value of structure generation, etc. can be associated with each training graph 30F. In machine learning, the error can be calculated using the correct answer label.
[0374] In the 5-2 specific example, during the machine learning of the model generation device 1 and during the inference process of the inference device 2F, the vertex feature map 70F is processed as the input map. The first feature weaving layer 5001 of the feature weaving network 50 is configured to receive a first input tensor 601F and a second input tensor 602F from the vertex feature map 70F. Each input tensor (601F, 602F) can be obtained by the same method as each input tensor (601E, 602E) in the 5-1 specific example. Except for these points, the structure of the 5-2 specific example can be the same as that of the above-described embodiment.
[0375] (Model generation device)
[0376] In the 5-2 specific example, the model generation device 1 can generate a trained inference module 5 through the same processing flow as the above-described embodiment, and the trained inference module 5 has the ability to obtain the solution to the task of inferring the vertex feature map 70F.
[0377] That is, in step S101, the control unit 11 acquires a plurality of training graphs 30F. Each training graph 30F includes the vertex feature map 70F. The vertices and attributes in each training graph 30F can be appropriately given in a manner that can appropriately represent the conditions of the training object. In step S102, the control unit 11 performs machine learning of the inference module 5 using the acquired plurality of training graphs 30F. Through the machine learning, a trained inference module 5 that has obtained the ability to solve the inference task for the vertex feature map 70F can be generated. In step S103, the control unit 11 generates learning result data 125 representing the generated trained inference module 5 and stores the generated learning result data 125 in a specified storage area. The learning result data 125 can be provided to the inference device 2F at any time.
[0378] (Inference device)
[0379] The hardware structure and software structure of the inference device 2F can be the same as those of the inference device 2 in the above-described embodiment. In the 5-2 specific example, the inference device 2F can infer the solution to the task of the vertex feature map 70F through the same processing flow as the inference device 2.
[0380] That is, in step S201, the control unit of the inference device 2F operates as an acquisition unit and acquires the object graph 221F. The object graph 221F includes the vertex feature map 70F. The vertices and attributes in the object graph 221F can be appropriately given in a manner that can appropriately represent the conditions of the inference object. In the case of structure generation, the value of each property (that is, the attribute value of each vertex) can also be given by a random number or can be given by the designation of an operator.
[0381] In step S202, the control unit operates as a speculation unit and uses the trained speculation module 5 to speculate on the solution to the task of the acquired object graph 221F. Specifically, the control unit inputs the object graph 221F to the trained speculation module 5 and executes the forward arithmetic processing of the trained speculation module 5. As a result of executing the arithmetic processing, the control unit can obtain, from the speculator 55, the result of speculating on the solution to the task of the object graph 221F.
[0382] In step S203, the control unit operates as an output unit and outputs information related to the speculation result. As an example, the control unit may also directly output the obtained speculation result to an output device. As another example, the control unit may execute information processing corresponding to the obtained speculation result. For example, in the case where the object graph 221F is configured to correspond to an image of a detected object, the control unit may execute information processing corresponding to the result of image recognition. As a specific application scenario, the 5-2 specific example can be applied to the recognition of a captured image obtained by photographing the external condition of a vehicle using a camera. In this scenario, when it is determined based on the result of image recognition that an obstacle is approaching the vehicle, the control unit may also give an instruction to the vehicle to avoid the obstacle (for example, stop, change lanes, etc.). Moreover, for example, in the case where the object graph 221F is configured to correspond to an image of detected feature points of a person, the control unit may execute information processing corresponding to the result of motion analysis. As a specific application scenario, the 5-2 specific example can be applied to the motion analysis of a person reflected in a captured image obtained by using a camera installed on a platform of a railway station. In this scenario, when it is determined based on the result of motion analysis that danger is approaching a person on the platform (for example, about to fall off the platform), the control unit may also output an alarm informing of this situation.
[0383] (Feature)
[0384] According to the 5-2 specific example, in a scenario where the task of speculating on the relationship between vertices is expressed by the vertex feature graph 70F, an improvement in the speculation accuracy of the speculation module 5 can be achieved. In the model generation device 1, the generation of a trained speculation module 5 that can be expected to accurately speculate on the solution to the task of speculating on the relationship between vertices of the vertex feature graph 70F is expected. In the speculation device 2F, by using such a trained speculation module 5, an accurate speculation on the solution to the task of speculating on the relationship between vertices of the vertex feature graph 70F can be expected.
[0385] (G) Case of adopting a hypergraph
[0386] Figure 18Schematically illustrate an example of the applicable scenario of the speculation system 100G in the sixth specific example. The sixth specific example is an example in which the described embodiment is applied in a scenario where a hypergraph is used as a graph expressing given conditions (objects of a solving task). The speculation system 100G in the sixth specific example includes a model generation device 1 and a speculation device 2G. The speculation device 2G is an example of the speculation device 2.
[0387] In the sixth specific example, each training graph 30G and object graph 221G (input graph) is a K-piece hypergraph 70G. K is an integer of 3 or more. The multiple vertices constituting the hypergraph 70G are divided into any one of K subsets such that there are no branches between the vertices within each subset. Each training graph 30G and object graph 221G corresponds to each training graph 30 and object graph 221 in the described embodiment. In addition, the number of vertices in the hypergraph 70G, the presence or absence of branches between vertices, and the value of K can be appropriately determined in a manner that appropriately expresses the object of the solving task.
[0388] Similar to the described embodiment, in the sixth specific example, the speculation module 5 includes a feature weaving network 50 and a speculator 55. The feature weaving network 50 is configured to receive the input of the input graph and output feature information related to the input graph by including a plurality of feature weaving layers 500. The speculator 55 is configured to speculate on the solution to the task of the input graph based on the feature information output from the feature weaving network. Each feature weaving layer 500 includes an encoder 505. The encoder 505 is configured to derive the relative feature amounts of the respective elements based on the feature amounts of all the input elements. On the other hand, in the sixth specific example, each feature weaving layer 500 and various information (input tensors, etc.) of the speculation module 5 are extended from the described embodiment in order to process the hypergraph 70G.
[0389] Figure 19 Schematically illustrate an example of the arithmetic processing of each feature weaving layer 500 in the sixth specific example. Figure 19 In it, N i represents the number of the i-th vertex (i is an integer from 1 to K). Each feature weaving layer 500 is configured to receive the input of K (K + 1)-layer input tensors (Z 1 l ,..., Z K l ). The i-th input tensor (Z i l) It is configured such that in the first axis, the i-th vertex belonging to the i-th subset of the input graph is arranged as an element, and in the j-th axis from the second axis to the K-th axis, cyclically starting from the i-th subset, the vertices belonging to the (j - 1)-th subset are arranged as elements, and in the (K + 1)-th axis, the characteristic quantities related to the branches flowing out from the combination of the i-th vertex belonging to the i-th subset to the vertices of (K - 1) (other than the i-th) other subsets are arranged as elements. For example, the elements of each axis of the first input tensor (Z 1 l ) are expressed as N1×N2×…×N K ×D l . The elements of each axis of the second input tensor are expressed as N2×N3×…×N K ×N1×D l . The elements of each axis of the K-th input tensor are expressed as N K ×N1×…×N K-1 ×D l .
[0390] Each feature weaving layer 500 is configured to perform the following operations. That is, for the i-th input tensor (Z i l ) from the first to the K-th, the (K + 1)-th axis of each of the other (K - 1) input tensors other than the i-th is fixed, and the axes other than the (K + 1)-th axis of each of the other input tensors are cycled in a manner consistent with the axes of the i-th input tensor. For example, in the processing path of the first input tensor, the axes other than the (K + 1)-th axis of the second input tensor to the K-th input tensor other than the first input tensor are cycled in a manner consistent with the axes of the first input tensor. Specifically, in the second input tensor, the K-th axis corresponds to the first axis of the first input tensor. Therefore, the first axis to the K-th axis of the second input tensor are cycled one by one in the right direction in the expression of the axis elements (N2×N3×…×N K ×N1×D l ) (P1(Z 2 l )). Similarly, the first axis to the K-th axis of the third input tensor are cycled two by two in the right direction. The first axis to the K-th axis of the K-th input tensor are cycled K by K in the right direction (P K (Z K l )). Moreover, for example, in the processing path of the K-th input tensor, the axes other than the (K + 1)-th axis of the first input tensor to the (K - 1)-th input tensor other than the K-th input tensor are cycled in a manner consistent with the axes of the K-th input tensor. Specifically, the first axis to the K-th axis of the first input tensor are cycled one by one in the right direction (P1(Z 1l ))。Similarly, in the first to K-th axes of the (K-1)-th input tensor, K elements are cycled K by K in the right direction (P K (Z K-1 l ))。In the processing paths of the second to (K-1)-th input tensors, the same method is also used to cycle the axes of each input tensor.
[0391] Subsequently, each feature weaving layer 500 connects the feature amounts of the elements of the i-th input tensor and each of the other input tensors whose axes are cycled. Thus, each feature weaving layer 500 generates K (K+1)-layer connection tensors. For example, in the processing path of the first input tensor, each feature weaving layer 500 connects the first input tensor (Z 1 l ) and the feature amounts of the elements of the second to K-th input tensors (P1(Z 2 l ))~P K (Z K l )) whose axes are cycled, thereby generating the first connection tensor (cat(Z 1 l , P1(Z 2 l ), …, P K (Z K l ))). Moreover, for example, in the processing path of the K-th input tensor, each feature weaving layer 500 connects the K-th input tensor (Z K l ) and the feature amounts of the elements of the first to (K-1)-th input tensors (P1(Z 1 l ))~P K (Z K -1 l )) whose axes are cycled, thereby generating the K-th connection tensor (cat(Z K l , P1(Z 1 l ), …, P K (Z K-1 l ))). In the processing paths of the second to (K-1)-th input tensors, the same method is also used to generate each connection tensor.
[0392] Moreover, each feature weaving layer 500 divides the generated K connection tensors into each element of the first axis respectively and inputs them to the encoder 505 to perform the operations of the encoder 505. Each feature weaving layer 500 obtains the operation results of the encoder 505 corresponding to each element, thereby generating K output tensors of (K + 1) layers respectively corresponding to the K input tensors. Each feature weaving layer 500 is configured to perform each of the above operations. Other structures of each feature weaving layer 500 may be the same as those in the above embodiment.
[0393] Similar to the above embodiment, in one example, the encoder 505 may include a single encoder commonly used among the elements of the first axis of each connection tensor. In another example, the encoder 505 may include K encoders used respectively for each connection tensor. Furthermore, in another example, the encoder 505 may be provided respectively corresponding to each element of each connection tensor.
[0394] Return Figure 18 , as long as the task can be expressed by a hypergraph, its content may not be particularly limited and can be appropriately selected according to the embodiment. In one example, the task may be a matching task among K groups and is a task of determining the best combination of objects belonging to each group. At this time, each subset of the hypergraph 70G corresponds to each group in the matching task. The vertices belonging to each subset of the hypergraph 70G correspond to the objects belonging to each group. The feature quantity related to the branches flowing out from the i-th vertex belonging to the i-th subset to the combinations of the vertices belonging to (K - 1) (other than the i-th) other subsets respectively may correspond to the cost or reward for matching the combination of the object corresponding to the i-th vertex and the objects corresponding to the vertices belonging to the other subsets other than the i-th. Alternatively, the feature quantity may correspond to the degree of expectation of the object corresponding to the i-th vertex for the combinations of the objects corresponding to the vertices belonging to the other subsets other than the i-th.
[0395] The matching objects can be appropriately selected according to the embodiment. The matching task can be performed, for example, for the purpose of determining the best combination of three or more parts; determining the best combination of three or more persons; determining the best combination of three items for learning for triplet loss calculation in distance learning; determining the combination of feature points belonging to the same person when simultaneously inferring the poses of multiple persons in an image; and determining the best combination of delivery trucks, luggage, and routes, etc. Correspondingly, the matching objects can be, for example, parts, persons, data that are objects for calculating triplet loss, feature points of persons detected in an image, delivery trucks / luggage / routes, etc. Determining the best combination of persons can be performed, for example, for team building of managers / programmers / business, front-end programmers / back-end programmers / data scientists of a web page, etc.
[0396] The speculator 55 can be appropriately configured to derive a speculation result of a solution to a task such as the matching task for the hypergraph 70G from the feature information. The correct answer in machine learning can be appropriately given according to the content of the task for the hypergraph 70G. In one example, when the task is the matching task, similar to the first specific example, etc., the correct answer in machine learning can be appropriately given in such a way as to obtain a matching result that satisfies the desired criterion. In another example, a correct answer label indicating the correct answer of speculation such as the true value of the match can be associated with each training graph 30G. In machine learning, the error can be calculated using the correct answer label. In addition to this, the correct answer in machine learning can be appropriately given in such a way as to obtain a speculation result that satisfies the desired criterion. As an example, in the case of the team building, a specified criterion can be specified, for example, to evaluate the achievements achieved in the real world such as the benefits of money, the decrease in the turnover rate, the reduction of the development process, the increase in the amount of progress reports, the improvement of psychological indicators (such as questionnaires, sleep time, blood pressure, etc.), and the correct answer can be given according to the evaluation value.
[0397] In the sixth specific example, during the machine learning of the model generation device 1 and during the speculation process of the speculation device 2G, the hypergraph 70G is processed as an input graph. The first feature weaving layer 5001 among the plurality of feature weaving layers 500 is configured to receive K input tensors 601G to 60KG (Z 1 0,..., Z K 0) from the hypergraph 70G (input graph). Similar to the above-described embodiment, in one example, each input tensor 601G to 60KG can be directly obtained from the input graph. In another example, each input tensor 601G to 60KG can be obtained based on the result of a specified arithmetic process performed on the input graph. Further, in another example, the input graph can be given in the form of each input tensor 601G to 60KG.
[0398] In one example, the feature weaving layers 500 can be continuously arranged in series. At this time, the other feature weaving layers 500 other than the first feature weaving layer 5001 are configured to receive each output tensor generated by the feature weaving layer 500 arranged immediately before it as each input tensor. That is, the feature weaving layer 500 arranged in the (l + 1)-th position is configured to receive each output tensor generated by the l-th feature weaving layer as each input tensor. However, similar to the above-described embodiment, in the sixth specific example, the arrangement of the feature weaving layers 500 is not limited to this example. For example, other types of layers (such as a convolutional layer, a fully connected layer, etc.) can be inserted between adjacent feature weaving layers 500.
[0399] The feature information includes K output tensors output from the last feature weaving layer 500L among a plurality of feature weaving layers 500. The speculator 55 can be appropriately configured to derive the speculation result from these K output tensors. Except for these points, the structure of the sixth specific example can be the same as that of the embodiment.
[0400] (Model generation device)
[0401] In the sixth specific example, the model generation device 1 can generate a trained speculation module 5 through the same processing flow as that of the embodiment. The trained speculation module 5 has the ability to obtain the solution to the task of the hypergraph 70G.
[0402] That is, in step S101, the control unit 11 acquires a plurality of training graphs 30G. Each training graph 30G includes the hypergraph 70G. The settings of the vertices and branches in each training graph 30G can be appropriately given in a manner that can appropriately represent the conditions of the training object. In step S102, the control unit 11 performs machine learning on the speculation module 5 using the acquired plurality of training graphs 30G. Similar to the embodiment, the machine learning is configured in the following manner: each training graph 30G is input into the feature weaving network 50 as an input graph, and thus the speculation module 5 is trained so that the speculation result obtained by the speculator 55 conforms to the correct solution (true value) of the task for each training graph 30G. Through the machine learning, a trained speculation module 5 that has the ability to obtain the solution to the speculation task of the hypergraph 70G can be generated. In step S103, the control unit 11 generates learning result data 125 representing the generated trained speculation module 5 and stores the generated learning result data 125 in a specified storage area. The learning result data 125 can be provided to the speculation device 2G at any time.
[0403] (Speculation device)
[0404] The hardware structure and software structure of the speculation device 2G can be the same as those of the speculation device 2 in the embodiment. In the sixth specific example, the speculation device 2G can speculate on the solution to the task of the hypergraph 70G through the same processing flow as that of the speculation device 2.
[0405] That is, in step S201, the control unit of the speculation device 2G operates as an acquisition unit to acquire the object graph 221G. The object graph 221G includes the hypergraph 70G. The settings of the vertices and branches in the object graph 221G can be appropriately given in a manner that can appropriately represent the conditions of the speculation object.
[0406] In step S202, the control unit operates as an estimation unit, and uses the trained estimation module 5 to estimate the solution to the task for the acquired object graph 221G. Similar to the above-described embodiment, the process of using the trained estimation module 5 to estimate the solution to the task for the object graph 221G is configured as follows: the object graph 221G is input as an input graph to the feature weaving network 50, and the result of estimating the solution to the task is obtained from the estimator 55. Specifically, the control unit inputs the object graph 221G to the trained estimation module 5 and executes the forward arithmetic process of the trained estimation module 5. As a result of executing the arithmetic process, the control unit can obtain from the estimator 55 the result of estimating the solution to the task for the object graph 221G.
[0407] In step S203, the control unit operates as an output unit and outputs information related to the estimation result. As an example, the control unit may also directly output the obtained estimation result to the output device. As another example, in the case where the above-described matching estimation result is obtained, similar to the first specific example and the like, the control unit may also cause at least a part of the estimation result to be matched. At this time, similar to the first specific example and the like, the estimation device 2G may be configured to perform matching online and in real time.
[0408] (Feature)
[0409] According to the sixth specific example, in the scenario where the task is expressed by the hypergraph 70G, it is possible to improve the estimation accuracy of the estimation module 5. In the model generation device 1, it is possible to expect the generation of a trained estimation module 5 that can accurately estimate the solution to the task expressed by the hypergraph 70G. In the estimation device 2G, by using such a trained estimation module 5, it is possible to expect an accurate estimation of the solution to the task expressed by the hypergraph 70G.
[0410] <4.2>
[0411] The method of machine learning is not limited to the examples of the above-described training process, and the omission, change, substitution, and addition of the process can be appropriately performed according to the embodiment. As another example, in the object of machine learning, in addition to the estimation module 5, an identification model may further be included. Correspondingly, in addition to the above-described training process, machine learning may further include performing adversarial learning between the estimation module 5 (especially the estimator 55) and the identification model.
[0412] Figure 20Schematically illustrate an example of the processing procedure of machine learning in this modification. Similar to the above-described embodiment, the speculation module 5 is configured to receive the input of the input graph Z and output the speculation result x of the solution to the task of the input graph Z. In contrast, the recognition model 59 is configured to receive the input of the correct answer (true value) X of the task corresponding to the input graph Z or the speculation result x of the speculation module 5, and identify the source of the input data (that is, whether the input data is the correct answer X or the speculation result x). The type of the machine learning model constituting the recognition model 59 is not particularly limited and can be appropriately selected according to the embodiment. The recognition model 59 can include, for example, a neural network or the like.
[0413] In an example of adversarial learning, the recognition model 59 is trained using the correct answer X for each training graph 30 (input graph Z) and the speculation result x of the speculation module 5 for each training graph 30 as input data, so as to reduce the error of the result of identifying the source of the input data. In the training step of the recognition model 59, the value of the parameter of the recognition model 59 is adjusted. In contrast, the value of the parameter of the speculation module 5 is fixed. Moreover, the speculation module 5 is trained to output the speculation result x that the recognition model 59 misidentifies. In the training step of the speculation module 5, the value of the parameter of the recognition model 59 is fixed. In contrast, the value of the parameter of at least a part (for example, the speculator 55) in the speculation module 5 is adjusted.
[0414] For each training method, a known method such as the error backpropagation method can be adopted. Each training step can be repeatedly executed alternately. Or, by inserting a gradient reversal layer between the speculation module 5 and the recognition model 59, the two training steps can be executed simultaneously. Each gradient reversal layer can be appropriately configured to directly pass the value during the forward propagation operation and reverse the value during the backward propagation.
[0415] According to this modification, corresponding to the speculation ability of the speculation module 5, the recognition model 59 obtains the ability to identify the source of the input data. On the other hand, corresponding to the recognition ability of the recognition model 59, the speculation module 5 obtains the ability to generate a speculation result approximate to the correct answer. By repeatedly executing each training step, the abilities of the speculation module 5 and the recognition model 59 are improved mutually. Thereby, the speculation module 5 can obtain the ability to generate a speculation result extremely close to the correct answer. As an example, this modification can be adopted in the case of generating the structure of the 5-2 specific example. Thereby, it can be expected to generate a trained speculation module 5 that can more accurately speculate the structure of the component satisfying the specified properties.
[0416] <4.3>
[0417] Figure 21 Schematically illustrate another example of the input form of the speculation module 5. As Figure 21As shown, in the described embodiment, the number of graphs input to the speculation module 5 at one time during the solution task may not be limited to one, and multiple input graphs (the first graph, the second graph, etc. in the figure) may be input to the speculation module 5 simultaneously. At this time, the structure of the speculation module 5 may be appropriately changed according to the number of input graphs.
[0418] Figure 22 An example of the structure of the speculation module 5H of this modification is schematically illustrated. The speculation module 5H is configured to receive the input of two graphs and perform the task of searching for another input graph (query) within one of the input graphs (pattern). The speculation module 5H is an example of the speculation module 5 of the described embodiment. For convenience of explanation, it is assumed that the first graph is the query and the second graph is the pattern. The roles of the first graph and the second graph can also be swapped.
[0419] Figure 22 In the example of, the speculation module 5H includes a first feature extractor 581, a second feature extractor 582, an interaction speculator 583, and a feature speculator 584. The first feature extractor 581 is configured to receive the input of the first graph and extract the features of the input first graph. The second feature extractor 582 is configured to receive the input of the second graph and extract the features of the input second graph. The first feature extractor 581 and the second feature extractor 582 are arranged in parallel on the input side. In one example, the first feature extractor 581 and the second feature extractor 582 may include a single feature extractor used in common (that is, the functions of the first feature extractor 581 and the second feature extractor 582 may also be performed by one feature extractor). In another example, the first feature extractor 581 and the second feature extractor 582 may be provided separately.
[0420] The interaction speculator 583 is configured to receive the results of the feature extraction of each graph by each feature extractor (581, 582) and speculate on the mutual relationship of the features between the graphs. The feature speculator 584 is configured to receive the speculation result of the interaction speculator 583 on the mutual relationship and output the result of searching for the first graph within the second graph. The form of the search result may be appropriately determined according to the embodiment. The search may include, for example: determining whether the pattern contains the query using a binary value, speculating on the number of queries contained in the pattern, searching for all parts of the pattern that contain the query, and searching for the largest common subgraph of the two graphs.
[0421] In one example, each feature extractor (581, 582) may be the feature weaving network 50. Correspondingly, the interaction speculator 583 and the feature speculator 584 may correspond to the speculator 55. At this time, each feature extractor (581, 582) may also receive the input of each graph in any form. In one example, each feature extractor (581, 582) may also receive the input of each graph in the form of the third specific example or the fourth specific example. That is, each graph may also be a general directed graph or a general undirected graph.
[0422] In another example, the interaction speculator 583 may be the feature weaving network 50. Correspondingly, the feature speculator 584 may correspond to the speculator 55. At this time, as long as each feature extractor (581, 582) can extract the features of each graph, its structure may not be particularly limited and can be appropriately determined according to the implementation manner. Each feature extractor (581, 582) may include any model. The interaction speculator 583 may receive the feature extraction results of each feature extractor (581, 582) in the graph form of the fifth - 2 specific example.
[0423] In this modification example, each training graph and the object graph include two graphs. One of the two graphs is a pattern and the other is a query. Except for these points, this modification example may be configured in the same manner as the above - mentioned embodiment. The model generation device 1 may be configured to implement the machine learning of the speculation module 5H. The speculation device 2 is configured to perform graph search using the trained speculation module 5H. The methods of machine learning and speculation processing may be the same as those in the above - mentioned embodiment.
[0424] According to this modification example, it is possible to improve the search accuracy of the speculation module 5H, which is configured to receive the input of two graphs and search for one graph within the other graph. In the model generation device 1, it is possible to expect the generation of the trained speculation module 5H for graph search with high accuracy. In the speculation device 2, by using such a trained speculation module 5H, it is possible to expect the execution of graph search with high accuracy.
[0425] <4.4>
[0426] In the above-described embodiment, machine learning may include reinforcement learning. At this time, the correct answer may be given by a prescribed rule in the same manner as in the above-described embodiment. A reward is calculated based on the correct answer given by the prescribed rule, and the estimation module 5 may be trained to maximize the total sum of the calculated rewards. The prescribed rule may be given in the real world, and the reward may also be obtained in the real world. Training the estimation module 5 so that the estimation result obtained by the estimator 55 conforms to the correct answer may include training by means of such reinforcement learning. At this time, the training of the estimation module 5 may be repeated while the estimation module 5 is being used in the estimation device 2. The estimation device 2 may include a learning processing unit 112, and the object graph 221 may be used as the training graph 30.
[0427] <4.5>
[0428] In the above-described embodiment, the encoder 505 may include a single encoder. However, the inventors of the present application have found the following problem: If the first input tensor and the second input tensor belong to different distributions (that is, the distributions of the feature quantities related to the first set and the second set are different from each other), there may be a situation where the domains of the inputs to the single encoder are twofold. Whenever the encoders are stacked, the bias caused by the difference in the distributions is amplified. As a result, the single encoder may fall into a state of solving two different tasks instead of a common task, and as a result, the accuracy of the encoder 505 may deteriorate, leading to deterioration of the overall accuracy of the estimation module 5.
[0429] By simply setting an encoder for each connection tensor as described above, the above problem can be eliminated. However, the method for eliminating the above problem is not limited to such an example. In the case where the encoder 505 includes a single encoder, the above problem may also be eliminated by adopting at least one of the following two methods.
[0430] (First method)
[0431] Figure 23An example of the structure of the characteristic weaving layer 500 that schematically illustrates a modified example is shown. The first method is a method of using the normalizer 508 to eliminate the differences in distribution. In this modified example, the encoder 505 of each characteristic weaving layer 500 may include a single encoder, and each characteristic weaving layer 500 may further include a normalizer 508 corresponding to each connection tensor. The normalizer 508 is configured to normalize the operation result of the encoder 505. That is, in the above-described embodiment, each characteristic weaving layer 500 may include two normalizers 508 corresponding to each connection tensor (2L normalizers 508 for the characteristic weaving layer 500 of the L-th layer). The method of normalization can be appropriately determined according to the embodiment. For the method of normalization, for example, a known method such as batch normalization can be adopted. As an example, each normalizer 508 may be configured to normalize the operation result of the encoder 505 through the operation shown in Equation 16 below.
[0432] [Equation 16]
[0433]
[0434] x represents the operation result of the encoder 505 on each connection tensor, u represents the average value of the operation result, and v represents the variance of the operation result. By providing the normalizer 508 corresponding to each connection tensor, the average value u and the variance v are calculated for each connection tensor (that is, for the first connection tensor and the second connection tensor respectively). According to the operation of the normalizer 508, the operation result of the encoder 505 on each connection tensor can be normalized in a manner that follows a distribution with an average value of 0 and a variance of 1. Thereby, the difference in distribution between the first input tensor and the second input tensor for the next characteristic weaving layer 500 can be eliminated.
[0435] Therefore, according to this modified example, even if the encoder 505 is shared among the connection tensors (that is, even if the encoder 505 includes a single encoder), the amplification of the bias caused by the difference in the distribution to which the first input tensor and the second input tensor belong can be suppressed by the action of the normalizer 508. As a result, the above problem can be eliminated, and the deterioration of the accuracy of the inference module 5 can be prevented.
[0436] In addition, in each of the above specific examples, each characteristic weaving layer 500 may further include a normalizer 508 corresponding to each connection tensor. In the sixth specific example (the case of a hypergraph), each characteristic weaving layer 500 may include K normalizers 508 corresponding to each connection tensor.
[0437] (Second method)
[0438] In the described embodiment, the encoder 505 of each feature weaving layer 500 may include a single encoder, and each feature weaving layer 500 may be further configured to assign a condition code to the third axis of each of the first input tensor and the second input tensor before performing the operation of the encoder 505. That is, the second method is a method of conditioning the input to the encoder 505. The second method may replace the first method or be used together with the first method.
[0439] The condition code may be assigned at any time from when each input tensor is obtained until the operation of the encoder 505 is performed. Moreover, the structure of the condition code may be appropriately determined according to the embodiment. In one example, the condition code may include a prescribed value corresponding to each input tensor. Each prescribed value may be appropriately selected in a manner capable of identifying each input tensor. As a specific example, the condition code assigned to the first input tensor may be a first value (e.g., 1), and the condition code assigned to the second input tensor may be a second value (e.g., 0). The portion to which it is assigned may be arbitrarily determined. In one example, each condition code may be uniformly appended to the very end of the third axis with respect to each element of the first axis and the second axis of each input tensor.
[0440] In this modification, when encoding the feature amount, the encoder 505 of each feature weaving layer 500 may identify the source of the data based on the condition code. Therefore, even when the first input tensor and the second input tensor belong to different distributions, during the machine learning process, the encoder 505 can be trained through the action of the condition code to eliminate the bias caused by the difference in distributions. Therefore, according to this modification, it is possible to prevent the accuracy deterioration of the speculation module 5 caused by the different distributions to which each input tensor belongs.
[0441] In addition, in each of the specific examples, each feature weaving layer 500 may further be configured to assign a condition code to the third axis of each of the first input tensor and the second input tensor before performing the operation of the encoder 505. In the sixth specific example (the case of a hypergraph), the condition codes for the respective input tensors may be appropriately determined in such a manner as to be able to identify at least one input tensor. As an example, the condition code assigned to the first input tensor may be a first value (e.g., 1), and the condition codes assigned to the second to K-th input tensors (i.e., input tensors other than the first input tensor) may be a second value (e.g., 0). At this time, a condition code may be assigned to the (K + 1)-th axis of each input tensor before the loop operation. Thus, in the part for assigning condition codes to each connection tensor, the number of loop iterations of the first input tensor can be determined based on the position of the first value. As a result, during the operation of the encoder 505, the input tensor (the i-th input tensor) that does not loop within each connection tensor can be determined. Therefore, by using the method described above to assign condition codes, when the encoder 505 performs operations on each of the K connection tensors, it is possible to identify the source of the data based on the condition codes. In addition, in the case of adopting the method described above, by assigning the first value to any one of the first to K-th input tensors and assigning the second value to the remaining K - 1 input tensors, it is possible to identify the source of the data based on the condition codes (the position of the first value) in the same manner as described above. Therefore, the object to which the first value is assigned is not limited to the first input tensor.
[0442] §5 Embodiment
[0443] In order to verify the effectiveness of the feature extraction performed by the feature weaving network 50, inference modules for embodiments and comparative examples are generated in the following four experiments. However, the present invention is not limited to the following embodiments.
[0444] [Method for Generating Samples]
[0445] First, as a graph expressing the given conditions, assume the directed bipartite graph of the first specific example. For the task, an example of the matching task, i.e., the stable marriage problem, is adopted. Assume that the number of matching objects (subjects) in each group is the same (M = N). Under this assumption, the method proposed in Reference 1 "Nikolaos Tziavelis, Ioannis Giannakopoulos, Katerina Doka, Nectarios Koziris, and Panagiotis Karras, "Equitable stable matchings in quadratic time", In Proceedings of Advances in Neural Information Processing Systems, pages 457–467, 2019." is adopted to create a list of the expected degrees of each object (vertex) for the other party according to the following distributions. Thus, samples for machine learning and evaluation are generated.
[0446] · First distribution: Uniform distribution (U)
[0447] The expected degree of each subject for the matching candidate (the other party) is random and is defined by the uniform distribution U(0, 1) (the larger the value, the higher the expected degree).
[0448] · Second distribution: Discrete distribution (D)
[0449] Each subject has an expected degree defined by U(0.5, 1) for a part of the matching candidates (the candidates that are the largest integer less than or equal to 0.4N), and an expected degree defined by U(0, 0.5) for the remaining candidates.
[0450] · Third distribution: Gaussian distribution (G)
[0451] The expected degree of each subject for the i-th matching candidate is defined by the Gaussian distribution N(i / N, 0.4).
[0452] · Fourth distribution: LibimSeTi (Lib)
[0453] Based on the two-dimensional distribution of the frequencies of each pair of probabilities inferred from the data (matching activity) obtained on the Czech web service "LibimSeTi", the expected degree of each subject of each group is defined (Reference 2 "Lukas Brozovsky and Vaclav Petricek, 'Recommender system for online dating service', Proceedings of the Conference Znalosti 2007, Ostrava, 2007. (Lukas Brozovsky and Vaclav Petricek, 'Recommender system for online dating service', In Proceedings of Conference Znalosti 2007, Ostrava, 2007.)").
[0454] By respectively selecting any one of the four distributions for the first group (first subset) and the second group (second subset), five different settings are produced (i.e., UU, DD, GG, UD, and Lib). For example, "UU" means that both groups define the expected degree of each subject with a uniform distribution (U). "UD" means that the expected degree of each subject belonging to one group is defined with a uniform distribution (U), and the expected degree of each subject belonging to the other group is defined with a discrete distribution (D). And for the five distribution settings and the sizes (N×N) specified in each of the following experiments, 1000 training samples (training data) and 1000 validation samples (validation data) are randomly generated respectively.
[0455] [Experiment I: Comparison with Other Machine Learning Models]
[0456] In Experiment I, the effectiveness of the present invention is verified by comparing the matching success rates between the solvers of the following embodiments and the solvers of the following comparative examples that include learned machine learning models.
[0457] <Embodiment>
[0458] By adopting the structure of the above-described embodiment, inference modules of the first embodiment (WN-6), the second embodiment (WN-18), the third embodiment (WN-18f), and the fourth embodiment (WN-18b) are produced.
[0459] In the first embodiment, the number of feature weaving layers is set to six layers. In the second to fourth embodiments, the number of feature weaving layers is set to eighteen layers. In each feature weaving layer of the first to fourth embodiments, a single encoder that is commonly used among the elements of the first axis of each connection tensor is adopted. For the structure of the encoder of the first to fourth embodiments, Figure 6BStructure. For the structure of the speculator in the first to fourth embodiments, the structure of the first specific example (including preprocessing) (Formulas 1 to 4) is adopted. In the first to fourth embodiments, as the operation for obtaining the final speculation result in the speculation stage (speculation device) where a match is obtained, the argmax operation is adopted. In the feature weaving network of the first embodiment, the residual structure is omitted (that is, no shortcut is provided). On the other hand, in the feature weaving networks of the second to fourth embodiments, shortcuts are provided in a manner connected to the inputs of the feature weaving layers for each even number.
[0460] In the first embodiment and the second embodiment, as the loss function in machine learning, a loss function λ for suppressing blocking pairs is adopted s and a loss function λ for making the speculation result converge to a symmetric doubly stochastic matrix c , and a loss function for making the speculation result converge in a manner that satisfies other criteria is not adopted. On the other hand, in the third embodiment, in addition to the loss functions (λ s , λ c ), a loss function λ for promoting fairness is further adopted f . In the fourth embodiment, in addition to the loss functions (λ s , λ c ), a loss function λ for promoting the balance between fairness and satisfaction is adopted b . Specifically, in the first to fourth embodiments, in the loss function λ of Formula 15, 0.7 is substituted into w s and 1 is substituted into w c . In the first embodiment and the second embodiment, 0 is substituted into the weights of other loss functions. In the third embodiment, 0.01 is substituted into w f , and 0 is substituted into the weights of other loss functions (w e , w b ). In the fourth embodiment, 0.01 is substituted into w b , and 0 is substituted into the weights of other loss functions (w f , w e ).
[0461] <Comparative Example>
[0462] In the first comparative example (MLP-O), the model proposed in Non-Patent Document 2 was adopted in the inference module. Specifically, the feature extractor that extracts features from a graph includes three intermediate layers, and the inference unit that infers a matching result based on the features obtained by the feature extractor includes one output layer. The machine learning method of the inference module also adopted the method proposed in Non-Patent Document 2. In the second comparative example (MLP), the same model structure as that of the first comparative example was adopted in the inference module. On the other hand, in the second comparative example, instead of the term corresponding to the following Equation 17 in the terms of the loss function proposed in Non-Patent Document 2, λ of the above Equations 8 to 9 was adopted c .
[0463] [Equation 17]
[0464]
[0465] In the third comparative example (GIN), the model proposed in Reference 3 “Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka, “How powerful are graph neural networks?”, Proceedings of the International Conference on Learning Representations, 2019.” was adopted in the inference module. Specifically, in the third comparative example, the inference module of the third comparative example is configured to include a feature extractor, a feature concatenator, and an inference unit. The feature extractor is configured to output a feature amount corresponding to each vertex through a two-layer model proposed in Reference 3. The feature concatenator is configured to concatenate the features of all vertices output from the feature extractor into a single flat vector. Further, the inference unit is configured to obtain an inference result of a match with N×M (in this example, M = N) flat vectors through the same one output layer as that of the first comparative example. By regarding the flat vector obtained from the inference unit as a matrix of N×M size, the third comparative example is configured to obtain an output having the same structure as the inference result of the match in each embodiment. For the machine learning method of the inference module, the method proposed in Reference 3 was adopted. For the loss function, the same loss function as that of the first embodiment was adopted.
[0466] In the fourth comparative example (DBM-6), the model proposed in Reference 4 “Daniel Gibbons, Cheng-Chew Lim, and Peng Shi, “Deep learning for bipartite assignment problems”, In Proceedings of IEEE International Conference on Systems, Man and Cybernetics, pages 2318-2325, 2019.” was adopted in the inference module. In the fourth comparative example, the number of layers constituting the feature extractor in the inference module was set to six layers. For the machine learning method of the inference module, the method proposed in Reference 4 was adopted. For the loss function, the same loss function as in the first embodiment was adopted. In the fifth comparative example (SSWN-6), the same model as in the fourth comparative example was adopted in the inference module. However, in the fifth comparative example, the encoder provided in each layer constituting the feature extractor was changed to Figure 6B the structure. The other structures and the machine learning method of the fifth comparative example adopted the same structures and methods as in the fourth comparative example.
[0467] <Machine learning and evaluation method>
[0468] For the sizes of the verification samples, four different sizes of 3×3, 5×5, 7×7, and 9×9 were adopted. The models of the first comparative example and the second comparative example do not accept inputs of different sizes. Therefore, in the training of the first comparative example and the second comparative example, training samples of each size were used. On the other hand, in the training of the first to fourth embodiments and the third to fifth comparative examples, 10×10 training samples were used. The number of training repetitions was set to 200,000 times, and the batch size was set to 8. Based on each of the distributions, training samples were randomly generated in each repetition, and the generated training samples were used for training. For the optimization algorithm, Adam was adopted (Reference 5 “Diederik P. Kingma and Jimmy Ba, “Adam: A Method for Stochastic Optimization”, Proceedings of the International Conference on Learning Representations, 2015. (Diederik P. Kingma and Jimmy Ba, “Adam: A method for stochastic optimization”, In Proceedings of International Conference on Learning Representations, 2015.)). The learning rate was set to 0.0001. Through these conditions, the trained inference modules (solvers) of the first to fourth embodiments and the first to fifth comparative examples were generated respectively.
[0469] Subsequently, using each of the generated trained inference modules, matching inference results were obtained for the verification samples of each size. Moreover, the success of the matching was set to obtaining a match without blocking pairs. Based on the obtained inference results, the probability of obtaining a match without blocking pairs was calculated as the success rate of the matching. Furthermore, the degree (cost) of the fairness (Sex equality, SEq(m; I)) of each matching result and the balance (Balance, Bal(m; I)) between fairness and satisfaction were calculated by the following Equations 18 to 21. And the probability of obtaining the best matching result for each index was calculated as the success rate of the matching that satisfies each index. Furthermore, for the second embodiment and the third embodiment, the average value of the fairness cost of the matching results for all the obtained verification samples was calculated. For the second embodiment and the fourth embodiment, the average value of the balance cost of the matching results for all the obtained verification samples was calculated.
[0470] [Equation 18]
[0471] SEq(m; I) = |P(m; A) - P(m; B)|…(Equation 18)
[0472] [Equation 19]
[0473] Bal(m; I) = max(P(m; A), P(m; B)) … (Equation 19)
[0474] [Equation 20]
[0475]
[0476] [Equation 21]
[0477]
[0478] M represents a match that holds based on the speculation result.
[0479] <Result>
[0480] Figure 24A And Figure 24B represents the result of calculating the success rate of the match for the verification samples of each size for the first to fourth embodiments ( Figure 24A ) and the first to fifth comparative examples ( Figure 24B ). Figure 25A And Figure 25B represents the result of calculating the success rate of the fair match for the verification samples of each size for the first to fourth embodiments ( Figure 25A ) and the first to fifth comparative examples ( Figure 25B ). Figure 25C represents the result of calculating the average value of the fairness cost of the match results for all verification samples for the second and third embodiments. Figure 26A And Figure 26B represents the result of calculating the success rate of the balanced match for the verification samples of each size for the first to fourth embodiments ( Figure 26A ) and the first to fifth comparative examples ( Figure 26B ). Figure 26C represents the result of calculating the average value of the balanced cost of the match results for all verification samples for the second and fourth embodiments. Figure 25C And Figure 26C Furthermore, it represents the ideal value (Ideal) of the average value of each index.
[0481] As Figure 24A And Figure 24B shown, compared with each comparative example, each embodiment can accurately speculate a stable match. In the first to third comparative examples, when the number of subjects is 5 or more, it is difficult to speculate a stable match. Conventionally, as a method for solving the task of the graph, a network model like the first to third comparative examples was considered suitable. However, according to Figure 24A And Figure 24BAs can be seen from the results shown, these network models are not suitable for solving the task of bilateral matching expressed by a directed bipartite graph. In contrast, according to the embodiments, stable matching results can be obtained with higher accuracy than in the comparative examples. In particular, the number of parameters in the first embodiment is approximately 25,000, while the number of parameters in the first comparative example and the second comparative example is approximately 28,000, the number of parameters in the third comparative example and the fourth comparative example is approximately 29,000, and the number of parameters in the fifth comparative example is approximately 25,000. From this result, it can be seen that according to the present invention, even if the number of parameters is of the same degree, the success rate of stable matching can be improved.
[0482] Moreover, as Figure 25A , Figure 25B , Figure 26A and Figure 26B show, in each of the embodiments, the success rate of stable matching is higher than that in each of the comparative examples in terms of the fairness and balance metrics. From this result, it can be seen that according to the present invention, the success rate of stable matching that meets the specified criteria can also be improved. Furthermore, as Figure 25C and Figure 26C show, it can be seen that by adding fairness and balance metrics to the loss function, the average value of the cost of each metric can be made closer to the ideal value. From this result, it can be seen that in the present invention, by changing the loss function, the ability obtained by the inference module can be changed. That is, it can be seen that the present invention can be applied not only to the stable marriage problem but also to various other tasks of graphs by appropriately changing the loss function. For example, it is implied that even if the type of graph is changed, by setting a loss function corresponding to the task, a trained inference module that has obtained the ability to solve the corresponding task can be generated.
[0483] [Experiment II: Comparison between Embodiments and Existing Algorithms]
[0484] In Experiment II, the effectiveness of the present invention was verified by comparing the matching results between the solvers of the following embodiments and the solvers of the existing algorithms of the following comparative examples.
[0485] <Embodiments>
[0486] By changing the number of feature weaving layers in the third embodiment to 60 layers, an inference module of the fifth embodiment (WN-60f) was fabricated. For other structures of the fifth embodiment, the same structure as that of the third embodiment was adopted. Moreover, by changing the number of feature weaving layers in the fourth embodiment to 60 layers, an inference module of the sixth embodiment (WN-60b) was fabricated. For other structures of the sixth embodiment, the same structure as that of the fourth embodiment was adopted. In the training of the fifth embodiment and the sixth embodiment, 30×30 training samples were used. For other conditions of machine learning, the same conditions as in Experiment I were adopted.
[0487] <Comparative Example>
[0488] In the sixth comparative example (GS), the GS algorithm proposed in Non-Patent Document 1 was adopted in the solver. In the seventh comparative example (PolyMin), the algorithm disclosed in Reference 6 “Dan Gusfield and Robert W. Irving, “The stable marriage problem: structure and algorithms”, MIT press, 1989. (Dan Gusfield and Robert W Irving, “The stable marriage problem: structure and algorithms”, MIT press, 1989.)” was adopted in the solver. Specifically, in the seventh comparative example, an algorithm that minimizes the costs of the following equations 22 and 23 (Reg(m; I), Egal(m; I)) was adopted in the solver instead of the costs of fairness (SEq(m; I)) and balance (Bal(m; I)).
[0489] [Equation 22]
[0490]
[0491] [Equation 23]
[0492] Egal(m; I) = P(m; A) + P(m; B) … (Equation 23)
[0493] In the eighth comparative example (DACC), the algorithm proposed in Reference 7 “Piotr Dworczak, “Deferred acceptance with compensation chains”, In Proceedings of the 2016 ACM Conference on Economics and Computation, pages 65 - 66, 2016. (Piotr Dworczak, “Deferred acceptance with compensation chains”, In Proceedings of the 2016 ACM Conference on Economics and Computation, pages 65 - 66, 2016.)” was adopted in the solver. In the ninth comparative example (Power Balance), the algorithm proposed in the above Reference 1 was adopted in the solver. In addition, the computational complexity of the algorithm in the eighth comparative example is O(n 4 ), and the computational complexity of the algorithm in the ninth comparative example is O(n 2 ).
[0494] <Evaluation Method>
[0495] Regarding the sizes of the verification samples, two different sizes of 20×20 and 30×30 were adopted. Using the trained speculation modules of the respective embodiments and the solvers of the respective comparative examples, the matching results for the verification samples of each size were obtained. And, using the same method as in Experiment I, for the fifth embodiment and the respective comparative examples, the average value of the fairness cost of the matching results for all the obtained verification samples was calculated. Similarly, for the sixth embodiment and the respective comparative examples, the average value of the balance cost of the matching results for all the obtained verification samples was calculated. Moreover, for each embodiment, a machine learning model was adopted in the speculation module, and it was not always possible to accurately derive the correct solution. Therefore, the matching success rate was calculated using the same method as in Experiment I. Furthermore, in each embodiment, since the speculation results of failed matches were included, it was considered that the average value of the cost of each index deteriorated. Therefore, between the comparative example with the best average value of each cost (the sixth comparative example for UD and the ninth comparative example for others) and each embodiment, the costs of each index of the matching results for each verification sample were compared, and the probabilities of the cases where each embodiment was better than the comparative example (win), the same (tie), and the cases where each embodiment was worse than the comparative example (loss+unstable) were calculated.
[0496] <Results>
[0497] [Table 1]
[0498]
[0499] [Table 2]
[0500]
[0501] Table 1 shows the results of calculating the respective evaluation indexes for the fifth embodiment and the respective comparative examples. Table 2 shows the results of calculating the respective evaluation indexes for the sixth embodiment and the respective comparative examples. As shown in Table 1 and Table 2, for the verification samples of UD, in the sixth comparative example, the best results for the costs of each index were obtained. For the other verification samples, in the ninth comparative example, the best results for the costs of each index were obtained. However, each embodiment also obtained results equivalent to those of the respective comparative examples. According to the comparison results of the costs of each index in the respective matches with the comparative example having the best results, in each embodiment, it was possible to obtain a match with good (win) or the same (tie) cost of each index with a probability of 75.7% to 95%. Furthermore, according to each embodiment, for the verification samples of 20×20 and 30×30, stable matches could be accurately speculated. From these results, it can be known that according to the present invention, it is possible to generate a trained speculation module that is not inferior to the existing best algorithms and has the ability to accurately speculate stable matches with high precision.
[0502] [Experiment III: Verification of the effectiveness for large-sized graphs]
[0503] In Experiment III, the sizes of the training samples and the verification samples were changed to 100×100, the distribution of each sample was changed to only UU, and through the same process as in Experiment II, the matching results were compared between the solvers of the following respective embodiments and the solvers of the existing algorithms of the comparative examples (the sixth to ninth comparative examples), thereby verifying the effectiveness of the present invention.
[0504] <Embodiment>
[0505] The number of feature weaving layers in the third embodiment was changed to 80 layers, and the operation for obtaining the inference result was changed from the argmax operation to the Hungarian algorithm, thereby fabricating the speculation module of the seventh embodiment (WN-80f + Hungarian). For other structures of the seventh embodiment, the same structure as that of the third embodiment was adopted. Moreover, the number of feature weaving layers in the fourth embodiment was changed to 80 layers, and the operation for obtaining the inference result was changed from the argmax operation to the Hungarian algorithm, thereby fabricating the speculation module of the eighth embodiment (WN-80b + Hungarian). For other structures of the eighth embodiment, the same structure as that of the fourth embodiment was adopted. In the training of the seventh and eighth embodiments, 100×100 training samples were used. The number of training repetitions was set to 300,000 times. Regarding other conditions of machine learning, the same conditions as those in Experiment I were adopted.
[0506] <Evaluation method>
[0507] Similar to Experiment II, the trained speculation modules of the respective embodiments and the solvers of the respective comparative examples were used to obtain the matching results for the verification samples of various sizes. And for the seventh embodiment and the respective comparative examples, the average value of the fairness cost of the matching results for all the obtained verification samples was calculated. For the eighth embodiment and the respective comparative examples, the average value of the balance cost of the matching results for all the obtained verification samples was calculated. Furthermore, for each embodiment, the success rate of matching was calculated.
[0508] <Results>
[0509] [Table 3]
[0510]
[0511] Table 3 shows the results of calculating the respective evaluation indices for each of the embodiments and comparative examples. From the results shown in Table 3, for a size of 100×100, the success rate of stable matching deteriorated, but in each of the embodiments, stable matching could be predicted with relatively high accuracy. Also, as in Experiment II, for the cost of each index, in each of the embodiments, results comparable to those of the comparative example with the best results could be obtained. Therefore, it can be understood that according to the present invention, even for a large-size (100×100) problem, a trained prediction module can be generated that is comparable to existing best algorithms and has the ability to predict stable matching with relatively high accuracy.
[0512] [Experiment IV: Verification of the effectiveness of each modification]
[0513] In Experiment IV, samples of the same size as in Experiment II were used, and the success rate of matching was compared between the following embodiments and comparative examples to verify the effectiveness of the form of the encoder 505, the form of the normalizer, and the form of the conditional code assigned, each separately set corresponding to each connection tensor.
[0514] <Embodiment>
[0515] After the encoder of each feature weaving layer in the fifth embodiment, a normalizer configured to normalize the operation result of the encoder is provided in a manner shared by each connection tensor, thereby producing a prediction module of the ninth embodiment (WN-60f-symmetric). After the encoder of each feature weaving layer in the fifth embodiment, a normalizer configured to normalize the operation result of the encoder is provided corresponding to each connection tensor, and an operation of assigning a conditional code to each input tensor is added, thereby producing a prediction module of the tenth embodiment (WN-60f-asymmetric). The encoder of each feature weaving layer in the fifth embodiment is changed from a single encoder to an encoder used separately for each connection tensor, thereby producing a prediction module of the eleventh embodiment (WN-60f-dual). A prediction module of the twelfth embodiment (WN-60f(20)) having the same structure as the prediction module of the fifth embodiment is produced. The operation result at the time of obtaining the prediction result in the fifth embodiment is changed from an argmax operation to a Hungarian algorithm, thereby producing a prediction module of the thirteenth embodiment (WN-60f + Hungarian). For the other structures of the ninth to eleventh embodiments and the twelfth embodiment, the same structure as that of the fifth embodiment was adopted.
[0516] Similarly, after the encoder of each characteristic braided layer in the sixth embodiment, a normalizer configured to normalize the operation result of the encoder is set in a manner shared by each connection tensor, thereby producing the inference module of the fourteenth embodiment (WN-60b-symmetric). After the encoder of each characteristic braided layer in the sixth embodiment, a normalizer configured to normalize the operation result of the encoder is set corresponding to each connection tensor, and an operation of assigning a conditional code to each input tensor is added, thereby producing the inference module of the fifteenth embodiment (WN-60b-asymmetric). The encoder of each characteristic braided layer in the sixth embodiment is changed from a single encoder to an encoder used for each connection tensor, thereby producing the inference module of the sixteenth embodiment (WN-60b-double). The inference module of the seventeenth embodiment (WN-60b (20)) having the same structure as the inference module of the sixth embodiment is produced. The operation result when obtaining the inference result in the sixth embodiment is changed from the argmax operation to the Hungarian algorithm, thereby producing the inference module of the eighteenth embodiment (WN-60b + Hungarian). As for the other structures of the fourteenth to sixteenth embodiments and the eighteenth embodiment, the same structures as those of the sixth embodiment are adopted.
[0517] <Comparative Example>
[0518] The number of layers constituting the feature extractor in the fifth comparative example is changed to 60 layers, and the fairness index (λ f , weight is 0.01) is added to the loss function, thereby producing the inference module of the tenth comparative example (SSWN-60f). In addition, the number of layers constituting the feature extractor in the fifth comparative example is changed to 60 layers, and the fairness index (λ b , weight is 0.01) is added to the loss function, thereby producing the estimation module of the eleventh comparative example (SSWN-60b). The other structures of the tenth comparative example and the eleventh comparative example adopt the same structure as the fifth comparative example.
[0519] <Machine Learning and Evaluation Methods>
[0520] In the training of the twelfth and seventeenth embodiments, 20×20 training samples were used. In the training of the other embodiments and comparative examples, 30×30 training samples were used. The other conditions for machine learning were the same as those of Experiment I.
[0521] Regarding the sizes of the verification samples, two different sizes, 20×20 and 30×30, were adopted. Using the trained inference modules of each embodiment and each comparative example, the matching results for the verification samples of each size were obtained. And, using the same method as in Experiment I, the success rate of matching was calculated for each embodiment and each comparative example. For the ninth to thirteenth embodiments and the tenth comparative example, the average value of the fairness cost of the matching results for all the obtained verification samples was calculated. For the fourteenth to eighteenth embodiments and the eleventh comparative example, the average value of the balance cost of the matching results for all the obtained verification samples was calculated. Additionally, for UU, DD, and GG with unbiased distributions, the verification of the tenth embodiment, the eleventh embodiment, the fifteenth embodiment, and the sixteenth embodiment was omitted.
[0522] <Results>
[0523] [Table 4]
[0524]
[0525] [Table 5]
[0526]
[0527] Table 4 shows the results of calculating the average values of the success rate of matching and the fairness cost for the ninth to thirteenth embodiments and the tenth comparative example. Table 5 shows the results of calculating the average values of the success rate of matching and the balance cost for the fourteenth to eighteenth embodiments and the eleventh comparative example. For the success rate and the cost of each index, better results were obtained for each embodiment than for each comparative example. Based on this result, the effectiveness of the present invention can be presented.
[0528] Moreover, it is known that in the case where there is a bias in the input distribution, in the structure where the encoder is commonly used (for example, the ninth embodiment, the fourteenth embodiment), the success rate of matching and the cost of each index may deteriorate. In contrast, according to the form in which the encoder is separately set corresponding to the connection tensor (the eleventh embodiment, the sixteenth embodiment) and the form adopting the first method and the second method of <4.5> (the tenth embodiment, the fifteenth embodiment), it is known that the success rate of matching and the cost of each index can be improved.
[0529] Furthermore, if these forms are compared, the possibility that the results of the form adopting the first method and the second method of <4.5> are better is higher than that of the form where the encoder is separately set. Based on this result, it is known that even without separately setting the encoder (that is, while suppressing the increase in parameters), the first method and the second method of <4.5> are effective as methods for suppressing the amplification of the bias caused by the difference in the input distribution.
Claims
1. A model generation device, comprising: an acquisition unit configured to acquire a plurality of training graphs; and a learning processing unit configured to perform machine learning of a speculation module using the acquired plurality of training graphs. In the model generation device, the speculation module includes a feature weaving network and a speculator, the feature weaving network is configured to receive an input of an input graph and output feature information related to the input graph by including a plurality of feature weaving layers, the speculator is configured to speculate a solution to a task of the input graph based on the feature information output from the feature weaving network, each feature weaving layer receives an input of a three-layer first input tensor and a second input tensor, the first input tensor is configured to have: a first axis arranged with first vertices belonging to a first set of the input graph as elements, a second axis arranged with second vertices belonging to a second set of the input graph as elements, and a third axis arranged with feature amounts related to branches flowing from the first vertices to the second vertices as elements, the second input tensor is configured to have: a first axis arranged with the second vertices as elements, a second axis arranged with the first vertices as elements, and a third axis arranged with feature amounts related to branches flowing from the second vertices to the first vertices as elements, each feature weaving layer includes an encoder, each feature weaving layer is configured to: fix the third axis of the second input tensor and cycle through the other axes of the second input tensor one by one, and connect the feature amounts of each element of the first input tensor and the cycled second input tensor, thereby generating a three-layer first connection tensor, fix the third axis of the first input tensor and cycle through the other axes of the first input tensor one by one, and connect the feature amounts of each element of the second input tensor and the cycled first input tensor, thereby generating a three-layer second connection tensor, divide the generated first connection tensor into each element of the first axis and input it to the encoder, perform the operation of the encoder, thereby generating a three-layer first output tensor corresponding to the first input tensor, and divide the generated second connection tensor into each element of the first axis and input it to the encoder, perform the operation of the encoder, thereby generating a three-layer second output tensor corresponding to the second input tensor, the encoder is configured to derive relative feature amounts of each element based on the feature amounts of all the input elements, the first feature weaving layer among the plurality of feature weaving layers is configured to receive the first input tensor and the second input tensor from the input graph, the feature information includes the first output tensor and the second output tensor output from the last feature weaving layer among the plurality of feature weaving layers, and The machine learning is configured as follows: each training graph is input into the feature weaving network as the input graph, thereby training the inference module so that the inference result obtained by the inferencer conforms to the correct answer of the task for each training graph.
2. The model generation device according to claim 1, wherein The feature weaving network has a residual structure.
3. The model generation device according to claim 1 or 2, wherein The encoder of each feature weaving layer includes a single encoder commonly used among elements of the first axis of the first connection tensor and the second connection tensor.
4. The model generation device according to claim 3, wherein Each feature weaving layer further includes a normalizer corresponding to each connection tensor, and the normalizer is configured to normalize the operation result of the encoder.
5. The model generation device according to claim 3, wherein Each feature weaving layer is further configured to assign conditional codes to the third axes of the first input tensor and the second input tensor respectively before performing the operation of the encoder.
6. The model generation device according to claim 1 or 2, wherein Each training graph is a directed bipartite graph, A plurality of vertices constituting the directed bipartite graph are divided into any one of two subsets, The first set of the input graph corresponds to one of the two subsets of the directed bipartite graph, The second set of the input graph corresponds to the other of the two subsets of the directed bipartite graph.
7. The model generation device according to claim 6, wherein The task is a matching task between two groups, and is a task of determining the best pairing of objects belonging to each group, The two subsets of the directed bipartite graph correspond to the two groups in the matching task, The vertices belonging to each subset of the directed bipartite graph correspond to the objects belonging to each group, The feature quantity related to the branch flowing from the first vertex to the second vertex corresponds to the degree of expectation of an object belonging to one of the two groups for an object belonging to the other group, The feature quantity related to the branch flowing from the second vertex to the first vertex corresponds to the degree of expectation of an object belonging to the other of the two groups for an object belonging to one of them.
8. The model generation device according to claim 1 or 2, wherein Each training graph is an undirected bipartite graph, A plurality of vertices constituting the undirected bipartite graph are divided into any one of two subsets, The first set of the input graph corresponds to one of the two subsets of the undirected bipartite graph, The second set of the input graph corresponds to the other of the two subsets of the undirected bipartite graph.
9. The model generation device according to claim 8, wherein The task is a matching task between two groups, and is a task of determining the best pairing of objects belonging to each group, The two subsets of the undirected bipartite graph correspond to the two groups in the matching task, The vertices of the undirected bipartite graph belonging to each subset correspond to the objects belonging to each of the groups. The feature quantity related to the branch flowing from the first vertex to the second vertex and the feature quantity related to the branch flowing from the second vertex to the first vertex both correspond to the cost or reward for pairing an object belonging to one of the two groups with an object belonging to the other.
10. The model generation device according to claim 1 or 2, wherein Each of the training graphs is a directed graph. The first vertex of the input graph belonging to the first set corresponds to the starting point of the directed branch constituting the directed graph. The second vertex belonging to the second set corresponds to the end point of the directed branch. The branch flowing from the first vertex to the second vertex corresponds to the directed branch flowing from the starting point to the end point. The branch flowing from the second vertex to the first vertex corresponds to the directed branch flowing from the starting point into the end point.
11. The model generation device according to claim 10, wherein The directed graph is configured as a network. The task is a task of inferring a state of affairs occurring on the network. The feature quantity related to the branch flowing from the first vertex to the second vertex and the feature quantity related to the branch flowing from the second vertex to the first vertex both correspond to the connection attributes between the respective vertices constituting the network.
12. The model generation device according to claim 1 or 2, wherein Each of the training graphs is an undirected graph. The first vertex of the input graph belonging to the first set and the second vertex belonging to the second set respectively correspond to the respective vertices constituting the undirected graph. The first input tensor and the second input tensor input to each of the feature weaving layers are the same as each other. The first output tensor and the second output tensor generated by each of the feature weaving layers are the same as each other. The process of generating the first output tensor and the process of generating the second output tensor are executed jointly.
13. The model generation device according to claim 1 or 2, wherein Each of the training graphs is a vertex feature graph including a plurality of vertices, and is a vertex feature graph in which each vertex has an attribute. The first vertex of the input graph belonging to the first set and the second vertex belonging to the second set respectively correspond to the respective vertices constituting the vertex feature graph. The feature quantity related to the branch flowing from the first vertex to the second vertex corresponds to the attribute of the vertex of the vertex feature graph corresponding to the first vertex. The feature quantity related to the branch flowing from the second vertex to the first vertex corresponds to the attribute of the vertex of the vertex feature graph corresponding to the second vertex.
14. The model generation device according to claim 13, wherein For each element of the third axis of each of the first input tensor and the second input tensor from the input graph, information representing the relationality between the corresponding vertices of the vertex feature graph is concatenated or multiplied. The task is a task of inferring a state of affairs derived from the vertex feature graph.
15. The model generation device according to claim 13, wherein The task is to infer the relationships between the vertices that make up the vertex feature map.
16. A speculation device, comprising: An acquisition unit configured to acquire an object graph; A speculation unit configured to use a speculation module trained by machine learning to speculate on the solution to the task of the acquired object graph; and An output unit configured to output information related to the result of speculating on the solution to the task. In the speculation device, The speculation module includes a feature weaving network and a speculator, The feature weaving network is configured to receive the input of the input graph and output feature information related to the input graph by including a plurality of feature weaving layers, The speculator is configured to speculate on the solution to the task of the input graph based on the feature information output from the feature weaving network, Each feature weaving layer receives the input of a three-layer first input tensor and a second input tensor, The first input tensor is configured to have: a first axis arranged with the first vertices belonging to the first set of the input graph as elements, a second axis arranged with the second vertices belonging to the second set of the input graph as elements, and a third axis arranged with the feature amounts related to the branches flowing from the first vertices to the second vertices as elements, The second input tensor is configured to have: a first axis arranged with the second vertices as elements, a second axis arranged with the first vertices as elements, and a third axis arranged with the feature amounts related to the branches flowing from the second vertices to the first vertices as elements, Each of the feature weaving layers includes an encoder, Each of the feature weaving layers is configured to: Fix the third axis of the second input tensor and cycle through the other axes of the second input tensor one by one, and connect the feature amounts of each element of the first input tensor and the cycled second input tensor, thereby generating a three-layer first connection tensor, Fix the third axis of the first input tensor and cycle through the other axes of the first input tensor one by one, and connect the feature amounts of each element of the second input tensor and the cycled first input tensor, thereby generating a three-layer second connection tensor, Divide the generated first connection tensor into each element of the first axis and input it to the encoder, and perform the operation of the encoder, thereby generating a three-layer first output tensor corresponding to the first input tensor, and Divide the generated second connection tensor into each element of the first axis and input it to the encoder, and perform the operation of the encoder, thereby generating a three-layer second output tensor corresponding to the second input tensor, The encoder is configured to derive the relative feature amounts of each element based on the feature amounts of all the input elements, The first feature weaving layer among the plurality of feature weaving layers is configured to receive the first input tensor and the second input tensor from the input graph, The feature information includes the first output tensor and the second output tensor output from the last feature weaving layer among the plurality of feature weaving layers, and The process of using the trained inference module to infer the solution to the task for the object graph is configured as follows: the object graph is input into the feature weaving network as the input graph, and the result of inferring the solution to the task is obtained from the inferencer.
17. A model generation method, executed by a computer to perform the following steps: Obtain a plurality of training graphs; and Use the obtained plurality of training graphs to perform machine learning on the inference module. In the model generation method, the inference module includes a feature weaving network and an inferencer, the feature weaving network is configured to receive the input of the input graph and output the feature information related to the input graph by including a plurality of feature weaving layers, the inferencer is configured to infer the solution to the task for the input graph according to the feature information output from the feature weaving network, each feature weaving layer receives the input of a three-layer first input tensor and a second input tensor, the first input tensor is configured to have: a first axis arranged with the first vertices belonging to the first set of the input graph as elements, a second axis arranged with the second vertices belonging to the second set of the input graph as elements, and a third axis arranged with the feature quantities related to the branches flowing from the first vertices to the second vertices as elements, the second input tensor is configured to have: a first axis arranged with the second vertices as elements, a second axis arranged with the first vertices as elements, and a third axis arranged with the feature quantities related to the branches flowing from the second vertices to the first vertices as elements, each feature weaving layer includes an encoder, each feature weaving layer is configured to: Fix the third axis of the second input tensor and cycle through the other axes of the second input tensor one by one, and connect the feature quantities of each element of the first input tensor and the cycled second input tensor, thereby generating a three-layer first connection tensor, Fix the third axis of the first input tensor and cycle through the other axes of the first input tensor one by one, and connect the feature quantities of each element of the second input tensor and the cycled first input tensor, thereby generating a three-layer second connection tensor, Divide the generated first connection tensor into each element of the first axis and input it into the encoder, and perform the operation of the encoder, thereby generating a three-layer first output tensor corresponding to the first input tensor, and Divide the generated second connection tensor into each element of the first axis and input it into the encoder, and perform the operation of the encoder, thereby generating a three-layer second output tensor corresponding to the second input tensor, the encoder is configured to derive the relative feature quantity of each element according to the feature quantities of all the input elements, the first feature weaving layer among the plurality of feature weaving layers is configured to receive the first input tensor and the second input tensor from the input graph, The feature information includes the first output tensor and the second output tensor output from the last feature weaving layer among the plurality of feature weaving layers, and The machine learning is configured such that each training graph is input as the input graph to the feature weaving network, thereby training the inference module so that the inference result obtained by the inferencer conforms to the correct answer of the task for each training graph.
18. A computer-readable storage medium that can be read and store a model generation program by a computer, the model generation program causing the computer to execute the following steps: Obtain a plurality of training graphs; and Perform machine learning of the inference module using the obtained plurality of training graphs. In the model generation program, The inference module includes a feature weaving network and an inferencer, The feature weaving network is configured to receive an input of an input graph and output feature information related to the input graph by including a plurality of feature weaving layers, The inferencer is configured to infer a solution to the task of the input graph based on the feature information output from the feature weaving network, Each feature weaving layer receives an input of a three-layer first input tensor and a second input tensor, The first input tensor is configured to have: a first axis arranged with first vertices belonging to a first set of the input graph as elements, a second axis arranged with second vertices belonging to a second set of the input graph as elements, and a third axis arranged with feature amounts related to branches flowing from the first vertices to the second vertices as elements, The second input tensor is configured to have: a first axis arranged with the second vertices as elements, a second axis arranged with the first vertices as elements, and a third axis arranged with feature amounts related to branches flowing from the second vertices to the first vertices as elements, Each feature weaving layer includes an encoder, Each feature weaving layer is configured to: Fix the third axis of the second input tensor and cycle through the other axes of the second input tensor one by one, and connect the feature amounts of each element of the first input tensor and the cycled second input tensor, thereby generating a three-layer first connection tensor, Fix the third axis of the first input tensor and cycle through the other axes of the first input tensor one by one, and connect the feature amounts of each element of the second input tensor and the cycled first input tensor, thereby generating a three-layer second connection tensor, Divide the generated first connection tensor into each element of the first axis and input it to the encoder, and execute the operation of the encoder, thereby generating a three-layer first output tensor corresponding to the first input tensor, and Divide the generated second connection tensor into each element of the first axis and input it to the encoder, and execute the operation of the encoder, thereby generating a three-layer second output tensor corresponding to the second input tensor, The encoder is configured to derive the relative feature amount of each element based on the feature amounts of all the input elements. The first feature weaving layer among the plurality of feature weaving layers is configured to receive the first input tensor and the second input tensor from the input graph. The feature information includes the first output tensor and the second output tensor output from the last feature weaving layer among the plurality of feature weaving layers, and The machine learning is configured in such a way that each training graph is input as the input graph to the feature weaving network, thereby training the inference module so that the inference result obtained by the inferencer conforms to the correct answer of the task for each training graph.
19. A model generation device, comprising: An acquisition unit configured to acquire a plurality of training graphs; And A learning processing unit configured to perform machine learning of an inference module using the acquired plurality of training graphs. In the model generation device, Each of the training graphs is K hypergraphs, K is 3 or more, The plurality of vertices constituting the hypergraph are divided into any one of K subsets, The inference module includes a feature weaving network and an inferencer, The feature weaving network is configured to receive the input of the input graph by including a plurality of feature weaving layers and output feature information related to the input graph, The inferencer is configured to infer the solution of the task for the input graph based on the feature information output from the feature weaving network, Each feature weaving layer receives the input of K (K + 1)-layer input tensors, The i-th input tensor among the K input tensors is configured as follows: In the first axis, the i-th vertex belonging to the i-th subset of the input graph is arranged as an element, In the j-th axis from the second axis to the K-th axis, cyclically starting from the i-th, the vertices belonging to the (j - 1)-th subset are arranged as elements, and In the (K + 1)-th axis, the feature quantities related to the branches flowing out from the i-th vertex belonging to the i-th subset to the combinations of the vertices belonging to (K - 1) other subsets are arranged as elements, Each feature weaving layer includes an encoder, Each feature weaving layer is configured to: For the i-th input tensor from the first to the K-th, fix the (K + 1)-th axis of the other (K - 1) input tensors other than the i-th, and cycle the other axes other than the (K + 1)-th axis of the other input tensors in a manner consistent with each axis of the i-th input tensor, and connect the feature quantities of each element of the i-th input tensor and the cycled other input tensors, thereby generating K (K + 1)-layer connection tensors, and Divide each of the generated K connection tensors into each element of the first axis and input it to the encoder, execute the operation of the encoder, thereby generating K (K + 1)-layer output tensors respectively corresponding to the K input tensors, The encoder is configured to derive the relative feature quantity of each element based on the feature quantities of all the input elements, The first feature weaving layer among the plurality of feature weaving layers is configured to receive the K input tensors from the input graph. The feature information includes the K output tensors output from the last feature weaving layer among the multiple feature weaving layers, and the machine learning is configured such that each of the training graphs is input as the input graph to the feature weaving network, thereby training the inference module so that the inference result obtained by the inferencer conforms to the correct answer of the task for each of the training graphs.
Citation Information
Patent Citations
Method and system for scheduling requests in a portable computing device
CN104303149A
Image region positioning method and device and medical image processing device
CN109978838A