Edge identification device and edge identification method
The edge identification apparatus addresses the challenge of identifying important edges in graphs by swapping nodes and edges to identify critical nodes in a derived graph, enabling effective edge identification and network optimization.
Patent Information
- Application Number
- JP2021181007
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-11-05
AI Technical Summary
Existing techniques struggle to identify important edges in a graph, which is crucial for maintaining or improving the overall characteristics of a network.
An edge identification apparatus and method that involves acquiring a first graph, deriving a second graph by swapping nodes and edges, identifying important nodes in the second graph, and then determining the corresponding edges in the original graph.
This approach effectively identifies important edges in the original graph, allowing for targeted modifications to maintain or improve the network's overall characteristics, such as suppressing the spread of infectious diseases or optimizing social network communication.
Smart Images

Figure 0007697350000001 
Figure 0007697350000002 
Figure 0007697350000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for identifying important edges in a graph having nodes and edges.
Background Art
[0002] Patent Document 1 discloses a graph analysis apparatus that identifies related properties related to an edge by evaluating whether or not it affects other nodes using edge evaluation means. The edge evaluation means calculates the edge importance based on the number of related properties, edge difference degree, edge strength, and the like.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The inventors of the present invention have recognized that it is desirable to identify important edges in order to maintain or improve the overall characteristics of a network represented by a graph.
[0005] An object of the present invention is to provide a technique capable of identifying important edges in a graph.
Means for Solving the Problems
[0006] In order to solve the above problems, an edge identification apparatus according to an aspect of the present invention includes an acquisition unit that acquires a first graph including nodes and edges, a derivation unit that derives a second graph by swapping the nodes and edges of the acquired first graph, a first identification unit that identifies a plurality of nodes that satisfy a predetermined condition regarding the degree of importance in the derived second graph, and a second identification unit that identifies the edges of the first graph corresponding to each node of the identified second graph.
[0007] Another aspect of the present invention is an edge identification method. This method includes An edge identification method executed by a computer, obtaining a first graph including nodes and edges, deriving a second graph by swapping the nodes and edges of the obtained first graph, identifying a plurality of nodes that satisfy a predetermined condition regarding the degree of importance in the derived second graph, and identifying the edges of the first graph corresponding to each node of the identified second graph.
Advantages of the Invention
[0008] According to the present invention, it is possible to provide a technique for identifying important edges in a graph.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Embodiments for Carrying Out the Invention
[0010] Before specifically describing the embodiments, an overview will be explained. Many people and things in the world interact with their surroundings and form a network. To improve not only specific parts within the network but also the overall characteristics of the network, it is effective to identify important nodes or edges within the network and act on them.
[0011] For example, considering a network for analyzing the spread of an infectious disease with cities as nodes and the number of movers between cities as edges, in order to effectively suppress the spread of the infectious disease through the edges in the entire network, it is necessary to identify the important edges whose number of movers should be controlled. In addition to the network related to the spread of infectious diseases, in various networks, it is desirable to identify important edges for maintaining or improving the overall characteristics of the network.
[0012] However, although techniques for identifying important nodes are known, it is difficult to identify important edges. Therefore, in the embodiments, in a Line Graph (line graph) obtained by swapping the nodes and edges of a graph representing a network, important nodes are identified, and the edges of the original graph corresponding to the identified important nodes are identified.
[0013] FIG. 1 shows the configuration of an edge identification device 10 according to an embodiment. The edge identification device 10 includes an acquisition unit 12, a derivation unit 14, a first identification unit 16, and a second identification unit 18.
[0014] The configuration of the edge identification device 10 can be realized hardware-wise by the CPU, memory, and other LSIs of an arbitrary computer, and software-wise by a program loaded into the memory, etc. Here, however, functional blocks realized by their cooperation are depicted. Therefore, it is understood by those skilled in the art that these functional blocks can be realized in various forms by hardware only, software only, or a combination thereof.
[0015] The acquisition unit 12 acquires a first graph G1 including a plurality of nodes and a plurality of edges, and supplies the data of the acquired first graph to the derivation unit 14. Each edge connects two nodes. A signal value is set for each node, and a weight is set for each edge. If there is no edge between two nodes, it can be said that the weight of that edge is zero. Each node does not necessarily have node position information, that is, node coordinate information.
[0016] The first graph G1 represents a predetermined network. The network is not particularly limited, and examples include networks related to the spread of infectious diseases, social networks for analyzing human communication with people as nodes and the amount of conversation between people as edges, sensor networks, networks of data for machine learning, neural networks, traffic networks, and various other networks.
[0017] The acquisition unit 12 may acquire the first graph G1 generated by a server device (not shown) or the like via a network such as the Internet, or may acquire it from a recording medium (not shown). Further, the acquisition unit 12 may acquire information of predetermined nodes, acquire signal values detected by a sensor (not shown), etc., derive the signal values of the nodes and the weights of the edges based on the acquired signal values, and derive the first graph G1.
[0018] FIGS. 2(a) to 2(e) are diagrams for explaining the edge identification process by the edge identification device 10 of FIG. 1. FIG. 2(a) shows an example of the first graph G1. The first graph G1 includes seven nodes N1 to N7 and ten edges E1 to E10. For example, the node N1 is connected to the node N2 by the edge E1, connected to the node N3 by the edge E8, and connected to the node N7 by the edge E7.
[0019] In the illustrated example, a first graph G1 with the number of nodes being "7" and the number of edges being "10" is shown to prevent complication of the explanation. However, for example, the number of nodes may be "several tens" or more, or may be "100" or more.
[0020] The derivation unit 14 derives a second graph G2 by swapping the nodes and edges of the acquired first graph G1, and supplies the data of the derived second graph G2 to the first specifying unit 16. The second graph G2 is a known Line Graph.
[0021] Specifically, for each edge of the first graph G1, the derivation unit 14 sets a node of the second graph G2. Next, for each pair of two edges having a common node in the first graph G1, the derivation unit 14 sets an edge between the corresponding two nodes in the second graph G2. In this way, the derivation unit 14 derives the second graph G2, which is a Line Graph, from the first graph G1.
[0022] FIG. 2(b) shows the relationship between the first graph G1 in FIG. 2(a) and the second graph G2 derived from the first graph G1. In FIG. 2(b), the illustration of the edge symbols is omitted. FIG. 2(c) shows the derived second graph G2. The nodes of the second graph G2 are shown as unfilled circles, and the edges are shown as dashed lines. The second graph G2 includes ten nodes NL1 to NL10 and twenty-one edges EL1 to EL21.
[0023] For the edge E1 of the first graph G1, the derivation unit 14 sets the node NL1 of the second graph G2. Similarly, for the edges E2 to E10 of the first graph G1, the derivation unit 14 sets the nodes NL2 to NL10 of the second graph G2. The plurality of nodes of the second graph G2 correspond one-to-one to the plurality of edges of the first graph G1, and the number of nodes of the second graph G2 is equal to the number of edges of the first graph G1.
[0024] Regarding the pair of two edges E1 and E2 having a common node N2 in the first graph G1, the derivation unit 14 sets an edge EL3 between the corresponding two nodes NL1 and NL2 in the second graph G2.
[0025] The derivation unit 14 sets an edge EL1 between two corresponding nodes NL1 and NL8 in the second graph G2 with respect to a pair of two edges E1 and E7 having a common node N1 in the first graph G1.
[0026] The derivation unit 14 sets an edge EL2 between two corresponding nodes NL1 and NL3 in the second graph G2 with respect to a pair of two edges E1 and E8 having a common node N1 in the first graph G1.
[0027] The derivation unit 14 sets an edge EL7 between two corresponding nodes NL8 and NL3 in the second graph G2 with respect to a pair of two edges E7 and E8 having a common node N1 in the first graph G1. The derivation unit 14 performs similar processing to set a plurality of edges in the second graph G2.
[0028] The first specifying unit 16 specifies, in the derived second graph G2, a plurality of nodes that satisfy a predetermined condition regarding the degree of importance according to a predetermined algorithm, and supplies the data of the specified plurality of nodes to the second specifying unit 18. The number of nodes to be specified is predetermined by the user. The specified plurality of nodes are important nodes in the second graph G2.
[0029] Depending on what kind of network the first graph G1 represents, etc., the predetermined condition can be set as appropriate, and thus the algorithm for specifying a plurality of nodes that satisfy the predetermined condition can be set as appropriate.
[0030] As an algorithm for specifying important nodes, although not limited to these, for example, known node sampling techniques such as FastGSSS, MaxCutoff, or MinSpec / MinTrac can be used.
[0031] FastGSSS is an algorithm that selects a plurality of nodes so that the spatial distribution of the selected plurality of nodes is generally uniform. In this case, satisfying a predetermined condition means that the spatial distribution of the specified plurality of nodes is generally uniform.
[0032] Since the way information spreads is different for each node in the graph, FastGSSS selects a plurality of important nodes as follows by taking advantage of this difference in the way information spreads for each node.
[0033] The first specifying unit 16 first selects a node that is most strongly connected to other nodes in the graph. A node that is most strongly connected to other nodes is a node that is connected to the most other nodes via a predetermined number or fewer edges. The predetermined number is a natural number and is set in advance by the user. In the example of FIG. 2(c), if the predetermined number is "1", the other nodes to which the node NL1 is connected via one edge are the nodes NL8, NL3, and NL2, and the number of other nodes is "3". If the predetermined number is "2", for example, the other nodes to which the node NL1 is connected via one or two edges are the nodes NL8, NL3, NL2, NL7, NL10, NL9, and NL4, and the number of other nodes is "7".
[0034] The first specifying unit 16 selects a node that has the least overlap in the way information spreads with the already selected important nodes and is most strongly connected to the unselected nodes. A node that is most strongly connected to the unselected nodes is a node that is connected to the most unselected nodes via a predetermined number or fewer edges. The first specifying unit 16 repeats this process to select a plurality of nodes.
[0035] The way information spreads for each node is obtained as follows. [First Process] The first specifying unit 16 sets only the signal value on the target node that is the target for obtaining the way information spreads to an arbitrary non-zero value, and sets the signal values on the nodes other than the target node to zero.
[0036] [Second Process] The first specifying unit 16 causes the signal value on the target node to propagate through the edges connected to the target node. Since the edges connected according to the node are different, the results of propagation are also different. Assume that as a result of propagation, the signal value of a node connected to the target node via a predetermined number or fewer edges becomes a non-zero value. After propagation, the set of nodes having non-zero signal values, that is, the set of nodes connected to the target node via a predetermined number or fewer edges, is defined as the way the information of the target node spreads.
[0037] [Third Process] The first specifying unit 16 executes all of the above-described first process and second process for each of all the nodes, and acquires the way the information spreads for each node.
[0038] Since the algorithm of FastGSSS is a known technique, further description thereof is omitted. For more detailed processing, for example, it is disclosed in the reference document "A. Sakiyama, Y. Tanaka, T. Tanaka, and A. Ortega, “Eigendecomposition-Free Sampling Set Selection for Graph Signals,” IEEE Transactions on Signal Processing, vol. 67, no. 10, pp. 2679-2692, May 2019."
[0039] MaxCutoff is an algorithm that selects a plurality of nodes such that the bandwidth of the data that can be restored from the selected plurality of nodes is maximized. For the specific processing of MaxCutoff, for example, it is disclosed in the reference "A. Anis, A. Gadde, and A. Ortega, “Efficient sampling set selection for bandlimited graph signals using graph spectral proxies,” IEEE Trans. Signal Process., vol. 64, no. 14, pp. 3775-3789, Jul. 2016."
[0040] MinSpec / MinTrac is an algorithm that selects a plurality of nodes such that the bandwidth of the data that can be restored from the selected plurality of nodes is maximized. For the specific processing of MinSpec / MinTrac, for example, it is disclosed in the reference "S. Chen, R. Varma, A. Sandryhaila, and J. Kovacevic, “Discrete signal processing on graphs: Sampling theory,” IEEE Trans. Signal Process., vol. 63, no. 24, pp. 6510-6523, Dec. 2015."
[0041] FIG. 2(d) shows the nodes identified in the second graph G2 of FIG. 2(c). As an example, it is assumed that four nodes NL3, node NL4, node NL6, and node NL8 that satisfy a predetermined condition are identified.
[0042] The second specifying unit 18 specifies the edges of the first graph G1 corresponding to each node of the specified second graph G2.
[0043] FIG. 2(e) shows the edges of the first graph G1 corresponding to the specified nodes in FIG. 2(d). The second specifying unit 18 specifies the edge E8 of the first graph G1 corresponding to the node NL3 of the second graph G2, the edge E3 corresponding to the node NL4, the edge E5 corresponding to the node NL6, and the edge E7 corresponding to the node NL8. These edges E3, E5, E7, and E8 correspond to important edges in the first graph G1.
[0044] The important edges can also be said to be the edges that have a relatively large influence on the overall characteristics of the network represented by the first graph G1. That is, by reducing the weights of these important edges, setting them to zero, or increasing them, the overall characteristics of the network can be greatly changed. Also, by leaving only these important edges and deleting the edges other than the specified edges, the change in the overall characteristics of the network can be suppressed and the characteristics can be easily maintained.
[0045] Next, an example where the first graph G1 has more nodes will be described. FIG. 3 shows an example of the first graph G1 having 100 nodes. From FIG. 3 to FIG. 5, the nodes are represented by filled circles and the edges are represented by solid lines. The first graph G1 in FIG. 3 includes seven clusters, which are groups of a plurality of nodes connected by a plurality of edges, and the clusters are also connected by edges.
[0046] FIG. 4 shows a graph obtained by deleting a plurality of edges specified by the edge specifying device 10 of FIG. 1 in the first graph G1 of FIG. 3. In this example, since the FastGSSS algorithm is used, the plurality of specified edges, that is, the plurality of deleted edges, are generally spatially evenly distributed in the graph.
[0047] Therefore, in the graph of FIG. 4, although a plurality of edges have been deleted, the number of clusters is the same as that in FIG. 3, and the connections between clusters are also maintained to a certain extent. That is to say, in FIG. 4, compared with FIG. 3, the relationship between clusters is maintained to a certain extent and does not change significantly. Thus, since the plurality of identified edges are spatially evenly distributed, even if these plurality of edges are deleted, the number of clusters and the relationship between clusters, that is, the structural characteristics of the original first graph G1 can be generally maintained.
[0048] FIG. 5 shows a graph obtained by deleting a plurality of edges identified by the edge identification method of the comparative example in the first graph G1 of FIG. 3. The number of edges deleted in FIG. 5 is the same as that in the example of FIG. 4.
[0049] The edge identification method of the comparative example does not convert the first graph G1 into a Line Graph, but applies an algorithm called MaxDegree to the first graph G1, and generally identifies a predetermined number of edges in descending order of degree in the first graph G1. The degree of an edge is the sum of the number of nodes connected to the nodes connected to one end of the target edge and the number of nodes connected to the nodes connected to the other end of the edge.
[0050] In the graph of FIG. 5, compared with FIG. 3, the clusters have disappeared because a plurality of nodes that formed the clusters have become isolated. This is because the edge identification method of the comparative example does not consider maintaining the structure of the original first graph G1. Therefore, in the graph of FIG. 5, the characteristics of the original first graph G1 are greatly damaged.
[0051] For example, consider the case where the first graph G1 of FIG. 3 represents a network for analyzing the spread of an infectious disease. As shown in FIG. 4, by deleting the identified edges that are generally evenly distributed spatially, or by reducing the number of movers, which is the weight of the identified edges, the number of infected people can be effectively reduced, and since the spatial structure of the network of the original first graph G1 can be maintained to a certain extent, the economic effects due to people's movement can be maintained to a certain extent.
[0052] On the other hand, as shown in FIG. 5, when a plurality of edges of the first graph G1 are identified using the MaxDegree algorithm of the comparative example and the identified edges are deleted, isolated nodes appear and clusters cannot be maintained. Therefore, compared with the present embodiment, the economic effect due to human movement is significantly reduced. This is the same even when the number of movers, which is the weight of the identified edges, is reduced. That is, in the comparative example, important edges cannot be identified to maintain or improve the overall characteristics of the network.
[0053] Also, although different from FIG. 3, it is also conceivable that the first graph G1 represents a social network in which members of a group are nodes, state values representing the current states of the members such as the stress state value and happiness level of the members are used as the signal values of the nodes, and the current communication volume such as the conversation volume between members is used as the edges. In the embodiment, when the current communication volume between members is zero, important edges are identified assuming that there are temporarily edges between these members. In the embodiment, since a plurality of edges having a relatively large influence on the overall characteristics of the social network can be identified, it is possible to identify which communication between members is important. Therefore, by reducing, setting to zero, or increasing the communication volume, which is the weight of these important edges, the future states of a plurality of members of the group can be effectively improved. Regarding the plurality of identified edges, which edge weight should be changed and how to change it to most improve the future states of the members of the group can be identified by simulation.
[0054] FIG. 6 is a flowchart showing the edge identification process of the edge identification device 10 in FIG. 1. The acquisition unit 12 acquires the first graph G1 (S10), and the derivation unit 14 derives a second graph G2 by swapping the nodes and edges of the acquired first graph G1 (S12). The first identification unit 16 identifies a plurality of nodes satisfying a predetermined condition in the derived second graph (S14), and the second identification unit 18 identifies the edges of the first graph corresponding to each node of the identified second graph (S16).
[0055] According to the embodiment, in the second graph G2 obtained by swapping the nodes and edges of the first graph G1, important nodes can be identified. Since the edges of the first graph G1 corresponding to each node of the identified second graph G2 are identified, important edges of the original first graph G1 can be identified. In this way, important edges can be identified in order to maintain or improve the overall characteristics of the network represented by the first graph G1.
[0056] The present invention has been described based on the embodiments. It should be understood by those skilled in the art that the embodiments are merely examples, and various modifications are possible for the combination of each component and each processing process, and such modifications are also within the scope of the present invention.
Description of Reference Numerals
[0057] G1... first graph, G2... second graph, 10... edge identification device, 12... acquisition unit, 14... derivation unit, 16... first identification unit, 18... second identification unit.
Claims
1. An acquisition unit that acquires a first graph including nodes and edges; A derivation unit that derives a second graph by swapping the nodes and edges of the acquired first graph; A first identification unit that identifies a plurality of nodes that satisfy a predetermined condition regarding the degree of importance in the derived second graph; A second identification unit that identifies the edges of the first graph corresponding to each node of the identified second graph; An edge identification device, characterized by comprising the above.
2. An edge identification method executed by a computer, comprising: A step of acquiring a first graph including nodes and edges; A step of deriving a second graph by swapping the nodes and edges of the acquired first graph; A step of identifying a plurality of nodes that satisfy a predetermined condition regarding the degree of importance in the derived second graph; A step of identifying the edges of the first graph corresponding to each node of the identified second graph; An edge identification method, characterized by comprising the above.
Citation Information
Patent Citations
Device, method and program for analyzing graph
JP2011028454A
Graph simplification method, graph simplification program, and information processing system
JP2020077299A
Information processing apparatus, information processing method, and information processing program
JP2020101893A
Information processor, method for processing information, and information processing program
JP2020187644A