System for updating most influential nodes of dynamic graphs (computer-implemented method, system, and computer program)
The method addresses the inefficiency of calculating subgraph centrality in dynamic graphs by using partial spectral factorization and update data to approximate the dominant eigenpair, achieving linear computational complexity and reduced memory needs.
Patent Information
- Application Number
- JP2024206348
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-12
AI Technical Summary
Existing methods for calculating subgraph centrality in dynamic graphs are inefficient due to high computational complexity and memory requirements, especially when graphs are large and constantly updated.
A computer-implemented method that uses partial spectral factorization data and update data to approximate the dominant eigenpair of the adjacency matrix for a dynamic graph, allowing for efficient calculation of subgraph centrality and identification of influential nodes.
The method reduces computational complexity from cubic to linear, enabling efficient handling of large dynamic graphs and reducing memory requirements by storing only update data, thus improving performance and resource utilization.
Smart Images

Figure 2025089276000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to graph theory, and more specifically, to using subgraph centrality measures and sub-spectrum factorization to determine influential nodes within a graph.
Background Art
[0002] Graph theory is the study of graphs, which are mathematical structures used to model pairwise relationships between objects. In this context, a graph consists of vertices or nodes, and lines called edges that connect them. Graphs are widely used in applications that model the dynamic characteristics of many types of relationships and processes in physical, biological, social, and information systems. Thus, many practical problems in modern technology, science, and business applications can usually be represented by graphs.
[0003] Node centrality is a widely used measure for determining the relative importance of nodes within a complete network or graph. Centrality metrics assign numbers or rankings to nodes within a graph corresponding to their positions on the network. Node centrality can be used to determine which nodes are important within a complex network, to understand influencers, or to find hot-spot links. For example, node centrality is typically used to determine how influential a person is within a social network, or in the theory of space syntax, how important a room is within a building, or how frequently a road is used within an urban network.
[0004] Subgraph centrality is a widely used centrality measure for identifying the most influential nodes of a graph. When the graph is static and can fit into system memory, the subgraph centrality of each node is equal to the corresponding diagonal entry of the exponent of the adjacency matrix.
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, many real-world graphs need to be either accessed from secondary storage or updated over time because they are dynamic graphs.
Means for Solving the Problems
[0006] According to an aspect of the present disclosure, a computer-implemented method is provided that includes obtaining an adjacency matrix corresponding to a first subgraph of a graph. The adjacency matrix is used to calculate a function of the adjacency matrix for a second subgraph of the graph based on previously stored partial spectral factorization data of the adjacency matrix and update data associated with additional nodes of the first subgraph. The second subgraph is a subset of the first subgraph.
[0007] According to another aspect of the present disclosure, a computer-implemented method for updating a dynamic graph is provided that includes obtaining an adjacency matrix for a first subgraph of the dynamic graph at time step t+1. The method also includes accessing a projection matrix for a previously stored adjacency matrix for a second subgraph of the dynamic graph. The accessed projection matrix corresponds to the previous time step t together with the previously stored adjacency matrix. The accessed projection matrix is used, together with update data for the dynamic graph received at time step t+1, to approximate a projection matrix for the adjacency matrix at time step t+1. The approximated projection matrix is then used to approximate a dominant eigenpair for the adjacency matrix at time step t+1. Next, the dynamic graph is updated based on the approximated dominant eigenpair, and then the updated dynamic graph is stored.
[0008] According to another aspect of the present disclosure, a system is provided. The system includes a circuit configured to obtain an adjacency matrix corresponding to a first sub-graph of a graph. The circuit is also configured to calculate a function of the adjacency matrix for a second sub-graph of the graph based on pre-stored partial spectral factorization data of the adjacency matrix and update data associated with additional nodes of the first sub-graph. The second sub-graph is a subset of the first sub-graph.
[0009] According to another aspect of the present disclosure, a computer program product for calculating a function of an adjacency matrix is provided. The computer program product includes a computer-readable storage medium having program instructions embodied thereon. The program instructions are executable by a system to cause the system to obtain an adjacency matrix corresponding to a first sub-graph of a graph. The program instructions are executable by a system to cause the system to calculate a function of the adjacency matrix for a second sub-graph of the graph based on pre-stored partial spectral factorization data of the adjacency matrix and update data associated with additional nodes of the first sub-graph.
[0010] Additional technical features and advantages are realized through the techniques of the present disclosure. Embodiments and aspects of the present disclosure are described in detail herein and are considered a part of the claimed subject matter. For a better understanding, refer to the detailed description and the drawings.
Brief Description of the Drawings
[0011] In the following description, details of preferred embodiments are provided with reference to the following drawings.
[0012]
Figure 1
[0013]
Figure 2
[0014]
Figure 3
[0015]
Figure 4
[0016]
Figure 5
[0017]
Figure 6
[0018]
Figure 7
[0019]
Figure 8
[0020]
Figure 9
DETAILED DESCRIPTION OF THE INVENTION
[0021] In graph theory and network analysis, centrality metrics assign numbers or rankings to nodes in a graph corresponding to their positions in the network. Subgraph centrality is a widely used centrality measure for identifying the most influential nodes (or vertices) in a graph. If the graph is static and can fit into system memory, the subgraph centrality of each node is equal to the corresponding diagonal entry of the exponent of the adjacency matrix. However, many real-world graphs are either accessed from secondary storage or updated over time. Also, in the case of dynamic graphs, which are graphs that are created and updated dynamically, their corresponding adjacency matrices may not fit into system memory due to their large and dynamically changing sizes.
[0022] A common operation involving any graph, specifically a dynamic graph, is to identify the most influential nodes of the graph (network) in order to perform graph analysis. Each node of the graph is associated with a real-valued scalar known as a centrality score, and the goal is to identify the N ∈ N nodes of the graph with the maximum centrality scores (the top N recommendations). Some applications of the top N centralities include identifying the most influential nodes in a social network, identifying the major hubs in a road and urban network, and identifying the most important proteins in a cell network. Different types of centrality metrics exist, and these are used as computational metrics for quantifying the relative importance of each node in a graph within an application. Some of these types of centrality metrics include degree centrality, betweenness centrality, eigenvector centrality, closeness centrality, PageRank centrality, and the like. Another centrality metric is subgraph centrality, which quantifies the number of closed walks of nodes starting and ending at a node. Subgraph centrality gives a measure of the involvement of a node in the number of subgraphs within a graph.
[0023] In one or more embodiments of the present disclosure, an exponential subgraph centrality measure for a graph is calculated, which is further used to calculate the most influential nodes of the graph. Subgraph centrality measures are used in the analysis of networks from a variety of applications, including but not limited to biology, neuroscience, and economics. Given an undirected graph of n ∈ N nodes associated with an n×n adjacency matrix A, the subgraph centrality of the i-th graph vertex is the magnitude of the i-th diagonal entry of the matrix exponential e A equal to.
[0024] For small graphs, the calculation of the subgraph centrality of each vertex can be achieved at a cost of O(n 3 ) by matrix diagonalization. However, graphs arising from modern data analysis tasks often feature a vast number of vertices and edge connections, and thus, due to the cubic complexity of the exact subgraph centrality calculation, the cost becomes very high. Also, this gives rise to the problem that the graph does not fit into system memory or becomes out-of-core due to its large size and high computational complexity.
[0025] In one or more embodiments of the present disclosure, a method is provided for approximately identifying the most influential nodes of a graph G when only a small subset of the rows of the adjacency matrix A of the graph is present in system memory at any given time.
[0026] In one or more embodiments of the present disclosure, a method is provided for calculating centrality scores, specifically, subgraph centrality scores for different nodes of a graph, in applications involving graphs that evolve dynamically over different time steps.
[0027] In one or more embodiments of the present disclosure, a method is provided for identifying the most influential nodes of a graph (i.e., top N recommendations) by accessing each submatrix of an adjacency matrix (for a subgraph) only once. To achieve this, information on matrix updates is accumulated by updating sub-spectrum factorization data for the adjacency matrix through Rayleigh-Ritz projection.
[0028] One or more embodiments of the present disclosure disclosed herein provide a reduction in the computational complexity associated with graph update calculations. In one or more embodiments, the computational complexity is of the order of linear complexity rather than the cubic complexity of conventional techniques, leading to savings in the computing resources of the application and performance enhancement. Further, in one or more embodiments of the present disclosure, the memory requirements associated with storing large and complex, dynamically evolving graphs are simplified because only the update data for each graph update is accessed only once.
[0029] According to an aspect of the present disclosure, a computer-implemented method is provided. The computer-implemented method includes obtaining an adjacency matrix corresponding to a first subgraph of a graph. The computer-implemented method includes calculating a function of the adjacency matrix for a second subgraph of the graph based on previously stored sub-spectrum factorization data of the adjacency matrix and update data associated with nodes added to the second subgraph to obtain the first subgraph. The second subgraph is a subset of the first subgraph. Calculating the function of the adjacency matrix based on the sub-spectrum factorization data requires accessing each adjacency matrix only once, which is computationally less expensive and enables storing graphs that are also very large in system memory because only the update data for such graphs needs to be updated in the form of sub-spectrum factors.
[0030] In various embodiments, the adjacency matrix of the first sub-graph is associated with the state of the graph at the current time instance, and the previously stored adjacency matrix of the second sub-graph is associated with the state of the graph at a previous time instance, such that the previous time instance is prior to the current time instance. As a result, in order to calculate any function on the adjacency matrix, only the calculations regarding the updated data need to be newly performed, while the previously stored data is stored compactly using the partial spectrum factors.
[0031] In various embodiments, the calculated function is used to calculate sub-graph centrality data about the graph. Sub-graph centrality is used to determine the top N set of nodes of the graph. The top N set of nodes are the nodes of the graph having indices associated with the highest sub-graph centrality data. The top N set of nodes provides a dataset of nodes that are connected to most of the other nodes of the graph stored in the database. This dataset of the top N nodes may be stored in the database as the most influential nodes of the graph and may be used in different types of computational applications related to accessing data of the influential nodes of the graph, such as protein analysis computational applications, social network platforms, author-oriented content platforms, and the like.
[0032] In various embodiments, the graph is a dynamic graph and the top N set of nodes are the most influential nodes of the dynamic graph. A dynamic graph includes data that is frequently updated and thus requires efficient computations such as recalculating the most influential nodes upon each update. Thus, using the top N set of nodes calculated through sub-graph centrality calculated based on the partial spectrum factorization data is faster and computationally more efficient.
[0033] In various embodiments, the dynamic graph corresponds to a graph approximation of a social network, and thus, the nodes of the dynamic graph represent users of the social network, the edges of the dynamic graph represent connections between users of the social network, and further, thus, the most influential nodes of the dynamic graph represent users of the influencer user type in the social network. The social network may be implemented as a social networking platform and store data of users of the social network in a database. The graph approximation of the dynamic graph is stored in this database and accessed each time a request for access to data of the most influential nodes of the graph is received. Such a request may be received from a computing device through which a user accesses a social networking application hosted by the social networking platform.
[0034] In various embodiments, the partial spectrum factorization data is determined based on a Rayleigh-Ritz projection calculation and based on a stage of calculating a projection matrix for an adjacency matrix of a first subgraph. The Rayleigh-Ritz projection calculation has a stage of calculating a Ritz pair, which is further used as an approximation of a dominant eigenvalue of the adjacency matrix. The problem of k-subgraph centrality is essentially equivalent to calculating k algebraic dominant eigenpairs of the adjacency matrix. By using the Rayleigh-Ritz projection calculation, the complexity of calculating the subgraph centrality is transformed into the calculation of the dominant approximate eigenpairs that can be performed using a sparse eigenvalue solver. Therefore, the computational complexity of calculating the subgraph centrality is reduced by using the Rayleigh-Ritz projection calculation.
[0035] In various embodiments, the computer-implemented method comprises receiving updated data associated with nodes added to a second subgraph at time instance t+1. Further, based on a Rayleigh-Ritz projection calculation, a projection matrix for an adjacency matrix at time instance t+1 is calculated. For this projection matrix, k leading eigenpairs are calculated. The computer-implemented method comprises determining k leading Ritz pairs for the k leading eigenpairs of the projection matrix and then updating t+1. The steps described above are repeated until a specified end condition is reached. By this repetition, since the graph data is continuously updated dynamically, an iterative calculation of the k leading eigenpairs is provided. This helps to maintain up-to-date data about the k leading eigenpairs of the graph and can be used to perform further calculations in real time to calculate subgraph centrality, for example.
[0036] In various embodiments, the specified end condition is based on the availability of updated data for the graph. As a result, as long as updates to the graph are received, the latest influential node data is also efficiently calculated and stored in a database or memory that stores the graph data.
[0037] In various embodiments, based on the step of determining the k leading Ritz pairs, the subgraph centrality for the graph is calculated at the time when the end condition is reached. This ensures that the problem of calculating subgraph centrality is a convergence problem.
[0038] In various embodiments, the step of calculating a projection matrix for the adjacency matrix has the step of calculating k leading Ritz pairs for the projection matrix such that they are the same as the k leading eigenpairs of the adjacency matrix at time instance t+1. Thereby, one approach is provided for identifying the k leading Ritz pairs for the projection matrix, which will be used for subgraph centrality calculation. The step of calculating the k leading Ritz pairs in this manner has a high-precision value associated with the identification of the most influential nodes of the graph.
[0039] In various embodiments, the step of calculating a projection matrix for an adjacency matrix has a step of calculating k leading eigenpairs of the adjacency matrix based on a Schur component. By this approach, an accuracy value close to 70% is provided when identifying the most influential nodes of a graph.
[0040] In various embodiments, the step of calculating a projection matrix for an adjacency matrix has a step of calculating a matrix formed by calculating orthonormal basis components for k leading eigenpairs of the adjacency matrix. By this approach, a high accuracy value associated with the identification of the most influential nodes of a graph is provided.
[0041] In various embodiments, the step of calculating a projection matrix for an adjacency matrix has a step of approximating k leading eigenpairs of the adjacency matrix using k dominant eigenvectors of a symmetric matrix. By this approach, a high accuracy value associated with the identification of the most influential nodes of a graph is provided.
[0042] In various embodiments, a previously stored adjacency matrix corresponds to a previous time instance t, and the adjacency matrix corresponds to a subsequent time instance t + 1. This provides once-limited access to each submatrix of the adjacency matrix.
[0043] According to aspects of the present disclosure, a computer-implemented method for updating a dynamic graph is provided. The computer-implemented method comprises obtaining an adjacency matrix for a first subgraph of the dynamic graph at time step t+1. The computer-implemented method comprises accessing a projection matrix for a previously stored adjacency matrix for a second subgraph of the dynamic graph. The computer-implemented method comprises approximating a projection matrix for the adjacency matrix at time step t+1 based on the accessed projection matrix for the previously stored adjacency matrix at time step t and update data for the dynamic graph received at time step t+1. The computer-implemented method comprises approximating a dominant eigenpair for the adjacency matrix at time step t+1 based on the approximated projection matrix. The computer-implemented method comprises updating the dynamic graph based on the approximated dominant eigenpair. The computer-implemented method comprises storing the updated dynamic graph. The computer-implemented method provides a computationally efficient calculation of the most influential nodes of the dynamic graph, which is also kept up-to-date by accessing the memory storing the dynamic graph in an efficient manner.
[0044] According to aspects of the present disclosure, a system is provided. The system comprises a circuit configured to obtain an adjacency matrix corresponding to a first subgraph of a graph. The circuit is also configured to calculate a function of the adjacency matrix for a second subgraph of the graph based on previously stored partial spectral factorization data of the adjacency matrix and update data associated with nodes added to the second subgraph to obtain the first subgraph, where the second subgraph is a subset of the first subgraph. The system provides a computing device and / or software program configured to efficiently calculate data for the most influential nodes of the graph, which is a computational approximation of real-world applications such as social networking applications, protein analysis computational applications, and the like.
[0045] According to an aspect of the present disclosure, a computer program product for calculating a function of an adjacency matrix is provided. The computer program product includes a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by the system to cause the system to obtain an adjacency matrix corresponding to a first sub-graph of a graph. Further, for a second sub-graph of the graph, the system is caused to calculate a function of the adjacency matrix based on previously stored partial spectral factorization data of the adjacency matrix and update data associated with nodes added to the second sub-graph to obtain the first sub-graph, where the second sub-graph is a subset of the first sub-graph. The computer program product thus provides a computing device and / or software program configured to efficiently calculate data of the most influential nodes for a graph, which is a computational approximation of real-world applications such as social networking applications, protein analysis calculation applications, and the like.
[0046] Various aspects of the present disclosure are illustrated by descriptions, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or at least partially in a time-overlapping manner.
[0047] An embodiment of a computer program product (referred to herein as a "CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media") collectively included within a set of one or more storage devices, which collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. A computer-readable storage medium can be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media are floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the major surfaces of disks), or any suitable combination of the foregoing. A computer-readable storage medium should not be construed as storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals communicated through a wire, and / or other transmission media, as the term is used in this disclosure. As will be understood by those skilled in the art, data is typically moved during normal operation of a storage device at some irregular points in time, such as during access, defragmentation, or garbage collection, but the data is not transient while it is stored, and thus the storage device is not considered to be transient for this reason.
[0048] FIG. 1 is a block diagram showing a computing environment for the calculation of a graph function according to an embodiment of the present disclosure. Referring to FIG. 1, a computing environment 100 is shown that includes an example of an environment for the execution of at least a portion of the computer code involved in the execution of the method of the present invention, such as an improved graph function calculator 120B. In addition to block 120B, the computing environment 100 includes, for example, a computer 102, a wide area network (WAN) 104, an end user device (EUD) 106, a remote server 108, a public cloud 110, and a private cloud 112. In this embodiment, the computer 102 includes a processor set 114 (including a processing circuit 114A and a cache 114B), a communication fabric 116, a volatile memory 118, a persistent storage 120 (including an operating system 120A and block 120B as specified above), a peripheral device set 122 (including a user interface (UI) device set 122A, a storage 122B, and an Internet of Things (IoT) sensor set 122C), and a network module 124. The remote server 108 includes a remote database 108A. The public cloud 110 includes a gateway 110A, a cloud orchestration module 110B, a host physical machine set 110C, a virtual machine set 110D, and a container set 110E.
[0049] Computer 102 may take the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch, or any other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device that is currently known or that may be developed in the future and that is capable of executing a program, accessing a network, or querying a database such as remote database 108A. As is well understood in the field of computer technology and depending on the technology, the execution of the computer implementation method may be distributed among multiple computers and / or between multiple locations. On the other hand, in the present presentation of computing environment 100, for the sake of keeping the presentation as simple as possible, the detailed discussion focuses on a single computer, specifically computer 102. Although computer 102 is not shown within the cloud in FIG. 1, it may be located within the cloud. On the other hand, computer 102 is not required to be present within the cloud except within any arbitrarily shown scope.
[0050] Processor set 114 includes one or more computer processors of any type that are currently known or that may be developed in the future. Processing circuitry 114A may be distributed among multiple packages, for example, multiple conditioned integrated circuit chips. Processing circuitry 114A may implement multiple processor threads and / or multiple processor cores. Cache 114B may be memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 114. Cache memory is typically organized into multiple levels depending on its relative proximity to processing circuitry 114A. Alternatively, some or all of cache 114B for processor set 114 may be located "off-chip". In some computing environments, processor set 114 may be designed to operate using qubits to perform quantum computing.
[0051] Computer-readable program instructions are typically loaded onto computer 102 and executed by a set of processors 114 of computer 102 through a series of operational steps, thereby enabling a computer-implemented method. As a result, the instructions thus executed will instantiate the method (collectively referred to as "the method of the present invention") specified in the flowchart and / or description of the computer-implemented method included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media such as cache 114B and other storage media discussed below. The program instructions and associated data are accessed by the set of processors 114 to control and direct the execution of the method of the present invention. In computing environment 100, at least some of the instructions for executing the method of the present invention may be stored in block 120B within persistent storage 120.
[0052] Communication fabric 116 is a signal conduction path that enables various components of computer 102 to communicate with each other. Typically, this fabric is created with switches and conductive paths such as buses, bridges, physical input / output ports, and switches and conductive paths that make up similar components. Other types of signal communication paths such as optical fiber communication paths and / or wireless communication paths may be used.
[0053] Volatile memory 118 is any type of volatile memory that is currently known or will be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 118 is characterized by random access, but this is not essential unless affirmatively indicated. In computer 102, volatile memory 118 is located within a single package and exists inside computer 102. Alternatively or additionally, volatile memory 118 may be distributed across multiple packages and / or located externally to computer 102.
[0054] The persistent storage 120 is any form of non-volatile storage for a computer that is currently known or will be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is being supplied directly to the computer 102 and / or to the persistent storage 120. The persistent storage 120 can be read-only memory (ROM), but usually at least a portion of the persistent storage 120 enables writing, deleting, and rewriting of data. Some well-known forms of the persistent storage 120 include magnetic disks and solid-state storage devices. The operating system 120A can take several forms, such as various known proprietary operating systems or open-source Portable Operating System Interface (POSIX)-type operating systems that employ a kernel. The improved graph function calculator has the code included in block 120B, which typically includes at least a portion of the computer code involved in the execution of the method of the present invention.
[0055] The peripheral device set 122 includes a set of peripheral devices of the computer 102. Data communication connections between the peripheral devices of the computer 102 and other components can be implemented in various ways, such as Bluetooth (registered trademark) connections, Near-Field Communication (NFC) connections, connections via cables (such as universal serial bus (USB) type cables), insertion type connections (e.g., secure digital (SD) cards), connections through local area communication networks, and even connections through wide area networks such as the Internet. In various embodiments, the UI device set 122A may include components such as a display screen, speakers, microphones, wearable devices (such as Google glasses and smartwatches), keyboards, mice, printers, touch pads, game controllers, and haptic devices. The storage 122B is external storage such as an external hard drive or insertable storage such as an SD card. The storage 122B can be persistent and / or volatile. In some embodiments, the storage 122B can take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 102 is required to have a large amount of storage (e.g., when the computer 102 locally stores and manages a large-scale database), in this case, this storage can be provided by a peripheral storage device designed to store a very large amount of data, such as a storage area network (SAN) shared by a plurality of geographically dispersed computers. The IoT sensor set 122C is composed of sensors that can be used in Internet of Things applications. For example, one sensor can be a thermometer, and another sensor can be a motion detector.
[0056] The network module 124 is an assembly of computer software, hardware, and firmware that enables the computer 102 to communicate with other computers through the WAN 104. The network module 124 may include hardware such as a modem or a Wi-Fi (registered trademark) signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data via the Internet. In some embodiments, the network control function and the network transfer function of the network module 124 are executed on the same physical hardware device. In other embodiments (e.g., embodiments that utilize Software-Defined Networking (SDN)), the control function and the transfer function of the network module 124 are executed on physically separate devices such that the control function manages several different network hardware devices. The computer-readable program instructions for executing the method of the present invention can usually be downloaded to the computer 102 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 124.
[0057] The WAN 104 is any wide area network (e.g., the Internet) that can communicate computer data over non-local distances by any technology for communicating computer data that is currently known or will be developed in the future. In some embodiments, the WAN 104 can be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. The WAN 104 and / or the LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.
[0058] The end-user device (EUD) 106 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating the computer 102), and can take any of the forms discussed above in relation to the computer 102. The EUD 106 typically receives beneficial and useful data from the operation of the computer 102. For example, in a virtual case where the computer 102 is designed to provide recommendations to the end user, this recommendation will typically be communicated from the network module 124 of the computer 102, through the WAN 104, to the EUD 106. In this way, the EUD 106 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 106 can be a client device such as a thin client, a thick client, a mainframe computer, a desktop computer, and the like.
[0059] The remote server 108 is any computer system that provides at least some data and / or functions to the computer 102. The remote server 108 can be controlled and used by the same entity that operates the computer 102. The remote server 108 represents a machine that collects and stores data that is beneficial and useful for use by other computers such as the computer 102. For example, in a virtual case where the computer 102 is designed and programmed to provide recommendations based on past data, in that case, this past data can be provided from the remote database 108A of the remote server 108 to the computer 102.
[0060] The public cloud 110 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. The direct active management of the computing resources of the public cloud 110 is performed by the computer hardware and / or software of the cloud orchestration module 110B. The computing resources provided by the public cloud 110 are typically implemented by a virtual computing environment that runs on various computers that make up the host physical machine set 110C, which is the universe of physical computers within and / or available in the public cloud 110. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 110D and / or containers from the container set 110E. It is understood that these VCEs can be stored as images and transferred either as images or after instantiation of the VCE among and within various physical machine hosts. The cloud orchestration module 110B manages the transfer and storage of the images, deploys new instantiations of the VCE, and manages the active instantiation of the VCE deployment. The gateway 110A is an aggregate of computer software, hardware, and firmware that enables the public cloud 110 to communicate through the WAN 104.
[0061] Here, some further explanations of virtual computing environments (VCEs) are provided. A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system where the kernel enables the existence of multiple isolated instances of user space, called containers. These isolated instances of user space typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices allocated to the container, and this feature is known as containerization.
[0062] The private cloud 112 is similar to the public cloud 110, except that computing resources are only available for use by a single enterprise. The private cloud 112 is shown as being in communication with the WAN 104, but in other embodiments, the private cloud may be completely disconnected from the Internet and only accessible through a local / private network. A hybrid cloud is a composite of multiple different types of clouds (e.g., private cloud, community cloud, or public cloud types), and is often implemented by different vendors. Each of the multiple clouds remains a separate discrete entity, but the larger hybrid cloud architecture is coupled by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both the public cloud 110 and the private cloud 112 are part of a larger hybrid cloud.
[0063] In one or more embodiments of the present disclosure, the computer 102 is used to store an improved graph calculation function 120B that provides for the calculation of functions for graphs. The function includes, but is not limited to, measures of centrality defined previously, specifically measures of subgraph centrality.
[0064] FIG. 2 is a block diagram of an environment for the calculation of graph functions, according to an embodiment of the present disclosure. FIG. 2 is described in conjunction with the elements from FIG. 1. Referring to FIG. 2, a diagram of a network environment 200 is shown. The network environment 200 includes a system 202, a display screen 204, a server 206, and a user 208. The network environment 200 may further include the EUD 106 and the WAN 104 of FIG. 1. The system 202 may be an example of the computer 102 of FIG. 1 in one embodiment.
[0065] In an embodiment of the present disclosure, the system 202 includes an application installed on the computer 102 and is accessed by a user associated with the EUD 106.
[0066] In an embodiment of the present disclosure, the system 202 may include suitable logic, circuitry, interfaces, and / or code configured to calculate a graph function.
[0067] Examples of the system 202 may include, but are not limited to, a computing device, a virtual computing device, a mainframe machine, a server, a computer workstation, a smartphone, a cellular phone, a mobile phone, a gaming device, a consumer electronic (CE) device, and / or any other device having trace calculation capabilities.
[0068] The EUD 106 may include suitable logic, circuitry, interfaces, and / or code configured to provide an adjacency matrix to the system 202 as user input. In another embodiment, the EUD 106 may be configured to output a calculated graph function of the adjacency matrix onto the display screen 204. Specifically, the system 202 may control the display screen 204 of the EUD 106 to display the calculated graph function of the adjacency matrix on the display screen 204. The EUD 106 may be associated with a user 208 who may wish to calculate a graph function to generate a solution to a graph analysis problem. Examples of the EUD 106 may include, but are not limited to, a computing device, a mainframe machine, a server, a computer workstation, a smartphone, a cellular phone, a mobile phone, a gaming device, a consumer electronic (CE) device, and / or any other device having graph function calculation capabilities.
[0069] The display screen 204 may include suitable logic, circuitry, and interfaces configured to display the graph function of the graph analysis problem. In an embodiment, the display screen 204 may further display one or more user interface elements from which the user 208 may provide user input. In some embodiments, the display screen 204 may be an external display device associated with the EUD 106. The display screen 204 may be a touch screen that enables a user to provide user input via the display screen 204. The touch screen may be at least one of a resistive film touch screen, a capacitive touch screen, or a thermal touch screen. The display screen 204 may be realized through some known technologies, such as, but not limited to, at least one of a liquid crystal display (LCD), a light emitting diode (LED) display, a plasma display, or an organic LED (OLED) display technology, or other display devices. According to an embodiment, the display screen 204 may refer to a display screen of a head mounted device (HMD), a smart glass device, a see-through display, a projection display, an electrochromic display, or a transparent display.
[0070] Server 206 may include suitable logic, circuitry, and interfaces, and / or code configured to store graph 200 and its data. Server 206 may be further configured to store the results of calculations of graph functions such as projection matrices, approximate graphs, graph update data, and the like for graph 200. Server 206 may be implemented as a cloud server and may perform operations through web applications, cloud applications, HTTP requests, repository operations, file transfers, and the like. Other exemplary implementations of server 206 may include, but are not limited to, database servers, file servers, web servers, media servers, application servers, mainframe servers, or cloud computing servers.
[0071] In at least one embodiment, server 206 may be implemented as a plurality of distributed cloud-based resources through the use of some techniques well known to those skilled in the art. Those skilled in the art will understand that the scope of the present disclosure need not be limited to the implementation of server 206 and system 202 as two separate entities. In certain embodiments, without departing from the scope of the present disclosure, the functions of server 206 may be wholly or at least partially incorporated into system 202, or vice versa.
[0072] During operation, user 208 may wish to solve a graph analysis problem. A graph analysis problem may correspond to a category of computational problems that may involve studying a graph and drawing insights from the graph.
[0073] In an embodiment, a graph analysis problem may be associated with real-world problems because it may provide a powerful framework for modeling, understanding, and solving complex problems where relationships, connections, and interactions between entities may be involved.
[0074] In embodiments of the present disclosure, when the graph represents a dynamic graph, the graph analysis problem may include the analysis of the dynamic graph.
[0075] In embodiments of the present disclosure, the dynamic graph corresponds to a graph approximation of a social network. Therefore, the nodes of the dynamic graph represent users of the social network, the edges of the dynamic graph represent the connections between users of the social network, and furthermore, therefore, the most influential nodes of the dynamic graph represent users of the influencer user type in the social network. The edges of the dynamic graph represent the friendship between users of the social network. In this dynamic graph, adding a user to the network corresponds to adding a new node, while adding a friend implies inserting a new edge.
[0076] In embodiments of the present disclosure, the dynamic graph corresponds to a road network. The nodes of such a dynamic graph are destinations, and the connections between destinations via different routes are the edges of the dynamic graph. The most influential nodes of the dynamic graph may, in this case, represent the main hubs of the road network that experience high traffic volume and a large number of users.
[0077] In embodiments of the present disclosure, the dynamic graph corresponds to a collaboration network. The nodes of the dynamic graph may be authors, and the edges indicate papers (or collaborations between different authors). The most influential nodes of the dynamic graph may, in this case, represent the most influential authors.
[0078] In embodiments of the present disclosure, the dynamic graph corresponds to a protein network. The nodes of the dynamic graph represent proteins, and the edges of the dynamic graph indicate the interactions between proteins. The most influential nodes of the dynamic graph may, in this case, represent the most important proteins.
[0079] In one or more embodiments of the present disclosure, a dynamic graph evolves dynamically over different time steps. Since nodes are added to (or removed from) the dynamic graph at different time steps, the dynamic graph has different states at different time steps. For real-time applications using a dynamic graph, information regarding the state of the dynamic graph at different time steps forms graph update data. For most real-time applications, the graph update data needs to be accumulated accurately and efficiently. FIG. 3 shows such a graph that evolves dynamically over time.
[0080] FIG. 3 is a diagram of a graph 300 that evolves dynamically over time. In one or more embodiments of the present disclosure, graph 300 represents a dynamic graph (hereinafter, the terms “dynamic graph” and “graph 300” shall be used interchangeably within the scope of the present disclosure).
[0081] As shown in FIG. 3, graph 300 has a state 302 at time step t, which is represented by a sub-graph G t Furthermore, graph 300 has a state 304 at time step t + 1, which is represented by a sub-graph G t+1 The sub-graph G t+1 is the first sub-graph of graph 300, and the sub-graph G t is the second sub-graph of graph 300. Each of the first sub-graph and the second sub-graph has one or more nodes.
[0082] The sub-graph G t has four nodes, namely node a, node b, node c, and node d. The sub-graph G t+1 has seven nodes, namely the four previous nodes, node a, node b, node c, and node d, and three additional nodes, node e, node f, and node g. The connections between the nodes represent edges. The sub-graph G t+1To obtain [it], three additional nodes, namely node e, node f, and node g, are added to the partial graph G at time step t + 1 t to it. Also, the partial graph G t⊆ G t+1 is [as follows]. The data of the three additional nodes is the updated data associated with the (three) additional nodes added to the second partial graph G to obtain the first partial graph G t+1 ; therefore, the second partial graph G t is a subset of the first partial graph G t . t+1
[0083] It can be understood by those skilled in the art that the number of nodes shown in the graph of FIG. 3 is limited for the sake of simplicity. However, such a graph shown in FIG. 3 can have any number of nodes in equivalent applications without departing from the scope of the present disclosure.
[0084] In an embodiment of the present disclosure, the graph 300 is stored in the computer 102, for example, in the persistent storage 120 of the computer 102.
[0085] In another embodiment of the present disclosure, the graph 300 is stored in the remote server 108, for example, in the remote database 108A, and is accessed from the computer 102.
[0086] In an embodiment of the present disclosure, the computer 102 accesses applications such as social networking sites, navigation and mapping services, bio-compound analysis applications, business analysis applications, analysis platforms hosting analysis services, enterprise resource management applications, applications related to the management of data in data centers, financial applications, and the like. The applications store the data required to implement different functions of the applications in the form of the graph 300, specifically, in the form of a dynamic graph.
[0087] In an embodiment of the present disclosure, graph 300 is used for analysis, and thus, the data of graph 300 is managed using a schema and a data mapping algorithm associated with the graph platform. The graph platform can be implemented by, for example, either computer 102 or remote server 108.
[0088] In an embodiment of the present disclosure, the graph platform is a relational graph platform, and the data within the relational graph platform is processed using data mining algorithms.
[0089] In an embodiment of the present disclosure, graph 300 is accessed to calculate the functions of graph 300 for system 200, which may be either an application (an "app") or a product embodied in the form of either software code or dedicated hardware, thereby enabling different types of analysis operations on graph 300. Some of these analysis operations are those described above. The analysis is performed by executing the improved graph function calculator 120B described previously with respect to FIG. 1.
[0090] In one embodiment of the present disclosure, each of the subgraphs of graph 300, namely, the second subgraph G t and the first subgraph G t+1is represented by an adjacency matrix. The adjacency matrix can be a square matrix that can represent the connections or relationships between nodes (or vertices) in a graph. Each entry in the adjacency matrix may correspond to a potential edge between two nodes, and its value indicates whether a connection exists between those nodes. Typically, in a graph, the adjacency matrix may be symmetric, having a "1" in entry (i, j) if nodes i and j are connected and a "0" if they are not. Thus, the adjacency matrix can provide a structured way to encode the connectivity of a graph and can be a fundamental tool for performing one or more graph-related calculations and algorithms, which is indispensable in various graph analysis problems from social network analysis to transportation route optimization and beyond. In an embodiment of the present disclosure, graph 300 is an undirected graph of nodes or vertices n ∈ N associated with an n×n adjacency matrix A, and the subgraph centrality of the i-th graph node is the matrix exponential e A equal to the magnitude of the i-th diagonal entry of.
[0091] In one embodiment of the present disclosure, system 202 is configured to obtain an adjacency matrix A t+1 corresponding to a first subgraph G of graph 300. Further, system 300 is configured to calculate a function of adjacency matrix A t+1 for a second subgraph of graph G t based on the previously stored partial spectral factorization data of adjacency matrix A t and updated data associated with additional nodes added to the second subgraph G t+1 to obtain the first subgraph G t . t+1
[0092] In an embodiment of the present disclosure, the adjacency matrix A t+1 of the first subgraph G t+1 is associated with the state of graph 300 at the current time instance t + 1, and the previously stored adjacency matrix A t of the second subgraph G tis associated with the state of graph 300 at the previous time instance t, and thus the previous time instance is before the current time instance. The second subgraph is, therefore, a subset of the first subgraph, i.e., G t ⊆G t+1 is.
[0093] In an embodiment of the present disclosure, the improved graph function calculator 120B of FIG. 1 calculates a function of the adjacency matrix A t+1 of the subgraph G t+1 , which is further used by the improved graph function calculator 120B to calculate subgraph centrality data for graph 300. Further, the improved graph function calculator 120B is configured to determine a top N set of nodes of graph 300 based on the calculated subgraph centrality, and thus the top N set of nodes are the nodes of graph 300 having indices associated with the highest subgraph centrality data. These top N nodes are then used to identify the most influential nodes of graph 300 as part of a graph analysis problem.
[0094] In an embodiment of the present disclosure, graph 300, which may be referred to as graph G, is stored external to the main system memory, such as external to the persistent storage 120 of computer 102, and is analyzed over several time steps, and thus the partial graph G t' processed up to time step "t" is an induced subgraph of the partial graph G t+1 processed up to time step "t + 1".
[0095] A ∈ R nXn is shown as representing the adjacency matrix associated with the undirected graph G := {V, ε}, where V represents the set of vertices of G n , and ε ⊆ {(i, j)|(i, j) ∈ V 2i ≠ j} denotes the corresponding set of edges in the graph. The subgraph centrality of a vertex (or, equivalently, a node) Φ ∈ V measures the number of closed walks that start in and end in Φ, while each closed walk is weighted inversely proportional to its length. j Φ th The diagonal entries indicate the number of closed walks of length j that start at vertex Φ and end at vertex Φ, and the subgraph centrality of vertex Φ is: SC(Φ) = Σ j=0 A j ΦΦ / j!=e A ΦΦ It is expressed as:
[0096] Considering all n vertices, the subgraph centrality of a vertex in G is given as a vector of length n:
number
[0097] Equation 1 requires the computation of all n eigenpairs of A, and is therefore impractical for anything but small-scale graphs. In one or more embodiments of the present disclosure, the application of subgraph centrality generally involves the computation of λ i We only need to find the few most influential vertices of each rank-1 term x i x T i The contribution of is the scalar e λi , which means that the contribution of each eigenpair is weighted by λ i This means that it decreases rapidly compared to the algebraic value of
[0098] Let k∈N satisfy 1≦k≦n, and let Λ k =diag(λ 1 ,...,λ k ) and X k =[x1 、...、x k is defined as follows. Next, the k subgraph centralities of graph G are
Number
[0099] A sufficient value of k for determining the top N centrality scores can be determined on-the-fly by setting k = 1 and increasing the value until the top N recommendations achieved by two consecutive k subgraph centrality approximations are equal up to machine epsilon.
[0100] In some embodiments, the top N recommendations include identifying the
Number
[0101] FIG. 4 is a flowchart of a method showing exemplary operations for graph function calculation according to an embodiment of the present disclosure. FIG. 4 is described in conjunction with the elements from FIGS. 1, 2, and 3. Referring to FIG. 4, a flowchart of a method 400 showing exemplary operations from 402 to 406 is shown as described herein. The exemplary operations shown in method 400 may be initiated at 402 and may be executed by any computing system, apparatus, or device, such as computer 102 of FIG. 1 or system 202 of FIG. 2. Although shown in discrete blocks, the exemplary operations associated with one or more blocks of method 400 may be divided into additional blocks, combined into fewer blocks, or removed, depending on a particular implementation.
[0102] At 402, method 400 for graph function calculation is initiated. At 404, an adjacency matrix for a first sub-graph of a graph is obtained. For example, an adjacency matrix A t+1 corresponding to a first sub-graph G t+1 is obtained. In one embodiment, the adjacency matrix A t+1 is obtained from out-of-core storage, for example, from remote server 108. In one embodiment, the adjacency matrix A t+1 is calculated by processor set 114 of computer 102.
[0103] At 406, a function of the adjacency matrix A t+1 , also referred to as a graph function, is calculated. The calculation of the graph function is performed based on previously stored partial spectral factorization data of the adjacency matrix for a second sub-graph of the graph and updated data associated with nodes added to the second sub-graph to obtain the first sub-graph. For example, referring to FIG. 3, the graph function is calculated based on the previously stored adjacency matrix A t of a second sub-graph G t , and updated data for nodes e, f, and g, which are additional nodes added to the second sub-graph G t+1 to obtain the first sub-graph G t .
[0104] In an embodiment of the present disclosure, the graph function is the k sub - graph centralities SC k (G) defined in Equation (2). The nodes of graph G are [Number] divided into k non - overlapping subsets V 1 ,..., V f , so that V 1 ∪... ∪ V f = V. Similarly, the set of edges of G is divided into f non - overlapping sets ε 1 ,..., ε f , so that ε 1 ∪... ∪ ε f = ε. Starting from the initial edge set ε 1 ={(i, j)|(i, j) ∈ V 1 2 , i ≠ j}, ε q ={(i, j)|i ∈ V g && j ∈ V q ∀g ≤ q} is defined by Equation (3).
[0105] According to Equation (3), the edge set ε q includes all the edges of G for G when a) both endpoints are in V q , or b) one endpoint is in V q and the other endpoint is in V g , where g < q.
[0106] With the above division, it is possible to access G in a continuous order consisting of f time steps, where during the time step "t + 1", equivalently by the system 202 or by the improved graph function calculator 120B, information regarding the sets V t+1 and ε t+1 is accessed.
[0107] In some embodiments of the present disclosure, the graph information accessed before the current time step is overwritten by low-rank partial spectral factorization data. The partial spectral factorization data is based on a Rayleigh-Ritz projection calculation and is determined based on the step of calculating a projection matrix for the adjacency matrix A t+1 of the first partial graph G t+1 . The calculation of the Rayleigh-Ritz projection is further described with reference to FIG. 5.
[0108] FIG. 5 is a diagram showing the calculation of projection data for calculating a graph function by an improved graph function calculator 120B according to an embodiment of the present disclosure.
[0109] Referring to FIG. 5, G t is an induced subgraph of G processed up to time step "t", where the matrix pair (X t,k , Λ t,k ) corresponds to the k leading eigenpairs of the adjacency matrix A t . Also, G t+1 is an induced subgraph of G processed up to time step "t + 1". To calculate the graph function at time step t + 1, the improved graph function calculator 120B is first configured to obtain the adjacency matrix A t+1 of the first partial graph G t+1 . The obtained adjacency matrix A t+1 of the first partial graph G t+1 is a surrogate adjacency matrix. A matrix pair (X t+1,k , Λ t+1,k ) is calculated. The matrix pair (X t+1,k , Λ t+1,k ) approximates the k leading eigenpairs of the surrogate adjacency matrix A t+1 .
[0110]
Equation
[0111] As shown in FIG. 5, W t+1 , W T t+1 and C t+1obtains the first partial graph G t+1 to form updated data associated with the nodes added to the second partial graph G t .
[0112] [Number] and s t+1 = |V t+1 | As an initial condition, A 1 = X 1,k Λ 1,k X T 1,K .
[0113] The matrix W t+1 contains data corresponding to the connections between the nodes n t of the second partial graph G t (node a, node b, node c), and the set V t+1 of nodes S t added to the second partial graph G t+1 to obtain the first partial graph G t+1 (node e, node f, node g). If the i-th node of G t is connected to the j-th node of the set V t+1 , the (i, j) entry of the matrix W T t+1 is non-zero.
[0114] The matrix C t+1 contains data corresponding to the connections between the nodes of the set V t+1 . The edge connections in the graph G t+1 are formed by the set ε t+1 = {(e, f), (f, g), (b, e), (c, g)}.
[0115] Next, for example, using the method shown in FIG. 6, the surrogate adjacency matrix A t+1 is used to calculate the graph function.
[0116] FIG. 6 is a flowchart of a method 600 for calculating a graph function according to an embodiment of the present disclosure. FIG. 6 is described in conjunction with the elements from FIGS. 1, 2, 3, 4, and 5. Referring to FIG. 6, a flowchart of a method 600 showing exemplary operations from 602 to 612 is shown as described herein. The exemplary operations shown in method 600 may begin at 602 and may be executed by any computing system, device, or apparatus, such as computer 102 of FIG. 1 or system 202 of FIG. 2. Although shown in discrete blocks, the exemplary operations associated with one or more blocks of method 600 may be divided into additional blocks, combined into fewer blocks, or removed, depending on a particular implementation.
[0117] At 602, updated data associated with a node added to the first sub-graph is received. For example, referring to FIG. 5, at time step (equally also referred to as a time instance) t+1, three additional nodes, node e, node f, and node g, are added to the second sub-graph G t This addition of these additional nodes leads to the formation of the first sub-graph G t+1 at time step t+1. Also, as discussed in conjunction with FIG. 5, the updated data includes matrices W t+1 ,W T t+1 and C t+1 .
[0118] At 604, a projection matrix for the adjacency matrix is calculated based on a Rayleigh-Ritz projection calculation. The updated data is used to calculate a projection matrix for the adjacency matrix at time instance t+1 based on the Rayleigh-Ritz projection calculation.
[0119] At time instance t+1, the first sub-graph G t+1 has an adjacency matrix A t+1 . To calculate A t+1 , a projection matrix Z t+1 is calculated.
[0120] The problem of k partial graph centralities is essentially equivalent to computing the k algebraic dominant eigenpairs of an adjacency matrix such as A t+1 . The sparse eigenvalue solver computes part of the eigenpairs by applying the Rayleigh–Ritz technique to a subspace Z ∈ R t+1 that conceptually encapsulates the invariant subspace associated with the desired eigenvalues of matrix A n . In the absence of rounding errors, the eigenpairs of matrix A t+1 can be recovered as a subset of the Ritz pairs of matrix Z T AZ, where matrix Z represents an orthonormal basis of Z. Since A t+1 is sparse, a standard approach to forming the projection matrix Z is to employ the Krylov subspace method and a certain integer T scaled such that υ
Number
Number
Number
[0121] When the specified termination condition is reached and there is no further available updated data, at 612, the graph function is calculated.
[0122] In one embodiment of the present disclosure, the graph function is a subgraph centrality function. The calculation of the subgraph centrality function is shown in FIG. 7.
[0123] Figure 7 is a flowchart of a method 700 for calculating a graph function according to an embodiment of the present disclosure. Figure 7 is described in conjunction with elements from Figures 1, 2, 3, 4, 5, and 6. Referring to Figure 7, a flowchart of method 700 showing exemplary operations from 702 to 704 is shown as described herein. The exemplary operations shown in method 700 may begin at 702 and may be executed by any computing system, device, or apparatus, for example, by computer 102 of Figure 1 or system 202 of Figure 2. Although shown as discrete operations, the exemplary operations of method 700 may be divided into additional sub-operations, combined into fewer operations, and / or deleted, depending on the particular implementation.
[0124] In one embodiment of the present disclosure, the graph function calculated in method 600 is used to calculate the subgraph centrality for graph G. Accordingly, at 702, the subgraph centrality for graph G is calculated based on the calculated graph function. For example, the k principal Lit pairs determined at the time of reaching the end condition are used to calculate the subgraph centrality of graph G at the time instance when the end condition is reached. For example, at time instance t+1, the i-th principal eigenpair (λt+1,i, xt+1,i) of A t+1 is approximated by the i-th principal Lit pair (τt+1,i, Zt+1rt+1,i).
[0125] In an embodiment of the present disclosure, the Rayleigh-Ritz eigenvalue problem is solved by a dense eigenvalue solver. Accordingly, the computational complexity of this task is cubic with respect to the number of columns of the projection matrix Z t+1 . In one embodiment of the present disclosure, the step of calculating the projection matrix Zt+1 for the adjacency matrix A t+1 at time instance t+1 has the step of calculating the k principal Lit pairs for the projection matrix such that they are the same as the k principal eigenpairs of the adjacency matrix. Ran(Z t+1 ) is the matrix Z T t+1 A t+1 Z t+1k dominant Ritz pairs of t+1 are set in such a way that they become the exact k dominant eigenpairs of matrix A t+1 . In exact arithmetic, by performing the Rayleigh–Ritz projection using matrix Z t+1 such that Ran(Z t+1 ) ≡ Ran(A t+1 ), the exact eigenvectors associated with all non-zero eigenvalues of matrix A t+1 are returned. Thus, the projection matrix Z
[0126]
Number
[0127] In one embodiment of the present disclosure, the step of calculating the projection matrix for the adjacency matrix has the step of calculating a matrix formed by calculating the orthonormal basis components for the k dominant eigenpairs of the adjacency matrix.
[0128]
Number
Number
[0129] The projection matrix thus forms an orthonormal basis for the column space of matrix A t+1 .
[0130] In one embodiment of the present disclosure, matrix Q t+1 is replaced by a low-rank approximation of matrix F t+1 . The target rank is
Number
Number
[0131] In one embodiment of the present disclosure, the step of calculating the projection matrix for the adjacency matrix includes the step of calculating k principal eigenpairs of the adjacency matrix based on the Schur components. In this case, the projection matrix Z t+1 is calculated as follows
Number
Number
Number
Number
[0132] In one embodiment of the present disclosure, the step of calculating the projection matrix for the adjacency matrix includes the step of approximating k principal eigenpairs of the adjacency matrix using k dominant eigenvectors of the symmetric matrix. Therefore, the projection matrix Z t+1 is calculated as follows
Number
[0133] After calculating the projection matrix Z t+1 in any of the manners described above, A t+1Using the k leading eigenpairs, k subgraph centralities for the graph G are computed as previously explained in Equation (2). Next, at 704, using the k subgraph centralities computed for the graph G, the top N set of nodes of the graph G is determined. The top N set of nodes are the nodes of the graph G having indices associated with the highest subgraph centrality data.
[0134] In one embodiment of the present disclosure, these top N nodes are then used to identify the most influential nodes of the graph G.
[0135] In an embodiment, as explained above, each time the time step t is updated, an updated adjacency matrix is computed. Further, the previously stored adjacency matrix corresponding to a previous time instance t, such as A t' is replaced with an adjacency matrix corresponding to a subsequent time instance t + 1, such as A t+1 Thus, each adjacency matrix corresponding to a subgraph is accessed only once, leading to computational efficiency and storage savings. Also, as explained above, the approximation of A using projection matrix calculations has, in some embodiments, the complexity of a linear calculation and provides a very efficient algorithm for graph updates and subsequent determination of the most influential nodes of the graph. t+1
[0136] FIG. 8 is a diagram illustrating an exemplary algorithm 800 for identifying the most influential nodes of a dynamic graph according to an embodiment of the present disclosure.
[0137] At 802, the algorithm 800 begins with the matrix A 1 and at 804, k leading eigenpairs of A 1 are computed.
[0138] The matrix A 1 then processes the additional nodes (and corresponding edges) added to A 2 to obtain A 1 2 is extended. This matrix is now, at 804, replaced by the action of (X 1,k , Λ 1,k X T 1,k ), so there is no longer a need to hold A 1 in the memory of system 202. This matrix is held in product form and is never explicitly formed. At 806, the k dominant eigenpairs of A 2 are approximated by updating (X 1,k , Λ 1,k ) to (X 2,k , Λ 2,k ) by means of the Rayleigh–Ritz projection calculation described above. This procedure is repeated until all nodes and edges of graph G have been processed.
[0139] In an embodiment, the k dominant eigenpairs of the entire adjacency matrix A are approximated in this process. Next, at 808, these k eigenpairs are used to approximate the graph function SC k (G) as in Equation (2), and at 810, the indices associated with its N largest entries are returned.
[0140] For each time instance, algorithm 800, as discussed above, iteratively updates t to t + 1 and, using the projection matrix Z t+1 , updates the k dominant eigenpairs of matrix A t to the k dominant eigenpairs of matrix A t+1 and computes its value. A method for performing such an update is shown in FIG. 9.
[0141] Figure 9 is a flowchart of a method 900 for updating a dynamic graph according to an embodiment of the present disclosure. Figure 9 is described in conjunction with elements from the previous figures. Referring to Figure 9, a flowchart of method 900 is shown, depicting exemplary operations from 902 to 912 as described herein. The exemplary operations shown in method 900 may begin at 902 and may be executed by any computing system, apparatus, or device, such as computer 102 of FIG. 1 or system 202 of FIG. 2. Although shown in discrete blocks, the exemplary operations associated with one or more blocks of method 900 may be divided into additional blocks, combined into fewer blocks, and / or deleted, depending on a particular implementation.
[0142] In one embodiment of the present disclosure, at 902, an adjacency matrix for a first sub-graph of the dynamic graph is obtained. For example, from the memory of system 202, etc., an adjacency matrix A t+1 for graph G t+1 is obtained. At 904, a set of k leading eigenpairs for the previously stored adjacency matrix for a second sub-graph is obtained. For example, the previously stored adjacency matrix A t has a set of leading eigenpairs X t,k , Λ t,k stored in the memory of system 202. At 906, a projection matrix for the adjacency matrix at time instance t+1 is approximated. For example, the projection matrix Z t+1 is approximated using any of equations (4) to (7) previously described for matrix A t+1 .
[0143] At 908, k leading eigenpairs for matrix A t+1 are approximated. For example, k leading eigenpairs of Z T t+1 A t+1 Z t+1 are calculated. The k leading Ritz pairs are the eigenpairs of the matrix (X t+1,k , Λ t+1,k) is further formed using. At 910, k major Ritz pairs are used to update the dynamic graph, and at 912, the updated dynamic graph is then stored in the memory of system 202.
[0144] In the embodiments described above, since the graph data is approximated in the form of a matrix product constructed using Rayleigh-Ritz projection, the storage of data about the dynamic graph (also referred to as dynamic graph data in the same sense) is efficient and bandwidth-saving. Further, the dynamic graph data is extended using rank-k partial spectral factorization, which is further used to compute graph functions such as subgraph centrality and to identify the most influential nodes of the dynamic graph required in many real-world applications such as social networks, road networks, biochemical compound analysis, and the like.
[0145] In one or more embodiments of the present disclosure, the pass for subgraphs and their adjacency matrices is executed only once, and thus is attractive for large graphs that are either not compatible with system memory or become dynamically available.
[0146] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best explain the principles of the embodiments, the practical application, or the technical improvements made to the technology found in the marketplace, or to enable other skilled artisans to understand the embodiments disclosed herein.
Claims
1. obtaining an adjacency matrix corresponding to a first subgraph of the graph; and calculating a function of the adjacency matrix for a second subgraph of the graph based on the previously stored partial spectral factorization data of the adjacency matrix and update data associated with nodes added to the second subgraph to obtain the first subgraph, where the second subgraph is a subset of the first subgraph. A computer-implemented method comprising:
2. 2. The computer-implemented method of claim 1, wherein the adjacency matrix of the first subgraph is associated with a state of the graph at a current time instance and the previously stored adjacency matrix of the second subgraph is associated with a state of the graph at a prior time instance, such that the prior time instance is earlier than the current time instance.
3. calculating subgraph centrality data for the graph based on the calculated function; and determining a top-N set of nodes of the graph based on the calculated subgraph centralities, where the top-N set of nodes are the nodes of the graph having indices associated with highest subgraph centrality data. The computer-implemented method of claim 1 , further comprising:
4. 4. The computer-implemented method of claim 3, wherein the graph is a dynamic graph and the top-N set of nodes are the most influential nodes of the dynamic graph.
5. 5. The computer-implemented method of claim 4, wherein the dynamic graph corresponds to a graph approximation of a social network, such that nodes of the dynamic graph represent users of the social network, edges of the dynamic graph represent connections between the users of the social network, and further such that the most influential nodes of the dynamic graph represent the users of an influencer user type in the social network.
6. 2. The computer-implemented method of claim 1, wherein the partial spectral factorization data is determined based on computing a projection matrix for the adjacency matrix of the first subgraph based on a Rayleigh-Ritz projection calculation.
7. a) receiving the update data associated with the node added to the second subgraph at time instance t+1; b) calculating the projection matrix for the adjacency matrix at the time instance t+1 based on the Rayleigh-Ritz projection calculation; c) computing k principal eigenpairs for the projection matrix at time instance t+1; d) determining k principal Ritz pairs for the k principal eigenpairs of the projection matrix; e) updating t+1; and f) repeating steps a through e until a specified termination condition is reached. The computer-implemented method of claim 6 , further comprising:
8. The computer-implemented method of claim 7 , wherein the specified termination condition is based on the availability of updated data for the graph.
9. 9. The computer-implemented method of claim 8, further comprising: computing a subgraph centrality for the graph based on determining the k principal Ritz pairs upon reaching the termination condition.
10. 8. The computer-implemented method of claim 7, wherein computing the projection matrix for the adjacency matrix comprises computing the k principal Ritz pairs for the projection matrix to be the same as the k principal eigenpairs of the adjacency matrix at the time instance t+1.
11. The computer-implemented method of claim 7 , wherein computing the projection matrix for the adjacency matrix comprises computing the k principal eigenpairs of the adjacency matrix based on Schur components.
12. 8. The computer-implemented method of claim 7, wherein computing the projection matrix for the adjacency matrix comprises computing a matrix formed by computing orthonormal basis elements for the k principal eigenpairs of the adjacency matrix.
13. 8. The computer-implemented method of claim 7, wherein computing the projection matrix for the adjacency matrix comprises approximating the k principal eigenpairs of the adjacency matrix using k dominant eigenvectors of a symmetric matrix.
14. 8. The computer-implemented method of claim 7, further comprising replacing a previously stored adjacency matrix corresponding to a previous time instance t with the adjacency matrix corresponding to a subsequent time instance t+1.
15. 1. A computer-implemented method for updating a dynamic graph, comprising: obtaining an adjacency matrix for a first subgraph of the dynamic graph at time step t+1; accessing a projection matrix for the previously stored adjacency matrix for a second subgraph of the dynamic graph; approximating a projection matrix for the adjacency matrix at time step t+1 based on the accessed projection matrix for the previously stored adjacency matrix at time step t and update data for the dynamic graph received at time step t+1; approximating a principal eigenpair for the adjacency matrix at the time step t+1 based on the approximated projection matrix; updating the dynamic graph based on the approximated principal eigenpairs; and Storing the updated dynamic graph. A computer-implemented method comprising:
16. calculating subgraph centrality data for the stored updated dynamic graph; and determining the most influential nodes of the dynamic graph based on the calculated subgraph centrality data; The computer-implemented method of claim 15 further comprising:
17. 16. The computer-implemented method of claim 15, wherein the updating of the dynamic graph based on the approximated principal eigenpairs is performed iteratively until a specified termination condition is met.
18. 20. The computer-implemented method of claim 17, wherein the specified exit condition comprises a determination of availability of updated data for the dynamic graph.
19. 13. A circuit comprising: Obtain an adjacency matrix corresponding to a first subgraph of the graph; and For a second subgraph of the graph, calculate a function of the adjacency matrix based on the previously stored partial spectral factorization data of the adjacency matrix and update data associated with nodes added to the second subgraph to obtain the first subgraph, where the second subgraph is a subset of the first subgraph. A circuit constructed as follows: A system comprising:
20. 20. The system of claim 19, wherein the adjacency matrix of the first subgraph is associated with a state of the graph at a current time instance and the previously stored adjacency matrix of the second subgraph is associated with a state of the graph at a prior time instance, such that the prior time instance is earlier than the current time instance.
21. The circuit comprises: Calculating subgraph centrality data for the graph based on the calculated function; and determining a top-N set of nodes of the graph based on the calculated subgraph centralities, where the top-N set of nodes are the nodes of the graph having indices associated with the highest subgraph centrality data; 20. The system of claim 19, further configured to:
22. 20. The system of claim 19, wherein the graph is a dynamic graph and the top-N set of nodes are the most influential nodes of the dynamic graph.
23. 20. The system of claim 19, wherein the partial spectral factorization data is determined based on a Rayleigh-Ritz projection calculation, the procedure being to calculate a projection matrix for the adjacency matrix of the first subgraph.
24. The circuit comprises: a) receiving the update data associated with the added node of the first subgraph at time instance t+1; b) calculating a projection matrix for said adjacency matrix at said time instance t+1 based on a Rayleigh-Ritz projection calculation; c) computing k principal eigenpairs for said projection matrix at said time instance t+1; d) determining k principal Ritz pairs for the k principal eigenpairs of the projection matrix; e) Update t+1; and f) Repeat steps a through e until the specified end condition is reached.
20. The system of claim 19, further configured to:
25. 1. A computer program for computing a function of an adjacency matrix, the computer program comprising: obtaining an adjacency matrix corresponding to a first subgraph of the graph; and calculating a function of the adjacency matrix for a second subgraph of the graph based on the previously stored partial spectral factorization data of the adjacency matrix and update data associated with nodes added to the second subgraph to obtain the first subgraph, where the second subgraph is a subset of the first subgraph. A computer program for executing the above.