Large-scale graph generation method and system based on divide-and-conquer

Through a partition-based method, large-scale graphs are divided into community graphs and deep generative models are used to generate connections within and between communities, which solves the shortcomings of the graph generation method in the existing technology in scalability and complex network feature capture, and realizes the generation and optimization of high-quality and large-scale graphs.

CN120123552APending Publication Date: 2025-06-10TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510031658.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing graph generation methods have shortcomings in scalability and capturing complex network characteristics, especially when generating large graphs, with reduced quality and lack of good scalability.

Method used

Using a large-scale graph generation method based on division and governance, undirected graph balance is divided into community graphs through the Metis algorithm. A deep generative model is used to combine spectral conditional variational autoencoder and graph autoencoder to generate connections within and between communities, and finally, the joint optimization of the loss function generated by the community generation and the bridge generation loss function is generated to generate high-quality large-scale graphs.

Benefits of technology

It realizes high-quality large-scale graph generation, improves the scalability of generated graphs, can effectively capture multiple characteristics in complex networks, meets the needs of large-scale graph generation, and optimizes memory and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123552A_ABST
    Figure CN120123552A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale graph generation method and system based on divide-and-conquer. The method comprises the steps of obtaining an undirected graph; performing balanced division on the undirected graph by using a Metis algorithm to obtain a community graph; inputting the community graph into a depth generative model for coding to obtain a generated graph; and jointly optimizing the depth generative model by using the community generation loss function and the bridging generation loss function. According to the method, a dividing and conquer graph generation framework BTGAE is introduced, the requirement of high-quality large-scale graph generation is met, and meanwhile, the memory and the calculation efficiency are optimized; a spectrum-guided conditional variation auto-encoder is adopted to learn and generate a community structure, and a global sparse bridging structure is processed in combination with a decomposition-based graph auto-encoder, so that a large graph with a real topological structure can be efficiently generated; according to the method, a joint optimization strategy is provided, so that information sharing between a community and a bridge generator in the training process becomes possible, and generation of a large graph is learned in an end-to-end manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph generation, and in particular, to a large-scale graph generation method and system based on divide and conquer. Background Art

[0002] Graph generation methods are rapidly becoming a key frontier field with a wide range of applications, covering many fields such as molecular design, network analysis, and social network modeling. This method can generate and simulate large-scale graph data such as global social networks and financial transaction graphs, which is crucial for data augmentation tasks and diverse application scenarios. In the field of social network analysis, graph generation methods can help researchers explore the paths of information dissemination, the evolution of group behavior, and the distribution of social influence. With the strengthening of data privacy regulations, the application of graph generation methods in privacy protection has received extensive attention. It provides an effective solution: by creating simulated data to replace real data, it not only protects sensitive information but also enables data analysis and utilization, meeting the growing privacy protection requirements. In practical applications such as power grid data centers and financial institutions, due to privacy protection needs, sensitive data often cannot be made public. To solve this problem, third-party agencies usually use simulated data to ensure privacy security when developing software or conducting stress tests. Therefore, graph generation methods are crucial in many fields.

[0003] Through a literature search of the existing technology, it is found that the current graph generation methods are mainly divided into traditional methods based on rule generation and generation methods based on deep learning. In the paper "A Scalable Generative Graph Model with Community Structure" in the top scientific computing journal SISC, TAMARA G. KOLDA et al. proposed the BTER model. This model can generate graph data with community structure by adjusting model parameters to match the degree distribution and clustering coefficient of real-world data, and it has good scalability. However, the BTER model may not be able to capture network characteristics more complex than community structure, such as hierarchical structure or other complex behaviors, which limits its application in more complex network analysis. In the top machine learning conference NeurIPS in 2024, the work "Exact Representation of Sparse Networks with Symmetric Nonnegative Embeddings" published by Sudhanshu Chanpuriya et al. proposed a graph model based on symmetric nonnegative embedding for representing sparse networks and performing node clustering. Through an extension of logical principal component analysis (LPCA), this model can accurately represent graphs with bounded treewidth and provides a new perspective to explain the connection probability between nodes. However, when extended to large graphs, the quality of the generated graphs will rapidly decline, and it does not have good scalability. Summary of the Invention

[0004] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this part, as well as in the abstract and title of the specification of this application, to avoid obscuring the purpose of this part, the abstract of the specification, and the title of the invention. However, such simplifications or omissions cannot be used to limit the scope of the present invention.

[0005] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides a divide-and-conquer-based large-scale graph generation method to solve the deficiencies of existing graph generation methods in terms of scalability and capturing complex network characteristics.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a divide-and-conquer-based large-scale graph generation method, including:

[0008] Obtain an undirected graph;

[0009] Use the Metis algorithm to perform a balanced partition on the undirected graph to obtain a community graph;

[0010] Input the community graph into a deep generative model for encoding to obtain a generated graph;

[0011] Jointly optimize the deep generative model using a community generation loss function and a bridging generation loss function.

[0012] As a preferred solution of the divide-and-conquer-based large-scale graph generation method of the present invention, where: inputting the community graph into a deep generative model for encoding to obtain a generated graph includes,

[0013] The deep generative model includes a first generation stage and a second generation stage;

[0014] The first generation stage is to use a spectral-based conditional variational autoencoder to generate the internal connections of each community in the community graph, and obtain the corresponding community graph and community set;

[0015] The second generation stage is to use a graph autoencoder to generate the connections between each community in the community graph, and obtain a community edge set;

[0016] Combine the community set and the community edge set to obtain a generated graph.

[0017] As a preferred solution of the divide-and-conquer-based large-scale graph generation method of the present invention, where: using a spectral-based conditional variational autoencoder to generate the internal connections of each community in the community graph includes,

[0018] Use a graph neural network to calculate the embedding of the community graph, expressed as:

[0019]

[0020] where C i is the corresponding community graph, H l is the l-th hidden layer of the conditional variational autoencoder, L enc is the number of layers of the encoder, A i is the adjacency matrix of C i , use the degree matrix D i as the feature matrix X i of C i , and use it to initialize the first hidden layer of the mean and standard deviation, GNN is the graph neural network, σ is the spectral conditional generation variance, and μ is the spectral conditional generation mean;

[0021] Use the concatenated graph embedding and the spectral conditional generation mean and variance, expressed as:

[0022]

[0023] where MLP represents a multi-layer perceptron, S i is the eigenvalue, Lenc is the number of layers of the encoder, σ is the spectral conditional generation variance, and μ is the spectral conditional generation mean;

[0024] Use L dec transposed convolutional layers to model the conditional distribution p θ (C|z), expressed as:

[0025] H (l) = ReLU(D T H (l-1) W (l-1) + b (l-1) )

[0026] for l = 1,…, L dec H (0) = [z ∥ S]

[0027] where H l is the l-th hidden layer of the decoder of the conditional variational autoencoder, D T is the transpose of the degree matrix D, z ∼ N(μ, σ), W and b are learnable parameters, ReLU(·) is the activation function, and S is the spectral set of all communities.

[0028] As a preferred embodiment of the divide-and-conquer-based large-scale graph generation method of the present invention, wherein: generating the connections between communities in the community graph by using a graph autoencoder includes,

[0029] calculating the convex combination of the community graph and the generated corresponding community graph, expressed as:

[0030]

[0031] where C i is the community graph, λ f and λ 0 are hyperparameters, is the corresponding community graph;

[0032] For each community graph, a specific label is assigned to facilitate matching the corresponding encoder and spectral conditions;

[0033] The graph autoencoder uses a two-layer graph convolutional neural network to learn the node embeddings in the latent space, as follows:

[0034]

[0035] where, N i represents the number of nodes in the i-th community, d represents the embedding size, is the adjacency matrix of the convex combination graph, is The degree matrix is used as the feature matrix.

[0036] As a preferred solution of the divide-and-conquer based large-scale graph generation method described in the present invention, wherein: combining the community set and the community edge set, the generated graph includes

[0037] Using a spectral-based conditional variational autoencoder to generate the community set. First, sample z ∼ N(0, I), and then use the spectral set as a condition to guide the generation of the community set, denoted as:

[0038]

[0039] Wherein, is the community set, S is the spectral set, is the conditional variational autoencoder, and z is a random variable sampled from a multivariate Gaussian distribution;

[0040] Using a graph autoencoder to generate the edges between communities to obtain the community edge set, denoted as:

[0041]

[0042] Wherein, is the community edge set, is the community set, is the graph autoencoder;

[0043] Taking the union of the community set and the community edge set to obtain the final generated graph.

[0044] As a preferred solution of the divide-and-conquer based large-scale graph generation method described in the present invention, wherein: the community generation loss function is expressed as:

[0045]

[0046] Wherein, k is the total number of communities, C i is the community graph, S i is the eigenvalue, z is the latent variable, D KL is the KL divergence, is the conditional variational autoencoder, is the encoder part of, and c is the first letter of community. The subscript with c represents a certain part of the community encoder.

[0047] As a preferred solution of the divide-and-conquer based large-scale graph generation method described in the present invention, wherein: the bridging generation loss function is expressed as:

[0048]

[0049] Wherein, k is the total number of communities, Bij and respectively represent the real and generated edges between community i and community j.

[0050] In a second aspect, the present invention provides a system for generating large-scale graphs based on divide-and-conquer, including:

[0051] An acquisition module, configured to acquire an undirected graph;

[0052] A partitioning module, configured to use the Metis algorithm to perform balanced partitioning on the undirected graph to obtain a community graph;

[0053] A generation module, configured to input the community graph into a deep generative model for encoding to obtain a generated graph;

[0054] An optimization module, configured to jointly optimize the deep generative model by using a community generation loss function and a bridging generation loss function.

[0055] In a third aspect, the present invention provides a computing device, including:

[0056] A memory and a processor;

[0057] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for generating large-scale graphs based on divide-and-conquer are implemented.

[0058] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the method for generating large-scale graphs based on divide-and-conquer are implemented.

[0059] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention introduces a divide-and-conquer graph generation framework BTGAE, which meets the requirements of high-quality large-scale graph generation, and at the same time optimizes memory and computing efficiency; specifically, since a large network is divided into dense community structures and edges between communities, a spectrum-guided conditional variational autoencoder is used to learn to generate community structures, and then a decomposition-based graph autoencoder is combined to process the globally sparse bridging structure. Therefore, the present invention can efficiently generate large graphs with realistic topological structures. In addition, due to using two-stage generation, the present invention proposes a joint optimization strategy, which makes it possible to share information between the community and bridging generators during the training process, and to learn the generation of large graphs end-to-end. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:

[0061] Figure 1 It is a schematic diagram of the overall process of the divide-and-conquer-based large-scale graph generation method according to an embodiment of the present invention;

[0062] Figure 2 It is a schematic diagram of the training phase process framework of the divide-and-conquer-based large-scale graph generation method according to an embodiment of the present invention;

[0063] Figure 3 It is a schematic diagram of the inference phase process of the divide-and-conquer-based large-scale graph generation method according to an embodiment of the present invention. Detailed implementation manners

[0064] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific implementation manners of the present invention in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0065] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0066] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that excludes other embodiments.

[0067] The present invention is described in detail in conjunction with the schematic diagrams. When describing the embodiments of the present invention in detail, for the convenience of explanation, the cross-sectional views showing the device structure will be enlarged locally out of the general proportion, and the schematic diagrams are only examples, which should not limit the protection scope of the present invention here. In addition, in actual production, three-dimensional spatial dimensions including length, width, and depth should be included.

[0068] Meanwhile, in the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper, lower, inner, and outer" is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first, second, or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0069] Unless otherwise clearly defined and limited in the present invention, the terms "installation, connection, and coupling" shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can also be a mechanical connection, an electrical connection, or a direct connection, and can also be indirectly connected through an intermediate medium, or it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0070] Embodiment 1

[0071] Referring to Figure 1 , an embodiment of the present invention provides a large-scale graph generation method based on divide and conquer, including:

[0072] S100: Obtain an undirected graph;

[0073] In the embodiment of the present application, an unweighted undirected graph G(V, E) is input, where V represents the set of nodes, E represents the set of edges, and

[0074] In the embodiment of the present application, the normalized Laplacian matrix of graph G is defined as where is the adjacency matrix, D n×n is the diagonal degree matrix of the graph. Since L norm is a symmetric positive semi-definite matrix, it can be diagonalized. L norm = UΛU T , where U = [u 0 , …, u n is an orthogonal matrix containing eigenvectors, and Λ = diag(λ 0 , …, λ n ) is a diagonal matrix composed of corresponding eigenvalues;

[0075] It should be noted that the objective of the present invention is to train a deep generative model g θ to synthesize a generated graph such that it is close enough to the original input graph G in terms of structural distribution.

[0076] S102: Use the Metis algorithm to perform balanced partitioning on the undirected graph to obtain a community graph;

[0077] Preferably, use the Metis algorithm to balance the input undirected graph G into communities and the connections between communities;

[0078] In the embodiments of the present application, using the Metis algorithm, while ensuring that the sizes of the generated partitions are as equal as possible, minimize the number of edges between different partitions, that is, minimize At the same time, it is necessary to satisfy the constraint:

[0079]

[0080] where V 1 and V 2 are two different partitions of V, w(u,v) is the weight of the edge connecting two nodes, and ∈ is a very small constant;

[0081] In the embodiments of the present application, the input graph G(V,E) is divided into k communities, where V i and E i respectively represent the node set and edge set corresponding to the community C i ;

[0082] In the embodiments of the present application, then consider the connections between communities. The edges between a "community pair" (C i , C j ) can be represented by a bipartite graph B ij =(V I , V j , ε(B ij )) where ε(B ij ) represents the set of edges connecting the nodes of the two communities C i and C j . Therefore, the edges between all communities can be represented by a set of bipartite graphs .

[0083] In the embodiments of the present application, the graph G can be divided into k communities and the connections between communities, that is

[0084] It should be noted that using the Metis algorithm to decompose in this way provides convenience for the subsequent algorithms to generate communities and connections separately.

[0085] S104: Input the community graph into the deep generative model for encoding to obtain a generated graph;

[0086] Preferably, the deep generative model includes a first generation stage and a second generation stage;

[0087] Preferably, in the first stage of generation, the internal connections of each community in the community graph are generated by using a spectrum-based conditional variational autoencoder, and the corresponding community graph and community set are obtained;

[0088] Preferably, in the second stage of generation, the connections between the communities in the community graph are generated by using a graph autoencoder, and the community edge set is obtained;

[0089] Preferably, the generated graph is obtained by combining the community set and the community edge set.

[0090] In the embodiment of the present application, the above step S104 further includes the following sub-steps A1-A3;

[0091] In A1, the internal connections of the community are generated by using a spectrum-based conditional variational autoencoder;

[0092] 1) Calculate the spectral condition. The Laplacian matrix is spectrally decomposed to obtain L norm = UΛU T , and the meanings of the relevant symbols have been explained at the beginning. Consider the first p non-zero eigenvalues S i ={λ 1 , λ 2 , …, λ p} as the condition for the generation of each community C i , where λ represents the eigenvalue. Therefore, the spectral set of all communities is expressed as

[0093] 2) Use a graph neural network to calculate the embeddings of the community graph C i , which is expressed as:

[0094]

[0095] Among them, H l is the l-th hidden layer of the decoder of the conditional variational autoencoder, L enc is the number of layers of the encoder, A i is the adjacency matrix of C i , and the degree matrix D i is used as the feature matrix X i of C i , and it is used to initialize the first hidden layer of the mean and standard deviation.

[0096] 3) Use the cascaded graph embedding and spectral condition to generate the mean and variance, as follows:

[0097]

[0098] Among them, MLP represents a multi-layer perceptron, which is a two-layer fully connected layer. Cascade the spectral condition and the graph embedding to utilize the spectral information to guide the graph generation process.

[0099] 4) To better capture the locally dense edge topology within each community, use L dec deconvolutional layers to model the conditional distribution p θ (C|z):

[0100] H (l) = ReLU(D T H (l-1) W (l-1) + b (l-1) )

[0101] for l = 1, …, L dec H (0) = [z||S]

[0102] where D T is the transpose of the degree matrix D, z ~ N(μ, σ), W and b are learnable parameters, and ReLU(·) is the activation function.

[0103] 5) Finally, to ensure that the output of the decoder is the probability of the connections between graph nodes, use the Sigmoid activation function:

[0104]

[0105] where is the adjacency matrix of the generated probability representation, and σ(·) represents the Sigmoid activation function. The probability can be converted to 0 or 1 by applying the argmax operation to the probability graph to obtain the target discrete format of the generated graph. In step 2, for each community graph C i an adjacency matrix corresponding to the community graph is generated. Next, step 3 generates the connections between community graphs.

[0106] In A2, use a graph autoencoder (GAE) to generate the connections between communities;

[0107] 1) Calculate the convex combination of the community graph C i and the generated corresponding community graph :

[0108]

[0109] where λ f and λ 0 are hyperparameters, with λ 0 > λ fTherefore, λ(t) is a linear scheduler that gradually decreases over time, which is to make GAE gradually learn more from the generated graph to learn encoding and decoding.

[0110] 2) For each community graph C i ∈ C, assign a specific label π i to match the corresponding encoder and the spectral condition S i .

[0111] 3) The encoder employs a two-layer graph convolutional neural network (GCN) to learn node embeddings in the latent space, as follows:

[0112]

[0113] where N i represents the number of nodes in the i-th community, d represents the embedding size, is the adjacency matrix of the convex combination graph, is 's degree matrix, serving as the feature matrix.

[0114] 4) In the decoding stage, use the latent variables of different communities to perform matrix inner product, as follows:

[0115]

[0116] where, is the reconstructed bipartite graph between the i-th community and the j-th community, that is, the generated connections between communities. σ(·) represents the Sigmoid function; thus, the connections between communities are generated.

[0117] In A3, use the spectral-based conditional variational autoencoder to generate the community set. First, sample z ∼ N(0, I), and then use the spectral set as a condition to guide the generation of the community set, denoted as:

[0118]

[0119] where, is the community set, S is the spectral set, is the conditional variational autoencoder, and z is a random variable sampled from a multivariate Gaussian distribution;

[0120] Use the graph autoencoder to generate the connections between communities to obtain the community connection set, denoted as:

[0121]

[0122] Among them, is the community edge set, is the community set, is the graph autoencoder,

[0123] The union of the community set and the community edge set is taken to obtain the final generated graph.

[0124] It should be noted that through the two-stage generation process of the above deep generative model, the present invention successfully encodes and transforms the community graph into a high-quality generated graph; specifically, in the first stage, the spectral-based conditional variational autoencoder accurately captures the complex connection structure within each community, and by combining graph neural network embedding and spectral conditions, effectively guides the generation of the graph within the community; in the second stage, the graph autoencoder further generates the connections between communities. By calculating the convex combination of the community graph and its generated graph, and using a specific encoder to learn node embeddings, the accurate generation of the edges between communities is finally achieved; by introducing the spectral-based conditional variational autoencoder to generate the community set and the community edge set, and taking their union, we obtain the final generated graph containing rich community structure and edge information. The present invention not only improves the quality and scalability of the generated graph, but also successfully captures various characteristics in complex networks, providing new ideas for the generation and analysis of graph data.

[0125] S108: Jointly optimize the deep generative model using the community generation loss function and the bridging generation loss function;

[0126] In the embodiments of the present application, in order to optimize the spectral-based conditional variational autoencoder in the first stage of the deep generative model, a community generation loss function is proposed, and the community generation loss function is expressed as:

[0127]

[0128] Among them, k is the total number of communities, C i is the community graph, S i is the eigenvalue, z is the latent variable, D KL is the KL divergence, is the conditional variational autoencoder, is the encoder part of, c is the first letter of community, and the subscript with c represents a certain part of the community encoder.

[0129] In the embodiments of the present application, in order to optimize the graph autoencoder in the second stage of the deep generative model, a bridging generation loss function is proposed, and the bridging generation loss function is expressed as:

[0130]

[0131] where k is the total number of communities, B ij and represent the real and generated edges between community i and community j, respectively.

[0132] Preferably, for joint optimization, the final generation loss is expressed as:

[0133]

[0134] where is the community generation loss function, is the bridging generation loss function.

[0135] It should be noted that through the joint optimization strategy of the above community generation loss function and bridging generation loss function, the present invention can comprehensively improve the performance of the deep generative model; the community generation loss function focuses on accurately capturing and reconstructing the internal connection structure of each community to ensure that the generated community graph is highly similar to the original community graph; at the same time, the bridging generation loss function focuses on optimizing the connections between communities, and by minimizing the difference between the real edges and the generated edges, it ensures that the generated edges between communities are both accurate and conform to the characteristics of the original network; this joint optimization method of the present invention not only enhances the model's understanding of the relationships within and between communities, but also promotes the improvement of the overall quality and consistency of the generated graph, making the generated graph data more truly reflect the structure and characteristics of the original complex network.

[0136] The above is a schematic solution of a divide-and-conquer based large-scale graph generation method in this embodiment. It should be noted that the technical solution of the divide-and-conquer based large-scale graph generation system belongs to the same concept as the technical solution of the above divide-and-conquer based large-scale graph generation method. For the details not described in detail in the technical solution of the divide-and-conquer based large-scale graph generation system in this embodiment, reference can be made to the description of the technical solution of the above divide-and-conquer based large-scale graph generation method.

[0137] The divide-and-conquer based large-scale graph generation system in this embodiment includes:

[0138] An acquisition module, configured to acquire an undirected graph;

[0139] A partitioning module, configured to perform balanced partitioning on the undirected graph by using the Metis algorithm to obtain a community graph;

[0140] A generation module, configured to input the community graph into a deep generative model for encoding to obtain a generated graph;

[0141] An optimization module, configured to jointly optimize the deep generative model by using the community generation loss function and the bridging generation loss function.

[0142] This embodiment also provides a computing device, which is applicable to the case of large-scale graph generation based on divide-and-conquer, including:

[0143] A memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for large-scale graph generation based on divide-and-conquer as proposed in the above embodiment.

[0144] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for large-scale graph generation based on divide-and-conquer as proposed in the above embodiment.

[0145] The storage medium proposed in this embodiment and the method for large-scale graph generation based on divide-and-conquer proposed in the above embodiment belong to the same inventive concept. For technical details not described in detail in this embodiment, reference can be made to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0146] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disc of a computer, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.

[0147] Embodiment 2

[0148] Refer to Figures 2-3 and Table 1-3. This is an embodiment of the present invention, which provides a method for large-scale graph generation based on divide-and-conquer. In order to verify its beneficial effects, the comparison results of two schemes are provided.

[0149] As Figure 2 shown, a divide-and-conquer graph generation method uses a dataset (YelpChi) of merchant review scores as its generation object, generates a graph dataset similar to the original dataset by using this method, and the generated graph achieves the data desensitization effect, which can better protect the individual privacy and security in the dataset.

[0150] First, the input undirected graph is evenly partitioned into communities and the connections between communities;

[0151] For the YelpChi dataset, the number of nodes in the network is 45,900, and the number of edges is 3,846,910. The network is partitioned using the Metis algorithm. Let k be 150, then the input graph G(V, E) is partitioned into 150 communities, with each community containing approximately 300 nodes. Where V i and E i represent the node set and edge set of the corresponding community C i respectively.

[0152] Considering the connections between communities, the edges between a pair of communities (C i , C j ) can be represented by a bipartite graph B ij =(V i , V j , ε(B ij )) where ε(B ij ) represents the set of edges between the two communities C i and C j . Therefore, all the edges between communities can be represented by the bipartite graph set .

[0153] Therefore, the undirected graph G can be partitioned into 150 communities and the edges between them, that is the network complexity is greatly reduced, and such decomposition provides convenience for the subsequent algorithms to generate communities and edges separately.

[0154] Secondly, the internal connections of the communities are generated using a spectral-based conditional variational autoencoder;

[0155] Denote the spectral-based conditional variational autoencoder network as SpCVAE(·). Its encoder model is a graph convolutional neural network (GCN), the decoder is a transposed convolutional layer, and the output layers of both the encoder and decoder are multi-layer perceptrons (MLP). The number of layers of the GCN is denoted as L enc , the number of layers of the transposed convolutional layer is denoted as L de , the number of layers of the MLP is L mlp , and the dimension size of the latent space is d hidden , where L enc =L dec =2, L mlp =2, d hidde =32.

[0156] Input the adjacency matrix A i of community C i and the degree matrix D i into the network, and sample to generate the adjacency matrix of the desensitized network of C i

[0157]

[0158] Thus, the edges within the community are generated, and then the edges between communities need to be generated.

[0159] Then, a graph autoencoder (GAE) is used to generate the connections between communities;

[0160] Denote the graph autoencoder for generating the edges between communities as brigeGAE(·). Its encoder is a graph convolutional neural network (GCN), and the decoder directly takes the inner product of the latent variables encoded by the encoder. Denote the number of layers of the GCN as L, and the dimension size of the encoded latent variables as d z , where L = 2, d z = 32.

[0161] Convex combination The coefficients in where T is the total number of training rounds, t represents the t-th round of training, λ 0 = 0.8, λ f = 0.2. Therefore, λ(t) will decrease as t increases, and the model will increasingly tend to learn from the community network generated by SpCVAE(·) Learn.

[0162] Input the adjacency matrix of the convex combination network and the degree matrix into the bridging graph autoencoder to generate the edges between communities:

[0163]

[0164] Therefore, 150 corresponding and 11,175 can be generated for 150 communities to represent the generated graph.

[0165] Finally, jointly optimize the quality of the graph generated by the model based on the evidence lower bound (ELBO) of the log-likelihood of the observed community set and the cross-entropy loss function, that is, the community generation loss and the bridging loss function;

[0166] Set the total number of training rounds T = 300. In each round, use the above loss function to optimize the model to continuously improve the generation ability of the model. When the model converges, stop training and start the inference stage.

[0167] As Figure 3 shown, the graph generation inference is divided into two stages, namely the first generation stage and the second generation stage described in the above method of the present invention.

[0168] To measure the quality of the model's graph generation network, the following statistics are used: Degree Maximum (DM), Normalized Triangle Count (NTC), Power Law Exponent (PLE), Gini Index (Gini), Relative Edge Distribution Entropy (REDE), Affinity Coefficient (AC), Global Clustering Coefficient (CC), Characteristic Path Length (CPL), and the number of overlapping edges (EO). A good graph generation model should produce a synthetic graph that matches the input graph in the above metrics, while having a relatively small edge overlap ratio. As shown in Table 1, we compared the differences between the graphs generated by BTGAE and other methods in the above statistics with the original dataset.

[0169] Table 1 Comparison of BTGAE and Other Methods

[0170]

[0171] As can be seen from Table 1, the bolded data indicates the closest to the statistics of the original dataset, and the underlined data comes second. By comparison, it can be found that the BTGAE method proposed in the present invention is closer to the original dataset in all metrics, and the number of overlapping edges (EO) is relatively small, indicating that the model does not rely on "memory" to restore the original graph, achieving a good desensitization effect.

[0172] In terms of algorithm complexity, the method of the present invention is more efficient, with shorter training time and smaller memory footprint.

[0173] Table 2 Training Time of Each Method for Generating Graphs with 50,000 Nodes

[0174]

[0175] As shown in Table 2, the training time of the method of the present invention for generating graphs with 50,000 nodes is less than 1 hour, which is 8.5 times the training speed of NetGAN! Thanks to the divide-and-conquer generation method, the method of the present invention also has an advantage over other methods in terms of memory footprint.

[0176] Table 3 Memory Footprint of Each Method during Training for Generating Graphs with 50,000 Nodes

[0177]

[0178] As shown in Table 3, the BTGAE proposed in the present invention is the only method that occupies less than 6GB of memory when generating graphs with 50,000 nodes, and the method of the present invention only occupies 11GB of memory when generating graphs with 300,000 nodes, while most deep learning-based methods are not capable of generating graphs with 300,000 nodes!

[0179] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A large-scale graph generation method based on divide-and-conquer, characterized in that: include: Get an undirected graph; The undirected graph is balancedly partitioned using the Metis algorithm to obtain a community graph; Inputting the community graph into a deep generative model for encoding to obtain a generated graph; The deep generative model is jointly optimized using a community generation loss function and a bridge generation loss function.

2. The divide-and-conquer large-scale graph generation method according to claim 1, characterized in that: Inputting the community graph into a deep generative model for encoding to obtain a generated graph includes: The deep generative model includes a first generation stage and a second generation stage; The first generation stage is to generate the internal connections of each community in the community graph using a spectrum-based conditional variational autoencoder to obtain the corresponding community graph and community set; The second generation stage is to use a graph autoencoder to generate connections between communities in the community graph to obtain a community edge set; The community set and the community edge set are combined to obtain a generated graph.

3. The divide-and-conquer large-scale graph generation method according to claim 1 or 2, characterized in that: The internal connections of each community in the community graph are generated using a spectrum-based conditional variational autoencoder, including: The embedding of the community graph is calculated using a graph neural network, expressed as: Among them, C i is the corresponding community graph, H l is the lth hidden layer of the conditional variational autoencoder, L enc is the number of encoder layers, A i It is C i The adjacency matrix of i As C i The feature matrix X i , and use it to initialize the first hidden layer of the mean and standard deviation, GNN is a graph neural network, σ is the spectral conditional generated variance, and μ is the spectral conditional generated mean; The mean and variance are generated using the concatenated graph embedding and spectral conditions, expressed as: Among them, MLP represents multi-layer perceptron, S i is the eigenvalue, L enc is the number of encoder layers, σ is the spectral conditional generated variance, and μ is the spectral conditional generated mean; Use L dec deconvolutional layers to model the conditional distribution p θ (C|z), expressed as: A (l) =ReLU(D T A (l-1) W (l-1) +b (l-1) ) forl=1,…,L dec H (0) =[z∥S] Among them, H l is the lth hidden layer of the decoder of the conditional variational autoencoder, D T is the transpose of the degree matrix D, z ∼ N (μ, σ), W and b are learnable parameters, ReLU (·) is the activation function, and S is the spectral set of all communities.

4. The divide-and-conquer large-scale graph generation method according to claim 3, characterized in that: The connections between the communities in the community graph generated by the graph autoencoder include: The convex combination of the computed community graph and the generated corresponding community graph is expressed as: Among them, C i For the community graph, λ f and λ0 are hyperparameters, is the corresponding community graph; For each community graph, a specific mark is given to match the corresponding encoder and spectral condition; The graph autoencoder uses a two-layer graph convolutional neural network to learn node embedding in the latent space, as shown below: in, N i represents the number of nodes in the ith community, d represents the embedding size, is the adjacency matrix of the convex composite graph, yes The degree matrix of is used as the feature matrix.

5. The divide-and-conquer large-scale graph generation method according to claim 4, characterized in that: Combining the community set and the community edge set, the generated graph includes: The community set is generated using the spectrum-based conditional variational autoencoder. First, z~N(0,I) is sampled, and then the spectrum set is used as a condition to guide the generation of the community set, which is recorded as: in, is the community set, S is the spectrum set, is a conditional variational autoencoder, z is a random variable sampled from a multivariate Gaussian distribution; Use the graph autoencoder to generate the edges between communities and obtain the community edge set, which is recorded as: in, To connect the community, Gather for the community, is a graph autoencoder; The community set and the community edge set are combined to obtain a final generated graph.

6. The method for generating a large-scale graph based on divide-and-conquer as claimed in claim 1, wherein the community The generation loss function is expressed as: Among them, k is the total number of communities, C i is the community graph, S i is the eigenvalue, z is the latent variable, D KL is the KL divergence, is a conditional variational autoencoder, for The encoder part of , c is the first letter of community community, and the part marked with c in the lower corner represents a part of the community encoder.

7. The divide-and-conquer large-scale graph generation method according to claim 6, characterized in that: The bridge generation loss function is expressed as: Among them, k is the total number of communities, B ij and They represent the real and generated edges between community i and community j respectively.

8. A system for large-scale graph generation based on divide-and-conquer, characterized in that: include, An acquisition module, used to obtain an undirected graph; A partitioning module, used for performing balanced partitioning of the undirected graph using the Metis algorithm to obtain a community graph; A generation module, used for inputting the community graph into a deep generative model for encoding to obtain a generated graph; An optimization module is used to jointly optimize the deep generative model using a community generation loss function and a bridge generation loss function.

9. An electronic device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the divide-and-conquer large-scale graph generation method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the divide-and-conquer large-scale graph generation method described in any one of claims 1 to 7.