Controllable community graph data generation method and system
Through the combination of knowledge distillation and a dual-channel decoupling generator, the node characteristics of multi-domain graph data are extracted and community graphs are generated, which solves the problems of coordinated optimization of controllability, structural fidelity and computing efficiency of multi-domain community graph generation in the existing technology, and achieves efficient and controllable community graph generation.
Patent Information
- Application Number
- CN202510577750.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The prior art is difficult to achieve coordinated optimization of controllability, structural fidelity and computing efficiency in the generation of multi-field community graphs, especially in the maintenance of community structure, computing resource consumption and attribute distribution control.
The student model is trained by knowledge distillation, node feature representations are extracted from multi-domain graph data, and a dual-channel decoupling generator is used to cooperate with the community-perceived graph convolution operator to generate node attributes and topological information respectively. Through joint loss optimization and multi-objective label vectors, controllable generation of graph structures and attributes is achieved.
It significantly improves the fidelity of the community structure, reduces the consumption of computing resources, and realizes the decoupling generation of node attributes and graph structures, making the generated community graph better than the existing technology in terms of community structure similarity and node attribute fidelity.
Smart Images

Figure CN120107408A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a graph data generation method, and in particular to a controllable community graph data generation method and system. Background Art
[0002] As a core tool for complex network modeling, graph generation technology has important value in social network analysis, biomolecular interaction simulation and other fields. However, the acquisition of real network data is often limited by privacy leakage risks (such as medical data sharing), incomplete observations (such as insufficient sensor network coverage) or policy constraints (such as restrictions on sensitive information in financial transaction networks), making synthetic data generation a key requirement for alternative solutions.
[0003] Existing generation methods can be divided into two categories: traditional models and deep learning models, but they all have significant defects in maintaining community structure, computational efficiency and controllability. Traditional random graph generation methods represented by the Erdős–Rényi (ER) model and the Barabási–Albert (BA) model can efficiently generate large-scale networks, but their connection mechanisms that rely on preset rules make it difficult to capture the heterogeneous community structure of real graphs. For example, in social network generation, such models cannot simulate the community aggregation phenomenon driven by user interests, resulting in the clustering coefficient of the generated graph being generally lower than the real data by more than 30%. Although the Stochastic Block Model (SBM) attempts to generate a hierarchical structure through community partitioning parameters, its community edge probability needs to be manually set, resulting in a community overlap error of up to 0.42 in the generated biological protein interaction network, which cannot adapt to dense communities and community overlaps.
[0004] Deep learning methods based on Generative Adversarial Network (GAN) and Variational Auto-Encoder (VAE), such as the Generative Adversarial Network for Learning Graph Topology (NetGAN) and Graph Variational Auto-Encoder (GraphVAE), have improved the generation quality through data-driven, but still face the following problems: (1) Insufficient fidelity of community structure: Existing models, such as variational graph auto-encoders (VGAE), focus on global topological features and ignore the local dense connection characteristics of nodes in the community. Experiments show that in the task of generating academic collaboration networks, the community structure similarity of mainstream models, namely the Normalized Mutual Information (NMI) index, is only 0.58, which is 28% lower than the real data; (2) Excessive consumption of computing resources: Recursive generation architectures, such as the deep autoregressive model (GraphRNN), need to generate edge sequences node by node, resulting in the graphics processing unit (GPU) memory occupying more than 48GB when generating a million-node graph, and the generation time increases exponentially; (3) Single control dimension: The existing controllable generation method, Content-Parsing Generative Adversarial Networks (CPGAN), only supports community size adjustment and cannot achieve decoupled control of attribute distribution (such as user age labels) and topological structure (such as community density). As a result, when adjusting the overlap of social network communities, the fidelity of node attributes is lost by 41%.
[0005] Although improved models for maintaining community structure (such as DCSBM and SBMGNN) introduce dynamic parameter optimization, their generation process still relies on static community division and is difficult to adapt to the distribution shift of cross-domain data. For example, when the model trained on the citation network is transferred to the generation of e-commerce user graphs, the similarity of community structure decreases by 53%. In addition, although graph convolution methods based on hierarchical pooling, such as Differentiable Pooling (DiffPool), can extract multi-scale features, the matrix calculation complexity of its pooling layer is ,The training time for processing a graph with 100,000 nodes exceeds 72 hours, which seriously restricts practical applications.
[0006] In summary, existing technologies have not yet solved the problem of coordinated optimization of controllability, structural fidelity, and computational efficiency in multi-domain community graph generation. Specifically: in terms of community structure preservation, existing models are difficult to achieve both community protection generation and community structure control at the same time; in terms of computational efficiency, although methods based on graph neural networks (GNNs) have advantages in representation learning, their high computational and parameter-intensive characteristics limit practical applications; in terms of controllability, existing methods cannot achieve decoupled generation of node attributes and graph structures. Therefore, there is an urgent need for a graph generation method that supports community control, is cross-domain adaptable, and is resource-efficient. Summary of the invention
[0007] The purpose of the present invention is to provide a controllable community graph data generation method and system.
[0008] The purpose of the present invention can be achieved by the following technical solutions: According to one aspect of the present invention, a method for generating controllable community graph data is provided, and the method steps include: S1. Train the student model through knowledge distillation and use the trained student model to extract node feature representation from multi-domain graph data; S2. Use the dual-channel decoupling generator and the community-aware graph convolution operator to perform dual-channel decoupling generation based on node feature representation to obtain node attributes and topology information respectively; then use the node attributes and topology information to construct a generated graph; S3. Construct a joint loss optimization based on the generated graph and design a multi-objective label vector. Use the joint loss optimization and multi-objective label vector to reversely optimize the dual-channel decoupled generator. S4. Use the optimized dual-channel decoupled generator to process the input graph and output the final generated graph, thereby completing the graph data generation.
[0009] As a preferred technical solution, the specific process of training the student model through knowledge distillation in S1 is as follows: first, the teacher model is used to extract node feature representations from multi-domain graph data through a cross-domain graph neural network; then, the intermediate layer output and final output probability of the teacher model are used as supervision signals; finally, the student model is trained through a cross-domain graph distillation algorithm to enable the student model to learn the mapping relationship between multi-domain graph data and node feature representations. When the distillation loss is less than the preset loss, the training is stopped to obtain the trained student model; among them, the teacher model adopts an improved version of the 3-layer graph convolutional network (Graph Convolutional Network v2, GCNv2) architecture, and the student model is a 2-layer multi-layer perceptron (Multi-Layer Perceptron, MLP); The specific formula for distillation loss is as follows: , , in, is the distillation loss function; D is the domain set; For the field d Variational encoder of z is a hidden variable; p ( z ) is the standard Gaussian prior distribution; λ is a hyperparameter; is the domain feature encoder; ψ is the shared distiller; For the field d Graph data.
[0010] As a preferred technical solution, in the cross-domain graph distillation algorithm, the weight of each domain increases according to the complexity of the graph. The specific formula is as follows: , in, For the field dThe weights are used to weight the distillation losses of each field in the cross-field graph distillation algorithm; β is a parameter that increases with the number of iterations; for d complexity rating of the domain; t is the current iteration number, is any domain within the domain set D.
[0011] As a preferred technical solution, the dual-channel decoupling generation in S2 includes node attribute reconstruction and topology structure reconstruction.
[0012] As a preferred technical solution, node attribute reconstruction is based on a multi-head graph attention network, and the attribute projection matrix is generated from the node feature representation as the node attribute; the topological structure reconstruction adopts sparse graph convolution, firstly generating the adjacency probability from the node feature representation, and then calculating the topological feature representation of the node based on the adjacency probability as the topological information. The specific formula is as follows: , , in, For Node i Attribute feature representation of ; is a multi-head graph attention encoder; is the input node feature vector; N ( i ) is a node i The set of neighbor nodes of is the attribute projection matrix; is the adjacency probability, indicating that the node i and nodes j The probability of the existence of an edge between them; σ is the sigmoid activation function; SGC is the sparse graph convolution operation; and Node i and nodes j The topological features of .
[0013] As a preferred technical solution, in the process of node attribute reconstruction and topology structure reconstruction, an orthogonal constraint loss is constructed to prevent feature coupling. The specific formula is as follows: , in, is the orthogonal loss term; is the attribute projection matrix; is the topological projection matrix; is the Frobenius norm.
[0014] As a preferred technical solution, the process of constructing a generated graph using node attributes and topological information in S2 includes local subgraph sampling, linear decoding and progressive generation process, wherein the local subgraph sampling process is specifically to select the central node based on the node degree probability and construct an egocentric network with a radius of 2; linear decoding is used to reduce the complexity of adjacency prediction; the progressive generation process is specifically to load the student model into the graphics processing unit (GPU) in stages to adapt to the scale of the expanded graph, and increase the number of nodes in each stage. , k is the stage number.
[0015] As a preferred technical solution, the joint loss optimization in S3 integrates the structural reconstruction term, the attribute fidelity term and the control constraint term; among them, the structural reconstruction term is calculated using the binary cross entropy of the adjacency matrix, the attribute fidelity term is calculated using the maximum mean difference, and the control constraint term is calculated using the mean square error of the control parameter; its specific formula is as follows: , in, is the total loss; It is a structural reconstruction item; is the attribute fidelity item; is the control constraint; α is the weight coefficient of the attribute fidelity item; β is the weight coefficient of the control constraint; is the generated adjacency matrix; is the true adjacency matrix; To generate the attribute distribution of data; is the attribute distribution of real data; M is the generated control parameter; is the target control parameter.
[0016] As a preferred technical solution, the multi-target label vector in S3 includes community control coding labels and attribute control coding labels, each label is 4 bits, consisting of 0 and 1; among them, the community control coding label is used to characterize the community density characteristics, and the attribute control coding label is used to characterize the attribute dispersion characteristics.
[0017] According to another aspect of the present invention, a controllable community graph data generation system is provided, the system works based on a controllable community graph data generation method as described above, the system includes a cross-domain feature distillation module, a dual-channel decoupling generation module, a dynamic control optimization module and a graph generation module; Among them, the cross-domain feature distillation module trains the student model through knowledge distillation, and uses the trained student model to extract node feature representations from multi-domain graph data; The dual-channel decoupling generation module uses the dual-channel decoupling generator and the community-aware graph convolution operator to perform dual-channel decoupling generation based on node feature representation to obtain node attributes and topology information respectively; then the node attributes and topology information are used to construct the generated graph; The dynamic control optimization module constructs a joint loss optimization based on the generation graph and designs a multi-objective label vector. It uses the joint loss optimization and multi-objective label vector to reversely optimize the dual-channel decoupled generator. The graph generation module uses the optimized dual-channel decoupled generator to process the input graph and output the final generated graph, thereby completing the graph data generation.
[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. In the present invention, a dual-channel decoupling generator is used to perform dual-channel decoupling generation based on node feature representation, and node attributes and topology information are obtained respectively to construct a generated graph; then the dual-channel decoupling generator is reversely optimized based on the generated graph; finally, the input graph is processed by the optimized dual-channel decoupling generator, and the final generated graph is output, thereby completing the graph data generation. Compared with the existing coupled generation method, the dual-channel design ensures the co-evolution of attributes and topology through the semantic alignment loss function, and cooperates with the community-aware graph convolution operator to make the generated social network node interaction behavior conform to the real semantic laws. Through the decoupled generation of node attributes and graph structure, the fidelity of the community structure generated by the graph is significantly enhanced.
[0019] 2. In the present invention, the teacher model is first used to extract node feature representations from multi-domain graph data, and then the student model is trained through a cross-domain graph distillation algorithm to learn the mapping relationship from data to features, and a graph generation framework for double distillation of community latent variables and model parameters is implemented. The large model representation is migrated to a lightweight architecture through dynamic knowledge compression technology. In resource-constrained scenarios (such as mobile devices or edge computing), the lightweight architecture can run efficiently while maintaining high performance, and is suitable for real-time graph data processing and large-scale graph analysis tasks.
[0020] 3. The process of constructing the generated graph in the present invention includes local subgraph sampling, linear decoding and progressive generation process, which significantly reduces the consumption of computing resources in the graph generation process. Especially when generating large-scale graphs, the progressive generation strategy and local subgraph sampling strategy effectively reduce the memory usage and computing time, so that the model can efficiently generate high-quality community graphs under limited resources, and its computing efficiency and resource utilization are significantly improved.
[0021] 4. The present invention uses a cross-domain graph distillation algorithm to extract node feature representations from multi-domain graph data and construct a feature space shared by multiple domains, so that the model can adapt to graph generation tasks in different domains. The weights of each domain are dynamically adjusted according to the complexity of the graph to ensure the generation effect of the model on complex graph data. Compared with the prior art, the present invention shows stronger adaptability and generalization ability in cross-domain graph generation tasks, and can effectively respond to graph generation needs in different fields.
[0022] 5. The multi-target label vector in the present invention includes community control coding labels and attribute control coding labels, each label is 4 bits, composed of 0 and 1; it creates an attribute constraint-oriented community graph generation paradigm, encoding domain knowledge into geometric constraints of the generation space through a differentiable logic layer. Compared with traditional retraining methods, this process supports dynamic injection of multi-dimensional attribute constraints (such as degree distribution, clustering coefficient, etc.), achieving fine-grained control while maintaining the integrity of the community structure. This allows the method to achieve controllable generation of graph structure and attributes in a variety of graph data generation.
[0023] 6. The present invention integrates the structural reconstruction term, attribute fidelity term and control constraint term through joint loss optimization, ensuring the high fidelity of the generated graph in topological structure and node attributes. The structural reconstruction term uses the adjacency matrix binary cross entropy, the attribute fidelity term uses the maximum mean difference, and the control constraint term uses the control parameter mean square error, so that the generated graph is highly consistent with the real graph in terms of community structure, node attributes and control parameters. Experiments show that the generated community graph is superior to the existing technology in indicators such as community structure similarity and node attribute fidelity. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A step diagram of a controllable community graph data generation method in the present invention; Figure 2 It is a flow chart of the synthesis diagram in the embodiment; Figure 3 A structural schematic diagram is generated for the controllable graph in the embodiment. DETAILED DESCRIPTION
[0025] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0026] Existing technologies have not yet solved the problem of collaborative optimization of controllability, structural fidelity, and computational efficiency in the generation of multi-domain community graphs. Specifically: (1) In terms of maintaining community structure, it is difficult for existing models to simultaneously achieve community protection generation and community structure control; (2) In terms of computational efficiency, although the graph neural network (GNN)-based method has advantages in representation learning, its high computational complexity and high number of parameters limit its practical application; (3) In terms of controllability, existing methods cannot achieve the decoupled generation of node attributes and graph structure. Therefore, there is an urgent need for a graph generation method that supports community control, cross-domain adaptation, and resource efficiency.
[0027] Example 1 In this embodiment, a controllable community graph data generation method is applied, and the method steps are as follows: Figure 1 As shown, specifically including: S1. Train the student model through knowledge distillation and use the trained student model to extract node feature representation from multi-domain graph data; S2. Use the dual-channel decoupling generator and the community-aware graph convolution operator to perform dual-channel decoupling generation based on node feature representation to obtain node attributes and topology information respectively; then use the node attributes and topology information to construct a generated graph; S3. Construct a joint loss optimization based on the generated graph and design a multi-objective label vector. Use the joint loss optimization and multi-objective label vector to reversely optimize the dual-channel decoupled generator. S4. Use the optimized dual-channel decoupled generator to process the input graph and output the final generated graph, thereby completing the graph data generation.
[0028] This solution uses three core modules, namely dynamic distillation, hierarchical generation and pluggable control, to achieve community structure preservation and efficient generation in cross-domain scenarios. Specifically, it includes: ① Cross-domain feature distillation layer: A domain-independent encoder is used to extract common features of graph structures (such as community overlap and degree distribution pattern), construct a transferable latent space, and use a student model to obtain graph encoding features at a low computational cost and achieve fast generation; ② Dual-channel decoupled generator: separates the attribute generation channel (processing node labels and features) from the topology generation channel (modeling adjacency matrix), and achieves independent regulation of structure-attribute through parameter isolation; ③ Dynamic control interface: Design differentiable mask matrix and community guided loss function to support real-time adjustment of parameters such as community density and attribute discreteness.
[0029] In this embodiment, the specific process of training the student model by knowledge distillation in S1 is: first, the teacher model is used to extract node feature representations from multi-domain graph data through a cross-domain graph neural network; then, the intermediate layer output and final output probability of the teacher model are used as supervision signals; finally, the student model is trained by a cross-domain graph distillation algorithm, so that the student model learns the mapping relationship between multi-domain graph data and node feature representations. When the distillation loss is less than the preset loss, the training is stopped to obtain the trained student model; wherein the teacher model adopts a 3-layer GCNv2 architecture, and the student model is a 2-layer MLP; The specific formula for distillation loss is as follows: , , in, is the distillation loss function; D is the domain set; For the field d Variational encoder of z is a hidden variable; p ( z ) is the standard Gaussian prior distribution; λ is a hyperparameter; is the domain feature encoder; ψ is the shared distiller; For the field d Graph data.
[0030] In this embodiment, in the cross-domain graph distillation algorithm, the weight of each domain increases according to the complexity of the graph, and the specific formula is as follows: , in, For the field d The weights are used to weight the distillation losses of each field in the cross-field graph distillation algorithm; β is a parameter that increases with the number of iterations; for d complexity rating of the domain; t is the current iteration number, is any domain within the domain set D.
[0031] In this embodiment, when the number of iterations is t =800, the distillation error converges to 0.12 (cosine similarity of high-dimensional data).
[0032] In this embodiment, the dual-channel decoupling generation in S2 includes node attribute reconstruction and topology structure reconstruction.
[0033] In this embodiment, node attribute reconstruction is based on a multi-head graph attention network, and an attribute projection matrix is generated from the node feature representation as the node attribute; topological structure reconstruction uses sparse graph convolution, firstly generating the adjacency probability from the node feature representation, and then calculating the topological feature representation of the node based on the adjacency probability as the topological information. The specific formula is as follows: , , in, For Node i Attribute feature representation of ; is a multi-head graph attention encoder; is the input node feature vector; N ( i ) is a node i The set of neighbor nodes of is the attribute projection matrix; is the adjacency probability, indicating that the node i and nodes j The probability of the existence of an edge between them; σ is the sigmoid activation function; SGC is the sparse graph convolution operation; and Node i and nodes j The topological features of .
[0034] In this embodiment, during the node attribute reconstruction and topology structure reconstruction process, an orthogonal constraint loss is constructed to prevent feature coupling, and its specific formula is as follows: , in, is the orthogonal loss term; is the attribute projection matrix; is the topological projection matrix; is the Frobenius norm.
[0035] In this embodiment, the process of constructing a generated graph using node attributes and topological information in S2 includes local subgraph sampling, linear decoding and progressive generation process, wherein the local subgraph sampling process is specifically to select the central node based on the node degree probability to construct an egocentric network with a radius of 2; linear decoding is used to reduce the complexity of adjacency prediction; the progressive generation process is specifically to load the student model into the GPU in stages to adapt to the scale of the expanded graph, and increase the number of nodes in each stage. , k is the stage number.
[0036] The process of constructing a generated graph (synthetic graph) and generating node attributes (synthetic attributes) by processing node-related information (corresponding to adjacency matrix and graph attribute-related processing) is as follows: Figure 2As shown, starting from the adjacency matrix, on the one hand, the community density is calculated and community control coding is performed; on the other hand, subgraph sampling is performed and latent variables are obtained through graph neural network encoding. Starting from the graph attributes, the attribute discreteness is calculated and attribute control coding is performed. Community control coding and attribute control coding work together with latent variables, and synthetic attributes are obtained through attribute decoding, and synthetic graphs are obtained through graph decoding. It embodies the steps of extracting features from the relevant information of the original graph and then generating them. In this embodiment, when the n=106 node graph is finally generated, the video memory occupies only 12.8GB.
[0037] The linear decoder reduces the complexity of neighbor prediction from Down to .
[0038] , in, is the generated adjacency matrix; Output feature matrix for encoder; is the decoder weight matrix; is the decoder bias term.
[0039] In this embodiment, the joint loss optimization in S3 integrates the structural reconstruction term, the attribute fidelity term and the control constraint term; wherein the structural reconstruction term is calculated using the binary cross entropy of the adjacency matrix, the attribute fidelity term is calculated using the maximum mean difference, and the control constraint term is calculated using the mean square error of the control parameter; the specific formula is as follows: , in, is the total loss; It is a structural reconstruction item; is the attribute fidelity item; is the control constraint; α is the weight coefficient of the attribute fidelity item; β is the weight coefficient of the control constraint; is the generated adjacency matrix; is the true adjacency matrix; To generate the attribute distribution of data; is the attribute distribution of real data; M is the generated control parameter; is the target control parameter.
[0040] In this embodiment, the multi-target label vector in S3 includes a community control coding label and an attribute control coding label. Each label is 4 bits, consisting of 0 and 1. There are three possibilities; among them, the community control coding label is used to characterize the community density characteristics, and the attribute control coding label is used to characterize the attribute dispersion characteristics.
[0041] In this example, the community density of data A is high (the community density label is 1,1,1,1) and the attribute discreteness is low (the attribute discreteness label is 0,0,0,0), so the concatenated label vector is: , in, t is the concatenated label vector; for community density labels; is the attribute dispersion label.
[0042] During the training process, the label vector t is used as a conditional input, and the generator parameters are optimized through back-propagation to achieve controllable generation of graph structure and attributes.
[0043] In this embodiment, the optimizer uses AdamW, and the learning rate , weight decay λ =0.01.
[0044] In summary, this scheme uses a dual-channel decoupled generator, in conjunction with a community-aware graph convolution operator, to perform dual-channel decoupled generation based on node feature representation, and obtains node attributes and topology information respectively to construct a generated graph; then reversely optimizes the dual-channel decoupled generator based on the generated graph; finally, the optimized dual-channel decoupled generator is used to process the input graph, and the final generated graph is output, thereby completing the graph data generation. Compared with the existing coupled generation method, the dual-channel design ensures the co-evolution of attributes and topology through a semantic alignment loss function, and in conjunction with a community-aware graph convolution operator, the generated social network node interaction behavior conforms to the real semantic laws. Through the decoupled generation of node attributes and graph structure, the fidelity of the community structure generated by the graph is significantly enhanced.
[0045] Example 2 In this embodiment, a controllable community graph data generation system is applied, which includes a cross-domain feature distillation module, a dual-channel decoupling generation module, a dynamic control optimization module and a graph generation module; Among them, the cross-domain feature distillation module trains the student model through knowledge distillation, and uses the trained student model to extract node feature representation from multi-domain graph data; the dual-channel decoupling generation module uses the dual-channel decoupling generator, in conjunction with the community-aware graph convolution operator, to perform dual-channel decoupling generation based on node feature representation, and obtains node attributes and topological information respectively; then the node attributes and topological information are used to construct the generated graph; the dynamic control optimization module constructs a joint loss optimization based on the generated graph, and designs a multi-objective label vector, and uses the joint loss optimization and multi-objective label vector to reversely optimize the dual-channel decoupling generator; the graph generation module uses the optimized dual-channel decoupling generator to process the input graph, and outputs the final generated graph, thereby completing the graph data generation.
[0046] In this embodiment, firstly, the input graph is given ;in, V is a node set; E is the edge set; each node All have corresponding node feature vectors .
[0047] The teacher model extracts high-level feature representations of graph structure and node attributes from the raw data and trains the student model.
[0048] Specifically, the teacher model generates feature representations in a shared latent space via a cross-domain graph neural network. Subsequently, the intermediate layer output and final output probability of the teacher model are used as supervisory signals to train a lightweight student model (MLP). The student model learns the mapping relationship from input data to the output features of the teacher model, thereby achieving efficient graph generation capabilities.
[0049] The method first performs cross-domain feature distillation, uses an encoder to extract common features of graph structures (such as community overlap and degree distribution pattern), constructs a transferable latent space, and uses a student model to obtain graph encoding features at a low computational cost and achieve fast generation. (1.1) Using the cross-domain graph distillation algorithm, the encoder is used to extract the common features of the graph structure (such as community overlap and degree distribution pattern), and a transferable latent space is constructed. The student model learns and constructs a mapping relationship from input data to the output features of the teacher model until the distillation error converges and the weights of each field are obtained. That is, the student model completes the acquisition of graph encoding features.
[0050] Construct a multi-domain shared feature space. The constructed multi-domain shared feature space and distillation error are: , , in, is the distillation loss function; D is the domain set; For the field d Variational encoder of z is a hidden variable; p ( z ) is the standard Gaussian prior distribution; λ is a hyperparameter; is the domain feature encoder; ψ is the shared distiller; For the field d Graph data.
[0051] The domain weight increases with the domain complexity score. The specific formula for the weight of each domain is: , in, For the field d The weights are used to weight the distillation losses of each field in the cross-field graph distillation algorithm; β is a parameter that increases with the number of iterations; for d complexity rating of the domain; t is the current iteration number, is any domain within the domain set D.
[0052] In this embodiment, when the number of iterations is t =800, the distillation error converges to 0.12 (cosine similarity of high-dimensional data).
[0053] (1.2) Perform dual-channel decoupling generation. The decoupling generation in this scheme includes two main stages: node attribute reconstruction and topology structure reconstruction; (1.2.1) Node attributes are reconstructed into node features based on the input graph, and the semantic information of the nodes is restored through the attribute generation channel; The attribute generation channel is based on the multi-head graph attention network, and the attribute projection matrix of the node is modeled. The specific formula is: , in, For Node i Attribute feature representation of ; is a multi-head graph attention encoder; is the input node feature vector; N ( i ) is a node i The set of neighbor nodes of is the attribute projection matrix.
[0054] In this embodiment, The attribute projection matrix of the node, output dimension .
[0055] (1.2.2) At the same time, the topological structure of the graph is gradually constructed by using the projection of node attributes and the conditions of the dynamic control module. The topological generation channel uses sparse graph convolution (SGC) to generate adjacency probability, and then calculates the topological projection matrix of the node based on the adjacency probability of the edges between nodes. The specific formula is: , in, is the adjacency probability, indicating that the node i and nodes jThe probability of the existence of an edge between them; σ is the sigmoid activation function; SGC is the sparse graph convolution operation; and Node i and nodes j The topological features of .
[0056] (1.2.3) In order to prevent the coupling of the characteristics of the node attribute reconstruction process and the topology structure reconstruction process, an orthogonal constraint loss is constructed, and its specific formula is: , in, is the orthogonal loss term; is the attribute projection matrix; is the topological projection matrix; is the Frobenius norm.
[0057] (1.3) Perform dynamic control and establish multiple target tags (each tag is 4 bits, total For example, if the community density of data A is high (the community density label is 1,1,1,1) and the attribute discreteness is low (the attribute discreteness label is 0,0,0,0), then the concatenated label vector is: ,
[0058] in, t is the concatenated label vector; for community density labels; is the attribute dispersion label.
[0059] During the training process, the label vector t As conditional input, the generator parameters are optimized through back-propagation to achieve controllable generation of graph structure and attributes.
[0060] The process of extracting features based on graph-related information, building relevant codes, and then generating the final graph data (synthetic graph) and attributes (synthetic attributes) through optimization processing is as follows: Figure 3 As shown in the figure, starting from the initial graph, the custom community density is determined and community control coding is performed; at the same time, the initial graph and graph attributes are mixedly coded with structural attributes. The attribute generation discreteness is calculated from the graph attributes and attribute control coding is performed. The community control coding, attribute control coding and structural attribute mixed coding together obtain the controlled latent variables, and the controlled latent variables are obtained by calculating the mean and standard deviation. , and finally generate synthetic attributes and synthetic graphs, which shows the key links from input to output in the generation of controllable community graph data.
[0061] (1.4) To reduce computational complexity, a lightweight generation pipeline is designed: (1.4.1) First, local subgraph sampling is performed, that is, based on the node degree probability Select the central node and construct an ego-graph with a radius of 2; (1.4.2) Reusing the linear decoder, the complexity of neighbor prediction is reduced from Down to .
[0062] , in, is the generated adjacency matrix; Output feature matrix for encoder; is the decoder weight matrix; is the decoder bias term.
[0063] (1.4.3) Progressive generation: The student model is put into the GPU in stages to adapt to the scale of the expanded graph, and the number of nodes is increased at each stage (k is the stage number), and when the n=106-node graph is finally generated, the video memory occupies only 12.8GB.
[0064] (1.5) Establish a joint optimization objective, namely the total loss function, which integrates structure reconstruction, attribute fidelity and control constraints: , in, is the total loss; It is a structural reconstruction item; is the attribute fidelity item; is the control constraint; α is the weight coefficient of the attribute fidelity item; β is the weight coefficient of the control constraint; is the generated adjacency matrix; is the true adjacency matrix; To generate the attribute distribution of data; is the attribute distribution of real data; M is the generated control parameter; is the target control parameter.
[0065] During the learning and training process, the optimizer uses AdamW (Adam with Weight Decay, Adam optimizer with weight decay), and the learning rate is , weight decay λ =0.01.
[0066] In summary, this solution realizes dual distillation of community latent variables and model parameters, and migrates large model representation to a lightweight architecture through dynamic knowledge compression technology. Compared with the existing coupled generation method, the dual-channel design ensures the co-evolution of attributes and topology through the semantic alignment loss function, and cooperates with the community-aware graph convolution operator to make the generated social network node interaction behavior conform to the real semantic laws. The dynamic control module pioneered the attribute constraint-oriented community graph generation paradigm, encoding domain knowledge into geometric constraints of the generation space through a differentiable logic layer. Compared with traditional retraining methods, this component supports dynamic injection of multi-dimensional attribute constraints (such as degree distribution, clustering coefficient, etc.), achieving fine-grained control while maintaining the integrity of the community structure. Experiments show that this method achieves efficient and controllable attribute graph generation in various graph data generation.
Claims
1. A controllable community graph data generation method, characterized in that: The method steps include: S1. Train the student model through knowledge distillation and use the trained student model to extract node feature representation from multi-domain graph data; S2. Use the dual-channel decoupling generator and the community-aware graph convolution operator to perform dual-channel decoupling generation based on node feature representation to obtain node attributes and topology information respectively; then use the node attributes and topology information to construct a generated graph; S3. Construct a joint loss optimization based on the generated graph and design a multi-objective label vector. Use the joint loss optimization and multi-objective label vector to reversely optimize the dual-channel decoupled generator. S4. Use the optimized dual-channel decoupled generator to process the input graph and output the final generated graph, thereby completing the graph data generation.
2. A controllable community graph data generation method according to claim 1, characterized in that: The specific process of training the student model by knowledge distillation in S1 is as follows: first, the teacher model is used to extract node feature representations from multi-domain graph data through a cross-domain graph neural network; then, the intermediate layer output and final output probability of the teacher model are used as supervision signals; finally, the student model is trained by a cross-domain graph distillation algorithm, so that the student model learns the mapping relationship between multi-domain graph data and node feature representations, and when the distillation loss is less than the preset loss, the training is stopped to obtain the trained student model; wherein, the teacher model adopts a 3-layer GCNv2 architecture, and the student model is a 2-layer MLP; The specific formula of the distillation loss is as follows: , , in, is the distillation loss function; D is the domain set; For the field d Variational encoder of z is a hidden variable; p ( z ) is the standard Gaussian prior distribution; λ is a hyperparameter; is the domain feature encoder; ψ is the shared distiller; For the field d Graph data.
3. A controllable community graph data generation method according to claim 1, characterized in that: In the cross-domain graph distillation algorithm, the weight of each domain increases according to the complexity of the graph. The specific formula is as follows: , in, For the field d The weights are used to weight the distillation losses of each field in the cross-field graph distillation algorithm; β is a parameter that increases with the number of iterations; for d complexity rating of the domain; t is the current iteration number, is any domain within the domain set D.
4. A controllable community graph data generation method according to claim 1, characterized in that: The dual-channel decoupling generation in S2 includes node attribute reconstruction and topology structure reconstruction.
5. A controllable community graph data generation method according to claim 1, characterized in that: The node attribute reconstruction is based on a multi-head graph attention network, and the attribute projection matrix is generated from the node feature representation as the node attribute; the topological structure reconstruction adopts sparse graph convolution, firstly generates the adjacency probability from the node feature representation, and then calculates the topological feature representation of the node based on the adjacency probability as the topological information. The specific formula is as follows: , , in, For Node i Attribute feature representation of ; is a multi-head graph attention encoder; is the input node feature vector; N ( i ) is a node i The set of neighbor nodes of is the attribute projection matrix; is the adjacency probability, indicating that the node i and nodes j The probability of the existence of an edge between them; σ is the sigmoid activation function; SGC is the sparse graph convolution operation; and Node i and nodes j The topological features of .
6. A controllable community graph data generation method according to claim 1, characterized in that: In the process of node attribute reconstruction and topology structure reconstruction, an orthogonal constraint loss is constructed to prevent feature coupling, and its specific formula is as follows: , in, is the orthogonal loss term; is the attribute projection matrix; is the topological projection matrix; is the Frobenius norm.
7. A controllable community graph data generation method according to claim 1, characterized in that: The process of constructing a generated graph using node attributes and topological information in S2 includes local subgraph sampling, linear decoding and progressive generation process, wherein the local subgraph sampling process is specifically to select the central node based on the node degree probability and construct an egocentric network with a radius of 2; linear decoding is used to reduce the complexity of adjacency prediction; the progressive generation process is specifically to load the student model into the GPU in stages to adapt to the scale of the expanded graph, and increase the number of nodes in each stage , k is the stage number.
8. A controllable community graph data generation method according to claim 1, characterized in that: The joint loss optimization in S3 integrates the structural reconstruction term, the attribute fidelity term and the control constraint term; wherein the structural reconstruction term is calculated using the binary cross entropy of the adjacency matrix, the attribute fidelity term is calculated using the maximum mean difference, and the control constraint term is calculated using the mean square error of the control parameter; the specific formula is as follows: , in, is the total loss; It is a structural reconstruction item; is the attribute fidelity item; is the control constraint; α is the weight coefficient of the attribute fidelity item; β is the weight coefficient of the control constraint; is the generated adjacency matrix; is the true adjacency matrix; To generate the attribute distribution of data; is the attribute distribution of real data; M is the generated control parameter; is the target control parameter.
9. A controllable community graph data generation method according to claim 1, characterized in that: The multi-objective label vector in S3 includes a community control coding label and an attribute control coding label, each label is 4 bits, and is composed of 0 and 1; wherein, the community control coding label is used to characterize the community density feature, and the attribute control coding label is used to characterize the attribute dispersion feature.
10. A controllable community graph data generation system, characterized in that: The system works based on a controllable community graph data generation method as described in any one of claims 1 to 9, the system comprising a cross-domain feature distillation module, a dual-channel decoupling generation module, a dynamic control optimization module and a graph generation module; The cross-domain feature distillation module trains a student model through knowledge distillation, and uses the trained student model to extract node feature representations from multi-domain graph data; The dual-channel decoupling generation module uses a dual-channel decoupling generator and cooperates with a community-aware graph convolution operator to perform dual-channel decoupling generation based on node feature representation to obtain node attributes and topology information respectively; and then uses the node attributes and topology information to construct a generation graph; The dynamic control optimization module constructs a joint loss optimization based on the generation graph, designs a multi-objective label vector, and uses the joint loss optimization and the multi-objective label vector to reversely optimize the dual-channel decoupling generator; The graph generation module uses the optimized dual-channel decoupled generator to process the input graph and outputs the final generated graph, thereby completing the graph data generation.
Citation Information
Patent Citations
Node classification method based on dual-channel knowledge distillation
CN113869425A
Knowledge distillation method, system, medium and equipment for cross-domain passive domain data
CN118298279A
Controllable hierarchical road topology generation system and method based on AIGC
CN118397449A
Feature pre-training and knowledge distillation combined time sequence diagram learning method
CN118780346A
Night-day image conversion method based on lightweight cyclic generative adversarial network
CN119399301A
Cited By
Industrial fan structure optimization design method and system based on AI assistance
CN120524827A