Unsupervised evolutionary graph neural network searching method based on proxy model
Through the unsupervised evolutionary graph neural network search method based on proxy model, the problem of traditional methods relying on label information and ignoring network layer relationships is solved, and a graph neural network with superior efficient search performance in label-free scenarios is realized.
Patent Information
- Application Number
- CN202510084492.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional graph neural network search methods rely heavily on label information and cannot handle label-free scenarios. They ignore the topological relationships and feature fusion strategies of the network layer in unsupervised search, resulting in low search efficiency and low efficiency of the generated graph neural network.
The unsupervised evolutionary graph neural network search method based on the proxy model is adopted. By designing the search space including the aggregation relationship, attention function, node message aggregation function, etc. of the network layer, and using genetic algorithms to search the graph neural network, the proxy model is used to evaluate the classification accuracy of the graph neural network, adapt to the label-free information scenario, and consider the topological relationship and feature fusion strategy of the network layer.
This method is adapted to the label-free information scenario, considers the topological relationship and feature fusion strategy of the network layer, reduces the model evaluation time, shortens the time for searching the model, and generates a graph neural network with better performance, with a wider search space and better generalization capabilities.
Smart Images

Figure CN120012823A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural architecture search technology, and in particular to an unsupervised evolutionary graph neural network search method based on a proxy model. Background Art
[0002] Graph neural networks have been extensively studied on graph data located in non-Euclidean space, and have flourished and shown good performance in different graph task applications, such as recommendation systems, traffic prediction, and fraud detection. As the scenarios become more and more complex, it becomes increasingly difficult to manually design graph neural networks. Neural network search, which automatically designs neural networks based on target tasks, has been widely studied.
[0003] In traditional graph neural network search methods, the search process is heavily dependent on label information and cannot handle unlabeled scenarios. For unsupervised graph neural network search methods, either a hypernetwork needs to be trained and only a limited number of network layers can be sampled, or the topological relationship of the network layers or the feature fusion strategy is ignored. The traditional use of evolutionary algorithms to search graph neural networks requires the evaluation of a large number of graph neural networks during the search process, and invalid graph neural networks will be generated, which is time-consuming. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides an unsupervised evolutionary graph neural network search method based on a proxy model;
[0005] An unsupervised evolutionary graph neural network search method based on a proxy model, the specific steps are as follows:
[0006] Step 1: Design the search space;
[0007] The search space includes aggregation relationships between network layers, attention functions, node message aggregation functions, activation function types, multi-head attention mechanisms, and enhancement strategies;
[0008] The enhancement strategy includes graph topology connection enhancement and graph node feature enhancement, as follows:
[0009] The topological connection enhancement of the graph is specifically implemented as follows:
[0010] Sample the modified subset from the original edge set ε with a set probability The calculation is as follows:
[0011]
[0012] Where (u,v) represents the edge between adjacent nodes u and v, (u,v)∈ε, is the probability of deleting (u,v), represents the set of edges in the generated view, The calculation formula is as follows:
[0013]
[0014] where p e is a hyperparameter, and yes The maximum and average values of Used to mitigate the impact of nodes with highly dense connections, p τ is the threshold, indicating the highest probability of deleting edge (u,v), The calculation formula is as follows:
[0015]
[0016] in represents the centrality score of the edge between adjacent nodes u and v, which is calculated as follows:
[0017]
[0018] Wherein formula (1) is the calculation of undirected graph, formula (2) is the calculation of directed graph, and Represents the node centrality scores of nodes u and v. There are three scoring methods for node centrality, namely degree centrality, eigenvector centrality, and PageRank centrality;
[0019] For node feature enhancement of the graph, noise is added, the importance of all features of each node is calculated, and a set of masks is generated according to the Bernoulli distribution to mask some features of the node to achieve node feature enhancement. The specific implementation is as follows:
[0020] First, sample a random vector Where F represents the feature dimension, each dimension Independently extracted from the Bernoulli distribution, that is, Generated node features x N represents the Nth feature of the node, N represents the number of node features, ° represents element-by-element multiplication, [·;·] represents the connection operation, reflecting the importance of the i-th dimension of the node feature Similar to the topological connection enhancement of the graph, the calculation is as follows:
[0021]
[0022] in for The maximum and average values of Used to mitigate the impact of nodes with highly dense connections, represents the weight of the i-th feature dimension of node u, p f is a hyperparameter used to control the overall magnitude of feature enhancement, p τ is the highest probability of deleting the i-th dimension, where The calculation is as follows:
[0023]
[0024] Wherein formula (3) represents the calculation of sparsely connected node feature weights of node u, x ui ∈{0,1},x ui represents the feature of the i-th dimension of node u, represents the node centrality score, V is the node set, and formula (4) is the calculation of the densely connected node feature weight of node u;
[0025] Step 2: Use genetic algorithms to search for graph neural networks and use proxy models to evaluate the classification accuracy of graph neural networks.
[0026] Step 2.1, randomly initialize several graph neural networks as the initial population individuals of the genetic algorithm, set the number of initialized individuals and the number of network layers of each individual; the encoding method is [[network layer 1, network layer 2, ..., network layer n], enhancement strategy], [network layer 1, network layer 2, ..., network layer n] represents the graph neural network, and the subsequent enhancement strategy represents the selected enhancement strategy type. The encoding method of each network layer is [t, f, h, a, c, m], where t represents the topological connection between the current network layer of the graph neural network and the previous network layer, represented by a one-dimensional array generated by 0 and 1, f represents the feature fusion strategy between network layers, h represents the number of multi-heads, a represents the activation function, c represents the node correlation coefficient, and m represents the node message aggregation method;
[0027] Use genetic algorithm search to search the graph neural network. For the individuals that meet any of the following rules generated during the search process, perform repair operations. The specific operations are as follows:
[0028] Rule 1: If f, h, a, c, m are all -1 in the network layer contained in the individual, all layers connected to it are disconnected and t=0 is set, i.e., an empty layer;
[0029] Rule 2: If t=[0,0,...,0], all layers connected to it are disconnected, f,h,a,c,m are set to invalid values -1, and t=0 is set, i.e., an empty layer;
[0030] Rule 3: If the current layer is valid but not connected to the next layer, randomly select a next layer to force a connection to it;
[0031] According to the encoding construction of each individual, the corresponding graph neural network is trained and the graph contrast learning method is used to train and record the classification accuracy of each individual. The number of generations is p. The graph data Cora is selected as the experimental data set, which is divided into training set, test set and validation set according to the set ratio. The category information of the node is not used in the training process, but is used in the testing and validation process.
[0032] When constructing the corresponding graph neural network, it is necessary to delete the network layer with [0,-1,-1,-1,-1,-1] in the individual encoding method, that is, the individual encoding length is fixed, but the corresponding graph neural network has an unfixed number of layers;
[0033] Step 2.2, randomly select N individuals, use multi-point crossover operation to generate offspring individuals, select individuals to cross each other to generate two offspring individuals, where the multi-point is set as an interval, the interval length is not fixed, the interval index and the interval length are less than the maximum number of layers of the network, and the offspring individuals generated in the crossover process that meet rules 1-rule 3 are repaired individually. The weight sharing method is used for the generated offspring, and the offspring initialization parameters are the parameters of the parent generation;
[0034] Step 2.3, perform mutation operation on the generated offspring; specifically, multiple mutation operations are used, and the initial number of mutations is set to 1. First, a number is generated to represent the type to be mutated, 0 represents the mutated topological connection, 1 represents the mutation operation type, and 2 represents the enhancement strategy. If 0 and 1 are selected, then another number is generated to represent the position to be mutated, and a number is randomly generated to represent the probability of the upcoming mutation. The mutation probability is set to 0.5. If the generated probability is greater than 0.5, a mutation operation is performed. If the selected mutation operation is 0, one of the topological connection numbers is randomly selected at the current position for mutation. If the selected mutation operation is 1, one of the layer feature fusion strategy, multi-head mechanism, activation function, attention function, and node message aggregation type is mutated. If the selected mutation operation is 2, other enhancement strategies are used. The offspring generated during the mutation process that meet rules 1-3 are individually repaired;
[0035] Step 2.4: Use the proxy model to predict the classification accuracy of the offspring, as follows:
[0036] The proxy model includes four machine learning models: multilayer perceptron, random forest, Gaussian process, and AdaBoost. The input is the individual code, and the output is the classification accuracy of the graph neural network corresponding to the individual. The Spearman correlation coefficients of the four models are compared in each generation, and the highest one is selected as the proxy model for predicting the current generation.
[0037] Compare the Spearman correlation coefficient calculated by the currently selected proxy model with the Spearman coefficient calculated by the proxy model selected in the previous generation. If the Spearman correlation coefficient calculated by the current generation is higher, the number of mutations in the current generation is 1.5 times the number of mutations in the previous generation. If the calculated number of mutations is not an integer, it is rounded up, otherwise the number of mutations starts from 1;
[0038] Step 2.5, select individuals to form a new generation population, the number of generations p = p + 1, each generation will be generated by merging all offspring and parents, using the elite selection strategy, the individuals with high classification accuracy will be retained until the number of initial population size is selected as the total population of the next generation;
[0039] Step 2.6: Determine whether the current generation has reached the maximum value. If so, end the evolution process, otherwise return to step 2.2;
[0040] Step 3: Test this method on the Cora dataset and select the individual with the highest classification accuracy on the Cora dataset as the required graph neural network.
[0041] The beneficial effects of adopting the above technical solution are:
[0042] The present invention provides an unsupervised evolutionary graph neural network search method based on a proxy model. In view of the shortcomings of the prior art, the present invention's method is suitable for unlabeled information scenarios, while considering the topological relationship of the network layers and the feature fusion strategy. It does not require pre-training of a super network, does not fix the number of network layers, and has a wider search space, thus including more graph neural networks with superior performance.
[0043] This method introduces a predictor model and a weight sharing strategy to reduce the time of model evaluation, thereby shortening the time of searching the model.
[0044] This method makes full use of the predictor model and adopts a multi-mutation strategy to evaluate as many graph neural networks as possible to obtain a graph neural network with better performance.
[0045] In this method, the neural network search on one data set can be generalized to multiple other data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is an overall flow chart of the community recommendation method of the present invention;
[0047] Figure 2 A flowchart of searching for a graph neural network according to the present invention;
[0048] Figure 3 This is an example of coding of the present invention. DETAILED DESCRIPTION
[0049] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0050] An unsupervised evolutionary graph neural network search method based on a proxy model, the overall framework is as follows Figure 1 As shown, the specific steps are as follows:
[0051] Step 1: Design the search space;
[0052] The search space includes aggregation relationships between network layers, attention functions, node message aggregation functions, activation function types, multi-head attention mechanisms, and enhancement strategies;
[0053] The enhancement strategy includes graph topology connection enhancement and graph node feature enhancement, as follows:
[0054] The topological connection enhancement of the graph is specifically as follows: consider whether to delete the edge. Since the two ends of the edge are connected to two nodes, for a directed graph, the centrality of the edge is the centrality score of the pointed node. For an undirected graph, the centrality of the edge is the average of the centrality scores of the two nodes. Three scoring methods of node centrality are considered, namely degree centrality, eigenvector centrality, and PageRank centrality. The specific implementation is as follows:
[0055] Sample the modified subset from the original edge set ε with a set probability The calculation is as follows:
[0056]
[0057] Where (u,v) represents the edge between adjacent nodes u and v, (u,v)∈ε, is the probability of deleting (u,v), reflecting the importance of edge (u,v), represents the set of edges in the generated view, The calculation formula is as follows:
[0058]
[0059] where p e is a hyperparameter that controls the total probability of deleting an edge. and yes The maximum and average values of Used to mitigate the impact of nodes with highly dense connections, p τ is the threshold, indicating the highest probability of deleting edge (u,v), The calculation formula is as follows:
[0060]
[0061] in represents the centrality score of the edge between adjacent nodes u and v, which is calculated as follows:
[0062]
[0063] Wherein formula (1) is the calculation of undirected graph, formula (2) is the calculation of directed graph, and represents the node centrality scores of nodes u and v;
[0064] For node feature enhancement of the graph, noise is added to destroy unimportant node features, the importance of all features of each node is calculated, and a set of masks is generated according to the Bernoulli distribution to mask some features of the node to achieve node feature enhancement. The specific implementation is as follows:
[0065] First, sample a random vector Where F represents the feature dimension, each dimension Independently extracted from the Bernoulli distribution, that is, f∈F, generated node features x N represents the Nth feature of the node, N represents the number of node features, ° represents element-by-element multiplication, [·;·] represents the connection operation, reflecting the importance of the i-th dimension of the node feature Similar to the topological connection enhancement of graph data, the calculation is as follows:
[0066]
[0067] in for The maximum and average values of Used to mitigate the impact of nodes with highly dense connections, represents the weight of the i-th feature dimension of node u, p f is a hyperparameter used to control the overall magnitude of feature enhancement, p τ is the highest probability of masking the i-th dimension, where The calculation is as follows:
[0068]
[0069] Wherein formula (3) represents the calculation of sparsely connected node feature weights of node u, x ui ∈{0,1},x ui represents the feature of the i-th dimension of node u, represents the node centrality score, V is the node set, and formula (4) is the calculation of the densely connected node feature weight of node u;
[0070] The probability p of setting one enhanced graph and the other enhanced graph in the two enhanced graphs is e and p f different.
[0071] Compute the contrastive loss objective, where for any node u, its generated embedding representation u in an augmented graph i , any other node embedding representation u k , whose embedding representation v generated in another augmented graph i Form a positive sample pair, and embed any other node to represent v k , each positive sample pair (u i , v i ) is defined as:
[0072]
[0073] represents the calculation of positive sample pairs, represents the calculation of negative sample pairs in the enhanced graph, The embedding representation of one augmented graph node and the embedding representation of another augmented graph node form a negative sample pair calculation, where τ is the temperature parameter for adjusting node similarity, s(·,·) represents cosine similarity, g(·) is a nonlinear projection, and the projection function in the method is implemented using a two-layer perceptron model. The loss function of the total positive sample pair is
[0074]
[0075] The node centrality scores include three types, namely degree centrality, eigenvector centrality, and PageRank centrality, namely There are three ways to calculate , and the edge score calculation is highly dependent on the node score. Therefore, this method considers these three enhancement strategies. Based on the above representation, the three enhancement strategies are represented by degree, eigenvector and pagerank respectively.
[0076] The search space is shown in Table 1.
[0077] Table 1 Search space
[0078]
[0079] Table 2 Attention function
[0080]
[0081] Step 2: Use genetic algorithms to search for graph neural networks and use proxy models to evaluate the classification accuracy of graph neural networks. For the specific process, see Figure 2, including the following steps:
[0082] Step 2.1, randomly initialize several graph neural networks as the initial population individuals of the genetic algorithm, set the number of initialized individuals and the number of network layers of each individual; in this embodiment, the number of initialized individuals is set to 200, and the maximum number of network layers of each individual is set to 8. The encoding method is [[network layer 1, network layer 2, ..., network layer n], enhancement strategy] (n<=8), [network layer 1, network layer 2, ..., network layer n] represents the graph neural network, and the subsequent enhancement strategy represents the selected enhancement strategy type. The encoding method of each network layer is [t, f, h, a, c, m], where t represents the topological connection between the current network layer of the graph neural network and the previous network layer, represented by a one-dimensional array generated by 0 and 1, f represents the feature fusion strategy between network layers, h represents the number of multi-heads, a represents the activation function, c represents the node correlation coefficient, and m represents the node message aggregation method. For example Figure 3 As shown;
[0083] Use genetic algorithm search to search the graph neural network. For the individuals that meet any of the following rules generated during the search process, perform repair operations. The specific operations are as follows:
[0084] Rule 1: If t=0 appears in the network layer contained in the individual, all layers connected to it are disconnected, and f, h, a, c, m are set to invalid values -1, that is, empty layers;
[0085] Rule 2: If t=[0,0,...,0], all layers connected to it are disconnected, f,h,a,c,m are set to invalid values -1, and t=0 is set, i.e., an empty layer;
[0086] Rule 3: If the current layer is valid but not connected to the next layer, randomly select a next layer to force a connection to it;
[0087] According to the encoding construction of each individual, the corresponding graph neural network is trained and the graph contrast learning method is used to train and record the classification accuracy of each individual. The number of generations is p=1. The graph data Cora is selected to verify the effectiveness of this method. The training set, test set, and validation set account for 20%, 20%, and 60% respectively. The category information of the node is not used in the training process, but is used in the testing and verification process.
[0088] When decoding into a graph neural network, the network layer containing [0,-1,-1,-1,-1,-1] in the individual encoding method needs to be deleted. That is, the individual encoding length is fixed, but the corresponding graph neural network has an unfixed number of layers.
[0089] Step 2.2, randomly select N individuals, use multi-point crossover operation to generate offspring individuals, the value of N is set to 100, select individuals and cross each other to generate two offspring individuals, here the multi-point crossover method is used, where the multi-point is set to an interval, the interval length is not fixed, the interval index and the interval length are less than the maximum number of layers of the network, such as [3,6], which indicates the position to be exchanged, and the offspring individuals generated in the crossover process that meet rules 1-rule 3 are repaired. For the generated offspring, since they are independent of the parent individuals, the weight sharing method is adopted, and the offspring initialization parameters are the parameters of the parent;
[0090] Step 2.3, perform mutation operation on the generated offspring; specifically, multiple mutation operations are used, the number of mutations is not fixed, and the initial number of mutations is set to 1. First, generate a number to represent the type to be mutated, 0 represents the mutated topological connection, 1 represents the mutation operation type, and 2 represents the enhancement strategy. If 0 and 1 are selected, then generate another number to represent the position to be mutated, and randomly generate a number to represent the probability of the upcoming mutation. Set the mutation probability to 0.5. If the generated probability is greater than 0.5, perform a mutation operation. If the selected mutation operation is 0, randomly select one of the topological connection numbers at the current position for mutation. If the selected mutation operation is 1, mutate one of the layer feature fusion strategy, multi-head mechanism, activation function, attention function, and node message aggregation type. If the selected mutation operation is 2, switch to other enhancement strategies, and perform individual repair on the offspring that meet rules 1-3 generated during the mutation process;
[0091] Step 2.4: Use the proxy model to predict the classification accuracy of the offspring, as follows:
[0092] The proxy model includes four machine learning models: multilayer perceptron, random forest, Gaussian process, and AdaBoost. The input is the individual code, and the output is the classification accuracy of the graph neural network corresponding to the individual. Each generation of individuals compares the Spearman correlation coefficients of the four models and selects the highest one as the proxy model for predicting the current generation of offspring.
[0093] Compare the Spearman correlation coefficient calculated by the currently selected proxy model with the Spearman coefficient calculated by the proxy model selected in the previous generation. If the Spearman correlation coefficient calculated by the current generation is higher, the number of mutations in the current generation is 1.5 times the number of mutations in the previous generation. If the calculated number of mutations is not an integer, it is rounded up, otherwise the number of mutations starts from 1;
[0094] By using a proxy model, only a small number of models need to be trained in each generation. In addition, since the offspring individuals and the parent individuals are independent and have high encoding similarity, a weight sharing strategy is adopted. The initial parameters of the offspring inherit the parameters of the parent, and the rest are predicted using the proxy model. These two methods can reduce the evaluation time.
[0095] Step 2.5, select individuals to form a new generation population, the generation number p = p + 1, the evolutionary generation number is set to 50 generations, and all the generated offspring and parent generations are merged in each generation. The elite selection strategy is adopted to retain individuals with high classification accuracy until the initial population size, that is, 200 individuals, is selected as the total population of the next generation;
[0096] Step 2.6: Determine whether the current generation has reached the maximum value. If so, end the evolution process, otherwise return to step 2.2;
[0097] Step 3: Test this method on the Cora dataset and select the individual with the highest classification accuracy on the Cora dataset as the required graph neural network.
[0098] This embodiment is applied to digital community recommendation. This method is used to search for the optimal graph neural network on the constructed social network graph, learn the node representation vector z of the graph, cluster z using the KMeans clustering algorithm, take the clustered clusters as communities, and generate community division results.
[0099] For graph construction, we first construct a social network G = {V, E, A, X} based on the user’s social records, where V is the set of nodes in the social network, V = {v1, v2, v3, …, v n}, represents the user individual, E represents the combination of the edges of the social network, e i,j =(v i ,v j )∈E, represents node v i and v j There are edges between them. If there are edges, it means that users have social connections. The matrix A∈R n×n Represents the adjacency matrix of a social network, when e i,j =1, A i,j =1, otherwise A i,j =0,X∈R n×m is the node feature matrix in the social network, i.e., user features, m represents the dimension of node features, X i,j Represents the j-th dimension feature value of node i;
[0100] For users, based on their social and characteristics, we can find communities with similar behaviors for them to achieve community recommendations.
[0101] In this embodiment, the generalization of the model is demonstrated. The neural network searched on the Cora dataset is tested with the PubMed dataset and the CiteSeer dataset. The optimal network and parameters are fine-tuned on the CiteSeer dataset and the PubMed dataset. The Cora, PubMed, and CiteSeer datasets are commonly used citation datasets. All three datasets consist of an undirected graph. Among them, the nodes represent paper documents, and the edges represent the citation relationship between paper documents. The node feature is a word vector. Each element of the word vector corresponds to a word, with a value of 0 or 1. The element value of 0 indicates that the word corresponding to the element does not appear in the paper, and the element value of 1 indicates that the word corresponding to the element appears in the paper. Each paper cites at least one other paper or is cited by other papers. The information of the three data sets is shown in Table 3:
[0102] Table 3 Dataset information
[0103] Dataset Number of nodes Number of edges Number of tags Feature Dimension Figure Task Category Cora 2078 5429 7 1433 Node Classification PubMed 19717 44338 3 500 Node Classification CiteSeer 3312 4723 6 3703 Node Classification
[0104] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.
Claims
1. An unsupervised evolutionary graph neural network search method based on a proxy model, characterized in that: The following steps are involved: Step 1: Design a search space; the search space includes aggregation relationships between network layers, attention functions, node message aggregation functions, activation function types, multi-head attention mechanisms, and enhancement strategies; Step 2: Use genetic algorithms to search for graph neural networks and use proxy models to evaluate the classification accuracy of graph neural networks. Step 3: Test this method on the Cora dataset and select the individual with the highest classification accuracy on the Cora dataset as the required graph neural network.
2. The unsupervised evolutionary graph neural network search method based on a proxy model according to claim 1, characterized in that: The enhancement strategy described in step 1 includes graph topology enhancement and graph node feature enhancement, as follows: The topological connection enhancement of the graph is specifically implemented as follows: Sample the modified subset from the original edge set ε with a set probability The calculation is as follows: Where (u,v) represents the edge between adjacent nodes u and v, (u,v)∈ε, is the probability of deleting (u,v), represents the set of edges in the generated view, The calculation formula is as follows: where p e is a hyperparameter, and yes The maximum and average values of Used to mitigate the impact of nodes with highly dense connections, p τ is the threshold, indicating the highest probability of deleting edge (u,v), The calculation formula is as follows: in represents the centrality score of the edge between adjacent nodes u and v, which is calculated as follows: Wherein formula (1) is the calculation of undirected graph, formula (2) is the calculation of directed graph, and Represents the node centrality scores of nodes u and v. There are three scoring methods for node centrality, namely degree centrality, eigenvector centrality, and PageRank centrality; The node feature enhancement of the graph is specifically to add noise, calculate the importance of all features of each node, generate a set of masks according to Bernoulli distribution, and mask some features of the node to achieve node feature enhancement. The specific implementation is as follows: First, sample a random vector Where F represents the feature dimension, each dimension Independently extracted from the Bernoulli distribution, i.e. Generated node features x N represents the Nth feature of the node, N represents the number of node features, ° represents element-by-element multiplication, [·;·] represents the connection operation, reflecting the importance of the i-th dimension of the node feature The calculation is as follows: in for The maximum and average values of Used to mitigate the impact of nodes with highly dense connections, represents the weight of the i-th feature dimension of node u, p f is a hyperparameter used to control the overall magnitude of feature enhancement, p τ is the highest probability of deleting the i-th dimension, where The calculation is as follows: Wherein formula (3) represents the calculation of sparsely connected node feature weights of node u, x ui ∈{0,1},x ui represents the feature of the i-th dimension of node u, represents the node centrality score, V is the node set, and formula (4) is the calculation of the densely connected node feature weight of node u.
3. The unsupervised evolutionary graph neural network search method based on a proxy model according to claim 1, characterized in that: The step 2 comprises the following steps: Step 2.1, randomly initialize several graph neural networks as the initial population individuals of the genetic algorithm, set the number of initialized individuals and the number of network layers of each individual; the encoding method is [[network layer 1, network layer 2, ..., network layer n], enhancement strategy], [network layer 1, network layer 2, ..., network layer n] represents the graph neural network, and the subsequent enhancement strategy represents the selected enhancement strategy type. The encoding method of each network layer is [t, f, h, a, c, m], where t represents the topological connection between the current network layer of the graph neural network and the previous network layer, represented by a one-dimensional array generated by 0 and 1, f represents the feature fusion strategy between network layers, h represents the number of multi-heads, a represents the activation function, c represents the node correlation coefficient, and m represents the node message aggregation method; Step 2.2, randomly select N individuals, use multi-point crossover operation to generate offspring individuals, select individuals to cross each other to generate two offspring individuals, where the multi-point is set as an interval, the interval length is not fixed, the interval index and the interval length are less than the maximum number of layers of the network; Step 2.3, perform mutation operation on the generated offspring; Step 2.4: Use the proxy model to predict the classification accuracy of the offspring; Step 2.5, select individuals to form a new generation population, the number of generations p = p + 1, each generation will be generated by merging all offspring and parents, using the elite selection strategy, the individuals with high classification accuracy will be retained until the number of initial population size is selected as the total population of the next generation; Step 2.6: Determine whether the current generation has reached the maximum value. If so, end the evolution process; otherwise, return to step 2.
2.
4. The unsupervised evolutionary graph neural network search method based on a proxy model according to claim 3 is characterized in that: Step 2.1 is as follows: Use the genetic algorithm to search the graph neural network, and perform repair operations on the individuals that meet any of the following rules generated during the search process. The specific operations are as follows: Rule 1: If f, h, a, c, m are all -1 in the network layer contained in the individual, all layers connected to it are disconnected and t=0 is set, i.e., an empty layer; Rule 2: If t=[0,0,...,0], all layers connected to it are disconnected, f,h,a,c,m are set to invalid values -1, and t=0 is set, i.e., an empty layer; Rule 3: If the current layer is valid but not connected to the next layer, randomly select a next layer to force a connection to it; According to the coding of each individual, the corresponding graph neural network is constructed, and the graph contrast learning method is used to train and record the classification accuracy of each individual. The number of generations is p. The graph data Cora is selected as the experimental data set, which is divided into training set, test set, and validation set according to the set ratio. The category information of the node is not used in the training process, but is used in the testing and validation process; When constructing the corresponding graph neural network, it is necessary to delete the network layer containing [0,-1,-1,-1,-1,-1] in the individual encoding method. That is, the individual encoding length is fixed, but the corresponding graph neural network has an unfixed number of layers.
5. The unsupervised evolutionary graph neural network search method based on a proxy model according to claim 4, characterized in that: In step 2.2, individual repair is performed on the offspring individuals generated during the crossover process that meet rules 1 to 3, and a weight sharing method is adopted for the generated offspring, and the offspring initialization parameters are the parameters of the parent generation.
6. The unsupervised evolutionary graph neural network search method based on a proxy model according to claim 4, characterized in that: The step 2.3 is specifically as follows: multiple mutation operations are adopted, the initial number of mutations is set to 1, first a number is generated to represent the type to be mutated, 0 represents the mutated topological connection, 1 represents the mutation operation type, 2 represents the enhancement strategy, if the mutation operations 0 and 1 are selected, then another number is generated to represent the position to be mutated, a number is randomly generated to represent the probability of the upcoming mutation, the mutation probability is set to 0.5, if the generated probability is greater than 0.5, a mutation operation is performed, if the selected mutation operation is 0, then one of the numbers of topological connections is randomly selected at the current position for mutation, if the selected mutation operation is 1, then one of the layer feature fusion strategy, multi-head mechanism, activation function, attention function, node message aggregation type is mutated, if the selected mutation operation is 2, then other enhancement strategies are used, and the offspring generated during the mutation process that meet rules 1-rule 3 are individually repaired.
7. The unsupervised evolutionary graph neural network search method based on a proxy model according to claim 4, characterized in that: The step 2.4 is specifically as follows: The proxy model includes four machine learning models: multi-layer perceptron, random forest, Gaussian process, and AdaBoost. The input is the encoding of the individual, and the output is the classification accuracy of the graph neural network corresponding to the individual. The Spearman correlation coefficients of the four models are compared in each generation, and the highest one is selected as the proxy model for predicting the current generation of offspring; Compare the Spearman correlation coefficient calculated by the currently selected proxy model with the Spearman coefficient calculated by the proxy model selected in the previous generation. If the Spearman correlation coefficient calculated by the current generation is higher, the number of mutations in the current generation is 1.5 times the number of mutations in the previous generation. If the calculated number of mutations is not an integer, it is rounded up. Otherwise, the number of mutations starts from 1.