Ensemble learning method based on evolutionary graph neural architecture search and application of ensemble learning method in graph mining
By introducing niche strategies and TPE integrated learning methods in GNAS, Ensemble-GNAS solves the problem of limited generalization capabilities of existing GNAS algorithms in large-scale graph mining tasks, achieving higher accuracy and stability.
Patent Information
- Application Number
- CN202510077731.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-27
AI Technical Summary
The existing graph neural architecture search (GNAS) algorithms have limited generalization capabilities, insufficient robustness and lack of diversity in large-scale graph mining tasks, making it difficult to discover the optimal neural network architecture.
Using an integrated learning method based on evolutionary graph neural architecture search, the Ensemble-GNAS is designed, and niche strategies and tree Parzen estimator (TPE) are used to improve local search capabilities and diversity of candidate networks, and the weight of the base learner in the integrated model is optimized.
It improves the accuracy and stability of the algorithm, enhances the generalization ability and robustness of the model, and can show better performance in tasks such as node classification, link prediction and graph classification.
Smart Images

Figure CN120218165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph mining, in particular to an ensemble learning method based on evolutionary graph neural architecture search and its application in graph mining. Background Art
[0002] GRAPH is a general data form for simulating complex relationships, such as social networks and biological networks. Graph data analysis is an important research topic in artificial intelligence. In recent years, graph neural networks (GNNs) have become the dominant paradigm for mining potential information in graph data, such as node classification, link prediction, and graph classification. The applications of GNNs involve social network analysis, network security, and medical diagnosis. Popular GNNs, including but not limited to GCN, GAT, and Graphsage, have achieved good results in graph data mining. In fact, in order to obtain the expected performance for a specific task, it is usually tricky even for neural network experts to design GNN architectures based on task-specific features. Neural architecture search (NAS), as a method for automatically designing neural network architectures, is an important topic in deep learning research in recent years. In addition, many researchers have extended NAS to the field of GNNs and designed graph neural architecture search (GNAS) methods for optimizing the architectures of GNNs when solving graph data mining tasks.
[0003] However, most existing GNAS algorithms are based on reinforcement learning, differentiable search methods, and evolutionary algorithms. These algorithms are restricted by the high dimensionality of data and ignore the distribution of different solutions in the search space, resulting in limited generalization ability, insufficient robustness, and lack of diversity. Designing more effective techniques to discover the optimal neural network architecture for large-scale graph mining tasks is a great challenge.
[0004] Recently, many studies have shown that ensemble learning (EL) can complete learning tasks by constructing and combining multiple GNNs, thereby improving learning performance and enhancing model generalization ability. The key points of these methods are that multiple GNN models with different initializations or architectures are trained and combined to improve the overall accuracy, reduce bias and variance, and mitigate the impact of noisy data. However, most existing GNN-based EL methods use manually designed graph neural networks as candidate networks for EL, which is still troubled by the problems of requiring a large amount of manpower and rich domain knowledge. Therefore, it is promising to automatically generate multiple GNNs simultaneously to construct an ensemble of GNNs. Summary of the Invention
[0005] To solve the problems existing in the prior art, the object of the present invention is to provide an ensemble learning method based on evolutionary graph neural architecture search and its application in graph mining. The present invention has better accuracy and stability.
[0006] To achieve the above object, the technical solution adopted by the present invention is: an ensemble learning method based on evolutionary graph neural architecture search, comprising the following steps:
[0007] Step 1, design an ensemble model of an ensemble learning EL method based on evolutionary graph neural architecture search GNAS, denoted as Ensemble-GNAS, and perform graph learning by retaining and integrating multiple solutions with different architectures but similar performance;
[0008] Step 2, use the evolutionary graph neural architecture search GNAS method based on the niche strategy to improve the local search ability of the algorithm and the diversity of the ensemble learning EL candidate networks;
[0009] Step 3, use an ensemble fusion strategy based on the tree-structured Parzen estimator TPE to optimize the weights of the base learners in the ensemble model.
[0010] As a further improvement of the present invention, in Step 1, in the architecture search stage, given the search space of the graph neural architecture S a the niche size N, and the graph data D, perform evolutionary graph neural architecture search GNAS to find the best individual in each niche; is the best individual in niche i, which maximizes the expected accuracy Acc val on the validation set D val (X), that is:
[0011]
[0012] After the GNAS search algorithm, obtain M best individuals as the EL candidate networks based on TPE; the final ensemble output is shown as:
[0013]
[0014] where out i represents the output of the i-th candidate network, and w i represents the weight optimized by TPE.
[0015] As a further improvement of the present invention, in Step 2, the evolutionary graph neural architecture search GNAS method based on the niche strategy is specifically as follows:
[0016] The evolutionary process includes crossover and mutation. Use uniform crossover to generate new offspring. The d-th dimensional element of the generated offspring O d is shown as follows:
[0017]
[0018] where X d and Y dRepresents the d-th dimensional element of two parents, R d Represents the d-th dimensional element of a random binary array with the same dimension as the parents; the mutation operation randomly generates a new individual;
[0019] When evaluating an individual, first train the individual network architecture on the training set, and then verify the trained network on the validation set; the j-th individual X i,j The fitness in the i-th niche is as follows:
[0020] fitess(X i,j ) = Acc val (X i,j ) (4)
[0021] Divide the dataset into a training set, a validation set, and a test set, then initialize the population based on the search space, and divide the individuals in the population into different niches; evaluate each individual in the population and record their fitness values; during the architecture evolution process, each individual crosses with another individual to produce offspring, and if mutation occurs, the offspring will become a completely new individual, compare the fitness values of the parent and the offspring, and retain the excellent individuals.
[0022] As a further improvement of the present invention, in step 2, the expression of the diversity OD of the ensemble learning EL candidate network is as follows:
[0023]
[0024] Where is the total number of samples for validation, n is the number of candidate networks; for each sample s in the validation set, the outputs generated by any two different networks i and j are respectively denoted as O i (s) and O j (s).
[0025] As a further improvement of the present invention, it further includes: dividing the network architectures with the same number of hidden units in the population into the same niche.
[0026] As a further improvement of the present invention, it further includes: according to the search space, the GNN architecture designs the network layer in a stacked structure, each layer consists of four components, and the four components jointly affect the performance and output of the network. The GNN architecture is described as an ordered list, and each element represents the value of an architecture component.
[0027] As a further improvement of the present invention, when initializing the population, the population is divided into M niches, each individual in a single niche has a specific number of hidden units; the number of individuals in each niche is N, that is, the number of individuals in the entire population is MN.
[0028] As a further improvement of the present invention, in step 3, the integration and fusion strategy based on the tree-structured Parzen estimator (TPE) is specifically as follows:
[0029] TPE is based on the weighted search space S w Initialize the probability model; then, in each iteration, generate a candidate parameter configuration c based on the existing probability model, evaluate the performance of the candidate parameter configuration and update the model, and continue until the preset number of iterations is reached or other stopping conditions are met;
[0030] Take the weights of the candidate network as parameters, with the range positioned as continuous numbers from 0 to 1; use the validation accuracy as the objective function, and obtain the optimal weight c by using TPE * ; use the optimal weight c * Weightedly integrate the candidate networks to obtain an integrated model; finally, apply the integrated model to the test set to obtain the final result.
[0031] The present invention also provides an application of the above-mentioned ensemble learning method based on evolutionary graph neural architecture search in graph mining.
[0032] The beneficial effects of the present invention are:
[0033] The Ensemble-GNAS of the present invention improves the local search ability of the algorithm and the diversity of the EL candidate networks by using the evolutionary GNAS method based on the niche strategy. Secondly, Ensemble-GNAS optimizes the weights of the base learners in the integrated model by using the integration and fusion strategy based on TPE. The proposed algorithm is tested not only on node classification datasets, including citation networks of bioinformatics data and cancer-driven gene identification networks, but also on problems related to link prediction and graph classification. Experimental results show that the algorithm has better accuracy and stability compared with the state-of-the-art methods. Description of the Drawings
[0034] Figure 1 It is the main framework diagram of Ensemble-GNAS in the embodiment of the present invention;
[0035] Figure 2 It is the schematic diagram of the principle of the EL algorithm in the embodiment of the present invention;
[0036] Figure 3 It is the schematic diagram of the individual encoding and decoding method in the embodiment of the present invention;
[0037] Figure 4 It is the schematic diagram of the output difference of two datasets in the embodiment of the present invention;
[0038] Figure 5 It is the schematic diagram of the comparison of search efficiency in the embodiment of the present invention;
[0039] Figure 6 This is a schematic diagram for parameter analysis in an embodiment of the present invention. Detailed implementation manners
[0040] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0041] Embodiment
[0042] An ensemble learning method based on evolutionary graph neural architecture search, called Ensemble-GNAS. A described EL method based on evolutionary GNAS, called Ensemble-GNAS, includes two improved strategies, an evolutionary GNAS method based on a niche strategy and an EL strategy based on a tree-structured Parzen estimator (TPE).
[0043] As Figure 1 shown, this embodiment provides an ensemble learning method based on evolutionary graph neural architecture search, called Ensemble-GNAS, which performs graph learning by retaining and integrating multiple solutions with different architectures but similar performance. It includes the following:
[0044] Content 1: Ensemble-GNAS algorithm framework
[0045] This method mainly consists of two parts: niche-based evolutionary GNAS and TPE-based EL. In the architecture search stage, given the search space of the graph neural architecture S a niche size N, and graph data D, evolutionary GNAS is executed to find the best individual in each niche. is the best individual in niche i, which maximizes the expected accuracy Acc val on the validation set D val (X), that is
[0046]
[0047] After the above GNAS search algorithm, M best individuals can be obtained as candidate networks for TPE-based EL. The final ensemble output is shown as:
[0048]
[0049] where represents the output of the i-th candidate network, and w i represents the weight optimized by TPE.
[0050] The entire process of Ensemble-GNAS is summarized in Algorithm 1. The niche-based evolutionary GNAS is used to search for candidate networks, and then the TPE-based EL is used to integrate the candidate networks.
[0051] Algorithm 1: Ensemble-GNAS
[0052]
[0053] Content 2: Evolutionary GNAS Method Based on Niche Strategy
[0054] The base learners of EL have an important impact on the performance of the ensemble model. The base learners not only require high accuracy but also maintain high diversity. However, existing GNAS methods focus on searching for a single network architecture with the highest accuracy, making it difficult to meet the requirements of EL. Therefore, an evolutionary architecture search strategy based on the niche strategy is considered, where each niche focuses on searching for network architectures with the same number of hidden units. The advantage of this is that it can not only improve the local search ability of the algorithm but also maintain a high diversity of candidate networks.
[0055] The evolutionary process includes crossover and mutation. Uniform crossover is used to generate new offspring. The d-th element of the generated offspring O d is shown as follows:
[0056]
[0057] where X d and Y d represent the d-th elements of two parents, and R d represents the d-th element of a random binary array with the same dimension as the parents. The mutation operation randomly generates a new individual.
[0058] When evaluating an individual, first train the individual network architecture on the training set, and then verify the trained network on the validation set. The fitness of the j-th individual X i,j in the i-th niche is shown as follows:
[0059] fitess(X i,j ) = Acc val (X i,j ) (4)
[0060] Algorithm 2 shows the detailed search algorithm. First, the dataset is divided into training set, validation set and test set. Then the population is initialized based on the search space, and the individuals in the population are divided into different niches. Each individual in the population is evaluated and their fitness values are recorded. During the architecture evolution process, each individual crosses with another individual to produce offspring. If a mutation occurs, the offspring will become a completely new individual. Compare the fitness values of the parent and offspring, and retain the excellent individuals. It is worth noting that the algorithm only allows the best individuals to cross with individuals in other niches, which improves the local search ability and population diversity of the algorithm. After the search algorithm, the best individual in each individual's field and its network output will be retained as candidates for the next EL operation.
[0061] Algorithm 2: Architecture evolution process
[0062]
[0063] Many studies have shown that the diversity of candidate networks is conducive to the improvement of EL performance. Therefore, it is necessary to ensure that the candidate networks have a certain diversity of outputs. The expression of output diversity (OD) is as follows:
[0064]
[0065] in is the total number of samples used for validation, and n is the number of candidate networks. For each sample s in the validation set, the outputs generated by any two different networks i and j are represented as O i (s) and O j (s). OD quantifies the average difference between the outputs produced by each pair of networks in the candidate set. Specifically, OD is a measure of the deformability or inconsistency of the predictions made by different networks given the same data points. By calculating this average output difference, one can gain insight into the degree of agreement or disagreement within a set of candidate networks, thereby informing decisions related to model selection, ensemble building, or further refinement of the network architecture and training process.
[0066] New niche partitioning strategy: Due to the particularity of the network architecture encoding strategy, it is difficult to calculate the crowding distance between different individuals, so it is necessary to design a new niche strategy. In order to adapt to the encoding strategy of the network architecture, a new niche partitioning strategy is proposed. In order to improve the diversity of the best individuals and promote the selection of networks of different sizes for different tasks, the network architectures with the same number of hidden units in the population are divided into the same niche. Since each niche focuses on searching for networks with the same number of hidden units during the architecture search process, there is no need to search the number of hidden units in the network architecture.
[0067] New search space: An architecture evolution algorithm that requires designing a new search space to adapt to. The improved search space (see Table I). According to the search space, the encoding and decoding methods of the network layer are as Figure 3 shown. The GNN architecture is designed with a stack structure, and each layer consists of four specific components, which jointly affect the performance and output of the network. This modular hierarchical framework allows for easy understanding and adjustment because each component can be understood independently and modified as needed to optimize the network behavior. An example of a single-layer GNN structure is represented as a serialized list, where each element in the list corresponds to a specific component of the layer. By clearly depicting the internal structure and understanding the role of each component, it becomes easier to understand how the GNN processes information and produces meaningful outputs. In summary, the stack architecture of GNNs with a modular design allows for the effective implementation and fine-tuning of graph neural networks. The clear description of the components in each layer and the ability to represent them as a serialized list further enhance the understandability and adaptability of these powerful models. The network layer shown in the figure uses GAT as the attention function and has 2 attention heads. It aggregates information by summarizing the information of adjacent nodes and uses elu for non-linear activation.
[0068] Table I Search space of structural components
[0069]
[0070]
[0071] New population initialization strategy: An architecture evolution algorithm that requires designing a new population initialization strategy to adapt to. In the population initialization method, the population is divided into M niches, and each individual in a single niche has a specific number of hidden units. The number of individuals in each niche is N, that is, the number of individuals in the entire population is MN. The population initialization method is shown in Table II.
[0072] Table II Population initialization method
[0073]
[0074] Content 3: Ensemble fusion strategy based on Tree-structured Parzen Estimator (TPE)
[0075] Traditional EL methods assign fixed weights to the outputs of the base learners of EL. Due to this ensemble strategy, base learners with insufficient performance may disrupt the entire ensemble model. Optimizing the weights can significantly reduce the risks brought by low-quality base learners and improve the accuracy and robustness of the ensemble model.
[0076] The EL method based on classification labels may suffer from information loss, leading to performance degradation and weakened robustness. The ensemble method based on classification confidence scores can make more full use of the prediction information of the model by considering the classification probabilities output by the sub-learners rather than just the classification labels. This method can capture the uncertainties of different category models, reduce information loss, and achieve more accurate predictions during the ensemble process.
[0077] TPE is used to optimize the ensemble weights of the candidate networks. The main steps for optimizing the ensemble weights are shown in Algorithm 3. TPE is based on the weight search space S w Initialize the probability model. Then, in each iteration, based on the existing probability model, generate a candidate parameter configuration c, and the algorithm evaluates the performance of the candidate parameter configuration and updates the model. This process continues until a preset number of iterations is reached or other stopping conditions are met.
[0078] Take the weights of the candidate networks as parameters and locate their ranges as continuous numbers from 0 to 1. Use the validation accuracy as the objective function. By using TPE, the optimal weight c can be obtained * . Use c * Weight the ensemble candidate networks to obtain an ensemble model. Finally, apply the ensemble model to the test set to obtain the final result.
[0079] Algorithm 3: Ensemble Strategy of Ensemble-GNAS
[0080]
[0081] In this embodiment, the performance of Ensemble-GNAS on citation networks and cancer-driven gene identification networks is used to evaluate its performance on node classification problems.
[0082] Citation networks are a popular type of dataset used in the study of graph neural networks. These datasets include Cora, Citeseer, and Pubmed, which vary in size and function. Each dataset consists of research papers that are classified into specific categories or classes. The dataset is divided into three different subsets: training, validation, and test. This stratification ensures a robust evaluation framework for the model being developed. Specifically, in the training phase, 20 nodes are assigned to each class, enabling the model to learn from different but representative data point samples of each class. Entering the validation phase, the dataset consists of 500 nodes to evaluate the generalization ability of the model. Finally, the test set has a large number of 1000 nodes, providing a comprehensive test platform for evaluating the final performance of the model on unknown data.
[0083] Another dataset used in GNN research is the cancer-driven gene identification network, including KEGG, DawnNet, and Regnetwork. These datasets have different numbers of positive and negative samples, making them suitable for evaluating the performance of GNNs in predicting cancer-related genes. Overall, these datasets provide diverse options for researchers to explore and develop GNN models for various applications. The statistics of the datasets are shown in Table III.
[0084] Table III Statistics of the Datasets
[0085]
[0086] Experiments
[0087] Parameter Settings
[0088] The parameter settings are not arbitrary but are based on extensive experiments and rigorous sensitivity analysis of various parameters. The primary goals are twofold: First, to ensure that the parameter settings make the comparison between Ensemble-GNAS and other methods fair. Second, to maintain the best performance of Ensemble-GNAS in the accuracy and efficiency experiments. In the NAS part, for the citation network node classification problem, the number of niches (M) is 7, and each niche corresponds to 4, 8, 16, 32, 64, 128, 256 hidden units. The number of individuals (N) in each niche is 7, which means the total number of individuals in the population is 49. For the cancer-driven gene identification node classification problem, the number of niches (M) is 5, and each niche corresponds to 4, 8, 16, 32, 64 hidden units. The number of individuals within the territory of each individual (N) is 10, which means the total number of individuals in the group is 50. For the hyperparameters during network training, the probability p of applying dropout is 0.6, the learning rate lr is 0.005, and the L2 regularization λ is 0.0005 as the default parameters. The number of evolutionary generations (G) is set to 50, and the number of training generations is set to 300. For link prediction and graph classification, each search algorithm uses the same search parameters as the node classification task.
[0089] To ensure the fairness of performance comparison, the test set accuracy is used as the evaluation metric on the cora, Citeseer, and Pubmed datasets. For cancer driver gene networks, AUROC (Area Under the Receiver Operating Characteristic Curve) and AUPRC (Area Under the Precision-Recall Curve) are used to evaluate the algorithms. AUROC focuses on the trade-off between the true positive rate and the false positive rate of the model at different probability thresholds and is applicable to imbalanced class distribution problems, especially when the model involves misclassification of negative samples. On the other hand, AUPRC focuses on balancing the accuracy and recall of the model and is especially applicable to scenarios where the positive class samples are few or very important. For link prediction and graph classification, AUC and test set accuracy are used as metrics for performance comparison.
[0090] Baseline methods
[0091] Ensemble-GNAS was compared with handcrafted GNN methods and GNAS methods.
[0092] Handcrafted GNN methods:
[0093] · Chebyshev is used to adapt CNNs to graph learning, which involves using Chebyshev polynomial bases for spectral filtering. This method allows CNNs to be applied to graph data, enabling effective aggregation of information from neighboring nodes.
[0094] · GCN is a two-layer architecture that uses spectral-based convolutional filters to aggregate information from neighbors.
[0095] The GCN model has been widely applied to various graph-related tasks and has shown impressive performance in many applications.
[0096] · GAT also uses Chebyshev polynomial bases for spectral filters but includes an attention mechanism to weigh the importance of different neighbors when aggregating information. This attention mechanism helps the model focus on relevant neighbors and improve its performance in various tasks.
[0097] · LGCN adopts a different approach, converting graph data into a grid-like structure and applying traditional CNN models. This method enables the use of standard convolutional layers on graph data, thus achieving efficient feature extraction and classification.
[0098] GNAS methods:
[0099] · GraphNAS uses reinforcement learning to generate variable-length strings describing GNN architectures. It searches for the optimal architecture by maximizing a reward signal based on model performance.
[0100] · Auto-GNN is another reinforcement learning method similar to GraphNAS, but with a parameter sharing strategy to reduce computational costs.
[0101] · Genetic-GNN uses an evolutionary algorithm as a search strategy for GNN model structures and hyperparameters. This method iteratively evolves a population of models by selecting the best-performing models and generating new models through mutation and crossover. This process leads to the discovery of high-performance GNN architectures customized for specific tasks.
[0102] To verify the effectiveness of the TPE-based fusion strategy, TPE was compared with the following hyperparameter optimization methods.
[0103] Hyperparameter optimization methods:
[0104] · Random search involves randomly selecting parameter combinations from the hyperparameter space for experiments and recording the performance of each combination to determine the best parameter set.
[0105] · Genetic algorithms simulate the process of natural evolution and use operations such as mutation, crossover, and selection to iteratively optimize hyperparameter combinations.
[0106] · Particle swarm optimization mimics the behavior of a flock of birds or fish searching for food, where each particle represents a hyperparameter combination and iteratively adjusts itself to find the optimal solution.
[0107] Experimental results
[0108] Overall results
[0109] Ensemble-GNAS was compared with handcrafted GNNs and other GNAS methods. Table IV summarizes the overall performance comparison in terms of test set accuracy.
[0110] Table IV Performance comparison
[0111]
[0112] The results of Ensemble-GNAS are the average of 10 independent runs for each dataset. By comparing the results in the table, it can be seen that Ensemble-GNAS achieves better performance compared to the manual graph neural network architectures. This is because the manual models require manual adjustment of the GNN structure for different datasets, making it difficult to obtain the optimal structure. It is worth noting that on the citation network, Ensemble-GNAS is competitive compared to other GNAS methods, while on the cancer driver gene identification network, Ensemble-GNAS has a significant advantage. The experimental results show that Ensemble-GNAS has superior performance in dealing with large-scale datasets. This is because the performance of a single network may not have sufficient expressive power in large-scale problems, while Ensemble-GNAS fully utilizes the information provided by multiple candidate networks and thus has better potential in dealing with large-scale problems.
[0113] Effectiveness of the niche strategy
[0114] To verify the effectiveness of the proposed search strategy, ablation experiments were designed to study the impact of whether the search strategy includes the niche strategy on performance.
[0115] Figure 4 The output differences between the two algorithms on the validation set are shown. OD represents the diversity of the candidate networks output by the two algorithms. As can be seen from the figure, without the niche strategy, as the number of generations increases, the output diversity gradually decreases. As the algorithm tends to converge, the candidate networks start to produce increasingly similar outputs. This decreasing variation in the network outputs may harm the efficacy of the ensemble because different output sets are usually crucial for enhancing the overall performance and robustness of the system. In contrast, the output diversity of the candidate networks generated based on the niche search strategy remains at a high level. Compared with the candidate networks with poor diversity, such candidate networks may have higher EL potential.
[0116] The best individuals and ensemble results of the two algorithms are shown in Table V. As can be seen from the table, compared with not using the niche strategy, the performance advantage of using the niche strategy to search for the best network architecture is not significant. This is because the search strategy based on the niche strategy searches for network architectures in multiple local regions simultaneously, while the search strategy without using the niche strategy may get stuck in a single local region. With the same number of samplings, the search strategy without using the niche strategy may be more likely to find a local optimal solution, and the EL of the candidate networks using the niche strategy can significantly improve performance.
[0117] Table V Performance comparison of search strategies
[0118]
[0119]
[0120] As can be seen from the table, without using the niche strategy, the performance of the network searched by the search strategy did not show significant changes before and after EL. This is because the diversity of candidate networks is poor, and EL cannot extract different information from multiple candidate networks. On the other hand, the search strategy using the niche strategy has better diversity of candidate networks, so it shows higher integration potential. This also confirms the previous conclusion that the better the diversity of candidate networks, the better the integration result may be.
[0121] Effectiveness of the TPE strategy
[0122] To verify the performance of the TPE-based fusion strategy, TPE was compared with random search (RS), genetic algorithm (GA), and particle swarm optimization (PSO). As shown in Table VI, in terms of average test set accuracy and stability, TPE outperforms the comparison algorithms. The highest performance is highlighted in bold. In summary, the experimental results prove the rationality and effectiveness of using TPE to optimize the integration weights.
[0123] Table VI Performance comparison of fusion strategies
[0124]
[0125] By maintaining the exploration and exploration models, TPE achieves a balance between known good regions and new regions. This balance strategy helps the algorithm avoid premature convergence while gradually focusing on potential optimal regions. Although evolutionary algorithms also attempt to maintain population diversity through mutation and crossover (exploration) and utilize the current best solution through selection pressure, this balance may not be as refined and efficient as the explicit modeling in TPE. And TPE can adaptively adjust the search strategy according to the specific situation of the problem, while GA and PSO usually require manual adjustment or preset strategies, which may limit their applicability and flexibility.
[0126] Efficiency comparison
[0127] Compared with the more complex and resource-intensive architecture search part, the time and computing resources required for the EL part of the algorithm are minimal. This means that the EL part of the algorithm is negligible in terms of efficiency compared to the architecture search part. The search efficiency of Ensemble-GNAS was studied by comparing the accuracy of its top 10 validation sets with other methods such as random search, GraphNAS, and Genetic-GNN. To ensure a fair comparison, 2,000 GNN architectures of each method were studied in the same search space. The settings of Ensemble-GNAS are detailed in this experiment, while the configurations of GraphNAS and Genetic-GNN follow their respective papers.
[0128] As Figure 5 shown, Ensemble-GNAS outperforms other methods in efficiently sampling GNN architectures with good performance. This means that Ensemble-GNAS can identify high-performance GNN architectures more effectively than other methods, thus obtaining better model performance. Because the search strategy based on the niche strategy enables different subpopulations to search independently, avoiding the entire population concentrating in a specific area. The decentralized search method can explore multiple potential optimal solutions in parallel, improving the search efficiency.
[0129] Multi-task performance
[0130] To demonstrate the generality of Ensemble-GNAS, link prediction and graph classification experiments were conducted to compare the performance of Ensemble-GNAS with RS, GraphNAS, and Genetic-GNN in terms of test set accuracy. Graph classification is a task that involves predicting the class of unseen graphs. This is very useful in various applications such as social network analysis, bioinformatics, and recommendation systems. The goal is to learn a model from labeled graph data and apply it to new unlabeled graphs to predict their classes. On the other hand, link prediction is a different task that focuses on predicting whether there is an edge between two nodes in a graph. This is useful in recommendation systems, where the goal is to suggest potential connections between users or items based on their interactions. Link prediction can also be applied to social networks to suggest potential friends or followers for users.
[0131] When preparing a dataset for the graph classification task, the dataset is usually divided into a training set, a validation set, and a test set. In this case, 60% of the data is used for training, 20% for validation, and 20% for testing. This division ensures that the model has enough data to learn patterns and generalizes well to unseen data. For the link prediction task, a balanced set of positive and negative edges is usually used. Positive edges represent actual connections between two nodes, while negative edges represent non-existent connections. In this case, 20% of the data is used for validation and another 20% for testing. This division helps to evaluate the performance of the model in predicting true and false links. The performance results are shown in Table VII.
[0132] Table VII Performance Comparison of Graph Classification and Link Prediction
[0133]
[0134] The experimental results show that Ensemble-GNAS can achieve competitive performance on different graph tasks compared with other search algorithms. EL reduces the risk of overfitting and the impact of a single outlier on the ensemble model by combining predictions from multiple models. It also takes advantage of the strengths of different models on different data subsets to improve the overall performance.
[0135] Sensitivity Analysis of Search Parameters
[0136] As Figure 6 shown, the sensitivity of some parameters used in Ensemble-GNAS is analyzed. (i) C is the number of candidate networks participating in EL. By default, the candidate networks generated in each niche will participate in EL. As C increases, the performance of the ensemble model gradually improves, but the size of the model also increases. Therefore, on a small-scale dataset, a larger C can be maintained to obtain the best ensemble. On a large-scale dataset, C can be appropriately reduced to balance the performance and size of the ensemble model. (ii) N is the number of individuals in each niche. As N increases, the performance of the ensemble network does not always improve because an overly large N may cause the search process not to converge completely within the set number of generations. G is the number of generations of the algorithm. As G increases, the overall performance of the algorithm shows an upward trend, but as the algorithm gradually converges, the overall performance tends to stabilize.
[0137] This embodiment proposes an integration method based on evolutionary GNAS, called Ensemble-GNAS. Ensemble-GNAS enhances the potential of EL by using a search algorithm based on the niche strategy to improve the diversity of candidate networks. When optimizing the ensemble weights, Ensemble-GNAS uses TPE to improve the algorithm's ability to handle high-dimensional continuous problems. Experimental results show that Ensemble-GNAS has an advantage in classification accuracy. In addition, ablation experiments are designed to demonstrate the effectiveness of the Ensemble-GNAS strategy.
[0138] Due to algorithm limitations, GNNs of different scales will be used as candidate networks for the ensemble, and the average complexity of the candidate networks is relatively high. Therefore, designing a multi-objective strategy to balance the complexity and accuracy of candidate networks is future work.
[0139] There are several high-order and complex data structures that require specialized methods for analysis. One such structure is the hypergraph, which is a generalization of a graph where an edge can connect any number of nodes. Studying how to optimize the search and integration of GNNs on hypergraphs is extremely important for unleashing the full potential of these powerful models in solving real-world challenges.
[0140] The above-described embodiments merely represent the specific implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. An ensemble learning method based on evolutionary graph neural architecture search, characterized in that: The following steps are involved: Step 1: Design an ensemble model of the ensemble learning EL method based on evolutionary graph neural architecture search (GNAS), denoted as Ensemble-GNAS, to perform graph learning by retaining and integrating multiple solutions with different architectures but similar performance. Step 2: Use the evolutionary graph neural architecture search (GNAS) method based on the niche strategy to improve the local search ability of the algorithm and the diversity of the ensemble learning EL candidate network; Step 3: Use the ensemble fusion strategy based on the Tree Parzen Estimator (TPE) to optimize the weights of the base learners in the ensemble model.
2. The ensemble learning method based on evolutionary graph neural architecture search according to claim 1 is characterized in that: In step 1, in the architecture search phase, given a graph neural architecture S a The search space, niche size N and graph data D are used to perform evolutionary graph neural architecture search (GNAS) to find the best individuals in each niche. is the best individual in niche i, which maximizes the validation set D val The expected accuracy Acc val (X), that is: After the GNAS search algorithm, the M best individuals are obtained as the TPE-based EL candidate network; the final integrated output is shown as: where out i represents the output of the i-th candidate network, w i represents the weights optimized by TPE.
3. The ensemble learning method based on evolutionary graph neural architecture search according to claim 2 is characterized in that: In step 2, the evolutionary graph neural architecture search (GNAS) method based on the niche strategy is as follows: The evolution process includes crossover and mutation. Uniform crossover is used to generate new offspring. d The d-th dimension element of is as follows: Where X d and Y d represents the d-th element of the two parents, R d Represents the d-th dimension element of a random binary array with the same dimension as the parent; the mutation operation randomly generates a new individual; When evaluating individuals, first train the individual network architecture on the training set, and then verify the trained network on the validation set; the jth individual X i,j The fitness in the i-th niche is as follows: fitess(X i,j )=Acc val (X i,j ) (4) The dataset is divided into training set, validation set and test set, then the population is initialized based on the search space, and the individuals in the population are divided into different microhabitats; each individual in the population is evaluated and their fitness values are recorded; during the architectural evolution process, each individual crosses with another individual to produce offspring. If a mutation occurs, the offspring will become a completely new individual. The fitness values of the parent and offspring are compared, and the excellent individuals are retained.
4. The ensemble learning method based on evolutionary graph neural architecture search according to claim 3 is characterized in that: In step 2, the expression of the diversity OD of the EL candidate network for ensemble learning is as follows: in is the total number of samples used for verification, n is the number of candidate networks; for each sample s in the verification set, the outputs generated by any two different networks i and j are represented as O i (s) and O j (s).
5. The ensemble learning method based on evolutionary graph neural architecture search according to claim 3 or 4, characterized in that: Also includes: Network architectures with the same number of hidden units in the population are grouped into the same niche.
6. The ensemble learning method based on evolutionary graph neural architecture search according to claim 5, characterized in that: Also includes: According to the search space, the GNN architecture uses a stacked structure to design network layers. Each layer consists of four components, which together affect the performance and output of the network. The GNN architecture is described as an ordered list, where each element represents the value of an architecture component.
7. The ensemble learning method based on evolutionary graph neural architecture search according to claim 5, characterized in that: When the population is initialized, it is divided into M small habitats, and each individual in a single small habitat has a specific number of hidden units; the number of individuals in each small habitat is N, that is, the number of individuals in the entire population is MN.
8. The ensemble learning method based on evolutionary graph neural architecture search according to claim 7 is characterized in that: In step 3, the ensemble fusion strategy based on the tree-shaped Parzen estimator TPE is as follows: TPE is based on the weighted search space S w Initialize the probability model; then, in each iteration, generate candidate parameter configurations c based on the existing probability model, evaluate the performance of the candidate parameter configurations and update the model, and continue until the preset number of iterations is reached or other stopping conditions are met; The weight of the candidate network is used as a parameter, ranging from 0 to 1; the validation accuracy is used as the objective function, and the optimal weight c is obtained by using TPE. * ; Use the best weight c * The candidate networks are weighted and integrated to obtain an integrated model; finally, the integrated model is applied to the test set to obtain the final result.
9. An application of an integrated learning method based on evolutionary graph neural architecture search as described in any one of claims 1 to 8 in graph mining.
Citation Information
Cited By
Link prediction method, system and device of graph network and storage medium
CN121682148A