Graph neural network architecture optimization method based on grammar genetic programming

Through the method based on grammatical genetic programming and Gaussian process proxy model, GNN architecture design is optimized, and the problems of GNN architecture design in the existing technology relying on expert experience, low optimization efficiency, and insufficient generalization capabilities are insufficient, and efficient, automated, and knowledge-driven GNN architecture optimization is achieved, improving the performance and generalization capabilities of the model.

CN120046701AActive Publication Date: 2025-05-27SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202510444959.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-27
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the prior art, GNN architecture design relies on expert experience, has low optimization efficiency and insufficient generalization capabilities, especially in complex application scenarios such as drug discovery, bioinformatics and financial risk control, it is difficult to meet the needs of high performance and high generalization.

Method used

A grammatical genetic programming method is used to construct a grammatical rule library of GNN architecture, and the initial GNN architecture population is generated through genetic programming, and the Gaussian process proxy model is used to predict the performance of individual GNN architecture in the population. In combination with genetic algorithms, search for the optimal GNN architecture in the syntax tree space, guide the search process through expected improvement functions, and realize knowledge-driven efficient GNN architecture optimization.

Benefits of technology

It significantly improves the automation level and efficiency of GNN architecture design, reduces the dependence on expert experience, improves the performance and generalization capabilities of the model, and is suitable for complex application scenarios such as drug discovery, bioinformatics and financial risk control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046701A_ABST
    Figure CN120046701A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of graph neural networks, in particular to a graph neural network architecture optimization method based on grammar genetic programming, and the method comprises the steps: S1, constructing a grammar rule library of a graph neural network architecture GNN; s2, based on the grammar rule base, generating an initial GNN architecture population through genetic programming; s3, predicting the performance of the individual GNN architecture in the population by using a Gaussian process proxy model; s4, searching an optimal GNN architecture in the syntax tree space through a genetic algorithm; s5, for the individual GNN architecture in the population, calculating a training probability based on an expected improvement function of the individual GNN architecture, if the training probability is higher than a preset evaluation threshold, selecting the individual GNN architecture to carry out actual training evaluation, obtaining the authenticity performance of the individual GNN architecture, and updating the Gaussian process agent model by using the authenticity performance; and S6, judging whether a preset optimization termination condition is met or not, if so, outputting the optimized GNN architecture, and otherwise, returning to the step S3. According to the method, the automation degree, efficiency and model performance of GNN model architecture search can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph neural networks, and more specifically, to a method for optimizing a graph neural network architecture based on grammar genetic programming. Background Art

[0002] Graph Neural Networks (GNNs), as an emerging deep learning model, have shown powerful capabilities in processing graph-structured data. By effectively combining graph computing and neural networks, GNNs have achieved remarkable results in many fields such as node classification, graph classification, link prediction, recommendation systems, risk control, etc.

[0003] However, the practical application and wide deployment of GNNs still face many challenges. First, designing a high-performance GNN architecture is a time-consuming process that relies on expert experience, especially in fields that require dealing with complex relationships and fine-grained patterns, such as predicting molecular activity in drug discovery, analyzing protein interaction networks in bioinformatics, or identifying complex fraud rings in financial risk control. Traditional manual design methods require researchers to repeatedly debug and optimize the network structure and hyperparameters according to specific problems, which is inefficient and difficult to guarantee finding the optimal architecture. Second, the hyperparameter optimization process of GNN architectures is complex. Manually adjusting hyperparameters is not only time-consuming but also prone to falling into local optima, making it difficult to fully exploit the performance potential of GNNs. In addition, existing automated graph machine learning techniques generally lack knowledge-driven guidance in the GNN architecture search process, resulting in a large search space, long search time, low optimization efficiency, and being prone to overfitting to specific datasets, with insufficient model generalization ability. This is unacceptable in key applications such as drug research and development, precision medicine, and financial security that require high-precision prediction and reliability.

[0004] In recent years, Neural Architecture Search (NAS) techniques have provided new ideas for automating the design of neural network architectures. In the field of graph neural networks, some NAS methods have also emerged. However, these methods still have limitations in terms of search efficiency and generalization ability. For example, NAS methods based on reinforcement learning have an unstable training process and low search efficiency; although NAS methods based on evolutionary algorithms can find architectures with better performance, they lack effective guidance for the search process, are prone to blind search, and have a high computational cost.

[0005] Therefore, there is an urgent need for a method that can efficiently, automatically, and knowledge-drivenly optimize GNN architectures to reduce the dependence on expert experience, shorten the R & D cycle, and improve the performance and generalization ability of GNN models. Summary of the Invention

[0006] In view of this, the present invention provides an optimization method for a graph neural network architecture based on grammar genetic programming, which can solve the problems in the prior art that the design of the GNN architecture depends on expert experience, the optimization efficiency is low, and the generalization ability is insufficient. Especially in complex application scenarios such as drug discovery, bioinformatics, and financial risk control that require high-performance graph neural network models, the automation degree, efficiency, and model performance of GNN model architecture search are improved.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] An optimization method for a graph neural network architecture based on grammar genetic programming, comprising the following steps:

[0009] S1. Construct a grammar rule library for the graph neural network architecture GNN;

[0010] S2. Based on the grammar rule library, generate an initial GNN architecture population through genetic programming. Each individual GNN architecture in the population is represented as a grammar tree, and the grammar tree is recursively generated according to production rules, and the network structure and hyperparameters of each layer of the network are encoded synchronously;

[0011] S3. Use a Gaussian process surrogate model to predict the performance of individual GNN architectures in the population;

[0012] S4. Based on the predicted performance of individual GNN architectures, search for the optimal GNN architecture in the grammar tree space through genetic algorithms;

[0013] S5. For the individual GNN architectures in the population, calculate the training probability P train (x) based on their expected improvement function EI(x). If the training probability P train (x) is higher than a preset evaluation threshold, then select this individual GNN architecture for actual training evaluation, obtain its true performance, and update the Gaussian process surrogate model with the true performance;

[0014] S6. Determine whether the preset optimization termination condition is satisfied. If so, output the optimized GNN architecture; otherwise, return to step S3. The optimized GNN architecture is used to process graph-structured data to perform classification tasks.

[0015] Further, in S1, the grammar rule library is defined based on the BNF paradigm, and includes a non-terminal symbol set N, a terminal symbol set T, a production rule set P, and a start symbol S;

[0016] Among them, the non-terminal symbol set N contains extensible network modules, which are divided into a hyperparameter non-terminal symbol set N H and a module non-terminal symbol set N M ;

[0017] The set of terminals \(T\) contains specific network operators and hyperparameter values, which are divided into the hyperparameter terminal set \(T\) H and the module terminal set \(T\) M ;

[0018] The production rule set \(P\) defines the derivation relationship from non-terminals to terminals and / or non-terminals;

[0019] The start symbol \(S\) represents the starting derivation symbol of the network architecture.

[0020] Furthermore, each production rule \(p\in P\) in the production rule set \(P\) is in the form of \(l\rightarrow r\), where \(l\in N\) is the symbol to be derived, and \(r = r\) H r M is the derivation result, and \(r\) H \(\in(N\) H \cup T\) H )^*\) represents the hyperparameter part, and \(r\) M \(\in(N\) M \cup T\) M )\) * represents the sub-module part, is the hyperparameter non-terminal set, is the hyperparameter terminal set, is the module non-terminal set, is the module terminal set.

[0021] Furthermore, in \(S2\), the process of generating the syntax tree includes:

[0022] Initialization: Create an initial syntax tree with the start symbol \(S\) as the root node;

[0023] Recursive expansion: For each non-terminal leaf node \(l\) k \(\in N\cup S\) in the current syntax tree, randomly select a production rule from the production rule subset \(P(l\) k )=\{p:l\) k \rightarrow r\) k ,p\in P\}\) for expansion, where \(r\) k represents the right part of the production rule and is a sequence composed of hyperparameter symbols and module symbols;

[0024] Termination condition: Repeat the recursive expansion step until all leaf nodes of the syntax tree are terminals.

[0025] Furthermore, \(S3\) includes:

[0026] S31. Define the Gaussian process surrogate model based on the following formula:

[0027] f\sim GP(m(x),k(x,x'))

[0028] Among them, f represents the fitness function, m(x) represents the mean function, k(x, x') represents the covariance function, and the mean function is defined and calculated based on the following formula:

[0029]

[0030] Among them, represents the expected value; f(x) represents the random function in the Gaussian process surrogate model, which is used to approximate the true fitness function;

[0031] The covariance function k(x, x') is calculated based on the tree kernel function defined as follows:

[0032]

[0033] Among them, D Tree (x, x') represents the tree edit distance between the syntax trees x and x’, and σ is the kernel function width parameter;

[0034] S32. Train the Gaussian process surrogate model based on the evaluated training data (X, y), where X is the set of evaluated GNN architectures and y is the corresponding true performance;

[0035] S33. Use the trained Gaussian process surrogate model to predict the fitness function f * of the new unevaluated GNN architecture X * The distribution of, and the result follows a normal distribution:

[0036]

[0037] Among them, μ * is the predicted mean; σ * is the predicted standard deviation, indicating the uncertainty of the prediction.

[0038] Furthermore, in S4, the genetic algorithm includes: selection, crossover, and mutation operations;

[0039] Among them, the selection operation adopts the tournament selection method. Each time, a certain number of individuals are randomly selected from the population for comparison, and the individual with the optimal predicted performance is selected as the Parent. The selection operation is repeated until the number of Parent individuals meets the requirements of crossover and mutation;

[0040] The crossover operation includes: randomly selecting a non-terminal node from the syntax trees of two Parent individuals respectively, ensuring that the non-terminal types of the two nodes are the same; swapping the subtrees with the two selected nodes as the roots to generate the syntax trees of two new offspring individuals;

[0041] The mutation operation includes: randomly selecting a non-terminal node in the individual's syntax tree; randomly generating a new subtree according to the production rule with this non-terminal as the left side in the syntax rule library; replacing the selected node with the newly generated subtree to complete the mutation operation.

[0042] Furthermore, in S5, the expected improvement function EI(x) is used to evaluate the potential improvement value of the individual GNN architecture, and its calculation result is used to determine the training probability P train (x), and is also used to guide the search direction of the genetic algorithm or select among multiple individuals that meet the evaluation conditions; among them, guiding the search direction of the genetic algorithm includes: giving priority to individuals with higher improvement value in the selection operation of the genetic algorithm, or selecting the individual with the highest improvement value from multiple GNN architecture individuals that meet the actual training evaluation conditions for actual training evaluation;

[0043] The expected improvement function EI(x) is defined based on the following formula:

[0044]

[0045] where Imp represents the improvement amount, Z represents the standardized improvement amount, and Imp = μ(x) - y best -ξ, μ(x) and σ(x) are respectively the mean and standard deviation predicted by the Gaussian process surrogate model, and y best is the optimal performance in the currently evaluated GNN architectures, ξ is the exploration factor, Φ and are respectively the cumulative distribution function and probability density function of the standard normal distribution.

[0046] Furthermore, S5 also includes:

[0047] Calculating the prediction probability P predict (x) and the training probability P train (x), and determining whether to conduct actual training evaluation on the individual GNN architecture according to the training probability P train (x), where the prediction probability P predict (x) and the training probability P train (x) are calculated according to the following formulas:

[0048] P predict (x) = max(P 0 , min(1, EI(x)))

[0049] P train (x) = 1 - P predict (x)

[0050] where P 0 is the lower bound of the prediction probability. If the training probability P trainIf (x) is higher than the preset value, then the individual GNN architecture x is actually trained and evaluated to obtain its true performance y on the validation set, and the Gaussian process surrogate model is updated using the new data point (x, y), that is, the new data point is added to the evaluated data set, and the Gaussian process surrogate model is retrained.

[0051] Further, in S6, the optimization termination conditions include at least one of the following:

[0052] The number of iterations of the genetic algorithm reaches the preset maximum number of iterations;

[0053] The improvement amplitude of the predicted performance of the population-optimal individual GNN architecture is lower than the preset threshold within a continuous preset number of generations;

[0054] The preset upper limit of computing resources is reached.

[0055] Further, in S6, the optimized GNN architecture is used for node classification in a citation network or for chemical compound graph classification to predict molecular properties; in the node classification task in the citation network, the input is the graph structure data of the citation network, the nodes represent documents, the edges represent citation relationships, and the output is the class label of each node; in the graph classification task of chemical compounds, the input is a set of graph structure data representing chemical compounds, the nodes represent atoms, the edges represent chemical bonds, and the output is the class label of each graph (for example, indicating whether the compound has specific bioactivity or toxicity properties).

[0056] As can be seen from the above technical solutions, compared with the prior art, the present invention has the following beneficial effects:

[0057] 1. By constructing a grammar rule library, the present invention realizes the structured representation and automatic generation of GNN architectures, reduces the dependence on expert experience, and improves the automation degree of GNN architecture design.

[0058] 2. The present invention introduces a Gaussian process surrogate model for GNN architecture performance prediction, avoids expensive actual training evaluations for each candidate architecture, significantly improves the architecture search efficiency, and shortens the optimization time.

[0059] 3. The present invention balances the exploration and exploitation of the search process through the expected improvement (EI) function, effectively guides the genetic algorithm to perform efficient search in the GNN architecture space, realizes knowledge-driven efficient GNN architecture optimization, accelerates the discovery of the optimal GNN architecture, and improves the quality of the search results, which enables the optimized GNN model to show better performance in key tasks such as chemical molecular property prediction (such as toxicity or activity prediction in graph classification tasks), biological network analysis (such as protein function prediction in node classification tasks), etc.

[0060] 4. The present invention accelerates the evaluation through automated search and surrogate models, significantly shortening the R & D cycle for customizing and deploying high-performance GNN models for specific complex applications (such as new drug screening, precision medical analysis, dynamic financial risk control), effectively reducing the dependence on computing resources and expert experience, and improving the generalization ability and application value of the models on real-world complex graph data. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.

[0062] Figure 1 Flowchart of the graph neural network architecture optimization method based on grammar genetic programming provided by the present invention;

[0063] Figure 2 Flowchart of evaluating the fitness of an individual GNN architecture based on a Gaussian process surrogate model provided by the present invention;

[0064] Figure 3 Example chart of grammar rules used in the GNN architecture search space provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0066] As Figure 1 shown, the embodiments of the present invention disclose a graph neural network architecture optimization method based on grammar genetic programming, which is characterized by including the following steps:

[0067] S1. Construct a grammar rule library for the graph neural network architecture GNN;

[0068] S2. Based on the grammar rule library, generate an initial population of GNN architectures through genetic programming. Each individual GNN architecture in the population is represented as a grammar tree, which is recursively generated according to production rules and synchronously encodes the network structure and hyperparameters of each layer of the network;

[0069] S3. Use a Gaussian process surrogate model to predict the performance of individual GNN architectures in the population;

[0070] S4. Search for the optimal GNN architecture in the syntax tree space through a genetic algorithm based on the prediction performance of the individual GNN architecture;

[0071] S5. For the individual GNN architectures in the population, calculate the training probability P train (x) based on their expected improvement function EI(x). If the training probability P train (x) is higher than the preset evaluation threshold, select this individual GNN architecture for actual training evaluation, obtain its true performance, and update the Gaussian process surrogate model using the true performance. Among them, in addition to being used to calculate P train (x), the expected improvement function EI(x) can also assist in guiding the search direction of the genetic algorithm or making a selection among multiple individuals that meet the evaluation conditions; the relative uncertainty index P train (x) comprehensively reflects the prediction performance improvement potential and prediction uncertainty of architecture x, and is calculated by the Gaussian process model based on the evaluated data and the expected improvement function EI(x).

[0072] S6. Determine whether the preset optimization termination condition is met. If so, output the optimized GNN architecture; otherwise, return to step S3. The optimized GNN architecture is used for graph data classification tasks.

[0073] The core idea of the present invention is: construct a syntax rule library for the GNN architecture, use grammar-based genetic programming (GBGP) to automatically generate the syntax tree representation of the GNN network architecture, and synchronously encode the network structure and hyperparameters. To improve the search efficiency, the present invention introduces a Gaussian process (GP) surrogate model to predict the performance of the GNN architecture, and combines a genetic algorithm to efficiently search for the optimal architecture in the syntax tree space. By balancing the exploration and exploitation of the search process through the expected improvement (EI) function, knowledge-driven efficient optimization of the GNN architecture is achieved.

[0074] Next, the above steps will be further described.

[0075] S1. Construct a syntax rule library for the graph neural network architecture GNN.

[0076] The syntax rule library is defined based on the BNF paradigm and includes a non-terminal symbol set N, a terminal symbol set T, a production rule set P, and a start symbol S, which is formally represented as a quadruple G = (N, T, P, S).

[0077] Among them, the non-terminal symbol set N contains extensible network modules, such as <net>(Network module), <gconv>(Graph Convolution Module), <residual>(Residual module), <attention>(Attention module), <fc>(Fully connected layer), <conv>(Convolutional layer), <gat>(Graph Attention Layer), <act>(Activation function), etc. The non-terminal symbol set is further divided into the hyperparameter non-terminal symbol set N H (such as <dim> 、、 <k>) and the set of module non-terminals N M (e.g. <net> 、 <gconv> 、 <fc>)。

[0078] The terminal symbol set T contains specific network operators and hyperparameter values, such as GCNConvLayer, ChebConvLayer, GATConvLayer, Relu, Sigmoid, Tanh, and specific values of hyperparameters (such as 16, 32, 64, 128, 0.4, 0.5, 0.6, 1, 2, 4, 1, 2, 3), etc. The terminal symbol set is further divided into the hyperparameter terminal symbol set T H (such as 16, 32, 64, 128, 1, 2, 4, 1, 2, 3, 0.4, 0.5, 0.6) and the module terminal symbol set T M (such as GCNConvLayer, Relu).

[0079] The production rule set P defines the derivation relationship from non-terminal symbols to terminal symbols and / or non-terminal symbols; each production rule p ∈ P in the production rule set P has the form l → r, where l ∈ N is the symbol to be derived, and r = r H r M is the derivation result, and r H ∈ (N H ∪ T H ) * represents the hyperparameter part, and r M ∈ (N M ∪ T M ) * represents the sub-module part, is the hyperparameter non-terminal symbol set, is the hyperparameter terminal symbol set, is the module non-terminal symbol set, is the module terminal symbol set. For example, the production rule of the graph convolution module can be <gconv>→ <dim> <conv> <fc>, indicating that a graph convolutional module consists of a hidden layer dimension <dim>, Convolutional layer <conv>and fully connected layer <fc>Composition.

[0080] The start symbol S represents the start derivation symbol of the network architecture, for example <net>, representing the starting derivation symbol of the network architecture.

[0081] The grammar rule library constructed by the present invention has good flexibility and scalability. Users can conveniently add new network modules (non-terminals), specific operations (terminals) or modify production rules according to prior knowledge or specific application requirements, so as to guide the search space to tilt towards a more expected direction, further combining expert experience with automated search.

[0082] S2. Based on the grammar rule library, generate an initial population of GNN architectures through genetic programming. Each individual GNN architecture in the population is represented as a syntax tree, and the syntax tree is recursively generated according to production rules, and the network structure and hyperparameters of each layer of the network are encoded synchronously. The generation process of the syntax tree includes representing each individual GNN architecture in the population as a syntax tree, and the generation process of the syntax tree includes:

[0083] Initialization: Create an initial syntax tree with the starting symbol S as the root node;

[0084] Recursive expansion: For each non-terminal leaf node l in the current syntax tree k ∈N∪S, randomly select a production rule from the subset of production rules P(l k ) = {p: l k →r k , p∈P} for expansion, where r k represents the right part of the production rule, refer to Figure 3 , which is a sequence composed of hyperparameter symbols and module symbols. For example, if the current node is <gconv>, then from P(· <gconv>)Randomly select a rule, for example <gconv>\rightarrow <dim> <conv> <fc>, will <gconv>The node is expanded to <dim> 、 <conv>And <fc>The subtree of the child node.

[0085] Termination condition: Repeat the recursive expansion step until all leaf nodes of the syntax tree are terminals.

[0086] Through the above process, each individual GNN architecture is encoded as a syntax tree, which not only describes the structure of the network but also contains the hyperparameter information of each layer of the network.

[0087] S3. Use the Gaussian process surrogate model to predict the performance of individual GNN architectures in the population. The specific process is as Figure 2 shown, including:

[0088] S31. To avoid the huge computational overhead caused by actual training and evaluation for each GNN architecture, the present invention introduces a Gaussian process surrogate model to predict the performance (fitness) of the GNN architecture. The Gaussian Process (GP) surrogate model is defined based on formula (1):

[0089] f ∼ GP(m(x), k(x, x')) (1);

[0090] where f represents the fitness function (e.g., the classification accuracy of the GNN architecture on the validation set), m(x) represents the mean function, and k(x, x') represents the covariance function, which is used to measure the similarity between two GNN architectures x and x'. The mean function is defined and calculated based on the following formula:

[0091]

[0092] where represents the Expected Value, is the expected value of the fitness function f(x); f(x) represents the random function in the Gaussian process surrogate model, which is used to approximate the true fitness function, such as the classification accuracy of the GNN architecture on the validation set;

[0093] The covariance function k(x, x') is calculated based on the tree kernel function defined as follows:

[0094]

[0095] where D Tree (x, x') represents the tree edit distance between the syntax trees x and x’, and σ is the kernel function width parameter, which controls the radial basis function width of the kernel function. The tree edit distance D Tree (x, x') is calculated based on the minimum edit operation cost. The edit operations include node insertion, node deletion, and node replacement, and are calculated using an existing tree edit distance algorithm (such as the ZSS algorithm).

[0096] S32. Train the Gaussian process surrogate model based on the evaluated training data (X, y), where X is the set of evaluated GNN architectures and y is the corresponding true performance (validation set accuracy).

[0097] S33. Use the trained Gaussian process surrogate model to predict the fitness function f of a new unevaluated GNN architecture X * of * which follows a normal distribution:

[0098]

[0099] where μ * is the predicted mean; σ * is the predicted standard deviation, representing the uncertainty of the prediction.

[0100] Use the covariance matrix ∑ of the training dataset, the covariance matrix ∑ between the training set and the test set * , and the autocovariance matrix ∑ of the test set ** . To improve computational stability, perform a Cholesky decomposition on the covariance matrix ∑ as shown in formula (5):

[0101] ∑ = LL T (5);

[0102] where L represents the lower triangular matrix obtained from the Cholesky decomposition of the covariance matrix ∑. Here, ∑ is the covariance matrix of the training set in the Gaussian process model, and the lower triangular property of L ensures the uniqueness and computational stability of the decomposition.

[0103] The predicted mean μ * represents the predicted performance value, and the predicted variance reflects the uncertainty of the prediction. Among them, the predicted mean μ * and the predicted variance are calculated through the following formulas:

[0104] μ * = ∑ * L -T L -1 y (6);

[0105]

[0106] Based on a small number of evaluated GNN architectures and their true performance data, train the Gaussian process surrogate model. For the unevaluated individual GNN architectures in the population, use the trained Gaussian process surrogate model to predict their performance and output the predicted mean μ(x) and the predicted variance σ 2 (x). The predicted variance σ 2 (x) characterizes the uncertainty of performance prediction. The greater the variance, the higher the uncertainty.

[0107] S4. Based on the prediction performance of the individual GNN architecture, search for the optimal GNN architecture in the syntax tree space through the genetic algorithm. The genetic algorithm includes: selection, crossover, and mutation operations;

[0108] Among them, the selection operation adopts the tournament selection method. Each time, a certain number (for example, 3) of individuals are randomly selected from the population for comparison, and the individual with the optimal prediction performance (mean μ(x)) is selected as the Parent. The selection operation is repeated until the number of Parent individuals meets the requirements of crossover and mutation.

[0109] The crossover operation includes: subtree crossover operation. Randomly select a non-terminal node from the syntax trees of two Parent individuals respectively, ensuring that the non-terminal types of the two nodes are the same (for example, both are <conv>) Exchange the subtrees rooted at two selected nodes to generate two new offspring individual syntax trees.

[0110] The mutation operation includes: randomly selecting a non-terminal node in the individual syntax tree; randomly generating a new subtree according to the production rule with this non-terminal as the left side in the syntax rule library; replacing the selected node with the newly generated subtree to complete the mutation operation.

[0111] S5. For the individual GNN architectures in the population, calculate the training probability P train (x) based on their expected improvement function EI(x); if the training probability P train (x) is higher than the preset evaluation threshold, then select this individual GNN architecture for actual training evaluation, obtain its true performance, and use this true performance to update the Gaussian process surrogate model; where the expected improvement function EI(x) is used to calculate P train (x) in addition to assisting in guiding the search direction of the genetic algorithm or selecting among multiple individuals that meet the evaluation conditions;

[0112] For the determination of the preset evaluation threshold, through cross-validation or preliminary experiments, 0.5 or 0.6 is a reasonable initial trial value and medium threshold.

[0113] Specifically, in order to balance the exploration and exploitation of the search process, the present invention uses the expected improvement (EI) function to assist in guiding the search direction of the genetic algorithm or selecting among multiple individuals that meet the evaluation conditions. The expected improvement function EI(x) is defined based on the following formula:

[0114]

[0115] where EI(x) represents the expected improvement, used to evaluate the possible performance improvement by sampling at the point x

[0116] , Imp represents the improvement amount, Z represents the standardized improvement amount, Imp = μ(x) - y best - ξ, μ(x) and σ(x) are respectively the mean and standard deviation predicted by the Gaussian process surrogate model, y best is the optimal performance in the currently evaluated GNN architectures, ξ is the exploration factor (used to balance exploration and exploitation, for example, ξ = 0), Φ and They are the cumulative distribution function and probability density function of the standard normal distribution respectively. The value calculated by EI(x) represents the magnitude of the expected performance improvement. The larger its value, the greater the potential performance improvement by sampling at this point. In the genetic algorithm, EI(x) is used to guide the search direction: by preferentially selecting individuals with higher EI(x) for crossover and mutation, regions more likely to improve performance are explored.

[0117] Calculate the predicted probability P predict (x) and the training probability P train (x), and determine whether to conduct an actual training evaluation on the individual GNN architecture according to the training probability P train (x). Among them, the predicted probability P predict (x) and the training probability P train (x) are calculated as follows:

[0118] P predict (x) = max(P 0 , min(1, EI(x))) (9)

[0119] P train (x) = 1 - P predict (x) (10)

[0120] Among them, P 0 is the lower bound of the predicted probability (for example, P 0 = 0.1). If the training probability P train (x) is higher than the preset value (for example, 0.5), then conduct actual training and evaluation on the individual GNN architecture x, obtain its true performance y on the validation set, and update the Gaussian process surrogate model with the new data point (x, y), that is, add the new data point to the evaluated data set and retrain the Gaussian process surrogate model.

[0121] S6. Determine whether the preset optimization termination condition is satisfied. If so, output the optimized GNN architecture (i.e., the GNN architecture with the highest predicted performance) as the optimization result. Otherwise, return to step S3 for the next generation of genetic evolution.

[0122] Among them, the optimization termination conditions include at least one of the following:

[0123] The number of iterations of the genetic algorithm reaches the preset maximum number of iterations (for example, 50 generations);

[0124] The improvement amplitude of the predicted performance of the population's optimal individual GNN architecture is lower than the preset threshold (for example, 0.001) within a continuous preset number of generations (for example, 10 generations);

[0125] The preset calculation resource limit is reached (for example, the total training time exceeds 24 hours).

[0126] The optimized GNN architecture can be applied to graph data analysis tasks with high requirements for performance and efficiency. Typical application scenarios include, but are not limited to:

[0127] Drug discovery and molecular property prediction: In this scenario, the input is graph data representing the molecular structure of chemical compounds (nodes represent atoms, edges represent chemical bonds, and nodes and edges can carry features such as atom types and chemical bond types). The optimized GNN architecture is used to perform graph classification tasks to predict specific properties of molecules, such as biological activity, toxicity, solubility, or drug-likeness. This method can automatically discover the optimal GNN architecture suitable for capturing complex molecular structure-property relationships, accelerate the candidate drug screening process, and reduce R & D costs. For example, in the graph classification task of chemical compounds, the output is the class label of each graph (molecule) (such as whether it has mutagenicity or high activity).

[0128] Bioinformatics and precision medicine: The input is biological network graph data, such as protein-protein interaction (PPI) networks (nodes are proteins, edges are interactions) or gene regulatory networks. The optimized GNN architecture can be used for node classification tasks (such as predicting protein functions and identifying disease-related genes) or graph classification tasks (such as classifying cancer subtypes based on gene expression networks). This method helps to discover GNN models that can effectively integrate biological network topologies and node features, promoting precision medicine and biological research. For example, in the node classification task in a citation network (which can be analogous to certain biological networks), the input is the graph structure data of the network, and the output is the class label of each node.

[0129] Financial risk control and fraud detection: The input is a financial transaction network (nodes are accounts / users, edges are transactions / transfers) or a user relationship network. The optimized GNN architecture is used for node classification (such as identifying fraud accounts or high-risk users) or subgraph detection (identifying fraud rings). This method can automatically optimize a GNN architecture that can effectively capture complex and dynamic fraud patterns, improving the risk identification ability and anti-fraud efficiency of financial institutions.

[0130] Node classification in a citation network, where the input is the graph structure data of the citation network, nodes represent documents, edges represent citation relationships, and the output is the class label of each node (such as research topic).

[0131] Next, a specific example of the method of the present invention will be described.

[0132] S1. Construct a GNN architecture grammar library: The GNN architecture grammar library is constructed based on Figure 3 the rules shown. This grammar library is defined using the BNF paradigm and includes non-terminals, terminals, and production rules.

[0133] For example, non-terminal <net>Represents a network module, and its production rules

[0134] <net> → <dim> <gconv> <fc>Indicates that a network can be represented by a dimensional hyperparameter <dim>, a graph convolution module <gconv>and a fully connected layer <fc>They are connected in series. The terminal GCNConv represents a specific GCN convolution operation, Relu represents the ReLU activation function, and 16, 32, 64, 128, etc. represent the specific values of the hidden layer dimensions.

[0135] S2. Generate the initial GNN architecture population: Based on Figure 3 the grammar rule library, use the genetic programming tool of the DEAP library to generate the initial population, and set the population size to 50. Each individual GNN architecture is represented as a syntax tree, and the generation process of the syntax tree is as described above. First, use the starting symbol <net>is the root node, and then according to the grammar rules, recursively expand the non-terminal nodes until all leaf nodes are terminal symbols. For example, according to the rule <net> → <dim> <gconv> <fc> , <net>The node is expanded into three child nodes <dim> 、 <gconv>and <fc>。For <gconv>Nodes can also be based on rules

[0136] <gconv> → <dim> <conv> <fc>Or <gconv> → <conv>Continue to expand until all leaf nodes are terminators such as GCNConvLayer, Relu, 16, etc., finally forming a complete syntax tree, which synchronously encodes the network structure and hyperparameter information of the GNN.

[0137] S3. Performance prediction based on Gaussian process surrogate model: Use the Gaussian process surrogate model to predict the performance of the GNN architecture. The process framework is as Figure 2 shown. The Gaussian process surrogate model is defined based on formula (1), and its covariance function k(x, x') is calculated based on the tree kernel function defined by formula (3). Use the scikit-learn and SciPy libraries in Python to build the Gaussian process surrogate model. The tree kernel function is implemented using the custom TreeKernel class, which calculates the similarity between syntax trees based on the tree edit distance algorithm provided by the zss library. The kernel width parameter σ is set to 1.0, and the regularization parameter is set to 1e -6 .

[0138] In the initial stage, randomly select 5 GNN architecture individuals for actual training and evaluation, and obtain their classification accuracy on the validation set as the true performance, which is used to initialize the Gaussian process surrogate model. Then, for the subsequently generated GNN architecture individuals, use the trained Gaussian process surrogate model to predict their performance and estimate the prediction uncertainty (prediction variance).

[0139] S4. Genetic algorithm optimization: Use the genetic algorithm tool provided by the DEAP library to perform optimization search in the syntax tree space. The selection operation adopts tournament selection, and the tournament size is set to 3. The crossover operation adopts subtree crossover, and the crossover probability is set to 0.5. The mutation operation adopts subtree mutation, and the mutation probability is set to 0.2.

[0140] S5. Actual training evaluation and model update: Calculate the expected improvement value EI(x), prediction probability P predict (x), and training probability P train (x) using formulas (8), (9), and (10). The exploration factor ξ is set to 0, and the lower bound of the prediction probability P 0 is set to 0.1. If the training probability P train If (x) is greater than 0.5, then the PyTorchGeometric library is used to actually train and evaluate this GNN architecture. The training optimizer is Adam, the learning rate is set to 0.01, the cross-entropy loss function is used, and the number of training epochs is set to 50 or 100. Experiments are conducted on the Cora and Citeseer node classification datasets, as well as the MUTAG and PTC graph classification datasets. The experimental hardware environment is a 64-bit system, an Intel(R) Core(TM) i5-11320H processor, 16GB of memory, and an NVIDIA GeForce MX450 graphics card.

[0141] The true performance of the GNN architecture obtained from the actual training and evaluation is used as a new data point and added to the training dataset of the Gaussian process surrogate model, and the Gaussian process surrogate model is retrained using all the evaluated data to improve the prediction accuracy.

[0142] S6. Optimization termination: The genetic algorithm terminates after 50 generations of iteration. The GNN architecture with the optimal prediction performance in the last generation of the population is output as the final optimization result. Then this optimized GNN architecture can be applied to specific graph data analysis tasks. For example, in the chemical compound graph classification task, it can be used to predict the mutagenicity of molecules (verified using the MUTAG dataset). The input is a molecular graph, the nodes are atoms, the edges are chemical bonds, and the model outputs a binary classification result indicating whether the molecule has mutagenicity. The architecture optimized by this method aims to improve the accuracy of such molecular property prediction tasks and demonstrate its potential in the auxiliary design of drug discovery. Similarly, the optimized architecture can also be applied to node classification in citation networks (such as the Cora dataset) or other graph data analysis tasks in the fields of bioinformatics, financial risk control, etc.

[0143] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0144] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.< / conv> < / gconv> < / fc> < / conv> < / dim> < / gconv> < / gconv> < / fc> < / gconv> < / dim> < / net> < / fc> < / gconv> < / dim> < / net> < / net> < / fc> < / gconv> < / dim> < / fc> < / gconv> < / dim> < / net> < / net> < / conv> < / fc> < / conv> < / dim> < / gconv> < / fc> < / conv> < / dim> < / gconv> < / gconv> < / gconv> < / net> < / fc> < / conv> < / dim> < / fc> < / conv> < / dim> < / gconv> < / fc> < / gconv> < / net> < / k> < / dim> < / act> < / gat> < / conv> < / fc> < / attention> < / residual> < / gconv> < / net>

Claims

1. A graph neural network architecture optimization method based on grammar genetic programming, characterized in that: The following steps are involved: S1. Build a grammatical rule base for the graph neural network architecture GNN; S2. Generate the initial GNN architecture population through genetic programming based on the grammar rule base. Each individual GNN architecture in the population is represented as a grammar tree. The grammar tree is recursively generated according to the production rules and synchronously encodes the network structure and the hyperparameters of each layer of the network. S3, using Gaussian process surrogate models to predict the performance of individual GNN architectures in the population; S4, based on the prediction performance of individual GNN architectures, search for the optimal GNN architecture in the syntax tree space through genetic algorithm; S5. For individual GNN architectures in the population, calculate the training probability P based on its expected improvement function EI(x) train (x), if the training probability P train (x) is higher than the preset evaluation threshold, the individual GNN architecture is selected for actual training evaluation to obtain its true performance, and the Gaussian process proxy model is updated using the true performance; S6. Determine whether the preset optimization termination condition is met. If so, output the optimized GNN architecture. Otherwise, return to step S3. The optimized GNN architecture is used to process graph structure data to perform classification tasks.

2. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 1, characterized in that: In S1, the grammar rule base is defined based on the BNF paradigm, including a non-terminal symbol set N, a terminal symbol set T, a production rule set P, and a start symbol S; Among them, the non-terminal symbol set N contains an extensible network module, which is divided into a hyperparameter non-terminal symbol set N H and module non-terminal set N M ; The terminal symbol set T contains specific network operators and hyperparameter values, which are divided into hyperparameter terminal symbol set T H and module terminator set T M ; The set of production rules P defines the inference relations from non-terminal symbols to terminals and / or non-terminal symbols; The starting symbol S represents the starting derivation symbol of the network architecture.

3. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 2 is characterized in that: Each production rule p∈P in the production rule set P is of the form l→r, where l∈N is the derived symbol and r=r H r M To derive the results, r H ∈(N H ∪T H ) * Represents the hyperparameter part, r M ∈(N M ∪T M ) * Represents the submodule part, is the set of hyperparameter non-terminal symbols, is the set of hyperparameter terminators, is the set of module non-terminal symbols, A collection of module terminators.

4. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 1, characterized in that: In S2, the syntax tree generation process includes: Initialization: Create an initial syntax tree with the starting symbol S as the root node; Recursive expansion: For each non-terminal leaf node l in the current syntax tree k ∈N∪S, from the production rule subset P(l k )={p:l k →r k ,p∈P} randomly selects a production rule to expand, where r k The right side of the production rule is a sequence of hyperparameter symbols and module symbols; Termination condition: Repeat the recursive expansion steps until all leaf nodes of the syntax tree are terminals.

5. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 1, characterized in that S3 include: S31. Define the Gaussian process proxy model based on the following formula: f~GP(m(x),k(x,x′)) Where f represents the fitness function, m(x) represents the mean function, k(x,x') represents the covariance function, and the mean function is defined and calculated based on the following formula: in, represents the expected value; f(x) represents the random function in the Gaussian process agent model, which is used to approximate the true fitness function; The covariance function k(x,x') is calculated based on the tree kernel function defined as follows: Among them, D Tree (x,x') represents the tree edit distance between syntax trees x and x', σ is the kernel function width parameter; S32, training the Gaussian process proxy model based on the evaluated training data (X, y), where X is the evaluated GNN architecture set and y is the corresponding true performance; S33. Use the trained Gaussian process surrogate model to predict the new un-evaluated GNN architecture X * The fitness function f * The distribution of , the result follows a normal distribution: Among them, μ * is the predicted mean; σ * is the prediction standard deviation, which indicates the uncertainty of the prediction.

6. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 1, characterized in that: In S4, the genetic algorithm includes: selection, crossover, and mutation operations; The selection operation adopts the tournament selection method, which randomly selects a certain number of individuals from the population for comparison each time, selects the individual with the best prediction performance as the Parent, and repeats the selection operation until the number of Parent individuals meets the requirements of crossover and mutation; The crossover operation includes: randomly selecting a non-terminal symbol node from each of the two Parent individual syntax trees, ensuring that the non-terminal symbol types of the two nodes are the same; exchanging the subtrees with the two selected nodes as root nodes to generate two new child individual syntax trees; The mutation operation includes: randomly selecting a non-terminal symbol node in the individual syntax tree; randomly generating a new subtree according to the production rule with the non-terminal symbol as the left side in the grammar rule base; replacing the selected node with the newly generated subtree to complete the mutation operation.

7. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 1, characterized in that: In S5, the expected improvement function EI(x) is used to evaluate the potential improvement value of the individual GNN architecture, and its calculation result is used to determine the training probability P train (x) is also used to guide the search direction of the genetic algorithm or to select from multiple individuals that meet the evaluation conditions; wherein guiding the search direction of the genetic algorithm includes: giving priority to individuals with higher improvement values ​​in the selection operation of the genetic algorithm, or selecting individuals with the highest improvement value from multiple individual GNN architectures that meet the actual training evaluation conditions for actual training evaluation; The expected improvement function EI(x) is defined based on the following formula: Where Imp represents the improvement amount, Z represents the standardized improvement amount, Imp = μ(x)-y best -ξ, μ(x) and σ(x) are the mean and standard deviation of the predictions of the Gaussian process surrogate model, respectively. best is the best performance among the currently evaluated GNN architectures, ξ is the exploration factor, Φ and are the cumulative distribution function and probability density function of the standard normal distribution, respectively.

8. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 7 is characterized in that: The S5 also includes: Calculate the predicted probability P predict (x) and training probability P train (x), and according to the training probability P train (x) Decide whether to actually train and evaluate the individual GNN architecture, where the prediction probability P predict (x) and training probability P train The calculation formula for (x) is: P predict (x)=max(P0,min(1,EI(x))) P train (x)=1-P predict (x) Among them, P0 is the lower bound of the prediction probability. If the training probability P train If (x) is higher than the preset value, the individual GNN architecture x is actually trained and evaluated to obtain its actual performance y on the validation set, and the Gaussian process proxy model is updated using the new data point (x, y), that is, the new data point is added to the evaluated data set and the Gaussian process proxy model is retrained.

9. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 1, characterized in that: In S6, the optimization termination condition includes at least one of the following: The number of iterations of the genetic algorithm reaches the preset maximum number of iterations; The prediction performance of the optimal individual GNN architecture in the population is improved by less than the preset threshold within the continuous preset number of generations; The preset computing resource limit is reached.

10. The method for optimizing graph neural network architecture based on grammar genetic programming according to claim 1, characterized in that: In S6, the optimized GNN architecture is used for node classification in citation networks or for chemical compound graph classification to predict molecular properties; in the node classification task in citation networks, the input is the graph structure data of the citation network, the nodes represent documents, the edges represent citation relationships, and the output is the category label of each node; in the graph classification task of chemical compounds, the input is a set of graph structure data representing chemical compounds, the nodes represent atoms, the edges represent chemical bonds, and the output is the category label of each graph.

Citation Information

Patent Citations

  • Particle orbit discovery method based on graph neural network cellular automaton

    CN115563841A

  • Water conservancy big data service analysis and evaluation model construction method and system

    CN119740759A

  • Diffusion for realistic scene generation

    US20240300527A1

Cited By

  • Network dynamics symbol regression method based on LLM and Bayesian optimization

    CN120471099A

  • A symbolic regression method for network dynamics based on LLM and Bayesian optimization

    CN120471099B