Method and device for graph neural architecture search under distributional shift
The method addresses the challenge of distribution shifts in GNAS by using a disentangled graph encoder and adaptive architecture blending, enabling GNNs to tailor unique architectures for each graph, enhancing generalization and accuracy across varying graph distributions.
Patent Information
- Application Number
- DE112022007534
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-07-14
- Publication Date
- 2025-06-18
AI Technical Summary
Existing graph neural architecture search (GNAS) methods fail to handle distribution shifts between training and test graphs, leading to overfitting and inaccurate predictions due to the assumption of independent and identical distribution (IID), which is not applicable in real-world scenarios with varying graph distributions.
A method involving a self-supervised disentangled graph encoder to project graphs into a disentangled latent space, followed by an architecture adaptation module that tailors specialized GNN architectures based on graph representations and a supernetwork module to blend operations in a continuous space, optimizing the GNN architecture for each graph instance.
Enhances the generalizability of GNNs under distributional shifts, allowing them to adapt to different graph structures and improve predictive accuracy by learning unique architectures and weights for each graph, demonstrated through superior performance on synthetic and real-world datasets.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
REGIONThe present disclosure relates generally to the art of artificial intelligence and, more particularly, to graph neural architecture search technology.BACKGROUNDGraph structured data has attracted much attention in recent years because of its flexible display capability in various areas. Graph neural network (GNN) models have been proposed that have achieved great success in many graph tasks. To minimize human effort in designing GNN architectures for various tasks and automatically design more powerful GNNs, the Graph Neural Architecture Search (GNAS) is used to search for an optimal GNN architecture. These automatically designed architectures have achieved competitive or better performance in comparison to manually designed GNNs for data sets with the same distributions assuming the independent and identical distribution (IID), i.e., the training and test graphs are taken independently from each other from the identical distribution.However, distribution offsets are ubiquitous and unavoidable in real graph applications because there are a large number of unpredictable and uncontrollable hidden factors. The existing GNAS approaches, under the IID assumption, only search a single fixed GNN architecture based on the training set before applying the selected architecture directly to the test set. In this case, they cannot process the different distribution displacements in the out-of-distribution setting. Because the single GNN architecture discovered with existing methods may be over-matched to the distributions of the training graph data, it is possible that accurate predictions cannot be made for test graph data having different distributions that are different from the training graph data.Therefore, there is a need for an improved method and apparatus for graph neural architecture searching involving distribution shifts between training graph data and test graph data.SUMMARYThe following provides a simplified summary of one or more aspects in accordance with the present disclosure to provide a basic understanding of such aspects. This summary is not a comprehensive overview of all aspects considered and is not intended to identify key or critical elements of all aspects, nor to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in simplified form as a prelude to the more detailed description presented below.In one aspect of the disclosure, a method of designing a graph neural network (GNN) by a graph neural architecture search under distribution shifts between training graphs and test graphs is disclosed. The method comprises: obtaining, by a graph encoder module, a graph representation of an input graph in a sophisticated latent space; obtaining, by an architecture adjustment module, a searched GNN architecture for the input graph based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture, the probability of the operation being a function of the similarity between the obtained graph representation and a trainable prototype vector representation of the operation; and obtaining, by a super network module, weights for the searched GNN architecture in which different operations in the set of candidate operations are mixed in a continuous space.In another aspect of the disclosure, an apparatus for designing a GNN by graph neural architecture search under distribution shifts between training graphs and test graphs is disclosed. The apparatus comprises: a graph encoder module to obtain a graph representation of an input graph in a sophisticated latent space; an architecture adaptation module to obtain a searched GNN architecture for the input graph based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture, the probability of the operation being dependent on the similarity between the obtained graph representation and a trainable prototype vector representation of the operation; and a super network module to obtain weights for the searched GNN architecture, wherein different operations in the set of candidate operations are mixed in a continuous space in the super network module.In another aspect of the disclosure, an apparatus for designing a GNN by graph neural architecture search under distribution shifts between training graphs and test graphs is disclosed. The apparatus may include a memory and at least one processor coupled to the memory. The at least one processor may be configured to: obtain a graph representation of an input graph in a disentangled latent space by a graph encoder module; obtain a searched GNN architecture for the input graph by an architecture adjustment module based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture, the probability of the operation being a function of the similarity between the obtained graph representation and a trainable prototype vector representation of the operation; and obtain weights for the searched GNN architecture by a super network module in which different operations in the set of candidate operations are mixed in a continuous space.In another aspect of the disclosure, a computer readable medium is disclosed that stores computer code for designing a GNN by graph neural architecture search under distribution shifts between training graphs and test graphs. When executed by a processor, the computer code may cause the processor to obtain a graph representation of an input graph in a disentangled latent space by a graph encoder module; obtain a searched GNN architecture for the input graph by an architecture adjustment module based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture, the probability of the operation being dependent on the similarity between the obtained graph representation and a trainable prototype vector representation of the operation; and obtain weights for the searched GNN architecture by a super network module in which different operations in the set of candidate operations are mixed in a continuous space.In another aspect of the disclosure, a computer program product for designing a GNN by graph neural architecture search under distribution shifts between training graphs and test graphs is disclosed. The computer program product may include processor executable computer code to: obtain, by a graph encoder module, a graph representation of an input graph in a disentangled latent space; obtain, by an architecture adaptation module, a searched GNN architecture for the input graph based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture, the probability of the operation being dependent on the similarity between the obtained graph representation and a trainable prototype vector representation of the operation; and obtain, by a super network module, weights for the searched GNN architecture in which different operations in the set of candidate operations are mixed in a continuous space.Other aspects or variations of the disclosure will become apparent upon consideration of the following detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGSThe following figures illustrate various embodiments of the present disclosure for illustrative purposes only. One skilled in the art will readily recognize from the following description that alternative embodiments of the methods and structures disclosed herein may be implemented without departing from the spirit and principles of the disclosure described herein. FIG. 1 illustrates a simplified example of a graph in accordance with an aspect of the present disclosure. FIG. 2 illustrates a schematic model of a graph neural architecture search under distribution displacements according to an aspect of the present disclosure. FIG. 3 illustrates a flowchart of a method for designing GNNS by GNAS under distribution offsets according to an aspect of the present disclosure. FIG. 4 illustrates a block diagram of an apparatus for designing GNNS by GNAS under distribution offsets according to an aspect of the present disclosure. FIG. 5 illustrates a block diagram of an apparatus for designing GNNS by GNAS under distribution offsets according to an aspect of the present disclosure.DETAILED DESCRIPTIONBefore explaining embodiments of the present disclosure in detail, it is to be understood that the disclosure is not limited in its application to the details of construction and arrangement of features set forth in the following description. The disclosure is capable of other embodiments and of being practiced or carried out in various ways.FIG. 1 illustrates a simplified example of a graph in accordance with an aspect of the present disclosure. A graph is a non-linear data structure consisting of nodes and edges. The nodes may also be referred to as vertices, and the edges are lines or arcs connecting any two nodes in the graph. For example, as shown in FIG. 1, a simple graph 100 consists of nodes n1to n7and edges e1to e6, where edge e1connects nodes n1and n3, edge e2connects nodes n2and n4, and so on. Graph structured data can be used in various fields including social networks, information networks, biological networks, infrastructure networks, etc., which cannot be structured in Euclidean space. In one example, a graph 100 may be a production line layout graph, each node in the graph may represent a workshop, and each edge may represent the connection between two workshops. In another example, a graph 100 may be a circuit design layout graph, each node in the graph may represent an electron device or circuit module, and each edge may represent the connection between two electron devices or circuit modules.Graph neural networks (GNNs) can learn node representations through a recursive messaging scheme in which nodes iteratively bundle information of their neighbors. Taking the graph classification task as an example, GNNs may use pooling methods to derive graph-level representations. Different GNN architectures differ mainly in their messaging mechanism, i.e. in the way information is exchanged to adapt to the requirements of different graph scenarios. Graph neural architecture search (GNAS) can be used to automatically design GNN architectures for various graph tasks. However, when a distribution shift occurs between training and test graphs, the existing approaches fail to adapt problematics to unknown test graph structures because they only search for a fixed architecture for all graphs. Using the example of drug discovery, it can be seen that only a limited amount of training data is available for experiments and that the interaction mechanisms differ greatly in different molecules on account of their complex chemical properties. Therefore, the GNN models designed for drug research must frequently be tested from data with distribution offsets.In this disclosure, an improved approach for graph neural architecture searching under distribution displacements is provided. Such an approach for graph neural architecture searching may possibly capture important information about graphs with widely different distributions under the out-of-distribution conditions by custom designing a unique GNN architecture for each graph instance.In particular, a self-supervised, sophisticated graph coder is designed that can project graphs into a sophisticated latent space, where each sophisticated factor in the space is trained simultaneously by the supervised task and the corresponding self-supervised learning task. With this design, the graph hidden key information can be better controlled and acquired via the self-supervised, sophisticated graph representation. In this way, the generalization of the representations under distribution displacements is improved. Then, prototype architecture adaptation is adopted to tailor specialized GNN architectures for graphs based on similarities of their representations to prototype vectors in latent space, where each prototype vector corresponds to a different operation. Next, a customized supernetwork with differentiable weights is designed for merging different operations, which provides great flexibility to compile different combinations of operations and to enable simple end-to-end optimization of the disclosed GNSS model by gradient-based methods. The developed graph representation and trainable prototype operation mapping designs can improve the generalizability of the disclosed GNAS model under distribution offsets. Extensive experiments with both synthetic and real graph datasets also demonstrate the superiority of the disclosed GNAS model over existing GNAS baselines.For simplicity, a graph space as a label space is referred to as a training graph dataset as the corresponding training label dataset as a test graph dataset and the corresponding test label dataset as. The goal of GNAS under distribution displacements is to design a model using G tr and Y tr that works well with G te and Y te assuming that P(G tr, Y tr) ≠ P(G te, Y te), where P(G tr, Y tr) denotes the probability distribution of the training graph dataset and P(Gte, Yte) denotes the probability distribution of the test graph dataset, i.e., wherein is a loss function. In a common but challenging environment, neither Y te nor unlabeled G te is available in the training phase. F could be GNNs for graph machine learning. A typical GNN is made up of two parts: an architecture and trainable weights, where and designate the architectural space and the weight space, respectively. Therefore, GNNs may be referred to as the following mapping function.This disclosure focuses mainly on different GNN layers, i.e. messaging functions, for searching the GNN architecture. Therefore, a search space is adopted with layered architectures without complex connections such as residue or hop connections, although the disclosed method can be easily generalized. In one embodiment, five widely used GNN layers may be used as the operation candidate set, including Graph Convolutional Network (GCN), Graph Attention Network (GAT), Graph Isomorphism Network (GIN), Graph Sample and Ageate (SAGE), and GraphCon. In addition, a multi-layer perceptron (MLP) can also be adopted which does not take into account graphene structures. A pooling layer may also be defined as standard global mean pooling at the end of the GNN architecture.In this disclosure, not as in the existing GNAS methods, a fixed GNN architecture is used for all graphs, but an individual GNN architecture can be adapted for each graph. In this way, the disclosed GNAS method is more flexible and can better process distribution shifted test graphs because different GNN architectures are known to fit different graphs. Therefore, it is necessary to learn an architecture mapping function and a weight mapping function in order that these functions can automatically generate the optimal GNN for various graphs including the architecture and their weights. Since the architecture depends only on the graph, the weight mapping function can be simplified further than Therefore, the equation (1) can be transformed into the following objective function: where the regularizer and γ is a hyperparameter representing the weight of the main loss function. Specific embodiments for appropriately designing them will be described in detail below in connection with FIG. 2 such that the disclosed GNAS method may be generalized among distribution offsets.FIG. 2 illustrates a schematic model of a graph neural architecture search under distribution displacements according to an aspect of the present disclosure. As shown in FIG. 2, the GNAS model 200 includes three cascaded modules, i.e., a self-supervised, sophisticated graph encoder module 210, a prototype strategy architecture adaptation module 220, and an adapted super network module 230 to tailor a unique GNN architecture for each graph instance, thereby enabling the model 200 to handle generalizations among distribution offsets with non-IID settings. The graph encoder module 210 may detect various graph structures through self-monitored and monitored loss. Subsequently, the architecture adaptation module 220 may tailor the most suitable GNN architecture based on the learned graph representation. Finally, the customized super network module 230 may enable efficient training through weight distribution. Each of these modules will be described in detail below.As shown in FIG. 2, graphs g 1, g 2 and g 3 having different structures are input to the graph encoder module 210. It should be appreciated that many more input graphs may be input to the graph encoder module 210 during the training phase. These input graphs may have different graph structures from different distributions. To detect such different graph structures, the graph encoder module 210 may learn low-dimensional representations of graphs. In one embodiment, K GNNs may be adopted to learn K-block graph representations: where the k-th block is the node representation in the l-th layer, A is the adjacent matrix of the graph, and || represents the concatenation. By using these developed GNN layers, various latent factors of the input graphs can be detected. A readout layer can then be adopted in order to bundle node-level representations into a graph-level representation:To learn the parameters of the self-supervised, sophisticated graph encoder module 210, both the graph-supervised learning task and the self-supervised learning task may be used simultaneously.The downstream target graph task, of course, provides supervisory signals for learning the graph encoder module 210. Therefore, after the obtained graph representation, a classification layer may be placed to obtain the prediction for the graph classification task. The plot for g_i may be referred to as h i. The supervised learning loss is as follows: wherein C(·) is the classification layer.Graph Self-Supervised Learning (SSL) aims to learn informative graph representations through front wall tasks, which has shown several advantages, including less label dependency, improved robustness, and model generic capability. Therefore, graph SSL may also be used as a supplement to the task of supervised learning. In particular, an SSL auxiliary task may be determined by generating pseudo-labels from graph structures, and the pseudo-labels may be used as additional monitoring signals. In addition, different pseudo labels may be adopted for different blocks of the evolved GNN, so that the evolved graph encoder 210 may detect different factors of the graph structure. In one embodiment, graph coder 210 may focus on the degree distribution of graphs as a representative and clearable structural feature, while generalization to other graph structures is straightforward. In particular, the pseudo-labels for the kthGN block may be generated by calculating the ratio of the nodes having exactly the degree k. Then, the SSL objective function may be formulated as follows: wherein the pseudo label is and may be obtained by applying a regression function, such as a linear layer followed by an activation function, to the kth block of the graph representation h_i. In one embodiment, the final block may be left without SSL tasks to allow more flexibility in learning the developed graph representations.As shown in FIG. 2, after receipt, the graph representations h 1, h 2 and h 3 may be input to the architecture adaptation module 220, which maps the prototype strategy representations to various customized GNN architectures. In particular, the probability of selecting an operation o in the i-th layer of a searched architecture may be referred to as p_o^i, where i∈{1,2,...,N}, N is the number of layers and. The probability may be calculated as follows: wherein a trainable prototype vector representation of the operation is o. I 2- normalization to q may be adopted to ensure numerical stability and fair competition between different operations. In the architecture adaptation module 220, a prototype vector may be learned for each candidate operation, and operations may be selected based on the preferences of the graph, i.e., if the graph representation has a large projection on a prototype vector, the corresponding operation is more likely to be selected. In addition, using the exponential function, the length of h can determine the shape of, i.e., the greater ||h|| 2, the more likely it is that a few values will be dominated, indicating that the graph requires certain operations.In one embodiment, to avoid the problem of mode collapse, i.e., vectors of different operations are similar and therefore no longer distinguishable, based on cosine distances between vectors, the following regularizer may be adopted to maintain the variety of operations:Prototype architecture adaptation in module 220 may tailor the most suitable GNN architectures for different input graphs based on the graph representations. In addition to GNN architectures, the weights of the architectures must also be learned.As shown in FIG. 2, a super network module 230 may be adopted to obtain the weights of architectures. In particular, in the super network all possible operations are taken into account together by mixing different operations in a continuous space as follows: where x is the input of the ithlayer and f i( x) is the output. Subsequently, all weights can be optimized using gradient descent methods. In addition, since the weights of different architectures are shared, training is substantially more efficient than separately training the weights for different architectures.Note that in most NAS models, the architecture is discretized at the end of the search phase by selecting the operation with the largest for all layers. The weights of the selected architecture are subsequently retrained. Retraining is not feasible in the disclosed novel GNAS model, however, as test graphs with architectures other than the training graphs can be adapted. Therefore, the weights of the super network can be used directly as weights in the searched architecture. In addition, the continuous architecture is maintained without the discretized step, which increases flexibility in architecture adaptation and simplifies the optimization strategy. Moreover, the adapted super network can serve as a strong ensemble model, wherein the ensemble weights are, which can also benefit generalization outside the distribution.As shown in FIG. 2, the GNAS model 200 may be optimized by using gradient descent methods based on the following loss function: where L___main" is the monitoring loss of the customized architectures in equation (2), i.e., the monitoring loss of the final prediction given by the super network module 230, are the monitoring loss and the self-monitoring loss of the self-monitored evolved graph encoder module 210, the cosine distance loss of the architecture matching module 230, β 1 and β 2 are hyperparameters. three additional loss functions introduced as regularizer in equation (2), and γ is the hyperparameter for controlling the contribution of the regularizer.For overall optimization, there are two groups of loss functions: classification loss and regularizer. It is possible that the self-supervised, sophisticated graph coder was not properly trained at an early stage of the training operation and the learned graph representation is also not meaningful, resulting in unstable architecture adaptation. Therefore, a larger weight may be set first for the regularizer, i.e., a smaller initial γ in equation (2), to force the self-supervised, sophisticated graph coder to learn through its supervised learning and SSL tasks. As the training process proceeds, the focus may be placed more stepwise on the training of the architecture adaptation module and the super network module by increasing γ as follows: where γ t is the hyperparameter value at the t-th time and Δγ is a small constant.The entire training procedure is as shown below. The most suitable GNN architecture with its parameters can be generated directly for the test graphs without retraining.Input: training dataset G tr and Y tr, hyperparameters γ 0, Δγ t, β 1, β 2initializing all learning parameters and setting γ = γ 0while not converging, performcalculating the plots h using equations (3) and (4)Calculate (5) and Eq using equations (5) and (6). (6) (6)Calculating Architectural Probability Using Equation (7) (7)Calculating Using Equation (8)obtaining the parameters from the supernet based on equation (9)Calculating Total Loss in Equation (2)updating the parameters using the gradient descentUpdating of γ = γ - ΔγEnd duringIn one embodiment, different learning rates may be used for the three modules 210-230. For example, the learning rate of the self-supervised sophisticated encoder module may be 1.5 e- 4. The learning rate of the architecture adaptation module may be 1 e- 4. The training process of these two modules is planned according to the Consensus Annealing method. The learning rate of the adjusted super network module may be 2e-3. γ may be initialized to 0.07 and linearly increased to 0.5. Moreover, β 1 may be set to 0.05 and β 2 may be set to 0.002. The number of layers may be set to 2 or 3. The disclosed GNAS approach is not limited to such settings.FIG. 3 illustrates a flow diagram of a method 300 for designing GNNs by GNAS with distribution relocation, in accordance with an aspect of the present disclosure. In one embodiment, method 300 may be used to design a graph neural network for a graph classification task by graph neural architecture search under distribution shifts between training graphs and test graphs. The method 300 may also be used for other graph machine learning tasks. In one embodiment, the method 300 may be a computer-implemented method.At block 310, the method 300 may obtain, by a graph encoder module, a graph representation of an input graph in a sophisticated latent space. The input graph may be a conveyor layout graph. The input graph can also be a circuit design layout graph, for example a printed circuit board design or a chip design. GNNs designed by method 300 may be used to make classifications of the input pipeline layout graphs or circuit design layout graphs under distribution offsets. For example, the GNNs may classify an input pipeline layout graph or circuit design layout graph according to whether the pipeline layout is meaningful or whether the circuit is efficient, etc. The graph encoder module may be a self-supervised sophisticated graph encoder that may characterize invariant factors hidden in different graph structures. The graph encoder module may calculate the graph representation of the input graph using equation (3) and equation (4). The graph encoder module may be trained by a supervised learning task and a self-supervised learning task. The graph encoder module may calculate β 2 and using equation (5) and equation (6).At block 320, the method 300 may obtain, by an architecture adaptation module, a searched GNN architecture for the input graph based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture. The probability of the operation is dependent on the similarity between the obtained graph representation and a trainable prototype vector representation of the operation. In one embodiment, the probability of operation may be calculated using equation (7) for each layer of the searched GNN architecture. The set of candidate operations may include at least one of a graph convolutional network (GCN), a graph attention network (GAT), a graph isomorphism network (GIN), and a graph sample and age (SAGE). The set of candidate operations may include other GNN layers, such as GraphCon. Depending on the various graph tasks, base graph shapes, and / or records, MLP may be adopted that does not take into account graph structures, a pooling layer may be set as global mean pooling at the end of the searched GNN architecture, or GIN may be set in the first layer for the Laneious Motif record.At block 330, the method 300 may obtain weights for the searched GNN architecture by a super network module in which different operations in the set of candidate operations are mixed in a continuous space, for example using equation (9). In one embodiment, the super network weights may be used directly as weights in the searched GNN architecture. The weights of the super network may be shared among different GNN architectures, making training substantially more efficient than if the weights were separately trained for different architectures. Then, in the test phase, the trained weights may be retrieved directly from the supernetwork for the GNN architecture being searched.Although not shown in FIG. 3, in the training phase, the method 300 may include optimizing the GNN based on a loss-of-main function and a regularizer, which may include repeating blocks 310- 330 until convergence. The main loss function may be a supervisory loss of the searched GNN architecture as in equation (2), and the regularizer may be based on a supervised learning loss function as in equation (5) and a self-supervised learning objective function as in equation (6) for the graph coder and a cosine distance loss function between trainable prototype vector representations of the various operations as in equation (8). In addition, since it is possible at the early stage of the training operation that the self-supervised, sophisticated graph coder was not properly trained and the learned graph representation is also not meaningful, a smaller initial weight may be set for the main loss function and the weight is incremented during the training operation.FIG. 4 illustrates a block diagram of an apparatus 400 for designing GNNs by GNAS under distribution offsets, in accordance with an aspect of the present disclosure. As shown in FIG. 4, the apparatus 400 may include a graph encoder module 410, an architecture adaptation module 420, and a super network module 430.The graph encoder module 410 may be used to obtain a graph representation of an input graph in a sophisticated latent space. The architecture adaptation module 420 may be used to obtain a searched GNN architecture for the input graph based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture. The probability of the operation is a function of the similarity between the obtained graph representation and a trainable prototype vector representation of the operation, as in equation (7), for example. The super network module 430 may be used to obtain weights for the searched GNN architecture. Various operations in the set of candidate operations are mixed in a continuous space in the super network module 430. In the training phase, the operations performed by the graph encoder module 410, the architecture adaptation module 420, and the super network module 430 may be repeated to optimize the GNN based on a loss-of-main function and a regularizer using equation (2).The apparatus 400 may also be used for a GNN designed by a graph neural architecture search under distribution shifts between training graphs and test graphs. For example, in the test phase, the graph encoder module 410 may obtain a graph representation of an input graph with the parameters learned for a particular graph task. The architecture adaptation module 420 may obtain different GNN architectures for different input graphs based on a probability or weight of an operation in each layer. The super network module 430 may obtain learned weights in the training phase and share these weights for different GNN architectures.FIG. 5 illustrates a block diagram of an apparatus 500 for designing GNNs by GNAS under distribution offsets, in accordance with an aspect of the present disclosure. The apparatus 500 may include a memory 510 and at least one processor 520. The processor 520 may be connected to the memory 510 and configured to perform the method 300 as described above with reference to FIG. 3. Processor 520 may be a general purpose processor or may be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. The memory 510 may store the input data, output data, data generated by the processor 520, and / or instructions executed by the processor 520. The apparatus 500 may also be used for a GNN designed by a graph neural architecture search under distribution shifts between training graphs and test graphs, in accordance with the present disclosure.The various operations, modules, and networks described herein in connection with the disclosure may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. According to an embodiment of the disclosure, a computer program product for computer vision processing may include computer code executable by a processor for performing the method 300 described above with reference to FIG. 3. According to another embodiment of the disclosure, a computer readable medium may store computer code for computer vision processing, wherein the computer code, when executed by a processor, may cause the processor to perform the method 300 described above with reference to FIG. 3. Computer readable media includes both non-transitory computer readable storage media and communication media including any media that supports the transfer of a computer program from one location to another. Each connection may be referred to as a computer readable medium. Other embodiments and implementations are within the scope of the disclosure.The foregoing description of the disclosed embodiments is provided to enable one skilled in the art to make or use the various embodiments. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the scope of the various embodiments. Thus, the claims are not intended to be limited to the embodiments shown herein, but are to be accorded the broadest scope consistent with the following claims and the principles and novel features disclosed herein.
Claims
A method for designing a graph neural network (GNN) by graph neural architecture search under distribution shifts between training graphs and test graphs, comprising: obtaining a graph representation of an input graph in a sophisticated latent space by a graph encoder module; obtaining a searched GNN architecture for the input graph by an architecture adaptation module based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture, wherein the probability of the operation is dependent on the similarity between the obtained graph representation and a trainable prototype vector representation of the operation; obtaining weights for the searched GNN architecture by a super network module in which different operations in the set of candidate operations are mixed in a continuous space.The method of claim 1, wherein the input graph is a pipeline layout graph or a circuit design layout graph, and the designed GNN is used to classify the pipeline layout graph or the circuit design layout graph among distribution offsets.The method of claim 1, wherein the graph encoder module is trained by a supervised learning task and a self-supervised learning task.The method of claim 3, wherein the GNN is optimized based on a loss-of-main function and a regularizer, and wherein the loss-of-main function is a loss-of-monitoring of the searched GNN architecture, and the regularizer is based on a monitored loss-of-learning function and a self-monitored learning objective function for the graph coder, as well as a cosine distance-loss function between trainable prototype vector representations of the various operations.The method of claim 4, wherein a weight for the main loss function is incremented by a training operation.The method of claim 1, wherein the set of candidate operations comprises at least one of a graph convolutional network (GCN), a graph attention network (GAT), a graph isomorphism network (GIN), and a graph sample and ageate (SAGE).The method of claim 1, wherein a pooling layer at the end of the searched GNN architecture is set as global mean pooling.An apparatus for a graph neural network (GNN) designed by graph neural architecture search under distribution shifts between training graphs and test graphs, comprising: a graph encoder module to obtain a graph representation of an input graph in a sophisticated latent space; an architecture adaptation module to obtain a searched GNN architecture for the input graph based on a probability of an operation selected from a set of candidate operations in a layer of the searched GNN architecture, wherein the probability of the operation is dependent on the similarity between the obtained graph representation and a trainable prototype vector representation of the operation; and a super network module for obtaining weights for the searched GNN architecture, wherein different operations in the set of candidate operations are mixed in a continuous space in the super network module.The apparatus of claim 8, wherein the input graph is a pipeline layout graph or a circuit design layout graph, and the designed GNN is used to classify the pipeline layout graph or the circuit design layout graph among distribution offsets.The apparatus of claim 8, wherein the graph encoder module is trained by a supervised learning task and a self-supervised learning task.The apparatus of claim 10, wherein the GNN is optimized based on a loss-of-main function and a regularizer, and wherein the loss-of-main function is a loss-of-monitoring of the searched GNN architecture, and the regularizer is based on a monitored loss-of-learning function and a self-monitored learning objective function for the graph coder, as well as a cosine distance-loss function between trainable prototype vector representations of the various operations.The apparatus of claim 11, wherein a weight for the main loss function is incremented by a training operation.The apparatus of claim 8, wherein the set of candidate operations comprises at least one of: a graph convolutional network (GCN), a graph attention network (GAT), a graph isomorphism network (GIN), and a graph sample and age (SAGE).The apparatus of claim 8, wherein a pooling layer at the end of the searched GNN architecture is set as global mean pooling.An apparatus for designing a graph neural network (GNN) by graph neural architecture search under distribution shifts between training graphs and test graphs, comprising: a memory; and at least one processor coupled to the memory and configured to perform the method of any one of claims 1 to 7.A computer readable medium storing computer code for designing a graph neural network (GNN) by graph neural architecture search under distribution shifts between training graphs and test graphs, wherein the computer code, when executed by a processor, causes the processor to perform the method of any one of claims 1 to 7.A computer program product for designing a graph neural network (GNN) by graph neural architecture search under distribution shifts between training graphs and test graphs, comprising: processor executable computer code for performing the method of any one of claims 1 to 7.A computer-implemented method for designing a graph neural network (GNN) by graph neural architecture search under distribution shifts between training graphs and test graphs, comprising steps of the method according to any one of claims 1 to 7.