Processing equipment, data processing method thereof and method for training graph convolutional network model
By generating compressed graphs and aggregating supernodes and hyperedges using GCN model, the problem of low efficiency in large-scale graph data processing in the prior art is solved, and efficient GCN model training and inference are achieved.
Patent Information
- Application Number
- CN202411591830.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-11-08
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art is difficult to efficiently process large-scale graph data, resulting in extended training time and limited scalability of graph convolutional network (GCN) models.
By generating compressed graphs, including hypernodes and hyperedges, the GCN model is used to aggregate hypernodes and hyperedges, and then new nodes are added to achieve efficient processing of large-scale graph data.
Effectively reduce the number of operations, improve the efficiency of large-scale GCN processing, and solve the computing complexity and scalability challenges.
Smart Images

Figure CN120020818A_ABST
Abstract
Description
[0001] This application is based on and claims the benefit of priority from Korean Patent Application No. 10-2023-0161216 filed in the Korean Intellectual Property Office on November 20, 2023, the contents of which are incorporated herein in their entirety by reference. Technical Field
[0002] The following disclosure relates to a processing device and a data processing method thereof, as well as a method for training a graph convolutional network (GCN) model. Background Art
[0003] Modern computing systems face increasing challenges in efficiently processing large-scale graph data, which is prevalent in various fields such as social networks, recommender systems, etc. Existing methods usually cannot handle the irregularity and sparsity of graph structures, resulting in suboptimal performance and scalability issues. Graph convolutional networks (GCNs) have demonstrated capabilities in tasks such as node classification, link prediction, and graph generation.
[0004] However, the training and inference processing of GCNs can be computationally intensive, especially for large-scale graphs with millions or billions of nodes and edges. Traditional computing architectures can have difficulty efficiently handling the computational requirements of GCNs, resulting in extended training times and limited scalability. Therefore, there is a need in the art for systems and methods that can efficiently train and deploy GCN models on large-scale graph data while addressing the computational complexity and scalability challenges associated with existing methods. Summary of the invention
[0005] The present disclosure describes systems and methods for data processing. Disclosed embodiments include methods for training large-scale graph convolutional network (GCN) models based on compressed graphs. In some cases, the compressed graph includes multiple supernodes and multiple hyperedges generated by grouping nodes of the graph. One or more embodiments include adding new nodes based on aggregating supernodes and / or hyperedges using a GCN model.
[0006] According to one aspect, a data processing method of a processing device is provided, the data processing method comprising: obtaining embedding information of a first node to be added to a graph and connection information between the graph and the first node, receiving supernode information of a supernode of a compressed graph corresponding to the graph, wherein the supernode includes multiple nodes from the graph, and using a graph convolutional network (GCN) model to generate modified embedding information of the first node based on the supernode information, embedding information and connection information.
[0007] The step of generating the embedded information may include: obtaining a first result by performing a first aggregation on the embedded information, the supernode information, and the connection information, and obtaining a second result by performing a second aggregation on the embedded information and additional supernode information of neighboring supernodes connected to the supernode when the supernode to which the second node belongs does not have a self-edge.
[0008] The step of generating the embedding information may include correcting the first result based on correction information, wherein the correction information indicates a difference between a connection relationship of the compressed graph and a connection relationship of the graph.
[0009] Correcting the first result may include adding additional embedding information of one or more nodes on the first edge to the first result when the first edge is removed from the compressed graph.
[0010] Correcting the first result may include, when the second edge is added to the compressed graph, subtracting embedding information of one or more nodes on the second edge from the first result.
[0011] The step of generating the embedding information may include determining whether the neighboring supernode has a self-edge, wherein the second aggregation is based on the determination.
[0012] According to one aspect, a processing device is provided, comprising: a first buffer configured to store supernode information of a compressed graph, wherein the compressed graph comprises supernodes corresponding to a plurality of nodes of the graph; a second buffer configured to store embedding information of a first node of the graph; and an operation circuit configured to obtain the supernode information from the first buffer, obtain the embedding information from the second buffer, obtain connection information between the first node and the second node, and generate modified embedding information of the first node based on the supernode information and the embedding information using a GCN model.
[0013] The operation circuit may be further configured to obtain a first result by performing a first aggregation on the embedded information, the supernode information, and the embedded information, and obtain a second result by performing a second aggregation on the embedded information and additional supernode information of neighboring supernodes connected to the supernode.
[0014] The operation circuit may be further configured to correct the first result based on correction information, wherein the correction information indicates a difference between a connection relationship of the compressed graph and a connection relationship of the graph.
[0015] The operation circuit may be further configured to: when the first edge is removed from the compressed graph, add additional embedding information of one or more nodes on the first edge to the first result.
[0016] The operation circuit may be further configured to: when a second edge is added to the compressed graph, subtract the embedding information of one or more nodes on the second edge from the first result.
[0017] The operational circuitry may be further configured to determine whether the neighboring supernode has a self-edge, wherein the second aggregation is based on the determination.
[0018] The operating circuit may be further configured to determine the embedded information of the first node based on the first result and the second result.
[0019] The determined embedding information may correspond to the embedding information represented when the first node is connected to the graph.
[0020] The processing apparatus may be included in a near memory processing (PNM) device.
[0021] According to one aspect, a method for training a GCN model is provided, the method comprising: compressing a graph to obtain a compressed graph, wherein the compressed graph may include supernodes representing multiple nodes of the graph and hyperedges representing multiple edges of the graph; performing aggregation based on the supernodes and hyperedges of the compressed graph; and correcting a result of the aggregation based on correction information, wherein the correction information indicates a difference between a connection relationship of the compressed graph and a connection relationship of the graph.
[0022] The step of performing the aggregation may include: obtaining embedding information of a node in the plurality of nodes in a supernode; determining supernode information of the supernode based on the embedding information; and updating the supernode information based on a hyperedge.
[0023] The step of correcting the result of the aggregation may include adding additional embedding information to the supernode information when correction information indicates that an edge of the graph is removed from the compressed graph.
[0024] The step of correcting the result of the aggregation may include: when the correction information indicates that the edge is added to the compressed graph, subtracting the embedding information of one or more nodes on the second edge from the supernode information.
[0025] According to one aspect, a method is provided, the method comprising: obtaining embedding information of a node of a graph; compressing the graph to obtain a compressed graph, wherein the graph is compressed by grouping a plurality of nodes of the graph to form a supernode of the compressed graph; and generating modified embedding information of the node based on the embedding information and the compressed graph using a graph convolutional network (GCN) model.
[0026] The method further includes iteratively updating the compressed graph by repeatedly calculating a memory requirement of the compressed graph and grouping additional nodes of the graph if the memory requirement exceeds a memory capacity.
[0027] The method for compressing the graph further includes: performing homogeneity-based node partitioning, and performing node merging of the plurality of nodes according to the homogeneity-based node partitioning.
[0028] Additional aspects of the embodiments will be set forth in part in the description which follows and, in part, will be apparent from the description, or may be learned by practice of the disclosure.
[0029] Embodiments may enable large-scale GCN training in hardware environments with large constraints on memory capacity.
[0030] Embodiments may provide operations that maximize data reuse of GCN operations based on compressed graphs, thereby improving the efficiency of large-scale GCN processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] These and / or other aspects, features and advantages of the invention will become clear and more easily understood based on the following description of the embodiments in conjunction with the accompanying drawings:
[0032] Figure 1 and Figure 2 A method for training a graph convolutional network (GCN) model according to an embodiment is shown;
[0033] Figures 3 to 6 illustrates a graph compression process according to an embodiment;
[0034] Figure 7 and Figure 8 A compression graph-based aggregation method according to an embodiment is shown;
[0035] Figures 9 to 11 An example of performing inference using a GCN model using a processing device is shown;
[0036] Fig.12 is a block diagram illustrating a dual in-line memory module (DIMM) according to an embodiment;
[0037] Fig.13 is a block diagram illustrating a near memory processing (PNM) module according to an embodiment;
[0038] Fig.14 is a flowchart illustrating an operating method of a processing device according to an embodiment; and
[0039] Fig.15 is a block diagram showing an example of a configuration of a processing device according to an embodiment.
[0040] Fig.16 is a flowchart illustrating an operating method of a processing device according to an embodiment. DETAILED DESCRIPTION
[0041] The present disclosure describes systems and methods for data processing. Disclosed embodiments include a method for training a large-scale graph convolutional network (GCN) based on a compressed graph. In some cases, the compressed graph includes multiple supernodes and multiple superedges generated by grouping nodes of the graph. One or more embodiments include adding new nodes based on aggregating supernodes and / or superedges using a GCN model.
[0042] Existing industrial-grade large-scale graphs (e.g., social graphs, network (web) graphs, etc.) include large-scale (e.g., billion-scale) nodes and large-scale (e.g., trillion-scale) edges. Therefore, such graphs may have a huge size ranging from hundreds of gigabytes (GB) to tens of terabytes (TB). Training such large-scale graphs may require a large amount of memory. Therefore, existing large-scale graph convolutional network (GCN) model training techniques need to expand the memory space to meet the training requirements. In addition, such techniques use multiple physically independent memories.
[0043] In the case of existing methods for learning large-scale graphs (such as network graphs or social network service (SNS) graphs) that require large memory capacity, the graph may be partitioned into multiple computing nodes or devices, and then the partitioning operation is performed on the subgraphs using a storage device. Such operations result in a large amount of data communication in GCN training. In some cases, the GCN training method may not consider the edges of the input graph that include a significant proportion.
[0044] Embodiments of the present disclosure include a method for training a GCN model, which is designed to overcome the limitations of conventional computing architectures when processing large-scale graph data. According to an embodiment, the training method considers the provided memory size and compresses the input graph based on node embeddings and edges, thereby minimizing operations using the characteristics of the compressed graph.
[0045] In some cases, the approach enables efficient processing of graph data by integrating processing units directly within or in close proximity to memory modules. As a result, the PNM approach significantly improves the performance and scalability of GCN training and inference tasks by reducing latency and improving energy efficiency.
[0046] The present disclosure describes a method for training a large-scale GCN using a compressed graph. In some cases, a compressed graph can be generated by compressing a large-scale graph including multiple nodes and multiple edges using a lossless compression method. In some cases, the generated compressed graph includes super nodes and super edges. For example, super nodes are generated based on grouping nodes of the large-scale graph.
[0047] According to an embodiment, the method for training the GCN includes performing an aggregation operation based on nodes included in each supernode of the compressed graph. For example, the aggregation operation may be referred to as supernode aggregation. In some cases, supernode items that can be iteratively used in the process of training the CGN model may be generated. In some cases, the method for training the GCN also includes an aggregation operation based on a hyperedge and actively using the supernode items. For example, the aggregation operation may be referred to as hyperedge aggregation.
[0048] In addition, after aggregation, addition or deletion of (remaining) correction edges is performed to obtain the final value. In some cases, the large-scale GCN training method can be divided into: a method of storing the entire graph in a host memory with a large capacity, a method of receiving a subgraph to be calculated from the host memory and processing the subgraph, and a method of distributing the graph to multiple devices and processing the graph.
[0049] Therefore, by using a compressed graph as input and using supernode items to perform multi-step aggregation operations including supernode aggregation and superedge aggregation, the embodiments of the present disclosure can effectively reduce the number of operations. In addition, by combining such graph-based computations, the embodiments can effectively handle the computational requirements of GCNs that lead to extended training time and limited scalability.
[0050] Embodiments of the present disclosure include a data processing method of a processing device, the method comprising obtaining embedding information of a first node to be added to a graph and connection information between the graph and the first node. The method also includes receiving supernode information of a supernode of a compressed graph corresponding to the graph, wherein the supernode includes multiple nodes from the graph. In addition, a graph convolutional network (GCN) model generates modified embedding information of the first node based on the supernode information, the embedding information, and the connection information.
[0051] Therefore, a method is provided, the method comprising obtaining embedding information of a node of a graph, and then compressing the graph to obtain a compressed graph. For example, the graph is compressed by grouping multiple nodes of the graph to form a supernode of the compressed graph. An embodiment includes a graph convolutional network (GCN) model that generates modified embedding information of nodes based on the embedding information and the compressed graph.
[0052] The following detailed structural or functional description is provided only as an example, and various changes and modifications may be made to the embodiments. Here, the embodiments should not be interpreted as limited to the disclosure, and should be understood to include all changes, equivalents and replacements within the spirit and technical scope of the disclosure.
[0053] Terms such as first, second, etc. may be used herein to describe components. Each of these terms is not used to define the nature, order, or sequence of the corresponding component, but is only used to distinguish the corresponding component from other components. For example, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.
[0054] It should be noted that if a first component is described as being “connected,” “coupled” or “engaged” to a second component, although the first component may be directly connected, coupled or engaged to the second component, a third component may be “connected,” “coupled” or “engaged” between the first and second components.
[0055] Unless the context clearly indicates otherwise, as used herein, the singular also includes the plural. It will also be understood that when used herein, the terms "include" and / or "comprise" indicate the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0056] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those commonly understood by those skilled in the art to which the present disclosure belongs. Unless explicitly defined as such herein, terms (such as those defined in commonly used dictionaries) should be interpreted as having a meaning consistent with their contextual meaning in the relevant art and should not be interpreted in an idealized or overly formal sense.
[0057] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. When describing the embodiments with reference to the accompanying drawings, the same reference numerals refer to the same elements, and any repeated description related to the same reference numerals will be omitted.
[0058] Figure 1 and Figure 2 A diagram illustrating a method of training a graph convolutional network (GCN) model according to an embodiment.
[0059] According to an embodiment, the method for training a GCN model may be performed by a training device (or a computing device). The GCN model may be a large-scale GCN, but is not limited thereto.
[0060] A machine learning model includes machine learning parameters (also called model parameters or weights), which are variables that provide the behavior and characteristics of the machine learning model. Machine learning parameters can be learned or estimated from training data and used to make predictions or perform tasks based on patterns and relationships learned in the data. Machine learning parameters are typically adjusted during the training process to minimize a loss function or maximize a performance metric. The goal of the training process is to find the optimal values of the parameters that allow the machine learning model to make accurate predictions or perform well for a given task.
[0061] For example, during the training process, the algorithm adjusts machine learning parameters according to an optimization technique (such as gradient descent, stochastic gradient descent, or other optimization algorithms) to minimize the error or loss between the predicted output and the actual target. Once the machine learning parameters are learned from the training data, the machine learning parameters are used to make predictions on new, unseen data.
[0062] Artificial neural networks (ANNs) have many parameters, including weights and biases associated with each neuron in the network, which control the degree of connectivity between neurons and affect the ability of the neural network to capture complex patterns in data. An ANN is a hardware or software component that includes multiple connected nodes (i.e., artificial neurons) that broadly correspond to neurons in the human brain. Each connection or edge sends a signal from one node to another (similar to a physical synapse in the brain). When a node receives a signal, it processes the signal and then sends the processed signal to other connected nodes.
[0063] In some cases, the signals between nodes include real numbers, and the output of each node is calculated by a function that sums its inputs. In some examples, nodes may determine their outputs using other mathematical algorithms (such as selecting the maximum value from the inputs as the output) or any other suitable algorithm for activating nodes. Each node and each edge is associated with one or more node weights that determine how to process and send signals.
[0064] In an ANN, a hidden (or intermediate) layer includes hidden nodes and is located between the input layer and the output layer. The hidden layer performs a nonlinear transformation of the input entering the network. Each hidden layer is trained to produce a defined output that contributes to the joint output of the output layer of the ANN. The hidden representation is a machine-readable data representation of the input learned from the hidden layer of the ANN and produced by the output layer. As the ANN's understanding of the input improves as the ANN is trained, the hidden representation gradually distinguishes from earlier iterations.
[0065] During the training process of an ANN, node weights are adjusted to improve the accuracy of the results (i.e., by minimizing the loss that corresponds in a particular way to the difference between the current result and the target result). The weights of the edges increase or decrease the strength of the signal sent between the nodes. In some cases, the nodes have a threshold below which no signal is sent at all. In some examples, the nodes are aggregated into layers. Different layers perform different transformations on their inputs. The initial layer is called the input layer, and the last layer is called the output layer. In some cases, the signal traverses a particular layer multiple times.
[0066] A GCN is a neural network that defines convolutional operations on graphs and uses the structural information of the graph. For example, GCN can be used for node classification (e.g., literature) in a graph (e.g., citation network), where labels are available for a subset of nodes using a semi-supervised learning approach. The feature description of each node is summarized in a matrix and a form of pooling operation is used to produce node-level outputs. In some cases, GCN uses enriched representation vectors for aspect terms and searches for dependency trees for the sentiment polarity of the input phrase / sentence.
[0067] The training or computing device for GCN can use dedicated hardware optimized for the unique computing requirements of graph-based neural networks. The device is characterized by high parallelism to efficiently process large-scale graph data and perform iterative graph convolution. The device integrates advanced memory management and data caching techniques to handle the irregularity and sparsity of the graph structure. In addition, the device may include dedicated accelerators for graph-specific operations (such as node aggregation and message passing). In some cases, the device facilitates rapid training and inference of GCN, thereby enabling applications in various fields (such as social network analysis, recommendation systems, etc.).
[0068] In operation 110, the training device may compress the graph. The training device may compress the graph to a level that avoids memory shortages in the memory capacity of the hardware. In some examples, the graph may be an industrial-grade graph (e.g., an industrial-scale large-scale graph) (e.g., a social graph, a network graph, etc.). In some examples, the graph may include a plurality of nodes (e.g., a billion-scale node) and a plurality of edges (e.g., a trillion-scale edge). As Figure 2 In the example shown in FIG. 210 , the graph 210 may include 10 nodes v0 to v9 (or v 0 to v 9 ) and 15 edges. However, Figure 2 The diagram 210 shown in FIG. 2 is merely an exemplary diagram for convenience of description, and the embodiment is not limited thereto.
[0069] In some cases, a compressed graph (or summary graph) may include multiple supernodes and multiple hyperedges (in some cases, hyperedges are used to connect supernodes). Figure 3Describes the process of graph compression.
[0070] In the context of graph convolutional networks (GCNs), a "supernode" refers to a central node or a group of nodes within a graph. In some cases, a supernode has a significant impact on the overall structure and dynamics of the graph. In some cases, a supernode may have a high degree of connectivity with many other nodes in the graph. Similarly, a "hyperedge" as described herein represents an edge connecting nodes (e.g., supernodes) within a graph. Hyperedges can represent relationships or dependencies between nodes, significantly contributing to the overall connectivity and information flow within the graph. When training GCNs, analyzing hyperedges requires attention because hyperedges have the potential to shape network behavior and learn dynamics.
[0071] In operation 120 , the training device may train a GCN model based on the compressed graph.
[0072] According to an embodiment, in operation 120, the training device may perform aggregation on the embedding information (e.g., embedding matrix) of the nodes belonging to each supernode of the compressed graph. For example, the aggregation may include an operation of summing the given objects (e.g., the embedding information of the nodes belonging to the supernode). For example, if the given objects are A and B, the aggregation may include an operation of determining (or calculating) A+B by summing A and B.
[0073] Aggregation of embedding information of nodes belonging to supernodes in the lth layer of the GCN model (e.g., embedding information output by the (l-1)th layer of the GCN model) by the training device may be referred to as "supernode aggregation". The training device may determine supernode information of each supernode using the supernode aggregation in the lth layer of the GCN model. As used herein, supernode information of a supernode may represent an aggregation result of embedding information of nodes belonging to the supernode.
[0074] In some cases, the training device may determine iteratively used (or reused) supernode information (hereinafter referred to as "supernode item") among the supernode information of multiple supernodes. The training device may perform aggregation based on the hyperedges of each supernode in the lth layer of the GCN model, and such aggregation may be referred to as "hyperedge aggregation". Therefore, the training device may use the supernode item while performing hyperedge aggregation, and thus may reduce the number of operations when training the GCN model and improve the training speed.
[0075] Existing compression methods for GCN models compress the node embedding matrix. However, such methods fail to compress the GCN side information. In the case of GCN, the side information occupies a larger proportion than the node embedding matrix. Therefore, even if the node embedding matrix is compressed, the overall compression rate of GCN may be low. According to an embodiment of the present disclosure, a training device may obtain a compressed graph by compressing the node embedding matrix (i.e., the embedding information of the node) and the side information of FIG210, and train a GCN model based on the compressed graph. Therefore, the memory requirement problem and data communication overhead that may occur when training a large-scale GCN model may be eliminated. Reference will be made to Figure 7 and Figure 8 Describe the training method of the GCN model.
[0076] Figures 3 to 6 A graph compression method according to an embodiment is shown. Figure 3 The steps in the above correspond to executing Figure 1 The described graph is compressed (ie, operation 110).
[0077] Reference Figure 3 In operation 310, the training device may perform lossless edge compression on the graph (or input graph) 210. As described herein, lossless edge compression may be an operation that reduces the number of edges in the graph 210. Figure 4 A flowchart for performing lossless edge compression is shown in FIG.
[0078] Lossless edge compression may refer to a data compression technique that reduces the size of data without losing any information. Specifically, in the context of a graph or network, lossless edge compression aims to reduce the storage space required to represent the edges of the graph while preserving all connection information. In some cases, lossless edge compression may be implemented using, but not limited to, run-length encoding (RLE), data encoding, or variable-length encoding (VLE) methods. By reducing the size of edge data while preserving connection information, lossless edge compression enables more efficient storage, transmission, and accurate processing of graph data.
[0079] Figure 4 Shown as Figure 3 The method for performing lossless edge compression is described in operation 310 of FIG. Figure 4 In operation 410, the training device may perform node partitioning based on homogeneity on the graph 210. In some cases, the graph 210 may correspond to a data set with homogeneity. As used herein, homogeneity refers to the property that nodes with similar tendencies are connected to each other. In some cases, homogeneity refers to the tendency that nodes with similar attributes or characteristics are more frequently connected to each other or form links than nodes with different attributes.
[0080] Therefore, the training device may group the nodes in the graph 210 based on the homogeneity of the graph 210. The training device may compress the graph 210 by dividing the nodes in the graph 210 into a plurality of groups, thereby reducing the time for compressing the graph 210.
[0081] For example, nodes v0 to v9 in the graph 210 may have corresponding (e.g., different) labels. The training device may group nodes v0 to v9 by category based on the corresponding labels of the nodes v0 to v9 in the graph 210. Nodes with the same category may belong to the same group.
[0082] According to an embodiment, some of the formed groups having a size greater than or equal to a predetermined level are referred to as super-large groups. In some cases, the training device may divide the super-large group based on the distance (or similarity) between the embedding information (e.g., embedding matrix) corresponding to each of the plurality of nodes in the super-large group.
[0083] In operation 420, the training device may perform node merging on the formed groups. In some cases, the training device may determine the node pairs that minimize the cost among the nodes in each formed group by performing a greedy algorithm on each formed group. The greedy algorithm is used to solve the optimization problem by making a local optimal choice at each step in the hope of finding a global optimal solution. The greedy algorithm iteratively establishes a solution by selecting the best available option at each stage without reconsidering the previous choice. For example, the greedy algorithm can be applied to tasks such as finding a minimum spanning tree, finding the shortest path, or solving a vertex cover problem.
[0084] In some cases, the cost may be, for example, |P| + |C + |+|C - Here, |P| may represent the number of hyperedges. In some cases, there may be edges (hereinafter referred to as “C edges”) that are represented (or defined) in the graph 210 but not represented (or defined) in the compressed graph. + edge”).|C + | can represent C + In some cases, there may be edges (hereinafter referred to as “C - edge”).|C - | can represent C - The training device may determine the number of hyperedges in each formed group, C + The number of edges and C - The node pairs for which the sum of the number of edges is minimized.
[0085] In operation 430, the training device may convert the graph 210 into a compressed form based on the determined node pair. For example, the training device may determine the determined node pair as a supernode (hereinafter referred to as an "initial supernode"). The training device may generate a compressed graph by connecting the determined initial supernodes. In the case of the first iteration (e.g., iteration=1), the compressed graph generated by connecting the initial supernodes is referred to as a "compressed graph". Figure 1 The training device can be found in FIG. 210 but in the compressed Figure 1 The edges not represented in (hereinafter, C 1 + side), and C 1 + Add edges to compression Figure 1 The training device may be found not shown in FIG. 210 but in the compressed Figure 1 The edge represented by (hereinafter referred to as C 1 - side).
[0086] In a second iteration (eg, iteration=2), the training device may perform operations 410 to 430. The training device may perform the following operations: Figure 1 Perform node partitioning based on homogeneity to form a plurality of groups, and perform node merging on the formed groups. Each formed group may include an initial super node. The training device may determine a super node pair that minimizes the cost between the initial super nodes in each formed group. The training device may compress the super node pair based on the determined super node pair. Figure 1 The training device may iterate operations 410 to 430 to group the nodes in the graph 210 .
[0087] Figure 5 Nodes v0 to v9 in graph 210 are shown to be grouped using an iterative process. Based on the iteration of operations 410 to 430 to determine supernode pairs for cost minimization, nodes v0 to v2 may form a first group 510, nodes v3 and v4 may form a second group 520, node v5 may be a third group 530, and nodes v6 to v9 may form a fourth group 540. Groups 510 to 540 may correspond to supernodes, and the training device may determine (or generate) a compressed graph based on groups 510 to 540. In some cases, the training device may represent (or add) edges (e.g., C edges) that are represented in graph 210 but not represented in the compressed graph in (or to) the compressed graph. + In addition, the training device may find edges that are not represented in the graph 210 but are represented in the compressed graph (e.g., C - side).
[0088] Refer again Figure 3In operation 320, the training device may determine (or estimate) the memory requirement based on the GCN model and the configuration of the hardware (eg, memory). For example, the training device may determine the memory requirement using Equation 1.
[0089] [Equation 1]
[0090] bit in |V|d 0 +bit inter |V|∑d 1 +bit edge (|P|+|C + |+|C - |)
[0091] As used herein, d 0 The dimension of the output information (e.g., output embedding matrix) of the input layer of the GCN model can be represented, and d 1 It can represent the dimension of the output information (e.g., output embedding matrix) of the lth layer of the GCN model. in bit may represent the bit precision (e.g., 32 bits, etc.) of the input embedding information (e.g., the embedding matrices of nodes v0 to v9 in graph 210) provided to the input layer. inter The bit precision (e.g., 32 bits, etc.) of the intermediate embedding information (e.g., the intermediate embedding matrix) corresponding to the operation result from the intermediate layer (e.g., the input layer to the (l-1)th layer) of the GCN model can be represented. edge The bit precision with which a hyperedge can be represented.
[0092] In Equation 1, |V| may represent the number of nodes in the graph 210, and |P| may represent the number of hyperedges in the compressed graph. + | can represent C + The number of edges, and |C - | can represent C - The number of edges.
[0093] In operation 330, the training device may determine whether the memory requirement is greater than the memory capacity. The training device may determine whether the memory capacity of a given environment is sufficient to meet the memory requirement of the training (eg, the memory requirement calculated using Equation 1).
[0094] In the case where the memory requirement is greater than the memory capacity, the training device may perform node embedding compression in operation 340. Node embedding compression may be compression of the embedding information (e.g., embedding matrix) of each node in the graph 210. In some cases, node embedding compression refers to a technique for reducing the dimensionality or storage requirement of node embeddings while preserving the information content of the node embeddings. For example, node embedding compression may be performed using various techniques (including but not limited to quantization, dimensionality reduction, clustering-based compression, sparse representation, etc.).
[0095] According to one embodiment, the training device may in Reduce to a level where the memory requirements match the memory capacity. Training devices can reduce bit in , so that bit in Keep 2 k For example, in bit in If the bit is 32, the training device may set bit in In addition, because errors in the input embedding information can have a greater impact on the accuracy of the GCN model than errors in the intermediate embedding information, the training device can ensure that the bit in ≥bit inter For example, if the bit in operation 320 in and bit inter Each is 32 bits, then the training device can convert bit in From 32 bits to 16 bits. The training device can inter Reduce from 32 bits to 16 bits to meet the bit in ≥bit inter .
[0096] In operation 350, the training device may perform lossy edge compression. In some cases, lossy edge compression is a data compression method that reduces the size of graph edge data to achieve a higher compression ratio by sacrificing some information (usually non-critical or redundant details). Lossy compression intentionally discards certain data in the compression process. The compression method includes transforming the edge data in a way that preserves necessary features and structure while minimizing storage requirements. Lossy edge compression can be used when a certain loss of fidelity is acceptable, such as in large-scale graph storage, transmission, or processing, where reducing data size takes precedence over maintaining absolute accuracy.
[0097] Reference Figure 4 The details of operation 350 are further described. That is, the training device may perform the following steps in operation 350: Figure 4 That is, the training device can remove one or more C from the compression graph at this time.+ Perform lossy edge compression.
[0098] Figure 6 An example of a compressed graph generated using graph compression in operation 110 is shown. Figure 6 In the example shown in FIG. 6 , the compressed graph 610 may include supernodes (eg, represented by S) A0 to A3, hyperedges (eg, represented by P) 621 to 624, C + Edge 605 and C - Edge 607. In this context, A0 to A3 may also be represented as A 0 To A 3 .exist Figure 6 In the example shown in , supernode A0 may include nodes v0, v1, and v2 of graph 210, and supernode A1 may include nodes v3 and v4 of graph 210. Supernode A2 may include node v5 of graph 210, and supernode A3 may include nodes v6, v7, v8, and v9 of graph 210.
[0099] As Figure 6 In the example shown in , supernode A0 may have a self-edge (or self-loop) 621, and supernode A3 may have a self-edge (or self-loop) 624. In certain cases, a self-edge or self-loop refers to a connection that originates from a node and terminates at the same node. A self-edge represents a relationship or interaction between a node and itself, representing a self-referential characteristic or property.
[0100] According to one embodiment, the training device may partition the graph 210. For example, the training device may use two or more physically independent memories. In this case, the training device may place nodes of the same category in the same memory, thereby minimizing data communication between nodes of the same category.
[0101] Figure 7 and Figure 8 A compressed graph-based aggregation method according to an embodiment is shown.
[0102] Reference Figure 7 , GCN model training may include super-node aggregation, super-edge aggregation, and correction of the results of super-edge aggregation. For example, GCN model training refers to Figure 1 Model training described in operation 120.
[0103] In operation 710, the training device may perform super-node aggregation in a layer (eg, a convolutional layer) of the GCN model. Figure 8As described in detail, the training device may perform supernode aggregation based on the embedding information of the nodes belonging to the supernodes in the lth layer of the GCN model. In some examples, the supernode aggregation may be performed based on the embedding matrix of the nodes output by the (l-1)th layer of the GCN model. The training device may determine the supernode information of the supernodes in the lth layer through supernode aggregation. In some examples, the training device may determine the sum of the embedding matrices of the nodes output by the (l-1)th layer.
[0104] In operation 720, the training device may perform hyperedge aggregation in the layers of the GCN model. Figure 8 As described in detail, the training device may perform hyperedge aggregation based on one or more hyperedges and supernode information of a supernode. By performing hyperedge aggregation based on one or more hyperedges and supernode information of a supernode, the embedding information of the node (e.g., the embedding matrix of the node output by the (l-1)th layer) may be updated in the lth layer.
[0105] In operation 730, the training device may + Edge and / or C - For example, the training device can send C + The updated embedding information of one of the nodes on the edge is added to form C + The training device can form C - The updated embedding information of one of the nodes on the edge is subtracted to form C - The embedding information of another node in the edge node (for example, the embedding information output by the (l-1)th layer). Figure 8 Further details are described regarding correction of the results of performing hyperedge aggregation.
[0106] Equation 2 below shows the embedding information of node v determined by operations 710 to 730 in the lth layer:
[0107] [Equation 2]
[0108]
[0109] As used herein, S(v) may represent the supernode to which node v belongs, C + (v) can represent the C of node v + Edge, C - (v) can represent the C of node v - In Equation 2, N(S(v)) may represent the neighboring supernodes connected to the supernode to which the node v belongs, and Si may represent the i-th supernode among neighboring supernodes.
[0110] Will refer to Figure 8 Further details and examples corresponding to operations 710 through 730 are described.
[0111] Figure 8 A compressed graph (e.g., compressed graph 610) generated using graph compression (e.g., a graph compression operation corresponding to operation 110) is depicted. Figure 8 As shown in , supernode A0 may include nodes v0, v1, and v2, supernode A1 may include nodes v3 and v4, supernode A2 may include node v5, and supernode A3 may include nodes v6, v7, v8, and v9.
[0112] exist Figure 8 In operation 710 of the GCN model, the training apparatus may determine the supernode information (i.e., Figure 8 Similarly, the training apparatus may determine the supernode information ( A0←v0+v1+v2) of the supernode A1 in the lth layer of the GCN model by summing the embedding information of nodes v3 and v4 output by the (l-1)th layer of the GCN model. Figure 8 The training device may determine the embedding information of the node v5 output by the (l-1)th layer of the GCN model as the supernode information of the supernode A2 in the lth layer of the GCN model ( Figure 8 The training apparatus may determine the supernode information (A2←v5) of the supernode A3 in the lth layer of the GCN model by summing the embedding information of the nodes v6, v7, v8, and v9 output by the (l-1)th layer of the GCN model. Figure 8 A3←v6+v7+v8+v9) as shown in operation 710 of .
[0113] Operation 720 may include performing hyperedge aggregation processing. Figure 8 As shown in operation 720 of , among the supernode information of supernodes A0, A1, A2, and A3, the supernode information of supernodes A2 and A3 may be iteratively used during operation processing (e.g., superedge aggregation processing). The training device may determine the supernode information of supernodes A2 and A3 as supernode items. The training device may iteratively use the supernode information of supernodes A2 and A3 corresponding to the supernode items. Therefore, the number of operations corresponding to the number of iterations (or reuses) of using the supernode items may be reduced, and the training speed of the GCN model may be increased.
[0114] exist Figure 8In operation 720, since the super node A0 forms a self-loop (or self-edge), the training device may update (or determine) the embedding information of nodes v0, v1, and v2 to the super node information ( Figure 8 v0′,v1′,v2′←A0).
[0115] Since supernode A1 is connected to supernode A2 through a hyperedge, the training device can determine the sum of the embedding information of node v3 output by the (l-1)th layer and the supernode information of supernode A2 as the embedding information of node v3 in the lth layer. Therefore, the training device can update the embedding information of node v3 to the sum of the embedding information of node v3 output by the (l-1)th layer and the supernode information of supernode A2 ( Figure 8 Similarly, the training device may update the embedding information of node v4 to the sum of the embedding information of node v4 output by the (l-1)th layer and the supernode information of supernode A2 ( Figure 8 v4′←v4+A2).
[0116] Since supernode A2 is connected to each of supernodes A1 and A3 through a hyperedge, the training device may determine the sum of the embedding information of node v5 output by the (l-1)th layer, the supernode information of supernode A1, and the supernode information of supernode A3 as the embedding information of node v5 in the lth layer. Therefore, the training device may update the embedding information of node v5 to the sum of the embedding information of node v5 output by the (i-1)th layer, the supernode information of supernode A1, and the supernode information of supernode A3 ( Figure 8 v5'←v5+A1+A3).
[0117] Since supernode A3 forms a self-loop and is connected to supernode A2 through a hyperedge, the training device may update the embedding information of nodes v6, v7, v8, and v9 to the sum of the supernode information of supernodes A2 and A3 determined in the lth layer of the GCN model ( Figure 8 v6', v7', v8', v9'←A3+A2).
[0118] In operation 810, the training device may + The result of applying the edge to the hyperedge aggregation. Figure 8 In the example shown in + The edge may indicate a connection between node v2 and node v3 (also as Figure 2 The training device may add the embedding information of node v3 output by the (l-1)th layer to the updated embedding information of node v2 (in Figure 8In operation 810, v2'←v2'+v3). The training device may add the embedding information of node v2 output by the (l-1)th layer to the updated embedding information of node v3 (in Figure 8 In operation 810, v3′←v3′+v2).
[0119] In operation 820, the training device may apply C-edges to the results of the hyperedge aggregation. Figure 8 In the example shown in - The edge may indicate that node v7 and node v9 are not connected in graph 210 (e.g. Figure 2 The training device may subtract the embedding information of node v9 output by the (l-1)th layer (indicated in FIG. 210 of FIG. 211 ) from the updated embedding information of node v7. Figure 8 In operation 820, v7'←v7'-v9). The training device may subtract the embedding information of node v7 output by the (l-1)th layer from the updated embedding information of node v9 (in Figure 8 In operation 820, v9'←v9'-v7).
[0120] Figure 8 Operations 810 and 820 in may include Figure 7 Operation 730 or corresponding to Figure 7 Operation 730. The training device may train the GCN model by iteratively performing operations 710, 720, and 730 on each of the multiple layers of the GCN model.
[0121] Figures 9 to 11 An example of performing inference using a GCN model by a processing device is shown.
[0122] According to an embodiment, the processing device 910 may correspond to a deep learning accelerator. For example, the processing device 910 may be a graphics processing unit (GPU) or a neural processing unit (NPU), but is not limited thereto.
[0123] Reference Fig. 9 According to an embodiment, the processing device 910 may "add to a large-scale graph (e.g., Figure 2 The connection information of the new node in the graph 210 (in the GCN model 920), the embedding information of the new node (e.g., the embedding matrix), and the compressed graph 610″ are input into the GCN model 920. In some cases, the input connection information may refer to the connection information about the nodes of the graph to which the new node will be connected. The compressed graph 610 may include, for example, supernode information of supernodes A0, A1, A2, and A3, hyperedges, C + Edge and C - In some examples, based on the shape of the compressed graph 610 (e.g. Figure 6 and Figure 8 described in ), C + Edge and / or C- The edges may not be included in the compressed graph 610. The GCN model 920 may be based on the reference Figures 1 to 8 The GCN model trained by the GCN model training method described in detail.
[0124] Processing device 910 may obtain an inference result from GCN model 920. For example, the inference result may include, but is not limited to, embedding information according to a connection of a new node to graph 210 (or compressed graph 610).
[0125] In the following, reference will be made to Fig.10 and Fig.11 The operation of the processing device 910 is described.
[0126] Fig.10 The new node v11 1010, the compressed graph 610 and the GCN model 920 are shown. The GCN model 920 may include multiple layers. For ease of description, it is assumed that the GCN model 920 includes Fig.10 The processing device 910 may perform super node aggregation and / or super edge aggregation (as shown in FIG. 1 ) through each of the layers 921 and 922. Figures 3 to 8 described).
[0127] The node v11 1010 may include a node v11 1010 indicating that the node v11 1010 is to be connected to the graph 210 (see Figure 2 Each of the nodes v5 and v8 described above may be referred to as a target node of the node v11 1010 .
[0128] Processing equipment (e.g., see Fig. 9 The processing device 910 described above may determine the embedding information of nodes v5 and v8 as input data of the GCN model 920.
[0129] The processing device 910 may determine input data of the GCN model 920 based on the hyperedge of the supernode to which the node to be connected to the node v11 1010 belongs and the connection information of the node v11 1010. For example, the processing device 910 may identify the supernode A2 to which the node v5 to be connected to the node v11 1010 belongs, and determine the supernode information of the supernodes A1 and A3 forming the hyperedge with the identified supernode A2 as the input data of the GCN model 920. The processing device 910 may identify the supernode A3 to which the node v8 to be connected to the node v11 1010 belongs.
[0130] In addition, since the identified supernode A3 forms a self-loop (or self-edge), the processing device 910 may determine the supernode information of the supernode A3 as input data. In some cases, since the identified supernode A3 forms a hyperedge with the supernode A2, the processing device 910 may determine the supernode information of the supernode A2 as input data of the GCN model 920. The processing device 910 may determine the embedding information of the nodes v5 and v8 to which the node v11 1010 is to be connected as input data of the GCN model 920.
[0131] Therefore, if Fig.10 As shown in the example of , the embedding information of node v11 1010, the embedding information of node v8, the supernode information of supernode A3, the supernode information of supernode A2, and the embedding information of supernode A1 may be input to the first layer 921. Since supernode A2 only includes node v5, the embedding information of node v5 may be the same as the supernode information of supernode A2. Fig.10 As shown in, but C + Edge and C - Edges may be input into the first layer 921 .
[0132] In the case where the node v11 1010 is added to the graph 210, the processing device 910 may recognize (or predict) that the supernode information of the supernodes A2 and A3 is iteratively used (or reused) in the operation of the GCN model 920. The processing device 910 may store the supernode information of the supernodes A2 and A3 in a buffer (e.g., to be referred to later). Fig.13 When the supernode information of supernodes A2 and A3 is used as an operand, the processing device 910 may receive the supernode information of supernodes A2 and A3 from the buffer. Therefore, the processing device 910 may reduce access to a dynamic random access memory (DRAM) and improve operation speed.
[0133] Examples of memory devices include random access memory (RAM), read-only memory (ROM) or hard disk. Examples of memory devices include solid-state memory and hard disk drive. In some examples, the memory is used to store computer-readable software, computer-executable software including instructions, which, when executed, cause the processor to perform various functions described herein. In some cases, the memory further includes a basic input / output system (BIOS), which controls basic hardware or software operations (such as, interaction with peripheral components or devices). In some cases, a memory controller operates a memory unit. For example, a memory controller may include a row decoder, a column decoder or both. In some cases, the memory unit in the memory stores information in the form of a logical state.
[0134] Dynamic random access memory (DRAM) is a type of semiconductor memory that stores data in cells composed of capacitors and transistors. DRAM offers high-density storage and fast access times, making it widely used in computer systems, mobile devices, and other electronic devices. The key advantage of DRAM is its ability to store data dynamically, requiring periodic refreshes to maintain the stored information. DRAM technology continues to evolve, and developments focus on increasing storage density, reducing power consumption, and improving memory access speeds, making DRAM a valuable subject for patent protection in the field of semiconductor memory technology.
[0135] The processing device 910 may update (or determine) the embedding information of the node v11 1010 and the embedding information of the nodes v5 and v8 to be connected to the node v11 1010 through the first layer 921. For example, the processing device 910 may update the embedding information of the node v11 1010 by summing the embedding information of the node v11 1010, the embedding information of the node v8, and the embedding information of the node v5 (=supernode information of supernode A2). In other words, v11'←v11+v8+A2.
[0136] Node v8 should form a connection with node v11 1010, and supernode A3 to which node v8 belongs has a self-edge, and supernode A3 is connected to supernode A2 (= node v5). Based on this connection relationship, processing device 910 can update the embedding information of node v8 by summing the embedding information of node v11 1010, the supernode information of supernode A3, and the embedding information of node v5 (= supernode information of supernode A2). In other words, v8'←v11+A3+A2.
[0137] Node v5 should form a connection with node v11 1010, and supernode A2 to which node v5 belongs is connected to each of supernodes A1 and A3. Based on this connection relationship, processing device 910 can update the embedding information of node v5 by summing the embedding information of node v11 1010, the supernode information of supernode A3, the embedding information of node v5 (=supernode information of supernode A2), and the supernode information of supernode A1. In other words, A2'←v11+A3+A2+A1.
[0138] The first layer 921 may output updated embedding information (hereinafter referred to as “intermediate embedding information”) of nodes v5 , v8 , and v11 .
[0139] The processing device 910 may update the intermediate embedding information of the node v11 by summing the intermediate embedding information of the nodes v5, v8, and v11. Therefore, v11″←v11'+v8'+A2'. The second layer 922 may output the updated intermediate embedding information of the node v11. The processing device 910 may obtain the intermediate embedding information output by the second layer 922 as an inference result.
[0140] Fig.11 New node v12 1110 , compressed graph 610 , and GCN model 920 are shown.
[0141] Node v12 1110 may have a header indicating that node v12 1110 is to be connected to graph 210 (eg, as shown in FIG. 210 ). Figure 2 Connection information for each of nodes v3 and v8 in Figure 210).
[0142] Processing equipment (such as, for example, Fig. 9 The processing device 910 described above may determine the input data of the GCN model 920 based on the hyperedge of the supernode to which the node to be connected to the node v12 1110 belongs and the connection information of the node v12 1110. For example, the processing device 910 may identify the supernode A1 to which the node v3 to be connected to the node v12 1110 belongs, and determine the supernode information of the supernode A2 forming a hyperedge with the identified supernode A1 as the input data of the GCN model 920.
[0143] In addition, the processing device may identify the supernode A3 to which the node v8 to be connected to the node v12 1110 belongs. Since the identified supernode A3 forms a self-loop (or self-edge), the processing device 910 may determine the supernode information of the supernode A3 as input data. Since the identified supernode A3 forms a hyperedge with the supernode A2, the processing device 910 may determine the supernode information of the supernode A2 as input data of the GCN model 920. The processing device 910 may determine the embedding information of the nodes v3 and v8 to which the node v12 1110 is to be connected as input data of the GCN model 920. Since the node v2 may form a CCN with the node v3 to be connected to the node v12 1110, the processing device 910 may determine the supernode information of the supernode A2 as input data of the GCN model 920. + edge, so the processing device 910 can determine the embedding information of the node v2 as the input data of the GCN model 920.
[0144] As Fig.11 In the example shown in , embedded information of node v12 1110 , embedded information of node v3 , embedded information of node v8 , supernode information of supernode A2 , supernode information of supernode A3 , and embedded information of node v2 may be input to the first layer 921 .
[0145] When node v12 1110 is added to graph 210, processing device 910 may identify (or predict) that supernode information of supernode A2 is iteratively used (or reused) in the operation of GCN model 920. Processing device 910 may store supernode information of supernode A2 in a buffer (e.g., to be referred to later). Fig.13 When the supernode information of the supernode A2 is used as an operand, the processing device 910 may receive the supernode information of the supernode A2 from the buffer. Therefore, the processing device 910 may reduce access to the DRAM and improve the operation speed.
[0146] The processing device 910 may update (or determine) the embedding information of the node v12 1110 and the embedding information of the nodes v3 and v8 to be connected to the node v12 1110 through the first layer 921. For example, the processing device 910 may update the embedding information of the node v12 1110 by summing the embedding information of the node v12 1110, the embedding information of the node v8, and the embedding information of the node v3. In other words, v12'←v12+v8+v3. The node v8 should form a connection with the node v12 1110, and the supernode A3 to which the node v8 belongs includes a self-edge, and the supernode A3 is connected to the supernode A2. Based on the connection, the processing device 910 may update the embedding information of the node v12 1110, the supernode information of the supernode A3 having the self-edge, and the supernode information of the supernode A2. In other words, v8'←v12+A2+A3. Node v3 should form a connection with node v12 1110, and supernode A1 to which node v3 belongs is connected to supernode A2. Based on this connection, processing device 910 may update the embedding information of node v3 by summing the embedding information of node v12 1110, the embedding information of node v3, and the supernode information of supernode A2. In other words, v3'←v3+v12+A2.
[0147] The first layer 921 may output updated embedding information (hereinafter referred to as “intermediate embedding information”) of nodes v12 , v8 , and v3 .
[0148] Since node v3 corresponds to node with C + The processing device 910 can be based on C + For example, the processing device 910 may add the embedding information of the node v2 to the intermediate embedding information of the node v3. In other words, v3'←v3'+v2 (ie, v3'←v3+v12+A2+v2).
[0149] The processing device 910 may update the intermediate embedding information of the node v12 by summing the intermediate embedding information of the nodes v12, v8, and v3. In other words, v12″←v12'+v8'+v3'. The second layer 922 may output the updated intermediate embedding information of the node v12. The processing device 910 may use the GCN model 920 to obtain the intermediate embedding information output by the second layer 922 as an inference result.
[0150] Fig.12 is a block diagram illustrating a dual in-line memory module (DIMM) according to an embodiment.
[0151] DIMMs are standardized modular components used in computer systems to provide additional random access memory (RAM) capacity. DIMMs typically consist of a small printed circuit board with multiple memory chips, connectors, and electrical traces. DIMMs are inserted into dedicated memory slots on the motherboard of a computer or server and provide expansion of memory capacity beyond that integrated into the motherboard. DIMMs come in a variety of form factors and speeds to accommodate different types of computer systems and memory requirements, and play a key role in improving system performance and scalability.
[0152] Thus, the embodiment describes DIMM 1200. Fig.12 As shown in FIG. 1 , a DIMM 1200 may include a buffer chip 1210 and a plurality of memory ranks 1220 and 1221. Each of the plurality of memory ranks 1220 and 1221 may include one or more DRAMs.
[0153] The buffer chip 1210 may include near memory processing (PNM) modules 1210-0 and 1210-1. The processing device 910 described above may be included in at least one of the PNM modules 1210-0 and 1210-1.
[0154] PNM refers to a computing architecture in which processing units are integrated directly into or in close proximity to memory modules. This arrangement provides computing tasks to be executed closer to data storage devices, minimizing data movement and alleviating bandwidth constraints. The PNM architecture can significantly improve the performance and energy efficiency of memory-constrained tasks such as data-intensive analysis, machine learning inference, and graph processing. By reducing the distance between processing and memory, the PNM architecture provides the potential for significant acceleration and improvement of overall system efficiency.
[0155] Reference Fig.12, PNM module 0 1210-0 may correspond to memory rank 0 1220, and PNM module 1 1210-1 may correspond to memory rank 1 1221. For example, a memory rank may refer to a grouping or subdivision of memory within a memory module. A memory rank represents a group of modules that can be accessed independently but share some common signals and resources. In some cases, memory ranks enable efficient organization and management of memory accesses. By dividing memory into memory ranks, the PNM architecture can exploit parallelism and reduce contention, thereby providing multiple memory operations occurring simultaneously, which improves overall system throughput and can improve performance of memory-bound tasks.
[0156] PNM module 0 1210-0 may store supernode information and hyperedge information of a portion of supernodes of the compressed graph 610. PNM module 1 1210-1 may store supernode information and hyperedge information of the remaining supernodes of the compressed graph 610. For example, PNM module 0 1210-0 may store supernode information and hyperedge information of supernodes A0 and A1 in the graph 610. According to an exemplary embodiment, PNM module 0 1210-0 may store information about C + PNM module 1 1210-1 may store information about supernodes A2 and A3 in the compressed graph 610 and information about superedges 605. - Information about edge 607.
[0157] Memory column 0 1220 may store embedding information and connection information of a portion of nodes of graph 210, and memory column 1 1221 may store embedding information and connection information of the remaining nodes of graph 210. For example, memory column 0 1220 may store embedding information and connection information of nodes v0, v1, v2, v3, and v4 of graph 210, and memory column 1 1221 may store embedding information and connection information of nodes v5, v6, v7, v8, and v9 of graph 210.
[0158] Each of PNM module 0 1210-0 and PNM module 1 1210-1 may perform as described in reference to Figures 9 to 11 The operation of processing device 910 is described.
[0159] Fig.13 FIG. 2 shows a near memory processing (PNM) module according to an embodiment. Fig.13 , illustrating multiple components that may be included in the PNM module 1300.
[0160] In some examples, reference Fig.12 Each of the PNM module 0 1210-0 and the PNM module 1 1210-1 described above may correspond to the PNM module 1300. Figures 9 to 11 The described processing device 910 may correspond to the PNM module 1300. The processing device 910 may include at least a part or all of the components of the PNM module 1300.
[0161] The PNM module 1300 may include a first buffer 1310 , a second buffer 1320 , an operation circuit 1330 , a control circuit (or control logic) 1340 , a DRAM controller 1350 , and double data rate (DDR) physical (PHY) interfaces 1360 and 1361 .
[0162] Considering the case when a new node (e.g., node v12) is added to the diagram 210, the PNM module 1300 may receive data required to perform an inference operation on the new node (e.g., node v12) from the DRAM of the memory rank (e.g., memory rank 1 1221) corresponding to the PNM module 1300 via the DDR PHY interface 1360. The received data may include, for example, Fig.11 Described: embedding information / connection information of the new node (e.g., node v12), supernode information / hyperedge information of supernodes A2 and A3, C + Edge information, embedding information / connection information of node v3 to be connected to node v12, and embedding information / connection information of node v8 to be connected to node v12.
[0163] The supernode information of the supernode A2 may correspond to the supernode item, so that the PNM module 1300 may store the supernode information of the supernode A2 in the first buffer 1310. When the supernode item (e.g., the supernode information of the supernode A2) is used as an operand of an operation (e.g., by the operation circuit 1330), the operation circuit 1330 may receive the supernode item from the first buffer 1310. Therefore, the number of DRAM accesses by the PNM module 1300 may be further reduced, which may increase the operation speed (or inference speed).
[0164] The PNM module 1300 may store the embedding information / connection information of the new node (eg, node v12), the supernode information of supernode A3, the hyperedge information of supernodes A2 and A3, and the hyperedge information of C + The edge information, the embedding information / connection information of the node v3 to be connected to the node v12 , and the embedding information / connection information of the node v8 to be connected to the node v12 are stored in the second buffer 1320 .
[0165] The operation circuit 1330 may include a plurality of processing elements (PEs), each of which may include one or more multiply-accumulate (MAC) operation circuits. The operation circuit 1330 may perform reference Fig.11 The computing operations of the processing device 910 described above. The operation circuit 1330 may perform the calculations described above. Figure 10 to Figure 11 The operation of the GCN model 920 is described.
[0166] For example, the operation circuit 1330 may receive the embedded information of the nodes v3, v8, and v12 from the second buffer 1320, and as shown in FIG. Fig.11 The embedded information of the node v12 1110 is updated according to v12′←v12+v8+v3. The operation circuit 1330 may store the updated embedded information of the node v12 1110 in the second buffer 1320.
[0167] The operation circuit 1330 may receive the supernode information of the supernode A2 from the first buffer 1310, and receive the embedded information of the node v12 and the supernode information of the supernode A3 from the second buffer 1320. The operation circuit 1330 may be as described with reference to Fig.11 The embedded information of the node v8 is updated according to v8′←v12+A2+A3 as described. The operation circuit 1330 may store the updated embedded information of the node v8 in the second buffer 1320 .
[0168] The operation circuit 1330 may receive the supernode information of the supernode A2 from the first buffer 1310, and receive the embedded information of the nodes v3 and v12 from the second buffer 1320. The operation circuit 1330 may Fig.11 Update the embedding information of node v3 according to v3'←v3+v12+A2 as described in + edge, so the operating circuit 1330 can be operated according to Fig.11 v3′←v3′+v2 described in 1330 adds the embedded information of the node v2 to the updated embedded information of the node v3 to correct the updated embedded information of the node v3. The operation circuit 1330 may store the corrected embedded information of the node v3 in the second buffer 1320.
[0169] The operating circuit 1330 may be configured as follows: Fig.11 The sum (eg, addition) of the updated embedding information of nodes v12, v8, and v3 is calculated by using v12″←v12′+v8′+v3′ described in . Therefore, the operation circuit 1330 may obtain the final embedding information of the node v12 as an inference result.
[0170] The DRAM controller 1350 may enable the PNM module 1300 to receive data from the DRAM and to write data to the DRAM. For example, the DRAM controller 1350 may write the inference result (e.g., the final embedded information of the node v12) to the DRAM (e.g., the DRAM at the memory rank 1 1221) through the DDR PHY interface 1361.
[0171] According to an embodiment, the control circuit 1340 may enable the PNM module 1300 to distinguish between side information, super side information, and correction side information (C + Side Information / C - For example, the side information, the super side information, and the corrected side information may each have the same format (e.g., a format of (integer, integer)). The control circuit 1340 may cause the PNM module 1300 (e.g., the operation circuit 1330) to distinguish the side information, the super side information, and the corrected side information during the operation process.
[0172] Fig.14 is a flowchart illustrating an operating method of a processing device according to an embodiment.
[0173] Reference Fig.14 In operation 1410, the processing device 910 may obtain a first node to be added to the graph 210 (eg, Fig.10 Node v11 or Fig.11 The initial embedding information of the first node may be the embedding information of the first node received by the processing device 910 from the DRAM (e.g., the embedding information when the first node is not connected to the graph 210).
[0174] In operation 1420, the processing device 910 may receive supernode information of a supernode of a compressed graph corresponding to the graph, wherein the supernode includes multiple nodes from the graph. For example, the processing device may receive supernode information for operation iteratively among multiple pieces of supernode information of multiple supernodes of the compressed graph 610 from a buffer (e.g., the first buffer 1310).
[0175] In operation 1430, the processing device may generate modified embedding information of the first node based on the supernode information, the embedding information, and the connection information using a graph convolutional network (GCN) model. In some examples, the processing device 910 may determine the embedding information of the first node based on the received supernode information, the obtained initial embedding information, the embedding information of the second node to be connected to the first node on the graph 210, and the GCN model 920.
[0176] According to an embodiment, in operation 1430, the processing device 910 may obtain the initial embedding information (eg, Fig.10 The embedded information of node v11), the received supernode information (for example, Fig.10 The supernode information of supernode A2) of the second node (for example, Fig.10The processing device 910 may perform aggregation (e.g., first aggregation) on the embedded information (i.e., connection information) of the node v8) to obtain a first result (or a first intermediate result). When the super node to which the second node belongs does not have a self-edge, the processing device 910 may obtain a second result (or a second intermediate result) by performing aggregation (e.g., second aggregation) on the obtained initial embedded information and "at least one of the (or additional) super node information about the neighboring super node connected to the super node to which the second node belongs and the received super node information". The processing device 910 may determine the embedded information of the first node based on the first result and the second result. The determined embedded information of the first node may correspond to, for example, the embedded information of the first node represented as being connected to the graph 210 (or the compressed graph 610).
[0177] According to an embodiment, when the supernode to which the second node belongs has correction information (e.g., C + Side information and / or C - When the first edge is removed from the compressed graph (e.g., the supernode to which the second node belongs has first correction information that the first edge represented on the graph 210 is not represented in the compressed graph 610 (e.g., C + The processing device 910 may correct the first result and / or the second result by adding the embedded information of a portion (or one or more) of the nodes on the first edge to the first result and / or the second result. As another example, when a second edge is added to the compressed graph (for example, the super node to which the second node belongs has the second correction information indicating that the second edge not defined on the graph 210 is represented in the compressed graph 610 (for example, C - The processing device 910 may correct the first result and / or the second result by subtracting the embedding information of a portion (or one or more) of the nodes on the second edge from the first result and / or the second result.
[0178] According to an embodiment, the step of generating the embedding information includes determining whether the supernode to which the second node belongs has a self-edge, wherein the second aggregation is based on the determination. Therefore, in the case where the supernode to which the second node belongs (e.g., supernode A3) has a self-edge, the processing device 910 may perform aggregation (e.g., second aggregation) on the obtained initial embedding information, supernode information about the supernode to which the second node belongs, and "at least one of supernode information about neighboring supernodes (e.g., neighboring supernode A2 of supernode A3) and received supernode information". In addition, in the case where the supernode to which the second node belongs (e.g., supernode A2) does not have a self-edge, the processing device 910 may perform aggregation (e.g., second aggregation) on the obtained initial embedding information, the embedding information of the second node, and "at least one of supernode information about neighboring supernodes (e.g., neighboring supernodes A1 and A3 of supernode A2) and received supernode information".
[0179] Fig.15 is a block diagram showing an example of a configuration of a processing device according to an embodiment.
[0180] According to an embodiment, the processing device 1500 (such as the processing device 910) may include a first buffer 1510 (eg, Fig.13 , the first buffer 1310 described in the specification), the second buffer 1520 (eg, Fig.13 1320) and the operation circuit 1530 (eg, Fig.13 Operating circuit 1330 described in .
[0181] The first buffer 1510 may store one or more pieces of super node information iteratively used for operation among a plurality of pieces of super node information of a plurality of super nodes of the compressed graph 610 of the graph 210 .
[0182] The second buffer 1520 may store the nodes on the graph 210 that are to be connected to the first node (eg, Fig.10 Node v11 as described in or Fig.11 The embedded information of the second node of the node v12) described in .
[0183] The operation circuit 1530 may perform the operation of the GCN model 920. The operation circuit 1530 may obtain initial embedding information of the first node and connection information between the first node and the second node. The operation circuit 1530 may receive the super node information stored in the first buffer 1510 and receive the embedding information of the second node from the second buffer 1520. The operation circuit 1530 may use the GCN model (e.g., as described in reference to Figures 9 to 11 The described GCN model 920) determines the embedding information of the first node based on the received supernode information, the obtained initial embedding information, and the received embedding information to generate modified embedding information.
[0184] According to an embodiment, the operation circuit 1530 may obtain a first result by performing aggregation (e.g., a first aggregation) on the obtained initial embedding information, the received supernode information, and the received embedding information. In some cases, when the supernode to which the second node belongs does not have a self-edge, the operation circuit 1530 may obtain a second result by performing aggregation (e.g., a second aggregation) on the obtained initial embedding information and "at least one of the (or additional) supernode information of the neighboring supernode connected to the supernode to which the second node belongs and the received supernode information."
[0185] According to an embodiment, when the supernode to which the second node belongs has correction information about the difference between the connection relationship of the compressed graph 610 and the connection relationship of the graph 210, the operation circuit 1530 may correct the first result and / or the second result based on the correction information.
[0186] According to an embodiment, when the supernode to which the second node belongs has first correction information indicating that a first edge on the graph 210 is not represented in the compressed graph 610, the operation circuit 1530 can correct the first result and / or the second result by adding the embedded information of a portion of the nodes on the first edge to the first result and / or the second result.
[0187] According to an embodiment, when the supernode to which the second node belongs has second correction information indicating that a second edge undefined on the graph 210 is represented in the compressed graph 610, the operation circuit 1530 can correct the first result and / or the second result by subtracting the embedded information of a portion of the nodes on the second edge from the first result and / or the second result.
[0188] According to an embodiment, when the supernode to which the second node belongs has a self-edge, the operation circuit 1530 may perform aggregation on the obtained initial embedding information, supernode information about the supernode to which the second node belongs, and "at least one of supernode information about neighboring supernodes and received supernode information."
[0189] According to an embodiment, the operation circuit 1530 may determine the embedded information of the first node based on the first result and the second result.
[0190] Fig.16 is a flowchart illustrating an operating method of a processing device according to an embodiment.
[0191] Reference Fig.16 In operation 1610, the processing device 910 may obtain embedding information of a node of a graph (eg, graph 210). The embedding information of the node (or the embedding information in a state where the node may or may not be connected to graph 210) may be received by the processing device 910 from the DRAM.
[0192] In operation 1620, the processing device 910 may compress the graph to obtain a compressed graph, wherein the graph is compressed by grouping a plurality of nodes of the graph to form a supernode of the compressed graph. Figure 2 The graph (or input graph) 210 described performs lossless edge compression to generate a reference Figure 6 The compressed graph 610 described above. In some cases, the compression process may include reducing the number of edges in the graph 210 to satisfy constraints related to memory capacity. For further details on the compression process, see Figures 3 to 6 is described.
[0193] In operation 1630, the processing device may generate modified embedding information of the node based on the embedding information and the compressed graph using a graph convolutional network (GCN) model. In some examples, the processing device 910 may use the GCN model 920 to determine the embedding information of the node based on the supernode information of the compressed graph and the initial embedding information of the graph. For further details on the process of generating the modified embedding information, refer to Figures 7 and 8 is provided.
[0194] According to an embodiment, Figure 3 to Figure 4 As shown in , the processing device 910 can iteratively update the compressed graph by repeatedly determining whether the memory requirement exceeds the memory capacity. In the case where the memory requirement exceeds the memory capacity (operation 330: yes), the processing device performs operations 340 and 350 (e.g., compressing supernodes and / or hyperedges of the compressed graph) before performing training of the GCN model based on the compressed graph.
[0195] According to an embodiment, the process of graph compression further includes: iteratively performing node partitioning based on homogeneity (such as referring to Figure 4 As described above), a node merging process of multiple nodes is then performed according to the node partitioning based on homogeneity. In some examples, nodes with similar trends or characteristics are connected to each other (for example, using a greedy algorithm) and merged to generate node pairs. For further details on graph compression, refer to Figures 3 to 6 is provided.
[0196] The units described herein can be implemented using hardware components, software components, and / or combinations thereof. The processing device (or processing equipment) can be implemented using one or more general or special computers (such as, for example, a processor, a controller and an arithmetic logic unit (ALU), a DSP, a microcomputer, an FPGA, a programmable logic unit (PLU), a microprocessor, or any other device capable of responding and executing instructions in a limited manner). The processing device can run an operating system (OS) and one or more software applications running on the OS. The processing device can also access, store, manipulate, process, and create data in response to the execution of the software. For simplicity, the description of the processing device is used as a singular; however, those skilled in the art will understand that the processing device may include multiple processing elements and multiple types of processing elements. For example, the processing device may include multiple processors, or a single processor and a single controller. In addition, different processing configurations (such as parallel processors) are possible.
[0197] Software may include computer programs, code segments, instructions, or specific combinations thereof to independently or collectively command or configure a processing device to operate as desired. Software and data may be implemented permanently or temporarily in any type of machine, component, physical or virtual device, computer storage medium or device, or in a propagating signal wave capable of providing instructions or data to or interpreted by a processing device. Software may also be distributed on networked computer systems so that the software is stored and executed in a distributed manner. Software and data may be stored on one or more non-transitory computer-readable recording media.
[0198] The method according to the above-described embodiment may be recorded in a non-transitory computer-readable medium, which includes program instructions for implementing various operations of the above-described embodiment. The medium may also include data files, data structures, etc., either alone or in combination with program instructions. The program instructions recorded on the medium may be program instructions specially designed and constructed for the purpose of the embodiment, or the program instructions recorded on the medium may be types known and available to technicians in the field of computer software. Examples of non-transitory computer-readable media include: magnetic media (such as hard disks, floppy disks, and tapes); optical media (such as CD-ROM disks, DVDs, and / or Blu-ray discs); magneto-optical media (such as optical discs); and hardware devices (such as read-only memory (ROM), random access memory (RAM), flash memory (e.g., USB flash drive, memory card, memory stick, etc.)) specially configured to store and execute program instructions. Examples of program instructions include both machine codes such as those generated by a compiler and files containing high-level codes that can be executed by a computer using an interpreter.
[0199] The above-mentioned apparatuses may be configured to act as one or more software modules in order to perform the operations of the above-mentioned examples, and vice versa.
[0200] Many embodiments have been described above. However, it should be understood that various modifications may be made to these embodiments. Suitable results may be achieved if the described techniques are performed in a different order, and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner, and / or replaced or supplemented by other components or their equivalents.
[0201] The processing discussed above is intended to be illustrative and not restrictive. Those skilled in the art will appreciate that the steps of the processing discussed herein may be omitted, modified, combined and / or rearranged, and any additional steps may be performed, without departing from the scope of the invention. More generally, the above disclosure is intended to be exemplary and not restrictive. Only the appended claims are intended to set limitations on what is included in the present invention. In addition, it should be noted that the features and limitations described in any one embodiment may be applied to any other embodiment herein, and the flowcharts or examples associated with one embodiment may be combined with any other embodiment in a suitable manner, performed in a different order, or performed in parallel. In addition, the systems and methods described herein may be executed in real time. It should also be noted that the above-described systems and / or methods may be applied to or used according to other systems and / or methods.
[0202] Accordingly, other implementations, other examples, and equivalents of the claims are within the scope of the appended claims.
Claims
1. A data processing method for a processing device, the data processing method comprising: Obtaining embedding information of a first node to be added to a graph and connection information between the graph and the first node; receiving supernode information of a supernode of a compressed graph corresponding to the graph, wherein the supernode includes a plurality of nodes from the graph; as well as A graph convolutional network model is used to generate modified embedding information of the first node based on the supernode information, the embedding information, and the connection information.
2. The data processing method according to claim 1, wherein: The steps to generate the modified embedding information include: Obtaining a first result by performing a first aggregation on the embedding information, the supernode information, and the connection information, wherein the connection information includes embedding information of a second node connected to the first node on the graph; and The second result is obtained by performing a second aggregation on the embedded information and the additional supernode information of the neighboring supernodes connected to the supernode to which the second node belongs.
3. The data processing method according to claim 2, wherein: The steps to generate the modified embedding information include: The first result and / or the second result is corrected based on the correction information, wherein the correction information indicates a difference between a connection relationship of the compressed graph and a connection relationship of the graph.
4. The data processing method according to claim 3, wherein: The step of correcting the first result and / or the second result comprises: When a first edge is removed from the compressed graph, additional embedding information of one or more nodes on the first edge is added to the first result and / or the second result.
5. The data processing method according to claim 3, wherein: The step of correcting the first result and / or the second result comprises: When a second edge is added to the compressed graph, embedding information of one or more nodes on the second edge is subtracted from the first result and / or the second result.
6. The data processing method according to any one of claims 2 to 5, wherein: The steps to generate the modified embedding information include: determining whether the supernode to which the second node belongs has a self-edge, wherein the second aggregation is based on the determination, When the supernode to which the second node belongs has a self-edge, performing a second aggregation on the embedded information, the additional supernode information of the neighboring supernodes, and the supernode information of the supernode to which the second node belongs, and When the super node to which the second node belongs has no self-edge, a second aggregation is performed on the embedded information, the additional super node information of the neighboring super nodes, and the embedded information of the second node.
7. A processing device comprising: a first buffer configured to store supernode information of a supernode of a compressed graph, wherein the compressed graph includes supernodes corresponding to a plurality of nodes of the graph; a second buffer configured to store embedding information of a first node of the graph; and The operating circuit is configured to obtain super node information from the first buffer, obtain embedding information from the second buffer, obtain connection information between the second node and the first node of the graph, and use a graph convolutional network model to generate modified embedding information of the first node based on the super node information, the embedding information and the connection information.
8. The processing device according to claim 7, wherein: The operation circuit is also configured to obtain a first result by performing a first aggregation on the embedding information, the supernode information, and the connection information, and obtain a second result by performing a second aggregation on the embedding information and the additional supernode information of the neighboring supernodes connected to the supernode to which the second node belongs, wherein the connection information includes the embedding information of the second node connected to the first node on the graph.
9. The processing device according to claim 8, wherein: The operation circuit is further configured to correct the first result and / or the second result based on correction information, wherein the correction information indicates a difference between a connection relationship of the compressed graph and a connection relationship of the graph.
10. The processing device according to claim 9, wherein: The operation circuit is further configured to: when the first edge is removed from the compressed graph, add additional embedded information of one or more nodes on the first edge to the first result and / or the second result.
11. The processing device according to claim 9, wherein: The operation circuit is further configured to: when a second edge is added to the compressed graph, subtract the embedding information of one or more nodes on the second edge from the first result and / or the second result.
12. The processing device according to claim 8, wherein: The operation circuit is further configured to: determine whether the super node to which the second node belongs has a self-edge, wherein the second aggregation is based on the determination, When the supernode to which the second node belongs has a self-edge, performing a second aggregation on the embedded information, the additional supernode information of the neighboring supernodes, and the supernode information of the supernode to which the second node belongs, and When the super node to which the second node belongs has no self-edge, a second aggregation is performed on the embedded information, the additional super node information of the neighboring super nodes, and the embedded information of the second node.
13. A processing device according to any one of claims 7 to 12, wherein: The processing device is included in a near memory processing arrangement.
14. A method for training a graph convolutional network model, the method comprising: Compressing a graph to obtain a compressed graph, wherein the compressed graph includes super nodes representing a plurality of nodes of the graph and hyper edges for connecting the super nodes; performing aggregation based on supernodes and hyperedges of the compressed graph; and The result of the aggregation is corrected based on the correction information, wherein the correction information indicates a difference between a connection relationship of the compressed graph and a connection relationship of the graph.
15. The method according to claim 14, wherein: The steps to perform aggregation are: Obtaining embedding information of nodes in the plurality of nodes in the supernode; determining supernode information of the supernode based on the embedded information; and Update supernode information based on hyperedges.
16. The method according to claim 15, wherein: The steps to correct the aggregated results include: When the correction information indicates that a first edge of the graph is removed from the compressed graph, additional embedding information of one or more nodes on the first edge is added to the supernode information.
17. The method according to claim 15, wherein: The steps to correct the aggregated results include: When the correction information indicates that a second edge is added to the compressed graph, embedding information of one or more nodes on the second edge is subtracted from the supernode information.
18. A data processing method, comprising: Get the embedding information of the nodes of the graph; compressing the graph to obtain a compressed graph, wherein the graph is compressed by grouping a plurality of nodes of the graph to form a supernode of the compressed graph; as well as A graph convolutional network model is used to generate modified embedding information of the node based on the embedding information and the compressed graph.
19. The data processing method according to claim 18, further comprising: The compressed graph is iteratively updated by repeatedly calculating a memory requirement of the compressed graph and compressing supernodes and / or hyperedges of the compressed graph if the memory requirement exceeds the memory capacity.
20. The data processing method according to claim 18, wherein: The step of compressing the graph comprises: Performing homogeneity-based node partitioning; and Node merging of the plurality of nodes is performed according to homomorphism-based node partitioning.
Citation Information
Patent Citations
Water supplying apparatus of mushroom spawn seed
KR1020230161216A