A cosine similarity-based CSIS-GCN model establishment method and system
Patent Information
- Application Number
- CN202410011893.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-01-03
AI Technical Summary
第一阶段是传统GCN,采用整个图作为输入,也称为全批量(full-batch)模式,通过矩阵计算的方式更新节点特征表示,这种学习方法被称为直推式学习,具有准确和稳定等优点,但是在添加新节点时,由于图结构发生变化,需要重新计算整个图中所有节点的特征信息,不仅无法快速预测新节点,还会导致大量冗余计算;第二阶段是基于归纳式学习的GCN,通常包含两个步骤:采样和聚合
[0017]采用本发明实施例可以包括以下有益效果:本发明实施例解决了大图全批次训练内存不足的问题,且避免了点采样逐层递归拓展邻域时采样数指数增长的问题,利用稀疏矩阵运算将图中任意两个节点间的相似度预先计算并保存,以降低多轮迭代训练造成的重复计算带来的开销,并综合考虑每个节点的活跃邻居和与其具有相似偏好的节点,以准确评估节点在图中的整体影响,从而更准确地表征节点在整个网络中的综合影响力。
Smart Images

Figure CN117808039B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular to a method and system for establishing a graph convolutional neural network model (CSIS-GCN) based on cosine similarity-weighted importance sampling. Background Technology
[0002] In terms of data processing, the development of Graph Convolutional Networks (GCNs) has undergone two stages of continuous improvement and innovation, aiming to improve the utilization of memory and computing resources, making it possible to train and infer large-scale graphs. The first stage is the traditional GCN, which uses the entire graph as input, also known as the full-batch mode. It updates the node feature representation through matrix calculation. This learning method is called transductive learning, which has advantages such as accuracy and stability. However, when adding a new node, due to the change in graph structure, it is necessary to recalculate the feature information of all nodes in the entire graph. This not only makes it impossible to quickly predict new nodes, but also leads to a lot of redundant computation. The second stage is GCN based on inductive learning, which usually includes two steps: sampling and aggregation. The purpose of sampling is to acquire the structural features of the entire graph by learning the structural features of a small number of representative nodes, hence it is also called importance sampling. It can effectively solve the computational efficiency problem in the training and inference process of large graphs, but its drawback is that it may introduce sampling bias. In order to further reduce the consumption of memory and computing resources, researchers have proposed dividing nodes into multiple subsets (mini-batch) and then training GCNs by importance sampling. This can effectively reduce the amount of computation and improve system efficiency. For example, the GraphSAGE algorithm based on point sampling calculates the embedding vector of a node by sampling a fixed number of neighboring nodes for each node. Since it only focuses on a subset of nodes and their neighboring nodes, and does not need to focus on the entire graph, it significantly alleviates the memory consumption problem. However, the GraphSAGE algorithm assumes that the sampling probability of all neighboring nodes is equal, which may cause the sampling complexity to increase with the number of graph layers. The FastGCN algorithm, based on layer sampling, avoids the drawback of the number of sampled nodes increasing exponentially with the number of layers in the graph. It uses local subgraphs instead of the global graph and samples the local subgraphs independently layer by layer to calculate the embedding vector of the nodes. However, the FastGCN algorithm only considers the weight influence of first-order active neighbors on the nodes, which makes it impossible to accurately express the comprehensive weight of the nodes in the entire network. This also means that the sampling probabilities will only show a significant difference when the number of neighboring nodes differs greatly. At the same time, since the sampling between layers is independent, it may ignore the high weight influence of nodes with the same preferences, as well as the inaccuracy of the computation graph caused by the loss of sparse connection information, thus reducing the training speed and generalization performance of the FastGCN algorithm.To address the issues of information loss and sampling bias that may arise from random node sampling in each training batch of the FastGCN algorithm, researchers have proposed a node hierarchy-based importance sampling method called the LADIES algorithm. This algorithm designs a sampling probability distribution based on the node's hierarchy information in the graph, ensuring that both high-level and low-level nodes are adequately sampled. This effectively avoids the message loss problem that exists in random sampling. However, the problem of insufficient expression of node importance information may still exist during the sampling process.
[0003] In summary, GCN-based node classification prediction technology still has the following problems:
[0004] 1) How to optimize the learning strategy for traditional full-batch training methods in order to alleviate the high memory overhead;
[0005] 2) In existing inductive learning schemes, the number of sampling nodes increases exponentially with the number of layers in the graph;
[0006] 3) Existing sampling strategies often measure the importance of nodes from only a single aspect, which cannot accurately express the comprehensive weight of nodes in the graph, resulting in insufficient message passing during model aggregation;
[0007] 4) Existing algorithms generally suffer from insufficient message passing during the training phase, making them difficult to train. Summary of the Invention
[0008] The purpose of this invention is to provide a method and system for building CSIS-GCN models based on cosine similarity, in order to solve the above-mentioned problems in the prior art.
[0009] This invention provides a method for establishing a CSIS-GCN model based on cosine similarity, comprising:
[0010] Calculate the importance sampling probability of each node in the large graph and divide all nodes in the large graph into multiple subsets. Remove isolated nodes in each subset and construct a connected subgraph for each subset. The importance sampling probability is a weighted sum of proximity and activity.
[0011] The input feature matrix and model parameters of the graph convolutional neural network model are initialized. The connected subgraph is randomly sampled according to the importance sampling probability to obtain input data. The input data is input into the graph convolutional neural network model for training. The input feature matrix is iteratively updated and the model parameters are updated through stochastic gradient descent until a trained CSIS-GCN model is obtained.
[0012] This invention provides a CSIS-GCN model building system based on cosine similarity, comprising:
[0013] The preprocessing module is used to calculate the importance sampling probability of each node in the large graph and divide all nodes in the large graph into multiple subsets, remove isolated nodes in each subset and construct a connected subgraph for each subset, wherein the importance sampling probability is a weighted sum of proximity and activity.
[0014] The module is used to initialize the input feature matrix and model parameters of the graph convolutional neural network model, randomly sample the connected subgraph according to the importance sampling probability to obtain input data, input the input data into the graph convolutional neural network model for training, iteratively update the input feature matrix and update the model parameters through stochastic gradient descent until a trained CSIS-GCN model is obtained.
[0015] This invention also provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the above-described method for establishing a CSIS-GCN model based on cosine similarity.
[0016] This invention also provides a computer-readable storage medium storing an information transmission implementation program, which, when executed by a processor, implements the steps of the above-described method for establishing a CSIS-GCN model based on cosine similarity.
[0017] The embodiments of the present invention can include the following beneficial effects: The embodiments of the present invention solve the problem of insufficient memory for full batch training of large graphs, and avoid the problem of exponential growth of the number of samples when recursively expanding the neighborhood layer by layer by point sampling. The similarity between any two nodes in the graph is pre-calculated and stored by using sparse matrix operations to reduce the overhead caused by repeated calculations in multiple rounds of iterative training. Furthermore, the active neighbors of each node and nodes with similar preferences are comprehensively considered to accurately evaluate the overall influence of the node in the graph, thereby more accurately representing the comprehensive influence of the node in the entire network. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of the CSIS-GCN model establishment method based on cosine similarity according to an embodiment of the present invention;
[0020] Figure 2 This is a framework diagram of the CSIS-GCN algorithm based on cosine similarity according to an embodiment of the present invention;
[0021] Figure 3 This is a schematic diagram of the microscopic F1 score under different sampling node number limitations in an embodiment of the present invention;
[0022] Figure 4 This is a schematic diagram of the microscopic F1 score under different training iterations in an embodiment of the present invention;
[0023] Figure 5 This is a schematic diagram of the CSIS-GCN model building system based on cosine similarity according to an embodiment of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0025] Method Implementation Examples
[0026] According to embodiments of the present invention, a method for establishing a CSIS-GCN model based on cosine similarity is provided. Figure 1 This is a flowchart of the CSIS-GCN model building method based on cosine similarity according to an embodiment of the present invention, as follows: Figure 1 As shown, the method for establishing a CSIS-GCN model based on cosine similarity according to an embodiment of the present invention specifically includes:
[0027] Step S101: Calculate the importance sampling probability of each node in the large graph and divide all nodes in the large graph into multiple subsets. Remove isolated nodes in each subset and construct a connected subgraph for each subset. The importance sampling probability is a weighted sum of proximity and activity.
[0028] The calculation of the importance sampling probability of each node in the large graph specifically includes:
[0029] The adjacency matrix of the large graph is normalized according to Formula 1 to obtain a normalized adjacency matrix. Based on the normalized adjacency matrix, cosine similarity is calculated according to Formula 2 to obtain the similarity between any two nodes in the large graph, and a similarity matrix is constructed. The activity of each node is calculated by the norm of the vector according to the normalized adjacency matrix, and the proximity of each node is calculated by the norm of the vector according to the similarity matrix. The activity and proximity are weighted and summed according to Formula 3 to obtain the importance sampling probability of each node in the large graph.
[0030]
[0031] in, Let A denote the normalized adjacency matrix, and D denote the degree matrix of the graph. D is a diagonal matrix, and its diagonal elements are... ii It is the sum of all elements in the i-th row of the adjacency matrix A, where i = 1, 2, ..., N, N represents the total number of nodes in the graph, and I represents the identity matrix;
[0032]
[0033] Among them, v i Let v represent the i-th node. j Let c represent the j-th node. ij Represents node v i and node v j The similarity between them, i.e., node v i and node v j Cosine similarity between sim(v) i ,v j ), Represents the normalized adjacency matrix. and These are the normalized adjacency matrices. The i-th and j-th column vectors, N represents the total number of nodes in the graph, and ||·||2 represents the 2-norm of the vector;
[0034]
[0035] Where, p i Represents any node v i Importance sampling probability, w0 represents the weight balancing parameter, C = (c ij ) represents the similarity matrix, C :,i This represents the similarity matrix C corresponding to the current node v. i The column vector, ||·||2 represents the 2-norm of the vector, ||·|| p Let p represent the p-norm of a vector, where p takes the value 1 or 2, ||·|| FLet N represent the Frobenius norm of the matrix, and N represent the total number of nodes in the graph. Represents the normalized adjacency matrix The i-th column corresponds to the current node v i column vectors, Represents the normalized adjacency matrix The j-th column corresponds to the current node v j Column vectors.
[0036] Step S102: Initialize the input feature matrix and model parameters of the graph convolutional neural network model. Randomly sample the connected subgraphs according to the importance sampling probabilities to obtain input data. Input the input data into the graph convolutional neural network model for training. Iteratively update the input feature matrix and update the model parameters using stochastic gradient descent until a trained CSIS-GCN model is obtained. Specifically, this includes:
[0037] The input data is input into the graph convolutional neural network model for training. The input feature matrix is iteratively updated according to Formula 4, and the model parameters are updated by stochastic gradient descent according to Formula 5 until a trained CSIS-GCN model is obtained.
[0038]
[0039]
[0040] Where σ represents a nonlinear activation function, Let L represent the normalized adjacency matrix, L represent the number of layers in the GCN model, and Θ represent the normalized adjacency matrix. (l) H represents the parameter matrix of the l-th layer. (l) H represents the matrix consisting of embedding vectors in the l-th layer. (l+1) This represents the feature matrix of the (l+1)th sampling layer node, whose update depends on the lth sampling layer, Θ (l) This represents the model parameters to be updated, and η represents the learning rate of the CSIS-GCN algorithm. The cross-entropy loss function L represents the loss function L with respect to the model parameters Θ. (l) The partial derivatives;
[0041] The method further includes:
[0042] The trained CSIS-GCN model is used to classify nodes in a large graph.
[0043] The following describes the specific implementation process of the CSIS-GCN model building method based on cosine similarity according to an embodiment of the present invention, and combines it with... Figure 2 The technical solutions described above in the embodiments of the present invention will be explained in detail.
[0044] Analysis of mainstream sampling algorithms reveals that a good GCN-based sampling strategy should possess the following three properties:
[0045] 1) Based on layer sampling, a fixed number of neighboring nodes are sampled layer by layer. Only a portion of the nodes and their neighbors are selected for calculation in each layer to avoid the phenomenon that the number of neighboring nodes grows exponentially with the number of layers.
[0046] 2) Ensure connectivity between layers to make the entire computation graph a connected graph, so as to avoid nodes where gradient propagation is interrupted during backpropagation, thereby ensuring that the network can eventually converge;
[0047] 3) By comprehensively considering the overall influence of nodes in the graph, including the influence of each node's active neighbors and the influence of each individual's preference weight, the sampling variance can be effectively reduced, thus accelerating the stable convergence of the model.
[0048] Therefore, the embodiments of the present invention address the shortcomings of the prior art based on the following points:
[0049] 1) Learn graph structure features through an inductive learning framework to reduce the memory shortage problem during full-batch training of large graphs;
[0050] 2) By using layer sampling, a fixed number of neighbor nodes are sampled at each layer, thus avoiding the problem of exponential growth in the number of samples when recursively expanding the neighborhood layer by layer in point sampling;
[0051] 3) During the data pre-training phase, for a given large graph G, its adjacency matrix A is a fixed value, and the entire large graph is usually a sparse graph. Therefore, before the iterative training begins, this embodiment of the invention uses sparse matrix operations to separate any two nodes v in the graph. i and v j similarity c ij Pre-calculate and save data to reduce the overhead of repetitive computation caused by multiple rounds of iterative training;
[0052] 4) This invention proposes a Cosine Similarity Based Weighted Importance Sampling for Training Graph Convolutional Networks (CSIS-GCN) algorithm, which comprehensively considers the active neighbors of each node and nodes with similar preferences to accurately evaluate the overall influence of the node in the graph. In other words, the CSIS-GCN algorithm is not limited to the influence of first-order active neighbors, but also considers the high-weight influence of similar preferences between nodes in the connected subgraph, thereby more accurately representing the comprehensive influence of the node in the entire network.
[0053] The CSIS-GCN algorithm proposed in this invention uses both node activity and closeness to characterize the comprehensive influence (CI), i.e., the sampling probability, which can be expressed by the following formula:
[0054] CI i =w0×Cn(C i )+(1-w0)×Cn(D i (1)
[0055] Where w0 represents the weight coefficient, Cn(.) represents a normalization method, and node v i Degree D i It is typically used to calculate the activity level Cn(D) of a node. i ), v i Similarity C with other nodes i Then it can be used to calculate v i proximity Cn(C i Obviously, the proximity Cn(C) i A reasonable definition of ) is essential for accurately reflecting v i Overall influence is crucial.
[0056] 1. Cosine similarity
[0057] This invention proposes a method for measuring the similarity between nodes using cosine similarity. Non-Euclidean graphs are typically represented using adjacency matrices, which represent the relationship between each node and its neighbors as a one-dimensional vector of edge connections. Given a graph G consisting of a node set V and an edge set E, suppose two nodes v... i and v j The edge connection vectors are respectively and Then v iand v j The cosine similarity between them can be calculated using the following formula:
[0058]
[0059] Where ||v||2 represents the 2-norm of vector v, and obviously, sim(v i ,v j The closer the value of ) is to 1, the more it means that node v i and v j The higher the similarity between them, the greater the similarity preference weight.
[0060] 2. Cosine similarity importance sampling
[0061] The CSIS-GCN algorithm proposed in this embodiment of the invention takes importance sampling as its core. When calculating the sampling probability of nodes in the l-th layer, it not only considers the weight influence of active first-order neighbor nodes, but also the constraint of the weight of highly favored common neighbors. Assuming that the adjacency matrix of graph G is A, it is preprocessed as follows:
[0062]
[0063] Where I is the identity matrix, D is a diagonal matrix, and its diagonal elements D ii It is the sum of all elements in the i-th row of the adjacency matrix A, where i = 1, 2, ..., N, and N represents the total number of nodes in the graph. Adjacency matrix normalization preprocessing transforms the relationships between nodes into relative weights, making the network applicable to different types of graph data. This helps improve the network's stability and effectiveness, thereby enhancing deep learning performance on graph data and playing a crucial role in graph data analysis and machine learning tasks. Let N represent the total number of nodes in the graph, and use the cosine similarity defined in (2) to measure the similarity between any two nodes v. i and v j Similarity between them:
[0064]
[0065] Where, C = (c ij Let represent the similarity matrix. Note that for a given large graph G, its adjacency matrix A is a fixed value, and the entire large graph is usually a sparse graph. Therefore, before iterative training begins, sparse matrix operations can be used to separate any two nodes v in the graph. i and v j similarity c ij As part of data preprocessing, computation does not incur significant computational overhead for learning large graph nodes.
[0066] According to (1), node v iThe final sampling probability is defined as v i The weighted sum of activity and proximity of nodes v in the CSIS-GCN algorithm is used to determine the proximity of nodes v. i The sampling probability is defined as:
[0067]
[0068] Where w0 represents the weight balancing parameter, C :,i This represents the similarity matrix C corresponding to the current node v. i The column vector, ||·|| p Let p represent the p-norm of a vector, where p takes the value 1 or 2, ||·|| F The Frobenius norm of the matrix is represented in (5), where the first term on the right indicates the number of nodes v calculated using the 2-norm of the vectors based on similarity. i The first term is the proximity, consistent with the FastGCN algorithm. The second term indicates that v is calculated based on the degree using the p-norm of the vector. i Activity level.
[0069] 3. Parameter update
[0070] In GCN, each node v i Each node has an embedding vector that can contain any of its attributes or features. The purpose of GCN's layer aggregation is to aggregate the embedding vectors of nodes in the previous layer into the embedding vectors of nodes in the current layer. This process can be represented by the following formula:
[0071]
[0072] Where σ is a nonlinear activation function, It is the normalized adjacency matrix, Θ (l) H is the parameter matrix to be updated. (l) H represents the matrix consisting of embedding vectors at the l-th layer. (0) H represents the input feature matrix of the GCN model. Taking a certain citation dataset as an example, H (0) The initialization method is as follows: Assuming there are N nodes in the citation dataset, and each node has a feature dimension of M, sort all words in the N articles according to their frequency, and extract the top M high-frequency words to reconstruct a sub-vocabulary. Then, any article can be represented by an M-dimensional one-hot encoded vector. These one-hot encoded vectors constitute the H... (0) .
[0073] This invention employs a semi-supervised learning approach to train the CSIS-GCN model, using cross-entropy to define the cost function L. Since the feature matrix H of the (l+1)th sampling layer node... (l+1)The update depends on the l-th sampling layer. In this embodiment of the invention, the parameter Θ is adjusted through backpropagation. (l) To minimize the loss function and improve model accuracy, specifically, the parameter Θ is updated using a Stochastic Gradient Descent (SGD) optimizer. (l) :
[0074]
[0075] Where η represents the learning rate of the CSIS-GCN algorithm, Figure 2 The specific implementation steps of the CSIS-GCN algorithm for forward computation and backpropagation in each round of iterative training are demonstrated.
[0076] It should be noted that the CSIS-GCN algorithm presets the sampling probability of isolated nodes to 0 in each subset in order to remove these isolated nodes and thus construct a connected subgraph G within that subset. B The purpose of this is to ensure that gradient propagation breaks due to graph sparsity during backpropagation.
[0077] 4. Experimental Results and Analysis
[0078] To verify the node classification performance of the CSIS-GCN sampling algorithm, this embodiment of the invention compares the CSIS-GCN sampling algorithm with several existing sampling algorithms, including four traditional machine learning algorithms labeled "ManiReg", "SeimEmb", "DeepWalk", and "Planetoid", and five graph convolutional neural network-based algorithms labeled "GraphSAGE", "FastGCN", "LADIES", "OGT", and "GRNN". This embodiment of the invention evaluates all sampling algorithms using the three datasets listed in Table 1, and ultimately uses the micro F1 score on the test dataset as the standard for measuring the performance of each algorithm.
[0079] Table 1 Dataset
[0080]
[0081] This invention provides the average micro F1 scores of the various graph node classification algorithms in 10 experiments, with a learning rate η = 0.01 and a sampling node count of K = 200. The calculation results are shown in Table 2. Table 2 shows that, firstly, under the same experimental settings, the CSIS-GCN sampling algorithm outperforms other algorithms on the CiteSeer and PubMed datasets, and its performance on the Cora dataset is second only to the "GRNN" algorithm. Secondly, compared to the GraphSAGE algorithm, which is based on point sampling and assumes equal sampling probabilities for all neighboring nodes, the FastGCN algorithm, based on layer sampling, performs better. Furthermore, compared to the FastGCN algorithm, which only considers node activity, the CSIS-GCN sampling algorithm, which comprehensively considers node activity and proximity, achieves micro F1 scores 2.56%, 2.25%, and 1.17% higher on the three datasets, respectively. This reflects the simultaneous consideration of direct neighbors and local high influence. Public neighbors play a positive role in updating the characteristics of nodes themselves and can more effectively address the problem of insufficient message passing during sampling in inductive learning methods. Finally, almost all sampling algorithms outperform the CiteSeer dataset in classification on the Cora and PubMed datasets. This is because the ratio of the number of edges to the number of nodes in the Cora and PubMed datasets is much larger than that in the CiteSeer dataset. This means that their nodes have richer features, which is beneficial for node classification tasks. In contrast, the CiteSeer dataset may suffer from insufficient message passing, which in turn reduces the overall classification accuracy. As can be seen from Table 2, the CSIS-GCN sampling algorithm has a more significant improvement on the CiteSeer dataset.
[0082] Table 2. Micro F1 scores (%) of various classification algorithms on three datasets.
[0083]
[0084] Subsequently, this embodiment of the invention illustrates the effectiveness and stability of the CSIS-GCN sampling algorithm from two aspects: the selection of the number of sampling nodes K and the iterative convergence of the node classification algorithm. On the one hand, to illustrate the impact of K on the performance of the sampling algorithm, this embodiment compares the micro F1 scores of the CSIS-GCN and FastGCN algorithms on the Cora dataset under different sampling number settings, taking K=2. k Let k = 3, 4, ..., 9, and the learning rate η = 0.03. The calculation results are as follows: Figure 3 As shown, it is clear that the performance of both algorithms increases with the increase of the number of sampling nodes K, but the CSIS-GCN sampling algorithm consistently outperforms the FastGCN algorithm. It is particularly noteworthy that when K is small, such as K≤32, the aggregated information is relatively limited. Figure 3 As can be seen, the performance of the CSIS-GCN sampling algorithm is significantly better than that of the FastGCN algorithm, indicating that the CSIS-GCN sampling algorithm can effectively alleviate the message delivery insufficiency problem caused by the limited number of samples. On the other hand, in this embodiment of the invention, the number of sampling nodes K = 30 and the learning rate η = 0.03 are selected, and... Figure 4 The convergence of the CSIS-GCN sampling algorithm and the FastGCN algorithm as the number of training iterations increases is shown. It can be seen that the CSIS-GCN sampling algorithm converges after about 210 training iterations, while the FastGCN algorithm always shows slight oscillations. This is because, under the premise of a small number of samples and limited aggregated information, the FastGCN algorithm only considers node activity and may sample irrelevant nodes, resulting in insufficient message passing and the algorithm cannot converge quickly and stably. Under the same conditions, the CSIS-GCN sampling algorithm accelerates the convergence speed of the model by making full use of effective nodes to update features through a strategy of weighted balancing of node activity and proximity.
[0085] System Implementation Examples
[0086] According to embodiments of the present invention, a CSIS-GCN model building system based on cosine similarity is provided. Figure 5 This is a schematic diagram of the CSIS-GCN model building system based on cosine similarity according to an embodiment of the present invention, as shown below. Figure 5 As shown, the CSIS-GCN model building system based on cosine similarity according to an embodiment of the present invention specifically includes:
[0087] Preprocessing module 50 is used to calculate the importance sampling probability of each node in the large graph, divide all nodes in the large graph into multiple subsets, remove isolated nodes in each subset, and construct a connected subgraph for each subset. The importance sampling probability is a weighted sum of proximity and activity. Specifically, it is used for:
[0088] The adjacency matrix of the large graph is normalized according to Formula 1 to obtain a normalized adjacency matrix. Based on the normalized adjacency matrix, cosine similarity is calculated according to Formula 2 to obtain the similarity between any two nodes in the large graph, and a similarity matrix is constructed. The activity of each node is calculated by the norm of the vector according to the normalized adjacency matrix, and the proximity of each node is calculated by the norm of the vector according to the similarity matrix. The activity and proximity are weighted and summed according to Formula 3 to obtain the importance sampling probability of each node in the large graph.
[0089]
[0090] in, Let A denote the normalized adjacency matrix, and D denote the degree matrix of the graph. D is a diagonal matrix, and its diagonal elements are... ii It is the sum of all elements in the i-th row of the adjacency matrix A, where i = 1, 2, ..., N, N represents the total number of nodes in the graph, and I represents the identity matrix;
[0091]
[0092] Among them, v i Let v represent the i-th node. j Let c represent the j-th node. ij Represents node v i and node v j The similarity between them, i.e., node v i and node v j Cosine similarity between sim(v) i ,v j ), Represents the normalized adjacency matrix. and These are the normalized adjacency matrices. The i-th and j-th column vectors, N represents the total number of nodes in the graph, and ||·||2 represents the 2-norm of the vector;
[0093]
[0094] Where, p i Represents any node v i Importance sampling probability, w0 represents the weight balancing parameter, C = (c ij ) represents the similarity matrix, C :,i This represents the similarity matrix C corresponding to the current node v. i The column vector, ||·||2 represents the 2-norm of the vector, ||·|| p Let p represent the p-norm of a vector, where p takes the value 1 or 2, ||·|| F Let N represent the Frobenius norm of the matrix, and N represent the total number of nodes in the graph. Represents the normalized adjacency matrix The i-th column corresponds to the current node v i column vectors, Represents the normalized adjacency matrix The j-th column corresponds to the current node v j Column vectors;
[0095] Module 52 is used to initialize the input feature matrix and model parameters of the graph convolutional neural network model. It randomly samples the connected subgraphs according to the importance sampling probability to obtain input data, inputs the input data into the graph convolutional neural network model for training, iteratively updates the input feature matrix, and updates the model parameters through stochastic gradient descent until a trained CSIS-GCN model is obtained. Specifically, it is used for:
[0096] The input data is fed into the graph convolutional neural network model for training. The input feature matrix is iteratively updated according to Formula 4, and the model parameters are updated by stochastic gradient descent according to Formula 5 until a trained CSIS-GCN model is obtained.
[0097]
[0098]
[0099] Where σ represents a nonlinear activation function, Let L represent the normalized adjacency matrix, L represent the number of layers in the GCN model, and Θ represent the normalized adjacency matrix. (l) H represents the parameter matrix of the l-th layer. (l) H represents the matrix consisting of embedding vectors in the l-th layer. (l+1) This represents the feature matrix of the (l+1)th sampling layer node, whose update depends on the lth sampling layer, Θ (l) This represents the model parameters to be updated, and η represents the learning rate of the CSIS-GCN algorithm. The cross-entropy loss function L represents the loss function L with respect to the model parameters Θ. (l) The partial derivatives;
[0100] The system further includes:
[0101] The classification module is used to classify nodes in a large image using the trained CSIS-GCN model.
[0102] The embodiments of the present invention are system embodiments corresponding to the above method embodiments. The specific operation of each module can be understood by referring to the description of the method embodiments, and will not be repeated here.
[0103] In summary, this invention proposes a high-precision classification algorithm based on weighted importance sampling using cosine similarity. The algorithm includes: in inductive learning, sampling is performed by calculating node proximity based on cosine similarity, and the sampling probability of node importance is weighted and balanced based on proximity and node degree. For a given large graph, its adjacency matrix is a fixed value, and the entire large graph is usually sparse. Therefore, before iterative training begins, the similarity between any two nodes in the graph can be calculated using sparse matrix operations, and the node similarity is stored as a result of data preprocessing for later use, reducing the huge additional computational overhead caused by repeated calculations during iterative training of large graphs. In inductive learning, node proximity is calculated based on cosine similarity as a dimension of importance sampling to improve sampling accuracy. Furthermore, a two-dimensional weighted and balanced importance sampling strategy is proposed: calculating node proximity based on cosine similarity and node activity based on node degree to more accurately express the influence of nodes in the entire graph, further weakening the insufficient message passing in the graph, improving the accuracy of message aggregation, and enhancing the model's generalization ability.
[0104] Device Example 1
[0105] This invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, performs the steps described in the method embodiment.
[0106] Device Example 2
[0107] This invention provides a computer-readable storage medium storing an information transmission implementation program, which, when executed by a processor, performs the steps described in the method embodiment.
[0108] The computer-readable storage media described in this embodiment include, but are not limited to, ROM, RAM, disk, or optical disk.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for establishing a weighted importance sampling graph convolutional neural network (CSIS-GCN) model based on cosine similarity, characterized in that, A node classification task applied to citation network graph data, wherein the citation network graph data includes multiple nodes and edges connecting the nodes, wherein each node corresponds to a scientific document, and each edge represents a citation relationship between two documents, the method includes: Calculate the importance sampling probability of each node in the large graph and divide all nodes in the large graph into multiple subsets. Remove isolated nodes from each subset and construct a connected subgraph for each subset. The importance sampling probability is a weighted sum of proximity and activity. Specifically, this includes: The adjacency matrix of the large graph is normalized according to Formula 1 to obtain a normalized adjacency matrix. Based on the normalized adjacency matrix, cosine similarity is calculated according to Formula 2 to obtain the similarity between any two nodes in the large graph, and a similarity matrix is constructed. The activity of each node is calculated by the norm of the vector according to the normalized adjacency matrix, and the proximity of each node is calculated by the norm of the vector according to the similarity matrix. The activity and proximity are weighted and summed according to Formula 3 to obtain the importance sampling probability of each node in the large graph. Formula 1; in, Represents the normalized adjacency matrix. Represents the adjacency matrix of a graph. The degree matrix of the graph is a diagonal matrix, and its diagonal elements D ii It is an adjacency matrix No. i The sum of all elements in the row, i =1,2,..., N , N This represents the total number of nodes in the graph. Represents the identity matrix; Official 2; in, Indicates the first 1 node Indicates the first 1 node Represents a node and nodes The similarity between nodes and nodes Cosine similarity between , Represents the normalized adjacency matrix. and These are the normalized adjacency matrices. The The and the first column vectors, This represents the total number of nodes in the graph. The 2-norm of a vector; Official 3; in, Represents any node Importance sampling probability, This represents the weight balancing parameter. Represents the similarity matrix. Representing the similarity matrix The middle corresponds to the current node column vectors, Describes the 2-norm of a vector. Representing vectors Norm, The value can be 1 or 2. Denotes the Frobenius norm of a matrix. This represents the total number of nodes in the graph. Represents the normalized adjacency matrix The The column corresponds to the current node. column vectors, Represents the normalized adjacency matrix The The column corresponds to the current node. Column vectors; The input feature matrix and model parameters of the graph convolutional neural network model are initialized. The connected subgraph is randomly sampled according to the importance sampling probability to obtain input data. The input data is input into the graph convolutional neural network model for training. The input feature matrix is iteratively updated and the model parameters are updated through stochastic gradient descent until a trained CSIS-GCN model is obtained.
2. The method according to claim 1, characterized in that, The method further includes: The trained CSIS-GCN model is used to classify nodes in a large graph.
3. The method according to claim 1, characterized in that, The input data is fed into the graph convolutional neural network model for training. The input feature matrix is iteratively updated, and the model parameters are updated using stochastic gradient descent until a trained CSIS-GCN model is obtained. Specifically, this includes: The input data is input into the graph convolutional neural network model for training. The input feature matrix is iteratively updated according to Formula 4, and the model parameters are updated by stochastic gradient descent according to Formula 5 until a trained CSIS-GCN model is obtained. Official 4; Official 5; in, Represents a non-linear activation function. Represents the normalized adjacency matrix. This indicates the number of layers in the GCN model. Indicates the first The parameter matrix of the layer, Indicates the first A layer is a matrix consisting of embedding vectors. Indicates the first l The feature matrix of the +1 sampling layer node, whose update depends on the first sampling layer node. l Sampling layer, This indicates the model parameters to be updated. This represents the learning rate of the CSIS-GCN algorithm. Represents the cross-entropy loss function For model parameters The partial derivatives of .
4. A system for building a weighted importance sampling graph convolutional neural network (CSIS-GCN) model based on cosine similarity, characterized in that, A node classification task applied to citation network graph data, wherein the citation network graph data includes multiple nodes and edges connecting the nodes, wherein each node corresponds to a scientific document, and each edge represents a citation relationship between two documents, the system comprising: The preprocessing module is used to calculate the importance sampling probability of each node in the large graph, divide all nodes in the large graph into multiple subsets, remove isolated nodes in each subset, and construct a connected subgraph for each subset. The importance sampling probability is a weighted sum of proximity and activity. Specifically, it is used for: The adjacency matrix of the large graph is normalized according to Formula 1 to obtain a normalized adjacency matrix. Based on the normalized adjacency matrix, cosine similarity is calculated according to Formula 2 to obtain the similarity between any two nodes in the large graph, and a similarity matrix is constructed. The activity of each node is calculated by the norm of the vector according to the normalized adjacency matrix, and the proximity of each node is calculated by the norm of the vector according to the similarity matrix. The activity and proximity are weighted and summed according to Formula 3 to obtain the importance sampling probability of each node in the large graph. Official 1; in, Represents the normalized adjacency matrix. Represents the adjacency matrix of a graph. The degree matrix of the graph is a diagonal matrix, and its diagonal elements D ii It is an adjacency matrix No. i The sum of all elements in the row, i =1,2,..., N , N This represents the total number of nodes in the graph. Represents the identity matrix; Official 2; in, Indicates the first 1 node Indicates the first 1 node Represents a node and nodes The similarity between nodes and nodes Cosine similarity between , Represents the normalized adjacency matrix. and These are the normalized adjacency matrices. The The and the first column vectors, This represents the total number of nodes in the graph. The 2-norm of a vector; Official 3; in, Represents any node Importance sampling probability, This represents the weight balancing parameter. Represents the similarity matrix. Representing the similarity matrix The middle corresponds to the current node column vectors, Describes the 2-norm of a vector. Representing vectors Norm, The value can be 1 or 2. Denotes the Frobenius norm of a matrix. This represents the total number of nodes in the graph. Represents the normalized adjacency matrix The The column corresponds to the current node. column vectors, Represents the normalized adjacency matrix The The column corresponds to the current node. Column vectors; The module is used to initialize the input feature matrix and model parameters of the graph convolutional neural network model, randomly sample the connected subgraph according to the importance sampling probability to obtain input data, input the input data into the graph convolutional neural network model for training, iteratively update the input feature matrix and update the model parameters through stochastic gradient descent until a trained CSIS-GCN model is obtained.
5. The system according to claim 4, characterized in that, The system further includes: The classification module is used to classify nodes in a large image using the trained CSIS-GCN model.
6. The system according to claim 4, characterized in that, The building module is specifically used for: The input data is input into the graph convolutional neural network model for training. The input feature matrix is iteratively updated according to Formula 4, and the model parameters are updated by stochastic gradient descent according to Formula 5 until a trained CSIS-GCN model is obtained. Official 4; Official 5; in, Represents a non-linear activation function. Represents the normalized adjacency matrix. This indicates the number of layers in the GCN model. Indicates the first The parameter matrix of the layer, Indicates the first A layer is a matrix consisting of embedding vectors. Indicates the first l The feature matrix of the +1 sampling layer node, whose update depends on the first sampling layer node. l Sampling layer, This indicates the model parameters to be updated. This represents the learning rate of the CSIS-GCN algorithm. Represents the cross-entropy loss function For model parameters The partial derivatives of .
7. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method for establishing a weighted importance sampling graph convolutional neural network (CSIS-GCN) model based on cosine similarity as described in any one of claims 1-3.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an implementation program for information transmission, which, when executed by a processor, implements the steps of the method for establishing a weighted importance sampling graph convolutional neural network (CSIS-GCN) model based on cosine similarity as described in any one of claims 1-3.
Citation Information
Patent Citations
Graph neural network sampling method for large-scale heterogeneous graph
CN115423073A
Node classification method for realizing heterogeneous primitive path aggregation based on graph convolution and self-attention mechanism
CN115828143A