Training method of graph node classification model, graph node classification method and related equipment

By dividing graph nodes into communities and combining a pre-trained language model with a community-node weight matrix to generate topological feature representations, the problems of computational complexity and accuracy in graph node classification methods are solved, achieving more efficient graph node classification.

CN120910329APending Publication Date: 2025-11-07CHINA MOBILE (XIONGAN) ICT CO LTD +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511001683.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing graph node classification methods suffer from increased computational complexity and low accuracy as the number of nodes increases. Graph neural networks ignore graph structure features when extracting node embeddings, resulting in node classification accuracy failing to meet expectations.

Method used

The graph nodes in the network relationship graph are divided into communities. The initial text embedding representation and the enhanced text embedding representation are determined by a pre-trained language model. During the training of the graph neural network, the community-node weight matrix is ​​used to generate the topological feature representation. The text embedding and the topological feature representation are fused to improve the classification accuracy.

Benefits of technology

By combining community tags and topological features, the system fully learns the multi-level textual information and topological relationships of graph nodes, thereby improving the accuracy of graph node classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910329A_ABST
    Figure CN120910329A_ABST
Patent Text Reader

Abstract

The invention discloses a graph node classification model training method, a graph node classification method and related equipment, and belongs to the technical field of graph representation learning. The method comprises the following steps: dividing a plurality of graph nodes in a network relation graph into communities, and determining community tags corresponding to the communities; according to the initial text of each graph node and a pre-training language model, determining an initial text embedding representation, an enhanced text embedding representation and a text prediction label of the graph node; obtaining a training sample set, wherein each training sample in the training sample set comprises an initial text embedding representation, an enhanced text embedding representation and a text prediction label corresponding to the same graph node, and a community label of a community to which the graph node belongs; and performing iterative training on the graph neural network by using the training sample set to obtain a graph node classification model. Through the mode, the graph neural network can fully learn the multi-level text information of the graph nodes and the topological relation of the graph nodes, so that the classification accuracy of the graph nodes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of graph representation learning, and in particular to a graph node classification model training method, a graph node classification method and related equipment. BACKGROUND

[0002] The graph node classification task is one of the basic tasks in graph analysis. Although the traditional graph convolutional network (GCN) performs well in the graph node classification task, when the number of nodes increases, the computational complexity of training also increases, resulting in an increased computational burden. In order to solve this problem, GraphSAGE reduces the computational complexity by neighbor sampling, but this also leads to the loss of some information.

[0003] At present, related research proposes a two-stage method for node classification tasks. In the first stage, the text attributes of the nodes in the graph are encoded using a pre-trained language model; in the second stage, the graph structure information of the nodes is further supplemented and enhanced through graph neural networks (GNNs). However, the graph neural network often ignores the graph structure features when extracting node embeddings. Although some graph structure feature extraction methods, such as matrix decomposition and random walk-based methods, have been applied in graph classification tasks, they are usually inefficient and redundant when combined with GNNs, resulting in the accuracy of node classification still failing to meet expectations. SUMMARY

[0004] The embodiments of the present application provide a graph node classification model training method, a graph node classification method and related equipment to at least solve the problem of low node classification accuracy in related methods in the graph node classification task.

[0005] In order to solve the above technical problems, the present application is implemented as follows: In a first aspect, the embodiments of the present application provide a method for training a graph node classification model, comprising: dividing a plurality of graph nodes in a network relationship graph into communities, and determining community labels corresponding to each community; determining an initial text embedding representation, an enhanced text embedding representation, and a text predicted label of each graph node according to an initial text of the graph node and a pre-trained language model; obtaining a training sample set, each training sample in the training sample set including the initial text embedding representation, the enhanced text embedding representation, and the text predicted label of a same graph node, and a community label of a community to which the graph node belongs; iteratively training a graph neural network using the training sample set, wherein, in the iterative training process, a community-node weight matrix is determined according to the community labels of the plurality of graph nodes; a topological feature representation is generated according to the community-node weight matrix and the community labels of the plurality of graph nodes; a node predicted result is determined according to the topological feature representation, and the initial text embedding representation, the enhanced text embedding representation of each graph node; a predicted loss value is determined according to the node predicted result and the community labels of the plurality of graph nodes; the community-node weight matrix is adjusted according to the predicted loss value until a preset first convergence condition is reached, and the graph neural network after convergence is taken as the graph node classification model.

[0006] In a second aspect, the embodiments of the present application provide a method for classifying graph nodes, comprising: dividing a plurality of target graph nodes in a target network relationship graph into communities, and determining target community labels corresponding to each community; inputting the target community labels and text embedding representations corresponding to the plurality of target graph nodes into a graph node classification model to obtain a graph node classification result; the graph node classification model is trained according to the method of the first aspect.

[0007] In a third aspect, an embodiment of the present application provides a device for training a graph node classification model, comprising: a community determination module configured to divide a plurality of graph nodes in a network relationship graph into communities, and determine community labels corresponding to the communities; a text prediction module configured to determine initial text embedding representations, enhanced text embedding representations, and text prediction labels of the graph nodes according to initial texts of the graph nodes and a pre-trained language model; a sample acquisition module configured to acquire a training sample set, wherein each training sample in the training sample set comprises initial text embedding representations, enhanced text embedding representations, and text prediction labels of a same graph node, and a community label of a community to which the graph node belongs; and a network training module configured to iteratively train a graph neural network using the training sample set, wherein in the iterative training process, a community-node weight matrix is determined according to community labels of the plurality of graph nodes; a topological feature representation is generated according to the community-node weight matrix and the community labels of the plurality of graph nodes; a node prediction result is determined according to the topological feature representation, and the initial text embedding representations and the enhanced text embedding representations of the graph nodes; a prediction loss value is determined according to the node prediction result and the community labels of the plurality of graph nodes; the community-node weight matrix is adjusted according to the prediction loss value until a preset first convergence condition is reached, and the graph neural network after convergence is taken as the graph node classification model.

[0008] In a fourth aspect, an embodiment of the present application provides a device for classifying graph nodes, comprising: a label determination module configured to divide a plurality of target graph nodes in a target network relationship graph into communities, and determine target community labels corresponding to the communities; and a node classification module configured to input target community labels and text embedding representations of the plurality of target graph nodes into a graph node classification model to obtain a graph node classification result, wherein the graph node classification model is trained according to the method of the first aspect.

[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the method of the first aspect or the second aspect.

[0010] In a sixth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores programs or instructions, and the programs or instructions are executed by a processor to implement the steps of the method of the first aspect or the second aspect.

[0011] In a seventh aspect, an embodiment of the present application provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions that, when executed by a computer, cause the computer to perform the steps of the method according to the first aspect or the second aspect.

[0012] In the embodiment of the present application, a plurality of graph nodes in the network relationship graph are divided into communities, and community labels corresponding to each community are determined; initial text embedding representations, enhanced text embedding representations and text prediction labels of the graph nodes are determined according to initial texts of each graph node and a pre-trained language model; a training sample set is obtained, each training sample in the training sample set including initial text embedding representations, enhanced text embedding representations and text prediction labels corresponding to the same graph node, and a community label of a community to which the graph node belongs; and the graph neural network is iteratively trained using the training sample set to obtain a graph node classification model. In this way, the multi-level text embedding representations of the graph nodes are determined by using the pre-trained language model, and in the training process of the graph neural network, the topological feature representations are generated according to the community-node weight matrix and the community labels of the plurality of graph nodes, and the text embedding representations and the topological feature representations are fused, so that the graph neural network fully learns the multi-level text information of the graph nodes and the topological relationship of the graph nodes, thereby improving the accuracy of the graph node classification.

[0013] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0014] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the present application.

[0015] Figure 1 A flowchart of a training method of a graph node classification model provided by an embodiment of the present application is shown; Figure 2 A flowchart of a community division method provided by an embodiment of the present application is shown; Figure 3 An architectural diagram of the training method of the graph node classification model provided by an embodiment of the present application is shown; Figure 4 A flowchart of pre-trained language model assisted data augmentation and fine-tuning provided by an embodiment of the present application is shown; Figure 5 A flowchart of GNN training and dynamic community-node weight matrix updating provided by an embodiment of the present application is shown; Figure 6 A flowchart of a graph node classification method provided by an embodiment of the present application is shown; Figure 7 A structural schematic diagram of a training device of a graph node classification model provided by an embodiment of the present application is shown. Figure 8 A structural schematic diagram of a graph node classification device provided by an embodiment of the present application is shown. Figure 9 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0016] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The following description is only exemplary and is not intended to limit the scope, applicability or configuration of the application. Rather, the following description is intended to describe some exemplary embodiments consistent with the present application. Alternative embodiments will become apparent to those skilled in the art to which the present application pertains. Furthermore, various omissions, substitutions and changes in the form of the detail of the methods and apparatuses as described herein can be made by those skilled in the art without departing from the spirit of the application.

[0017] Figure 1 A flowchart of a training method of a graph node classification model provided by an embodiment of the present application is shown. The execution subject of the method can be a terminal device or a server. The terminal device can be a device such as a personal computer, or a mobile terminal device such as a mobile phone or a tablet computer. The terminal device can be a terminal device used by a user. The server can be a standalone server, or a server cluster composed of multiple servers. The server can be a background server of a service, or a background server of a platform or an application (such as a cloud computing platform, a graph database and a graph analysis platform, etc.). In the present embodiment, the execution subject is taken as an example of a server. For the case of a terminal device, the relevant content described below can be processed, and thus the description is omitted here. As shown in the figure, the method 100 can include the following steps: Step 101: dividing a plurality of graph nodes in a network relationship graph into communities, and determining community labels corresponding to each community.

[0018] The network relationship graph in the step 101 can be a graph for describing interpersonal relationships in a social network, or a graph for describing relationships between researchers, papers and citations in an academic network. The graph nodes in the step 101 can be used to represent different entities, such as users in a social network, papers in an academic network, etc. The communities in the step 101 are subsets or groups divided according to the connection relationships between the plurality of graph nodes.

[0019] In this step, the topology of the graph can be extracted by a community detection algorithm, and the plurality of graph nodes in the graph can be divided into communities according to the topology of the graph, each community corresponding to a community label, the community label being a unique identifier for identifying a community, for example, in a community network, the community label is used to represent mutually related users; in the academic field, the community label is used to represent closely related research topics or papers. Here, the community detection algorithm can be louvain algorithm, Girvan-Newman algorithm, etc.

[0020] In an exemplary embodiment, as shown in Figure 2 , the above step 101 specifically includes the following steps: step 1011: graph data loading; step 1012: judging whether the loaded graph data is a DGL (Deep Graph Library) graph, if not, entering step 1013, if yes, entering step 1014; step 1013: converting the graph data and the label corresponding to the graph data into a DGL graph G; step 1014: using a community detection algorithm, such as louvain algorithm, to detect communities of the graph G, and through iteration to maximize the module degree gain each time to obtain the mapping of nodes to communities ; step 1015: extracting the community label corresponding to the node and saving, wherein, V ( G ) is the node set of the graph G , and the number of communities is calculated and saved, wherein, is the set of all community labels. The saved community label of the community to which the node belongs and the number of node communities are used as input for subsequent steps.

[0021] In this way, the community to which each graph node belongs can be distinguished by the community label, which is conducive to more accurate classification and prediction of the graph nodes in the subsequent steps.

[0022] Step 102: determining the initial text embedding representation, enhanced text embedding representation and text prediction label of the graph node according to the initial text of each graph node and the pre-trained language model.

[0023] The initial text in the step 102 can be a user description in a social network, article content, a review, etc. The initial text embedding representation in the step 102 is a vector form converted from the initial text. The initial text embedding representation can be generated by encoding the initial text using a pre-trained language model (such as BERT, GPT, etc.). The enhanced text embedding representation in the step 102 is a representation obtained by further enhancing the initial text embedding representation. The text prediction label in the step 102 is a prediction result obtained by classifying the initial text. For example, in a social network, the interest label of a user or the topic category of an article can be predicted.

[0024] In this step, the prompt word can be determined according to the initial text of the graph node, and the prompt word is input into the pre-trained language model. The initial text embedding representation, the enhanced text embedding representation, and the text prediction label of the graph node are obtained by encoding and predicting the initial text of the graph node using the pre-trained language model.

[0025] In this way, the pre-trained language model can obtain more rich text information of the graph node, so that the graph neural network can fully learn the text information of the graph node, and the accuracy of the graph node classification can be improved.

[0026] Step 103: Obtain a training sample set. Each training sample in the training sample set includes the initial text embedding representation, the enhanced text embedding representation, and the text prediction label of the corresponding same graph node, and the community label of the community to which the graph node belongs.

[0027] In this step, the training data set is constructed, including the initial text embedding representation, the enhanced text embedding representation, and the text prediction label of each graph node, and the community label of the community to which the graph node belongs. These training samples will be used to train the graph neural network, so as to obtain the graph node classification model.

[0028] Step 104: Iteratively train the graph neural network using the training sample set to obtain the graph node classification model.

[0029] In the iterative training process, the community-node weight matrix is determined according to the community labels of the plurality of graph nodes. The topological feature representation is generated according to the community-node weight matrix and the community labels of the plurality of graph nodes. The node prediction result is determined according to the topological feature representation, and the initial text embedding representation and the enhanced text embedding representation of each graph node. The prediction loss value is determined according to the node prediction result and the community labels of the plurality of graph nodes. The community-node weight matrix is adjusted according to the prediction loss value until a preset first convergence condition is reached. The converged graph neural network is used as the graph node classification model.

[0030] In an exemplary embodiment, as shown in FIG. 1, the method for training a graph node classification model includes the following steps.Figure 3 As shown, after obtaining a plurality of graph nodes, on one hand, the topological structure features of the graph are extracted in a community detection manner to obtain community labels corresponding to the graph nodes, and on the other hand, initial text embedding representations and text prediction labels are generated by a pre-trained language model LLM. Then, the initial text embedding representations and the text prediction labels are used to fine-tune a pre-trained language model with fewer parameters to obtain enhanced text embedding representations. Further, a community-node weight matrix is determined according to the community labels of the plurality of graph nodes: wherein, M is the number of community labels, N is the number of graph nodes, a topological feature representation is determined according to the community-node weight matrix W and the community labels of the plurality of graph nodes; a node prediction result is determined according to the topological feature representation, the initial text embedding representation and the enhanced text embedding representation of each graph node; a prediction loss value is determined according to the node prediction result and the community labels of the plurality of graph nodes; the community-node weight matrix is adjusted according to the prediction loss value until a preset first convergence condition is reached. The converged graph neural network is used as a graph node classification model. The first convergence condition can be that the prediction loss value converges to a preset range, or the number of iterations reaches a preset number threshold.

[0031] In this way, in the training process of the graph neural network, the topological feature representation is determined according to the community-node weight matrix and the community labels of the plurality of graph nodes, and the text embedding representation and the topological feature representation are fused, so that the graph neural network fully learns the multi-level text information of the graph nodes and the topological relationship between the graph nodes, thereby improving the accuracy of graph node classification.

[0032] In some possible implementations, in step 102, the initial text embedding representation, the enhanced text embedding representation and the text prediction label of each graph node are determined according to the initial text of each graph node and the pre-trained language model, including: a prompt word is generated according to the initial text of each graph node; the enhanced text and the text prediction label of the graph node are obtained by inputting the prompt word into the pre-trained language model; a node enhanced text representation is determined according to the initial text, the enhanced text and the text prediction label; the node enhanced text representation is input into a text classification model to obtain a prediction result, and a text prediction loss value is determined according to the prediction result and the text prediction label. The parameters of the text classification model are adjusted according to the text prediction loss value until a preset second convergence condition is reached, and the initial text embedding representation and the enhanced text embedding representation of the graph node are obtained; the text classification model includes the pre-trained language model and a linear output layer.

[0033] In an exemplary embodiment, as Figure 4As shown, first, the prompt word Prompt is generated as the input of the pre-training language model. Taking the ogbn-arxiv dataset as an example, first, the large model is used to generate the prompt keywords including Abstract, Title and Question ; wherein, An open question is provided to indicate that the pre-training language model predicts one or more class labels of the initial text. It requires the pre-training language model to list the resulting text prediction labels from most likely to least likely and explain why each label is selected. The pre-training language model response results in a text prediction label , the enhanced text , wherein, represents the prediction labels sorted from highest to lowest. In combination with the initial text , the node-enhanced text representation can be obtained for each node.

[0034] The initial text and the data augmentation text node embedding are obtained by fine-tuning a pre-training language model with fewer parameters. The initial text embedding is obtained by tokenizing the initial text words into the basic units of the pre-training language model input, and then adding a linear output layer to obtain , and then using a cross-entropy loss function to fine-tune the model to produce the initial text embedding representation. The enhanced text is obtained by tokenizing the enhanced text words into the basic units of the model input, and then adding a linear output head to obtain , and then using a cross-entropy loss function to fine-tune the model until a preset second convergence condition is reached, resulting in an enhanced text embedding representation. The second convergence condition can be that the number of iterations reaches a preset iteration threshold, or that the model's prediction loss converges within a preset threshold.

[0035] In some possible implementations, in step 104, the community-node weight matrix is determined according to the community labels of the plurality of graph nodes, including: According to the number of graph nodes and the number of community labels, an initial weight matrix is constructed; for each graph node, the target element corresponding to the community label of the graph node in the initial weight matrix is set to a first preset value, and the other elements in the initial weight matrix except the target element are set to a second preset value; the first preset value is greater than the second preset value; the initial weight matrix after the setting is determined as the community-node weight matrix.

[0036] Wherein, the first preset value in the above step can be set to 1, 2, 10, etc., which can be adjusted according to actual needs, and is not specifically limited here; the second preset value in the above step is less than the first preset value, which can be set to 0, 0.1, 1, etc., which can be adjusted according to actual needs, and is not specifically limited here.

[0037] In an exemplary embodiment, for the initial weight matrix: Each column is used to indicate the weight of a single node on all communities. In the initialization phase, the community to which the node obtained by community detection belongs is assigned to 1 in the corresponding W matrix, while other communities are assigned to a number much smaller than 1. Then, the initial weight matrix is dynamically updated during the training process to obtain the community-node weight matrix.

[0038] In some possible implementations, in the step 104, the node prediction result is determined according to the topological feature representation, and the initial text embedding representation and the enhanced text embedding representation of each graph node, including: According to the topological feature representation and the initial text embedding representation of each graph node, an initial text node is determined, and a first prediction result is obtained by inputting the initial text node into the graph neural network; according to the topological feature representation and the enhanced text embedding representation of each graph node, an enhanced text node is determined, and a second prediction result is obtained by inputting the enhanced text node into the graph neural network; according to the topological feature representation and the text prediction label of each graph node, a text prediction node is determined, and a third prediction result is obtained by inputting the text prediction node into the graph neural network; and according to the first prediction result, the second prediction result and the third prediction result, the node prediction result is determined.

[0039] In an exemplary embodiment, as shown in Figure 5 , for each community, a corresponding embedding representation is generated, and the embedding representations of all communities are , then for all nodes, the weighted sum of all communities is calculated as the topological feature representation of the node . The latest representation of the initial text node is obtained by combining the topological feature representation and the initial text embedding representation , as the input of the GNN model, the first prediction result is obtained by the GNN model Revgat, the enhanced text node and the text prediction label are trained with the initial text node, and the node embedding representations are and , and the corresponding second prediction result is , and the third prediction result is Finally, the node prediction result is determined according to the first prediction result, the second prediction result and the third prediction result, and specifically, the final prediction result can be obtained by averaging the three prediction results Finally, the model parameters and the community-node weight matrix W are calculated and updated by using a cross-entropy loss function.

[0040] Figure 6 A flowchart of a graph node classification method provided by an embodiment of the present application is shown, and the execution subject of the method can be a terminal device or a server. The terminal device can be a device such as a personal computer, or a mobile terminal device such as a mobile phone or a tablet computer. The terminal device can be a terminal device used by a user. The server can be a standalone server or a server cluster composed of multiple servers. The server can be a background server of a certain service, or a background server of a certain platform or application (for example, a cloud computing platform, a graph database and a graph analysis platform, etc.). In the embodiment of the present application, the execution subject is taken as an example of a server. For the case of a terminal device, the relevant content described below can be processed, and details are not repeated here. As shown in the figure, the method 600 can include the following steps: Step 601: dividing a plurality of target graph nodes in a target network relationship graph into communities, and determining target community labels corresponding to each community.

[0041] In this step, the topological structure of the graph can be extracted by a community detection algorithm, and the plurality of target graph nodes in the graph can be divided into communities according to the topological structure of the graph. Each community corresponds to a target community label, and the target community label is a unique identifier for identifying a community. For example, in a community network, a community label is used to represent mutually related users; in an academic field, a community label is used to represent closely related research topics or papers. Here, the community detection algorithm can be a louvain algorithm, a Girvan-Newman algorithm, etc.

[0042] Step 602: inputting the target community labels and text embedding representations corresponding to the plurality of target graph nodes into a graph node classification model to obtain a graph node classification result. The graph node classification model is trained according to the training method of the graph node classification model described above.

[0043] In this step, the text embedding representations of the plurality of target nodes can be obtained by a pre-trained language model, and the target community labels and text embedding representations corresponding to the plurality of target graph nodes can be input into the graph node classification model to obtain a graph node classification result.

[0044] Thus, by training the graph neural network through the training method of the graph node classification model, the graph neural network can sufficiently learn the multi-level text information of the graph nodes and the topological relationship between the graph nodes, and thus the obtained graph node classification model can effectively capture the text information and the topological relationship of multiple graph nodes, thereby improving the accuracy of graph node classification.

[0045] Figure 7 A structure diagram of a training device of a graph node classification model is shown, which can implement all or part of the embodiments of the present application. Figure 1 The training device 700 includes: A community determination module 710 is configured to divide multiple graph nodes in a network relationship graph into communities and determine community labels corresponding to the communities. A text prediction module 720 is configured to determine initial text embedding representations, enhanced text embedding representations, and text prediction labels of the graph nodes according to initial texts of the graph nodes and a pre-trained language model. A sample acquisition module 730 is configured to acquire a training sample set, wherein each training sample in the training sample set includes initial text embedding representations, enhanced text embedding representations, and text prediction labels of a same graph node, and a community label of a community to which the graph node belongs. A network training module 740 is configured to iteratively train a graph neural network using the training sample set, wherein in the iterative training process, a community-node weight matrix is determined according to community labels of the multiple graph nodes, a topological feature representation is generated according to the community-node weight matrix and the community labels of the multiple graph nodes, a node prediction result is determined according to the topological feature representation and the initial text embedding representations and the enhanced text embedding representations of the graph nodes, a prediction loss value is determined according to the node prediction result and the community labels of the multiple graph nodes, the community-node weight matrix is adjusted according to the prediction loss value until a preset first convergence condition is reached, and the converged graph neural network is taken as a graph node classification model.

[0046] In some possible implementation manners, the text prediction module 720, when being configured to determine initial text embedding representations, enhanced text embedding representations, and text prediction labels of the graph nodes according to initial texts of the graph nodes and a pre-trained language model, is specifically configured to: generate prompt words according to the initial texts of the graph nodes; acquire enhanced texts and text prediction labels of the graph nodes by inputting the prompt words into the pre-trained language model; determine node enhanced text representations according to the initial texts, the enhanced texts, and the text prediction labels; inputting the node enhanced text representation into a text classification model to obtain a prediction result, determining a text prediction loss value according to the prediction result and the text prediction label, adjusting parameters of the text classification model according to the text prediction loss value until a preset second convergence condition is reached, and obtaining the initial text embedding representation and the enhanced text embedding representation of the graph node; the text classification model comprises a pre-trained language model and a linear output layer.

[0047] In some possible implementation manners, the network training module 740, when used for determining the community-node weight matrix according to the community labels of the plurality of graph nodes, is specifically configured to: construct an initial weight matrix according to the number of graph nodes and the number of community labels; for each graph node, set a target element corresponding to the community label of the graph node in the initial weight matrix as a first preset value, and set other elements in the initial weight matrix except the target element as a second preset value; the first preset value is greater than the second preset value; determine the initial weight matrix after the setting as the community-node weight matrix.

[0048] In some possible implementation manners, the network training module 740, when used for determining the node prediction result according to the topological feature representation and the initial text embedding representation and the enhanced text embedding representation of each graph node, is specifically configured to: determine an initial text node according to the topological feature representation and the initial text embedding representation of each graph node, and obtain a first prediction result by inputting the initial text node into the graph neural network; determine an enhanced text node according to the topological feature representation and the enhanced text embedding representation of each graph node, and obtain a second prediction result by inputting the enhanced text node into the graph neural network; determine a text prediction node according to the topological feature representation and the text prediction label of each graph node, and obtain a third prediction result by inputting the text prediction node into the graph neural network; determine the node prediction result according to the first prediction result, the second prediction result and the third prediction result.

[0049] The embodiment of the application provides a kind of graph node classification model training device, including community determination module, text prediction module, sample acquisition module and network training module;Community determination module divides multiple graph nodes in network relationship graph into community, determines the community label corresponding to each community;Text prediction module determines the initial text embedding representation of each graph node, enhanced text embedding representation and text prediction label according to the initial text of each graph node and pre-training language model;Sample acquisition module obtains training sample set, each training sample in the training sample set includes the initial text embedding representation, enhanced text embedding representation and text prediction label corresponding to the same graph node, and the community label of the community to which the graph node belongs;Network training module uses the training sample set to iteratively train graph neural network, and obtains graph node classification model.This way, the text embedding representation of graph node is determined by pre-training language model, in the training process of graph neural network, according to community-node weight matrix and the community label of multiple graph nodes, topological feature representation is determined, and text embedding representation and topological feature representation are fused, so that graph neural network fully learns multi-level text information of graph node and topological relationship between graph nodes, to improve the accuracy of graph node classification.

[0050] Figure 8 The structure schematic diagram of the graph node classification device provided by the embodiment of the application is shown, which can realize all or part of the contents in the embodiment shown Figure 6 The graph node classification device 800 includes: The label determination module 810 is configured to divide multiple target graph nodes in a target network relationship graph into communities, and determine target community labels corresponding to each community. The node classification module 820 is configured to input the target community labels and text embedding representations corresponding to the multiple target graph nodes into a graph node classification model to obtain a graph node classification result.

[0051] Therefore, the graph node classification model obtained can effectively capture the text information and topological relationship of multiple graph nodes, thereby improving the accuracy of graph node classification.

[0052] Figure 9A hardware structure schematic diagram of an electronic device is shown, and reference is made to the diagram. At the hardware level, the electronic device 900 includes a processor 910, optionally, an internal bus 920, a network interface 930, and a memory. The memory can include a memory 941, such as a high-speed random access memory (RAM), and can also include a non-volatile memory 942, such as at least one disk memory. Of course, the electronic device can also include other hardware required by the business.

[0053] The processor 910, the network interface 930, and the memory can be connected to each other through the internal bus 920, which can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one bidirectional arrow is shown in the diagram, but it does not mean that there is only one bus or only one type of bus.

[0054] The memory stores programs. Specifically, the programs can include program code, which includes computer operation instructions. The memory can include the memory 941 and the non-volatile memory 942, and provide instructions and data to the processor 910.

[0055] The processor 910 reads the corresponding computer program from the non-volatile memory 942 into the memory and then runs, and forms a device for positioning a target user at the logical level. The processor 910 executes the programs stored in the memory, and specifically executes: Figure 1 Or Figure 6 The method disclosed in the illustrated embodiment achieves the functions and beneficial effects of the methods described in the foregoing method embodiments, and will not be described here.

[0056] The above as described in the present application Figure 1 Or Figure 6The method disclosed by the embodiment shown can be applied to the processor 910 or implemented by the processor 910. The processor 910 can be an integrated circuit chip with processing capability. In the implementation process, each step of the above method can be completed by integrated logic circuits in hardware or instructions in software form in the processor 910. The processor 910 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiment of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiment of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above method.

[0057] The computer device can also execute the methods described in the foregoing method embodiments, and achieve the functions and beneficial effects of the methods described in the foregoing method embodiments, which will not be described here.

[0058] Of course, in addition to the software implementation, the electronic device of the present application does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or a logic device.

[0059] The embodiment of the present application also proposes a computer readable storage medium, the computer readable medium stores one or more programs, the one or more programs when being executed by the electronic device including a plurality of application programs, make the electronic device execute Figure 1 Or Figure 6 The method disclosed by the embodiment shown and the functions and beneficial effects of the methods described in the foregoing method embodiments are not described here.

[0060] The computer readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disc or an optical disc, etc.

[0061] Further, the embodiment of the present application further provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer readable storage medium, the computer program comprising program instructions which, when executed by a computer, implement the following processes: Figure 1 or Figure 6 The method disclosed in the embodiments and the functions and advantages of the methods described in the foregoing method embodiments are not described here again.

[0062] The embodiments of the present application can be applied to various electronic device cooperation or interconnection scenarios, including: mobile phone and notebook computer / tablet computer cooperation and interconnection; mobile terminal and smart television / display cooperation and interconnection; mobile phone, or tablet computer and vehicle entertainment system cooperation and interconnection; mobile terminal and smart conference system cooperation and interconnection, etc. Thus, the needs of users in smart home, smart office, smart travel and other diversified scenarios are met.

[0063] In summary, the above only describes the preferred embodiments of the present application, and does not limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0064] The system, device, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain function. A typical implementation device is a computer. Specifically, the computer may, for example, be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0065] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0066] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article or apparatus that includes a list of elements does not only include those elements, but also includes other elements not explicitly listed, or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0067] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

Claims

1. A method for training a graph node classification model, the method comprising: The method comprises the following steps: dividing a plurality of graph nodes in a network relationship graph into communities, and determining community labels corresponding to the communities; determining initial text embedding representations, enhanced text embedding representations and text prediction labels of the graph nodes according to initial texts of each of the graph nodes and a pre-trained language model; obtaining a training sample set, wherein each training sample in the training sample set comprises initial text embedding representations, enhanced text embedding representations and text prediction labels corresponding to a same graph node, and a community label of a community to which the graph node belongs; iteratively training a graph neural network using the training sample set, wherein, in the iterative training process, a community-node weight matrix is determined according to the community labels of the plurality of graph nodes; and a topological feature representation is generated according to the community-node weight matrix and the community labels of the plurality of graph nodes; determining a node prediction result according to the topological feature representation and the initial text embedding representations and the enhanced text embedding representations of each of the graph nodes; determining a prediction loss value according to the node prediction result and the community labels of the plurality of graph nodes; adjusting the community-node weight matrix according to the prediction loss value until a preset first convergence condition is reached, and taking the graph neural network after convergence as a graph node classification model.

2. The method of claim 1, wherein, The method comprises the following steps: generating prompt words according to the initial texts of each of the graph nodes; obtaining enhanced texts and text prediction labels of the graph nodes by inputting the prompt words into the pre-trained language model; determining a node enhanced text representation according to the initial texts, the enhanced texts and the text prediction labels; inputting the node enhanced text representation into a text classification model to obtain a prediction result, determining a text prediction loss value according to the prediction result and the text prediction labels, adjusting parameters of the text classification model according to the text prediction loss value until a preset second convergence condition is reached, and obtaining the initial text embedding representations and the enhanced text embedding representations of the graph nodes; the text classification model comprises the pre-trained language model and a linear output layer.

3. The method of claim 1, wherein, The method comprises the following steps: constructing an initial weight matrix according to the number of graph nodes and the number of community labels; for each graph node, setting a target element corresponding to the community label of the graph node in the initial weight matrix to a first preset value, and setting other elements in the initial weight matrix except the target element to a second preset value; the first preset value is greater than the second preset value; determining the initial weight matrix after the setting as a community-node weight matrix.

4. The method of claim 1, wherein, The method comprises the following steps: determine an initial text node according to the topological feature representation and the initial text embedding representation of each graph node, and obtain a first prediction result by inputting the initial text node into the graph neural network; determine an enhanced text node according to the topological feature representation and the enhanced text embedding representation of each graph node, and obtain a second prediction result by inputting the enhanced text node into the graph neural network; determine a text prediction node according to the topological feature representation and the text prediction label of each graph node, and obtain a third prediction result by inputting the text prediction node into the graph neural network; determine a node prediction result according to the first prediction result, the second prediction result and the third prediction result.

5. A method of classifying graph nodes, the method comprising: comprising: dividing a plurality of target graph nodes in a target network relationship graph into communities, and determining target community labels corresponding to each community; inputting the target community labels and text embedding representations corresponding to the plurality of target graph nodes into a graph node classification model to obtain a graph node classification result; the graph node classification model is trained according to the method of any one of claims 1 to 4. 6.A device for training a graph node classification model, the device comprising: comprising: a community determination module, configured to divide a plurality of graph nodes in a network relationship graph into communities, and determine community labels corresponding to each community; a text prediction module, configured to determine initial text embedding representations, enhanced text embedding representations and text prediction labels of the graph nodes according to initial texts of each of the graph nodes and a pre-trained language model; a sample acquisition module, configured to acquire a training sample set, wherein each training sample in the training sample set comprises initial text embedding representations, enhanced text embedding representations and text prediction labels corresponding to a same graph node, and a community label of a community to which the graph node belongs; a network training module, configured to iteratively train a graph neural network using the training sample set, wherein in the iterative training process, a community-node weight matrix is determined according to the community labels of the plurality of graph nodes, and a topological feature representation is generated according to the community-node weight matrix and the community labels of the plurality of graph nodes; determine a node prediction result according to the topological feature representation and the initial text embedding representation, the enhanced text embedding representation of each graph node; determine a prediction loss value according to the node prediction result and the community labels of the plurality of graph nodes; adjust the community-node weight matrix according to the prediction loss value until a preset first convergence condition is reached, and use the converged graph neural network as a graph node classification model.

7. A graph node classification apparatus characterized by comprising: comprising: a label determination module, configured to divide a plurality of target graph nodes in a target network relationship graph into communities, and determine target community labels corresponding to each community; a node classification module, configured to input target community labels and text embedding representations corresponding to the plurality of target graph nodes into a graph node classification model to obtain a graph node classification result; the graph node classification model is trained according to the method of any one of claims 1 to 4.

8. An electronic device, comprising: The electronic device comprises a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method according to any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a program or instructions, the program or instructions, when executed by a processor, implementing the steps of the method according to any one of claims 1 to 5.

10. A computer program product, characterised in that, The computer program product comprises a computer program stored on a non-transitory computer readable storage medium, the computer program comprising program instructions which, when executed by a computer, cause the computer to perform the steps of the method according to any one of claims 1 to 5.