A method and device for pushing information based on graph network

By determining the subgraph structure of the node to be processed in the graph network and training the node feature extraction model, adding graph structure feature vectors to constrain the training process, the problem of low rationality in graph structure learning and sample in the prior art is solved, and the training accuracy and information push accuracy are improved.

CN112069398BActive Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010856777.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-24
Publication Date
2025-05-13
Estimated Expiration
2040-08-24

AI Technical Summary

Technical Problem

The prior art cannot effectively learn the specified graph structure in the graph network, and the rationality of the positive and negative sample pairs formed by the connection relationship is low, resulting in low training accuracy and convergence speed, which in turn reduces the accuracy of node vector representation and information push.

Method used

By obtaining the node to be processed, its subgraph structure in the graph network is determined, and the subgraph structure is input to the node feature extraction model to obtain the aggregated feature vector. Based on these feature vectors, the matching nodes are filtered as information push nodes and information push is performed. At the same time, train node feature extraction models and add graph structure feature vectors to constrain the training process.

Benefits of technology

It improves the accuracy and convergence speed of graph network training, enhances the accuracy of node feature representation, and thus improves the accuracy of information push.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112069398B_ABST
    Figure CN112069398B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technology, and in particular to a method and device for pushing information based on a graph network, which comprises obtaining a node to be processed; determining a subgraph structure with the node to be processed as a central node in the graph network, wherein the subgraph structure includes nodes and connection relationships connected to the node to be processed; inputting the subgraph structure into a node feature extraction model to obtain an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model; based on the aggregate feature vectors of each node in the graph network, screening nodes matching the node to be processed as information pushing nodes; and pushing business information corresponding to the information pushing node, so that a graph structure feature vector is added on the basis of the aggregate feature vector during training, that is, graph structure information is added for constraint, thereby improving the convergence speed and accuracy of training, and further, the subgraph structure can be used as input during application, thereby improving the accuracy of node representation in the graph network and the accuracy of information pushing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and device for pushing information based on a graph network. Background Art

[0002] Graph Network (GN) is the most direct tool for describing community relationship chains. For example, in business scenarios such as finance and the Internet of Things, a graph network consists of nodes and edges. It is a generalized artificial neural network based on a graph structure. The graph embedding algorithm is the process of mapping a graph data into a low-dimensional dense vector, encoding the nodes so that they can be easily applied to specific downstream tasks, such as information push. In related technologies, an unsupervised graph embedding algorithm can be used to train the graph network. For example, the GraphSAGE algorithm samples the local neighbors of a node and aggregates them into features, and then learns based on the features of the node itself, so that some graph structure information of the graph network can be learned. However, this method cannot learn and train the specified graph structure, and generally only forms positive and negative sample pairs through connection relationships, which easily leads to low rationality of positive and negative samples. It only considers the adjacency relationship between nodes and lacks the graph structure description between nodes, which reduces the training accuracy and convergence speed, thereby reducing the accuracy of node vector representation and the accuracy of information push. Summary of the invention

[0003] The embodiments of the present application provide a method and device for information push based on a graph network to improve the accuracy of information push.

[0004] The specific technical solutions provided by the embodiments of this application are as follows:

[0005] An embodiment of the present application provides an information push method based on a graph network, including:

[0006] Get the nodes to be processed;

[0007] Determine a subgraph structure with the node to be processed as a central node in the graph network, wherein the subgraph structure includes nodes and connection relationships connected to the node to be processed;

[0008] Inputting the subgraph structure into a node feature extraction model, obtaining an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model; the aggregate feature vector represents the node to be processed in the graph network;

[0009] Based on the aggregated feature vectors of each node in the graph network, selecting nodes matching the node to be processed as information pushing nodes;

[0010] The business information corresponding to the information pushing node is pushed.

[0011] Another embodiment of the present application provides a graph network training method, including:

[0012] For a target node in a graph network, determining a positive sample node and a negative sample node of the target node in the graph network, wherein a correlation degree between the target node and the positive sample node is greater than a correlation degree between the target node and the negative sample node;

[0013] Extracting respectively the aggregate feature vector and the graph structure feature vector of the target node and the associated positive sample nodes and negative sample nodes, wherein the graph structure feature vector represents the features of the subgraph structure containing a preset order;

[0014] A node feature extraction model is trained based on the target node, and the aggregate feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes. The node feature extraction model is used to extract the aggregate feature vectors of the nodes in the graph network.

[0015] Another embodiment of the present application provides a graph network training device, including:

[0016] A first determination module is used to determine, for a target node in a graph network, a positive sample node and a negative sample node of the target node in the graph network, wherein a correlation degree between the target node and the positive sample node is greater than a correlation degree between the target node and the negative sample node;

[0017] A second determination module is used to extract the aggregate feature vector and graph structure feature vector of the target node and the associated positive sample nodes and negative sample nodes, respectively, wherein the graph structure feature vector represents the features of the subgraph structure containing a preset order;

[0018] The training module is used to train a node feature extraction model based on the target node and the aggregate feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes, and the node feature extraction model is used to extract the aggregate feature vectors of the nodes in the graph network.

[0019] Another embodiment of the present application provides an information push device based on a graph network, including:

[0020] An acquisition module, used to acquire nodes to be processed;

[0021] A determination module, used to determine a subgraph structure with the node to be processed as a central node in the graph network, wherein the subgraph structure includes nodes and connection relationships connected to the node to be processed;

[0022] An extraction module, used for inputting the subgraph structure into a node feature extraction model, and obtaining an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model; the aggregate feature vector represents the node to be processed in the graph network;

[0023] A screening module, used for screening nodes matching the nodes to be processed as information push nodes based on the aggregated feature vectors of each node in the graph network;

[0024] The push module is used to push the business information corresponding to the information push node.

[0025] Another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the processor implements the steps of any one of the above-mentioned graph network training or graph network-based information push methods.

[0026] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned graph network training or graph network-based information push methods.

[0027] In an embodiment of the present application, a node to be processed is obtained; a subgraph structure with the node to be processed as a central node in a graph network is determined, and the subgraph structure is input into a node feature extraction model to obtain an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model; and then based on the aggregate feature vectors of each node in the graph network, a node matching the node to be processed is screened as an information push node; and the business information corresponding to the information push node is pushed. In this way, since a graph structure feature vector is added on the basis of the aggregate feature vector during training, that is, the graph structure information is added for constraint, the irrationality of positive and negative samples can be reduced, the training convergence speed can be accelerated, and the specified motif structure can be learned and recalled, which can be generalized to new nodes, reducing the repeated calculation of the graph structure feature vector, and making up for the description of the topological structure information of the graph network, thereby improving the accuracy of training. Then, the node feature extraction model obtained after training takes the subgraph structure of the node to be processed as input, and can obtain the aggregate feature vector representation of the node to be processed, and can also improve the accuracy of the node representation in the graph network, and then applied to a specific application scenario, the accuracy of information push can be improved based on the obtained aggregate feature vector representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A schematic diagram of an application architecture provided in an embodiment of the present application;

[0029] Figure 2This is a flow chart of the network training method in the embodiment of the present application;

[0030] Figure 3 This is an example diagram of a graph network structure in an embodiment of the present application;

[0031] Figure 4 A schematic diagram of the principle of determining an aggregated feature vector in an embodiment of the present application;

[0032] Figure 5 This is a schematic diagram of a sub-graph structure in an embodiment of the present application;

[0033] Figure 6 This is an example diagram of nodes with different graph structures in the embodiments of the present application;

[0034] Figure 7 This is a schematic diagram of the principle of the network training method in the embodiment of the present application;

[0035] Figure 8 This is a flow chart of a method for representing nodes in a network in an embodiment of the present application;

[0036] Fig. 9 This is a flow chart of the information push method based on the graph network in the embodiment of the present application;

[0037] Fig.10 This is a schematic diagram of the structure of a graph network training device in an embodiment of the present application;

[0038] Fig.11 This is a schematic diagram of the structure of an information push device based on a graph network in an embodiment of the present application;

[0039] Fig.12 Schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0041] To facilitate understanding of the embodiments of the present application, several concepts are briefly introduced below:

[0042] Graph network: Graph network is the most direct tool for describing community relationship chains. It is composed of nodes and edges. Nodes represent relationship chain objects, and the connections between nodes are called edges. Edges represent the degree of connection between two objects. The properties of nodes and edges in graph networks are the same as graph structures. They can be divided into directed graphs and undirected graphs. Directed graphs include recursive neural networks and recurrent neural networks, and undirected graphs include Hopfield networks and Markov networks. Graph networks can be used to process data with graph structures, such as knowledge graphs, social networks, molecular networks, etc.

[0043] Node attributes: In a graph network, each node contains inherent information, which is called node attributes. In the embodiments of the present application, this information is called the attribute information of the node when used. The initial feature vector of the node can be determined based on the attribute information of the node. For example, in a product graph network, the brand and price of the product are all attribute information.

[0044] Graph Embedding Algorithms: It is a process of mapping graph data (usually high-dimensional dense matrices) into low-dimensional dense vectors (i.e., embedding features). Graph embedding needs to capture the topological structure of the graph network, the relationship between nodes, and other information (such as subgraphs, edges, etc.). The purpose of graph embedding algorithms is to learn the structure of the graph or the adjacency relationship between nodes, encode the nodes (or reduce the dimensionality of inherent features), and map all nodes into vectors of equal dimensions so that they can be easily applied to downstream clustering, classification, association analysis or visualization tasks. Therefore, in practical applications, graph embedding is usually a preprocessing task, and most graph embedding algorithms are unsupervised learning algorithms.

[0045] Graph embedding algorithms have gradually transitioned to the era of neural networks, such as SDNE and GraphSAGE. Graph embedding algorithms based on neural networks are no longer limited to the adjacency information of nodes, but begin to incorporate the characteristics of the nodes themselves into the model considerations, and gradually evolve from static transductive learning to dynamic inductive learning, greatly improving both fitting and generalization capabilities.

[0046] Transductive learning: In graph embedding training, the embedding features of each node are learned directly on a fixed graph. When the graph network structure changes and new nodes appear, transductive learning needs to be retrained to obtain new embedding features.

[0047] Inductive learning: Corresponding to direct learning, it generalizes unknown nodes by learning the connections between nodes, from specific to general. When the graph network structure changes and new nodes appear, new embedding features can be obtained without retraining. In the embodiments of the present application, the unsupervised graph embedding algorithm GraphSAGE based on big data inductive learning is mainly improved, and a motif Attention learning module is added to realize the motif feature learning of nodes, and the constraint information of the motif feature is added to the loss function.

[0048] Graph counting: also known as motif counting, refers to a certain connection pattern that frequently appears in a graph network, and counts the number of occurrences of this connection form. After statistics, the structural representation of the graph network is obtained. In the embodiment of the present application, the connection pattern that appears in the graph network is called a subgraph structure, which is represented by a graph structure feature vector (i.e., motif feature).

[0049] Noise Contrastive Estimation (NCE) loss function: It is a commonly used loss function in unsupervised graph embedding algorithms. The calculation formula of the existing NCE loss function is:

[0050]

[0051] Among them, z u ,z v , represents nodes u, v, v n embedding features, σ represents the sigmoid function, v is the positive sample of node u, v n is the negative sample of u, Q is the number of negative samples, P n It is the probability distribution of negative sampling. In the embodiment of the present application, the embedding feature is also called an aggregated feature vector.

[0052] In the related technology, for example, the GraphSAGE algorithm can be used to train the graph network. For a target node in the graph network, the positive sample nodes and the negative sample nodes are determined, and the neighbor nodes of the target node, the positive sample node and the negative sample node are sampled and aggregated respectively to obtain the feature representation of the target node, the positive sample node and the negative sample node. Specifically, for example, taking the target node and sampling the second order as an example, the first-order node is represented by the feature representation of the corresponding second-order neighbor node, and the features of the first-order nodes are merged to form the neighbor features of the target node. The initial features of the target node itself are combined with the neighbor features of the target node to obtain the final feature representation of the target node. Then, the final feature representation of the target node, the positive sample node and the negative sample node obtained can be input into the NCE loss function for gradient back propagation. Although this method can learn some topological information of the graph network by aggregation, it has the following defects in the learning process: 1) The current GraphSAGE cannot recall the corresponding group by specifying a specific graph structure, while special motif structures in graph networks usually represent special groups; 2) During GraphSAGE training, positive and negative sample pairs are generally formed only through connection relationships, which can easily lead to poor accuracy of positive and negative samples and fail to reflect differences. For example, two nodes that are actually not connected may have similar topological structures. Such negative samples are of low rationality and hinder training convergence; 3) GraphSAGE only considers the adjacency relationship between nodes and lacks structural descriptions between small communities of node subgraphs, thereby reducing the accuracy and convergence speed of graph network training, which in turn leads to low accuracy in feature representation of nodes and the accuracy of information push based on feature representation of nodes.

[0053] Therefore, in response to the above problems, an information push method based on a graph network is provided in an embodiment of the present application, which obtains a node to be processed, determines a subgraph structure with the node to be processed as the central node in the graph network, and inputs the subgraph structure into a node feature extraction model to obtain an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model, and then based on the aggregate feature vector of each node in the graph network, selects nodes matching the node to be processed as information push nodes, and pushes the business information corresponding to the information push nodes. In this way, during training, the graph structure feature vector is added to perform motif feature learning. Through learning, the motif features of each node can be obtained and generalized to new nodes, and the specified motif structure can be learned. Learning and recall, and can reduce the repeated calculation of motif features, and adding motif features to the loss function as a constraint during training can reduce the irrationality of negative sample nodes and accelerate the convergence of training. At the same time, the addition of motif features makes it possible to consider both adjacency and community structure relationships on the basis of aggregated feature vectors, making up for the description of the topological structure information of the graph network, thereby improving the accuracy of training and the accuracy of node representation in the graph network. The node feature extraction model obtained after training can obtain the aggregated feature vector representation of the node to be processed by taking the subgraph structure of the node to be processed as input. Since the obtained aggregated feature vector is more accurate, the accuracy of information push is also improved.

[0054] See also Figure 1 As shown, it is a schematic diagram of an application architecture provided in an embodiment of the present application, including a terminal 100 and a server 200.

[0055] The terminal 100 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.

[0056] Various applications can be installed on the terminal 100. Different applications can correspond to different application business scenarios. For example, in a product recommendation application, the application can be a shopping website, and users can search, browse, click, buy, and other operations on the shopping website. The server 200 can obtain these operation information, and then establish a graph network and perform different business processing.

[0057] The server 200 is the background server of the terminal 100 and can provide various network services for the terminal 100. For different applications, the server 200 can be considered as the corresponding background server. For example, the server 200 can obtain the user's click behavior and purchase behavior, etc., and use different commodities as nodes. Different commodities can be connected according to the user's click behavior and purchase behavior, etc., and the connection relationship between different commodities is determined as an edge, thereby forming a graph network. The server 200 can train the graph network. After the training is completed, the aggregated feature vector of the commodity can be determined and compared with the similarity, so as to recommend highly similar commodities to the user. For example, when a user searches for a commodity, or a user clicks to view a commodity, the server 200 can recommend similar commodities to the user through graph network calculation, thereby improving the user's shopping experience.

[0058] Among them, server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.

[0059] The terminal 100 and the server 200 may be connected directly or indirectly via wired or wireless communication, which is not limited in this application. For example Figure 1 In the example, the terminal 100 and the server 200 are connected via the Internet to achieve mutual communication.

[0060] Optionally, the above-mentioned Internet uses standard communication technology, protocol, or a combination of the two. The Internet is usually the Internet, but it can also be any network, including but not limited to any combination of local area network (Local Area Network, LAN), metropolitan area network (Metropolitan Area Network, MAN), wide area network (Wide Area Network, WAN), mobile, wired or wireless network, dedicated network or virtual private network. In some embodiments, the data exchanged through the network is represented by technology and / or format including Hyper Text Mark-up Language (Hyper Text Mark-up Language, HTML), Extensible Markup Language (Extensible Markup Language, XML), etc. In addition, conventional encryption technologies such as Secure Socket Layer (Secure Socket Layer, SSL), Transport Layer Security (Transport Layer Security, TLS), Virtual Private Network (Virtual Private Network, VPN), Internet Protocol Security (Internet Protocol Security, IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technology can also be used to replace or supplement the above-mentioned data communication technology.

[0061] It should be noted that the graph network training method and the information push method based on the graph network in the embodiment of the present application are mainly executed by the server 200. The present embodiment is explained by applying the method to the server 200. For example, the server 200 can obtain relevant behavior data from the terminal 100 to establish a graph network, and perform training based on the graph network training method in the embodiment of the present application to obtain a trained node extraction model. For another example, the server 200 can also calculate the aggregated feature vector of the newly added node based on the trained node extraction model, and then perform subsequent business processing and information push. Of course, the information push method based on the graph network in the embodiment of the present application can also be executed by the terminal 100, which is not limited in the embodiment of the present application.

[0062] In addition, the graph network training in the embodiment of the present application needs to be trained in advance, and its training process is usually performed by the server 200 side due to the performance limitation of the terminal 100.

[0063] It is worth noting that the application architecture diagram in the embodiment of the present application is intended to more clearly illustrate the technical solution in the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. For other application architectures and business applications, the technical solution provided in the embodiment of the present application is also applicable to similar problems. In each embodiment of the present application, the application is applied to Figure 1 The application architecture shown is used as an example for schematic description.

[0064] To facilitate understanding of the information push method based on the graph network in the embodiment of the present application, the training of the node feature extraction model in the embodiment of the present application, that is, the training method of the graph network, is first described below. Based on the above embodiment, refer to Figure 2 As shown, it is a flow chart of the network training method in the embodiment of the present application, and the method specifically includes:

[0065] Step 200: For a target node in a graph network, determine positive sample nodes and negative sample nodes of the target node in the graph network.

[0066] Among them, the correlation between the target node and the positive sample node is greater than the correlation between the target node and the negative sample node.

[0067] In the embodiment of the present application, when training a graph network, a graph embedding algorithm is used, and it is necessary to first determine the positive and negative sample pairs, and a random walk method can be used. Specifically, the execution of step 200 includes:

[0068] S1. Using the random walk method, starting from the target node, determine the correlation between each node in the graph network and the target node.

[0069] The random walk method is a completely random walk. Through the random walk, the correlation between each node in the graph network can be obtained, wherein the target node is the node selected from the graph network. A certain number of nodes can be randomly selected as the target node. This is not limited in the embodiment of the present application. When calculating the correlation between other nodes and a certain target node, it is necessary to start from the certain target node each time. Through multiple walking iterations, the correlation between each other node and the certain target node will tend to a fixed value, that is, the correlation between each other node and the certain target node is obtained. The number of iterations can be pre-set according to actual experience or circumstances.

[0070] For example, see Figure 3As shown, it is an example diagram of a graph network structure in an embodiment of the present application. There are several nodes A, B, C, a, b, c, and d in the graph network. A, B, and C represent users, and a, b, c, and d represent products. The connecting line represents that the user likes the product. Assuming that the correlation between the target node A and other products is calculated, each random walk starts from A, and randomly walks between the ABCabcd nodes through multiple iterations, so that the correlation between other nodes and A can be determined.

[0071] The present application provides a possible implementation method for calculating the correlation degree:

[0072] 1) Set the initial value of the association degree to PR(A)=1, and the others to 0.

[0073] 2) Start walking outward from node A. Assuming the probability of walking out of node A is α, the probability of staying at node A is 1-α. Node a will get a correlation of 1*α*1 / 2 with respect to node A. At this time, because the correlations of other nodes are all 0, the correlation of node a is 1*α*1 / 2. Similarly, the final correlation of node c is also 1*α*1 / 2. At this time, the correlation of node A itself is 1-α. The first iteration ends.

[0074] 3) In the second iteration, in addition to node A, a and c also have correlation. Starting from these nodes, continue to wander, calculate the correlation of other nodes, and repeat the above process. Because each time starts from node A, node A should be added with 1-α at the end.

[0075] 4) When iterates to a certain number of times, the association degree of each node to A will tend to a fixed value, that is, the association degree of each node with the target node A is obtained. Figure 3 A high correlation in the scene indicates that user A prefers the product, and the most relevant products to A can be found and recommended.

[0076] The iteration formula can be summarized as follows:

[0077]

[0078] Among them, u is the target node, that is, user A in the above example, PR(i) is the current relevance of node i, out(i) is the number of nodes connected to node i, and α is the probability of not staying at the current node.

[0079] Of course, other methods may also be used to calculate the association between each node and the target node, which is not limited in the embodiments of the present application.

[0080] S2. Nodes with a correlation degree greater than or equal to the correlation threshold are used as positive sample nodes of the target node, and nodes with a correlation degree less than the correlation threshold are used as negative sample nodes of the target node.

[0081] Among them, the association threshold is not limited in the embodiment of the present application and can be set according to actual conditions. In this way, node pairs with high association can be used as positive sample node pairs, and node pairs with low association can be used as negative sample node pairs. In the graph network, adjacent nodes are usually positive sample node pairs, and those that are not connected or have distant connection relationships are negative sample pairs.

[0082] However, in a graph network, two unconnected nodes may correspond to similar subgraph structures, and the training of negative sample pairs may be inaccurate because the feature vector representation between negative sample pairs with the same or similar subgraph structures should be more similar than that of negative sample pairs with dissimilar subgraph structures. In addition, for positive sample pairs, their second-order or third-order subgraph structures may be different, and only the first-order subgraph structures may be the same. To be more accurate, the subgraph structure differences between the positive sample pairs should also be reflected. However, the related art only relies on connection relationships for training without considering subgraph structure information, and is unable to reflect subgraph structure differences. Therefore, in an embodiment of the present application, subgraph structure information is added during training. By adding the graph structure feature vector of the node, the network structure description is more complete, and the training process is constrained, thereby improving accuracy and convergence speed.

[0083] Step 210: extracting the aggregate feature vector and graph structure feature vector of the target node and the associated positive sample nodes and negative sample nodes respectively, wherein the graph structure feature vector represents the features of the subgraph structure containing a preset order.

[0084] As discussed above, during training in the embodiment of the present application, motif features are added on the basis of the aggregated feature vector, that is, the graph structure feature vector is used as a label guide and constraint. When step 210 is specifically performed, it can be divided into the following two aspects:

[0085] The first aspect is to extract the aggregate feature vectors of the target node and the associated positive sample nodes and negative sample nodes respectively.

[0086] In the embodiment of the present application, the existing GraphSAGE algorithm may be used to determine the aggregated feature vector. Of course, other algorithms, such as the PinSAGE algorithm, may also be used without limitation.

[0087] Among them, for the target node, or the associated positive sample node and negative sample node, the aggregated feature vector is determined in the following manner:

[0088] S1. Obtain neighbor nodes of a first node to be processed by random sampling, wherein the first node to be processed is any one of the following: a target node, a positive sample node, or a negative sample node.

[0089] Among them, when obtaining the neighbor nodes of the first node to be processed, the neighbor nodes of the first node to be processed are randomly sampled, and the number of neighbor nodes to be obtained can be set to reduce the complexity of the calculation. Of course, all neighbor nodes can also be obtained, which is not limited in the embodiment of the present application.

[0090] Moreover, the neighbor nodes are not limited to the first order, and the order can be set without restriction, for example, see Figure 4 As shown, it is a schematic diagram of the principle of determining the aggregated feature vector in an embodiment of the present application, such as Figure 4 As shown in (1), the second-order neighbor nodes of the first node A to be processed are randomly sampled. Figure 4 The arrow direction in (1) indicates the randomly sampled neighbor nodes, and it can be seen that the sampled neighbor nodes of the first node to be processed are obtained.

[0091] S2. According to the attribute information of the first node to be processed and the obtained neighboring nodes, determine the initial feature vectors of the first node to be processed and the obtained neighboring nodes respectively.

[0092] Among them, the attribute information is the inherent information of the first node to be processed. Based on different business scenarios and the different objects corresponding to the first node to be processed, different attribute information corresponds accordingly. For example, in a product map network, the attribute information includes price, brand, color, style, etc.

[0093] In the embodiment of the present application, the initial feature vectors of the first node to be processed and the neighboring nodes themselves may be determined based on the attribute information.

[0094] S3. Aggregate the obtained initial feature vectors of the neighboring nodes.

[0095] S4. Concatenate the initial feature vectors of the aggregated neighbor nodes and the initial feature vector of the first node to be processed to obtain an aggregated feature vector of the first node to be processed.

[0096] Specifically, during aggregation, aggregation can be performed one order at a time. For example, the initial feature vectors of the k-th order neighbor nodes corresponding to the k-1-th order neighbor nodes from the first node to be processed are first aggregated, and the aggregated initial feature vectors are concatenated with the corresponding k-1-th order neighbor nodes to obtain the final aggregated feature vector of the k-1-th order neighbor nodes. The k-1-th order neighbor nodes are aggregated again until the final aggregated feature vector of the first node to be processed is obtained.

[0097] Among them, when aggregating, average aggregation, long short-term memory (LSTM) network aggregation, pooling aggregation, etc. can be used, which is not limited in the embodiments of the present application.

[0098] For example, the average aggregation algorithm averages each dimension of the initial feature vector of the k-th order neighbor node, and concatenates it with the corresponding k-1-th order neighbor node before performing nonlinear transformation. It is also possible to directly average each dimension of the initial feature vector of the first node to be processed and all the obtained neighbor nodes before performing nonlinear transformation.

[0099] For another example, the LSTM aggregation algorithm, because LSTM does not conform to the property of "sort invariant", it is necessary to first randomly sort the neighbor nodes, and then use the initial feature vector of the randomly sorted neighbor node sequence as the LSTM input.

[0100] For example, the Pooling aggregation algorithm first performs a nonlinear transformation on the initial feature vector of the k-1 layer of each neighbor node, and then applies the maximum pooling operation or the average pooling operation by dimension to capture the outstanding or comprehensive performance of the neighbor node set in some aspect, thereby obtaining the aggregated feature vector of the first node to be processed.

[0101] For example, Figure 4 As shown in (2), taking the average aggregation algorithm as an example, Figure 4 The arrow direction in (2) indicates the aggregation direction, that is, the neighbor nodes sampled from the second order are aggregated and concatenated with the corresponding first order neighbor nodes to obtain the aggregate feature vector of the first order neighbor nodes, and then the first order neighbor nodes are aggregated and concatenated with the initial feature vector of the first node to be processed, as shown in Figure 4 As shown in Figure (3), the aggregated feature vector of the first node to be processed can be obtained.

[0102] The second aspect: extract the graph structure feature vectors of the target node and the associated positive sample nodes and negative sample nodes respectively.

[0103] In the embodiments of the present application, the graph structure feature vector may also be referred to as a motif feature. The motif feature is a means for statistically analyzing graph structure information, which represents different connection structure forms in the graph network. There are multiple motif structure forms, which can be defined according to different business scenarios. Therefore, the motif structure form usually carries rich business prior information.

[0104] Among them, for the target node, or the associated positive sample node and negative sample node, the graph structure feature vector is determined in the following way:

[0105] S1. Determine a subgraph structure of a preset order with a second node to be processed as a central node, wherein the second node to be processed is any one of the following: a target node, a positive sample node, or a negative sample node.

[0106] The order of the subgraph structure corresponding to the second node to be processed should be the same as the order of the preset subgraph structure, and the preferred preset order is greater than or equal to 2.

[0107] In an embodiment of the present application, during training, the subgraph structure of each second node to be processed can be counted, the neighboring nodes of the second node to be processed, for example, the second order, are selected to form a subgraph, and the motif count feature of the subgraph is calculated.

[0108] S2. Based on each preset subgraph structure, count the number of occurrences of each preset subgraph structure in the subgraph structure determined corresponding to the second node to be processed.

[0109] S3. Determine the graph structure feature vector of the second node to be processed according to the statistical number of occurrences.

[0110] For example, see Figure 5 As shown, it is a schematic diagram of a sub-graph structure in an embodiment of the present application. Figure 5 As shown in the figure, 16 second-order subgraph structures are listed. Of course, different subgraph structures can be customized according to business scenarios without restriction. For example, in the financial graph network, Figure 5 The fourth subgraph structure can be used to represent the state of fund distribution. Assuming that the 16 second-order subgraph structures are preset subgraph structures, a vector with a dimension of 16 can be formed. The number of occurrences of the 16 subgraph structures in the subgraph structure corresponding to the second node to be processed is counted respectively. The number of occurrences of the 16 subgraph structures counted is the element in the 16-dimensional vector. Specifically, for Figure 5 The first subgraph structure is counted, and the number of occurrences of the subgraph structure is taken as the first element in the vector. The number of occurrences of the second subgraph structure is counted, and the number of occurrences of the second subgraph structure is taken as the second element in the vector. 16 occurrences can be obtained in sequence to form a 16-dimensional vector, which is the graph structure feature vector of the second node to be processed.

[0111] Step 220: Based on the target node, and the aggregate feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes, a node feature extraction model is trained to extract the aggregate feature vectors of the nodes in the graph network.

[0112] In the embodiment of the present application, when training the graph network, the graph structure information is added, but in practice, it is difficult to intuitively obtain the corresponding graph structure information in the graph embedding algorithm. For example, it is difficult to directly observe the motif structure form from the embedding feature of the node. Without the intuitive motif structure form, it is difficult to complete the community screening of such forms in the downstream business. In addition, for the graph network with a large amount of data, the calculation amount and resource consumption of the motif features of the entire graph network are huge, and it is difficult to implement. Therefore, in the embodiment of the present application, in order to solve this problem, on the one hand, the graph embedding algorithm in the related technology is improved, and the motif features are added during the training process. On the other hand, in order to reduce the amount of calculation of the motif features, a motif attention module is added, and the real motif features are obtained only by calculating the number of occurrences during training, and the motif attention module is trained based on this, so that after the training is completed, the motif features of the node can be directly predicted and obtained without repeating the calculation by counting the number of occurrences, which can effectively reduce the amount of calculation. In addition, in the embodiment of the present application, when adding motif features for training, the loss function can also be constrained to improve accuracy.

[0113] Specifically, executing step 220 includes: training a node feature extraction model according to the target node, and the aggregated feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes, so that the target loss function of the node feature extraction model converges.

[0114] Therefore, in the embodiment of the present application, when training the graph network, it can be understood as a multi-task training process. On the one hand, the training makes the representation of the node's graph structure feature vector more accurate, and on the other hand, the training makes the representation of the node's aggregate feature vector more accurate. The target loss function includes a first loss function and a second loss function. The first loss function is the loss function between the graph structure feature vector predicted by the node feature extraction model and the determined graph structure feature vector. The second loss function makes the graph structure feature similarity and aggregate feature similarity of the target node and the associated positive sample nodes and negative sample nodes change in direct proportion to the similarity.

[0115] Specifically, the determination methods of the first loss function and the second loss function are described below respectively:

[0116] Part 1: The first loss function.

[0117] In the embodiment of the present application, a possible implementation method is provided for determining the first loss function:

[0118] 1) Normalize the graph structure feature vector of the determined target node.

[0119] That is, the graph structure feature vector of the target node determined by counting the number of occurrences can be normalized by the softmax function, and the elements in the graph structure feature vector can be normalized to between 0 and 1 to obtain a pseudo-label with motif attention information. The motif attention module can be used to learn the subgraph structure with the largest number of occurrences corresponding to the target node. At this time, the graph structure feature vector determined based on the number of occurrences can be understood as the true value.

[0120] 2) The initial feature vector of the target node is passed through a fully connected layer to obtain a graph structure feature vector of the target node prediction, wherein the initial feature vector of the target node is determined according to the attribute information of the target node.

[0121] That is, the graph structure feature vector obtained by the fully connected layer is the predicted value. The first loss function is determined by the difference between the predicted value and the true value. Continuous training can make the predicted value closer to the true value and improve the accuracy of the predicted graph structure feature vector.

[0122] 3) Determine a first loss function based on the normalized graph structure feature vector of the target node and the predicted graph structure feature vector.

[0123] For example, the first loss function is: L motif =|f s (V M )-V′ M |.

[0124] Among them, f s is the softmax normalization formula, V M is the graph structure feature vector determined based on the number of occurrences, V′ M is the graph structure feature vector predicted by the fully connected layer.

[0125] In the embodiment of the present application, the first loss function can be determined only based on the target node, without having to be determined based on all associated positive sample nodes and negative sample nodes, which can reduce complexity.

[0126] Part 2: Second loss function.

[0127] In the embodiment of the present application, a possible implementation method is provided for determining the second loss function:

[0128] 1) According to the target node and the graph structure feature vectors of the associated positive sample nodes, first graph structure feature similarity weight parameters of the target node and the associated positive sample nodes are determined respectively.

[0129] 2) According to the graph structure feature vectors of the target node and the associated negative sample node, second graph structure feature similarity weight parameters of the target node and the associated negative sample node are determined respectively.

[0130] 3) Determine a second loss function according to the first graph structure feature similarity weight parameter, the second graph structure feature similarity weight parameter, the target node, and the aggregated feature vectors of the associated positive sample nodes and negative sample nodes.

[0131] During the training process, it is usually to make the embedding features of adjacent nodes more similar, and the features of unconnected or distantly connected nodes more different. Therefore, this type of training considers more the closeness of the relationship between the two nodes, while in real graph networks, it is often necessary to consider the similarity of the community topological structures of the subgraphs where the two nodes are located. For example, as two nodes in a positive sample pair, if the community topological structures of the subgraphs they are in are similar, the similarity of the positive sample pair should be higher than that of the positive sample pair with dissimilar community topological structures; and the negative sample pairs with similar community topological structures should weaken their dissimilarity to each other. Therefore, in an embodiment of the present application, in order to achieve this purpose, in an embodiment of the present application, the NCE loss function in the related technology is improved, and graph structure information is added as a constraint condition to obtain a global NCE loss function based on motif feature information, which can be applied to any unsupervised graph embedding algorithm.

[0132] For example, the target node u, v is the positive sample node of u, v n is the negative sample node of u, and the graph structure feature vector of u is V Mu , the graph structure feature vector of v is V Mv ,v n The graph structure feature vector of z u ,z v , represents nodes u, v, v n The aggregate feature vector of .

[0133] The first graph structural feature similarity weight parameter is:

[0134]

[0135] The similarity weight parameter of the second graph structure feature is:

[0136]

[0137] Then the second loss function determined according to the first graph structure feature similarity weight parameter, the second graph structure feature similarity weight parameter, and the aggregated feature vector is:

[0138]

[0139] Among them, σ represents the sigmoid function, Q is the number of negative samples, P n is the probability distribution of negative sampling, and γ is a hyperparameter used to adjust the weight size.

[0140] In this way, by adding graph structure information, the similarity between positive sample nodes, negative sample nodes and target nodes can be adjusted and constrained, so that the aggregate feature vector is affected by the supervised gradient of the motif feature. For example, if the subgraph topological structures of nodes u and v are similar, the corresponding motif features are similar, ω motif (z u ,z v ) is smaller, the corresponding weight is larger. At this time, the u and v nodes are positive sample pairs, and the dot product result of their embedding features should be smaller, then the corresponding embedding features are more similar, and u and v n For negative sample pairs, if the subgraph topological structures are similar, the dissimilarity of the corresponding embedding features can be weakened through the similarity weight parameter of the second graph structure feature.

[0141] In an embodiment of the present application, the positive sample nodes and negative sample nodes of the target node in the graph network are respectively determined, and the aggregated feature vectors and graph structure feature vectors of the target node and the associated positive sample nodes and negative sample nodes are respectively determined, and then the node feature extraction model can be trained based on the aggregated feature vector and the graph structure feature vector. In this way, when training the graph network, the graph structure feature vector is added on the basis of the aggregated feature vector, that is, the graph structure information is added, the representation of the network structure is more complete, and the impact of unreasonable positive and negative samples on the training can be reduced, and the convergence speed and accuracy of the training are improved. In addition, through training, the complex operation of motif features can be converted into simple predictions, that is, the graph structure feature vector of the node can be obtained by prediction through the node feature extraction model, which reduces the amount of calculation, especially for large data graph networks. The effect is better, and the motif directional characteristics can be intuitively given, which is convenient for subsequent downstream business applications.

[0142] Based on the above embodiment, the calculation of the aggregate feature vector by using the GraphSAGE algorithm is taken as an example for explanation. In the GraphSAGE algorithm commonly used in large data graph networks in the related art, in order to adapt to the calculation of large networks, the community structure information is omitted, resulting in the loss of structural information of the graph network. GraphSAGE obtains part of the topological structure pattern of the subgraph by aggregating the embedding features of the first-order and second-order neighbor nodes. However, GraphSAGE only considers the connection status between the central node and the neighbor nodes on the subgraph (that is, whether the central node is connected to the neighbor nodes), and does not consider the relationship between the last-order neighbor nodes and the neighbor nodes. For example, Figure 6As shown, it is an example diagram of nodes of different graph structures in the embodiment of the present application, such as Figure 6 As shown, in the subgraph containing second-order neighbor nodes, Figure 6 The graph structures of the left and right images are completely different. However, in the existing GraphSAGE algorithm, if the attribute information of each node is similar, the embedding features of nodes A and B will also be similar, which cannot reflect the difference in graph structure.

[0143] Therefore, in the embodiment of the present application, the GraphSAGE algorithm in the related art can be improved, and motif features can be added as pseudo-label guidance during the training process to improve the accuracy of node feature vector representation and also improve the training convergence speed.

[0144] Specific reference Figure 7 As shown, it is a schematic diagram of the principle of the network training method in the embodiment of the present application. The target node is A. In order to simplify and facilitate the explanation of the method in the embodiment of the present application, the neighbor nodes of A include nodes B and C as an example. In addition, for the positive sample nodes and negative sample nodes associated with the target node A, the principles are the same during calculation. Here, node D is used to represent the positive sample node or the negative sample node for explanation. One iterative training needs to be trained based on the target node A, the positive sample node and the negative sample node D. Here, the positive sample node and the negative sample node are unified into node D as an example for explanation.

[0145] 1) If Figure 7 In the GraphSAGE module, the initial feature vectors of the target node A and neighbor nodes B and C are determined. The initial feature vectors of A, B and C are determined according to the attribute information of A, B and C respectively. The initial feature vectors of B and C are aggregated, called the neighbor feature vector of the target node A, and concatenated with the initial feature vector of the target node A to obtain the aggregated feature vector of the target node A.

[0146] like Figure 7 The motif attention module in , for the target node A, determines the graph structure feature vector of A, namely the motif feature, by counting the number of occurrences of each preset subgraph structure in the subgraph structure corresponding to A.

[0147] The motif features of the target node A are normalized by the softmax function to obtain a pseudo-label with motifattention information. The initial feature vector of the target node A obtained by GraphSAGE is passed through a fully connected layer to obtain a predicted graph structure feature vector. The first loss function is determined based on the normalized graph structure feature vector of the target node A and the predicted graph structure feature vector. By learning the motif attention mechanism, the graph structure information with the most occurrences corresponding to the target node A can be learned, thereby improving the accuracy of the graph structure feature vector predicted by the fully connected layer.

[0148] 2) Node D can be a positive sample node or a negative sample node of the target node A. Similarly, the aggregate feature vector of node D can be determined by the GraphSAGE module, and the graph structure feature vector of node D can be determined by the motif attention module. However, at this time, node D does not need to learn the attention mechanism, and does not need to be used for the first loss function.

[0149] 3) Then, according to the graph structure feature vectors of the target node A and the node D (including the associated positive sample nodes and the negative sample nodes), the first graph structure feature similarity weight parameter and the second graph structure feature similarity weight parameter can be determined respectively, and according to the first graph structure feature similarity weight parameter, the second graph structure feature similarity weight parameter, and the aggregated feature vector of the target node A and the node D, the second loss function is determined, and the motif feature constraint is added to the second loss function. Through the training and learning of the second loss function, the accuracy of the vector representation of the node is improved, that is, the accuracy of the aggregated feature vector is obtained.

[0150] In an embodiment of the present application, GraphSAGE is improved by adding an end-to-end motifattention module to the algorithm. The motif feature can be calculated for each subgraph through the motif attention module. The motif feature representation of each node is obtained through learning and can be generalized to new nodes, thereby reducing the repeated calculation of motif features. Motif feature constraints are added when calculating the NCE loss function to weaken the discrimination of negative sample pairs with similar motif features, reduce the confrontation between unconnected node pairs with the same graph structure, and accelerate the convergence of graph network training. GraphSAGE lacks a description of the graph topology, and the motif feature can be used to make up for the missing graph structure information of GraphSAGE, thereby improving the accuracy and reliability of graph network training.

[0151] The graph network training method in the embodiment of the present application can be applied to any application related to the relationship chain, for example, business scenarios such as product recommendation and financial relationship chain. After the graph network training is completed, the feature vector representation of the nodes in the graph network can be determined based on the node feature extraction model obtained through training, and then the feature vector representation of the nodes can be applied to subsequent downstream tasks in the business scenario for subsequent processing.

[0152] Based on the above embodiments, see Figure 8 As shown, it is a flow chart of a method for representing nodes in a network in the figure in an embodiment of the present application, and the method specifically includes:

[0153] Step 800: Obtain the node to be processed.

[0154] Step 810: Determine a subgraph structure with the node to be processed as a central node in the graph network, wherein the subgraph structure includes nodes and connection relationships connected to the node to be processed.

[0155] The preset order of the subgraph structure is preferably the same as the order of the subgraph structure obtained during training. A subgraph structure may generally include each node in the subgraph structure and the connection relationship between each node.

[0156] Step 820: Input the subgraph structure into the node feature extraction model to obtain an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model.

[0157] Furthermore, the attribute information of the node to be processed and the nodes included in the subgraph structure may be simultaneously used as input to obtain the aggregated feature vector of the node to be processed.

[0158] Among them, the node feature extraction model is obtained by training based on the graph network training method in the above-mentioned embodiment of the present application. Specifically, the node feature extraction model is trained according to the determined target node, and the aggregated feature vectors and graph structure feature vectors of the positive sample nodes and the negative sample nodes, so as to converge the target loss function of the node feature extraction model. Among them, the target loss function includes a first loss function and a second loss function. The first loss function is the loss function between the graph structure feature vector predicted by the node feature extraction model and the determined graph structure feature vector. The second loss function makes the graph structure feature similarity of the target node and the associated positive sample nodes and negative sample nodes change in direct proportion to the aggregate feature similarity.

[0159] In the embodiment of the present application, in actual application, it is only necessary to input the subgraph of the node to be processed and the attribute information of the nodes contained in the subgraph, and then the aggregated feature vector of the node to be processed can be obtained based on the trained node feature extraction model. Furthermore, the graph structure feature vector of the node to be processed can also be obtained.

[0160] Specifically, based on the graph network training method in the above-mentioned embodiment, the subgraph structure and attribute information of the node to be processed are input, and the attribute information of the node to be processed and other nodes included in the subgraph structure can be used to respectively determine the initial feature vectors of the node to be processed and other nodes included in the subgraph structure, and based on the aggregation function (also called aggregation algorithm) in the trained node feature extraction model, the aggregated feature vector of the node to be processed is obtained by aggregation. At the same time, the graph structure feature vector of the node to be processed can also be obtained based on the fully connected layer in the trained node feature extraction model.

[0161] In this way, in the embodiment of the present application, based on the node feature extraction model obtained by training the graph network training method in the embodiment of the present application, when a new node arrives, it is only necessary to perform forward calculation to obtain the motif features of the subgraph where the node is located, without the need to repeat the calculation by counting the number of occurrences. The predicted motif features and aggregated feature vectors are combined to improve the description of the topological structure of the graph network, and the obtained aggregated feature vectors and graph structure feature vectors are more accurate, thereby improving the accuracy of other business functions.

[0162] Based on the above embodiments, the present application also provides an example of a specific business application scenario. The graph network training method and the node representation method in the graph network in the present application can be applied to all applications related to relationship chains. For example, taking the application in the commodity recommendation scenario as an example, please refer to Fig. 9 As shown, it is a flow chart of the information push method based on the graph network in an embodiment of the present application, and the specific method includes:

[0163] Step 900: Obtain the node to be processed.

[0164] Among them, the node to be processed is the node in the graph network, which can be a newly added node or any other node. The nodes in the graph network are associated with their business application scenarios. For example, in the product recommendation scenario, each node in the graph network is a product, and the products are connected through certain association relationships to form a graph network.

[0165] Step 910: Determine a subgraph structure with the node to be processed as a central node in the graph network, wherein the subgraph structure includes nodes and connection relationships connected to the node to be processed.

[0166] Step 920: Input the subgraph structure into the node feature extraction model to obtain an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model.

[0167] The aggregated feature vector represents the node to be processed in the graph network, that is, the vector representation of the node to be processed.

[0168] Step 930: Based on the aggregated feature vectors of each node in the graph network, select nodes that match the node to be processed as information push nodes.

[0169] Furthermore, through the node representation method in the graph network of the above embodiment, the aggregated feature vectors of other nodes in the graph network can also be obtained, which will not be described in detail here.

[0170] Specifically, the node matching the node to be processed is selected as the information push node in step 930. A possible implementation method is provided in the embodiment of the present application:

[0171] 1) Determine the similarity between the node to be processed and each node in the graph network based on the aggregated feature vector of the node to be processed and the aggregated feature vector of each node in the graph network.

[0172] For example, the cosine values ​​between the node to be processed and the aggregated feature vector and the aggregated feature vectors of other nodes may be calculated respectively, and the cosine similarity may be used to measure the similarity between two nodes.

[0173] 2) Determine the information push node based on the similarity.

[0174] There are two ways:

[0175] The first method is to use nodes whose similarity with the node to be processed is greater than a threshold as information push nodes.

[0176] The second method is to use the node with the greatest similarity to the node to be processed as the information push node.

[0177] Step 940: Push the business information corresponding to the information push node.

[0178] For example, if the business information is a commodity, the target commodity corresponding to the node to be processed can be pushed. Commodities with greater similarity can improve the accuracy of the push, thereby improving the user's shopping experience and helping merchants increase the user's purchase rate.

[0179] Of course, the embodiments of the present application are not limited to this business scenario, but can also be applied to other business scenarios related to relationship chains. The embodiments of the present application are not limited thereto. It should be noted that as long as they are based on the graph network training method in the embodiments of the present application and the node representation method in the graph network, they should fall within the scope of protection of the present application.

[0180] Based on the same inventive concept, the present application also provides a graph network training device in an embodiment. The graph network training device may be, for example, the server in the aforementioned embodiment. The graph network training device may be a hardware structure, a software module, or a hardware structure plus a software module. Fig.10As shown, a graph network training device in an embodiment of the present application specifically includes:

[0181] A first determination module 1000 is used to determine, for a target node in a graph network, a positive sample node and a negative sample node of the target node in the graph network, wherein a correlation degree between the target node and the positive sample node is greater than a correlation degree between the target node and the negative sample node;

[0182] The second determination module 1010 is used to extract the aggregate feature vector and graph structure feature vector of the target node and the associated positive sample nodes and negative sample nodes, respectively, wherein the graph structure feature vector represents the features of the subgraph structure containing a preset order;

[0183] The training module 1020 is used to train a node feature extraction model based on the target node and the aggregate feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes. The node feature extraction model is used to extract the aggregate feature vectors of the nodes in the graph network.

[0184] Optionally, when determining the positive sample nodes and negative sample nodes associated with the target node in the graph network, the first determination module 1000 is specifically used to:

[0185] Using random walk method, starting from the target node, determine the correlation between each node in the graph network and the target node;

[0186] The nodes whose association degree is greater than or equal to the association threshold are regarded as the positive sample nodes of the target node, and the nodes whose association degree is less than the association threshold are regarded as the negative sample nodes of the target node.

[0187] Optionally, when respectively extracting the aggregate feature vectors of the target node and the associated positive sample nodes and negative sample nodes, the second determination module 1010 is used to:

[0188] For the target node, or the associated positive sample node and negative sample node, the aggregate feature vector is determined in the following way:

[0189] A random sampling method is adopted to obtain neighbor nodes of a first node to be processed, wherein the first node to be processed is any one of the following: a target node, a positive sample node, or a negative sample node;

[0190] Determining initial feature vectors of the first node to be processed and the obtained neighboring nodes respectively according to the attribute information of the first node to be processed and the obtained neighboring nodes;

[0191] Aggregate the obtained initial feature vectors of neighbor nodes;

[0192] The initial feature vectors of the aggregated neighbor nodes and the initial feature vector of the first node to be processed are concatenated to obtain an aggregated feature vector of the first node to be processed.

[0193] Optionally, when respectively extracting the graph structure feature vectors of the target node and the associated positive sample nodes and negative sample nodes, the second determination module 1010 is used to:

[0194] For the target node, or the associated positive sample node and negative sample node, the graph structure feature vector is determined in the following way:

[0195] Determine a subgraph structure of a preset order with a second node to be processed as a central node, wherein the second node to be processed is any one of the following: a target node, a positive sample node, or a negative sample node;

[0196] Based on each preset subgraph structure, respectively counting the number of occurrences of each preset subgraph structure in the subgraph structure determined corresponding to the second node to be processed;

[0197] According to the statistical number of occurrences, a graph structure feature vector of the second node to be processed is determined.

[0198] Optionally, when a node feature extraction model is obtained by training according to the target node and the aggregate feature vector and graph structure feature vector of the associated positive sample node and negative sample node, the training module 1020 is specifically used for:

[0199] According to the target node, and the aggregated feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes, a node feature extraction model is trained to converge a target loss function of the node feature extraction model, wherein the target loss function includes a first loss function and a second loss function, the first loss function being a loss function between a graph structure feature vector predicted by the node feature extraction model and a determined graph structure feature vector, and the second loss function causing the graph structure feature similarity and the aggregated feature similarity of the target node and the associated positive sample nodes and negative sample nodes to change in direct proportion to the similarity.

[0200] Optionally, for the method of determining the first loss function, a first calculation module 1030 is further included, which is used to:

[0201] Normalize the graph structure feature vector of the determined target node;

[0202] The initial feature vector of the target node is passed through a fully connected layer to obtain a graph structure feature vector of the target node prediction, wherein the initial feature vector of the target node is determined according to the attribute information of the target node;

[0203] A first loss function is determined according to the normalized graph structure feature vector of the target node and the predicted graph structure feature vector.

[0204] Optionally, for the method of determining the second loss function, a second calculation module 1040 is further included, which is used to:

[0205] According to the target node and the graph structure feature vectors of the associated positive sample nodes, respectively determine the first graph structure feature similarity weight parameters of the target node and the associated positive sample nodes;

[0206] According to the target node and the graph structure feature vectors of the associated negative sample node, respectively determine the second graph structure feature similarity weight parameters of the target node and the associated negative sample node;

[0207] A second loss function is determined according to the first graph structure feature similarity weight parameter, the second graph structure feature similarity weight parameter, the target node, and the aggregated feature vectors of the associated positive sample nodes and negative sample nodes.

[0208] In an embodiment of the present application, when training a graph network, the target node and the aggregated feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes are respectively determined, and based on the target node and the aggregated feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes, the graph network is trained so that the target loss function of the graph network converges. During training, graph structure information is added as a constraint condition, which can reduce the irrationality of positive and negative samples, accelerate the convergence speed, and improve the accuracy of training.

[0209] Based on the same inventive concept, the embodiment of the present application also provides an information push device based on a graph network, which can be, for example, the server in the aforementioned embodiment, and can be a hardware structure, a software module, or a hardware structure plus a software module. Fig.11 As shown, in an embodiment of the present application, an information push device based on a graph network specifically includes:

[0210] An acquisition module 1100 is used to acquire a node to be processed;

[0211] A determination module 1110 is used to determine a subgraph structure with the node to be processed as a central node in the graph network, wherein the subgraph structure includes nodes and connection relationships connected to the node to be processed;

[0212] Extraction module 1120, used to input the subgraph structure into the node feature extraction model, and obtain the aggregate feature vector corresponding to the node to be processed output by the node feature extraction model; the aggregate feature vector represents the node to be processed in the graph network;

[0213] A screening module 1130 is used to screen nodes matching the nodes to be processed as information push nodes based on the aggregated feature vectors of each node in the graph network;

[0214] The push module 1140 is used to push the business information corresponding to the information push node.

[0215] In this way, since the aggregated feature vector representation of the node is more accurate, the accuracy of information push can also be improved in business scenarios.

[0216] Based on the above embodiments, see Fig.12 Shown is a schematic diagram of the structure of an electronic device in an embodiment of the present application.

[0217] An embodiment of the present application provides an electronic device, which may be the server in the aforementioned embodiment, and may include a processor 1210 (Center Processing Unit, CPU), a memory 1220, an input device 1230, an output device 1240, and the like.

[0218] The memory 1220 may include a read-only memory (ROM) and a random access memory (RAM), and provides the program instructions and data stored in the memory 1220 to the processor 1210. In an embodiment of the present application, the memory 1220 may be used to store a program of any graph network training, a node representation method in a graph network, or an information push method based on a graph network in an embodiment of the present application.

[0219] The processor 1210 calls the program instructions stored in the memory 1220, and the processor 1210 is used to execute any graph network training, node representation method in the graph network, or information push method based on the graph network in the embodiments of the present application according to the obtained program instructions.

[0220] Based on the above embodiments, in an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the graph network training, the node representation method in the graph network, or the information push method based on the graph network in any of the above method embodiments is implemented.

[0221] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0222] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0223] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0224] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0225] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0226] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the embodiments of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for pushing information based on a graph network, characterized in that: include: Get the nodes to be processed; Determine a subgraph structure with the node to be processed as a central node in a graph network, wherein the subgraph structure includes nodes and connection relationships connected to the node to be processed, wherein the graph network is composed of different business information as nodes, and connection relationships between different business information as edges, and the connection relationships between different business information are determined according to user operation behaviors on the business information; The subgraph structure is input into a node feature extraction model to obtain an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model; the aggregate feature vector represents the node to be processed in the graph network; wherein the node feature extraction model is trained based on the aggregate feature vectors and graph structure feature vectors of the target node in the graph network and the associated positive sample nodes and negative sample nodes, the association degree between the target node and the positive sample node is greater than the association degree between the target node and the negative sample node, and the graph structure feature vector represents the characteristics of the subgraph structure containing a preset order; Based on the aggregated feature vectors of each node in the graph network, selecting nodes matching the node to be processed as information pushing nodes; The business information corresponding to the information pushing node is pushed.

2. The method according to claim 1, characterized in that The training method of the node feature extraction model is: For a target node in a graph network, determining a positive sample node and a negative sample node of the target node in the graph network; Respectively extracting the target node, and the aggregate feature vector and graph structure feature vector of the associated positive sample node and negative sample node; A node feature extraction model is trained based on the target node, and the aggregate feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes. The node feature extraction model is used to extract the aggregate feature vectors of the nodes in the graph network.

3. The method according to claim 2, characterized in that Determining the positive sample nodes and the negative sample nodes of the target node in the graph network specifically includes: Using a random walk method, taking the target node as a starting point, determining the correlation between each node in the graph network and the target node; Nodes with a correlation degree greater than or equal to a correlation threshold are used as positive sample nodes of the target node, and nodes with a correlation degree less than the correlation threshold are used as negative sample nodes of the target node.

4. The method according to claim 2, characterized in that Extracting the target node and the aggregated feature vectors of the associated positive sample nodes and negative sample nodes respectively, including: For the target node, or the associated positive sample node and negative sample node, the aggregated feature vector is determined in the following manner: A random sampling method is adopted to obtain neighbor nodes of a first node to be processed, wherein the first node to be processed is any one of the following: the target node, a positive sample node, or a negative sample node; Determining initial feature vectors of the first node to be processed and the acquired neighboring node respectively according to the attribute information of the first node to be processed and the acquired neighboring node; Aggregating the obtained initial feature vectors of the neighboring nodes; The initial feature vectors of the aggregated neighbor nodes and the initial feature vector of the first node to be processed are concatenated to obtain an aggregated feature vector of the first node to be processed.

5. The method according to any one of claims 2 to 4, characterized in that: Extracting graph structure feature vectors of the target node and associated positive sample nodes and negative sample nodes respectively, including: For the target node, or the associated positive sample node and negative sample node, the graph structure feature vector is determined in the following manner: Determine a subgraph structure of a preset order with a second node to be processed as a central node, wherein the second node to be processed is any one of the following: the target node, a positive sample node, or a negative sample node; Based on each preset subgraph structure, respectively counting the number of occurrences of each preset subgraph structure in the subgraph structure determined corresponding to the second node to be processed; According to the counted number of occurrences, a graph structure feature vector of the second node to be processed is determined.

6. The method according to claim 5, characterized in that According to the target node, and the aggregated feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes, a node feature extraction model is trained, specifically including: According to the target node, and the aggregated feature vectors and graph structure feature vectors of the associated positive sample nodes and negative sample nodes, a node feature extraction model is trained to converge a target loss function of the node feature extraction model, wherein the target loss function includes a first loss function and a second loss function, the first loss function being a loss function between a graph structure feature vector predicted by the node feature extraction model and a determined graph structure feature vector, and the second loss function causing the graph structure feature similarity and the aggregated feature similarity of the target node and the associated positive sample nodes and negative sample nodes to change in direct proportion to the similarity.

7. The method according to claim 6, characterized in that The first loss function is determined as follows: Normalizing the determined graph structure feature vector of the target node; Passing the initial feature vector of the target node through a fully connected layer to obtain a graph structure feature vector predicted by the target node, wherein the initial feature vector of the target node is determined according to attribute information of the target node; The first loss function is determined according to the normalized graph structure feature vector of the target node and the predicted graph structure feature vector.

8. The method according to claim 6, characterized in that The second loss function is determined as follows: Determining first graph structure feature similarity weight parameters of the target node and the associated positive sample node respectively according to the graph structure feature vectors of the target node and the associated positive sample node; Determining second graph structure feature similarity weight parameters of the target node and the associated negative sample node respectively according to the graph structure feature vectors of the target node and the associated negative sample node; The second loss function is determined according to the first graph structure feature similarity weight parameter, the second graph structure feature similarity weight parameter, and the aggregated feature vectors of the target node, the associated positive sample nodes, and the negative sample nodes.

9. An information push device based on a graph network, characterized in that: include: An acquisition module, used to acquire nodes to be processed; A determination module, used to determine a subgraph structure with the node to be processed as a central node in a graph network, wherein the subgraph structure includes nodes and connection relationships connected to the node to be processed, wherein the graph network is composed of different business information as nodes, and the connection relationships between different business information as edges, and the connection relationships between different business information are determined according to the user's operation behavior on the business information; An extraction module is used to input the subgraph structure into a node feature extraction model to obtain an aggregate feature vector corresponding to the node to be processed output by the node feature extraction model; the aggregate feature vector represents the node to be processed in the graph network; wherein the node feature extraction model is trained based on the aggregate feature vectors and graph structure feature vectors of the target node in the graph network and the associated positive sample nodes and negative sample nodes, the association degree between the target node and the positive sample node is greater than the association degree between the target node and the negative sample node, and the graph structure feature vector represents the characteristics of the subgraph structure containing a preset order; A screening module, used for screening nodes matching the nodes to be processed as information push nodes based on the aggregated feature vectors of each node in the graph network; The push module is used to push the business information corresponding to the information push node.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Method and device for processing interactive sequence data

    CN111258469A

  • Object recommendation method and device based on intelligence, and storage medium

    CN111382190A