Method and device for training vectorization model of graph network, method and device for vectorization of graph network, equipment and computer program product

Through unsupervised comparative learning methods and graph network enhancement strategies, vectorized models are trained on the graph network, solving the problem of difficult and few samples in the graph network training, and achieving better training effects and diversity.

CN119940403APending Publication Date: 2025-05-06中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510054944.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the vectorized model training of graph networks has problems such as difficulty in annotating graph samples and poor training results.

Method used

Unsupervised contrast learning method is adopted to enhance the original graph network data through preset graph network enhancement strategy, generate enhanced graph network data, and use the vectorized model of graph network for vectorization. At the same time, multi-dimensional loss value (node ​​loss value, edge loss value and global loss value) is calculated and model parameters are optimized to improve training effect.

Benefits of technology

It effectively solves the problem of difficult and few samples for graph network training, improves the learning effect of the model, and meets the vectorization needs of graph networks in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940403A_ABST
    Figure CN119940403A_ABST
Patent Text Reader

Abstract

The invention discloses a training method and device of a vectorization model of a graph network, a vectorization method and device of the graph network, equipment and a computer program product. The training method comprises the steps of obtaining original graph network data; performing enhancement processing on the original graph network data by using a preset graph network enhancement strategy to obtain enhanced graph network data; according to the original graph network data and the enhanced graph network data, performing vectorization processing by using a vectorization model of the graph network to obtain a multi-dimensional vectorization result of the graph network data; according to a multi-dimensional vectorization result, calculating multi-dimensional loss values, including a node loss value, an edge loss value and a global loss value, of a vectorization model of the graph network; and optimizing model parameters according to the multi-dimensional loss value. According to the method, the model is trained by using an unsupervised contrast learning technology, the problems of difficulty in graph network sample labeling and few samples are solved, and the model training adopts multi-dimensional loss to jointly optimize the model parameters, so that the model training effect and the model application range are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of graph network technology, and in particular to a training method and device for a vectorized model of a graph network, a vectorized method and device, equipment and computer program product for a graph network. Background Art

[0002] Graph networks have a wide range of application scenarios in the computer field, such as social networks, biological proteins, and chemical structure modeling, so the digital representation of graph networks has also become a hot topic of research. In the early computer field, adjacency matrix vectors were often used to represent the structure of a graph, with the points of the graph representing each node and the edges representing the connection relationship between nodes. However, this representation method is not conducive to graph analysis tasks, such as graph similarity analysis, correlation analysis, and other tasks.

[0003] If some method can be used to vectorize graphs, then the similarity analysis of graphs can be transformed into the similarity analysis of two vectors. However, the current training of vectorized models of graph networks has problems such as difficulty in labeling graph samples and poor training results. Summary of the invention

[0004] The embodiments of the present application provide a training method and apparatus for a vectorized model of a graph network, a vectorized method and apparatus for a graph network, a device and a computer program product to improve the training effect of the vectorized model of the graph network and meet the vectorization requirements of different scenarios.

[0005] The present application embodiment adopts the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a method for training a vectorized model of a graph network, and the method for training a vectorized model of a graph network includes:

[0007] Get the original graph network data;

[0008] Performing enhancement processing on the original graph network data using a preset graph network enhancement strategy to obtain enhanced graph network data;

[0009] According to the original graph network data and the enhanced graph network data, vectorization processing is performed using a graph network vectorization model to obtain a multi-dimensional vectorization result of the graph network data;

[0010] According to the multi-dimensional vectorization result of the graph network data, a multi-dimensional loss value of the vectorization model of the graph network is calculated, wherein the multi-dimensional loss value includes a node loss value, an edge loss value, and a global loss value;

[0011] Optimize the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

[0012] Optionally, the using a preset graph network enhancement strategy to enhance the original graph network data to obtain enhanced graph network data includes:

[0013] Generate network structure data of the original graph network according to the original graph network data, wherein the network structure data includes node data and edge data associated with the nodes;

[0014] Performing statistics on the degree of each node according to the network structure data of the original graph network;

[0015] Determine the probability of each node being masked and the probability of the edge associated with the node being masked according to the degree of each node;

[0016] According to the probability of each node being masked and the probability of the edge associated with the node being masked, mask processing is performed on the original graph network data to obtain enhanced graph network data.

[0017] Optionally, performing vectorization processing using a graph network vectorization model according to the original graph network data and the enhanced graph network data to obtain a multi-dimensional vectorization result of the graph network data includes:

[0018] Performing vectorization processing on the original graph network data using the vectorization model of the graph network to obtain a multi-dimensional feature vector of the original graph network;

[0019] The enhanced graph network data is vectorized using the vectorization model of the graph network to obtain a multi-dimensional feature vector of the enhanced graph network.

[0020] Optionally, the multidimensional vectorization result of the graph network data includes a multidimensional feature vector of the original graph network and a multidimensional feature vector of the enhanced graph network, the multidimensional feature vectors of the original graph network and the enhanced graph network both include a node feature vector, an edge feature vector and a global feature vector, and the multidimensional loss value of the vectorization model of the graph network is calculated according to the multidimensional vectorization result of the graph network data, including:

[0021] Calculating a node loss value of a vectorized model of the graph network according to the node feature vector of the original graph network and the node feature vector of the enhanced graph network;

[0022] Calculating an edge loss value of a vectorized model of the graph network according to the edge feature vector of the original graph network and the edge feature vector of the enhanced graph network;

[0023] According to the global feature vector of the original graph network and the global feature vector of the enhanced graph network, a global loss value of the vectorized model of the graph network is calculated.

[0024] Optionally, optimizing the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain the trained vectorization model of the graph network includes:

[0025] Fusing the node loss value, the edge loss value, and the global loss value to obtain a fused loss value;

[0026] The fusion loss value is used to optimize the parameters of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

[0027] In a second aspect, an embodiment of the present application further provides a method for vectorizing a graph network, the method for vectorizing a graph network comprising:

[0028] Get the graph network data to be vectorized;

[0029] Performing vectorization processing on the graph network data to be vectorized by using a graph network vectorization model to obtain a multi-dimensional feature vector of the graph network;

[0030] Among them, the vectorization model of the graph network is trained based on any of the training methods for the vectorization model of the graph network described above.

[0031] Optionally, after vectorizing the graph network data to be vectorized by using the graph network vectorization model to obtain a multi-dimensional feature vector of the graph network, the graph network vectorization method further includes:

[0032] A correlation analysis is performed on the feature vectors of any two graph networks to obtain correlation analysis results of the two graph networks, wherein the correlation analysis results include at least one of a node correlation analysis result, an edge correlation analysis result, and an overall correlation analysis result.

[0033] In a third aspect, an embodiment of the present application further provides a training device for a vectorized model of a graph network, wherein the training device for a vectorized model of a graph network comprises:

[0034] A first acquisition unit, used to acquire original graph network data;

[0035] An enhancement processing unit, used to perform enhancement processing on the original graph network data using a preset graph network enhancement strategy to obtain enhanced graph network data;

[0036] A first vectorization unit, configured to perform vectorization processing on the original graph network data and the enhanced graph network data using a graph network vectorization model to obtain a multi-dimensional vectorization result of the graph network data;

[0037] A calculation unit, used to calculate a multi-dimensional loss value of the vectorization model of the graph network according to the multi-dimensional vectorization result of the graph network data, wherein the multi-dimensional loss value includes a node loss value, an edge loss value, and a global loss value;

[0038] The optimization unit optimizes the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

[0039] In a fourth aspect, an embodiment of the present application further provides a vectorization device for a graph network, wherein the vectorization device for a graph network comprises:

[0040] A second acquisition unit, used to acquire graph network data to be vectorized;

[0041] A second vectorization unit, used for performing vectorization processing on the graph network data to be vectorized by using a vectorization model of the graph network to obtain a multi-dimensional feature vector of the graph network;

[0042] Among them, the vectorization model of the graph network is trained based on the training device of the aforementioned vectorization model of the graph network.

[0043] In a fifth aspect, an embodiment of the present application further provides a device, including:

[0044] A processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform any of the aforementioned methods for training a vectorized model of a graph network and to perform any of the aforementioned methods for vectorizing a graph network.

[0045] In a sixth aspect, an embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements a training method for a vectorization model of any of the aforementioned graph networks, and implements a vectorization method for any of the aforementioned graph networks.

[0046] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: the training method of the vectorization model of the graph network in the embodiment of the present application first obtains the original graph network data; then uses the preset graph network enhancement strategy to enhance the original graph network data to obtain enhanced graph network data; then uses the vectorization model of the graph network to perform vectorization processing based on the original graph network data and the enhanced graph network data to obtain a multi-dimensional vectorization result of the graph network data; then calculates the multi-dimensional loss value of the vectorization model of the graph network based on the multi-dimensional vectorization result of the graph network data, and the multi-dimensional loss value includes a node loss value, an edge loss value, and a global loss value; finally, optimizes the parameters of the vectorization model of the graph network based on the multi-dimensional loss value of the vectorization model of the graph network to obtain a trained vectorization model of the graph network. The training method of the vectorization model of the graph network in the embodiment of the present application uses an unsupervised contrastive learning graph neural network to solve the problem of difficult annotation of training samples and small number of samples of the graph network, and the model training uses node loss, edge loss, and global loss to jointly optimize the model parameters, which improves the learning effect of the model and meets the needs of graph network vectorization in subsequent different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0048] Figure 1 A schematic diagram of a flow chart of a method for training a vectorized model of a graph network in an embodiment of the present application;

[0049] Figure 2 This is a schematic diagram of the structure of a graph network data in an embodiment of the present application;

[0050] Figure 3 A schematic diagram of a training process of a vectorization model of a graph network in an embodiment of the present application;

[0051] Figure 4 A schematic diagram of a flow chart of a method for vectorizing a graph network in an embodiment of the present application;

[0052] Figure 5 A schematic diagram of the structure of a training device for a vectorized model of a graph network in an embodiment of the present application;

[0053] Figure 6 A schematic diagram of the structure of a graph network vectorization device in an embodiment of the present application;

[0054] Figure 7 This is a schematic diagram of the structure of a device in an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0056] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0057] The main technical terms involved in this application are:

[0058] (1) Contrastive Learning

[0059] Contrastive learning is an unsupervised learning method that generates positive and negative sample pairs through data augmentation. The goal of training graph neural networks is to make data of the same type similar (data before and after augmentation belong to the same type) and data of different types dissimilar. The goal of contrastive learning is to make the original image more similar to the augmented image, and to make the original image less similar to other images.

[0060] (2) Graph Neural Networks

[0061] A graph is a data structure consisting of nodes and edges, used to represent the relationship between objects. Nodes are also called vertices, and edges represent the connection relationship between nodes.

[0062] Graph Neural Networks is a deep learning method based on graph structure. Traditional neural networks are mainly used to process data with regular structures, such as images and texts, while graph neural networks are specially designed to process graph structured data, such as social networks and molecular structures. The core idea of ​​graph neural networks is to enrich the representation of nodes by using the relationship between nodes. By defining the connection relationship between nodes on the graph, graph neural networks can use the node's neighbor information (first order and second order) to update the node's representation, thereby realizing information transmission and learning of the entire graph.

[0063] (3) Mask

[0064] Mask is a technique used to select, filter, or hide specific areas in an image. Mask is usually a matrix with the same dimensions as the original image, in which the elements determine whether to keep, modify, or block the image pixels at the corresponding position.

[0065] At present, the training of vectorized models of graph networks mainly has the following problems:

[0066] (1) Graph sample labeling is difficult. Training a deep learning model is a simple way to obtain graph vectors, but graph sample labeling is very difficult. The human eye can intuitively judge whether two images are similar, but it is very difficult to judge whether a graph network consists of a large number of edges and nodes.

[0067] (2) The effect of graph sample enhancement is not good. In deep learning, some enhancement methods are often used to improve the robustness of the model. For example, in the field of images, enhancement methods include flipping, translating, and cropping images. On the one hand, this increases the number of training samples, and on the other hand, it allows the model to learn richer information. In the field of graph network enhancement, the random mask node method currently used does not consider the connection relationship between nodes and edges, which often leads to a large difference between the enhanced graph network and the original graph network.

[0068] (3) The versatility of the model is poor. The currently trained graph network vectorization model only uses a single loss function to optimize the model, and the graph network vector output by the model is also relatively simple, which makes it difficult to meet the vectorization requirements of different usage scenarios.

[0069] Based on this, the present application embodiment provides a training method for a vectorized model of a graph network, such as Figure 1 As shown, a flow chart of a method for training a vectorized model of a graph network in an embodiment of the present application is provided, and the method for training a vectorized model of a graph network at least includes the following steps S110 to S150:

[0070] Step S110, obtaining original graph network data.

[0071] When training a vectorized model of a graph network, you need to first obtain the training sample data of the model. The training sample data can be the collected original graph network data, which can include various graph structures with node and edge relationships, such as social networks, knowledge graphs, and transportation networks. The quality and diversity of the original graph network data are crucial to the subsequent training effect.

[0072] In order to make the acquired original graph network data meet the structural requirements of subsequent model input, the relationships between all nodes and edges contained in each graph network sample can be mapped, and the relationship between edges can be represented by a binary tuple of nodes. Figure 2 As shown, a schematic diagram of the structure of a graph network data in an embodiment of the present application is provided. The graph network includes nodes n1, n2, n3, n4, and n5, wherein the edge associated with node n1 can be represented as E<n1,n2> ,E<n1,n3> In the same way, we can find the edges associated with nodes n2, n3, n4, and n5, including E<n2,n3> , E<n2,n5> , E<n1,n2> ,….

[0073] Step S120, using a preset graph network enhancement strategy to enhance the original graph network data to obtain enhanced graph network data.

[0074] The training of the vectorized model of the graph network in the embodiment of the present application is carried out in an unsupervised, contrastive learning manner, without the need to annotate samples in advance, and only requires the use of a graph network enhancement strategy such as the Mask technology to perform graph enhancement processing on the above-mentioned original graph network data to obtain an enhanced graph network. The graph network enhancement strategy used here needs to consider the importance of nodes and edges in the graph network, and adaptively determine the possibility of each node and edge being Masked, so that the enhanced graph network does not differ too much from the original graph network, thereby affecting the effect of subsequent model learning.

[0075] Similarly, the enhanced graph network is the same as the original graph network. It is necessary to map the relationships between all nodes and edges in the enhanced graph network, and use node binary tuples to represent edge relationships.

[0076] Step S130, performing vectorization processing using a graph network vectorization model according to the original graph network data and the enhanced graph network data, to obtain a multi-dimensional vectorization result of the graph network data.

[0077] Based on the original graph network data and enhanced graph network data obtained in the above steps, the original graph network data and enhanced graph network data are respectively vectorized using the graph network vectorization model. The graph network vectorization model of the embodiment of the present application is a graph neural network, which is trained to output multi-dimensional graph network vectors, such as node vectors, edge vectors, and graph network overall vectors. The specific type of graph neural network to be used can be flexibly selected by those skilled in the art according to actual needs, and is not specifically limited here.

[0078] Step S140, based on the multi-dimensional vectorization result of the graph network data, calculate the multi-dimensional loss value of the vectorization model of the graph network, wherein the multi-dimensional loss value includes a node loss value, an edge loss value, and a global loss value.

[0079] After obtaining the multi-dimensional vectorization results of the original graph network data and the enhanced graph network data, it is necessary to calculate the multi-dimensional loss values ​​of the graph network vectorization model, including node loss value, edge loss value and global loss value, according to the multi-dimensional vectorization results of the original graph network data and the enhanced graph network data. The size of the node loss value reflects the learning effect of the model on the node features in the graph network data, the size of the edge loss value reflects the learning effect of the model on the edge features in the graph network data, and the size of the global loss value reflects the learning effect of the model on the overall features of the graph network data.

[0080] The loss function used to calculate the loss value above may be, for example, a cross entropy loss function. Of course, those skilled in the art may also flexibly select other loss functions according to actual needs, which is not specifically limited here.

[0081] Step S150, optimizing the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

[0082] The training goal of the vectorized model of the graph network in the embodiment of the present application is to enable the model to have good learning capabilities for node features, edge features, and overall features of the graph network. Therefore, it is necessary to jointly optimize the model parameters based on the multi-dimensional loss values ​​calculated above until the model meets the training requirements.

[0083] The training method of the graph network vectorization model in the embodiment of the present application uses an unsupervised contrastive learning graph neural network to solve the problems of difficult labeling of graph network training samples and small number of samples. The model training uses node loss, edge loss and global loss to jointly optimize model parameters, which improves the learning effect of the model and meets the needs of graph network vectorization in subsequent different application scenarios.

[0084] In some embodiments of the present application, the use of a preset graph network enhancement strategy to enhance the original graph network data to obtain enhanced graph network data includes: generating network structure data of the original graph network based on the original graph network data, the network structure data including node data and data of edges associated with the nodes; counting the degree of each node based on the network structure data of the original graph network; determining the probability of each node being masked and the probability of an edge associated with the node being masked based on the degree of each node; and performing masking processing on the original graph network data based on the probability of each node being masked and the probability of an edge associated with the node being masked to obtain enhanced graph network data.

[0085] like Figure 3 As shown, a schematic diagram of the training process of a vectorized model of a graph network in an embodiment of the present application is provided. In the graph enhancement stage, it mainly involves enhancing the nodes and edges of the graph network using a preset graph network enhancement strategy.

[0086] Specifically, we can first count the degree d of each node contained in each graph network. i , the degree of a node represents the number of edges associated with the node, for example, Figure 2 , the degree of node n1 is d1=2.

[0087] Then, we use the graph enhancement technique with degree weighted fusion to calculate the probability of each node and each edge being masked. For example, after considering the “weight” of the node, node n iThe probability of being Masked is:

[0088]

[0089] Among them, d i Represents node n i N represents the degree of a node, and N represents a hyperparameter (which can be customized, such as 10). The larger the degree of a node, that is, the larger the weight, the lower the probability that the node is masked. By setting hyperparameters, the overall mask probability can be adjusted.

[0090] Edge E <n i , n j >The probability of being Masked is:

[0091]

[0092] Among them, d i and d j Represents node n i and node n j M represents the hyperparameter (which can be customized, such as 10). If the degree of the node associated with the current edge is larger, that is, the weight of the edge is larger, the probability of the edge being masked is lower.

[0093] From the above formulas (1) and (2), it can be seen that the probability of a node being masked mainly depends on the number of edges associated with the node. The more edges associated with the node, the more important the node is in the entire graph network, that is, the greater the weight, and the lower the probability of being masked. The probability of an edge being masked mainly depends on the importance of the two nodes associated with the edge. The greater the weight of the two nodes associated with the edge, the lower the probability of this edge being masked.

[0094] After calculating the probability of each node and each edge being masked, the probability of each node and each edge being masked can be compared with a set probability threshold. If it is greater than the threshold, the node or edge will be masked, otherwise it will not be masked. The size of the probability threshold can be customized, mainly playing a dynamic role in regulating the degree of masking.

[0095] The embodiment of the present application analyzes the importance of nodes and edges in the entire graph network and adaptively performs enhancement processing on the graph network, thereby improving the enhancement effect of the graph network and avoiding the problem of excessive difference before and after enhancement.

[0096] In some embodiments of the present application, the vectorization processing is performed on the original graph network data and the enhanced graph network data using a vectorization model of the graph network to obtain a multi-dimensional vectorization result of the graph network data, including: vectorization processing is performed on the original graph network data using the vectorization model of the graph network to obtain a multi-dimensional feature vector of the original graph network; and vectorization processing is performed on the enhanced graph network data using the vectorization model of the graph network to obtain a multi-dimensional feature vector of the enhanced graph network.

[0097] When using the vectorization model of the graph network for vectorization processing, the original graph network data and the corresponding enhanced graph network data are constituted into sample pairs for comparative learning, and the vectorization model of the graph network performs vectorization processing separately, thereby obtaining the multi-dimensional feature vector of the original graph network and the multi-dimensional feature vector of the enhanced graph network respectively. The multi-dimensional feature vector may include node feature vectors, edge feature vectors and global feature vectors, which serve as the basis for the subsequent calculation of the multi-dimensional loss value.

[0098] In some embodiments of the present application, the multidimensional vectorization result of the graph network data includes the multidimensional feature vector of the original graph network and the multidimensional feature vector of the enhanced graph network. The multidimensional feature vectors of the original graph network and the enhanced graph network both include node feature vectors, edge feature vectors and global feature vectors. Calculating the multidimensional loss value of the vectorization model of the graph network based on the multidimensional vectorization result of the graph network data includes: calculating the node loss value of the vectorization model of the graph network based on the node feature vector of the original graph network and the node feature vector of the enhanced graph network; calculating the edge loss value of the vectorization model of the graph network based on the edge feature vector of the original graph network and the edge feature vector of the enhanced graph network; calculating the global loss value of the vectorization model of the graph network based on the global feature vector of the original graph network and the global feature vector of the enhanced graph network.

[0099] Continue to refer Figure 3 In order to improve the training effect and application scope of the model, the embodiment of the present application adopts node loss, edge loss and global loss to constitute the joint loss of the network, so as to ensure that the training using contrastive learning can better fit the real graph network information.

[0100] The node loss is calculated by comparing the node feature vectors of the graph network data before and after enhancement. The node loss function mainly ensures that the nodes before and after the mask are the most similar, that is, the model can learn the characteristics of the nodes through continuous training. The edge loss can be calculated by comparing the edge feature vectors of the graph network data before and after enhancement. The edge loss function can estimate the connection relationship of the edge and strengthen the model's ability to learn the edge. Therefore, the global loss can be calculated by comparing the global feature vectors of the graph network data before and after enhancement. The function is to ensure that the difference between the graph network before and after enhancement is the lowest. Therefore, through comparative learning, the vectorized model of the graph network can fully learn the node features and edge features of the graph network before and after enhancement.

[0101] In some embodiments of the present application, optimizing the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain a trained vectorization model of the graph network includes: fusing the node loss value, the edge loss value and the global loss value to obtain a fused loss value; and optimizing the parameters of the vectorization model of the graph network using the fused loss value to obtain a trained vectorization model of the graph network.

[0102] By fusing the node loss values, edge loss values, and global loss values ​​obtained in the above embodiments, a fused loss value can be obtained. The fusion method can be a simple weighted summation or a more complex function combination. The choice of weights mainly depends on the requirements of the specific task and the characteristics of the data. The fused loss value is intended to comprehensively reflect the feature learning ability of the model at the node, edge, and global levels.

[0103] The obtained fusion loss value is used to update the parameters of the vectorized model of the graph network through optimization algorithms (such as gradient descent, Adam, etc.). The goal of optimization is to minimize the fusion loss value to obtain a trained model. During the optimization process, the model parameters will be continuously adjusted according to the change of the loss value until the preset stop condition is reached (such as loss value convergence, reaching the maximum number of iterations, etc.). When the optimization process ends, the trained vectorized model of the graph network is output, which can generate high-quality node and edge vector representations and is suitable for various graph analysis tasks.

[0104] In summary, this solution comprehensively evaluates the performance of the model by fusing node, edge, and global loss values, and continuously adjusts the model parameters through the optimization algorithm to minimize these loss values, thereby obtaining a vectorized model of the trained graph network. This method helps to improve the accuracy and robustness of the model in various graph analysis tasks.

[0105] The present application also provides a method for vectorizing a graph network. Figure 4As shown, a schematic diagram of a flow chart of a method for vectorizing a graph network in an embodiment of the present application is provided, and the method for vectorizing a graph network at least includes the following steps S410 to S420:

[0106] Step S410, obtaining graph network data to be vectorized;

[0107] Step S420, using a vectorization model of a graph network to perform vectorization processing on the graph network data to be vectorized, to obtain a multi-dimensional feature vector of the graph network;

[0108] Among them, the vectorization model of the graph network is trained based on any of the training methods for the vectorization model of the graph network described above.

[0109] In the vectorization stage of the graph network, the binary representation of the nodes and edges contained in the graph network can be first extracted from the graph network data to be vectorized, and then the binary representation of the nodes and edges contained in the graph network can be input into the trained graph network vectorization model to output a multi-dimensional feature vector of the graph network, which may include, for example, node feature vectors, edge feature vectors, and global feature vectors.

[0110] Of course, it should be noted that the vectorized model of the graph network trained in this application has the ability to simultaneously output node feature vectors, edge feature vectors and global feature vectors. As for which dimension or dimensions of feature vectors to output, it can be flexibly set according to the needs of the actual application scenario and is not specifically limited here.

[0111] In some embodiments of the present application, after using the vectorization model of the graph network to vectorize the graph network data to obtain a multi-dimensional feature vector of the graph network, the graph network vectorization method also includes: performing correlation analysis on the feature vectors of any two graph networks to obtain correlation analysis results of the two graph networks, and the correlation analysis results include at least one of a node correlation analysis result, an edge correlation analysis result, and an overall correlation analysis result.

[0112] Based on the needs of actual application scenarios such as fake web page identification and identification of similar people in social networks, the feature vectors of any two graph networks output by the model can be subjected to correlation analysis. For example, algorithms such as cosine similarity can be used to calculate the similarity score. The higher the similarity score, the more related the two graph networks are.

[0113] When performing specific correlation analysis, the node feature vectors, edge feature vectors, and global feature vectors of the two graph networks can be selectively analyzed according to actual needs. For example, if the scenario focuses more on the correlation of nodes in the graph network, similarity calculation can be performed only on the node feature vectors; if the scenario focuses more on the correlation of edges, similarity calculation can be performed only on the edge feature vectors.

[0114] The present application also provides a training device 500 for a vectorized model of a graph network, such as Figure 5 As shown, a schematic diagram of the structure of a training device for a vectorized model of a graph network in an embodiment of the present application is provided. The training device 500 for the vectorized model of a graph network includes: a first acquisition unit 510, an enhancement processing unit 520, a first vectorization unit 530, a calculation unit 540, and an optimization unit 550, wherein:

[0115] A first acquisition unit 510 is used to acquire original graph network data;

[0116] An enhancement processing unit 520, configured to perform enhancement processing on the original graph network data using a preset graph network enhancement strategy to obtain enhanced graph network data;

[0117] A first vectorization unit 530 is used to perform vectorization processing on the original graph network data and the enhanced graph network data using a graph network vectorization model to obtain a multi-dimensional vectorization result of the graph network data;

[0118] A calculation unit 540 is used to calculate a multi-dimensional loss value of the vectorization model of the graph network according to the multi-dimensional vectorization result of the graph network data, wherein the multi-dimensional loss value includes a node loss value, an edge loss value, and a global loss value;

[0119] The optimization unit 550 optimizes the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

[0120] In some embodiments of the present application, the enhancement processing unit 520 is specifically used to: generate network structure data of the original graph network based on the original graph network data, the network structure data including node data and data of edges associated with the nodes; perform statistics on the degree of each node based on the network structure data of the original graph network; determine the probability of each node being masked and the probability of the edge associated with the node being masked based on the degree of each node; perform mask processing on the original graph network data based on the probability of each node being masked and the probability of the edge associated with the node being masked to obtain enhanced graph network data.

[0121] In some embodiments of the present application, the first vectorization unit 530 is specifically used to: use the vectorization model of the graph network to vectorize the original graph network data to obtain a multi-dimensional feature vector of the original graph network; use the vectorization model of the graph network to vectorize the enhanced graph network data to obtain a multi-dimensional feature vector of the enhanced graph network.

[0122] In some embodiments of the present application, the multi-dimensional vectorization result of the graph network data includes the multi-dimensional feature vector of the original graph network and the multi-dimensional feature vector of the enhanced graph network. The multi-dimensional feature vectors of the original graph network and the enhanced graph network both include node feature vectors, edge feature vectors and global feature vectors. The calculation unit 540 is specifically used to: calculate the node loss value of the vectorization model of the graph network according to the node feature vector of the original graph network and the node feature vector of the enhanced graph network; calculate the edge loss value of the vectorization model of the graph network according to the edge feature vector of the original graph network and the edge feature vector of the enhanced graph network; calculate the global loss value of the vectorization model of the graph network according to the global feature vector of the original graph network and the global feature vector of the enhanced graph network.

[0123] In some embodiments of the present application, the optimization unit 550 is specifically used to: fuse the node loss value, the edge loss value and the global loss value to obtain a fused loss value; use the fused loss value to optimize the parameters of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

[0124] It can be understood that the above-mentioned training device for the vectorized model of the graph network can implement the various steps of the training method for the vectorized model of the graph network provided in the aforementioned embodiments. The relevant explanations on the training method for the vectorized model of the graph network are applicable to the training device for the vectorized model of the graph network and will not be repeated here.

[0125] The present application embodiment also provides a graph network vectorization device 600, such as Figure 6 As shown, a schematic diagram of the structure of a vectorization device for a graph network in an embodiment of the present application is provided, wherein the vectorization device 600 for a graph network includes: a second acquisition unit 610 and a second vectorization unit 620, wherein:

[0126] A second acquisition unit 610 is used to acquire graph network data to be vectorized;

[0127] A second vectorization unit 620, configured to perform vectorization processing on the graph network data to be vectorized by using a graph network vectorization model to obtain a multi-dimensional feature vector of the graph network;

[0128] Among them, the vectorization model of the graph network is trained based on the training device of the aforementioned vectorization model of the graph network.

[0129] In some embodiments of the present application, the graph network vectorization device 600 also includes: a correlation analysis unit, which is used to vectorize the graph network data to be vectorized using the graph network vectorization model to obtain a multi-dimensional feature vector of the graph network, and then perform correlation analysis on the feature vectors of any two graph networks to obtain correlation analysis results of the two graph networks, wherein the correlation analysis results include at least one of a node correlation analysis result, an edge correlation analysis result, and an overall correlation analysis result.

[0130] It can be understood that the above-mentioned graph network vectorization device can implement the various steps of the graph network vectorization method provided in the aforementioned embodiments. The relevant explanations about the graph network vectorization method are applicable to the graph network vectorization device and will not be repeated here.

[0131] Figure 7 Schematic diagram of the structure of a device in the embodiment of the present application. Figure 7 As shown, the device includes one or more processors (or processing units), may further include one or more memories coupled to the processors, and may further include a communication module coupled to the processors.

[0132] The communication module can be used to communicate with other devices or apparatuses, such as the transmission or reception of data and / or signals. The communication module can have at least one communication module for communication. The communication module can include any interface necessary for communicating with other devices. Exemplarily, the communication module can be a transceiver, a circuit, a bus, a module, or other types of communication modules.

[0133] The processor may include, but is not limited to, at least one of the following: a general-purpose computer, a special-purpose computer, a microcontroller, a digital signal controller (DSP), or one or more of a controller-based multi-core controller architecture. The device may have multiple processors, such as application-specific integrated circuit chips, which are time-dependent and synchronized with a clock of a main processor.

[0134] The memory may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, at least one of the following: read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, hard disk, compact disc (CD), digital video disc (DVD), or other magnetic storage and / or optical storage. Examples of volatile memories include, but are not limited to, at least one of the following: random access memory (RAM), or other volatile memories that do not persist during the duration of a power outage.

[0135] The computer program includes computer executable instructions executed by an associated processor. The program can be stored in ROM. The processor can perform any suitable actions and processes by loading the program into RAM.

[0136] The possible implementation of the present application can be implemented by means of a program, so that the communication device can perform any process discussed in the above embodiments. The possible implementation of the present application can also be implemented by hardware or by a combination of software and hardware.

[0137] In some embodiments, the program may be tangibly contained in a computer-readable storage medium, which may be included in the device (such as in a memory) or other storage device accessible by the device. The program may be loaded from the computer-readable storage medium to the RAM for execution. The computer-readable storage medium may include any type of tangible non-volatile memory, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc.

[0138] The present application embodiment also provides a computer-readable storage medium, on which computer instructions or program codes are stored, and when the processor runs the instructions or the program codes, the processor executes the methods and functions involved in any of the above embodiments. Computer-readable media can be any tangible medium containing or storing programs for or related to instruction execution systems, devices or equipment. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or devices, or any suitable combination thereof. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrations. More detailed examples of computer-readable storage media include electrical connections with one or more wires, magnetic media (e.g., disks, floppy disks, hard disks, tapes, magnetic storage devices), optical media (e.g., optical storage devices, DVDs), semiconductor media (e.g., solid-state hard drives), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), or any suitable combination thereof, etc.

[0139] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The embodiment of the present application also provides at least one computer program product tangibly stored on a non-temporary computer-readable storage medium. The computer program product includes one or more computer executable instructions, such as instructions included in a program module, which are executed in a device on a real or virtual processor of the target to perform the process, method and function involved in any of the above embodiments. When the computer program instruction is loaded and executed on a computer, a process or function according to an embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instruction can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instruction can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.

[0140] The present application embodiment also proposes a computer program product, including a computer program or instruction, when the computer program or instruction is run on a computer, the computer is made to perform the process, method and function in the above-mentioned embodiment. Usually, a program module includes routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or realize specific abstract data types. In various embodiments, the functions of program modules can be combined or divided between program modules as needed. Machine executable instructions for program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.

[0141] In general, various embodiments of the present application may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be performed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of the present disclosure are shown and described as block diagrams, flow charts, or using some other graphical representations, it should be understood that the boxes, devices, systems, techniques, or methods described herein may be implemented as, for example, non-limiting examples, hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0142] It should be noted that although the embodiments of the present application are described above in conjunction with the accompanying drawings, the above embodiments are not independent of each other, and they can also be combined to obtain other embodiments. The division of the modes, situations, categories and embodiments in the embodiments of the present application is only for the convenience of description and should not constitute a special limitation. The features in the various modes, categories, situations and embodiments can be combined with each other in a logical manner. The various implementation methods of the present application can be combined arbitrarily to achieve different technical effects. The embodiments of the present application no longer list various combinations.

[0143] In addition, although the operation of the method of the present disclosure is described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flow chart can change the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution. It should also be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of a device described above can be further divided into being embodied by multiple devices.

[0144] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0145] The above is only the embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A training method for a vectorized model of a graph network, characterized in that: The training method of the vectorized model of the graph network includes: Get the original graph network data; Performing enhancement processing on the original graph network data using a preset graph network enhancement strategy to obtain enhanced graph network data; According to the original graph network data and the enhanced graph network data, vectorization processing is performed using a graph network vectorization model to obtain a multi-dimensional vectorization result of the graph network data; According to the multi-dimensional vectorization result of the graph network data, a multi-dimensional loss value of the vectorization model of the graph network is calculated, wherein the multi-dimensional loss value includes a node loss value, an edge loss value, and a global loss value; Optimize the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

2. The training method of the vectorized model of the graph network according to claim 1, characterized in that: The enhancing process of the original graph network data by using a preset graph network enhancement strategy to obtain enhanced graph network data includes: Generate network structure data of the original graph network according to the original graph network data, wherein the network structure data includes node data and edge data associated with the nodes; Performing statistics on the degree of each node according to the network structure data of the original graph network; Determine the probability of each node being masked and the probability of the edge associated with the node being masked according to the degree of each node; According to the probability of each node being masked and the probability of the edge associated with the node being masked, mask processing is performed on the original graph network data to obtain enhanced graph network data.

3. The training method of the vectorized model of the graph network according to claim 1, characterized in that: The vectorization processing is performed using a vectorization model of a graph network according to the original graph network data and the enhanced graph network data to obtain a multi-dimensional vectorization result of the graph network data, including: Performing vectorization processing on the original graph network data using the vectorization model of the graph network to obtain a multi-dimensional feature vector of the original graph network; The enhanced graph network data is vectorized using the vectorization model of the graph network to obtain a multi-dimensional feature vector of the enhanced graph network.

4. The training method of the vectorized model of the graph network according to claim 1, characterized in that: The multi-dimensional vectorization result of the graph network data includes a multi-dimensional feature vector of the original graph network and a multi-dimensional feature vector of the enhanced graph network, and the multi-dimensional feature vectors of the original graph network and the enhanced graph network both include a node feature vector, an edge feature vector, and a global feature vector. The multi-dimensional loss value of the vectorization model of the graph network is calculated according to the multi-dimensional vectorization result of the graph network data, including: Calculating a node loss value of a vectorized model of the graph network according to the node feature vector of the original graph network and the node feature vector of the enhanced graph network; Calculating an edge loss value of a vectorized model of the graph network according to the edge feature vector of the original graph network and the edge feature vector of the enhanced graph network; According to the global feature vector of the original graph network and the global feature vector of the enhanced graph network, a global loss value of the vectorized model of the graph network is calculated.

5. The method for training a vectorized model of a graph network according to claim 1, characterized in that: Optimizing the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain the trained vectorization model of the graph network includes: Fusing the node loss value, the edge loss value, and the global loss value to obtain a fused loss value; The fusion loss value is used to optimize the parameters of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

6. A method for vectorizing a graph network, characterized in that: The vectorization method of the graph network includes: Get the graph network data to be vectorized; Performing vectorization processing on the graph network data to be vectorized by using a graph network vectorization model to obtain a multi-dimensional feature vector of the graph network; The vectorized model of the graph network is trained based on the training method of the vectorized model of the graph network according to any one of claims 1 to 5.

7. The method for vectorizing a graph network according to claim 6, characterized in that: After the graph network data to be vectorized is vectorized using the graph network vectorization model to obtain a multi-dimensional feature vector of the graph network, the graph network vectorization method further includes: A correlation analysis is performed on the feature vectors of any two graph networks to obtain correlation analysis results of the two graph networks, wherein the correlation analysis results include at least one of a node correlation analysis result, an edge correlation analysis result, and an overall correlation analysis result.

8. A training device for a vectorized model of a graph network, characterized in that: The training device of the vectorized model of the graph network includes: A first acquisition unit, used to acquire original graph network data; An enhancement processing unit, used to perform enhancement processing on the original graph network data using a preset graph network enhancement strategy to obtain enhanced graph network data; A first vectorization unit, configured to perform vectorization processing on the original graph network data and the enhanced graph network data using a graph network vectorization model to obtain a multi-dimensional vectorization result of the graph network data; A calculation unit, used to calculate a multi-dimensional loss value of the vectorization model of the graph network according to the multi-dimensional vectorization result of the graph network data, wherein the multi-dimensional loss value includes a node loss value, an edge loss value, and a global loss value; The optimization unit optimizes the parameters of the vectorization model of the graph network according to the multi-dimensional loss value of the vectorization model of the graph network to obtain a trained vectorization model of the graph network.

9. A graph network vectorization device, characterized in that: The vectorization device of the graph network includes: A second acquisition unit, used to acquire graph network data to be vectorized; A second vectorization unit, used for performing vectorization processing on the graph network data to be vectorized by using a vectorization model of the graph network to obtain a multi-dimensional feature vector of the graph network; Wherein, the vectorization model of the graph network is trained based on the training device of the vectorization model of the graph network described in claim 8.

10. A device comprising: processor; And a memory arranged to store computer executable instructions, which, when executed, cause the processor to execute the training method of the vectorized model of the graph network of any one of claims 1 to 5, and execute the vectorization method of the graph network of any one of claims 6 to 7.

11. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, they implement a method for training a vectorized model of a graph network as claimed in any one of claims 1 to 5, and a method for vectorizing a graph network as claimed in any one of claims 6 to 7.