Method for Measuring Node Similarity in Heterogeneous Knowledge Network Based on Multi-View Consistency
Through node characterization learning and multi-view consistency methods, the complexity problem of node similarity measurement in heterogeneous knowledge networks is solved, and more accurate and robust similarity measurements are achieved.
Patent Information
- Application Number
- CN202510135009.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-07
AI Technical Summary
Traditional node similarity measurement methods encounter difficulties in the face of heterogeneous knowledge networks, especially when dealing with multimodal, multi-source data and complex connection relationships, it is difficult to accurately reflect the similarity of nodes.
Through node characterization learning, node characterization is constructed for different types of nodes in heterogeneous knowledge networks, and path sampling is performed based on predefined multiple metapaths to build multiple views. Then, a common representation of the nodes in multiple different views is determined and the similarity between the nodes is calculated based on these common representations.
This method can more accurately reflect node similarity in heterogeneous knowledge networks, improve the comprehensiveness and accuracy of similarity measurements, and reduce the impact of node type or relationship type differences, and improve robustness.
Smart Images

Figure CN119577480B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graphs, and particularly to a method for measuring the similarity of heterogeneous knowledge network nodes based on multi-view Figure 1 consistency. Background Art
[0002] Node similarity refers to the degree of similarity between two nodes in a graph or network. When understanding, mining, and applying a knowledge network, node similarity measurement is a key factor supporting data mining tasks. For example, when making recommendations based on a knowledge network, if the object that strictly meets the recommendation requirements is not available in the current scenario, by analyzing the similarity of nodes, similar objects can be recommended to users, thereby improving the performance of the recommendation system. Another example is when completing a knowledge graph. By measuring node similarity, the missing information in the knowledge graph can be better understood, and thus the knowledge graph can be completed and relationship prediction can be performed.
[0003] Traditional node similarity measurements are usually based on feature similarity and node connectivity. However, when faced with heterogeneous knowledge networks, these traditional measurement methods face many challenges. On the one hand, the nodes in heterogeneous knowledge networks involve multi-modal and multi-source data, which makes the similarity measurement between nodes more complex; on the other hand, the connection relationships between nodes in heterogeneous knowledge networks are complex and there are many types of edges, posing new challenges to methods based on node connectivity. Therefore, an innovative node similarity measurement method is needed to more accurately reflect the node similarity in heterogeneous knowledge networks. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method for measuring the similarity of heterogeneous knowledge network nodes based on multi-view Figure 1 consistency that can more accurately reflect the node similarity in heterogeneous knowledge networks.
[0005] A method for measuring the similarity of heterogeneous knowledge network nodes based on multi-view Figure 1 consistency provided by this application includes:
[0006] Construct node representations for different types of nodes in the heterogeneous knowledge network through node representation learning;
[0007] Based on a plurality of predefined meta-paths, sample paths from the heterogeneous knowledge network respectively to construct views corresponding to each meta-path;
[0008] For each node in the heterogeneous knowledge network, determine the common representation of the node under multiple different views based on the node representations of the node in each view;
[0009] For any two nodes in the heterogeneous knowledge network, determine the similarity between the two nodes based on the common representations of each node.
[0010] In some embodiments, construct node representations for different types of nodes in the heterogeneous knowledge network through node representation learning, including:
[0011] For each node in the heterogeneous knowledge network, obtain the node type corresponding to the node;
[0012] Determine the corresponding target representation model based on the node type;
[0013] Input the node attributes corresponding to the node into the target representation model to obtain the node representation output by the target representation model.
[0014] In some embodiments, determining the corresponding target representation model based on the node type includes:
[0015] Based on the node type of the node, determine the target representation model corresponding to the node type from the pre-configured correspondence relationship, where the correspondence relationship represents the correspondence relationship between the node type and the representation model.
[0016] In some embodiments, the node type includes one or more of images, texts, and entities;
[0017] The representation model includes one or more of a vision Transformer model, a word embedding model, and a multi-layer perceptron.
[0018] In some embodiments, based on a plurality of predefined meta-paths, respectively perform path sampling on the heterogeneous knowledge network to construct views corresponding to each meta-path, including:
[0019] Based on a plurality of predefined meta-paths, respectively perform path sampling on the heterogeneous knowledge network to obtain a node relationship sequence corresponding to each meta-path;
[0020] Construct a view corresponding to each meta-path based on the node relationship sequence.
[0021] In some embodiments, based on a plurality of predefined meta-paths, respectively perform path sampling on the heterogeneous knowledge network to obtain a node relationship sequence corresponding to each meta-path, including:
[0022] Define a plurality of meta-paths according to different node types and relationship types between nodes;
[0023] For each meta-path, extract path instances from the heterogeneous knowledge network based on the node type and relationship type represented by the meta-path;
[0024] Obtain the node relationship sequence corresponding to the meta-path according to the extracted path instances.
[0025] In some embodiments, for each node in the heterogeneous knowledge network, based on the node representations of the node on each view, determining the common representation of the node under multiple different views includes:
[0026] For each node in the heterogeneous knowledge network, randomly select two views from multiple views to form a view pair;
[0027] Input the node representations of the node in the view pair into a pre-trained heterogeneous graph neural network model to obtain the common representation output by the heterogeneous graph neural network model.
[0028] In some embodiments, for any two nodes in the heterogeneous knowledge network, determining the similarity between the two nodes based on the common representation of each node includes:
[0029] For any two nodes in the heterogeneous knowledge network, calculate the cosine similarity or Euclidean distance between the two nodes based on the common representation of each node to obtain the similarity.
[0030] In the method of the embodiments of the present application, by constructing views with multiple different relationship patterns, different types of information in the heterogeneous knowledge network can be comprehensively utilized, the features and relationships of nodes in different aspects can be captured, the association relationships between nodes can be considered more comprehensively, and the comprehensiveness and accuracy of similarity measurement can be improved. Moreover, by defining the common representation of nodes in different views, the influence caused by differences in node types or relationship types in the heterogeneous knowledge network can be reduced, and the robustness and accuracy of node similarity measurement can be improved. Brief Description of the Drawings
[0031] Figure 1 It is a flowchart of a similarity measurement method in an embodiment of the present application;
[0032] Figure 2 It is a flowchart of a similarity measurement method in an embodiment of the present application;
[0033] Figure 3 It is a schematic diagram of a similarity measurement method in an embodiment of the present application;
[0034] Figure 4 It is a schematic diagram of a similarity measurement method in an embodiment of the present application;
[0035] Figure 5 It is a schematic diagram of a similarity measurement method in an embodiment of the present application;
[0036] Figure 6Flowchart of the similarity measurement method in an embodiment of the present application;
[0037] Figure 7 Schematic diagram of the similarity measurement method in an embodiment of the present application;
[0038] Figure 8 Schematic diagram of the similarity measurement method in an embodiment of the present application. Detailed implementation manners
[0039] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present application and are not used to limit the present application.
[0040] Unmanned aerial vehicle (UAV) swarm confrontation refers to the collaborative operation of a swarm of UAVs. Due to the limitations of a single UAV, such as being easily damaged and having limited task execution capabilities, researchers have turned their research direction to UAV swarms. UAV swarm confrontation requires UAVs to have the capabilities of autonomous judgment, planning and decision-making, and to achieve information interaction and collaborative actions among swarms.
[0041] Regarding the confrontation problem between UAV swarms with different maneuvering capabilities, currently it mainly relies on establishing a collaborative confrontation model between UAV swarms, using the collaborative confrontation model to simulate the collaborative confrontation behaviors of UAVs in two-dimensional and three-dimensional spaces, and providing the autonomous decision-making capabilities of UAVs.
[0042] To realize the construction of the collaborative confrontation model, it is necessary to construct a multi-source heterogeneous knowledge network. Multi-source means diverse knowledge sources, and heterogeneous means diverse knowledge structures. By recommending, fusing and applying multi-source heterogeneous knowledge, it provides a data basis for the collaborative confrontation model of UAV swarms. Among them, the heterogeneous knowledge network needs to associate knowledge by calculating the node similarity between knowledge.
[0043] Node similarity refers to the degree of similarity between two nodes in a graph or network. When understanding, mining and applying a knowledge network, node similarity measurement is a key factor supporting data mining tasks. For example, when making recommendations based on a knowledge network, if the object that strictly meets the recommendation requirements is not available in the current scenario, by analyzing the node similarity, similar objects can be recommended to the user, thereby improving the performance of the recommendation system. Another example is when performing knowledge graph completion, by measuring the node similarity, the missing information in the knowledge graph can be better understood, thereby performing knowledge graph completion and relationship prediction.
[0044] Traditional node similarity measures are usually based on feature similarity and node connectivity. However, when faced with heterogeneous knowledge networks, these traditional measures face many challenges. A heterogeneous knowledge network refers to a complex network composed of different types of data sources and different kinds of data objects. In a heterogeneous knowledge network, nodes can represent different types of information, such as text, images, entities, etc., and edges represent the association relationships between these different types of information.
[0045] On the one hand, the nodes in a heterogeneous knowledge network involve multi-modal and multi-source data, which makes the similarity measure between nodes more complex. Traditional similarity measure methods often rely on single-type data features and are difficult to handle the complex data structures in heterogeneous knowledge networks. On the other hand, the connection relationships between nodes in a heterogeneous knowledge network are complex and there are many types of edges, bringing new challenges to methods based on node connectivity.
[0046] In related technologies, in order to measure the similarity of nodes in a heterogeneous knowledge network, some methods rely on manual analysis. Manual analysis can use artificial knowledge and combine the semantics, background, and actual application scenarios of nodes to measure the similarity of nodes in a heterogeneous knowledge network. Although analysts can make more flexible and intuitive judgments according to specific scenarios, which helps to handle some complex problems that are difficult to solve by computational models, the method based on manual analysis has certain limitations. First, manual analysis is affected by individual subjective judgments, and different analysts may produce different results, leading to subjectivity and inconsistency in the analysis. Second, the scale of heterogeneous knowledge networks is huge, and manual analysis is restricted by time and resources when dealing with large-scale networks, and analysts cannot effectively process all nodes and relationships in large-scale networks.
[0047] Another part of the methods is based on the similarity of node features and directly calculates the feature similarity as the similarity of nodes, which can intuitively understand the similarity between nodes. At the same time, compared with some complex deep learning methods, the methods based on node feature similarity usually have higher computational efficiency. However, the methods based on node feature similarity usually ignore the topological structure and relationship information between nodes. In some cases, the structural information of the network may contain important relationships, and this method may not be able to fully capture these relationships. In addition, this method is sensitive to the quality and expressive ability of node features. If the features of nodes are not sufficient to accurately reflect their positions or attributes in the network, the accuracy of similarity measurement is low.
[0048] Based on the problems and adjustments in the above related technologies, this application provides a method for measuring the similarity of heterogeneous knowledge network nodes based on multi-view Figure 1 consistency, starting from the consistency of nodes in a heterogeneous knowledge network under multiple views, and improving the accuracy and robustness of node similarity measurement.
[0049] As Figure 1 shown, the method for measuring the similarity of heterogeneous knowledge network nodes based on multi-view consistency provided by this application includes: Figure 1 S110. Construct node representations for different types of nodes in the heterogeneous knowledge network through node representation learning.
[0050] It can be understood that a heterogeneous knowledge network refers to a graph data network composed of different types of data sources or different kinds of data objects. Therefore, the nodes in a heterogeneous knowledge network often correspond to multiple types, and different types of nodes may have different attributes and relationships.
[0051] For example, in one example, the node types include but are not limited to images, texts, entities, etc. The node attributes corresponding to image nodes can be pixel points, the node attributes corresponding to text nodes can be word sequences, and the node attributes corresponding to entity nodes can be entity features.
[0052] At the same time, in a heterogeneous knowledge network, the relationship types of the connection relationships between different types of nodes are also equally complex. For example, a text node may be connected to an image node through a certain relationship type, and connected to another text node through another relationship type.
[0053] In the embodiments of this application, node representation (Entities Representation) refers to the representation of the original node features in the vector space, which is a feature vector with a length of L. The purpose of constructing node representations for the nodes in the heterogeneous knowledge network is to map the original node features of different types of nodes included in the heterogeneous knowledge network into the same representation space, so that these different types of nodes can be represented by feature vectors in the same space.
[0054] In some embodiments, multiple representation models can be pre-trained, and each representation model corresponds to a type of node. Thus, according to the different node types, a suitable representation model can be selected to construct the corresponding node representation. The following will be described in conjunction with
[0055] As Figure 2 shown, in some embodiments, in the similarity measurement method of this application, constructing node representations for different types of nodes in the heterogeneous knowledge network through node representation learning includes:
[0056] S111. For each node in the heterogeneous knowledge network, obtain the node type corresponding to the node. Figure 2 S112. Determine the corresponding target representation model based on the node type.
[0057]
[0058]
[0059] S113. Input the node attributes corresponding to the node into the target representation model to obtain the node representation output by the target representation model.
[0060] In the embodiments of the present application, first, a corresponding representation model can be constructed for each node type. The representation model is a neural network model based on deep learning. The input of the representation model is the node attributes, and the output is the node representation corresponding to the entity node.
[0061] For the convenience of description, in the following examples, the node types are taken as images, texts, and entities as examples to illustrate the process of constructing the node representation.
[0062] For image nodes, a vision Transformer model can be used. For example, in one example, a Transformer (ViT) model can be used to map the image corresponding to the image node from the pixel space to a vector space of length L, so as to obtain the corresponding node representation.
[0063] For text nodes, a word embedding model can be used to map the text information of the text node to a vector space of length L, so as to obtain the corresponding node representation. In one example, after mapping the text information of the text node to a vector space of length L, the Transformer model can be further used to learn the dependencies in the text sequence, so as to obtain the node representation.
[0064] For entity nodes, a multi-layer perceptron (MLP) can be used to map the entity features of the entity node to a vector space of length L, so as to obtain the corresponding node representation.
[0065] Here, L refers to the pre-set vector length of the node representation, which is a hyperparameter of the model. Its specific value can be selected according to requirements, and the present application does not limit this.
[0066] In the examples of the present application, after constructing the representation models for each type of node respectively, based on the correspondence between the node type and the representation model, a correspondence relationship as shown in Table 1 below can be generated:
[0067] Table 1
[0068]
[0069] In the embodiments of the present application, when processing the nodes in the heterogeneous knowledge network, taking any node as an example, first, the node type of the node can be obtained, and then based on the above-configured correspondence relationship, the representation module corresponding to the node type of the node can be determined by looking up the table, that is, the target representation model in the present application, and then the node representation corresponding to the node can be constructed by using the target representation model.
[0070] For example, in one example, taking the node type as "image" for instance. First, through the corresponding relationship shown in Table 1 above, it can be determined that the target representation model corresponding to the image node is the "Vision Transformer model". Then, the Vision Transformer model can be called to process this image node. Specifically, the node attribute of this image node (i.e., the image corresponding to the image node) can be input into the Vision Transformer model, and then the Vision Transformer model maps the original pixel features to a vector space of length L to obtain a feature vector of length L, and this feature vector is the node representation corresponding to the image node.
[0071] For example, in another example, taking the node type as "entity" for instance. First, through the corresponding relationship shown in Table 1 above, it can be determined that the target representation model corresponding to the entity node is the "Multi-Layer Perceptron", and then the Multi-Layer Perceptron can be called to process this entity node. Specifically, the node attribute of this entity node (i.e., the original entity features corresponding to the entity node) can be input into the Multi-Layer Perceptron, and then the Multi-Layer Perceptron maps the original entity features to a vector space of length L to obtain a feature vector of length L, and this feature vector is the node representation corresponding to the entity node.
[0072] The above takes one node in the heterogeneous knowledge network as an example to illustrate the method process of constructing node representations. For each node in the heterogeneous knowledge network, the above method process is sequentially repeated to complete the node representation learning of the entire heterogeneous knowledge network, which will not be elaborated in this application.
[0073] According to the above, through node representation learning, all different types of nodes in the heterogeneous knowledge network can be uniformly mapped to a common representation space, so as to obtain the node representation corresponding to each node in the same space. The node representation is the feature vector corresponding to the node, providing a data basis for subsequent similarity measurement.
[0074] S120. Based on a plurality of predefined meta-paths, respectively perform path sampling on the heterogeneous knowledge network to construct views corresponding to each meta-path.
[0075] A meta-path (Meta Path) is a path defined in the heterogeneous knowledge network, used to describe the relationship between different types of nodes. Utilizing meta-paths can help extract the structured relationships between nodes, form a structured description of node relationships, and thus better understand the relevance between nodes. Meta-paths can not only capture the topological structure of the knowledge network but also introduce semantic information through the combination of node types and relationship types in the path.
[0076] In the embodiments of the present application, multiple meta-paths can be predefined based on node types and relationship types, and multiple relationship views can be constructed by sampling the multiple meta-paths for subsequent multi-view Figure 1 consistency learning.
[0077] In some embodiments, meta-paths can first be defined according to node types and relationship types in a heterogeneous knowledge network. A meta-path consists of a sequence of node types and relationship types. For example, a meta-path of "A-R-B" means reaching a node of type B from a node of type A through a relationship type R. Such a meta-path can capture the relationship patterns between different types of nodes.
[0078] Then, for each defined meta-path, path sampling can be defined in the heterogeneous knowledge network, that is, path instances that conform to the meta-path definition are extracted from the heterogeneous knowledge network, and these path instances form a node relationship sequence consisting of a series of nodes and edges.
[0079] Finally, after obtaining the node relationship sequence, a subgraph can be constructed as a view based on the nodes and edges included in the node relationship sequence, and this view contains the nodes and edges that conform to the meta-path definition.
[0080] The above is only an example of the method process of sampling one meta-path. For multiple predefined meta-paths, the above method process is repeatedly executed in sequence, and a view corresponding to each meta-path can be obtained, thereby constructing a view set containing multiple different views.
[0081] S130. For each node in the heterogeneous knowledge network, determine the common representation of the node under multiple different views based on the node representations of the node in each view.
[0082] In the embodiments of the present application, the common representation of a node refers to the consistent representation learned by the node on multiple views, which reflects the consistency of the node among multiple views. The better the multi-view Figure 1 consistency of the node, the higher the accuracy of the similarity calculated between two nodes for the heterogeneous knowledge network, and the stronger the robustness and generalization ability of the model.
[0083] In some embodiments, a heterogeneous graph neural network model can be pre-constructed through deep learning. The heterogeneous graph neural network model is used to evaluate the consistency of nodes in different views, and the model parameters of the heterogeneous graph neural network model are optimized through self-supervised learning to minimize the representation differences of nodes in multiple views. After the model training is completed, at a preset stage, the node representations of each node can be input into the heterogeneous graph neural network model, and the common representations of each node in different views can be predicted through the heterogeneous graph neural network model. The following embodiments of the present disclosure will illustrate this, and will not be elaborated here for the time being.
[0084] S140. For any two nodes in the heterogeneous knowledge network, determine the similarity between the two nodes based on the common representation of each node.
[0085] In the embodiments of the present application, through the foregoing method process, the common representation corresponding to each node in the heterogeneous knowledge network is obtained. The common representation can reflect the consistency of the node in multiple different views. Therefore, when calculating the similarity between any two nodes, the similarity between the two nodes can be calculated based on the common representations of the two nodes.
[0086] For example, in an example, when calculating the similarity between node A and node B, the cosine similarity or Euclidean distance can be calculated according to the common representation of node A and the common representation of node B, and the cosine similarity or Euclidean distance is determined as the similarity between node A and node B. For the specific calculation process, the following embodiments of the present application will illustrate this, and will not be elaborated here.
[0087] As can be seen from the above, the method of the embodiments of the present application can comprehensively utilize different types of information in the heterogeneous knowledge network by constructing views with multiple different relationship patterns, capture the characteristics and relationships of nodes in different aspects, consider the association relationships between nodes more comprehensively, and improve the comprehensiveness and accuracy of similarity measurement. Moreover, by defining the common representation of nodes in different views, the influence caused by differences in node types or relationship types in the heterogeneous knowledge network can be reduced, and the robustness and accuracy of node similarity measurement can be improved.
[0088] For the convenience of further understanding and illustration, in the following of the present application, taking the academic social network scenario as an example, the method for measuring the similarity of nodes in the heterogeneous knowledge network of the present application will be described.
[0089] The academic social network is a special type of heterogeneous knowledge network, which mainly focuses on researchers, scholars, papers in the academic community and the relationships between them. The academic social network is a multi-level heterogeneous knowledge network, involving multiple types of nodes and edges, and there are complex association relationships between these types.
[0090] For example Figure 3is an example other than academic social networks, see Figure 3 As can be seen, an academic social network contains various node types and relationship types (edge types). For example, node types include author nodes, paper nodes, conference nodes, keyword nodes, etc., and relationship types include cooperation relationships, publication relationships, citation relationships, location relationships, and keyword relationships, etc.
[0091] In actual application scenarios, a series of real-world functions can be realized by performing similarity measurement on the nodes in this academic social network. For example, by measuring the similarity of authors, the system can recommend potential partners similar to author A in terms of academic interests, thus promoting academic cooperation. Another example is that in academic search, by understanding articles similar to article A, the system can provide more personalized paper recommendations and improve the information retrieval effect. The node similarity measurement means in the related art are difficult to be applied to such complex heterogeneous knowledge networks. Therefore, the similarity measurement method of this application aims to provide a comprehensive, accurate, and robust similarity measurement method for such complex heterogeneous knowledge networks, providing a reliable basis for analysis tasks based on heterogeneous knowledge networks.
[0092] First, in combination with Figure 2 According to the embodiments, since there are various node types and different forms of node attributes in the heterogeneous knowledge network, it is impossible to use a unified model for the representation learning of nodes. Therefore, in the embodiments of this application, a corresponding representation model can be constructed for each node type, and then the nodes can be uniformly mapped to a vector space using the corresponding representation model of the nodes.
[0093] In Figure 3 the example, for an author node, it can be regarded as an entity node, and its features include the author's research field, affiliated unit, cooperation method, etc. The multi-layer perceptron in Table 1 above can be used for mapping to obtain the node representation corresponding to the author node. For a paper node, its features are often text sequences such as abstracts, and the word embedding model and Transformer model in Table 1 above can be used for mapping. Figure 4 shows the model structure of the multi-layer perceptron in an example of this application, Figure 5 shows the model structures of the word embedding model and Transformer model in an example of this application. Those skilled in the art can understand and fully implement it with reference to Figure 4 and Figure 5 shown, and details are not described herein again.
[0094] In the example of this application, after processing each node in the heterogeneous knowledge network through the foregoing process, for any node in the heterogeneous knowledge network, its corresponding node representation can be obtained.
[0095] After obtaining the node representations of each node in the heterogeneous knowledge network, path sampling can be performed based on meta-paths in the heterogeneous knowledge network. The following will be described in conjunction with Figure 6 this.
[0096] As Figure 6 shown, in some embodiments, in the similarity measurement method exemplified in the present application, the process of obtaining a node relationship sequence by meta-path sampling includes:
[0097] S610. Define multiple meta-paths according to different node types and relationship types between nodes.
[0098] S620. For each meta-path, extract path instances from the heterogeneous knowledge network based on the node types and relationship types represented by the meta-path.
[0099] S630. Obtain the node relationship sequence corresponding to the meta-path according to the extracted path instances.
[0100] In the embodiments of the present application, multiple meta-paths can be predefined based on node types and relationship types. A meta-path consists of a sequence of node types and relationship types. Then, for each defined meta-path, path sampling can be performed in the heterogeneous knowledge network based on the definition of the meta-path, that is, path instances that conform to the definition of the meta-path are extracted from the heterogeneous knowledge network, and these path instances form a node relationship sequence composed of a series of nodes and edges.
[0101] In some embodiments, the formal description of meta-path sampling is as follows: Assume that the heterogeneous knowledge network contains multiple node types and edge types, where represents the node set of the heterogeneous knowledge network, represents the edge set of the heterogeneous knowledge network, and the node set can be divided into different types of nodes, and the edge set contains heterogeneous relationships connecting different types of nodes. A meta-path is a sequence of nodes, where the nodes are connected according to a predefined pattern. Formally, a meta-path can be represented as , where represents the node type, represents the edge type, is the length of the path. In one example, Figure 7 shows examples of meta-paths of different lengths that may exist in the academic social network of the present application. For example, for an academic social network, the meta-path may be .
[0102] After defining the meta - paths, meta - path sampling is to select specific meta - paths from all possible meta - paths. For example, meta - paths can be selected by random sampling, sampling based on heuristic rules, or other methods. Formally, let be the set of meta - paths, and meta - path sampling can be expressed as , where is the sampled meta - path.
[0103] After obtaining the meta - path set through sampling, path sampling can be performed on the heterogeneous knowledge network according to each meta - path in the meta - path set . Specifically, for each node, taking it as the starting point, random or regular sampling is performed from the meta - path set to obtain multiple meta - paths. These meta - paths represent different types of association relationships and will be used to construct multiple views.
[0104] For each meta - path, relevant information of the nodes is extracted, including node attributes, neighbor nodes, edge weights, etc., to construct a corresponding node view. This view can be regarded as a sub - graph extracted from the knowledge network, which contains the target node and its relevant neighbor nodes. Formally, the node view can be expressed as , where is the set of nodes on the meta - path, and is the set of edges connecting these nodes. Assuming views are constructed for each node, a multi - view set can be obtained for each node. Figure 8 shows an example of multi - view construction with author A as the target node. Each view in the figure is a path instance extracted, that is, the node relationship sequence corresponding to the meta - path.
[0105] After constructing the multi - view set for each node, an heterogeneous graph neural network model can be used to learn the corresponding common representation for each node, and self - supervised training can be performed on this heterogeneous graph neural network model based on multi - view Figure 1 consistency.
[0106] Specifically, first, an heterogeneous graph neural network model is constructed. This neural network model can handle different node types and relationship types in the heterogeneous knowledge network and optimize the network parameters through forward and backward propagation to learn the representation of each node. In one example, to train this heterogeneous graph neural network model, a multi - view Figure 1 consistency self - supervised learning task is introduced, and the network parameters are optimized by maximizing or minimizing the consistency of node representations on multiple views. For example:
[0107] Positive Consistency: For each node, its representations in different views should be consistent. This can be achieved by maximizing their similarity, and metrics such as cosine similarity and Euclidean distance can be used;
[0108] Negative Consistency: For each node, its representations should be inconsistent with those of other nodes. Therefore, corresponding negative samples can be extracted from the network, and by minimizing the similarity between the node and the negative samples, the model is encouraged to generate inconsistent node representations.
[0109] On this basis, by designing the corresponding loss function and using backpropagation to train the network parameters, the representations learned by the model for each node can meet the multi-view Figure 1 consistency requirements. This process is described as follows: Taking the node as an example of the target node, the multi-view set of this node is obtained through the aforementioned method process. Let be a heterogeneous graph neural network model with parameters , and then the node representation learned by the corresponding node in view . The corresponding training process is as follows:
[0110] (1)
[0111] In formula (1), , are two randomly sampled views of node , and is a random view of node , serving as a negative sample. The function is a custom distance function, which can be cosine similarity or Euclidean distance.
[0112] It can be understood that through the loss function shown in formula (1), the network training process of the heterogeneous graph neural network model can be realized, that is, the network parameters of the heterogeneous graph neural network model are continuously optimized until the convergence condition is met, and then the model training process can be stopped to complete the model training.
[0113] In the prediction stage, after obtaining the multi-view set of node through the aforementioned method process, two views can be randomly or regularly selected from multiple views to form a view pair, and then the node representations of the nodes in the view pair are input into the heterogeneous graph neural network model, and the common representation of the node preset to be output by the model under different views can be obtained.
[0114] After obtaining the common representations of nodes in the heterogeneous knowledge network under multiple views, for any two nodes and , the similarity can be calculated based on their common representations. In some embodiments, cosine similarity or Euclidean distance can be used to measure the similarity between nodes, expressed as:
[0115] (2)
[0116] (3)
[0117] For cosine similarity, the value is between [-1, 1], where 1 indicates complete similarity and -1 indicates complete dissimilarity. For Euclidean distance, the smaller the value, the more similar, and the larger the value, the more dissimilar.
[0118] Through the above formulas (2) and (3), the similarity between any two nodes in the heterogeneous knowledge network can be calculated, and then tasks such as recommendation can be performed based on the similarity. This application will not elaborate further on this.
[0119] As can be seen from the above, in the method of the embodiments of this application, in response to the problems of low efficiency, low accuracy, and insufficient robustness in the existing means, corresponding solutions are proposed. First, since a large amount of manual judgment is not required, this method can be applied to large-scale heterogeneous knowledge networks; second, the method of meta-path sampling introduces topological information into the learning of node representations, and the multi-view representation based on meta-paths realizes the full utilization of heterogeneous graph data, improving the accuracy of node representations; finally, the self-supervised training based on multi-view Figure 1 consistency improves the robustness of node representations. In summary, the present invention has the advantages of high efficiency, high accuracy, and strong robustness compared with the existing means.
[0120] The above-described embodiments merely represent several embodiments of this application, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application shall be subject to the appended claims.
Claims
1. A method for measuring node similarity in heterogeneous knowledge networks based on multi-view consistency, characterized in that: include: Construct node representations for different types of nodes in heterogeneous knowledge networks through node representation learning; Based on a plurality of predefined meta-paths, path sampling is performed on the heterogeneous knowledge network respectively, and a view corresponding to each meta-path is constructed; For each node in the heterogeneous knowledge network, determining a common representation of the node in multiple different views based on the node representation of the node in each view; For any two nodes in the heterogeneous knowledge network, determining the similarity between the two nodes based on the common representation of each node; Node representation learning is used to construct node representations for different types of nodes in heterogeneous knowledge networks, including: For each node in the heterogeneous knowledge network, obtaining a node type corresponding to the node; Based on the node type of the node, determining a target representation model corresponding to the node type from a preconfigured correspondence relationship, wherein the correspondence relationship represents a correspondence relationship between the node type and the representation model; Inputting the node attribute corresponding to the node into the target representation model to obtain the node representation output by the target representation model; Based on a plurality of predefined meta-paths, the heterogeneous knowledge network is sampled respectively, and a view corresponding to each meta-path is constructed, including: Define multiple meta-paths based on different node types and relationship types between nodes; For each meta-path, extracting a path instance from the heterogeneous knowledge network based on the node type and the relationship type represented by the meta-path; Obtaining a node relationship sequence corresponding to the meta-path according to the extracted path instance; A view corresponding to each meta-path is constructed based on the node relationship sequence.
2. The method according to claim 1, characterized in that The node type includes one or more of image, text, and entity; The representation model includes one or more of a visual Transformer model, a word embedding model, and a multi-layer perceptron.
3. The method according to claim 1, characterized in that For each node in the heterogeneous knowledge network, based on the node representation of the node in each view, determining a common representation of the node in multiple different views includes: For each node in the heterogeneous knowledge network, two views are randomly selected from multiple views to form a view pair; The node representations of the nodes in the view pair are input into a pre-trained heterogeneous graph neural network model to obtain the common representation output by the heterogeneous graph neural network model.
4. The method according to claim 1, characterized in that For any two nodes in the heterogeneous knowledge network, determining the similarity between the two nodes based on the common representation of each node includes: For any two nodes in the heterogeneous knowledge network, the cosine similarity or Euclidean distance between the two nodes is calculated based on the common representation of each node to obtain the similarity.
Citation Information
Patent Citations
Relation perception news recommendation method, system and equipment based on news heterogeneous network
CN115422470A
Multi-view fusion heterogeneous node representation learning method based on large model technology
CN118296554A