Knowledge graph representation method and device based on similarity diffusion
By using a similarity diffusion-based approach, a knowledge graph is constructed using a multimodal large model and a human-in-the-loop mechanism. Similarity diffusion and manifold geometric transmission under local clustering constraints are performed, which solves the problem of insufficient mining of entity relationships and improves the mining, fusion and retrieval capabilities of the knowledge graph.
Patent Information
- Application Number
- CN202411318514.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-09-20
AI Technical Summary
The broader potential relationships between entities in existing technologies have not been fully explored, resulting in insufficient mining, fusion, reasoning, and retrieval capabilities of knowledge graphs.
A similarity-based diffusion approach is adopted to construct a knowledge graph through a multimodal large model and a human-in-the-loop mechanism. Similarity diffusion under local clustering constraints is performed using an adjacency graph matrix, and geometric transmission is carried out in the manifold graph to mine potential relationships between entities.
It improves the mining, fusion, reasoning, and retrieval capabilities of knowledge graphs, resulting in more effective and robust relation representations, and guiding the construction and updating of knowledge graphs.
Smart Images

Figure CN119443223B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of knowledge graph, and particularly relates to a knowledge graph representation method and device based on similarity diffusion. BACKGROUND
[0002] A knowledge graph is a structured semantic knowledge base, which organizes knowledge entities and their relationships in the form of a graph, and can effectively represent complex knowledge in an information management system. Specifically, each entity in a knowledge graph network is represented as a node, and various connections between entities are represented by edges, so that the graph can systematically store rich entity information and their mutual relationships, greatly improving the accessibility and understandability of information.
[0003] The construction of a knowledge graph usually relies on a large amount of data mining and information extraction techniques, including entity recognition, relationship extraction, and knowledge fusion steps. Although there are many modeling methods for knowledge graphs at present, the more extensive potential relationships between entities have not been fully mined, so the mining, fusion, reasoning, and retrieval capabilities of the knowledge graph still need to be improved. SUMMARY
[0004] The present application provides a knowledge graph representation method and device based on similarity diffusion, which solves the defect that the more extensive potential relationships between entities have not been fully mined in the prior art, and can better guide the construction and update of the knowledge graph, and enhance its data mining, fusion, reasoning, and retrieval capabilities.
[0005] The present application provides a knowledge graph representation method based on similarity diffusion, comprising the following steps:
[0006] Constructing a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism;
[0007] Based on the similarity relationship between each entity in the target knowledge graph, constructing an adjacency graph matrix of the target knowledge graph;
[0008] Based on the adjacency graph matrix, performing similarity diffusion under local clustering constraints to obtain a first similarity matrix;
[0009] Based on the first similarity matrix, performing geometric transmission of the target knowledge graph in a manifold graph to obtain a relationship representation of the target knowledge graph.
[0010] According to the knowledge graph representation method based on similarity diffusion provided by the present application, the adjacency graph matrix of the target knowledge graph is constructed based on the similarity relationship between each entity in the target knowledge graph, specifically comprising:
[0011] obtain an initial distance matrix of the entities based on the Euclidean distances between the entities in the target knowledge graph;
[0012] construct an adjacency graph matrix of the target knowledge graph based on the initial distance matrix and the nearest neighbor edges of the entities.
[0013] According to the knowledge graph representation method based on similarity diffusion provided by the application, the relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in a manifold graph based on the first similarity matrix, and specifically includes:
[0014] The first similarity matrix is subjected to neighborhood-guided smoothing enhancement based on the neighbor nodes of the initial distance matrix, and a second similarity matrix is obtained;
[0015] The relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in a manifold graph based on the second similarity matrix.
[0016] According to the knowledge graph representation method based on similarity diffusion provided by the application, the first similarity matrix is obtained by performing similarity diffusion under local clustering constraints based on the adjacency graph matrix, and specifically includes:
[0017] The local clustering graph of each entity is constructed based on the adjacency graph matrix and a k- mutual neighbor strategy;
[0018] The first similarity matrix is obtained by performing diffusion process on the local clustering graph of each entity through a bidirectional propagation mechanism.
[0019] According to the knowledge graph representation method based on similarity diffusion provided by the application, the relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in a manifold graph based on the second similarity matrix, and specifically includes:
[0020] The third similarity matrix is obtained by performing information propagation on the second similarity matrix;
[0021] The relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in a manifold graph based on the third similarity matrix.
[0022] According to the knowledge graph representation method based on similarity diffusion provided by the application, the relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in a manifold graph based on the first similarity matrix, the second similarity matrix or the third similarity matrix, and specifically includes:
[0023] The geodesic distance matrix containing manifold space information is obtained based on the initial distance matrix and the adjacency graph matrix.
[0024] based on the geodesic distance matrix, an optimal transport matrix between distribution representations of each entity in the target knowledge graph in manifold space is calculated; wherein the distribution representations of each entity in the target knowledge graph in manifold space are obtained based on the first similarity matrix, the second similarity matrix or the third similarity matrix;
[0025] based on the optimal transport matrix and the distribution representations of each entity in manifold space, geometric transport distances between each entity are determined;
[0026] based on the geometric transport distances, a relationship representation of the target knowledge graph is obtained.
[0027] The application also provides a knowledge graph representation method based on similarity diffusion.
[0028] The construction module is configured to construct a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism.
[0029] The adjacency module is configured to construct an adjacency graph matrix of the target knowledge graph based on similarity relationships between each entity in the target knowledge graph.
[0030] The diffusion module is configured to perform similarity diffusion under local clustering constraints based on the adjacency graph matrix to obtain a first similarity matrix.
[0031] The representation module is configured to perform geometric transport on the target knowledge graph in a manifold graph based on the first similarity matrix to obtain a relationship representation of the target knowledge graph.
[0032] The application also provides an electronic device including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the knowledge graph representation method based on similarity diffusion.
[0033] The application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the knowledge graph representation method based on similarity diffusion.
[0034] The application also provides a computer program product including a computer program, wherein the computer program is executable by a processor to implement the knowledge graph representation method based on similarity diffusion.
[0035] The application provides a knowledge graph representation method and device based on similarity diffusion, which constructs a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism, constructs an adjacency graph matrix of the target knowledge graph according to the similarity relationship between each entity in the target knowledge graph, then performs similarity diffusion under local clustering constraints based on the adjacency graph matrix, can mine more extensive potential relationships between entities, obtains a first similarity matrix, and performs geometric transmission on the target knowledge graph in a manifold graph according to the first similarity matrix, so as to obtain a more effective and robust relationship representation of the target knowledge graph. Through this relationship representation, the construction and updating of the knowledge graph can be guided, and the mining, fusion, reasoning and retrieval capabilities of the knowledge graph can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0037] Figure 1 The flowchart of the knowledge graph representation method based on similarity diffusion provided by the present application.
[0038] Figure 2 The flowchart of obtaining more robust relationships between knowledge graph entities based on similarity diffusion provided by the present application.
[0039] Figure 3 The structural diagram of the knowledge graph representation device based on similarity diffusion provided by the present application.
[0040] Figure 4 The structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor belong to the scope of protection of the present application.
[0042] Figure 1 The flowchart of the knowledge graph representation method based on similarity diffusion provided by the present application is shown in FIG. 1, which includes the following steps: Figure 1
[0043] Step 100, based on the multi-modal large model and the human-in-the-loop mechanism, constructing the target knowledge graph.
[0044] Specifically, the target knowledge graph in the embodiment of the present application is the knowledge graph to be constructed by the user.
[0045] The multi-modal large model refers to an artificial intelligence model that can process and understand multiple types of data, such as text, images, audio, video, etc. The design of these models aims to integrate information from different modalities to facilitate more comprehensive and accurate task processing. The human-in-the-loop mechanism refers to the interaction and feedback mechanism between humans and automated systems in certain systems.
[0046] Therefore, the text description information of the image can be first extracted using the human-in-the-loop prompt learning technology combined with the multi-modal large model, and the method of combining the neural network-based named entity recognition and relationship classification model can be used to realize the extraction of triple information from the text description of all database images, so as to construct the target knowledge graph.
[0047] In the process of constructing the target knowledge graph using the multi-modal large model, the text description of the image library can be automatically extracted using the multi-modal large model, and the attribute knowledge graph can be constructed based on these descriptions, which requires obtaining the entity, relationship, and attribute information of the target knowledge graph.
[0048] The human-in-the-loop algorithm is to manually filter and extract images, corpora, and triple information based on intelligent algorithms. In the field of knowledge extraction, the human-in-the-loop mechanism can retain the high generalization and robustness of the multi-modal large model, and integrate human prior knowledge to improve the reliability of the knowledge graph knowledge base.
[0049] Step 101, based on the similarity relationship between each entity in the target knowledge graph, constructing the adjacency graph matrix of the target knowledge graph.
[0050] Specifically, after constructing the target knowledge graph, for each entity in the target knowledge graph, the entity set can be defined as , where is the number of entities in the target knowledge graph, and for each entity in the set , a knowledge representation can be constructed for it, using the powerful expression ability of the deep learning network to encode it into a dimensional feature , which is used in the subsequent knowledge graph mining, fusion, reasoning, and retrieval process.
[0051] The similarity between any pair of entities and in the set The similarity between entities can be calculated by, for example, Euclidean distance, which is used to measure the distance between two entities in the feature space, and a smaller distance indicates a higher similarity; or, for example, cosine similarity, which is used to represent the similarity between entity features by calculating the cosine value of the angle between two vectors, and a value closer to 1 indicates a higher similarity.
[0052] After calculating the similarity between each entity, the adjacency graph matrix of the target knowledge graph can be constructed according to the similarity calculation results. The adjacency graph can be represented as In the adjacency graph, the vertex set corresponds to the feature encoding of each entity in , and the edge set represents the connection relationship between entity vertices. The edge weight of the adjacency graph can be determined according to the similarity calculation results between each entity, and then the weight matrix of the adjacency graph is used as the adjacency graph matrix.
[0053] According to the knowledge graph representation method based on similarity diffusion provided by the present application, the adjacency graph matrix of the target knowledge graph is constructed based on the similarity relationship between each entity in the target knowledge graph, and specifically includes:
[0054] Based on the Euclidean distance between each entity in the target knowledge graph, an initial distance matrix of each entity is obtained.
[0055] Based on the initial distance matrix and the nearest neighbor edge of each entity, the adjacency graph matrix of the target knowledge graph is constructed.
[0056] Specifically, in the feature encoding space of the target knowledge graph, the similarity between any pair of entities and can be calculated by the Euclidean distance of their feature vectors, and the calculation formula is as follows:
[0057]
[0058] After the distance between each pair of entities is calculated, an initial distance matrix can be obtained, where the matrix element is defined as . Each row of the matrix represents the distance between entity and the data set , and can be used to sort the data set.
[0059] Based on the prior knowledge that similar entities tend to be clustered in the low-dimensional manifold structure implied by the distance matrix , the present embodiment can cluster The mapping to a new space enhances the expressiveness of this structure. This mapping ensures that entities with similar semantics are not only close to each other in the traditional Euclidean space, but also exhibit higher similarity in the manifold space.
[0060] To approximate this manifold structure, embodiments of the present application can employ a k-nearest neighbor (k-NN) method to select the k nearest neighbors of each node and assign weights to the edges between them according to the similarity, thereby constructing a k-NN graph. In this graph, the vertex set corresponds to the feature encoding of each entity in , and the edge set represents the connection relationship between entity vertices, each edge connects vertices and , and the edge weight is the similarity between entities and , and the weight matrix can be represented as:
[0061]
[0062] where the matrix is an indicator matrix used to connect the k-NN edges, formally, when , , , represents the set of the k nodes most similar to node . is a parameter that controls the decay of similarity, so that the similarity between entities with smaller distances is higher, and the weight is larger. The constructed adjacency graph can be represented by a weight matrix
[0063] , i.e., an adjacency graph matrix. Each row of the adjacency graph matrix represents the similarity relationship between an entity and all other entities, constituting the basic connection structure in the knowledge graph. Step 102, based on the adjacency graph matrix, similarity diffusion is performed under local clustering constraints to obtain a first similarity matrix.
[0064] Specifically, after obtaining the adjacency graph matrix, similarity information can be diffused under the constraint of local clustering to obtain a more robust similarity matrix, avoiding the influence of noise and outliers on the construction and updating of the overall graph.
[0065]
[0066] Local clustering refers to defining a local group containing similar neighbors for each entity. Thus, local clustering of each entity can be determined first, and the diffusion process is performed only among the entities in the local clustering, so that the similarity is smoothly transferred within the local group.
[0067] According to the knowledge graph representation method based on similarity diffusion provided by the application, the similarity diffusion is performed under the local clustering constraint based on the adjacency matrix, and a first similarity matrix is obtained, and specifically includes:
[0068] Based on the adjacency matrix and the k-reciprocal neighbors strategy, a local clustering graph of each entity is constructed.
[0069] The local clustering graph of each entity is diffused through a bidirectional propagation mechanism to obtain a first similarity matrix.
[0070] Specifically, the most trusted neighbors of each entity in the adjacency matrix can be found through the k-reciprocal neighbors strategy first. This strategy ensures the bidirectional relationship between entities, that is, if entity is a neighbor of entity , and entity is also a neighbor of entity , then the similarity between them is more reliable. This process avoids the negative impact of untrusted neighbors and constructs a local clustering set with higher confidence .
[0071] The definition of the k-reciprocal neighbors is as follows:
[0072]
[0073] Specifically, given an entity , its -neighbor set is defined as , and for any element in this set, only when entity is also in the -neighbor set of the element, the element belongs to the - reciprocal neighbor set of entity , and this set can be regarded as the local clustering of entity .
[0074] After obtaining the local clustering graph of each entity, the local clustering graph of each entity is diffused through a bidirectional propagation mechanism to obtain a first similarity matrix.
[0075] Traditional similarity propagation methods usually propagate similarity in the whole graph, which is easily disturbed by outliers and noisy samples. To this end, a bidirectional similarity propagation under local clustering constraints is proposed, which only propagates similarity in local graphs (i.e. local clustering sets ). The goal of the propagation process is to achieve smooth propagation of similarity in local clustering, and ensure the symmetry of the similarity matrix through the bidirectional propagation mechanism.
[0076] In the embodiments of the present application, a target function is defined, which aims to minimize the difference between the similarity matrix and the local graph structure, and at the same time, the local similarity between entities is considered, and the definition is as follows:
[0077]
[0078]
[0079] Wherein, is a smoothness control parameter. represents an adjacency matrix, for the sake of simplicity, a symmetric matrix may be used instead of the original adjacency matrix in construction. The matrix is a diagonal matrix, and the first diagonal element is equal to the sum of the first row of In the regularization term guided by the weight , the matrix is positive definite, which can be used to prevent from being too smooth, and different strategies can be taken to initialize according to requirements. For the entity triplets and in the local clustering, the adjacency weight is used to constrain the similarity between and . At the same time, the similarity of the pair and is also considered, so that the smoothing strategy is bidirectional, and the obtained similarity matrix has symmetry.
[0080] For the target function with the constraint in the foregoing, it is very difficult to directly solve the optimization target, therefore, in some embodiments, in order to simplify the optimization problem, the complex target function can be decomposed into two more easily solvable sub-problems. First, for bidirectional propagation without clustering constraints, the optimization target is:
[0081]
[0082] Then, only the entities belonging to the corresponding clustering The similarity within the cluster is used as an effective approximation of performing similarity diffusion within the local cluster. To more conveniently solve the optimization problem in the above equation, this embodiment of the invention assumes the optimization objective is... Using mathematical tools from graph theory, the objective expression can be reconstructed as follows:
[0083]
[0084] in It is an identity matrix. The regularization matrix is represented as ,matrix It is about The mean Kronecker product is defined as , operator It is a vectorization function that can concatenate the column vectors of a matrix to form a new column vector, and That is The inverse function of . It can be proven that the objective function after relaxing the constraints is . It is a convex function, meaning its Hessian matrix is positive definite. This can be determined by... Taking the partial derivative, we can obtain:
[0085]
[0086] Setting the value of this expression to zero allows us to obtain a closed-form solution. as follows:
[0087]
[0088] This can be used Replace hyperparameters In simplified form, without any... A similar expression can be derived for the more general case when performing symmetry transformation. Inverting the matrix in a closed-form solution is a very time-consuming operation. To find the optimal solution more quickly within a finite amount of time, the problem can be transformed into solving the following Lyapunov equation:
[0089]
[0090] This equation is equivalent to the graph Considered as a The asymptotically stable linear system is defined, where Representing the state of each vertex, by solving... The induced Lyapunov energy can reveal manifold information contained in the adjacency graph. This equation can be solved numerically, and through iterative approximation, the computational time complexity is reduced to a minimum. Thus, the whole process is more feasible and efficient, and the approximate optimal solution is obtained by the following iterative process:
[0091]
[0092] After the iteration converges, the similarity matrix with relaxed constraints is obtained . Next, the similarity between each point and its local cluster is optionally retained, thereby approximately obtaining the similarity matrix under the local cluster constraint . Specifically, for an entity , only the items in that belong to are retained to construct . Subsequently, the normalization operation is applied to each row of to obtain the final smoothed similarity matrix (i.e., the first similarity matrix).
[0093] Step 103, based on the first similarity matrix, performing geometric transmission on the target knowledge graph in the manifold graph to obtain a relationship representation of the target knowledge graph.
[0094] Specifically, in the embodiments of the present application, after obtaining the first similarity matrix, the target knowledge graph can be geometrically transmitted in the manifold space based on the first similarity matrix to more accurately represent the similarity relationship between each entity of the target knowledge graph. By considering the geometric characteristics of the data in the manifold space, the robustness of the entity relationship in the knowledge graph can be ensured.
[0095] The knowledge graph representation method based on similarity diffusion provided by the present application constructs a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism, constructs an adjacency graph matrix of the target knowledge graph according to the similarity relationship between each entity in the target knowledge graph, then performs similarity diffusion under local cluster constraints based on the adjacency graph matrix, which can mine more extensive potential relationships between entities, obtain a first similarity matrix, and perform geometric transmission on the target knowledge graph in a manifold graph according to the first similarity matrix, thereby obtaining a more effective and robust relationship representation of the target knowledge graph. Through this relationship representation, the construction and updating of the knowledge graph can be guided, and the mining, fusion, reasoning, and retrieval capabilities of the knowledge graph can be improved.
[0096] According to the knowledge graph representation method based on similarity diffusion provided by the present application, the target knowledge graph is geometrically transmitted in the manifold graph based on the first similarity matrix to obtain a relationship representation of the target knowledge graph, which specifically includes:
[0097] The first similarity matrix is smoothed and enhanced by neighborhood guidance based on the neighbor nodes of the initial distance matrix, to obtain a second similarity matrix;
[0098] The target knowledge graph is geometrically transmitted in the manifold graph based on the second similarity matrix, to obtain a relationship representation of the target knowledge graph.
[0099] Specifically, local neighbors can provide more accurate estimation of inter-class relationship than single nodes, and in order to further enhance the expression ability of the similarity matrix, the similarity consistency between entities and neighborhoods can be constrained by using heuristic information from the local neighbor set , so as to further smooth and enhance the similarity matrix.
[0100] Therefore, after obtaining the first similarity matrix, the first similarity matrix can be first smoothed and enhanced by neighborhood guidance based on the neighbor nodes found from the initial distance matrix, to obtain a second similarity matrix, and then the target knowledge graph is geometrically transmitted in the manifold graph based on the second similarity matrix, to obtain a relationship representation of the target knowledge graph.
[0101] In the embodiment of the application, the process can be represented as first smoothing the first similarity matrix by using the neighbor nodes to obtain a smoothed neighborhood-consistent matrix , and then obtaining an enhanced representation thereof, i.e., a second similarity matrix .
[0102] In the embodiment of the application, the local neighbors can be determined by - a mutual neighbor strategy, and specifically, for an entity , the neighborhood of the entity is defined as , where . In the neighborhood space of the entity , the embodiment of the application defines as the average similarity of the local neighbors of the entity to the entity , and defines the average similarity between the paired entities in the local neighborhood as , where a larger means that the neighbor is more likely to belong to the same cluster, reflecting the reliability of the local neighbor, and the similarity of the entity to the entity should be consistent with the factor . After assigning the weight and adding a regularization term, the optimization target can be expressed as:
[0103]
[0104]
[0105]
[0106]
[0107] where the operator multiplies the corresponding elements of two vectors to get a new vector, is a local clustering set used to constrain the similarity diffusion, and the weight is The regular term can be used to constrain the value of the target from deviating too much from the initial value, and a relatively low value can be set to obtain better numerical stability. In addition, when constructing the vector , the embodiment of the present application artificially ensures that all its values are not greater than by a truncation operation.
[0108] After solving the KKT conditions (Karush-Kuhn-Tucker Conditions) of the optimization problem, the elements of the matrix consistent in the neighborhood are as follows:
[0109]
[0110] In addition to increasing consistency, local neighbors can further enhance the representation of the similarity matrix in other ways. For example, when constructing the nearest neighbor graph, higher weights can be assigned to edges related to , thereby emphasizing the role of local neighbors in the diffusion process. Formally, when , the term of the indicator matrix becomes , where the hyperparameter is the importance weight. In addition, since entities within the same neighborhood are likely to belong to the same category, the embodiment of the present application can aggregate entities within the neighborhood to obtain a neighborhood-enhanced representation , which can be implemented by weighted averaging the manifold distribution of entities in
[0111]
[0112] where is used to enhance information from neighborhood information , which can balance the relative weights of neighborhood-guided similarity distribution and original similarity distribution, and is used for normalization, thereby ensuring that each row of the target similarity matrix is Normalized. By adjusting the hyperparameters , the influence of the neighborhood and the nearest neighbors on the final similarity matrix can be controlled, ensuring that certain global smoothness is maintained on the basis of local enhancement.
[0113] By combining the similarity information of the node with the similarity distribution of its neighbors and the nearest neighbors , the local similarity consistency is enhanced. The second similarity matrix obtained in this way can better represent the relationship between the node and the surrounding nodes, thereby improving the accuracy of the similarity matrix.
[0114] According to the knowledge graph representation method based on similarity diffusion provided by the application, based on the second similarity matrix, the target knowledge graph is geometrically transmitted in the manifold graph to obtain the relationship representation of the target knowledge graph, specifically comprising:
[0115] The second similarity matrix is propagated to obtain a third similarity matrix;
[0116] Based on the third similarity matrix, the target knowledge graph is geometrically transmitted in the manifold graph to obtain the relationship representation of the target knowledge graph.
[0117] Specifically, after completing the local similarity enhancement to obtain the second similarity matrix, if it is necessary to better improve the global similarity, an additional step of global information propagation can be performed on the entire graph to further smooth the similarity matrix of the entire graph.
[0118] The second similarity matrix can be regarded as the best representation of the data manifold in local clustering, and in order to need better global expression ability, the embodiment of the application can perform an additional information propagation on the entire graph to obtain a third similarity matrix , the third similarity matrix can be obtained by multiplying the transition matrix calculated by transposing the second similarity matrix with itself:
[0119]
[0120] In this process, the embodiment can also combine the sparsification operation to process the transition matrix , filter out the too low similarity value, and retain the most significant similarity relationship to obtain the third similarity matrix , thereby avoiding noise interference and better reducing the influence of outliers.
[0121] According to the knowledge graph representation method based on similarity diffusion provided by the application, the target knowledge graph is geometrically transmitted in a manifold graph based on a first similarity matrix, a second similarity matrix or a third similarity matrix, and a relationship representation of the target knowledge graph is obtained, and the method specifically comprises the following steps:
[0122] Based on the initial distance matrix and the adjacency graph matrix, a geodesic distance matrix containing manifold space information is obtained.
[0123] Based on the geodesic distance matrix, an optimal transport matrix between distribution representations of each entity in the target knowledge graph in the manifold space is calculated, wherein the distribution representations of each entity in the target knowledge graph in the manifold space are obtained based on the first similarity matrix, the second similarity matrix or the third similarity matrix.
[0124] Based on the optimal transport matrix and the distribution representations of each entity in the manifold space, geometric transmission distances between each entity are determined.
[0125] Based on the geometric transmission distances, the relationship representation of the target knowledge graph is obtained.
[0126] Specifically, the traditional similarity metric (such as Euclidean distance) cannot fully express the complex relationship between entities, especially when the data is located in a manifold structure. The manifold-aware geometric transmission method based on Wasserstein distance or optimal transport theory is used to find a more suitable similarity metric in the manifold space. This method not only considers the direct distance between entities, but also fully utilizes the structure information of the manifold space to calculate the optimal transmission path.
[0127] Firstly, the geodesic distance matrix containing manifold space information can be obtained according to the initial distance matrix and the adjacency graph matrix (the constructed nearest neighbor graph).
[0128] In this embodiment, the target is modeled as a transmission of each entity in the manifold space . The first similarity matrix, the second similarity matrix or the third similarity matrix can be used, so that the calculated Wasserstein distance can fully consider the manifold information of the whole space. In order to simplify the process of optimal transport, the application embodiment defines the transmission path as the shortest path between two nodes in the graph. For the nodes and in the manifold graph , the shortest path between them is:
[0129]
[0130] where belongs to the set Each path in the graph is composed of a sequence of nodes, as defined in , which starts at and ends at . The edge between nodes and is defined as . The geodesic between and on the manifold graph is equivalent to the shortest distance between the two points, and this path contains the overall information in the data manifold space. The geodesic distance between and is calculated as follows:
[0131]
[0132] The distance is calculated by summing the weights of the edges on the geodesic path, and the geodesic distances between all pairs of entity nodes form the geodesic matrix . In graph theory, calculating the shortest distance between all pairs of nodes on a graph is a mature problem that has been discussed for decades, and many solutions can be found to efficiently calculate the geodesic distance matrix .
[0133] After obtaining the geodesic distance matrix, the optimal transport matrix between the distribution representations of each entity in the target knowledge graph in the manifold space can be calculated according to it.
[0134] The geodesic distance matrix can be used as the cost of optimal transport, and the optimal transport matrix between two manifold distributions can be obtained by solving the following optimization objective:
[0135]
[0136]
[0137] where the matrix represents the optimal transport matrix from the distribution to the distribution in the manifold space, is the distribution representation of entity in the manifold space, is the distribution representation of entity in the manifold space. To more efficiently calculate the optimal distance, an entropy regularization term is added to the optimization objective, so that this optimal transport problem can be solved iteratively using the Sinkhorn-Knopp algorithm.
[0138] After the optimal transport matrix is obtained, the geometric transport distance between the distribution representations of the entities in the manifold space can be determined according to the optimal transport matrix and the Euclidean distance between the entities.
[0139] First, after the optimal transport matrix from to is obtained, the distance between them in the manifold graph can be calculated by the following formula:
[0140]
[0141] In order to preserve the important neighbor relationship in the original distance space, the initial Euclidean distance is also taken into account in the embodiment of the application, so that the geometric transport distance between the two distribution representations finally obtained can be expressed as:
[0142]
[0143] In the formula, the initial Euclidean distance is , and the optimal transport distance on the manifold is , and the two distances are balanced by a hyperparameter .
[0144] Finally, the relationship representation of the target knowledge graph can be obtained based on the geometric transport distance. The geometric transport distance obtained in this way can more highly measure the similarity relationship between entities, so as to be used to guide the fusion construction of the knowledge graph and improve the retrieval and reasoning efficiency of the knowledge graph.
[0145] The similarity diffusion-based knowledge graph representation method provided by the application is further explained through the embodiments in specific application scenarios.
[0146] Figure 2 The process schematic diagram for obtaining more robust relationships between entities of a knowledge graph based on similarity diffusion provided by the application is shown in FIG. 1. As shown in the figure, the embodiment first constructs a knowledge graph based on a multi-modal large model using a human-in-the-loop mechanism, then constructs an adjacency graph matrix using the similarity relationship between entities, and then performs similarity diffusion under local clustering constraints. Next, the distribution construction is enhanced using neighborhood information, and finally a more effective relationship representation for modeling the knowledge graph is obtained with the help of a manifold-aware geometric transport module. Figure 2
[0147] The steps are described in detail as follows.
[0148] I. Knowledge graph construction based on human-in-the-loop
[0149] The purpose of human-in-the-loop (HIL) knowledge graph construction is to automatically extract text descriptions from an image database using a multimodal large-scale model, and then construct an attribute knowledge graph based on these descriptions. Taking a pedestrian description knowledge graph as an example, its core is to obtain entity, relation, and attribute information related to pedestrian descriptions. First, HIL-based cue learning technology combined with a multimodal large-scale model is used to extract text description information from images. Then, a method combining neural network-based named entity recognition and relation classification models is employed to extract triple information from text descriptions in all database images. It should be understood that the HIL-based algorithm manually filters and uniformly extracts image, corpus, and triple information based on intelligent algorithms. In the field of knowledge extraction, the HIL mechanism can enhance the reliability of the knowledge graph knowledge base by integrating prior human knowledge while preserving the high generalization and robustness of the multimodal large-scale model.
[0150] To ensure targeted descriptions, a human-in-the-loop mechanism can be used to generate biased descriptions of pedestrian images using a manually designed prompt template ('the person is wearing [xxxx]'). Leveraging the text extraction capabilities of multimodal large models, algorithms using manually designed prompt templates can extract relevant descriptive information from images. During this process, since images often contain a large number of noisy or low-quality images, the image library can be manually filtered to remove low-quality or information-poor images, retaining only high-quality, information-rich images. Then, a multimodal large model, such as a Large Language Model (LLM), can be used to generate corresponding text descriptions for each image.
[0151] After generating text descriptions for all images, a knowledge graph semantic network for pedestrian descriptions can be constructed based on entity, relation, and attribute information networks, forming triples of "entity-relationship-entity" and "entity-attribute-attribute value" categories. Specifically, an intelligent word segmentation algorithm can be used to extract entities and relations from each description using the "entity-relationship-entity" pattern. Given the text description of an image, the word segmentation algorithm performs structured analysis on the content of the attribute description corpus, decomposing sentence components by combining the part-of-speech and semantics of each word or phrase, and identifying proper nouns, colors, entities, and other related terms in the sentence that relate to the pedestrian attribute description.
[0152] Since the generated "entity" or "relationship" contains a large number of invalid or low-quality meaningless phrases (for example, adverbs, auxiliary words, pronouns), some irrelevant or low-quality "entity" and "relationship" phrases can be manually further screened out to further improve the accuracy of the knowledge graph corpus. It should be understood that the attribute text description of the pedestrian attribute knowledge graph is converted in form, processed by word segmentation, and invalid words are deleted to construct a relatively clean pedestrian attribute knowledge graph semantic library. Through the above operations, the triple knowledge of the text description library containing "entity-relation-entity" data can be finally generated, and the knowledge graph is constructed. Therefore, the embodiment as a whole constructs a pedestrian attribute knowledge graph containing 112323 entities and 2877187 relationships. An "entity-relation-entity" mode relationship in the knowledge graph is as follows: 'woman'- 'is wearing'- 'black jacket with jeans'.
[0153] II. Adjacency relationship graph construction in knowledge graph.
[0154] The definition of the entity set can be wherein is the number of entities in the knowledge graph, and for each entity in the set , a knowledge representation can be constructed for it. By using the powerful expression ability of the deep learning network, it can be encoded into a dimensional feature , which is used in the subsequent knowledge graph mining, fusion, reasoning and retrieval process. In this feature encoding space, the similarity between any pair of entities and can be calculated by the Euclidean distance of their feature vectors as follows:
[0155]
[0156] After calculating the distance between each pair of entities, an initial distance matrix is obtained, wherein the matrix element is defined as . Each row of the matrix represents the distance between the entity and the data set , which can be used to sort the data set. Based on the prior knowledge that similar entities tend to be clustered in the low-dimensional manifold structure implied by the distance matrix , the embodiment can use the manifold learning method to reduce the dimension of the distance matrix This structure is enhanced by mapping to a new space. This mapping ensures that entities with similar knowledge semantics are not only close to each other in the traditional Euclidean space, but also exhibit higher similarity in the manifold space. To approximate this manifold structure, a feature set can be used to construct a... Nearest neighbor graph In this graph, the vertex set correspond The feature encoding of each entity, and the edge set Represents the connection relationship between entity vertices, each edge Connect vertices and Its weight is expressed as:
[0157]
[0158] Where the matrix This is an indicator matrix used for connection. Nearest neighbor, formally, when hour, .
[0159] III. Bidirectional similarity diffusion under local clustering constraints.
[0160] After constructing the affinity weight matrix as follows of Nearest neighbor graph Subsequently, in order to obtain more robust similarity relationships by utilizing the manifold information contained in the nearest neighbor graph, this embodiment employs a diffusion-based method to transmit information in the graph, thereby obtaining a similarity matrix. Traditional diffusion processes propagate similarity across the entire image, resulting in a similarity matrix. It is susceptible to the negative effects of outliers and nearby manifolds, leading to a decrease in the obtained similarity matrix. Insufficient expressive ability.
[0161] To address this issue, this embodiment proposes bidirectional similarity diffusion under local clustering constraints, aiming to perform similarity propagation on the local graph rather than the entire graph. This approach reduces the propagation of misleading information from other noisy samples, thereby improving the accuracy of similarity inference. For each entity, its local graph can be viewed as a local cluster containing the entity itself and its similar neighbors. For example, entities Local clustering is represented as To obtain a smooth similarity matrix This embodiment defines an objective function that aims to minimize the difference between the similarity matrix and the local graph structure, while also taking into account the local similarity between entities. Its definition is as follows:
[0162]
[0163]
[0164] in The weight matrix representing the adjacency graph can be constructed using a symmetric matrix for simplicity. Replace the original adjacency matrix. Furthermore, the matrix... It is a diagonal matrix, its first... The diagonal elements equal to The The sum of the rows, in weight In the guided regularization terms, the matrix It is positive definite, and it can be used to prevent Too smooth; different strategies can be adopted depending on the needs. Initialization is performed. For entity triples in local clusters... and Use adjacency weight To constrain and The similarities between them were also considered. and The similarity between these pairs makes the smoothing strategy bidirectional, ensuring that the obtained similarity matrix is consistent. It has symmetry.
[0165] In addition, in order to construct local clusters This embodiment utilizes The nearest neighbor strategy is used to find the neighbor node with the highest confidence. This method takes into account reverse information and can be regarded as... - A more stringent case of nearest neighbors, which constructs local clusters that can effectively exclude erroneous neighbors, thereby minimizing their negative impact on the final diffusion process. -The formal definition of mutual nearest neighbors is as follows:
[0166]
[0167] Specifically, given an entity Its - The nearest neighbor set is defined as For any element in this set, only if the entity Also in this element - An element only belongs to the nearest neighbor set if it is within the nearest neighbor set. of - sets of nearest neighbors This set can be considered as a local cluster. To avoid ambiguity and reduce the number of parameters, the embodiment sets the value of the number of clusters to be built as In addition, when the number of elements in the set reaches a certain threshold, the entity - the nearest neighbor reciprocal expansion to a larger set to improve the expressive power of the model.
[0168] However, for the objective function with the preceding constraints, it is very difficult to directly solve the optimization objective. In the bidirectional similarity diffusion process under the local clustering constraint proposed in the embodiment, the optimization problem is simplified into two easily solvable sub-problems by relaxing the constraint condition. First, consider the bidirectional diffusion process without clustering constraints, the objective function is as follows:
[0169]
[0170] Then, only the similarity belonging to the corresponding cluster is retained, which is used as an effective approximation for performing similarity diffusion within the local cluster. In order to more conveniently solve the optimization problem in the above formula, the embodiment assumes that the optimization objective is , and using the mathematical tools in graph theory, the objective formula can be reconstructed as:
[0171]
[0172] where is an identity matrix, the regularization matrix of , the matrix is the mean Kronecker product about , defined as , the operator is a vectorization function that can concatenate the column vectors of a matrix to form a new column vector, and is the inverse function of . It can be proved that the objective function after relaxing the constraint is a convex function, that is, its Hessian matrix is positive definite, and by taking the partial derivative of , we can get:
[0173]
[0174] Setting the value of this expression to zero, the closed-form solution can be obtained as follows:
[0175]
[0176] Here the function Replace hyperparameters In simplified form, without any... A similar expression can be derived for the more general case when performing symmetry transformation. Inverting the matrix in a closed-form solution is a very time-consuming operation. To find the optimal solution more quickly within a finite amount of time, the problem can be transformed into solving the following Lyapunov equation:
[0177]
[0178] This equation is equivalent to the graph Considered as a The asymptotically stable linear system is defined, where Representing the state of each vertex, by solving... The induced Lyapunov energy can reveal manifold information contained in the adjacency graph. This equation can be solved numerically, and through iterative approximation, the computational time complexity is reduced to a minimum. This makes the entire process more feasible and efficient. The iterative process for the approximate optimal solution is as follows:
[0179]
[0180] After iterative convergence, the similarity matrix of the relaxed constraints can be obtained. Next, each point is selectively preserved along with its local cluster. The similarity between them is used to approximate the similarity matrix under local clustering constraints. Specifically, for entities In other words, only The corresponding row in the middle belongs to The items are reserved for construction. Subsequently, on Each line of application Normalization is performed to obtain the final smooth similarity matrix.
[0181] IV. Neighborhood-guided distribution structure is smoothed and enhanced.
[0182] Local neighbors can provide a more accurate estimate of inter-class relationships than a single node. This embodiment, to further enhance the expressive power of the similarity matrix, utilizes data from the local neighbor set. Heuristic information is used to constrain the similarity consistency between entities and their neighborhoods, thereby further enhancing the smoothness of the similarity matrix. It is then adjusted as needed to achieve better global properties; this step is called neighborhood-guided distribution construction smoothness enhancement. This process can be summarized as follows: first, based on the initially obtained similarity matrix... Smoothing yields a neighborhood-consistent matrix Then, its enhanced representation is obtained using neighboring nodes. Finally, according to the needs, Extend to full image generation 'In order to achieve better global characteristics.'
[0183] This embodiment can be re-tested. - The mutual nearest neighbor strategy is used to determine local neighbors. Specifically, for entities... its neighborhood Defined as , here In the entity In the neighborhood space, this embodiment defines For entities To the entity The average similarity of local neighbors defines the local neighborhood. The average similarity between paired entities is Larger This means that neighbors are more likely to belong to the same cluster, reflecting the reliability of local neighbors. With entity The similarity should be related to the factor Consistency, in assigning weights After adding regularization terms, the optimization objective can be expressed as:
[0184]
[0185]
[0186]
[0187]
[0188] Operators You can multiply corresponding elements of two vectors to obtain a new vector. It is the local cluster set introduced in the previous section to constrain similarity diffusion, with weights of . Regular terms can be used to constrain the target. To prevent it from deviating too much from the initial value, a lower value can usually be set. To achieve better numerical stability, the value is determined. Furthermore, in constructing vectors... In this embodiment, a truncation operation is used to manually ensure that all values are not greater than a certain value. After solving the KKT conditions for the optimization problem, a neighborhood-consistent matrix can be obtained. The elements are as follows:
[0189]
[0190] In addition to increasing consistency, local neighbors can further enhance the representation of by other means. For example, when constructing the nearest neighbor graph, higher weights can be assigned to edges related to , thus emphasizing the role of local neighbors in the diffusion process. Formally, when , the entry of the indicator matrix becomes , where the hyperparameter is the importance weight. Furthermore, since entities within the same neighborhood are likely to belong to the same class, the present embodiment can aggregate entities within the neighborhood to obtain a neighborhood-enhanced representation . This can be implemented by a weighted average of the manifold distribution of entities in :
[0191]
[0192] The similarity matrix obtained in this step can be regarded as the best representation of the data manifold in local clusters. If better global expressiveness is needed, the present embodiment can perform an additional information propagation on the whole graph to obtain the final similarity matrix . The propagation matrix can be computed by . In this process, the present embodiment can also combine a sparsification operation to handle the transition matrix, thus better reducing the impact of outliers.
[0193] Five, manifold-aware geometric transport.
[0194] In the next step, the goal of the present embodiment is to establish a distance metric in an efficient way, which can be used to better measure the mutual relationship of entities in the manifold space. Some traditional methods lack consideration of the geometric properties of the entire manifold space when calculating the distance between pairs of samples. To solve this problem, the present embodiment models the problem as a transport of the distribution of each entity in the manifold space to represent . The Wasserstein distance thus calculated can fully consider the manifold information of the entire space. To simplify this optimal transport process, the present embodiment defines the transport path as the shortest path between two nodes in the graph. Formally, for nodes and in the manifold graph , the shortest path between them is defined as :
[0195]
[0196] where each path in the set is composed of a sequence of nodes, as shown in the definition , the sequence of nodes starts from and ends at . The edge between node and is defined as , the geodesic line on the manifold graph from to is equivalent to the shortest distance between the two points, and this path contains the overall information in the data manifold space. Formally, the geodesic distance from to is defined as:
[0197] The calculation of this distance is equivalent to summing the weights of the edges on the geodesic path. The set of geodesic distances between all pairs of entity nodes is
[0198] . In the field of graph theory, calculating the shortest distance between all pairs of nodes on a graph is a mature problem that has been discussed for decades. For different sizes of data, many solutions can be found to efficiently obtain the geodesic distance matrix . Next, the geodesic distance matrix is taken as the cost of optimal transport, and the optimal transport matrix between two manifold distributions can be obtained by solving the following optimization objective:
[0199]
[0200]
[0201] where the matrix represents the optimal transport matrix from distribution to distribution on the manifold space. To more efficiently calculate the optimal distance, an entropy regularization term is added to the optimization objective, so that this optimal transport problem can be solved iteratively by the Sinkhorn-Knopp algorithm. After obtaining the optimal transport matrix from to , their distance in the manifold graph can be calculated by the following formula:
[0202]
[0203] Further, in order to preserve the important neighbor relationship in the original distance space, the initial Euclidean distance is also taken into account in this embodiment, so that the geometric transport distance between the two distributions finally obtained can be expressed as:
[0204]
[0205] The initial Euclidean distance in the formula is , and the optimal transport distance on the manifold is , which are balanced by a hyperparameter . The geometric transport distance finally obtained can more highly measure the similarity relationship between entities, and can be used to guide the fusion construction of the knowledge graph or improve the retrieval and reasoning efficiency of the knowledge graph.
[0206] The method proposed in this embodiment has the advantages of low time complexity, strong generalization, and parallel implementation on a graphics processing unit (GPU), and can effectively improve the construction, updating, mining, fusion, reasoning, and retrieval capabilities of the knowledge graph.
[0207] The similarity diffusion-based knowledge graph representation device provided by the present application is described below. The similarity diffusion-based knowledge graph representation device described below can be mutually referenced with the similarity diffusion-based knowledge graph representation method described above.
[0208] Figure 3 The structure diagram of the similarity diffusion-based knowledge graph representation device provided by the present application is shown in Figure 3 , which comprises the following modules:
[0209] The construction module 300 is configured to construct a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism.
[0210] The adjacency module 310 is configured to construct an adjacency graph matrix of the target knowledge graph based on the similarity relationship between each entity in the target knowledge graph.
[0211] The diffusion module 320 is configured to perform similarity diffusion under local clustering constraints based on the adjacency graph matrix to obtain a first similarity matrix.
[0212] The representation module 330 is configured to perform geometric transport on the target knowledge graph in a manifold graph based on the first similarity matrix to obtain a relationship representation of the target knowledge graph.
[0213] According to the similarity diffusion-based knowledge graph representation device provided by the present application, the adjacency graph matrix of the target knowledge graph is constructed based on the similarity relationship between each entity in the target knowledge graph, and specifically comprises:
[0214] Obtaining an initial distance matrix of each entity based on the Euclidean distance between each entity in the target knowledge graph;
[0215] Constructing an adjacency graph matrix of the target knowledge graph based on the initial distance matrix and the nearest neighbor edge of each entity.
[0216] According to the knowledge graph representation device based on similarity diffusion provided by the application, the relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in the manifold graph based on the first similarity matrix, and specifically includes:
[0217] The first similarity matrix is subjected to neighborhood-guided smoothing enhancement based on the neighbor nodes of the initial distance matrix, and the second similarity matrix is obtained;
[0218] The relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in the manifold graph based on the second similarity matrix.
[0219] According to the knowledge graph representation device based on similarity diffusion provided by the application, the first similarity matrix is obtained by performing similarity diffusion under local clustering constraint based on the adjacency graph matrix, and specifically includes:
[0220] The local clustering graph of each entity is constructed based on the adjacency graph matrix and the k- mutual neighbor strategy;
[0221] The first similarity matrix is obtained by performing diffusion process on the local clustering graph of each entity through a bidirectional propagation mechanism.
[0222] According to the knowledge graph representation device based on similarity diffusion provided by the application, the relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in the manifold graph based on the second similarity matrix, and specifically includes:
[0223] The third similarity matrix is obtained by performing information propagation on the second similarity matrix;
[0224] The relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in the manifold graph based on the third similarity matrix.
[0225] According to the knowledge graph representation device based on similarity diffusion provided by the application, the relationship representation of the target knowledge graph is obtained by performing geometric transmission on the target knowledge graph in the manifold graph based on the first similarity matrix, the second similarity matrix or the third similarity matrix, and specifically includes:
[0226] The geodesic distance matrix containing manifold space information is obtained based on the initial distance matrix and the adjacency graph matrix;
[0227] An optimal transmission matrix between distribution representations of each entity in the target knowledge graph in the manifold space is calculated based on the geodesic distance matrix; wherein the distribution representations of each entity in the target knowledge graph in the manifold space are obtained based on the first similarity matrix, the second similarity matrix or the third similarity matrix;
[0228] A geometric transmission distance between each entity is determined based on the optimal transmission matrix and the distribution representations of each entity in the manifold space.
[0229] A relationship representation of the target knowledge graph is obtained based on the geometric transmission distance.
[0230] Figure 4 The electronic device provided by the present application provides a structural schematic diagram of an electronic device, as shown in Figure 4 The electronic device can include a processor (processor) 410, a communication interface (communications interface) 420, a memory (memory) 430 and a communication bus 440, wherein the processor 410, the communication interface 420, the memory 430 and the communication bus 440 complete mutual communication through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the knowledge graph representation method based on similarity diffusion, which includes:
[0231] Based on the multi-modal large model and the human-in-the-loop mechanism, a target knowledge graph is constructed.
[0232] Based on the similarity relationship between each entity in the target knowledge graph, an adjacency graph matrix of the target knowledge graph is constructed.
[0233] Based on the adjacency graph matrix, similarity diffusion is performed under local clustering constraints to obtain a first similarity matrix.
[0234] Based on the first similarity matrix, geometric transmission is performed on the target knowledge graph in the manifold graph to obtain a relationship representation of the target knowledge graph.
[0235] Moreover, the logic instructions in the memory 430 described above can be implemented in the form of software function units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0236] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the knowledge graph representation method based on similarity diffusion provided by the above-mentioned methods, the method comprising:
[0237] constructing a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism;
[0238] constructing an adjacency graph matrix of the target knowledge graph based on the similarity relationship between each entity in the target knowledge graph;
[0239] performing similarity diffusion under local clustering constraints based on the adjacency graph matrix to obtain a first similarity matrix;
[0240] performing geometric transmission of the target knowledge graph in a manifold graph based on the first similarity matrix to obtain a relationship representation of the target knowledge graph.
[0241] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the knowledge graph representation method based on similarity diffusion provided by the above-mentioned methods, the method comprising:
[0242] constructing a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism;
[0243] constructing an adjacency graph matrix of the target knowledge graph based on the similarity relationship between each entity in the target knowledge graph;
[0244] performing similarity diffusion under local clustering constraints based on the adjacency graph matrix to obtain a first similarity matrix;
[0245] Based on the first similarity matrix, the target knowledge graph is geometrically transmitted in the manifold graph to obtain a relationship representation of the target knowledge graph.
[0246] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0247] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0248] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A knowledge graph representation method based on similarity diffusion, characterized in that, The method comprises the steps of: constructing a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism; constructing an adjacency graph matrix of the target knowledge graph based on the similarity relationship between each entity in the target knowledge graph; based on the adjacency graph matrix, similarity diffusion is performed under local clustering constraints to obtain a first similarity matrix; based on the first similarity matrix, geometric transmission is performed on the target knowledge graph in a manifold graph to obtain a relationship representation of the target knowledge graph; based on the similarity relationship between each entity in the target knowledge graph, an adjacency graph matrix of the target knowledge graph is constructed, specifically comprising: based on the Euclidean distance between each entity in the target knowledge graph, an initial distance matrix of the entity is obtained; based on the initial distance matrix and the nearest neighbor edge of the entity, the adjacency graph matrix of the target knowledge graph is constructed; based on the first similarity matrix, geometric transmission is performed on the target knowledge graph in a manifold graph to obtain a relationship representation of the target knowledge graph, specifically comprising: based on the neighbor nodes of the initial distance matrix, the first similarity matrix is subjected to neighborhood-guided smoothing enhancement to obtain a second similarity matrix; based on the second similarity matrix, geometric transmission is performed on the target knowledge graph in a manifold graph to obtain a relationship representation of the target knowledge graph. 2.The similarity diffusion based knowledge graph representation method of claim 1, wherein, based on the adjacency graph matrix, similarity diffusion is performed under local clustering constraints to obtain a first similarity matrix, specifically comprising: based on the adjacency graph matrix and the k-neighbor strategy, a local clustering graph of each entity is constructed; the local clustering graph of each entity is subjected to a diffusion process through a bidirectional propagation mechanism to obtain a first similarity matrix. 3.The similarity diffusion based knowledge graph representation method of claim 1, wherein, based on the second similarity matrix, geometric transmission is performed on the target knowledge graph in a manifold graph to obtain a relationship representation of the target knowledge graph, specifically comprising: information propagation is performed on the second similarity matrix to obtain a third similarity matrix; based on the third similarity matrix, geometric transmission is performed on the target knowledge graph in a manifold graph to obtain a relationship representation of the target knowledge graph. 4.The similarity diffusion based knowledge graph representation method according to claim 1 or 3, characterized in that, based on the first similarity matrix, the second similarity matrix or the third similarity matrix, geometric transmission is performed on the target knowledge graph in a manifold graph to obtain a relationship representation of the target knowledge graph, specifically comprising: based on the initial distance matrix and the adjacency graph matrix, a geodesic distance matrix containing manifold space information is obtained; based on the geodesic distance matrix, an optimal transmission matrix between the distribution representations of each entity in the target knowledge graph in the manifold space is calculated; wherein the distribution representations of each entity in the target knowledge graph in the manifold space are obtained based on the first similarity matrix, the second similarity matrix or the third similarity matrix; based on the optimal transmission matrix and the distribution representations of each entity in the manifold space, the geometric transmission distance between each entity is determined; based on the geometric transmission distance, the relationship representation of the target knowledge graph is obtained. 5.A similarity diffusion based knowledge graph representation apparatus, characterized in that, The method comprises the steps of: constructing a target knowledge graph based on a multi-modal large model and a human-in-the-loop mechanism; An adjacency module is configured to construct an adjacency graph matrix of the target knowledge graph based on similarity relationships between entities in the target knowledge graph; A diffusion module is configured to perform similarity diffusion under local clustering constraints based on the adjacency graph matrix to obtain a first similarity matrix; A representation module is configured to perform geometric transmission of the target knowledge graph in a manifold graph based on the first similarity matrix to obtain relationship representation of the target knowledge graph; The adjacency graph matrix of the target knowledge graph is constructed based on similarity relationships between entities in the target knowledge graph, and specifically includes: An initial distance matrix of the entities is obtained based on Euclidean distances between the entities in the target knowledge graph; The adjacency graph matrix of the target knowledge graph is constructed based on the initial distance matrix and nearest neighbor edges of the entities; The relationship representation of the target knowledge graph is obtained by performing geometric transmission of the target knowledge graph in a manifold graph based on the first similarity matrix, and specifically includes: The first similarity matrix is subjected to neighborhood-guided smoothing enhancement based on neighbor nodes of the initial distance matrix to obtain a second similarity matrix; The relationship representation of the target knowledge graph is obtained by performing geometric transmission of the target knowledge graph in a manifold graph based on the second similarity matrix.
6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the knowledge graph representation method based on similarity diffusion according to any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the knowledge graph representation method based on similarity diffusion according to any one of claims 1 to 4.
8. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the knowledge graph representation method based on similarity diffusion according to any one of claims 1 to 4.
Citation Information
Patent Citations
Storage method and system of common sense knowledge graph
CN116910276A