A knowledge graph search method and system based on local semantics and global community

By adopting local semantic global community-based methods in knowledge graph search, the problems of insufficient local information processing, insufficient semantic processing capabilities, and insufficient global information expression and utilization in the prior art are solved, and more accurate and comprehensive search results are achieved.

CN119577098BActive Publication Date: 2025-05-16JIANGXI NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510134422.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-16
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

The existing knowledge graph search technology has problems such as insufficient local information processing, insufficient semantic processing capabilities, and insufficient global information expression and utilization when processing complex data, resulting in lack of integrity and accuracy of search results.

Method used

A knowledge graph search method based on local semantic global communityization is proposed. By extracting entities and relationships, undirected graphs are constructed, community division is performed, entity attributes within the community are extracted, community information tables and community reports are generated, and entities, relationships, data sources and community contexts are generated through vectorization processing and similarity calculations, and finally these context information are fused to improve the relevance and accuracy of search results.

Benefits of technology

Through global perspective and deep semantic analysis, this method can more accurately match user queries, improve the integrity and accuracy of search results, and meet complex user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119577098B_ABST
    Figure CN119577098B_ABST
Patent Text Reader

Abstract

The present application relates to the field of knowledge graph technology, and discloses a knowledge graph search method and system based on local semantic global communityization, the method comprising: extracting entities and their relationships from the knowledge graph to construct an undirected graph; dividing the undirected graph into multi-level community information to obtain a community report by generating a community information table; selecting instance triples containing data source information from the knowledge graph, constructing an instance triple information source table, and extracting the ID, name and attributes of the entity to generate an entity attribute table to obtain an entity attribute vector table; by extracting key entities in the query question and performing vector representation, calculating the similarity with the entity attribute vector table, selecting the top N related entities to generate context (including entity context, relationship context, data source context and community context), and finally obtaining the search results for the query question. The method can accurately retrieve information related to the query question from the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graph technology, and specifically to a knowledge graph search method and system based on local semantics and global community. Background Art

[0002] Knowledge Graph (KG), as a complex network structure for organizing and expressing data, has become a key tool for efficient use of data elements. Knowledge graphs achieve the transformation from data to knowledge through structured representation of entities and their relationships, and are widely used in scenarios such as information retrieval, personalized recommendations, and intelligent question-and-answer services. However, with the continuous growth of data scale and complexity, existing KG technologies face many challenges at the search and application levels, including uneven information distribution, lack of global structure, and insufficient deep semantic expression.

[0003] Although current KG search technology has made some progress, it still has the following key problems in the search process: (1) Limited to local information processing and lack of global perspective. Existing KG search usually focuses on processing single entities or local relationships, ignoring the overall structure of the graph and the complex relationship network. This one-sided processing method makes the search results lack integrity and cannot help users gain comprehensive knowledge insights. (2) Insufficient semantic processing capabilities and difficulty in responding to complex semantic needs. Most current technologies remain at the processing of surface features of entities and fail to deeply analyze the semantic attributes and relationships of entities. This shallow processing method leads to inaccurate search results, especially in complex scenarios. It is difficult to meet user needs. (3) Insufficient expression and utilization of global information. Many KG tools fail to effectively organize and analyze global information, making it difficult for users to understand the hierarchical structure and subject classification between data. This not only limits the potential of global information, but also creates obstacles in multi-level application scenarios.

[0004] Therefore, the current KG search technology needs to further improve the global perspective, semantic processing capabilities, global information expression and utilization, etc. to better meet the complex needs of users. Summary of the invention

[0005] Based on this, this application proposes a knowledge graph search method and system based on local semantics and global community, aiming to further improve the global perspective, semantic processing capabilities, global information expression and utilization in KG search technology, so as to better meet complex user needs.

[0006] The first aspect of the present application provides a knowledge graph search method based on local semantic global communityization, the method comprising:

[0007] Extract entities and relationships between them from the target semantic knowledge graph to construct an undirected graph and for data preprocessing;

[0008] Performing community division on the undirected graph to obtain multi-level community information;

[0009] Based on the multi-level community information, extract the attributes and attribute values ​​of the entities in each community from the target semantic knowledge graph to generate a community information table, and generate a community report according to the community information table;

[0010] Select <instance triple-type triple> containing data source information from the target semantic knowledge graph to construct an instance triple information source table;

[0011] Extracting the entity ID, name, attribute and attribute value thereof from the <instance triplet-type triplet> to construct an entity attribute table, and performing vectorization processing on the entity attribute table to generate an entity attribute vector table;

[0012] Extract key entities from the query question, convert the key entities into vector representations, calculate similarity between the vector representations and entity attribute vectors in the entity attribute vector table, and select N related entities with the highest similarity;

[0013] Generating entity context according to the similarity ranking of the N related entities;

[0014] Generating a relationship context according to the relationship weights and relationship degree sorting of the N related entities in the instance triple information source table;

[0015] Find the corresponding information source in the instance triple information source table according to the N related entities to form a data source context;

[0016] Generating a community context according to the quantity and level ranking of the N related entities in the community report;

[0017] fusing the entity context, the relationship context, the data source context and the community context into a final context;

[0018] Based on the final context, a search result corresponding to the question to be queried is obtained.

[0019] As an optional implementation of the first aspect, extracting entities and relationships between the entities from the target semantic knowledge graph to construct an undirected graph and the steps for data preprocessing include: extracting entities and relationships between the entities from the target semantic knowledge graph, taking the entities as nodes and the relationships between the entities as edges; constructing the undirected graph based on the nodes and the edges; the data preprocessing includes: calculating the edge weights between the nodes and calculating the edge degrees of the nodes; the edge weight calculation formula is: ,in, represents the edge weight, that is, the relationship weight, Indicates the number of times a relationship occurs. Represents the total number of occurrences of all relationships; the calculation formula for the edge degree of the node is: ,in, Representation Node i The edge degree, that is, the relationship degree, Representation Node i The set of neighbor nodes of j Representation and Node i An adjacent node, Represents each neighbor node to node i The degree of contributes one unit.

[0020] As an optional implementation of the first aspect, the step of dividing the undirected graph into communities to obtain multi-level community information includes: performing preliminary clustering based on the nodes and edges in the undirected graph to generate primary community information; if the number of levels n set by the user has not been reached, n>0, taking the undirected graph of the previous level as the parent graph, extracting and clustering subgraphs from the parent graph, and combining the community information separated from the community information of the previous level to generate the next level of community information to obtain multi-layer non-primary community information; integrating the primary community information and the multi-layer non-primary community information into the multi-level community information.

[0021] As an optional implementation manner of the first aspect, the step of performing preliminary clustering according to the nodes and edges in the undirected graph to generate primary community information includes: calculating the modularity of the edge weight and edge degree of each node in the undirected graph according to the Leiden algorithm; performing preliminary community division on all nodes in the undirected graph according to the modularity to obtain the primary community information; the calculation formula of the modularity is: ,in, represents modularity, Representation Node i and nodes j The edge weights between represents the sum of weights of all edges, Representation Node iThe edge degree of i The sum of the weights of the connected edges, is the indicative function, when the node i and nodes j When they belong to the same community, =1, otherwise =0 .

[0022] As an optional implementation of the first aspect, extracting key entities from the question to be queried, converting the key entities into vector representations, calculating similarity between the vector representations and entity attribute vectors in the entity attribute vector table, and selecting the top N related entities in similarity includes: extracting key entities from the question to be queried using a large language model; converting the key entities into vector representations using a text-embedding-3-small model, wherein the vector representations and the entity attribute vectors in the entity attribute vector table are vectors of the same dimension; calculating similarity between the vector representations and the entity attribute vectors, sorting the similarity scores from high to low, and selecting the top N related entities in similarity scores; the similarity calculation formula is: ,in, represents the similarity score, represents the inner product of two vectors, They are vector representations and entity attribute vector The second norm of .

[0023] As an optional implementation of the first aspect, the step of generating a relationship context according to the relationship weights and relationship degrees of the N related entities in the instance triple information source table includes: according to the N related entities, extracting entities that have a direct relationship with the N related entities from the instance triple information source table to form an internal network, and extracting entities that have an indirect relationship with the N related entities to form an external network; wherein the direct relationship is relationship information listed in descending order according to the relationship weights in the instance triple information source table, and the indirect relationship is relationship information listed in descending order according to the relationship degrees in the instance triple information source table; and constructing the relationship context according to the internal network and the external network.

[0024] As an optional implementation of the first aspect, the step of generating a community context according to the number and level sorting of the N related entities in the community report includes: obtaining corresponding communities from the community report as candidate communities according to the N related entities; counting the number of related entities contained in each candidate community, and sorting them from high to low in terms of number; if the candidate communities contain the same number of related entities, arranging the candidate communities from low to high in terms of level; and constructing the community context according to the community information of the sorted candidate communities.

[0025] The second aspect of the present application provides a knowledge graph search system based on local semantic global communityization, the system comprising:

[0026] A global information extraction module is used to extract entities and the relationships between the entities from the target semantic knowledge graph to construct an undirected graph and for data preprocessing; divide the undirected graph into communities to obtain multi-level community information; based on the multi-level community information, extract the attributes and attribute values ​​of the entities in each community from the target semantic knowledge graph to generate a community information table, and generate a community report based on the community information table;

[0027] A local information extraction module is used to select <instance triple-type triple> containing data source information from the target semantic knowledge graph to construct an instance triple information source table; extract the ID, name, attribute and attribute value of the entity from the <instance triple-type triple> to construct an entity attribute table, and vectorize the entity attribute table to generate an entity attribute vector table;

[0028] A query processing module is used to extract key entities from the query question, convert the key entities into vector representations, calculate the similarity between the vector representations and the entity attribute vectors in the entity attribute vector table, select N related entities with the highest similarity, sort the N related entities according to their similarity, generate entity contexts, sort the N related entities according to their relationship weights and relationship degrees in the instance triple information source table, generate relationship contexts, find the corresponding information sources in the instance triple information source table according to the N related entities to form data source contexts, sort the N related entities according to their number and level in the community report, generate community contexts, and merge the entity contexts, the relationship contexts, the data source contexts and the community contexts into final contexts.

[0029] The query result generating module is used to obtain the information source of the question to be queried based on the final context.

[0030] The third aspect of the present application provides an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the above-mentioned knowledge graph search method based on local semantics and global communityization.

[0031] The fourth aspect of the present application provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above-mentioned knowledge graph search method based on local semantics and global communityization.

[0032] Compared with the prior art, the present application provides a knowledge graph search method based on local semantic global communityization: first, entities and relationships are extracted to construct an undirected graph, and then community division is performed to decompose the complex knowledge graph into subsets with intrinsic connections, which helps to understand and explore the structure of the knowledge graph, and also facilitates targeted information extraction for each community; by extracting entity attributes in the community, community information tables and community reports are generated. This step strengthens the understanding of each community and helps locate knowledge areas related to the query question; constructing an instance triple information source table and an entity attribute table provides detailed information on the data source and entity attributes, respectively, laying the foundation for subsequent vectorization processing and similarity calculation; through vectorization processing and similarity calculation, entity attributes are converted In vector form, it is convenient to calculate the similarity between the key entity and other entities in the knowledge graph, thereby narrowing the search scope; by generating entity context, relationship context, data source context and community context, these contextual information reflect the position, relationship, data source and importance of the community to which the key entity belongs in the knowledge graph, providing comprehensive background information for the final search results; then through context fusion, various context information are integrated together to form the final context. This step comprehensively considers multiple dimensions such as entities, relationships, data sources and communities, and improves the relevance and accuracy of search results; finally, based on the final context, search results are obtained. Using the fused context information, the system can more accurately match the query question and thus reply to the most relevant search results. Overall, this technical solution achieves efficient and accurate retrieval of information related to the query question from the knowledge graph through the above steps, reflecting the search advantages of local semantics and global communityization.

[0033] Additional aspects and advantages of the present application will be given in part in the following description, and in part will become apparent from the following description, or will be understood through the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A flowchart of a knowledge graph search method based on local semantic global communityization proposed in the first embodiment of the present application;

[0035] Figure 2 This is a flowchart of dividing the knowledge graph search method into offline and online parts in the first embodiment of the present application;

[0036] Figure 3 This is a flowchart of local semantic depth analysis and global multi-level community division in the first embodiment of the present application;

[0037] Figure 4 This is a specific implementation flow chart of global multi-level community division in the first embodiment of the present application;

[0038] Figure 5 This is a flow chart of intelligent question answering using global information and local information in the first embodiment of the present application;

[0039] Figure 6 A structural diagram of a knowledge graph search system based on local semantics and global community proposed in the second embodiment of the present application.

[0040] The following specific implementation methods will further illustrate the present application in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0042] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.

[0043] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.

[0044] Example 1

[0045] See also Figure 1 , which is a flowchart of a knowledge graph search method based on local semantics and global community proposed in the first embodiment of the present application.

[0046] like Figure 2As shown, the method proposed for the first embodiment of this application can be divided into two parts: offline and online. It should be noted that the offline part performs knowledge graph preprocessing and flattening to generate global indexes and local indexes, and the online part uses global indexes and local indexes to perform multi-scenario intelligent trusted question answering. In the offline part, global and local flattening of the knowledge graph involves local semantic depth analysis and global multi-level community division. For details, see Figure 3 and Figure 4 The large language model is used in the online part of "extracting key entities from questions" and "generating answers". At the end of each answer in the answer results, there will be <...> to refer to the global and local context, clearly stating the specific source of the answer.

[0047] like Figure 3 As shown, it is a flowchart of the local semantic deep analysis and global multi-level community division of this application. It should be noted that the two steps of local semantic deep analysis and global multi-level community division are performed in parallel. The main steps of local semantic deep analysis are to extract and embed entity attributes and extract instance triple information sources. The global multi-level community division mainly converts the semantic knowledge graph into an undirected graph, divides the community according to the user-specified level, improves the entity information in each community, and uses a large language model to generate a community report.

[0048] like Figure 4 As shown, this is a specific implementation flowchart of the global multi-level community division of this application. It should be noted that the mining of global multi-level communities is mainly introduced. The Leiden algorithm is first used to divide the community of the undirected graph to generate the 0th level community. Each community at the 0th level is mapped to an undirected subgraph, and then the Leiden algorithm is used to divide the community. It is recursive repeatedly until the user-specified level is reached, and finally the community information of each level is summarized.

[0049] like Figure 5 As shown, this is a flowchart of the application using global information and local information for intelligent question answering. It should be noted that community reports, instance triplet information sources, relationships, and entity information are retrieved based on similar entities. The retrieved information is reasonably sorted under predetermined rules (community ranking is based on the number and level of matching entities, information source ranking is based on descending order of entity scores, relationship ranking is based on descending order of relationship degree and weight, and entity rankings are sorted according to similarity scores when matching), and global and local contexts are formed under the limit of the number of tokens. The proportion and size of the tokens can be customized to facilitate multi-scenario intelligent question answering.

[0050] The method proposed in the first embodiment of the present application may specifically include the following steps S01 to S04.

[0051] S01: Extract entities and relationships between entities from the target semantic knowledge graph to construct an undirected graph and use it for data preprocessing; divide the undirected graph into communities to obtain multi-level community information; based on the multi-level community information, extract the attributes and attribute values ​​of entities in each community from the target semantic knowledge graph to generate a community information table, and generate a community report based on the community information table.

[0052] In some embodiments, entities and relationships between entities are extracted from the target semantic knowledge graph, with entities as nodes and relationships between entities as edges; an undirected graph is constructed based on the nodes and edges; data preprocessing includes: calculating edge weights between nodes, and calculating edge degrees of nodes.

[0053] Specifically, load entity and relationship data from the target semantic knowledge graph to build an undirected graph: Extract all entities and relationships from the semantic KG and build an undirected graph:

[0054] ,

[0055] Among them, V represents the node set, corresponding to the unique ID of the entity, E represents the edge set, each edge Representation Node i and nodes j The relationship between ,and .

[0056] The calculation formula for the edge weight of the calculation node is:

[0057] ,

[0058] in, represents the edge weight, that is, the relationship weight, Indicates the number of times a relationship occurs. represents the total number of occurrences of all relations;

[0059] The calculation formula for the edge degree of a node is:

[0060] ,

[0061] in, Representation Node i The edge degree, that is, the relationship degree, Representation Node i The set of neighbor nodes of j Representation and Node i An adjacent node, Represents each neighbor node to node i The degree of contributes one unit.

[0062] In this step, by analyzing the frequency of each relationship in the relational data, the relationship weights are calculated, which is transformed into calculating the edge weights of nodes in an undirected graph. And the number of edges connected to each node is calculated as the edge degree of the node to measure the importance and centrality of the node, preparing data for community division.

[0063] Next, in some embodiments, based on the nodes and edges in the undirected graph, preliminary clustering is performed to generate primary community information; if the user-specified level number n, n>0, has not been reached, the undirected graph of the previous level is used as the mother graph, subgraphs are extracted from the mother graph and clustered, and combined with the community information separated from the community information of the previous level to generate the community information of the next level; the primary community information and non-primary community information are integrated into multi-level community information.

[0064] Specifically, according to the undirected weighted graph G u and the level number n, community division is performed, and the following process is executed:

[0065] First, initial community division is performed.

[0066] Use the Leiden algorithm to perform preliminary clustering on the nodes in the undirected graph G u to generate the community information of level 0. At this level, the format of the community information is <C, {ID and e}>, where C is the community name, ID is the ID of the entity, e is the entity name, and each community consists of several entities. In the implementation process of the Leiden algorithm, the edge weights play an important role in community division. For the undirected weighted graph , each edge has a weight , reflecting the strength or importance of the relationship between node i and node j. In this case, when calculating the modularity Q, the Leiden algorithm replaces the standard adjacency matrix with the weighted adjacency matrix to more accurately characterize the contribution degree of different relationships in the network structure. Specifically, the weighted modularity formula is expressed as follows:

[0067] ,

[0068] where, represents the modularity, represents the edge weight between node i and node j , represents the sum of the weights of all edges, represents the edge degree of node i , that is, the sum of the weights of the edges connected to node i , is the indicative function, when the node i and nodes j When they belong to the same community, =1, otherwise =0 .

[0069] It should be noted that in a weighted undirected graph, edge weights directly affect the community division results. Node pairs with larger edge weights have stronger connections and are therefore more likely to be classified into the same community during modularity optimization; while node pairs with smaller edge weights have weaker connections and are more likely to be divided into different communities. This mechanism ensures that community division not only focuses on the number of connections between nodes, but also reflects the strength of the edges, thereby improving the accuracy and stability of the division.

[0070] Secondly, multi-level community division is carried out.

[0071] If the number of levels n set by the user has not yet been reached, all entities contained in each community at the current level must be mapped from the undirected graph G u Extract the corresponding undirected subgraph from G us , and cluster the nodes in the subgraph to generate the next level of community information. Summarize all the next level of community information collection to form new community information. If the number of levels n set by the user has been reached, stop further division.

[0072] Subsequently, the community information is collated. All community information from level 0 to level n is collected and collated to form complete community information. Based on this complete community information, the attributes and attribute values ​​of each entity in the community are extracted from the semantic knowledge graph to generate a community information table. The format of the community information table is<C、ID、e、Attr / Attrv、L> , where Attr / Attrv represents the attributes and attribute values ​​of the entity, and L represents the community level.

[0073] Finally, a community report is generated using a large language model based on the community information table. The content of the community report includes detailed information about each community as follows. Exemplarily, the community report includes a title, summary, impact severity rating, scoring explanation, and detailed findings. Among them, Title: represents its key entities-the title should be short but specific, and if possible, include representative named entities in the title. Summary: An executive summary of the overall structure of the community, the relationships between its entities, and important information related to its entities. Impact severity rating: A floating point score between 0-10 that indicates the severity of the impact constituted by the entities within the community. Impact is the scoring importance of the community. Scoring explanation: A single-sentence explanation of the impact severity rating. Detailed findings: A list of 5-10 key insights about the community. Each insight should have a short summary followed by multiple paragraphs of explanatory text. Entity ID: The ID of the entity contained in the community.

[0074] S02: Select <instance triple-type triple> containing data source information from the target semantic knowledge graph to construct an instance triple information source table; extract the entity ID, name, attributes and their attribute values ​​from <instance triple, type triple> to construct an entity attribute table, vectorize the entity attribute table, and generate an entity attribute vector table.

[0075] It should be noted that in the semantic knowledge graph, for each <(h, r, t)-(hT, rT, tT)> containing data source information, the following two tasks are performed simultaneously. Among them, (h, r, t) is an instance triple, h is the head entity, t is the tail entity, r is the relationship, (hT, rT, tT) is a type triple, hT is the head entity type, tT is the tail entity type, and rT is the relationship type.

[0076] The first task is to extract the data source information of the corresponding entity from these <(h, r, t)-(hT, rT, tT)> and generate records in the format of <(hT, rT, tT), tr>, where tr is the instance triple information source. These records are aggregated to form an instance triple information source table.

[0077] The second task is to extract the entity ID, name and its attributes / attribute values ​​from the same <(h, r, t)-(hT, rT, tT)> and construct an entity attribute table. Then, the text-embedding-3-small model is used to vectorize the entire entity attribute table to generate { E(e 1 ) , E(e 2 ) ,... , E(e n}, E(ei )=W e ·e i +b e , E(e i ) Is Entity e i The vector representation of W e is the weight matrix of the embedder, b e is the bias term and is stored in the entity attribute vector table. The entity attribute vector table contains not only the entity attribute vector of each entity, but also the ID, name, attribute / attribute value of each entity.

[0078] S03: Extract key entities from the query question, convert the key entities into vector representations, calculate the similarity between the vector representations and the entity attribute vectors in the entity attribute vector table, and select the top N related entities with the highest similarity; sort the N related entities according to their similarity to generate entity contexts; sort the N related entities according to their relationship weights and relationship degrees in the instance triple information source table to generate relationship contexts; find the corresponding information sources in the instance triple information source table based on the N related entities to form data source contexts; sort the N related entities according to their number and level in the community report to generate community contexts; merge the entity contexts, relationship contexts, data source contexts and community contexts into the final context.

[0079] In some embodiments, a large language model is used to extract key entities from the query question; a text-embedding-3-small model is used to convert the key entities into vector representations, wherein the vector representations and the entity attribute vectors in the entity attribute vector table are vectors of the same dimension; similarity is calculated between the vector representations and the entity attribute vectors, the similarity scores are sorted from high to low, and N related entities with the top similarity scores are selected.

[0080] For example, a large language model is used to extract key entities from the query question { e k1, e k2, ... , e kn}, and use the text-embedding-3-small model to convert these key entities into vector representations { E(e 1 ) , E(e 2 ) , ... ,E (e n )}, and transform these vectors into vectors of the same dimension, so that , then, use the formula Represent this vector E(e kj ) Calculate the similarity with each entity attribute vector in the entity attribute vector table, where is the inner product of two vectors, They are vector representations and entity attribute vector The second norm (i.e., vector length) of the entities is calculated and sorted from high to low according to their similarity scores to obtain the top 20 entities, which are recorded as ={e 1 , e 2 ,... e i}.

[0081] Furthermore, contexts are constructed for the first 20 entities, mainly including constructing entity context, relationship context, data source context and community context.

[0082] First, build entity context:

[0083] For example, according to the sorted (similarity scores from high to low) When the generated context information reaches the specified token number threshold (20% of the maximum token number), it stops immediately.

[0084] Second, build relationship context:

[0085] In some embodiments, based on N related entities, entities that are directly related to the N related entities are extracted from the instance triple information source table to form an inner network, and entities that are indirectly related to the N related entities are extracted to form an outer network; wherein the direct relationship is the relationship information listed in descending order according to the relationship weight in the instance triple information source table, and the indirect relationship is the relationship information listed in descending order according to the relationship degree in the instance triple information source table; a relationship context is constructed based on the inner network and the outer network.

[0086] For example, according to the sorted , extract entities with direct relationships from the semantic knowledge graph to form an inner network, and then extract entities with indirect relationships to form an outer network. In the inner network, according to the relationship weight Sort and list these relationship information from large to small; in the external network, calculate the relationship degree directly connected to the entity , and press The relationship information is sorted from largest to smallest. The relationship context is generated based on the sorted relationship information. When the generated context information reaches the specified token number threshold (20% of the maximum token number), it stops immediately.

[0087] The third aspect is to build the data source context:

[0088] It should be noted that since the instance triple (h, r, t) is directly related to the information source, the sorted The corresponding information source is sorted by entity order. The data source context is generated based on the sorted information source. When the generated context information reaches the specified token number threshold (40% of the maximum token number), it stops immediately.

[0089] Fourth, build community context:

[0090] In some implementations, according to N related entities, corresponding communities are obtained from community reports as candidate communities; the number of related entities contained in each candidate community is counted, and the candidate communities are sorted from high to low by number; if the number of related entities contained in the candidate communities is the same, the candidate communities are sorted from low to high by level; and community contexts are constructed according to the community information of the sorted candidate communities.

[0091] For example, the sorted Find the corresponding communities C in the community report and use them as candidate communities C pre Then count each C pre The number of selected entities contained in the list is sorted from highest to lowest. C pre If the number of entities contained is the same, C pre The levels are arranged from low to high (such as level l0, level l1, level l2, etc.). Generate community context based on the sorted community information. When the generated context information reaches the specified token number threshold (20% of the maximum token number), stop immediately. Finally, the generated entity context, relationship context, data source context and community context are merged into the final context. The token ratio of these four contexts can be adjusted independently to flexibly adapt to different application scenarios.

[0092] S04: Based on the final context, obtain search results corresponding to the query question.

[0093] Specifically, based on the constructed generated context, a reply in markdown format is generated for the query question, that is, the search result, and each search result can provide a source of information based on the basis.

[0094] In summary, the implementation method of the present application first extracts entities and relationships to construct an undirected graph, and then performs community division to decompose the complex knowledge graph into subsets with intrinsic connections, which helps to understand and explore the structure of the knowledge graph, and also facilitates targeted information extraction for each community; by extracting entity attributes in the community, a community information table and a community report are generated. This step strengthens the understanding of each community and helps locate the knowledge area related to the query question; the two steps of constructing the instance triple information source table and the entity attribute table provide detailed information on the data source and entity attributes, respectively, laying the foundation for subsequent vectorization processing and similarity calculation; through vectorization processing and similarity calculation, the entity attributes are converted into vector form, which is convenient for calculating key entities. The similarity between the entity and other entities in the knowledge graph is used to narrow the search scope; by generating entity context, relationship context, data source context and community context, these contextual information reflect the position, relationship, data source and importance of the community to which the key entity belongs in the knowledge graph, providing comprehensive background information for the final search results; then through context fusion, various context information is integrated together to form the final context. This step comprehensively considers multiple dimensions such as entities, relationships, data sources and communities, and improves the relevance and accuracy of search results; finally, based on the final context, search results are obtained. Using the fused context information, the system can more accurately match the query question and thus reply with the most relevant search results. Overall, this technical solution achieves efficient and accurate retrieval of information related to the query question from the knowledge graph through the above steps, reflecting the search advantages of local semantics and global communityization.

[0095] Example 2

[0096] See also Figure 6 , shown is a schematic diagram of the structure of a knowledge graph search system based on local semantic global community proposed in the second embodiment of the present application, the system comprising:

[0097] The global information extraction module 100 is used to extract entities and the relationships between the entities from the target semantic knowledge graph to construct an undirected graph and for data preprocessing; perform community division on the undirected graph to obtain multi-level community information; based on the multi-level community information, extract the attributes and attribute values ​​of the entities in each community from the target semantic knowledge graph to generate a community information table, and generate a community report based on the community information table;

[0098] The local information extraction module 200 is used to select <instance triple-type triple> containing data source information from the target semantic knowledge graph to construct an instance triple information source table; extract the ID, name, attribute and attribute value of the entity from the <instance triple-type triple> to construct an entity attribute table, and vectorize the entity attribute table to generate an entity attribute vector table;

[0099] The query processing module 300 is used to extract key entities from the query question, convert the key entities into vector representations, calculate the similarity between the vector representations and the entity attribute vectors in the entity attribute vector table, and select N related entities with the highest similarity; generate entity contexts according to the similarity sorting of the N related entities; generate relationship contexts according to the relationship weights and relationship degrees sorting of the N related entities in the instance triple information source table; generate data source contexts according to the information source sorting of the N related entities in the instance triple information source table; generate community contexts according to the number and level sorting of the N related entities in the community report; merge the entity contexts, the relationship contexts, the data source contexts and the community contexts into the final context;

[0100] The query result generating module 400 is used to obtain the information source of the question to be queried based on the final context.

[0101] On the other hand, the present application also proposes an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the above-mentioned knowledge graph search method based on local semantics and global communityization.

[0102] On the other hand, the present application also proposes a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above-mentioned knowledge graph search method based on local semantics and global community.

[0103] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0104] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0105] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A knowledge graph search method based on local semantics and global community, characterized in that: The method comprises: Extract entities and relationships between them from the target semantic knowledge graph to construct an undirected graph and for data preprocessing; Performing community division on the undirected graph to obtain multi-level community information; Based on the multi-level community information, extract the attributes and attribute values ​​of the entities in each community from the target semantic knowledge graph to generate a community information table, and generate a community report according to the community information table; Select <instance triple-type triple> containing data source information from the target semantic knowledge graph to construct an instance triple information source table; Extracting the entity ID, name, attribute and attribute value thereof from the <instance triplet-type triplet> to construct an entity attribute table, and performing vectorization processing on the entity attribute table to generate an entity attribute vector table; Extract key entities from the query question, convert the key entities into vector representations, calculate similarity between the vector representations and entity attribute vectors in the entity attribute vector table, and select N related entities with the highest similarity; Generating entity context according to the similarity ranking of the N related entities; Generating a relationship context according to the relationship weights and relationship degree sorting of the N related entities in the instance triple information source table; Find the corresponding information source in the instance triple information source table according to the N related entities to form a data source context; Generating a community context according to the quantity and level ranking of the N related entities in the community report; fusing the entity context, the relationship context, the data source context and the community context into a final context; Based on the final context, a search result corresponding to the question to be queried is obtained.

2. According to claim 1, a knowledge graph search method based on local semantics and global communityization is characterized in that: The steps of extracting entities and the relationships between the entities from the target semantic knowledge graph to construct an undirected graph and for data preprocessing include: Extracting entities and relationships between entities from the target semantic knowledge graph, taking the entities as nodes, and taking the relationships between the entities as edges; Constructing the undirected graph according to the nodes and the edges; The data preprocessing includes: calculating the edge weights between the nodes, and calculating the edge degrees of the nodes; The calculation formula of the edge weight is: , in, represents the edge weight, that is, the relationship weight, Indicates the number of times a relationship occurs. represents the total number of occurrences of all relations; The calculation formula of the edge degree of the node is: , in, Representation Node i The edge degree, that is, the relationship degree, Representation Node i The set of neighbor nodes of j Representation and Node i An adjacent node, Represents each neighbor node to node i The degree of contributes one unit.

3. According to claim 2, a knowledge graph search method based on local semantics and global communityization is characterized in that: The steps of dividing the undirected graph into communities and obtaining multi-level community information include: Performing preliminary clustering according to the nodes and edges in the undirected graph to generate primary community information; If the number of levels n set by the user has not been reached, and n>0, the undirected graph of the previous level is used as the parent graph, and subgraphs are extracted and clustered from the parent graph, and the community information separated from the community information of the previous level is combined to generate the community information of the next level, so as to obtain multi-layer non-primary community information; The primary community information and the multi-layer non-primary community information are integrated into the multi-layer community information.

4. According to claim 3, a knowledge graph search method based on local semantics and global communityization is characterized in that: The steps of performing preliminary clustering according to the nodes and edges in the undirected graph to generate primary community information include: According to the Leiden algorithm, the modularity is calculated for the edge weight and edge degree of each node in the undirected graph; According to the modularity, all nodes in the undirected graph are preliminarily divided into communities to obtain the primary community information; The calculation formula of the modularity is: , in, represents modularity, Representation Node i and nodes j The edge weights between represents the sum of weights of all edges, Representation Node i The edge degree of i The sum of the weights of the connected edges, is an indicative function, when the node i and nodes j When they belong to the same community, =1, otherwise =0 .

5. According to claim 1, a knowledge graph search method based on local semantics and global communityization is characterized in that: The steps of extracting key entities from the query question, converting the key entities into vector representations, calculating similarity between the vector representations and entity attribute vectors in the entity attribute vector table, and selecting N related entities with the highest similarity include: Extract key entities from the query question using a large language model; The key entity is converted into a vector representation using a text-embedding-3-small model, wherein the vector representation is a vector of the same dimension as the entity attribute vector in the entity attribute vector table; Calculate the similarity between the vector representation and the entity attribute vector, sort the similarity scores from high to low, and select N related entities with the highest similarity scores; The calculation formula of the similarity is: , in, represents the similarity score, represents the inner product of two vectors, They are vector representations and entity attribute vector The second norm of .

6. According to claim 1, a knowledge graph search method based on local semantics and global communityization is characterized in that: The step of generating a relationship context according to the relationship weights and relationship degrees of the N related entities in the instance triple information source table comprises: According to the N related entities, entities directly related to the N related entities are extracted from the instance triple information source table to form an inner network, and entities indirectly related to the N related entities are extracted to form an outer network; The direct relationship is relationship information listed in descending order according to the relationship weight in the example triple information source table, and the indirect relationship is relationship information listed in descending order according to the relationship degree in the example triple information source table; The relationship context is constructed according to the inner network and the outer network.

7. According to claim 1, a knowledge graph search method based on local semantics and global communityization is characterized in that: According to the quantity and level ranking of the N related entities in the community report, the step of generating the community context includes: According to the N related entities, obtaining a corresponding community from the community report as a candidate community; Counting the number of related entities contained in each candidate community and sorting them from high to low; If the candidate communities contain the same number of related entities, the candidate communities are arranged from low to high according to their levels; The community context is constructed according to the sorted community information of the candidate communities.

8. A knowledge graph search system based on local semantics and global community, characterized by: The system comprises: A global information extraction module is used to extract entities and the relationships between the entities from the target semantic knowledge graph to construct an undirected graph and for data preprocessing; divide the undirected graph into communities to obtain multi-level community information; based on the multi-level community information, extract the attributes and attribute values ​​of the entities in each community from the target semantic knowledge graph to generate a community information table, and generate a community report based on the community information table; A local information extraction module is used to select <instance triple-type triple> containing data source information from the target semantic knowledge graph to construct an instance triple information source table; extract the ID, name, attribute and attribute value of the entity from the <instance triple-type triple> to construct an entity attribute table, and vectorize the entity attribute table to generate an entity attribute vector table; A query processing module is used to extract key entities from the query question, convert the key entities into vector representations, calculate the similarity between the vector representations and the entity attribute vectors in the entity attribute vector table, select N related entities with the highest similarity, sort the N related entities according to their similarity, generate entity contexts, sort the N related entities according to their relationship weights and relationship degrees in the instance triple information source table, generate relationship contexts, find the corresponding information sources in the instance triple information source table according to the N related entities to form data source contexts, sort the N related entities according to their number and level in the community report, generate community contexts, and merge the entity contexts, the relationship contexts, the data source contexts and the community contexts into final contexts. The query result generating module is used to obtain the information source of the question to be queried based on the final context.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the executable instructions to implement a knowledge graph search method based on local semantics and global community as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute a knowledge graph search method based on local semantics and global community as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Scientific knowledge discovery method and system based on knowledge graph

    CN117786122A

  • RAG question and answer method and system based on knowledge graph and medium

    CN118673126A