Community discovery method and device for mineral product knowledge graph, and storage medium

Through a new mineral knowledge graph community discovery method, the problems of large dependence on computing power, limitations and insufficient model universality in traditional methods are solved, and more efficient knowledge graph construction and community distinction are achieved, reducing the uncertainty of large model output.

CN120144780AActive Publication Date: 2025-06-13INST OF MINERAL RESOURCES CHINESE ACAD OF GEOLOGICAL SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510214034.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The traditional mineral resource knowledge graph construction method has problems such as high computing power dependence, limitations, difficulty in extracting cross-text entity relationships, slow data update speed, insufficient model universality, and difficulty in application development.

Method used

A community discovery method for mineral knowledge graph is proposed. By obtaining the knowledge base, text blocking is performed, entities are extracted, entity relationship graph is constructed, distance and similarity between entities are calculated, edge weights are determined, module degree is calculated, and nodes are aggregated to obtain multiple communities.

Benefits of technology

Reliance on high-performance large models is reduced, the ability to answer questions is improved, the ability to answer questions is obtained, the community distinction is obtained, and the uncertainty of large model output is reduced, providing a logical basis for the formation of geoscience communities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144780A_ABST
    Figure CN120144780A_ABST
Patent Text Reader

Abstract

The invention relates to a community discovery method and device for a mineral knowledge graph and a storage medium. According to the community discovery method of the mineral product knowledge graph, discrete knowledge base documents are aggregated, the dependence on a high-performance large model is reduced, and the ability of answering questions is improved. Compared with the traditional technology of exploring the weight value from the cue word, the method provided by the invention has the advantages that the quantitative mode of weight analysis is provided, the performance requirement is reduced, and better community distinguishing is also obtained. And finally, community aggregation and community data recombination are realized through modularity, a logic basis is provided for formation of geoscience communities, and the uncertainty of large model output is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer information processing, and more particularly, to a method, apparatus, and storage medium for community discovery of a mineral knowledge graph. Background Art

[0002] Large language models can quickly process and analyze a vast amount of literature related to mineral resources, accurately extract key information, and provide strong support for decision-makers. Specifically, large language models can assist researchers in quickly organizing core data such as the distribution, reserves, and exploitation status of mineral resources, and constructing a detailed knowledge graph in the field of mineral resources. Through intelligent analysis, the model can accurately identify and distinguish direct and indirect complex relationships in the knowledge graph, that is, achieve community discovery, thereby helping decision-makers comprehensively grasp the situation and providing a solid technical guarantee for the intelligent management of mineral resources.

[0003] However, the inventors found that traditional methods for constructing mineral resource knowledge graphs have problems such as high computing power dependence and high limitations. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, apparatus, and storage medium for community discovery of a mineral knowledge graph that can accurately obtain a small computing power dependence and good generalization in real time.

[0005] To achieve the above object, on the one hand, an embodiment of the present application provides a method for community discovery of a mineral knowledge graph, including the steps of:

[0006] Obtain a knowledge base, and perform text chunking on the knowledge base to obtain each text chunk;

[0007] Based on each text chunk, obtain the entities of the knowledge base;

[0008] According to each entity and the relationships between entities, construct a first entity relationship graph;

[0009] Obtain the first distance between any two entities within each text chunk, and based on the first distance, confirm the second distance between any two entities in the knowledge base;

[0010] Obtain the similarity between any two entities in the concept library;

[0011] The sum of the similarity and the second distance is confirmed as the weight value of the edge in the first entity relationship graph; where the edge is formed according to any two entities;

[0012] According to the weight value of the edge, obtain the modularity;

[0013] Based on the increment of the modularity, aggregate the nodes of the first entity relationship graph to obtain multiple communities.

[0014] In one embodiment, the step of confirming the second distance between any two entities in the knowledge base based on the first distance includes:

[0015] Set a truncation function;

[0016] Based on the truncation function and the first distance, obtain the second distance.

[0017] In one embodiment, in the step of obtaining the second distance based on the truncation function and the first distance, the second distance is obtained based on the following formula:

[0018]

[0019] Where WD is the second distance; d w is the first distance; δ c is the truncation function; all i is the maximum number of times any two entities appear simultaneously within a preset range.

[0020] In one embodiment, the step of obtaining the similarity between any two entities in the concept library includes:

[0021] Obtain the description texts corresponding to any two entities and the concept library;

[0022] Perform vectorization calculation on the description texts corresponding to any two entities to obtain vectorized entities;

[0023] Based on the vectorized entities, obtain the vector of the target concept in the concept library;

[0024] Based on the vectorized entities and the vector of the target concept, obtain the similarity between any two entities in the concept library.

[0025] In one embodiment, in the step of obtaining the similarity between any two entities in the concept library based on the vectorized entities and the vector of the target concept, the similarity is obtained based on the following formula:

[0026]

[0027] Where CD is the similarity; i is the category code of the target concept; allc represents the number of categories; is one of any two vectorized entities; is the other of any two vectorized entities; is the vector of the target concept of category i extracted according to one of the vectorized entities; according to is the vector of the target concept of category i extracted according to the other of the vectorized entities. W ci is the weight of the target concept with category code i.

[0028] In one embodiment, it further includes the steps of:

[0029] Obtain the description text corresponding to each community;

[0030] Summarize the description text corresponding to each community to obtain a community summary.

[0031] In one embodiment, it further includes the steps of:

[0032] Receive a user request;

[0033] Convert the request into a request vector;

[0034] Perform similarity matching on the request vector and the community summary to obtain a target community;

[0035] In the target community, extract entities that match the request vector.

[0036] In one embodiment, the step of performing text chunking on the knowledge base to obtain each text chunk includes:

[0037] Split the knowledge base to obtain multiple statements;

[0038] Perform vectorization processing on each statement to obtain multiple vectors;

[0039] Based on the cosine similarity of the vectors of adjacent statements, group the adjacent statements into the same text chunk or different text chunks.

[0040] On the one hand, the embodiment of the present invention provides a community discovery device for a mineral knowledge graph, including:

[0041] A chunking module, configured to obtain a knowledge base and perform text chunking on the knowledge base to obtain each text chunk;

[0042] An entity extraction module, configured to obtain the entities of the knowledge base based on each text chunk;

[0043] A construction module, configured to construct a first entity relationship graph according to each entity and the relationships between the entities;

[0044] A distance acquisition module, configured to obtain a first distance between any two entities within each text chunk, and confirm a second distance between any two entities in the knowledge base based on the first distance;

[0045] A similarity acquisition module, configured to obtain the similarity between any two entities in a concept library;

[0046] A weight value calculation module, configured to confirm the weight value of the edge in the first entity relationship graph as the sum of the similarity and the second distance; wherein, the edge is formed according to any two entities;

[0047] A modularity calculation module, configured to obtain modularity according to the weight values of edges;

[0048] An aggregation module, configured to aggregate the nodes of the first entity relationship graph based on the increment of modularity to obtain multiple communities.

[0049] On the other hand, the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps of the above method when running.

[0050] One of the above technical solutions has the following advantages and beneficial effects:

[0051] The above community discovery method for the mineral knowledge graph aggregates discrete knowledge base documents, reduces the dependence on high-performance large models, and improves the ability to answer questions. Compared with the traditional method of groping for weight values from prompt words, this method proposes a quantitative method for weight analysis, which not only reduces the performance requirements but also obtains better community differentiation. Finally, modularity is used to achieve the aggregation of communities and the reorganization of community data, providing a logical basis for the formation of geological communities and reducing the uncertainty of the output of large models. Description of the Drawings

[0052] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0053] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0054] Figure 1 It is a first schematic flowchart of the community discovery method for the mineral knowledge graph in an embodiment;

[0055] Figure 2 It is a schematic flowchart of the steps for obtaining the similarity of any two entities in the concept library in an embodiment;

[0056] Figure 3 It is a schematic flowchart of the steps for confirming the second distance of any two entities in the knowledge base based on the first distance in an embodiment;

[0057] Figure 4 It is a second schematic flowchart of the community discovery method for the mineral knowledge graph in an embodiment;

[0058] Figure 5A schematic flowchart of the steps for text chunking of a knowledge base to obtain each text chunk in an embodiment. Detailed implementation manners

[0059] To facilitate the understanding of the present application, the present application will be described more comprehensively below with reference to the relevant accompanying drawings. Embodiments of the present application are shown in the accompanying drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided so that the disclosure of the present application is thorough and comprehensive.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0061] In subsequent descriptions, suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of description of this application and have no specific meaning in themselves. Therefore, "module" and "component" can be used interchangeably.

[0062] It can be understood that in the following embodiments, "connection", if there is an electrical signal or data transmission between the connected circuits, modules, units, etc., should be understood as "electrical connection", "communication connection", etc.

[0063] As used herein, the singular forms "a", "an", and "the" may also include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprise / include" or "have" etc. specify the presence of the stated features, wholes, steps, operations, components, parts, or combinations thereof, but do not preclude the possibility of the presence or addition of one or more other features, wholes, steps, operations, components, parts, or combinations thereof.

[0064] Currently, traditional means for constructing a mineral resource knowledge graph mainly rely on relational databases (such as trade databases) or are based on small-scale datasets and small deep learning models. However, the inherent limitations of these methods have led to the following series of problems:

[0065] 1. Limitations in cross-text entity relationship extraction: When dealing with information across paragraphs, sentences, and even documents, these methods are difficult to accurately capture and establish the connections between core knowledge points. This defect greatly hinders the discovery of general knowledge, making it difficult for us to comprehensively extract key information from multiple documents.

[0066] Data update lag: Due to the cumbersome and time-consuming training process of deep learning models and the complex construction process, the data update speed is slow. This means that a batch of data often takes a long time cycle from collection to completion of model training and deployment.

[0067] 2. Insufficient model generality: Given that traditional methods usually adopt small models, their application scope is often limited to specific fields. In the current technical context where large models are emerging continuously, small models are difficult to meet the needs of cross-domain and general knowledge graph construction.

[0068] 3. Application development challenges: Once the knowledge graph is constructed, subsequent application development also faces many challenges. In particular, when implementing the question-answering query function based on natural language, due to the technical limitations of traditional methods, it is often difficult to quickly and efficiently develop applications that meet user needs.

[0069] In summary, the traditional methods for constructing mineral resource knowledge graphs have significant deficiencies in cross-text relationship extraction, data timeliness, model generality, and application development. The community discovery method for mineral knowledge graphs provided in this application can effectively solve the above problems.

[0070] In one embodiment, as Figure 1 shown, a community discovery method for a mineral knowledge graph is provided, including the steps:

[0071] S110, obtain a knowledge base, and perform text chunking on the knowledge base to obtain each text chunk;

[0072] Specifically, the knowledge base can be obtained by quickly processing and analyzing a large number of mineral resource-related documents through a language model. The content of the knowledge base is divided into smaller and more manageable text chunks according to a certain logic or rule. This helps with information extraction and processing in subsequent steps. In one example, natural paragraphs can also be obtained through format features to form text chunks. In another specific example, a vector model can be used for text chunk division, which can make the text chunks aggregate semantically, avoiding the phenomenon of interrupted semantics in text chunks to a certain extent, and also avoiding division errors caused by the often very diverse formats, layouts, and punctuation styles of paragraph division.

[0073] S120, obtain the entities of the knowledge base based on each text chunk;

[0074] Specifically, extract key entities from each text block. Entities can be noun phrases, personal names, place names, organization names, etc. Further, in this step, through the understanding ability of the large language model, entity extraction is carried out one by one with the text block as the unit to obtain entity texts. After extraction, the entity texts are first cleaned through regular matching, and then the entity model is secondarily cleaned through the vector model to completely eliminate duplicate information to obtain the final entities. The entity extraction of the knowledge base determines the division of the community.

[0075] S130. Construct a first entity relationship graph according to each entity and the relationships between entities.

[0076] Specifically, analyze the relationships between entities in the text block, which can be identified and confirmed through co-occurrence, semantic similarity, or specific relationship templates. Construct a preliminary entity relationship graph based on the identified entities and their relationships, where nodes represent entities and edges represent relationships.

[0077] S140. Obtain the first distance between any two entities within each text block, and confirm the second distance between any two entities in the knowledge base based on the first distance.

[0078] Among them, the first distance refers to the distance between two entities within different text blocks, and the second distance refers to the sum of the distances between two entities in the knowledge base. It should be noted that the relevance of entities can be continuously improved as follows: a. Two entities appear in the same article. b. Two entities appear multiple times in the same article. c. Two entities appear in the same article but in different text blocks. d. Two entities appear in the same article, in different text blocks, and appear multiple times in each text block. e. Two entities appear in the same article, in different text blocks, and appear multiple times in each paragraph, and also appear simultaneously in multiple sentences in each text block. Further, the distance between any two entities can be calculated based on the lexical distance or sentence-level distance within the text block, or can be calculated based on other technical means in this field.

[0079] S150. Obtain the similarity between any two entities in the concept library.

[0080] Specifically, in some cases, two entities do not appear in the same paragraph of text, but they belong to the concepts of a certain category. And in this category, there is a certain connection between the two entities. Further, a specialized geoscience professional knowledge base is introduced as the concept library. In a specific example, such as Figure 2As shown, the steps to obtain the similarity between any two entities in the concept library include: S210, obtaining the description texts corresponding to any two entities and the concept library; S220, performing vectorization calculation on the description texts corresponding to any two entities to obtain vectorized entities; S230, based on the vectorized entities, obtaining the vector of the target concept in the concept library; S240, based on the vectorized entities and the vector of the target concept, obtaining the similarity between any two entities in the concept library. That is, the entities are generalized to obtain entity generalizations (i.e., the description texts corresponding to the above entities), and then the content in the entity generalizations is vectorized by an embedding model. The similarities between them and related concepts are extracted from the concept library, and then the similarities of the concepts they extract are compared. Finally, the comparisons in all categories are combined together.

[0081] Specifically, in the step of obtaining the similarity between any two entities in the concept library based on the vectorized entities and the vector of the target concept, the similarity is obtained based on the following formula:

[0082]

[0083] where CD is the similarity; i is the category code of the target concept; allc represents the number of categories; is one of any two vectorized entities; is the other of any two vectorized entities; is the vector of the target concept of category i extracted according to one of the vectorized entities; according to is the vector of the target concept of category i extracted according to the other vectorized entity. W ci is the weight of the target concept with category code i.

[0084] S160, confirming the sum of the similarity and the second distance as the weight value of the edge in the first entity relationship graph; where the edge is formed according to any two entities;

[0085] Specifically, adding the similarity and the second distance to obtain the weight value of each edge in the first entity relationship graph, which reflects the degree of association between entities.

[0086] S170, obtaining the modularity according to the weight value of the edge;

[0087] where the modularity is an index to measure the quality of community division. It reflects the degree of connection tightness between nodes within a community relative to the degree of connection tightness between nodes outside the community. In this step, the modularity of the entire entity relationship graph is calculated according to the weight value of the edge.

[0088] S180, aggregating the nodes of the first entity relationship graph based on the increment of the modularity to obtain multiple communities.

[0089] Specifically, the nodes of the first entity relationship diagram are aggregated through the modularity formula and the modularity increment formula.

[0090] Modularity:

[0091] Modularity increment:

[0092] Among them, ∑ in represents the sum of the weights of all edges in community C; ∑ tot represents the sum of the weights of all variable edges pointing to community C; k i represents the sum of the weights of all edges pointing to node i; k i,in represents the sum of the weights of the edges between node i and community C; m represents the sum of the weights of all edges.

[0093] In S160, the weight of each edge has been obtained. With the weight of each edge, the sum of the weights of all edges m can be calculated. Then, among all the nodes, the node with the largest △Q is found and aggregated into a community. In the first step of community aggregation, each node is an independent community, so ∑ tot , which is the sum of the weights of the relationships between a node and all other nodes, that is, the sum of the weights of the edges between its other nodes:

[0094]

[0095] From the formula of △Q, we can see that in order to complete the aggregation, the target node k i,in and the weight value of the initial node The difference between k determines the chance of these two nodes forming a community. i It represents the sum of the weights of all edges pointing to node i, which is the target node. Therefore, when the weight of the relationship between the target node and other nodes is small, and the weight of the relationship with the initial node is large enough, the initial node and the target node are easy to form a community. One round of iteration can gather some nodes into a community. The community acts as a new node and continues to aggregate with the nodes formed by other communities according to the formula of △Q, iterating back and forth until no new community can be formed, that is, △Q is less than a certain threshold, which is usually 0. For a community with multiple nodes, the ki,in value is relatively large, so it is easier to aggregate with new nodes into a new community. In this way, the community gradually completes aggregation.

[0096] The above community discovery method for the mineral knowledge graph aggregates discrete knowledge base documents, reduces the dependence on high-performance large models, and improves the ability to answer questions. Compared with the traditional method of exploring weight values from prompt words, this method proposes a quantitative method for weight analysis, which not only reduces the performance requirements but also obtains better community differentiation. Finally, modularity is used to achieve community aggregation and community data reorganization, providing a logical basis for the formation of the geoscience community and reducing the uncertainty of the large model output.

[0097] In one embodiment, as Figure 3 shown, the step of confirming the second distance between any two entities in the knowledge base based on the first distance includes:

[0098] S310, set a truncation function;

[0099] Specifically, the purpose of setting the truncation function is that when d w is greater than the preset threshold, this distance is not used. Since the previous summation operation has taken into account the situation where entities appear in different text blocks, this truncation function needs to be added to eliminate the interference caused by the cross-text distance of entities.

[0100] S320, based on the truncation function and the first distance, obtain the second distance.

[0101] In one of the embodiments, in the step of obtaining the second distance based on the truncation function and the first distance, the second distance is obtained based on the following formula:

[0102]

[0103] where WD is the second distance; d w is the first distance; δ c is the truncation function; all i is the maximum number of times any two entities appear simultaneously within the preset range.

[0104] Furthermore, the weight value W l of the edge in the first entity relationship graph is:

[0105]

[0106] In one of the embodiments, it further includes the steps of:

[0107] Obtain the description text corresponding to each community;

[0108] Summarize the description text corresponding to each community to obtain a community summary.

[0109] Specifically, after the knowledge base is abstracted and divided into several communities, these initially aggregated communities are only composed of entities and lack textual descriptions, making it difficult to perform subsequent operations such as retrieval-augmented generation. To solve this problem, descriptive texts of each entity are extracted. On this basis, the capabilities of large models are also utilized to deeply summarize and integrate the textual descriptions of all entities within the community, thereby forming a summary description for each community. In this way, each community will have its own exclusive descriptive text, providing a solid foundation and convenience for subsequent knowledge graph searches. This step not only enriches the information content of the community but also greatly improves the practicality and operability of the knowledge graph.

[0110] In one embodiment, as Figure 4 shown, it further includes the steps:

[0111] S410, receiving a user's request;

[0112] S420, converting the request into a request vector;

[0113] S430, performing a similarity match between the request vector and the community summary to obtain a target community;

[0114] S440, extracting entities that match the request vector within the target community.

[0115] Specifically, the aggregated communities and their corresponding entities together constitute the core structure of the knowledge graph, where the communities serve as tree nodes and the entities serve as leaf nodes, with a clear hierarchy and logical clarity. When performing RAG (Retrieval-Augmented Generation) operations, we first use an advanced vector model to convert the user's query request into a request vector, and then perform a cosine similarity match between this vector and the summary descriptions of each community, thereby accurately extracting the communities that are highly similar to the user's request. Immediately afterwards, within these selected communities, we use similarity matching techniques again to carefully extract the entities that best match the user's request.

[0116] Compared with naive RAG operations and even RAG equipped with a re-ranker, our method can extract text data more extensively, significantly improving the quality and depth of the answers of large models when dealing with comprehensive questions. In addition, the aggregation of communities not only enhances the efficiency and accuracy of data retrieval but also opens up new thinking paths for researchers, helping them deeply explore the potential laws and internal connections of the research object. This innovative method not only optimizes the application experience of the knowledge graph but also brings new inspirations and possibilities to the fields of academic research and data analysis.

[0117] In one embodiment, as Figure 5As shown in the figure, the steps of text chunking the knowledge base to obtain each text chunk include:

[0118] S510, split the knowledge base to obtain multiple statements;

[0119] S520, perform vectorization processing on each statement to obtain multiple vectors;

[0120] S530, based on the cosine similarity of the vectors of adjacent statements, group adjacent statements into the same text chunk or different text chunks.

[0121] Specifically, first chunk the knowledge base text by sentence, vectorize the previous and subsequent sentences through an embedding model, and then calculate the cosine similarity of these two vectors. Then set a preset threshold to determine whether these two sentences are similar. If they are similar, aggregate the two sentences together. If they are not similar, the subsequent sentence forms a new text chunk, and then repeat this step until the text chunking is completed. Through the above method, the text chunks can be aggregated semantically, and to a certain extent, the phenomenon of interrupted semantics in the text chunks can be avoided.

[0122] In one embodiment, a community discovery device for a mineral knowledge graph includes:

[0123] A chunking module for obtaining a knowledge base and performing text chunking on the knowledge base to obtain each text chunk;

[0124] An entity extraction module for obtaining the entities of the knowledge base based on each text chunk;

[0125] A construction module for constructing a first entity relationship graph according to each entity and the relationships between entities;

[0126] A distance acquisition module for obtaining a first distance between any two entities within each text chunk and confirming a second distance between any two entities in the knowledge base based on the first distance;

[0127] A similarity acquisition module for obtaining the similarity between any two entities in the concept library;

[0128] A weight value calculation module for confirming the weight value of the edge in the first entity relationship graph as the sum of the similarity and the second distance; where the edge is formed according to any two entities;

[0129] A modularity calculation module for obtaining the modularity according to the weight value of the edge;

[0130] An aggregation module for aggregating the nodes of the first entity relationship graph based on the increment of the modularity to obtain multiple communities.

[0131] For the specific limitations of the community discovery device for the mineral knowledge graph, reference can be made to the limitations of the community discovery method for the mineral knowledge graph in the above text, which will not be elaborated here. Each module in the above community discovery device for the mineral knowledge graph can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0132] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0133] Obtain a knowledge base, and perform text chunking on the knowledge base to obtain each text chunk;

[0134] Based on each text chunk, obtain the entities of the knowledge base;

[0135] According to each entity and the relationships between entities, construct a first entity relationship graph;

[0136] Obtain the first distance between any two entities within each text chunk, and based on the first distance, confirm the second distance between any two entities in the knowledge base;

[0137] Obtain the similarity between any two entities in the concept library;

[0138] Confirm the sum of the similarity and the second distance as the weight value of the edge in the first entity relationship graph; where the edge is formed according to any two entities;

[0139] According to the weight value of the edge, obtain the modularity;

[0140] Based on the increment of the modularity, aggregate the nodes of the first entity relationship graph to obtain multiple communities.

[0141] In one embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are implemented:

[0142] Obtain a knowledge base, and perform text chunking on the knowledge base to obtain each text chunk;

[0143] Based on each text chunk, obtain the entities of the knowledge base;

[0144] According to each entity and the relationships between entities, construct a first entity relationship graph;

[0145] Obtain the first distance between any two entities within each text block, and confirm the second distance between any two entities in the knowledge base based on the first distance;

[0146] Obtain the similarity between any two entities in the concept library;

[0147] Confirm the sum of the similarity and the second distance as the weight value of the edge in the first entity relationship graph; wherein, the edge is formed according to any two entities;

[0148] Obtain the modularity according to the weight value of the edge;

[0149] Aggregate the nodes of the first entity relationship graph based on the increment of the modularity to obtain multiple communities.

[0150] In the specific implementation of the embodiments of the present application, reference may be made to the above-mentioned various embodiments, and corresponding technical effects are achieved.

[0151] It can be understood that these embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, or other electronic units for performing the functions described in the present application, or a combination thereof.

[0152] For software implementation, the technologies described herein can be implemented by units that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented inside or outside the processor.

[0153] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented with electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0154] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0155] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0156] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0158] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A community discovery method for mineral knowledge graph, characterized in that: include: Acquire a knowledge base, and divide the knowledge base into text blocks to obtain text blocks; Based on each of the text blocks, obtaining entities of the knowledge base; Constructing a first entity relationship graph according to the entities and the relationships between the entities; Acquire a first distance between any two of the entities within each of the text blocks, and determine a second distance between any two of the entities in the knowledge base based on the first distance; Obtaining the similarity between any two entities and the concept library; Confirming the sum of the similarity and the second distance as the weight value of the edge in the first entity relationship graph; wherein the edge is formed based on any two of the entities; According to the weight value of the edge, the modularity is obtained; Based on the increment of the modularity, the nodes of the first entity relationship graph are aggregated to obtain a plurality of communities.

2. The community discovery method of mineral knowledge graph according to claim 1, characterized in that: The step of confirming a second distance between any two entities in the knowledge base based on the first distance comprises: Set the truncation function; The second distance is obtained based on the truncation function and the first distance.

3. The community discovery method of mineral knowledge graph according to claim 2 is characterized in that: In the step of obtaining the second distance based on the cutoff function and the first distance, the second distance is obtained based on the following formula: Wherein, WD is the second distance; d w is the first distance; δ c is the truncation function; all i is the maximum number of times any two entities appear simultaneously within a preset range.

4. The community discovery method of mineral knowledge graph according to claim 1, characterized in that: The step of obtaining the similarity between any two entities and the concept library includes: Obtaining description texts and concept libraries corresponding to any two of the entities; Performing vectorization calculation on the description texts corresponding to any two entities to obtain vectorized entities; Based on the vectorized entity, obtaining a vector of a target concept from the concept library; Based on the vectorized entities and the vectors of the target concepts, the similarity between any two entities and the concept library is obtained.

5. The community discovery method of mineral knowledge graph according to claim 4 is characterized in that: In the step of obtaining the similarity between any two entities and the concept library based on the vectorized entity and the vector of the target concept, the similarity is obtained based on the following formula: Wherein, CD is the similarity; i is the category code of the target concept; allc represents the number of categories; is one of any two of the vectorized entities; is the other of any two of the vectorized entities; is the vector of the target concept of category i extracted according to one of the vectorized entities; is the vector of the target concept of category i extracted from another vectorized entity; W ci is the weight of the target concept with category code i.

6. The community discovery method of mineral knowledge graph according to claim 1, characterized in that: Also includes the steps: Obtaining description text corresponding to each of the communities; The description texts corresponding to the communities are summarized to obtain a community summary.

7. The community discovery method of mineral knowledge graph according to claim 6, characterized in that: Also includes the steps: Receiving a user's request; converting the request into a request vector; Performing similarity matching on the request vector and the community summary to obtain a target community; In the target community, entities matching the request vector are extracted.

8. The community discovery method of mineral knowledge graph according to claim 1, characterized in that: The step of dividing the knowledge base into text blocks to obtain each text block includes: Splitting the knowledge base to obtain multiple statements; Vectorizing each of the statements to obtain multiple vectors; Based on the cosine similarity of their vectors, adjacent sentences are classified as the same text block or different text blocks.

9. A community discovery device for a mineral knowledge graph, characterized in that: include: A block segmentation module is used to obtain a knowledge base and segment the knowledge base into text blocks to obtain text blocks; An entity extraction module, used for obtaining entities of a knowledge base based on each of the text blocks; A construction module, used for constructing a first entity relationship graph according to the entities and the relationships between the entities; A distance acquisition module, used for acquiring a first distance between any two entities in each text block, and determining a second distance between any two entities in the knowledge base based on the first distance; A similarity acquisition module, used to acquire the similarity between any two entities in the concept library; A weight value calculation module, used to confirm the sum of the similarity and the second distance as the weight value of the edge in the first entity relationship graph; wherein the edge is formed based on any two of the entities; A modularity calculation module, used for obtaining the modularity according to the weight value of the edge; An aggregation module is used to aggregate the nodes of the first entity relationship graph based on the increment of the modularity to obtain multiple communities.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the steps of the method described in any one of claims 1 to 8 when executed.

Citation Information

Patent Citations

  • Community discovery method and device based on entity similarity in knowledge graph

    CN108959370A

  • RAG question and answer method and system based on knowledge graph and medium

    CN118673126A

  • Knowledge discovery method and device based on large language model and knowledge graph, and medium

    CN119226466A

  • Method and apparatus for interactive evolutionary optimization of concepts

    WO2014143729A1

  • Method and apparatus for acquiring character, page processing method, method for constructing knowledge graph, and medium

    WO2021102632A1