Knowledge question-answering method and system combining enhanced retrieval and knowledge graph
By adopting a knowledge question-and-answer method that enhances the combination of search and knowledge graphs in the coal industry, the shortcomings of traditional search systems in handling complex queries and semantic understanding are solved, and more efficient and accurate information retrieval and analysis are achieved.
Patent Information
- Application Number
- CN202510473342.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional search systems are difficult to cope with the complex query needs and semantic understanding in the coal industry, especially when processing a large number of professional terms and unstructured data, it is difficult to accurately identify and recall synonyms and synonyms, and lack in-depth causal analysis.
The knowledge question-and-answer method is adopted to combine enhanced search and knowledge graphs, and the entities are expanded through the pre-constructed data mapping model and the knowledge graph of the coal preparation plant, and the entities to be searched are obtained, and the search enhancement technology is used for searching. Finally, the search results are input into the big model to generate query results.
It improves the accuracy and efficiency of information retrieval, can accurately identify and recall synonyms, synonyms and colloquial entities, enhances semantic understanding and correlation analysis capabilities, supports complex queries and in-depth causal relationship analysis.
Smart Images

Figure CN119990338A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a knowledge question answering method and system combining enhanced retrieval and knowledge graph. Background Art
[0002] In the coal industry, the amount of information is huge and complex, including coal quality data, production reports, production records, safety regulations, etc. This information usually exists in unstructured or semi-structured form. Traditional retrieval systems rely on keyword matching and are difficult to cope with complex query requirements and semantic understanding, which makes it difficult for traditional retrieval systems to provide the required information efficiently and accurately. Summary of the invention
[0003] In view of this, the purpose of the present invention is to provide a knowledge question answering method and system that enhances the combination of retrieval and knowledge graph, so as to improve the accuracy and efficiency of information retrieval.
[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a knowledge question-answering method combining enhanced retrieval and knowledge graph, which is applied to the coal industry, and includes: obtaining a query question or query document input by a user, and extracting entities of the query question or query document; wherein the entity includes one or more of the following: equipment name, process flow and fault description; expanding the entity based on a pre-built data mapping model and a knowledge graph of a coal preparation plant to obtain an entity to be retrieved; wherein the entity to be retrieved includes: an entity, one or more of a standardized entity corresponding to the entity, a synonym entity, a near-synonymous entity and a colloquial entity, as well as an associated equipment entity and an associated process flow entity of the entity; using retrieval enhancement to retrieve the entity to be retrieved to obtain a retrieval result; inputting the retrieval result, the query question or query document and the constructed prompt word into a large model to generate a query result of the query question or query document.
[0005] Optionally, the data mapping model includes: mapping relationships between standardized entities, synonym entities, near-synonymous entities and colloquial entities of equipment, process flow, and faults; the knowledge graph includes: process flow network, entity triple relationship, and preset constraint rules; among which, the entity triple relationship is used to characterize the upstream and downstream relationships, causal relationships, and timing relationships of equipment; the preset constraint rules include at least: fault propagation rules and operation rules.
[0006] Optionally, the entities are expanded based on a pre-built data mapping model and the knowledge graph of the coal preparation plant to obtain entities to be retrieved, including: based on the pre-built data mapping model, searching for one or more of the standardized entities, synonym entities, near-synonymous entities and colloquial entities corresponding to the entity; searching for the associated process flow entities of the entity based on the process flow network in the knowledge graph, and searching for the associated equipment entities of the entity based on the entity triple relationship and preset constraint rules in the knowledge graph.
[0007] Optionally, retrieval enhancement is used to search the entity to be searched to obtain retrieval results, including: performing vector retrieval, keyword retrieval and knowledge graph retrieval on the entity to be searched respectively to obtain multiple recall results; re-ranking the multiple recall results based on the relevance between the multiple recall results and the query question or query document, and selecting a preset number of recall results as retrieval results based on the re-ranking results.
[0008] Optionally, extracting entities of the query question or query document includes: extracting entities of the query question or query document using natural language processing technology or based on a predefined vocabulary.
[0009] Optionally, after extracting the entities of the query question or query document, the method further includes: aligning the entities and obtaining entities related to the entities based on preset entity relationships.
[0010] Optionally, the knowledge graph also includes: equipment abnormalities and processing steps; the method also includes: obtaining equipment abnormality information, and determining the causes and processing steps of the equipment abnormality information based on historical data and the knowledge graph, and generating an equipment failure report.
[0011] In the second aspect, the present invention provides a knowledge question and answer system combining enhanced retrieval and knowledge graph, which is applied to the coal industry, including: an entity extraction module, used to obtain query questions or query documents input by users, and extract entities of the query questions or query documents; wherein the entities include one or more of the following: equipment name, process flow and fault description; an entity expansion module, used to expand the entities based on a pre-built data mapping model and the knowledge graph of the coal preparation plant to obtain entities to be retrieved; wherein the entities to be retrieved include: entities, standardized entities corresponding to the entities, synonym entities, near-synonymous entities and colloquial entities, as well as the entity's associated equipment entities and associated process flow entities; a retrieval module, used to use retrieval enhancement to retrieve the entities to be retrieved and obtain retrieval results; a question and answer module, used to input the retrieval results, query questions or query documents and constructed prompt words into a large model to generate query results of the query questions or query documents.
[0012] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of any one of the methods provided in the first aspect above.
[0013] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any one of the methods provided in the first aspect are executed.
[0014] The present invention brings the following beneficial effects: The knowledge question-answering method and system combining the enhanced retrieval and knowledge graph provided by the present invention are applied to the coal industry. First, the query question or query document input by the user is obtained, and the entities of the query question or query document are extracted; wherein the entity includes one or more of the following: equipment name, process flow and fault description; then, based on the pre-constructed data mapping model and the knowledge graph of the coal preparation plant, the entity is expanded to obtain the entity to be retrieved; wherein, the entity to be retrieved includes: entity, standardized entity corresponding to the entity, synonym entity, near-synonymous entity and one or more of the colloquial entity, as well as the entity's associated equipment entity and associated process flow entity; then, the entity to be retrieved is retrieved using retrieval enhancement to obtain the retrieval result; finally, the retrieval result, the query question or query document and the constructed prompt word are input into the large model to generate the query result of the query question or query document. In the above method, entities are expanded through the pre-built data mapping model and the knowledge graph of the coal preparation plant, which can accurately identify and recall synonym entities, near-synonymous entities and colloquial entities, thereby improving the accuracy of information retrieval; at the same time, by enhancing the combination of retrieval and knowledge graph, complex queries can be processed, the semantic understanding and association analysis capabilities are improved, and the accuracy and efficiency of retrieval are thereby improved.
[0015] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the preferred embodiments are specifically mentioned below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A flowchart of a knowledge question answering method combining enhanced retrieval and knowledge graph provided by an embodiment of the present invention; Figure 2 An information processing flow of a knowledge question answering method combining enhanced retrieval and knowledge graph provided by an embodiment of the present invention; Figure 3 A system framework schematic diagram provided for an embodiment of the present invention; Figure 4 A schematic diagram of the structure of a knowledge query system combining enhanced retrieval and knowledge graph provided by an embodiment of the present invention; Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] At present, traditional retrieval systems mainly rely on keyword matching for information retrieval. Although this method is simple and direct, it does not work well when faced with complex query requirements and unstructured data. With the development of information technology, knowledge graph technology has gradually attracted attention because of its ability to represent information graphically. In the coal industry, some systems have tried to use knowledge graphs to improve the quality of information retrieval. However, these systems often lack deep semantic understanding and association analysis capabilities and cannot fully meet user needs. Therefore, the existing retrieval methods have the following problems: for the coal industry, which contains a large number of professional terms, it is difficult for existing systems to accurately identify and recall synonyms and antonyms; it is unable to effectively handle domain-specific language variants, which limits the scope of application of the system; due to the lack of sufficient background knowledge support, the existing question-answering system is prone to produce incorrect answers; it is unable to provide detailed causal analysis and the display of potential influencing factors, which is not conducive to in-depth understanding and decision support.
[0021] Based on this, an embodiment of the present invention provides a knowledge question answering method and system that combines enhanced retrieval and knowledge graph, which can improve the accuracy and efficiency of information retrieval.
[0022] In the embodiment of the present invention, a data mapping model and a knowledge graph of a coal preparation plant can be pre-constructed. The data mapping model includes: mapping relationships between standardized entities, synonym entities, near-synonymous entities and colloquial entities of equipment, process flow, and faults; the knowledge graph includes: process flow network, entity triple relationship, and preset constraint rules; the entity triple relationship is used to characterize the upstream and downstream relationship, causal relationship, and temporal relationship of the equipment; the preset constraint rules include at least: fault propagation rules and operation rules.
[0023] In specific implementation, the various production process links in the coal production process (such as mining, transportation, processing, etc.) and their interrelationships can be constructed into the knowledge graph to form a complete process flow network, which is convenient for users to understand and optimize the entire production chain.
[0024] The knowledge graph also stores entity triple relationships and preset constraint rules. The entity triple relationship (subject-relationship-object) includes: upstream and downstream relationships, causal relationships, and temporal relationships of devices. For example: (1) Entity triple relationship: Vibrating screen - [connection] -> crusher, crusher - [connection] -> jig, vibrating screen blockage - [may lead to] -> crusher idling, crusher overload - [may lead to] -> jig feed shortage, heavy medium system - [includes] -> three-product heavy medium cyclone, heavy medium system - [includes] -> two-product heavy medium cyclone, three-product heavy medium cyclone - [includes] -> pressurized three-product heavy medium cyclone, three-product heavy medium cyclone - [includes] -> non-pressurized three-product heavy medium cyclone, two-product heavy medium cyclone - [includes] -> pressurized two-product heavy medium cyclone, two-product heavy medium cyclone - [includes] -> non-pressurized two-product heavy medium cyclone.
[0025] (2) Example of preset constraint rules: Fault propagation rule: IF vibrating screen is clogged THEN crusher feed rate drops by 50%; IF crusher current > rated value THEN jig separation efficiency drops by 30%.
[0026] Operation rules: IF the vibrating screen stops THEN stop the crusher immediately (to avoid idling wear).
[0027] In addition, the knowledge graph also includes: equipment anomalies and processing steps. When querying equipment anomalies, the cause of the equipment anomaly and the processing steps can be analyzed based on the knowledge graph to generate a fault handling report.
[0028] To facilitate understanding of this embodiment, a knowledge question answering method combining enhanced retrieval and knowledge graph disclosed in an embodiment of the present invention is first introduced in detail. The method can be executed by an electronic device, such as a smart phone, a computer, a tablet computer, etc. Figure 1 The flowchart of a knowledge question answering method combining enhanced retrieval and knowledge graph is shown, indicating that the method mainly includes the following steps S101 to S104: Step S101: obtaining a query question or a query document input by a user, and extracting entities of the query question or the query document.
[0029] In one implementation, key entities are extracted from a query question or query document input by a user, and these entities include one or more of the following: equipment name, process flow, fault description, or professional terminology, etc. In an embodiment of the present invention, natural language processing technology or a predefined vocabulary can be used to extract entities of the query question or query document.
[0030] After the entities are extracted, the entities are aligned, and the preset entity relationships are used to obtain entities related to the entities. In the specific implementation, the entities extracted from different sources are compared and matched to solve the problem of inconsistent representation of the same entity in different documents or data sources. For example, "clean coal with gangue" and "coal with gangue" are regarded as the same entity, and "belt" and "belt conveyor" are regarded as the same entity. Furthermore, based on the extracted entities, other entities related to them are automatically obtained and included in the retrieval scope, which helps to broaden the search field of view and increase the chances of finding relevant information.
[0031] For example: the user's query question is: possible causes and solutions for the presence of gangue in the refined coal of a three-product heavy medium cyclone. In this question, the key entity parts in the question are first identified, which generally include the equipment name, process flow and fault phenomenon. Corresponding to this question, the extracted entities include: three-product heavy medium cyclone, and gangue in refined coal. Entity alignment is to align the professional terms in the document. For example, the document is generally described as refined coal with gangue, and the entity will automatically align to the standardized entity "gangue in coal". In this process, the three-product heavy medium cyclone actually also includes pressurized three-product heavy medium cyclone and unpressurized three-product heavy medium cyclone. After entity expansion, this information can be used as a supplementary entity answer.
[0032] Step S102: Expand the entity based on the pre-built data mapping model and the knowledge graph of the coal preparation plant to obtain the entity to be retrieved.
[0033] In one embodiment, the entities to be retrieved include: an entity, a standardized entity corresponding to the entity, one or more of a synonym entity, a near-synonymous entity and a colloquial entity, and an associated equipment entity and an associated process flow entity of the entity.
[0034] In the specific implementation, firstly, based on the pre-built data mapping model, one or more standardized entities, synonym entities, near-synonymous entities and colloquial entities corresponding to the entity are searched. Specifically, the extracted entity is first compared with the pre-built data mapping model to determine the type of the entity, that is, which expression of the extracted entity belongs to the standardized entity, synonym entity, near-synonymous entity and colloquial entity; then the corresponding entities of other types are obtained to expand the entity. For example: if the extracted entity is a standardized entity, the corresponding synonym entity, near-synonymous entity and colloquial entity are searched from the pre-built data mapping model as the entity to be detected; if the extracted entity is a colloquial entity, the corresponding synonym entity, near-synonymous entity and standardized entity are searched from the pre-built data mapping model as the entity to be detected.
[0035] Then, the associated process flow entities of the entity are searched based on the process flow network in the knowledge graph, and the associated equipment entities of the entity are searched based on the entity triple relationship and preset constraint rules in the knowledge graph.
[0036] In specific implementation, the process flow network, triple relationship and preset constraint rules in the knowledge graph can be used to expand the original query content and increase the diversity and depth of the query questions. For example, if a user asks about the function of a certain device, the process flow involved in the device and the collaboration information with other devices can be provided according to the relevant process flow network, triple relationship and preset constraint rules. For example, if a user inquires about the function of a crusher, the upstream and downstream devices of the crusher can be determined according to the entity triple relationship: vibrating screen-[connection]->crusher, crusher-[connection]->jig (i.e., associated device entities); the process flow in which the crusher is located and the related process flows (i.e., associated process flow entities) can also be determined according to the process flow network; and the fault propagation and operation rules of the crusher can be determined according to the preset constraint rules (i.e., associated device entities), such as: IF the vibrating screen is blocked THEN the feed amount of the crusher decreases by 50%, IF the crusher current>rated value THEN the jig sorting efficiency decreases by 30%, IF the vibrating screen is stopped THEN the crusher is stopped immediately.
[0037] In an embodiment of the present invention, standardized entities, synonym entities, near-synonymous entities and colloquial entities found through a data mapping model, as well as associated equipment entities and associated process flow entities obtained through knowledge topology expansion can be added to the original entities to obtain entities to be retrieved.
[0038] Step S103: Use search enhancement to search for the entity to be searched and obtain search results.
[0039] In one embodiment, a retrieval-augmented generation (RAG) technique is used, including: retrieving knowledge fragments through vector recall (big2small) and word segmentation recall, fusing the recall results after rough recall, and then finely ranking the recall results through a rerank model, and obtaining the retrieval results in combination with a topN algorithm, wherein N is dynamic and its value is determined by a fragment relevance threshold, thereby achieving enhanced retrieval.
[0040] Step S104: input the search results, query questions or query documents and constructed prompt words into the big model to generate query results of the query questions or query documents.
[0041] In one implementation, the obtained search results, query questions or query documents, and constructed prompt words are input into a large model to obtain complete query results.
[0042] The knowledge question-answering method combining the enhanced retrieval and knowledge graph provided by the present invention can accurately identify and recall synonym entities, near-synonymous entities and colloquial entities by expanding entities through a pre-constructed data mapping model and the knowledge graph of a coal preparation plant, thereby improving the accuracy of information retrieval; at the same time, by combining enhanced retrieval with the knowledge graph, it can handle complex queries, improve semantic understanding and association analysis capabilities, and thereby improve the accuracy and efficiency of retrieval.
[0043] In one embodiment, for the aforementioned step S103, that is, when using retrieval enhancement to search for the entity to be searched and obtain the retrieval results, the following methods can be used, including but not limited to: first, vector retrieval, keyword retrieval and knowledge graph retrieval are performed on the entity to be searched respectively to obtain multiple recall results; then, the multiple recall results are rearranged based on the relevance between the multiple recall results and the query question or query document, and a preset number of recall results are selected as retrieval results based on the rearranged results.
[0044] In the specific implementation, relevant databases of the coal industry are established in advance, including: vector index, inverted index, graph database and relational database for unstructured data; vector database, relational database and graph database for structured data, and then retrieval is performed based on the established database.
[0045] For vector retrieval, first convert the entity to be retrieved into a vector form. This can be achieved through a trained model (such as Word2Vec, BERT, etc.), which can map text into a high-dimensional vector space. Then, in the pre-built index library, find the vector that is most similar to the vector of the entity to be retrieved. Cosine similarity or Euclidean distance can be used as a similarity measure; finally, sort by similarity and return a list of entities most relevant to the entity to be retrieved. Among them, the index library includes: vector index and inverted index, and the index structure includes KD tree, LSH (local sensitive hashing), IVF (inverted file), etc.
[0046] For keyword search, first search for documents or entities containing the entity to be searched in the pre-built inverted index. The inverted index is a data structure that records in which documents each word appears. Then calculate the relevance score of each matching document, which can be done using TF-IDF weights, BM25 algorithms, etc. Finally, sort by relevance score and return the most relevant entity or document list.
[0047] For knowledge graph retrieval, first identify the entity to be retrieved in the user's query and its location in the knowledge graph. Then, based on the identified entity, traverse along the relationship path in the knowledge graph to find other entities or information related to the query. Combined with the ontology and semantic rules in the knowledge graph, we understand the true intent of the query and ensure the accuracy of the retrieval results. Finally, organize and return the knowledge graph fragment or entity set that best matches the query.
[0048] Furthermore, the multiple recall results obtained through vector retrieval, keyword retrieval and knowledge graph retrieval are fused, and then the relevance between the recall results and the query questions or query documents is calculated, and the multiple recall results are rearranged according to the relevance. Finally, the TopN (i.e., a preset number) recall results are selected as the final retrieval results based on the rearrangement results.
[0049] For example, the query question is: possible causes and solutions for the presence of gangue in the refined product of a three-product heavy medium cyclone. The expanded entities to be retrieved may include: three-product heavy medium cyclone, gangue in refined product, gangue in coal, pressurized three-product heavy medium cyclone, and unpressurized three-product heavy medium cyclone; after the retrieval, the top N knowledge fragments with the highest relevance to the query question are output. The embodiments of the present invention can enhance the ability to understand unstructured data, thereby improving the relevance and accuracy of the retrieval results.
[0050] In one embodiment, the above method also includes: obtaining equipment abnormality information, determining the cause of the equipment abnormality information based on historical data and knowledge graphs, and generating an equipment failure report.
[0051] In specific implementation, a model can be established based on the actual operation data in the coal preparation plant to simulate and analyze the production process, including monitoring and evaluation of factors such as equipment operating status, load status, and production yield. When an abnormality in the equipment is detected, the possible cause of the failure can be automatically inferred based on the historical data and the associated information in the knowledge graph, and a detailed failure report can be generated. This not only speeds up problem location, but also provides a basis for maintenance decisions.
[0052] For ease of understanding, the embodiment of the present invention also provides an information processing flow of a knowledge question answering method combining enhanced retrieval and knowledge graph, see Figure 2 As shown in the figure, after the user enters a query question, the query question is first extracted through the entity extraction algorithm, and then the entity is generalized through colloquial language (i.e., the colloquial entity is recognized), synonym search, upstream and downstream related equipment entity search and graph reasoning are performed to expand the entity. Then, the entity, synonym entity, equipment entity subgraph, process subgraph and the set thinking chain prompt words are combined with the context for retrieval and reasoning, and GAR multi-way retrieval and rerank are used to obtain multi-way recall results, which are returned to the user.
[0053] The embodiment of the present invention also provides a knowledge question answering system combining enhanced retrieval and knowledge graph, see Figure 3 As shown, it includes three parts: database construction, retrieval and generation.
[0054] Among them, database construction includes: for unstructured data (Word, PPT, PDF, web pages, etc.), first perform document parsing, including text extraction, table parsing, formula parsing, image parsing, etc.; then perform data processing, including Chunk segmentation, entity relationship extraction, original data extraction, data enhancement, etc.; finally, build an index or DB, including vector index, inverted index, graph database, relational database, etc. For structured data (data tables, etc.), build a knowledge graph, including vector database, relational database, graph database, etc.
[0055] The retrieval includes: first obtaining the user query, then performing preprocessing, including: query entity alignment, query rewriting, synonym expansion, redundant word removal, etc.; then performing mixed retrieval, including: keyword retrieval, vector retrieval and knowledge graph retrieval; finally, using post-processing strategies to process multi-way recall results, including: rerank, small2big, etc.
[0056] The generation includes: selecting the top K results from the recall results based on relevance, combining the user query and prompts, analyzing through a large model, and outputting the final results.
[0057] The coal preparation knowledge question and answer method provided by the embodiment of the present invention, which integrates enhanced retrieval technology and coal preparation knowledge graph, can efficiently handle complex query requirements in the coal industry, especially questions involving synonyms, near-synonyms and colloquial entities. The expansion of knowledge is achieved by combining the established knowledge spectrum graph, thereby effectively reducing model hallucinations. For example: by searching for coal preparation plant equipment, the model segmentation of equipment is increased, and the search scope is increased by recalling synonyms of entities. It has the following beneficial effects: (1) By constructing and utilizing an industry-specific coal preparation knowledge graph, we solve the problem of recalling synonyms and antonyms of professional terms in the coal preparation industry, thereby improving the accuracy and relevance of retrieval results.
[0058] (2) It can recognize and understand the colloquial expressions of the coal industry, thereby improving the system's generalization capabilities and user experience.
[0059] (3) Utilize a multi-channel recall mechanism to retrieve information from multiple dimensions (such as text content, structured data, relationship paths, etc.) and integrate these results to provide the final answer, effectively reducing incorrect answers or illusions caused by insufficient information from a single source.
[0060] (4) Integrate the production process flow into the knowledge graph to support classified knowledge recall, and use the thinking chain to analyze the process operations to achieve more accurate and comprehensive answer generation.
[0061] (5) Use in-plant production data to model production processes, support automatic reasoning of equipment failures and generation of detailed failure reports, and optimize the maintenance decision-making process.
[0062] (6) Ability to extract key entities from queries or documents and perform entity alignment and expansion to broaden the search horizon and increase the chance of finding relevant information while ensuring consistency across different data sources.
[0063] For the knowledge question answering method combining enhanced retrieval and knowledge graph provided in the above embodiment, the embodiment of the present invention also provides a knowledge question answering system combining enhanced retrieval and knowledge graph, which is applied to the coal industry, see Figure 4 The schematic diagram of the structure of a knowledge query system combining enhanced retrieval and knowledge graph is shown in FIG. 1 , which shows that the system mainly includes the following parts: The entity extraction module 401 is used to obtain the query question or query document input by the user and extract the entity of the query question or query document; wherein the entity includes one or more of the following: equipment name, process flow and fault description; The entity expansion module 402 is used to expand the entity based on the pre-built data mapping model and the knowledge graph of the coal preparation plant to obtain the entity to be retrieved; wherein the entity to be retrieved includes: one or more of the entity, the standardized entity corresponding to the entity, the synonym entity, the near-synonymous entity and the colloquial entity, and the entity's associated equipment entity and associated process flow entity; A retrieval module 403, used to use retrieval enhancement to retrieve the entity to be retrieved and obtain a retrieval result; The question-answering module 404 is used to input the search results, query questions or query documents and constructed prompt words into the big model to generate query results of the query questions or query documents.
[0064] The knowledge question and answer device combining the enhanced retrieval and knowledge graph provided by the present invention can accurately identify and recall synonym entities, near-synonymous entities and colloquial entities by expanding entities through a pre-constructed data mapping model and the knowledge graph of a coal preparation plant, thereby improving the accuracy of information retrieval; at the same time, by combining enhanced retrieval with the knowledge graph, it can handle complex queries, improve semantic understanding and association analysis capabilities, and thereby improve the accuracy and efficiency of retrieval.
[0065] In one embodiment, the above-mentioned data mapping model includes: mapping relationships between standardized entities, synonym entities, near-synonymous entities and colloquial entities of equipment, process flow, and faults; the knowledge graph includes: process flow network, entity triple relationship, and preset constraint rules; wherein the entity triple relationship is used to characterize the upstream and downstream relationship, causal relationship, and timing relationship of the equipment; the preset constraint rules include at least: fault propagation rules and operation rules.
[0066] In one embodiment, the above-mentioned entity expansion module 402 is specifically used to: search for one or more standardized entities, synonym entities, near-synonymous entities and colloquial entities corresponding to the entity based on a pre-built data mapping model; search for the associated process flow entities of the entity based on the process flow network in the knowledge graph, and search for the associated equipment entities of the entity based on the entity triple relationship and preset constraint rules in the knowledge graph.
[0067] In one embodiment, the above-mentioned retrieval module 403 is specifically used to: perform vector retrieval, keyword retrieval and knowledge graph retrieval on the search entity to obtain multiple recall results; re-arrange the multiple recall results based on the relevance between the multiple recall results and the query question or query document, and select a preset number of recall results as retrieval results based on the re-arrangement results.
[0068] In one implementation, the entity extraction module 401 is specifically used to extract entities of a query question or a query document using natural language processing technology or based on a predefined vocabulary.
[0069] In one embodiment, the system further includes: an entity processing module, which is used to: align entities and obtain entities related to the entities based on preset entity relationships.
[0070] In one embodiment, the knowledge graph also includes: equipment anomalies and processing steps; the above system also includes: a fault report generation module, which is used to: obtain equipment anomaly information, and determine the cause of the equipment anomaly information based on historical data and the knowledge graph, and generate an equipment fault report.
[0071] The device provided in the embodiment of the present invention has the same implementation principle and technical effects as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned method embodiment.
[0072] It should be noted that the specific numerical values provided in the implementation of the present invention are only exemplary and are not limited here.
[0073] An embodiment of the present invention further provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above implementation methods.
[0074] Figure 5 A structural diagram of an electronic device provided in an embodiment of the present invention, the electronic device 100 includes: a processor 50, a memory 51, a bus 52 and a communication interface 53, wherein the processor 50, the communication interface 53 and the memory 51 are connected via the bus 52; the processor 50 is used to execute an executable module stored in the memory 51, such as a computer program.
[0075] The memory 51 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 53 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0076] The bus 52 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0077] Among them, the memory 51 is used to store programs, and the processor 50 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 50 or implemented by the processor 50.
[0078] The processor 50 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 50. The above processor 50 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present invention can be directly embodied as a hardware decoding processor to be executed, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 51, and the processor 50 reads the information in the memory 51 and completes the steps of the above method in combination with its hardware.
[0079] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be referred to the previous method embodiments, which will not be repeated here.
[0080] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0081] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A knowledge question answering method combining enhanced retrieval and knowledge graph, characterized in that: Applications in the coal industry include: Obtaining a query question or query document input by a user, and extracting entities of the query question or query document; wherein the entities include one or more of the following: equipment name, process flow, and fault description; The entity is expanded based on a pre-built data mapping model and a knowledge graph of a coal preparation plant to obtain an entity to be retrieved; wherein the entity to be retrieved includes: the entity, one or more of a standardized entity, a synonym entity, a near-synonymous entity and a colloquial entity corresponding to the entity, and an associated equipment entity and an associated process flow entity of the entity; Using search enhancement to search the entity to be searched, and obtaining a search result; The retrieval result, the query question or query document and the constructed prompt words are input into the big model to generate the query result of the query question or query document.
2. The method according to claim 1, characterized in that The data mapping model includes: mapping relationships between standardized entities, synonym entities, near-synonymous entities and colloquial entities of equipment, process flow and fault; The knowledge graph includes: a process flow network, entity triple relationships, and preset constraint rules; wherein the entity triple relationships are used to characterize the upstream and downstream relationships, causal relationships, and timing relationships of equipment; the preset constraint rules include at least: fault propagation rules and operation rules.
3. The method according to claim 2, characterized in that The entities are expanded based on the pre-built data mapping model and the knowledge graph of the coal preparation plant to obtain entities to be retrieved, including: Based on a pre-built data mapping model, searching for one or more of a standardized entity, a synonym entity, a near-synonymous entity, and a colloquial entity corresponding to the entity; Based on the process flow network in the knowledge graph, the associated process flow entities of the entity are searched, and based on the entity triple relationship and preset constraint rules in the knowledge graph, the associated equipment entities of the entity are searched.
4. The method according to claim 1, characterized in that The entity to be searched is searched by using search enhancement to obtain search results, including: Performing vector search, keyword search and knowledge graph search on the entity to be searched respectively to obtain multiple recall results; The multiple recall results are rearranged based on the relevance between the multiple recall results and the query question or the query document, and a preset number of recall results are selected as retrieval results based on the rearrangement result.
5. The method according to claim 1, characterized in that Extracting entities of the query question or query document includes: The entities of the query question or query document are extracted using natural language processing technology or based on a predefined vocabulary.
6. The method according to claim 1, characterized in that After extracting the entities of the query question or query document, the following is further included: The entities are aligned, and entities related to the entities are acquired based on preset entity relationships.
7. The method according to claim 1, characterized in that The knowledge graph also includes: device anomalies and processing steps; the method also includes: Acquire equipment abnormality information, determine the cause and processing steps of the equipment abnormality information based on historical data and the knowledge graph, and generate an equipment failure report.
8. A knowledge question answering system combining enhanced retrieval and knowledge graph, characterized in that: Applications in the coal industry include: An entity extraction module is used to obtain a query question or a query document input by a user, and extract entities of the query question or the query document; wherein the entity includes one or more of the following: equipment name, process flow and fault description; An entity expansion module is used to expand the entity based on a pre-built data mapping model and a knowledge graph of a coal preparation plant to obtain an entity to be retrieved; wherein the entity to be retrieved includes: the entity, one or more of a standardized entity, a synonym entity, a near-synonymous entity and a colloquial entity corresponding to the entity, and an associated equipment entity and an associated process flow entity of the entity; A retrieval module, used to retrieve the entity to be retrieved by using retrieval enhancement to obtain a retrieval result; The question-answering module is used to input the retrieval result, the query question or query document and the constructed prompt words into the big model to generate the query result of the query question or query document.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are executed.
Citation Information
Patent Citations
Document book semantic retrieval system based on knowledge graph
CN115563313A
Dynamic correlation enhancement retrieval generation system and method driven by intelligent knowledge graph
CN118839021A
Method and system for generating enhanced knowledge questions and answers for mixed retrieval of heterogeneous database
CN119311831A
Cited By
Airport maintenance knowledge question answering method and system, electronic equipment and storage medium
CN121094085A
Information retrieval method and electronic equipment
CN121434225A