Coal industry large model retrieval enhancement generation method and system based on knowledge graph

By building a knowledge graph in the coal industry and combining a large language model, the problem of poor search results in existing search technologies is solved, and more efficient and accurate search results are achieved.

CN119988600AActive Publication Date: 2025-05-13CHINA COAL RES INST +1

Patent Information

Application Number
CN202510031500.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-13
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The search enhancement generation technology in the existing coal industry has problems with poor search results, resulting in poor search accuracy and inefficiency.

Method used

Using a knowledge graph-based method, a knowledge graph in the coal industry is constructed. By identifying the user query statements intently, determining the target search range, searching, determining the reliability score of the target text, and inputting it into a large language model to generate a response answer.

Benefits of technology

The search algorithm is effectively optimized, the search accuracy and efficiency are improved, and more accurate and high-quality response answers are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988600A_ABST
    Figure CN119988600A_ABST
Patent Text Reader

Abstract

The invention provides a coal industry large model retrieval enhancement generation method and system based on a knowledge graph. The method comprises the following steps: constructing a coal industry knowledge graph; performing intention recognition on the user query statement, and determining a target retrieval range from the coal industry knowledge graph according to an intention recognition result, the target retrieval range comprising a plurality of candidate texts; performing retrieval based on the user query statement to determine a target text from the plurality of candidate texts; determining a reliability score corresponding to each target text; and inputting the target text into the large language model according to the reliability score to obtain a response answer text corresponding to the user query statement. By implementing the method disclosed by the invention, the retrieval algorithm can be effectively optimized, and the retrieval accuracy and the retrieval efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent coal mine technology, and in particular to a method and system for enhancing the generation of large-scale model retrieval in the coal industry based on a knowledge graph. Background Art

[0002] In the coal industry, in order to improve the intelligence level of coal mines, retrieval-enhanced generation technology has been introduced. Retrieval-enhanced generation means that when a large model generates text or answers questions, relevant information is first retrieved from a large document collection, and then the retrieved information is used to guide the generation of text, thereby improving the quality and accuracy of predictions.

[0003] In the related art, when performing retrieval based on retrieval enhancement generation technology, there are great limitations and the retrieval effect is not good. Summary of the invention

[0004] The present disclosure aims to solve one of the technical problems in the related art at least to some extent.

[0005] To this end, the purpose of the present invention is to propose a large-model retrieval enhancement generation method and system for the coal industry based on knowledge graph, which can effectively optimize the retrieval algorithm and improve the retrieval accuracy and efficiency.

[0006] To achieve the above-mentioned purpose, the first aspect of the present disclosure proposes a method for enhancing the generation of a large model retrieval for the coal industry based on a knowledge graph, including:

[0007] Construct a knowledge graph for the coal industry;

[0008] Performing intent recognition on the user query statement, and determining a target search scope from the coal industry knowledge graph according to the intent recognition result, wherein the target search scope includes multiple candidate texts;

[0009] Performing a search based on the user query statement to determine a target text from the plurality of candidate texts;

[0010] Determining a reliability score corresponding to each of the target texts;

[0011] The target text is input into a large language model according to the reliability score to obtain a response answer text corresponding to the user query statement.

[0012] To achieve the above-mentioned purpose, the second aspect of the present disclosure proposes a coal industry large model retrieval enhancement generation system based on knowledge graph, including:

[0013] Building modules for constructing knowledge graphs for the coal industry;

[0014] A first determination module is used to perform intent recognition on a user query statement, and determine a target search scope from the coal industry knowledge graph according to the intent recognition result, wherein the target search scope includes multiple candidate texts;

[0015] A second determination module, configured to perform a search based on the user query statement to determine a target text from the plurality of candidate texts;

[0016] A third determination module is used to determine the reliability score corresponding to each of the target texts;

[0017] An input module is used to input the target text into a large language model according to the reliability score to obtain a response answer text corresponding to the user query statement.

[0018] The disclosed method and system for coal industry large model retrieval enhancement generation based on knowledge graph builds coal industry knowledge graph; performs intent recognition on user query statements, and determines the target retrieval scope from the coal industry knowledge graph according to the intent recognition result, wherein the target retrieval scope includes multiple candidate texts; performs retrieval based on user query statements to determine the target text from multiple candidate texts; determines the reliability score corresponding to each target text; and inputs the target text into the large language model according to the reliability score to obtain the response answer text corresponding to the user query statement. Thus, the retrieval algorithm can be effectively optimized, and the retrieval accuracy and efficiency can be improved.

[0019] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description or learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and / or additional aspects and advantages of the present disclosure will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0021] Figure 1 It is a flow chart of a method for enhancing the generation of a large model retrieval for the coal industry based on a knowledge graph according to an embodiment of the present disclosure;

[0022] Figure 2 It is a flow chart of a method for enhancing the generation of a large model retrieval for the coal industry based on a knowledge graph according to another embodiment of the present disclosure;

[0023] Figure 3 It is a structural diagram of a large model retrieval and enhanced generation system for the coal industry based on a knowledge graph proposed in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] Embodiments of the present disclosure are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present disclosure, and are not to be construed as limitations of the present disclosure. On the contrary, the embodiments of the present disclosure include all changes, modifications, and equivalents that fall within the spirit and connotation of the appended claims.

[0025] Figure 1 It is a flow chart of a method for enhancing the generation of a large model retrieval for the coal industry based on a knowledge graph proposed in an embodiment of the present disclosure.

[0026] Among them, it should be noted that the executor of the large model retrieval and enhanced generation method for the coal industry based on the knowledge graph in this embodiment is the large model retrieval and enhanced generation system for the coal industry based on the knowledge graph. The system can be implemented by software and / or hardware. The system can be configured in a computer device. The computer device may include but is not limited to a terminal, a server, etc. For example, the terminal may be a mobile phone, a PDA, etc.

[0027] like Figure 1 As shown in the figure, the coal industry large model retrieval enhancement generation method based on knowledge graph includes:

[0028] S101: Build a knowledge graph for the coal industry.

[0029] The coal industry knowledge graph refers to a knowledge graph used to indicate information related to the coal industry. The types and quantities of coal industry knowledge graphs are not limited in the disclosed embodiments.

[0030] It is understandable that there is a large amount of relevant information in the coal industry, and it may be cumbersome to search directly based on this information. Therefore, in the embodiment of the present disclosure, a coal industry knowledge graph can be constructed to achieve structured organization of coal industry related texts to facilitate subsequent retrieval.

[0031] S102: Perform intent recognition on the user query statement, and determine a target search scope from the coal industry knowledge graph based on the intent recognition result, wherein the target search scope includes multiple candidate texts.

[0032] The user query statement refers to the text that the user queries.

[0033] Among them, the intention recognition result can be used to indicate the user's query intention, for example, it can include the query object, query scope, etc., without any restriction.

[0034] Among them, the target search scope refers to the search scope preliminarily determined from the coal industry knowledge graph based on the intent recognition results.

[0035] Among them, candidate text refers to the text related to the coal industry included in the target search scope.

[0036] In the disclosed embodiment, the intent of the user query statement is recognized, and the target search scope is determined from the coal industry knowledge graph according to the intent recognition result, thereby realizing a preliminary search to effectively reduce the computational cost of the subsequent search process.

[0037] S103: performing a search based on the user query statement to determine a target text from a plurality of candidate texts.

[0038] The target text refers to the relevant text retrieved and determined based on the user query statement in the embodiment of the present disclosure.

[0039] That is to say, in the embodiment of the present disclosure, after performing intent recognition on the user query statement and determining the target search scope from the coal industry knowledge graph based on the intent recognition results, a search can be performed based on the user query statement to determine the target text from multiple candidate texts, thereby providing reliable data support for the subsequent response answer text corresponding to the user query statement.

[0040] S104: Determine the reliability score corresponding to each target text.

[0041] Among them, the reliability score can be used to indicate the reliability of the target text in responding to the user's query statement.

[0042] In the disclosed embodiment, when determining the reliability score corresponding to each target text, the target text may be input into a pre-trained machine learning model to obtain the corresponding reliability score, or the reliability score corresponding to each target text may be determined based on a third-party scoring device, without limitation.

[0043] S105: Input the target text into the large language model according to the reliability score to obtain a response answer text corresponding to the user query statement.

[0044] The response answer text refers to the text used to answer the above user query statement.

[0045] In the disclosed embodiment, when the target text is input into the large language model according to the reliability score to obtain the response answer text corresponding to the user query statement, the target text can be reordered according to the reliability score, and then the reordered target text is input into the large language model, which will be sorted by the large language model to generate the corresponding response answer text.

[0046] In this embodiment, by constructing a coal industry knowledge graph; performing intent recognition on user query statements, and determining a target search scope from the coal industry knowledge graph based on the intent recognition results, wherein the target search scope includes multiple candidate texts; performing a search based on the user query statement to determine the target text from multiple candidate texts; determining the reliability score corresponding to each target text; and inputting the target text into a large language model based on the reliability score to obtain a response answer text corresponding to the user query statement. In this way, the retrieval algorithm can be effectively optimized, and the retrieval accuracy and efficiency can be improved.

[0047] Figure 2 It is a flow chart of a method for enhancing the generation of large model retrieval in the coal industry based on knowledge graph proposed in another embodiment of the present disclosure.

[0048] like Figure 2 As shown in the figure, the coal industry large model retrieval enhancement generation method based on knowledge graph includes:

[0049] S201: Obtain data related to the coal industry, wherein the data related to the coal industry includes a plurality of related texts.

[0050] Among them, coal industry-related data refers to data related to the coal industry, such as laws and regulations, journal documents, patents, standards, safety regulations, etc., without any restrictions.

[0051] Among them, relevant text refers to the text contained in the relevant data of the coal industry.

[0052] In the disclosed embodiment, when data related to the coal industry is obtained, reliable data support can be provided for the subsequent construction of a knowledge graph for the coal industry.

[0053] S202: performing data cleaning and preprocessing on relevant texts to obtain reference texts, wherein the candidate text belongs to multiple reference texts.

[0054] The reference text refers to the text obtained after data cleaning and preprocessing of the relevant text.

[0055] It is understandable that the initially acquired relevant texts may contain duplicate data and irrelevant data, which may affect the subsequent retrieval efficiency and cost. Therefore, in the implementation of the present disclosure, the relevant texts can be cleaned and preprocessed to effectively improve the quality and accuracy of the obtained reference texts.

[0056] S203: Determine label information corresponding to each reference text.

[0057] Among them, the label information can be used to indicate the relevant features corresponding to the reference text.

[0058] Optionally, in some embodiments, when determining the label information corresponding to each reference text, it can be to determine the industry business type to which the reference text belongs, and determine the primary label according to the industry business type; determine the industry application scenario to which the reference text belongs, and determine the secondary label according to the industry application scenario; determine the data source, constraint type and viewpoint type corresponding to the reference text, and determine other dimensional labels according to the data source, constraint type and viewpoint type; and combine the primary label, secondary label and other dimensional labels as the label information corresponding to the reference text. In this way, the indication effect of the obtained label information can be effectively improved, and the practicality and reliability of the label information can be guaranteed.

[0059] The industry business type may include, for example, foundation, production, safety, management, etc., without limitation. The primary label may be used to indicate the industry business type to which the reference text belongs.

[0060] Among them, the industry application scenario may include, for example, license information, institutions, geological and hydrological conditions, mining conditions, disaster conditions, IT infrastructure and others, etc. The secondary tag may be used to indicate the industry application scenario to which the reference text belongs.

[0061] Among them, other dimension labels can be used to indicate characteristics such as the data source of the reference text (laws and regulations, standards, patents, papers, etc.), constraint type (mandatory, recommended, etc.), and opinion type (subjective, objective, etc.).

[0062] In the disclosed embodiment, when the label information corresponding to each reference text is determined, reliable reference information can be provided for the subsequent construction of the coal industry knowledge graph.

[0063] S204: Construct a coal industry knowledge graph based on label information and reference text.

[0064] Optionally, in some embodiments, when constructing a coal industry knowledge graph based on label information and reference text, the ontology layer of the knowledge subgraph can be constructed based on the secondary label, wherein the knowledge subgraph is associated with the primary label; entity recognition is performed on the reference text to determine multiple reference entities of the knowledge subgraph, and the association information between the multiple reference entities is determined; the similarity between the multiple reference entities is determined, and the multiple reference entities are fused according to the similarity; the reference entity is attributed based on other dimensional labels; a data connection is established between the reference entity and the corresponding reference text; and the collection of knowledge subgraphs corresponding to multiple primary labels is used as the coal industry knowledge graph. In this way, the indication effect of the obtained coal industry knowledge graph can be effectively improved, and the practicality and reliability of the coal industry knowledge graph can be guaranteed.

[0065] Among them, the ontology layer refers to the level in the knowledge graph used to define and describe knowledge in a specific field, including entities, concepts, attributes, relationships and their constraints.

[0066] The reference entity refers to the entity determined by performing entity recognition on the reference text.

[0067] Among them, the knowledge subgraph refers to the knowledge graph constructed for the first-level tags.

[0068] That is to say, in the disclosed embodiment, a knowledge graph can be constructed for each first-level label respectively, and the collection of knowledge subgraphs corresponding to multiple first-level labels can be used as the coal industry knowledge graph to ensure the clarity of the indication of the obtained coal industry knowledge graph.

[0069] That is, in the disclosed embodiment, data related to the coal industry can be obtained, wherein the data related to the coal industry includes multiple related texts; data cleaning and preprocessing are performed on the related texts to obtain reference texts, wherein the candidate texts belong to multiple reference texts; label information corresponding to each reference text is determined; and a coal industry knowledge graph is constructed based on the label information and the reference texts. Thus, the reliability and practicality of the obtained coal industry knowledge graph can be effectively improved.

[0070] S205: Perform intent recognition on the user query statement, and determine a target search scope from the coal industry knowledge graph based on the intent recognition result, wherein the target search scope includes multiple candidate texts.

[0071] S206: Perform a search based on the user query statement to determine a target text from a plurality of candidate texts.

[0072] S207: Determine the reliability score corresponding to each target text.

[0073] S208: Input the target text into the large language model according to the reliability score to obtain a response answer text corresponding to the user query statement.

[0074] The description of S205-S208 can be specifically referred to the above embodiment, which will not be repeated here.

[0075] In this embodiment, by obtaining data related to the coal industry, wherein the data related to the coal industry includes multiple related texts; performing data cleaning and preprocessing on the related texts to obtain reference texts, wherein the candidate texts belong to multiple reference texts; determining the label information corresponding to each reference text; and constructing a coal industry knowledge graph based on the label information and the reference texts. Thus, the reliability and practicality of the obtained coal industry knowledge graph can be effectively improved.

[0076] Optionally, in some embodiments, when performing intent recognition on a user query statement and determining a target search scope from the coal industry knowledge graph based on the intent recognition result, the following steps may be performed: performing intent recognition on the user query statement to determine the primary and secondary tags related to the user query statement as the intent recognition result; determining a candidate entity from multiple reference entities based on the primary and secondary tags related to the user query statement; and using the reference text connected to the candidate entity as the candidate text, wherein multiple candidate texts jointly constitute the target search scope. In this way, the target search scope can be determined accurately and quickly.

[0077] The candidate entity refers to an entity determined from multiple reference entities based on the primary tags and secondary tags related to the user query statement.

[0078] Optionally, in some embodiments, when searching based on a user query to determine a target text from multiple candidate texts, the following steps may be performed: performing an unstructured search based on the user query to determine a first text from multiple candidate texts; performing a structured search based on the user query to determine a second text from multiple candidate texts; and combining the first text and the second text as the target text. Thus, unstructured search and structured search may be combined to ensure the accuracy of the search for the target text.

[0079] Optionally, in some embodiments, when performing unstructured retrieval based on a user query statement to determine a first text from multiple candidate texts, the following steps may be performed: optimizing and adjusting the user query statement to obtain a target query statement; converting the target query statement into a query statement vector, and converting the candidate texts into candidate text vectors; determining the cosine similarity between the query statement vector and each candidate text vector; and determining the first text from multiple candidate texts based on the cosine similarity. Thus, the accuracy of the determined first text may be guaranteed by combining the cosine similarity between the query statement vector and each candidate text vector.

[0080] Optionally, in some embodiments, when performing structured retrieval based on a user query statement to determine the second text from multiple candidate texts, the following steps may be performed: determining the entity to be queried corresponding to the target query statement; determining the matching result between the entity to be queried and each candidate entity; and determining the second text from the multiple candidate texts based on the matching result. In this way, the second text can be quickly determined.

[0081] Optionally, in some embodiments, when determining the reliability score corresponding to each target text, the following steps may be performed: according to the industry application scenario corresponding to the target text, determine the reference score and weight value corresponding to the target text and each other dimension label; perform weighted summation based on the reference score and weight value to obtain the reliability score. Thus, the reference score and weight value corresponding to the target text and each other dimension label may be determined in combination with the industry application scenario corresponding to the target text, thereby ensuring the adaptability of the obtained reliability score to the personalized application scenario.

[0082] In summary, the present disclosure relates to knowledge graph construction and knowledge retrieval of text data, and the specific implementation steps mainly include the following aspects:

[0083] 1. Constructing a knowledge graph for the coal industry based on the industry label system

[0084] Step 1: Data collection. In order to form a complete vertical field knowledge graph, industry-related knowledge is widely collected. First, the data sources required for the coal industry knowledge graph are sorted out, including but not limited to laws and regulations, journal documents, patents, standards, safety regulations and other data. Next, a combination of automated algorithms and manual review is used to extract data from the above multiple sources to form a unified format. The automated algorithm can capture and transform data in batches, and manual review ensures the accuracy and relevance of the data. Finally, the collected data is classified into structured, semi-structured and unstructured for subsequent processing and analysis. Structured and semi-structured knowledge have explicit structures and fixed formats, and are easy to extract information. Unstructured data is mostly plain text knowledge with no obvious format, which is relatively difficult to extract knowledge.

[0085] Step 2: Data cleaning. In order to improve the quality and accuracy of the data, the data needs to be cleaned and preprocessed, mainly including identifying and removing duplicate and irrelevant data in the data set, filling in missing data, and data standardization. Use relevant technologies of natural language processing to clean the data and improve data quality, laying a solid foundation for the subsequent construction of the knowledge graph.

[0086] Step 3: Data labeling. Based on the set industry labeling system, data labeling tools are used for labeling. Labels include primary labels, secondary labels, and other dimensional labels. The primary labels are divided into four categories according to industry business: basic, production, safety, and management. Based on the primary labels, the secondary labels further classify the data according to the actual application scenarios of the industry. For example, the basic category includes license information, institutions, geological and hydrological conditions, mining conditions, disaster conditions, IT infrastructure, and other seven categories of labels. Other dimensional labels will be classified according to data sources (laws, regulations, standards, patents, papers, etc.), constraint types (mandatory, recommended, etc.), and opinion types (subjective, objective, etc.). Manual labeling is the main method in the early stage. When the labeled data accumulates to a certain amount, deep learning models can be trained to achieve automatic labeling, such as BERT and textcnn models. Subsequent manual inspections will improve the accuracy and efficiency of labeling.

[0087] Step 4: Construct a knowledge graph. Based on the labeled data, the attribute graph model is used as the knowledge representation method, and the knowledge graph is constructed using the Neo4j graph database. For each first-level label, a knowledge subgraph is constructed separately, and the collection of subgraphs of all first-level labels forms a knowledge graph for the coal industry. Taking the production knowledge subgraph as an example, the detailed process of knowledge graph construction is described:

[0088] ① Ontology layer construction defines the concepts and relationships within the field. Due to the particularity of the coal industry, data-driven automated construction cannot clearly summarize and classify. Therefore, manual construction should be adopted to combine industry knowledge to build an ontology layer that conforms to the industry business logic. When constructing the production subgraph, the secondary label of the production class is used as the ontology layer. Since the data has been labeled in step 3, the ontology layer can be constructed quickly, which not only saves effort but also conforms to the industry business logic.

[0089] ②Entity recognition and relationship extraction: Entity objects extracted from knowledge are extracted as nodes of the knowledge graph. In the coal industry, this may involve identifying professional terms such as mines, equipment models, and mining technologies. Relationship extraction is to identify the semantic connection between entities, such as "the mine uses a certain type of equipment." These steps usually require natural language processing technology, including named entity recognition (NER) and relationship extraction algorithms.

[0090] ③ Knowledge fusion, which integrates synonyms and similar entities, and resolves conflicts and ambiguities between different data sources. In the coal industry, the same equipment has different names in different documents. For example, the two entity names "coal mining machine" and "coal cutting machine" refer to the same equipment. Knowledge fusion needs to unify these synonyms to ensure the consistency of the knowledge graph. Since the algorithm model lacks certain background knowledge of the coal industry, it is difficult to achieve the effect of industry professionalism with the traditional knowledge fusion strategy. It is necessary to combine machine learning algorithms with the synonym word list precipitated by the coal industry to ensure the accuracy of the fusion results. In addition, knowledge fusion also includes entity linking and knowledge merging to ensure that knowledge from different sources can be integrated together. To achieve this goal, it is necessary to use other dimensional tags to confirm more authoritative information. For example, when there is a conflict between subjective and objective, objective content should be given priority; when there is a conflict between mandatory and recommended categories, mandatory content should be selected to ensure that the merged knowledge is of higher quality.

[0091] ④ Attribute filling: use labels of other dimensions as attributes of the entity and embed them into the knowledge graph.

[0092] ⑤ Connect each entity node to its original document to facilitate data tracking, unstructured retrieval, and context understanding.

[0093] ⑥Store knowledge. Finally, store the extracted knowledge in the graph database.

[0094] Through the above steps, a structured coal industry knowledge sub-graph with rich semantic information can be constructed. Following the same process, sub-graphs can be constructed for other first-level tags.

[0095] 2. Identify the user’s query intent and narrow the search scope

[0096] Perform intent recognition on the query statement entered by the user and confirm the primary and secondary tags of the query statement. Specifically, the identified primary tags can narrow the search scope to the sub-graph, and the secondary tags can narrow the search scope to the related entities in the sub-graph, so as to improve the subsequent search efficiency. The intent recognition of query statements can be achieved in three ways: 1. Intent recognition of query statements based on the automatic labeling model trained in the early stage; 2. Intent recognition of query statements based on a large model, that is, by building a prompt template, guiding the large model to answer the label of the query statement; 3. Intent recognition based on vector matching, vectorize the query and label and their description information respectively, and perform similarity matching.

[0097] 3. Vector-based unstructured retrieval and knowledge graph-based structured retrieval

[0098] After confirming the search scope, this step combines the advantages of unstructured and structured retrieval to form a hybrid retrieval mechanism to provide comprehensive and accurate search results. First, for the user query statement, you can selectively rewrite / expand it to optimize the expression of the query statement and facilitate more accurate retrieval. Then, a vector-based unstructured retrieval method is introduced. In this method, the system converts the query statement into a high-dimensional vector and calculates its cosine similarity with the text vector in the search scope locked in the second step. Through this step, the system can quickly identify the n texts that are semantically closest to the query statement, reducing large-scale similarity calculations, and use this retrieval result as an unstructured search result. Next, structured retrieval technology is used to identify entities in the query statement through an entity extraction model. Using the structured information in the knowledge graph, the system can match relevant entities and attributes within the locked search scope and retrieve neighborhood information that is highly relevant to the user query. This structured retrieval process not only provides accurate retrieval results, but also reveals the complex relationships between entities through the links of the knowledge graph. The query result will be used as a structured query result.

[0099] 4. Retrieve and rearrange to generate context

[0100] Retrieval re-ranking and context generation are important steps to improve the quality of retrieval results. This step re-ranks the retrieval results based on the authority scores of the attributes in the knowledge graph to ensure that more authoritative information is ranked first. The authoritative score can be achieved by calculating the weighted average. Different weights are assigned to each label dimension, and scores are given in each dimension according to the degree of authority. For example, in the constraint type dimension, mandatory labels > recommended labels; in the opinion type dimension, objective labels > subjective labels. The authority of the retrieval content is calculated in this way. According to the authority score of the retrieval content, the retrieval results are re-ranked, and the retrieval results with high scores are ranked first. A comprehensive context is generated to provide input for large language models.

[0101] 5. Large language models generate response answers

[0102] Generate response answers using a large language model. Insert the text content retrieved and rearranged in the previous step into the pre-built prompt template as the input of the large model, and guide the large language model to generate accurate, logical, and professional answers.

[0103] The present invention narrows the search scope and improves query efficiency by constructing a knowledge graph based on an industry tag system. At the same time, it improves the accuracy and authority of answers through attribute-based search rearrangement. Finally, it utilizes the rich relationships in the knowledge graph to generate logically coherent and highly relevant text, thereby improving the quality of semantic understanding and text generation.

[0104] Figure 3 It is a structural diagram of a large model retrieval and enhanced generation system for the coal industry based on a knowledge graph proposed in an embodiment of the present disclosure.

[0105] like Figure 3 As shown, the coal industry large model retrieval enhancement generation system 30 based on knowledge graph includes:

[0106] Construction module 301, used to construct a knowledge graph of the coal industry;

[0107] The first determination module 302 is used to perform intent recognition on the user query statement, and determine a target search scope from the coal industry knowledge graph according to the intent recognition result, wherein the target search scope includes multiple candidate texts;

[0108] A second determination module 303 is used to perform a search based on a user query statement to determine a target text from a plurality of candidate texts;

[0109] The third determination module 304 is used to determine the reliability score corresponding to each target text;

[0110] The input module 305 is used to input the target text into the large language model according to the reliability score to obtain the response answer text corresponding to the user query statement.

[0111] It should be noted that the aforementioned explanation of the large model retrieval and enhanced generation method for the coal industry based on the knowledge graph is also applicable to the large model retrieval and enhanced generation system for the coal industry based on the knowledge graph in this embodiment, and will not be repeated here.

[0112] In this embodiment, by constructing a coal industry knowledge graph; performing intent recognition on user query statements, and determining a target search scope from the coal industry knowledge graph based on the intent recognition results, wherein the target search scope includes multiple candidate texts; performing a search based on the user query statement to determine the target text from multiple candidate texts; determining the reliability score corresponding to each target text; and inputting the target text into a large language model based on the reliability score to obtain a response answer text corresponding to the user query statement. In this way, the retrieval algorithm can be effectively optimized, and the retrieval accuracy and efficiency can be improved.

[0113] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0114] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

[0115] It should be noted that, in the description of the present disclosure, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present disclosure, unless otherwise specified, the meaning of "plurality" is two or more.

[0116] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.

[0117] It should be understood that the various parts of the present disclosure can be implemented in hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0118] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0119] In addition, each functional unit in each embodiment of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0120] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0121] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0122] Although the embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.

Claims

1. A large model retrieval enhancement generation method for the coal industry based on knowledge graph, characterized in that: include: Construct a knowledge graph for the coal industry; Performing intent recognition on the user query statement, and determining a target search scope from the coal industry knowledge graph according to the intent recognition result, wherein the target search scope includes multiple candidate texts; Performing a search based on the user query statement to determine a target text from the plurality of candidate texts; Determining a reliability score corresponding to each of the target texts; The target text is input into a large language model according to the reliability score to obtain a response answer text corresponding to the user query statement.

2. The method according to claim 1, characterized in that The construction of the coal industry knowledge graph includes: Acquire coal industry related data, wherein the coal industry related data includes a plurality of related texts; Performing data cleaning and preprocessing on the relevant texts to obtain reference texts, wherein the candidate text belongs to a plurality of the reference texts; Determine label information corresponding to each of the reference texts; Based on the label information and the reference text, the coal industry knowledge graph is constructed.

3. The method according to claim 2, characterized in that The determining of the label information corresponding to each of the reference texts includes: Determine the industry business type to which the reference text belongs, and determine a primary label according to the industry business type; Determine the industry application scenario to which the reference text belongs, and determine the secondary label according to the industry application scenario; Determine the data source, constraint type and viewpoint type corresponding to the reference text, and determine other dimension labels according to the data source, constraint type and viewpoint type; The primary label, the secondary label and the other dimensional labels are combined as the label information corresponding to the reference text.

4. The method according to claim 3, characterized in that The step of constructing the coal industry knowledge graph based on the tag information and the reference text includes: Based on the secondary tags, constructing an ontology layer of a knowledge subgraph, wherein the knowledge subgraph is associated with the primary tags; Performing entity recognition on the reference text to determine multiple reference entities of the knowledge subgraph, and determining association information between the multiple reference entities; Determining similarities between the multiple reference entities, and fusing the multiple reference entities according to the similarities; Filling attributes of the reference entity based on the other dimension tags; Establishing a data connection between the reference entity and the corresponding reference text; The collection of the knowledge subgraphs corresponding to the multiple first-level labels is used as the coal industry knowledge graph.

5. The method according to claim 4, characterized in that The performing of intent recognition on the user query statement and determining the target search scope from the coal industry knowledge graph according to the intent recognition result includes: Performing intent recognition on the user query statement to determine a primary tag and a secondary tag related to the user query statement as the intent recognition result; Determine a candidate entity from the multiple reference entities according to the primary tags and the secondary tags related to the user query statement; The reference text connected to the candidate entity is used as the candidate text, wherein a plurality of the candidate texts jointly constitute the target search scope.

6. The method according to claim 5, characterized in that The retrieving based on the user query statement to determine the target text from the multiple candidate texts includes: Performing unstructured retrieval based on the user query statement to determine a first text from the multiple candidate texts; Performing structured retrieval based on the user query statement to determine a second text from the multiple candidate texts; The first text and the second text are combined as the target text.

7. The method according to claim 6, characterized in that The performing unstructured retrieval based on the user query statement to determine the first text from the multiple candidate texts includes: Optimizing and adjusting the user query statement to obtain a target query statement; Convert the target query sentence into a query sentence vector, and convert the candidate text into a candidate text vector; Determining the cosine similarity between the query sentence vector and each of the candidate text vectors; The first text is determined from the plurality of candidate texts according to the cosine similarity.

8. The method according to claim 6, characterized in that The performing structured retrieval based on the user query statement to determine the second text from the multiple candidate texts includes: Determine the entity to be queried corresponding to the target query statement; Determine a matching result between the entity to be queried and each of the candidate entities; The second text is determined from the multiple candidate texts according to the matching result.

9. The method according to claim 3, characterized in that Determining the reliability score corresponding to each target text includes: According to the industry application scenario corresponding to the target text, determine the reference score and weight value corresponding to the target text and each of the other dimension labels; A weighted sum is performed based on the reference score and the weight value to obtain the reliability score.

10. A large model retrieval enhancement generation system for the coal industry based on knowledge graph, characterized in that: include: Building modules for constructing knowledge graphs for the coal industry; A first determination module is used to perform intent recognition on a user query statement, and determine a target search scope from the coal industry knowledge graph according to the intent recognition result, wherein the target search scope includes multiple candidate texts; A second determination module, configured to perform a search based on the user query statement to determine a target text from the plurality of candidate texts; A third determination module is used to determine the reliability score corresponding to each of the target texts; An input module is used to input the target text into a large language model according to the reliability score to obtain a response answer text corresponding to the user query statement.

Citation Information

Patent Citations

  • Natural language query method and device

    CN118312604A

  • Question and answer method and system based on tin smelting knowledge graph

    CN118779424A

  • Generative question answering method and system based on knowledge graph and document retrieval integration

    CN119046448A

  • Method and system for predicting biological entities

    GB202402771D0

  • Multi-source hybrid question answering method and system thereof

    KR101662450B1

Cited By

  • Stamp retrieval method and system based on knowledge graph, electronic equipment and medium

    CN120386887A

  • Digital human interaction method, device, equipment and program product

    CN121999777A

  • Data desensitization method based on large language model

    CN122197055A