Knowledge base construction method and device and knowledge question and answer processing method and device
By hierarchical division of the original document and vector data construction, the search path of the knowledge question-and-answer system is optimized, the problem of low retrieval efficiency is solved, the relevant information is quickly positioned and the recall rate is improved, and the accuracy and credibility of the answer is enhanced.
Patent Information
- Application Number
- CN202510842962.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The search efficiency based on search enhancement generation technology in the existing knowledge Q&A system needs to be improved, especially in the power industry, target information cannot be quickly positioned, and the recall rate of related documents is low.
By hierarchically dividing the original document, texts at different search levels are constructed and corresponding vector data are constructed. The search order is assigned to different levels of texts according to preset standards, the index path is optimized, and hierarchical filtering is realized to avoid full-store scanning.
It improves the speed of knowledge retrieval, improves the efficiency of key documents recall, enhances the system's ability to understand user intentions, and improves the accuracy and credibility of answers.
Smart Images

Figure CN120373439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of retrieval optimization, and particularly to a method and device for knowledge base construction and knowledge question and answer processing. Background Art
[0002] With the rapid development of intelligent equipment and large language models, knowledge question and answer systems based on information retrieval technology have been widely used, which can help users improve the efficiency of problem-solving to a certain extent. Ensuring the answer accuracy of the knowledge question and answer system is of great significance for users to solve problems quickly and accurately.
[0003] In the related art, in order to improve the answer accuracy of the knowledge question and answer system, retrieval enhanced generation technology is adopted to enhance the answering ability of the language model. However, the retrieval efficiency based on the retrieval enhanced generation technology in the related art needs to be improved. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems in the related art to some extent. For this purpose, the present invention proposes a method and device for knowledge base construction and knowledge question and answer processing, which optimize the index path by performing hierarchical division processing on the original document, thereby improving the retrieval efficiency.
[0005] To achieve the above object, a first aspect embodiment of the present invention provides a method for constructing a knowledge base, the method comprising: obtaining an original document; dividing the original document according to a preset hierarchical division method to obtain a first retrieval hierarchical text and a second retrieval hierarchical text; wherein the retrieval order of the first retrieval hierarchical text is before that of the second retrieval hierarchical text; constructing first vector data based on the first retrieval hierarchical text, and constructing second vector data based on the second retrieval hierarchical text; and constructing a knowledge base according to the first vector data and the second vector data.
[0006] According to an embodiment of the present invention, the dividing the original document according to a preset hierarchical division method to obtain a first retrieval hierarchical text and a second retrieval hierarchical text comprises: determining a first classification domain and a second classification domain corresponding to the original document according to the preset hierarchical division method; parsing the original document to determine the first retrieval hierarchical text under the first classification domain and the second retrieval hierarchical text under the second classification domain; wherein the granularity of the second retrieval hierarchical text is smaller than that of the first retrieval hierarchical text.
[0007] According to an embodiment of the present invention, the first retrieval level text corresponds to a first position index, and / or, the second retrieval level text corresponds to a second position index; the first position index is used to indicate the position of the original text content corresponding to the first retrieval level text in the original document; the second position index is used to indicate the position of the original text content corresponding to the second retrieval level text in the original document.
[0008] According to an embodiment of the present invention, there is a first mapping relationship between the first vector data and the second vector data; before constructing the knowledge base according to the first vector data and the second vector data, the knowledge base construction method further includes: determining a third classification domain corresponding to the original document according to the preset hierarchical division method; parsing the original document to obtain a third retrieval level text under the third classification domain; wherein, the retrieval order of the second retrieval level text is before that of the third retrieval level text; the granularity of the third retrieval level text is smaller than that of the second retrieval level text; constructing third vector data based on the third retrieval level text; wherein, there is a second mapping relationship between the second vector data and the third vector data; constructing the knowledge base according to the first vector data and the second vector data includes: constructing the knowledge base according to the first vector data, the second vector data, the first mapping relationship, the third vector data and the second mapping relationship.
[0009] According to an embodiment of the present invention, the preset hierarchical division method is determined according to the field to which the original document belongs; the first retrieval level text is determined according to the document-level content of the original document; the second retrieval level text is determined according to the chapter-level content of the original document; the third retrieval level text is determined according to the paragraph-level content of the original document.
[0010] According to an embodiment of the present invention, the original document corresponds to value level information, and the value level information is used to determine the retrieval priority corresponding to the original document; the knowledge base contains the value level information corresponding to the original document.
[0011] According to an embodiment of the present invention, the original document also corresponds to a confidentiality level; constructing the knowledge base according to the first vector data and the second vector data includes: constructing the knowledge base according to the first vector data, the second vector data and the confidentiality level.
[0012] To achieve the above object, a second aspect embodiment of the present invention provides a method for processing knowledge questions and answers. The method includes: obtaining question text data; retrieving the question text data in a knowledge base constructed according to the knowledge base construction method described in any one of the foregoing embodiments in accordance with the retrieval order of the retrieval level text to obtain a first retrieval result; generating a first answer text corresponding to the question text data based on the first retrieval result.
[0013] According to an embodiment of the present invention, the method for processing knowledge questions and answers further includes: determining a target original document corresponding to the first retrieval result according to the first retrieval result; generating a text attachment corresponding to the first retrieval result according to the first retrieval result and the target original document.
[0014] According to an embodiment of the present invention, the first retrieval result corresponds to a first retrieval range; the method for processing knowledge questions and answers further includes: evaluating the first answer text to obtain an evaluation result; if the evaluation result does not meet a preset condition, determining a second retrieval range according to the first retrieval range; retrieving the question text data in the knowledge base based on the second retrieval range to obtain a second retrieval result, so as to generate a second answer text according to the second retrieval result.
[0015] According to an embodiment of the present invention, the question text data is obtained by performing vectorization processing on original question data, and the original question data includes at least one of text data, image data, and voice data.
[0016] To achieve the above object, a third aspect embodiment of the present invention provides a knowledge base construction device. The device includes: an original document acquisition module for acquiring an original document; a retrieval level text acquisition module for dividing the original document according to a preset level division method to obtain a first retrieval level text and a second retrieval level text; wherein, the retrieval order of the first retrieval level text is before that of the second retrieval level text; a vector data construction module for constructing first vector data based on the first retrieval level text and constructing second vector data based on the second retrieval level text; a knowledge base construction module for constructing a knowledge base according to the first vector data and the second vector data.
[0017] To achieve the above object, a fourth aspect embodiment of the present invention provides a knowledge Q&A processing device, the device comprising: a question text acquisition module for acquiring question text data; a retrieval result acquisition module for retrieving the question text data in a knowledge base constructed according to the knowledge base construction method described in any of the foregoing embodiments in accordance with the retrieval order of retrieval level texts to obtain a first retrieval result; and an answer text generation module for generating a first answer text corresponding to the question text data based on the first retrieval result.
[0018] According to multiple embodiments provided by the present invention, by performing hierarchical division processing on the original document according to a preset standard or rule, for example, performing hierarchical division processing according to document importance or authority or timeliness, or performing hierarchical division processing according to the structure of the document itself, and assigning different retrieval orders (or retrieval priorities) to texts at different retrieval levels, so that retrieval is performed in accordance with the retrieval order of different retrieval levels during the knowledge retrieval process, realizing hierarchical filtering, avoiding full database scanning, optimizing the index path, facilitating faster positioning to the most relevant information, thereby improving the retrieval speed and enhancing the recall efficiency of key documents.
[0019] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart of a knowledge base construction method provided according to an embodiment of the present specification.
[0021] Figure 2 It is a flowchart of a process for dividing an original document provided according to an embodiment of the present specification.
[0022] Figure 3a It is a flowchart of another knowledge base construction method provided according to an embodiment of the present specification.
[0023] Figure 3b It is a schematic diagram of a hierarchical structure of domain layering provided according to an embodiment of the present specification.
[0024] Figure 4 It is a flowchart of a knowledge Q&A processing method provided according to an embodiment of the present specification.
[0025] Figure 5 It is a flowchart of another knowledge Q&A processing method provided according to an embodiment of the present specification.
[0026] Figure 6 It is a flowchart of yet another knowledge Q&A processing method provided according to an embodiment of the present specification.
[0027] Figure 7a It is a schematic structural diagram of an improved power scenario RAG system provided according to the embodiments of this specification.
[0028] Figure 7b It is a schematic diagram of a scoring - retrieval expansion process provided according to the embodiments of this specification.
[0029] Figure 7c It is a schematic diagram of a federated data processing method provided according to the embodiments of this specification.
[0030] Figure 7d It is a schematic flow diagram of an improved power scenario RAG method provided according to the embodiments of this specification. Specific Embodiments
[0031] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.
[0032] With the rapid development of artificial intelligence technology, especially the continuous maturity and popularization of intelligent equipment and large language models, artificial intelligence has gradually penetrated into all walks of life and people's daily lives. In this context, artificial intelligence not only improves the efficiency of information processing but also significantly enhances the ability to solve complex problems. For example, in multiple industries such as power, transportation, medical care, and finance, knowledge - based question - answering systems based on information retrieval technology have been used for auxiliary decision - making, intelligent question - answering, fault diagnosis, etc., which can greatly improve work efficiency.
[0033] In the application of knowledge - based question - answering systems, ensuring the accuracy of the answers of the knowledge - based question - answering systems is of great significance for users to solve problems quickly and accurately. In related technologies, the retrieval - augmented generation technology is applied in the knowledge - based question - answering systems to enhance the answering ability of the language model.
[0034] Retrieval-Augmented Generation (RAG) is an artificial intelligence technology framework that combines a retrieval system with a generation model. Its core lies in enhancing the answering ability of a language model through the dynamic retrieval of an external knowledge base, and it is applicable to scenarios that require processing large-scale unstructured data and generating accurate responses. The RAG system has many advantages. For example, it relies on a vector database or a document retrieval system to store external knowledge, supports the model to call real-time or historical data, and solves the problem of insufficient update of dynamic information by traditional large models; it separates the "information retrieval" and "content generation" processes. First, the retrieval module locates relevant document fragments, and then the generation model integrates the information to form an answer. By retrieving real documents to support the answer, it reduces the hallucination of large models and improves the accuracy and traceability of the results; it can be compatible with multi-domain knowledge bases and supports specific adjustments to meet the needs of vertical industries. It has been initially applied and implemented in fields such as healthcare, law, and finance.
[0035] Taking the power industry as an example, electricity is ubiquitous in people's modern lives, and it is crucial to ensure the safe and stable operation of the power system. The daily operation and maintenance tasks faced by power workers are complex and diverse, and they need to quickly obtain accurate technical support and troubleshooting guidance. Applying an artificial intelligence question-and-answer system to the power industry can help front-line production personnel quickly solve the problems they encounter. By applying RAG automated text-enhanced dialogue to the large model in the power industry, the accuracy of the answers can be greatly improved, providing more convenience for power industry workers and thus improving the overall efficiency.
[0036] However, in related technologies, RAG usually adopts a centralized knowledge base architecture and a flat vector retrieval method based on Chunk Embedding, lacking refined hierarchical management and retrieval mechanisms, resulting in poor timeliness and limited coverage of knowledge retrieval, inability to quickly locate target information during retrieval, and low recall rate of relevant documents. Therefore, the knowledge retrieval efficiency based on retrieval-augmented generation technology needs to be improved.
[0037] To improve the knowledge retrieval efficiency, it is necessary to propose a knowledge base construction, knowledge question-and-answer processing method and device. In the knowledge base construction method provided in this specification, the original document is divided into the first retrieval level text and the second retrieval level text at different retrieval levels according to a preset hierarchical division method, and then corresponding vector data are constructed based on the text at different retrieval levels, that is, the corresponding first vector data are constructed based on the first retrieval level text, and the corresponding second vector data are constructed based on the second retrieval level text, so as to construct a knowledge base according to the first vector data and the second vector data. Among them, the retrieval order of the first retrieval level text is before the second retrieval level text, that is, the retrieval order of the first vector data is before the second vector data.
[0038] Thus, by hierarchically partitioning and processing the original document according to preset criteria or rules, for example, partitioning and processing according to document importance, authority, or timeliness, or partitioning and processing according to the structure of the document itself, and assigning different retrieval orders (or retrieval priorities) to texts at different retrieval levels, during the knowledge retrieval process, retrieval is performed according to the retrieval orders of different retrieval levels, achieving hierarchical filtering, avoiding full database scanning, optimizing the index path, facilitating faster location of the most relevant information, thereby improving the retrieval speed and enhancing the recall efficiency of key documents.
[0039] The embodiments of this specification provide a knowledge base construction method. Refer to Figure 1 As shown, the knowledge base construction method may include the following steps.
[0040] S110. Obtain the original document.
[0041] S120. Partition the original document according to a preset hierarchical partitioning method to obtain first retrieval level text and second retrieval level text; wherein, the retrieval order of the first retrieval level text is before that of the second retrieval level text.
[0042] S130. Construct first vector data based on the first retrieval level text, and construct second vector data based on the second retrieval level text.
[0043] S140. Construct a knowledge base according to the first vector data and the second vector data.
[0044] Among them, the first retrieval level text is the text at the first retrieval level, and the second retrieval level text is the text at the second retrieval level. The retrieval order of the first retrieval level text being before that of the second retrieval level text means that during the information retrieval process, retrieval is first performed on the first retrieval level text, and then on the second retrieval level text.
[0045] Partitioning the original document according to a preset hierarchical partitioning method may be to partition different original documents into different retrieval levels, in which case the retrieval level text can be obtained according to the corresponding entire original document, or to partition the content of the same original document into different retrieval levels, in which case the retrieval level text can be obtained according to the content of the corresponding part of the original document.
[0046] Vector data is a data form used to represent text features and can describe the semantic information in the text.
[0047] Specifically, first, collect and organize the original documents to be incorporated into the knowledge base. These original documents can include various forms of text materials, such as operation manuals, technical documents, guides, papers, logs, etc. Divide the original documents according to a preset hierarchical division method to obtain texts at different retrieval levels, including the first retrieval level text at the first retrieval level and the second retrieval level text at the second retrieval level. The preset division method can be determined based on the importance, authority, usage frequency, update frequency, etc. of the documents, or can be determined according to the document structure, etc. Among them, the retrieval order of the first retrieval level text is before the second retrieval level text, which means that the first retrieval level text has a higher retrieval priority and will be retrieved prior to the second retrieval level text during the information retrieval process.
[0048] Perform vectorization processing on the first retrieval level text to construct the first vector data, and perform vectorization processing on the second retrieval level text to construct the second vector data. Build the final knowledge base based on the first vector data and the second vector data. This knowledge base can serve as the basis for functions such as information query and question answering in the knowledge-based question answering system.
[0049] In some embodiments, the preset division method can be determined based on the importance, authority, timeliness, usage frequency, update cycle, etc. of the documents.
[0050] Exemplarily, determine the preset division method based on authority, divide the corresponding documents with an authority level greater than or equal to the threshold into the first retrieval level, and divide the corresponding documents with an authority level less than the threshold into the second retrieval level; determine the preset division method based on importance and update frequency, divide the corresponding documents with an importance level greater than or equal to the importance level threshold and an update cycle less than or equal to the update cycle threshold into the first retrieval level, and divide the corresponding documents with an importance level less than the importance level threshold and an update cycle greater than the update cycle threshold into the second retrieval level.
[0051] Furthermore, in some cases, any one of methods such as a dynamic weight tree can be used to determine the importance corresponding to different documents, or any one of methods such as a time decay model can be used to determine the timeliness corresponding to different documents, which is used as the basis for hierarchical division, so as to achieve hierarchical division of the original documents based on the dynamic weight tree or the time decay model.
[0052] In some other embodiments, the preset division method can be determined according to the technical fields included in the documents.
[0053] Exemplarily, for the technical standard document of transmission line detection, the technical fields it includes are equipment-related fields, operation-related fields, and safety-related fields. Each technical field can further include lower-level technical sub-domains. Taking the equipment-related field as an example, it can further include equipment failure sub-domains and equipment maintenance sub-domains. Then, according to the domain division method, the original document can be divided, and the corresponding first retrieval level texts can be determined according to the equipment-related content, operation-related content, and safety-related content respectively. Taking the equipment-related field as an example, the corresponding second retrieval level texts can be further determined according to the equipment failure-related content and equipment maintenance-related content respectively.
[0054] It should be noted that by dividing the original document according to the preset hierarchical division method, third retrieval level texts, fourth retrieval level texts, etc. can also be obtained. The specific number of levels can be determined according to the actual application scenario or requirements, etc., as long as there is a sequential relationship in the retrieval order between different retrieval level texts. The preset hierarchical division method can be specifically determined according to the actual application requirements, etc., and is not specifically limited in this specification.
[0055] In the above embodiments, by performing hierarchical division processing on the original document according to the preset hierarchical division method, texts at different retrieval levels are obtained, and the retrieval orders of the texts at different retrieval levels are different, so that in the knowledge retrieval process, retrieval can be performed according to the retrieval order of the retrieval levels, realizing hierarchical indexing of the document, facilitating faster positioning to the most relevant information, and thus improving the retrieval speed. At the same time, the method of using hierarchical division helps to achieve hierarchical management of data, and through the vector representation method, it helps the system to more accurately understand the user's query intention and retrieve the most relevant answers from the knowledge base, improving the quality and credibility of knowledge retrieval.
[0056] In some embodiments, as shown in Figure 2 According to the preset hierarchical division method, dividing the original document to obtain the first retrieval level text and the second retrieval level text may include the following steps.
[0057] S210. Determine the first classification domain and the second classification domain corresponding to the original document according to the preset hierarchical division method.
[0058] S220. Parse the original document to determine the first retrieval level text under the first classification domain and the second retrieval level text under the second classification domain; wherein, the granularity of the second retrieval level text is smaller than that of the first retrieval level text.
[0059] Among them, the classification domain refers to the classification standard or the category of knowledge or data divided according to specific rules, which can be used for semantic classification of the document content.
[0060] The first classification domain corresponds to the first retrieval level, and the second classification domain corresponds to the second retrieval level.
[0061] Granularity refers to the fineness of the text content or the richness of the details contained in the text. The granularity of the second retrieval level text is smaller than that of the first retrieval level text, which means that the second retrieval level text has more details and finer content compared to the first retrieval level text. That is, the first retrieval level text contains coarse-grained knowledge units, and the second retrieval level text contains fine-grained specific knowledge content.
[0062] Specifically, for any original document, first, the first classification domain and the second classification domain corresponding to the original document can be determined according to the preset hierarchical division method. Second, by parsing the original document, coarse-grained information related to the first classification domain can be extracted from the original document as the first retrieval level text corresponding to the original document, and fine-grained (fine-grained compared to the granularity of the first retrieval level text) information related to the second classification domain can be extracted from the original document as the second retrieval level text.
[0063] Furthermore, there is a first mapping relationship between the first vector data constructed based on the first retrieval level text and the second vector data constructed based on the second retrieval level text (that is, the association relationship between the contents in the original document), that is, the corresponding data in the second vector data can be found from a certain first vector data. Constructing a knowledge base based on the first vector data and the second vector data further includes: constructing a knowledge base according to the first vector data, the second vector data, and the first mapping relationship.
[0064] In some embodiments, the preset division method can be determined according to the technical field to which the document belongs.
[0065] Exemplarily, for a technical standard document for transmission line detection, the technical field to which it belongs is the transmission field, then the transmission field can be used as the first classification domain. The transmission field also includes sub-fields, such as the equipment sub-field, the operation sub-field, the safety sub-field, etc., then the equipment sub-field, the operation sub-field, the safety sub-field, etc. can be used as the second classification domains. By parsing this document, coarse-grained information can be extracted from at least one of the document title, introduction, abstract, etc. to obtain the corresponding first retrieval level text, and fine-grained information can be extracted from the equipment-related content, operation-related content, safety-related content, etc. in the document to obtain the corresponding second retrieval level text, or the equipment-related content, operation-related content, safety-related content, etc. in the document can be used as the second retrieval level text.
[0066] In some other embodiments, the preset hierarchical division method may be to first perform a first retrieval hierarchical division according to the title in the original document or the field to which the original document belongs, and then perform a second hierarchical division according to the structure of the original document (chapters, paragraphs, charts, formulas, etc.). Then, the first classification domain can be determined according to the document title of the original document or the field to which the original document belongs, and the second classification domain can be determined according to the chapters or paragraphs of the original document, etc.
[0067] Exemplarily, the original document is a document about the emergency plan for the main transformer tripping in a substation. By performing a first retrieval hierarchical division according to the document title, the first classification domain can be determined as "Emergency Disposal". Or, by performing a first retrieval hierarchical division according to the field to which the document belongs, the first classification domain can be determined as "Substation". By parsing this document, coarse-grained information can be extracted from at least one of the document title, introduction, abstract, etc., to obtain the corresponding first retrieval hierarchical text.
[0068] Assume that the document contains chapters on accident handling principles and the main transformer tripping handling process. Then, the second classification domain can be determined to include "Handling Principles" and "Handling Process". By parsing this document, fine-grained information can be extracted from the content of the corresponding chapters to obtain the corresponding second retrieval hierarchical text, or the content of the chapters can be used as the corresponding second retrieval hierarchical text.
[0069] Exemplarily, the original document is a substation equipment maintenance manual, which includes content such as "Chapter 1, Transformer Maintenance Regulations", "Chapter 2, Circuit Breaker Operation Specifications", and "Chapter 3, Fault Emergency Handling Process". By performing a first retrieval hierarchical division according to the titles in the document, the first classification domain can be determined to include "Transformer", "Circuit Breaker", and "Emergency Handling". Assume that the content of the chapter on transformer maintenance regulations in this document includes "Insulating Oil Sampling Operation Steps" and "Bushing Infrared Temperature Measurement Standards". Then, the second classification domain can be determined to include "Insulating Oil Sampling" and "Bushing Temperature Measurement".
[0070] In still some other embodiments, the preset hierarchical division method may be to first perform a first retrieval hierarchical division according to the document type of the original document, and then perform a second hierarchical division according to the structure of the original document. Then, the first classification domain can be determined according to the document type of the original document, and the second classification domain can be determined according to the chapters or paragraphs of the original document, etc.
[0071] Exemplarily, if the original document is a technical standard document, the first classification domain can be determined as the technical standard category, and the second classification domain can include general principles, technical requirements, test methods, etc.; if the original document is a journal paper, the first classification domain can be determined as the paper category, and the second classification domain can be obtained by chapter division; if the original document is a log document, the first classification domain can be determined as the log category, and the second classification domain can be obtained by dividing according to the specific event type or operation type recorded in the log.
[0072] It should be noted that the specific method for determining the first classification domain and the second classification domain corresponding to the original document according to the preset hierarchical division method can be determined according to the actual application scenario or requirements, etc., and is not specifically limited in this specification. The parsing method can be implemented by means such as title recognition and semantic segmentation, and is not specifically limited in this specification.
[0073] In the above embodiments, by converting the original document into a well-structured knowledge system, the system can first quickly locate relevant coarse-grained information at a high level, and then delve into the low level to find the corresponding fine-grained information, reducing the performance loss caused by the flat full-scale search in the related art, enhancing the system's ability to understand the user's intention and match the correct answer, thereby helping to improve the efficiency of knowledge retrieval and solve the balance problem between retrieval accuracy and timeliness. At the same time, this hierarchical design makes the content of the knowledge base easy to expand, replace, and version control, and can meet different knowledge update requirements.
[0074] In some embodiments, the first retrieval level text corresponds to a first position index, and / or the second retrieval level text corresponds to a second position index. The first position index is used to indicate the position of the original text content corresponding to the first retrieval level text in the original document; the second position index is used to indicate the position of the original text content corresponding to the second retrieval level text in the original document.
[0075] Among them, the position index records the physical or logical position of the retrieval level text in the corresponding original document, and is used to achieve fast traceability from the retrieval result to the original text.
[0076] Specifically, for each first retrieval level text, a corresponding first position index can be established to record the physical or logical position of this level text in the original document where it is located, such as page number, paragraph number, coordinates, etc. And / or, for each second retrieval level text, a corresponding second position index can be established to record the physical or logical position of this level text in the original document where it is located.
[0077] If the first search level text corresponds to a first position index, or the second search level text corresponds to a second position index, then during knowledge retrieval, the search results can be traced back to the position in the original document pointed to by the first position index or the second position index, thereby locating the corresponding original text.
[0078] If the first retrieval level text corresponds to a first position index, and the second retrieval level text corresponds to a second position index, then during knowledge retrieval, if the second retrieval level is retrieved, the retrieval results can be traced back to the position in the original document pointed to by the second position index; if only the first retrieval level is retrieved, the retrieval results can be traced back to the position in the original document pointed to by the first position index, thereby locating the corresponding original text.
[0079] Furthermore, in the knowledge question and answer system, based on the location index, the content of the original text in the original document pointed to by the location index can be returned to the user as part of the answer in the form of reference or attachment. For example, the content of the original text can be displayed in the form of screenshots or PDF (Portable Document Format).
[0080] In the above implementation, a traceability mechanism can be implemented by introducing a location index. When the constructed knowledge base is used for knowledge retrieval, it supports rapid tracing back to the source document content based on the search results, and builds a trusted link from the original knowledge to the answer generation. Compared with the non-traceable solutions in related technologies that cannot locate the original documents, this has higher credibility.
[0081] In some implementations, the first vector data and the second vector data have a first mapping relationship. Figure 3a As shown, before constructing the knowledge base according to the first vector data and the second vector data, the knowledge base construction method may further include: S310: Determine a third classification domain corresponding to the original document according to a preset hierarchical division method.
[0082] S320, parsing the original document to obtain a third retrieval level text under a third classification domain; wherein the retrieval order of the second retrieval level text is before the third retrieval level text; and the granularity of the third retrieval level text is smaller than the granularity of the second retrieval level text.
[0083] S330, constructing third vector data based on the third retrieval level text; wherein there is a second mapping relationship between the second vector data and the third vector data.
[0084] Accordingly, constructing a knowledge base according to the first vector data and the second vector data may include: S340. Construct a knowledge base based on the first vector data, the second vector data, the first mapping relationship, the third vector data, and the second mapping relationship.
[0085] Among them, the first mapping relationship is used to represent the association between the first vector data and at least one second vector data, that is, at least one second vector data can be found from the specified first vector data.
[0086] The first mapping relationship can be determined according to the relationship between the first retrieval level text and the second retrieval level text in the document hierarchy (or the logical relationship in content).
[0087] The third classification domain corresponds to the third retrieval level.
[0088] The second mapping relationship is used to represent the association between the second vector data and at least one third vector data.
[0089] Specifically, for any original document, on the basis of determining the first classification domain and the second classification domain corresponding to the original document according to the preset hierarchical division method, the third classification domain can be further determined. By parsing the original document, the coarse-grained information related to the first classification domain can be extracted from the original document as the first retrieval level text corresponding to the original document, the medium-grained information related to the second classification domain can be extracted to obtain the second retrieval level text, and the fine-grained information related to the third classification domain can be extracted to obtain the third retrieval level text. It should be understood that the first retrieval level text, the second retrieval level text, and the third retrieval level text have a logical relationship in content.
[0090] Construct the first vector data based on the first retrieval level text, construct the second vector data based on the second retrieval level text, and construct the third vector data based on the third retrieval level text. According to the logical relationship in content among the first retrieval level text, the second retrieval level text, and the third retrieval level text, the first mapping relationship between the first vector data and the second vector data, and the second mapping relationship between the second vector data and the third vector data can be determined. Then, the knowledge base can be constructed based on the first vector data, the second vector data, the first mapping relationship, the third vector data, and the second mapping relationship.
[0091] Exemplarily, the original document is a maintenance manual for substation equipment, which includes contents such as "Chapter 1, Transformer Maintenance Regulations", "Chapter 2, Circuit Breaker Operation Specifications", "Chapter 3, Fault Emergency Handling Procedures", etc. Based on the first classification domain being "substation operation and maintenance" and the second classification domain including "transformer", "circuit breaker", and "fault emergency", the third classification domain is further determined. Taking the "transformer" classification domain as an example, assuming that the sub-domains of insulating oil sampling and bushing infrared temperature measurement are included in this domain, the third classification domain can be determined to include "insulating oil sampling" and "bushing temperature measurement".
[0092] When parsing this document, it can be divided according to the document structure, and then coarse-grained information can be extracted from the title, abstract, general principles, etc. of this document to obtain the first-level retrieval text under the first classification domain; taking "transformer" in the second classification domain as an example, medium-grained information is extracted from the content of this chapter to obtain the second-level retrieval text under this second classification domain; taking "insulating oil sampling" in the third classification domain as an example, fine-grained information is extracted from this paragraph to obtain the third-level retrieval text under this third classification domain, or this paragraph can be used as the third-level retrieval text.
[0093] Exemplarily, the original document is a technical standard document. Based on the first classification domain being technical standards and the second classification domain being general principles, technical requirements, and test methods, the third classification domain is further determined. Taking technical requirements as an example, assuming that the content of this part includes technical requirements corresponding to different equipment, the third classification domain can include different equipment types. The description of parsing the document to obtain each retrieval-level text is similar to that in the above example and will not be elaborated here.
[0094] Furthermore, the first-level retrieval text corresponds to a first position index, and / or, the second-level retrieval text corresponds to a second position index, and / or, the third-level retrieval text corresponds to a third position index. The first position index is used to indicate the position of the original text content corresponding to the first-level retrieval text in the original document; the second position index is used to indicate the position of the original text content corresponding to the second-level retrieval text in the original document; the third position index is used to indicate the position of the original text content corresponding to the third-level retrieval text in the original document.
[0095] It should be noted that the specific method for determining the first classification domain, second classification domain, and third classification domain corresponding to the original document according to the preset hierarchical division method can be determined according to the actual application scenario or requirements, etc., and is not specifically limited in this specification.
[0096] In the above embodiments, by hierarchically dividing the original document, multi-level vector data and mapping relationships are constructed to form a "pyramid-style" knowledge acquisition path, providing a knowledge base solution that combines retrieval efficiency and depth. It can effectively balance retrieval accuracy and timeliness. Compared with the flat knowledge base in related technologies, it can effectively solve the problems of poor retrieval timeliness and limited coverage, and effectively improve the efficiency of knowledge retrieval.
[0097] In some embodiments, the preset hierarchical division method is determined according to the field to which the original document belongs; the first retrieval level text is determined according to the document-level content of the original document. The second retrieval level text is determined according to the chapter-level content of the original document. The third retrieval level text is determined according to the paragraph-level content of the original document.
[0098] Among them, the document-level content refers to the content that can reflect the overall theme or main idea of the document, such as the abstract, document theme name, etc.
[0099] The chapter-level content refers to the content that reflects the main idea of the chapter after the document is structurally divided by chapters. It has a more detailed description than the document-level content, such as chapter titles, chapter text contents, etc.
[0100] The paragraph-level content refers to the content of the paragraphs after the document (or chapters in the document) is structurally divided by paragraphs, and may include paragraph text, charts, formulas, etc.
[0101] Specifically, by setting the preset hierarchical division method based on the professional field or subject field involved in the original document, the first classification domain, the second classification domain, and the third classification domain can be determined according to the professional field or subject field involved in the original document. The second classification domain is a sub-domain under the first classification domain, and the third classification domain is a sub-domain under the second classification domain.
[0102] For the same original document, by structurally dividing the document, the document-level content of the document (such as the document theme title, abstract, table of contents, or overview chapter content, etc.), the chapter-level content (such as chapter titles, chapter specific contents, etc.), and the paragraph-level content (such as specific paragraphs, sentences, charts, formulas, etc. in the document) can be obtained. By extracting information from the document-level content, the first retrieval level text can be obtained; by extracting information from the chapter-level content, the second retrieval level text can be obtained; by extracting information from the paragraph-level content, the third retrieval level text can be obtained.
[0103] Exemplarily, perform domain stratification and structuring on the documents to build a knowledge base based on the hierarchical structure of the knowledge base. Set domain subdivision rules, establish a three-level sub-domain system, divide the documents in the knowledge base according to the major industries into the first-level domain (corresponding to the first classification domain in this specification), and then further refine it into the second-level sub-domain (corresponding to the second classification domain in this specification) and the third-level fine-grained domain (corresponding to the third classification domain in this specification), and build a knowledge base with a pyramid-like structure in the form similar to the first-level directory, second-level directory, and third-level directory.
[0104] Taking the power industry as an example, the first-level domain can include the power transmission domain and the power distribution domain. Taking the power transmission domain as an example, referring to Figure 3b As shown, the sub-domains under this domain include facilities and equipment, management and operation, safety and risk, etc. The sub-domain of facilities and equipment further includes the fine-grained domains of equipment failure, aging and maintenance. The sub-domain of management and operation further includes the fine-grained domains of project management and resource allocation. The sub-domain of safety and risk further includes the fine-grained domain of risk assessment. Then, a three-level sub-domain system of "major professional domain - medium-grained sub-domain - fine-grained domain" can be established according to the above domain division method.
[0105] Among them, taking the equipment sub-domain as an example, the corresponding medium-grained sub-domain can be defined as the equipment type sub-domain, and the equipment types are such as insulators, conductors, etc. That is, the content related to equipment types such as insulators and conductors in the relevant original documents can be extracted to obtain the retrieval-level text under the equipment type sub-domain. The corresponding fine-grained domain under the equipment type sub-domain can be defined as the problem type fine-grained domain, and the problem types are such as corrosion detection, mechanical damage, etc.
[0106] Perform long document structuring on the original documents in the power transmission domain. For example, for technical standard documents, the chapters can be segmented according to the structure of "general rules - technical requirements - test methods", and for research reports, the paragraph units can be automatically extracted according to the structure of "problem description - mathematical model - conclusion". Information extraction can be performed on the content of the corresponding parts after structuring to obtain the retrieval-level text under each of the above sub-domains.
[0107] According to the retrieved hierarchical text in each sub-domain, a vector database with a pyramid-like structure is constructed using a multi-granularity vectorization strategy and a hierarchical indexing architecture method. Specifically, the multi-granularity levels include document level, section level, and paragraph level. The object of vectorization corresponding to the document level is the entire technical standard or research report, etc. (specifically, the first retrieved hierarchical text obtained by the above method), and the vector dimension can be set to 128, which is achieved using the BERT-CLS vector mean pooling technique; the object of vectorization corresponding to the section level is the section-level technical module (such as the "insulation coordination" and other technical modules in a technical standard document), and the vector dimension can be set to 256, which is achieved using the BiLSTM-Attention (Bidirectional Long Short-Term Memory with Attention Mechanism) semantic encoding technique; the object of vectorization corresponding to the paragraph level is a single knowledge point, formula, or chart, etc. in the paragraph, and the vector dimension can be set to 512, which is achieved using multi-modal fusion (such as text semantic analysis + formula parsing + graph convolutional processing, etc.) techniques.
[0108] Among them, special processing is required for some special content. For example, for formulas, the MathBERT can be used to extract symbolic logical relationship vectors; for charts, the ResNet-50 (Residual Network-50 layers) features and the Graph Attention Network (GAT) encoding topology can be used. It should be noted that in the above vectorization process, the vector dimension and the specific implementation technology can be determined according to the actual application scenario or requirements, etc.
[0109] A pyramid-style vector database is constructed based on the vectorized data to implement a hierarchical indexing architecture. Specifically, the hierarchical indexing architecture includes a top layer, a middle layer, and a bottom layer. The content stored in the top layer is the document-level vector (corresponding to the first vector data in this specification), and the retrieval function is to achieve wide-area retrieval, with a recall rate of over 95%; the content stored in the middle layer is the section-level vector (corresponding to the second vector data in this specification), and the retrieval function is to achieve exact matching (top-ten screening method); the content stored in the bottom layer is the paragraph-level vector and formulas or charts, etc., and the retrieval function is to achieve answer positioning with millimeter-level accuracy.
[0110] There is an associated mapping relationship between each level in the above vector database, that is, all chapters of the corresponding medium-grained subdomain are associated with the top-level document-level vector (for example, for a certain technical standard document in the power transmission field, all chapter-level vectors in the equipment type subdomain are associated with its corresponding top-level document-level vector); the middle-level chapter-level vectors link the bottom-level paragraph vector clusters; the bottom-level paragraph-level vectors and formulas or charts are inversely mapped to the original document positions. By using this hierarchical model to implement the cross-level retrieval path optimization algorithm, the task search scope is locked layer by layer, the search process is faster, and the search results are better.
[0111] Thus, domain indexing can be achieved according to the document type (here the document type can refer to the domain type to which the document belongs), and hierarchical indexing can be performed on the documents within the same large domain, improving the document recall rate, retrieval timeliness, and retrieval coverage.
[0112] In some embodiments, the original document corresponds to value level information, and the value level information is used to determine the retrieval priority corresponding to the original document; the knowledge base contains the value level information corresponding to the original document.
[0113] Among them, the value level information can be used to measure at least one attribute such as the importance, authority, timeliness, reliability, and content quality of the document.
[0114] Specifically, for each original document, the value level information corresponding to the document can be determined according to at least one of the importance, authority, timeliness, reliability, content quality, etc. of the document, and the document can be marked according to the value level information, so that the knowledge base contains the value level information. When performing knowledge retrieval, the vector data corresponding to the documents with higher value levels can be preferentially retrieved according to the value level information.
[0115] Furthermore, in knowledge retrieval, multiple retrieval results may be obtained, and it can also be determined according to the value level information which results correspond to the documents with higher value levels, so that the proportion of these results in answer generation can be increased.
[0116] Exemplarily, the value level information corresponding to a document is determined according to the document authority. Taking the power industry as an example, assuming that the authority of enterprise standard documents is the highest, the authority of industry standard documents is the second highest, the authority of paper documents is the third highest, and the authority of manual documents is the lowest, the weight coefficient of enterprise standard documents can be set to 0.6, and it can be used as the value level information of enterprise standard documents; the weight coefficient of industry standard documents is set to 0.5 and can be used as the value level information of industry standard documents; the weight coefficient of paper documents is set to 0.3 and can be used as the value level information of paper documents; the weight coefficient of manual documents is set to 0.2 and can be used as the value level information of manual documents. Further, taking paper documents as an example, the authority of SCI (Science Citation Index) papers is higher than that of general journal papers, so different value levels can also be set for papers in different journals.
[0117] Exemplarily, the value level information corresponding to a document is determined according to the document timeliness. Taking real-time operation and maintenance log documents as an example, a time decay factor can be set (for example, the time decay factor is set to 0.8 or 0.6, etc.), and the value level information of operation and maintenance log documents generated at different times is calculated based on the time decay factor.
[0118] The description of determining the value level information corresponding to a document according to other attributes is similar to the description in the above example, and will not be elaborated here.
[0119] It should be noted that the determination method of the value level information can be set according to the actual application scenario or requirements, etc., and is not specifically limited in this specification.
[0120] In some embodiments, the original document also corresponds to a confidentiality level; constructing a knowledge base according to the first vector data and the second vector data may include: constructing a knowledge base according to the first vector data, the second vector data, and the confidentiality level.
[0121] Among them, the confidentiality level can be used to identify the sensitivity degree or access permission level of the document.
[0122] Specifically, for each original document, the confidentiality level corresponding to the document can be determined according to at least one of the sensitivity degree or access permission level of the document, and the document is marked according to the confidentiality level. When constructing a knowledge base according to the first vector data, the second vector data, and the confidentiality level, the vector data corresponding to each original document can be subjected to confidentiality processing or non-confidentiality processing according to the confidentiality level.
[0123] It should be noted that for the data of the documents updated to the knowledge base subsequently, corresponding processing can be performed in the knowledge base according to the confidentiality level of the updated documents, without the need to reconstruct the knowledge base to achieve incremental updates.
[0124] Furthermore, there is a first mapping relationship between the first vector data and the second vector data; before constructing the knowledge base based on the first vector data and the second vector data, the knowledge base construction method may further include: determining a third classification domain corresponding to the original document according to a preset hierarchical division method; parsing the original document to obtain a third retrieval hierarchical text under the third classification domain; constructing third vector data based on the third retrieval hierarchical text; wherein, there is a second mapping relationship between the second vector data and the third vector data. Correspondingly, constructing the knowledge base according to the first vector data and the second vector data may include: constructing the knowledge base according to the first vector data, the second vector data, the first mapping relationship, the third vector data, the second mapping relationship, and the confidentiality level.
[0125] In some embodiments, the original documents whose confidentiality levels meet certain conditions and their corresponding vector data can be stored in the private domain in the knowledge base, or corresponding access permissions can be set for the original documents whose confidentiality levels meet certain conditions and their corresponding vector data, so that only users with access permissions can access and retrieve these data.
[0126] Exemplarily, the public database and the privacy database can be distinguished in the form of federated data. Taking the power industry as an example, considering the data privacy and security issues in the power industry and referring to the form of federated data, the materials can be divided into public materials and private domain materials. Public materials are materials that allow public access and query; while private domain materials are the internal materials of different enterprises respectively, and the private domain materials between different companies are not allowed to be accessed. Correspondingly, a public database for storing public materials and a privacy database for storing private domain materials are set up. The public database is a database that the public can normally access, and stores information such as industry standards and industry knowledge; the privacy database is opposite to the public database, stores the private data within the enterprise, and cannot be accessed through the public network, with a certain degree of confidentiality.
[0127] For the documents whose confidentiality levels are determined, the confidentiality levels can be directly assigned to the corresponding documents; for some unclassified documents, or the continuously generated internal enterprise documents, the confidentiality levels can be dynamically assigned to the documents. The confidentiality levels are mainly applied within the enterprise, and different departments or groups have different accessible confidential levels. In response to the questions of different users, the returned answers cannot involve the occurrence of information leakage, so it is necessary to store the private data confidentially.
[0128] For example, the confidentiality levels can be divided into top-secret level, internal level, and public level. The confidentiality level can be determined according to the document content. For example, for a top-secret document or a document with a classified document number, its confidentiality level can be set to the top-secret level; for a document containing sensitive words in the power industry (such as "load curve", "relay setting value", etc.), its confidentiality level can be set to the internal level; for a document containing temporary risk words such as fault events (such as insulator icing flashover accidents, etc.), its confidentiality level can be set to the internal level; for industry technical standard documents, its confidentiality level is set to the public level.
[0129] For top-secret level documents and their corresponding vector data, they can be stored in an encrypted manner using physical isolation + national cryptography (such as the SM4 symmetric encryption algorithm); for internal level documents and their corresponding vector data, they can be transmitted using in-domain access control + SSL (Secure Sockets Layer) protocol; for public level documents and their corresponding vector data, they are stored in a standard storage manner in a public database. The specific isolation storage method, encrypted storage method, access control method, etc. can be determined according to the actual application scenario and requirements, etc.
[0130] Furthermore, the risk value of the document can be automatically calculated by means of automatic identification of sensitive words, etc., so as to automatically assign a confidentiality level to the document, and the storage strategy of the corresponding document can be determined by means of a confidentiality level classification decision matrix, etc. The classification decision matrix can be specifically designed with reference to Table 1 below:
[0131] In some other embodiments, the original documents and their corresponding vector data with confidentiality levels meeting certain conditions can be encrypted and stored in the knowledge base.
[0132] It should be noted that the specific determination method of the confidentiality level can be set according to the actual application scenario or requirements, etc., and is not specifically limited in this specification.
[0133] The embodiments of this specification provide a knowledge question and answer processing method. Referring to Figure 4 as shown, the knowledge question and answer processing method may include the following steps.
[0134] S410. Obtain the question text data.
[0135] S420. In the knowledge base constructed according to the knowledge base construction method described in any of the foregoing embodiments, retrieve the question text data in the retrieval order of the retrieval level text to obtain a first retrieval result.
[0136] S430. Generate a first answer text corresponding to the question text data based on the first retrieval result.
[0137] Specifically, obtain the corresponding question text data according to the input content of the user's question. In the knowledge base constructed by the knowledge base construction method in any of the foregoing embodiments, retrieve the question text data according to the retrieval order of the retrieval level text in the multi-level knowledge structure, that is, first retrieve in the first vector data corresponding to the first retrieval level text, and continue to retrieve in the second vector data corresponding to the second retrieval level text after retrieving relevant content. If there is a third retrieval level text, continue to retrieve in the third vector data corresponding to the third retrieval level text after retrieving relevant content, and so on, to obtain the first retrieval result. Generate an answer according to the first retrieval result to generate the first answer text corresponding to the question text data and return it to the user.
[0138] It can be understood that the first retrieval level text corresponds to the first retrieval level, the second retrieval level text corresponds to the second retrieval level, and the retrieval order of the retrieval level text is the retrieval order of the retrieval level. Retrieving the question text data according to the retrieval order of the retrieval level text means retrieving the question text data according to the retrieval order of the retrieval level.
[0139] Exemplarily, taking the three-layer pyramid vector database constructed by the foregoing knowledge base construction method as an example, it is composed of a top layer storing document-level vectors, a middle layer storing chapter-level vectors, and a bottom layer storing paragraph-level vectors as well as formulas or charts. In this database, retrieve the question text data according to the retrieval order of the retrieval level text, that is, adopt a progressive retrieval mechanism to achieve three-stage precise positioning.
[0140] First, perform coarse-grained retrieval to quickly lock the relevant field through document-level vectors. Secondly, perform medium-grained matching to calculate the cosine similarity between the chapter vectors and the question text data. Then, perform fine-grained positioning to calculate the hybrid similarity in the associated paragraphs to obtain the corresponding retrieval result that meets the similarity condition.
[0141] Among them, the response time of the coarse-grained retrieval is within 50 milliseconds. According to the result of the medium-grained matching, the retrieval range can be further limited to the chapter vectors with a cosine similarity greater than a certain threshold with the question text data. For example, the threshold can be 0.75 or 0.85 or 0.9, etc. In the fine-grained positioning, the calculation method of the hybrid similarity S_final can refer to the following formula: S_final = 0.6 Text similarity + 0.3 Formula structure similarity + 0.1 Chart feature similarity Among them, the calculation methods of text similarity, formula structure similarity, and chart feature similarity can be specifically determined according to the actual application scenario or requirements, etc., and are not specifically limited in this specification.
[0142] It should be noted that the similarity matching and retrieval methods for the above different retrieval levels can be specifically determined according to the actual application scenario or requirements, etc., and are not specifically limited in this specification.
[0143] Furthermore, the timeliness weight corresponding to each original document in the knowledge base can be dynamically calculated, so that during retrieval, for the retrieval level text of the same retrieval level, the retrieval level text for priority retrieval can be determined in this retrieval level according to the timeliness weight corresponding to the corresponding original document, thereby further improving the retrieval timeliness. The calculation method of the timeliness weight W can refer to the following formula: W = α log(importance level) + β timeliness factor Among them, the coefficients α and β can be set according to the actual situation; for the description of the importance level in the formula, reference can be made to the description of the authority level in this specification, and for the description of the timeliness factor, reference can be made to the description of the time decay factor in this specification.
[0144] Alternatively, the calculation method of the timeliness weight can also refer to the following formula: W = 0.5 log2(importance level) + 0.3 (1 + e^(-0.2 Δt)) + 0.2 update_frequency Among them, Δt is the document update time difference, with the unit of month; update_frequency is the number of updates in the past six months.
[0145] In some cases, before retrieving the question text data according to the retrieval order of the retrieval level texts, the knowledge Q&A processing method may further include: determining the retrieval depth according to the question complexity corresponding to the question text data; determining the target retrieval level text in the knowledge base according to the retrieval depth; accordingly, retrieving the question text data according to the retrieval order of the retrieval level texts to obtain a first retrieval result, which may include: retrieving the question text data according to the retrieval order of each retrieval level text in the target retrieval level text to obtain a first retrieval result. The retrieval depth is the number of retrieval levels, starting from the first retrieval level and including at least one retrieval level. Through this method, dynamic depth adjustment during retrieval can be achieved. For simple questions, only the top-level retrieval can be triggered, such as simple questions like "What are the transmission line standards?"; for complex questions, it can automatically penetrate to the bottom layer for multi-level sequential retrieval, such as complex questions like "Derivation of the calculation formula for the ice flash probability of ±800kV insulators". Regarding the determination method of question complexity, at least one of methods such as text feature analysis, syntax structure analysis, and intention recognition can be adopted, which is not specifically limited in this specification.
[0146] In some embodiments, referring to Figure 5 as shown, the knowledge Q&A processing method may further include: S510. Determine the target original document corresponding to the first retrieval result according to the first retrieval result.
[0147] S520. Generate a text attachment corresponding to the first retrieval result according to the first retrieval result and the target original document.
[0148] The target original document is the source document containing the information of the first retrieval result.
[0149] The text attachment can be used to be returned to the user together with the first answer text.
[0150] Specifically, the first retrieval result may be associated with the information of the corresponding original document. According to the first retrieval result, the original document containing the information of the first retrieval result can be located and used as the target original document. According to the first retrieval result, the content related to the first retrieval result can be screened out from the target original document, and the text attachment corresponding to the first retrieval result can be generated according to the screened content.
[0151] In some embodiments, generating the text attachment corresponding to the first retrieval result according to the first retrieval result and the target original document can be achieved through the foregoing position index. Specifically, the part closely related to the first retrieval result can be located in the target original document according to the position index corresponding to the first retrieval result, and then the text attachment can be generated according to the relevant part.
[0152] In some other embodiments, to generate a text attachment corresponding to the first retrieval result based on the first retrieval result and the target original document, it is possible to parse and extract the part corresponding to the first retrieval result from the target original document to generate the corresponding text attachment.
[0153] It should be noted that the text attachment can be a screenshot, or a PDF file, an editable electronic document, etc.
[0154] In the above embodiments, by generating a text attachment corresponding to the first retrieval result, on the basis of the conventional output of the answer text, the data content of the original document on which the output answer content is based is increased. The text attachment can be presented in the form of a screenshot, PDF, etc., thereby realizing the output of the traceability mechanism and improving the credibility of the output result.
[0155] In some embodiments, the first retrieval result corresponds to a first retrieval range; as Figure 6 shown, the knowledge question and answer processing method may further include: S610. Evaluate the first answer text to obtain an evaluation result.
[0156] S620. If the evaluation result does not meet the preset conditions, determine a second retrieval range according to the first retrieval range.
[0157] S630. Based on the second retrieval range, retrieve the question text data in the knowledge base to obtain a second retrieval result, so as to generate a second answer text according to the second retrieval result.
[0158] Among them, the retrieval range is a vector search space, which can be determined according to the search depth and semantic radius. The first retrieval range is smaller than the second retrieval range.
[0159] The evaluation result not meeting the preset conditions means that the quality of the first answer text does not reach the set quality standard, and the evaluation result can be in the form of a score or a grade, etc.
[0160] Specifically, set a preset evaluation rule for evaluating whether the answer text generated according to the retrieval result meets the requirements corresponding to the question input by the user. Evaluate the first answer text obtained after retrieval based on the first retrieval range according to the preset evaluation rule to obtain the corresponding evaluation result. If the evaluation result corresponding to the first answer text meets the preset conditions, it means that the quality of the first answer text meets the user's needs, and then the first answer text can be returned to the user.
[0161] If the evaluation result does not meet the preset conditions, it means that the quality of the first answer text does not meet the user's needs, then a larger second retrieval range is determined according to the first retrieval range, so as to retrieve the question text data in the knowledge base again based on the second retrieval range to obtain a second retrieval result, and generate a second answer text according to the second retrieval result.
[0162] Furthermore, evaluate the second retrieval result according to a preset evaluation rule to obtain its corresponding evaluation result. If the evaluation result corresponding to the second retrieval result meets the preset conditions, return the second answer text to the user; if it does not meet the preset conditions, determine a larger third retrieval range according to the second retrieval range, and re-retrieve the question text data in the knowledge base based on the third retrieval range to obtain a third retrieval result, and generate a third answer text according to the third retrieval result. And so on, using this loop verification mechanism (or, called self-reflection and error correction mechanism), until an answer text that meets the user's needs is obtained, or until the number of loop verification times reaches a preset number threshold.
[0163] In some embodiments, to evaluate the first answer text, a Reward (reward) scoring model can be used, and the evaluation result is the score output by the model. Correspondingly, the evaluation result not meeting the preset conditions can specifically be that the score corresponding to the first answer text is less than a preset score threshold.
[0164] Exemplarily, introducing the Reward model scoring mechanism to score and evaluate the question-and-answer output result and the user's input question-and-answer requirements, a 0-5 score range can be set. The training process of the Reward model can include: inputting question samples into the knowledge question-and-answer system to obtain the answers generated by the knowledge question-and-answer system as answer samples. Manually score the answer samples, and use the scoring results as the annotations of the answer samples. This stage is recorded as the annotation process. Input the question text and the corresponding answer samples into an untrained initial model for scoring, so that the initial model outputs the initial scoring result corresponding to the answer samples. Train the initial model according to the initial scoring result output by the model and the scoring annotation corresponding to the answer samples, converge the model parameters until the training stop condition is met, obtain the trained Reward model, and apply this Reward model to the actual scoring scenario.
[0165] For the first answer text, when the score value output by the Reward model is greater than or equal to 3 points, it means that the judgment model determines that the output result is relatively in line with the question answer, and then the answer text can be directly returned to the user; when the score value is less than 3 points, enter the loop generation link, expand the retrieval range on the basis of the first retrieval range, determine the second retrieval range, realize automatically expanding the vector search space, and enhance the granularity of the prompt (prompt word) content. In addition, a threshold for the number of loop verifications can be set. For example, the maximum number of loop verifications can be default set to 3 times (or 5 times or 6 times, etc.), to find a balance point in ensuring the accuracy and timeliness of the output content. The threshold for the number of loop verifications can be specifically determined according to the actual application scenario or requirements, etc.
[0166] The scoring dimensions for different score ranges can be referred to in Table 2 below:
[0167] The retrieval range can be determined according to the search depth (e.g., the number of retrieved texts or the retrieval volume) and the semantic radius. The initial value D0 of the search depth can be set to 50, and the initial value R0 of the semantic radius can be set to 0.8.
[0168] When the Reward score S is less than 3, the search range can be expanded according to the multi-level expansion retrieval parameter table shown in Table 3 below:
[0169] For example, the user's question is: "How to calculate the annual failure rate of insulators?" The first-round answer is: "According to operation experience, the failure rate of insulators is usually between 0.1% and 0.3%." The Reward score is 2.1 (lacking calculation formula and standard basis).
[0170] The score of the first-round answer is less than 3, so the first expansion retrieval is performed. Parameter adjustment is carried out, the search depth is adjusted from 50 to 75, the semantic radius is adjusted from 0.8 to 0.9, and the Prompt enhancement strategy is to append "The calculation formula in DL / T436 standard is required". On this basis, the second-round answer is obtained: "According to Article 5.3.2 of DL / T436-2018, the calculation formula for the annual failure rate λ is λ = N_fault / (N_total L), where N_fault is the number of faulty pieces, N_total is the total number, and L is the operation years. It is recommended to take 0.05% - 0.2%". At this time, the Reward score is 4.3, and it is considered to meet the user's needs, so the answer of this round of answer is returned to the user.
[0171] In some cases, the search depth can also be defined as D = log2(N) × 0.8, the initial value of N is 50, and N is expanded in multiples of 1.5, 2.2, and 3.0 as described above.
[0172] Furthermore, the retrieval range can also be determined according to the search depth, semantic radius, and time range. Among them, the time range represents the publication or update time range of the document. For example, the time range of 6 months means that the documents within the retrieval range include the documents published or updated in the past 6 months. The time range can be defined as T = min(1, Δt / 36), and the initial value of Δt is 6 months.
[0173] By introducing the above self-reflection and error-correction mechanism, the hallucination rate is less than 1.2%, the average number of response iterations is 1.8 times, and the knowledge base update delay is less than 2 hours (from the release of the new standard to storage), effectively improving the overall performance of the system. In addition, an optimization mechanism can be set. For example, when the same type of question triggers extended retrieval 3 times in a row, a special review of the knowledge base can be automatically initiated, and a standard timeliness warning can be established to prompt the specification documents that will be abolished 6 months in advance, etc.
[0174] Thus, through the loopback verification mechanism, adopting the scoring - extended retrieval process, the system's self-reflection and continuous improvement capabilities are improved. Through error pattern recognition, user feedback learning, and performance monitoring and analysis, the model performance and knowledge base quality are continuously optimized. This mechanism ensures a high accuracy rate for professional Q&A in related fields, effectively reducing the hallucination risk in professional fields and providing reliable and authoritative professional support for users.
[0175] In some embodiments, the question text data is obtained by vectorizing the original question data, and the original question data includes at least one of text data, image data, and voice data.
[0176] Among them, in this embodiment, the question text data is the standard input text finally used for knowledge base retrieval.
[0177] The original question data can be multi-modal data.
[0178] Specifically, the original question data input by the user can be at least one of different modal data such as text, image, and voice. Vectorizing the original question data can be transformed into question text data in a standard format for knowledge base retrieval. In this way, the problem that the single text input interaction mode in related technologies does not support multi-modal input, resulting in an increase in the operation complexity of front-line personnel, can be solved, and the interaction threshold can be reduced.
[0179] Exemplarily, taking the power industry as an example, this specification provides an improved power scenario RAG method and system. Refer to Figure 7a As shown, for the structure diagram of the system, the system mainly includes an input module, a data classification and vectorization module, a base model, a loopback generation module, and an output module.
[0180] In the input module, multiple modal inputs are supported, and the input types can include: text type, picture type, and voice type. In the picture type, special formats such as color drawings (design drawings), infrared images (equipment thermal imaging), text images (scanned documents), etc. are supported, as well as pictures taken on-site routinely; in the voice type, multiple spoken language inputs such as Mandarin Chinese, English, Cantonese, Sichuan dialect, etc. are covered. In this module, through technologies such as image recognition (such as OCR + infrared feature extraction, etc.) and dialect voice-to-text, natural interaction in the front-line operation scenario can be achieved.
[0181] The data classification and vectorization module is mainly used to vectorize the hierarchical structure of the knowledge base. The knowledge base is divided according to major industries and further refined to construct a pyramid-like structure in the form of first-level directories, second-level directories, third-level directories, etc. For specific descriptions, please refer to the description of knowledge base construction in this specification. Additionally, considering data privacy and security issues in the power industry, as shown in Figure 7b , the knowledge base refers to the form of federated data for federated data processing. Specifically, the data is divided into public data and private domain data. Public data is used to construct a shared vector library, allowing public access and query; private domain data is used to construct a private domain vector library, and private domain data and private domain vector libraries between different departments are not allowed to be accessed (for example, Figure 7b the vector library of Unit A constructed based on the private domain data of Unit A can only be accessed by Unit A, and the vector library of Unit B constructed based on the private domain data of Unit B can only be accessed by Unit B). Thus, a cross-domain access control protocol based on the zero-knowledge proof verification process can be realized.
[0182] The base model mainly uses the DeepSeek large model to achieve question-and-answer output, and uses this model to realize input-output mapping. The main processes include multi-modal input, joint feature space mapping, vector retrieval results, and prompt engineering optimization. The output forms include text answers (which can include citation annotations of papers / books / specification documents) and traceability attachments in the form of PDFs or screenshots.
[0183] The scoring-expansion retrieval process is implemented based on the loop generation module. As shown in Figure 7c , the scoring-expansion retrieval process mainly includes: performing Reward scoring based on the answer text output by the DeepSeek large model. If the score is greater than or equal to 3 points, the system outputs the text and traceability attachments in the form of PDFs or screenshots; if the score is less than 3 points, the vector search space is expanded, the Prompt granularity is enhanced, and answers are retrieved and generated again. Among them, the scoring model can realize a hybrid evaluation based on semantic similarity (such as measured by the BERTScore metric) and domain rules (such as the matching degree of power terms). This scoring-expansion retrieval mechanism can also be called the Agent verification loop.
[0184] As shown in Figure 7dAs shown in the figure, the main process of the power scenario RAG method includes: First, perform modality recognition based on the user input data. If text data is recognized, perform semantic parsing; if image data is recognized, perform visual feature extraction; if voice data is recognized, perform speech recognition. Second, perform joint vector retrieval. Through hierarchical indexing, perform document-level vector retrieval in the first-level index (i.e., the first retrieval level), perform chapter-level vector retrieval in the second-level index (i.e., the second retrieval level), and perform paragraph-level vector retrieval in the third-level index (i.e., the third retrieval level). Generate an answer based on the retrieval results, then enter the Agent verification loop, and finally output the answer and the traceable PDF or screenshot.
[0185] The power scenario RAG method and system provided in this specification have the following advantages: (1) Adopt federated data to solve the problem in related technologies that the centralized knowledge base architecture mixes private data and public data for storage and is easily accessible by unauthorized parties, thus resolving the contradiction between data privacy and sharing needs in the power industry and reducing the risk of cross-unit leakage of private data; (2) Solve the balance problem between retrieval accuracy and timeliness through hierarchical vectorization and pyramid indexing, and improve the recall rate of relevant documents; (3) Support multi-modal input. Aiming at the problem in related technologies that the interaction form is single and cannot meet the scenario requirements such as infrared image recognition and multi-language voice input for front-line personnel, it can improve the usability of multi-modal interaction in complex on-site environments; (4) Introduce Agent loop verification to evaluate the generated answer and adaptively increase the search scope. Compared with the lack of self-reflection and optimization mechanisms in related technologies, it can effectively improve the answer quality; (5) Implement a traceable mechanism. By returning the function of the original document PDF / image fragment, it is closer to the requirements of industrial scenarios and increases the interaction experience. Compared with the method in related technologies that cannot return the original document for users to quickly consult, it reduces the doubts of users about the returned results and improves the credibility of the answer.
[0186] The power scenario RAG method and system provided in this specification comprehensively consider the characteristics of data privacy, authority, traceability, multi-modal input, etc. in the power industry itself, and specifically give a unique and innovative solution, which has great promotion and application value.
[0187] The embodiments of this specification provide a knowledge base construction device, which may include: an original document acquisition module, a retrieval level text acquisition module, a vector data construction module, and a knowledge base construction module.
[0188] The original document acquisition module is used to acquire the original document.
[0189] A retrieval level text acquisition module, which is used to divide the original document according to a preset level division method to obtain a first retrieval level text and a second retrieval level text; wherein, the retrieval order of the first retrieval level text is before that of the second retrieval level text.
[0190] A vector data construction module, which is used to construct first vector data based on the first retrieval level text and construct second vector data based on the second retrieval level text.
[0191] A knowledge base construction module, which is used to construct a knowledge base according to the first vector data and the second vector data.
[0192] The embodiments of this specification also provide a knowledge question and answer processing device, which may include: a question text acquisition module, a retrieval result acquisition module, and an answer text generation module.
[0193] A question text acquisition module, which is used to acquire question text data.
[0194] A retrieval result acquisition module, which is used to retrieve the question text data in the knowledge base constructed according to the knowledge base construction method described in any of the foregoing embodiments in accordance with the retrieval order of the retrieval level text to obtain a first retrieval result.
[0195] An answer text generation module, which is used to generate a first answer text corresponding to the question text data based on the first retrieval result.
[0196] For the specific limitations on the knowledge base construction device, reference may be made to the limitations on the knowledge base construction method in the foregoing text. For the specific limitations on the knowledge question and answer processing device, reference may be made to the limitations on the knowledge question and answer processing method in the foregoing text, which will not be elaborated here. Each module in the above-mentioned knowledge base construction device and knowledge question and answer processing device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0197] The embodiments of this specification also provide a computer device, which may include a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the knowledge base construction method or the knowledge question and answer processing method in any of the foregoing embodiments.
[0198] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, it implements the knowledge base construction method or the knowledge question and answer processing method in any of the foregoing embodiments.
[0199] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0200] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.
[0201] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0202] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0203] In the present invention, unless otherwise clearly specified and defined, terms such as "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0204] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for constructing a knowledge base, characterized in that, The method includes: Obtaining an original document; Dividing the original document according to a preset hierarchical division method to obtain a first retrieval hierarchical text and a second retrieval hierarchical text; wherein, the retrieval order of the first retrieval hierarchical text is before that of the second retrieval hierarchical text; Constructing first vector data based on the first retrieval hierarchical text, and constructing second vector data based on the second retrieval hierarchical text; Constructing a knowledge base according to the first vector data and the second vector data.
2. The knowledge base construction method according to claim 1, wherein The dividing the original document according to a preset hierarchical division method to obtain a first retrieval hierarchical text and a second retrieval hierarchical text includes: Determining a first classification domain and a second classification domain corresponding to the original document according to the preset hierarchical division method; Parsing the original document to determine the first retrieval hierarchical text under the first classification domain and the second retrieval hierarchical text under the second classification domain; wherein, the granularity of the second retrieval hierarchical text is smaller than that of the first retrieval hierarchical text.
3. The knowledge base construction method according to claim 2, wherein The first retrieval hierarchical text corresponds to a first position index, and / or, the second retrieval hierarchical text corresponds to a second position index; The first position index is used to indicate the position of the original text content corresponding to the first retrieval hierarchical text in the original document; The second position index is used to indicate the position of the original text content corresponding to the second retrieval hierarchical text in the original document.
4. The knowledge base construction method according to claim 2, characterized in that There is a first mapping relationship between the first vector data and the second vector data; Before constructing the knowledge base according to the first vector data and the second vector data, the method further includes: Determining a third classification domain corresponding to the original document according to the preset hierarchical division method; Parsing the original document to obtain a third retrieval hierarchical text under the third classification domain; wherein, the retrieval order of the second retrieval hierarchical text is before that of the third retrieval hierarchical text; the granularity of the third retrieval hierarchical text is smaller than that of the second retrieval hierarchical text; Constructing third vector data based on the third retrieval hierarchical text; wherein, there is a second mapping relationship between the second vector data and the third vector data; The constructing the knowledge base according to the first vector data and the second vector data includes: Constructing the knowledge base according to the first vector data, the second vector data, the first mapping relationship, the third vector data and the second mapping relationship.
5. The knowledge base construction method according to claim 4, wherein The preset hierarchical division method is determined according to the field to which the original document belongs; the first retrieval hierarchical text is determined according to the document-level content of the original document; The second retrieval hierarchical text is determined according to the chapter-level content of the original document; The third retrieval hierarchical text is determined according to the paragraph-level content of the original document.
6. The knowledge base construction method according to any one of claims 1 to 5, characterized in that The original document corresponds to value level information, and the value level information is used to determine the retrieval priority corresponding to the original document; The knowledge base contains the value level information corresponding to the original document.
7. The knowledge base construction method according to any one of claims 1 to 5, characterized in that The original document also corresponds to a confidentiality level; constructing the knowledge base according to the first vector data and the second vector data includes: Constructing the knowledge base according to the first vector data, the second vector data, and the confidentiality level.
8. A method for processing knowledge questions and answers, characterized in that The method includes: Obtaining problem text data; Retrieving the problem text data in the knowledge base constructed by the knowledge base construction method according to any one of claims 1 to 7 in accordance with the retrieval order of the retrieval level text to obtain a first retrieval result; Generating a first answer text corresponding to the problem text data based on the first retrieval result.
9. The knowledge question and answer processing method according to claim 8, wherein The method further includes: Determining a target original document corresponding to the first retrieval result according to the first retrieval result; Generating a text attachment corresponding to the first retrieval result according to the first retrieval result and the target original document.
10. The knowledge question and answer processing method according to claim 8, characterized in that The first retrieval result corresponds to a first retrieval range; the method further includes: Evaluating the first answer text to obtain an evaluation result; If the evaluation result does not meet the preset conditions, determining a second retrieval range according to the first retrieval range; Retrieving the problem text data in the knowledge base based on the second retrieval range to obtain a second retrieval result, so as to generate a second answer text according to the second retrieval result.
11. The knowledge Q&A processing method according to any one of claims 8 to 10, characterized in that, The problem text data is obtained by performing vectorization processing on original problem data, and the original problem data includes at least one of text data, image data, and voice data.
12. A knowledge base construction device, characterized in that The device includes: An original document acquisition module, configured to acquire an original document; A retrieval level text acquisition module, configured to divide the original document according to a preset level division method to obtain a first retrieval level text and a second retrieval level text; wherein, the retrieval order of the first retrieval level text is before that of the second retrieval level text; A vector data construction module, configured to construct first vector data based on the first retrieval level text and construct second vector data based on the second retrieval level text; A knowledge base construction module, configured to construct a knowledge base according to the first vector data and the second vector data.
13. A knowledge Q&A processing device, characterized in that, The device includes: A problem text acquisition module, configured to acquire problem text data; A retrieval result acquisition module, configured to retrieve the problem text data in the knowledge base constructed by the knowledge base construction method according to any one of claims 1 to 7 in accordance with the retrieval order of the retrieval level text to obtain a first retrieval result; An answer text generation module, configured to generate a first answer text corresponding to the problem text data based on the first retrieval result.
Citation Information
Patent Citations
Literature retrieval method and device based on knowledge graph, electronic equipment and medium
CN113590845A
Document duplicate checking method based on hierarchical feature vector search
CN117951256A
Text data processing method and device, electronic equipment and nonvolatile storage medium
CN118797011A
Construction method, system and device of retrieval enhancement generation system and medium
CN118797060A
Domain speech recognition method and system based on RAG
CN119296516A