Engineering knowledge base construction and deep retrieval method
By constructing a multi-level tree structure and embedding vectors, the retrieval process of the engineering knowledge base is optimized, solving the problems of granularity mismatch and information ambiguity in engineering knowledge retrieval and achieving high-quality knowledge retrieval results.
Patent Information
- Application Number
- CN202511407762.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-29
AI Technical Summary
Knowledge retrieval in the engineering field faces problems such as granularity mismatch and information ambiguity. Traditional RAG technology is difficult to meet the needs of high-quality retrieval, including insufficient quality of knowledge chunks, insufficient retrieval relevance, and insufficient retrieval completeness.
An engineering knowledge base is constructed by performing multi-level tree structure recognition and heterogeneous content parsing on engineering knowledge documents to generate embedding vectors. Combining path embedding and content embedding, the relationship between knowledge fragments and nodes is stored. Furthermore, a list of knowledge gaps is generated through query rewriting to optimize the retrieval process.
This improved the quality of knowledge blocks in the engineering knowledge base, enhanced the relevance and completeness of retrieval, and improved the accuracy and efficiency of knowledge retrieval.
Smart Images

Figure CN120893548A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent retrieval, in particular to an engineering knowledge base construction and deep retrieval method. BACKGROUND
[0002] In the whole life cycle of an engineering project, from the early design, the intermediate construction to the later operation and maintenance, relevant technical personnel will raise a large number of professional problems around the core dimensions of cost control, quality control, safety specification, etc. The answers to these problems are often scattered in industry standards, technical specifications, project reports, inspection records and other types of engineering documents, so there is an urgent need for efficient and accurate knowledge retrieval technology to support business decision-making.
[0003] Compared with the knowledge retrieval scene in general fields, the knowledge retrieval in the engineering field faces more prominent technical challenges, and the core contradictions are reflected in two aspects: one is the granularity mismatch problem. Engineering documents are usually written according to a systematized chapter structure (such as specification clauses, technical chapters, etc.), which is lengthy and covers multi-dimensional content, while a single query usually focuses on specific technical details, scenario-based problems or specific indicators, resulting in a natural mismatch between the overall structure of the document and the local needs of the query. The second is the information expression ambiguity problem. In the engineering field, the same technical concept, indicator or process often has multiple dimensional expression methods (such as professional terms, colloquial names, scenario-based descriptions, etc.), which significantly increases the difficulty of information matching in the retrieval process.
[0004] Retrieval Augmentation Generation (RAG) technology, as an important retrieval paradigm in the era of large language models, provides a new path for domain knowledge retrieval. Its core mechanism is to integrate domain knowledge into large language models through the process of document blocking-vector mapping-similar retrieval-generating augmentation: first, the domain document is divided into several knowledge segments and converted into vectors, then the relevant content is retrieved based on the semantic similarity between the query and the segment, and finally the results with accuracy and summary are output by the generative large model. In theory, RAG technology can alleviate the granularity mismatch problem through domain knowledge injection, and reduce the impact of information ambiguity with the semantic understanding ability of large language models.
[0005] However, the traditional RAG technology still has significant limitations in the practical application of the engineering field, making it difficult to meet the high-quality retrieval needs:
[0006] (1) Insufficient quality of knowledge blocking
[0007] Traditional RAG relies on simple segmentation of fixed length or punctuation order, completely ignoring the inherent hierarchical structure of engineering documents (such as chapter titles, clause levels, etc.). This segmentation mode is prone to two problems: one is semantic break, which splits logically related content into different fragments, or combines unrelated content into the same fragment, resulting in a lack of semantic integrity of knowledge fragments; the second is single feature utilization, which only relies on text semantic features for retrieval, ignoring key structural features such as document chapter relationships and logical structures, thereby interfering with retrieval accuracy and affecting the quality of the final answer.
[0008] (2) Insufficient search relevance
[0009] Traditional RAG directly uses the user's original question as the search query, lacking depth analysis of the query intent. This "surface matching" mode does not distinguish the relevance of paragraph content and the core needs of the query, which may lead to the fact that the key knowledge related to the user's intent is not retrieved, and irrelevant information is included in the results, reducing the accuracy of the search.
[0010] (3) Insufficient search completeness
[0011] On the one hand, traditional RAG uses a one-time search generation mode, lacking iterative optimization and thinking mechanisms for the search process, making it difficult to discover potential knowledge gaps, resulting in necessary knowledge related to implicit demand being missed; on the other hand, the semantic dimension of a single knowledge fragment is limited, often lacking context information such as pre-clause and associated technical descriptions that support its understanding, and these contexts may contain key related knowledge, ultimately resulting in insufficient semantic integrity of the search results. SUMMARY
[0012] To solve the above technical problems, the technical solution adopted by the present application is:
[0013] The engineering knowledge base construction and deep retrieval method provided by the embodiments of the present application comprises the following steps:
[0014] S1, constructing an engineering knowledge base based on engineering knowledge documents.
[0015] S2, when receiving a user question, generating a knowledge blank question list through query rewriting, traversing the list to search the engineering knowledge base, grouping knowledge, sorting relevance, and judging the termination of the cycle, and after the termination of the cycle, generating a reply content for the user question based on the accumulated existing knowledge in the cycle.
[0016] Among them, S1 specifically includes:
[0017] S11, identifying the knowledge level of each input engineering knowledge document to form a multi-level tree structure, wherein each chapter title and its corresponding chapter content in the engineering knowledge document constitute a node of the multi-level tree structure.
[0018] S12, respectively analyze the heterogeneous content of each node to generate knowledge fragments.
[0019] S13, generate the embedding vector of each knowledge fragment based on the combination of path embedding and content embedding.
[0020] S14, store the association relationship of knowledge fragments and nodes, the hierarchical relationship between nodes, the mapping relationship between nodes and source documents, and the vector data of knowledge fragments, related information and engineering knowledge document original information, and complete the construction of the engineering knowledge base.
[0021] The engineering knowledge base construction and deep retrieval method provided by the embodiment of the application includes two core processes of knowledge base construction and deep retrieval: in the knowledge base construction stage, the knowledge hierarchy of the engineering knowledge document is recognized to form a multi-level tree structure with chapter titles and corresponding contents as nodes; the heterogeneous content (text, table, image, etc.) of each node is analyzed to generate knowledge fragments; the embedding vector of the knowledge fragment is generated by combining path embedding and content embedding, and the association relationship between knowledge, vector data and original information are stored to complete the construction of the knowledge base. In the deep retrieval stage, when the user question is received, the knowledge blank question list is generated by query rewriting, the knowledge base is retrieved by traversing the list, the retrieval process is optimized by knowledge grouping, relevance sorting and loop termination judgment, and finally the reply content verified by traceability is generated and output based on the accumulated existing knowledge. The structured knowledge organization, deep analysis of heterogeneous content and intelligent retrieval optimization improve the reuse efficiency and retrieval accuracy of engineering knowledge.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the application, nor is it intended to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0024] Figure 1 The flowchart of the engineering knowledge base construction and deep retrieval method provided by the embodiment of the application;
[0025] Figure 2 The construction schematic diagram of the engineering knowledge base in the embodiment of the application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0028] It should be noted that some of the example embodiments are described as processes or methods depicted as flow diagrams. Although the flow diagrams describe the processes as a sequential process, many of the steps can be performed in parallel, concurrently or simultaneously. In addition, the order of the steps can be re-arranged. The processes can be terminated when their operations are completed, but can also have additional steps not included in the figure, which can also be performed after the operations of the processes are completed. The processes can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0029] The present application aims to realize the establishment of a high-quality knowledge base in the engineering field and high-quality retrieval of knowledge in the engineering field, and help to improve the retrieval efficiency and quality of knowledge in the engineering field for relevant personnel in the engineering field.
[0030] The key technical problems to be solved by the present application include: quality problems of the engineering knowledge base, better analysis technical means need to be proposed for the engineering field knowledge documents, the quality of knowledge segment segmentation and blocking is improved, and the semantic integrity of the knowledge segment is improved; the relevance problem of knowledge retrieval, considering implementing which technology to ensure that the knowledge retrieval result is highly relevant to the user query; the completeness problem of knowledge retrieval, better knowledge retrieval means need to be proposed and applied to ensure the knowledge completeness and sufficiency of the retrieval result, and then the quality of knowledge question and answer is improved.
[0031] The present application takes the basic principle of RAG technology as a baseline, proposes a new high-quality knowledge base construction and deep retrieval technology for the engineering field, improves the knowledge blocking quality problem of the engineering knowledge base, solves the problem of insufficient quality of knowledge retrieval caused by the mismatch of granularity between the query and the engineering knowledge and the inherent fuzziness of information, and improves the relevance, completeness and accuracy of the knowledge retrieval query of the engineering knowledge base.
[0032] Further, the present application discloses an engineering knowledge base construction and deep retrieval method, such asFigure 1 As shown, the method comprises the following steps:
[0033] S1, constructing an engineering knowledge base based on engineering knowledge documents.
[0034] When receiving a user question, a knowledge blank question list is generated by query rewriting, the engineering knowledge base is searched by traversing the list, knowledge grouping, relevance sorting, and loop termination judgment are performed, and after the loop termination, reply content for the user question is generated based on the accumulated existing knowledge in the loop process.
[0035] S2, when receiving a user question, a knowledge blank question list is generated by query rewriting, the engineering knowledge base is searched by traversing the list, knowledge grouping, relevance sorting, and loop termination judgment are performed, and after the loop termination, reply content for the user question is generated based on the accumulated existing knowledge in the loop process.
[0036] Further, as shown, Figure 2 S1 specifically comprises:
[0037] S11, identifying the knowledge level of each input engineering knowledge document to form a multi-level tree structure, wherein each chapter title and its corresponding chapter content in the engineering knowledge document constitute a node of the multi-level tree structure.
[0038] Engineering knowledge documents refer to various structured or unstructured materials generated in the whole life cycle (design, research and development, construction, operation and maintenance, acceptance, technical summary, etc.) in the engineering field (such as mechanical engineering, civil engineering, electrical engineering, software engineering, etc.) for carrying, recording, and transferring engineering professional knowledge. From the content type, engineering knowledge documents usually cover heterogeneous forms, including but not limited to:
[0039] Text: technical specifications, design specifications, construction plans, process manuals, fault analysis reports, technical disclosure documents, academic papers, patent documents, etc.
[0040] Tables: engineering parameter tables, material performance data tables, test record tables, progress plan tables, etc.
[0041] Graphics / images: design drawings (CAD drawings, BIM model views), flowcharts, schematics, equipment structure diagrams, experimental data visualization charts, and on-site working condition photos.
[0042] Other structured data: such as parameter configuration files exported by engineering databases, data conclusions in simulation analysis reports, etc.
[0043] From the perspective of knowledge attributes, engineering knowledge documents carry content with professionalism, hierarchy, and correlation: professionalism is reflected in the focus on engineering technical principles, standards, and practical experience; hierarchy is reflected in the logical division of documents into chapters, items, etc. (such as the "chapter-section-item" structure of a manual), forming a natural knowledge hierarchy; correlation is reflected in the technical dependence between different documents and chapters (such as the correlation between design specifications and construction plans, and the correlation between fault cases and maintenance manuals).
[0044] In the present application, engineering knowledge documents are the core raw materials for building an engineering knowledge base. Through knowledge level identification, heterogeneous content analysis, and vector representation, the structured deposition and efficient reuse of knowledge are achieved, providing basic data support for knowledge retrieval and decision support in the engineering field.
[0045] Specifically, S11 specifically includes:
[0046] S111, extracting chapter titles from engineering knowledge documents through heuristic rules, the heuristic rules including: determining titles based on document built-in label attributes, determining titles based on title starting character labels, and determining titles based on text style features.
[0047] Among them, determining titles based on document built-in label attributes includes: for docx, doc documents, parsing the document structure through the Python package docx to extract the built-in label attributes of the text line; if the built-in label attributes of the text line are Heading1, Heading2, etc. hierarchical title labels, then directly determine that the text line is a chapter title of the corresponding level.
[0048] Determining titles based on title starting character labels includes: matching the starting characters of the text line through regular expressions, if the starting characters are preset title identifiers (such as Arabic numerals + decimal points "1.1", Chinese character numbers "one," "(1)", English letters "A." "a)", etc.), then determine that the text line is a chapter title.
[0049] Determining titles based on text style features includes: for docx, doc documents, reading the text style through the Python package docx, and for pdf files, extracting the text style through the Python package pdfplumber; if the font size of the text line is greater than a preset threshold (such as greater than 12-point font), and has features such as bold, center alignment, and there is a clear difference in style with adjacent text lines, then it is determined to be a chapter title.
[0050] S112, inputting the extracted chapter titles into a large language model, and constructing a multi-level tree structure based on the relative granularity of the chapter titles by the large language model, wherein the content correlation between adjacent chapter titles is associated with the corresponding upper chapter title, forming a node containing chapter titles and their corresponding chapter content.
[0051] In the embodiment of the application, the large language model can be, for example, moonshot-v1-auto. The prompt word explicitly requires the task: "The following is a list of chapter titles of a document (each title occupies a line), please identify the hierarchical relationship based on the hierarchical logic of the title, and output the tree structure in JSON format. Requirements: the root node key name is 'root', which contains the first-level title; the first-level title is nested with the second-level title, and so on; the title level is determined by the serial number and granularity (for example, '1.1' is the sub-title of '1'). Example: [input example] → [output example]."
[0052] The large language model parses the hierarchical relationship between the titles according to the prompt word, and generates a JSON structure containing the hierarchical relationship of the titles; based on the JSON structure, the content between adjacent chapter titles is associated with the corresponding upper title node to form a node of "chapter title + corresponding chapter content", and finally a complete multi-level tree structure is constructed.
[0053] S12, the heterogeneous content of each node is parsed and processed to generate knowledge fragments.
[0054] S12 specifically includes:
[0055] S121, for the pure text content in the node, if the text length exceeds p word units, first based on paragraph splitting, if the paragraph length still exceeds p word units, then based on paragraph semantic splitting into multiple semantically independent knowledge fragments; if the text length does not exceed p word units, or the length of a single paragraph after paragraph splitting does not exceed p word units, then the text or paragraph is directly taken as a knowledge fragment.
[0056] In the embodiment of the application, p=512. Through S121, all pure text content in a node can be divided into one or split into multiple semantically independent knowledge fragments.
[0057] S122, for the table content in the node, the column name, row name and unmerged cell structure in the table content are extracted, the table content is reconstructed into a standardized Python dictionary format and converted into JSON data to form a knowledge fragment.
[0058] In the embodiment of the application, restoring the unmerged cell structure fills the content of the merged cell to maintain data integrity. The key of the Python dictionary format is a combination of "column name-row name" identifier, and the value is the corresponding cell data.
[0059] S123, for the image content in the node, the semantic information in the image content is extracted through a visual large model, and the extracted semantic information is taken as a knowledge fragment.
[0060] If there is content in the form of a picture in a node, the step is used for processing, specifically: the picture is extracted separately, a prompt word is designed, a visual large model (for example, qwen2.5-vl-72b) is used to extract semantic information of the picture, and the original picture is converted into semantic information of the picture. In an illustrative embodiment, an example of the prompt word is: "You are an image semantic extraction assistant, and you need to complete the following tasks: 1. Describe the image content in Chinese in full, including core elements, layout, and visual highlights; 2. Accurately identify and extract all text (including formulas, labels, and data) in the image; 3. Explain the semantic logic expressed by the image (for example, if it is a flowchart, explain the technical process step by step; if it is a data table, explain the data correlation; if it is a schematic diagram, explain the principle or structure); 4. If there is key information such as time, place, and event, mark it clearly."
[0061] S13, generating an embedding vector of each knowledge segment in a combined manner based on path embedding and content embedding.
[0062] In the embodiment of the application, the embedding vector of each knowledge segment is generated based on the title level of the multi-level tree structure. Further, S13 specifically includes:
[0063] S131, starting from the root title of the multi-level tree structure, performing hierarchical traversal sorting according to the parent-child membership relationship between nodes to obtain a node sequence arranged in hierarchical order, and ensuring that the upper parent node of the node can be completely traced back.
[0064] S132, traversing each node in the node sequence, calling an embedding model to respectively perform vectorization processing on the chapter titles of each node to generate corresponding chapter title embedding vectors.
[0065] In the embodiment of the application, the embedding model uses bge-large-zh-v1.5.
[0066] S133, for any target node, extracting chapter title embedding vectors of all upper parent nodes to which the node belongs, calculating the arithmetic mean of the chapter title embedding vectors of the upper parent nodes element by element according to the vector dimension, obtaining the path embedding vector of the target node, and the path embedding vector is used to represent the hierarchical association context of the target node in the multi-level tree structure.
[0067] S134, for each knowledge segment in the target node, using an embedding model to perform vectorization processing to generate a content embedding vector of the knowledge segment, which is used to represent the core semantic features of the knowledge segment.
[0068] In the embodiments of the present application, the target node refers to a specific node to be processed in a multi-level tree structure, and specifically refers to a node as an operation object in the steps of path embedding calculation and knowledge fragment association. Specifically, in the generation process of the knowledge fragment embedding vector (such as the path embedding calculation step), the target node is a specific node for which a path embedding vector needs to be generated: by extracting the chapter title embedding vectors of all superior parent nodes (from the direct parent node to the root node) to which the node belongs, the path embedding is calculated by averaging to represent the hierarchical association context of the node in the tree structure. Subsequently, the knowledge fragments contained in the node will be fused with the path embedding and the content embedding to generate the final embedding vector.
[0069] S135, the content embedding vector of each knowledge fragment and the path embedding vector of the target node are calculated by dimension by dimension and element by element to obtain an average vector as the final embedding vector of the knowledge fragment.
[0070] In the present application, vectorization processing refers to converting text information (including chapter titles and knowledge fragment contents) into fixed-dimensional numerical vectors through a preset embedding model, so that the semantic information of the text can be quantified through the distance in the vector space.
[0071] The vectorization processing of the chapter title includes the following steps:
[0072] Text preprocessing: standardizing the chapter title of the node, including removing redundant symbols such as special punctuation, format markers, and unifying English and Chinese punctuation such as converting “,” to “,”, correcting errors or non-standard expressions in the title such as “concrete” to “concrete”, and ensuring consistent title text format.
[0073] Model input encoding: input the preprocessed chapter title text into the embedding model, and the model encodes the text through the internal Transformer network structure, specifically including:
[0074] Tokenizing the title text: split the continuous text into a word sequence that can be recognized by the model, such as “1.2 concrete maintenance” into “1.2” “concrete” “maintenance”;
[0075] Convert the word into an initial vector through the word embedding layer, and then capture the semantic association between the words such as the collocation relationship between “concrete” and “maintenance” through the multi-layer attention mechanism;
[0076] Vector generation: The model output layer generates a fixed-dimensional vector (such as the 768-dimensional vector output by bge-large-zh-v1.5), which is the embedding vector of the chapter title and can represent the core semantics and hierarchical features of the title. For example, the vector of "1.2 Concrete Maintenance" is highly similar to the vector of "1.1 Concrete Proportioning" in space, reflecting the association of the same "Concrete Construction" theme.
[0077] The vectorization of knowledge fragments includes the following steps:
[0078] Text preprocessing: Special processing is performed on different types of knowledge fragments:
[0079] Plain text fragment: Remove redundant line breaks and spaces in paragraphs, and merge short sentences into coherent text, such as correcting "Maintenance temperature. ≥5℃" to "Maintenance temperature ≥5℃".
[0080] Table JSON data: Convert structured dictionaries to natural language description text, such as {"material": "cement", "label": "C30"} to "Material is cement, label is C30".
[0081] Image semantic description: Keep the complete description text generated by the visual large model, such as "Flowchart shows concrete pouring steps: 1. Template installation; 2. Steel bar binding; 3. Pouring and vibrating".
[0082] Model input encoding: Use the same embedding model as the chapter title vectorization to encode the preprocessed knowledge fragment text:
[0083] For fragments longer than the maximum input limit of the model (such as bge-large-zh-v1.5 supports 512 tokens), use the sliding window truncation method (window size 512 tokens, step size 256 tokens) to generate multiple sub-fragments, encode each sub-fragment and take the average as the final vector.
[0084] For short text fragments (≤512 tokens), directly input the model for encoding.
[0085] Vector generation: The model output is an embedding vector of the same dimension as the chapter title vector (such as 768 dimensions), which is the content embedding vector of the knowledge fragment and can accurately represent the core semantics of the fragment. For example, the vector of "Maintenance temperature ≥5℃" is close to the query vector of "Construction environment temperature requirement" in space.
[0086] Through the above steps, the fusion of path embedding and content embedding is realized: the path embedding retains the hierarchical association information of knowledge in the tree structure (such as chapter affiliation), and the content embedding captures the core semantics of the knowledge fragment; the combination of the two makes the vector representation of the knowledge fragment contain not only what the content is, but also where the content is located in the knowledge system, effectively improving the semantic matching accuracy of subsequent retrieval, on the one hand, reducing the matching error caused by engineering field term ambiguity through content embedding, and on the other hand, distinguishing the differences of similar contents in different knowledge levels through path embedding, while providing traceable hierarchical context for the retrieval results, enhancing the transparency and reliability of knowledge retrieval.
[0087] S14, store the association relationship of the knowledge fragment and the node, the hierarchical relationship between the nodes, the mapping relationship of the node and the source document, and the vector data, the related information of the knowledge fragment and the original information of the engineering knowledge document, and complete the construction of the engineering knowledge base.
[0088] In the embodiment of the application, the original information of the engineering knowledge document includes an original file of the engineering knowledge document and an image data file extracted from the engineering knowledge document.
[0089] Further, S14 specifically includes:
[0090] S141, store the association relationship of the knowledge fragment and the node, the hierarchical relationship between the nodes, and the mapping relationship of the node and the source document through a MySQL relational database.
[0091] In the embodiment of the application, the association relationship of the knowledge fragment and the node is used to establish the mapping of the knowledge fragment and the node to which it belongs, and to clearly define the context attribution of the knowledge fragment. The knowledge fragment and the node can be associated through a unique identification field. A unique ID such as fragment_id and node_id is allocated to the knowledge fragment and the node in the database respectively, and an external key node_id is set in the table corresponding to the knowledge fragment, and is associated with the primary key node_id of the node table. The association relationship is specifically used to record to which node each knowledge fragment such as a pure text fragment, a table converted JSON fragment, and an image semantic extraction fragment specifically belongs, so as to ensure that the source node of the knowledge fragment can be traced back, and to provide a basis for subsequent knowledge grouping and context supplement based on the node.
[0092] The hierarchical relationship between nodes is used to maintain the parent-child membership relationship of nodes in the multi-level tree structure, and support the hierarchical tracing and reconstruction of the tree structure. The parent node and the child node in the same hierarchical structure are associated by the parent node identification field. The parent_node_id field is set in the node table as a foreign key associated with the primary key node_id of the table: for the root node of the tree structure, parent_node_id is empty (or set to a preset value indicating no parent node); for non-root nodes, parent_node_id points to the node_id of the direct superior parent node. The hierarchical relationship between nodes is specifically used to record the direct parent node information of each node, and the complete hierarchical path of the node such as “root title→first chapter→second section→current node” can be traced through recursive query of parent_node_id, thereby supporting the generation of path embedding vectors and the complete restoration of the tree structure.
[0093] The mapping relationship between nodes and source documents refers to the mapping association between each node and its source original engineering document in the multi-level tree structure of the engineering knowledge document. Specifically, during the construction of the engineering knowledge base, a single engineering document (such as a standard specification or an inspection report) is parsed into multiple nodes (each node corresponds to a chapter title and its own content in the document), and the relationship between the node and the source document records the specific source of each node from which original document, including but not limited to the unique identification of the document (such as document ID), document name, storage path, and other associated information. The core role of this relationship is to establish the traceability association between the knowledge node and the original document, and in the subsequent retrieval process, the source document of the node can be located reversely to ensure the traceability of the knowledge, and to provide basic data support for the integrity verification of the knowledge, document version management, etc.
[0094] S142, store the embedding vectors of each knowledge fragment and the related information of the knowledge fragment in the ElasticSearch vector database.
[0095] The related information of each knowledge fragment refers to the metadata and content description information related to the knowledge fragment in ElasticSearch in addition to the embedding vector of the knowledge fragment. These information are the basic attributes of the knowledge fragment, used to assist the accuracy of vector retrieval and the understandability of the results, specifically including:
[0096] Basic identification information: unique ID of the knowledge fragment, ID of the node it belongs to (used to associate to the node in the tree structure);
[0097] Content description information: original content of the knowledge fragment (such as specific text of a pure text fragment, JSON data after table parsing, semantic description text after image parsing);
[0098] Type attribute information: the source type of the knowledge fragment (such as "pure text", "table", "image semantics"), used to distinguish different heterogeneous data fragments;
[0099] Correlation context information: briefly record the location of the knowledge fragment in the node or the correlation identifier of the adjacent fragment (assist in restoring the complete semantics when grouping knowledge later).
[0100] These information and embedding vectors are stored in ElasticSearch together, and when retrieving, not only the knowledge fragments are matched through vector similarity, but also the retrieval results can be optimized by combining content description, type attribute and other information, to ensure that the retrieved knowledge fragments not only meet the semantic similarity, but also meet the content correlation requirements.
[0101] S143, store the engineering knowledge document original file and the image data file extracted from the engineering knowledge document through the MINIO distributed object storage system.
[0102] In the embodiment of the application, the knowledge gap question is a sub-question required to answer the original question, specifically, a knowledge gap that needs to be filled before answering the main question, and a sub-question of the necessary knowledge base constructed for answering and solving the original question.
[0103] Further, in S2, the knowledge gap question list is generated by query rewriting, specifically comprising:
[0104] S211, arrange the current existing knowledge, the executed retrieval list and the user original question into prompt words.
[0105] Specifically, the current existing knowledge (accumulated relevant knowledge fragments and node content), the executed retrieval list (historical sub-query records) and the user original question can be arranged into structured prompt words according to a preset format, wherein each part of information is separated by clear identifiers such as "[existing knowledge]", "[executed retrieval list]", "[user original question]", to ensure that the large language model can clearly identify the input boundary.
[0106] S212, call the large language model to analyze the prompt words and generate an initial knowledge gap question list related to the user question.
[0107] In the embodiment of the present application, the targeted prompt word guide model focuses on knowledge gap analysis. The prompt word example is: "You are a professional query rewriting expert, and you need to generate subqueries based on the following information: 1. Analyze the gap between
existing knowledge
user original question
executed search list
Executed search list
Existing knowledge
User original question
[0108] In S213, the repeated sub-questions in the initial knowledge gap question list are filtered and removed by vectorizing the semantic similarity, and the final knowledge gap question list is obtained.
[0109] Specifically, the sub-questions in the initial knowledge gap question list are vectorized by the vectorization model to obtain sub-question vectors, and the cosine similarity of any two sub-question vectors is calculated; if the similarity of two sub-queries exceeds the repetition threshold, it is determined that the semantics are repeated, and one of them is retained; after the de-duplication processing, the final knowledge gap question list is obtained.
[0110] In the embodiment of the present application, the vectorization model uses the same model as the knowledge segment embedding, such as bge-large-zh-v1.5. The repetition threshold can be 0.8.
[0111] Through the above steps, the query rewriting of the user's original question is completed, and the generated knowledge gap question list accurately locates the gap between the current existing knowledge and the user's question. Subsequently, the knowledge base will be searched by traversing the list to supplement and improve the existing knowledge.
[0112] Further, in S2, the knowledge segment retrieval in the engineering knowledge base is retrieved by traversing the list, specifically including:
[0113] In S221, for each sub-question in the knowledge gap question list, the sub-question is converted into an embedding vector by an embedding model.
[0114] In S222, an ElasticSearch query statement is constructed based on the embedding vector, and vector similarity retrieval is performed in the ElasticSearch vector database to obtain knowledge segments related to the semantics of the sub-question.
[0115] In the query statement, the retrieval target is specified as the embedding vector field of the knowledge fragments in the knowledge base. The semantically related knowledge fragments can be matched by vector distance calculation (such as cosine similarity). After performing retrieval in the ElasticSearch vector database, the results are sorted in descending order of similarity scores, a similarity threshold (such as 0.6) is set to filter valid results, and the number of knowledge fragments returned by a single sub-question retrieval is limited (such as Top50) to avoid redundant data; finally, the knowledge fragments and their associated information (such as node ID, hierarchical path) that meet the conditions are obtained, providing a basis for subsequent knowledge grouping and sorting. The main purpose of knowledge grouping is that the retrieved direct knowledge fragments are relatively short and have single semantics, and the necessary context content may be missing, and these context contents may contain potential related knowledge. By grouping knowledge fragments, the knowledge fragments are grouped and restored to nodes, i.e. the original complete chapter content of engineering documents, which can restore complete semantic information.
[0116] Further, in S2, the knowledge grouping specifically includes:
[0117] S231, based on the identification information of the retrieved knowledge fragments, querying the node to which the knowledge fragments belong from the MySQL database.
[0118] Each retrieved knowledge fragment carries unique identification information (including knowledge fragment ID and node ID), based on which the knowledge fragment-node association table pre-stored in the MySQL database is queried, and the original node (i.e. the corresponding chapter node in the engineering document) to which the knowledge fragment belongs is accurately located through the node ID.
[0119] S232, reading the complete content of the node in the multi-level tree structure, and taking the complete content of the node as the context information of the knowledge fragment.
[0120] Specifically, the complete content of the node in the multi-level tree structure is read from the knowledge base, including: the chapter title and hierarchical path of the node (such as “1. Engineering design → 1.2 Structure parameters”), all knowledge fragments contained by the node (covering pure text fragments, table JSON data, image semantic description, and other heterogeneous contents), and the arrangement order of the knowledge fragments in the node (maintaining the same logical order as the original document chapter).
[0121] S233, grouping the retrieved knowledge fragments according to the node dimension.
[0122] Specifically, all knowledge segments belonging to the same node and the complete content of the node are grouped together to form a node-level knowledge unit. Through this grouping operation, originally isolated knowledge segments are restored to their chapter context environment, supplementing the semantic association, logical order and hierarchical background between segments, and avoiding semantic fragmentation or information loss caused by single short segments.
[0123] Further, in S2, the relevance ranking specifically includes:
[0124] S241, calculate the cosine similarity between the embedding vector of each retrieved knowledge segment and the embedding vector of the corresponding knowledge blank question, to obtain the semantic relevance.
[0125] For each sub-question in the knowledge blank question list, calculate the cosine similarity between its embedding vector and the embedding vector of all knowledge segments retrieved for the sub-question, to obtain the semantic relevance of each knowledge segment with respect to the current sub-question, with a value range of 0-1, and a higher value indicating closer semantic.
[0126] S242, calculate the mean of the semantic relevance of all knowledge segments, and solve the interpolation of the semantic relevance of each knowledge segment and the mean.
[0127] The interpolation of the semantic relevance of each knowledge segment and the mean is the difference between the semantic relevance of a single knowledge segment and the average relevance of all segments, which is used to measure the deviation of the relevance of the segment from the overall average level.
[0128] Specifically, for all knowledge segments corresponding to the current sub-question, the arithmetic mean of their semantic relevance is calculated; for each knowledge segment, the interpolation of its semantic relevance and the mean, i.e. the segment relevance, is calculated, if the segment relevance is higher than the mean, the interpolation is positive (strengthening its importance); if it is lower than the mean, the interpolation is negative (weakening its influence), to highlight the contribution of high-relevance segments.
[0129] S243, accumulate the sum of the interpolation of the knowledge segments contained in the node as the node weight, sort the nodes by weight from high to low, and select the contents of the top-ranked nodes as the knowledge and context information of the corresponding knowledge blank question, and add them to the existing knowledge for cumulative storage.
[0130] Specifically, the interpolation of all knowledge segments contained in the node is accumulated, i.e., the sum of the segment relevance of all segments under the same node, to obtain the weight of the node. The higher the weight, the closer the semantic association between the node as a whole and the sub-problem. The nodes are sorted in descending order of node weight, and the top nodes (such as Top5) are selected. The complete content of these nodes, including chapter titles, hierarchical paths, and all knowledge segments, i.e., the "node-level knowledge unit" formed after knowledge grouping, is extracted as the core knowledge and context corresponding to the current sub-problem. These node contents are added to the existing knowledge for cumulative storage, while the node ID is used for deduplication to avoid duplicate storage of the same node.
[0131] Further, in S2, the cycle termination judgment specifically includes:
[0132] S251, judge whether the executed search round exceeds the preset threshold, if it exceeds, terminate the cycle; if it does not exceed, construct a prompt word based on the executed search list, the current existing knowledge and the user's original question, call the large language model to judge whether the current existing knowledge is sufficient to answer the user's question, if it is sufficient to answer, terminate the cycle, otherwise return to the query rewriting link in step S2 to continue the cycle.
[0133] Specifically, the system has a search round counter, the initial value is 0, and it is automatically incremented by 1 after completing the "query rewriting -> search -> grouping -> sorting -> knowledge accumulation" process. First, judge whether the current counter value exceeds the preset threshold, if it exceeds the threshold, regardless of whether the current existing knowledge is sufficient, it is forced to terminate the cycle to avoid low search efficiency caused by infinite generation of knowledge blank problems. If it does not exceed the round threshold, it enters the knowledge sufficiency judgment link:
[0134] Prompt word construction: generate a prompt word according to the structure of "[executed search list] + [current existing knowledge] + [user's original question]", where "current existing knowledge" needs to be arranged into a structured abstract of node-level knowledge unit (including core conclusions, key data and hierarchical association), and "executed search list" needs to list all historical sub-query texts to ensure that the large language model has a comprehensive understanding of the search progress and knowledge accumulation state.
[0135] Large language model judgment: call a large language model (such as moonshot-v1-auto), the prompt word example is: "Based on the historical queries in [executed search list] and [current existing knowledge], judge whether enough information has been obtained to answer [user's original question]. Requirements: 1. If the knowledge covers the core elements of the user's question (such as technical principles, parameter ranges, implementation steps, etc.) and there is no missing key information, return'sufficient'; 2. If there are core elements not covered or key logical gaps, return 'insufficient'; 3. Only return'sufficient' or 'insufficient', do not add other explanations."
[0136] Loop decision: if the model returns "sufficient", terminate the loop and enter the subsequent reply generation stage; if it returns "insufficient", add 1 to the round counter and return to the query rewriting link in S2 to generate a new knowledge blank question list for continuous retrieval.
[0137] In the embodiment of the application, the preset threshold is set to 3-5 rounds, which can be dynamically adjusted according to the complexity of engineering knowledge.
[0138] Further, after the loop is terminated, the existing knowledge accumulated in the loop is used to generate a reply content for the user question, specifically including:
[0139] S252, based on the executed retrieval list, the accumulated existing knowledge at the time of loop termination, and the user's original question, fill in the prompt word template, and call a large language model to generate an original answer based on the existing knowledge.
[0140] In the embodiment of the application, the prompt word template needs to be clear: input elements: executed retrieval list (presented in the form of a list of all sub-queries), accumulated existing knowledge at the time of loop termination (organized into structured content, including node level path, core knowledge fragment and heterogeneous data summary, such as "1.2 structure parameter table data: concrete strength grade C30"), user's original question; output constraints: call a large language model (such as moonshot-v1-auto), require it to generate an answer only based on "existing knowledge", prohibit the introduction of external information; if it is determined that the existing knowledge is irrelevant to the user's question, it needs to strictly return the preset phrase "Sorry, the answer you want is not found in the knowledge base"; if it is relevant, it needs to use the structure of "conclusion + knowledge base basis" (such as "According to the content of section 1.2, the concrete curing temperature should be ≥5℃, and the specific data can be found in table row 3"), to ensure that the correspondence between the answer and the knowledge fragment can be traced back.
[0141] S253, the original answer generated by the large model is processed by sentence, and each sentence is vectorized by embedding the model.
[0142] Among them, the sentence processing adopts the combination of punctuation symbol splitting and semantic verification, and the original answer is split into independent sentences according to punctuation such as period, question mark, semicolon, etc., and then the semantic integrity of the sentence is verified by a large language model (to avoid ambiguity caused by splitting, such as splitting "if the temperature is <5℃, take temperature protection measures" into two valid sentences). The vectorization conversion uses the same model as the knowledge fragment embedding (such as bge-large-zh-v1.5) to vectorize each valid sentence, generating a sentence embedding vector, ensuring that the vector space is unified with the knowledge fragments in the knowledge base.
[0143] S254, calculate the cosine similarity of each sub-sentence vector and the knowledge fragment vector in the engineering knowledge base, if the similarity exceeds the preset threshold, it is determined that the content of the sub-sentence comes from the corresponding knowledge fragment, and the relevance with the knowledge base is verified.
[0144] Specifically, the cosine similarity of each sub-sentence vector and all knowledge fragment vectors in the engineering knowledge base is calculated, and the top 3 knowledge fragments and their scores are recorded. If the highest similarity score exceeds the preset threshold, it is determined that the content of the sub-sentence comes from the corresponding knowledge fragment, and if there are multiple fragments in the top 3 that exceed the threshold, the one with the highest score is taken as the main source. If all scores are below the threshold, it is marked as "not verified" and determined as "content without basis". The preset threshold can be 0.6.
[0145] S255, integrate all sub-sentences that pass the relevance verification to generate the final reply content; if there are sub-sentences that do not pass the verification, modify them based on existing knowledge and then integrate and output.
[0146] Specifically, for the verified sub-sentence, the original expression is retained, integrated in logical order, and the source identification (such as "[derived from node 1.2.3]") is marked after the sub-sentence to enhance traceability. For the sub-sentence that does not pass the verification, call the large language model to modify, and prompt the word to clearly require "restate the core information of the sub-sentence based on existing knowledge, and delete if you cannot modify"; after modification, S23-S24 verification needs to be performed again to ensure that the modified content is associated with the knowledge base. Integrate all sub-sentences that pass the verification (including the modified content) to form a reply content with complete structure and logical coherence, and attach a traceability list (list all cited knowledge fragment IDs and corresponding node paths) for user to check the knowledge source.
[0147] The embodiment of the application also provides an electronic device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are arranged to execute the method described in the embodiment of the application.
[0148] The embodiment of the application also provides a computer readable storage medium storing computer executable instructions, and the computer executable instructions are used to execute the method described in the embodiment of the application.
[0149] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, the steps described in the present application can be executed in parallel, sequentially or in different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, which is not limited herein.
[0150] The above detailed description does not limit the scope of the application. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed embodiment within the scope of the application. Any modification, equivalent replacement and improvement made without departing from the spirit and principle of the application shall fall within the scope of the application.
Claims
1. A method for constructing and deeply retrieving an engineering knowledge base, characterized in that, The method includes the following steps: S1, Building an engineering knowledge base based on engineering knowledge documents; S2, upon receiving a user's question, generates a list of blank knowledge questions by querying and rewriting, traverses the list to search the engineering knowledge base, and performs knowledge grouping, relevance sorting, and loop termination judgment. After the loop terminates, it generates a response to the user's question based on the existing knowledge accumulated during the loop. Specifically, S1 includes: S11, identify the knowledge level of each input engineering knowledge document and form a multi-level tree structure, wherein the chapter titles and their corresponding chapter contents in the engineering knowledge document constitute the nodes of the multi-level tree structure. S12, the heterogeneous content of each node is parsed and processed to generate knowledge fragments; S13, Based on a combination of path embedding and content embedding, the embedding vectors of each knowledge fragment are generated; S14, store the association relationship between knowledge fragments and nodes, the hierarchical relationship between nodes, the mapping relationship between nodes and source documents, as well as the vector data of knowledge fragments, related information and original information of engineering knowledge documents, to complete the construction of the engineering knowledge base.
2. The method according to claim 1, characterized in that, S11 specifically includes: Chapter titles are extracted from engineering knowledge documents using heuristic rules, which include: determining titles based on built-in document tag attributes, determining titles based on the title's starting character tag, and determining titles based on text style features. The extracted chapter titles are input into the large language model, which constructs a multi-level tree structure based on the relative granularity of the chapter titles. The content between adjacent chapter titles is associated with the corresponding parent chapter title, forming nodes that contain chapter titles and their corresponding chapter content.
3. The method according to claim 1, characterized in that, S12 specifically includes: For plain text content in a node, if the text length exceeds p words, it is first split into paragraphs. If the paragraph length still exceeds p words, it is then split into multiple semantically independent knowledge fragments based on paragraph semantics. If the text length does not exceed p words, or the length of a single paragraph after paragraph splitting does not exceed p words, then the text or paragraph is directly treated as a knowledge fragment. For the table content in the node, extract the column names and row names from the table content and restore the unmerged cell structure. Reconstruct the table content into a standardized Python dictionary format and then convert it into JSON data to form knowledge fragments. For the image content in the node, semantic information is extracted from the image content using a large visual model, and the extracted semantic information is used as knowledge fragments.
4. The method according to claim 1, characterized in that, S13 specifically includes: Starting from the root title of the multi-level tree structure, the nodes are traversed and sorted hierarchically according to the parent-child relationship between the nodes to obtain a sequence of nodes arranged in hierarchical order; Traverse each node in the node sequence, call the embedding model to vectorize the chapter title of each node, and generate the corresponding chapter title embedding vector; For any target node, extract the chapter title embedding vectors of all parent nodes to which the node belongs, and calculate the arithmetic mean of the chapter title embedding vectors of all parent nodes element by element according to the vector dimension to obtain the path embedding vector of the target node. For each knowledge fragment in the target node, the embedding model is used to vectorize it, generating the content embedding vector of that knowledge fragment; The arithmetic mean of the content embedding vector of each knowledge fragment and the path embedding vector of its target node is calculated element-wise along each dimension, and the resulting mean vector is used as the final embedding vector of the knowledge fragment.
5. The method according to claim 1, characterized in that, S14 specifically includes: The relationship between knowledge fragments and nodes, the hierarchical relationship between nodes, and the mapping relationship between nodes and source documents are stored in a MySQL relational database. The embedding vectors of each knowledge fragment and related information of the knowledge fragment are stored in the ElasticSearch vector database. The MINIO distributed object storage system stores the original files of the engineering knowledge document and the image data files extracted from the engineering knowledge document.
6. The method according to claim 1, characterized in that, In S2, the process of generating a knowledge gap question list through query rewriting specifically includes: Organize existing knowledge, the list of executed searches, and the user's original question, and combine them into suggestion words; The large language model is invoked to parse the prompt words and generate an initial list of knowledge gap questions related to the user's question. The knowledge gap questions are sub-questions required to answer the original question. By calculating semantic similarity using vectorization, duplicate sub-problems in the initial knowledge gap problem list are filtered and removed to obtain the final knowledge gap problem list.
7. The method according to claim 6, characterized in that, In S2, the traversal of the list to retrieve knowledge fragments in the engineering knowledge base specifically includes: For each sub-problem in the knowledge gap problem list, the sub-problem is converted into an embedding vector using an embedding model; Based on the embedded vector, an ElasticSearch query statement is constructed, and vector similarity retrieval is performed in the ElasticSearch vector database to obtain knowledge fragments that are semantically related to the sub-question.
8. The method according to claim 7, characterized in that, In S2, the knowledge grouping specifically includes: Based on the identification information of the retrieved knowledge fragments, query the node to which the knowledge fragment belongs from the MySQL database; Read the complete content of the node in the multi-level tree structure and use the complete content of the node as the context information of the knowledge fragment; The retrieved knowledge fragments are grouped according to the node dimension.
9. The method according to claim 8, characterized in that, In S2, the correlation ranking specifically includes: The semantic relevance is obtained by calculating the cosine similarity between the embedding vector of each knowledge fragment retrieved and the embedding vector of the corresponding knowledge gap question. Calculate the mean of the semantic relevance of all knowledge fragments, and solve the interpolation between the semantic relevance of each knowledge fragment and the mean. The interpolated sum of knowledge fragments contained in a node is used as the node weight. The nodes are sorted from high to low weight, and a preset number of the top-ranked nodes are selected as the knowledge and context information for the corresponding knowledge gap problem. These are then added to the existing knowledge for cumulative storage.
10. The method according to claim 9, characterized in that, In S2, the loop termination determination specifically includes: Determine whether the number of retrieval rounds executed has exceeded a preset threshold. If it has, terminate the loop. If it has not, construct prompt words based on the executed retrieval list, existing knowledge, and the user's original question. Call the large language model to determine whether the existing knowledge is sufficient to answer the user's question. If it is sufficient, terminate the loop. Otherwise, return to the query rewriting step in step S2 to continue the loop.
Citation Information
Patent Citations
Large language model RAG optimization method based on tree neighbor context
CN119293195A
Knowledge retrieval method, system and equipment based on document structure context enhancement and medium
CN119807359A
Intelligent question answering system method for air traffic control communication business knowledge
CN120407745A
Hierarchical semantic-driven retrieval enhancement generation method and system
CN120632119A
Intelligent verbal skill generation method and device based on multi-modal knowledge
CN120688629A
Cited By
Retrieval enhancement generation method suitable for building structure field
CN121301589A
A search enhancement generation method suitable for the field of building structures
CN121301589B
Knowledge base construction and retrieval method and system based on multi-source text in building field
CN121681810A