Intensive operation and maintenance question answering method based on large language model and dynamic knowledge graph
Patent Information
- Application Number
- CN202610919999.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-29
AI Technical Summary
[0009]本发明的目的在于提供一种基于大语言模型与动态知识图谱的集约化运维问答方法,旨在解决集约化运维场景下运维知识构建慢、检索不准、大模型易产生幻觉以及缺乏个性化上下文的技术难题
1. 运维知识库的构建:摒弃了传统人工维护QA对的落后模式,通过文档向量化与图谱自动抽取,直接盘活了企业沉睡的数万份历史文档和工单,新系统上线时的知识接入时间从传统的按“周”计算缩短为按“分钟”计算。
Smart Images

Figure CN122840232A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and in particular to an intensive operation and maintenance question answering method based on a large language model and dynamic knowledge graph. Background Technology
[0002] As enterprise informatization continues to advance, various business systems (such as finance, audit, OA, and procurement) are gradually merging into a massive, centralized platform. The daily operation and maintenance of such a large system (such as troubleshooting, configuration modification, permission granting, and execution of inspection scripts) faces significant knowledge barriers.
[0003] Traditional operations and maintenance (O&M) knowledge is often fragmented and scattered across Word documents and PDF manuals from different teams. When a system malfunctions or a new O&M staff member takes over, they often have to spend a lot of time flipping through documents or asking around, resulting in a persistently high time to repair (MTTR). Meanwhile, frontline customer service (L1 / L2 Support) staff have to handle a large number of repetitive O&M inquiries every day, leading to extremely low work efficiency.
[0004] In recent years, while Large Language Models (LLMs) have demonstrated powerful general question-answering capabilities, their direct application in enterprise internal operations and maintenance (O&M) scenarios often results in severe "illusions" (providing incorrect restart commands or configuration parameters) due to a lack of knowledge about the enterprise's proprietary architecture. Furthermore, they fail to offer personalized, secure, and compliant O&M solutions tailored to the specific role of the questioner (e.g., R&D, DBA, general staff). Therefore, quickly revitalizing dormant O&M assets and building a "personal O&M knowledge assistant" that understands the enterprise's proprietary architecture and can provide personalized solutions is a critical pain point that centralized O&M teams urgently need to address.
[0005] Existing operation and maintenance knowledge bases and question-and-answer systems mainly suffer from the following technical deficiencies: 1. Fragmented knowledge bases, high construction costs, and low retrieval accuracy: Traditional knowledge bases primarily rely on manual QA (question-answer pairs) input, resulting in extremely high maintenance costs and easy expiration with system iterations. Retrieval primarily depends on keyword-matching inverted indexing technologies like ElasticSearch. If the terminology used by the user's question does not perfectly match the keywords in the document (e.g., a user asks "OA login failure," while the document states "Unified authentication gateway error"), a valid answer cannot be retrieved.
[0006] 2. General-purpose large models lack enterprise-specific context and are prone to dangerous "illusions": When a general-purpose LLM is directly introduced for question answering, since the model has not seen the enterprise's internal microservice topology and private script library during the pre-training stage, it is very easy to generate seemingly reasonable but completely wrong troubleshooting commands based on probability (such as generating incorrect database clearing SQL). If the operation and maintenance personnel blindly execute them, it will lead to disastrous consequences.
[0007] 3. The question-and-answer system lacks personalized context awareness: existing question-and-answer robots treat all users the same. In a centralized operations and maintenance scenario, the same question, "How to handle the failure of the accounting system's expense report?", should result in different responses depending on the user profile. For a regular employee, the assistant should provide "instructions on how to resubmit the form"; while for a backend developer or DBA, the assistant should provide "command line scripts for viewing specific microservice logs and troubleshooting deadlocks." Existing systems cannot achieve role awareness and dynamic answer routing based on user profiles. Summary of the Invention
[0008] The present invention aims to at least partially solve one of the technical problems in the related art.
[0009] The purpose of this invention is to provide an intensive operation and maintenance question-answering method based on a large language model and dynamic knowledge graph, aiming to solve the technical problems of slow operation and maintenance knowledge construction, inaccurate retrieval, large models being prone to illusions, and lack of personalized context in intensive operation and maintenance scenarios.
[0010] To achieve the above objectives, a first aspect of this invention proposes an intensive operation and maintenance question-answering method based on a large language model and dynamic knowledge graph, comprising: S1. Automatically extract operation and maintenance knowledge from multi-source heterogeneous data sources, transform unstructured documents into structured knowledge fragments through a multimodal document parsing model, and convert the structured knowledge fragments into semantic vectors and store them in a vector database. At the same time, extract entities and relationships from the text to construct an operation and maintenance knowledge graph. S2, obtain the user's profile tags, input the profile tags, the current context dialogue and the user's original question into the large language model for intent rewriting, and generate the reconstructed query statement; S3, using the reconstructed query statement, perform semantic similarity retrieval in the vector database and entity relationship reasoning retrieval in the operation and maintenance knowledge graph, fuse and deduplicate the knowledge fragments recalled by the two retrieval paths and reorder them based on relevance, and generate an answer with source citation based on the ordered knowledge fragments; S4, output the answer and capture the user's interaction with the answer, and inject the user's corrected experience data back into the personal knowledge base and the global knowledge base to achieve continuous updating of the knowledge base.
[0011] Optionally, the step of automatically extracting operational knowledge from multi-source heterogeneous data sources and converting unstructured documents into structured knowledge fragments through a multimodal document parsing model includes: By integrating with internal enterprise data sources such as Confluence, Jira historical work orders, GitLab operation and maintenance script libraries, and PDF operation manuals, a multimodal document parsing model based on a vision-language pre-trained architecture is used to automatically identify titles, paragraphs, code blocks, and tables in documents. A region proposal network is used to segment document pages into title regions, body paragraph regions, code block regions, table regions, and image regions. For titles and body paragraphs, a combination of OCR text recognition and semantic role labeling is used to extract text content and its hierarchical relationships. For code blocks, code language automatic detection and structured extraction based on grammatical features are used. For tables, row and column relationship reasoning and cell merging restoration techniques are used to maintain the integrity of the table's structured information, transforming unstructured document content into structured knowledge fragments with semantic tags.
[0012] Optionally, the automatic identification of titles, paragraphs, code blocks, and tables in a document using a multimodal document parsing model based on a vision-language pre-trained architecture includes: using the LayoutLMv3 architecture as the multimodal document parsing model, which simultaneously integrates input information from three modalities: text semantic features, two-dimensional layout features, and document image visual features, and performs joint encoding through a multimodal Transformer encoder to analyze the layout of the document page.
[0013] Optionally, the step of converting the structured knowledge fragments into semantic vectors and storing them in a vector database includes: using a dual-tower Sentence-BERT architecture based on contrastive learning training as the embedding model. This architecture contains two BERT encoders with shared weights, which encode the query text and the knowledge document text respectively, mapping the text to a 768-dimensional dense semantic vector space, so that semantically similar texts are closer in the vector space.
[0014] Optionally, the semantic similarity retrieval in the vector database includes: converting the reconstructed query statement into a query vector using an Embedding model; calculating the cosine similarity between the query vector and the knowledge fragment vector using an approximate nearest neighbor algorithm in the vector database; and recalling the top-K most semantically relevant operation manual fragments and historical solution work orders.
[0015] Optionally, the approximate nearest neighbor algorithm uses the HNSW index, and the K value of the Top-K is 10.
[0016] Optionally, the entity relationship reasoning retrieval in the operation and maintenance knowledge graph includes: extracting key entities from the reconstructed query statement using named entity recognition technology, performing multi-hop traversal reasoning in the knowledge graph with the key entities as starting nodes, and automatically recalling upstream and downstream dependency information, historical fault cases and corresponding solutions related to the system in question.
[0017] Optionally, generating answers with source citations based on sorted knowledge fragments includes: the large language model synthesizing answers strictly according to the citation generation pattern based on recalled and verified private knowledge slices, and automatically attaching links to reference sources at the end of the answer.
[0018] Optionally, the output of the answer includes: for operational questions, directly outputting an executable Shell script or SQL code block.
[0019] Optionally, capturing user interaction behavior with the answer and injecting the user's corrected experience data back into the personal knowledge base and the global knowledge base includes: capturing user likes or dislikes for the answer, or user interaction behavior of modifying or correcting the script generated by the assistant on the interface; packaging the original question, the user's corrected script, and the explanation of the correction into a high-quality QA pair; and injecting it into the user's personal knowledge base and the global operation and maintenance knowledge base after review.
[0020] The embodiments of the present invention have the following beneficial effects: 1. Construction of the Operation and Maintenance Knowledge Base: Abandoning the outdated model of traditional manual maintenance of QA pairs, by using document vectorization and automatic graph extraction, tens of thousands of dormant historical documents and work orders of the enterprise were directly revitalized. The knowledge access time when the new system goes live has been shortened from the traditional calculation in "weeks" to the calculation in "minutes".
[0021] 2. Personalized assistance significantly reduces time to repair (MTTR): Based on user profile-based intent reconstruction, the system can accurately determine the user's true technical needs and push troubleshooting suggestions that match their permissions and skill level. For L1 / L2 level front-line operations support, the problem resolution rate (FCR) has increased by more than 60%, significantly reducing the need to escalate work orders to backend core R&D.
[0022] 3. Secure and reliable enterprise-grade AIOps application that eliminates illusions: The rigorous RAG architecture combined with knowledge graph tracing ensures that every troubleshooting step and every line of script answered by the assistant comes from the enterprise's officially verified documentation library, giving LLM extremely high security and availability in demanding production environments. Attached Figure Description
[0023] The above-described and additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which: Figure 1 A simplified flowchart of the intensive operation and maintenance question-answering method based on a large language model and dynamic knowledge graph provided in this embodiment of the invention; Figure 2 A detailed flowchart of the intensive operation and maintenance question-answering method based on a large language model and dynamic knowledge graph provided in the embodiments of the present invention. Detailed Implementation
[0024] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] The following description, with reference to the accompanying drawings, describes an intensive operation and maintenance question-answering method based on a large language model and dynamic knowledge graph, according to an embodiment of the present invention.
[0027] Example 1 like Figure 1 and Figure 2 As shown, the intensive operation and maintenance question answering method based on large language models and dynamic knowledge graphs includes the following steps: S1. Automatically extract operation and maintenance knowledge from multi-source heterogeneous data sources, transform unstructured documents into structured knowledge fragments through a multimodal document parsing model, and convert the structured knowledge fragments into semantic vectors and store them in a vector database. At the same time, extract entities and relationships from the text to construct an operation and maintenance knowledge graph.
[0028] In the intensive intelligent Q&A method for operations and maintenance, the first step is to automatically extract operations and maintenance knowledge and build a knowledge base. This step aims to automatically identify and extract operations and maintenance-related knowledge information from various heterogeneous data sources within the enterprise, transforming it from unstructured raw document form into a structured knowledge representation that can be used for subsequent retrieval and reasoning.
[0029] To achieve this objective, embodiments of the present invention access multiple data sources, including but not limited to enterprise internal document management systems, work order systems, code repositories, and various operation manuals. For unstructured documents from these sources, which vary in format and have mixed content, embodiments of the present invention employ a multimodal document parsing model. This model can simultaneously utilize the document's textual semantic features, layout features, and visual image features to comprehensively segment the document page into semantic blocks, thereby accurately identifying different structural units such as heading levels, body paragraphs, code blocks, tables, and images. For different types of blocks, the model employs corresponding recognition strategies. For example, semantic role labeling is performed on text content to extract hierarchical relationships; syntactic feature detection is performed on code blocks to maintain their structural integrity; and row and column relationship reasoning is performed on tables to reconstruct their structured information. Through the above multimodal fusion and partitioning recognition strategies, embodiments of the present invention transform unstructured document content into structured knowledge fragments with semantic tags.
[0030] Subsequently, this embodiment of the invention transforms these structured knowledge fragments into high-dimensional semantic vectors through an embedding model and stores them in a vector database to support subsequent semantic similarity-based retrieval. Simultaneously, this embodiment of the invention also utilizes a large language model to automatically extract entities and semantic relationships between entities from the text, constructing a global operations and maintenance knowledge graph. This graph reflects structured knowledge in the operations and maintenance domain, including system architecture, component dependencies, and fault correlations. The vector database and the knowledge graph coexist, jointly forming a hybrid knowledge base supporting subsequent intelligent question answering.
[0031] As one implementation, the multimodal document parsing model can adopt a vision-language pre-training architecture, such as the LayoutLMv3 architecture. This architecture integrates input information from three modalities: text semantic features, two-dimensional layout features, and document image visual features, and performs joint encoding through a multimodal Transformer encoder. The embedding model can adopt a dual-tower Sentence-BERT architecture based on contrastive learning training. This architecture contains two BERT encoders with shared weights, which encode the query text and the knowledge document text respectively, mapping the text to a 768-dimensional dense semantic vector space.
[0032] This step, through automated multimodal document parsing and knowledge extraction, enables the rapid and efficient construction of a structured knowledge base from massive heterogeneous enterprise operation and maintenance documents. This significantly reduces the time and manpower costs required for traditional manual knowledge base construction, providing a rich, accurate, and searchable knowledge foundation for subsequent intelligent question answering.
[0033] S2, obtain the user's profile tags, input the profile tags, the current context dialogue, and the user's original question into the large language model for intent rewriting, and generate the reconstructed query statement.
[0034] When a user initiates a query, this embodiment of the invention first obtains a profile tag associated with the user. The profile tag is used to characterize the user's identity attributes, technical background and scope of authority. Furthermore, it obtains the context information of the current conversation and the original query statement entered by the user.
[0035] Subsequently, in this embodiment of the invention, the aforementioned user profile tags, the current context dialogue, and the original query statement are input into a large language model. This model then performs semantic understanding and intent rewriting on the original query statement based on the acquired multi-dimensional information to generate a reconstructed query statement. This reconstruction process aims to eliminate potential ambiguities in the original query, supplement missing contextual information, and precisely and personally adjust the query intent according to the user role and permission scope indicated by the user profile tags. This ensures that the reconstructed query statement more accurately reflects the user's actual information needs in a specific scenario.
[0036] As one implementation method, embodiments of the present invention can extract the user's department, technical rank, responsible microservice module, and historical question records as profile tags, and input these tags, along with the current dialogue context and the original question, into a large language model. The model will automatically complete the missing specific system names, environmental information, or technical details in the original question, and generate a reconstructed query statement containing more complete semantics.
[0037] By introducing user profile tags and contextual information for intent rewriting, this step can significantly improve the semantic accuracy and personalization of query statements, enabling subsequent retrieval and answer generation processes to be based on more accurate query input. This effectively avoids retrieval bias and answer mismatch caused by vague original query statements or lack of context, thereby improving the overall response quality and user satisfaction of the question-and-answer system.
[0038] S3. Using the reconstructed query statement, perform semantic similarity retrieval in the vector database and entity relationship reasoning retrieval in the operation and maintenance knowledge graph. Merge, deduplicate, and reorder the knowledge fragments recalled by the two retrieval paths. Generate an answer with source citation based on the sorted knowledge fragments.
[0039] When generating answers using the reconstructed query statements, this embodiment of the invention employs a hybrid retrieval mechanism that executes two independent retrieval paths in parallel to retrieve relevant information from the constructed knowledge base.
[0040] The first pathway is vector retrieval based on semantic similarity. It maps the reconstructed query statement to the same semantic vector space as the knowledge fragments and calculates the similarity between the query vector and the fragment vectors in the knowledge base, thereby recalling the most semantically relevant knowledge fragments. The second pathway is reasoning retrieval based on knowledge graphs. It identifies key entities from the reconstructed query statement and uses these entities as starting nodes to perform multi-hop traversal reasoning along the relationships between entities in the constructed operational knowledge graph, thereby recalling upstream and downstream dependency information, historical cases, or solutions that are logically related to the query content.
[0041] The knowledge fragments retrieved by each of the two pathways are then sent to a fusion processing module. This module deduplicates the fragments from different pathways, eliminates redundant information, and reorders them according to their relevance to the query statement in order to select the most valuable knowledge fragments.
[0042] Ultimately, the large language model synthesizes answers based on these sorted and verified knowledge fragments using a citation-based approach, and automatically includes source information for the cited knowledge fragments in the answers, thereby ensuring the traceability and authenticity of the generated content.
[0043] As one implementation method, the vector retrieval path can use the approximate nearest neighbor algorithm to calculate cosine similarity to recall the Top-K fragments, while the knowledge graph reasoning path extracts entities through named entity recognition and performs multi-hop traversal. After the fragments recalled by the two paths are fused, deduplicated, and reordered, the large language model generates the answer strictly according to the citation pattern and attaches the reference source link.
[0044] This step significantly improves the comprehensiveness and accuracy of knowledge retrieval by integrating semantic similarity retrieval and knowledge graph reasoning retrieval, avoiding omissions or biases that may be caused by a single retrieval method. At the same time, by integrating, deduplicating, and re-ranking the recalled fragments and combining them with citation generation mode, it effectively suppresses the risk of generating false or erroneous information by large language models, ensuring the reliability and traceability of the generated answers.
[0045] S4, output the answer and capture the user's interaction with the answer, and inject the user's corrected experience data back into the personal knowledge base and the global knowledge base to achieve continuous updating of the knowledge base.
[0046] After generating and outputting the answer, this embodiment of the invention further performs continuous updates and personalized evolution of the knowledge base. Specifically, this embodiment not only presents the user with the final answer including text explanations and executable scripts, but also actively captures the user's interactive feedback behavior on the answer. This interactive feedback behavior includes, but is not limited to, explicit user evaluations of the answer (such as liking or disliking), direct modifications to the answer content, corrections or replacements of the generated script, and subsequent follow-up questions or supplementary information from the user. This embodiment of the invention treats these interactive behaviors as high-value experience feedback data and automatically identifies the correction logic and optimization experience contained within them based on preset rules or models.
[0047] Subsequently, this embodiment of the invention structures and packages the user's original question, the user's corrected answer or script, and the explanation of the correction reason, forming a high-quality experience data record. This experience data record is injected into two levels of knowledge bases: the first is a personal knowledge base, which is associated with a specific user or user profile and is used to store the user's personalized experience accumulated in historical interactions. This allows this embodiment of the invention to prioritize the user's personal historical correction experience when generating answers for that user in the future, thereby providing more targeted answers. The second is a global knowledge base, which is shared by all users and is used to precipitate verified correction experiences with universal value into public knowledge, so that other users can also benefit from the optimization experience when encountering similar problems.
[0048] Through the above mechanism, the embodiments of the present invention realize the continuous dynamic updating and self-evolution of the knowledge base. That is, as the frequency of user use increases, the experience data in the knowledge base is continuously enriched and optimized, thereby enabling the embodiments of the present invention to "become smarter with use" and gradually improve the accuracy and personalization of the answers.
[0049] As one implementation method, when a user modifies the operation and maintenance script generated by the embodiment of the present invention, the modification behavior is automatically captured, and the original question, the modified script, and the explanation of the modification reason are packaged into a high-quality question and answer pair, which are injected into the user's personal knowledge base and the global operation and maintenance knowledge base, respectively, so as to realize the personalized accumulation and global sharing of knowledge.
[0050] This technical step achieves continuous dynamic evolution of the knowledge base by constructing a closed-loop mechanism for user interaction feedback and knowledge base updates, significantly improving the adaptability of the embodiments of the present invention to users' personalized needs. At the same time, by injecting high-value correction experience into the global knowledge base, it promotes knowledge sharing and reuse, thereby continuously optimizing the accuracy and practicality of the answers in long-term use.
[0051] Example 2 Based on the above embodiments, this embodiment provides a detailed description of the specific implementation of step S1 in the intensive operation and maintenance intelligent question-answering method: "Automatically extract operation and maintenance knowledge from multi-source heterogeneous data sources, transform unstructured documents into structured knowledge fragments through a multimodal document parsing model, and convert the structured knowledge fragments into semantic vectors and store them in a vector database, while extracting entities and relationships from the text to construct an operation and maintenance knowledge graph."
[0052] In this embodiment, the specific implementation method of automatically extracting operation and maintenance knowledge from multi-source heterogeneous data sources and converting unstructured documents into structured knowledge fragments through a multimodal document parsing model in step S1 is as follows.
[0053] First, this embodiment of the invention accesses internal enterprise data sources such as Confluence, Jira historical work orders, GitLab operation and maintenance script library, and PDF operation manuals to obtain raw unstructured document data. Then, it uses a multimodal document parsing model based on a vision-language pre-trained architecture to automatically identify titles, paragraphs, code blocks, and tables in the documents.
[0054] Specifically, the multimodal document parsing model adopts the LayoutLMv3 architecture, which integrates input information from three modalities: text semantic features, two-dimensional layout features, and document image visual features. These are then jointly encoded by a multimodal Transformer encoder to perform layout analysis on the document page.
[0055] In the document structure recognition stage, the model first uses a region proposal network to segment the document page into different semantic blocks such as title regions, body paragraph regions, code block regions, table regions, and image regions. For titles and body paragraphs, this embodiment of the invention uses a combination of OCR text recognition and semantic role annotation to extract text content and its hierarchical relationships. That is, the text is first recognized by OCR technology, and then the hierarchical affiliation of each text fragment in the document structure is determined by semantic role annotation, such as recognizing chapter titles, section titles, and body paragraphs. For code blocks, this embodiment of the invention uses automatic code language detection and structured extraction based on grammatical features. That is, by analyzing the grammatical features of the code block (such as keywords, indentation, comment symbols, etc.), the programming language (such as Shell, SQL, Python, etc.) is automatically determined, and the complete code fragment and its grammatical structure are extracted. For tables, this embodiment of the invention uses row and column relationship reasoning and cell merging restoration technology. By analyzing the visual layout and text content of the table, the row and column structure of the table is inferred, and the original relationship of merged cells is restored, thereby maintaining the integrity of the table's structured information.
[0056] Through the aforementioned multimodal fusion and partitioning recognition strategies, this embodiment of the invention transforms unstructured document content into structured knowledge fragments with semantic tags, outputting structured data that includes heading levels, paragraph affiliation, code block boundaries, and table structure, providing high-quality input for subsequent vectorized storage and knowledge graph construction.
[0057] This specific implementation introduces a multimodal document parsing model based on the LayoutLMv3 architecture and combines it with a region proposal network for semantic block segmentation. This enables automated and high-precision parsing of multi-source heterogeneous operation and maintenance documents, significantly reducing the cost of manually building a knowledge base and improving the completeness and accuracy of knowledge extraction. It also lays a solid data foundation for subsequent semantic retrieval and graph reasoning.
[0058] In this embodiment, the specific implementation method of converting structured knowledge fragments into semantic vectors and storing them in the vector database in step S1 is as follows.
[0059] This invention first uses the structured knowledge fragments obtained through multimodal document parsing model processing as the text input to be encoded. Then, this invention employs a dual-tower Sentence-BERT architecture based on contrastive learning training as the embedding model. This architecture includes two BERT encoders with shared weights, used to encode the query text and the knowledge document text respectively.
[0060] Specifically, for each knowledge fragment, this embodiment of the invention inputs it into one of the BERT encoders. This encoder first performs contextual semantic modeling on the input text using a multi-layer Transformer structure, extracting a deep semantic representation for each token. Then, it aggregates the representation of the entire sequence into a fixed-length semantic vector using a pooling strategy (such as mean pooling or CLS token pooling). This architecture, through a shared weight design, ensures that the query text and the knowledge document text are encoded in the same semantic space, thus making semantically similar texts closer together in the vector space. After encoding, each knowledge fragment is mapped to a 768-dimensional dense semantic vector space, forming a corresponding semantic vector. This embodiment of the invention stores these semantic vectors, along with their corresponding original knowledge fragment text and metadata (such as source document, title level, timestamp, etc.), into a vector database. In the vector database, this embodiment of the invention constructs an index structure for these semantic vectors, for example, using an HNSW (Hierarchical Navigable Small World) index to support subsequent efficient approximate nearest neighbor retrieval. Through the above processing, the embodiments of the present invention transform unstructured operation and maintenance knowledge into a vectorized representation that can be used for semantic retrieval, providing a high-quality semantic retrieval foundation for subsequent hybrid retrieval steps.
[0061] This specific implementation adopts a dual-tower Sentence-BERT architecture based on contrastive learning training, which maps knowledge fragments to a dense semantic vector space of 768 dimensions, significantly improving the accuracy of semantic similarity retrieval. This enables the embodiments of the present invention to accurately recall knowledge fragments that are semantically similar to the user's query, thereby effectively overcoming the technical defect of traditional keyword matching-based retrieval methods that cannot recall effective answers when terms are inconsistent.
[0062] Example 3 Based on the above embodiments, this embodiment provides a detailed description of the specific implementation of step S3 in the intensive operation and maintenance intelligent question answering method: "Using the reconstructed query statement, perform semantic similarity retrieval in the vector database, and perform entity relationship reasoning retrieval in the operation and maintenance knowledge graph. Merge, deduplicate, and reorder the knowledge fragments recalled by the two retrieval paths, and generate an answer with source citation based on the ordered knowledge fragments."
[0063] In this embodiment, the specific implementation process of semantic similarity retrieval in the vector database in step S3 is as follows.
[0064] First, this embodiment of the invention obtains the reconstructed query statement generated in step S2 as input. This query statement has integrated user profile tags, contextual dialogue, and original question information. Then, this embodiment of the invention inputs the reconstructed query statement into the Embedding model. This model adopts a dual-tower Sentence-BERT architecture based on contrastive learning training, containing two BERT encoders with shared weights. It can map the input text to a 768-dimensional dense semantic vector space, thereby generating a corresponding query vector. This query vector represents the semantic features of the reconstructed query statement. Next, this embodiment of the invention uses this query vector as the retrieval basis to perform semantic similarity retrieval in a pre-constructed vector database. This vector database stores structured knowledge fragments parsed and quantized from multi-source heterogeneous data sources using a multimodal document parsing model in step S1, including operation manual fragments and historical solution work orders. During the retrieval process, this embodiment of the invention employs an approximate nearest neighbor algorithm, specifically, this algorithm is implemented based on the HNSW (Hierarchical Navigable SmallWorld) index structure. HNSW indexes, by constructing a multi-layered graph structure, can efficiently locate neighboring nodes with semantic similarity to the query vector in a high-dimensional vector space, thereby significantly improving retrieval speed.
[0065] This invention utilizes the HNSW index to calculate the cosine similarity between the query vector and each knowledge fragment vector in the vector database. A higher cosine similarity value indicates a closer semantic relationship between the two vectors. This invention sorts the vectors according to their cosine similarity scores from highest to lowest and recalls the top-K most semantically relevant knowledge fragments, where K is set to 10. Therefore, this invention ultimately outputs 10 operation manual fragments and historical solution work orders that best match the semantics of the reconstructed query statement, serving as the recall results for this retrieval path, for subsequent fusion processing with the results of the knowledge graph reasoning retrieval path.
[0066] Through the aforementioned approximate nearest neighbor retrieval and cosine similarity calculation based on the HNSW index, this embodiment can achieve millisecond-level semantic retrieval response in a large-scale vector database, significantly improving the recall efficiency and accuracy of knowledge fragments, and providing a reliable data foundation for the subsequent generation of high-quality, low-latency operation and maintenance question-and-answer answers.
[0067] In this embodiment, the specific implementation method of entity relationship reasoning and retrieval in the operation and maintenance knowledge graph described in step S3 is as follows.
[0068] First, this embodiment of the invention receives the reconstructed query statement generated in step S2 as input. This query statement incorporates user profile tags, contextual dialogue, and the original question, such as "How to safely clean up the Redis cluster cache in the production environment of this embodiment of the invention". This embodiment of the invention uses named entity recognition technology to parse the reconstructed query statement and extract key entities, such as "this embodiment of the invention" and "Redis". This named entity recognition technology can employ sequence labeling methods based on pre-trained language models, such as the BERT-BiLSTM-CRF architecture, to accurately identify entity names belonging to the operations and maintenance domain in the query text. Subsequently, this embodiment of the invention uses the extracted key entities as starting nodes and performs multi-hop traversal reasoning in the operations and maintenance knowledge graph constructed in step S1. This knowledge graph stores various entities in the operations and maintenance architecture (such as systems, microservices, databases, middleware, and hardware devices) and their dependencies, calls, deployments, and other relationships. For example, when the key entity is "Business and Finance System", this embodiment of the invention performs a multi-hop traversal along the relationship chain in the knowledge graph of "Business and Finance System (Invention Embodiment) → Dependency → Payment Gateway Microservice → Call → Bank Interface", automatically recalling upstream and downstream dependency information related to the involved embodiment of the invention, such as the status of the payment gateway microservice and the call logs of the bank interface. At the same time, this embodiment of the invention will also traverse the historical failure case nodes associated with this entity in the knowledge graph, such as "Payment Gateway Memory Leak Failure Case", and recall the corresponding solution nodes, such as "Steps to Investigate JVM Heap Memory Usage and Export Heap Dump Files".
[0069] Through the aforementioned multi-hop traversal reasoning, this embodiment of the invention can output a set of structured knowledge fragments. These fragments not only contain information directly related to the query entity but also cover its upstream and downstream dependencies and historical fault handling experience, thus providing rich contextual support for subsequent answer generation. These knowledge fragments recalled from the knowledge graph will be fused with the Top-K operation manual fragments and historical solution work orders recalled by the vector semantic retrieval path for deduplication and relevance reordering. Finally, the large language model generates an answer with source citations based on the ordered knowledge fragments.
[0070] This specific implementation introduces a multi-hop traversal reasoning mechanism based on knowledge graphs, enabling the embodiments of the present invention to extract deep-level operational knowledge from the relationships between entities, such as inter-system dependency links and historical failure cases, thereby significantly improving the comprehensiveness and accuracy of the answers and effectively avoiding incorrect answers due to a lack of context.
[0071] In this embodiment, step S3 generates an answer with source citations based on the sorted knowledge fragments, which is specifically implemented in the following way.
[0072] First, this embodiment of the invention uses knowledge fragments obtained after fusion deduplication and relevance reordering as input. These knowledge fragments are private knowledge slices that have undergone authenticity verification. Their sources include semantically relevant operation manual fragments and historical solution work orders recalled from vector databases, as well as upstream and downstream dependency information, historical fault cases, and corresponding solutions recalled from the operation and maintenance knowledge graph through multi-hop traversal reasoning. After receiving these private knowledge slices, the large language model synthesizes the answer strictly according to the citation generation pattern. Specifically, when processing the user's reconstructed query statement, the model treats each recalled knowledge slice as an independent evidence unit. When generating each sentence or key argument in the answer, the model points to the specific knowledge slice number or source identifier it depends on, ensuring that each part of the answer is supported by corresponding officially verified enterprise documents or graph data. For example, when the model generates an answer to "Investigating memory leaks in the payment gateway of the business and finance system," for the step of "checking the JVM heap memory usage of the payment gateway," the model will associate it with the corresponding operation manual fragment retrieved from the vector database; for the system relationship description of "the payment gateway depends on the bank interface," the model will associate it with the relationship links inferred from the knowledge graph. After completing the synthesis of the answer text, this embodiment of the invention automatically appends links to reference sources at the end of the answer. These links point to the specific location of the original document in Confluence, GitLab, or the PDF operation manual, or to the detailed description page of the corresponding entity in the knowledge graph.
[0073] Through the above processing, the answers output by the embodiments of the present invention not only include accurate operation and maintenance guidance, but also provide traceable and verifiable reference information, thereby completely eliminating the illusion problem caused by the lack of enterprise private context in large models.
[0074] This specific implementation significantly improves the reliability and traceability of operation and maintenance Q&A results by combining the citation generation mode with the source link, ensuring that users can obtain safe and accurate troubleshooting suggestions based on verified private knowledge slices in harsh production environments, effectively reducing operational risks caused by erroneous information.
[0075] Example 4 Based on the above embodiments, this embodiment provides a detailed description of the specific implementation of step S4 in the centralized operation and maintenance intelligent question-answering method: "output the answer, capture the user's interaction behavior with the answer, and inject the user's corrected experience data back into the personal exclusive knowledge base and the global knowledge base to achieve continuous updating of the knowledge base."
[0076] In this embodiment, the specific implementation of outputting the answer in step S4 includes: when the question raised by the user is identified by the embodiment of the present invention as an operation-related question, such as a scenario involving system fault diagnosis, configuration modification, script execution, etc. that requires specific command line operations, the embodiment of the present invention not only generates an answer containing textual descriptions, but also automatically generates a Shell script or SQL code block that can be directly executed in the operation and maintenance environment based on the retrieved knowledge fragments and user profiles.
[0077] Specifically, in generating answers based on sorted knowledge fragments, this embodiment of the invention first performs semantic analysis on the recalled knowledge fragments using a large language model to determine whether the question belongs to the operation category. If it is determined to be an operation category question, the large language model further extracts structured information related to the operation steps from the knowledge fragments, such as command parameters, execution order, and dependency conditions, and generates well-formatted and clearly commented script code based on this information.
[0078] For example, when a user asks, "How to troubleshoot memory leaks in the payment gateway of the business and finance system?", this embodiment of the invention, during the answer generation process, retrieves the JVM configuration information of the payment gateway microservice, troubleshooting steps from historical failure cases, and relevant Shell command templates from the knowledge graph. Then, the large language model integrates this information into an executable Shell script fragment, including the command "jmap -heap" to view JVM heap memory usage. <pid>The command to export a heap dump file is "jmap -dump:format=b,file=heap_dump.hprof". <pid>"And the command to check the frequency of Full GC in the GC log is "grep Full GC / var / log / app / gc.log | tail -20".
[0079] For example, when a DBA asks, "How do I find transactions that haven't been committed for a long time in the business and finance system?", this embodiment of the invention retrieves database monitoring-related knowledge fragments from the knowledge base and generates corresponding SQL query statements, such as "SELECT trx_id, trx_state, trx_started, trx_mysql_thread_id FROM information_schema.innodb_trx WHERE TIMESTAMPDIFF(MINUTE, trx_started, NOW())>10 ORDERBY trx_started ASC". The generated script or code block is embedded in the answer text as an independent code block, along with security prompts and parameter descriptions before execution, ensuring that users can directly copy and execute it safely in the appropriate environment. In this way, this embodiment of the invention directly transforms the retrieved knowledge into operable instructions, significantly improving the efficiency of operations and maintenance personnel in handling operational issues.
[0080] The beneficial effects of this specific implementation method are as follows: by automatically generating executable Shell scripts or SQL code blocks, the retrieved operation and maintenance knowledge is directly transformed into operable instructions, avoiding the tedious process of users manually consulting documents and writing commands, greatly shortening the response time for troubleshooting and handling, reducing the risk caused by errors in manually entering commands, and improving the accuracy and security of operation and maintenance operations.
[0081] In this embodiment, the process of capturing the user's interaction behavior with the answer in step S4 and injecting the user's corrected experience data back into the personal knowledge base and the global knowledge base is specifically implemented in the following way.
[0082] When the embodiments of the present invention output an answer containing a text explanation or an executable script, the user can evaluate or modify the answer.
[0083] Specifically, this embodiment of the invention provides "like" and "dislike" evaluation buttons on the answer display interface, allowing users to express their approval or disapproval of the answer's quality by clicking the corresponding button. Simultaneously, for answers containing executable scripts, this embodiment provides a script editing area on the interface, where users can directly modify and correct the script content generated by the assistant. This embodiment captures these interactions in real time, including user "like" or "dislike" click records and user modifications to the script within the script editing area. When this embodiment detects that a user has modified the script, it automatically associates and packages the original question, the user-corrected script, and the user-inputted explanation of the correction into a high-quality QA pair. The data structure of this QA pair includes three fields: the original question text, the corrected script content, and the correction reason text.
[0084] Subsequently, this embodiment of the invention submits the QA pair to the review process. After its correctness and validity are confirmed by a preset review mechanism (e.g., manual review by a senior operations engineer or system administrator), this embodiment of the invention injects the QA pair into two target knowledge bases. The first target knowledge base is the user's personal knowledge base, which stores personalized knowledge fragments associated with the user's profile tags (such as department, technical level, and microservice module they are responsible for). This allows the user to prioritize retrieving high-quality experience data that has been corrected by the user when encountering similar problems in the future, thereby providing more accurate personalized answers. The second target knowledge base is a global operations and maintenance knowledge base, which stores general operations and maintenance knowledge shared by all users, allowing other users to benefit from the corrected optimization experience when making relevant queries. Through the above mechanism, this embodiment of the invention achieves continuous updating and self-optimization of the knowledge base. That is, as user interaction accumulates, the content in the knowledge base is continuously enriched and corrected, thereby improving the accuracy and practicality of subsequent Q&A.
[0085] This specific implementation captures user likes or dislikes for answers and script modification behavior, and injects the revised experience data into the personal knowledge base and the global knowledge base after review. This achieves continuous updates to the knowledge base in both personalized and global ways, significantly improving the accuracy of the system's question and answer for individual users and all users, and enabling the system to have the self-evolutionary ability to "become smarter with use".
[0086] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0087] In the description of this specification, the references to "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0088] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.< / pid> < / pid>
Claims
1. An intensive operation and maintenance question-answering method based on a large language model and dynamic knowledge graph, characterized in that, Includes the following steps: S1. Automatically extract operation and maintenance knowledge from multi-source heterogeneous data sources, transform unstructured documents into structured knowledge fragments through a multimodal document parsing model, and convert the structured knowledge fragments into semantic vectors and store them in a vector database. At the same time, extract entities and relationships from the text to construct an operation and maintenance knowledge graph. S2, obtain the user's profile tags, input the profile tags, the current context dialogue and the user's original question into the large language model for intent rewriting, and generate the reconstructed query statement; S3, using the reconstructed query statement, perform semantic similarity retrieval in the vector database and entity relationship reasoning retrieval in the operation and maintenance knowledge graph, fuse and deduplicate the knowledge fragments recalled by the two retrieval paths and reorder them based on relevance, and generate an answer with source citation based on the ordered knowledge fragments; S4, output the answer and capture the user's interaction with the answer, and inject the user's corrected experience data back into the personal knowledge base and the global knowledge base to achieve continuous updating of the knowledge base.
2. The method as described in claim 1, characterized in that, The automatic extraction of operational knowledge from multi-source heterogeneous data sources, and the transformation of unstructured documents into structured knowledge fragments through a multimodal document parsing model, includes: By integrating with internal enterprise data sources such as Confluence, Jira historical work orders, GitLab operation and maintenance script libraries, and PDF operation manuals, a multimodal document parsing model based on a vision-language pre-trained architecture is used to automatically identify titles, paragraphs, code blocks, and tables in documents. A region proposal network is used to segment document pages into title regions, body paragraph regions, code block regions, table regions, and image regions. For titles and body paragraphs, a combination of OCR text recognition and semantic role labeling is used to extract text content and its hierarchical relationships. For code blocks, code language automatic detection and structured extraction based on grammatical features are used. For tables, row and column relationship reasoning and cell merging restoration techniques are used to maintain the integrity of the table's structured information, transforming unstructured document content into structured knowledge fragments with semantic tags.
3. The method as described in claim 2, characterized in that, The method of automatically identifying titles, paragraphs, code blocks, and tables in a document using a multimodal document parsing model based on a vision-language pre-trained architecture includes: using the LayoutLMv3 architecture as the multimodal document parsing model, which simultaneously integrates input information from three modalities: text semantic features, two-dimensional layout features, and document image visual features. This information is then jointly encoded using a multimodal Transformer encoder to perform layout analysis on the document page.
4. The method as described in claim 1, characterized in that, The step of converting the structured knowledge fragments into semantic vectors and storing them in a vector database includes: using a dual-tower Sentence-BERT architecture based on contrastive learning training as the embedding model. This architecture contains two BERT encoders with shared weights, which encode the query text and the knowledge document text respectively, mapping the text to a 768-dimensional dense semantic vector space, so that semantically similar texts are closer in the vector space.
5. The method as described in claim 1, characterized in that, The semantic similarity retrieval in the vector database includes: converting the reconstructed query statement into a query vector using an Embedding model; calculating the cosine similarity between the query vector and the knowledge fragment vector using an approximate nearest neighbor algorithm in the vector database; and recalling the top-K most semantically relevant operation manual fragments and historical work orders.
6. The method as described in claim 5, characterized in that, The approximate nearest neighbor algorithm uses the HNSW index, and the K value of the Top-K is 10.
7. The method as described in claim 1, characterized in that, The entity relationship reasoning retrieval in the operation and maintenance knowledge graph includes: using named entity recognition technology to extract key entities from the reconstructed query statement, using the key entities as starting nodes to perform multi-hop traversal reasoning in the knowledge graph, and automatically recalling upstream and downstream dependency information, historical fault cases and corresponding solutions related to the system in question.
8. The method as described in claim 1, characterized in that, The process of generating answers with source citations based on sorted knowledge fragments includes: the large language model synthesizes answers strictly according to the citation generation pattern based on recalled and verified private knowledge slices, and automatically attaches links to reference sources at the end of the answer.
9. The method as described in claim 1, characterized in that, The output of the answer includes: for operation-related questions, directly outputting an executable Shell script or SQL code block.
10. The method as described in claim 1, characterized in that, The process of capturing user interaction with the answers and injecting the user's corrected experience data back into the user's personal knowledge base and the global knowledge base includes: capturing user likes or dislikes for the answers, or user interaction with the script generated by the correction assistant on the interface; packaging the original question, the user's corrected script, and the explanation of the correction into a high-quality QA pair; and injecting it into the user's personal knowledge base and the global operation and maintenance knowledge base after review.