Operation and maintenance method and device based on multi-source dynamic retrieval mechanism, equipment and medium

Through the multi-source dynamic retrieval mechanism and causal reasoning chain mechanism, the problems of low efficiency and poor accuracy of traditional operation and maintenance methods in multi-source, multi-dimensional, and multi-cause problems are solved, and the causes are quickly positioned and logical solutions are generated, which improves operation and maintenance efficiency and accuracy.

CN120429154APending Publication Date: 2025-08-05GUANGDONG ESHORE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510584064.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Traditional operation and maintenance methods have low efficiency and poor accuracy when dealing with multi-source, multi-dimensional and multi-caus problems, and cannot quickly locate the root cause of the problem. They lack effective information integration and experience reuse, making it difficult to meet the needs of enterprises for integrated intelligent operation and maintenance.

Method used

A multi-source dynamic search mechanism is adopted to collect multiple data sources for preprocessing, a multi-source vector database is generated, a preset big model is used to judge the problem type, and a causal reasoning chain mechanism is used to disassemble the problem into sub-problems, perform multi-source dynamic search, and generate diagnostic results and solutions.

Benefits of technology

It improves the accuracy and efficiency of problem diagnosis, avoids misdiagnosis or overdiagnosis, provides in-depth intelligent support, and meets the needs of enterprises for integrated intelligent operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429154A_ABST
    Figure CN120429154A_ABST
Patent Text Reader

Abstract

The invention relates to an operation and maintenance method and device based on a multi-source dynamic retrieval mechanism, equipment and a medium. The method comprises the following steps: collecting a plurality of data sources related to operation and maintenance, and preprocessing the plurality of data sources to obtain a multi-source vector database; obtaining a question input by a user, and judging a question type of the question based on a preset large model; if the problem type of the problem is a diagnosis type problem, disassembling the problem into a plurality of sub-problems with a causal relationship by adopting a causal reasoning chain mechanism; and performing multi-source dynamic retrieval on the multi-source vector database based on the sub-questions, calling answers corresponding to the sub-questions, and generating a diagnosis result and a solution of the question. According to the scheme provided by the invention, related information can be intelligently extracted from different data sources according to the type and intention of the problem, a more accurate retrieval result is ensured, a causal reasoning mechanism chain is adopted for reasoning and decomposing complex fault problems, potential reasons are gradually analyzed according to the causal relationship, and a logical solution is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an operation and maintenance method, apparatus, device, and medium based on a multi-source dynamic retrieval mechanism. Background Art

[0002] With the rapid development of information technology, enterprise IT (Information Technology) system architectures are becoming increasingly complex, encompassing numerous hardware devices, software applications, and network infrastructure. Traditional operations and maintenance methods are gradually exposing their limitations when dealing with multi-source, multi-dimensional, and multi-cause problems.

[0003] Currently, most operations and maintenance personnel still rely on scheduled scripts, log polling, and manual analysis to perform daily inspections, troubleshooting, and experience-based analysis. However, these methods present numerous challenges. For one thing, operations and maintenance information is scattered across multiple sources, such as alarm systems, logging platforms, knowledge bases, and work order systems, lacking effective unified integration and efficient access mechanisms. This forces operations and maintenance personnel to switch back and forth between multiple systems when faced with problems, consuming significant time and effort to collect and organize relevant information, severely impacting operational efficiency. Furthermore, traditional operations and maintenance methods exhibit significant shortcomings in terms of system response speed, problem location accuracy, and experience reuse. Slow system response prevents timely detection of potential problems; ambiguous problem location makes it difficult to quickly and accurately identify the root cause; and low experience reuse prevents the effective use of past experience to solve similar problems. These issues severely impact fault response efficiency and business continuity. Although some intelligent operations and maintenance platforms have introduced basic anomaly detection and alerting capabilities in recent years, these platforms still struggle to meet enterprises' needs for integrated intelligent operations and maintenance in complex scenarios. They cannot effectively achieve the goal of "quickly locating the cause + proposing a solution path" and cannot provide comprehensive, accurate and timely decision-making support for operation and maintenance personnel.

[0004] Therefore, there is an urgent need for a new operation and maintenance technology that can break through the limitations of traditional operation and maintenance methods, effectively integrate multi-source operation and maintenance information, quickly and accurately locate problems and provide solutions, so as to meet the growing IT system operation and maintenance needs of enterprises and ensure the efficient and stable operation of enterprise business. Summary of the Invention In order to solve or partially solve the problems existing in the related technologies, the present application provides an operation and maintenance method, device, equipment and medium based on a multi-source dynamic retrieval mechanism, which can decompose the problem, gradually analyze the cause and generate corresponding solutions, effectively improve the accuracy of problem diagnosis, and avoid misdiagnosis or overdiagnosis.

[0005] In a first aspect, the present application provides an operation and maintenance method based on a multi-source dynamic retrieval mechanism, comprising: Collecting multiple data sources related to operation and maintenance, and preprocessing the multiple data sources to obtain a multi-source vector database; Obtain the question input by the user and determine the type of question based on the preset large model; If the problem type is a diagnostic problem, the causal reasoning chain mechanism is used to decompose the problem into multiple sub-problems with causal relationships; Based on the sub-problems, a multi-source dynamic search is performed on the multi-source vector database, the answers corresponding to the sub-problems are retrieved, and the diagnosis results and solutions of the problems are generated.

[0006] Preferably, the use of a causal reasoning chain mechanism to decompose the problem into multiple causally related sub-problems includes: The preset large model calls the inference chain template; The problem is decomposed into multiple sub-problems with causal relationships based on the reasoning chain template.

[0007] Preferably, performing a multi-source dynamic search on the multi-source vector database based on the sub-question, retrieving the answer corresponding to the sub-question, and generating a diagnosis result and solution for the problem includes: Dynamically determine the data source to be searched in the multi-source vector database based on the problem; Retrieve answers corresponding to the plurality of sub-questions from the data source to be retrieved; Integrate the answers corresponding to the multiple sub-questions and input the integrated results into the preset large model; Generate diagnostic results and solutions for the problem based on the preset inspection report template.

[0008] Preferably, dynamically determining the data source to be searched in the multi-source vector database based on the question includes: Identify data sources relevant to the stated problem and prioritize them for retrieval; Based on the problem and the current operating status, the data source to be retrieved is dynamically determined.

[0009] Preferably, the step after determining the type of the question based on the preset large model further includes: If the question type is a general knowledge question, dynamically determining a data source to be searched in the multi-source vector database based on the question; Retrieving the answer corresponding to the question from the data source to be retrieved, and inputting the answer into the preset large model; Generate diagnostic results and solutions for the problem based on the preset inspection report template.

[0010] Preferably, the preprocessing of the plurality of data sources includes: Cleaning, classifying, and labeling the data source; Based on the type of the data source, different methods are used to segment the data source to obtain data blocks; vectorizing the data source; The data blocks and the vectorized data source are stored.

[0011] A second aspect of the present application provides an operation and maintenance device based on a multi-source dynamic retrieval mechanism, comprising: A collection module, used to collect multiple data sources related to operation and maintenance, and pre-process the multiple data sources to obtain a multi-source vector database; A judgment module is used to obtain a question input by a user and determine the type of the question based on a preset large model; A decomposition module is used to decompose the problem into multiple sub-problems with causal relationships using a causal reasoning chain mechanism if the problem type is a diagnostic problem; The retrieval module is used to perform a multi-source dynamic search on the multi-source vector database based on the sub-problem, retrieve the answer corresponding to the sub-problem, and generate a diagnosis result and solution for the problem.

[0012] Preferably, the calling module includes: a determination submodule, configured to dynamically determine a data source to be retrieved in the multi-source vector database based on the question; A retrieval submodule, configured to retrieve answers corresponding to the plurality of sub-questions from the data source to be retrieved; An integration submodule, configured to integrate the answers corresponding to the plurality of sub-questions and input the integrated results into the preset large model; The generation submodule is used to generate the diagnosis results and solutions of the problem based on the preset inspection report template.

[0013] A third aspect of the present application provides an electronic device, including: processor; and The memory stores executable codes thereon, and when the executable codes are executed by the processor, the processor is caused to execute the method described above.

[0014] A fourth aspect of the present application provides a computer-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method described above.

[0015] The technical solution provided by this application may include the following beneficial results: The present application provides an operation and maintenance method based on a multi-source dynamic retrieval mechanism, comprising: collecting multiple data sources related to operation and maintenance, preprocessing the multiple data sources to obtain a multi-source vector database; obtaining a question input by a user and determining the problem type based on a preset large model; if the problem type is a diagnostic problem, using a causal reasoning chain mechanism to decompose the problem into multiple sub-problems with causal relationships; performing a multi-source dynamic search on the multi-source vector database based on the sub-problems, retrieving the answers corresponding to the sub-problems, and generating a diagnostic result and solution for the problem. The above method, through the multi-source dynamic retrieval mechanism, can intelligently extract relevant information from different data sources (such as operation and maintenance logs, warning logs, real-time performance data, etc.) based on the type and intent of the problem; secondly, when the problem is a diagnostic problem, a causal reasoning chain mechanism is used for reasoning. During the problem diagnosis process, the causal reasoning chain mechanism can not only decompose complex fault problems, but also gradually analyze potential causes based on causal relationships and generate logical solutions. It can provide deep intelligent support in multiple links such as fault diagnosis, solution recommendation, and inspection report generation, avoiding misdiagnosis or overdiagnosis, thereby greatly improving operation and maintenance efficiency and accuracy.

[0016] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above and other objects, features and advantages of the present application will become more apparent by describing in more detail exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.

[0018] Figure 1 This is a flow chart of an operation and maintenance method based on a multi-source dynamic retrieval mechanism according to an embodiment of the present application; Figure 2 This is another flowchart of an operation and maintenance method based on a multi-source dynamic retrieval mechanism according to an embodiment of the present application; Figure 3 This is a flowchart of an operation and maintenance method based on a multi-source dynamic retrieval mechanism according to an embodiment of the present application; Figure 4 Schematic diagram of the structure of an operation and maintenance device based on a multi-source dynamic retrieval mechanism according to an embodiment of the present application; Figure 5 It is a structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although the accompanying drawings illustrate embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0020] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0021] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0022] The following are explanations of the relevant terms involved in the embodiments of this application: Intelligent Operations: Leverage artificial intelligence (AI) technologies, particularly machine learning, deep learning, and data mining, to automate, optimize, and improve IT operations management.

[0023] Large Language Model (LLM): A deep learning-based natural language processing (NLP) model designed to understand and generate natural language.

[0024] Retrieval-augmented Generation (RAG): a natural language processing technology that combines information retrieval and text generation.

[0025] Multi-source dynamic retrieval: Dynamically select multiple heterogeneous knowledge sources for information retrieval based on the user's question type, and provide a mechanism for high-quality contextual support for large model generation through unified strategy scheduling, result fusion and summary processing.

[0026] Causal reasoning: This mechanism automatically breaks down user questions into a series of causally related sub-questions and performs chain logical reasoning based on the search results to gradually analyze and locate the cause of the problem and generate explanatory answers.

[0027] Currently, most operations and maintenance personnel still rely on scheduled scripts, log polling, and manual analysis to perform daily inspections, troubleshooting, and experience-based analysis. However, these methods present numerous challenges. For one thing, operations and maintenance information is scattered across multiple sources, such as alarm systems, logging platforms, knowledge bases, and work order systems, lacking effective unified integration and efficient access mechanisms. This forces operations and maintenance personnel to switch back and forth between multiple systems when faced with problems, consuming significant time and effort to collect and organize relevant information, severely impacting operational efficiency. Furthermore, traditional operations and maintenance methods exhibit significant shortcomings in terms of system response speed, problem location accuracy, and experience reuse. Slow system response prevents timely detection of potential problems; ambiguous problem location makes it difficult to quickly and accurately identify the root cause; and low experience reuse prevents the effective use of past experience to solve similar problems. These issues severely impact fault response efficiency and business continuity. Although some intelligent operations and maintenance platforms have introduced basic anomaly detection and alerting capabilities in recent years, these platforms still struggle to meet enterprises' needs for integrated intelligent operations and maintenance in complex scenarios. They cannot effectively achieve the goal of "quickly locating the cause + proposing a solution path" and cannot provide comprehensive, accurate and timely decision-making support for operation and maintenance personnel.

[0028] In response to the above problems, the embodiments of the present application provide an operation and maintenance method, device, electronic device and storage medium based on a multi-source dynamic retrieval mechanism, which can decompose the problem, gradually analyze the cause and generate corresponding solutions, effectively improve the accuracy of problem diagnosis, and avoid misdiagnosis or overdiagnosis.

[0029] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0030] Figure 1 It is a flowchart of an operation and maintenance method based on a multi-source dynamic retrieval mechanism shown in an embodiment of the present application.

[0031] See also Figure 1 , the method comprising: Step 110 : Collect multiple data sources related to operation and maintenance, pre-process the multiple data sources, and obtain a multi-source vector database.

[0032] In the embodiment of the present application, multiple data sources related to operation and maintenance can be collected, such as operation and maintenance logs, warning logs, document data, real-time performance indicator monitoring data, fault databases, etc. The collected data sources are preprocessed to obtain a multi-source vector database related to operation and maintenance.

[0033] Operation and maintenance logs are management objects used to store system event and operation information. They primarily record detailed data such as various events, status changes, and error messages during system operation. Operation and maintenance logs are crucial for understanding program execution. By analyzing logs, we can locate system failures, optimize system performance, and meet compliance requirements. Warning logs indicate potential issues or unrecommended operations that may not immediately impact system operation, but are noteworthy and may require future action. Warning logs are typically used to alert administrators or developers of potential risks or situations requiring attention so that they can take timely preventative or remedial measures. Document data refers to data in the form of documents, typically including text, images, tables, and other content. It is persistent and can be read by humans or machines. Performance monitoring involves real-time monitoring of the performance and health of systems, applications, or networks by collecting, analyzing, and reporting key performance indicators. Performance monitoring can help identify potential performance issues, bottlenecks, and optimize performance.

[0034] Step 120: Obtain the question input by the user and determine the question type based on the preset large model.

[0035] Users can enter their questions on the system's interactive interface. For example, if a user asks, "My online service is responding slowly, what's going on?" A pre-set large model can be used to determine the type of question, whether it's a diagnostic or general knowledge question. The pre-set large model can be a large language model.

[0036] In step 130 , if the problem type is a diagnostic problem, a causal reasoning chain mechanism is used to decompose the problem into multiple sub-problems with causal relationships.

[0037] If the problem type is determined to be diagnostic, after receiving the user's question, a pre-set large model is invoked and a causal chain mechanism is used to break the problem down into a set of causally related sub-questions, thereby constructing a causal chain. For example, if the user enters a question like "My online service is responding slowly. What's going on?", the question can be broken down into the following sub-questions: Are high concurrent requests causing increased service pressure? Is the network connection slow? Are database queries slow? Have there been any recent code releases or configuration changes? Is there any resource contention (CPU (Central Processing Unit), memory, and disk I / O))? There is a causal relationship between these sub-problems, such as "slow response" ← "database slowdown" ← "missing index or full connection number".

[0038] Step 140 , based on the sub-questions, a multi-source dynamic search is performed on the multi-source vector database, the answers corresponding to the sub-questions are retrieved, and a diagnosis result and solution to the problem are generated.

[0039] For each sub-question, retrieval enhancement generation technology can be used to perform multi-source dynamic retrieval on the data sources in the multi-source vector database. That is, multiple heterogeneous knowledge sources can be dynamically selected for information retrieval based on the user's question type, and relevant evidence or answers can be retrieved from multiple data sources to support each sub-question, forming a causal reasoning chain analysis.

[0040] For example, the user enters the question: "My online service is responding slowly, what's going on?" Break the problem down into sub-questions like the following: Are there high concurrent requests causing stress on the service? Is the network connection slow? Are database queries slow? Has there been a recent code release or configuration change? Is there resource contention (CPU, memory, disk IO)? The results of a multi-source dynamic search in a multi-source vector database were as follows: performance monitoring data revealed that CPU utilization had reached 95%, indicating resource contention; database logs revealed a sudden increase in the execution time of a certain SQL (Structured Query Language) query, involving a large, unindexed table; and the release log database revealed that a database structure adjustment had been performed just the previous night.

[0041] Based on retrieval-enhanced generation technology, diagnostic structures and solutions to problems can be generated based on the retrieved answers.

[0042] The present application provides an operation and maintenance method based on a multi-source dynamic retrieval mechanism, comprising: collecting multiple data sources related to operation and maintenance, preprocessing the multiple data sources to obtain a multi-source vector database; obtaining a question input by a user and determining the problem type based on a preset large model; if the problem type is a diagnostic problem, using a causal reasoning chain mechanism to decompose the problem into multiple sub-problems with causal relationships; performing a multi-source dynamic search on the multi-source vector database based on the sub-problems, retrieving the answers corresponding to the sub-problems, and generating a diagnostic result and solution for the problem. The above method, through the multi-source dynamic retrieval mechanism, can intelligently extract relevant information from different data sources (such as operation and maintenance logs, warning logs, real-time performance data, etc.) based on the type and intent of the problem; secondly, when the problem is a diagnostic problem, a causal reasoning chain mechanism is used for reasoning. During the problem diagnosis process, the causal reasoning chain mechanism can not only decompose complex fault problems, but also gradually analyze potential causes based on causal relationships and generate logical solutions. It can provide deep intelligent support in multiple links such as fault diagnosis, solution recommendation, and inspection report generation, avoiding misdiagnosis or overdiagnosis, thereby greatly improving operation and maintenance efficiency and accuracy.

[0043] Figure 2 This is another flowchart of the operation and maintenance method based on the multi-source dynamic retrieval mechanism shown in an embodiment of the present application.

[0044] See also Figure 2 , the method comprising: Step 210 : Collect multiple data sources related to operation and maintenance, pre-process the multiple data sources, and obtain a multi-source vector database.

[0045] In the embodiment of the present application, multiple data sources related to operation and maintenance can be collected, such as operation and maintenance logs, warning logs, document data, real-time performance indicator monitoring data, etc. The collected data sources are preprocessed to obtain a multi-source vector database.

[0046] In an optional embodiment of the present application, step 210 includes: Sub-step 211, cleaning, classifying, and labeling the data source; Sub-step 212, based on the type of data source, using different methods to segment the data source to obtain data blocks; Sub-step 213, vectorizing the data source; Sub-step 214: storing the data block and the vectorized data source.

[0047] Uniformly extract, clean, deduplicate, and label data from different sources and in different formats. Remove invalid data, such as blank records, duplicate records, and irrelevant content; unify time, date, and numerical units; use jieba for word segmentation on text data, remove stop words (such as meaningless words like "de", "shi", "le", etc.), and delete noise (such as special characters, HTML (HyperText Markup Language) tags, etc.); classify data according to themes or domains, such as "storage management", "operation and maintenance optimization", "performance optimization", etc., and assign labels to each classification for subsequent retrieval optimization.

[0048] For different types of data sources, different methods can be used to split them into data chunks (Chunks). For example: If the data source is document data (such as operation and maintenance manuals, regulations, and fault case documents), use LangChain (an LLM programming framework) to split it into smaller segments at the paragraph level or semantic complete units. Each unit represents an independent semantic chunk (Chunk), and add meta-information (Metadata) to each segment, such as titles and labels. If the data source is log data (such as application logs, alarm logs, and exception records), split it by single log records, and perform a sliding window combination on consecutive log chunks to form context log segments for convenient problem chain reasoning. If the data source is real-time performance metric monitoring data (such as CPU load, response time, database connection count, etc.), split it by each metric such as "CPU usage rate" according to a time window (such as one record per minute). If the data source is a fault knowledge base, automatically extract the four-section structure of problem - cause - solution - result for each knowledge entry as a data unit.

[0049] For easier understanding, the following Table 1 is an example of the document data to be split, and Table 2 is an example after the document data is split.

[0050] Table 1

[0051] Table 2

[0052] Document data, log data, and fault knowledge bases can be vectorized using BERT (Bidirectional Encoder Representations from Transformers, a two-way encoder representation based on Transformer). Real-time performance metric monitoring data can use the time series model LSTM (Long Short-Term Memory) to model the time series part of the monitoring data and extract features for data vectorization.

[0053] MongoDB (distributed document storage database) is used as the document database to store text snippets and metadata, and FAISS (Facebook AI Similarity Search) is used as the vector database to store document vectors.

[0054] Step 220: Obtain the question input by the user and determine the question type based on the preset large model.

[0055] Users can enter the questions they want to ask on the system's interactive interface. For example, if the question the user enters is: "My online service response is slow, what's going on?" The preset large model can be called to determine the type of problem and determine whether the problem is a diagnostic problem or a general knowledge problem.

[0056] In step 230 , if the problem type is a diagnostic problem, the preset large model calls the inference chain template and decomposes the problem into multiple sub-problems with causal relationships based on the inference chain template.

[0057] If the problem type is determined to be a diagnostic issue, after receiving the user's question, the system will call the preset large model and the causal reasoning module. Using the reasoning chain template automatically generated by the internal call to the preset large model, the problem will be broken down into multiple sub-problems. Based on the causal relationships between the sub-problems, the system will conduct a step-by-step analysis to build a complete causal reasoning chain and locate the potential root cause of the problem. For example, if the user inputs the question "My online service response is slow. What's going on?", after receiving the user's question, the system will call the preset large model and generate a set of causal sub-questions: Are high concurrent requests causing increased service pressure? Is the network connection slow? Are database queries slow? Has there been a recent code release or configuration change? Is there any resource contention (CPU, memory, disk IO)? There is a causal relationship between these sub-problems, such as "slow response" ← "database slowdown" ← "missing index or full connection number".

[0058] Step 240: Dynamically determine the data source to be searched in the multi-source vector database based on the question.

[0059] According to the question, the data source most relevant to the question can be determined in the multi-source vector database, and the relevant data source is determined as the data source to be retrieved.

[0060] In an optional embodiment of the present application, step 240 includes: Sub-step 241, determining data sources related to the question and determining the search priority; Sub-step 242 , dynamically determining the data source to be retrieved based on the question and the current operating status.

[0061] After a user enters a question, a question classifier is invoked to determine the type of problem (e.g., network, database, or security). Based on the question, the most relevant data source in the multi-source vector database is identified, and the search priority among data sources is determined. A search strategy is then developed to determine how to prioritize searches. The question type classifier is implemented using a machine learning classification algorithm.

[0062] Based on the problem and the current state of the system, the system dynamically selects the data sources to search. For example, if a user reports a database performance issue, the system prioritizes extracting information from data sources such as database logs, query performance, and configuration files. If a network connectivity issue is reported, the system prioritizes searching network performance metrics and network fault logs.

[0063] Step 250: Retrieve answers corresponding to the multiple sub-questions from the data source to be retrieved.

[0064] Based on the selected retrieval strategy, the system retrieves relevant data for each sub-problem from each data source. Each data source has its own dedicated retrieval module: Log Data Retrieval: Uses log analysis tools or regular expressions to quickly retrieve relevant error and warning information; Performance Monitoring Data Retrieval: Obtains the latest resource usage information through real-time performance indicator monitoring systems; Document Data Retrieval: Uses text retrieval algorithms to search the document library for knowledge base entries related to the current problem; Configuration File Data Retrieval: Extracts relevant configuration items through configuration management tools or scripts to check for any configuration changes; Historical Fault Data Retrieval: Searches the historical fault case library to identify similar faults and their solutions.

[0065] Step 260: Integrate the answers corresponding to the multiple sub-questions and input the integrated results into the preset large model.

[0066] Because data comes from different sources, the answers to multiple sub-questions may contain duplicate information, different formats, or redundant content. The system needs to remove duplicates, filter, and sort, integrating and cleaning the retrieved content from multiple data sources to ensure the high value of the information returned. After integration, the integrated results, questions, and prompts are input into the pre-set master model.

[0067] Step 270: Generate a diagnosis result and solution for the problem based on a preset inspection report template.

[0068] The system will automatically generate a structured report based on the analysis process and generate an inspection report based on the preset inspection report template. The inspection report will display the diagnosis results and solutions of the problem. The preset inspection report can display detailed inspection content, such as inspection purpose, inspection time, inspection items such as system health status, network status, application services, abnormal conditions and risk points, optimization suggestions, etc. Table 3 shows an example of a preset inspection report template: Table 3

[0069] The above table content is only an example of a preset inspection report. This application does not limit the specific content and report format of the preset inspection report.

[0070] In an optional embodiment of the present application, the steps after determining the question type based on the preset large model further include: If the question type is a general knowledge question, the data source to be searched in the multi-source vector database is dynamically determined based on the question; Retrieve the answer to the question from the data source to be retrieved, and input the answer into the preset large model; Generate diagnostic results and solutions to problems based on preset inspection report templates.

[0071] In one example, if the question type is a general knowledge question, a multi-source dynamic search is performed within the multi-source vector database. The data source to be searched is dynamically determined based on the question. The answer to the question is retrieved from the data source and input into a pre-set macro model. Using search-enhanced generation technology, the pre-set macro model generates a diagnosis and solution for the problem based on a pre-set inspection report template.

[0072] like Figure 3As shown, it is a flow chart of the operation and maintenance method based on the multi-source dynamic retrieval mechanism shown in the embodiment of the present application. The user enters the question "My disk is full, how can I solve the problem of insufficient disk space?". It is determined whether the problem is a general knowledge problem. If so, a multi-source dynamic search is performed on the various data resources in the multi-source vector database. If not (that is, a diagnostic problem), the LLM is called to decompose the problem into multiple sub-problems with causal relationships, and a causal reasoning chain is constructed to perform a multi-source dynamic search on the various categories of data resources in the multi-source vector database. When performing the multi-source dynamic search, the problem classifier is called to determine what type of problem the problem belongs to (such as network problems, database problems, security problems, etc.), and the most relevant data source in the multi-source vector database is determined according to the problem, and the search priority between the data sources is determined. A search strategy is formulated to decide how to prioritize the search. If the question is general knowledge, the search results are directly entered into the LLM. If the question is non-general knowledge, a causal inference chain mechanism is used to retrieve relevant evidence from multiple data sources to support each sub-question. This allows for evidence correlation analysis within the inference chain and integrates the answers to multiple sub-questions to produce an integrated result: "chunk 1: Delete temporary or junk files. Chunk 2: Check if log files are too large and archive them regularly. Chunk 3;..." A prompt is set: "Please answer the above question based on the following information:" The question, integrated results, and prompt are entered into the LLM. If RAG is used, the LLM provides an answer, generating a diagnostic result and solution based on the question: "First, delete unnecessary temporary or junk files: 1. In Windows, you can open the Disk Cleanup tool, select "Temporary Files," and click Clean. 2. In Linux, run the command rmrf / tmp / * to clear temporary files...." If RAG is not used, the LLM responds: "Insufficient disk space may be caused by a variety of reasons, and I don't know the specifics." The generated diagnostic result and solution are then fed back to the user.

[0073] In addition, the solution of the embodiment of the present application can use an adaptive retrieval and reasoning mechanism. After receiving the generated diagnostic results and solutions, users can provide feedback, and dynamic optimization can be performed based on user feedback to further improve the accuracy of problem diagnosis, avoiding the lack of real-time feedback and optimization in traditional operations and maintenance.

[0074] An embodiment of the present application provides an operation and maintenance method based on a multi-source dynamic retrieval mechanism, including: collecting multiple data sources related to operation and maintenance, preprocessing the multiple data sources to obtain a multi-source vector database, obtaining a question input by a user, and judging the problem type of the question based on a preset large model. If the problem type of the question is a diagnostic problem, the preset large model calls an inference chain template; based on the inference chain template, the problem is decomposed into multiple sub-problems with causal relationships, the data source to be retrieved in the multi-source vector database is dynamically determined based on the problem, answers corresponding to multiple sub-problems are retrieved from the data source to be retrieved, the answers corresponding to the multiple sub-problems are integrated, and the integrated results are input into the preset large model, and a diagnostic result and solution to the problem are generated based on a preset inspection report template. Through the above solution, relevant information can be intelligently extracted from different data sources (such as operation and maintenance logs, warning logs, real-time performance data, etc.) according to the type of problem and user intention. The priority of the data source can be adjusted according to the actual situation to ensure more accurate retrieval results. Secondly, when the problem is a diagnostic problem, a causal reasoning mechanism chain is used for reasoning. During the problem diagnosis process, the causal reasoning chain mechanism can not only decompose complex fault problems, but also gradually analyze potential causes based on cause-and-effect relationships and generate logical solutions. It can provide deep intelligent support in multiple links such as fault diagnosis, solution recommendation and inspection report generation, avoid misdiagnosis or overdiagnosis, and thus greatly improve operation and maintenance efficiency and accuracy.

[0075] Corresponding to the aforementioned application function implementation method embodiment, the present application also provides an operation and maintenance device, electronic device and corresponding embodiments based on a multi-source dynamic retrieval mechanism.

[0076] Figure 4 It is a structural diagram of an operation and maintenance device based on a multi-source dynamic retrieval mechanism shown in an embodiment of the present application.

[0077] See also Figure 4 , the device comprises: The acquisition module 410 is used to acquire multiple data sources related to operation and maintenance, and pre-process the multiple data sources to obtain a multi-source vector database.

[0078] In the embodiment of the present application, multiple data sources related to operation and maintenance can be collected, such as operation and maintenance logs, warning logs, document data, real-time performance indicator monitoring data, fault databases, etc. The collected data sources are preprocessed to obtain a multi-source vector database.

[0079] The judgment module 420 is used to obtain the question input by the user and judge the question type based on the preset large model.

[0080] Users can enter their questions on the system's interactive interface. For example, if a user asks, "My online service is responding slowly. What's going on?" A pre-set large model can be used to determine the type of problem, whether it's a diagnostic or general question. The pre-set large model can be a large language model.

[0081] The decomposition module 430 is used to decompose the problem into multiple sub-problems with causal relationships by using a causal reasoning chain mechanism if the problem type is a diagnostic problem.

[0082] If the problem type is determined to be diagnostic, after receiving the user's question, a pre-set large model is invoked and a causal chain mechanism is used to break the problem down into a set of causally related sub-questions, thereby constructing a causal chain. For example, if the user enters a question like "My online service is responding slowly. What's going on?", the question can be broken down into the following sub-questions: Are high concurrent requests causing increased service pressure? Is the network connection slow? Are database queries slow? Have there been any recent code releases or configuration changes? Is there any resource contention (CPU (Central Processing Unit), memory, and disk I / O))? There is a causal relationship between these sub-problems, such as "slow response" ← "database slowdown" ← "missing index or full connection number".

[0083] The retrieval module 440 is used to perform a multi-source dynamic search on the multi-source vector database based on the sub-questions, retrieve the answers corresponding to the sub-questions, and generate diagnostic results and solutions to the problems.

[0084] For each sub-question, multi-source dynamic retrieval can be performed on the data sources in the multi-source vector database, that is, multiple heterogeneous knowledge sources can be dynamically selected for information retrieval based on the user's question type, and relevant evidence or answers can be retrieved from multiple data sources to support each sub-question, forming a causal reasoning chain analysis.

[0085] For example, the user enters the question: "My online service is responding slowly, what's going on?" Break the problem down into sub-questions like the following: Are there high concurrent requests causing stress on the service? Is the network connection slow? Are database queries slow? Has there been a recent code release or configuration change? Is there resource contention (CPU, memory, disk IO)? The results of a multi-source dynamic search in a multi-source vector database were as follows: performance monitoring data revealed that CPU utilization had reached 95%, indicating resource contention; database logs revealed a sudden increase in the execution time of a certain SQL (Structured Query Language) query, involving a large, unindexed table; and the release log database revealed that a database structure adjustment had been performed just the previous night.

[0086] Retrieval-enhanced generation technology can generate diagnostic structures and solutions to problems based on the retrieved answers.

[0087] In an optional embodiment of the present application, the disassembly module 430 includes: The calling submodule is used to preset the large model calling inference chain template; The disassembly submodule is used to decompose the problem into multiple sub-problems with causal relationships based on the reasoning chain template.

[0088] In an optional embodiment of the present application, the calling module 440 includes: A determination submodule, for dynamically determining the data source to be retrieved in the multi-source vector database based on the question; A retrieval submodule, configured to retrieve answers corresponding to the multiple sub-questions from the data source to be retrieved; The integration submodule is used to integrate the answers corresponding to multiple sub-questions and input the integrated results into the preset large model; The generation submodule is used to generate diagnostic results and solutions to problems based on the preset inspection report template.

[0089] In an optional embodiment of the present application, the determining submodule is further configured to: Identify data sources relevant to the problem and prioritize them for retrieval; Dynamically determine the data source to be retrieved based on the problem and current operating status.

[0090] In an optional embodiment of the present application, the device further includes: A determination module, for dynamically determining the data source to be searched in the multi-source vector database based on the question type if the question is a general knowledge question; The input module is used to retrieve the answer corresponding to the question from the data source to be retrieved and input the answer into the preset large model; The generation module is used to generate diagnostic results and solutions for problems based on preset inspection report templates.

[0091] In an optional embodiment of the present application, the device further includes: Cleaning module, used to clean, classify, and label data sources; The segmentation module is used to segment the data source into data blocks using different methods based on the type of data source; Vectorization module, used to vectorize data sources; The storage module is used to store data blocks and vectorized data sources.

[0092] The embodiment of the present application provides an operation and maintenance device based on a multi-source dynamic retrieval mechanism, which can intelligently extract relevant information from different data sources (such as operation and maintenance logs, warning logs, real-time performance data, etc.) according to the type of problem and user intention, and can adjust the priority of the data source according to the actual situation to ensure more accurate retrieval results; secondly, when the problem is a diagnostic problem, a causal reasoning mechanism chain is used for reasoning. During the problem diagnosis process, the causal reasoning chain mechanism can not only decompose complex fault problems, but also gradually analyze potential causes based on cause-and-effect relationships, and generate logical solutions. It can provide deep intelligent support in multiple links such as fault diagnosis, solution recommendation, and inspection report generation, avoid misdiagnosis or overdiagnosis, and thus greatly improve operation and maintenance efficiency and accuracy.

[0093] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.

[0094] Figure 5 It is a structural diagram of an electronic device shown in an embodiment of the present application.

[0095] See also Figure 5 , the electronic device 500 includes a memory 510 and a processor 520.

[0096] The processor 520 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. Memory 510 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage. ROM may store static data or instructions required by processor 520 or other computer modules. Permanent storage may be a readable and writable storage device. Permanent storage may be a non-volatile storage device that retains stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device utilizes a mass storage device (e.g., a magnetic or optical disk, flash memory). In other embodiments, the permanent storage device may be a removable storage device (e.g., a floppy disk, optical drive). System memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory (DRAM). System memory may store some or all instructions and data required by the processor during operation. Furthermore, memory 510 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), as well as magnetic disks and / or optical disks. In some embodiments, the memory 510 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.

[0097] The memory 510 stores executable codes. When the executable codes are processed by the processor 520 , the processor 520 may execute part or all of the above-mentioned methods.

[0098] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.

[0099] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium), which stores executable code (or computer program or computer instruction code) and, when executed by a processor of an electronic device (or server, etc.), enables the processor to perform part or all of the steps of the above-mentioned method according to the present application.

[0100] The present application also provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the above method is implemented.

[0101] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

Claims

1. An operation and maintenance method based on a multi-source dynamic retrieval mechanism, characterized in that: include: Collecting multiple data sources related to operation and maintenance, and preprocessing the multiple data sources to obtain a multi-source vector database; Obtain the question input by the user and determine the question type based on the preset large model; If the problem type is a diagnostic problem, the causal reasoning chain mechanism is used to decompose the problem into multiple sub-problems with causal relationships; Based on the sub-problems, a multi-source dynamic search is performed on the multi-source vector database, answers corresponding to the sub-problems are retrieved, and a diagnosis result and solution for the problem are generated.

2. The method according to claim 1, characterized in that The causal reasoning chain mechanism is used to decompose the problem into multiple causal sub-problems, including: The preset large model calls the inference chain template; The problem is decomposed into multiple sub-problems with causal relationships based on the reasoning chain template.

3. The method according to claim 1, characterized in that The performing of a multi-source dynamic search on the multi-source vector database based on the sub-question, retrieving the answer corresponding to the sub-question, and generating a diagnosis result and a solution for the problem includes: Dynamically determine the data source to be searched in the multi-source vector database based on the question; Retrieve answers corresponding to the plurality of sub-questions from the data source to be retrieved; Integrate the answers corresponding to the multiple sub-questions and input the integrated results into the preset large model; Generate diagnostic results and solutions for the problem based on the preset inspection report template.

4. The method according to claim 3, characterized in that The dynamically determining the data source to be searched in the multi-source vector database based on the question includes: Identify data sources relevant to the stated problem and prioritize them for retrieval; Based on the problem and the current operating status, the data source to be retrieved is dynamically determined.

5. The method according to claim 1, wherein The step after determining the type of the problem based on the preset large model also includes: If the question type is a general knowledge question, dynamically determining a data source to be searched in the multi-source vector database based on the question; Retrieving the answer corresponding to the question from the data source to be retrieved, and inputting the answer into the preset large model; Generate diagnostic results and solutions for the problem based on the preset inspection report template.

6. The method according to claim 1, characterized in that The preprocessing of the plurality of data sources comprises: Cleaning, classifying, and labeling the data source; Based on the type of the data source, different methods are used to segment the data source to obtain data blocks; vectorizing the data source; The data blocks and the vectorized data source are stored.

7. An operation and maintenance device based on a multi-source dynamic retrieval mechanism, characterized in that: include: A collection module, used to collect multiple data sources related to operation and maintenance, and pre-process the multiple data sources to obtain a multi-source vector database; A judgment module is used to obtain a question input by a user and determine the type of the question based on a preset large model; A decomposition module is used to decompose the problem into multiple sub-problems with causal relationships using a causal reasoning chain mechanism if the problem type is a diagnostic problem; The retrieval module is used to perform a multi-source dynamic search on the multi-source vector database based on the sub-problem, retrieve the answer corresponding to the sub-problem, and generate a diagnosis result and solution for the problem.

8. The device according to claim 7, characterized in that The calling module includes: a determination submodule, configured to dynamically determine a data source to be retrieved in the multi-source vector database based on the question; A retrieval submodule, configured to retrieve answers corresponding to the plurality of sub-questions from the data source to be retrieved; An integration submodule, configured to integrate the answers corresponding to the plurality of sub-questions and input the integrated results into the preset large model; The generation submodule is used to generate the diagnosis results and solutions of the problem based on the preset inspection report template.

9. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having executable codes stored thereon, wherein when the executable codes are executed by a processor of an electronic device, the processor is caused to execute the method according to any one of claims 1 to 6.