Database anomaly diagnosis method and system based on large language model

Through the database exception diagnosis method based on large language models, the shortcomings of existing tools in dealing with complex and new exceptions are solved, efficient and accurate abnormal diagnosis and detailed report generation are achieved, and continuous optimization is carried out through user feedback.

CN120104385APending Publication Date: 2025-06-06TSINGHUA UNIVERSITY +1

Patent Information

Application Number
CN202510029571.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing database exception diagnosis tools lack effective inference capabilities and the ability to automatically generate detailed reports, making them difficult to deal with complex exceptions and are unable to handle new exceptions.

Method used

Design a database abnormal diagnosis method based on large language models. By splitting the diagnostic documents according to chapters, extracting knowledge fragments, clustering to generate prompt word templates, matching exception information and inputting the large language model for parallel processing, combining the tree search algorithm to optimize the diagnostic path, and supporting iterative optimization of user feedback.

Benefits of technology

It improves the accuracy and efficiency of database abnormal diagnosis, can effectively handle complex abnormalities and new abnormalities, generate detailed diagnostic reports, and continuously optimize diagnostic results through user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104385A_ABST
    Figure CN120104385A_ABST
Patent Text Reader

Abstract

The invention discloses a database anomaly diagnosis method and system based on a large language model. The method comprises the steps that a diagnosis document is split according to chapters to extract diagnosis knowledge fragments so as to construct a diagnosis knowledge base; clustering the diagnosis knowledge fragments according to a preset topic, and generating a corresponding prompt word template for each cluster; judging whether the monitored index value in the diagnosis knowledge base exceeds a preset threshold value or not, and matching related diagnosis knowledge from the diagnosis database according to the prompt word template based on a judgment result that the index value exceeds the preset threshold value so as to collect corresponding abnormal information; and inputting the exception information into a plurality of large language models for parallel processing to obtain corresponding database exception diagnosis results, and synthesizing the diagnosis results of the large language models through a tree search algorithm to carry out diagnosis path optimization. The method aims at effectively analyzing abnormal root causes of different databases in a real scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information retrieval, and in particular to a database anomaly diagnosis method and system based on a large language model. Background Art

[0002] Database anomaly diagnosis refers to the process of detecting, analyzing, and identifying abnormal behaviors or abnormal states in database systems, identifying potential root causes of failures, and proposing solutions. Database anomaly diagnosis is critical to ensuring high availability of databases. Most companies rely on manual diagnosis by database administrators to deal with daily anomalies, which usually takes a lot of time and effort. At the same time, with the increasing number of database instances, manual diagnosis methods are increasingly difficult to meet scale requirements. Currently, most semi-automatic diagnostic tools rely on preset rules and small-scale machine learning models. The main limitations of such tools are that they lack effective reasoning capabilities and the ability to automatically generate detailed reports, and they are unable to cope with new or complex anomalies (such as root causes involving multiple modules). Therefore, the anomaly diagnosis scenarios that such tools can usually support are very limited.

[0003] The recent development of large language models has given them excellent capabilities in natural language understanding, reasoning, and generation, making them a potential new solution for database anomaly diagnosis. Diagnostic systems based on large language models can automate database diagnosis in the following ways: (1) understanding abnormal situations; (2) using relevant knowledge and database tools for analysis; (3) following user feedback, such as personal needs, professional advice, etc.; and (4) generating readable reports.

[0004] However, using large language models to build database diagnosis systems still faces the following three challenges. First, how to effectively integrate the diverse knowledge of database administrators to enhance the diagnostic capabilities of large language models? Experienced database administrators can provide a variety of diagnostic experiences (such as documents, diagnostic cases, etc.), which can effectively make up for the shortcomings of large language models in the task of database anomaly diagnosis. However, without unified processing, existing large language models find it difficult to effectively utilize these diverse experiences.

[0005] Secondly, how to improve the ability of large language models to reason and solve complex abnormality diagnosis tasks? Complex diagnostic tasks often require tedious exploratory analysis, and large language models have hallucination problems and early stopping problems. Therefore, it is crucial to design an effective reasoning mechanism to help large language models improve their ability to solve complex diagnostic tasks.

[0006] Finally, how can we effectively combine user feedback to improve diagnostic accuracy and meet user needs? Large language models may have omissions during the diagnosis process (such as misunderstanding monitoring logs, etc.), and users may also have customized needs (such as the language category of generated reports). Therefore, it is also very important to leverage user experience, allow users to provide corresponding feedback on diagnostic results (such as screening out some causes), and make improvements based on feedback to improve diagnostic accuracy.

[0007] Currently, there is no database anomaly diagnosis method based on large language models. Summary of the invention

[0008] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0009] Therefore, in order to fill this gap and overcome the shortcomings and challenges of existing methods, the present invention intends to design a database anomaly diagnosis method based on a large language model to effectively analyze the root causes of different database anomalies in real scenarios.

[0010] Another object of the present invention is to provide a database anomaly diagnosis system based on a large language model.

[0011] To achieve the above object, the present invention proposes a database anomaly diagnosis method based on a large language model, comprising:

[0012] Split the diagnostic documents by chapters to extract diagnostic knowledge fragments to build a diagnostic knowledge base;

[0013] Clustering the diagnostic knowledge fragments according to preset themes, and generating corresponding prompt word templates for each cluster;

[0014] Determine whether the indicator value monitored in the diagnostic knowledge base exceeds a preset threshold, and based on the judgment result of exceeding the preset threshold, match relevant diagnostic knowledge from the diagnostic database according to the prompt word template to collect corresponding abnormal information;

[0015] The abnormal information is input into multiple large language models for parallel processing to obtain the corresponding database abnormality diagnosis result, and the diagnosis results of each large language model are integrated through a tree search algorithm to optimize the diagnosis path.

[0016] The database anomaly diagnosis method based on a large language model in the embodiment of the present invention may also have the following additional technical features:

[0017] In one embodiment of the present invention, the diagnosis document is split into chapters to extract diagnostic knowledge fragments, including:

[0018] Extract preliminary knowledge fragments from existing diagnostic documents based on the document's chapter structure;

[0019] Diagnostic knowledge fragments are retrieved from the preliminary knowledge fragments through metadata filtering and similarity search.

[0020] In one embodiment of the present invention, the abnormal information is input into multiple large language models for parallel processing to obtain corresponding database abnormality diagnosis results, and the diagnosis results of each large language model are integrated through a tree search algorithm to optimize the diagnosis path, including:

[0021] Identify potential types of abnormal information when receiving a diagnosis task, and assign the abnormal information to a related large language model agent based on the identified abnormal information type;

[0022] The large language model agent performs calculations through a tree search algorithm and shares the intermediate diagnosis results of the large language model agent. It also searches multiple reasoning paths and records the failed paths encountered to select the optimal path as the final diagnosis result.

[0023] The final diagnosis results of each large language model agent are aggregated to generate a comprehensive anomaly diagnosis report.

[0024] In one embodiment of the present invention, when receiving a diagnosis task, identifying the potential type of abnormal information, and assigning the abnormal information to a related large language model agent based on the identified abnormal information type, includes:

[0025] Filter diagnostic knowledge fragments based on metadata, and then use vector similarity as an indicator to search for the knowledge fragment that is most similar to the abnormal description;

[0026] The abnormality-related information and the most similar knowledge fragment retrieved are embedded into the prompt word template of the agent of the corresponding topic.

[0027] In one embodiment of the present invention, after generating the comprehensive abnormality diagnosis report, the method further includes:

[0028] Obtain user feedback information;

[0029] Iteratively optimize the final diagnosis result according to user feedback information to obtain an optimized diagnosis result;

[0030] A corresponding optimization model is generated according to the optimized diagnosis result, and the optimization model is added to a diagnosis knowledge base.

[0031] In one embodiment of the present invention, the final diagnosis result is iteratively optimized according to the user feedback information to obtain an optimized diagnosis result, including:

[0032] Use a large language model to extract key points from user feedback information and generate corresponding test programs for each key point;

[0033] Execute the test program on the generated final diagnostic results, and adjust and optimize according to the test error log until all tests pass or the preset upper limit of the number of iterations is reached

[0034] In one embodiment of the present invention, generating a corresponding optimization pattern according to the optimized diagnosis result and adding the optimization pattern to a diagnosis knowledge base includes:

[0035] Comparing the diagnosis results before and after the optimization to obtain a comparison result;

[0036] According to the comparison results and using a large language model, an effective optimization pattern is identified and extracted;

[0037] The extracted optimization patterns are stored in a diagnostic knowledge base.

[0038] To achieve the above object, the present invention further provides a database anomaly diagnosis system based on a large language model, comprising:

[0039] A diagnostic knowledge base construction module is used to split diagnostic documents by chapters to extract diagnostic knowledge fragments to construct a diagnostic knowledge base;

[0040] A fragment clustering module, used to cluster the diagnostic knowledge fragments according to preset themes and generate a corresponding prompt word template for each cluster;

[0041] The abnormal information collection module is used to determine whether the indicator value monitored in the diagnostic knowledge base exceeds a preset threshold, and based on the judgment result of exceeding the preset threshold, match the relevant diagnostic knowledge from the diagnostic database according to the prompt word template to collect the corresponding abnormal information;

[0042] The abnormal result diagnosis module is used to input the abnormal information into multiple language models for parallel processing to obtain the corresponding database abnormality diagnosis results, and optimize the diagnosis path by integrating the diagnosis results of each large language model through a tree search algorithm.

[0043] The database anomaly diagnosis method and system based on a large language model of the embodiment of the present invention implements knowledge extraction through a hybrid retrieval strategy for different anomaly alarms, uses multi-model collaborative reasoning to analyze the root cause of the anomaly, and continuously optimizes the diagnosis results in combination with user feedback.

[0044] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0046] Figure 1 is a flowchart of a database anomaly diagnosis method based on a large language model according to an embodiment of the present invention;

[0047] Figure 2 is a general architecture diagram of a database anomaly diagnosis method based on a large language model according to an embodiment of the present invention;

[0048] Figure 3 is a specific flow chart of a database anomaly diagnosis method based on a large language model according to an embodiment of the present invention;

[0049] Figure 4 It is a structural diagram of a database anomaly diagnosis system based on a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0050] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0051] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0052] The following describes a database anomaly diagnosis method and system based on a large language model according to an embodiment of the present invention with reference to the accompanying drawings.

[0053] Figure 1 is a flow chart of a database anomaly diagnosis method based on a large language model according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0054] S1, split the diagnostic documents by chapters to extract diagnostic knowledge fragments to build a diagnostic knowledge base;

[0055] S2, clustering the diagnostic knowledge fragments according to preset themes, and generating corresponding prompt word templates for each cluster;

[0056] S3, judging whether the indicator value monitored in the diagnostic knowledge base exceeds a preset threshold, and matching relevant diagnostic knowledge from the diagnostic database according to the prompt word template based on the judgment result of exceeding the preset threshold to collect corresponding abnormal information;

[0057] S4, inputting the abnormal information into multiple large language models for parallel processing to obtain corresponding database abnormality diagnosis results, and optimizing the diagnosis path by integrating the diagnosis results of each large language model through a tree search algorithm.

[0058] It is understandable that the database anomaly diagnosis method proposed in the present invention extracts key knowledge from existing diagnostic documents through offline knowledge extraction, adopts a hybrid retrieval strategy to achieve knowledge matching, and fills the matched knowledge into the prompt words of the large language model to support online diagnosis, and supports the abnormal monitoring function, which can detect anomalies with the help of indicators and collect detailed information of anomalies for further analysis and diagnosis by the large language model. The present invention also introduces a multi-large language model collaborative diagnosis mechanism, in which a large language model agent that masters diagnostic knowledge in different fields performs accurate diagnosis through tree search algorithms and cross-examination. At the same time, a user feedback improvement mechanism is designed to create demand tests based on user feedback and iteratively optimize the response of the large language model, and supports the extraction of optimization patterns and storage in the knowledge base for subsequent use.

[0059] Figure 2 and Figure 3 This is the overall architecture diagram of the database anomaly diagnosis method based on the large language model of the present invention, and its overall process is as follows:

[0060] Firstly, concise knowledge fragments are extracted from existing diagnostic documents based on the chapter structure of the documents, and relevant knowledge fragments are actively retrieved through metadata filtering and similarity search to guide the large language model to generate content.

[0061] In one embodiment of the present invention, database diagnostic knowledge in existing diagnostic documents is extracted offline. The documents are split according to chapter structure, and refined knowledge fragments (such as "relevant indicators include...; the steps are as follows...") are extracted for use in subsequent diagnosis.

[0062] Then, the diagnostic knowledge fragments are clustered according to preset themes, and a corresponding prompt word template is generated for each cluster.

[0063] In one embodiment of the present invention, the agent role is generated and assigned: the database diagnostic knowledge extracted in the above steps is clustered by topics (such as CPU, load, etc.), and a corresponding prompt word template is generated for each cluster to reflect the characteristics of the large model agent.

[0064] Then, it is determined whether the indicator value monitored in the diagnosis knowledge base exceeds a preset threshold, and based on the determination result of exceeding the preset threshold, relevant diagnosis knowledge is matched from the diagnosis database according to the prompt word template to collect corresponding abnormal information.

[0065] In one embodiment of the present invention, the database is automatically detected to see if an abnormality has occurred by monitoring whether certain indicators exceed normal thresholds (e.g., available memory is less than 10%). Then, the abnormality details (e.g., slow SQL query logs) are collected from the database information and provided to the large language model for further diagnosis.

[0066] Finally, the abnormal information is input into multiple large language models for parallel processing to obtain the corresponding database abnormality diagnosis results, and the diagnosis results of each large language model are integrated through a tree search algorithm to optimize the diagnosis path.

[0067] In one embodiment of the present invention, a multi-language model collaborative diagnosis mechanism is designed, in which each large language model agent masters diagnostic knowledge of different topics (such as CPU, load, etc.), explores multi-step reasoning paths through a tree search algorithm, and cross-examines each other to improve the quality of diagnosis.

[0068] It is understandable that complex diagnostic tasks often require tedious exploratory analysis, and large language models have hallucination problems and early stopping problems. Therefore, the present invention introduces a multi-large language model collaborative diagnosis mechanism. Different large language model agents are formed for various abnormal topics. For abnormal information, each agent uses a tree search algorithm to generate an inference path, and exchanges and cross-examines information to further improve the accuracy of diagnosis.

[0069] Specifically, when a diagnosis task is received, the potential type of abnormal information is identified, and the abnormal information is assigned to the relevant large language model agent based on the identified type of abnormal information; the large language model agent performs calculations through a tree search algorithm, and shares the intermediate diagnosis results of the large language model agent, while searching multiple reasoning paths to record the failed paths encountered to select the optimal path as the final diagnosis result; the final diagnosis results of each large language model agent are summarized to generate a comprehensive abnormal diagnosis report.

[0070] In order to let each large language model focus on a specific topic, the present invention equips them with document libraries of different topics. Therefore, when receiving a diagnosis task, the present invention first identifies the potential possible types of abnormalities and assigns them to the relevant large language model agents based on the type.

[0071] For the relevant information of the anomaly, a hybrid retrieval method is used to improve the accuracy of the retrieved knowledge fragments: first filter the fragments based on metadata (such as topics, indicators), and then search for the knowledge fragments most similar to the anomaly description based on vector similarity (such as using Chroma).

[0072] When all necessary knowledge and abnormal information are ready, the abnormality-related information and the retrieved knowledge block are embedded into the prompt word template of the agent of the corresponding topic for further diagnosis.

[0073] In the multi-step large language model agent reasoning process (e.g., calling tools, retrieving relevant knowledge, etc.), hallucinations and unstable responses are common failures. Therefore, the present invention uses a tree search algorithm to explore multiple possible reasoning paths, record the failed paths encountered, and select the optimal path as the result.

[0074] Multi-agent group discussion. Based on independent diagnosis using the tree search algorithm, each language model agent regularly shares intermediate diagnosis results to avoid misjudgment and optimizes diagnosis quality through mutual review.

[0075] The diagnostic results of each large language model agent are aggregated to generate a comprehensive anomaly diagnosis report that is traceable and easy to read for users.

[0076] Afterwards, a user feedback improvement mechanism was introduced. Based on user feedback, demand tests were created and the responses of the large language model were iteratively optimized until the demand tests were met; at the same time, optimization patterns were extracted (such as "if the response suggests reducing the number of processes, provide specific tools and their usage to achieve this") and stored in the knowledge base for subsequent use.

[0077] Specifically, user feedback information is obtained; the final diagnostic result is iteratively optimized according to the user feedback information to obtain an optimized diagnostic result; a corresponding optimization mode is generated according to the optimized diagnostic result, and the optimization mode is added to the diagnostic knowledge base. The key points are extracted from the user feedback information using a large language model, and a corresponding test program is generated for each key point; the test program is executed on the generated final diagnostic result, and adjustments and optimizations are made according to the test error log until all tests pass or the preset upper limit of the number of iterations is reached. The diagnostic results before and after optimization are compared to obtain a comparison result; according to the comparison result and using a large language model, effective optimization modes are identified and extracted; and the extracted optimization modes are stored in the diagnostic knowledge base.

[0078] In the diagnosis process of the embodiment of the present invention, the large language model may have omissions (such as misunderstanding the monitoring log), and the user may also put forward customized requirements (such as specifying the language category of the report). The method proposed in the present invention supports interaction with the user and iteratively optimizes the diagnosis results according to user feedback to further improve the diagnosis accuracy, including:

[0079] Test generation: In order to ensure that the optimized diagnostic report can achieve the corresponding requirements of feedback, the present invention uses a large language model to analyze and extract key points from user feedback, and generates corresponding tests for each key point.

[0080] Iterative optimization: The above tests are performed on the generated diagnostic results. The large language model will adjust and optimize itself according to the test error log until all tests pass or the preset upper limit of the number of iterations (for example, 3 times) is reached.

[0081] Furthermore, the self-iterative upgrade based on feedback optimization. The present invention can learn from the feedback optimization history of the above steps, generate corresponding optimization patterns, and add them to the knowledge base. For example, the user may require the large language model agent to follow his or her preferences (such as changing the language) or provide professional advice (such as guiding possible diagnostic directions), and the optimized large language model agent response reflects these adaptive changes, including:

[0082] Optimization mode extraction. The present invention compares the diagnosis results before and after optimization and uses a large language model to identify and extract effective optimization modes (for example, if user feedback requires the report to be output in Chinese, the optimization mode "If the user requires output in a specified language, respond in that language for easy understanding by the user" is extracted).

[0083] Knowledge base maintenance of optimization patterns. The extracted optimization patterns will be stored in the feedback library. Before a new pattern is stored, the system will check whether there is a conflicting pattern. If a conflict is found, the new pattern will be retained first and the old conflicting pattern will be removed.

[0084] In summary, the method of the present invention can automatically retrieve diagnostic knowledge similar to abnormal information, and assist the large language model to perform abnormal analysis in combination with existing knowledge like a human database administrator, thereby improving the accuracy of diagnosis. Through the collaborative work of multiple large language models of different topics, the model is gradually guided to find the root cause of the abnormality and propose solutions, while generating a traceable diagnostic report to achieve comprehensive and systematic abnormality troubleshooting. The present invention allows users to provide feedback on diagnostic results and continuously optimize the diagnostic effect to make it more in line with user needs. In addition, the improvement strategies of user feedback can be summarized as rules. These rules can be directly applied to subsequent diagnostic processes to achieve system self-evolution.

[0085] The present invention extracts effective knowledge fragments from existing diagnostic documents, clusters them by topic, generates corresponding prompt word templates, and forms a multi-model intelligent agent system. The system realizes automatic anomaly detection and collects anomaly details by monitoring key indicators, specifies relevant intelligent agents based on anomaly categories, and uses hybrid retrieval strategies to match anomaly-related knowledge fragments, embeds relevant knowledge into large language model prompt words, and uses the large language model for further analysis and use. A multi-large language model collaborative diagnosis mechanism is introduced, and the intelligent agent models in each field independently generate multi-path reasoning and cross-examine through a tree search algorithm. The system supports test-driven optimization based on user feedback, and can extract and generate adaptive optimization patterns from the feedback optimization history, store them in the knowledge base for subsequent use, and continuously iterate the model diagnosis performance.

[0086] It can automatically monitor abnormal information in the database and generate a reasonable and reliable diagnostic report within an acceptable time (for example, within ten minutes), including the identified abnormal causes and recommended solutions. First, the present invention extracts diagnostic knowledge based on the chapter structure of the document. Subsequently, through the collaborative diagnosis mechanism of multiple language models, each language model independently performs reasoning on complex diagnostic tasks and improves the accuracy of diagnosis through cross-examination. Finally, the present invention combines test-driven optimization, generates test points based on user feedback, and continuously iterates and optimizes the diagnostic results.

[0087] According to the database anomaly diagnosis method based on a large language model according to an embodiment of the present invention, for different abnormal alarms, the method realizes knowledge extraction through a hybrid retrieval strategy, uses multi-model collaborative reasoning to analyze the root cause of the anomaly, and continuously optimizes the diagnosis results in combination with user feedback. Ultimately, the present invention can accurately identify the cause of the anomaly, provide corresponding solutions, generate traceable diagnostic reports, and summarize user feedback to form improvement strategies, thereby achieving iterative optimization and self-evolution.

[0088] In order to implement the above embodiment, Figure 4 As shown, this embodiment also provides a database anomaly diagnosis system 10 based on a large language model, including:

[0089] A diagnosis knowledge base construction module 100 is used to split the diagnosis document by chapter to extract the diagnosis knowledge fragments to construct the diagnosis knowledge base;

[0090] The segment clustering module 200 is used to cluster the diagnostic knowledge segments according to preset themes and generate corresponding prompt word templates for each cluster;

[0091] The abnormal information collection module 300 is used to determine whether the indicator value monitored in the diagnostic knowledge base exceeds a preset threshold, and based on the judgment result of exceeding the preset threshold, match the relevant diagnostic knowledge from the diagnostic database according to the prompt word template to collect the corresponding abnormal information;

[0092] The abnormal result diagnosis module 400 is used to input the abnormal information into multiple large language models for parallel processing to obtain the corresponding database abnormality diagnosis results, and optimize the diagnosis path by integrating the diagnosis results of each large language model through a tree search algorithm.

[0093] Furthermore, the diagnostic knowledge base construction module 100 is also used for:

[0094] Extract preliminary knowledge fragments from existing diagnostic documents based on the document's chapter structure;

[0095] Diagnostic knowledge fragments are retrieved from the preliminary knowledge fragments through metadata filtering and similarity search.

[0096] Furthermore, the abnormal result diagnosis module 400 is also used for:

[0097] Identify potential types of abnormal information when receiving a diagnosis task, and assign the abnormal information to a related large language model agent based on the identified abnormal information type;

[0098] The large language model agent performs calculations through a tree search algorithm and shares the intermediate diagnosis results of the large language model agent. It also searches multiple reasoning paths and records the failed paths encountered to select the optimal path as the final diagnosis result.

[0099] The final diagnosis results of each large language model agent are aggregated to generate a comprehensive anomaly diagnosis report.

[0100] According to the database anomaly diagnosis system based on a large language model according to an embodiment of the present invention, for different anomaly alarms, the method realizes knowledge extraction through a hybrid retrieval strategy, uses multi-model collaborative reasoning to analyze the root cause of the anomaly, and continuously optimizes the diagnosis results in combination with user feedback. Ultimately, the present invention can accurately identify the cause of the anomaly, provide corresponding solutions, generate traceable diagnostic reports, and summarize user feedback to form improvement strategies, thereby achieving iterative optimization and self-evolution.

[0101] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0102] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

Claims

1. A database anomaly diagnosis method based on a large language model, characterized in that: include: Split the diagnostic documents by chapters to extract diagnostic knowledge fragments to build a diagnostic knowledge base; Clustering the diagnostic knowledge fragments according to preset themes, and generating corresponding prompt word templates for each cluster; Determine whether the indicator value monitored in the diagnostic knowledge base exceeds a preset threshold, and based on the judgment result of exceeding the preset threshold, match relevant diagnostic knowledge from the diagnostic database according to the prompt word template to collect corresponding abnormal information; The abnormal information is input into multiple large language models for parallel processing to obtain the corresponding database abnormality diagnosis result, and the diagnosis results of each large language model are integrated through a tree search algorithm to optimize the diagnosis path.

2. The method according to claim 1, characterized in that The diagnostic documents are split by chapters to extract diagnostic knowledge fragments, including: Extract preliminary knowledge fragments from existing diagnostic documents based on the document's chapter structure; Diagnostic knowledge fragments are retrieved from the preliminary knowledge fragments through metadata filtering and similarity search.

3. The method according to claim 1, characterized in that The abnormal information is input into multiple language models for parallel processing to obtain the corresponding database abnormality diagnosis results, and the diagnosis results of each large language model are integrated through a tree search algorithm to optimize the diagnosis path, including: Identify potential types of abnormal information when receiving a diagnosis task, and assign the abnormal information to a related large language model agent based on the identified abnormal information type; The large language model agent performs calculations through a tree search algorithm and shares the intermediate diagnosis results of the large language model agent. It also searches multiple reasoning paths and records the failed paths encountered to select the optimal path as the final diagnosis result. The final diagnosis results of each large language model agent are aggregated to generate a comprehensive anomaly diagnosis report.

4. The method according to claim 3, characterized in that When receiving a diagnosis task, it identifies the potential type of abnormal information and assigns the abnormal information to the relevant large language model agent based on the identified abnormal information type, including: Filter diagnostic knowledge fragments based on metadata, and then use vector similarity as an indicator to search for the knowledge fragment that is most similar to the abnormal description; The abnormality-related information and the most similar knowledge fragment retrieved are embedded into the prompt word template of the agent of the corresponding topic.

5. The method according to claim 3, characterized in that: After generating the comprehensive abnormality diagnosis report, the method further includes: Obtain user feedback information; Iteratively optimize the final diagnosis result according to user feedback information to obtain an optimized diagnosis result; A corresponding optimization model is generated according to the optimized diagnosis result, and the optimization model is added to a diagnosis knowledge base.

6. The method according to claim 5, characterized in that The final diagnosis result is iteratively optimized according to user feedback information to obtain the optimized diagnosis result, including: Use a large language model to extract key points from user feedback information and generate corresponding test programs for each key point; The test program is executed on the generated final diagnosis result, and is adjusted and optimized according to the test error log until all tests pass or a preset upper limit of the number of iterations is reached.

7. The method according to claim 5, characterized in that Generating a corresponding optimization mode according to the optimized diagnosis result, and adding the optimization mode to a diagnosis knowledge base, including: Comparing the diagnosis results before and after the optimization to obtain a comparison result; According to the comparison results and using a large language model, an effective optimization pattern is identified and extracted; The extracted optimization patterns are stored in a diagnostic knowledge base.

8. A database anomaly diagnosis system based on a large language model, characterized in that: include: A diagnostic knowledge base construction module is used to split diagnostic documents by chapters to extract diagnostic knowledge fragments to construct a diagnostic knowledge base; A fragment clustering module, used to cluster the diagnostic knowledge fragments according to preset themes and generate a corresponding prompt word template for each cluster; The abnormal information collection module is used to determine whether the indicator value monitored in the diagnostic knowledge base exceeds a preset threshold, and based on the judgment result of exceeding the preset threshold, match the relevant diagnostic knowledge from the diagnostic database according to the prompt word template to collect the corresponding abnormal information; The abnormal result diagnosis module is used to input the abnormal information into multiple language models for parallel processing to obtain the corresponding database abnormality diagnosis results, and optimize the diagnosis path by integrating the diagnosis results of each large language model through a tree search algorithm.

9. The system according to claim 8, characterized in that Diagnostic knowledge base building blocks are also used to: Extract preliminary knowledge fragments from existing diagnostic documents based on the document's chapter structure; Diagnostic knowledge fragments are retrieved from the preliminary knowledge fragments through metadata filtering and similarity search.

10. The system according to claim 8, characterized in that The abnormal result diagnosis module is also used for: Identify potential types of abnormal information when receiving a diagnosis task, and assign the abnormal information to a related large language model agent based on the identified abnormal information type; The large language model agent performs calculations through a tree search algorithm and shares the intermediate diagnosis results of the large language model agent. It also searches multiple reasoning paths and records the failed paths encountered to select the optimal path as the final diagnosis result. The final diagnosis results of each large language model agent are aggregated to generate a comprehensive anomaly diagnosis report.

Citation Information

Patent Citations

  • Method for constructing equipment fault diagnosis and maintenance knowledge base based on large language model

    CN116861189A

  • Fault diagnosis method and device, equipment, storage medium and program product

    CN118819941A

  • Database operation and maintenance method and system and storage medium

    CN119166610A

Cited By

  • Slow query optimization method and device

    CN120994699A

  • Dynamic troubleshooting problem processing method and device, electronic equipment and readable storage medium

    CN121681206A

  • Dynamic problem troubleshooting method and device, electronic equipment and readable storage medium

    CN121681206B