Distributed database fault positioning method based on operation and maintenance knowledge graph

By building an operation and maintenance knowledge graph and combining intelligent models and automation analysis tools, the problems of complex technical components and complex business indicators in the operation and maintenance of financial institutions' information systems are solved, and efficient fault location and operation and maintenance automation are achieved.

CN119988348APending Publication Date: 2025-05-13SHANGHAI RURAL COMML BANK
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411993363.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Financial institutions' information system operation and maintenance face the problems of complex technical components, complex business indicators, and operation and maintenance knowledge that are difficult to digitize and automate, resulting in high cost and low efficiency of fault location.

Method used

The distributed database fault location method based on the operation and maintenance knowledge graph is adopted, and the automatic application and intelligent operation and maintenance knowledge are realized by building the operation and maintenance knowledge graph, combining intelligent models and automated analysis tools.

Benefits of technology

It reduces the work difficulty and fault positioning costs of operation and maintenance personnel, improves operation and maintenance efficiency and fault positioning accuracy, and is suitable for dealing with complex and unknown problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988348A_ABST
    Figure CN119988348A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of databases, and discloses a distributed database fault positioning method based on an operation and maintenance knowledge graph. Through operation and maintenance experiences settled for many years by operation and maintenance experts and an operation and maintenance knowledge graph constructed by operation and maintenance cases continuously accumulated by users in a production environment, an automatic analysis tool and an intelligent algorithm are combined, so that previous operation and maintenance experiences are settled; operation and maintenance personnel do not need to pay attention to threshold values and actual changes of certain baseline indexes, can monitor changes of related baseline indexes or index groups through operation and maintenance experience, and once an alarm is triggered by certain operation and maintenance experience, related analysis is carried out, and the operation and maintenance personnel do not need to pay attention to a large number of unfamiliar indexes and only need to pay attention to a small number of known problems. And the working difficulty and the fault positioning cost are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of database technology, and in particular to a distributed database fault locating method based on an operation and maintenance knowledge graph. Background Art

[0002] The operation and maintenance of information systems is facing the pressure of digital transformation. With the rapid development of information systems, the scale and complexity of systems have reached unprecedented heights. The traditional model of relying on experts and a large amount of manpower will no longer be sustainable. It is imperative to adapt to the new IT scale development through digital transformation of operation and maintenance.

[0003] At present, financial institutions face several major difficulties in the digital transformation of information system operation and maintenance:

[0004] 1. The various technical components of the application system are increasing in number and complexity, and the various business indicators are too numerous and difficult to understand, let alone to collect accurately;

[0005] 2. With the increase in business volume and data volume, a large amount of operation and maintenance experience and knowledge is difficult to digitize, and even more difficult to automate. It is technically difficult and costly for a single enterprise to complete this work;

[0006] 3. The R&D team lacks real operation and maintenance experts, making it difficult to develop operation and maintenance automation products that can replace humans.

[0007] Therefore, we need to propose a distributed database fault location method based on the operation and maintenance knowledge graph. A system with operation and maintenance knowledge, expert experience and intelligent models as the core is more suitable for dealing with unknown and complex problems. Summary of the invention

[0008] The purpose of the present invention is to provide a distributed database fault location method based on an operation and maintenance knowledge graph, build intelligent operation and maintenance capabilities with knowledge automation as the core, and build an operation and maintenance knowledge graph through the operation and maintenance experience accumulated by operation and maintenance experts over the years and the operation and maintenance cases continuously accumulated by users in the production environment, combined with automated analysis tools and intelligent algorithms, so that previous operation and maintenance experience can be accumulated; operation and maintenance personnel do not need to pay attention to the thresholds and actual changes of certain baseline indicators, and can monitor the changes of related baseline indicators or indicator groups through operation and maintenance experience. Once an operation and maintenance experience alarm is triggered, relevant analysis can be performed. There is no need to pay attention to a large number of unfamiliar indicators, but only a small number of known problems, which reduces the work difficulty and fault location cost, so as to solve the problems raised in the above background technology.

[0009] To achieve the above object, the present invention provides the following technical solution: a distributed database fault location method based on operation and maintenance knowledge graph, comprising the following steps:

[0010] S1. Provide database support: optimize and adapt distributed databases, and effectively monitor, analyze and handle database issues;

[0011] S2. Create intelligent models: Intelligent models include health models, performance models, and load models. They keep track of the database's operating status and the current load and performance status of the database unified management system. They also use AI models to predict model indicators for the next three time periods, providing early status warnings for operation and maintenance personnel.

[0012] S3. Create a fault model: monitor the status of the unified database management system through preset expert experience and user-defined operation and maintenance experience. When a fault is found, an alarm is triggered in time, the cause of the problem is analyzed, and an analysis report is provided;

[0013] S4. In-depth log diagnosis: The unified database management system filters out log information with problems and doubts from the logs and displays them to the operation and maintenance personnel. It also automatically conducts in-depth analysis of the logs through the database log, system log, and application log analysis knowledge point tools accumulated in the system to discover possible system failures and hidden dangers, and generate system alarms at the same time.

[0014] S5. Automated inspection: Regularly perform automated inspection and testing on the database to keep it in good operating condition, detect potential problems in time, and avoid them from turning into serious failures. It also automatically generates daily inspection reports, monthly inspection reports, capacity forecast reports, and SQL audit reports.

[0015] S6. Key SQL tracking: Real-time monitoring and audit analysis of SQL statements that are frequently executed in the database and have a great impact on performance, to understand the performance bottleneck of the database and to provide early warning of potential problem points and risks;

[0016] S7. Performance optimization: Utilize the optimization analysis tools provided by the unified database management system and the results of the system intelligent evaluation to conduct targeted diagnosis and analysis of the database operation data, identify the performance risks of the database, and adjust and optimize the configuration, structure, and index of the database;

[0017] S8. Build an operation and maintenance knowledge graph: Integrate the knowledge of database architecture, components, configuration, and troubleshooting as a basic support tool for intelligent diagnostic tools, pan-routing knowledge points, waiting event analysis, and deep log analysis, providing operation and maintenance personnel with convenient knowledge query and reasoning tools.

[0018] Preferably, in step S1, the distributed database has a unique architecture and characteristics, and the database supports the underlying storage engine, transaction processing mechanism, concurrency control strategy, common failure modes and performance bottlenecks of the database, and builds specialized monitoring and diagnostic tools to achieve effective monitoring and fault location of the database.

[0019] Preferably, in step S2, the workflow of the intelligent model is as follows:

[0020] S21. Data collection: collect operation data from distributed databases. The operation data includes performance indicators, log information, and user behavior.

[0021] S22, feature extraction: extracting features from the collected data to extract features that have a significant impact on the database operation status;

[0022] S23, model training: using machine learning algorithms to train the extracted features and build intelligent models. During the training process, the model learns the normal operation mode and abnormal operation mode of the database;

[0023] S24. Prediction and early warning: The intelligent model predicts the future state of the database based on real-time operation data. When the error between the prediction result and the actual operation data is greater than the threshold, the intelligent model triggers an alarm.

[0024] Preferably, in step S3, the workflow of the fault model is as follows:

[0025] S31. Fault data collection: Collect fault data from historical fault records. The fault data includes fault type, fault occurrence time, fault impact scope, and fault solution;

[0026] S32, fault scenario construction: construct multiple fault scenarios and corresponding fault manifestations based on the collected fault data;

[0027] S33, Solution Arrangement: Arrange solutions and preventive measures for each failure scenario;

[0028] S34, Matching and recommendation: When a problem occurs in the database, the fault model will be matched with the current symptoms to quickly locate the cause of the fault and provide an analysis report.

[0029] Preferably, in step S4, the workflow of log in-depth diagnosis is as follows:

[0030] S41, log collection: collect log files from the database, the log files include system logs, application logs, and error logs;

[0031] S42, Log filtering: Use natural language processing technology to automatically parse and classify log information, extract key error information and abnormal behavior;

[0032] S43, anomaly detection: perform anomaly detection on the parsed log information, analyze the abnormal patterns and error trends in the log information, and discover potential faults and problems;

[0033] S44, Problem location: Locate the root cause of the problem based on the anomaly detection results, trigger the alarm mechanism, and notify the operation and maintenance personnel.

[0034] Preferably, in step S5, the workflow of the automated inspection is as follows:

[0035] S51. Formulate inspection plan: set the time, frequency and content of database inspection;

[0036] S52, perform inspection tasks: perform inspection tasks according to the inspection plan, the inspection tasks include database performance indicator inspection, security configuration inspection, and disk space inspection;

[0037] S53, Inspection screening: During the inspection process, when potential problems or abnormalities are found, an inspection report is generated and sent to the operation and maintenance personnel;

[0038] S54. Problem handling and tracking: Operation and maintenance personnel handle and track the problems in the inspection reports and solve the problems and risks in a timely manner.

[0039] Preferably, in step S6, the workflow of key SQL tracking is as follows:

[0040] S61, identifying key SQL statements, and identifying SQL statements that are frequently executed and have poor performance through query logs and performance analysis tools;

[0041] S62, Real-time monitoring: Enable performance monitoring and logging functions to collect SQL execution data in real time;

[0042] S63, Audit analysis: parse the collected SQL statements, evaluate the performance of SQL statements, identify performance bottlenecks and potential risks, compare current performance with historical data, and analyze performance change trends;

[0043] S64. Risk warning: Based on the analysis results, set performance warning rules and use monitoring tools to issue real-time warnings.

[0044] Preferably, in step S7, the workflow of performance optimization is as follows:

[0045] S71. Performance bottleneck identification: Identify database performance bottlenecks by analyzing database operation data;

[0046] S72. Formulate optimization plan: Formulate optimization plan according to performance bottlenecks. The optimization plan includes adjusting database configuration, optimizing table structure, and adding indexes.

[0047] S73. Implement and test the optimization plan: implement the optimization plan and test it until the optimization effect reaches the expected result;

[0048] S74. Continuous optimization and monitoring: Continuously monitor and optimize the database's performance to keep it in the best operating state.

[0049] Preferably, in step S8, the workflow for constructing the operation and maintenance knowledge graph is as follows:

[0050] S81. Knowledge collection and organization: Collect database operation and maintenance knowledge from various sources, organize and classify it to form a structured knowledge base;

[0051] S82. Knowledge graph construction: Use graph database technology to build an operation and maintenance knowledge graph, and represent and store the collected knowledge in the form of a graph;

[0052] S83, Knowledge query and reasoning: Operation and maintenance personnel use the operation and maintenance knowledge graph to perform knowledge query and reasoning, and make inferences and predictions based on known information to provide accurate fault location and solutions;

[0053] S84. Knowledge updating and maintenance: adding new knowledge entities and relationships, and deleting outdated or erroneous knowledge.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] 1. The present invention builds intelligent operation and maintenance capabilities with knowledge automation as the core. It uses the operation and maintenance experience accumulated by operation and maintenance experts over the years and the operation and maintenance cases accumulated by users in the production environment to build an operation and maintenance knowledge map, combined with automated analysis tools and intelligent algorithms, to accumulate previous operation and maintenance experience;

[0056] 2. The operation and maintenance personnel of the present invention do not need to pay attention to the thresholds and actual changes of certain baseline indicators. They can monitor the changes of related baseline indicators or indicator groups through operation and maintenance experience. Once a certain operation and maintenance experience alarm is triggered, they can conduct relevant analysis. They do not need to pay attention to a large number of unfamiliar indicators, but only need to pay attention to a small number of known problems, which reduces the work difficulty and fault location cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a flowchart of the present invention;

[0058] Figure 2 It is a flowchart of the workflow of the intelligent model of the present invention;

[0059] Figure 3 This is a flowchart of the fault model of the present invention.

[0060] Figure 4 This is a flowchart of the workflow of the log in-depth diagnosis of the present invention;

[0061] Figure 5 This is a flowchart of the automated inspection process of the present invention;

[0062] Figure 6 This is a flowchart of the key SQL tracking process of the present invention.

[0063] Figure 7 A flowchart of the performance optimization of the present invention;

[0064] Figure 8 A workflow flowchart for constructing an operation and maintenance knowledge graph for the present invention. DETAILED DESCRIPTION

[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] See also Figure 1-8 , the present invention provides a technical solution: a distributed database fault location method based on operation and maintenance knowledge graph, which uses the existing operation and maintenance knowledge base to build an operation and maintenance knowledge graph, combines artificial intelligence, big data analysis and other technologies to process and analyze various types of operation and maintenance data such as data monitoring of information systems, and combines the operation and maintenance knowledge of operation and maintenance experts and actual fault cases to form a set of deep operation and maintenance tools with operation and maintenance knowledge as the core capability, which is used to assist manual operation and maintenance, gradually realize the purpose of replacing manual work with tools, and help the digital transformation of information system operation and maintenance;

[0067] The following steps are involved:

[0068] S1. Provide database support: optimize and adapt distributed databases, and effectively monitor, analyze and handle database issues;

[0069] In step S1, the distributed database has a unique architecture and characteristics. The database supports the underlying storage engine, transaction processing mechanism, concurrency control strategy, common failure modes and performance bottlenecks of the database, and builds specialized monitoring and diagnostic tools to achieve effective monitoring and fault location of the database.

[0070] S2. Create intelligent models: Intelligent models include health models, performance models, and load models. They keep track of the database's operating status and the current load and performance status of the database unified management system. They also use AI models to predict model indicators for the next three time periods, providing early status warnings for operation and maintenance personnel.

[0071] In step S2, the workflow of the intelligent model is as follows:

[0072] S21. Data collection: collect operation data from distributed databases. The operation data includes performance indicators, log information, and user behavior.

[0073] S22, feature extraction: extracting features from the collected data to extract features that have a significant impact on the database operation status;

[0074] S23, model training: using machine learning algorithms to train the extracted features and build intelligent models. During the training process, the model learns the normal operation mode and abnormal operation mode of the database;

[0075] S24. Prediction and early warning: The intelligent model predicts the future state of the database based on real-time operation data. When the error between the prediction result and the actual operation data is greater than the threshold, the intelligent model triggers an alarm.

[0076] Health model: It is a key component of the unified database management system and is used to evaluate the overall health of the database. The model ensures the stability and reliability of the database by monitoring key indicators of the database, such as data integrity, data consistency, and system error rate. When potential problems occur in the database, the health model can detect and issue early warnings in a timely manner, thereby helping operation and maintenance personnel to quickly locate and solve the problem and prevent the problem from further deteriorating.

[0077] Performance model: used to evaluate the operating performance of the database, including key indicators such as query speed, response time, and throughput. The model comprehensively tests and analyzes the performance of the database by simulating the actual operating scenario of the database. The performance model can help operation and maintenance personnel understand the performance of the database under different loads, thereby optimizing database configuration and query statements and improving the operating efficiency of the database.

[0078] Load model: used to evaluate the current load of the database, including the use of key resources such as CPU usage, memory usage, and disk I / O. This model helps operation and maintenance personnel understand the load status of the database by monitoring the resource usage of the database in real time, so as to perform reasonable resource allocation and load balancing operations. The load model can prevent problems such as database performance degradation or crash caused by resource overload.

[0079] AI model: Use machine learning algorithms to learn and analyze the historical data of the database to predict the future operation status and performance indicators of the database. The AI ​​model can predict model indicators in the next three time intervals, such as health index, performance score, load status, etc., to provide early status warnings for operation and maintenance personnel.

[0080] S3. Create a fault model: monitor the status of the unified database management system through preset expert experience and user-defined operation and maintenance experience. When a fault is found, an alarm is triggered in time, the cause of the problem is analyzed, and an analysis report is provided;

[0081] In step S3, the workflow of the fault model is as follows:

[0082] S31. Fault data collection: Collect fault data from historical fault records. The fault data includes fault type, fault occurrence time, fault impact scope, and fault solution;

[0083] Fault type: describes the specific nature of the fault, such as hardware fault, software fault, network fault, etc.

[0084] Fault occurrence time: Recording the exact time when the fault occurs helps to analyze the frequency and periodicity of the fault;

[0085] Fault impact scope: describes the impact of the fault on the database system, application, or user;

[0086] Troubleshooting solution: Record the specific steps and methods for solving the problem, including the repair process, tools and techniques used, etc.

[0087] S32, fault scenario construction: construct multiple fault scenarios and corresponding fault manifestations based on the collected fault data;

[0088] Fault scenarios and corresponding fault manifestations include:

[0089] Fault trigger conditions: describe the specific conditions or events that cause the fault to occur;

[0090] Fault symptoms: List the abnormal performance that may occur in the system when the fault occurs, such as performance degradation, data loss, service interruption, etc.;

[0091] Failure impact analysis: Assess the potential impact of a failure on system stability, data integrity, and user experience.

[0092] S33, Solution Arrangement: Arrange solutions and preventive measures for each failure scenario;

[0093] S34, Matching and recommendation: When a problem occurs in the database, the fault model will be matched with the current symptoms to quickly locate the cause of the fault and provide an analysis report.

[0094] Symptom analysis: Collect abnormal performance of the current system and compare it with the symptoms in the fault scenario;

[0095] Fault location: Determine the most likely cause of the fault based on the results of symptom matching;

[0096] Analysis report generation: Provides a detailed report including fault location, impact analysis, and recommended solutions.

[0097] S4. In-depth log diagnosis: The unified database management system filters out log information with problems and doubts from the logs and displays them to the operation and maintenance personnel. It also automatically conducts in-depth analysis of the logs through the database log, system log, and application log analysis knowledge point tools accumulated in the system to discover possible system failures and hidden dangers, and generate system alarms at the same time.

[0098] In step S4, the workflow of log deep diagnosis is as follows:

[0099] S41. Log collection: Collect log files from the database. The log files include system logs, application logs, and error logs. The log information records key information such as the operation status of the database system, user operations, and system anomalies.

[0100] S42. Log filtering: Use natural language processing technology to automatically parse and classify log information, extract key error information and abnormal behavior; filtering rules can be set based on multiple dimensions such as log content, log level, and log source.

[0101] S43, anomaly detection: perform anomaly detection on the parsed log information, analyze the abnormal patterns and error trends in the log information, and discover potential faults and problems;

[0102] S44, Problem Location: Based on the abnormal detection results, locate the root cause of the problem, trigger the alarm mechanism, and notify the operation and maintenance personnel. The alarm information will notify the operation and maintenance personnel in a variety of ways, such as console prompts, email notifications, SMS reminders, etc. The alarm information will contain key information such as the specific description of the fault, possible causes, and recommended solutions.

[0103] S5. Automated inspection: Regularly perform automated inspection and testing on the database to keep it in good operating condition, detect potential problems in time, and avoid them from turning into serious failures. It also automatically generates daily inspection reports, monthly inspection reports, capacity forecast reports, and SQL audit reports.

[0104] The contents of automated inspection include:

[0105] Database performance check: Check the database's CPU usage, memory usage, disk I / O and other key performance indicators, analyze the database's query efficiency, and identify SQL statements with long execution times for optimization;

[0106] Database integrity check: Verify the data integrity of the database, including the integrity of the data table and index, and check the database backup and recovery strategy to ensure that the data can be quickly restored when lost or damaged;

[0107] Database security check: Check the database security settings, such as user permissions, encrypted connections, etc., and monitor database security events, such as unauthorized access, data leakage, etc.;

[0108] Database configuration check: Check the database configuration parameters, such as the number of connections, buffer pool size, etc., to ensure that they match the current system resources and business requirements.

[0109] In step S5, the workflow of the automated inspection is as follows:

[0110] S51. Formulate inspection plan: set the time, frequency and content of database inspection;

[0111] S52, perform inspection tasks: perform inspection tasks according to the inspection plan, the inspection tasks include database performance indicator inspection, security configuration inspection, and disk space inspection;

[0112] S53, Inspection screening: During the inspection process, when potential problems or abnormalities are found, an inspection report is generated and sent to the operation and maintenance personnel;

[0113] S54. Problem handling and tracking: Operation and maintenance personnel handle and track the problems in the inspection reports and solve the problems and risks in a timely manner.

[0114] S6. Key SQL tracking: Real-time monitoring and audit analysis of SQL statements that are frequently executed in the database and have a great impact on performance, to understand the performance bottleneck of the database and to provide early warning of potential problem points and risks;

[0115] In step S6, the workflow of key SQL tracking is as follows:

[0116] S61. Identify key SQL statements. Use query logs and performance analysis tools (such as MySQL's PerformanceSchema and Oracle's AWR reports) to identify SQL statements that are frequently executed and have poor performance. Set thresholds for key indicators such as execution time and execution frequency based on business needs to screen out SQL statements that require special attention.

[0117] S62, Real-time monitoring: Enable performance monitoring and logging functions to collect SQL execution data in real time;

[0118] S63, Audit Analysis: Parse the collected SQL statements, analyze their execution plans, access paths, index usage, etc., evaluate the performance of SQL statements, identify performance bottlenecks and potential risks, compare current performance with historical data, analyze performance change trends, and find out the reasons for performance degradation;

[0119] S64. Risk warning: Based on the analysis results, set performance warning rules. For example, when the execution time of a SQL statement exceeds the set threshold, trigger a warning. Use monitoring tools to issue real-time warnings. When the warning conditions are met, notify relevant personnel via email, text message, etc.

[0120] Optionally, review key SQLs regularly to ensure the effectiveness of optimization measures and identify new performance issues. Upgrade and optimize the database system in a timely manner according to business development needs and technology development trends. Use automated monitoring tools to achieve real-time monitoring and early warning of key SQLs, reduce manual intervention, and use machine learning and other technologies to perform intelligent analysis of SQL execution data to improve the efficiency of problem discovery and resolution.

[0121] S7. Performance optimization: Utilize the optimization analysis tools provided by the unified database management system and the results of the system intelligent evaluation to conduct targeted diagnosis and analysis of the database operation data, identify the performance risks of the database, and adjust and optimize the configuration, structure, and index of the database;

[0122] Database operation data includes key indicators such as CPU usage, memory usage, disk I / O, and network throughput, and identifies SQL statements with long execution time and high resource consumption;

[0123] With the help of the optimization analysis tools provided by the database (such as Oracle's SQL Tuning Advisor and MySQL's EXPLAIN), the execution plan of SQL statements is analyzed to find performance bottlenecks, and the overall performance of the database is evaluated to obtain optimization suggestions.

[0124] Based on the collected performance data and the results of the optimization analysis tools, the database's operating status is diagnosed and analyzed to identify performance risks such as improper database configuration, unreasonable structure, and index failure. Based on the diagnostic analysis results, the database's memory allocation, buffer size, connection pool size and other configuration parameters are adjusted to adapt to different workloads, and the database's locking mechanism, query cache and other configurations are optimized to improve concurrent processing capabilities and query efficiency.

[0125] In addition, the database table structure should be reasonably designed to avoid redundant data and complex joint queries, and large tables should be partitioned. Horizontal partitioning or vertical partitioning should be selected according to business needs to improve query performance, avoid excessive or unnecessary indexes to reduce storage space usage and maintenance costs, and regularly optimize and rebuild indexes to prevent index fragmentation and improve index query efficiency.

[0126] In step S7, the workflow of performance optimization is as follows:

[0127] S71. Performance bottleneck identification: Identify database performance bottlenecks by analyzing database operation data;

[0128] S72. Formulate optimization plan: Formulate optimization plan according to performance bottlenecks. The optimization plan includes adjusting database configuration, optimizing table structure, and adding indexes.

[0129] S73. Implement and test the optimization plan: implement the optimization plan and test it until the optimization effect reaches the expected result;

[0130] S74. Continuous optimization and monitoring: Continuously monitor and optimize the database's performance to keep it in the best operating state.

[0131] S8. Build an operation and maintenance knowledge graph: Integrate the knowledge of database architecture, components, configuration, and troubleshooting as a basic support tool for intelligent diagnostic tools, pan-routing knowledge points, waiting event analysis, and deep log analysis, providing operation and maintenance personnel with convenient knowledge query and reasoning tools.

[0132] The elements of the operation and maintenance knowledge graph include:

[0133] Entity: represents objects or concepts in the real world, such as databases, servers, network equipment, operation and maintenance tools, etc.

[0134] Attributes: describe the characteristics or properties of an entity, such as the type, version, and capacity of a database; configuration information such as the server's CPU, memory, and hard disk;

[0135] Relationship: represents the connection or association between entities, such as the connection relationship between a database and a server, the communication relationship between a server and a network device, etc.

[0136] Semantic type: describes the category or type of an entity or attribute, which helps to better organize and classify knowledge;

[0137] Metadata: Provides information about the knowledge graph itself and its content, such as the creation time of the knowledge graph, update frequency, data source, etc.

[0138] In step S8, the workflow for constructing the operation and maintenance knowledge graph is as follows:

[0139] S81. Knowledge collection and organization: Collect database operation and maintenance knowledge from various sources, organize and classify it to form a structured knowledge base;

[0140] S82. Knowledge graph construction: Use graph database technology to build an operation and maintenance knowledge graph, and represent and store the collected knowledge in the form of a graph;

[0141] S83, Knowledge query and reasoning: Operation and maintenance personnel use the operation and maintenance knowledge graph to perform knowledge query and reasoning, and make inferences and predictions based on known information to provide accurate fault location and solutions;

[0142] S84. Knowledge updating and maintenance: adding new knowledge entities and relationships, and deleting outdated or erroneous knowledge.

[0143] Intelligent diagnosis tool: Use the fault handling knowledge in the knowledge graph to perform intelligent diagnosis on faults that occur during the operation and maintenance process, and improve the accuracy and efficiency of fault diagnosis by associating relevant fault data in the fault knowledge base;

[0144] Pan-routing knowledge points: Integrate knowledge on network equipment configuration, performance, troubleshooting, etc., and provide comprehensive network equipment knowledge query and reasoning tools for operation and maintenance personnel;

[0145] Waiting event analysis: Use knowledge graphs to analyze the causes and influencing factors of system waiting events, optimize system performance and provide decision support;

[0146] Deep log analysis: Combine the entity and relationship information in the knowledge graph to conduct in-depth analysis of log files and explore potential patterns and failure modes in log files.

[0147] This method mainly provides intelligent operation and maintenance capabilities with knowledge automation as the core. Unlike the current mainstream AIOPS concept, the theoretical basis of the unified database management system is not a complex data analysis algorithm, but an operation and maintenance knowledge graph built by operation and maintenance experts' years of operation and maintenance experience and users' continuous accumulation of operation and maintenance cases in the production environment. Based on the knowledge graph and knowledge reasoning, the basic framework of the unified database management system is constructed. The introduction of automated analysis tools and intelligent algorithms has allowed the previous operation and maintenance experience to accumulate. The algorithmic intelligence composed of the operation and maintenance knowledge graph, machine learning, and deep learning algorithms forms the dual core functions of the unified database management system.

[0148] In a unified database management system, under normal circumstances, operation and maintenance personnel do not need to pay attention to the thresholds and actual changes of certain baseline indicators. They can use operation and maintenance experience to monitor the changes of related baseline indicators or indicator groups, and once an operation and maintenance experience alarm is triggered, they can conduct relevant analysis. Operation and maintenance experience allows operation and maintenance personnel to not need to pay attention to hundreds of unfamiliar indicators, but only need to pay attention to dozens of known problems.

[0149] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A distributed database fault location method based on operation and maintenance knowledge graph, characterized in that: The following steps are involved: S1. Provide database support: optimize and adapt distributed databases, and effectively monitor, analyze and handle database issues; S2. Create intelligent models: Intelligent models include health models, performance models, and load models. They keep track of the database's operating status and the current load and performance status of the database unified management system. They also use AI models to predict model indicators for the next three time periods, providing early status warnings for operation and maintenance personnel. S3. Create a fault model: monitor the status of the unified database management system through preset expert experience and user-defined operation and maintenance experience. When a fault is found, an alarm is triggered in time, the cause of the problem is analyzed, and an analysis report is provided; S4. In-depth log diagnosis: The unified database management system filters out log information with problems and doubts from the logs and displays them to the operation and maintenance personnel. It also automatically conducts in-depth analysis of the logs through the database log, system log, and application log analysis knowledge point tools accumulated in the system to discover possible system failures and hidden dangers, and generate system alarms at the same time. S5. Automated inspection: Regularly perform automated inspection and testing on the database to keep it in good operating condition, detect potential problems in time, and avoid them from turning into serious failures. It also automatically generates daily inspection reports, monthly inspection reports, capacity forecast reports, and SQL audit reports. S6. Key SQL tracking: Real-time monitoring and audit analysis of SQL statements that are frequently executed in the database and have a great impact on performance, to understand the performance bottlenecks of the database and to provide early warning of potential problem points and risks; S7. Performance optimization: Utilize the optimization analysis tools provided by the unified database management system and the results of the system intelligent evaluation to conduct targeted diagnosis and analysis of the database operation data, identify the performance risks of the database, and adjust and optimize the configuration, structure, and index of the database; S8. Build an operation and maintenance knowledge graph: Integrate the knowledge of database architecture, components, configuration, and troubleshooting as a basic support tool for intelligent diagnostic tools, pan-routing knowledge points, waiting event analysis, and deep log analysis, providing operation and maintenance personnel with convenient knowledge query and reasoning tools.

2. According to claim 1, a distributed database fault location method based on operation and maintenance knowledge graph is characterized in that: In step S1, the distributed database has a unique architecture and characteristics. The database supports the underlying storage engine, transaction processing mechanism, concurrency control strategy, common failure modes and performance bottlenecks of the database, and builds specialized monitoring and diagnostic tools to achieve effective monitoring and fault location of the database.

3. According to claim 1, a distributed database fault location method based on operation and maintenance knowledge graph is characterized in that: In step S2, the workflow of the intelligent model is as follows: S21. Data collection: collect operation data from distributed databases. The operation data includes performance indicators, log information, and user behavior. S22, feature extraction: extracting features from the collected data to extract features that have a significant impact on the database operation status; S23, model training: using machine learning algorithms to train the extracted features and build intelligent models. During the training process, the model learns the normal operation mode and abnormal operation mode of the database; S24. Prediction and early warning: The intelligent model predicts the future state of the database based on real-time operation data. When the error between the prediction result and the actual operation data is greater than the threshold, the intelligent model triggers an alarm.

4. According to claim 1, a distributed database fault location method based on operation and maintenance knowledge graph is characterized in that: In step S3, the workflow of the fault model is as follows: S31. Fault data collection: Collect fault data from historical fault records. The fault data includes fault type, fault occurrence time, fault impact scope, and fault solution; S32, fault scenario construction: construct multiple fault scenarios and corresponding fault manifestations based on the collected fault data; S33, Solution Arrangement: Arrange solutions and preventive measures for each failure scenario; S34, Matching and recommendation: When a problem occurs in the database, the fault model will be matched with the current symptoms to quickly locate the cause of the fault and provide an analysis report.

5. According to claim 1, a distributed database fault location method based on operation and maintenance knowledge graph is characterized in that: In step S4, the workflow of log deep diagnosis is as follows: S41, log collection: collect log files from the database, the log files include system logs, application logs, and error logs; S42, Log filtering: Use natural language processing technology to automatically parse and classify log information, extract key error information and abnormal behavior; S43, anomaly detection: perform anomaly detection on the parsed log information, analyze the abnormal patterns and error trends in the log information, and discover potential faults and problems; S44, Problem location: Locate the root cause of the problem based on the anomaly detection results, trigger the alarm mechanism, and notify the operation and maintenance personnel.

6. According to claim 1, a distributed database fault location method based on operation and maintenance knowledge graph is characterized by: In step S5, the workflow of the automated inspection is as follows: S51. Formulate inspection plan: set the time, frequency and content of database inspection; S52, perform inspection tasks: perform inspection tasks according to the inspection plan, the inspection tasks include database performance indicator inspection, security configuration inspection, and disk space inspection; S53, Inspection screening: During the inspection process, when potential problems or abnormalities are found, an inspection report is generated and sent to the operation and maintenance personnel; S54. Problem handling and tracking: Operation and maintenance personnel handle and track the problems in the inspection reports and solve the problems and risks in a timely manner.

7. The method for locating faults in a distributed database based on an operation and maintenance knowledge graph according to claim 1 is characterized in that: In step S6, the workflow of key SQL tracking is as follows: S61, identifying key SQL statements, and identifying SQL statements that are frequently executed and have poor performance through query logs and performance analysis tools; S62, Real-time monitoring: Enable performance monitoring and logging functions to collect SQL execution data in real time; S63, Audit analysis: parse the collected SQL statements, evaluate the performance of SQL statements, identify performance bottlenecks and potential risks, compare current performance with historical data, and analyze performance change trends; S64. Risk warning: Based on the analysis results, set performance warning rules and use monitoring tools to issue real-time warnings.

8. The method for locating faults in a distributed database based on an operation and maintenance knowledge graph according to claim 1, characterized in that: In step S7, the workflow of performance optimization is as follows: S71. Performance bottleneck identification: Identify database performance bottlenecks by analyzing database operation data; S72. Formulate optimization plan: Formulate optimization plan according to performance bottlenecks. The optimization plan includes adjusting database configuration, optimizing table structure, and adding indexes. S73. Implement and test the optimization plan: implement the optimization plan and test it until the optimization effect reaches the expected result; S74. Continuous optimization and monitoring: Continuously monitor and optimize the database's performance to keep it in the best operating state.

9. The method for locating faults in a distributed database based on an operation and maintenance knowledge graph according to claim 1, characterized in that: In step S8, the workflow for constructing the operation and maintenance knowledge graph is as follows: S81. Knowledge collection and organization: Collect database operation and maintenance knowledge from various sources, organize and classify it to form a structured knowledge base; S82. Knowledge graph construction: Use graph database technology to build an operation and maintenance knowledge graph, and represent and store the collected knowledge in the form of a graph; S83, Knowledge query and reasoning: Operation and maintenance personnel use the operation and maintenance knowledge graph to perform knowledge query and reasoning, and make inferences and predictions based on known information to provide accurate fault location and solutions; S84. Knowledge updating and maintenance: adding new knowledge entities and relationships, and deleting outdated or erroneous knowledge.

Citation Information

Cited By

  • Database autonomous service processing method and device based on machine learning, and terminal

    CN120994644A