Database parameter anomaly handling
Patent Information
- Application Number
- PCT/IB2024/062967
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-17
AI Technical Summary
Due to its complexity and the number of large-scale database instances, the management of database parameter exceptions becomes extremely difficult, making it difficult to efficiently discover and repair abnormal database parameters.
A database parameter exception management system is designed, including multiple data collection agents, real-time message queues, online analysis services, inspection services and parameter repair agents. The system collects and stores database parameter data in real time, conducts online analysis and decision-making, and realizes real-time repair of abnormal parameters.
The system can efficiently and accurately detect abnormal database parameters, greatly shorten the survival time of abnormal parameters, improve the stability of database instances, and reduce the difficulty of operating and maintaining database parameters.
Smart Images

Figure IB2024062967_17072025_PF_FP_ABST
Abstract
Description
Technical Field of Database Parameter Anomaly Governance
[0001] This application relates to the field of database technologies, and in particular to database parameter anomaly governance. Background Art
[0002] A cloud database system refers to a database system optimized or deployed in a virtual computer environment, which can achieve the advantages of pay-on-demand, scale-on-demand, high availability, and storage consolidation. Currently, the market has increasing requirements for the characteristics of cloud database systems. To meet various user needs, cloud database systems are becoming more and more complex, and performance adaptation needs to be carried out through parameter adjustment for different user application scenarios.
[0003] In the daily operation and maintenance of cloud database systems, promptly discovering and fixing abnormal database parameters is an important operation and maintenance task. With the rapid development of cloud database systems, some cloud database systems have more than 1 million database instances, and the number of parameters of database instances reaches more than 700. In addition, users also have the need for independent optimization of database parameters. The distribution process of independently optimized database parameters involves multiple parameter synchronizations, which easily leads to abnormal database behavior. The superposition of multiple factors greatly increases the operation and maintenance difficulty of database parameters and poses a greater challenge to the database parameter anomaly governance solution. Summary of the Invention
[0004] Multiple aspects of this application provide a database parameter anomaly governance system, method, electronic device, and storage medium to provide a better database parameter anomaly governance solution.
[0005] An embodiment of the present application provides a database parameter anomaly governance system, including: a plurality of data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly patrol service, and a plurality of parameter repair agents deployed on multiple database instances; the data collection agents are used to collect real-time parameter data of the database instances in real time and report it to the first real-time message queue for storage; the parameter anomaly online analysis service is used to obtain the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and take a snapshot to obtain a real-time parameter snapshot at the most recent collection time; call a big data processing component to analyze parameter anomaly events for the real-time parameter snapshot at the most recent collection time, and send the parameter anomaly event analysis results to the parameter anomaly patrol service, and the parameter anomaly event analysis results include the abnormal database instance and its abnormal database parameters; the parameter anomaly patrol service is used to make a decision based on the parameter anomaly event analysis results to obtain a first decision result, and the first decision result includes the normal parameter values corresponding to the abnormal database parameters of the abnormal database instance, and send the first decision result to the parameter repair agent deployed on the abnormal database instance; the parameter repair agent deployed on the abnormal database instance is used to hot-fix the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
[0006] The embodiment of the present application further provides a method for managing abnormal database parameters, which is applied to a system for managing abnormal database parameters. The system includes: a plurality of data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly patrol service, and a plurality of parameter repair agents deployed on multiple database instances; the method includes: the data collection agents collect real-time parameter data of the database instances in real time and report it to the first real-time message queue for storage; the parameter anomaly online analysis service obtains the real-time parameter data of at least one database instance reported at the most recent collection time in the first real-time message queue and takes a snapshot to obtain a real-time parameter snapshot at the most recent collection time; call a big data processing component to analyze parameter anomaly events for the real-time parameter snapshot at the most recent collection time, and send the parameter anomaly event analysis results to the parameter anomaly patrol service. The parameter anomaly event analysis results include the abnormal database instances and their abnormal database parameters; the parameter anomaly patrol service makes a decision based on the parameter anomaly event analysis results to obtain a first decision result. The first decision result includes the normal parameter values corresponding to the abnormal database parameters of the abnormal database instances, and the first decision result is sent to the parameter repair agents deployed on the abnormal database instances; the parameter repair agents deployed on the abnormal database instances hot-fix the real-time parameter values of the abnormal database parameters of the abnormal database instances to the corresponding normal parameter values in the first decision result.
[0007] The embodiment of the present application further provides an electronic device, including: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to execute the steps in the method for managing abnormal database parameters.
[0008] The embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to be able to implement the steps in the method for managing abnormal database parameters.
[0009] The embodiments of the present application provide a database parameter anomaly governance system, method, electronic device, and storage medium. Among them, the database parameter anomaly governance system includes: multiple data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly inspection service, and multiple parameter repair agents deployed on multiple database instances. The data collection agents transmit the collected real-time parameter data to the first real-time message queue with low latency. Each data collection agent does not interfere with each other and has high robustness, which can greatly improve the timeliness and accuracy of real-time parameter data. The parameter anomaly online analysis service calls the big data processing component to analyze the parameter anomaly events of the real-time parameter snapshot at the most recent collection time, and sends the analysis results of the parameter anomaly events to the parameter anomaly inspection service, so that the parameter anomaly inspection service can make immediate decisions and send the decision results to the parameter repair agents. The parameter repair agents perform repairs in a hot repair manner based on the decision results. Thus, a new database parameter anomaly governance solution is provided. This solution reduces the operation and maintenance difficulty of database parameters, discovers abnormal database parameters efficiently and accurately, greatly shortens the survival time of abnormal database parameters, and improves the stability of database instances. Further optionally, the database parameter anomaly governance system also uses the parameter anomaly offline analysis service to intelligently analyze the parameter anomaly problems with large data volume and complex multiple data sources by combining the offline full-scale data obtained by integrating data of multiple dimensions, and uses big data computing power to support parameter anomaly detection. The parameter anomaly online analysis service is responsible for discovering abnormal database parameters in real time, and the parameter anomaly offline analysis service serves as a backup for the anomaly detection of offline full-scale data. The combination of the parameter anomaly online analysis service and the parameter anomaly offline analysis service can discover more comprehensive abnormal database parameters and conduct more comprehensive governance, and basically will not miss the discovery of abnormal database parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0011] FIG. 1 is an architecture diagram of a traditional database parameter anomaly governance solution;
[0012] FIG. 2 is an architecture diagram of a database parameter anomaly governance system provided by an embodiment of the present application;
[0013] FIG. 3 is an exemplary application scenario diagram provided by an embodiment of the present application;
[0014] FIG. 4 is a process diagram of exemplary online analysis and offline analysis provided by an embodiment of the present application;
[0015] FIG. 5 is a flowchart of a method for governing database parameter anomalies provided by an embodiment of the present application;
[0016] FIG. 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0017] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0018] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the access relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular In addition, in the embodiments of the present application, "first", "second", "third", etc. are only used to distinguish the content of different objects and have no other special meanings.
[0019] Some terms related to the present application are introduced below.
[0020] Database instance: It can be an independent database service process that occupies physical memory, and different memory sizes, disk spaces, and database types can be set for the database instance.
[0021] Database parameters refer to the parameters that affect database performance and can also be understood as the instance parameters of a database instance. For example, but not limited to: Buffer Pool size (cache pool size), query_cache_size (query cache size), max_user_connections (maximum number of user connections), max_connections (maximum number of connections), etc. In this embodiment, the parameter values of database parameters can be divided into required parameter values, configuration parameter values, and real-time parameter values. The required parameter value refers to the parameter value set by the user according to needs, that is, the parameter value that meets the user's needs; the configuration parameter value refers to the parameter value written into the configuration file of the database instance; the real-time parameter value refers to the parameter value of the database parameters generated during the operation of the database instance collected in real time. In practical applications, some database parameters need to restart the database instance to take effect. Generally, they are first sent to the configuration file, and then the database instance is restarted, so as to load the configuration parameter value of the database parameters in the configuration file into the memory to form the effective operation state parameters. The operation state parameters refer to the database parameters with real-time parameter values. That is to say, some real-time parameter values are obtained after the configuration parameter values in the configuration file are loaded into the memory and take effect. The real-time parameter value is the effective configuration parameter value.
[0022] Big data processing component: It can be any streaming processing and batch processing framework. For example, but not limited to: Flink streaming computing platform, which is used for distributed computing and processing of real-time data streams and large-scale data batches.
[0023] Big data development and governance platform: It is a data governance platform that provides services such as data integration, data development, data map, data quality, and data services. For example, but not limited to: DataWorks big data governance platform. DataWorks is based on multiple big data engines and provides a unified full-link big data development and governance platform for solutions such as data warehouses, data lakes, and lakehouse integration.
[0024] Offline data warehouse: It refers to a data warehouse that stores offline data. After the data in different data sources are processed through various processes such as cleaning, transformation, and processing, they are stored in the offline data warehouse and become offline data. The offline data in the offline data warehouse is used for subsequent data analysis or mining. With the help of the offline data warehouse, the data processing and storage are separated, the load of a single system is reduced, the efficiency and accuracy of data processing are improved, and better decision support is provided.
[0025] Figure 1 is an architecture diagram of a traditional database parameter anomaly governance solution. Referring to Figure 1, in the traditional database parameter anomaly governance solution, the required parameter values of the database parameters configured by users according to their needs are stored in the meta-database. The centralized inspection service links to each database instance, and the centralized inspection service senses the real-time parameter values of the database parameters of each database instance, and compares the sensed real-time parameter values of the database parameters with the required parameter values of the database parameters in the meta-database. According to the comparison results, abnormal database instances and their abnormal database parameters are found, and the abnormal database parameters of the abnormal database instances are repaired. When the number of database parameters or the number of database instances expands, the traditional solution will have a great delay, resulting in an extended existence time of database parameter anomalies, which is very likely to cause online failures of the database system. In addition, when the centralized inspection service deals with complex parameter anomaly problems, its ability to sense various application data is limited, and there is a bottleneck in computing power, and it is unable to perform batch complex calculations on all database instances. In the traditional solution, when the repair system repairs the abnormal database parameters of an abnormal database instance, it often obtains the required parameter values of the database parameters from the meta-database, and distributes the required parameter values of the database parameters to the configuration file of the database instance. If the distribution to the configuration file is successful, the required parameter values are converted into configuration parameter values. The repair system first restarts the abnormal database instance, and loads the configuration parameter values of the database parameters obtained from the configuration file into the memory to take effect as the real-time parameter values of the database parameters, thereby completing the repair of the abnormal database parameters of the abnormal database instance. However, the distribution of the required parameter values of the database parameters in the meta-database to the configuration file of the database instance often involves multiple complex parameter synchronizations, which are very likely to result in inconsistent situations between the configuration parameter values, real-time parameter values of the database parameters and the required parameter values in the meta-database.
[0026] In summary, the traditional database parameter anomaly governance solution will more or less miss some abnormal database parameters that are not discovered or it is difficult to accurately repair the parameter values of the database parameters.
[0027]
[0028] In the daily operation and maintenance of a cloud database system, promptly detecting and fixing abnormal database parameters is an important operation and maintenance task. With the rapid development of cloud database systems, some cloud database systems have more than 1 million database instances, and the number of parameters for database instances reaches more than 700. In addition, users also have the need for independent optimization of database parameters. The process of distributing the independently optimized database parameters involves multiple parameter synchronizations, which can easily lead to abnormal database behavior. The superposition of multiple factors has greatly increased the operation and maintenance difficulty of database parameters, posing a greater challenge to the solution for governing abnormal database parameters.
[0029] Therefore, the embodiments of this application provide a system, method, electronic device, and storage medium for governing abnormal database parameters. Among them, the system for governing abnormal database parameters includes: multiple data collection agents deployed on multiple database instances, a first real-time message queue, an online parameter anomaly analysis service, a parameter anomaly inspection service, and multiple parameter repair agents deployed on multiple database instances. The data collection agents transparently transmit the collected real-time parameter data to the first real-time message queue with low latency. Each data collection agent does not interfere with each other and has high robustness, which can greatly improve the timeliness and accuracy of real-time parameter data. The online parameter anomaly analysis service calls the big data processing component to analyze parameter anomaly events for the real-time parameter snapshot at the most recent collection time, and sends the analysis results of parameter anomaly events to the parameter anomaly inspection service, so that the parameter anomaly inspection service can make immediate decisions and send the decision results to the parameter repair agents. The parameter repair agents perform repairs in a hot-fix manner based on the decision results. Thus, a new solution for governing abnormal database parameters is provided. This solution reduces the operation and maintenance difficulty of database parameters, efficiently and accurately discovers abnormal database parameters, greatly shortens the survival time of abnormal database parameters, and improves the stability of database instances. Further optionally, the system for governing abnormal database parameters also uses the offline parameter anomaly analysis service to intelligently analyze parameter anomaly problems with large amounts of data and complex multiple data sources by combining offline full-scale data obtained by integrating data from multiple dimensions, and uses big data computing power to support parameter anomaly detection. The online parameter anomaly analysis service is responsible for real-time discovery of abnormal database parameters, and the offline parameter anomaly analysis service serves as a backup for anomaly detection of offline full-scale data. The combination of the online parameter anomaly analysis service and the offline parameter anomaly analysis service can discover more comprehensive abnormal database parameters and conduct more comprehensive governance, and basically will not miss the discovery of abnormal database parameters.
[0030] The following will describe in detail the technical solutions provided by the embodiments of this application with reference to the accompanying drawings.
[0031] Figure 2 is an architecture diagram of a database parameter anomaly governance system provided by an embodiment of the present application. Referring to Figure 2, the system may include: multiple data collection agents 10 deployed on multiple database instances, a first real-time message queue 20, a parameter anomaly online analysis service 30, a parameter anomaly inspection service 40, and multiple parameter repair agents 50 deployed on multiple database instances.
[0032] In this embodiment, the data collection agent 10 is used to collect real-time parameter data of the database instance in real time and report it to the first real-time message queue for storage.
[0033] Specifically, the data collection agent 10 is a network agent tool with data collection capabilities. A data collection agent 10 responsible for collecting real-time parameter data is deployed for each database instance. The data collection agent 10 can collect real-time parameter data of the database instance it is deployed on in real time according to a set sampling period. The real-time parameter data may include not only the real-time parameter values of at least one database parameter generated during the operation of the database instance, but also the configuration parameter values of at least one database parameter obtained from the configuration file of the database instance.
[0034] Multiple data collection agents 10 distributed on multiple different database instances form a distributed data collection system. Compared with the method of a centralized inspection service alone centrally perceiving data of multiple database instances, the perception requirements of real-time parameter data are dispersed to each data collection agent 10 in the distributed data collection system. The data collection agent 10 transmits the collected real-time parameter data to the first real-time message queue with low latency, and each data collection agent 10 does not interfere with each other, with high robustness, which can greatly improve the timeliness and accuracy of real-time parameter data.
[0035] In this embodiment, the real-time parameter data of the database instance collected by the data collection agent 10 in real time is reported to the first real-time message queue for storage by the first real-time message queue. The first real-time message queue can be any message queue, for example, including but not limited to: RocketMQ. RocketMQ is a distributed message queue system with low latency, high reliability, and strong scalability, supporting features such as distributed deployment, horizontal expansion, and disaster recovery.
[0036] In this embodiment, the online parameter anomaly analysis service 30 is configured to obtain the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and take a snapshot to obtain the real-time parameter snapshot at the most recent collection time; call the big data processing component to analyze the parameter anomaly events in the real-time parameter snapshot at the most recent collection time, and send the analysis results of the parameter anomaly events to the parameter anomaly patrol service 40. The analysis results of the parameter anomaly events include the abnormal database instances and their abnormal database parameters.
[0037] Specifically, the online parameter anomaly analysis service 30 is a service that provides anomaly analysis of database parameters in an online analysis manner and belongs to an online analysis service. The real-time parameter data reported by the data collection agent 10 to the first real-time message queue also includes the collection time of the real-time parameter data. First, the online parameter anomaly analysis service 30 obtains the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time, and the most recent collection time is the collection time closest to the current time. Optionally, the online parameter anomaly analysis service 30 can perform data cleaning on the obtained real-time parameter data to improve data quality. Then, the online parameter anomaly analysis service 30 takes a snapshot of the real-time parameter data of the at least one database instance obtained, to obtain the real-time parameter snapshot at the most recent collection time, and the real-time parameter snapshot at the most recent collection time includes the real-time parameter data of at least one database instance reported at the most recent collection time. Then, the online parameter anomaly analysis service 30 calls the big data processing component to analyze the parameter anomaly events in the real-time parameter snapshot at the most recent collection time. Calling the big data processing component for online analysis can improve the timeliness and accuracy of the online analysis results.
[0038] The online analysis service 30 for parameter anomalies analyzes parameter anomaly events based on database parameters. The recommended parameter values of the database parameters are the parameter values of the database parameters that can ensure the normal operation of the database instance. Compare the real-time parameter values of the database parameters in the real-time parameter snapshot with the corresponding recommended parameter values. If the real-time parameter values of the database parameters in the real-time parameter snapshot are consistent with the corresponding recommended parameter values, it is confirmed that the database parameter is normal; if the real-time parameter values of the database parameters in the real-time parameter snapshot are inconsistent with the corresponding recommended parameter values, it is confirmed that the database parameter is abnormal. In practical applications, if the difference between the real-time parameter value of the database parameter in the real-time parameter snapshot and the corresponding recommended parameter value is within the specified numerical range, the real-time parameter value of the database parameter in the real-time parameter snapshot is consistent with the corresponding recommended parameter value; if the difference between the real-time parameter value of the database parameter in the real-time parameter snapshot and the corresponding recommended parameter value does not fall within the specified numerical range, the real-time parameter value of the database parameter in the real-time parameter snapshot is inconsistent with the corresponding recommended parameter value.
[0039] In this embodiment, the analysis result of the parameter anomaly event output by the online analysis service 30 for parameter anomalies includes the abnormal database instance and its abnormal database parameters. An abnormal database instance refers to a database instance with abnormal database parameters. An abnormal database parameter refers to a database parameter with an abnormal parameter value, that is, the real-time parameter value of the abnormal database parameter is inconsistent with the corresponding recommended parameter value.
[0040] In this embodiment, the online analysis service 30 for parameter anomalies sends the analysis result of the parameter anomaly event to the parameter anomaly inspection service 40 for immediate decision-making by the parameter anomaly inspection service 40, greatly shortening the survival time of the abnormal database parameters and improving the stability of the database instance.
[0041] In this embodiment, the parameter anomaly inspection service 40 is used to make a decision based on the analysis result of the parameter anomaly event to obtain a first decision result. The first decision result includes the normal parameter values corresponding to the abnormal database parameters of the abnormal database instance, and the first decision result is sent to the parameter repair agent 50 deployed on the abnormal database instance.
[0042] Specifically, the parameter anomaly inspection service 40 is an inspection service that can determine the normal parameter values that can repair the abnormal database parameters to normal database parameters. The normal parameter values can be understood as the parameter values of the database parameters when the database instance is running normally. The parameter anomaly inspection service 40 can make a decision based on the analysis result of the parameter anomaly event, and the obtained decision result is the first decision result. In practical applications, the parameter anomaly inspection service 40 can notify the operation and maintenance personnel of the parameter anomaly event, and the parameter anomaly inspection service 40 obtains the first decision result determined by the operation and maintenance personnel. Alternatively, the parameter anomaly inspection service 40 runs a pre-trained machine learning model with a decision-making function, and the machine learning model makes the first decision result, and there is no limitation on this. Alternatively, the parameter anomaly inspection service 40 selects the recommended parameter value corresponding to the abnormal database parameter of the abnormal database instance from the recommended parameter values of each pre-stored database parameter as the corresponding normal parameter value, and generates the first decision result based on the selected normal parameter value. Of course, there is no limitation on the decision-making method of the parameter anomaly inspection service 40.
[0043] The parameter anomaly inspection service 40 sends the first decision result to the parameter repair agent 50 deployed on the abnormal database instance, so that the parameter repair agent 50 can repair the abnormal database parameters of the abnormal database instance.
[0044] In this embodiment, the parameter repair agent 50 deployed on the abnormal database instance is used to hot-fix the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
[0045] Specifically, the parameter repair agent 50 is a network proxy tool with a data repair function. A parameter repair agent 50 responsible for the parameter repair task is deployed on each database instance. The multiple parameter repair agents 50 distributed on multiple different database instances form a distributed data repair system. Compared with the method of centrally repairing the parameters of multiple database instances by only one centralized inspection service, it can greatly improve the timeliness and accuracy of parameter repair.
[0046] The parameter repair agent 50 performs the repair in a hot-fix manner, so that the parameter repair of the database instance can be completed during the operation of the database instance, which improves the parameter repair efficiency and also ensures the stability of the database instance.
[0047] The database parameter anomaly governance system provided by the embodiment of the present application includes: multiple A data collection agent 10, a first real-time message queue, a parameter anomaly online analysis service 30, a parameter anomaly patrol service 40, and multiple parameter repair agents 50 deployed on multiple database instances. The data collection agent 10 transmits the collected real-time parameter data to the first real-time message queue with low latency. Each data collection agent 10 does not interfere with each other and has high robustness, which can greatly improve the timeliness and accuracy of the real-time parameter data. The parameter anomaly online analysis service 30 calls the big data processing component to analyze the parameter anomaly events of the real-time parameter snapshot at the most recent collection time, and sends the analysis results of the parameter anomaly events to the parameter anomaly patrol service 40, so that the parameter anomaly patrol service 40 can make immediate decisions and send the decision results to the parameter repair agent 50. The parameter repair agent 50 performs repairs in a hot repair manner based on the decision results. Thus, a new database parameter anomaly governance solution is provided. This solution reduces the operation and maintenance difficulty of database parameters, efficiently and accurately discovers abnormal database parameters, greatly shortens the survival time of abnormal database parameters, and improves the stability of database instances.
[0048] In some alternative embodiments, the parameter anomaly online analysis service 30 further has a parameter change event analysis function. Referring to FIG. 2, the database parameter anomaly governance system may further include: a second real-time message queue 60. The second real-time message queue 60 can be any message queue, such as including but not limited to: RocketMQ.
[0049] In this embodiment, the parameter anomaly online analysis service 30 is further configured to report the real-time parameter snapshot at the most recent collection time to the second real-time message queue 60 for storage; obtain the real-time parameter snapshot at the previous collection time from the second real-time message queue 60; call the big data processing component to perform parameter change event analysis based on the real-time parameter snapshots at the most recent collection time and the previous collection time, and send the analysis results of the parameter change events to the parameter anomaly patrol service 40. The analysis results of the parameter change events include the changed database instance and its changed database parameters.
[0050] Specifically, the second real-time message queue 60 stores the real-time parameter snapshots at each collection time. The parameter anomaly online analysis service 30 reports the real-time parameter snapshot at the most recent collection time to the second real-time message queue 60 for storage.
[0051] The online parameter anomaly analysis service 30 analyzes parameter change events based on real-time parameter snapshots at two adjacent collection times. When analyzing parameter change events, for the same database parameter of the same database instance, it is judged whether the real-time parameter value of the database parameter in the real-time parameter snapshot at the most recent collection time is consistent with the real-time parameter value of the database parameter in the real-time parameter snapshot at the previous collection time. If the difference between the real-time parameter value of the database parameter in the real-time parameter snapshot at the most recent collection time and the real-time parameter value of the database parameter in the real-time parameter snapshot at the previous collection time falls within the specified data range, the judgment result is consistent, and it is confirmed that no parameter change event has occurred; if the difference between the real-time parameter value of the database parameter in the real-time parameter snapshot at the most recent collection time and the real-time parameter value of the database parameter in the real-time parameter snapshot at the previous collection time does not fall within the specified data range, the judgment result is inconsistent, and it is confirmed that a parameter change event has occurred.
[0052] The online parameter anomaly analysis service 30 sends the parameter change event analysis result to the parameter anomaly inspection service 40. The parameter change event analysis result includes the changed database instance and its changed database parameter. The changed database instance is the database instance where the parameter change event occurs, and the changed database parameter is the database parameter where the parameter change event occurs.
[0053] The parameter anomaly inspection service 40 is used to make a decision based on the parameter change event analysis result to obtain a second decision result. The second decision result includes the normal parameter value corresponding to the changed database parameter of the changed database instance, and the decision result is sent to the parameter repair agent 50 deployed on the changed database instance.
[0054] Specifically, the parameter anomaly inspection service 40 can make a decision based on the parameter change event analysis result, and the obtained decision result is the second decision result. In practical applications, the parameter anomaly inspection service 40 can notify the operation and maintenance personnel of the parameter change event, and the parameter anomaly inspection service 40 obtains the second decision result of the operation and maintenance personnel's decision. Or, the parameter anomaly inspection service 40 runs a pre-trained machine learning model with a decision-making function, and the machine learning model makes the second decision result. Or, the parameter anomaly inspection service 40 selects the recommended parameter value corresponding to the abnormal database parameter of the abnormal database instance from the pre-stored recommended parameter values of each database parameter as the corresponding normal parameter value, and generates the second decision result based on the selected normal parameter value. Of course, the decision-making method of the parameter anomaly inspection service 40 is not limited.
[0055] The parameter anomaly inspection service 40 sends the second decision result to the parameter repair agent 50 deployed on the changed database instance, so that the parameter repair agent 50 repairs the changed database parameters of the changed database instance.
[0056] In this embodiment, the parameter repair agent 50 deployed on the changed database instance is used to hot-fix the real-time parameter value of the changed database parameters of the changed database instance to the corresponding normal parameter value.
[0057] It is worth noting that the database parameter anomaly governance system also has the ability to analyze parameter change events, can instantly repair database parameters, and improve the effect of database parameter anomaly governance.
[0058] In some alternative embodiments, referring to FIG. 2, the database parameter anomaly governance system may further include: a parameter snapshot retrieval service 70.
[0059] Specifically, the parameter snapshot retrieval service 70 is a retrieval service that can retrieve real-time parameter snapshots in the second real-time message queue 60. The parameter snapshot retrieval service 70 can also provide a parameter snapshot retrieval API (Application Programming Interface) for retrieval.
[0060] In practical applications, during the decision-making process, if the parameter anomaly inspection service 40 needs to obtain more real-time parameter snapshots to assist in decision-making or view more real-time parameter snapshots, it requests the parameter snapshot retrieval service 70 to retrieve real-time parameter snapshots that meet the retrieval conditions.
[0061] Optionally, the parameter anomaly inspection service 40 is used to send a retrieval request including a retrieval time range to the parameter snapshot retrieval service 70, and receive real-time parameter snapshots returned by the parameter snapshot retrieval service 70 that fall within the retrieval time range. The parameter snapshot retrieval service 70 is used to obtain real-time parameter snapshots that fall within the retrieval time range from the second real-time message queue 60 according to the retrieval request.
[0062] Specifically, the retrieval condition refers to retrieving real-time parameter snapshots that fall within the retrieval time range. Since the real-time parameter snapshots in the second real-time message queue 60 are stored with the collection time as the time granularity, the parameter snapshot retrieval service 70 can accurately retrieve real-time parameter snapshots that fall within the retrieval time range.
[0063] In some alternative embodiments, the database parameter anomaly governance system may further include: a meta-database 80, an offline data warehouse 90, and a parameter anomaly offline analysis service 100.
[0064] In this embodiment, the metadata database 80 is used to store demand parameter data that meets the user's requirements.
[0065] Specifically, the demand parameter data includes demand parameter values of at least one database parameter of at least one database instance, and the demand parameter values are the parameter values of the database parameters configured by the user as needed. The demand parameter values of the database parameters configured by the user are first saved in the metadata database 80. As time goes by, the demand parameter values of the database parameters in the metadata database 80 are read and written into the configuration file of the database instance to be converted into configuration parameter values.
[0066] In this embodiment, the offline data warehouse 90 is used to store offline full-volume data. The offline full-volume data includes, for example, but is not limited to at least one of the following: real-time parameter snapshots of at least one collection time obtained from the first real-time message queue, demand parameter data obtained from the metadata database 80, a parameter value suggestion table, a list of key database parameters, instance performance data, and instance log data.
[0067] In this embodiment, the parameter value suggestion table includes suggested parameter values of at least one database parameter; the list of key database parameters includes at least one key parameter. A key parameter refers to a database parameter that has a greater impact on the database performance and can be flexibly selected as needed. The instance performance data refers to data related to the performance of the database instance, such as, for example, but not limited to: disk usage rate, memory usage rate, and CPU (Central Processing Unit) usage rate. The instance log data includes, for example, but is not limited to: instance error log data and instance slow log data. The instance error log data is the log data recording various errors that occur in the instance, such as, for example, but not limited to: instance creation failure information, instance startup failure information, or instance connection failure information. The instance slow log data refers to the slow query log (Slow Log), which is used to record commands in the database instance whose execution time exceeds a specified threshold.
[0068] In this embodiment, the parameter anomaly offline analysis service 100 is used to call the big data development and governance platform to perform complex correlation analysis on the offline full-volume data in the offline data warehouse 90 to obtain an offline analysis result; and send the offline analysis result to the parameter anomaly inspection service 40.
[0069] Specifically, the parameter anomaly offline analysis service 100 is an offline analysis service that performs complex correlation analysis on the offline full-volume data by virtue of the big data mining and analysis capabilities of the big data development and governance platform. The complex correlation analysis can The solution is to perform correlation analysis on complex offline full-scale data. The offline analysis results include, for example, but are not limited to at least one of the following.
[0070] ①. A detailed list of inconsistent master-slave database parameters. Among them, the detailed list of inconsistent master-slave database parameters includes the detailed information of at least one database parameter. The detailed information of the database parameter includes: the master database instance and the standby database instance associated with the database parameter, and the real-time parameter values of the database parameter of the master database instance are inconsistent with the real-time parameter values of the database parameter of the standby database instance.
[0071] ②. A detailed list of inconsistent configuration parameter values and real-time parameter values. Among them, the detailed list of inconsistent configuration parameter values and real-time parameter values includes the detailed information of at least one database parameter. The detailed information of the database parameter includes: the real-time parameter value of the database parameter and its configuration parameter value in the configuration file, and the real-time parameter value of the database parameter is inconsistent with its configuration parameter value in the configuration file.
[0072] ③. A detailed list of inconsistent requirement parameter values and real-time parameter values. Among them, the detailed list of inconsistent requirement parameter values and real-time parameter values includes the detailed information of at least one database parameter. The detailed information of the database parameter includes: the requirement parameter value and the real-time parameter value of the database parameter, and the requirement parameter value of the database parameter is inconsistent with the real-time parameter value.
[0073] ©. A detailed list of instance performance indicators exceeding the recommended parameter values. The detailed list of instance performance indicators exceeding the recommended parameter values includes at least one instance performance indicator exceeding the corresponding recommended parameter value.
[0074] In this embodiment, the parameter anomaly inspection service 40 is used to make a decision based on the offline analysis results to obtain a third decision result. The third decision result includes the normal parameter values corresponding to the abnormal database parameters of the abnormal database instance, and the decision result is sent to the parameter repair agent 50 deployed on the abnormal database instance.
[0075] Specifically, the parameter anomaly patrol service 40 can make a decision based on the offline analysis results, and the obtained decision result is the third decision result. In practical applications, the parameter anomaly patrol service 40 can notify the operation and maintenance personnel of a parameter change event, and the parameter anomaly patrol service 40 obtains the third decision result of the operation and maintenance personnel's decision. Alternatively, the parameter anomaly patrol service 40 runs a pre-trained machine learning model with a decision-making function, and the machine learning model makes the third decision result. Alternatively, the parameter anomaly patrol service 40 selects the recommended parameter value corresponding to the abnormal database parameter of the abnormal database instance from the recommended parameter values of each database parameter stored in advance as the corresponding normal parameter value, and generates the third decision result based on the selected normal parameter value. Of course, the decision-making method of the parameter anomaly patrol service 40 is not limited.
[0076] The parameter anomaly patrol service 40 sends the third decision result to the parameter repair agent 50 deployed on the abnormal database instance, so that the parameter repair agent 50 repairs the abnormal database parameter of the abnormal database instance.
[0077] In this embodiment, the parameter repair agent 50 deployed on the abnormal database instance is used to hot-fix the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value 0
[0078] In this embodiment, through the parameter anomaly offline analysis service 100, combined with the offline full-scale data obtained by fusing data of multiple dimensions, intelligent analysis is performed on the parameter anomaly problems with a large amount of data and complex data sources of multiple data sources, and the parameter anomaly detection is supported by big data computing power. The parameter anomaly online analysis service 30 is responsible for real-time discovery of abnormal database parameters, and the parameter anomaly offline analysis service 100 provides a backup for the anomaly detection of offline full-scale data. The combination of the parameter anomaly online analysis service 30 and the parameter anomaly offline analysis service 100 can discover more comprehensive abnormal database parameters and perform more comprehensive governance, and basically no abnormal database parameters will be missed.
[0079] A specific database parameter anomaly governance solution will be introduced below with reference to FIGS. 3 and 4.
[0080] Specifically, a data collection agent and a parameter repair agent are deployed on each database instance. The data collection agents distributed on multiple database instances form a distributed data collection system; the parameter repair agents distributed on multiple database instances form a distributed data repair system.
[0081] In this embodiment, each data collection agent is responsible for collecting in real time the real-time parameter data of the database instance it deploys. The real-time parameter data includes the real-time parameter values and / or configuration parameter values of at least one database parameter. As shown in ① of Figure 3, the real-time parameter data of the database instance collected in real time by the data collection agent is reported to the first real-time message queue. As shown in ② of Figure 3, the parameter anomaly online analysis service obtains the real-time parameter data in the first real-time message queue for online analysis and outputs the online analysis result. Referring to Figure 4, the online analysis result includes, for example, but is not limited to: the analysis result of the parameter anomaly event and the analysis result of the parameter change event. Referring to Figure 4, the parameter anomaly online analysis service includes: the parameter anomaly analysis task, the transaction issuance task, the parameter snapshot aggregation service, the rule configuration center, the parameter anomaly event retrieval API, the algorithm center, and so on.
[0082] Among them, the parameter anomaly analysis task is used to collect the real-time parameter data from the first real-time message queue for data cleaning, and call the parameter snapshot aggregation service to perform snapshot processing on the cleaned real-time parameter data to obtain a real-time parameter snapshot; and perform online analysis based on the real-time parameter snapshot.
[0083] The transaction issuance task is used to issue the analysis result of the parameter anomaly event or the analysis result of the parameter change event to the parameter anomaly patrol service in the form of an event.
[0084] The parameter snapshot aggregation service is used to aggregate the real-time parameter data reported by all data collection agents within the time granularity into a snapshot data; the rule configuration center is responsible for maintaining the generation rules of events such as parameter anomaly events and parameter update events, and issuing rules for different usage scenarios.
[0085] The parameter anomaly event retrieval APL is used to retrieve the analysis result of the parameter anomaly event that meets the user's needs.
[0086] The algorithm center is a common library that precipitates event generation and analysis algorithms, and provides a unified SDK (Software Development Kit) for online analysis. package, Software Development Kit) for online analysis.
[0087] In this embodiment, as shown in ③ of Figure 3, the parameter anomaly online analysis service sends the online analysis result to the parameter anomaly patrol service. The parameter anomaly patrol service makes a decision and outputs a decision result. The decision result includes the normal parameter values corresponding to the database parameters of the database instance that needs to be repaired.
[0088] In this embodiment, as shown in ④ of FIG. 3, the parameter anomaly patrol service sends the decision result to the corresponding parameter repair agent, so that the parameter repair agent performs hot repair on the database parameters of the database instance to be repaired based on the normal parameter values in the decision result. Thus, through the mutual cooperation of the parameter anomaly online analysis service and the parameter anomaly patrol service, abnormal database parameters are discovered in real time and repaired in a timely manner.
[0089] As shown in ⑤ of FIG. 3, the parameter anomaly online analysis service can also report the real-time parameter snapshot to the second real-time message queue for storage. As shown in ⑥ and ⑦ of FIG. 3, the parameter anomaly patrol service can also send a retrieval request to the parameter snapshot retrieval service. The parameter snapshot retrieval service retrieves the real-time parameter snapshots that meet the retrieval conditions in the second real-time message queue and returns them to the parameter anomaly patrol service for the parameter anomaly patrol service to analyze and use.
[0090] In this embodiment, as shown in ⑧ of FIG. 3, the parameter anomaly offline analysis service obtains the offline full-volume data from the offline data warehouse. Referring to FIG. 4, the offline full-volume data includes, for example, but is not limited to: real-time parameter snapshots, demand parameter data, parameter value suggestion tables, key database parameter lists, instance performance data, and instance log data, etc. The offline analysis results output by the parameter anomaly offline analysis service include, for example, but are not limited to: the detailed list of inconsistent primary and standby database parameters, the detailed list of inconsistent configuration parameter values and real-time parameter values, the detailed list of inconsistent demand parameter values and real-time parameter values, and the detailed list of instance performance indicators exceeding the recommended parameter values.
[0091] In this embodiment, as shown in ⑨ of FIG. 3, the parameter anomaly offline analysis service sends the offline analysis result to the parameter anomaly patrol service for decision-making and outputs the decision result. The decision result includes the normal parameter values corresponding to the database parameters of the database instance to be repaired. The parameter anomaly patrol service sends the decision result to the corresponding parameter repair agent, so that the parameter repair agent performs hot repair on the database parameters of the database instance to be repaired based on the normal parameter values in the decision result. Thus, through the mutual cooperation of the parameter anomaly offline analysis service and the parameter anomaly patrol service, it provides a backup for the anomaly detection of the offline full-volume data.
[0092] The database parameter anomaly governance solution provided by the embodiments of this application uses multiple data collection agents and multiple parameter repair agents deployed on different database instances to perform data collection and data repair respectively, and the Agents with different functions operate independently of each other, with high robustness.
[0093] The data collection agent collects the real-time parameter data of the database instance with low latency to achieve maximum timeliness and accuracy; the centralized data perception requirements are dispersed to each data collection agent, and the centralized parameter anomaly inspection service no longer undertakes the data perception task and only serves as the decision-maker for database parameter anomaly governance, reducing the load pressure of the parameter anomaly inspection service. Especially in the face of the increasing number of database instances, the inspection efficiency for abnormal database parameters is improved.
[0094] The parameter anomaly online analysis service can call the big data processing component for online analysis, greatly shortening the online survival time of abnormal database parameters and improving the stability of the database instance. The parameter anomaly offline analysis service can call the big data development and governance platform for offline analysis. During offline analysis, combined with the offline full-scale data obtained by integrating data from multiple dimensions, it conducts intelligent analysis on the parameter anomaly problems with large data volume and complex data sources from multiple data sources, and uses big data computing power to support parameter anomaly detection. The parameter anomaly online analysis service is responsible for real-time discovery of abnormal database parameters, and the parameter anomaly offline analysis service serves as a backup for the anomaly detection of offline full-scale data. The combination of the parameter anomaly online analysis service and the parameter anomaly offline analysis service can discover more comprehensive abnormal database parameters and conduct more comprehensive governance, basically not missing the discovery of abnormal database parameters, and realizing full-process and multi-dimensional support for the parameter governance of cloud databases (especially distributed cloud databases).
[0095] Figure 5 is a flowchart of a method for database parameter anomaly governance provided by an embodiment of the present application. It is applied to a database parameter anomaly governance system, which includes: multiple data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly inspection service, and multiple parameter repair agents deployed on multiple database instances.
[0096] Referring to Figure 5, the method for database parameter anomaly governance may include the following steps 101 to step 104.
[0097] 101. The data collection agent collects the real-time parameter data of the database instance in real time and reports it to the first real-time message queue for storage.
[0098] 102. The parameter anomaly online analysis service obtains the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and takes a snapshot to obtain the real-time parameter snapshot at the most recent collection time; calls the big data processing component to analyze the parameter anomaly events for the real-time parameter snapshot at the most recent collection time, and sends the parameter anomaly event analysis results to the parameter anomaly patrol service. The parameter anomaly event analysis results include the abnormal database instance and its abnormal database parameters.
[0099] 103. The parameter anomaly patrol service makes a decision based on the parameter anomaly event analysis results to obtain the first decision result. The first decision result includes the normal parameter values corresponding to the abnormal database parameters of the abnormal database instance, and sends the first decision result to the parameter repair agent deployed on the abnormal database instance.
[0100] 104. The parameter repair agent deployed on the abnormal database instance hot-fixes the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
[0101] Further optionally, the above system further includes: a second real-time message queue; the above method further includes: the parameter anomaly online analysis service reports the real-time parameter snapshot at the most recent collection time to the second real-time message queue for storage; obtains the real-time parameter snapshot at the previous collection time from the second real-time message queue; calls the big data processing component to analyze the parameter change events based on the real-time parameter snapshots at the most recent collection time and the previous collection time, and sends the parameter change event analysis results to the parameter anomaly patrol service. The parameter change event analysis results include the changed database instance and its changed database parameters; the parameter anomaly patrol service makes a decision based on the parameter change event analysis results to obtain the second decision result. The second decision result includes the normal parameter values corresponding to the changed database parameters of the changed database instance, and sends the decision result to the parameter repair agent deployed on the changed database instance; the parameter repair agent deployed on the changed database instance hot-fixes the real-time parameter value of the changed database parameter of the changed database instance to the corresponding normal parameter value.
[0102] Further optionally, the above system further includes: a parameter snapshot retrieval service; the above method further includes: the parameter anomaly patrol service sends a retrieval request including the retrieval time range to the parameter snapshot retrieval service, and receives the real-time parameter snapshots within the retrieval time range returned by the parameter snapshot retrieval service; the parameter snapshot retrieval service obtains the real-time parameter snapshots within the retrieval time range from the second real-time message queue according to the retrieval request.
[0103] Further optionally, the above system further includes: a meta database, an offline data warehouse, and a parameter anomaly offline analysis service; the above method further includes: the meta database stores demand parameter data that meets user requirements; the offline data warehouse stores offline full-volume data; the parameter anomaly offline analysis service calls the big data development and governance platform to perform complex correlation analysis on the offline full-volume data in the offline data warehouse to obtain an offline analysis result; and sends the offline analysis result to the parameter anomaly inspection service; the parameter anomaly inspection service makes a decision based on the offline analysis result to obtain a third decision result, where the third decision result includes the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and sends the decision result to the parameter repair agent deployed on the abnormal database instance; the parameter repair agent deployed on the abnormal database instance hot-fixes the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value.
[0104] Further optionally, the offline full-volume data includes at least two of the following: real-time parameter snapshots at at least one collection time obtained from the first real-time message queue, demand parameter data obtained from the meta database, a parameter value suggestion table, a key database parameter list, instance performance data, and instance log data, where the parameter value suggestion table includes suggested parameter values for at least one database parameter, and the key database parameter list includes at least one key parameter.
[0105] Further optionally, the offline analysis result includes at least one of the following: a master-slave database parameter inconsistency details list, where the master-slave database parameter inconsistency details list includes detailed information about at least one database parameter, and the detailed information about the database parameter includes: the master database instance and the standby database instance associated with the database parameter, and the real-time parameter values of the database parameter of the master database instance and the real-time parameter values of the database parameter of the standby database instance are inconsistent; the configuration parameter value and the real-time parameter value inconsistency details list, where the configuration parameter value and the real-time parameter value inconsistency details list includes detailed information about at least one database parameter, and the detailed information about the database parameter includes: the real-time parameter value of the database parameter and its configuration parameter value in the configuration file, and the real-time parameter value of the database parameter and its configuration parameter value in the configuration file are inconsistent; the demand parameter value and the real-time parameter value inconsistency details list, where the demand parameter value and the real-time parameter value inconsistency details list includes detailed information about at least one database parameter, and the detailed information about the database parameter includes: the demand parameter value and the real-time parameter value of the database parameter, and the demand parameter value and the real-time parameter value of the database parameter are inconsistent; the configuration parameter value and the real-time parameter value inconsistency details list, where the configuration parameter value and the real-time parameter value inconsistency details list includes detailed information about at least one database parameter, and the detailed information about the database parameter includes: the real-time parameter value of the database parameter and its configuration parameter value in the configuration file, and the real-time parameter value of the database parameter and its configuration parameter value in the configuration file are inconsistent; the demand parameter value and the real-time parameter value inconsistency details list, where the demand parameter value and the real-time parameter value inconsistency details list includes detailed information about at least one database parameter, and the detailed information about the database parameter includes: the demand parameter value and the real-time parameter value of the database parameter, and the demand parameter value and the real-time parameter value of the database parameter are inconsistent;
[0106] Detailed list of instance performance indicators exceeding the recommended parameter values. The detailed list of instance performance indicators exceeding the recommended parameter values includes at least one instance performance indicator exceeding the corresponding recommended parameter value.
[0107] For the detailed implementation process of each step in the method shown in Embodiment of FIG. 5, reference may be made to the relevant descriptions in the foregoing system embodiment, which will not be elaborated herein.
[0108] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0109] FIG. 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As shown in FIG. 6, the electronic device includes: a memory 61 and a processor 62; the memory 61 is used for storing computer programs and can be configured to store various other data to support operations on a computing platform. Examples of these data include instructions for any application program or method for operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.
[0110] The memory 61 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0111] The processor 62 is coupled to the memory 61 and is used for executing the computer program in the memory 61 to: execute the steps in the method for governing database parameter anomalies.
[0112] Further optionally, as shown in FIG. 6, the electronic device further includes other components such as a communication component 63, a display 64, a power supply component 65, an audio component 66, etc. Only some components are schematically shown in FIG. 6, which does not mean that the electronic device only includes the components shown in FIG. 6. In addition, the components within the dashed box in FIG. 6 are optional components, rather than essential components, and can be determined according to the product form of the electronic device. The electronic device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT (Internet of Things) device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the electronic device in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include the components within the dashed box in FIG. 6; if the electronic device in this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include the components within the dashed box in FIG. 6. For the detailed implementation process of the processor executing each action, reference can be made to the relevant descriptions in the foregoing method embodiment or device embodiment, and details are not described herein again.
[0113] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step executable by the electronic device in the foregoing method embodiment.
[0114] Correspondingly, an embodiment of the present application further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the processor can be caused to implement each step executable by the electronic device in the foregoing method embodiment.
[0115]
[0116] The above-mentioned communication component is configured to facilitate communication, either wired or wireless, between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi (Wireless Fidelity), 2G (2 Generation), 3G (3 Generation), 4G (4 Generation) / LTE (Long Term Evolution), 5G (5 Generation), etc., or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, The Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0117] The above-mentioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations.
[0118] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0119] The above-described audio component can be configured to output and / or input an audio signal. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in a memory or transmitted via a communication component. In some embodiments, the audio component further includes a speaker for outputting an audio signal.
[0120] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.
[0121] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0122] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0124] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and memory.
[0125] The memory may include non-permanent memory in the form of computer-readable media, random access memory (Random Access Memory, RAM) and / or non-volatile memory such as read only memory (Read Only Memory, ROM) or flash RAM. Memory is an example of computer-readable media.
[0126] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change RAM (PRAM), static random access memory (Static Random-Access Memory, SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), other types of random access memory (Random Access Memory, RAM), read only memory (Read Only Memory, ROM), electrically erasable programmable read-only memory (Electrically-Erasable Programmable Read-Only Memory, EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disc (Digital versatile disc, DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0127] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising a ....." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0128] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
Claims 1. A database parameter anomaly management system, comprising: Multiple data collection agents deployed in multiple database instances, a first real-time message queue, an abnormal parameter online analysis service, an abnormal parameter inspection service, and multiple parameter repair agents deployed in multiple database instances; the data collection agent is used to collect real-time parameter data of the database instance in real time and report it to the first real-time message queue for storage; The parameter anomaly online analysis service is used to obtain the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and take a snapshot to obtain the real-time parameter snapshot at the most recent collection time; call the big data processing component to perform parameter anomaly event analysis on the real-time parameter snapshot at the most recent collection time, and send the parameter anomaly event analysis result to the parameter anomaly inspection service, wherein the parameter anomaly event analysis result includes an abnormal database instance and its abnormal database parameters; The parameter abnormality inspection service is used to make a decision according to the parameter abnormality event analysis result to obtain a first decision result, wherein the first decision result includes a normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and send the first decision result to a parameter repair agent deployed on the abnormal database instance; The parameter repair agent deployed on the abnormal database instance is used to hot-repair the real-time parameter value of the abnormal database parameter of the abnormal database instance to a normal parameter value corresponding to the first decision result.
2. The system according to claim 1, further comprising: Second real-time message queue; The parameter anomaly online analysis service is further used to report the real-time parameter snapshot at the most recent acquisition time to the second real-time message queue for storage; Obtaining a real-time parameter snapshot of the last collection time from the second real-time message queue; Call the big data processing component to perform parameter change event analysis based on the real-time parameter block of the most recent acquisition time and the previous acquisition time, and send the parameter change event analysis result to the parameter abnormality inspection service, wherein the parameter change event analysis result includes the changed database instance and its changed database parameter; The parameter abnormality inspection service is used to make a decision according to the parameter change event analysis result to obtain a second decision result, wherein the second decision result includes a normal parameter value corresponding to the changed database parameter of the changed database instance, and send the decision result to a parameter repair agent deployed on the changed database instance; The parameter repair agent deployed on the change database instance is used to hot-repair the real-time parameter value of the change database parameter of the change database instance to a corresponding normal parameter value.
3. The system according to claim 2, further comprising: Parameter snapshot retrieval service; The parameter abnormality inspection service is used to send a search request including a search time range to the parameter snapshot search service, and receive a real-time parameter snapshot within the search time range returned by the parameter snapshot search service; The parameter snapshot retrieval service is used to obtain the real-time parameter snapshot within the retrieval time range from the second real-time message queue according to the retrieval request.
4. The system according to claim 1, further comprising: Metadatabase, offline data warehouse, and parameter anomaly offline analysis services; The metadata database is used to store demand parameter data that meets user needs; The offline data warehouse is used to store the full amount of offline data; The parameter anomaly offline analysis service is used to call the big data development and governance platform to perform complex correlation analysis on the offline full data in the offline data warehouse to obtain offline analysis results; and sending the offline analysis result to a parameter abnormality inspection service; The parameter abnormality inspection service is used to make a decision based on the offline analysis result to obtain a third decision result, wherein the third decision result includes a normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and send the decision result to a parameter repair agent deployed on the abnormal database instance; The parameter repair agent deployed on the abnormal database instance is used to hot-repair the real-time parameter value of the abnormal database parameter of the abnormal database instance to a corresponding normal parameter value.
5. The system according to claim 4, wherein: The offline full data includes at least two of the following: a real-time parameter snapshot of at least one collection time obtained from the first real-time message queue, demand parameter data obtained from the metadata database, a parameter value recommendation table, a key database parameter list, instance performance data, and instance log data, wherein the parameter value recommendation table includes a recommended parameter value of at least one database parameter, and the key database parameter list includes at least one key parameter.
6. The system according to claim 4, wherein: The offline analysis result includes at least one of the following: a master-slave database parameter inconsistency detailed table, wherein the master-slave database parameter inconsistency detailed table includes detailed information of at least one database parameter, and the detailed information of the database parameter includes: the master database instance and the standby database instance associated with the database parameter, and the real-time parameter value of the database parameter of the master database instance is inconsistent with the real-time parameter value of the database parameter of the standby database instance; a configuration parameter value and real-time parameter value inconsistency detailed table, wherein the configuration parameter value and real-time parameter value inconsistency detailed table includes detailed information of at least one database parameter, and the detailed information of the database parameter includes: the real-time parameter value of the database parameter and its configuration parameter value in the configuration file, and the real-time parameter value of the database parameter and its configuration parameter value in the configuration file are inconsistent; a requirement parameter value and real-time parameter value inconsistency detailed table, wherein the requirement parameter value and real-time parameter value inconsistency detailed table includes detailed information of at least one database parameter, and the detailed information of the database parameter includes: the requirement parameter value and the real-time parameter value of the database parameter, and the requirement parameter value and the real-time parameter value of the database parameter are inconsistent consistent; an instance performance indicator exceeding a recommended parameter value list, wherein the instance performance indicator exceeding a recommended parameter value list includes at least one instance performance indicator exceeding a corresponding recommended parameter value.
7. A database parameter anomaly management method, applied to a database parameter anomaly management system, the system comprising: A plurality of data acquisition agents deployed on a plurality of database instances, a first real-time message queue, an online parameter anomaly analysis service, an anomaly parameter inspection service, and a plurality of parameter repair agents deployed on a plurality of database instances; the method comprises: the data acquisition agent collects the real-time parameter data of the database instance in real time and reports it to the first real-time message queue for storage; the online parameter anomaly analysis service obtains the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent acquisition time and takes a snapshot to obtain the real-time parameter snapshot of the most recent acquisition time; the big data processing component is called to perform an abnormal parameter event analysis on the real-time parameter snapshot of the most recent acquisition time, and the abnormal parameter event analysis result is sent to the abnormal parameter inspection service, the abnormal parameter event analysis result includes the abnormal database instance and its abnormal database parameter; the abnormal parameter inspection service makes a decision according to the abnormal parameter event analysis result to obtain a first decision result, the first decision result includes the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and sends the first decision result to the parameter repair agent deployed on the abnormal database instance; The parameter repair agent deployed on the abnormal database instance hot-repairs the real-time parameter value of the abnormal database parameter of the abnormal database instance to a normal parameter value corresponding to the first decision result.
8. The method according to claim 7, wherein: The system also includes: a second real-time message queue; the method also includes: the parameter anomaly online analysis service reports the real-time parameter snapshot of the most recent collection time to the second real-time message queue for storage; obtains the real-time parameter snapshot of the last collection time from the second real-time message queue; calls the big data processing component to perform parameter change event analysis based on the real-time parameter snapshot of the most recent collection time and the last collection time, and sends the parameter change event analysis result to the parameter anomaly inspection service, the parameter change event analysis result includes the change database instance and its change database parameter; the parameter anomaly inspection service makes a decision based on the parameter change event analysis result to obtain a second decision result, the second decision result includes the normal parameter value corresponding to the change database parameter of the change database instance, and sends the decision result to the parameter repair agent deployed on the change database instance; the parameter repair agent deployed on the change database instance hot repairs the real-time parameter value of the change database parameter of the change database instance to the corresponding normal parameter value.
9. The method according to claim 8, wherein: The system also includes: a parameter snapshot retrieval service; The parameter anomaly inspection service sends a retrieval request including a retrieval time range to the parameter snapshot retrieval service, and receives a real-time parameter snapshot returned by the parameter snapshot retrieval service that falls within the retrieval time range; the parameter snapshot retrieval service obtains the real-time parameter snapshot that falls within the retrieval time range from the second real-time message queue according to the retrieval request.
10. The method according to claim 7, wherein: The system also includes: a metadata database, an offline data warehouse and a parameter anomaly offline analysis service; the metadata database stores demand parameter data that meets user needs; the offline data warehouse stores offline full data; the parameter anomaly offline analysis service calls the big data development and governance platform to perform complex correlation analysis on the offline full data in the offline data warehouse to obtain an offline analysis result; and sends the offline analysis result to the parameter anomaly inspection service; the parameter anomaly inspection service is used to make a decision based on the offline analysis result to obtain a third decision result, the third decision result includes a normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and sends the decision result to a parameter repair agent deployed on the abnormal database instance; the parameter repair agent deployed on the abnormal database instance hot repairs the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value.
11. An electronic device, comprising: Memory and processor; The memory is used to store a computer program; the processor is coupled to the memory, and is used to execute the computer program to perform the steps in the method according to any one of claims 7 to 10.
12. A computer-readable storage medium storing a computer program, wherein: When the computer program is executed by a processor, the processor is enabled to implement the steps in any one of the methods of claims 7 to 10.
Citation Information
Patent Citations
Method and system for testing performance
CN106855844A
Execution plan processing method, device and system
CN113312371A
Monitoring method and system for monitoring and alarming multiple distributed MPP clusters
CN113419925A
Parameter adjusting method and device of distributed database and electronic equipment
CN115658663A
Database migration method and device, computer equipment and storage medium
CN116185991A