Database parameter exception management system and method, electronic equipment and storage medium
By deploying data acquisition agents, real-time message queues, online analysis services, inspection services and repair agents in the cloud database system, real-time acquisition, analysis and repair of database parameters is achieved, solving the problem of increasing difficulty in operating and maintaining database parameters in the cloud database system and improving system stability.
Patent Information
- Application Number
- CN202311773450.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-24
AI Technical Summary
The difficulty of operating and maintaining database parameters in cloud database systems increases, resulting in difficulty in discovering and repairing abnormal database parameters in a timely manner, affecting system stability.
Design a database parameter exception management system, including data collection agent, real-time message queue, online parameter exception analysis service, parameter exception inspection service and parameter repair agent, and achieve rapid response and automatic repair by collecting, analyzing and repairing database parameters in real time.
It effectively reduces the difficulty of operating and maintaining database parameters, improves the efficiency of discovering and repairing abnormal database parameters, shortens the survival time of abnormal parameters, and improves the stability of database instances.
Smart Images

Figure CN120196610A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of database technologies, and in particular, to a database parameter anomaly governance system, method, electronic device, and storage medium. Background Art
[0002] A cloud database system refers to a database system optimized or deployed in a virtual computer environment, which can realize the advantages of pay-per-use, on-demand scaling, high availability, and storage integration. Currently, the market has increasing requirements for the characteristics of cloud database systems. To meet various user needs, cloud database systems are becoming more and more complex, and performance adaptation needs to be carried out through parameter adjustment for different user application scenarios.
[0003] In the daily operation and maintenance of cloud database systems, promptly discovering and fixing abnormal database parameters is an important operation and maintenance task. With the rapid development of cloud database systems, some cloud database systems have more than 1 million database instances, and the number of parameters of database instances reaches more than 700. In addition, users also have the need for independent tuning of database parameters. The process of issuing independently tuned database parameters involves multiple parameter synchronizations, which easily leads to abnormal database behavior. The superposition of multiple factors greatly increases the operation and maintenance difficulty of database parameters, posing a greater challenge to the database parameter anomaly governance solution. Summary of the Invention
[0004] Multiple aspects of this application provide a database parameter anomaly governance system, method, electronic device, and storage medium to provide a better database parameter anomaly governance solution.
[0005] An embodiment of the present application provides a database parameter anomaly governance system, including: a plurality of data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly patrol inspection service, and a plurality of parameter repair agents deployed on multiple database instances; the data collection agents are used to collect real-time parameter data of the database instances in real time and report it to the first real-time message queue for storage; the parameter anomaly online analysis service is used to obtain the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and take a snapshot to obtain a real-time parameter snapshot at the most recent collection time; call a big data processing component to perform parameter anomaly event analysis on the real-time parameter snapshot at the most recent collection time, and send the parameter anomaly event analysis result to the parameter anomaly patrol inspection service, and the parameter anomaly event analysis result includes the abnormal database instance and its abnormal database parameters; the parameter anomaly patrol inspection service is used to make a decision according to the parameter anomaly event analysis result to obtain a first decision result, and the first decision result includes the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and send the first decision result to the parameter repair agent deployed on the abnormal database instance; the parameter repair agent deployed on the abnormal database instance is used to hot-fix the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
[0006] An embodiment of the present application also provides a database parameter anomaly governance method, which is applied to a database parameter anomaly governance system. The system includes: a plurality of data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly patrol inspection service, and a plurality of parameter repair agents deployed on multiple database instances; the method includes: the data collection agents collect real-time parameter data of the database instances in real time and report it to the first real-time message queue for storage; the parameter anomaly online analysis service obtains the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and takes a snapshot to obtain a real-time parameter snapshot at the most recent collection time; call a big data processing component to perform parameter anomaly event analysis on the real-time parameter snapshot at the most recent collection time, and send the parameter anomaly event analysis result to the parameter anomaly patrol inspection service, and the parameter anomaly event analysis result includes the abnormal database instance and its abnormal database parameters; the parameter anomaly patrol inspection service makes a decision according to the parameter anomaly event analysis result to obtain a first decision result, and the first decision result includes the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and sends the first decision result to the parameter repair agent deployed on the abnormal database instance; the parameter repair agent deployed on the abnormal database instance hot-fixes the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
[0007] An embodiment of the present application further provides an electronic device, including: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to execute the steps in the database parameter exception governance method.
[0008] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the database parameter exception governance method.
[0009] An embodiment of the present application provides a database parameter exception governance system, method, electronic device and storage medium. Among them, the database parameter exception governance system includes: a plurality of data collection agents deployed in multiple database instances, a first real-time message queue, a parameter exception online analysis service, a parameter exception patrol service, and a plurality of parameter repair agents deployed in multiple database instances. The data collection agents transparently transmit the collected real-time parameter data to the first real-time message queue with low latency. Each data collection agent does not interfere with each other and has high robustness, which can greatly improve the timeliness and accuracy of real-time parameter data. The parameter exception online analysis service calls the big data processing component to analyze the parameter exception events of the real-time parameter snapshot at the most recent collection time, and sends the analysis results of the parameter exception events to the parameter exception patrol service for immediate decision-making, and the decision results are sent to the parameter repair agent. The parameter repair agent performs repairs in a hot repair manner based on the decision results. Thus, a new database parameter exception governance solution is provided. This solution reduces the operation and maintenance difficulty of database parameters, discovers abnormal database parameters efficiently and accurately, greatly shortens the survival time of abnormal database parameters, and improves the stability of database instances. Further optionally, the database parameter exception governance system also uses the parameter exception offline analysis service to intelligently analyze the parameter exception problems with large data volume and complex multiple data sources by combining the offline full-scale data obtained by integrating data of multiple dimensions, and uses big data computing power to support parameter exception detection. The parameter exception online analysis service is responsible for real-time discovery of abnormal database parameters, and the parameter exception offline analysis service serves as a backup for the exception detection of offline full-scale data. The combination of the parameter exception online analysis service and the parameter exception offline analysis service can discover more comprehensive abnormal database parameters and conduct more comprehensive governance, and basically will not miss the discovery of abnormal database parameters. Description of the Drawings
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0011] Figure 1 is an architecture diagram of a traditional database parameter exception governance solution;
[0012] Figure 2 It is an architecture diagram of a database parameter anomaly governance system provided by an embodiment of the present application;
[0013] Figure 3 It is an exemplary application scenario diagram provided by an embodiment of the present application;
[0014] Figure 4 It is a process diagram of exemplary online analysis and offline analysis provided by an embodiment of the present application;
[0015] Figure 5 It is a flowchart of a database parameter anomaly governance method provided by an embodiment of the present application;
[0016] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0017] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0018] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the access relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural. In the written description of the present application, the character " / " generally represents an "or" relationship between the associated objects before and after. In addition, in the embodiments of the present application, "first", "second", "third", etc. are only used to distinguish the contents of different objects and have no other special meanings.
[0019] Some terms related to the present application are introduced below:
[0020] Database instance: It can be an independent database service process that occupies physical memory, and different memory sizes, disk spaces, and database types can be set for the database instance.
[0021] Database parameters refer to the parameters that affect database performance and can also be understood as the instance parameters of a database instance. For example, but not limited to: Buffer Pool size (cache pool size), query_cache_size (query cache size), max_user_connections (maximum number of user connections), max_connections (maximum number of connections), etc. In this embodiment, the parameter values of database parameters can be divided into required parameter values, configured parameter values, and real-time parameter values. The required parameter value refers to the parameter value set by the user as needed, that is, the parameter value that meets the user's needs; the configured parameter value refers to the parameter value written into the configuration file of the database instance; the real-time parameter value refers to the parameter value of the database parameters generated during the operation of the database instance collected in real time. In practical applications, some database parameters need to restart the database instance to take effect. Generally, they are first sent to the configuration file, and then the database instance is restarted, so as to load the configured parameter values of the database parameters in the configuration file into the memory to form the effective operation state parameters. The operation state parameters refer to the database parameters with real-time parameter values. That is to say, some real-time parameter values are obtained after the configured parameter values in the configuration file are loaded into the memory and take effect. The real-time parameter value is the effective configured parameter value.
[0022] Big data processing component: It can be any streaming processing and batch processing framework. For example, but not limited to: Flink streaming computing platform, which is used for distributed computing and processing of real-time data streams and large-scale data batches.
[0023] Big data development and governance platform: It is a data governance platform that provides services such as data integration, data development, data map, data quality, and data services. For example, but not limited to: DataWorks big data governance platform. DataWorks is based on multiple big data engines and provides a unified full-link big data development and governance platform for solutions such as data warehouses, data lakes, and lakehouse integration.
[0024] Offline data warehouse: It refers to a data warehouse that stores offline data. After data from different data sources undergoes various processes such as cleaning, transformation, and processing, it is stored in the offline data warehouse as offline data. The offline data in the offline data warehouse is used for subsequent data analysis or mining. With the help of the offline data warehouse, data processing and storage are separated, the load of a single system is reduced, the efficiency and accuracy of data processing are improved, and better decision support is provided.
[0025] Figure 1 It is an architecture diagram of a traditional database parameter anomaly governance solution. See Figure 1In the traditional database parameter anomaly management solution, the required parameter values of the database parameters configured by the user on demand are stored in the metadata database. The centralized inspection service links each database instance. The centralized inspection service perceives the real-time parameter values of the database parameters of each database instance, and compares the perceived real-time parameter values of the database parameters with the required parameter values of the database parameters in the metadata database. According to the comparison results, the abnormal database instance and its abnormal database parameters are found, and the abnormal database parameters of the abnormal database instance are repaired. When the number of database parameters expands or the number of database instances expands, the traditional solution will experience a huge delay, resulting in a prolonged existence time of database parameter anomalies, which can easily cause online failures of the database system. In addition, when dealing with complex parameter anomaly problems, the centralized inspection service has limited ability to perceive multiple application data and bottlenecks in computing power, and cannot perform complex batch calculations on all database instances.
[0026] In traditional solutions, when the repair system repairs the abnormal database parameters of the abnormal database instance, it often obtains the required parameter values of the database parameters from the metadata database, and sends the required parameter values of the database parameters to the configuration file of the database instance. If the values are successfully sent to the configuration file, the required parameter values are converted into configuration parameter values. The repair system first restarts the abnormal database instance, and loads the configuration parameter values of the database parameters obtained from the configuration file into the memory to take effect as the real-time parameter values of the database parameters, thereby completing the repair of the abnormal database parameters of the abnormal database instance. However, sending the required parameter values of the database parameters in the metadata database to the configuration file of the database instance often involves multiple complex parameter synchronizations, which can easily lead to inconsistencies between the configuration parameter values and real-time parameter values of the database parameters and the required parameter values in the metadata database.
[0027] In summary, traditional database parameter anomaly management solutions will more or less miss some abnormal database parameters and fail to find them or have difficulty in accurately repairing the parameter values of database parameters.
[0028] In the daily operation and maintenance of cloud database systems, timely detection and repair of abnormal database parameters is an important operation and maintenance task. With the rapid development of cloud database systems, some cloud database systems have more than 1 million database instances, and the number of database instance parameters has reached more than 700. In addition, users also have the need to self-tune database parameters. The process of issuing self-tuned database parameters involves multiple parameter synchronizations, which can easily lead to abnormal database behavior. The superposition of multiple factors has greatly increased the difficulty of database parameter operation and maintenance, and posed greater challenges to the database parameter anomaly management solution.
[0029] To this end, the embodiments of the present application provide a database parameter anomaly governance system, method, electronic device, and storage medium. Among them, the database parameter anomaly governance system includes: multiple data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly patrol inspection service, and multiple parameter repair agents deployed on multiple database instances. The data collection agents transmit the collected real-time parameter data to the first real-time message queue in a low-latency manner. Each data collection agent does not interfere with each other and has high robustness, which can greatly improve the timeliness and accuracy of real-time parameter data. The parameter anomaly online analysis service calls the big data processing component to analyze the parameter anomaly events in the real-time parameter snapshot of the most recent collection time, and sends the analysis results of the parameter anomaly events to the parameter anomaly patrol inspection service for immediate decision-making. The decision-making results are sent to the parameter repair agent, and the parameter repair agent performs repairs in a hot repair manner based on the decision-making results. Thus, a new database parameter anomaly governance solution is provided. This solution reduces the operation and maintenance difficulty of database parameters, discovers abnormal database parameters efficiently and accurately, greatly shortens the survival time of abnormal database parameters, and improves the stability of database instances. Further optionally, the database parameter anomaly governance system also uses the parameter anomaly offline analysis service to intelligently analyze the parameter anomaly problems with large data volume and complex multiple data sources by combining the offline full-scale data obtained by integrating data of multiple dimensions, and uses big data computing power to support parameter anomaly detection. The parameter anomaly online analysis service is responsible for real-time discovery of abnormal database parameters, and the parameter anomaly offline analysis service serves as a backup for the anomaly detection of offline full-scale data. The combination of the parameter anomaly online analysis service and the parameter anomaly offline analysis service can discover more comprehensive abnormal database parameters and conduct more comprehensive governance, and basically no abnormal database parameters will be missed in discovery.
[0030] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0031] Figure 2 It is an architecture diagram of a database parameter anomaly governance system provided by an embodiment of the present application. Refer to Figure 2 This system may include: multiple data collection agents 10 deployed on multiple database instances, a first real-time message queue 20, a parameter anomaly online analysis service 30, a parameter anomaly patrol inspection service 40, and multiple parameter repair agents 50 deployed on multiple database instances.
[0032] In this embodiment, the data collection agent 10 is used to collect the real-time parameter data of the database instance in real time and report it to the first real-time message queue for storage.
[0033] Specifically, the data collection agent 10 is a network agent tool with data collection capabilities. One data collection agent 10 responsible for real-time parameter data collection is deployed for each database instance. The data collection agent 10 can collect the real-time parameter data of the database instance it is deployed on in real time according to the set sampling period. The real-time parameter data includes not only the real-time parameter values of at least one database parameter generated during the operation of the database instance, but also the configuration parameter values of at least one database parameter obtained from the configuration file of the database instance.
[0034] Multiple data collection agents 10 distributed across multiple different database instances form a distributed data collection system. Compared with the method of a centralized inspection service alone centrally perceiving data of multiple database instances, the perception requirements of real-time parameter data are dispersed to each data collection agent 10 in the distributed data collection system. The data collection agent 10 transmits the collected real-time parameter data to the first real-time message queue with low latency. Each data collection agent 10 does not interfere with each other and has high robustness, which can greatly improve the timeliness and accuracy of real-time parameter data.
[0035] In this embodiment, the real-time parameter data of the database instance collected by the data collection agent 10 in real time is reported to the first real-time message queue for storage by the first real-time message queue. The first real-time message queue can be any message queue, such as including but not limited to: RocketMQ. RocketMQ is a distributed message queue system with low latency, high reliability, and strong scalability, and supports features such as distributed deployment, horizontal expansion, and disaster recovery.
[0036] In this embodiment, the parameter anomaly online analysis service 30 is used to obtain and snapshot the real-time parameter data of at least one database instance reported in the most recent collection time in the first real-time message queue to obtain a real-time parameter snapshot at the most recent collection time; call the big data processing component to analyze the parameter anomaly events for the real-time parameter snapshot at the most recent collection time, and send the parameter anomaly event analysis results to the parameter anomaly inspection service 40. The parameter anomaly event analysis results include the abnormal database instance and its abnormal database parameters.
[0037] Specifically, the online parameter anomaly analysis service 30 provides a service for analyzing database parameter anomalies in an online analysis manner, which belongs to an online analysis service. The real-time parameter data reported by the data collection agent 10 to the first real-time message queue also includes the collection time of the real-time parameter data. First, the online parameter anomaly analysis service 30 obtains the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time, and the most recent collection time is the collection time closest to the current time. Optionally, the online parameter anomaly analysis service 30 can perform data cleaning on the obtained real-time parameter data to improve data quality. Then, the online parameter anomaly analysis service 30 takes a snapshot of the real-time parameter data of at least one database instance obtained, obtaining a real-time parameter snapshot at the most recent collection time. The real-time parameter snapshot at the most recent collection time includes the real-time parameter data of at least one database instance reported at the most recent collection time. Then, the online parameter anomaly analysis service 30 calls the big data processing component to perform parameter anomaly event analysis on the real-time parameter snapshot at the most recent collection time. Invoking the big data processing component for online analysis can improve the timeliness and accuracy of the online analysis results.
[0038] The online parameter anomaly analysis service 30 performs parameter anomaly event analysis based on the database parameters. The recommended parameter value of the database parameter is the parameter value of the database parameter that can ensure the normal operation of the database instance. The real-time parameter value of the database parameter in the real-time parameter snapshot is compared with the corresponding recommended parameter value. If the real-time parameter value of the database parameter in the real-time parameter snapshot is consistent with the corresponding recommended parameter value, it is confirmed that the database parameter is normal; if the real-time parameter value of the database parameter in the real-time parameter snapshot is inconsistent with the corresponding recommended parameter value, it is confirmed that the database parameter is abnormal. In practical applications, if the difference between the real-time parameter value of the database parameter in the real-time parameter snapshot and the corresponding recommended parameter value falls within the specified numerical range, the real-time parameter value of the database parameter in the real-time parameter snapshot is consistent with the corresponding recommended parameter value; if the difference between the real-time parameter value of the database parameter in the real-time parameter snapshot and the corresponding recommended parameter value does not fall within the specified numerical range, the real-time parameter value of the database parameter in the real-time parameter snapshot is inconsistent with the corresponding recommended parameter value.
[0039] In this embodiment, the parameter anomaly event analysis result output by the online parameter anomaly analysis service 30 includes the abnormal database instance and its abnormal database parameter. The abnormal database instance refers to the database instance with abnormal database parameters; the abnormal database parameter refers to the database parameter with abnormal parameter values, that is, the real-time parameter value of the database parameter of the abnormal database parameter is inconsistent with the corresponding recommended parameter value.
[0040] In this embodiment, the online analysis service 30 for parameter anomalies sends the analysis results of parameter anomaly events to the parameter anomaly patrol service 40 for immediate decision-making by the parameter anomaly patrol service 40, greatly shortening the survival time of abnormal database parameters and improving the stability of the database instance.
[0041] In this embodiment, the parameter anomaly patrol service 40 is used to make a decision based on the analysis results of parameter anomaly events to obtain a first decision result. The first decision result includes the normal parameter values corresponding to the abnormal database parameters of the abnormal database instance, and the first decision result is sent to the parameter repair agent 50 deployed on the abnormal database instance.
[0042] Specifically, the parameter anomaly patrol service 40 is a patrol service that can decide the normal parameter values that can repair the abnormal database parameters to normal database parameters. The normal parameter values can be understood as the parameter values of the database parameters when the database instance is running normally. The parameter anomaly patrol service 40 can make a decision based on the analysis results of parameter anomaly events, and the obtained decision result is the first decision result. In practical applications, the parameter anomaly patrol service 40 can notify the operation and maintenance personnel of the parameter anomaly event, and the parameter anomaly patrol service 40 obtains the first decision result of the operation and maintenance personnel's decision. Or, the parameter anomaly patrol service 40 runs a pre-trained machine learning model with a decision-making function, and the machine learning model decides the first decision result, which is not limited here. Or, the parameter anomaly patrol service 40 selects the recommended parameter value corresponding to the abnormal database parameter of the abnormal database instance from the pre-stored recommended parameter values of each database parameter as the corresponding normal parameter value, and generates the first decision result based on the selected normal parameter value. Of course, the decision-making method of the parameter anomaly patrol service 40 is not limited.
[0043] The parameter anomaly patrol service 40 sends the first decision result to the parameter repair agent 50 deployed on the abnormal database instance, so that the parameter repair agent 50 can repair the abnormal database parameters of the abnormal database instance.
[0044] In this embodiment, the parameter repair agent 50 deployed on the abnormal database instance is used to hot-fix the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
[0045] Specifically, the parameter repair agent 50 is a network proxy tool with a data repair function. Each database instance deploys a parameter repair agent 50 responsible for parameter repair tasks. Multiple parameter repair agents 50 distributed on multiple different database instances form a distributed data repair system. Compared with the method of only a centralized patrol service centrally repairing the parameters of multiple database instances, it can greatly improve the timeliness and accuracy of parameter repair.
[0046] The parameter repair agent 50 performs repair in a hot repair manner, so that the parameter repair of the database instance can be completed during the operation of the database instance, which improves the parameter repair efficiency and ensures the stability of the database instance.
[0047] The database parameter anomaly governance system provided by the embodiments of the present application includes: a plurality of data collection agents 10 deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service 30, a parameter anomaly inspection service 40, and a plurality of parameter repair agents 50 deployed on multiple database instances. The data collection agent 10 transmits the collected real-time parameter data to the first real-time message queue in a low-latency manner. Each data collection agent 10 does not interfere with each other and has high robustness, which can greatly improve the timeliness and accuracy of the real-time parameter data. The parameter anomaly online analysis service 30 calls the big data processing component to analyze the parameter anomaly events of the real-time parameter snapshot at the most recent collection time, and sends the analysis results of the parameter anomaly events to the parameter anomaly inspection service 40 for immediate decision-making, and the decision-making results are sent to the parameter repair agent 50, and the parameter repair agent 50 performs repair in a hot repair manner based on the decision-making results. Thus, a new database parameter anomaly governance solution is provided. This solution reduces the operation and maintenance difficulty of database parameters, discovers abnormal database parameters efficiently and accurately, greatly shortens the survival time of abnormal database parameters, and improves the stability of the database instance.
[0048] In some optional embodiments, the parameter anomaly online analysis service 30 further has a parameter change event analysis function. Refer to Figure 2 , the database parameter anomaly governance system may further include: a second real-time message queue 60. The second real-time message queue 60 can be any message queue, for example, including but not limited to: RocketMQ.
[0049] In this embodiment, the parameter anomaly online analysis service 30 is further configured to report the real-time parameter snapshot at the most recent collection time to the second real-time message queue 60 for storage; obtain the real-time parameter snapshot at the previous collection time from the second real-time message queue 60; call the big data processing component to perform parameter change event analysis based on the real-time parameter snapshots at the most recent collection time and the previous collection time, and send the analysis results of the parameter change events to the parameter anomaly inspection service 40. The analysis results of the parameter change events include the changed database instance and its changed database parameters.
[0050] Specifically, the second real-time message queue 60 stores the real-time parameter snapshots at each collection time. The parameter anomaly online analysis service 30 reports the real-time parameter snapshot at the most recent collection time to the second real-time message queue 60 for storage.
[0051] The online analysis service for parameter anomalies 30 analyzes parameter change events based on real-time parameter snapshots at two adjacent acquisition times. When analyzing parameter change events, for the same database parameter of the same database instance, it is judged whether the real-time parameter value of this database parameter in the real-time parameter snapshot at the most recent acquisition time is the same as the real-time parameter value of this database parameter in the real-time parameter snapshot at the previous acquisition time. If the difference between the real-time parameter value of this database parameter in the real-time parameter snapshot at the most recent acquisition time and the real-time parameter value of this database parameter in the real-time parameter snapshot at the previous acquisition time falls within the specified data range, the judgment result is consistent, and it is confirmed that no parameter change event has occurred; if the difference between the real-time parameter value of this database parameter in the real-time parameter snapshot at the most recent acquisition time and the real-time parameter value of this database parameter in the real-time parameter snapshot at the previous acquisition time does not fall within the specified data range, the judgment result is inconsistent, and it is confirmed that a parameter change event has occurred.
[0052] The online analysis service for parameter anomalies 30 sends the analysis result of the parameter change event to the parameter anomaly inspection service 40. The analysis result of the parameter change event includes the changed database instance and its changed database parameter. The changed database instance is the database instance where the parameter change event occurs, and the changed database parameter is the database parameter where the parameter change event occurs.
[0053] The parameter anomaly inspection service 40 is used to make a decision based on the analysis result of the parameter change event to obtain a second decision result. The second decision result includes the normal parameter value corresponding to the changed database parameter of the changed database instance, and the decision result is sent to the parameter repair agent 50 deployed on the changed database instance.
[0054] Specifically, the parameter anomaly inspection service 40 can make a decision based on the analysis result of the parameter change event, and the obtained decision result is the second decision result. In practical applications, the parameter anomaly inspection service 40 can notify the operation and maintenance personnel of the parameter change event, and the parameter anomaly inspection service 40 obtains the second decision result of the operation and maintenance personnel's decision. Or, the parameter anomaly inspection service 40 runs a pre-trained machine learning model with a decision-making function, and the machine learning model makes the second decision result. Or, the parameter anomaly inspection service 40 selects the recommended parameter value corresponding to the abnormal database parameter of the abnormal database instance from the recommended parameter values of each database parameter stored in advance as the corresponding normal parameter value, and generates the second decision result based on the selected normal parameter value. Of course, the decision-making method of the parameter anomaly inspection service 40 is not limited.
[0055] The parameter anomaly inspection service 40 sends the second decision result to the parameter repair agent 50 deployed on the changed database instance, so that the parameter repair agent 50 can repair the changed database parameter of the changed database instance.
[0056] In this embodiment, the parameter repair agent 50 deployed on the changed database instance is used to hot-fix the real-time parameter value of the changed database parameter of the changed database instance to the corresponding normal parameter value.
[0057] It should be noted that the database parameter anomaly governance system also has the ability to analyze parameter change events, can instantly repair database parameters, and improve the effect of database parameter anomaly governance.
[0058] In some alternative embodiments, refer to Figure 2 , the database parameter anomaly governance system may further include: a parameter snapshot retrieval service 70.
[0059] Specifically, the parameter snapshot retrieval service 70 is a retrieval service that can retrieve real-time parameter snapshots in the second real-time message queue 60. The parameter snapshot retrieval service 70 can also provide a parameter snapshot retrieval API (Application Programming Interface) for retrieval.
[0060] In practical applications, during the decision-making process of the parameter anomaly inspection service 40, if it is necessary to obtain more real-time parameter snapshots to assist in decision-making or view more real-time parameter snapshots, the parameter snapshot retrieval service 70 is requested to retrieve real-time parameter snapshots that meet the retrieval conditions.
[0061] Optionally, the parameter anomaly inspection service 40 is used to send a retrieval request including a retrieval time range to the parameter snapshot retrieval service 70, and receive the real-time parameter snapshots that fall within the retrieval time range returned by the parameter snapshot retrieval service 70. The parameter snapshot retrieval service 70 is used to obtain the real-time parameter snapshots that fall within the retrieval time range from the second real-time message queue 60 according to the retrieval request.
[0062] Specifically, the retrieval condition refers to retrieving real-time parameter snapshots that fall within the retrieval time range. Since the real-time parameter snapshots in the second real-time message queue 60 are stored at the time granularity of the collection time, the parameter snapshot retrieval service 70 can accurately retrieve the real-time parameter snapshots that fall within the retrieval time range.
[0063] In some alternative embodiments, the database parameter anomaly governance system may further include: a meta-database 80, an offline data warehouse 90, and a parameter anomaly offline analysis service 100.
[0064] In this embodiment, the meta-database 80 is used to store demand parameter data that meets user requirements.
[0065] Specifically, the demand parameter data includes the demand parameter values of at least one database parameter of at least one database instance, and the demand parameter values are the parameter values of the database parameters configured by the user according to needs. The demand parameter values of the database parameters configured by the user are first saved to the meta-database 80, and over time, the demand parameter values of the database parameters in the meta-database 80 are read and written into the configuration file of the database instance to be converted into configuration parameter values.
[0066] In this embodiment, the offline data warehouse 90 is used to store offline full-volume data. The offline full-volume data includes, for example, but is not limited to at least one of the following: real-time parameter snapshots of at least one collection time obtained from the first real-time message queue, demand parameter data obtained from the meta-database 80, parameter value suggestion tables, key database parameter lists, instance performance data, and instance log data.
[0067] In this embodiment, the parameter value suggestion table includes the suggested parameter values of at least one database parameter; the key database parameter list includes at least one key parameter. The key parameter refers to a database parameter that has a greater impact on the database performance and is flexibly selected according to needs. The instance performance data refers to the data related to the performance of the database instance, and includes, for example, but is not limited to: disk usage rate, memory usage rate, and CPU (Central Processing Unit) usage rate. The instance log data includes, for example, but is not limited to: instance error log data and instance slow log data. The instance error log data is the log data recording various errors that occur in the instance, and includes, for example, but is not limited to: instance creation failure information, instance startup failure information, or instance connection failure information. The instance slow log data refers to the slow query log (Slow Log), which is used to record the commands in the database instance whose execution time exceeds the specified threshold.
[0068] In this embodiment, the parameter anomaly offline analysis service 100 is used to call the big data development and governance platform to perform complex correlation analysis on the offline full-volume data in the offline data warehouse 90 to obtain the offline analysis result; and send the offline analysis result to the parameter anomaly inspection service 40.
[0069] Specifically, the parameter anomaly offline analysis service 100 is an offline analysis service that performs complex correlation analysis on the offline full-volume data by leveraging the big data mining and analysis capabilities of the big data development and governance platform. The complex correlation analysis can be understood as performing correlation analysis on the complex offline full-volume data. The offline analysis result includes, for example, but is not limited to at least one of the following:
[0070] ①. Master-slave database parameter inconsistency details list. Among them, the master-slave database parameter inconsistency details list includes the detailed information of at least one database parameter. The detailed information of the database parameter includes: the master database instance and the standby database instance associated with the database parameter, and the real-time parameter value of the database parameter of the master database instance is inconsistent with the real-time parameter value of the database parameter of the standby database instance.
[0071] ②. Configuration parameter value and real-time parameter value inconsistency details list. Among them, the configuration parameter value and real-time parameter value inconsistency details list includes the detailed information of at least one database parameter. The detailed information of the database parameter includes: the real-time parameter value of the database parameter and its configuration parameter value in the configuration file, and the real-time parameter value of the database parameter is inconsistent with its configuration parameter value in the configuration file.
[0072] ③. Requirement parameter value and real-time parameter value inconsistency details list. Among them, the requirement parameter value and real-time parameter value inconsistency details list includes the detailed information of at least one database parameter. The detailed information of the database parameter includes: the requirement parameter value and the real-time parameter value of the database parameter, and the requirement parameter value of the database parameter is inconsistent with the real-time parameter value.
[0073] ④. Instance performance index exceeds recommended parameter value details list. The instance performance index exceeds recommended parameter value details list includes at least one instance performance index that exceeds the corresponding recommended parameter value.
[0074] In this embodiment, the parameter anomaly inspection service 40 is used to make a decision based on the offline analysis result to obtain a third decision result. The third decision result includes the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and the decision result is sent to the parameter repair agent 50 deployed on the abnormal database instance.
[0075] Specifically, the parameter anomaly inspection service 40 can make a decision based on the offline analysis result, and the obtained decision result is the third decision result. In practical applications, the parameter anomaly inspection service 40 can notify the operation and maintenance personnel of a parameter change event, and the parameter anomaly inspection service 40 obtains the third decision result of the operation and maintenance personnel's decision. Or, the parameter anomaly inspection service 40 runs a pre-trained machine learning model with a decision-making function, and the machine learning model makes the third decision result. Or, the parameter anomaly inspection service 40 selects the recommended parameter value corresponding to the abnormal database parameter of the abnormal database instance from the recommended parameter values of each database parameter stored in advance as the corresponding normal parameter value, and generates the third decision result based on the selected normal parameter value. Of course, the decision-making method of the parameter anomaly inspection service 40 is not limited.
[0076] The parameter anomaly inspection service 40 sends the third decision result to the parameter repair agent 50 deployed on the anomaly database instance, so that the parameter repair agent 50 repairs the abnormal database parameters of the anomaly database instance.
[0077] In this embodiment, the parameter repair agent 50 deployed on the anomaly database instance is used to hot-fix the real-time parameter value of the abnormal database parameters of the anomaly database instance to the corresponding normal parameter value.
[0078] In this embodiment, through the parameter anomaly offline analysis service 100, combined with the offline full-scale data obtained by integrating data of multiple dimensions, intelligent analysis is performed on the parameter anomaly problems with large data volume and complex multiple data sources, and big data computing power is used to support parameter anomaly detection. The parameter anomaly online analysis service 30 is responsible for real-time discovery of abnormal database parameters, and the parameter anomaly offline analysis service 100 serves as a backup for the anomaly detection of offline full-scale data. The combination of the parameter anomaly online analysis service 30 and the parameter anomaly offline analysis service 100 can discover more comprehensive abnormal database parameters and conduct more comprehensive governance, and basically no abnormal database parameters will be missed.
[0079] The following combines Figure 3 and Figure 4 to introduce a specific database parameter anomaly governance solution.
[0080] Specifically, a data collection agent and a parameter repair agent are deployed on each database instance, and the data collection agents distributed on multiple database instances form a distributed data collection system; the parameter repair agents distributed on multiple database instances form a distributed data repair system.
[0081] In this embodiment, each data collection agent is responsible for real-time collection of the real-time parameter data of the database instance on which it is deployed, and the real-time parameter data includes at least the real-time parameter value and / or configuration parameter value of one database parameter. As shown in ① of Figure 3 , the real-time parameter data of the database instance collected by the data collection agent in real time is reported to the first real-time message queue. As shown in ② of Figure 3 , the parameter anomaly online analysis service obtains the real-time parameter data in the first real-time message queue for online analysis and outputs the online analysis result. As shown in Figure 4 , the online analysis result includes, for example, but is not limited to: parameter anomaly event analysis result and parameter change event analysis result. As shown in Figure 4 , the parameter anomaly online analysis service includes: parameter anomaly analysis task, transaction issuance task, parameter snapshot aggregation service, rule configuration center, parameter anomaly event retrieval API, algorithm center, and so on.
[0082] Among them, the parameter anomaly analysis task is used to collect real-time parameter data from the first real-time message queue for data cleaning, and call the parameter snapshot aggregation service to perform snapshot processing on the cleaned real-time parameter data to obtain a real-time parameter snapshot; and perform online analysis based on the real-time parameter snapshot.
[0083] The transaction issuance task is used to issue the analysis result of the parameter anomaly event or the analysis result of the parameter change event to the parameter anomaly patrol service in the form of an event.
[0084] The parameter snapshot aggregation service is used to aggregate the real-time parameter data reported by all data collection agents within the time granularity into a snapshot data;
[0085] The rule configuration center is responsible for maintaining the generation rules of events such as parameter anomaly events and parameter update events, and issuing rules for different usage scenarios.
[0086] The parameter anomaly event retrieval API is used to retrieve the analysis results of parameter anomaly events that meet the user's needs.
[0087] The algorithm center is a common library that precipitates event generation and analysis algorithms, and provides a unified SDK (Software Development Kit) for online analysis.
[0088] In this embodiment, as shown in Figure 3 ③, the parameter anomaly online analysis service sends the online analysis result to the parameter anomaly patrol service. The parameter anomaly patrol service makes a decision and outputs a decision result, and the decision result includes the normal parameter value corresponding to the database parameter of the database instance that needs to be repaired.
[0089] In this embodiment, as shown in Figure 3 ④, the parameter anomaly patrol service sends the decision result to the corresponding parameter repair agent, so that the parameter repair agent performs hot repair on the database parameter of the database instance that needs to be repaired based on the normal parameter value in the decision result. Thus, based on the mutual cooperation of the parameter anomaly online analysis service and the parameter anomaly patrol service, abnormal database parameters are discovered in real time and repaired in time.
[0090] As shown in Figure 3 ⑤, the parameter anomaly online analysis service can also report the real-time parameter snapshot to the second real-time message queue for storage. As shown in Figure 3 ⑥ and ⑦, the parameter anomaly patrol service can also send a retrieval request to the parameter snapshot retrieval service, and the parameter snapshot retrieval service retrieves the real-time parameter snapshot that meets the retrieval conditions in the second real-time message queue and returns it to the parameter anomaly patrol service for the parameter anomaly patrol service to analyze and use.
[0091] In this embodiment, as shown in Figure 3 , the parameter anomaly offline analysis service obtains the offline full data from the offline data warehouse. As shown in Figure 4 , the offline full data includes, for example, but is not limited to: real-time parameter snapshots, required parameter data, parameter value suggestion tables, key database parameter lists, instance performance data, and instance log data, etc. The offline analysis results output by the parameter anomaly offline analysis service include, for example, but are not limited to: the detailed list of inconsistent primary and standby database parameters, the detailed list of inconsistent configuration parameter values and real-time parameter values, the detailed list of inconsistent required parameter values and real-time parameter values, and the detailed list of instance performance metrics exceeding the suggested parameter values.
[0092] In this embodiment, as shown in Figure 3 , the parameter anomaly offline analysis service sends the offline analysis results to the parameter anomaly inspection service for decision-making and outputs the decision results. The decision results include the normal parameter values corresponding to the database parameters of the database instances that need to be repaired. The parameter anomaly inspection service sends the decision results to the corresponding parameter repair agent, so that the parameter repair agent performs hot repair on the database parameters of the database instances that need to be repaired based on the normal parameter values in the decision results. Thus, based on the mutual cooperation of the parameter anomaly offline analysis service and the parameter anomaly inspection service, it provides a backup for the anomaly detection of the offline full data.
[0093] The database parameter anomaly governance solution provided by the embodiments of this application uses multiple data collection agents and multiple parameter repair agents deployed on different database instances to perform data collection and data repair respectively, and the Agents with different functions operate independently of each other, with high robustness.
[0094] The data collection agent collects the real-time parameter data of the database instance with low latency to achieve maximum timeliness and accuracy; the centralized data perception requirements are dispersed to each data collection agent, and the centralized parameter anomaly inspection service no longer undertakes the data perception task and only serves as the decision-maker for the database parameter anomaly governance, reducing the load pressure on the parameter anomaly inspection service. Especially in the face of the increasing number of database instances, it improves the inspection efficiency for abnormal database parameters.
[0095] The online analysis service for parameter anomalies can call the big data processing component for online analysis, which greatly shortens the survival time of abnormal database parameters online and improves the stability of database instances. The offline analysis service for parameter anomalies can call the big data development and governance platform for offline analysis. During offline analysis, it combines the offline full-scale data obtained by integrating data from multiple dimensions to perform intelligent analysis on the parameter anomaly problems with large data volumes and complex data sources from multiple data sources, and uses big data computing power to support parameter anomaly detection. The online analysis service for parameter anomalies is responsible for real-time discovery of abnormal database parameters, and the offline analysis service for parameter anomalies serves as a backup for anomaly detection of offline full-scale data. The combination of the online analysis service for parameter anomalies and the offline analysis service for parameter anomalies can discover more comprehensive abnormal database parameters and conduct more comprehensive governance, with basically no omission of abnormal database parameters, realizing full-process and multi-dimensional support for the parameter governance of cloud databases (especially distributed cloud databases).
[0096] Figure 5 It is a flowchart of a method for governing database parameter anomalies provided by an embodiment of this application. It is applied to a database parameter anomaly governance system, which includes: multiple data collection agents deployed on multiple database instances, a first real-time message queue, an online analysis service for parameter anomalies, a parameter anomaly inspection service, and multiple parameter repair agents deployed on multiple database instances;
[0097] See Figure 5 , the method for governing database parameter anomalies may include the following steps:
[0098] 101. The data collection agent collects real-time parameter data of the database instance in real time and reports it to the first real-time message queue for storage.
[0099] 102. The online analysis service for parameter anomalies obtains the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and takes a snapshot to obtain a real-time parameter snapshot at the most recent collection time; calls the big data processing component to analyze parameter anomaly events for the real-time parameter snapshot at the most recent collection time, and sends the parameter anomaly event analysis results to the parameter anomaly inspection service. The parameter anomaly event analysis results include the abnormal database instance and its abnormal database parameters;
[0100] 103. The parameter anomaly inspection service makes a decision based on the parameter anomaly event analysis results to obtain a first decision result. The first decision result includes the normal parameter values corresponding to the abnormal database parameters of the abnormal database instance, and sends the first decision result to the parameter repair agent deployed on the abnormal database instance;
[0101] 104. The parameter repair agent deployed on the abnormal database instance hot-fixes the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
[0102] Further optionally, the above system further includes: a second real-time message queue; the above method further includes: the parameter anomaly online analysis service reporting the real-time parameter snapshot at the most recent collection time to the second real-time message queue for storage; obtaining the real-time parameter snapshot at the previous collection time from the second real-time message queue; calling the big data processing component to perform parameter change event analysis based on the real-time parameter snapshots at the most recent collection time and the previous collection time, and sending the parameter change event analysis result to the parameter anomaly patrol service, the parameter change event analysis result including the changed database instance and its changed database parameter; the parameter anomaly patrol service making a decision according to the parameter change event analysis result to obtain a second decision result, the second decision result including the normal parameter value corresponding to the changed database parameter of the changed database instance, and sending the decision result to the parameter repair agent deployed on the changed database instance; the parameter repair agent deployed on the changed database instance hot-fixing the real-time parameter value of the changed database parameter of the changed database instance to the corresponding normal parameter value.
[0103] Further optionally, the above system further includes: a parameter snapshot retrieval service; the above method further includes: the parameter anomaly patrol service sending a retrieval request including a retrieval time range to the parameter snapshot retrieval service, and receiving the real-time parameter snapshot falling within the retrieval time range returned by the parameter snapshot retrieval service; the parameter snapshot retrieval service obtaining the real-time parameter snapshot falling within the retrieval time range from the second real-time message queue according to the retrieval request.
[0104] Further optionally, the above system further includes: a meta database, an offline data warehouse, and a parameter anomaly offline analysis service; the above method further includes: the meta database storing the demand parameter data that meets the user's requirements; the offline data warehouse storing the offline full-volume data; the parameter anomaly offline analysis service calling the big data development and governance platform to perform complex correlation analysis on the offline full-volume data in the offline data warehouse to obtain an offline analysis result; and sending the offline analysis result to the parameter anomaly patrol service; the parameter anomaly patrol service making a decision according to the offline analysis result to obtain a third decision result, the third decision result including the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and sending the decision result to the parameter repair agent deployed on the abnormal database instance; the parameter repair agent deployed on the abnormal database instance hot-fixing the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value.
[0105] Further optionally, the offline full data includes at least two of the following: a real-time parameter snapshot of at least one collection time obtained from the first real-time message queue, demand parameter data obtained from the metadata database, a parameter value recommendation table, a key database parameter list, instance performance data, and instance log data, wherein the parameter value recommendation table includes a recommended parameter value of at least one database parameter, and the key database parameter list includes at least one key parameter.
[0106] Further optionally, the offline analysis result includes at least one of the following: a master-standby database parameter inconsistency detailed table, wherein the master-standby database parameter inconsistency detailed table includes detailed information of at least one database parameter, and the detailed information of the database parameter includes: a master database instance and a standby database instance associated with the database parameter, and the real-time parameter value of the database parameter of the master database instance is inconsistent with the real-time parameter value of the database parameter of the standby database instance; a configuration parameter value and real-time parameter value inconsistency detailed table, wherein the configuration parameter value and real-time parameter value inconsistency detailed table includes detailed information of at least one database parameter, and the detailed information of the database parameter includes: the real-time parameter value of the database parameter and its configuration parameter value in the configuration file, and the real-time parameter value of the database parameter and its configuration parameter value in the configuration file are inconsistent; a requirement parameter value and real-time parameter value inconsistency detailed table, wherein the requirement parameter value and real-time parameter value inconsistency detailed table includes detailed information of at least one database parameter, and the detailed information of the database parameter includes: the requirement parameter value and the real-time parameter value of the database parameter, and the requirement parameter value and the real-time parameter value of the database parameter are inconsistent;
[0107] The detailed list of instance performance indicators exceeding recommended parameter values includes at least one instance performance indicator exceeding a corresponding recommended parameter value.
[0108] Figure 5 The detailed implementation process of each step in the method shown in the embodiment can be found in the relevant description in the aforementioned system embodiment, which will not be repeated here.
[0109] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0110] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 6 As shown, the electronic device includes: a memory 61 and a processor 62;
[0111] A memory 61 for storing computer programs and configurable to store various other data to support operations on a computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.
[0112] The memory 61 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disks or optical discs.
[0113] A processor 62, coupled to the memory 61, for executing the computer programs in the memory 61 to: perform the steps in the method for managing database parameter anomalies.
[0114] Further optionally, as Figure 6 shown, the electronic device further includes: a communication component 63, a display 64, a power supply component 65, an audio component 66, and other components. Figure 6 Only some components are schematically shown, and it does not mean that the electronic device only includes Figure 6 the components shown. Additionally, Figure 6 the components within the dashed box are optional components, not mandatory components, and can be determined according to the product form of the electronic device. The electronic device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an Internet of Things (IoT) device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the electronic device in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it can include Figure 6 the components within the dashed box; if the electronic device in this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include Figure 6 the components within the dashed box.
[0115] For the detailed implementation process of the processor to execute each action, reference may be made to the relevant descriptions in the foregoing method embodiments or device embodiments, which will not be elaborated herein.
[0116] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step executable by an electronic device in the foregoing method embodiment.
[0117] Correspondingly, an embodiment of the present application further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the processor can implement each step executable by an electronic device in the foregoing method embodiment.
[0118] The foregoing communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi (Wireless Fidelity), 2G (2 Generation), 3G (3 Generation), 4G (4 Generation) / LTE (long Term Evolution), 5G (5 Generation) and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on technologies such as Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.
[0119] The foregoing display includes a screen, and the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0120] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0121] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0122] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0123] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0124] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.
[0126] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and a memory.
[0127] The memory may include non-permanent memory in the form of computer-readable media, random access memory (Random Access Memory, RAM) and / or non-volatile memory, such as read-only memory (Read Only Memory, ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0128] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (Phase Change RAM, PRAM), static random access memory (Static Random-Access Memory, SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), other types of random access memory (Random Access Memory, RAM), read-only memory (Read Only Memory, ROM), electrically erasable programmable read-only memory (Electrically-Erasable Programmable Read-Only Memory, EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (Digital versatile disc, DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0129] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0130] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A database parameter anomaly governance system, characterized in that, Including: Multiple data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly patrol service, and multiple parameter repair agents deployed on multiple database instances; The data collection agent is used to collect the real-time parameter data of the database instance in real time and report it to the first real-time message queue for storage; The parameter anomaly online analysis service is used to obtain the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and take a snapshot to obtain a real-time parameter snapshot at the most recent collection time; call a big data processing component to perform parameter anomaly event analysis on the real-time parameter snapshot at the most recent collection time, and send the parameter anomaly event analysis result to the parameter anomaly patrol service, where the parameter anomaly event analysis result includes the abnormal database instance and its abnormal database parameters; The parameter anomaly patrol service is used to make a decision based on the parameter anomaly event analysis result to obtain a first decision result, where the first decision result includes the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and send the first decision result to the parameter repair agent deployed on the abnormal database instance; The parameter repair agent deployed on the abnormal database instance is used to hot-fix the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
2. The system according to claim 1, wherein Also including: A second real-time message queue; The parameter anomaly online analysis service is also used to report the real-time parameter snapshot at the most recent collection time to the second real-time message queue for storage; Obtain the real-time parameter snapshot at the previous collection time from the second real-time message queue; Call the big data processing component to perform parameter change event analysis based on the real-time parameter snapshots at the most recent collection time and the previous collection time, and send the parameter change event analysis result to the parameter anomaly patrol service, where the parameter change event analysis result includes the changed database instance and its changed database parameters; The parameter anomaly patrol service is used to make a decision based on the parameter change event analysis result to obtain a second decision result, where the second decision result includes the normal parameter value corresponding to the changed database parameter of the changed database instance, and send the decision result to the parameter repair agent deployed on the changed database instance; The parameter repair agent deployed on the changed database instance is used to hot-fix the real-time parameter value of the changed database parameter of the changed database instance to the corresponding normal parameter value.
3. The system according to claim 2, wherein Also including: A parameter snapshot retrieval service; The parameter anomaly patrol service is used to send a retrieval request including a retrieval time range to the parameter snapshot retrieval service and receive the real-time parameter snapshot that falls within the retrieval time range returned by the parameter snapshot retrieval service; The parameter snapshot retrieval service is used to obtain the real-time parameter snapshot that falls within the retrieval time range from the second real-time message queue according to the retrieval request.
4. The system according to claim 1, wherein Also including: A meta-database, an offline data warehouse, and a parameter anomaly offline analysis service; The metadata database is used to store requirement parameter data that meets user requirements; The offline data warehouse is used to store offline full-volume data; The parameter anomaly offline analysis service is used to call the big data development and governance platform to perform complex correlation analysis on the offline full-volume data in the offline data warehouse to obtain an offline analysis result; and send the offline analysis result to the parameter anomaly inspection service; The parameter anomaly inspection service is used to make a decision based on the offline analysis result to obtain a third decision result, where the third decision result includes the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and send the decision result to the parameter repair agent deployed on the abnormal database instance; The parameter repair agent deployed on the abnormal database instance is used to hot-fix the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value.
5. The system according to claim 4, wherein The offline full-volume data includes at least two of the following: real-time parameter snapshots at at least one collection time obtained from the first real-time message queue, requirement parameter data obtained from the metadata database, a parameter value suggestion table, a list of key database parameters, instance performance data, and instance log data, where the parameter value suggestion table includes suggested parameter values of at least one database parameter, and the list of key database parameters includes at least one key parameter.
6. The system according to claim 4, wherein The offline analysis result includes at least one of the following: A master-slave database parameter inconsistency detail list, where the master-slave database parameter inconsistency detail list includes detail information of at least one database parameter, and the detail information of the database parameter includes: the master database instance and the standby database instance associated with the database parameter, and the real-time parameter value of the database parameter of the master database instance is inconsistent with the real-time parameter value of the database parameter of the standby database instance; A configuration parameter value and real-time parameter value inconsistency detail list, where the configuration parameter value and real-time parameter value inconsistency detail list includes detail information of at least one database parameter, and the detail information of the database parameter includes: the real-time parameter value of the database parameter and its configuration parameter value in the configuration file, and the real-time parameter value of the database parameter is inconsistent with its configuration parameter value in the configuration file; A requirement parameter value and real-time parameter value inconsistency detail list, where the requirement parameter value and real-time parameter value inconsistency detail list includes detail information of at least one database parameter, and the detail information of the database parameter includes: the requirement parameter value and the real-time parameter value of the database parameter, and the requirement parameter value of the database parameter is inconsistent with the real-time parameter value; An instance performance index exceeding the suggested parameter value detail list, and the instance performance index exceeding the suggested parameter value detail list includes at least one instance performance index exceeding the corresponding suggested parameter value.
7. A method for managing database parameter anomalies, characterized in that, Applied to a database parameter anomaly governance system, which includes: multiple data collection agents deployed on multiple database instances, a first real-time message queue, a parameter anomaly online analysis service, a parameter anomaly inspection service, and multiple parameter repair agents deployed on multiple database instances; the method includes: The data collection agent collects the real-time parameter data of the database instance in real time and reports it to the first real-time message queue for storage; The parameter anomaly online analysis service obtains the real-time parameter data of at least one database instance reported in the first real-time message queue at the most recent collection time and takes a snapshot to obtain the real-time parameter snapshot at the most recent collection time; calls the big data processing component to analyze parameter anomaly events for the real-time parameter snapshot at the most recent collection time, and sends the parameter anomaly event analysis results to the parameter anomaly patrol service, where the parameter anomaly event analysis results include the abnormal database instance and its abnormal database parameters; The parameter anomaly patrol service makes a decision based on the parameter anomaly event analysis results to obtain a first decision result, where the first decision result includes the normal parameter values corresponding to the abnormal database parameters of the abnormal database instance, and sends the first decision result to the parameter repair agent deployed on the abnormal database instance; The parameter repair agent deployed on the abnormal database instance hot-fixes the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value in the first decision result.
8. The method according to claim 7, wherein The system further includes: a second real-time message queue; the method further includes: The parameter anomaly online analysis service reports the real-time parameter snapshot at the most recent collection time to the second real-time message queue for storage; obtains the real-time parameter snapshot at the previous collection time from the second real-time message queue; calls the big data processing component to analyze parameter change events based on the real-time parameter snapshots at the most recent collection time and the previous collection time, and sends the parameter change event analysis results to the parameter anomaly patrol service, where the parameter change event analysis results include the changed database instance and its changed database parameters; The parameter anomaly patrol service makes a decision based on the parameter change event analysis results to obtain a second decision result, where the second decision result includes the normal parameter values corresponding to the changed database parameters of the changed database instance, and sends the decision result to the parameter repair agent deployed on the changed database instance; The parameter repair agent deployed on the changed database instance hot-fixes the real-time parameter value of the changed database parameter of the changed database instance to the corresponding normal parameter value.
9. The method according to claim 8, wherein The system further includes: a parameter snapshot retrieval service; The parameter anomaly patrol service sends a retrieval request including a retrieval time range to the parameter snapshot retrieval service, and receives the real-time parameter snapshots falling within the retrieval time range returned by the parameter snapshot retrieval service; The parameter snapshot retrieval service obtains the real-time parameter snapshots falling within the retrieval time range from the second real-time message queue according to the retrieval request.
10. The method according to claim 7, characterized in that, The system further includes: a meta-database, an offline data warehouse, and a parameter anomaly offline analysis service; The meta-database stores the demand parameter data that meets the user's requirements; The offline data warehouse stores the offline full-scale data; The parameter anomaly offline analysis service calls the big data development and governance platform to perform complex correlation analysis on the offline full-volume data in the offline data warehouse, and obtains an offline analysis result; and sends the offline analysis result to the parameter anomaly inspection service; The parameter anomaly inspection service is used to make a decision based on the offline analysis result to obtain a third decision result, where the third decision result includes the normal parameter value corresponding to the abnormal database parameter of the abnormal database instance, and sends the decision result to the parameter repair agent deployed on the abnormal database instance; The parameter repair agent deployed on the abnormal database instance hot-fixes the real-time parameter value of the abnormal database parameter of the abnormal database instance to the corresponding normal parameter value.
11. An electronic device, characterized in that, Comprising: A memory and a processor; The memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to perform the steps in the method according to any one of claims 7-10.
12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it causes the processor to be able to implement the steps in the method according to any one of claims 7-10.