Database cluster log collection method, system and device and medium
Through the automated log collection method, the problem of low efficiency in database cluster failure analysis is solved, and efficient and accurate log collection and fault analysis are achieved.
Patent Information
- Application Number
- CN202510669033.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The prior art is less efficient when analyzing the cause of failure of database clusters, especially when the number of database nodes is large, it is necessary to manually view the log files of each database node, which is time-consuming and inefficient.
Provides a database cluster log collection method, which automatically collects log files of database nodes to be checked by obtaining the types of troubleshooting, configuration file analysis and log collection component calls.
It realizes automatic log collection in the event of database cluster failure, improves log collection efficiency and accuracy, thereby improving the efficiency and accuracy of failure cause analysis.
Smart Images

Figure CN120196669A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of database technologies, and in particular, to a method, system, device, and medium for collecting database cluster logs. Background Art
[0002] A database cluster usually includes multiple database nodes. There is one or more database instances in each database node, and one database instance corresponds to one or more log files. In the prior art, when a database cluster fails, it is usually necessary to manually view the log files of all database instances of multiple database nodes in the database cluster for failure cause analysis. When the number of database nodes is large, manually viewing the log files of all database instances of each database node takes a lot of time, resulting in low efficiency in analyzing the failure causes of the database cluster. Summary of the Invention
[0003] This application provides a method, system, device, and medium for collecting database cluster logs to solve the problem of low efficiency in analyzing the failure causes of a database cluster in the prior art. Among them, the technical solutions provided by this application are as follows: On the one hand, this application provides a method for collecting database cluster logs, including: Obtain the type of fault to be investigated for the database cluster; Based on the first configuration file, according to the type of fault to be investigated, determine the information of the database nodes to be investigated corresponding to the database cluster and the information of the log types to be investigated corresponding to the information of the database nodes to be investigated; wherein, the first configuration file is configured with the information of the database nodes corresponding to different fault types and the information of the log types corresponding to the information of the database nodes. Based on the second configuration file, according to the information of the log types to be investigated corresponding to the information of the database nodes to be investigated, determine the information of the target log collection components corresponding to the information of the database nodes to be investigated; wherein, the second configuration file is configured with the information of the log collection components corresponding to different log type information. Call the target log collection component corresponding to the information of the target log collection component, and based on the information of the log types to be investigated, collect the log files of the database nodes to be investigated represented by the information of the database nodes to be investigated.
[0004] Optionally, based on the first configuration file, according to the type of fault to be investigated, determining the information of the database nodes to be investigated corresponding to the database cluster and the information of the log types to be investigated corresponding to the information of the database nodes to be investigated includes: Call the parameter parsing component to parse out the information of the database nodes corresponding to the type of fault to be investigated and the information of the log types corresponding to the information of the database nodes from the first configuration file; Determine the parsed database node information and the log type information corresponding to the database node information as the database node information to be troubleshot corresponding to the database cluster and the log type information to be troubleshot corresponding to the database node information to be troubleshot.
[0005] Optionally, based on the second configuration file, determine the target log collection component information corresponding to the database node information to be troubleshot according to the log type information to be troubleshot corresponding to the database node information to be troubleshot, including: Call the parameter parsing component to parse the log collection component information corresponding to the log type information to be troubleshot from the second configuration file; Determine the parsed log collection component information as the target log collection component information corresponding to the database node information to be troubleshot.
[0006] Optionally, call the target log collection component corresponding to the target log collection component information, and based on the log type information to be troubleshot, collect the log files of the database node to be troubleshot represented by the database node information to be troubleshot, including: Call the task generation component to generate a log collection task based on the database node information and the log type information to be troubleshot corresponding to the database node information; Call the task execution component to execute the log collection task, and during the execution of the log collection task, call the target log collection component corresponding to the target log collection component information to collect the log files corresponding to the log type information to be troubleshot from the database node to be troubleshot represented by the database node information to be troubleshot.
[0007] Optionally, call the task generation component to generate a log collection task based on the database node information and the log type information to be troubleshot corresponding to the database node information, including: Call the task generation component to obtain the user login account information, and generate a log collection task based on the user login account information, the database node information, and the log type information to be troubleshot corresponding to the database node information.
[0008] Optionally, after calling the target log collection component corresponding to the target log collection component information to collect the log files corresponding to the log type information to be troubleshot from the database node to be troubleshot represented by the database node information to be troubleshot, it further includes: Call the task execution component to save the log files to the temporary output directory of the database node to be troubleshot represented by the database node information to be troubleshot, and perform compression processing on all the log files in the temporary output directory to obtain a log file compression package; When it is determined that the log file download condition is met, the log file compressed package is downloaded from the temporary output directory to the local output directory, and the log file compressed package in the local output directory is decompressed to obtain the log file of the database node to be troubleshot represented by the database node information to be troubleshot.
[0009] Optionally, after calling the target log collection component corresponding to the target log collection component information and collecting the log file of the database node to be troubleshot represented by the database node information to be troubleshot based on the log type information to be troubleshot, it further includes: Generate a log collection result statistical report based on the component execution log during the log file collection process, and output the log collection result statistical report.
[0010] On the other hand, the present application provides a database cluster log collection system, including: A database cluster fault monitoring component for obtaining the fault type to be troubleshot of the database cluster; A database parameter acquisition component for determining the database node information to be troubleshot corresponding to the database cluster and the log type information to be troubleshot corresponding to the database node information to be troubleshot according to the fault type to be troubleshot based on the first configuration file; wherein, the first configuration file is configured with the database node information corresponding to different fault types and the log type information corresponding to the database node information. A tool parameter acquisition component for determining the target log collection component information corresponding to the database node information to be troubleshot according to the log type information to be troubleshot corresponding to the database node information to be troubleshot based on the second configuration file; wherein, the second configuration file is configured with the log collection component information corresponding to different log type information. A log collection execution component for calling the target log collection component corresponding to the target log collection component information and collecting the log file of the database node to be troubleshot represented by the database node information to be troubleshot based on the log type information to be troubleshot.
[0011] On the other hand, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned database cluster log collection method is implemented.
[0012] On the other hand, the present application provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions. When the computer instructions are executed by a processor, the above-mentioned database cluster log collection method is implemented.
[0013] The beneficial effects of the present application are as follows: Based on the first configuration file, this application determines the information of the database nodes to be troubleshot for the types of faults to be troubleshot and the information of the log types corresponding to the information of the database nodes to be troubleshot. After determining the information of the target log collection component corresponding to the information of the database nodes to be troubleshot based on the second configuration file, it calls the target log collection component corresponding to the information of the target log collection component, and collects the log files of the database nodes to be troubleshot represented by the information of the database nodes to be troubleshot based on the information of the log types to be troubleshot, which can realize the automatic log collection during the failure of the database cluster, thereby improving the efficiency and accuracy of log collection during the failure of the database cluster, and further improving the efficiency and accuracy of analyzing the cause of the failure of the database cluster.
[0014] Other features and advantages of this application will be described in the following description, and some of them can be made obvious from the description, or understood by implementing this application. The objectives and other advantages of this application can be achieved and obtained through the structures specifically pointed out in the written description and the drawings. Brief Description of the Drawings
[0015] The drawings described herein are used to provide a further understanding of this application, and constitute a part of this application. The schematic diagrams and descriptions thereof are used to explain this application and do not constitute an improper limitation to this application. In the drawings: Figure 1 It is a schematic diagram of the overall framework of the database cluster log collection method in the embodiment of this application; Figure 2 It is a schematic diagram of the general process of the database cluster log collection method in the embodiment of this application; Figure 3 It is a schematic diagram of the model structure of the fault node association prediction model in the embodiment of this application; Figure 4 It is a schematic diagram of the specific process of the database cluster log collection method in the embodiment of this application; Figure 5 It is a schematic diagram of the composition structure of the database cluster log collection system in the embodiment of this application; Figure 6 It is a schematic diagram of the hardware structure of the electronic device in the embodiment of this application. Detailed Embodiments
[0016] In order to make the objectives, technical solutions and beneficial effects of this application clearer and more understandable, the technical solutions of this application will be clearly and completely described below in conjunction with the embodiments and the drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application.
[0017] For better understanding of the present application by those skilled in the art, the technical terms involved in the present application are briefly introduced below.
[0018] Database node information refers to node attribute information such as node identification information, network address information, database version, and database deployment form corresponding to the database node.
[0019] Log type information refers to log attribute information such as the log type of the log to be troubleshot and the log storage address corresponding to the database node. In the present application, the logs to be troubleshot corresponding to database nodes with different deployment forms and / or different versions are different. For example, the logs to be troubleshot corresponding to the master-slave database node include dn logs and cm logs, etc., and the logs to be troubleshot corresponding to the distributed database node include dn logs, cn logs, gtm logs, and etcd logs, etc.
[0020] Log collection component information refers to tool attribute information such as the component name and call method of the log collection component.
[0021] The first configuration file is a file configured with database node information corresponding to different failure types and log type information corresponding to the database node information.
[0022] The second configuration file is a file configured with log collection component information corresponding to different log type information.
[0023] It should be noted that the "first", "second", etc. mentioned in the present application are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the "and / or" mentioned in the present application describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the front and back associated objects.
[0024] After introducing the technical terms involved in the present application, next, the application scenario and design concept of the present application are briefly introduced.
[0025] Currently, in the fault location of the database cluster, usually, the personnel need to manually log in to each database node of the database cluster and enter the directories of multiple database instances under each database node to view the log files, so as to perform fault location and fault cause analysis. For such log file viewing operations under the fault conditions of the database cluster, when the number of database nodes exceeds 3, the efficiency of viewing the log files drops rapidly, thus affecting the efficiency of fault location and fault cause analysis of the database cluster.
[0026] To this end, in this application, refer to Figure 1 As shown, the first stage is the configuration parameter stage, including: (1) Configure the first configuration file (such as the cluster information in Figure 1 ): The first configuration file at least includes the database version, the first correspondence between the database deployment form and the log type information and the log collection range; among them, the log collection range at least includes the failure type, and may also include one of the log file time range and the percentage of the number of log files in the total number; that is, the log file time range and the percentage of the number of log files in the total number can be selected for one configuration or multiple combinations of configurations; (2) Configure the second configuration file (such as the collection content in Figure 1 ): The second configuration file at least includes the second correspondence between the log type information and the log collection component information; among them, the log type information at least includes the following log types and log attribute information such as the log storage address of each log type: the log under the main directory of the database log; transaction log information; serial number log information for recording transaction submission; operating system log information; stack information after program crash; database main configuration file; database client authentication configuration file; main process file of the database; current running status information of the database; data table information, including database, system table, system view, ordinary table and ordinary view; directory where the main process file of the database is located; correspondingly, the log collection component information at least includes the following component names and tool attribute information such as the call method of the log collection component represented by each component name: database log collection component; database transaction log collection component; database serial number collection component for recording transactions; operating system log collection component; program stack information file component; database configuration file collection component; database client verification configuration file; system database and business database collection component; system table and business table collection component; system view and business view collection component; database executable program collection component; (3) Configure the user login account information (such as the node information in Figure 1 ): If collecting operating system-related log content, configure the operating system super administrator user name and password as the user login account information. If collecting database-related log content, configure the database installation administrator user name and password as the user login account information; (4) Configure the output path (such as Figure 1The output paths in it are: the temporary output directory and the local output directory. The second stage is the log collection stage, including: calling the parameter parsing component to obtain the tool parameters from the second configuration file and the database parameters from the first configuration file; calling the task generation component to generate a task list, where each task in the task list contains information such as the directory where the logs need to be collected, the node network address, and the login account; using the task execution component to execute each task in the task list in a multi-threaded manner to collect the log files of the database nodes, generating a statistical report on the log collection results based on the component execution logs during the log file collection process, and outputting the statistical report on the log collection results, thereby realizing the automatic log collection during the failure of the database cluster, improving the log collection efficiency during the failure of the database cluster, and further improving the efficiency of fault location and fault cause analysis of the database cluster.
[0027] After introducing the application scenario and design concept of this application, the technical solutions provided by this application will be described in detail below.
[0028] An embodiment of this application provides a method for collecting database cluster logs, which is applied to electronic devices such as computers and servers. Refer to Figure 2 As shown, the general process of the method for collecting database cluster logs provided by the embodiment of this application is as follows: Step 201: Obtain the type of fault to be investigated for the database cluster.
[0029] In the embodiment of this application, when the database cluster failure monitoring component monitors that the database cluster fails, obtain the failure code of the database cluster, and determine the type of fault to be investigated based on the failure code of the database cluster. For example, use the failure code of the database cluster as the type of fault to be investigated.
[0030] Step 202: Based on the first configuration file, determine the information of the database nodes to be investigated corresponding to the database cluster and the information of the log types to be investigated corresponding to the information of the database nodes to be investigated according to the type of fault to be investigated; where, in the first configuration file, the database node information corresponding to different fault types and the log type information corresponding to the database node information are configured.
[0031] In the embodiment of this application, the first configuration file is pre-generated based on the database node information corresponding to different fault types and the log type information corresponding to the database node information. The database nodes corresponding to different fault types and the log types corresponding to the database nodes can be predicted based on the fault node association prediction model. Specifically, when predicting the correspondence between the fault types, database nodes, and the log types of the database nodes in the database cluster based on the fault node association prediction model, the following methods can be used but are not limited to: Step 1: Collect each historical fault case, and perform text recognition and feature extraction on each historical fault case to obtain each training sample; where each training sample includes a historical fault type, an associated database node identifier, and historical feature data corresponding to the associated database node identifier for the historical fault type. Specifically, it includes: (1) Collect each historical fault case of the database cluster; where each historical fault case includes a fault type label (such as lock contention, query timeout, node downtime, etc.), an associated database node identifier, database node log text, and performance metrics (CPU, memory, network, etc.) during the fault period.
[0032] (2) Based on NLP technology, extract the associated database node identifier and the corresponding historical log data, historical fault data, and historical feature data from each historical fault case; where the historical log data includes slow query logs, error logs, transaction logs, lock wait logs, etc.; the historical fault data includes error codes, exception stacks, SQL statement patterns, etc. in the database node log text; the historical feature data includes structured features and topological features. The structured features include time series data of performance metrics such as resource utilization rate, transaction throughput, and lock wait time of the database node, and the topological features include a database node dependency relationship graph, and the database node dependency relationship graph includes call link weights between database nodes.
[0033] (3) Align the associated database node identifier with the historical log data, historical fault data, and historical feature data according to the timestamp to obtain each training sample.
[0034] Step 2: Based on each training sample, perform iterative training on the fault node association prediction model; where the fault node association prediction model includes a fault classification sub-model, a node localization sub-model, a log association sub-model, and a fusion analysis sub-model as shown in Figure 3 the following. Specifically, it includes: (1) Through the fault classification sub-model, using a supervised learning network (such as XGBoost, LSTM, etc.), predict the probability distribution of the fault type based on the historical log data and historical feature data corresponding to the associated database node identifier, and output the predicted fault type based on the probability distribution of the fault type; for example, when it is detected that "Deadlock detected" frequently appears in the historical log data of a certain associated database node and the lock wait time exceeds the threshold, it is predicted as a "deadlock fault".
[0035] (2)Through the node localization sub-model, using a graph neural network (such as GNN, etc.), based on the historical feature data and historical fault data corresponding to the associated database node identifiers, identify the associated database nodes that caused historical faults, that is, the first mapping relationship between the fault type and the associated database nodes; for example, by analyzing the sudden increase trend of the connection numbers between each associated database node and the query queue depth, locate the hot node with the highest load as the associated database node of the "query peak" fault.
[0036] (3)Through the log association sub-model, using a language analysis network (such as BERT, etc.), based on the historical log data and historical fault data corresponding to the associated database node identifiers, predict the second mapping relationship between the fault type and the required logs, where the second mapping relationship between the fault type and the required logs is shown in Table 1: Table 1.
[0037] (4)Through the fusion analysis sub-model, using a multi-modal fusion analysis network (such as a knowledge graph reasoning network, a multi-modal joint analysis network based on the attention mechanism, a dual-mode fusion network of a knowledge graph and the attention mechanism, an ensemble learning network, etc.), perform fusion analysis on the predicted fault type output by the fault classification sub-model, the first mapping relationship between the fault type and the associated database nodes output by the node localization sub-model, and the second mapping relationship between the fault type and the required logs output by the log association sub-model, to obtain the associated database nodes corresponding to the fault type and the log types of the associated database nodes. In the embodiments of the present application, the multi-modal fusion analysis network is preferably a dual-mode fusion network of a knowledge graph and the attention mechanism. For example, the multi-modal fusion analysis network constructs a triple of a knowledge graph with the fault type, database nodes, and log types, dynamically calculates the weights of the features of each modality through a cross-attention mechanism (such as focusing on the node localization result in the scenario where the fault type confidence is high), and generates a joint probability distribution of the fused fault type, database nodes, and log types through a graph traversal algorithm (such as bidirectional breadth-first search) to determine the complete association chain.
[0038] (5)Based on the associated database node identifiers and historical log data corresponding to the historical fault data in the training samples, as well as the associated database nodes corresponding to the predicted fault type and the log types of the associated database nodes, update the parameters of the fault node association prediction model until the iteration termination condition is met, and then determine the fault node association prediction model with the last parameter update as the trained fault node association prediction model.
[0039] Step 3: Perform fault node association prediction based on the trained fault node association prediction model to obtain the database nodes and log types corresponding to different fault types.
[0040] Step 4: Generate a first configuration file based on the database nodes and log types corresponding to different fault types; for example: (deadlock fault, associated node, gbase node) → (analysis basis, required logs, lock wait logs).
[0041] Furthermore, when determining the information of the database nodes to be investigated and the information of the log types corresponding to the database nodes to be investigated for the database cluster based on the first configuration file according to the fault type to be investigated, the following methods can be used but are not limited to: First, call the parameter parsing component to parse out the information of the database nodes corresponding to the fault type to be investigated and the information of the log types corresponding to the database node information from the first configuration file.
[0042] Then, determine the information of the database nodes to be investigated for the database cluster and the information of the log types corresponding to the database nodes to be investigated as the information of the database nodes to be investigated and the information of the log types corresponding to the database nodes to be investigated parsed out.
[0043] Step 203: Based on the second configuration file, determine the information of the target log collection component corresponding to the information of the database nodes to be investigated according to the information of the log types corresponding to the information of the database nodes to be investigated; wherein, the information of the log collection components corresponding to different log type information is configured in the second configuration file.
[0044] In the embodiment of the present application, the second configuration file can be a file configured and generated manually; when determining the information of the target log collection component corresponding to the information of the database nodes to be investigated according to the information of the log types corresponding to the information of the database nodes to be investigated based on the second configuration file, the following methods can be used but are not limited to: First, call the parameter parsing component to parse out the information of the log collection component corresponding to the information of the log type to be investigated from the second configuration file.
[0045] Then, determine the information of the parsed log collection component as the information of the target log collection component corresponding to the information of the database nodes to be investigated.
[0046] Step 204: Call the target log collection component corresponding to the information of the target log collection component, and collect the log files of the database nodes to be investigated represented by the information of the database nodes to be investigated based on the information of the log types to be investigated.
[0047] In the embodiment of the present application, when calling the target log collection component corresponding to the information of the target log collection component and collecting the log files of the database nodes to be investigated represented by the information of the database nodes to be investigated based on the information of the log types to be investigated, the following methods can be used but are not limited to: First, call the task generation component to generate a log collection task based on the database node information and the log type information to be investigated corresponding to the database node information. Specifically, call the task generation component to obtain the user login account information, and generate a log collection task based on the user login account information, the database node information, and the log type information to be investigated corresponding to the database node information.
[0048] Then, call the task execution component to execute the log collection task, and during the execution of the log collection task, call the target log collection component corresponding to the target log collection component information to collect the log files corresponding to the log type information to be investigated from the database node to be investigated represented by the database node information to be investigated. Specifically, first, call the task generation component to generate a log collection task list. Each log collection task in the log collection task list includes information such as the log type information to be investigated (such as the log directory to be collected), the database node information (such as the node network address), and the user login account information; then, call the task execution component to use multi-threading to parallelly execute each log collection task in the log collection task list; among them, when each thread executes the log collection task, based on the user login account information, log in to the database node represented by the database node information, and based on the log type information to be investigated, collect the log files of the database node from the logged-in database node.
[0049] In addition, in the embodiment of the present application, when calling the target log collection component corresponding to the target log collection component information to collect the log files corresponding to the log type information to be investigated from the database node to be investigated represented by the database node information to be investigated, the time range of the log files and / or the percentage of the number of log files in the total number parsed from the first configuration file can also be combined. Specifically, first, call the task generation component to generate a log collection task list; each log collection task in the log collection task list includes information such as the log type information to be investigated (such as the log directory to be collected), the database node information (such as the node network address), the log collection range information (such as the time range of the log files and / or the percentage of the number of log files in the total number), and the user login account information; call the task execution component to use multi-threading to parallelly execute each log collection task in the log collection task list; among them, when each thread executes the log collection task, based on the user login account information, log in to the database node represented by the database node information, and based on the log type information to be investigated, collect the log files of the database node from the logged-in database node according to the log collection range information (such as the time range of the log files and / or the percentage of the number of log files in the total number).
[0050] In the implementation of this application, in order to further improve the collection efficiency of log files, after calling the target log collection component corresponding to the target log collection component information to collect the log files corresponding to the to-be-investigated log type information from the to-be-investigated database nodes characterized by the to-be-investigated database node information, the task execution component can also be called first to save the log files to the temporary output directory of the to-be-investigated database node characterized by the to-be-investigated database node information, and compress all the log files in the temporary output directory to obtain a log file compression package. Then, when it is determined that the log file download condition is met, the log file compression package is downloaded from the temporary output directory to the local output directory, and the log file compression package in the local output directory is decompressed to obtain the log files of the to-be-investigated database node characterized by the to-be-investigated database node information; wherein, the log file download condition includes at least one of the following conditions: reaching the set download period (for example, downloading once every 2 minutes), receiving a download instruction, and all the log files of all the to-be-investigated log types of all the to-be-investigated database nodes corresponding to the to-be-investigated fault type have been collected. In this way, by first saving the log files to the temporary output directory and then compressing the log files through a file compression tool, the execution efficiency of the log collection task can be improved while reducing the file size of the log file transmission, thereby saving data transmission resources and improving the collection efficiency of the log files.
[0051] Furthermore, after calling the target log collection component corresponding to the target log collection component information to collect the log files of the to-be-investigated database node characterized by the to-be-investigated database node information based on the to-be-investigated log type information, a log collection result statistical report can also be generated based on the component execution logs during the log file collection process, and the log collection result statistical report can be output to facilitate subsequent tracing of the log file collection process and facilitate the optimization and upgrade of the database cluster log collection system.
[0052] Next, the database cluster log collection method provided by the embodiments of this application will be further described in detail. Refer to Figure 4 As shown, the specific process of the database cluster log collection method provided by the embodiments of this application is as follows: Step 301: When a failure of the database cluster is detected by the failure monitoring component, obtain the failure code of the database cluster, and determine the to-be-investigated failure type based on the failure code of the database cluster.
[0053] Step 302: Call the parameter parsing component, and parse out the to-be-investigated database node information corresponding to the database cluster and the to-be-investigated log type information corresponding to the to-be-investigated database node information from the first configuration file according to the to-be-investigated failure type, and parse out the target log collection component information corresponding to the to-be-investigated database node information from the second configuration file according to the to-be-investigated log type information corresponding to the to-be-investigated database node information.
[0054] Step 303: Invoke the task generation component to generate a log collection task based on the user login account information, database node information, and the log type information to be investigated corresponding to the database node information.
[0055] Step 304: Invoke the task execution component to execute the log collection task. During the execution of the log collection task, invoke the target log collection component corresponding to the target log collection component information to collect the log files corresponding to the log type information to be investigated from the database node to be investigated represented by the database node information to be investigated.
[0056] Step 305: Invoke the task execution component to save the log files collected by the target log collection component to the temporary output directory of the database node to be investigated represented by the database node information to be investigated, and compress all the log files in the temporary output directory to obtain a log file compression package.
[0057] Step 306: Invoke the task execution component. When it is determined that all the log files of all the log types to be investigated corresponding to all the database nodes to be investigated for the fault type to be investigated have been collected, download the log file compression packages from the temporary output directories of each database node to be investigated to the local output directory, and decompress each log file compression package in the local output directory to obtain all the log files of all the database nodes to be investigated corresponding to the fault type to be investigated.
[0058] Step 307: Invoke the task execution component to generate a log collection result statistical report based on the component execution logs during the log file collection process, and output the log collection result statistical report.
[0059] In the database cluster log collection provided by the embodiments of the present application, by configuring log collection range options, such as fault type, log file time range, and the percentage of the number of log files in a folder to the total number, etc., log collection can be performed according to fault requirements, thereby improving the flexibility of log collection. Moreover, by first saving the log files to the temporary output directory and then compressing the log files using a file compression tool, the execution efficiency of the log collection task can be improved while reducing the file size of the log file transmission, and thus data transmission resources can be saved and the collection efficiency of the log files can be improved. In addition, during the fault location of the database cluster, it is possible to view all the log files of all the database nodes related to the fault type in one stop, thereby improving the fault location efficiency and providing effective data support for database operation and maintenance.
[0060] Based on the above embodiments, the embodiments of the present application provide a database cluster log collection system. Refer to Figure 5As shown in the figure, the database cluster log collection system 400 provided by the embodiment of the present application at least includes: A database cluster fault monitoring component 401, configured to obtain the fault types to be investigated of the database cluster; A database parameter acquisition component 402, configured to determine the database node information to be investigated corresponding to the database cluster and the log type information corresponding to the database node information to be investigated according to the fault types to be investigated based on the first configuration file; wherein, the database node information corresponding to different fault types and the log type information corresponding to the database node information are configured in the first configuration file; A tool parameter acquisition component 403, configured to determine the target log collection component information corresponding to the database node information to be investigated according to the log type information corresponding to the database node information to be investigated based on the second configuration file; wherein, the log collection component information corresponding to different log type information is configured in the second configuration file; A log collection execution component 404, configured to call the target log collection component 405 corresponding to the target log collection component information, and collect the log files of the database node to be investigated represented by the database node information to be investigated based on the log type information to be investigated.
[0061] In a possible implementation manner, the database parameter acquisition component 402 is configured to parse out the database node information corresponding to the fault type to be investigated and the log type information corresponding to the database node information from the first configuration file; and determine the parsed database node information and the log type information corresponding to the database node information as the database node information to be investigated corresponding to the database cluster and the log type information to be investigated corresponding to the database node information to be investigated.
[0062] In a possible implementation manner, the tool parameter acquisition component 403 is configured to parse out the log collection component information corresponding to the log type information to be investigated from the second configuration file; and determine the parsed log collection component information as the target log collection component information corresponding to the database node information to be investigated.
[0063] In a possible implementation manner, the log collection execution component 404 includes a task generation component 4041 and a task execution component 4042; The task generation component 4041 is configured to generate a log collection task based on the database node information and the log type information to be investigated corresponding to the database node information; The task execution component 4042 is configured to execute the log collection task, and during the execution of the log collection task, call the target log collection component 405 corresponding to the target log collection component information, and collect the log files corresponding to the log type information to be investigated from the database node to be investigated represented by the database node information to be investigated.
[0064] In a possible implementation manner, the task generation component 4041 is configured to obtain user login account information, and generate a log collection task based on the user login account information, the database node information, and the log type information to be troubleshot corresponding to the database node information.
[0065] In a possible implementation manner, the task execution component 4042 is configured to save the log file to the temporary output directory of the database node to be troubleshot represented by the database node information to be troubleshot, and perform compression processing on all the log files in the temporary output directory to obtain a log file compression package; when it is determined that the log file download condition is satisfied, download the log file compression package from the temporary output directory to the local output directory, and perform decompression processing on the log file compression package in the local output directory to obtain the log file of the database node to be troubleshot represented by the database node information to be troubleshot.
[0066] In a possible implementation manner, the task execution component 4042 is further configured to generate a log collection result statistical report based on the component execution logs during the log file collection process, and output the log collection result statistical report.
[0067] It should be noted that the principle of the database cluster log collection system 400 provided by the embodiments of the present application to solve technical problems is similar to that of the database cluster log collection method provided by the embodiments of the present application. Therefore, for the implementation of the database cluster log collection system 400 provided by the embodiments of the present application, reference can be made to the implementation of the database cluster log collection method provided by the embodiments of the present application, and the repeated parts will not be described again.
[0068] After introducing the database cluster log collection method and system provided by the embodiments of the present application, next, the electronic device provided by the embodiments of the present application will be briefly introduced.
[0069] Refer to Figure 6 As shown, the electronic device 500 provided by the embodiments of the present application at least includes a processor 501, a memory 502, and a computer program stored on the memory 502 and executable on the processor 501. When the processor 501 executes the computer program, the database cluster log collection method provided by the embodiments of the present application is implemented.
[0070] The electronic device 500 provided by the embodiments of the present application may further include a bus 503 connecting different components (including the processor 501 and the memory 502). Among them, the bus 503 represents one or more of several types of bus structures, including a memory bus, a peripheral bus, a local area bus, etc.
[0071] The memory 502 may include a readable storage medium in the form of volatile memory, such as Random Access Memory (RAM) 5021 and / or cache memory 5022, and may further include Read Only Memory (ROM) 5023. The memory 502 may also include a program tool 5025 having a set (at least one) of program modules 5024. The program modules 5024 include, but are not limited to, an operating subsystem, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0072] The processor 501 may be a processing element or a collective term for multiple processing elements. For example, the processor 501 may be a Central Processing Unit (CPU), or one or more integrated circuits configured to implement the database cluster log collection method provided in the embodiments of the present application. Specifically, the processor 501 may be a general-purpose processor, including but not limited to a CPU, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0073] The electronic device 500 may communicate with one or more external devices 504 (such as a keyboard, a remote control, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 500 (such as a mobile phone, a computer, etc.), and / or communicate with a device that enables the electronic device 500 to communicate with one or more other electronic devices 500 (such as a router, a modem, etc.). Such communication may be performed through an Input / Output (I / O) interface 505. And, the electronic device 500 may also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through a network adapter 506. As Figure 6 shown, the network adapter 506 communicates with other modules of the electronic device 500 through a bus 503. It should be understood that although Figure 6Not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 500, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, redundant arrays of independent disks (RAID) subsystems, magnetic tape drives, and data backup storage subsystems, etc.
[0074] It should be noted that Figure 6 The shown electronic device 500 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0075] Next, the computer-readable storage medium provided by the embodiments of the present application will be introduced. The computer-readable storage medium provided by the embodiments of the present application stores computer instructions, and when the computer instructions are executed by a processor, the database cluster log collection method provided by the embodiments of the present application is implemented. Specifically, the computer instructions may be built-in or installed in the processor so that the processor can implement the database cluster log collection method provided by the embodiments of the present application by executing the built-in or installed computer instructions.
[0076] In addition, the database cluster log collection method provided by the embodiments of the present application may also be implemented as a computer program product. The computer program product includes program code, and when the program code runs on a processor, the database cluster log collection method provided by the embodiments of the present application is implemented.
[0077] The computer program product provided by the embodiments of the present application may adopt one or more computer-readable storage media, and the computer-readable storage media may be but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any suitable combination of the above. Specifically, more specific examples (non-exhaustive list) of the computer-readable storage media include electrical connections with one or more wires, portable disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0078] The computer program product provided by the embodiments of the present application may be a CD-ROM and include program code, and may also run on electronic devices such as servers and computers. However, the computer program product provided by the embodiments of the present application is not limited thereto. In the embodiments of the present application, the computer-readable storage medium may be any tangible medium that contains or stores program code, and the program code may be used by or in combination with an instruction execution system, apparatus, or device.
[0079] It should be noted that although several units or subunits of the apparatus are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units described above may be embodied in one unit. Conversely, the features and functions of one unit described above may be further divided and embodied by multiple units.
[0080] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0081] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0082] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the embodiments of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A method for collecting database cluster logs, characterized in that, Including: Obtain the fault type to be investigated for the database cluster; Based on the first configuration file, according to the fault type to be investigated, determine the database node information to be investigated corresponding to the database cluster and the log type information corresponding to the database node information to be investigated; wherein, the database node information corresponding to different fault types and the log type information corresponding to the database node information are configured in the first configuration file; Based on the second configuration file, according to the log type information corresponding to the database node information to be investigated, determine the target log collection component information corresponding to the database node information to be investigated; wherein, the log collection component information corresponding to different log type information is configured in the second configuration file; Call the target log collection component corresponding to the target log collection component information, and based on the log type information to be investigated, collect the log files of the database node to be investigated represented by the database node information to be investigated.
2. The method for collecting database cluster logs according to claim 1, wherein, Based on the first configuration file, according to the fault type to be investigated, determine the database node information to be investigated corresponding to the database cluster and the log type information corresponding to the database node information to be investigated, including: Call the parameter parsing component to parse out the database node information corresponding to the fault type to be investigated and the log type information corresponding to the database node information from the first configuration file; Determine the parsed database node information and the log type information corresponding to the database node information as the database node information to be investigated corresponding to the database cluster and the log type information corresponding to the database node information to be investigated.
3. The method for collecting database cluster logs according to claim 1, wherein Based on the second configuration file, according to the log type information corresponding to the database node information to be investigated, determine the target log collection component information corresponding to the database node information to be investigated, including: Call the parameter parsing component to parse out the log collection component information corresponding to the log type information to be investigated from the second configuration file; Determine the parsed log collection component information as the target log collection component information corresponding to the database node information to be investigated.
4. The method for collecting database cluster logs according to claim 1, wherein, Call the target log collection component corresponding to the target log collection component information, and based on the log type information to be investigated, collect the log files of the database node to be investigated represented by the database node information to be investigated, including: Call the task generation component to generate a log collection task based on the database node information and the log type information corresponding to the database node information; Call the task execution component to execute the log collection task, and during the execution of the log collection task, call the target log collection component corresponding to the target log collection component information to collect the log files corresponding to the log type information to be investigated from the database node to be investigated represented by the database node information to be investigated.
5. The method for collecting database cluster logs according to claim 4, wherein Call the task generation component to generate a log collection task based on the database node information and the log type information corresponding to the database node information, including: Call the task generation component to obtain the user login account information, and generate the log collection task based on the user login account information, the database node information, and the log type information to be investigated corresponding to the database node information.
6. The method for collecting database cluster logs according to claim 4, wherein, After calling the target log collection component corresponding to the target log collection component information to collect the log files corresponding to the log type information to be investigated from the database node to be investigated represented by the database node information to be investigated, it further includes: Call the task execution component to save the log files to the temporary output directory of the database node to be investigated represented by the database node information to be investigated, and compress all the log files in the temporary output directory to obtain a log file compression package; When it is determined that the log file download condition is met, download the log file compression package from the temporary output directory to the local output directory, and decompress the log file compression package in the local output directory to obtain the log files of the database node to be investigated represented by the database node information to be investigated.
7. The method for collecting database cluster logs according to any one of claims 1-6, characterized in that, After calling the target log collection component corresponding to the target log collection component information to collect the log files of the database node to be investigated represented by the database node information to be investigated based on the log type information to be investigated, it further includes: Generate a log collection result statistical report based on the component execution logs during the log file collection process, and output the log collection result statistical report.
8. A database cluster log collection system, characterized in that, It includes: A database cluster fault monitoring component for obtaining the fault type to be investigated of the database cluster; A database parameter acquisition component for determining the database node information to be investigated corresponding to the database cluster and the log type information to be investigated corresponding to the database node information to be investigated according to the fault type to be investigated based on the first configuration file; wherein, different fault type corresponding database node information and the log type information corresponding to the database node information are configured in the first configuration file; A tool parameter acquisition component for determining the target log collection component information corresponding to the database node information to be investigated according to the log type information to be investigated corresponding to the database node information to be investigated based on the second configuration file; wherein, the log collection component information corresponding to different log type information is configured in the second configuration file; A log collection execution component for calling the target log collection component corresponding to the target log collection component information to collect the log files of the database node to be investigated represented by the database node information to be investigated based on the log type information to be investigated.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the database cluster log collection method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by the processor, it implements the database cluster log collection method according to any one of claims 1-7.
Citation Information
Patent Citations
Fetch and display system of each node log under large-scale cluster
CN105553716A
Log collection method and device
CN111651324A
Log acquisition method and device, electronic equipment, chip and storage medium
CN114531340A
Log processing method and device, equipment and storage medium
CN114817192A
Distributed database fault diagnosis method and device, electronic equipment and storage medium
CN116048859A
Cited By
Business system and business execution method
CN120952950A