A python-based cluster log automatic analysis method
Patent Information
- Application Number
- CN202211372783.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-01
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-11-01
AI Technical Summary
建立了一个基于语义信息的完整日志分析模型,此发明解决了中间组件的日志收集归类,同样未涉及到多集群部署业务下的错误日志分析
[0038]1.通过集群账号配置文件中的集群主机ip和用户信息并发登录到集群服务器,结合kubectl命令实现指定关键字日志的分析,搭配jenkins服务和git服务使用,实现无人值守的自动化多集群日志分析,并根据日志分析形报告。
Smart Images

Figure CN115686675B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of log analysis technology, and in particular to an automated cluster log analysis method based on Python. Background Technology
[0002] As modern enterprises, especially software companies, continue to expand their business scale, many business operations are no longer conducted on a single server or cluster. Different business models often require deployment on different clusters, each with a large number of servers. This massive number of services generates a large amount of logs. To assess the stability of business systems, users often need to analyze the stability of services within the cluster by examining error messages in the logs. By analyzing the error messages in the logs, problems in the system can be identified and fixed. In practical applications, since clusters often deploy a large number of services, manually analyzing the logs of these services would be extremely time-consuming. Therefore, how to automate the analysis of error logs from all services across multiple clusters is a pressing issue that needs to be addressed.
[0003] Chinese patent CN114490164A discloses a log collection method, system, device, and computer storage medium. This method is applied to a client-side application, acquiring internal error logs and collecting them according to preset rules. While this method solves the problem of collecting client error logs according to specified rules, it does not address error log analysis in multi-cluster deployment environments.
[0004] Chinese Patent Publication No. CN113297051A discloses a log analysis and processing method and apparatus. This method focuses on analyzing abnormal logs of intermediate components to obtain log information of at least one log type; analyzing and processing error log information within the at least one log type to obtain at least one exception type; outputting at least one of the log types and exception types; classifying logs into normal logs and error logs based on semantic information; further detecting exception types contained in error logs based on semantic information; and analyzing logs in conjunction with log source patterns. A complete log analysis model based on semantic information is established. This invention solves the problem of log collection and classification for intermediate components, but it does not address error log analysis in multi-cluster deployment scenarios.
[0005] Therefore, how to provide an efficient method for analyzing error logs in a multi-cluster business model has become an urgent technical problem to be solved. Summary of the Invention
[0006] In view of this, the present invention mainly addresses the issue of providing an efficient method for analyzing error logs in a multi-cluster business mode, thereby improving the efficiency of error log analysis for users in a multi-cluster business mode and reducing the corresponding resource consumption.
[0007] This invention provides an automated cluster log analysis method based on Python, comprising:
[0008] Step S1: Install the Python package on the local Windows machine, and install the Jenkins and Git services on the Linux machine;
[0009] Step S2: Create a log analysis project on the Git service, download the created log analysis project to the Windows machine, create a Python file and a cluster account configuration file under the created log analysis project, and upload the created Python file and cluster account configuration file to the log analysis project of the Git service;
[0010] Step S3: Create a log analysis task on the Jenkins service based on the log analysis project, and configure the created log analysis task;
[0011] Step S4: Start the log analysis task created in step S3 by calling the Python file, perform log analysis on multiple clusters at regular intervals, and save the log analysis report to the Jenkins service.
[0012] Furthermore, step S2 of the Python-based automated cluster log analysis method of the present invention includes:
[0013] Step S21: Create a log analysis project on the git service, and download the created log analysis project to the Windows machine using the git clone command;
[0014] Step S22: Enter the log analysis project directory and create a cluster account configuration file. The cluster account configuration file stores cluster information in the form of key-value pairs, where the key is the English abbreviation of the cluster and the value is the user information of the corresponding cluster. The user information includes IP address, username, password and port.
[0015] Step S23: Enter the log analysis project directory, create a Python file, write the cluster list ip_list and task execution time run_time parameters for the logs to be analyzed, and create the functions parse_config, multi_connect, ssh_login and log_analysis.
[0016] Step S24: Upload the created Python file and cluster account configuration file to the log analysis project of the Git service.
[0017] Furthermore, step S3 of the Python-based automated cluster log analysis method of the present invention includes:
[0018] Step S31: Create a log analysis task on the Jenkins service based on the log analysis project;
[0019] Step S32: Configure the cluster list ip_list parameter in the Python file to the cluster abbreviation in the cluster account configuration file, and configure the task execution time run_time parameter in the Python file to 24 hours;
[0020] Step S33: Configure triggers, source code management, shell command execution, and report archive directory for the log analysis task.
[0021] Furthermore, step S33 of the Python-based automated cluster log analysis method of the present invention includes: configuring the trigger event of the trigger to 0:00 every day.
[0022] Furthermore, step S33 of the Python-based automated cluster log analysis method of the present invention further includes: configuring the source code management of the log analysis task by inputting the URL and authentication information of the log analysis project on the Git service.
[0023] Furthermore, step S4 of the Python-based automated cluster log analysis method of the present invention includes:
[0024] Step S41: Start the log analysis task by selecting the cluster abbreviation of the log cluster to be analyzed in the cluster list ip_list and the run_time parameter corresponding to the log cluster to be analyzed on the Jenkins service;
[0025] Step S42: Parse the cluster account configuration file and log in to the corresponding cluster based on the user information obtained from the parsing;
[0026] Step S43: Obtain all service names in the cluster by executing the kubectl command, traverse each service according to the service name, query the logs containing error information keywords in the past hour using the log_anlysis function, and save the logs named after the service name in the error log folder of the report archive directory.
[0027] Step S44: By comparing the current task's execution time with the task's execution time (run_time) parameter, the log analysis task can be terminated or continued based on the comparison result.
[0028] Furthermore, step S42 of the Python-based automated cluster log analysis method of the present invention includes:
[0029] Step S421: Parse the cluster account configuration file using the Parse_config function, and obtain the user information corresponding to the cluster to be analyzed based on the cluster's English abbreviation;
[0030] Step S422: Use the Multi_connect function to enable the corresponding number of threads based on the number of clusters in the cluster list ip_list;
[0031] Step S423: Remotely log in to the corresponding cluster using the SSH_login function based on the user information.
[0032] Furthermore, the error message keywords in step S43 of the Python-based automated cluster log analysis method of the present invention include: ERROR, Exception, error, and Error.
[0033] Furthermore, step S44 of the Python-based automated cluster log analysis method of the present invention includes:
[0034] If the execution time of the current task is less than the task execution time run_time parameter, continue to execute steps S42 and S43;
[0035] If the current task's execution time is greater than or equal to the task's execution time parameter `run_time`, end the log analysis task.
[0036] Furthermore, step S4 of the Python-based automated cluster log analysis method of the present invention also includes: when the log analysis task is completed, saving the report in the report archive directory to the Jenkins service.
[0037] The present invention provides an automated cluster log analysis method based on Python, which has the following advantages:
[0038] 1. Log in to the cluster server concurrently using the cluster host IP and user information in the cluster account configuration file. Combine this with the kubectl command to analyze logs for specified keywords. Use it in conjunction with Jenkins and Git services to achieve unattended automated multi-cluster log analysis and generate reports based on the log analysis.
[0039] 2. Using a JSON configuration file as input, Python automatically reads the configuration file, logs into the cluster server concurrently via multi-threading, executes the kubectl command to collect the context of exception logs, and saves them in the Jenkins server, achieving a simple and efficient analysis method. When cluster host information needs to be updated, only the cluster.json configuration file needs to be manually maintained, without any additional maintenance. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart of an exemplary first embodiment of the present invention, which describes an automated cluster log analysis method based on Python.
[0042] Figure 2 This is a flowchart of an exemplary second embodiment of the present invention, which describes an automated cluster log analysis method based on Python.
[0043] Figure 3 This is a flowchart of an exemplary third embodiment of the present invention, which describes an automated cluster log analysis method based on Python.
[0044] Figure 4 This is a flowchart of an exemplary fourth embodiment of the present invention, which describes an automated cluster log analysis method based on Python.
[0045] Figure 5 This is a flowchart of an exemplary fifth embodiment of the present invention, which describes an automated cluster log analysis method based on Python. Detailed Implementation
[0046] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0047] It should be noted that, unless otherwise specified, the following embodiments and features can be combined with each other; and, based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0048] It should be noted that various aspects of the embodiments described below are within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.
[0049] The following are explanations of the terms used in the embodiments:
[0050] Python is an object-oriented, interpreted computer programming language invented by Guido van Rossum in late 1989. Python has a concise and clear syntax, a rich and powerful library, and is widely used in fields such as automated testing, artificial intelligence, and big data processing.
[0051] Jenkins is an open-source software project, a continuous integration tool developed in Java. Jenkins offers a rich set of plugins and aims to provide an open and easy-to-use software platform that enables continuous integration and delivery of software projects. Jenkins' main functions are to continuously and automatically build software projects and execute scheduled tasks such as automated testing, version building, and deployment.
[0052] Git is an open-source distributed version control system used for agile and efficient version management of source code or files for projects of all sizes. Git was developed by Linus Torvalds to help manage the Linux kernel source code. Its main functions include cloning code and version information from a server to a local machine; creating branches and modifying code on a local machine; and uploading code to a server branch.
[0053] The `git clone` command is a command within the `git` tool. Its usage is `git clone <repository URL>`. This command creates a directory on the local machine with the same name as the remote server's repository, and then synchronizes the resource files and source code from the remote server's repository to the local directory.
[0054] Figure 1 This is a flowchart of an automated cluster log analysis method based on Python according to an exemplary first embodiment of the present invention, as shown below. Figure 1 As shown, the method in this embodiment includes:
[0055] Step S1: Install the Python package on the local Windows machine, and install the Jenkins and Git services on the Linux machine;
[0056] Step S2: Create a log analysis project on the Git service, download the created log analysis project to the Windows machine, create a Python file and a cluster account configuration file under the created log analysis project, and upload the created Python file and cluster account configuration file to the log analysis project of the Git service;
[0057] Step S3: Create a log analysis task on the Jenkins service based on the log analysis project, and configure the created log analysis task;
[0058] Step S4: Start the log analysis task created in step S3 by calling the Python file, perform log analysis on multiple clusters at regular intervals, and save the log analysis report to the Jenkins service.
[0059] In this embodiment, the Python package version 2.6 or higher is used.
[0060] In this embodiment, the Jenkins service is used to configure log analysis tasks, enabling automated task execution and report display; the Git service can host Python files and work with Jenkins to facilitate Jenkins downloading and calling Python files during automatic execution.
[0061] Figure 2 This is an example of an automated cluster log analysis method based on Python, according to a second exemplary embodiment of the present invention. This embodiment is... Figure 1 Preferred embodiments of the method shown are as follows: Figure 2 As shown, step S2 of the method in this embodiment includes:
[0062] Step S21: Create a log analysis project on the git service, and download the created log analysis project to the Windows machine using the git clone command;
[0063] Step S22: Enter the log analysis project directory and create a cluster account configuration file. The cluster account configuration file stores cluster information in the form of key-value pairs, where the key is the English abbreviation of the cluster and the value is the user information of the corresponding cluster. The user information includes IP address, username, password and port.
[0064] Step S23: Enter the log analysis project directory, create a Python file, write the cluster list ip_list and task execution time run_time parameters for the logs to be analyzed, and create the functions parse_config, multi_connect, ssh_login and log_analysis.
[0065] Step S24: Upload the created Python file and cluster account configuration file to the log analysis project of the Git service.
[0066] In this embodiment, the cluster account configuration file is created in JSON format.
[0067] In this embodiment, the `ip_list` parameter controls the list of clusters to be analyzed for the logs, and the `run_time` parameter controls the execution duration of the task. The `Parse_config` function parses the cluster account configuration file and retrieves the user information corresponding to the clusters in the logs to be analyzed based on their abbreviations. The `Multi_connect` function enables the corresponding number of threads based on the number of clusters in the `ip_list`. The `Ssh_login` function remotely logs into the corresponding cluster based on the user information. The `log_analysis` function queries logs containing error message keywords from the past hour and saves these logs, named after the service, in the error log folder of the report archive directory.
[0068] In practical applications, the method in this embodiment logs in to the cluster server concurrently using the cluster host IP and user information in the cluster account configuration file. It combines the kubectl command to analyze logs with specified keywords. When used with Jenkins and Git services, it achieves unattended automated multi-cluster log analysis and generates reports based on the log analysis.
[0069] Figure 3 This embodiment is a Python-based automated cluster log analysis method according to an exemplary third embodiment of the present invention. Figure 1 Preferred embodiments of the method shown are as follows: Figure 3 As shown, step S3 of the method in this embodiment includes:
[0070] Step S31: Create a log analysis task on the Jenkins service based on the log analysis project;
[0071] Step S32: Configure the cluster list ip_list parameter in the Python file to the cluster abbreviation in the cluster account configuration file, and configure the task execution time run_time parameter in the Python file to 24 hours;
[0072] Step S33: Configure triggers, source code management, shell command execution, and report archive directory for the log analysis task.
[0073] In practical applications, configuring the cluster list ip_list parameter in the Python file to the cluster abbreviation in the cluster account configuration file can enable multiple parameter selection at startup.
[0074] In practical applications, step S33 of the method in this embodiment includes: configuring the trigger event of the trigger to be executed at 0:00 every day, that is, the task is automatically executed at 0:00 every day by default, or the task can be started manually.
[0075] In practical applications, the shell command execution configuration can be performed as follows: python analysis_cluster.py –cluster$ip_list –run-time$run_time.
[0076] In practical applications, the source code management of log analysis tasks is configured by inputting the URL of the log analysis project on the Git service and the authentication information.
[0077] In practical applications, configuring the report archive directory as logs / ** means collecting all logs in the logs directory.
[0078] Figure 4 This embodiment is a Python-based automated cluster log analysis method according to an exemplary fourth embodiment of the present invention. Figure 1 Preferred embodiments of the method shown are as follows: Figure 4 As shown, step S4 of the method in this embodiment includes:
[0079] Step S41: Start the log analysis task by selecting the cluster abbreviation of the log cluster to be analyzed in the cluster list ip_list and the run_time parameter corresponding to the log cluster to be analyzed on the Jenkins service;
[0080] Step S42: Parse the cluster account configuration file and log in to the corresponding cluster based on the user information obtained from the parsing;
[0081] Step S43: Obtain all service names in the cluster by executing the kubectl command, traverse each service according to the service name, query the logs containing error information keywords in the past hour using the log_anlysis function, and save the logs named after the service name in the error log folder of the report archive directory.
[0082] Step S44: By comparing the current task's execution time with the task's execution time (run_time) parameter, the log analysis task can be terminated or continued based on the comparison result.
[0083] In practical applications, the error message keywords in step S43 of this embodiment include: ERROR, Exception, error, and Error.
[0084] In practical applications, step S44 of the method in this embodiment includes:
[0085] If the execution time of the current task is less than the task execution time run_time parameter, continue to execute steps S42 and S43;
[0086] If the current task's execution time is greater than or equal to the task's execution time parameter `run_time`, end the log analysis task.
[0087] In practical applications, when the log analysis task is completed, the report in the report archive directory is saved to the Jenkins service.
[0088] A specific application of the method in this embodiment is as follows:
[0089] In Jenkins, select env1 and env2 as the IP list, where env1 and env2 are the English abbreviations of the clusters to be analyzed. Configure the run-time to 24 hours and start the Jenkins log analysis task. After the task starts, call pythonanalysis_cluster.py –cluster env1, env2 –run-time 24. Obtain two copies of the corresponding IP, username, password, and port from cluster.json using env1 and env2. Then, start two threads to remotely log in to the env1 and env2 clusters using the IP, username, password, and port respectively. Use the kubectl command to obtain the names of all services in the cluster. Traverse each service and use the kubectl command to query the context of logs containing keywords such as ERROR, Exception, error, and Error for that service in the past hour. Save the logs named after the service in the logs / {ip} / {time} folder.
[0090] The system checks if the current task has been executed for longer than the set `run_time` parameter. If so, the task stops automatically. If not, the task waits for 1 hour before continuing to execute the log analysis task. When the task execution time exceeds the `run_time` parameter, the task eventually exits. After the task exits, Jenkins saves all log reports in the `logs` directory on the server. You can view the cluster log analysis report by clicking on the file generated by the current build in Jenkins.
[0091] This embodiment uses a JSON configuration file as input data, automatically reads the configuration file using Python, logs into the cluster server concurrently through multi-threading, executes the kubectl command to collect the context of exception logs, and saves them in the Jenkins server, thus achieving a simple and efficient analysis method. When the cluster host information needs to be updated, only the cluster.json configuration file needs to be manually maintained, without any additional maintenance.
[0092] Figure 5 This embodiment is a Python-based automated cluster log analysis method according to an exemplary fifth embodiment of the present invention. Figure 1 and Figure 4 Preferred embodiments of the method shown are as follows: Figure 5 As shown, step S42 of the method in this embodiment includes:
[0093] Step S421: Parse the cluster account configuration file using the Parse_config function, and obtain the user information corresponding to the cluster to be analyzed based on the cluster's English abbreviation;
[0094] Step S422: Use the Multi_connect function to enable the corresponding number of threads based on the number of clusters in the cluster list ip_list;
[0095] Step S423: Remotely log in to the corresponding cluster using the SSH_login function based on the user information.
[0096] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for automated cluster log analysis based on Python, characterized in that, The method includes: Step S1: Install the Python package on the local Windows machine, and install the Jenkins and Git services on the Linux machine; Step S2: Create a log analysis project on the Git service, download the created log analysis project to the Windows machine, create a Python file and a cluster account configuration file under the created log analysis project, and upload the created Python file and cluster account configuration file to the log analysis project of the Git service; Step S3: Create a log analysis task on the Jenkins service based on the log analysis project, and configure the created log analysis task; Step S4: Start the log analysis task created in step S3 by calling the Python file, perform log analysis on multiple clusters at regular intervals, and save the log analysis report to the Jenkins service; Step S2 includes: Step S21: Create a log analysis project on the git service, and download the created log analysis project to the Windows machine using the git clone command; Step S22: Enter the log analysis project directory and create a cluster account configuration file. The cluster account configuration file stores cluster information in the form of key-value pairs, where the key is the English abbreviation of the cluster and the value is the user information of the corresponding cluster. The user information includes IP address, username, password and port. Step S23: Enter the log analysis project directory, create a Python file, write the cluster list ip_list and task execution time run_time input parameters for the logs to be analyzed, and create the parse_config, multi_connect, ssh_login and log_analysis functions; Step S24: Upload the created Python file and cluster account configuration file to the log analysis project of the Git service; Step S4 includes: Step S41: Start the log analysis task by selecting the cluster abbreviation of the log cluster to be analyzed in the cluster list ip_list and the run_time parameter corresponding to the log cluster to be analyzed on the Jenkins service; Step S42: Parse the cluster account configuration file and log in to the corresponding cluster based on the user information obtained from the parsing; Step S43: Obtain all service names in the cluster by executing the kubectl command, traverse each service according to the service name, query the logs containing error information keywords in the past hour using the log_analysis function, and save the logs named after the service name in the error log folder of the report archive directory; Step S44: By comparing the current task's execution time with the task's execution time (run_time) parameter, end or continue the log analysis task based on the comparison result; Step S42 includes: Step S421: Parse the cluster account configuration file using the parse_config function, and obtain the user information corresponding to the cluster to be analyzed based on the cluster's English abbreviation; Step S422: Use the multi_connect function to enable the corresponding number of threads based on the number of clusters in the cluster list ip_list; Step S423: Remotely log in to the corresponding cluster using the ssh_login function based on the user information; The error message keywords in step S43 include: ERROR, Exception, error, and Error.
2. The automated cluster log analysis method based on Python according to claim 1, characterized in that, Step S3 includes: Step S31: Create a log analysis task on the Jenkins service based on the log analysis project; Step S32: Configure the cluster list ip_list parameter in the Python file to the cluster abbreviation in the cluster account configuration file, and configure the task execution time run_time parameter in the Python file to 24 hours; Step S33: Configure triggers, source code management, shell command execution, and report archive directory for the log analysis task.
3. The automated cluster log analysis method based on Python according to claim 2, characterized in that, Step S33 includes: configuring the trigger's trigger event to be at 0:00 every day.
4. The automated cluster log analysis method based on Python according to claim 2, characterized in that, Step S33 also includes: configuring source code management for log analysis tasks by inputting the URL and authentication information of the log analysis project on the git service.
5. The automated cluster log analysis method based on Python according to claim 1, characterized in that, Step S44 includes: If the execution time of the current task is less than the task execution time run_time parameter, continue to execute steps S42 and S43; If the current task's execution time is greater than or equal to the task's execution time parameter `run_time`, end the log analysis task.
6. The automated cluster log analysis method based on Python according to claim 1, characterized in that, Step S4 also includes: when the log analysis task is completed, saving the report in the report archive directory to the Jenkins service.
Citation Information
Patent Citations
Log analysis processing method and device
CN113297051A
Log collection method, system and equipment and computer storage medium
CN114490164A