Log-based service health inspection method and device
Through the log-based hierarchical inspection method, the health of the information system is automatically assessed, which solves the problems of low efficiency and poor accuracy of traditional manual inspections and realizes efficient and accurate system health assessment and problem location.
Patent Information
- Application Number
- CN202510845380.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional manual inspection methods are inefficient and inaccurate in large and complex information systems, pose a risk of misjudgment, and cannot meet the health requirements of modern information systems.
A log-based service health inspection method is adopted to conduct layered inspections, including basic environment data inspections, manager inspections, and product function inspections. Inspections are automated using log data and host performance data, and accurate system health assessments are provided through inspection reports.
It has achieved automated inspections, improved inspection efficiency and accuracy, eliminated inspection blind spots, clearly divided team responsibilities, timely discovered system stability and performance issues, and provided practical solutions.
Smart Images

Figure CN120743871A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of system security technology, and in particular to a log-based service health inspection method and device. Background Art
[0002] With the rapid development of information technology, modern information systems are becoming increasingly complex, and the demand for system service health is increasing.
[0003] Generally, professionals are required to inspect the system to determine its health. However, as the scale of business systems gradually increases and the diversity of applications gradually increases, traditional manual inspection methods are unable to cope with such a huge amount of information, and work efficiency needs to be improved. In addition, manual inspections are subjective to a certain extent, and the inspection results often depend on the experience and judgment of the inspectors. Different inspectors may come to different conclusions, which increases the risk of misjudgment.
[0004] Therefore, how to provide an automated service health inspection method, improve the efficiency of the inspection method and provide the accuracy of the inspection results is a key research topic for those skilled in the art. Summary of the Invention
[0005] In the first aspect, the present application provides a log-based service health inspection method, which is used to perform service health inspection on a log management system. The log management system is a cluster including at least one node, and each node includes at least one module. The method includes: obtaining log data of each node in the log management system, and host performance data of each node; based on the log data and the host performance data, performing basic environment data inspection, manager inspection, and product function inspection on the log management system; wherein, the basic environment data inspection includes: license information inspection of the log management system, inter-module version compatibility information inspection, module Block patch package version information check, cluster organizational structure update check; the manager inspection includes: host performance check of each node, database backup check, module health status check, detection program check, and preset important configuration item check; the product function inspection includes: log query function check, log collection module function check, log parsing function check, data visualization function check, data report function check, scheduled task function check, tenant management function check, tenant usage limit check, and node user login behavior check; based on the inspection results corresponding to the basic data inspection, the manager inspection, and the product function inspection, an inspection report is issued.
[0006] The log-based service health inspection method provided in this application, on the one hand, automates the inspection tasks, solving the problems of low efficiency and poor data accuracy of traditional manual inspection methods under the microservice architecture.
[0007] Furthermore, a structured inspection approach is employed, conducting layered inspections across three dimensions: basic environmental information inspections, manager inspections, and product function inspections. Basic data inspections ensure underlying data accuracy, manager inspections guarantee system health, and product function inspections verify business operations. This allows for more accurate identification of key risk points at different levels. Inspections comprehensively cover the entire system layer, from underlying environmental data to the top business application layer, eliminating blind spots and enabling a comprehensive assessment of the overall health and availability of the log management system.
[0008] Furthermore, clear dimensional divisions facilitate the clarification of each team's responsibilities, improving inspection efficiency and problem location. For example, basic information inspections are handled by the operations or infrastructure team, manager inspections are handled by administrators or the platform team responsible for cluster operations and performance management, and product functionality inspections are verified by the product support team, testing team, or end-user representatives. Different teams can conduct inspections of their respective responsibilities in parallel.
[0009] In some possible implementations, the license information check includes: checking the latest validity period and usage limit of the license based on the log data; the cluster module compatibility information check includes: checking the software version compatibility between the modules in the cluster based on preset compatibility conditions; the module patch package version information check includes: checking whether the patch package version of each module in the cluster is the latest version; the cluster organizational structure update check includes: obtaining the latest organizational structure of the cluster and determining the change information of the cluster organizational structure.
[0010] With this approach, basic information inspections include checking license information, preventing distorted inspection results due to expired system licenses. Basic information inspections also include checking cluster organizational structure updates, promptly identifying additions, deletions, and changes to the cluster organizational structure and preventing node blind spots during subsequent manager inspections. Basic information inspections also include checking cluster module compatibility information, promptly identifying broken dependency chains between modules and assisting in identifying unpatched modules.
[0011] In some possible implementations, the host performance check includes: checking the number of hosts in the cluster, the central processing unit (CPU) performance data of each host, memory performance data, disk performance data, disk input / output (IO) time ratio, traffic volume within a preset time period, and the number of currently open files; the database backup check includes: checking whether the database backup switch is turned on, and whether there is a backup file under the preset database backup path; the cluster module health status check includes: checking whether the module has unrecovered alarm information; the detection program check includes: the number of cluster detection programs, and whether each detection program is in a normal detection state; the preset important configuration item baseline check includes: checking at least two of the data cache queue backlog information, the data backlog information of the log data stream component, and the allocation delay information of the query service.
[0012] This method is used to inspect key indicators that affect system stability and performance. The inspection indicators are representative and accurate, which can better identify stability and performance issues of the log management system.
[0013] In some possible implementations, the log query function check includes: checking whether the log query function returns the query results normally; the log collection module function check includes: checking the number of log collection modules and checking whether the log collection modules collect active log streams normally; the log parsing function check includes: checking the number of field extraction modules in the log management system and checking whether the field extraction modules can parse logs normally; the data visualization function check includes: checking the number of dashboards in the log management system and checking whether the dashboards can display data normally; the data report function check includes: checking the number of data reports in the log management system and Whether the data report can output report data normally; the scheduled task function check includes: checking the number of scheduled tasks in the log management system, and whether the scheduled tasks are executed normally; the tenant management function check includes: checking whether the tenant quota is modified, the latest configuration of the number of over-limits allowed for the tenant, and at least two of the modification actions of the tenant's account and / or password; the usage limit check includes: checking at least two of the tenant quota, the tenant's actual traffic usage, idle quota, number of over-usages, the tenant's highest daily traffic usage this month, and the tenant's average daily log volume this month; the cluster user login behavior check includes: counting the number of logins of node users within a preset time period.
[0014] In some possible implementations, the basic environment data inspection, manager inspection, and product function inspection of the log management system include: performing the basic environment data inspection, the manager inspection, and the product function inspection on the log management system in sequence.
[0015] Using this method, basic environment information inspection is performed first, which can ensure that subsequent inspections are performed based on the latest and most accurate underlying data in the log management system, reducing the impact of underlying data errors on subsequent inspection results.
[0016] In some possible implementations, the method further includes: when it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes the addition of a first node, determining that the manager inspection module corresponding to the manager inspection adds a manager inspection task for the first node; when it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes the deletion of a second node, determining that the manager inspection module does not perform a manager inspection on the second node. When it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes the change of role information of a third node, determining that the manager inspection module keeps executing the manager inspection task of the third node and modifies the node name corresponding to the inspection result of the third node to the node name after the role change.
[0017] Using this method, when the basic data inspection finds a new node, it automatically triggers the manager inspection of the new node. When it finds a departing node, it automatically deletes the manager inspection task of the departing node. When it finds a role-changing node, it automatically modifies the node name corresponding to the inspection report of the role-changing node.
[0018] This further ensures the integrity and comprehensiveness of the inspection tasks. Compared with the manual discovery, creation, deletion or modification of inspection tasks, the manager inspection tasks for newly added nodes can be triggered more timely, accurately and efficiently.
[0019] In some possible implementations, the inspection report includes information for indicating the association relationship between inspection items, wherein, when the cluster module health status check finds the presence of alarm information related to log parsing, and the log parsing function check finds that the field value parsed by the field extraction module is null, the inter-module version compatibility information check finds that the version information of the upstream and downstream modules related to the field extraction module is abnormal, and the module patch package version information check finds that the patch package of the field extraction module is not the latest version, it is determined that the inspection item association information includes a first association map, and the first association map is used to display the problem location map relationship of the associated inspection items related to the log parsing alarm information.
[0020] This method analyzes the correlation information between independent inspection items and displays the correlation map of independent inspection items, which can better assist users in fault location.
[0021] In some possible implementations, the product function inspection includes a tenant management function inspection and a usage limit inspection. The tenant management function inspection includes checking the latest configuration of the number of over-limits allowed for the tenant, and the usage limit inspection includes: checking the tenant's actual traffic usage, average idle quota, and number of over-usages; issuing an inspection report includes: when it is determined that the tenant's average idle quota is greater than the preset traffic and the number of over-usages is equal to 0, based on the average value of the tenant's actual traffic usage, determining that the inspection report includes a quota reduction suggestion, and the quota reduction suggestion is used to indicate the tenant's average idle quota and the recommended quota; when it is determined that the tenant's over-usage number is greater than or equal to the tenant's corresponding allowed over-limit number, based on the average value of the tenant's actual traffic usage, determining that the inspection report includes a quota increase suggestion, and the quota increase suggestion is used to indicate the number of times the tenant's actual traffic usage exceeds the quota and the recommended quota.
[0022] This approach provides practical and effective solutions based on inspection data, better assisting users in solving related problems and improving the user experience of the inspection method.
[0023] In some possible implementations, before performing a basic environment data inspection, a manager inspection, and a product function inspection on the log management system based on the log data and the host performance data, the method further includes: performing data cleaning on the log data and the host performance data based on preset rules, and storing the data obtained after data cleaning in a search engine, wherein the preset rules are used to indicate that preset valid fields are retained; performing a basic environment data inspection, a manager inspection, and a product function inspection on the log management system based on the log data and the host performance data includes: performing a basic environment data inspection, a manager inspection, and a product function inspection on the log management system based on the data content stored by the search engine.
[0024] This method can convert unstructured text into queryable, alarmable, and visualized data, reducing the storage overhead of invalid data.
[0025] In a second aspect, the present application further provides a log-based service health inspection device, comprising a unit for executing any one of the log-based service health inspection methods in the first aspect.
[0026] In a third aspect, the present application further provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor according to any one of the log-based service health inspection methods in the first aspect.
[0027] In a fourth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when run on an electronic device, enables the electronic device to execute any one of the log-based service health inspection methods in the first aspect.
[0028] In a fifth aspect, an embodiment of the present application further provides a chip module, comprising a transceiver component and a chip, wherein the chip is used to execute any one of the log-based service health inspection methods in the first aspect.
[0029] It is understood that the log-based service health inspection device, computer storage medium, computer program, computer program product, and chip system provided above are all used to perform the method shown in any implementation of the first aspect of the embodiment of this application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a flow chart of a log-based service health inspection method provided by an embodiment of the present application;
[0031] Figure 2 This is a schematic diagram of an inspection report provided in an embodiment of the present application;
[0032] Figure 3 This is a schematic diagram of a log-based service health inspection device provided in an embodiment of the present application;
[0033] Figure 4 This is a schematic diagram of another log-based service health inspection device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solutions and advantages of this application clearer, this application will be further described below with reference to the accompanying drawings.
[0035] See also Figure 1 , Figure 1 A flow chart of a log-based service health inspection method provided in an embodiment of the present application. The method is used to perform service health inspection on a log management system, which is a cluster including at least one node, and each node in the cluster includes at least one module, such as Figure 1 As shown, the method includes the following steps:
[0036] S101: The electronic device obtains log data of each node in the log management system and host performance data of each node.
[0037] Exemplarily, the electronic device collects log data of each node through an agent process of a log collection module in the cluster.
[0038] Exemplarily, the electronic device obtains the performance data of each node host server under the host performance data collection log management system service by detecting the agent process of the host performance index.
[0039] In some possible implementations, after the electronic device obtains the above-mentioned log data and host performance data, it cleans the log data and the host performance data based on preset rules, and stores the data obtained after data cleaning in a search engine. The preset rules are used to indicate that preset valid fields are retained.
[0040] For example, the fields of the cleaned host performance data include, but are not limited to, host name, IP address, CPU usage, memory usage, disk usage, and disk input / output. The fields of each log event in the cleaned log data include, but are not limited to, time (timestamp), identification information (trace_id), event source host IP address (host_ip), event type (event_action), and processing result (event.outcome).
[0041] Therefore, log data and host performance data are cleaned, and unstructured text is converted into queryable, alarmable, and visualized data, reducing the storage overhead of invalid data.
[0042] S102, the electronic device performs a basic environmental data inspection on the log management system.
[0043] In the embodiments of this application, service health inspections include inspections of basic environment data. Basic environment data inspections are used to inspect the health and compliance of the underlying supporting environment (cluster organizational structure information, license information). Understandably, basic environment data is the cornerstone of stable system operation. Ensuring the correctness of the underlying environment data ensures that subsequent inspection results are more informative and accurate.
[0044] In an embodiment of the present application, the basic environment data inspection includes: license information inspection of the log management system, cluster organizational structure update inspection, inter-module version compatibility information inspection, and module patch package version information inspection.
[0045] In some possible implementations, the license information check includes: checking the latest validity period and usage limit of the license based on log data. Specifically, the latest validity period and usage limit of the license can be checked based on the license log file. The electronic device can also determine that the inspection result includes renewal prompt information for indicating license expiration information based on the latest validity period of the license, if the license validity period meets a preset condition (for example, the license expiration time is less than or equal to 15 days and greater than 0). The electronic device can also determine that the inspection result includes prompt information for indicating changes in the total usage limit of the license based on changes in the total usage limit of the license (for example, a decrease or increase in the total usage limit).
[0046] In some possible implementations, the electronic device can also obtain the latest quota information of the tenant corresponding to the license (one license can correspond to one or more tenants), and when the tenant's quota changes compared to the historical quota, determine that the inspection result includes quota change information indicating the tenant.
[0047] Exemplarily, the latest validity period of the license can be checked from the license log file license_info based on the following code statement.
[0048] "appname:rzyxj tag:license_info
[0049] eval expired_timestamp=parsedate(mes.expired_timestamp,"yyyy-MM-dd'T'HH:mm:ssZ")
[0050] / / Extract the expiration time from the license log
[0051] stats max(expired_timestamp)as expired_timestamp / / Get the latest expiration timestamp in the log
[0052] eval expired_timestamp=formatdate(tolong(expired_timestamp),"yyyy-MM-ddHH:mm:ss")
[0053] eval cnt=1
[0054] limit 1”
[0055] It should be noted that electronic devices can not only obtain the latest validity period and usage limit of the license based on the license log file, but also obtain the latest validity period and usage limit of the license through other means, such as by accessing the corresponding license information record database table. This article does not limit this.
[0056] Using this method, basic information inspection includes inspection of license information, which can avoid the problem of distorted inspection results caused by expiration of system licenses. For example, due to license expiration, the detection program stops reporting, the dashboard shows that everything is normal, and the actual node failure is not discovered. For example, after the license expires, the log management system may automatically downgrade to demonstration mode, and the performance data is controlled in a normal state by default, thereby covering up the real performance bottleneck.
[0057] In some possible implementations, electronic devices can also regularly simulate license expiration scenarios to verify the availability of inspection functions (including basic environment data inspection, manager inspection, and product function inspection) under license expiration scenarios, and discover in advance the hidden risk of inspection functions' dependence on license services (for example, the license service and the inspection function share the underlying library, resulting in interruption of inspection log collection and cessation of the inspection function), so as to remove the dependency in advance, reduce the coupling between the log management system function and the service health inspection function, thereby reducing the probability of common cause failures and improving the overall security of the system.
[0058] In some possible implementations, checking for cluster organizational structure update includes: obtaining the latest organizational structure of the cluster, and determining change information of the cluster organizational structure.
[0059] For example, the electronic device can obtain all key organizational structure metadata such as cluster nodes, modules, and configurations based on distributed key-value storage access, or can also obtain the latest organizational structure of the cluster based on a command line tool. After the electronic device obtains the latest organizational structure of the cluster, it can also compare the historical organizational structure data to find out the change information of the cluster organizational structure (such as node joining, leaving, role change, service expansion and contraction), and based on the change information of the cluster organizational structure, determine that the inspection results include information indicating the change of the cluster organizational structure.
[0060] In some possible implementations, when the change information of the cluster organizational structure includes a newly added node, the electronic device may further notify the manager inspection module to add a manager inspection task for the newly added node.
[0061] In the case that the change information of the cluster organizational structure includes deletion of a node, the electronic device may further notify the manager inspection module not to perform the manager inspection task on the deleted node.
[0062] If the change in the cluster organizational structure includes a node role change, the electronic device can also notify the manager inspection module to no longer perform manager inspections on the node before the role change, and to add manager inspection tasks for the node after the role change. Alternatively, if the node role changes but the node's unique identifier remains unchanged, the electronic device notifies the manager inspection module to modify the node name corresponding to the inspection result for the node to the name of the changed role while maintaining the inspection task for the node.
[0063] Understandably, in modern clusters, cluster organizational architecture information is highly dynamic. For example, cluster organizational architecture can be changed through node scaling (the cluster adjusts the number of nodes based on expansion strategies and current resource demand changes), role conversion, network topology reconstruction, and hybrid architecture addition.
[0064] Using this method, basic information inspection includes inspection of cluster organizational structure update information, timely discovery of cluster organizational structure information addition, deletion and change information, avoiding node blind spot problems in subsequent manager inspections, and reducing inspection costs for deleted nodes. This improves the accuracy of service health inspections and reduces inspection costs. In addition, allowing inspection users to know the latest organizational structure of the cluster in a timely manner can better assist inspection users in locating problems.
[0065] In some possible implementations, checking the cluster module compatibility information includes: checking the compatibility of software versions between modules in the cluster based on preset compatibility conditions.
[0066] For example, electronic devices can check the version compatibility of corresponding modules based on the dependencies and interface contracts between cluster modules. For example, cluster module compatibility can be checked by consulting the official version compatibility matrix of the cluster software to see if it meets the requirements. Another example is the module's own version documentation, which clearly states the version requirements of dependent components. Alternatively, cluster module compatibility inspections can be performed based on known information that specific version combinations may have compatibility issues.
[0067] Using this method, basic information inspection includes inspection of cluster module compatibility information, which can promptly discover broken dependency chains between modules and assist in identifying unpatched modules.
[0068] In some possible implementations, checking the module patch package version information includes: checking whether the patch package version of each module in the cluster is the latest version.
[0069] Exemplarily, the electronic device checks whether the patch package version of each module in the cluster is the latest version based on the module patch package version information table. When the electronic device determines that the patch package version of the target module is not the latest version, it can determine that the inspection result includes indication information for indicating that the patch package version of the target module is not the latest version. When the electronic device determines that the patch package version of the target module is not the latest version, it can also directly obtain the patch installation package of the target module, automatically update the patch package of the target module, and perform the module patch package version information check again after the update is completed. When it is determined that the abnormal patch problem no longer exists, it determines that the inspection result includes indication information for automatic upgrade of the patch package of the target module.
[0070] Using this method, basic information inspection includes inspection of module patch package version information, identifying software versions with vulnerabilities, and better assisting in locating fault problems so that users can discover and take remedial measures in a timely manner.
[0071] S103, the electronic device performs a manager inspection on the log management system.
[0072] In the embodiment of the present application, manager inspection is used to inspect the health status of the middle-layer physical resources and virtual resources, inspect key indicators that affect the stability and performance of the system, and ensure the stability and performance of the system operation.
[0073] In the embodiment of the present application, the manager inspection includes: host performance inspection of each node, database backup inspection, module health status inspection, cluster detection program inspection, and important configuration item inspection.
[0074] In some possible implementations, the host performance check includes: checking the number of cluster hosts, the central processing unit (CPU) performance data of each host, memory performance data, disk performance data, disk input / output IO time ratio, traffic size within a preset time period, and the number of currently open files.
[0075] In an embodiment of the present application, the electronic device can obtain the host performance data by accessing a search engine that stores log data and host performance data.
[0076] In some possible implementations, checking CPU performance data includes checking the IP address, host status, number of CPU cores, and average CPU load (including average load within 1 minute, average load within 5 minutes, and average load within 15 minutes). For example, a piece of CPU performance data obtained during an electronic device inspection is shown in Table 1 below.
[0077] Table 1
[0078] Inspection items IP address Host Status Number of CPU cores Average CPU load (1m, 5m, 15m) CPU 192.168.xxx OK 8 1.4,1.67,1.46
[0079] Exemplarily, the CPU performance data may be checked based on the following code statements.
[0080] "dbxquery connection="rizhiyi_manager"query="select ip,state,cpu_count,load_average from host"
[0081] eval checklist="CPU"
[0082] table checklist,ip,state,cpu_count,load_average
[0083] sort by+ip
[0084] streamstats count()as seq
[0085] fields seq,checklist,ip,state,cpu_count,load_average
[0086] rename seq as "sequence number", checklist as "check item", ip as "IP address", state as "host status", cpu_count as "number of CPU cores", load_average as "CPU average load (1m, 5m, 15m)"
[0087] In an embodiment of the present application, the average CPU load reflects the average number of tasks waiting for CPU time. If the average CPU load is close to the number of CPU cores in the system, it means that the CPU utilization rate is high; if the average CPU load is much smaller than the number of CPU cores, it means that the CPU utilization rate is low; if the average CPU load is greater than the number of CPU cores, it may mean that the system is beginning to be overloaded and tasks are queued for processing. Based on this, in some possible implementations, the electronic device can determine the indication information about the CPU utilization included in the inspection result based on the average CPU load. For example, if the average CPU load is greater than the number of CPU cores, it is determined that the inspection result includes indication information indicating that the average CPU load is greater than the number of CPU cores and the corresponding host system is beginning to be overloaded.
[0088] In some possible implementations, memory performance data includes items such as IP address, physical memory (GB), used memory (GB), memory usage, total swap partition (SWAP) space (GB), used SWAP space (GB), and SWAP space usage. For example, a piece of memory performance data obtained during an electronic device inspection is shown in Table 2 below.
[0089] Table 2
[0090]
[0091] In an embodiment of the present application, the ideal state of memory performance is that free memory accounts for between 20% and 70% of total memory, and swap space usage is low. Based on this, in some possible implementations, when the electronic device determines that the available memory is continuously less than 20%, the electronic device determines that the inspection result includes information indicating that the available memory is less than 20%, and that memory usage can be reduced by increasing physical memory or optimizing applications.
[0092] In some possible implementations, the disk performance data inspection items include IP address, root directory partition, root directory partition total space (GB), root directory partition used space (GB), root directory partition usage rate, data directory partition total space (GB), data directory partition used space (GB), and data directory partition usage rate. For example, a piece of disk performance data obtained by an electronic device inspection is shown in Table 3 below.
[0093] Table 3
[0094]
[0095] In some possible implementations, when the electronic device determines that the usage rate of the root directory partition exceeds 70% and / or the usage rate of the data partition exceeds 80%, it determines that the inspection result includes instruction information indicating that the usage rate of the corresponding partition is high and recommends cleaning up unnecessary files and logs. The electronic device may also clean up unnecessary files (such as files downloaded for more than 1 day) and logs (such as log files with a logging time of more than half a year) when it determines that the usage rate of the root directory partition exceeds 70% and / or the usage rate of the data partition exceeds 80%, and determine that the inspection result includes instruction information indicating that the files have been cleaned up, thereby reducing the usage rate of the corresponding partition.
[0096] In some possible implementations, the traffic volume check within the preset time period may be a traffic volume check for the most recent day.
[0097] In some possible implementations, if the electronic device determines that the current number of open files of the host is greater than the preset number of files, it determines that the inspection result includes indication information indicating that the current number of open files is greater than the preset number of files.
[0098] In some possible implementations, if the electronic device determines that the host's current number of open files is greater than a preset number of files, the average CPU load is greater than the number of CPU cores, and the disk IO time accounts for a large proportion, the inspection item association information in the inspection result is determined to include a second association map. This second association map is used to illustrate the problem location map relationship between the large average CPU load and the large number of currently open files. Using this method, discrete inspection items can be associated to better assist in problem location.
[0099] In some possible implementations, the database backup check includes checking whether the database backup switch is on and whether backup files exist in a preset database backup path. This approach can promptly detect issues such as accidentally turning off the database backup switch or accidentally deleting database backup files. In some possible implementations, upon determining that a node's database backup switch is off, the electronic device can turn the switch on and output a prompt indicating that the node's database backup switch has been turned on. If the device lacks permission to turn the switch on, it can output a prompt suggesting that the node's database backup switch be turned on to ensure data security.
[0100] In some possible implementations, checking the health status of the cluster module includes checking whether the module has any unrecovered alarm information.
[0101] Exemplarily, the electronic device can obtain unrecovered alarm information and the number of unrecovered alarm information by reading the cluster alarm information alarm table. The alarm table is used to record the alarm information of the modules detected by the cluster detection program. The alarm types may include but are not limited to host performance information alarms, database query rate alarms, message backlog alarms, request delay alarms, request error alarms, product function call error alarms, thread blocking alarms, etc. The detection program corresponding to the alarm table can also incrementally update new alarm detection types as the system version is upgraded.
[0102] Exemplarily, the electronic device may obtain the number of unrecovered alarm information through the following code statement: and when it is determined that unrecovered alarm information exists, determine that the inspection result includes prompt information for indicating unrecovered abnormal alarm information.
[0103] "dbxquery connection="rizhiyi_manager"query="select count(*)as cntfrom alarm where recover_time is NULL and module="zookeeper'"
[0104] eval des=if(cnt>0,"Abnormal","Healthy")”
[0105] Illustratively, some modules included in the cluster and the health status of the modules obtained by checking the health status of the cluster modules may be shown in Table 4 below.
[0106] Table 4
[0107] Check Module Health status Distributed coordination service module Zookeeper healthy Data acquisition module heka healthy Data cache module Kafka healthy Detection alarm module alert_manager healthy Field extraction module logriver healthy Data query module splserver healthy ... ...
[0108] In some possible implementations, the cluster detection program check includes: the number of cluster detections, and whether each detection program is in a normal detection state. Whether the detection program is in a normal detection state can be determined by whether the detection program switch is in an on state, and / or by whether the storage file path corresponding to the detection program contains file data. If the detection program switch is not turned on or the file content does not exist under the corresponding storage file path, it is determined that the inspection result includes indication information indicating that the detection program is not in a normal detection state. If the detection program switch is not turned on, the electronic device can enable the corresponding detection program and output a description that the detection program has been enabled.
[0109] In some possible implementations, the baseline check of important configuration items includes: checking at least two of the backlog information of the data cache queue (kafka), the data backlog information of the log data stream component (such as the field extraction module logriver), and the allocation delay information of the query service allocation query task (such as the data query module splserver allocation query task).
[0110] For example, Kafka backlog information may include but is not limited to: the maximum queue backlog time within a preset time (for example, within 24 hours or 48 hours) and the maximum backlog message queue size.
[0111] For example, the data backlog information of logriver includes but is not limited to: the maximum number of read threads (task queue) and the maximum number of write threads (sink queue) within a preset time (e.g., within 24 hours or 48 hours). Exemplarily, if the electronic device determines that the data backlog information of logriver is manifested as the sink queue being greater than the number of read threads task queue, it can be determined that the inspection result includes prompt information for indicating that logriver has a data backlog phenomenon.
[0112] For example, the allocation delay information of splserver includes but is not limited to: the average dispatch delay time of the service scheduling job and the average number of dispatch failures.
[0113] In some possible implementations, if an electronic device discovers a log parsing delay problem through a logriver data backlog inspection, the inspection result may include information indicating the logriver data backlog problem and suggesting that it be resolved by increasing the number of FluentdWorkers. If an electronic device discovers an uneven query task distribution problem through a splserver allocation delay inspection, the inspection result may include information indicating the uneven query task distribution problem and suggesting that it be resolved by improving the scheduling algorithm.
[0114] Using this approach, manager inspections include baseline inspections of important configuration items and core indicators that affect the health of the data pipeline. This helps predict changes in important configuration items in advance, prevent emergency failures, and accurately locate bottlenecks so that targeted performance optimization suggestions can be made.
[0115] S104, the electronic device performs a product function inspection on the log management system.
[0116] In an embodiment of the present application, product function inspection is used to inspect whether the core business functions facing users are working normally.
[0117] In an embodiment of the present application, product function inspections include: log query function inspection, log collection module function inspection, log parsing function inspection, data visualization function inspection, data report function inspection, scheduled task function inspection, tenant management function inspection, tenant usage limit inspection, and node user login behavior inspection.
[0118] In some possible implementations, checking the log query function includes checking whether the log query function normally returns query results. For example, the electronic device may send a query request to the log query function to check whether the log query function normally returns query results. In some possible implementations, if the electronic device determines that the log query function cannot normally return query results, it may determine that the inspection result includes information indicating that the log query function is abnormal.
[0119] In some possible implementations, the log collection module function check includes: checking the number of log collection modules, and checking whether the log collection module is collecting active log streams normally. Exemplarily, the electronic device can obtain the number of log collection modules in the log management system by whether the module name (such as the tag field) belongs to the preset log collection module name. Exemplarily, the electronic device can determine whether the log collection module is collecting active log streams normally by the distance between the timestamp of the last log event of the log file generated by the log collection module and the current time. For example, if the distance exceeds 1 hour during working hours, it is determined that the log collection module is malfunctioning. In some possible implementations, if the electronic device determines that the log collection module cannot collect active log streams normally, it can determine that the inspection result contains indication information for indicating that the log collection module is malfunctioning.
[0120] In some possible implementations, the log parsing function check includes: checking the number of field extraction modules in the log management system, and checking whether the field extraction module can parse the log normally. Exemplarily, the electronic device can parse the most recent log events (for example, the last 5 log events) by checking whether the field extraction module parses the most recent log events and whether the obtained field values are all null, thereby determining whether the field extraction module can parse the log normally. In some possible implementations, if the electronic device determines that the log parsing function cannot parse the log normally, it can determine that the inspection result includes indication information indicating that the log parsing function is abnormal.
[0121] In some possible implementations, the data visualization function check includes: checking the number of dashboards in the log management system, and checking whether the dashboards can display data normally.
[0122] In some possible implementations, the data report function check includes: checking the number of data reports in the log management system, and whether the data reports can normally output report data.
[0123] In some possible implementations, the scheduled task function check includes: checking the number of scheduled tasks in the log management system and whether the scheduled tasks are executed normally. Exemplarily, the scheduled task function may include but is not limited to collecting target indicator values on time at a preset set time (e.g., daily) and compiling statistics on the collected target indicator values at a preset set time (e.g., monthly).
[0124] In some possible implementations, the tenant management function check includes: checking at least two of: whether the tenant quota has been modified, the latest configuration of the number of times the tenant is allowed to exceed the limit, the modification of the tenant's account and / or password, and whether all tenant purchase features have been selected.
[0125] In some possible implementations, the usage limit check includes: checking at least two of the tenant quota, the tenant's actual traffic usage, the idle quota, the number of overuses, the tenant's highest daily traffic usage this month, and the tenant's average daily log increase this month.
[0126] For example, the following code statements can be used to check the actual traffic usage of the tenant and the number of times it exceeds the limit.
[0127] "dbxquery connection="rizhiyi_system"query="select name,round(limit_flow_quota / 1024 / 1024 / 1024,0)as limit_1from Domain where name="ops'"
[0128] eval a=1
[0129] join a[[
[0130] starttime="-1d / d"endtime="now"appname:rzyxj AND tag:license_info
[0131] |stats count()as cnt by mes.volume_in_GB / / / According to actual license traffic, group statistics, number of occurrences
[0132] |rename mes.volume_in_GB as xe
[0133] |eval xe=tolong(xe)
[0134] |eval a=1]]
[0135] |eval status=if(xe<=limit_1,"No","Yes")
[0136] |fields name,limit_1,xe,cnt,status”
[0137] For example, the data of the highest daily traffic usage of the tenant this month, the average daily log increase of the tenant this month, and the number of times the tenant exceeded the usage limit this month detected by the electronic device are shown in Table 5 below.
[0138] Table 5
[0139]
[0140] In some possible implementations, checking cluster user login behavior includes counting the number of user logins within a preset time period (e.g., 7 days). For example, the electronic device may count the number of admin users, the number of non-admin users, the number of non-admin user logins, and the number of system accesses in the log management system within 7 days.
[0141] S105 , issuing inspection reports based on the inspection results of the electronic equipment basic data inspection, manager inspection, and product function inspection.
[0142] In some possible implementations, the basic environment data inspection further includes: checking the latest version number of the log management system, and the inspection report includes the latest version number of the log management system.
[0143] For example, Figure 2 As shown, an inspection report is provided based on the inspection results corresponding to the basic data inspection, manager inspection, and product function inspection provided in an embodiment of the present application. The inspection report includes the health status of the system service, a description of the optimized problem, and optimization suggestions. For example, the health status of the system service includes: system usage (determined by checking the node user login behavior in the product function inspection), system operation status (determined by the corresponding inspection items in the product function inspection and the cluster detection program inspection in the manager inspection), host operation status (determined by the host performance inspection in the manager inspection), and manager operation status (determined by the cluster module health status check and the preset important configuration item baseline check in the manager inspection).
[0144] In an embodiment of the present application, the electronic device performs basic environment information inspection, manager inspection, and product function inspection on the log management system. The three inspection tasks can be executed in parallel or serially.
[0145] By adopting the log-based service health inspection method provided in the embodiment of the present application, on the one hand, the inspection task is automated, which solves the problems of low efficiency and poor data accuracy of traditional manual inspection methods under the microservice architecture.
[0146] On the other hand, a structured inspection approach is adopted to conduct layered inspections from three dimensions: basic environment information inspection, manager inspection, and product function inspection. Basic data inspection ensures the accuracy of underlying data, manager inspection ensures system health, and product function inspection verifies that business is normal. This can more accurately identify key risk points at different levels.
[0147] On the other hand, the inspection content covers everything from the underlying environmental data to the middle system layer and then to the top business application layer, eliminating inspection blind spots and enabling a comprehensive assessment of the overall health status and availability of the log management system.
[0148] Furthermore, clear dimensional divisions facilitate the clear responsibilities of each team, improving inspection efficiency and problem location. For example, basic information inspections are handled by the operations or infrastructure team, manager inspections are handled by administrators or the platform team responsible for cluster operations and performance management, and product functionality inspections are verified by the product support team, testing team, or end-user representatives. Different teams can conduct inspections of their respective responsibilities in parallel.
[0149] On the other hand, when an abnormality occurs in the system, by quickly checking the status of the three dimensions, the problem can be quickly located at a specific level, shortening the average time for troubleshooting and accelerating service restoration.
[0150] On the other hand, the clearly defined inspection items in each dimension make the inspection process highly standardized, ensuring that key inspection items are not missed, improving the reliability and consistency of the inspection results, and making the inspection results at different time points and between different clusters repeatable and comparable.
[0151] In some possible implementations, the electronic device may perform basic environment data inspection, manager inspection, and product function inspection on the log management system in sequence. That is, the basic environment data inspection is performed first, then the manager inspection, and finally the product function inspection.
[0152] Using this method, basic environment information inspection is carried out first, which can ensure that subsequent inspections are carried out based on the latest and most accurate underlying data of the log management system, reducing the impact of underlying data errors (such as system license expiration, organizational structure update) on subsequent inspection results (for example, license expiration, etc. causing subsequent inspection data distortion).
[0153] It should be noted that the electronic equipment performs basic environment data inspection, manager inspection, and product function inspection on the log management system in sequence in order to reduce the changes in the dependency of the next level inspection on the previous level inspection and the impact on the accuracy of the next level inspection, so as to better ensure the accuracy of the inspection tasks. However, the dependencies of the inspection tasks executed in sequence do not include strong dependencies. For example, it does not include a strong dependency that the next level inspection task cannot continue to execute if an error item is found in the previous level inspection. In other words, if an error item is found in the log management system at any level of inspection, it does not affect the continued execution of the next level inspection task.
[0154] In some possible implementations, if the electronic device's first inspection results indicate an abnormality, the electronic device still performs a second inspection in the inspection sequence. The first inspection is a basic environment data inspection or a manager inspection, and the second inspection is a manager inspection or a product function inspection.
[0155] In some possible implementations, when the inspection results of the first inspection indicate the presence of an abnormal item, the electronic device outputs a prompt indicating the presence of an abnormal item in the first inspection, and outputs a confirmation request requesting the user to confirm whether to continue with the subsequent inspection task or to suspend the inspection first. After receiving the user confirmation message confirming the continuation of the subsequent inspection task, the electronic device continues with the subsequent inspection task; or, if the electronic device does not receive the user confirmation message after a preset time period, it defaults to continuing with the subsequent inspection task. After receiving the user confirmation message to suspend the inspection first, the inspection is suspended, and the inspection is resumed after receiving the instruction to start the inspection again.
[0156] For example, when the electronic device determines that the inspection results of the basic environmental data inspection show that there are abnormal items, it can ask the user whether to continue to perform subsequent inspection tasks. After receiving the confirmation information from the user to continue, the electronic device continues to perform subsequent inspection tasks; or, if the electronic device does not receive the user's confirmation information after a preset period of time (for example, 30 seconds), it will continue to perform subsequent inspection tasks by default.
[0157] Using this method, when there are abnormal items in the previous inspection, the user is provided with the option of pausing the inspection so that the user can first solve the abnormal problems of the previous inspection, meet the user's diverse needs, and further reduce the overhead caused by the invalidation of subsequent inspection results due to abnormal problems.
[0158] In some possible implementations, a communication relationship exists between a basic environment data inspection module corresponding to the basic environment data inspection and a manager inspection module corresponding to the manager inspection. When it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes the addition of a first node, the basic environment data inspection module sends a first indication message to the manager inspection module (the first indication message is used to indicate the triggering of a manager inspection task for the newly added node), and the manager inspection module adds a manager inspection task for the first node based on the first indication message. The first indication message includes an identifier of the first node.
[0159] When it is found during the cluster organizational structure update check that the change information of the cluster organizational structure includes the deletion of the second node, the basic environment data inspection module sends a second indication information to the manager inspection module (the second indication information is used to indicate the manager inspection task for eliminating the departing node), and the manager inspection module determines not to perform a manager inspection on the second node based on the second indication information.
[0160] When it is found during the cluster organizational structure update check that the change information of the cluster organizational structure includes a change in the role information of the third node, the basic environment data inspection module sends a third indication information to the manager inspection module (the third indication information is used to indicate that the manager inspection task for the third node should be maintained, and the node name corresponding to the inspection result of the third node should be modified to the node name after the role change). Based on the third indication information, the manager inspection module maintains the execution of the manager inspection task of the third node, and modifies the node name corresponding to the inspection result of the third node to the node name after the role change.
[0161] This method further ensures the integrity and comprehensiveness of inspection tasks. Compared with manual discovery, creation, deletion, or modification of inspection tasks, it can trigger manager inspection tasks for newly added nodes more timely, accurately, and efficiently.
[0162] In some possible implementations, the inspection report issued by the electronic device may also include information for indicating the association relationship between inspection items, wherein, when the cluster module health status check finds the presence of alarm information related to log parsing, and the log parsing function check finds that the field value parsed by the field extraction module is null, the inter-module version compatibility information check finds that the version information of the upstream and downstream modules related to the field extraction module is abnormal, and the module patch package version information check finds that the patch package of the field extraction module is not the latest version, it is determined that the inspection item association information includes a first association map, and the first association map is used to display the problem location map relationship of the associated inspection items related to the log parsing alarm information.
[0163] This method analyzes the correlation information between independent inspection items and displays the correlation map of independent inspection items, which can better assist users in fault location.
[0164] In some possible implementations, the electronic device issues an inspection report including: when it is determined that the tenant's average idle quota is greater than the preset traffic and the number of excess usage is equal to 0, based on the average value of the tenant's actual traffic usage, determining that the inspection report includes a quota reduction suggestion, and the quota reduction suggestion is used to indicate the tenant's average idle quota and the recommended quota; when it is determined that the tenant's number of excess usage is greater than or equal to the tenant's corresponding allowed number of excesses, based on the average value of the tenant's actual traffic usage, determining that the inspection report includes a quota increase suggestion, and the quota increase suggestion is used to indicate the number of times the tenant's actual traffic usage exceeds the quota, and the recommended quota.
[0165] This approach provides practical and effective solutions based on inspection data, better assisting users in solving related problems and improving the user experience of the inspection method.
[0166] It can be understood that the embodiment of the present application uses an electronic device as an example to illustrate the execution subject of the log-based service health inspection method provided by the present application. The electronic device can also be understood as the log-based service health inspection device shown in the embodiment of the present application. In the embodiment of the present application, the electronic device can be a microprocessor or computer for executing program code, etc. The electronic device that can be used to execute the method provided by the embodiment of the present application belongs to the protection scope of the embodiment of the present application, and the present application does not impose any restrictions. For example, the electronic device can be a desktop computer, a portable notebook, a mobile terminal, a 32-bit microprocessor or a 64-bit microprocessor, etc., and the embodiment of the present application does not limit this.
[0167] An embodiment of the present application further provides a log-based service health inspection device, comprising a unit for executing any one of the log-based service health inspection methods in the method embodiments.
[0168] Please refer to Figure 3 , which provides a structural diagram of a log-based service health inspection device in an embodiment of the present application.
[0169] like Figure 3 As shown, the device may include:
[0170] The acquisition unit 301 is used to acquire log data of each node in the cluster and host performance data of each node;
[0171] The execution unit 302 is configured to perform a basic environment data inspection, a manager inspection, and a product function inspection on the log management system based on the log data and the host performance data;
[0172] The basic environment data inspection includes: the license information inspection of the log management system, the inter-module version compatibility information inspection, the module patch package version information inspection, and the cluster organizational structure update inspection;
[0173] The manager inspection includes: host performance inspection of each node, database backup inspection, module health status inspection, detection program inspection, and preset important configuration item inspection;
[0174] The product function inspection includes: log query function inspection, log collection module function inspection, log analysis function inspection, data visualization function inspection, data report function inspection, scheduled task function inspection, tenant management function inspection, tenant usage limit inspection, and node user login behavior inspection;
[0175] The output unit 303 is configured to issue an inspection report based on the inspection results corresponding to the basic data inspection, the manager inspection, and the product function inspection.
[0176] In some possible implementations, the execution unit 302 is configured to sequentially perform the basic environment data inspection, the manager inspection, and the product function inspection on the log management system.
[0177] In some possible implementations, the apparatus further includes:
[0178] The determination unit 304 is configured to determine that the manager inspection module corresponding to the manager inspection adds a manager inspection task for the first node when it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes the addition of a new first node; and to determine that the manager inspection module does not perform a manager inspection on the second node when it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes the deletion of a second node. When it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes the change of role information of a third node, the manager inspection module maintains executing the manager inspection task of the third node and determines to modify the node name corresponding to the inspection result of the third node to the node name after the role change.
[0179] In some possible implementations, the above-mentioned output unit 303 is specifically used to, when it is determined that the average idle quota of the tenant is greater than the preset traffic and the number of excess usage is equal to 0, determine based on the average value of the tenant's actual traffic usage, that the inspection report includes a quota reduction suggestion, and the quota reduction suggestion is used to indicate the average idle quota of the tenant and the recommended quota; when it is determined that the number of excess usage of the tenant is greater than or equal to the corresponding allowed number of excesses of the tenant, determine based on the average value of the tenant's actual traffic usage, that the inspection report includes a quota increase suggestion, and the quota increase suggestion is used to indicate the number of times the tenant's actual traffic usage exceeds the quota and the recommended quota.
[0180] For the description of the inspection items in the basic environment data inspection, manager inspection, and product function inspection, as well as the description of the inspection report, please refer to the above method embodiment and will not be described in detail here.
[0181] In the embodiments of the present application, any implementation method mentioned in the method embodiments is also applicable to the log-based service health inspection device provided by the present application. The specific execution steps can be found in the description of the aforementioned method embodiments and will not be described in detail here.
[0182] An embodiment of the present application further provides a log-based service health inspection device, comprising a processor, wherein the processor is configured to execute any one of the log-based service health inspection methods in the above method embodiments.
[0183] Please refer to Figure 4 , is a structural diagram of another log-based service health inspection device provided in an embodiment of the present application, such as Figure 4 As shown, the log-based service health inspection device 400 may include: at least one processor 401, such as a CPU, at least one communication interface 403, a memory 404, and at least one communication bus 402. The communication bus 402 is used to implement connection and communication between these components. The communication interface 403 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface or a Bluetooth interface, etc.). The memory 404 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 404 may optionally also be at least one storage device located away from the aforementioned processor 401. As Figure 4 As shown, the memory 404 as a computer storage medium may include an operating system, a network communication module, and program instructions.
[0184] exist Figure 4In the log-based service health inspection device 400 shown, the processor 401 can be used to load program instructions stored in the memory 404 and specifically perform the following operations:
[0185] Obtain log data of each node in the cluster, as well as host performance data of each node;
[0186] Based on the log data and the host performance data, perform basic environment data inspection, manager inspection, and product function inspection on the log management system;
[0187] The basic environment data inspection includes: the license information inspection of the log management system, the inter-module version compatibility information inspection, the module patch package version information inspection, and the cluster organizational structure update inspection;
[0188] The manager inspection includes: host performance inspection of each node, database backup inspection, module health status inspection, detection program inspection, and preset important configuration item inspection;
[0189] The product function inspection includes: log query function inspection, log collection module function inspection, log analysis function inspection, data visualization function inspection, data report function inspection, scheduled task function inspection, tenant management function inspection, tenant usage limit inspection, and node user login behavior inspection;
[0190] An inspection report is issued based on the inspection results corresponding to the basic data inspection, the manager inspection, and the product function inspection.
[0191] It should be noted that the specific execution process can be found in the specific description of the above method embodiment and will not be described in detail here.
[0192] The specific execution steps can be found in the description of the aforementioned method embodiment and will not be described in detail here.
[0193] An embodiment of the present application also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor and executing the log-based service health inspection method provided by an embodiment of the present application. The specific execution process can be found in the specific description of the method embodiment shown above, and will not be described in detail here.
[0194] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method steps of the method embodiment shown above.
[0195] An embodiment of the present application also provides a chip module, including a transceiver component and a chip, wherein the chip is used to execute the method steps of the method embodiment shown above.
[0196] It is understood that the log-based service health inspection system, log-based service health inspection device, computer storage medium, computer program, computer program product, and chip provided above are all used to execute the method shown in any implementation of the corresponding aspects of the embodiments of this application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding methods and will not be described in detail here.
[0197] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes the processes of the embodiments of the above-mentioned methods.
[0198] The at least one (item) involved in this application indicates one (item) or more (items). More than one (item) refers to two (items) or more than two (items). "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. In addition, it should be understood that although the terms first, second, etc. may be used to describe each object in this application, these objects should not be limited to these terms. These terms are only used to distinguish each object from each other.
[0199] As mentioned above, the terms "includes" and "having" and any variations thereof, are intended to cover a non-exclusive inclusion.
[0200] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A log-based service health inspection method, characterized in that: The method is used to perform a service health inspection on a log management system, where the log management system is a cluster including at least one node, each node including at least one module, and the method includes: Obtaining log data of each node in the log management system and host performance data of each node; Based on the log data and the host performance data, perform basic environment data inspection, manager inspection, and product function inspection on the log management system; The basic environment data inspection includes: the license information inspection of the log management system, the inter-module version compatibility information inspection, the module patch package version information inspection, and the cluster organizational structure update inspection; The manager inspection includes: host performance inspection of each node, database backup inspection, module health status inspection, detection program inspection, and preset important configuration item inspection; The product function inspection includes: log query function inspection, log collection module function inspection, log analysis function inspection, data visualization function inspection, data report function inspection, scheduled task function inspection, tenant management function inspection, tenant usage limit inspection, and node user login behavior inspection; An inspection report is issued based on the inspection results corresponding to the basic data inspection, the manager inspection, and the product function inspection.
2. The method according to claim 1, wherein The license information check includes: checking the latest validity period and usage limit of the license based on the log data; The cluster module compatibility information check includes: checking the compatibility of software versions between modules in the cluster based on preset compatibility conditions; The module patch package version information check includes: checking whether the patch package version of each module in the cluster is the latest version; The cluster organizational structure update check includes: obtaining the latest organizational structure of the cluster, and determining change information of the cluster organizational structure.
3. The method according to claim 1 or 2, wherein: The host performance check includes: checking the number of hosts in the cluster, the CPU performance data of each host, the memory performance data, the disk performance data, the disk input / output time ratio, the traffic volume within a preset time period, and the number of currently open files; The database backup check includes: checking whether the database backup switch is turned on and whether there is a backup file in the preset database backup path; The cluster module health status check includes: checking whether the module has unrecovered alarm information; The detection program check includes: the number of cluster detection programs, and whether each detection program is in a normal detection state; The preset important configuration item baseline check includes: checking at least two of the data cache queue backlog information, the data backlog information of the log data stream component, and the allocation delay information of the query service.
4. The method according to any one of claims 1 to 3, wherein The log query function check includes: checking whether the log query function returns the query result normally; The log collection module function check includes: checking the number of log collection modules and checking whether the log collection modules are collecting active log streams normally; The log parsing function check includes: checking the number of field extraction modules in the log management system, and checking whether the field extraction modules can parse logs normally; The data visualization function check includes: checking the number of dashboards in the log management system and checking whether the dashboards can display data normally; The data report function check includes: checking the number of data reports in the log management system and whether the data reports can output report data normally; The scheduled task function check includes: checking the number of scheduled tasks in the log management system and whether the scheduled tasks are executed normally; The tenant management function check includes: checking whether the tenant quota has been modified, the latest configuration of the number of times the tenant is allowed to exceed the limit, and at least two of the modification actions of the tenant's account and / or password; The usage limit check includes: checking at least two of the tenant's quota, the tenant's actual traffic usage, the idle quota, the number of overuses, the tenant's highest daily traffic usage this month, and the tenant's average daily log increase this month; The cluster user login behavior check includes: counting the number of user logins within a preset time period.
5. The method according to any one of claims 1 to 4, characterized in that The basic environment data inspection, manager inspection, and product function inspection of the log management system include: The basic environment data inspection, the manager inspection, and the product function inspection are performed on the log management system in sequence.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: When it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes a newly added first node, determining that the manager inspection module corresponding to the manager inspection adds a manager inspection task for the first node; If it is found in the cluster organizational structure update check that the change information of the cluster organizational structure includes deletion of the second node, it is determined that the manager inspection module does not perform manager inspection on the second node. When it is found during the cluster organizational structure update check that the change information of the cluster organizational structure includes a change in the role information of the third node, the manager inspection module is determined to continue executing the manager inspection task of the third node, and the node name corresponding to the inspection result of the third node is modified to the node name after the role change.
7. The method according to any one of claims 1 to 6, wherein: The inspection report includes information indicating the association relationship between inspection items. Among them, when the cluster module health status check finds the existence of alarm information related to log parsing, and the log parsing function check finds that the field value parsed by the field extraction module is null, the inter-module version compatibility information check finds that the version information of the upstream and downstream modules related to the field extraction module is abnormal, and the module patch package version information check finds that the patch package of the field extraction module is not the latest version, it is determined that the inspection item association information includes a first association map, and the first association map is used to display the problem location map relationship of the associated inspection items related to the log parsing alarm information.
8. The method according to any one of claims 1 to 7, wherein: The product function inspection includes tenant management function inspection and usage limit inspection. The tenant management function inspection includes checking the latest configuration of the number of times the tenant is allowed to exceed the limit. The usage limit inspection includes checking the tenant's actual traffic usage, average idle quota, and number of over-usage times; The inspection report includes: If it is determined that the average idle quota of the tenant is greater than the preset traffic volume and the number of excess usages is equal to 0, determining, based on the average value of the actual traffic usage of the tenant, that the inspection report includes a quota reduction suggestion, the quota reduction suggestion being used to indicate the average idle quota of the tenant and a recommended quota amount; When it is determined that the number of times the tenant exceeds the quota is greater than or equal to the allowed number of times the tenant exceeds the limit, the inspection report is determined to include a quota increase suggestion based on the average value of the tenant's actual traffic usage. The quota increase suggestion is used to indicate the number of times the tenant's actual traffic usage exceeds the quota, and the recommended quota amount.
9. A log-based service health inspection device, characterized in that: Comprising means for performing the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program. When the computer program is executed, the method according to any one of claims 1 to 8 is performed.