Lightweight standard monitoring method and system based on service error log
By using lightweight, non-invasive SDK intercepting and standardizing error log processing in the microservice architecture, the problem of log fragmentation and high transformation costs is solved, automated monitoring and efficient error management are realized, and system stability is improved.
Patent Information
- Application Number
- CN202510132283.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-10
AI Technical Summary
Under the microservice architecture, the diversity and inconsistency of error logs lead to fragmentation of log data, making it difficult to achieve centralized and automated monitoring, and the existing technology has problems such as high modification costs and error omissions.
The lightweight, non-invasive SDK is used to intercept the original error logs output by the service, collect and standardize the log information, realize centralized storage of logs and automated data collection, evaluate the error level through month-on-month and year-on-year analysis, and trigger the alarm mechanism.
It realizes standardized output and automated collection of logs, reduces transformation costs, improves the efficiency and accuracy of error management, and enhances the stability and reliability of the system.
Smart Images

Figure CN120123170A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer networks, and particularly to a lightweight standard monitoring method and system based on service error logs. Background Art
[0002] In the conventional processes of software development and operation and maintenance, continuous monitoring of newly deployed code by developers is a key link to ensure the stability and reliability of the system. Among them, error logs, as an important information source for diagnosing system anomalies and faults, their effective management and analysis are particularly important. Currently, the common practice in the industry relies on developers directly checking log files to identify potential problems.
[0003] However, this traditional method faces significant challenges in a microservices architecture. By splitting an application into multiple independent and loosely coupled services, the microservices architecture improves the flexibility and scalability of the system, but at the same time, due to the dispersion of services and being maintained by different teams, it is difficult to centrally manage log information. This characteristic increases the difficulty of cross-service error tracking, making log monitoring from a global perspective complex and inefficient.
[0004] To solve this problem, many existing technologies have provided different methods. For example, CN201610112551.6 discloses a page monitoring method, device, and system (publication date: 2020-09-11); CN201210421115.9 discloses a method and system for monitoring application logs (publication date: 2016-05-11); CN201510030017.6 discloses a method, device, and system for monitoring logs based on a software development kit (publication date: 2019-02-26);
[0005] However, these traditional methods all configure an alarm mechanism to trigger an alarm for specific error logs according to preset rules, or directly embed notification means such as email sending in the code. Although these measures can improve the timeliness of problem discovery to a certain extent, they still essentially belong to the category of manual management and have obvious limitations. It is difficult to adaptively cover all newly emerging error types, thus easily leading to the omission of key errors and affecting the overall stability and service quality of the system.
[0006] Analyzing more deeply, the root cause of the dilemmas encountered by the above two conventional monitoring methods lies in the diversity and inconsistency of error log outputs in a microservices environment. Since each service uses different log recording frameworks, formats, and storage strategies, this directly leads to the fragmentation of log data, setting obstacles for the implementation of centralized and automated monitoring. To achieve unified log monitoring across services, not only requires large-scale transformation and integration of existing log systems technically, but also involves coordination and standardization work at the organizational level, which undoubtedly increases the implementation cost and complexity.
[0007] To this end, the present invention proposes a lightweight standard monitoring method and system based on service error logs. Summary of the Invention
[0008] In view of this, the present invention hopes to provide a lightweight standard monitoring method and system based on service error logs to solve or alleviate the technical problems existing in the prior art, that is, in view of the difficulties in centralized and automated observation and the too large interference of historical error logs, as well as the large transformation cost caused by the inconsistent form of error logs, how to complete the standardization of logs with a lightweight and non-invasive SDK based on the log framework with low transformation cost, realize the centralized observation by means of an automated data collection method, and complete the error level classification by means of a year-on-year comparison warning method to exclude interference.
[0009] The technical solution of the present invention is realized as follows:
[0010] In a first aspect, a lightweight standard monitoring method based on service error logs:
[0011] (1) Overview:
[0012] The present invention aims to construct a lightweight and non-invasive log monitoring and alarm system to efficiently solve the problems of log centralization, automated observation difficulties and inconsistent error log forms under the microservice architecture. By providing a standardized error log output interception SDK, this solution can collect key information in the original error logs and perform standardized processing to ensure the consistency of log formats. Subsequently, the processed log information is classified and stored to provide data support for subsequent alarm calculation and query. The system regularly performs alarm calculation on the stored log data, evaluates the error level through year-on-year analysis, and automatically triggers the alarm mechanism when the error quantity or ratio reaches a preset threshold, and sends the alarm information to the specified team or personnel in a timely manner. In addition, this solution also has a loop monitoring function, which can judge whether to terminate the execution of the solution according to preset conditions, so as to realize continuous log monitoring and error management. This solution not only reduces the transformation cost, but also improves the efficiency and accuracy of error management, providing a strong guarantee for the stability and reliability of the system under the microservice architecture.
[0013] (2) Technical solution:
[0014] To achieve the above technical objectives, after the input log monitoring scheme activation instruction is received, the present invention selects to perform the following operation steps.
[0015] 2.1 Step S1, SDK initialization:
[0016] The SDK starts to intercept the original error logs output by the service and collect the standard information therein, including the service name, timestamp, error type, static error message and error location.
[0017] 2.1.1 Step S100, Load the SDK library:
[0018] Load the SDK library (such as JAR packages, DLL files, etc.) into the running environment of the application through the dependency management tool of the project (such as Maven, Gradle for Java projects, or NuGet for.NET projects).
[0019] 2.1.2 Step S101, Initialize the SDK instance:
[0020] Call the initialization method provided by the SDK, pass the log storage path and the network request timeout time, and create an instance object of the SDK in the application.
[0021] 2.1.3 Step S102, Configure the log interception rules:
[0022] Based on the log interception rules of the SDK that can identify and intercept the original error logs output by the service, so that it can identify and intercept the original error logs output by the service. It includes configuring the log level (ERROR or WARN), log format, and service identifier.
[0023] 2.1.4 Step S103, Set the standard information for log collection:
[0024] Define the standard information fields that the SDK needs to collect when intercepting logs, including service name, timestamp, error type, static error message, and error location.
[0025] 2.2 Step S2, Log standardization processing:
[0026] Perform standardization processing on the collected log information, and classify and store the standardized log information in the log system to provide data support for subsequent alarm calculation and query.
[0027] 2.2.1 Step S200, Log information cleaning:
[0028] Use regular expressions to remove irrelevant characters, blank lines, duplicate logs, etc. from the collected original log information to ensure the accuracy and validity of the log information.
[0029] 2.2.2 Step S201, Log format parsing:
[0030] Use a parser to parse the cleaned log information according to the predefined log format (such as JSON, XML, custom format, etc.), and extract the key fields, including service name, timestamp, error type, error message, and error location.
[0031] 2.2.3 Step S202, Log Information Standardization:
[0032] Use the mapping table to standardize the parsed log information fields, including timestamp formatting, error type coding unification, and error message normalization.
[0033] 2.2.4 Step S203, Log Information Classification:
[0034] Use classification rules, conditional judgment, hash table, or database query to classify the standardized log information according to business requirements or log attributes (such as service name or / and error type, etc.) for subsequent storage and query.
[0035] 2.2.5 Step S204, Log Information Storage:
[0036] Use a database, file system, or distributed storage system to store the classified log information in the log system; use a full-text search engine, inverted index, or database index to create an index for the stored log information to optimize query performance.
[0037] 2.3 Step S3, Alarm Calculation:
[0038] Perform alarm calculation on the stored log data, calculate the newly added errors compared with the previous execution cycle, and calculate the error ratio compared with the previous execution cycle to evaluate the error level. According to the alarm calculation result, determine whether to trigger the alarm mechanism. If the number or ratio of errors reaches the preset threshold, execute Step S4; otherwise, execute Step S5.
[0039] 2.3.1 Step S300, Data Preparation:
[0040] Use SQL query, API call, or data stream processing to filter the log data of the current execution cycle and the previous execution cycle from the log system according to the timestamp field.
[0041] 2.3.2 Step S301, New Error Calculation:
[0042] Compare the log data of the current execution cycle and the previous execution cycle, and calculate the number of newly added errors. The method is as follows:
[0043] S3010, Count the error types in the log data of the current execution cycle to obtain the quantity of each error type;
[0044] S3011, Count the error types in the log data of the previous execution cycle to obtain the quantity of each error type in the previous cycle;
[0045] S3012, Compare the quantity of error types in the two cycles, calculate the difference value, which is the number of newly added errors.
[0046] 2.3.3 Step S302, Error Ratio Calculation:
[0047] Calculate the error ratio for the current execution cycle and the previous execution cycle to evaluate the error trend. The method is as follows:
[0048] S3020, Calculate the ratio of the total number of errors to the total number of logs in the current execution cycle to obtain the error ratio for the current cycle;
[0049] S3021, Calculate the ratio of the total number of errors to the total number of logs in the previous execution cycle to obtain the error ratio for the previous cycle;
[0050] S3022, Compare the error ratios of the two cycles, calculate the difference or ratio change, and use it to evaluate the error level.
[0051] 2.3.4 Step S303, Error Level Evaluation:
[0052] Evaluate the current error level based on the change in the number of new errors and the error ratio. The method is as follows:
[0053] S3030, Set the threshold ranges for the number of new errors and the error ratio corresponding to "low", "medium", and "high";
[0054] S3030, Match the corresponding error level according to the calculated number of new errors and the error ratio.
[0055] 2.3.5 Step S304, Alarm Judgment:
[0056] Judge whether it is necessary to trigger the alarm mechanism according to the error level and the preset alarm conditions. The method is as follows:
[0057] If the error level reaches or exceeds the preset alarm level (such as the "high" level), trigger the alarm.
[0058] If the number of errors or the ratio reaches or exceeds the preset specific threshold (such as the number of new errors exceeds 100, or the error ratio increases by more than 10%), trigger the alarm.
[0059] 2.3.6 Step S305, Execute the Corresponding Steps
[0060] If the alarm is triggered, execute Step S4 (Alarm Triggered).
[0061] If the alarm is not triggered, execute Step S5 (Continuous Monitoring).
[0062] 2.4 Step S4, Alarm Notification:
[0063] Send the alarm information to the designated team or person.
[0064] 2.4.1 Step S400, Alarm Information Assembly:
[0065] According to the alarm calculation result, assemble an alarm message containing key information. This includes error type, error quantity, error ratio, scope of influence, and occurrence time, and perform filling according to a preset alarm template to form a complete alarm message.
[0066] 2.4.2 Step S401, Determine Alarm Recipients:
[0067] Based on the alarm type and urgency, determine the teams or personnel who need to receive the alarm information.
[0068] This includes a configuration table or database of alarm recipients, which records different alarm types and corresponding recipient information (email address, phone number, or instant messaging account). Query the configuration table according to the alarm type to obtain the recipient information.
[0069] 2.4.3 Step S402, Send Alarm Information:
[0070] For email alarms, use an email sending library or API to send the alarm message as the email content to the recipient's email address;
[0071] For SMS alarms, use an SMS sending service to send the alarm message as the SMS content to the recipient's mobile phone number;
[0072] For phone alarms, use an automatic dialing system or voice synthesis technology to call the recipient's phone number and play the alarm message;
[0073] For instant messaging alarms, use an instant messaging API or SDK to send the alarm message to the recipient's instant messaging account.
[0074] 2.5 Step S5, Loop Monitoring:
[0075] If a stop instruction is received or the system shuts down, terminate the execution; otherwise, perform loop monitoring.
[0076] (III) Mechanism for Solving Technical Problems:
[0077] 3.1 Log Interception and Collection:
[0078] Utilize a lightweight and non-invasive SDK to achieve the interception and collection of raw error logs without changing the existing microservice architecture and code. The SDK collects key information in the logs, such as service name, timestamp, error type, static error message, and error location, to ensure the comprehensiveness and accuracy of the information.
[0079] 3.2 Centralized Log Storage:
[0080] The classified and standardized log information is stored in the log system to achieve centralized management of log data. Centralized storage facilitates subsequent data mining and analysis work, and also helps to reduce storage costs and improve query efficiency.
[0081] 3.3 Alarm calculation and early warning analysis:
[0082] The system regularly performs alarm calculations on the stored log data, and evaluates the error level through month-on-month (such as comparing with yesterday) and year-on-year (such as comparing with last week) analysis. According to the preset alarm threshold, it judges whether it is necessary to trigger the alarm mechanism. If the number or proportion of errors reaches the preset threshold, alarm information will be automatically generated.
[0083] 3.4 Error level classification and interference elimination:
[0084] The error level classification is completed through month-on-month and year-on-year early warning methods, and the errors are divided into different levels (such as severe, general, minor, etc.), so that relevant personnel can adopt different processing strategies according to the error level. For interfering information in historical error logs (such as frequently occurring known errors, non-critical errors, etc.), the system should provide a filtering mechanism or an automatic ignoring function to reduce interference with normal monitoring work.
[0085] In the second aspect, a lightweight standard monitoring system based on service error logs:
[0086] As Figure 3 shown, it includes a memory storing program instructions, and a microservice node connected to the memory. When the microservice node executes the program instructions, the lightweight standard monitoring method based on service error logs described above can be implemented. Wherein the processor is connected with:
[0087] (1) A log collection module responsible for initializing the SDK and intercepting the original error logs output by the service: collecting standard information (service name, timestamp, error type, static error message, and error location).
[0088] Directly interact with the microservice node and receive the original log data.
[0089] (2) A log processing module for standardizing the collected log information: including format unification, field mapping, etc., and classifying and storing the processed log information in the log system.
[0090] Receive data from the log collection module, process and store it; at the same time, provide data support for the alarm calculation module.
[0091] (3) An alarm calculation module that periodically performs alarm calculations on the stored log data: Analyze the error quantity and ratio year-on-year, and evaluate the error level. Determine whether to trigger the alarm mechanism according to the calculation results.
[0092] Obtain data from the log processing module and perform calculation and analysis; if an alarm is required, call the alarm trigger module.
[0093] (4) An alarm delivery module responsible for sending alarm information to a specified team or person: Such as emails, text messages, instant messages, etc.
[0094] Receive the alarm instruction from the alarm calculation module and execute the sending of the alarm information.
[0095] (5) A monitoring and control module responsible for circularly monitoring the running status of the entire system: Determine whether to terminate the execution of the solution according to the preset conditions.
[0096] Monitor the running conditions of other modules, receive external stop instructions or system shutdown signals, and control the termination of the solution.
[0097] Compared with the prior art, the beneficial effects of the present invention are:
[0098] First, improve the log management efficiency: By introducing a lightweight and non-invasive SDK, the present invention realizes the standardized output and automated collection of logs, greatly improving the efficiency of log management. R & D personnel do not need to manually sort and analyze logs, and the system can automatically complete these tasks, thus saving time and effort.
[0099] Second, reduce the transformation cost: Due to the lightweight and non-invasive design of the SDK, the present invention does not require large-scale modification of the existing microservice architecture and code during implementation. The transformation cost is reduced, making the present invention easier to be accepted and promoted by enterprises and teams.
[0100] Third, improve the error location speed: The present invention can quickly locate and analyze errors through a centralized and automated log observation and alarm system. When an exception or fault occurs in the system, R & D personnel can quickly obtain relevant log information, so as to quickly locate the problem and take corresponding treatment measures.
[0101] Fourth, optimize the system stability: Through the real-time alarm and error level classification functions, the present invention can timely detect and respond to potential security threats and fault risks. This helps to optimize the system stability and reduce business interruptions and losses caused by faults.
[0102] Fifth, improve the team collaboration efficiency: The present invention provides a unified log monitoring and alarm platform, enabling different teams to share log information and strengthen collaboration. This helps to improve the team collaboration efficiency and jointly address system problems and challenges. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0104] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0105] Figure 2 It is a schematic diagram of the interactive steps of the method practice of the present invention;
[0106] Figure 3 It is a schematic diagram of the system composition of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0107] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific embodiments of the present invention with reference to the drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;
[0108] It should be noted that the embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0109] Explanation of related terms:
[0110] (1) SDK library: Software Development Kit, which provides specific functions for developers to call in applications.
[0111] (2) Error type: Error classification, which is used to identify and distinguish different types of errors.
[0112] (3) Static error message: Fixed error prompt information, which does not change with the error context.
[0113] (4) Error location: The code location where the error occurs, including file name, line number, function name, etc.
[0114] (5) JAR package: Java Archive file, which contains Java classes, resources and metadata, and is used for Java application distribution.
[0115] (6) DLL file: Dynamic Link Library, which contains functions and resources that can be called by a program during runtime and is used for the Windows platform.
[0116] (7) Log level (ERROR or WARN), log format, and service identifier: Define the importance and format of the log, as well as the service that identifies the log source.
[0117] (8) Parser: A program component used to parse and process data or files.
[0118] (9) Alarm template: A predefined alarm message format used to quickly generate alarm information.
[0119] Example 1: As Figures 1-2 shown, this example discloses a specific implementation solution in the online car-hailing background management system. To ensure the stability and efficiency of the service, it is necessary to centrally and automatically process and analyze the error logs generated in the system. Introduce a lightweight SDK to intercept the original error logs output by the service, and perform standardized processing, alarm calculation, and alarm reaching to achieve effective management and warning of error logs. The process is as follows:
[0120] In this example, regarding step S1: SDK initialization: Enable the SDK to start intercepting and processing the original error logs output by the service.
[0121] Specifically, load the SDK library (step S100): In the Java project of the online car-hailing background management system, use the Maven dependency management tool to load the JAR package of the SDK into the running environment of the application.
[0122] Specifically, initialize the SDK instance (step S101): When the application starts, call the initialization method provided by the SDK and pass the log storage path (such as / var / log / ride-hailing-backend) and the network request timeout (such as 5 seconds) to create an instance object of the SDK.
[0123] Specifically, configure the log interception rule (step S102): Set the log interception rule of the SDK so that it can identify and intercept the original error logs output by the service. Configure the log level to ERROR and WARN, the log format to JSON, and specify the service identifier as ride-hailing-backend.
[0124] Specifically, set the log collection standard information (step S103): Define the standard information fields that the SDK needs to collect when intercepting logs, including service name (ride-hailing-backend), timestamp, error type, static error message, and error location (such as file name, line number).
[0125] In this embodiment, regarding step S2: Log standardization processing: Clean, parse, standardize, and classify and store the collected log information to provide data support for subsequent alarm calculation and query.
[0126] Specifically, log information cleaning (step S200): Use regular expressions to remove irrelevant characters, blank lines, and duplicate logs from the collected original log information to ensure the accuracy and effectiveness of the log information.
[0127] Specifically, log format parsing (step S201): Use a JSON parser to parse the cleaned log information and extract key fields, such as service name, timestamp, error type, error message, and error location.
[0128] Specifically, log information standardization (step S202): Use a mapping table to standardize the parsed log information fields. For example, format the timestamp as YYYY-MM-DD HH:MM:SS, unify the error type encoding into a predefined encoding, and normalize the error message into a concise description.
[0129] Specifically, log information classification (step S203): Classify the standardized log information according to the service name and error type. For example, classify "order processing error" into one category and "payment error" into another category.
[0130] Specifically, log information storage (step S204): Use the Elasticsearch distributed storage system to store the classified log information in the log system and establish an inverted index to optimize query performance.
[0131] In this embodiment, regarding step S3: Alarm calculation: Perform alarm calculation on the stored log data, evaluate the error level, and determine whether to trigger the alarm mechanism.
[0132] Specifically, data preparation (step S300): Use the query API of Elasticsearch to filter the log data of the current execution cycle (such as the past hour) and the previous execution cycle (such as the previous hour) from the log system.
[0133] Specifically, new error calculation (step S301): Compare the log data of the two cycles and calculate the number of new errors. For example, it is found that 5 new error types have occurred in the current cycle.
[0134] Specifically, error ratio calculation (step S302): Calculate the error ratio of the current cycle and the previous cycle to evaluate the error trend. For example, the error ratio in the current cycle is 0.05%, and in the previous cycle it was 0.03%, indicating an increase in the error ratio.
[0135] Specifically, error level evaluation (step S303): Evaluate the current error level as "medium" based on the change in the number of new errors and the error ratio.
[0136] Specifically, alarm judgment (step S304): Determine whether to trigger the alarm mechanism according to the error level and the preset alarm conditions (such as triggering an alarm above the "medium" level).
[0137] Specifically, execute the corresponding steps (step S305): Since the alarm is triggered, execute step S4 (alarm trigger).
[0138] In this embodiment, regarding step S4: Alarm reach: Send the alarm information to the designated team or personnel in a timely manner so that they can respond and handle it in a timely manner.
[0139] Specifically, alarm information assembly (step S400): Assemble an alarm message containing key information according to the alarm calculation result, such as "There is a medium-level error in the online car-hailing background management system, 5 new error types have been added, and the error ratio has risen to 0.05%. Please handle it as soon as possible."
[0140] Specifically, determine the alarm recipient (step S401): Determine the team or personnel who need to receive the alarm information according to the alarm type and urgency, such as the background operation and maintenance team.
[0141] Specifically, send the alarm information (step S402): Use the email sending library to send the alarm message as the email content to the email address of the background operation and maintenance team; at the same time, use the SMS sending service to send the alarm message to the mobile phone number of the team leader to ensure that they can receive and handle the alarm information in a timely manner.
[0142] In this embodiment, regarding step S5: Loop monitoring: Continuously monitor the error logs of the online car-hailing background management system to detect and handle potential problems in a timely manner.
[0143] Specifically, in the absence of a stop instruction or system shutdown, the SDK will continuously intercept and process the original error logs output by the service, and perform log standardization processing, alarm calculation, and alarm reach according to the above process to achieve loop monitoring of the online car-hailing background management system.
[0144] In the solution provided in this embodiment: The original error logs output by the service are intercepted by the lightweight SDK and standardized, solving the problem of high transformation costs caused by inconsistent error log formats. The non-invasive design of the SDK reduces the changes and impacts on the application, making log collection simpler and more efficient. Through the automated data collection method of the SDK, centralized observation and management of error logs are achieved. This helps the operation and maintenance personnel to discover and handle potential problems in a timely manner, improving the stability and reliability of the system.
[0145] Furthermore, through the calculation methods of month-on-month and year-on-year, early warning analysis is carried out on the error logs. This helps the operation and maintenance personnel to accurately evaluate the error level and trend, and take corresponding measures for processing and prevention in a timely manner. The entire solution is based on the lightweight and non-invasive design concept of the log framework, realizing efficient log processing and early warning functions with low transformation costs. This enables the ride-hailing backend management system to better meet the operation and maintenance requirements and improve system stability while maintaining the original functions unchanged.
[0146] Embodiment 2: This embodiment further discloses a python execution program for the lightweight SDK interception service solution in the ride-hailing backend management system provided in Embodiment 1:
[0147] import json
[0148] import re
[0149] import datetime
[0150] import smtplib
[0151] from email.mime.text import MIMEText
[0152] from elasticsearch import Elasticsearch
[0153] # Elasticsearch configuration
[0154] ES_HOST = "localhost"
[0155] ES_PORT = 9200
[0156] ES_INDEX = "ride-hailing-backend-logs"
[0157] # Email sending configuration
[0158]
[0159]
[0160]
[0161]
[0162]
[0163]
[0164] In the above program, the sdk_init function executes the initialization process of the SDK, including loading the SDK library, creating instances, and configuring interception rules. The standardize_log function is responsible for cleaning, parsing, and standardizing logs. It uses regular expressions to remove extra spaces, parses JSON-formatted logs, and performs standardization processing (such as timestamp formatting, error type encoding, etc.). The store_log_to_es function stores the standardized logs in Elasticsearch for subsequent querying and analysis.
[0165] The alarm_calculation function executes the query of log data for the current cycle and the previous cycle from Elasticsearch, calculates the number of new errors, error ratio, evaluates the error level based on the calculation results, and determines whether to trigger an alarm.
[0166] The trigger_alarm function is responsible for assembling alarm information and sending it to the specified team or person via email and SMS.
[0167] The main function executes the process of loop monitoring, obtains the original error logs from the source, performs standardization processing, storage, and alarm calculation, and executes alarm triggering if an alarm is triggered.
[0168] All of the above embodiments only represent the implementation manners of the relevant practical applications of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
[0169] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0170] At the same time, those skilled in the art can understand that all or part of the processes of implementing the methods of all the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media provided in the present application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
Claims
1. A lightweight standard monitoring method based on service error logs, characterized in that: When the log monitoring solution activation command is entered, the following execution steps are included: S1, SDK starts to intercept the original error log output by the service and collects standard information, including service name, timestamp, error type, static error message and error location; S2, standardize the collected log information, and classify and store the standardized log information in the log system; S3, perform alarm calculation on the stored log data, compare the number of new errors in the calculation in the previous execution cycle, and compare the ratio of calculation errors in the previous execution cycle; if the number or ratio of errors reaches a preset threshold, execute S4, otherwise execute S5; S4, sending the alarm information to the designated team or personnel; S5: If a stop command is received or the system is shut down, the execution is terminated; otherwise, the monitoring is cyclic.
2. The standard monitoring method according to claim 1, characterized in that: The implementation method of S1 includes: S100, download the SDK into the application's running environment through the project's dependency management tool; S101, calling the initialization method provided by the SDK, passing the log storage path and network request timeout, and creating an instance object of the SDK in the application; S102, intercepting the original error log output by the service based on the log interception rule of the SDK that is set to be able to identify and intercept the original error log output by the service; S103, defining standard information fields that the SDK needs to collect when intercepting logs, including service name, timestamp, error type, static error message, and error location.
3. The standard monitoring method according to claim 1, characterized in that: The implementation method of S2 includes: S200, using regular expressions to remove irrelevant characters, blank lines, and duplicate logs from the collected original log information; S201, using a parser to parse the cleaned log information according to a predefined log format, extracting key fields, including service name, timestamp, error type, error message, and error location S202, using a mapping table to standardize the parsed log information fields, including timestamp formatting, error type coding unification, and error message normalization; S203, using classification rules, conditional judgment, hash table or database query, and classifying the standardized log information according to business requirements or log attributes.
4. The standard monitoring method according to claim 1, characterized in that: The implementation method of S3 includes: S300, use SQL query, API call or data stream processing to filter out the log data of the current execution cycle and the previous execution cycle based on the timestamp field from the log system S301, comparing the log data of the current execution cycle and the previous execution cycle, and calculating the number of new errors; S302, calculating the error ratio of the current execution cycle and the previous execution cycle to evaluate the error trend S303, evaluating the current error level based on the change in the number of new errors and the error ratio S304: If the error level reaches or exceeds a preset alarm level, an alarm is triggered; or if the error quantity or ratio reaches or exceeds a preset specific threshold, an alarm is triggered.
5. The standard monitoring method according to claim 4, characterized in that: The implementation method of S301 is: S3010, counting the error types in the log data of the current execution cycle to obtain the number of each error type; S3011, counting the error types in the log data of the previous execution cycle to obtain the number of each error type in the previous cycle; S3012, comparing the number of error types in two cycles and calculating the difference, which is the number of new errors.
6. The standard monitoring method according to claim 4, characterized in that: The implementation method of S302 is: S3020, calculating the ratio of the total number of errors in the current execution cycle to the total number of logs to obtain the error ratio of the current cycle; S3021, calculating the ratio of the total number of errors in the previous execution cycle to the total number of logs to obtain the error ratio of the previous cycle; S3022, comparing the error ratios of the two cycles and calculating the difference or ratio change.
7. The standard monitoring method according to claim 4, characterized in that: The implementation method of S303 is: S3030, setting the threshold ranges of the number of new errors and the error ratio to thresholds corresponding to "low", "medium" and "high"; S3030: Match corresponding error levels according to the calculated number of new errors and error ratio.
8. The standard monitoring method according to any one of claims 1 to 7, characterized in that: In S4, according to the alarm calculation result, an alarm message containing key information is assembled, including error type, error number, error ratio, impact range and occurrence time, and filled in according to a preset alarm template to form a complete alarm message.
9. A lightweight standard monitoring system based on service error logs, characterized in that: The system comprises: A memory storing program instructions, and a microservice node connected to the memory, wherein when the microservice node executes the program instructions, the standard monitoring method according to any one of claims 1 to 8 is implemented.
10. The system according to claim 9, characterized in that: The system further comprises: The log collection module is responsible for initializing the SDK and intercepting the original error logs output by the service; A log processing module that performs standardized processing on the collected log information; An alarm calculation module that periodically performs alarm calculations on stored log data; An alarm contact module responsible for sending alarm information to designated teams or personnel; A monitoring and control module responsible for cyclically monitoring the operating status of the entire system.
Citation Information
Patent Citations
Method and system for monitoring application logs
CN102981943A
A method, apparatus, and system for monitoring logs based on a software development kit.
CN105871574B
Page monitoring methods, devices and systems
CN107133240B