Log processing method, device and equipment
By decoupling exception determination and alarm determination, abnormal records and alarm information are generated, and fault repair tasks are generated according to conditions, the hardware failure duplication repair problem caused by massive equipment logs is solved, and the efficiency of log processing and fault repair is improved.
Patent Information
- Application Number
- CN202311447657.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-06
AI Technical Summary
The log data generated by massive equipment is large and complex, resulting in repeated repairs and processing of hardware failures, affecting maintenance efficiency.
By setting different alarm determination conditions, decoupling abnormal determination and alarm determination, generating abnormal records and generating alarm information based on alarm conditions. If no relevant fault repair tasks are found, fault repair tasks will be generated to avoid repeated establishment tasks.
It improves log processing efficiency, avoids fault repair efficiency affected by repeated repair processing, and meets the diverse detection and processing needs of a large number of servers.
Smart Images

Figure CN119938365A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a log processing method, device and equipment. Background Art
[0002] Hardware security is the bottom-level security of the server. Proper detection and processing of various hardware and timely handling of hardware failures are the basis for ensuring the normal operation of system services. Usually, logs are an important basis for diagnosing machine hardware failures. Logs record the performance status of the server itself, as well as the modules and processes it runs. Analyze the logs to extract information related to the hardware status and determine the status of the hardware.
[0003] At present, in the process of log analysis, the generated logs are often collected and stored before detecting anomalies in the logs; however, the logs generated by a large number of devices (such as servers) have a large amount of data and are complex, and different users may have diverse anomaly determination and fault alarm requirements for logs, which can easily lead to repeated maintenance of hardware faults, thereby affecting the efficiency of hardware fault maintenance. Summary of the invention
[0004] In view of this, the embodiments of the present application propose a log processing method, device and equipment, which can decouple abnormal judgment and alarm judgment, flexibly configure detection and processing requirements by setting different alarm judgments, and perform fault repair task queries through target triples and first triples, thereby avoiding repeated establishment of fault repair tasks based on abnormal logs, improving the efficiency of log processing, and further improving the efficiency of repair processing.
[0005] The embodiment of the present application is implemented by adopting the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a log processing method, the method comprising: determining an abnormal log in a set of equipment operation logs; generating an abnormal record according to target slot information of a target component in the abnormal log, target device information corresponding to a target device from which the abnormal log originates, and an event code corresponding to the abnormal log, wherein the target component refers to a component indicated by the abnormal log as having an abnormality; if the number of abnormal records including the same event code generated for a target component in a target device within a first time period reaches an alarm condition corresponding to the event code, generating alarm information for the target component based on a target fault code corresponding to the alarm condition; if a fault repair task associated with a target triplet is not found in an unfinished task set, and a fault repair task associated with a first triplet is not found, generating a fault repair task for the target component based on the alarm information, and adding the generated fault repair task to the unfinished task set; the target triplet comprises target device information, target slot information, and target fault code, and the first triplet is obtained by replacing a target fault code in the target triplet with a fault code mutually exclusive with the target fault code; and sending the fault repair task in the unfinished task set.
[0007] In a second aspect, an embodiment of the present application provides a log processing device, the device comprising: a detection module for determining an abnormal log in a device operation log set; an abnormal recording module for generating an abnormal record based on target slot information of a target component in the abnormal log, target device information corresponding to a target device from which the abnormal log originates, and an event code corresponding to the abnormal log, wherein the target component refers to a component indicated by the abnormal log as having an abnormality; an alarm module for, if the number of abnormal records including the same event code generated for a target component in a target device within a first time period reaches an alarm condition corresponding to the event code, based on a target fault code corresponding to the alarm condition, Generate alarm information for the target component; a task generation module, which is used to generate a fault repair task for the target component based on the alarm information if the fault repair task associated with the target triplet is not found in the unfinished task set and the fault repair task associated with the first triplet is not found, and add the generated fault repair task to the unfinished task set; the target triplet includes target device information, target slot information and target fault code, and the first triplet is obtained by replacing the target fault code in the target triplet with a fault code that is mutually exclusive with the target fault code; a sending module is used to send the fault repair tasks in the unfinished task set.
[0008] In some implementations, the detection module is used to determine, for each device operation log in the device operation log set, that the device operation log is an abnormal log if the device operation log includes an abnormal keyword.
[0009] In some embodiments, the detection module is also used to identify anomalies in each device operation log in the device operation log set through an abnormal log detection model, and obtain an abnormal classification result of the device operation log, wherein the abnormal classification result includes a first classification result and a second classification result, and the first classification result is used to indicate whether the device operation log is an abnormal log; the second classification result is used to indicate abnormal keywords in the device operation log; wherein the abnormal log detection model is trained in the following manner: obtaining multiple sample device operation logs and label information corresponding to each sample device operation log; the label information includes first label information and second label information, and the first label information is used to indicate whether the sample device operation log is The sample device operation log is an abnormal log, and the second label information is used to indicate abnormal keywords in the sample device operation log; the abnormal log detection model performs abnormal identification based on the sample device operation log to determine the sample classification result corresponding to the sample device operation log, the sample classification result includes a first sample classification result and a second sample classification result, the first sample classification result is used to indicate whether the sample device operation log is an abnormal log; the second sample classification result is used to indicate abnormal keywords in the sample device operation log; the model loss is determined according to the sample classification result corresponding to the sample device operation log and the label information corresponding to the sample device operation log; the weight parameters of the abnormal log detection model are reversely adjusted according to the model loss until the model training end condition is met.
[0010] In some embodiments, the exception record module includes a matching unit and an extraction unit; the matching unit is used to determine the event code corresponding to the exception keyword in the exception log based on the correspondence between the exception keyword and the event code, and use the event code corresponding to the exception keyword in the exception log as the event code corresponding to the exception log; the extraction unit is used to extract the target slot information of the target component from the exception log based on a preset slot extraction regular expression.
[0011] In some embodiments, the task generation module includes a shielding unit, a mutual exclusion unit and a query unit; the shielding unit is used to determine the target triplet based on the alarm information if it is determined according to a preset shielding strategy that the target device corresponding to the alarm information belongs to a non-shielded object; the mutual exclusion unit is used to replace the target fault code in the target triplet with a fault code that is mutually exclusive with the target fault code to obtain a first triplet; the query unit is used to perform task query in the unfinished task set based on the target triplet and the first triplet.
[0012] In some embodiments, the query unit is further used to determine that there is no need to create a new fault repair task for the target component based on the alarm information if a fault repair task associated with the target triplet is found in the unfinished task set, or a fault repair task associated with the first triplet is found.
[0013] In some embodiments, the unfinished task set includes task queues corresponding to multiple fault levels; the higher the fault level corresponding to the task queue where the fault repair task is located, the higher the sending priority of the fault repair task. The task generation module is also used to determine the target fault level corresponding to the target fault code based on the correspondence between the fault code and the fault level; and add the generated fault repair task to the task queue corresponding to the target fault level.
[0014] In some implementations, the log processing device further includes an update module, which is configured to receive repair completion information returned for the first fault repair task; and delete the first fault repair task from the uncompleted task set in response to the repair completion information.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above-mentioned method.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores program code, and the program code can be called by a processor to execute the above method.
[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0018] A log processing method, apparatus and device provided in an embodiment of the present application generates an exception record through an exception log, generates alarm information through an alarm condition corresponding to the exception record, and realizes the decoupling of exception judgment and alarm judgment, so that the detection and processing requirements can be flexibly configured by setting different alarm conditions, thereby meeting the diversified detection and processing requirements of massive servers; at the same time, since the target fault code in the target triplet and the fault code in the first triplet are mutually exclusive, the fault repair task associated with the target triplet and the fault repair task associated with the first triplet actually solves the fault caused by the same fault cause. In the present application, when the fault repair task associated with the target triplet and the fault repair task associated with the first triplet are not queried in the unfinished task set, a fault repair task is generated for the target component in the target device based on the target fault code, thereby avoiding repeated establishment of fault repair tasks for the same fault cause, resulting in invalid fault repair tasks affecting the processing efficiency of fault repair tasks, thereby improving the efficiency of fault detection to meet the detection and processing of massive equipment.
[0019] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 A schematic diagram of an implementation scenario involved in the present application is shown.
[0022] Figure 2 A flow chart of a log processing method provided in an embodiment of the present application is shown.
[0023] Figure 3 Another flow chart of the log processing method provided in the embodiment of the present application is shown.
[0024] Figure 4 Another flow chart of the log processing method provided in the embodiment of the present application is shown.
[0025] Figure 5 A schematic block diagram of a scenario of a log processing method provided in an embodiment of the present application is shown.
[0026] Figure 6 A flow chart of the data access / analysis service provided in an embodiment of the present application is shown.
[0027] Figure 7 A flow chart of an abnormal event detection service provided in an embodiment of the present application is shown.
[0028] Figure 8 A flow chart of the alarm judgment service provided in an embodiment of the present application is shown.
[0029] Fig. 9 A schematic diagram of the process of the fault repair service provided by the embodiment of the present application is shown.
[0030] Fig.10 A schematic block diagram of another scenario of the log processing method provided in an embodiment of the present application is shown.
[0031] Fig.11 A schematic block diagram of a log processing device provided in an embodiment of the present application is shown.
[0032] Fig.12 A schematic block diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0033] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0034] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0035] To facilitate the understanding of this solution, several terms are explained below:
[0036] In-band: refers to data transmission on the same channel or transmission medium. For example, when using a network protocol, data packets are transmitted on the same channel as control information.
[0037] Out-of-band: refers to the use of a channel or transmission medium different from the main data channel, such as using a dedicated management network for management or detection.
[0038] Sensor Data Record (SDR): It is mainly used for active and regular information collection. It collects the probe data of each component on the motherboard, and then judges the basic information and health status information of each component, and uses it in combination with the actual server operation needs.
[0039] Simple Network Management Protocol (SNMP): It is an application layer protocol in the five-layer TCP / IP protocol and is used for network management. SNMP is mainly used for the management of network devices and is currently the most widely used network management protocol. SNMP Trap is a part of SNMP. When a specific event occurs in the detected segment (such as performance problems, network device interface interruption, etc.), the agent will send an alarm event to the management station.
[0040] Baseboard Management Controller (BMC): A hardware chip or module installed on the computer motherboard that detects, manages, and controls computer hardware by communicating with hardware devices such as the motherboard, processor, and memory.
[0041] Event code: A code that describes a type of abnormal event. For example, if there is an event where the keyword "DID_SOFT_ERROR" appears in the dmesg log, the event code is defined as "dmesg_log_did_soft_error".
[0042] Fault code: A code that describes a type of fault. For example, if the event code "dmesg_log_did_soft_error" above exceeds 100 errors within 1 hour, it is considered a fault and the fault code is defined as "03-IN-01-1201".
[0043] In the existing hardware fault detection technology, due to the limited local disk storage capacity, too much historical data cannot be stored. With the advent of the cloud era, more and more individual users or enterprise users can own their own large-scale servers. In the scenario of massive devices (devices such as servers, edge computing devices, etc.), the logs generated by massive devices have a large amount of data, the data is complex, and different users may have diverse abnormality judgment and fault alarm requirements for the logs, which easily leads to repeated maintenance of hardware faults, thereby affecting the efficiency of maintenance of hardware faults.
[0044] In order to solve the above problems, the present application provides a log processing method, device and equipment, the log processing method comprising: determining an abnormal log in a set of equipment operation logs; generating an abnormal record according to the target slot information of the target component in the abnormal log, the target device information corresponding to the target device from which the abnormal log is derived, and the event code corresponding to the abnormal log, the target component refers to the component indicated by the abnormal log to have an abnormality; if the number of abnormal records including the same event code generated for the target component in the target device within a first time period reaches the alarm condition corresponding to the event code, based on the target fault code corresponding to the alarm condition, generate alarm information for the target component; if the fault repair task associated with the target triplet is not found in the unfinished task set, and the fault repair task associated with the first triplet is not found, based on the alarm information, generate a fault repair task for the target component, and add the generated fault repair task to the unfinished task set; the target triplet includes target device information, target slot information and target fault code, the first triplet is obtained by replacing the target fault code in the target triplet with a fault code mutually exclusive with the target fault code; send the fault repair task in the unfinished task set.
[0045] Through the log processing method provided in the present application, exception records are generated through exception logs, and alarm information is generated through alarm conditions corresponding to the exception records, thereby achieving decoupling of exception judgment and alarm judgment, so that the detection and processing requirements can be flexibly configured by setting different alarm judgments, thereby meeting the diversified detection and processing requirements of massive servers; at the same time, fault repair task queries are performed through the target triplet and the first triplet, thereby avoiding repeated establishment of fault repair tasks for the same fault cause, thereby improving the efficiency of log processing, and then improving the efficiency of fault detection to meet the detection and processing of massive servers.
[0046] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0047] like Figure 1 As shown, Figure 1 A schematic diagram of an implementation scenario involved in the present application is provided, including a device 10 to be managed (the device 10 may be a server, an edge computing device, an intelligent processing device, etc., which is not specifically limited here) and a log processing terminal 20, which may be a physical server or a cloud server, which is not specifically limited here. The device 10 communicates with the log processing terminal 20 via a wired or wireless network.
[0048] Multiple devices 10 can upload device operation logs (i.e., the device operation log set in this application) to the log processing end 20. After receiving the uploaded device operation logs, the log processing end 20 can execute the method of this application, thereby generating a fault repair task based on the device operation log and sending the fault repair task.
[0049] Exemplarily, the log processing terminal 20 generates an exception record according to the exception log in the equipment operation log, generates alarm information according to the alarm condition corresponding to the exception record, queries the fault repair task through the alarm information, and then generates the fault repair task according to the query result and the alarm information, and sends the fault repair task.
[0050] The log processing end 20 can be an independent physical server or a server cluster or distributed system composed of multiple physical servers. The log processing end 20 can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network) and big data and artificial intelligence platforms, and there is no restriction on this here.
[0051] like Figure 2 As shown, Figure 2 A flow chart of a log processing method provided in an embodiment of the present application is given, and the log processing method includes:
[0052] S110: Determine an abnormal log in the device operation log set.
[0053] Among them, the device operation log refers to the log generated during the operation of the device, which can be created and recorded by each process during the operation of the device; it can record the operation process and abnormal information of the device, and provide detailed information for quickly locating problems that occur during the operation of the device and program debugging problems during the development process. Among them, since the device includes multiple components (for example, the server includes the motherboard, the motherboard management controller (BMC), the processor, the storage, the memory, the hard disk, the flash memory, etc.), the operation of the device is achieved through the collaborative work between multiple components. Therefore, the device operation log can be understood as the log generated by each component in the device during the operation of the device.
[0054] The device operation log set may be a set of operation logs of a single device within a preset period of time, or a set of operation logs of multiple devices within a preset period of time.
[0055] An abnormal log refers to a log indicating that a component in a device has an abnormality during operation. In some embodiments, determining an abnormal log in a device operation log set may be to extract log contents of each device operation log in the device operation log set, and determine abnormal log contents by detecting the log contents, thereby determining the abnormal log.
[0056] Furthermore, the detection of log content can be achieved through semantic recognition or keyword matching. For example, semantic recognition of log content is performed through a pre-trained neural network model to obtain semantic features, and abnormal log content is determined based on the semantic features; or keyword extraction of log content can be performed through regular expressions of abnormal keywords, and abnormal log content is determined based on the extraction results. The specific implementation method is not limited here.
[0057] In some implementations, step S110 specifically includes: for each device operation log in the device operation log set, if the device operation log includes an abnormal keyword, determining that the device operation log is an abnormal log.
[0058] Different abnormal keywords can be set for different abnormal states; since different components may have different abnormal states, an abnormal keyword can be used to represent an abnormal state of a component. For example, for abnormal state A in the memory, the abnormal keyword in the log is a; for abnormal state B in the memory, the abnormal keyword in the log is b.
[0059] In some embodiments, a regular expression set may be set, the regular expression set including anomaly detection regular expressions corresponding to different anomaly keywords, the anomaly detection regular expressions including anomaly keywords, wherein the anomaly detection regular expressions are generated according to the format of anomaly logs in which corresponding anomaly keywords appear. Thus, based on the set anomaly detection regular expressions corresponding to different anomaly keywords, if a device operation log matches an anomaly detection regular expression (i.e., an anomaly keyword at a corresponding position in the anomaly detection regular expression exists in the device operation log), the device operation log may be determined to be an anomaly log.
[0060] In some embodiments, since different components may have multiple abnormal states, a first regular expression set corresponding to each component type can be pre-set. The first regular expression set corresponding to a component type includes multiple abnormality detection regular expressions for matching abnormal logs that may be generated by components of this component type, such as memory, hard disk, processor, motherboard, etc. On this basis, for each device operation log, the target component type to which the device operation log is derived can be determined based on the field representing the source component in the device operation log. After that, the device operation log is matched in the first regular expression set corresponding to the target component type. If the first regular expression set corresponding to the target component type has an abnormality detection regular expression that matches the device operation log, the device operation log is determined to be an abnormal log. Through this embodiment, the device operation log is matched in the first regular expression set corresponding to the component type of the component from which the device operation log is derived, without having to match it in all the first regular expression sets, thereby improving the efficiency of abnormal log identification.
[0061] In other embodiments, the abnormal log detection model can be used to identify abnormalities in the device operation log, and the abnormal classification result of the device operation log is output. The abnormal classification result includes a first classification result and a second classification result. The first classification result is used to indicate whether the device operation log is an abnormal log; the second classification result is used to indicate abnormal keywords in the device operation log. Wherein, the abnormal log detection model is constructed by one or more neural networks, such as recurrent neural networks (such as long short-term memory neural networks), fully connected neural networks, etc., which are not specifically limited here. Wherein, in the case where the first classification result indicates that the device operation log is not an abnormal log, the second classification result can be empty. Wherein, the abnormal keywords in the determined device operation log can be used to determine the event code corresponding to the device operation log, that is, in the case where the device operation log is an abnormal log, the event code corresponding to the abnormal keyword in the device operation log is used as the event code corresponding to the device operation log.
[0062] The anomaly log detection model can be trained through multiple sample device operation logs (in this application, the device operation logs used to train the anomaly log detection model are referred to as sample device operation logs) and label information corresponding to each sample device operation log. The label information corresponding to the sample device operation log includes first label information and second label information. The first label information is used to indicate whether the sample device operation log is an anomaly log, and the second label information is used to indicate anomaly keywords in the sample device operation log and the position of the anomaly keywords in the sample device operation log. It is worth mentioning that when the sample device operation log is not an anomaly log, the second label information of the sample device operation log can be empty.
[0063] During the training process, the sample device operation log can be segmented, and the segmented words in the device operation log can be embedded and position encoded to obtain the embedding vector corresponding to the device operation log. Then, the embedding vector corresponding to the device operation log is input into the abnormal log detection model, and the abnormal log detection model extracts features from the input embedding vector, and classifies according to the extracted features, and outputs the sample classification result, which includes a first sample classification result and a second sample classification result, wherein the first sample classification result is used to indicate the predicted classification result for indicating whether it is an abnormal log, and the second sample classification result is used to indicate the predicted abnormal keyword; then, the model loss can be calculated according to the label information corresponding to the sample device operation log and the corresponding sample classification result, and the weight parameter of the abnormal log detection model can be adjusted in reverse according to the model loss until the model training end condition is reached. Among them, the model training end condition can be that the number of iterations reaches a preset number threshold, or the loss function of the abnormal log detection model converges.
[0064] In some embodiments, a first loss can be determined based on the first label information corresponding to the sample device operation log and the first sample classification result corresponding to the sample device operation log; and a second loss can be determined based on the second label information corresponding to the sample device operation log and the second sample classification result corresponding to the sample device operation log; thereafter, the first loss and the second loss are weighted, and the weighted result is used as the model loss. Based on the trained abnormal log detection model, it is possible to accurately identify whether the device operation log is an abnormal log, and if the device operation log is an abnormal log, accurately output the abnormal keywords in the device operation log.
[0065] In other implementations, multiple message queues may be preset, and different types of device operation logs may be cached in different queues, and abnormal keywords corresponding to each type of device operation log may be set respectively; when performing abnormal keyword matching on the device operation log, for each message queue, only the corresponding abnormal keyword is used for matching, thereby improving matching efficiency.
[0066] Furthermore, the abnormal keyword matching can be performed on each message queue in turn, or on multiple message queues in parallel; at the same time, for each message queue, the abnormal keyword matching can be performed on each device operation log in the message queue in turn, or on multiple operation logs in the message queue in parallel, without specific limitations here.
[0067] By using different message queues to cache different types of device operation logs, when determining abnormal logs, only the corresponding abnormal keywords can be used for matching. At the same time, due to the existence of multiple message queues, matching operations can be performed concurrently through multiple work lines, thereby improving the efficiency of confirming abnormal logs and meeting the detection and processing needs of a large number of operation logs generated by massive servers.
[0068] It is worth mentioning that before step S110, it is necessary to first obtain a set of device operation logs.
[0069] In some implementations, the device operation log in the device operation log set may be actively reported by the device, or may be reported in response to a request from the log collection terminal, which is not specifically limited here.
[0070] Furthermore, in some embodiments, the device can transmit its own device operation log to the log collection end through both out-of-band and in-band communication methods; illustratively, for a pre-developed and deployed device agent, it collects data through operating system commands or manufacturer tools to obtain an operation log, which can be reported using an in-band communication method; for the device operation log recorded by the BMC on the device, an out-of-band communication method can be used for reporting.
[0071] In some implementations, considering that the data structures of operation logs generated by different devices or operation logs generated by different components of the same device may be different, in order to improve the processing efficiency of the device operation logs, the received device operation logs may be structured and then stored in the device operation log collection. By performing structured processing, device operation logs in different data formats can be converted into device operation logs with the same target data structure.
[0072] Among them, structured processing involves the transformation of data structure. For data of different structures, corresponding processing rules need to be configured to achieve structured transformation; for example, corresponding structured conversion rules can be configured for device operation logs of different data formats, and then the device operation log can be structured according to the structured conversion rules corresponding to the data format of the device operation log. Exemplarily, for the logs generated by SDR and SNMP Trap, they have a standard data format and can be structured, so the logs generated by SDR and SNMP Trap are structured, converted into JSON format (JavaScript Object Notation, a lightweight data interaction format) logs, and added to the device operation log collection.
[0073] In some embodiments, when the received device operation log does not have a standard format, such as a simple text format, the length and content of the text are uncertain, so the corresponding structured conversion rules cannot be configured, and the conditions for structuring are not met, and structured processing is not required.
[0074] S120: Generate an exception record according to the target slot information of the target component in the exception log, the target device information corresponding to the target device from which the exception log originates, and the event code corresponding to the exception log.
[0075] The target component refers to the component where the abnormality is indicated by the abnormality log, such as a hard disk, memory, processor, flash memory, etc. The target device from which the abnormality log comes refers to the device from which the abnormality log comes. For example, if the abnormality log is reported by device A, then device A is the source device of the abnormality log.
[0076] The target slot information refers to the slot information of the abnormal component in the target device from which the abnormal log originates. The slot information indicates the slot of the abnormal component in the target device from which the abnormal log originates. For example, when the target component corresponding to the abnormal log is a hard disk and the corresponding target device is A, the target slot information indicates the nth hard disk in device A where the abnormality occurs.
[0077] In the present application, the device operation log includes the slot information of the component from which the device operation log originates in the corresponding device. Therefore, the target slot information can be extracted from the exception log; further, the target slot information of the target component can be extracted from the exception log through the slot extraction regular expression.
[0078] The target device information may include the device identification of the target device from which the abnormal log originates, the network address of the device (eg, IP address), and may also include information such as the business to which the device IP belongs and the department to which the device belongs.
[0079] The device operation log may include the device identification of the target device from which the device operation log originates. Thus, based on the device identification of the device from which the exception log originates, device information such as the business and department to which the device belongs may be obtained from the Configuration Management Database (CMDB).
[0080] It should be noted that the event code is pre-configured; further, the event code is used to characterize the abnormal component and the abnormal state corresponding to the abnormal component; the event code and the abnormal keyword are one-to-one corresponding, and different abnormal states of the same abnormal component need to correspond to different abnormal keywords, that is, they need to correspond to different event codes. For example, the event code corresponding to the hard disk abnormality may include event codes A and B, where event code A represents a hard disk read error and event code B represents a hard disk startup error.
[0081] In some embodiments, before step S120, the method further includes: determining the event code corresponding to the abnormal keyword in the abnormal log based on the correspondence between the abnormal keyword and the event code, and using the event code corresponding to the abnormal keyword in the abnormal log as the event code corresponding to the abnormal log; and extracting the target slot information of the target component from the abnormal log based on a preset slot extraction regular expression.
[0082] In some implementations, in order to uniquely identify each exception record, the exception record may further include a record identifier, which may be a serial number assigned to the exception record; at the same time, the user may also query the generated exception record by the serial number.
[0083] In some embodiments, a timestamp may be extracted from the exception log, where the timestamp indicates the time when the exception indicated by the log exception occurs (ie, the time when the exception occurs). In the process of generating the exception record, the timestamp may also be added to the exception record.
[0084] In some embodiments, when log processing is performed for a large number of devices, due to the large number of devices and the fact that each device includes multiple components, the number of logs in the device operation log set is large, and the abnormal records that may be generated are also large. In order to reduce the storage pressure of the log processing end, the abnormal records can be stored in the abnormal record database. Afterwards, when detailed information in the corresponding abnormal record is needed, it can be queried from the abnormal record database. In some embodiments, in order to facilitate tracing, the abnormal record and the abnormal log corresponding to the abnormal record can be associated and stored in the abnormal record database.
[0085] Step S130: If the number of abnormal records including the same event code generated for the target component in the target device within the first time period reaches the alarm condition corresponding to the event code, generate alarm information for the target component based on the target fault code corresponding to the alarm condition.
[0086] Among them, the alarm condition corresponding to the event code can be limited to the number of times the event code appears within the first time period; the target fault code refers to the fault code corresponding to the alarm condition corresponding to the current event code. It can be understood that if the number of event codes generated for the target component in the target device within the first time period (the first time period is such as 1 minute, 5 minutes, 10 minutes, 30 minutes, etc.) is large, for example, reaching the number of times defined by an alarm condition, it indicates that the target component is more likely to fail. Therefore, in this case, alarm information for the target component can be generated based on the target fault code corresponding to the alarm condition, so as to timely issue an alarm reminder for the target component.
[0087] Furthermore, the alarm condition corresponding to the same event code may be one or more, and one alarm condition corresponding to the event code corresponds to a fault code; for example, for event code A, when the number of times it appears within 1 minute is greater than 5 and less than 10, the corresponding fault code is B; when the number of times it appears within 1 minute is greater than 10, the corresponding fault code is C.
[0088] In some implementations, when the same event code corresponds to multiple alarm conditions, the severity of the abnormality can be determined by the specific alarm conditions that are met, thereby facilitating determination of maintenance strategies such as expedited maintenance, deferred maintenance, etc. based on the severity of the abnormality.
[0089] It is worth mentioning that the equipment operation log often carries a timestamp. The timestamp can be used to determine the occurrence time of each exception log, and then the number of exception records with the same event code generated for the target component in the target device within the first time period can be calculated based on the occurrence time of each exception log.
[0090] Specifically, through the target device information in the exception record, all exception records in each target device can be counted; through the target slot information in the exception record, the specific location of the abnormal component can be determined; through the event code in the exception record, the specific abnormal state of the abnormal component in the target slot of the target device can be determined; obviously, in a target device, the number of exception records with the same event code and the same target slot information is the number of times the same abnormal state of the abnormal component (i.e. corresponding to the same event code) occurs.
[0091] S140: If no fault repair task associated with the target triplet is found in the unfinished task set, and no fault repair task associated with the first triplet is found, a fault repair task for the target component is generated based on the alarm information, and the generated fault repair task is added to the unfinished task set.
[0092] The target triplet includes target device information, target slot information and target fault code, and the first triplet is obtained by replacing the target fault code in the target triplet with a fault code mutually exclusive with the target fault code. For example, the target triplet can be device IP+fault code+slot information, and the first triplet can be device IP+mutually exclusive fault code+slot information.
[0093] Furthermore, a fault code that is mutually exclusive with a target fault code refers to other fault codes whose cause of failure is the same as that corresponding to the target fault code. In other words, if two fault codes correspond to the same cause of failure, the two fault codes are mutually exclusive, that is, one of the fault codes is mutually exclusive with the other fault code; exemplarily, the target fault code is A, which means that the hard disk cannot be started, and the cause of the failure is physical damage to the hard disk; for the fault code is B, which means that the hard disk cannot be read, and the cause of the failure is also physical damage to the hard disk. At this time, fault code B is also a fault code that is mutually exclusive with the target fault code A; obviously, for fault code A and fault code B, since they are both caused by physical damage to the hard disk, their corresponding maintenance strategies are the same. When one of the fault codes is repaired, the fault corresponding to the other fault code is also eliminated accordingly.
[0094] The unfinished task set refers to a set of multiple fault repair tasks that have been established and are waiting to be executed, or in other words, it refers to a set of fault repair tasks that have been established and whose task status is in an unfinished state. Among them, the fault repair task refers to a repair task generated for a component that has failed, that is, if it is determined that the number of times an event code appears in a component within a first time period reaches an alarm condition corresponding to the event code, it can be determined that the component has failed. In the present application, the fault repair tasks in the unfinished task set are stored in association with triples, wherein the triples include device information, slot information, and fault codes. Since device information can identify the device, slot information can locate the component in the device, and fault codes can characterize the fault that has occurred, the object targeted by the fault repair task (that is, which component in which device) and the fault targeted by the fault repair task can be accurately expressed based on the triple.
[0095] It can be understood that if the unfinished task set includes a fault repair task associated with the target triplet, it is equivalent to having established a fault repair task for the target component in the target device for the target fault code. At this time, if a new fault repair task is established, it will cause duplication of the fault repair task; similarly, since the fault code in the target triplet and the fault code in the first triplet are mutually exclusive, if the unfinished task set includes a fault repair task associated with the first triplet, then after the repair of the fault repair task associated with the first triplet is completed, the fault cause corresponding to the target fault code on the target component in the target device will also be eliminated accordingly. Therefore, in this case, there is no need to create different fault repair tasks for the same fault cause.
[0096] Therefore, in the present application, when no fault repair task associated with the target triplet is found in the unfinished task set, and no fault repair task associated with the first triplet is found, that is, when there is no fault repair task in the unfinished task set for the target component in the target device to solve the fault cause corresponding to the target fault code, a fault repair task for the target component is generated based on the alarm information to avoid repeatedly generating the same fault repair task for the same component, and also to avoid repeatedly generating fault repair tasks for the same component to solve the same fault cause.
[0097] Of course, if a fault repair task associated with the target triplet is queried in the unfinished task set, or a fault repair task associated with the first triplet is queried, it means that a fault repair task that can solve the cause of the fault already exists in the unfinished task set. Therefore, it is determined that there is no need to create a new fault repair task for the target component based on the alarm information, thereby avoiding the repeated generation of multiple fault repair tasks for the same fault cause, and the efficiency of fault repair task processing can be improved subsequently.
[0098] In some embodiments, when the generated fault repair task is added to the unfinished task collection, it can be added to the unfinished task collection in sequence according to the generation time of the fault repair task, thereby forming an unfinished task queue according to the generation time of the fault repair task, and when the fault repair task is sent subsequently, the fault repair task is sent based on the order of entering the unfinished task queue.
[0099] In other embodiments, taking into account the different urgency of different fault codes, the fault codes also correspond to fault levels, and the unfinished task set may also include task queues corresponding to multiple fault levels; the higher the fault level corresponding to the task queue where the fault repair task is located, the higher the sending priority of the fault repair task.
[0100] Furthermore, adding the generated fault repair task to the unfinished task set specifically includes: determining a target fault level corresponding to the target fault code based on a correspondence between the fault code and the fault level; and adding the generated fault repair task to a task queue corresponding to the target fault level.
[0101] Obviously, by setting task queues corresponding to multiple fault levels, fault repair tasks with high fault levels can be sent first and then processed first, which can improve the response efficiency for urgent fault repair tasks.
[0102] In some embodiments, considering the different detection and processing requirements of different devices, for example, for specific equipment in a specific department, such as a computer involved in account settlement in the finance office, considering its confidentiality and security, it needs to be repaired by a dedicated person, and there is no need to generate a fault repair task based on the log processing method of the present application. Therefore, before step S140, Figure 3 As shown, Figure 3 Another flow chart of the log processing method provided in the embodiment of the present application is given. The log processing method also includes the following S141-143:
[0103] S141: If it is determined according to a preset shielding policy that the target device corresponding to the alarm information is a non-shielded object, a target triplet is determined based on the alarm information.
[0104] Taking the device information as the device IP as an example, the preset blocking strategy can be to directly give the device IP to be blocked, or to give the business and / or department to which the device to be blocked belongs, and then determine the device IP corresponding to the business scenario or department by querying the CMDB. The specific method is not limited here.
[0105] Obviously, if the target device corresponding to the alarm information is a shielded object, no fault repair task will be generated based on the corresponding alarm information. Therefore, there is no need to determine its target triplet, and there is no need to query the fault repair task in the unfinished task set subsequently.
[0106] S142. Replace the target fault code in the target triplet with a fault code that is mutually exclusive with the target fault code to obtain a first triplet.
[0107] It should be noted that the mutually exclusive relationship between fault codes is pre-configured; further, one fault code may be mutually exclusive with multiple fault codes. Obviously, when one fault code is mutually exclusive with multiple fault codes, the corresponding first triples are also multiple.
[0108] S143: Perform task query in the unfinished task set based on the target triplet and the first triplet.
[0109] Specifically, if no fault repair task associated with the target triplet is found in the unfinished task set, and no fault repair task associated with the first triplet is found, a fault repair task for the target component is generated based on the alarm information; if a fault repair task associated with the target triplet is found in the unfinished task set, or a fault repair task associated with the first triplet is found, it is determined that there is no need to create a new fault repair task for the target component based on the alarm information.
[0110] Through the preset shielding strategy, the equipment that needs to establish fault repair tasks is screened, which not only realizes the differentiated management of different equipment, but also meets the differentiated needs in different scenarios. At the same time, the shielding strategy reduces the target triplets that need to be determined, avoids the establishment of useless fault repair tasks, and thus improves the efficiency of fault repair.
[0111] S150: Send the fault repair task in the unfinished task set.
[0112] In some implementations, a fault repair task in a set of unfinished tasks may be sent to a maintenance management terminal; wherein the maintenance management terminal is used to execute a specific maintenance process according to the fault repair task; such as establishing a fault ticket, dispatching a fault ticket, following up on the progress of the fault ticket, etc. according to the fault repair task, or the personnel of the maintenance management terminal repair the components in the corresponding equipment based on the fault repair task.
[0113] In some embodiments, after step S150, Figure 4 As shown, Figure 4 Another flow chart of the log processing method provided in the embodiment of the present application is given, and the log processing method further includes:
[0114] S160: Receive repair completion information returned for the first fault repair task.
[0115] The maintenance completion information may include the task order number corresponding to the first fault task, the corresponding equipment information, such as the equipment IP, and slot information, etc. After receiving the maintenance completion information, the corresponding fault maintenance task may be determined according to the specific information in the maintenance completion information.
[0116] S170. In response to the repair completion information, delete the first fault repair task from the uncompleted task set.
[0117] By deleting the first fault repair task from the unfinished task set, when the same fault occurs again, a corresponding fault repair task can be generated, thereby ensuring normal progress of the repair task.
[0118] Through the log processing method provided by the present application, an exception record is generated according to the exception log, and then the alarm information is generated according to the corresponding alarm condition, thereby realizing the decoupling of exception judgment and alarm judgment, so that the detection and processing requirements can be flexibly configured by setting different alarm conditions, thereby meeting the diversified detection and processing requirements of massive servers; at the same time, through the setting of mutually exclusive fault codes, the first triplet corresponding to the target triplet is determined, and then the fault repair task query is performed based on the target triplet and the first triplet, thereby avoiding the repeated establishment of repair tasks for the same fault cause, thereby resulting in invalid fault repair tasks, improving the efficiency of log processing, and then improving the efficiency of fault detection to meet the fault processing of massive servers.
[0119] For easier understanding, see Figure 5 , Figure 5 A schematic block diagram of a scenario of a log processing method provided in an embodiment of the present application is provided, including multiple device servers 30 ( Figure 5 ), a server hardware fault processing system 400 and a fault repair system 500 are exemplarily shown in the figure, wherein the log processing method provided by the application is executed by the server hardware fault processing system.
[0120] The device server 30 reports logs in both in-band and out-of-band ways. When facing a large number of servers, the amount of log reports can reach 1 billion logs per day.
[0121] The server hardware fault processing system 40 includes data access / analysis services, abnormal event detection services, alarm judgment services and fault ticket creation services.
[0122] The data access / analysis service needs to parse all device operation logs uploaded to the server and transmit the parsed device operation logs to the abnormal event detection service. Figure 6 As shown, Figure 6 A flow chart of data access / analysis service is given.
[0123] Taking massive servers as an example, the data access / analysis service uniformly receives in-band and out-of-band collected data, namely, the equipment operation log, and extracts the log content by parsing the private protocol adopted by the equipment operation log. The parsing rules need to be pre-configured by the fault operator through the web page and stored in the configuration system. When starting the analysis, it is necessary to load the timed synchronization to pull the latest parsing rules in the configuration system so that the parsing service can obtain the latest parsing rules.
[0124] For the operation logs with standard structure, such as the logs generated by sdr and snmptrap, they meet the structural conditions and are structured (structural processing can be performed by Figure 6The data structuring plug-in in the message queue is specifically executed) to convert it into a target structure, such as a JSON structure, and send the structured operation log to the message queue; the configuration rules for structured processing also need to be pre-configured by the fault operator through the web page and stored in the configuration system; for other types of logs, which are usually text-based and do not have structuring conditions, they are directly sent to the message queue; further, different types of logs need to be sent to corresponding message queues.
[0125] The abnormal event detection service is used to pull the device operation log in the message queue, and determine the abnormal log through analysis, and then determine the number of abnormal logs, and send a message to the next module; specifically, Figure 7 As shown, Figure 7 A processing diagram of the abnormal event detection service is given.
[0126] The abnormal event detection service includes a consumer thread pool, which contains multiple consumer threads, each of which is used to process a message queue; further, each consumer thread contains multiple work threads, and multiple work threads can process multiple device operation logs in parallel, thereby improving processing efficiency.
[0127] The work line matches each equipment operation log through the regular expression of the abnormal keyword, and determines the equipment operation log containing the abnormal keyword as the abnormal log. After determining it as the abnormal log, the slot extraction regular expression is used to extract the target slot information in the abnormal log, and generate the serial number corresponding to the abnormal log. The serial number, event code, matching event, fault occurrence time in the log, target slot information, machine information (IP, department, business, etc.), and log source file are stored to obtain the abnormal record. The abnormal record can also be stored in the abnormal record database so that the fault operation personnel can directly query the abnormal record through the query web page.
[0128] At the same time, in order to reduce the amount of data transmission, the abnormal event detection service can send the serial number corresponding to the abnormal record and the event code in the abnormal record to the alarm determination service.
[0129] It should be noted that in the anomaly detection service, event codes, the correspondence between event codes and anomaly keywords, faulty components corresponding to event codes (such as memory, hard disk), log types that need to be processed for each consumption line, regular expressions for extracting anomaly keywords and target slots, and other matching rules all need to be configured in advance by the fault operation personnel and stored in the configuration system; when the anomaly detection service is started, the latest configuration in the configuration system is pulled through loading and update synchronization.
[0130] The alarm judgment service is used to receive the serial number and event code of the abnormal log determined in the abnormal detection service, and calculate according to the pre-configured rules. If the alarm condition is met (such as meeting the preset number of thresholds), a corresponding alarm message is generated and sent to the next module, namely the fault ticket creation service; specifically, Figure 8 As shown, Figure 8 A flowchart of the alarm judgment service is given.
[0131] After the alarm judgment service receives the serial number and event code of the abnormal log determined in the abnormality detection service, it determines the corresponding alarm condition through the event code, and returns the abnormal record stored in the abnormality detection service according to the serial number and event code for query and statistics. When the number of times the same event code appears in the components in the same slot of the same target device reaches the alarm condition, a fault code corresponding to the alarm condition is generated, and an alarm message is generated based on the fault code. In some embodiments, when the alarm condition occurs once in one cycle, it is not necessary to query the abnormal record in the abnormal record database, and a fault code can be directly generated, thereby reducing the number of queries for abnormal records.
[0132] In order to reduce the amount of data transmission, only the alarm identifier (alarm identifier such as alarm ID) in the alarm information and the fault code corresponding to the alarm record can be sent to the next service, namely the fault order creation service, through the message queue; and the alarm judgment service generates an alarm flow table containing fault time, machine information, fault code, event code and alarm information, so that the fault operation personnel can directly query the alarm information, such as through a web page query.
[0133] It should be noted that in the alarm judgment service, the correspondence between event codes, alarm conditions and fault codes, and the correspondence between fault codes and alarm information need to be configured in advance by the fault operation personnel and stored in the configuration system; when the alarm judgment service is started, the latest configuration in the configuration system is pulled by loading update synchronization.
[0134] The fault repair service is used to receive the alarm information sent by the alarm judgment service, and filter out some alarms according to the preset strategy, and finally generate fault repair tasks for the faults that meet the conditions; specifically, Fig. 9 As shown, Fig. 9 A flowchart of fault repair service is given.
[0135] After receiving the alarm information, the fault repair service queries the previously stored abnormal records, alarm flow tables or configuration management databases (such as CMDB) for corresponding information, such as device information, slot information, business information or department information, etc. According to the alarm information, when the target device corresponding to the alarm information belongs to the shielded object, the processing of the alarm information is directly terminated; when the target device corresponding to the alarm information does not belong to the shielded object, a target triple is generated according to the alarm information, that is, the device IP+fault code+slot information, and the target triple is used to query whether there are associated unfinished tasks in the unfinished task set (the unfinished task set can be read from a storage database, such as Redis; the unfinished tasks in the unfinished task set can also be stored in the form of triples). When the unfinished task associated with the target triple is queried in the unfinished task set, it can also be called task convergence, and the processing of the alarm information is terminated; when the unfinished task associated with the target triple is not queried in the unfinished task set, a fault repair task is generated for the alarm information, that is, a fault order is created.
[0136] At the same time, it is also necessary to replace the fault code in the target triplet with a mutually exclusive fault code, that is, the first triplet in this application, and query again in the unfinished task set whether there are related unfinished tasks according to the first triplet. If there are related unfinished tasks, which can also be called task mutual exclusion, the processing of the alarm information is terminated; if there are no related unfinished tasks, a fault repair task is generated for the alarm information, that is, a fault order is created.
[0137] The generated fault repair task is sent to the subsequent fault repair system. At the same time, the fault time, machine information, fault code, order creation status and abnormal reason are stored accordingly to obtain an order creation flow table, so that the fault operation personnel can directly query the corresponding order creation status.
[0138] The fault repair system 50, which is also the repair management end in the present application, can generate a fault ticket according to the received fault repair task, and then perform a series of operations such as dispatching, following up and providing feedback on the fault ticket to complete the repair of the fault.
[0139] Taking the operation log of the order of 1 billion logs / day as an example, when the server hardware fault processing system 20 faces a data volume of 1 billion logs / day, since each log may need to be detected multiple times, after tens of billions of abnormal detections in the abnormal event detection service, abnormal events (that is, abnormal logs or abnormal records) of the order of 10 million abnormal events / day can be screened out, and then the screened abnormal events are input into the alarm judgment service, and the alarm information of the order of 100 million is matched / day according to different alarm conditions, and an alarm of the order of 1 million alarms / day can be obtained, and then through the convergence, mutual exclusion or shielding strategies in the fault order creation task, a fault maintenance task of the order of 1,000 fault orders / day is generated; it can be seen that when faced with massive operation logs, the log processing method provided by the present application can reduce the generated fault maintenance tasks by an exponential order, thereby being able to cope with the log processing and fault detection of millions of servers.
[0140] For other implementations, see Fig.10 , Fig.10 Another scenario diagram of an embodiment of the present application is given, in which the fault repair task output by the server hardware fault handling system 40 can be sent to the server operation and maintenance personnel, hardware fault handling personnel, spare parts management personnel and detection and processing strategy developers simultaneously while being sent to the fault repair system 50.
[0141] The fault repair system 50 is used to create a fault ticket according to the fault repair task, and manage various nodes of the fault repair process, such as fault ticket distribution, fault ticket task follow-up, fault ticket receipt reception, etc.
[0142] Server operation and maintenance personnel are used to determine whether to migrate services or authorize machine repair after receiving a fault repair task.
[0143] Hardware fault handling personnel are used to determine whether to replace server hardware (memory, hard disk, network cable, CPU, motherboard, etc.) after receiving a fault repair task.
[0144] The spare parts manager is responsible for managing and warehousing spare parts based on the faulty parts in the fault ticket and the judgment results of the hardware fault handler.
[0145] Detection and processing strategy developers: Analyze the accuracy of detection and processing strategies and optimize them.
[0146] In some implementations, according to the log processing method provided in the embodiments of the present application, the present application also provides a log processing device, such as Fig.11 As shown, Fig.11 A schematic block diagram of a log processing device provided in an embodiment of the present application is given, and the log processing device 200 includes:
[0147] The detection module 210 is used to determine abnormal logs in the device operation log set.
[0148] The exception record module 220 is used to generate an exception record according to the target slot information of the target component in the exception log, the target device information corresponding to the target device from which the exception log comes, and the event code corresponding to the exception log. The target component refers to the component indicated by the exception log to have an exception.
[0149] The alarm module 230 is used to generate alarm information for the target component based on the target fault code corresponding to the alarm condition if the number of abnormal records including the same event code generated for the target component in the target device within a first time period reaches the alarm condition corresponding to the event code.
[0150] The task generation module 240 is used to generate a fault repair task for a target component based on the alarm information if no fault repair task associated with the target triplet is found in the unfinished task set and no fault repair task associated with the first triplet is found, and add the generated fault repair task to the unfinished task set; the target triplet includes target device information, target slot information and target fault code, and the first triplet is obtained by replacing the target fault code in the target triplet with a fault code that is mutually exclusive with the target fault code.
[0151] The sending module 250 is used to send the fault repair task in the unfinished task set.
[0152] Specifically, in some implementations, the detection module 210 is used to determine, for each device operation log in the device operation log set, that the device operation log is an abnormal log if the device operation log includes an abnormal keyword.
[0153] In some embodiments, the detection module 210 is also used to identify the anomalies of each device operation log in the device operation log set through the abnormal log detection model, and obtain the abnormal classification result of the device operation log, the abnormal classification result includes a first classification result and a second classification result, the first classification result is used to indicate whether the device operation log is an abnormal log; the second classification result is used to indicate the abnormal keywords in the device operation log; wherein the abnormal log detection model is trained by: obtaining multiple sample device operation logs and label information corresponding to each sample device operation log; the label information first label information and second label information, the first label information is used to indicate whether the sample device operation log is If it is not an abnormal log, the second label information is used to indicate abnormal keywords in the sample device operation log; the abnormal log detection model performs abnormal identification based on the sample device operation log to determine the sample classification result corresponding to the sample device operation log, the sample classification result includes a first sample classification result and a second sample classification result, the first sample classification result is used to indicate whether the sample device operation log is an abnormal log; the second sample classification result is used to indicate abnormal keywords in the sample device operation log; the model loss is determined according to the sample classification result corresponding to the sample device operation log and the label information corresponding to the sample device operation log; the weight parameters of the abnormal log detection model are reversely adjusted according to the model loss until the model training end condition is met.
[0154] In some embodiments, the exception recording module 220 includes a matching unit and an extraction unit; the matching unit is used to determine the event code corresponding to the exception keyword in the exception log based on the correspondence between the exception keyword and the event code, and use the event code corresponding to the exception keyword in the exception log as the event code corresponding to the exception log; the extraction unit is used to extract the target slot information of the target component from the exception log based on a preset slot extraction regular expression.
[0155] In some embodiments, the task generation module 240 includes a shielding unit, a mutual exclusion unit and a query unit; the shielding unit is used to determine the target triplet based on the alarm information if it is determined according to a preset shielding strategy that the target device corresponding to the alarm information belongs to a non-shielded object; the mutual exclusion unit is used to replace the target fault code in the target triplet with a fault code that is mutually exclusive with the target fault code to obtain a first triplet; the query unit is used to perform task query in the unfinished task set based on the target triplet and the first triplet.
[0156] In some embodiments, the query unit is further used to determine that there is no need to create a new fault repair task for the target component based on the alarm information if a fault repair task associated with the target triplet is found in the unfinished task set, or a fault repair task associated with the first triplet is found.
[0157] In some embodiments, the unfinished task set includes task queues corresponding to multiple fault levels; the higher the fault level corresponding to the task queue where the fault repair task is located, the higher the sending priority of the fault repair task. The task generation module 240 is also used to determine the target fault level corresponding to the target fault code based on the correspondence between the fault code and the fault level; and add the generated fault repair task to the task queue corresponding to the target fault level.
[0158] In some implementations, the log processing device 200 further includes an updating module, which is configured to receive repair completion information returned for the first fault repair task; and delete the first fault repair task from the uncompleted task set in response to the repair completion information.
[0159] See also Fig.12 Based on the log processing method provided in the above embodiment, the embodiment of the present application also provides another electronic device including a processor that can execute the above method. The electronic device can be a terminal device, and the terminal device can be a smart phone, a tablet computer, a computer, or a portable computer. The present application also provides an electronic device, which includes: one or more processors; a memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs are configured for the log processing method provided in the present application.
[0160] Fig.12 The schematic diagram of the structure of the computer system of the electronic device suitable for implementing the embodiment of the present application is shown. The electronic device may be the terminal as described above, which is used to implement the log processing method provided by the present application. It should be noted that: Fig.12 The computer system 1400 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0161] like Fig.12 As shown, the computer system 1400 includes a central processing unit (CPU) 1401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1402 or the program loaded from the storage part 1408 to the random access memory (RAM) 1403, such as executing the method in the above embodiment. In RAM 1403, various programs and data required for system operation are also stored. CPU 1401, ROM 1402 and RAM 1403 are connected to each other through bus 1404. Input / output (I / O) interface 1405 is also connected to bus 1404.
[0162] The following components are connected to the I / O interface 1405: an input section 1406 including a keyboard, a mouse, a microphone, etc.; an output section 1407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1408 including a hard disk, etc.; and a communication section 1409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1409 performs communication processing via a network such as the Internet. A drive 1410 is also connected to the I / O interface 1405 as needed. A removable medium 1411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1410 as needed so that a computer program read therefrom is installed into the storage section 1408 as needed.
[0163] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 1409, and / or installed from a removable medium 1411. When the computer program is executed by a central processing unit (CPU) 1401, various functions defined in the system of the present application are executed.
[0164] It should be noted that the computer-readable medium shown in the embodiment of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (Erasable Programmable Read Only Memory, EPROM), flash memory, optical fiber, portable compact disk read-only memories (Compact Disc Read-Only Memory, CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as a part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0165] The log processing method, device and equipment provided by the present application, the method determines the abnormal log in the equipment operation log set; generates an abnormal record according to the abnormal log, and then determines the alarm information according to the alarm condition; if there is no unfinished fault maintenance task associated with the alarm information, generates a fault maintenance task for the target component based on the alarm information, and gives priority to sending the fault maintenance task with a high fault level, so that the maintenance management end can give priority to creating a fault ticket for the fault maintenance task with a high fault level, and then start the fault maintenance process; through the log processing method provided by the present application, the abnormal record and the alarm information are judged respectively, and the decoupling of the abnormal judgment and the alarm judgment is realized, so that the detection and processing requirements can be flexibly configured by setting different alarm conditions, thereby meeting the diversified detection and processing requirements of massive servers; at the same time, through the setting of mutually exclusive fault codes, the first triple corresponding to the target triple is determined, and then the fault maintenance task query is performed based on the target triple and the first triple, so as to avoid the maintenance task for the same fault cause being repeatedly established, thereby causing invalid fault maintenance tasks, improving the efficiency of log processing, and then improving the efficiency of fault detection to meet the detection and processing of massive servers.
[0166] The above are only preferred embodiments of the present application, and are not intended to limit the present application in any form. Although the present application has been disclosed as above with preferred embodiments, it is not intended to limit the present application. Any technical personnel in the field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present application. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.
Claims
1. A log processing method, characterized in that: include: Determine abnormal logs in the device operation log set; Generate an exception record according to the target slot information of the target component in the exception log, the target device information corresponding to the target device from which the exception log is sourced, and the event code corresponding to the exception log, wherein the target component refers to the component indicated by the exception log as having an exception; If the number of abnormal records including the same event code generated for the target component in the target device within the first time period reaches the alarm condition corresponding to the event code, based on the target fault code corresponding to the alarm condition, generate alarm information for the target component; If no fault repair task associated with the target triplet is found in the unfinished task set, and no fault repair task associated with the first triplet is found, a fault repair task for the target component is generated based on the alarm information, and the generated fault repair task is added to the unfinished task set; the target triplet includes the target device information, the target slot information and the target fault code, and the first triplet is obtained by replacing the target fault code in the target triplet with a fault code mutually exclusive with the target fault code; Sending the fault repair task in the unfinished task set.
2. The method according to claim 1, characterized in that If no fault repair task associated with the target triplet is found in the unfinished task set and no fault repair task associated with the first triplet is found in the unfinished task set, generating a fault repair task for the target component based on the alarm information and adding the generated fault repair task to the unfinished task set, the method further includes: If it is determined according to a preset shielding strategy that the target device corresponding to the alarm information belongs to a non-shielded object, the target triplet is determined based on the alarm information; Replacing a target fault code in the target triplet with a fault code that is mutually exclusive with the target fault code to obtain the first triplet; A task query is performed in the unfinished task set based on the target triplet and the first triplet.
3. The method according to claim 1, characterized in that The determining of abnormal logs in the device operation log set includes: For each device operation log in the device operation log set, an abnormality identification is performed on the device operation log by using an abnormal log detection model to obtain an abnormality classification result of the device operation log, wherein the abnormality classification result includes a first classification result and a second classification result, wherein the first classification result is used to indicate whether the device operation log is an abnormal log; and the second classification result is used to indicate an abnormal keyword in the device operation log; The abnormal log detection model is trained in the following way: Acquire multiple sample device operation logs and label information corresponding to each sample device operation log; the label information includes first label information and second label information, the first label information is used to indicate whether the sample device operation log is an abnormal log, and the second label information is used to indicate abnormal keywords in the sample device operation log; The abnormal log detection model performs abnormal identification according to the sample device operation log, and determines a sample classification result corresponding to the sample device operation log, wherein the sample classification result includes a first sample classification result and a second sample classification result, wherein the first sample classification result is used to indicate whether the sample device operation log is an abnormal log; and the second sample classification result is used to indicate an abnormal keyword in the sample device operation log; Determine the model loss according to the sample classification result corresponding to the sample device operation log and the label information corresponding to the sample device operation log; The weight parameters of the abnormal log detection model are reversely adjusted according to the model loss until the model training end condition is reached.
4. The method according to claim 1, characterized in that The determining of abnormal logs in the device operation log set includes: For each device operation log in the device operation log set, if the device operation log includes an abnormal keyword, the device operation log is determined to be an abnormal log.
5. The method according to claim 4, characterized in that Before generating the exception record according to the target slot information of the target component in the exception log, the target device information corresponding to the target device from which the exception log is derived, and the event code corresponding to the exception log, the method further includes: Based on the correspondence between the abnormal keyword and the event code, determine the event code corresponding to the abnormal keyword in the abnormal log, and use the event code corresponding to the abnormal keyword in the abnormal log as the event code corresponding to the abnormal log; Based on a preset slot extraction regular expression, the target slot information of the target component is extracted from the abnormal log.
6. The method according to claim 1, characterized in that The unfinished task set includes task queues corresponding to multiple fault levels; the higher the fault level corresponding to the task queue where the fault repair task is located, the higher the sending priority of the fault repair task; The step of adding the generated fault repair task to the unfinished task set includes: Determining a target fault level corresponding to the target fault code based on a correspondence between the fault code and the fault level; The generated fault repair task is added to the task queue corresponding to the target fault level.
7. The method according to claim 1, characterized in that After sending the fault repair task in the unfinished task set, the method further includes: Receiving maintenance completion information returned for the first fault maintenance task; In response to the maintenance completion information, the first fault maintenance task is deleted from the uncompleted task set.
8. A log processing device, characterized in that: include: A detection module, used to determine abnormal logs in a set of device operation logs; An exception record module, used to generate an exception record according to the target slot information of the target component in the exception log, the target device information corresponding to the target device from which the exception log is derived, and the event code corresponding to the exception log, wherein the target component refers to the component indicated by the exception log as having an exception; an alarm module, configured to generate alarm information for a target component in the target device based on a target fault code corresponding to the alarm condition if the number of abnormal records including the same event code generated for the target component in the target device within a first time period reaches an alarm condition corresponding to the event code; A task generation module, configured to generate a fault repair task for the target component based on the alarm information if no fault repair task associated with the target triplet is found in the unfinished task set and no fault repair task associated with the first triplet is found, and add the generated fault repair task to the unfinished task set; the target triplet includes the target device information, the target slot information and the target fault code, and the first triplet is obtained by replacing the target fault code in the target triplet with a fault code mutually exclusive with the target fault code; The sending module is used to send the fault repair tasks in the unfinished task set.
9. An electronic device, characterized in that: include: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: include: The computer-readable storage medium stores program codes, and the program codes can be called by a processor to execute the method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Log file analysis method and device based on eMMC, equipment and medium
CN120560917A