Log processing method and device, electronic equipment and storage medium
By calculating the similarity of log files and using hashing algorithm to identify the overlap of log contents, the problem of repeated log collection in server memory RMT test is solved, and storage resources are optimized and cost reduction is achieved.
Patent Information
- Application Number
- CN202510725006.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, due to multiple hard restarts in server memory RMT tests, the same logs are repeatedly collected, resulting in a high storage occupancy rate.
By calculating the similarity value between log files, the new log file is retained only when the similarity is less than the preset threshold, natural language processing and hashing algorithms are used to identify the overlap of log contents, avoiding redundant storage.
It reduces the consumption of storage resources, improves the effective utilization of log data, and reduces the storage management cost of data centers.
Smart Images

Figure CN120234178A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to a method and apparatus for processing logs, an electronic device, and a storage medium. Background Art
[0002] Server memory reliability maintenance and testing (RMT) is a core part to ensure the stable operation of a data center.
[0003] Currently, the method used for server memory RMT testing is based on automated log collection of SOL. After each log collection, a hard restart (Power Cycle) is performed to reset the server state. However, multiple restarts will cause the same logs to be collected repeatedly, resulting in a high storage occupancy rate. Summary of the Invention
[0004] This application provides a method and apparatus for processing logs, an electronic device, and a storage medium, so as to at least solve the problem in the related art that the same logs are collected repeatedly, resulting in a high storage occupancy rate.
[0005] This application provides a method for processing logs, including: Obtaining a first log file; Calculating a similarity value between the first log file and a second log file; wherein, the second log file and the first log file are log files of different repetition periods of the same server; When it is determined that the similarity value is less than a preset threshold, determining to retain the first log file.
[0006] This application further provides an apparatus for processing logs, including: An obtaining unit, configured to obtain a first log file; A calculating unit, configured to calculate a similarity value between the first log file and a second log file; wherein, the second log file and the first log file are log files of different repetition periods of the same server; A retaining unit, configured to determine to retain the first log file when it is determined that the similarity value is less than a preset threshold.
[0007] This application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above methods for processing logs when executing the computer program.
[0008] This application further provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program, when executed by a processor, implements the steps of any of the above methods for processing logs.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned log processing methods when executed by a processor.
[0010] Through the present application, due to the adoption of the log file content similarity analysis mechanism, the degree of information overlap between newly collected logs and historical logs can be identified, and only the new logs are retained when they are different, avoiding redundant storage of invalid logs with highly repetitive content. Therefore, the technical problems in the prior art caused by repeated hard restarts resulting in repeated collection of the same logs and excessive occupation of storage resources can be solved, achieving the technical effects of reducing storage capacity consumption, improving the effective utilization rate of log data, and reducing the storage management cost of the data center. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic flowchart of a log processing method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a log collection method provided by an embodiment of the present application; Figure 3 It is an architecture diagram of a log collection system provided by an embodiment of the present application; Figure 4 It is a schematic diagram of another log collection method provided by an embodiment of the present application; Figure 5 It is a schematic diagram of another log collection method provided by an embodiment of the present application; Figure 6 It is a schematic diagram of another log collection method provided by an embodiment of the present application; Figure 7 It is a schematic diagram of another log collection method provided by an embodiment of the present application; Figure 8 It is a schematic diagram of another log collection method provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0015] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0016] An embodiment of the present application provides a method for processing logs. In combination with the execution flow of the method for processing logs, the method will be described in detail.
[0017] Figure 1 It is a schematic flowchart of a method for processing logs provided by an embodiment of the present application.
[0018] As Figure 1 shown, the method includes the following steps: Step 101, obtain a first log file.
[0019] During the process of server memory reliability maintenance and testing (RMT), first, a first log file needs to be collected. In some embodiments, the first log file can be obtained based on the log storage directory specified by the server operating system.
[0020] In some embodiments, the server system interface (such as a file reading API or a command line tool) can be called to automatically scan and identify the newly generated log files within the current test cycle, that is, the first log file. In some possible implementation manners, according to the identification features such as the file generation timestamp and the cycle identifier in the file name, the latest log file corresponding to the current test cycle is filtered out for subsequent steps to process. The entire acquisition process is implemented through an automated program to ensure the real-time, accuracy, and integrity of the log file, providing a reliable data basis for subsequent log similarity calculation.
[0021] Step 102, calculate the similarity value between the first log file and the second log file; wherein, the second log file and the first log file are log files of the same server in different repetition cycles.
[0022] In some embodiments, according to the server unique identifier recorded in the first log file (such as server ID, MAC address) and the test cycle information (such as timestamp, cycle number), the corresponding second log file can be determined from the historical log repository, and the second log file is the log data generated by the same server in the previous test cycle.
[0023] After obtaining the first log file and the second log file, natural language processing (NLP) technology can be used to perform word segmentation on the log text, extract key information fields (such as error code, memory address, exception event description), and generate a feature vector containing the core content of the log. Through a preset text similarity algorithm (such as the cosine similarity algorithm based on term frequency-inverse document frequency (TF-IDF)), the feature vectors of the two log files are calculated to obtain a similarity value reflecting the degree of content repetition between the two. This value usually ranges from [0,1], and the larger the value, the higher the similarity of the log content.
[0024] Step 103, when it is determined that the similarity value is less than the preset threshold, determine to retain the first log file.
[0025] Compare the similarity value calculated in step 102 with the preset threshold of the system. In some embodiments, this preset threshold can be set in advance according to the business requirements of the test and the characteristics of the historical log data, and is used to define whether there is an effective difference in the content of the log file (for example, when the threshold is set to 0.7, it is considered that there is a significant difference when the similarity is lower than 70%). When it is determined that the similarity value is less than the preset threshold, it indicates that the first log file contains effective information that is significantly different from the second log file. At this time, the log retention mechanism is triggered. In a possible implementation, the first log file is moved from the temporary storage area to the long-term repository, and metadata tags (such as log generation time, server unique identifier, test cycle number, etc.) are added to it during the storage process for subsequent retrieval and analysis.
[0026] If the similarity value is greater than or equal to the preset threshold, it is determined that the content of the first log file is highly repetitive with the historical log, and the file is directly discarded to avoid redundant storage. The entire decision-making and execution process is implemented through an automated program to ensure the differential management of log files, while ensuring the validity of test data and controlling the consumption of storage resources.
[0027] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a log collection method provided by an embodiment of the present application. As Figure 2 shown, the first log file contains first time anchors injected at preset time intervals; the second log file contains second time anchors injected at the preset time intervals; before calculating the similarity value between the first log file and the second log file, the method further includes: Dividing the first log file based on the first time anchors to obtain at least one first log segment; Dividing the second log file based on the second time anchors to obtain at least one second log segment.
[0028] In some embodiments, both the first log file and the second log file contain time anchors injected at a unified preset time interval, such as every 1 minute or 2 minutes, etc. The time anchor is a specific marker used to identify the time interval of data records in the log file.
[0029] Before calculating the similarity value between the two, based on the positions of all the first time anchors in the first log file, in the order of the appearance of the time anchors and the preset time interval, the first log file is divided into at least one continuous first log segment. Each first log segment corresponds to an independent time interval and contains the log data generated within that interval. Using the same processing logic, the second log file is operated on: based on the positions of the second time anchors and the same preset time interval, the second log file is divided into at least one second log segment. Each second log segment has a corresponding relationship with the first log segment in terms of the time interval (such as the log data between two adjacent time anchors).
[0030] In some embodiments, during the log segment division process, the time anchors can be located through a text parsing algorithm, and the log content can be divided according to the time stamps or marking symbols of the anchors to ensure that the start and end positions of each log segment accurately match the preset time interval, forming structured log segment data, which provides a standardized processing unit for subsequent similarity calculation of log segments according to time intervals.
[0031] In some embodiments, calculating the similarity value between the first log file and the second log file includes: Determining the first log segment and the second log segment in the same time period in the first log file and the second log file; Calculating the similarity value between the first log segment and the second log segment in the same time period.
[0032] Continuing the narrative of the above application embodiments, each test cycle is divided into multiple preset time periods of equal length (for example, each cycle includes fixed time intervals such as 0 - 5 minutes, 6 - 10 minutes, etc.), and time anchors are injected at the starting positions of the corresponding time periods of each cycle according to a unified rule (for example, time points such as the 1st minute and the 6th minute of each cycle). By parsing the time anchors and cycle identifiers in the log files, the system identifies the first log segment in the first log file that is within the 1 - 5 minute time period of the current cycle, and the second log segment in the second log file that is within the 1 - 5 minute time period at the relative position in the previous cycle, ensuring that both occupy the same time interval position in their respective test cycles.
[0033] In some embodiments, for the log segments in the same selected time period, content standardization processing can be performed first. For example: removing dynamic metadata irrelevant to the test cycle from the log (such as real - time system time, temporary process ID), extracting core data fields such as memory error codes, address information, and exception event descriptions, and unifying the text encoding format and data representation method (such as converting 16 - hexadecimal addresses to a unified case form). Subsequently, the similarity value between the two log segments is calculated. The similarity value quantitatively reflects the consistency degree of the server memory state changes within the same relative time period in different test cycles, providing a data basis for identifying periodic repeated faults or progressive hardware degradation.
[0034] In some embodiments, calculating the similarity value between the first log segment and the second log segment in the same time period includes: Calculating the hash value of the first log segment and the hash value of the second log segment respectively.
[0035] Please continue to refer to Figure 2 , before calculating the hash value of the log segment, pre - processing of the log content can be performed. For example, removing variable metadata (such as real - time timestamps, dynamic IP addresses, etc.) and redundant format specifiers that do not affect the core information in the log segment, and retaining key data fields such as memory error codes, exception event descriptions, and address pointers to form a standardized log content text. After the pre - processing is completed, the hash algorithm module generates hash values for the standardized texts of the two log segments respectively. The hash algorithm can use one - way encryption hash functions such as MD5 (Message - Digest Algorithm 5), SHA - 1 (Secure Hash Algorithm 1), etc., to map the log segment text content into a hash string of a fixed length (such as MD5 generates a 128 - bit hash value). After generating the hash values, the system directly compares the consistency of the two hash strings: if the hash values are exactly the same, it is determined that the content of the first log segment is exactly the same as that of the second log segment, and the content of the first log segment is not retained; if the hash values are different, it is determined that the content of the first log segment is not exactly the same as that of the second log segment, and the log content corresponding to the first log segment is saved.
[0036] It should be noted that this description is only an exemplary description. In actual application, it is necessary to compare the first log segment of the first log file with the second log segment of the second log file in sequence. For the detailed comparison process, please refer to the description of the above-mentioned application embodiment, and the embodiment of this application will not be repeated here one by one.
[0037] The uniqueness of the hash value can be used to quickly determine the degree of duplication of the log segment content, avoiding complex text semantic analysis, improving the efficiency and accuracy of similarity calculation, and providing a lightweight technical means for the rapid screening of log files.
[0038] In some embodiments, when determining that the similarity value is less than a preset threshold, determining to retain the first log file further includes: When it is determined that the hash value of the first log segment is different from the hash value of the second log segment, it is determined to retain the first log segment.
[0039] After obtaining the hash values of the first log segment and the second log segment of the same time period, perform a hash value comparison operation. If the comparison result is that the hash value of the first log segment is different from the hash value of the second log segment, it is determined that there is a difference in the content of the two, and the log segment retention mechanism is triggered: the first log segment is transferred from the temporary storage area to the long-term storage repository, and a metadata tag is attached to it (including the time interval corresponding to the log segment, the server unique identifier, the test cycle to which it belongs, and other information) for subsequent retrieval and analysis by time period. If the hash values are the same, the first log segment is directly discarded to avoid redundant storage. The entire process is implemented through an automated program, which quickly identifies the differences in log segment content based on the uniqueness of the hash value, ensures that only log segment data with valid information is stored, and realizes refined management of log files and optimized configuration of storage resources.
[0040] In some embodiments, before obtaining the first log file, the method further includes: Creating a preset number of LAN serial session threads; wherein the preset number is the number of servers; A connection relationship between each of the LAN serial conversation threads and the server is established respectively, so as to send an injection instruction of a time anchor point to the server through the LAN serial conversation thread.
[0041] According to the number of servers to be tested, create a corresponding preset number of Local Area Network (LAN) serial session threads, with each thread corresponding to a unique server identifier. After creating the threads, establish the connection relationship between each LAN serial session thread and the target server through a network communication protocol (such as TCP / IP). During the connection process, server authentication (such as username / password verification or key exchange) needs to be completed to ensure the security and reliability of the communication link. After the connection is established, each thread injects time anchor instructions into the connected server according to the system preset time anchor injection strategy (such as triggering at fixed time intervals or test cycles). This instruction is used to indicate that when the server generates a log file, insert a time anchor in a specific format (such as a marker field containing a timestamp) into the log content at preset time intervals (such as every minute or every hour). Through the above process, parallel control of multiple servers is achieved, ensuring the consistency and standardization of time anchors in the log files of each server, and providing a standardized data basis for subsequent log segmentation and similarity calculation based on time anchors.
[0042] In some embodiments, the obtaining of the first log file further includes: After determining that the file capacity of the first log file reaches the preset log threshold, save a status snapshot and restart the server based on a lightweight kernel.
[0043] In some embodiments, by real-time monitoring the file capacity of the first log file, when it is detected that the file capacity reaches the preset log threshold (such as a pre-configured 50MB or 60MB), first trigger the status snapshot saving mechanism: use the process status saving tool provided by the server operating system (such as the coredump mechanism in the Linux system or the checkpoint technology in the virtualization environment) to capture snapshots of the key information such as the current memory status, running processes, and network connections of the server, and store the snapshot data in a specified location to ensure data integrity. After completing the status snapshot saving, the system calls the lightweight kernel restart interface to perform the server restart operation based on the lightweight kernel module (such as a streamlined kernel image containing minimized drivers and services). This process skips the full hardware initialization process in the traditional hard restart and only reloads the operating system kernel and necessary services to shorten the restart time and maintain the hot state of the hardware device. Through the above process, while ensuring the integrity of log data, the server status can be quickly reset, providing a stable operating environment for the next cycle of log collection.
[0044] In some embodiments, after determining to retain the first log file, the method further includes; Associate the log files corresponding to different repetition cycles based on the time sequence to generate a fault propagation link, and store the fault propagation link in a preset database; When it is determined that the fault type corresponding to the fault propagation link is a marked event, a fault report is generated and the fault report is pushed based on a preset push method; wherein, the fault report includes a fault path and a fault type.
[0045] After determining to retain the first log file, the system first establishes an association relationship between log files with different repetition periods in chronological order based on the timestamps carried by the time anchors in the log files or test cycle identifiers, and analyzes the fault event sequences in the logs of each period through an event association algorithm (such as a path tracing algorithm based on causal relationships) to identify the sequence and dependency of event occurrences, and generates a fault propagation link reflecting the fault development path. This link is stored in the form of a directed graph, with nodes being the key fault events in the logs of each period (such as memory error codes, abnormal address information), and edges being the causal association or sequential progression relationships between events, and the link data is stored in a preset fault database (such as a relational database or a graph database).
[0046] When the system detects through a preset fault type matching rule (such as regular expression matching or a machine learning classification model) that the fault type corresponding to the fault propagation link belongs to a pre-marked key event (such as marked events like "memory controller failure", "persistent bit flip error", etc.), it automatically triggers a fault report generation mechanism: extracts the key node information, event occurrence time sequence, and associated log file paths in the link, and generates a structured report containing the fault path (such as the propagation order of event A → event B → event C) and the fault type. The generated fault report is sent to the operation and maintenance personnel or the monitoring system through a preset push method (such as email notification, system message queue, or API interface call) for timely fault location and repair.
[0047] Please refer to Figure 3 , Figure 3 which is the architecture diagram of the log acquisition system provided by the embodiment of the present application. As Figure 3 shown, it includes: Central control node: Deploy SOL session management, task scheduling, and log analysis modules, and the hardware configuration needs to support high-concurrency processing (such as an Intel Xeon 8-core CPU and 64GB of memory).
[0048] Server cluster under test: Equipped with a BMC and an RMT memory module supporting the IPMI 2.0 standard and a local status cache.
[0049] Distributed storage system: Adopts MinIO object storage to save incremental logs and status snapshots.
[0050] Software architecture: Control layer: SOL session scheduler (supporting multi-threaded concurrency), restart control engine (integrating IPMI tools).
[0051] Data layer: Incremental log database, status snapshot storage.
[0052] Analysis layer: Log correlation analysis algorithm, visualization front-end.
[0053] Among them, the central control node includes: SOL session dynamic management module: Automatically controls the establishment, maintenance, and recovery of multi-node SOL connections, and supports multi-node parallel sessions.
[0054] Please refer to Figure 4 , Figure 4 , which is the flowchart of another log collection method provided by the embodiment of the present application. As Figure 4 shown, Session pool pre-allocation: The central control node pre-starts multiple SOL session threads, manages multi-node connections using the non-blocking I / O model, and each thread is bound to an independent port (such as IPMI port 623).
[0055] Disconnection self-healing mechanism: Real-time monitors the session status, automatically reconnects after detecting timeout (retry interval is 3 seconds, up to 5 times). During the disconnection period, the logs are temporarily stored in the BMC cache and are retransmitted to the central node after recovery.
[0056] Incremental log collection engine: Extracts valid logs based on the timestamp anchor point and hash deduplication algorithm.
[0057] Please continue to refer to Figure 2 , Timestamp anchor injection: Before each restart, send instructions through SOL, such as: [RMT_ANCHOR:SEQ=001,TIME=20231001120000] to the server serial port. The anchor point is used as the log segmentation identifier, and the central node splits the log file according to the anchor point.
[0058] Calculates the SHA-1 hash value of the logs between adjacent anchor points based on the hash deduplication algorithm. If it is the same as the previous cycle, it is discarded.
[0059] Intelligent restart control unit: Optimizes the server restart process and integrates the function of saving and restoring the memory status snapshot.
[0060] Please refer to Figure 5 , Figure 5 , which is the flowchart of another log collection and sending method provided by the embodiment of the present application. As Figure 5 shown, Memory status snapshot: Before restart, saves the current test status to the NVDIMM (non-volatile memory) through the IPMI command ipmitool chassis power cycle of the BMC, including the tested memory address range, error count, etc.
[0061] Fast restart link: The BMC loads a customized lightweight kernel (such as a Linux micro-image based on BusyBox, with a size ≤ 10MB), skipping the BIOS self-check and operating system initialization. After the kernel starts, it automatically restores the test state from the NVDIMM and continues to execute the unfinished RMT tasks.
[0062] Please refer to Figure 6 , Figure 6 which is the flowchart of another method for collecting and sending logs provided by the embodiments of this application. As Figure 6 shown, Error event chain modeling: Concatenate the multi-round test logs by anchor points to construct an error propagation path (such as "first error → recurrence after restart → final fix"). Use a database (such as Neo4j) to store event association relationships to support fast query.
[0063] Visualization dashboard: Mark key events (such as sudden temperature increase, ECC error correction exceeding the threshold), and associate them with specific memory modules (DIMM slots). Generate PDF / HTML test reports, and recommend replacing faulty hardware or adjusting the test strategy.
[0064] Log correlation analysis platform: Automatically associate the logs of multiple restarts according to the test stages to generate a complete error link.
[0065] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0066] The embodiments of this application also provide a log processing device. Figure 7 which is the structural schematic diagram of a log processing device provided by the embodiments of this application. As Figure 7 shown, it includes: An acquisition unit 21, configured to acquire a first log file; A calculation unit 22, configured to calculate the similarity value between the first log file and a second log file; wherein, the second log file and the first log file are log files of different repetition cycles of the same server; A retention unit 23, configured to determine to retain the first log file when it is determined that the similarity value is less than a preset threshold.
[0067] Furthermore, in a possible implementation manner of the embodiments of the present disclosure, as Figure 8 shown, the first log file includes first time anchors injected at preset time intervals; the second log file includes second time anchors injected at the preset time intervals; the device further includes: The first splitting unit 24 is configured to split the first log file based on the first time anchor before the computing unit 22 calculates the similarity value between the first log file and the second log file, so as to obtain at least one first log segment; The second splitting unit 25 is configured to split the second log file based on the second time anchor to obtain at least one second log segment.
[0068] Further, in a possible implementation manner of the embodiment of the present disclosure, the computing unit 22 is further configured to: Determine a first log segment and a second log segment in the same time period in the first log file and the second log file; Calculate the similarity value between the first log segment and the second log segment in the same time period.
[0069] Further, in a possible implementation manner of the embodiment of the present disclosure, the computing unit 22 is further configured to: Calculate the hash value of the first log segment and the hash value of the second log segment respectively.
[0070] Further, in a possible implementation manner of the embodiment of the present disclosure, the computing unit 22 is further configured to: When it is determined that the hash value of the first log segment is different from the hash value of the second log segment, determine to retain the first log segment.
[0071] Further, in a possible implementation manner of the embodiment of the present disclosure, as Figure 8 shown, the apparatus further includes: The creating unit 26 is configured to create a preset number of local area network serial session threads before the obtaining unit 21 obtains the first log file; wherein, the preset number is the number of servers; The establishing unit 27 is configured to establish a connection relationship between each of the local area network serial session threads and the server respectively, so as to send an injection instruction of the time anchor to the server through the local area network serial session thread.
[0072] Further, in a possible implementation manner of the embodiment of the present disclosure, as Figure 8 shown, the obtaining unit 21 is further configured to: After determining that the file capacity of the first log file reaches a preset log threshold, save a status snapshot and restart the server based on a lightweight kernel.
[0073] Further, in a possible implementation manner of the embodiment of the present disclosure, as Figure 8 shown, the apparatus further includes; A generating unit 28, configured to, after the retention unit 23 determines to retain the first log file, associate log files corresponding to different repetition periods in chronological order, generate a fault propagation link, and store the fault propagation link in a preset database; A generating unit, configured to generate a fault report when determining that the fault type corresponding to the fault propagation link is a marked event, and push the fault report based on a preset push method; wherein, the fault report includes a fault path and a fault type.
[0074] For the description of the features in the embodiments corresponding to the log processing device, reference can be made to the relevant description in the embodiments corresponding to the log processing method, which will not be elaborated here one by one.
[0075] An embodiment of the present application further provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the embodiments of the above log processing method.
[0076] An embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the steps in any one of the embodiments of the above log processing method when running.
[0077] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc, and other various media that can store computer programs.
[0078] An embodiment of the present application further provides a computer program product, where the computer program product includes a computer program, and the computer program implements the steps in any one of the embodiments of the above log processing method when executed by a processor.
[0079] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, where the non-volatile computer-readable storage medium stores a computer program, and the computer program implements the steps in any one of the embodiments of the above log processing method when executed by a processor.
[0080] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0081] The above has introduced in detail a method, apparatus, electronic device, and storage medium for processing a log provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for processing logs, characterized in that, Including: Obtain a first log file; Calculate a similarity value between the first log file and a second log file; wherein, the second log file and the first log file are log files of different repetition cycles of the same server; When it is determined that the similarity value is less than a preset threshold, determine to retain the first log file.
2. The method for processing a log according to claim 1, wherein The first log file includes first time anchors injected at preset time intervals; the second log file includes second time anchors injected at the preset time intervals; Before calculating the similarity value between the first log file and the second log file, the method further includes: Based on the first time anchors, split the first log file to obtain at least one first log segment; Based on the second time anchors, split the second log file to obtain at least one second log segment.
3. The method for processing a log according to claim 2, wherein The calculating the similarity value between the first log file and the second log file includes: Determine first log segments and second log segments in the same time period in the first log file and the second log file; Calculate the similarity value between the first log segment and the second log segment in the same time period.
4. The method for processing a log according to claim 3, wherein The calculating the similarity value between the first log segment and the second log segment in the same time period includes: Calculate the hash value of the first log segment and the hash value of the second log segment respectively.
5. The method for processing a log according to claim 4, wherein, The determining to retain the first log file when it is determined that the similarity value is less than a preset threshold further includes: When it is determined that the hash value of the first log segment is different from the hash value of the second log segment, determine to retain the first log segment.
6. The method for processing a log according to claim 1, wherein Before obtaining the first log file, the method further includes: Create a preset number of local area network serial session threads; wherein, the preset number is the number of servers; Establish connection relationships between each of the local area network serial session threads and the server respectively, so as to send injection instructions for time anchors to the server through the local area network serial session threads.
7. The method for processing a log according to claim 6, wherein The obtaining the first log file further includes: After it is determined that the file capacity of the first log file reaches a preset log threshold, save a state snapshot and restart the server based on a lightweight kernel.
8. The method for processing a log according to any one of claims 1-7, characterized in that, After it is determined to retain the first log file, the method further includes; Associate log files corresponding to different repetition cycles based on time sequence, generate a fault propagation link, and store the fault propagation link in a preset database; When it is determined that the fault type corresponding to the fault propagation link is a marked event, generate a fault report, and push the fault report based on a preset push method; wherein, the fault report includes a fault path and a fault type.
9. A processing device for logs, characterized in that, Including: An obtaining unit, configured to obtain a first log file; A calculating unit, configured to calculate a similarity value between the first log file and a second log file; wherein, the second log file and the first log file are log files of different repetition cycles of the same server; A retaining unit, configured to determine to retain the first log file when it is determined that the similarity value is less than a preset threshold.
10. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor for implementing the steps of the log processing method according to any one of claims 1 to 8 when executing the computer program.
11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the log processing method according to any one of claims 1 to 8 when executed by a processor.
12. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the log processing method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Log duplication eliminating processing method and device
CN106844143A
Log alarm information generation method and device, electronic equipment and storage medium
CN115348161A
Log processing method, control device and computer storage medium
CN117785823A