Log collection method, device, equipment and medium

Through the operating system kernel and management controller co-processing of hardware event logs, the problem of waste and inefficient resource collection of management controller logs is solved, and efficient and secure log data processing and fault diagnosis is achieved, reducing costs.

CN120353620BActive Publication Date: 2025-08-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510804698.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-22
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In the prior art, management controller log collection relies on out-of-band management, resulting in waste of resources, high cost and low efficiency, and cross-network transmission logs are susceptible to bandwidth limitations and high latency.

Method used

Through the operating system kernel and management controller, the driver module captures hardware events and writes them to the shared memory area. The management controller reads and parses log data, performs structured processing and encryption, and uploads it to the cloud to realize in-band communication.

Benefits of technology

It realizes efficient and low-overhead processing of hardware status monitoring and fault diagnosis, reduces dependence on dedicated hardware, reduces redundant storage requirements, improves system stability and fault diagnosis efficiency, and saves costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353620B_ABST
    Figure CN120353620B_ABST
Patent Text Reader

Abstract

The present invention discloses a log collection method, device, equipment and medium, which relate to the field of log processing technology. The method includes: receiving a target instruction from a driver module embedded in an operating system kernel; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area; in response to the target instruction, reading original log data from the shared memory area; parsing key fields of the original log data and converting it into structured log data; encrypting the structured log data in a pre-established security isolation area, and uploading the processed log data to the cloud so that the cloud returns a storage confirmation signal. The method is based on in-band communication and communicates with a management controller through the operating system kernel. It can reduce hardware dependence, achieve efficient collection and processing of hardware event log data, reduce redundant overhead in the data processing process, and improve the efficiency of hardware status monitoring and fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of log processing technology, and in particular to a log collection method, device, equipment and medium. Background Art

[0002] Management controller log collection typically relies on out-of-band management, which requires an independent network channel and has many drawbacks. For example, out-of-band management requires additional hardware resources, such as a dedicated network card for the management controller, resulting in wasted resources and high costs. Transmitting logs across networks is susceptible to bandwidth limitations, resulting in high latency and low efficiency. Summary of the Invention

[0003] The present invention provides a log collection method, apparatus, device and medium, which can realize efficient and low-overhead hardware status monitoring and fault diagnosis through the collaboration of an operating system kernel and a management controller.

[0004] The present invention provides a log collection method for managing a controller, comprising:

[0005] Receive target instructions from a driver module embedded in the operating system kernel; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area;

[0006] In response to the target instruction, reading original log data from the shared memory area;

[0007] Parsing the key fields of the original log data and converting it into structured log data;

[0008] The structured log data is encrypted in a pre-created secure isolation area, and the processed log data is uploaded to the cloud, so that the cloud returns a storage confirmation signal.

[0009] The present invention also provides a log collection device for managing a controller, comprising:

[0010] An instruction receiving module is used to receive target instructions from a driver module embedded in the operating system kernel; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area;

[0011] a data reading module, configured to read original log data from the shared memory area in response to the target instruction;

[0012] A log processing module, configured to parse the original log data for key fields and convert it into structured log data;

[0013] The encryption transmission module is used to encrypt the structured log data in a pre-created security isolation area and upload the processed log data to the cloud so that the cloud returns a storage confirmation signal.

[0014] The present invention also provides an electronic device, comprising: a server and a management controller; wherein the operating system kernel of the server is embedded with a driver module;

[0015] The driver module is configured to capture hardware events and write corresponding log data into a pre-established shared memory area, and send target instructions to the management controller;

[0016] The management controller is used to receive and respond to the target instruction, read the original log data from the shared memory area; parse the key fields of the original log data and convert it into structured log data; encrypt the structured log data in a pre-created secure isolation area, and upload the processed log data to the cloud, so that the cloud returns a storage confirmation signal.

[0017] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned log collection methods are implemented.

[0018] Through the present invention, since the log collection method provided by the present invention is based on in-band communication, it directly communicates with the management controller through the operating system kernel. First, the driver module embedded in the operating system kernel captures the hardware event and writes the corresponding log data into the shared memory area, which can avoid user-mode polling delay and reduce dependence on specific hardware devices; after the log is written, the target instruction is sent to the management controller; then the management controller reads the original log data from the shared memory area, and uploads it to the cloud after parsing, conversion and encryption processing, thereby realizing efficient collection and processing of hardware event log data; the structured processing and encrypted upload of log data make the hardware status monitoring and fault diagnosis process more standardized and secure. The coordinated operation of each link reduces the redundant overhead in the data processing process and improves the efficiency of hardware status monitoring and fault diagnosis, thereby achieving real-time monitoring of hardware status and rapid fault diagnosis in a low-overhead, highly coordinated manner, ensuring stable system operation; and can achieve dual savings in hardware and operation and maintenance, eliminating dedicated network cards and independent management networks, reducing redundant storage requirements, and saving costs.

[0019] In addition, the present invention also provides a corresponding log collection device, electronic device and computer-readable storage medium for the log collection method, which have the same or corresponding technical features as the above-mentioned log collection method and have the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 A flow chart of a log collection method provided by an embodiment of the present invention;

[0022] Figure 2 A schematic diagram of the interaction between the operating system kernel, management controller, and cloud provided by an embodiment of the present invention;

[0023] Figure 3 A schematic diagram of an interrupt trigger timing sequence provided by an embodiment of the present invention;

[0024] Figure 4 A schematic diagram of a framework corresponding to the log collection method provided in an embodiment of the present invention;

[0025] Figure 5 A specific flow chart of the log collection method provided by an embodiment of the present invention;

[0026] Figure 6 A schematic diagram of the structure of a log collection device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0030] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the log collection method depends, the specific application environment architecture or specific hardware architecture is described here.

[0031] An embodiment of the present invention provides a log collection method, which is described in detail in conjunction with the execution flow of the log collection method. Figure 1 A flow chart of a log collection method provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method is used to manage the controller, including:

[0032] S101, receiving a target instruction from a driver module embedded in an operating system kernel; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area.

[0033] It should be noted that the management controller used in the present invention may be a baseboard management controller (BMC) or other management controller, without limitation. The operating system may be an operating system for a server based on a microprocessor architecture. The microprocessor may be an Advanced RISC Machine (ARM) computer. With the explosive growth in demand for cloud computing, artificial intelligence, and edge computing, the penetration rate of ARM-based servers has increased significantly, and the development trend of micro data centers is becoming increasingly prominent. Therefore, the in-band log collection method provided by the present invention can achieve efficient and low-overhead hardware status monitoring and fault diagnosis by collaborating with the operating system kernel and the management controller of an ARM-based server. The present invention first embeds a driver module into the operating system kernel, which can be a lightweight driver module. In the driver module, a shared memory area with the management controller can be established through memory mapping (MMAP). The driver module is responsible for capturing hardware events (such as abnormal processor temperature, memory error checking and correction, etc.) and writing the corresponding log data to a pre-established shared memory area. The operating system of the present invention interacts with the management controller via shared memory and is primarily implemented as a driver kernel module.

[0034] The driver module of the present invention adheres to minimalist principles when writing code related to shared memory initialization and address mapping, interrupt handling, data packetization / unpacking, and protocol security verification. This approach retains only the code necessary to implement core functionality, eliminating unnecessary functionality, complex logical structures, redundant dependencies, and excessive design. This ensures concise, single-function, and efficient code. Protocol security verification can include cyclic redundancy checks (CRC) or hash-based message authentication codes (HMAC). From a development perspective, concise code offers clear logic, is easy to understand, and maintain, allowing developers to quickly identify and correct issues, reducing debugging time. Reduced code size also means fewer potential errors and vulnerabilities, improving system stability and reliability. Furthermore, the single, clear functional design reduces coupling between code modules, enhancing code reusability and scalability, and facilitating subsequent system upgrades or feature additions.

[0035] At the same time, the present invention can directly operate the pre-allocated shared memory area, avoid data copying from user space to kernel space, and map the physical memory address to the kernel virtual address space through the ioremap (input / output memory remapping) or memremap (memory remapping) function, so as to ensure that both parties can read and write the same memory segment, improve data transmission efficiency, reduce system resource consumption, support direct hardware access, and reduce memory usage costs.

[0036] The present invention also pre-allocates a fixed physical memory area (e.g., 64MB) through the Device Tree Source (DTS) or Unified Extensible Firmware Interface (UEFI) / Basic Input / Output System (BIOS) and marks it as reserved memory, ensuring that the operating system kernel does not occupy this area. A double-buffering mechanism is used to achieve efficient data exchange. Double buffering achieves lockless synchronization by alternating between two buffers and combining them with atomic variables (e.g., sequence counters), avoiding the context switching overhead and potential deadlock risk associated with traditional mutex locks.

[0037] S102 : Reading original log data from the shared memory area in response to the target instruction.

[0038] In implementation, after writing the log data, the driver module of the operating system kernel may send a target instruction to the management controller to notify the management controller to read the original log data from the shared memory area.

[0039] S103: parse the key fields of the original log data and convert it into structured log data.

[0040] In implementation, after reading the raw log data, the management controller can parse its key fields (such as error code, timestamp, hardware location, etc.) and convert them into structured log data.

[0041] S104: encrypt the structured log data in a pre-created secure isolation zone, and upload the processed log data to the cloud, so that the cloud returns a storage confirmation signal.

[0042] In implementation, the present invention can use TrustZone (hardware-level security isolation) technology to create a secure isolation zone (Trusted Execution Environment, TEE) to encrypt structured log data. The encrypted log can be uploaded to the cloud through an encryption protocol (Transport Layer Security, TLS) channel to prevent man-in-the-middle attacks, data leakage and tampering. After receiving the encrypted log, the cloud returns a storage confirmation signal to the management controller.

[0043] In the above-mentioned log collection method provided by the embodiment of the present invention, based on in-band communication, the operating system kernel directly communicates with the management controller. First, the driver module embedded in the operating system kernel captures the hardware event and writes the corresponding log data into the shared memory area, which can avoid user-mode polling delay and reduce dependence on specific hardware devices; after the log is written, the target instruction is sent to the management controller; then the management controller reads the original log data from the shared memory area, and uploads it to the cloud after parsing, conversion and encryption processing, thereby realizing efficient collection and processing of hardware event log data; the structured processing and encrypted upload of log data make the hardware status monitoring and fault diagnosis process more standardized and secure. The coordinated operation of each link reduces the redundant overhead in the data processing process and improves the efficiency of hardware status monitoring and fault diagnosis, thereby achieving real-time monitoring of hardware status and rapid fault diagnosis in a low-overhead, highly coordinated manner, ensuring stable system operation; and can achieve dual savings in hardware and operation and maintenance, eliminating dedicated network cards and independent management networks, reducing redundant storage requirements, and saving costs.

[0044] Figure 2 Schematic diagram of the interaction between the operating system kernel, management controller and cloud provided by the embodiment of the present invention. Figure 2As shown, the operating system kernel is responsible for low-level hardware interaction and resource management. It can perceive changes in server hardware status, failure events, and other events, preprocess hardware event log data, and transmit the preprocessed log data to the management controller via shared memory. The management controller can read log data from shared memory to achieve cross-module data transmission, classify logs, isolate and store related logs, and generate diagnostic reports. The cloud-based log aggregation center can receive encrypted logs from the secure isolation zone and perform storage confirmation; the cloud-based sharing platform can receive diagnostic reports and perform in-depth log analysis.

[0045] To prevent data loss or tampering, the present invention can limit the direct memory access (DMA) range of the management controller through the Input / Output Memory Management Unit (IOMMU) / System Memory Management Unit (SMMU), allowing only access to the shared memory area.

[0046] Furthermore, in a specific implementation, in the above-mentioned log collection method provided in an embodiment of the present invention, before executing step S101 to receive the target instruction from the driver module embedded in the operating system kernel, it may also include: monitoring the level status of the input and output pins, and when the input and output pins are pulled low by the driver module while writing log data, triggering an internal interrupt signal and executing an interrupt; or, monitoring the value of the setting register, and when the preset value is written into the setting register by the driver module while writing log data, triggering an internal interrupt signal and executing an interrupt.

[0047] In implementation, Figure 3 Schematic diagram of interrupt trigger timing provided by an embodiment of the present invention. Figure 3As shown in the figure, hardware detects an anomaly (such as a fault or an out-of-bounds value), triggering an interrupt request, which is then handled by the operating system kernel. The operating system kernel then writes the anomaly log information to a shared memory area. It can trigger a management controller interrupt by pulling a general-purpose input / output (GPIO) pin low or writing to a specific register, instructing the management controller to read the log data from the shared memory. The operating system can trigger a management controller interrupt using an Intelligent Platform Management Interface (IPMI) command or a memory-mapped input / output (MMIO) register write. After the management controller writes the data, it triggers a message signaled interrupt (MSI) or message signaled interrupt extended (MSI-X). The driver then registers a handler function using request_irq (the core system call used to register hardware interrupt handlers).

[0048] The management controller can also classify and compress log data in shared memory and upload the encrypted logs to the cloud via an encrypted channel. After receiving the logs, the cloud returns an acknowledgment (ACK) signal and automatically retransmits the data if a timeout occurs. The management controller can also notify the operating system kernel to clear the interrupt status and avoid repeated responses.

[0049] It should be noted that the present invention can use an asynchronous event-driven model to handle management controller interrupts, implementing non-blocking interrupt responses through tasklets (a lightweight interrupt bottom-half processing mechanism) or workqueues (a kernel-thread-based deferred processing mechanism) within the kernel. This avoids directly executing time-consuming operations (such as memory copies and complex logic) during interrupt processing, preventing blocking of other kernel threads and ensuring system real-time performance.

[0050] The present invention can also use non-cacheable memory to mark the shared memory area as non-cacheable; or, call a cache refresh instruction (such as clflush) or a DMA synchronization interface (dma_sync) after a critical operation to ensure that the processor and the management controller have the same view. Figure 1 To.

[0051] Furthermore, in a specific implementation, in the above-mentioned log collection method provided in an embodiment of the present invention, before executing step S103 to parse the key fields of the original log data, it can also include: detecting its own heartbeat signal to determine whether the heartbeat signal is within the set heartbeat range; if so, parsing the key fields of the original log data; if not, transferring the original log data to a local Secure Digital Card (SD) card or a redundant storage device.

[0052] During implementation, the present invention can monitor its own operating status in real time through the management controller, such as detecting the heartbeat signal. If an abnormality is detected in the management controller (such as heartbeat packet loss, heartbeat signal is not within the set heartbeat range), the unuploaded logs will be immediately transferred to the local secure digital card or redundant storage device to prevent data loss and exit the current processing flow.

[0053] Furthermore, in a specific implementation, in the above-mentioned log collection method provided in an embodiment of the present invention, step S103 parses the key fields of the original log data and converts it into structured log data, which may specifically include: performing word segmentation processing on the original log data, combining regular expression matching and semantic analysis to extract key fields; and converting the key fields into structured log data.

[0054] In practice, the management controller reads raw logs from shared memory and uses regular expressions and word segmentation techniques to parse key fields (such as error codes, timestamps, and hardware locations) and convert them into structured data in a predefined format. This format can be a lightweight data exchange format like JSON (JavaScript Object Notation). This replaces manual parsing, significantly improving log analysis efficiency and reducing operation and maintenance costs.

[0055] Furthermore, in a specific implementation, in the above-mentioned log collection method provided in an embodiment of the present invention, after executing step S103 to convert into structured log data, and before executing step S104 to encrypt the structured log data, it may also include: classifying the structured log data into key logs and non-key logs according to the characteristics of the structured log data.

[0056] In the above steps, structured log data is classified into critical logs and non-critical logs based on its characteristics. Specifically, the steps may include: building a log classification model; collecting key log information and oversampling it using the Synthetic Minority Over-sampling Technique (SMOTE) algorithm to obtain a sample set; training the log classification model using the sample set; optimizing the loss function using focal loss or weighted cross entropy during training, annotating samples whose confidence level output by the log classification model is lower than a set confidence level, combining the output of the log classification model with predefined rules to implement multi-level priority processing, and locating the root cause through correlation analysis; triggering a circuit breaker mechanism when log traffic exceeds a set threshold; classifying the structured log data using the trained log classification model based on its characteristics; initiating an alarm process when the log data is classified as a critical log; and compressing the non-critical log data using a compression algorithm.

[0057] In implementation, the log classification model can be a lightweight XGBoost (eXtreme GradientBoosting, gradient boosting tree algorithm) model, and the inference speed can be optimized through ARM Neon (single instruction multiple data under the ARM architecture) instructions.

[0058] This invention utilizes a lightweight XGBoost model to dynamically classify logs into critical and non-critical logs based on their characteristics. Critical logs can include P0-level faults (i.e., the highest priority and most severe faults). When classified as critical, an alert process can be initiated promptly. Non-critical logs can be compressed, with a compression rate of greater than 80%.

[0059] It should be noted that this invention dynamically filters critical logs based on machine learning. The main implementation methods are as follows: Machine learning model design primarily involves model selection. First, real-time classification is performed. This can predict error trends and detect unknown errors. During the sampling phase, critical logs (positive samples) are oversampled using a synthetic minority oversampling algorithm, focal loss, or weighted cross-entropy, to improve minority class recognition. Low-confidence samples are then requested to be annotated by operations personnel. Second, priority grading is implemented, employing a multi-level processing strategy that combines model output with predefined rules and correlates root causes. Third, system optimization is implemented, including the ability to convert FP32 (32-bit single-precision floating-point model) to INT8 (8-bit integer model) while maintaining over 95% accuracy. Preprocessing and inference logic are compiled into eBPF (extended Berkeley Packet Filter) programs or kernel modules, enabling inference on graphics processing units (GPUs) or neural network processing units (NPUs) for models like XGBoost. Fourth, resource isolation, including limiting the processor quota and memory bandwidth of log processing threads and setting up a circuit breaker mechanism: when the log spike exceeds the threshold, it will downgrade to basic rule filtering. Fifth, observability, using Local Interpretable Model-agnostic Explanations (LIME) or Shapley Additive Explanations (SHAP) to explain individual prediction results.

[0060] Figure 4 Schematic diagram of the framework corresponding to the log collection method provided by the embodiment of the present invention. Figure 4As shown, the system parses raw log data, converting unstructured data into structured data. Valuable features are extracted from the parsed log data, which serve as input for subsequent log classification model inference. A pre-trained log classification model is used to infer the extracted features and predict whether the log data contains anomalies. The model inference results are dynamically filtered to remove potential false positives and improve detection accuracy. Based on the model inference and dynamic filtering results, the system determines whether the log data is critical or an anomaly. Log data identified as critical is encrypted and stored to ensure data security and traceability. If an anomaly is identified, a real-time alert is triggered, notifying relevant personnel for prompt action. Relevant personnel evaluate the results and provide feedback, comparing the actual situation with the system's judgment. Online learning is conducted using human feedback (such as false positive labeling) to continuously optimize the model and features. Based on the online learning results, the local model weights are updated to improve model accuracy. The extracted features are optimized to better reflect the essential characteristics of the log data.

[0061] Furthermore, in a specific implementation, in the above-mentioned log collection method provided in an embodiment of the present invention, step S104 encrypts the structured log data in a pre-created secure isolation area, which may specifically include: using hardware-level secure isolation to divide the operating environment into a secure world and a normal world; establishing memory isolation between the secure world and the normal world, and creating a secure isolation area through an address space controller; encrypting key logs in the structured log data in the secure isolation area, and storing the key in the secure isolation area.

[0062] In practice, encryption can be performed using symmetric encryption, such as AES-256 and AES-GCM. AES-256 is a variant of the Advanced Encryption Standard (AES) that uses a 256-bit key for data encryption, while AES-GCM (Advanced Encryption Standard - Galois / Counter Mode) is an authenticated encryption mode. The encryption key can be injected through the secure boot process to prevent man-in-the-middle attacks. In practical applications, the present invention can encrypt only critical logs or both critical and non-critical logs, depending on the specific situation.

[0063] It should be noted that this invention utilizes ARM TrustZone technology to encrypt the log transmission path. The specific design is as follows: First, the TrustZone security partitioning design: world division, including the secure world and the normal world; memory isolation, using the TrustZone Address Space Controller (TZASC) to divide secure and non-secure memory areas, preventing the normal world from accessing the encrypted log buffer; hardware encryption acceleration, providing hardware-level Advanced Encryption Standard (AES) / Secure Hash Algorithm (SHA) acceleration to reduce encryption latency. Second, the end-to-end log encryption process: log generation and capture, such as kernel logs, can be redirected to the secure world buffer by modifying the printk function (which outputs messages to the kernel log buffer). User-mode logs can be transferred to the secure isolation zone via the ioctl (device control kernel interface resolution) interface or shared memory (which must be configured as secure memory). The secure world's encryption processing includes key management and signing. Secure transmission and storage include adding a monotonically increasing serial number to each log entry. The receiving end verifies serial number continuity and rejects data with outdated serial numbers. JTAG (Joint Test Action Group) access is disabled for the common world, allowing only authorized secure world users to decrypt logs through the secure debug channel. TrustZone I / O virtualization (SMMU configuration) ensures that the DMA transfer path cannot be tampered with by the common world. Session keys are rotated periodically (e.g., hourly), and old keys are immediately erased.

[0064] Figure 5 A specific flow chart of the log collection method provided by an embodiment of the present invention is as follows: Figure 5As shown, the operating system kernel first creates a shared memory area for data exchange with the management controller, facilitating data transfer and sharing. The kernel then monitors and captures various events generated by the system's hardware devices, checking the processor load to understand the system's operating status. If the processor load is too high, the hardware event sampling frequency is reduced to reduce system burden. If the processor load is normal, operations continue according to the current sampling frequency and processing strategy. The captured and processed hardware event data is written to the previously created shared memory area for the management controller to read, triggering an interrupt to the management controller. After reading the log data from the shared memory, the management controller checks its own heartbeat signal. If normal, it parses and segmentes the data and converts it into JSON format for easier processing and storage. Next, a log classification model is used to classify the logs, identifying critical and non-critical logs. The corresponding logs are compressed and stored, and real-time alerts are issued for abnormal conditions. Based on human feedback, the log classification model is updated online to improve processing accuracy. The logs are then encrypted and uploaded to the cloud via a secure network channel. The cloud then sends a confirmation message to the management controller. If the management controller heartbeat detection is abnormal, in order to prevent data loss, the data is urgently transferred to a secure digital card or other redundant storage device and the current processing flow is exited.

[0065] The entire process achieves the monitoring, processing and log management of hardware events through the collaborative work of the operating system kernel and the management controller. It also has the functions of real-time alarm and data security storage, and continuously optimizes the processing process through manual feedback and online model updates.

[0066] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0067] An embodiment of the present application also provides a log collection device. Figure 6 This is a schematic diagram of the structure of the log collection device provided by an embodiment of the present invention. This embodiment is based on the perspective of functional modules, such as Figure 6 As shown, the device is used to manage the controller, including:

[0068] The instruction receiving module 10 is used to receive target instructions from the driver module embedded in the operating system kernel; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area;

[0069] The data reading module 11 is used to read the original log data from the shared memory area in response to the target instruction;

[0070] The log processing module 12 is used to parse the key fields of the original log data and convert it into structured log data;

[0071] The encryption transmission module 13 is used to encrypt the structured log data in a pre-created security isolation area and upload the processed log data to the cloud so that the cloud returns a storage confirmation signal.

[0072] In the above-mentioned log collection device provided in the embodiment of the present invention, the interaction of the above-mentioned four modules can realize efficient collection and processing of hardware event log data based on in-band communication, reduce redundant overhead in the data processing process, and improve the efficiency of hardware status monitoring and fault diagnosis, thereby achieving real-time monitoring of hardware status and rapid fault diagnosis in a low-overhead and highly coordinated manner, ensuring stable operation of the system; and can achieve dual savings in hardware and operation and maintenance, eliminate dedicated network cards and independent management networks, reduce redundant storage requirements, and save costs.

[0073] Since the embodiments of the log collection device correspond to the embodiments of the log collection method, the description of the features in the corresponding embodiments of the log collection device can be found in the description of the corresponding embodiments of the log collection method, and will not be repeated here. The same beneficial effects as the aforementioned log collection method are achieved.

[0074] Furthermore, in a specific implementation, the above-mentioned log collection device provided in an embodiment of the present invention may also include: an interrupt module, which is used to monitor the level status of the input and output pins, and when the input and output pins are pulled low by the driving module while writing log data, an internal interrupt signal is triggered and an interrupt is executed; or, monitor the value of the setting register, and when the preset value is written into the setting register by the driven module while writing log data, an internal interrupt signal is triggered and an interrupt is executed.

[0075] Furthermore, in a specific implementation, the above-mentioned log collection device provided in an embodiment of the present invention may also include: a heartbeat detection module, which is used to detect the heartbeat signal of the operating system kernel and determine whether the operating system is online; if it is online, the processing process of the log processing module 12 is executed; if it is not online, the original log data is transferred to a local secure digital card or a redundant storage device.

[0076] Furthermore, in specific implementation, in the above-mentioned log collection device provided in the embodiment of the present invention, the log processing module 12 can be specifically used to perform word segmentation processing on the original log data, combine regular expression matching and semantic analysis, extract key fields; and convert the key fields into structured log data with a set format.

[0077] Furthermore, in a specific implementation, the log collection device provided in the embodiment of the present invention may further include: a log classification module, configured to classify the structured log data into key logs and non-key logs according to the characteristics of the structured log data.

[0078] In implementation, the log classification module can be specifically used to build a log classification model; collect key log information, and use the synthetic minority oversampling algorithm for oversampling to obtain a sample set; use the sample set to train the log classification model; during the training process, use focal loss or weighted cross entropy to optimize the loss function, and mark the samples whose confidence level of the log classification model output is lower than the set confidence level, and combine the log classification model output with predefined rules to achieve multi-level priority processing, and locate the root cause through correlation analysis; when the log traffic exceeds the set threshold, the circuit breaker mechanism is triggered; according to the characteristics of the structured log data, the trained log classification model is used to classify and process the structured log data; when it is classified as a critical log, the alarm process is started; when it is classified as a non-critical log, the compression algorithm is used to compress the non-critical log.

[0079] Furthermore, in a specific implementation, in the above-mentioned log collection device provided in an embodiment of the present invention, the encryption transmission module 13 can be specifically used to create a secure isolation zone; encrypt key logs in the structured log data in the secure isolation zone, and store the key in the secure isolation zone.

[0080] An embodiment of the present application also provides an electronic device, including a server and a management controller; wherein the operating system kernel of the server is embedded with a driver module; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area, and send target instructions to the management controller; the management controller is used to receive and respond to the target instructions, read the original log data from the shared memory area; parse the key fields of the original log data and convert it into structured log data; encrypt the structured log data in a pre-created secure isolation area, and upload the processed log data to the cloud so that the cloud returns a storage confirmation signal.

[0081] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned log collection method embodiments when running.

[0082] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0083] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned log collection method embodiments are implemented.

[0084] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned log collection method embodiments are implemented.

[0085] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0086] The above is a detailed introduction to the log collection method, device, equipment, and medium provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A log collection method, characterized in that: Used to manage controllers, including: Monitor the level status of input and output pins, and when the input and output pins are pulled low by the driver module embedded from the operating system kernel while writing log data, trigger an internal interrupt signal and execute an interrupt; or monitor the value of a setting register, and when the setting register is written with a preset value by the driver module embedded from the operating system kernel while writing log data, trigger an internal interrupt signal and execute an interrupt; Receive the target instruction of the driver module; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area; In response to the target instruction, reading original log data from the shared memory area; Parsing the key fields of the original log data and converting it into structured log data; Encrypting the structured log data in a pre-created secure isolation zone and uploading the processed log data to the cloud, so that the cloud returns a storage confirmation signal; In the log collection method, based on in-band communication, the operating system kernel directly communicates with the management controller.

2. The log collection method according to claim 1, characterized in that: Before parsing the key fields of the original log data, the following steps are also included: Detect its own heartbeat signal and determine whether the heartbeat signal is within the set heartbeat range; If yes, parse the key fields of the original log data; If not, the original log data is transferred to a local secure digital card or redundant storage device.

3. The log collection method according to claim 1, wherein: Parsing the key fields of the original log data and converting it into structured log data includes: Perform word segmentation on the original log data, combine regular expression matching and semantic analysis, and extract key fields; The key fields are converted into structured log data having a set format.

4. The log collection method according to claim 1, wherein: After converting into structured log data and before encrypting the structured log data, the method further includes: Classifying the structured log data into key logs and non-key logs according to characteristics of the structured log data; The structured log data is encrypted, including: The key log is encrypted.

5. The log collection method according to claim 4, characterized in that: Classifying the structured log data into key logs and non-key logs according to the characteristics of the structured log data includes: Build a log classification model; Collect key log information and use the synthetic minority oversampling algorithm to perform oversampling to obtain a sample set; The log classification model is trained using the sample set; during the training process, the loss function is optimized using focal loss or weighted cross entropy, and samples whose confidence level output by the log classification model is lower than a set confidence level are marked, and the output of the log classification model is combined with predefined rules to implement multi-level priority processing, and the root cause is located through correlation analysis; when the log traffic exceeds the set threshold, a circuit breaker mechanism is triggered; According to the characteristics of the structured log data, the trained log classification model is used to classify the structured log data; When a log is classified as a critical log, the alarm process is initiated; When a log is classified as a non-critical log, a compression algorithm is used to compress the non-critical log.

6. The log collection method according to claim 1, wherein: The structured log data is encrypted in a pre-created secure isolation zone, including: Use hardware-level security isolation to divide the operating environment into a secure world and a normal world; Establishing memory isolation between the secure world and the normal world, and creating a secure isolation zone through an address space controller; Key logs in the structured log data are encrypted in the secure isolation area, and the key is stored in the secure isolation area.

7. A log collection device, characterized in that: Used to manage controllers, including: An interrupt module is configured to monitor the level status of input and output pins, and trigger an internal interrupt signal and execute an interrupt when the input and output pins are pulled low by the driver module embedded in the operating system kernel while writing log data; or monitor the value of a setting register, and trigger an internal interrupt signal and execute an interrupt when the setting register is written with a preset value by the driver module embedded in the operating system kernel while writing log data; An instruction receiving module is used to receive target instructions from the driver module; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area; A data reading module, configured to read original log data from the shared memory area in response to the target instruction; A log processing module, configured to parse the original log data for key fields and convert it into structured log data; An encryption transmission module, configured to encrypt the structured log data in a pre-created secure isolation zone and upload the processed log data to the cloud, so that the cloud returns a storage confirmation signal; The log collection device directly communicates with the management controller through the operating system kernel based on in-band communication.

8. An electronic device, characterized in that: include: A server and a management controller; wherein the operating system kernel of the server is embedded with a driver module; in the electronic device, the operating system kernel directly communicates with the management controller based on in-band communication; The driver module is configured to capture hardware events and write corresponding log data into a pre-established shared memory area, and send target instructions to the management controller; The management controller is used to monitor the level status of the input and output pins, and when the input and output pins are pulled low by the driver module while writing log data, trigger an internal interrupt signal and execute an interrupt; or, monitor the value of the setting register, and when the preset value is written into the setting register by the driver module while writing log data, trigger an internal interrupt signal and execute an interrupt; receive and respond to the target instruction, read the original log data from the shared memory area; parse the key fields of the original log data and convert it into structured log data; encrypt the structured log data in a pre-created secure isolation area, and upload the processed log data to the cloud, so that the cloud returns a storage confirmation signal.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the log collection method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method for finishing software log by utilizing kernel module and application module under linux system

    CN103176891A

  • Log processing method and device, electronic equipment and storage medium

    CN116361106A