Log collection method and device, equipment and medium

Through the operating system kernel and management controller working in concert, hardware events are captured and structured processing and encrypted uploads are solved, and the management controller log collection problems are wasted and efficient and low-overhead hardware status monitoring and fault diagnosis are achieved.

CN120353620AActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510804698.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-22
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Existing management controller log collection relies on out-of-band management, resulting in waste of resources, high cost and low efficiency, and high cross-network transmission latency.

Method used

Through the operating system kernel and the management controller, the driver module captures hardware events and writes them to the shared memory area. The management controller reads and parses log data, performs structured processing and encryption, and uploads it to the cloud.

Benefits of technology

It realizes efficient and low-overhead processing of hardware status monitoring and fault diagnosis, reduces hardware dependence, reduces resource waste and costs, and improves system stability and fault diagnosis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353620A_ABST
    Figure CN120353620A_ABST
Patent Text Reader

Abstract

The invention discloses a log collection method, device and equipment and a medium, and relates to the technical field of log processing.The method comprises the steps that a target instruction from a driving module embedded in an operating system kernel is received; the driving module is used for capturing a hardware event and writing corresponding log data into a pre-established shared memory area; in response to the target instruction, reading original log data from the shared memory area; performing key field analysis on the original log data, and converting the original log data into structured log data; and performing encryption processing on the structured log data in a pre-created security isolation region, and uploading the processed log data to the cloud, so that the cloud returns a storage confirmation signal. The method is based on in-band communication and communicates with a management controller through an operating system kernel, hardware dependence can be reduced, efficient collection and processing of hardware event log data are achieved, redundant overhead in the data processing process is reduced, and the efficiency of hardware state monitoring and fault diagnosis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of log processing, and in particular, to a log collection method, device, equipment and medium. Background Art

[0002] The collection of management controller logs usually relies on out-of-band management, which requires an independent network channel and has many defects. For example, out-of-band management needs to occupy additional hardware resources, such as a dedicated network card for the management controller, resulting in waste of resources and high costs; cross-network transmission of logs is vulnerable to bandwidth limitations, with high latency and low efficiency. Summary of the Invention

[0003] The present invention provides a log collection method, device, equipment and medium, which can achieve efficient and low-overhead hardware status monitoring and fault diagnosis through the cooperation of the operating system kernel and the management controller.

[0004] The present invention provides a log collection method for a management controller, including: Receiving a target instruction from a driver module embedded in the operating system kernel; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area; Responding to the target instruction, reading the original log data from the shared memory area; Parsing key fields of the original log data and converting them into structured log data; Performing encryption processing on the structured log data in a pre-created secure isolation area and uploading the processed log data to the cloud so that the cloud returns a storage confirmation signal.

[0005] The present invention also provides a log collection device for a management controller, including: An instruction receiving module, configured to receive a target instruction from a driver module embedded in the operating system kernel; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area; A data reading module, configured to respond to the target instruction and read the original log data from the shared memory area; A log processing module, configured to parse key fields of the original log data and convert them into structured log data; An encryption transmission module, configured to perform encryption processing on the structured log data in a pre-created secure isolation area and upload the processed log data to the cloud so that the cloud returns a storage confirmation signal.

[0006] The present invention also provides an electronic device, including: a server and a management controller; wherein, a driver module is embedded in the operating system kernel of the server; The driving module is used to capture hardware events, write corresponding log data into a pre-established shared memory area, and send a target instruction to the management controller; The management controller is used to receive and respond to the target instruction, read the original log data from the shared memory area; parse the keyword fields of the original log data and convert them into structured log data; encrypt the structured log data in a pre-created secure isolation area, and upload the processed log data to the cloud to enable the cloud to return a storage confirmation signal.

[0007] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above log collection methods are implemented.

[0008] Through the present invention, since the log collection method provided by the present invention is based on in-band communication and directly communicates with the management controller through the operating system kernel. First, the driving module embedded in the operating system kernel captures hardware events and writes the corresponding log data into the shared memory area, which can avoid user-mode polling latency and reduce the dependence on specific hardware devices; after the log is written, a target instruction is sent to the management controller; then the management controller reads the original log data from the shared memory area, and after parsing, conversion and encryption processing, it uploads the data to the cloud, realizing the efficient collection and processing of hardware event log data; the structured processing and encrypted upload of log data make the hardware status monitoring and fault diagnosis processes more standardized and secure. Each link operates in coordination, reducing the redundant overhead in the data processing process, improving the efficiency of hardware status monitoring and fault diagnosis, and thus achieving real-time monitoring of hardware status and rapid fault diagnosis in a low-overhead and highly collaborative manner, ensuring the stable operation of the system; and it can achieve double savings in hardware and operation and maintenance, eliminating dedicated network cards and independent management networks, reducing redundant storage requirements, and saving costs.

[0009] In addition, the present invention also provides a corresponding log collection device, electronic device and computer-readable storage medium for the log collection method, which have the same or corresponding technical features as the above-mentioned log collection method, and the effects are the same. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a flowchart of the log collection method provided by the embodiment of the present invention; Figure 2 Interaction schematic diagram among the operating system kernel, management controller, and cloud provided by an embodiment of the present invention; Figure 3 Schematic diagram of the interrupt trigger timing provided by an embodiment of the present invention; Figure 4 Frame schematic diagram corresponding to the log collection method provided by an embodiment of the present invention; Figure 5 Specific flowchart of the log collection method provided by an embodiment of the present invention; Figure 6 Structural schematic diagram of the log collection device provided by an embodiment of the present invention. Detailed implementation manners

[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0013] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0014] To enable those skilled in the art of the present technology to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0015] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the log collection method depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0016] An embodiment of the present invention provides a log collection method, and the method will be described in detail in combination with the execution process of the log collection method. Figure 1 Flowchart of the log collection method provided by an embodiment of the present invention, as Figure 1 shown, this method is used for the management controller and includes: S101. Receive a target instruction from a driver module embedded in the operating system kernel; the driver module is used to capture hardware events and write the corresponding log data into a pre-established shared memory area.

[0017] It should be noted that the above-mentioned management controller used in the present invention can be a Baseboard Management Controller (BMC), or other management controllers can also be selected, which is not limited herein. The above-mentioned operating system can be an operating system for a server based on a microprocessor architecture. The microprocessor can be an Advanced RISC Machines (ARM). With the explosion of demands for cloud computing, artificial intelligence, and edge computing, the penetration rate of ARM architecture servers has increased significantly, and the development trend of micro data centers has become increasingly obvious. Therefore, the in-band log collection method provided by the present invention can achieve efficient and low-overhead hardware status monitoring and fault diagnosis through the cooperation of the operating system kernel of the ARM architecture server and the management controller. The present invention first embeds a driver module in the operating system kernel, and the driver module can be a lightweight driver module. In the driver module, a shared memory area with the management controller can be established through Memory-Mapped File (MMAP). The driver module can be responsible for capturing hardware events (such as abnormal processor temperature, memory error checking and correction, etc.) and writing the corresponding log data into the pre-established Shared Memory area. The operating system of the present invention interacts with the management controller through the shared memory, mainly implemented in the form of a driver kernel module.

[0018] When writing the code related to shared memory initialization and address mapping, interrupt handling, data packetization / depacketization, and protocol security verification in the driver module of the present invention, the Minimalist principle can be followed, only retaining the code necessary to implement the core functions, eliminating unnecessary functions, complex logical structures, redundant dependencies, and excessive designs, ensuring that the code is concise, has a single function, and is efficient. The protocol security verification can include Cyclic Redundancy Check (CRC) or Hash-based Message Authentication Code (HMAC). From a development perspective, concise code has clear logic, is easy to understand and maintain. Developers can quickly locate and modify problems, reducing debugging time; the reduction in code volume also means fewer potential errors and vulnerabilities, improving the stability and reliability of the system. In addition, the single and clear function design reduces the coupling degree between code modules, enhancing the reusability and extensibility of the code. When upgrading the system or adding new functions in the future, it can be more convenient to call and integrate.

[0019] Meanwhile, the present invention can directly operate on the pre-allocated shared memory area, avoiding data copying from user space to kernel space, and mapping the physical memory address to the kernel virtual address space through the ioremap (input / output memory remapping) or memremap (memory remapping) function, which can ensure that both parties can read and write the same memory segment, improve data transmission efficiency, reduce system resource consumption, support direct hardware access, and reduce memory usage costs.

[0020] The present invention can also pre-allocate a fixed physical memory area (such as 64MB) through the Device Tree Source (DTS) or the Unified Extensible Firmware Interface (UEFI) / Basic Input / Output System (BIOS), and mark it as reserved-memory to ensure that the operating system kernel does not occupy this area. The double-buffer mechanism is adopted to achieve efficient data exchange. The double buffer alternately uses two buffers and combines atomic variables (such as sequence number counters) to achieve lock-free synchronization, avoiding the context switching overhead and potential deadlock risks brought by traditional mutex locks.

[0021] S102. In response to the target instruction, read the original log data from the shared memory area.

[0022] In implementation, after writing the log data, the driver module of the operating system kernel can send a target instruction to the management controller to notify the management controller to read the original log data from the shared memory area.

[0023] S103. Parse the key fields of the original log data and convert it into structured log data.

[0024] In implementation, after reading the original log data, the management controller can parse its key fields (such as error code, timestamp, hardware location, etc.) and convert it into structured log data.

[0025] S104. Encrypt the structured log data in the pre-created secure isolation area and upload the processed log data to the cloud to enable the cloud to return a storage confirmation signal.

[0026] In implementation, the present invention can utilize TrustZone (hardware-level security isolation) technology to create a secure isolation area (Trusted Execution Environment, TEE), encrypt the structured log data, and upload the encrypted logs to the cloud through an encrypted protocol (Transport Layer Security, TLS) channel to prevent man-in-the-middle attacks, data leakage, and tampering. After receiving the encrypted logs, the cloud returns a storage confirmation signal to the management controller.

[0027] In the above log collection method provided by the embodiment of the present invention, based on in-band communication, the operating system kernel directly communicates with the management controller. First, a driver module embedded in the operating system kernel captures hardware events and writes the corresponding log data into a shared memory area, which can avoid user-mode polling latency and reduce the dependence on specific hardware devices. After the log is written, a target instruction is sent to the management controller. Then, the management controller reads the original log data from the shared memory area, uploads it to the cloud after parsing, conversion, and encryption processing, realizing the efficient collection and processing of hardware event log data. The structured processing and encrypted upload of log data make the hardware status monitoring and fault diagnosis processes more standardized and secure. Each link operates in coordination, reducing the redundant overhead in the data processing process, improving the efficiency of hardware status monitoring and fault diagnosis, and thus achieving real-time monitoring of hardware status and rapid fault diagnosis in a low-overhead and highly collaborative manner to ensure the stable operation of the system. And it can achieve double savings in hardware and operation and maintenance, eliminating the dedicated network card and independent management network, reducing the redundant storage requirement, and saving costs.

[0028] Figure 2 It is an interaction schematic diagram among the operating system kernel, the management controller, and the cloud provided by the embodiment of the present invention. As Figure 2 shown, the operating system kernel is responsible for underlying hardware interaction and resource management, can sense the state changes, faults, etc. of the server hardware, preprocess the log data of hardware events, and transfer the preprocessed log data to the management controller through the shared memory method. The management controller can read the log data from the shared memory to achieve cross-module data transfer, can also classify the logs, isolate and store the relevant logs, and can also generate a diagnostic report. The log aggregation center of the cloud can receive the encrypted logs of the secure isolation area and execute storage confirmation; the sharing platform of the cloud can receive the diagnostic report and perform in-depth analysis of the logs.

[0029] To prevent data loss or tampering, the present invention can limit the direct memory access (DMA) range of the management controller through an Input / Output Memory Management Unit (IOMMU) / System Memory Management Unit (SMMU), and only allow access to the shared memory area.

[0030] Further, in specific implementation, in the above-mentioned log collection method provided by the embodiments of the present invention, before executing step S101 to receive the target instruction from the driver module embedded in the operating system kernel, it may further include: monitoring the level state of the input / output pins, and when the input / output pins are pulled low by the driver module while writing log data, triggering an internal interrupt signal and executing an interrupt; or, monitoring the value of the set register, and when the set register is written with a preset value by the driver module while writing log data, triggering an internal interrupt signal and executing an interrupt.

[0031] In implementation, Figure 3 is a schematic diagram of the interrupt trigger timing provided by the embodiments of the present invention. As Figure 3 shown, when the hardware detects an abnormality (such as a fault, an out-of-bounds value), it triggers an interrupt request, which is processed by the operating system kernel; the operating system kernel writes the log information of the abnormality into the shared memory area, and can trigger a management controller interrupt by pulling low a General-Purpose Input / Output (GPIO) pin or writing to a specific register, notifying the management controller to read the log data in the shared memory. Among them, the operating system can trigger a management controller interrupt through an Intelligent Platform Management Interface (IPMI) command or a Memory-Mapped Input / Output (MMIO) write register. After the management controller writes data, it can trigger a Message Signaled Interrupts (MSI) / Message Signaled Interrupts eXtended (MSI-X) interrupt, and the driver registers a processing function through request_irq (a core system call for registering a hardware interrupt processing function).

[0032] The management controller can also share the log data of the memory for classification / compression, and upload the encrypted log to the cloud through an encrypted channel; after receiving the log, the cloud returns an acknowledgment signal (ACK); if it times out, it can be automatically retransmitted. The management controller can also clear the interrupt status by notifying the operating system kernel to avoid repeated responses.

[0033] It should be added that the present invention can adopt an asynchronous event-driven model to process management controller interrupts, and achieve non-blocking interrupt response through tasklet (a lightweight mechanism for handling the lower half of interrupts) or workqueue (a delayed processing mechanism based on kernel threads) in the kernel. This can avoid directly performing time-consuming operations (such as memory copying and complex logic) during interrupt processing, prevent blocking other kernel threads, and ensure the real-time performance of the system.

[0034] The present invention can also use non-cacheable memory to mark the shared memory area as non-cacheable; or, call a cache flush instruction (such as clflush) or a DMA synchronization interface (dma_sync) after a critical operation to ensure the consistency between the processor and the management controller. Figure 1 consistent.

[0035] Furthermore, in specific implementation, in the above-mentioned log collection method provided by the embodiment of the present invention, before performing step S103 to parse the key fields of the original log data, it may further include: detecting its own heartbeat signal, and determining whether the heartbeat signal is within a set heartbeat range; if so, parsing the key fields of the original log data; if not, storing the original log data in a local Secure Digital Card (SD card) or a redundant storage device.

[0036] In implementation, the present invention can monitor its own operating state in real time through the management controller. For example, by detecting the heartbeat signal, if it detects that the management controller is abnormal (such as the loss of the heartbeat packet and the heartbeat signal not being within the set heartbeat range), it immediately stores the unuploaded log in a local Secure Digital Card or a redundant storage device to prevent data loss, and exits the current processing flow.

[0037] Furthermore, in specific implementation, in the above-mentioned log collection method provided by the embodiment of the present invention, step S103 parses the key fields of the original log data and converts them into structured log data, which may specifically include: performing word segmentation on the original log data, combining regular expression matching and semantic analysis to extract key fields; converting the key fields into structured log data.

[0038] In implementation, the management controller reads the original logs from the shared memory, and can parse the key fields (such as error codes, timestamps, hardware locations, etc.) using regular expressions and tokenization techniques, and convert them into structured data with a set format. The set format can be a lightweight data interchange format such as JSON (JavaScript Object Notation). This replacement of manual parsing can significantly improve the efficiency of log analysis and reduce the operation and maintenance costs.

[0039] Furthermore, in specific implementation, in the above-mentioned log collection method provided by the embodiments of the present invention, after performing step S103 to convert to structured log data and before performing step S104 to encrypt the structured log data, it may further include: classifying the structured log data into critical logs and non-critical logs according to the characteristics of the structured log data.

[0040] In the above steps, classifying the structured log data into critical logs and non-critical logs according to the characteristics of the structured log data may specifically include: constructing a log classification model; collecting log key information and performing oversampling using the Synthetic Minority Over-sampling Technique (SMOTE) to obtain a sample set; training the log classification model with the sample set; during the training process, optimizing the loss function using Focal Loss or weighted cross-entropy, and annotating the samples with a confidence level lower than the set confidence level value output by the log classification model, and combining the output of the log classification model with predefined rules to achieve multi-level priority processing, and locating the root cause through correlation analysis; when the log traffic exceeds the set threshold, triggering a fuse mechanism; classifying the structured log data using the trained log classification model according to the characteristics of the structured log data; when classifying to critical logs, starting an alarm process; when classifying to non-critical logs, compressing the non-critical logs using a compression algorithm.

[0041] In implementation, the above-mentioned log classification model can be a lightweight XGBoost (eXtreme Gradient Boosting) model, and the inference speed can be specifically optimized through ARM Neon (Single Instruction Multiple Data under the ARM architecture) instructions.

[0042] The present invention can use a lightweight XGBoost model to dynamically classify into critical logs and non-critical logs according to log characteristics. Critical logs may include P0-level faults (i.e., faults with the highest priority and the most serious impact). When classifying to critical logs, the alarm process can be started in a timely manner. When classifying to non-critical logs, they can be compressed, and the compression ratio can be set to be greater than 80%.

[0043] It should be noted that the present invention dynamically filters key logs based on machine learning, and the main implementation methods are as follows: The design of the machine learning model mainly involves model selection. One is real-time classification, which can predict error trends, detect unknown errors, etc. In the sample sampling part, the synthetic minority over-sampling technique (SMOTE) is used to oversample the key logs (positive samples), and focal loss or weighted cross-entropy is adopted to improve the recognition of minority classes. For samples with low confidence, the operation and maintenance personnel are requested to label them. The second is priority grading, which adopts a multi-level processing strategy, combines the model output with predefined rules, and correlates the root causes. The third is system optimization, including the conversion of FP32 (32-bit single-precision floating-point model) to INT8 (8-bit integer model), maintaining an accuracy of more than 95%, compiling the preprocessing and inference logic into eBPF (extended Berkeley Packet Filter) programs or kernel modules, and using a graphics processing unit (GPU) or neural processing unit (NPU) for inference for models such as XGBoost. The fourth is resource isolation, including restricting the processor quota and memory bandwidth of the log processing thread, and setting a fuse mechanism: when the log peak exceeds the threshold, it degrades to basic rule filtering. The fifth is observability, using Local Interpretable Model-agnostic Explanations (LIME) or SHapley Additive exPlanations (SHAP) to explain individual prediction results.

[0044] Figure 4 It is a schematic diagram of the framework corresponding to the log collection method provided by the embodiment of the present invention. As Figure 4As shown, the original log data is parsed to convert unstructured log data into structured data. Valuable features are extracted from the parsed log data, and these features will be used as the input for subsequent inference of the log classification model. The pre-trained log classification model is used to infer the extracted features to predict whether there are anomalies in the log data. The results of the model inference are dynamically filtered to remove some possible false alarms and improve the detection accuracy. Based on the results of the model inference and dynamic filtering, it is determined whether the log data belongs to critical logs or abnormal events. The log data determined to be critical after judgment will be encrypted and stored to ensure data security and traceability. If it is determined to be an abnormal event, a real-time alarm will be triggered to notify relevant personnel for processing in a timely manner. Relevant personnel evaluate and give feedback on the processing results, comparing the actual situation with the system judgment. Online learning is carried out using the data of manual feedback (such as false alarm annotation) to continuously optimize the model and features. According to the results of online learning, the weights of the local model are updated to improve the accuracy of the model. The extracted features are optimized to better reflect the essential features of the log data.

[0045] Further, in specific implementation, in the above log collection method provided by the embodiments of the present invention, step S104 encrypts the structured log data in a pre-created secure isolation area, which may specifically include: dividing the running environment into a secure world and a normal world using a hardware-level secure isolation method; establishing memory isolation between the secure world and the normal world, and creating a secure isolation area through an address space controller; encrypting the critical logs in the structured log data in the secure isolation area and storing the encryption key in the secure isolation area.

[0046] In implementation, the encryption processing method may be symmetric encryption, such as AES-256, AES-GCM, etc.; among them, AES-256 is a variant of the Advanced Encryption Standard (AES), which uses a 256-bit key for data encryption; AES-GCM (Advanced Encryption Standard - Galois / Counter Mode) is an authenticated encryption mode. The encryption key can be injected through a secure boot process to prevent man-in-the-middle attacks. In practical applications, the present invention can encrypt only critical logs or encrypt both critical logs and non-critical logs, which can be adjusted according to the actual situation.

[0047] It should be noted that the present invention uses ARM TrustZone technology to encrypt the log transmission path, and the specific design is as follows: First, the TrustZone secure partition design: world division, including the secure world and the normal world; memory isolation, dividing the secure / non-secure memory areas through the TZASC (TrustZone Address Space Controller), preventing the normal world from accessing the encrypted log buffer; hardware encryption acceleration, providing hardware-level Advanced Encryption Standard (AES) / Secure Hash Algorithm (SHA) acceleration to reduce encryption latency. Second, the end-to-end log encryption process: log generation and capture, such as kernel logs, can redirect critical logs to the secure world buffer by modifying printk (a function used to output messages to the kernel log buffer), and user-state logs can be passed to the secure isolation area through the ioctl (the core interface parsing for device control) interface or shared memory (configured as secure memory). The encryption process in the secure world includes key management and signature. Secure transmission and storage, including adding a monotonically increasing sequence number to each log, verifying the continuity of the sequence number at the receiving end, rejecting data with old sequence numbers, disabling normal world JTAG (Joint Test Action Group) access, and only allowing the authorized secure world to decrypt the logs through the Secure Debug channel. Using TrustZone I / O virtualization (SMMU configuration) to ensure that the DMA transmission path cannot be tampered with by the normal world. Periodically rotating the session key (such as once an hour) and immediately erasing the old key, etc.

[0048] Figure 5 The specific flowchart of the log collection method provided by the embodiment of the present invention is as Figure 5As shown in the figure, first, the operating system kernel creates a shared memory area for data interaction with the management controller to facilitate data transfer and sharing. The kernel starts to monitor and capture various events generated by hardware devices in the system, and detects the load condition of the processor to understand the operating state of the system. If the processor load is too high, in order to reduce the system burden, measures are taken to reduce the sampling frequency of hardware events. If the processor load is normal, the operation continues according to the current sampling frequency and processing strategy. The captured and processed hardware event data is written into the previously created shared memory area for the management controller to read, triggering a management controller interrupt. After the management controller reads the log data in the shared memory, it detects its own heartbeat signal. If normal, it parses and tokenizes the read data, converts the parsed data into JSON format to make it easier to process and store. Next, a log classification model is used to classify the logs, identify critical logs and non-critical logs, compress and store the corresponding logs, and issue real-time warnings for abnormal situations. According to the information fed back by humans, the log classification model is updated online to improve the processing accuracy. Then the logs are encrypted and uploaded to the cloud through a secure network channel, and the cloud sends a confirmation message to the management controller. If the management controller heartbeat detection is abnormal, in order to prevent data loss, the data is urgently transferred and stored in a secure digital card or other redundant storage devices, and the current processing flow is exited.

[0049] Through the collaborative work of the operating system kernel and the management controller, the whole process realizes the monitoring, processing of hardware events and the management of logs, and at the same time has the functions of real-time warning and secure data storage, and continuously optimizes the processing process through human feedback and online model updates.

[0050] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0051] The embodiments of the present application also provide a log collection device. Figure 6 It is a schematic structural diagram of the log collection device provided by the embodiment of the present invention. From the perspective of functional modules, as Figure 6 shown, this device is for a management controller and includes: An instruction receiving module 10, configured to receive a target instruction from a driver module embedded in the operating system kernel; the driver module is configured to capture hardware events and write the corresponding log data into a pre-established shared memory area; A data reading module 11, configured to read original log data from the shared memory area in response to the target instruction; A log processing module 12 is configured to parse key fields from the original log data and convert it into structured log data; An encrypted transmission module 13 is configured to encrypt the structured log data in a pre-created secure isolation area and upload the processed log data to the cloud, so that the cloud returns a storage confirmation signal.

[0052] In the above log collection device provided by the embodiments of the present invention, through the interaction of the above four modules, efficient collection and processing of hardware event log data can be achieved based on in-band communication, reducing redundant overhead in the data processing process, improving the efficiency of hardware status monitoring and fault diagnosis, thereby achieving real-time monitoring of the hardware status and rapid fault diagnosis in a low-overhead and highly collaborative manner, ensuring the stable operation of the system; and it can achieve double savings in hardware and operation and maintenance, eliminating dedicated network cards and independent management networks, reducing redundant storage requirements, and saving costs.

[0053] Since the embodiments of the log collection device part correspond to the embodiments of the log collection method part, the descriptions of the features in the corresponding embodiments of the log collection device can refer to the relevant descriptions of the corresponding embodiments of the log collection method, which will not be elaborated here one by one. And it has the same beneficial effects as the above-mentioned log collection method.

[0054] Further, in specific implementation, in the above log collection device provided by the embodiments of the present invention, it may further include: an interrupt module, configured to monitor the level state of the input / output pins, and when the input / output pins are pulled low by the driving module while writing log data, trigger an internal interrupt signal and execute an interrupt; or, monitor the value of a set register, and when the driving module writes a preset value to the set register while writing log data, trigger an internal interrupt signal and execute an interrupt.

[0055] Further, in specific implementation, in the above log collection device provided by the embodiments of the present invention, it may further include: a heartbeat detection module, configured to detect the heartbeat signal of the operating system kernel to determine whether the operating system is online; if it is online, execute the processing process of the log processing module 12; if it is not online, transfer the original log data to a local secure digital card or a redundant storage device.

[0056] Further, in specific implementation, in the above log collection device provided by the embodiments of the present invention, the log processing module 12 may specifically be configured to perform word segmentation processing on the original log data, extract key fields by combining regular expression matching and semantic analysis; and convert the key fields into structured log data with a set format.

[0057] Further, in specific implementation, in the above-mentioned log collection device provided by the embodiments of the present invention, it may further include: a log classification module, configured to classify the structured log data into critical logs and non-critical logs according to the characteristics of the structured log data.

[0058] In implementation, the log classification module may specifically be configured to build a log classification model; collect log key information, and perform oversampling using the Synthetic Minority Over-sampling Technique (SMOTE) to obtain a sample set; use the sample set to train the log classification model; during the training process, optimize the loss function using focal loss or weighted cross-entropy, and label the samples with a confidence level lower than the set confidence level value output by the log classification model, and combine the output of the log classification model with predefined rules to achieve multi-level priority processing, and locate the root cause through correlation analysis; when the log traffic exceeds the set threshold, trigger a fusing mechanism; classify the structured log data using the trained log classification model according to the characteristics of the structured log data; when classifying critical logs, initiate an alarm process; when classifying non-critical logs, perform compression processing on the non-critical logs using a compression algorithm.

[0059] Further, in specific implementation, in the above-mentioned log collection device provided by the embodiments of the present invention, the encryption transmission module 13 may specifically be configured to create a secure isolation area; encrypt the critical logs in the structured log data in the secure isolation area, and store the key in the secure isolation area.

[0060] The embodiments of the present application further provide an electronic device, including a server and a management controller; wherein, a driver module is embedded in the operating system kernel of the server; the driver module is configured to capture hardware events and write the corresponding log data into a pre-established shared memory area, and send a target instruction to the management controller; the management controller is configured to receive and respond to the target instruction, read the original log data from the shared memory area; parse the key fields of the original log data and convert it into structured log data; encrypt the structured log data in a pre-created secure isolation area, and upload the processed log data to the cloud so that the cloud returns a storage confirmation signal.

[0061] The embodiments of the present invention further provide a computer-readable storage medium, in which a computer program is stored, and wherein the computer program is configured to execute the steps in any of the above-mentioned embodiments of the log collection method when running.

[0062] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0063] An embodiment of the present invention also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the log collection method are implemented.

[0064] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the log collection method are implemented.

[0065] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0066] The above has introduced in detail a log collection method, device, equipment, and medium provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A log collection method, characterized in that, For a management controller, including: Receiving a target instruction from a driver module embedded in the operating system kernel; the driver module is used to capture hardware events and write corresponding log data into a pre-established shared memory area; In response to the target instruction, reading the original log data from the shared memory area; Parsing key fields of the original log data and converting it into structured log data; Performing encryption processing on the structured log data in a pre-created secure isolation area and uploading the processed log data to the cloud so that the cloud returns a storage confirmation signal.

2. The log collection method according to claim 1, wherein Before receiving a target instruction from a driver module embedded in the operating system kernel, it further includes: Monitoring the level status of input / output pins, and when the input / output pins are pulled low by the driver module while writing log data, triggering an internal interrupt signal and executing an interrupt; Or, monitoring the value of a set register, and when a preset value is written to the set register by the driver module while writing log data, triggering an internal interrupt signal and executing an interrupt.

3. The log collection method according to claim 1, wherein Before parsing key fields of the original log data, it further includes: Detecting its own heartbeat signal and determining whether the heartbeat signal is within a set heartbeat range; If so, parsing key fields of the original log data; If not, transferring the original log data to a local secure digital card or a redundant storage device.

4. The log collection method according to claim 1, wherein Parsing key fields of the original log data and converting it into structured log data, including: Performing word segmentation on the original log data, combining regular expression matching and semantic analysis to extract key fields; Converting the key fields into structured log data with a set format.

5. The log collection method according to claim 1, characterized in that, After converting to structured log data and before performing encryption processing on the structured log data, it further includes: Classifying the structured log data into critical logs and non-critical logs according to the characteristics of the structured log data; Performing encryption processing on the structured log data, including: Performing encryption processing on the critical logs.

6. The log collection method according to claim 5, characterized in that, Classifying the structured log data into critical logs and non-critical logs according to the characteristics of the structured log data, including: Constructing a log classification model; Collecting log key information and using the synthetic minority over-sampling technique to perform over-sampling to obtain a sample set; Training the log classification model with the sample set; during the training process, optimizing the loss function using focal loss or weighted cross-entropy, annotating samples with a confidence level lower than a set confidence level value output by the log classification model, and combining the output of the log classification model with predefined rules to achieve multi-level priority processing and locating the root cause through association analysis; when the log traffic exceeds a set threshold, triggering a fusing mechanism; Classifying the structured log data according to the characteristics of the structured log data using the trained log classification model; When classifying to critical logs, starting an alarm process; When classifying to non-critical logs, performing compression processing on the non-critical logs using a compression algorithm.

7. The log collection method according to claim 1, wherein Performing encryption processing on the structured log data in a pre-created secure isolation area, including: Dividing the operating environment into a secure world and a normal world using a hardware-level secure isolation method; Establishing memory isolation between the secure world and the normal world, and creating a secure isolation area through an address space controller; Performing encryption processing on the critical logs in the structured log data in the secure isolation area, and storing the key in the secure isolation area.

8. A log collection device, characterized in that, For a management controller, including: An instruction receiving module, configured to receive a target instruction from a driver module embedded in the operating system kernel of the server; the driver module is configured to capture hardware events and write the corresponding log data into a pre-established shared memory area; A data reading module, configured to read the original log data from the shared memory area in response to the target instruction; A log processing module, configured to parse the keyword fields of the original log data and convert them into structured log data; An encryption transmission module, configured to perform encryption processing on the structured log data in a pre-created secure isolation area, and upload the processed log data to the cloud to enable the cloud to return a storage confirmation signal.

9. An electronic device, characterized in that, Including: A server and a management controller; wherein, a driver module is embedded in the operating system kernel of the server; The driver module is configured to capture hardware events and write the corresponding log data into a pre-established shared memory area, and send a target instruction to the management controller; The management controller is configured to receive and respond to the target instruction, read the original log data from the shared memory area; parse the keyword fields of the original log data and convert them into structured log data; perform encryption processing on the structured log data in a pre-created secure isolation area, and upload the processed log data to the cloud to enable the cloud to return a storage confirmation signal.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the log collection method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method for finishing software log by utilizing kernel module and application module under linux system

    CN103176891A

  • Log processing method and device, electronic equipment and storage medium

    CN116361106A

  • Log file storage method and device, storage medium and electronic device

    CN116991331A

  • Guest Operating System Buffer and Log Accesses by an Input-Output Memory Management Unit

    US20200387326A1

Cited By

  • Log data acquisition method and device, electronic equipment and storage medium

    CN121434166A