A lightweight distributed log collection method

CN122838362APending Publication Date: 2026-09-29SHANSHU TECH (BEIJING) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611024262.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0010]本发明的目的在于提出一种方案来解决低资源消耗、低代码侵入、灵活部署及可靠传输的问题,以适配现代分布式系统的复杂需求

Benefits of technology

[0033]一、通过轻量级标准化日志工具与边车组件解耦业务逻辑,显著降低代码侵入性,避免业务系统性能损耗;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838362A_ABST
    Figure CN122838362A_ABST
Patent Text Reader

Abstract

The application discloses a kind of lightweight distributed log collection methods, comprising: S1, the log data generated in the running process of business system is obtained, the log data is converted into the log file of pre-set format by the preset lightweight standardization log tool for standardization processing of log format, and output to local file system;S2, based on the sidecar component deployed in the host of business system, the write event about the log file in the local file system is monitored by inter-process communication mechanism, and when detecting the write event, the log data corresponding to the log file is obtained by asynchronous reading;S3, the log data read based on the sidecar component is packaged as transmission message, and the transmission message is transmitted to central log storage analysis platform for storage by message middleware module.The application realizes efficient and reliable log transmission by low-invasion lightweight deployment, and has strong adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer software processing technology, and in particular to a lightweight distributed log collection method. Background Technology

[0002] As software architecture evolves from monolithic applications to microservices and containerization (such as Docker and Kubernetes), a complete business system typically consists of dozens or even hundreds of independently deployed service instances. These service instances may run on different servers, containers, or virtual machines, be developed using different programming languages ​​(such as Java, Python, Go, etc.), and generate massive amounts of runtime logs. These logs are crucial for system monitoring, troubleshooting, business analysis, and security auditing. Therefore, how to efficiently, accurately, and uniformly collect these logs scattered across various nodes and transmit them to a centralized storage and analysis platform has become a key technical challenge in modern software operations and maintenance.

[0003] Currently, there are various log collection solutions in the industry:

[0004] One approach is direct code transmission. This method embeds log sending logic directly into the business logic code, sending logs to a centralized log server in real time via protocols such as HTTP and TCP. This approach requires no additional components and relies on the business system's own capabilities for log transmission.

[0005] The second option is a centralized proxy solution. This solution deploys an independent log proxy (such as Logstash) on the server, which is responsible for collecting local logs, parsing, filtering, and aggregating them before forwarding them to the target storage system. The proxy offers comprehensive functionality, supports complex processing logic, and is suitable for diverse log scenarios.

[0006] However, the code direct transfer solution is highly intrusive, coupled with business logic, and affects performance and reliability.

[0007] Centralized proxy solutions rely on heavy-duty proxies, consume a lot of resources, compete with business operations for resources, and are complex to operate and maintain.

[0008] Furthermore, the above solutions lack flexibility and reliability, making it difficult to meet the needs of cross-language and cross-platform microservice architectures. Especially in containerized or edge computing scenarios, resource constraints and dynamic scaling characteristics exacerbate the above contradictions, making it impossible to simultaneously meet the core requirements of low intrusion, lightweight, high reliability, and easy deployment.

[0009] Therefore, a new log collection solution is urgently needed to solve the above problems. Summary of the Invention

[0010] The purpose of this invention is to propose a solution to address the issues of low resource consumption, low code intrusion, flexible deployment, and reliable transmission, so as to adapt to the complex requirements of modern distributed systems.

[0011] To achieve this objective, the present invention adopts the following technical solution:

[0012] A lightweight distributed log collection method includes the following steps:

[0013] S1. Obtain log data generated during the operation of the business system, convert the log data into a log file of a preset format using a preset lightweight standardized log tool for standardizing log format processing, and output it to the local file system;

[0014] S2. Based on the sidecar component deployed on the host machine where the business system is located, the write events of the log file in the local file system are monitored through the inter-process communication mechanism, and when the write event is detected, the log data corresponding to the log file is obtained by asynchronous reading;

[0015] S3. Based on the sidecar component, the read log data is encapsulated into a transmission message, and the transmission message is transmitted to the central log storage and analysis platform for storage through the message middleware module.

[0016] Furthermore, the preset lightweight standardized logging tool in step S1 is a log formatting interface decoupled from the business system language, used to uniformly convert log data generated by different programming languages ​​into the log file of the preset format;

[0017] The preset format includes JSON format or structured text format, and the preset format contains at least one of the following information: log source, timestamp, log level, and log content.

[0018] Furthermore, the sidecar component in step S2 is deployed in an independent process on the host machine where the business system is located, and the sidecar component is isolated from the business system through an inter-process communication mechanism, which includes at least one of a file monitoring mechanism or a system call listening mechanism.

[0019] Furthermore, the message middleware module mentioned in step S3 specifically comprises:

[0020] Reuse the publish-subscribe mechanism of the Redis cache database already deployed in the business system; or,

[0021] A pluggable, highly reliable message queue, including at least one of Kafka, RabbitMQ, or RocketMQ.

[0022] Furthermore, when the message middleware module reuses the Redis cache database, the transmitted messages are persistently transmitted through Redis's Stream data structure, or transmitted in real time through Redis's publish-subscribe pattern.

[0023] Furthermore, the transmission message includes a message header and a message body. The message header contains a unique identifier of the log data, source host information, and transmission timestamp. The message body is the encrypted or compressed log data.

[0024] Furthermore, in step S2, after the step of obtaining the log data corresponding to the log file by asynchronous reading based on the sidecar component, at least one of the following operations is also included:

[0025] Perform a validity check on the log data;

[0026] The log data is parsed according to preset rules to extract key fields used to encapsulate the transmitted message.

[0027] Furthermore, in step S1, the step of the preset lightweight standardized log tool converting the log data into a log file of a preset format and outputting it to the local file system further includes:

[0028] A file rotation strategy is adopted to trigger the rolling creation and cleanup of the log files based on time or file size.

[0029] Furthermore, step S2 also includes:

[0030] Monitor the resource utilization rate of the sidecar component. When the resource utilization rate exceeds the preset threshold, trigger dynamic adjustment of the log collection frequency or enable the log cache queue.

[0031] Furthermore, the lightweight distributed log collection method is applied in a containerized environment or microservice architecture, where the business system and the sidecar component are deployed in different containers on the same host machine and share the mounted volume of the local file system.

[0032] The technical solution provided by this invention may include the following beneficial effects:

[0033] 1. By decoupling business logic from the sidecar component using lightweight, standardized logging tools, code invasiveness is significantly reduced, and performance degradation of the business system is avoided.

[0034] Second, reuse existing Redis resources or pluggable message queues in business systems to reduce deployment costs and resource consumption, and adapt to resource-constrained scenarios.

[0035] Third, based on inter-process communication and asynchronous reading mechanisms, physical isolation between log collection and business processing is achieved, thereby improving system reliability and throughput;

[0036] IV. Supports containerization and microservice architecture, with high flexibility and cross-environment adaptability;

[0037] Fifth, it simplifies operation and maintenance complexity, eliminates the need for additional agents or heavy components, and enables lightweight deployment and efficient log transmission, meeting the low-cost and high-reliability requirements of distributed systems for log collection. Attached Figure Description

[0038] Figure 1 This is a flowchart of the steps of a lightweight distributed log collection method provided in an embodiment of the present invention. Detailed Implementation

[0039] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that the specific embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the present invention.

[0040] Figure 1 This is a flowchart illustrating the steps of a lightweight distributed log collection method provided in this embodiment of the invention. The lightweight distributed log collection method includes the following steps:

[0041] S1. Obtain log data generated during the operation of the business system, convert the log data into a log file of a preset format using a preset lightweight standardized log tool for standardizing log format processing, and output it to the local file system.

[0042] Specifically, in this embodiment of the invention, the business system is the source of log data, and its specific form can be a microservice instance based on a microservice architecture, a web server application, a mobile application backend, or an algorithm processing program, etc. During operation, the business system continuously generates raw log data reflecting its operating status, user operation records, or error information.

[0043] The aforementioned lightweight standardized logging tool is specifically a software development kit (SDK) integrated into the business system. This SDK refers to a lightweight library or code package integrated into the business system code. It is typically introduced into the business system development project as a dependency package (such as a Java JAR file, a Python Wheel package, etc.). It encapsulates the underlying logic for log formatting and standardized output. This lightweight standardized logging tool provides a standardized logging interface decoupled from programming languages, supporting multiple mainstream development languages ​​such as Java and Python. The business system records logs by calling the chained functions provided by the SDK. It should be noted that this SDK is only responsible for log generation and local disk persistence, and does not contain complex network transmission or remote call logic, thus ensuring extremely low resource consumption by the business system.

[0044] The aforementioned lightweight standardized logging tool transforms the diverse raw log data generated by business systems into structured logs containing standard fields such as application identifier (appCode), module identifier (module), business identifier (business), trace ID (trace), and log content (data). This standardized format lays the foundation for subsequent automated parsing and analysis.

[0045] The converted logs are written to the local file system of the host machine where the business system is located in a preset format.

[0046] Preferably, the preset format is JSON. This format has good readability and machine parsability, facilitating subsequent extraction, conversion, and loading processes.

[0047] S2. Based on the sidecar component deployed on the host machine where the business system is located, the write events of the log file in the local file system are monitored through the inter-process communication mechanism, and when the write event is detected, the log data corresponding to the log file is obtained by asynchronous reading.

[0048] In this embodiment of the invention, the sidecar component is a daemon process or container that runs independently of the business system. During deployment, the sidecar component and the business system are deployed on the same host machine, and share the local file system directory of the business system's output logs through a mounted volume.

[0049] The host machine provides the hardware resources and operating system environment required for the log collection method to run. In one specific application scenario, the host machine can be a physical server; in another preferred application scenario (such as Docker containerized deployment), the host machine can be a virtualized computing node, where the business system and the sidecar component are scheduled and run as two independent containers on the same host machine. The host machine provides the basic environment for data interaction between the business system and the sidecar component.

[0050] The local file system serves as the medium for data interaction between the business system and the sidecar component. Specifically, the local file system is a storage directory mounted on the host disk and shared by the business system and the sidecar component. The business system writes standardized log data to a specified path within this local file system in the form of files.

[0051] The sidecar component in step S2 is deployed in an independent process on the host machine where the business system resides. The sidecar component is isolated from the business system through an inter-process communication mechanism, which includes at least one of a file monitoring mechanism or a system call listening mechanism. The sidecar component utilizes the file monitoring mechanism provided by the operating system (such as Linux's inotify) to monitor change events of log files in a specified directory in real time. When a file is added or modified (i.e., a write event) is detected, the sidecar component immediately reads the newly added log content in an asynchronous, non-blocking manner. This incremental collection method avoids repeated readings, ensuring the real-time performance and efficiency of the collection.

[0052] The sidecar component interacts with the business system through an inter-process communication mechanism using a shared file system, achieving physical decoupling. During implementation, even if the sidecar component fails, it will not affect the normal operation of the business system, thus improving the overall reliability of the system.

[0053] S3. Based on the sidecar component, the read log data is encapsulated into a transmission message, and the transmission message is transmitted to the central log storage and analysis platform for storage through the message middleware module.

[0054] In this embodiment of the invention, after reading the raw log data, the sidecar component encapsulates it into a standard transmission message. This message typically includes a message header and a message body. The message header may contain metadata such as the log's unique identifier, the source host IP, and a timestamp, while the message body is the log content itself, which can be compressed or encrypted as needed. Subsequently, the sidecar component, acting as a producer, sends the encapsulated message out through a message middleware module. The central log storage and analysis platform, acting as a consumer, subscribes to the corresponding message topics, receives and processes these log messages, and ultimately persists them to a central log storage and analysis platform composed of databases such as Elasticsearch, HBase, or MySQL for subsequent querying, analysis, and monitoring alerts.

[0055] Furthermore, the message middleware module mentioned in step S3 specifically comprises:

[0056] Reuse the publish-subscribe mechanism of the Redis cache database already deployed in the business system; or,

[0057] A pluggable, highly reliable message queue, including at least one of Kafka, RabbitMQ, or RocketMQ.

[0058] As a preferred implementation, this embodiment of the invention fully utilizes the Redis component typically deployed in business systems. The sidecar component acts as a publisher, publishing log messages to a designated Redis channel. The receiving service of the central log storage and analysis platform acts as a subscriber, listening to the channel to retrieve logs. This approach eliminates the need for heavyweight message queues like Kafka, significantly reducing system complexity and resource overhead, making it particularly suitable for scenarios where log reliability requirements are not extremely stringent but a lightweight approach is desired.

[0059] As another embodiment, for scenarios with large log data volumes and high reliability requirements, the message middleware module can be configured as a specialized message queue such as Kafka, RabbitMQ, or RocketMQ. The sidecar component sends logs to a designated topic, where they are consumed by the backend consumer group. This pluggable design allows the method proposed in this embodiment to flexibly adapt to different business scenarios and infrastructure environments.

[0060] When the message middleware module reuses the Redis cache database, the transmitted messages are persistently transmitted through Redis's streaming data structure, or transmitted in real time through Redis's publish-subscribe pattern.

[0061] During implementation, when using Redis, two modes can be selected based on requirements. If extreme real-time performance is desired and a small amount of log loss is acceptable, the publish-subscribe mode can be used. If log loss is guaranteed, the Stream data structure introduced in Redis 5.0 can be used. The Stream data structure provides message persistence and consumer group functionality, enabling message acknowledgment and backtracking mechanisms similar to Kafka, providing higher reliability while maintaining lightweight operation.

[0062] The transmission message includes a message header and a message body. The message header contains a unique identifier of the log data, source host information, and transmission timestamp. The message body is the encrypted or compressed log data.

[0063] Specifically, the message header carries routing and metadata information, such as message_id, host_ip, and ingest_timestamp, facilitating backend tracing and auditing. The message body carries the actual log content. To save network bandwidth and storage space, the message body can be compressed using compression algorithms or encrypted before sending to ensure data security.

[0064] Furthermore, in step S2, after the step of obtaining the log data corresponding to the log file by asynchronous reading based on the sidecar component, at least one of the following operations is also included:

[0065] The log data is validated for legality. After reading the logs, the sidecar component can perform simple validation, such as checking whether the log lines are empty or whether the JSON format is complete, to filter out obviously invalid dirty data and reduce the processing pressure on the backend.

[0066] The log data is parsed according to preset rules to extract key fields for encapsulation into the transmission message. The sidecar component supports configurable parsing rules, which can extract key information from the raw logs according to the rules, such as extracting fields like status_code and request_time from Nginx logs, and putting these fields into the transmission message as independent structured data, thus realizing log preprocessing and structuring.

[0067] Furthermore, in step S1, the step of the preset lightweight standardized log tool converting the log data into a log file of a preset format and outputting it to the local file system also includes:

[0068] A file rotation strategy is adopted to trigger the rolling creation and cleanup of the log files based on time or file size.

[0069] To prevent individual log files from becoming too large, the lightweight, standardized logging tool configures a log rotation strategy. For example, when a file exceeds 100MB or spans a day, a new log file is automatically created, and historical log files are archived or cleaned up. The sidecar component can automatically identify and monitor newly generated log files, ensuring the continuity of log collection.

[0070] Furthermore, step S2 also includes:

[0071] Monitor the resource utilization rate of the sidecar component. When the resource utilization rate exceeds the preset threshold, trigger dynamic adjustment of the log collection frequency or enable the log cache queue.

[0072] The sidecar component includes a self-monitoring module that monitors its own CPU and memory usage in real time. When resource usage exceeds a preset threshold (e.g., CPU utilization greater than 5%), the sidecar component automatically triggers protection mechanisms to prevent impact on the business systems on the host machine. These mechanisms may include reducing the frequency of file monitoring, enabling an internal cache queue to buffer logs, or temporarily discarding low-priority debug logs, thereby achieving self-protection and ensuring the stability of the business systems.

[0073] Furthermore, the lightweight distributed log collection method is applied in a containerized environment or microservice architecture, where the business system and the sidecar component are deployed in different containers on the same host machine and share the mount volume of the local file system.

[0074] In container orchestration platforms such as Kubernetes, this embodiment of the invention is deployed in Sidecar mode. Within the same Pod, two containers are defined: one running the business application, and the other running the sidecar component. By defining a shared volume of type emptyDir or hostPath and mounting it to the same path in both containers, log file sharing is achieved. The business container writes logs to the mounted directory, and the sidecar container reads logs from the same directory, perfectly realizing lightweight, decoupled log collection in a microservice architecture.

[0075] The technical solution provided by this invention may include the following beneficial effects:

[0076] 1. By decoupling business logic from the sidecar component using lightweight, standardized logging tools, code invasiveness is significantly reduced, and performance degradation of the business system is avoided.

[0077] Second, reuse existing Redis resources or pluggable message queues in business systems to reduce deployment costs and resource consumption, and adapt to resource-constrained scenarios.

[0078] Third, based on inter-process communication and asynchronous reading mechanisms, physical isolation between log collection and business processing is achieved, thereby improving system reliability and throughput;

[0079] IV. Supports containerization and microservice architecture, with high flexibility and cross-environment adaptability;

[0080] Fifth, it simplifies operation and maintenance complexity, eliminates the need for additional agents or heavy components, and enables lightweight deployment and efficient log transmission, meeting the low-cost and high-reliability requirements of distributed systems for log collection.

[0081] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0082] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0083] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a multimedia terminal device (which may be a mobile phone, computer, television receiver, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0084] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A lightweight distributed log collection method, characterized in that, Includes the following steps: S1. Obtain log data generated by the business system during operation, convert the log data into a log file of a preset format using a preset lightweight standardized log tool for standardizing log format processing, and output it to the local file system; S2. Based on the sidecar component deployed on the host machine where the business system is located, the write events of the log file in the local file system are monitored through the inter-process communication mechanism, and when the write event is detected, the log data corresponding to the log file is obtained by asynchronous reading; S3. Based on the sidecar component, the read log data is encapsulated into a transmission message, and the transmission message is transmitted to the central log storage and analysis platform for storage through the message middleware module.

2. The lightweight distributed log collection method according to claim 1, characterized in that, The preset lightweight standardized logging tool in step S1 is a log formatting interface decoupled from the business system language, used to uniformly convert log data generated by different programming languages ​​into the log file of the preset format; The preset format includes JSON format or structured text format, and the preset format contains at least one of the following information: log source, timestamp, log level, and log content.

3. The lightweight distributed log collection method according to claim 1, characterized in that, The sidecar component in step S2 is deployed in an independent process on the host machine where the business system is located, and the sidecar component is isolated from the business system through an inter-process communication mechanism, which includes at least one of a file monitoring mechanism or a system call listening mechanism.

4. The lightweight distributed log collection method according to claim 1, characterized in that, The message middleware module mentioned in step S3 specifically refers to: Reuse the publish-subscribe mechanism of the Redis cache database already deployed in the business system; or, A pluggable, highly reliable message queue, including at least one of Kafka, RabbitMQ, or RocketMQ.

5. The lightweight distributed log collection method according to claim 4, characterized in that, When the message middleware module reuses the Redis cache database, the transmitted messages are persistently transmitted through Redis's Stream data structure, or transmitted in real time through Redis's publish-subscribe pattern.

6. The lightweight distributed log collection method according to claim 5, characterized in that, The transmission message includes a message header and a message body. The message header contains a unique identifier of the log data, source host information, and transmission timestamp. The message body is the encrypted or compressed log data.

7. The lightweight distributed log collection method according to claim 1, characterized in that, In step S2, after the step of obtaining the log data corresponding to the log file by asynchronous reading based on the sidecar component, at least one of the following operations is also included: Perform a validity check on the log data; The log data is parsed according to preset rules to extract key fields used to encapsulate the transmitted message.

8. The lightweight distributed log collection method according to claim 1, characterized in that, In step S1, the step of the preset lightweight standardized log tool converting the log data into a log file of a preset format and outputting it to the local file system further includes: A file rotation strategy is adopted to trigger the rolling creation and cleanup of the log files based on time or file size.

9. The lightweight distributed log collection method according to claim 1, characterized in that, Step S2 also includes: Monitor the resource utilization rate of the sidecar component. When the resource utilization rate exceeds the preset threshold, trigger dynamic adjustment of the log collection frequency or enable the log cache queue.

10. The lightweight distributed log collection method according to claim 1, characterized in that, The lightweight distributed log collection method is applied in a containerized environment or microservice architecture, where the business system and the sidecar component are deployed in different containers on the same host machine and share the mount volume of the local file system.