Streaming method and system for processing massive syslogs
Through distributed streaming computing architecture and intelligent parsing technology, the real-time processing and analysis problems of massive syslog logs are solved, high-throughput, low-latency and high-availability log processing is achieved, and the accuracy and automation of log analysis are improved.
Patent Information
- Application Number
- CN202511166836.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional batch processing methods cannot meet real-time requirements. Existing stream processing solutions have performance bottlenecks, data loss, and low parsing accuracy when facing massive syslog log data. They are also difficult to cope with complex and diverse log formats and lack intelligent anomaly detection capabilities.
It adopts a distributed streaming computing architecture, uses Apache Kafka and Apache Flink to build a highly available data collection and processing engine, combines regular expressions and machine learning models for intelligent field extraction and anomaly detection, writes to the Elasticsearch cluster in real time and performs multi-dimensional indexing, supports custom rules and dynamic adjustment, and performs hierarchical storage and unified interface output.
It realizes real-time processing and analysis of massive syslog logs, improves data processing throughput and accuracy, reduces storage costs, provides convenient data access, and significantly improves the automation level of log processing and analysis efficiency.
Smart Images

Figure CN120763174A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and in particular to a streaming method and system for processing massive syslog logs. Background Art
[0002] As the scale of information systems continues to expand, the syslog log data generated by various network devices and different systems has exploded.
[0003] Traditional batch processing methods can no longer meet real-time requirements, while existing stream processing solutions face performance bottlenecks, data loss, and low parsing accuracy when dealing with massive amounts of data. Furthermore, in real-world applications, syslog log formats are complex and diverse, containing a large amount of unstructured information. Traditional parsing methods based on string matching struggle to cope with format changes and lack intelligent anomaly detection capabilities.
[0004] Therefore, a syslog log processing method based on streaming computing, distributed architecture and intelligent parsing is needed. Summary of the Invention
[0005] The purpose of the present invention is to provide a streaming method and system for processing massive syslog logs to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: a streaming method for processing massive syslog logs, comprising the following steps:
[0007] Log data collection and buffering: This layer collects various syslog log data from business systems, establishes a highly available data collection layer, uses Apache Kafka as a message queue, supports multiple log formats, ensures reliable data transmission and storage through a multi-level buffering mechanism, and supports data playback and fault-tolerant processing.
[0008] Streaming data preprocessing: Real-time cleaning and preprocessing of collected raw log data, including data deduplication, format standardization, and invalid data filtering. Each processing unit includes three components: validation, normalization, and filtering to ensure data quality.
[0009] Distributed streaming analysis: Apache Flink is used to build a distributed stream processing engine, achieving high-throughput processing through horizontal expansion. Flink's windowing mechanism and state management capabilities are used to implement complex event processing and session analysis, and support dynamic adjustment of parallelism to adapt to load changes.
[0010] Intelligent field extraction: Combines regular expressions and machine learning models to achieve intelligent field recognition and extraction. The pre-trained NLP model can automatically identify key information such as IP addresses, timestamps, user IDs, and error codes, and supports custom field extraction rules.
[0011] Subsequent processing and integration: The parsed structured data is written to the Elasticsearch cluster in real time to establish a multi-dimensional index structure; a streaming machine learning algorithm is used to detect abnormal patterns in the logs in real time. When anomalies are detected, alarm notifications are sent through multiple channels; historical log data is intelligently compressed and stored in layers; the above processing results are integrated to provide a unified query and analysis interface.
[0012] Preferably, the subsequent processing and integration steps are specifically as follows:
[0013] Real-time index building: Write the parsed structured data to the Elasticsearch cluster in real time to establish a multi-dimensional index structure, supporting time series indexing, full-text search indexing, and aggregate analysis indexing to optimize query performance;
[0014] Anomaly detection and alerting: Streaming machine learning algorithms are used to detect abnormal patterns in logs in real time, including statistical anomalies, pattern anomalies, and time series anomalies. Dynamic configuration of alert rules and threshold adjustment are supported. When an anomaly is detected, alert notifications are sent through multiple channels.
[0015] Data compression storage: Intelligently compresses and tiers historical log data, automatically adjusting storage strategies based on data access frequency. Hot data is stored in SSDs to ensure query performance, while warm and cold data are stored in different storage media to optimize costs.
[0016] Result output integration: Integrates the processing results of real-time index construction, anomaly detection and alarm, and data compression and storage, providing a unified REST API and GraphQL interface to support real-time query, historical analysis, and report generation.
[0017] Preferably, in the log data collection and buffering step, Apache Kafka is used as the message queue to establish a multi-level buffering mechanism, which can automatically adjust the buffering strategy according to data flow and system load to ensure that data is not lost in different scenarios. At the same time, it supports data playback function to reprocess data when needed, and has fault-tolerant processing capabilities to deal with system failures.
[0018] Preferably, in the intelligent field extraction step, the pre-trained NLP model is trained with a large amount of annotated syslog log data, and can accurately and automatically identify key information such as IP addresses, timestamps, user IDs, and error codes, and allows users to customize field extraction rules according to actual business needs. The model can be dynamically adjusted and optimized according to the new rules.
[0019] Preferably, in the result output integration step, the unified REST API and GraphQL interface provided have high concurrent processing capabilities, which can meet the diverse needs of different clients for real-time query, historical analysis and report generation. At the same time, the interface has good scalability and can be easily integrated into existing business systems.
[0020] A system for processing massive syslog logs in a streaming manner, including the following functional modules:
[0021] Log data collection and buffering module: This module receives massive syslog data streams through high-performance message queues, collects various syslog log data from business systems, establishes a highly available data collection layer, uses Apache Kafka as a message queue, and supports multiple log formats. It also builds a multi-level buffering mechanism to ensure reliable data transmission and storage, while also supporting data playback and fault-tolerant processing.
[0022] Streaming data preprocessing module: This module performs real-time cleaning and preprocessing on the collected raw log data, including data deduplication, format standardization, and invalid data filtering. Each processing unit includes three components: validation, normalization, and filtering to ensure data quality.
[0023] Distributed stream parsing module: This module uses Apache Kafka and Apache Flink to build a distributed stream processing architecture and Apache Flink to build a distributed stream processing engine. This module achieves high-throughput processing through horizontal scaling. Leveraging Flink's windowing mechanism and state management capabilities, it implements complex event processing and session analysis, and supports dynamic adjustment of parallelism to adapt to load changes, enabling parallel parsing and processing of logs.
[0024] Intelligent Field Extraction Module: This module combines regular expressions and machine learning models to implement intelligent field recognition and extraction. The pre-trained NLP model can automatically identify key information such as IP addresses, timestamps, user IDs, and error codes, while also supporting custom field extraction rules.
[0025] Subsequent processing and integration module: Writes the parsed structured data to the Elasticsearch cluster in real time and establishes a multi-dimensional index; uses streaming machine learning algorithms to detect abnormal patterns in logs in real time and trigger an alarm mechanism; intelligently compresses and hierarchically stores historical data; and integrates the above processing results to provide a unified query and analysis interface.
[0026] Preferably, the subsequent processing and integration module specifically includes:
[0027] Real-time index building submodule: Writes parsed structured data to the Elasticsearch cluster in real time, establishes a multi-dimensional index structure, and supports time series indexing, full-text search indexing, and aggregate analysis indexing to optimize query performance.
[0028] Anomaly Detection and Alerting Submodule: This module uses streaming machine learning algorithms to detect abnormal patterns in logs in real time, covering statistical anomalies, pattern anomalies, and time series anomalies. When an anomaly is detected, it sends alarm notifications through multiple channels and supports dynamic configuration of alarm rules and threshold adjustment.
[0029] Data compression and storage submodule: Intelligently compresses and tiers historical log data, automatically adjusts storage strategies based on data access frequency, stores hot data in SSDs to ensure query performance, and stores warm and cold data in different storage media to achieve cost optimization.
[0030] Result output integration sub-module: Integrates the processing results of real-time index construction, anomaly detection and alarm, and data compression and storage, provides a unified REST API and GraphQL interface, and supports real-time query, historical analysis, and report generation.
[0031] Preferably, in the log data collection buffer module, the multi-level buffer mechanism has the ability of adaptive adjustment, which can automatically optimize the buffer strategy according to the data flow size, system load and network conditions, ensuring that data loss can be effectively prevented in different scenarios, while ensuring the stability of the data playback function and the efficiency of fault-tolerant processing.
[0032] Preferably, in the intelligent field extraction module, the pre-trained NLP model is trained based on a large amount of annotated syslog log data, and the accuracy and recall rate of field recognition are improved by continuously optimizing the model parameters; for custom field extraction rules, the system provides a visual configuration interface to facilitate users to flexibly set them according to actual business needs, and can apply new rules to the field extraction process in real time.
[0033] Preferably, the unified REST API and GraphQL interface provided in the result output integration module has high concurrent processing capabilities and good scalability; the interface adopts a security authentication mechanism to ensure the security of data transmission and access; at the same time, the system provides complete interface documentation and sample code to facilitate developers to quickly integrate and use, meeting the diverse needs of different clients for real-time query, historical analysis and report generation.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] The streaming method and system for processing massive syslog logs proposed in the present invention adopt a distributed streaming computing architecture to achieve real-time processing and analysis of log data. The method has the characteristics of high throughput, low latency, and high availability, and can process millions of log data per second. Through intelligent field extraction and anomaly detection technology, the accuracy and practicality of log analysis are greatly improved. The tiered storage strategy effectively reduces storage costs, and the unified query interface provides users with a convenient data access method. The method can be widely used in scenarios such as enterprise-level log management, security monitoring, and operation and maintenance analysis, significantly improving the automation level and analysis efficiency of log processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0037] In order to clearly and completely describe the objectives and technical solutions of the present invention and make the advantages more clearly understood, the embodiments of the present invention are further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, not all of them, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0038] For example 1, please refer to Figure 1 The present invention provides a technical solution: a streaming method for processing massive syslog logs, comprising the following steps:
[0039] S1: Log Data Collection and Buffering: This layer collects various syslog log data from network devices or business systems to establish a highly available data collection layer. It uses Apache Kafka as a message queue, supports multiple log formats (RFC3164, RFC5424, etc.), and uses a multi-level buffering mechanism to ensure reliable data transmission and storage, including data playback and fault tolerance. It receives massive syslog data streams through a high-performance message queue and establishes a multi-level buffering mechanism to prevent data loss.
[0040] S2: Streaming Data Preprocessing: This process cleans and preprocesses collected raw log data in real time, including operations such as data deduplication, format standardization, and invalid data filtering. Each processing unit includes three components: validation, normalization, and filtering to ensure data quality. Raw logs are cleaned, formatted, and initially categorized in real time to remove invalid data.
[0041] S3: Distributed Stream Parsing: We use Apache Flink to build a distributed stream processing engine, achieving high-throughput processing through horizontal scalability. We leverage Flink's windowing mechanism and state management capabilities to implement complex event processing and session analysis, supporting dynamic adjustment of parallelism to accommodate load fluctuations. We use Apache Kafka and Apache Flink to build a distributed stream processing architecture, enabling parallel parsing and processing of logs.
[0042] S4: Intelligent Field Extraction: This combines regular expressions and machine learning models to achieve intelligent field recognition and extraction. The pre-trained NLP model automatically identifies key information such as IP addresses, timestamps, user IDs, and error codes, and supports custom field extraction rules. This automatically identifies and extracts key field information based on regular expressions and machine learning models.
[0043] S5: Real-time Index Building: Writes parsed structured data to the Elasticsearch cluster in real time, establishing a multi-dimensional index structure. Supports time series indexing, full-text search indexing, and aggregate analysis indexing to optimize query performance. Automatic data archiving and cleanup are achieved through index lifecycle management. Writes parsed structured data to the Elasticsearch cluster in real time, establishing a multi-dimensional index.
[0044] S6: Anomaly Detection and Alerting: Streaming machine learning algorithms are used to detect abnormal patterns in logs in real time, including statistical, pattern, and timing anomalies. When an anomaly is detected, alerts are sent via multiple channels (email, SMS, and webhooks), supporting dynamic configuration of alert rules and threshold adjustment. Streaming machine learning algorithms detect abnormal patterns in real time and trigger alerts.
[0045] S7: Data Compression Storage: Intelligently compresses and tiers historical log data, automatically adjusting storage policies based on data access frequency. Hot data is stored in SSDs to ensure query performance, while warm and cold data are stored in separate storage media for cost optimization. Intelligent compression and tiered storage of historical data optimizes storage costs.
[0046] S8: Result output integration: Integrates the processing results of S5, S6, and S7, provides a unified REST API and GraphQL interface, and supports real-time query, historical analysis, and report generation.
[0047] Example 2, based on Example 1, proposes a system for processing massive syslog logs in a streaming manner, including the following functional modules:
[0048] Log data collection and buffering module: Receives massive syslog data streams through high-performance message queues, collects various syslog log data from business systems, establishes a highly available data collection layer, uses Apache Kafka as a message queue, and supports multiple log formats; builds a multi-level buffering mechanism to ensure reliable data transmission and storage, while supporting data playback and fault-tolerant processing; the multi-level buffering mechanism has adaptive adjustment capabilities and can automatically optimize the buffering strategy based on data flow size, system load, and network conditions to ensure that data loss can be effectively prevented in different scenarios, while ensuring the stability of the data playback function and the efficiency of fault-tolerant processing.
[0049] Streaming data preprocessing module: This module performs real-time cleaning and preprocessing on the collected raw log data, including data deduplication, format standardization, and invalid data filtering. Each processing unit includes three components: validation, normalization, and filtering to ensure data quality.
[0050] Distributed stream parsing module: This module uses Apache Kafka and Apache Flink to build a distributed stream processing architecture and Apache Flink to build a distributed stream processing engine. This module achieves high-throughput processing through horizontal scaling. Leveraging Flink's windowing mechanism and state management capabilities, it implements complex event processing and session analysis, and supports dynamic adjustment of parallelism to adapt to load changes, enabling parallel parsing and processing of logs.
[0051] Intelligent field extraction module: Combines regular expressions and machine learning models to achieve intelligent field recognition and extraction. The pre-trained NLP model can automatically identify key information such as IP addresses, timestamps, user IDs, and error codes, and also supports custom field extraction rules. The pre-trained NLP model is trained based on a large amount of annotated syslog log data, and the accuracy and recall of field recognition are improved by continuously optimizing model parameters. For custom field extraction rules, the system provides a visual configuration interface to facilitate users to flexibly set them according to actual business needs, and can apply new rules to the field extraction process in real time.
[0052] Subsequent processing and integration module: Writes the parsed structured data to the Elasticsearch cluster in real time and establishes a multi-dimensional index; uses streaming machine learning algorithms to detect abnormal patterns in logs in real time and trigger an alarm mechanism; intelligently compresses and hierarchically stores historical data; and integrates the above processing results to provide a unified query and analysis interface. Specifically, it includes: real-time index construction submodule: writing the parsed structured data into the Elasticsearch cluster in real time, establishing a multi-dimensional index structure, supporting time series index, full-text search index and aggregate analysis index to optimize query performance; anomaly detection and alarm submodule: using streaming machine learning algorithms to detect abnormal patterns in logs in real time, covering statistical anomalies, pattern anomalies and timing anomalies; when an anomaly is detected, alarm notifications are sent through multiple channels, and dynamic configuration and threshold adjustment of alarm rules are supported; data compression and storage submodule: intelligently compressing and tiered storage of historical log data, automatically adjusting storage strategies based on data access frequency, storing hot data in SSD to ensure query performance, and storing warm data and cold data in different storage media to achieve cost optimization; result output integration submodule: integrating the processing results of real-time index construction, anomaly detection and alarm, and data compression storage, providing a unified REST API and GraphQL interface, supporting real-time query, historical analysis and report generation.
[0053] The unified REST API and GraphQL interface provided have high concurrent processing capabilities and good scalability; the interface adopts a security authentication mechanism to ensure the security of data transmission and access; at the same time, the system provides complete interface documentation and sample code to facilitate developers to quickly integrate and use, meeting the diverse needs of different clients for real-time query, historical analysis and report generation.
[0054] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A streaming method for processing massive syslog logs, characterized by: The following steps are involved: Log data collection and buffering: This layer collects various syslog log data from business systems, establishes a highly available data collection layer, uses Apache Kafka as a message queue, supports multiple log formats, ensures reliable data transmission and storage through a multi-level buffering mechanism, and supports data playback and fault-tolerant processing. Streaming data preprocessing: Real-time cleaning and preprocessing of collected raw log data, including data deduplication, format standardization, and invalid data filtering. Each processing unit includes three components: validation, normalization, and filtering to ensure data quality. Distributed streaming analysis: Apache Flink is used to build a distributed stream processing engine, achieving high-throughput processing through horizontal expansion. Flink's windowing mechanism and state management capabilities are used to implement complex event processing and session analysis, and support dynamic adjustment of parallelism to adapt to load changes. Intelligent field extraction: Combines regular expressions and machine learning models to achieve intelligent field recognition and extraction. The pre-trained NLP model can automatically identify key information such as IP addresses, timestamps, user IDs, and error codes, and supports custom field extraction rules. Subsequent processing and integration: The parsed structured data is written to the Elasticsearch cluster in real time to establish a multi-dimensional index structure. Streaming machine learning algorithms are used to detect abnormal patterns in logs in real time. When anomalies are detected, alert notifications are sent through multiple channels. Intelligently compress and hierarchically store historical log data; integrate the above processing results to provide a unified query and analysis interface.
2. A streaming method for processing massive syslog logs according to claim 1, characterized in that: Real-time index building: Write the parsed structured data to the Elasticsearch cluster in real time to establish a multi-dimensional index structure, supporting time series indexing, full-text search indexing, and aggregate analysis indexing to optimize query performance; Anomaly detection and alerting: Streaming machine learning algorithms are used to detect abnormal patterns in logs in real time, including statistical anomalies, pattern anomalies, and time series anomalies. Dynamic configuration of alert rules and threshold adjustment are supported. When an anomaly is detected, alert notifications are sent through multiple channels. Data compression storage: Intelligently compresses and tiers historical log data, automatically adjusting storage strategies based on data access frequency. Hot data is stored in SSDs to ensure query performance, while warm and cold data are stored in different storage media to optimize costs. Result output integration: Integrates the processing results of real-time index construction, anomaly detection and alarm, and data compression and storage, providing a unified REST API and GraphQL interface to support real-time query, historical analysis, and report generation.
3. A streaming method for processing massive syslog logs according to claim 2, characterized in that: In the log data collection and buffering step, Apache Kafka is used as the message queue to establish a multi-level buffering mechanism. This mechanism can automatically adjust the buffering strategy according to data flow and system load to ensure that data is not lost in different scenarios. It also supports data playback function to reprocess data when needed, and has fault-tolerant processing capabilities to deal with system failures.
4. A streaming method for processing massive syslog logs according to claim 3, characterized in that: In the intelligent field extraction step, the pre-trained NLP model is trained using a large amount of labeled syslog log data. It can accurately and automatically identify key information such as IP addresses, timestamps, user IDs, and error codes, and allows users to customize field extraction rules based on actual business needs. The model can be dynamically adjusted and optimized based on the new rules.
5. A streaming method for processing massive syslog logs according to claim 4, characterized in that: In the result output integration step, the unified REST API and GraphQL interface provided have high concurrent processing capabilities, which can meet the diverse needs of different clients for real-time query, historical analysis and report generation. At the same time, the interface has good scalability and can be easily integrated into existing business systems.
6. A system for processing a stream method of massive syslog logs according to claim 5, comprising a plurality of functional processing modules, characterized in that: Log data collection and buffering module: This module receives massive syslog data streams through high-performance message queues, collects various syslog log data from business systems, establishes a highly available data collection layer, uses Apache Kafka as a message queue, and supports multiple log formats. It also builds a multi-level buffering mechanism to ensure reliable data transmission and storage, while also supporting data playback and fault-tolerant processing. Streaming data preprocessing module: This module performs real-time cleaning and preprocessing on the collected raw log data, including data deduplication, format standardization, and invalid data filtering. Each processing unit includes three components: validation, normalization, and filtering to ensure data quality. Distributed stream parsing module: This module uses Apache Kafka and Apache Flink to build a distributed stream processing architecture and Apache Flink to build a distributed stream processing engine. This module achieves high-throughput processing through horizontal scaling. Leveraging Flink's windowing mechanism and state management capabilities, it implements complex event processing and session analysis, and supports dynamic adjustment of parallelism to adapt to load changes, enabling parallel parsing and processing of logs. Intelligent Field Extraction Module: This module combines regular expressions and machine learning models to implement intelligent field recognition and extraction. The pre-trained NLP model can automatically identify key information such as IP addresses, timestamps, user IDs, and error codes, while also supporting custom field extraction rules. Subsequent processing and integration module: Writes the parsed structured data to the Elasticsearch cluster in real time and establishes a multi-dimensional index; uses streaming machine learning algorithms to detect abnormal patterns in logs in real time and trigger an alarm mechanism; intelligently compresses and hierarchically stores historical data; and integrates the above processing results to provide a unified query and analysis interface.
7. A system according to claim 6, characterized in that: The subsequent processing and integration modules specifically include: Real-time index building submodule: Writes parsed structured data to the Elasticsearch cluster in real time, establishes a multi-dimensional index structure, and supports time series indexing, full-text search indexing, and aggregate analysis indexing to optimize query performance. Anomaly Detection and Alerting Submodule: This module uses streaming machine learning algorithms to detect abnormal patterns in logs in real time, covering statistical anomalies, pattern anomalies, and time series anomalies. When an anomaly is detected, it sends alarm notifications through multiple channels and supports dynamic configuration of alarm rules and threshold adjustment. Data compression and storage submodule: Intelligently compresses and tiers historical log data, automatically adjusts storage strategies based on data access frequency, stores hot data in SSDs to ensure query performance, and stores warm and cold data in different storage media to achieve cost optimization. Result output integration sub-module: Integrates the processing results of real-time index construction, anomaly detection and alarm, and data compression and storage, provides a unified REST API and GraphQL interface, and supports real-time query, historical analysis, and report generation.
8. A system according to claim 7, characterized in that: In the log data collection buffer module, the multi-level buffering mechanism has adaptive adjustment capabilities. It can automatically optimize the buffering strategy according to data traffic size, system load, and network conditions to ensure that data loss can be effectively prevented in different scenarios, while ensuring the stability of the data playback function and the efficiency of fault-tolerant processing.
9. A system according to claim 8, characterized in that: In the intelligent field extraction module, the pre-trained NLP model is trained based on a large amount of annotated syslog log data. By continuously optimizing the model parameters, the accuracy and recall of field recognition are improved. For custom field extraction rules, the system provides a visual configuration interface to facilitate users to flexibly set them according to actual business needs, and can apply new rules to the field extraction process in real time.
10. A system according to claim 9, characterized in that: The unified REST API and GraphQL interface provided in the result output integration module has high concurrent processing capabilities and good scalability; the interface adopts a security authentication mechanism to ensure the security of data transmission and access; at the same time, the system provides complete interface documentation and sample code to facilitate developers to quickly integrate and use, meeting the diverse needs of different clients for real-time query, historical analysis and report generation.