Log management and query method
By using dynamic topic classification and asynchronous message push technology in PostgreSQL log management and query methods, log data is stored in non-relational databases, solving the problems of irregular log format, low query efficiency, and insufficient log analysis and monitoring, and achieving efficient log query and analysis.
Patent Information
- Application Number
- CN202510607802.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
AI Technical Summary
When processing PostgreSQL logs, the prior art has problems such as irregular log format, low query efficiency, and insufficient log analysis and monitoring. Especially when processing high throughput log data, it cannot meet the requirements of performance, scalability and query efficiency.
A log management and query method is adopted, including log data extraction, dynamic topic classification, asynchronous message push and distributed storage processing. The classified log data is pushed asynchronously to a non-relational database for storage through message queue middleware, achieving efficient structured storage and query.
It improves log query efficiency and system throughput, ensures that log data can still be transmitted stably under high load conditions, and realizes efficient structured storage and query, which facilitates subsequent analysis and monitoring.
Smart Images

Figure CN120123307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method for log management and query. Background Art
[0002] Unstructured log files (such as the pg_log file of PostgreSQL) record detailed information about system operations, such as query execution, error logs, transaction information, etc. These logs are usually unstructured text data, making it difficult to directly perform efficient queries and analysis. Currently, some log management solutions have carried out log transmission and storage through centralized log management tools (such as ELK, Splunk, etc.), but they usually lack an efficient structured storage method, resulting in low log query efficiency, especially in cases where the log volume is large and certain events need to be quickly located.
[0003] PostgreSQL logs are recorded in text format and stored in local log files. These logs usually contain a large amount of raw text information and lack a structured storage form. Therefore, the following problems exist in querying and analyzing:
[0004] Irregular log format: The log data of PostgreSQL is usually stored in plain text, lacking a unified structure, making it difficult to automatically extract information from the logs or perform efficient queries. For example, information such as query time, SQL statements, error codes, execution status, etc. are mixed in the plain text, resulting in difficult and inefficient manual analysis.
[0005] Low query efficiency: Since PostgreSQL logs are unstructured text, queries need to be performed through full-text retrieval, and specific events or data cannot be quickly located. This is particularly inconvenient for large-scale log data management and fault diagnosis, especially in scenarios where the log volume is extremely large or queries are frequent, and the performance bottleneck is very obvious.
[0006] Insufficient log analysis and monitoring: Taking ELK as an example, the Logstash component in ELK can process data through various plugins, but for complex log formats or irregular data structures, users need to write a large number of configurations and scripts for data cleaning and formatting. Complex log parsing and cleaning operations may cause delays in the Logstash processing process, thus affecting the real-time nature of the logs. ELK is mainly suitable for the storage and analysis of log and event data. Although it can process some structured data, it is not suitable for use as a large-scale transactional database. If the logs contain a large amount of structured data or complex relational data, ELK is not the best choice.
[0007] In summary, the existing technical solutions still have significant deficiencies in the efficient storage, structured processing, and query analysis of logs. Especially when dealing with high-throughput log data, the existing solutions often fail to meet the requirements in terms of performance, scalability, and query efficiency. Summary of the Invention
[0008] To solve the problems of non-standard log format, low query efficiency, and insufficient log analysis and monitoring in the existing PostgreSQL logs, the present invention provides a log management and query method.
[0009] The technical solution adopted by the present invention is as follows:
[0010] A log management and query method includes the following steps:
[0011] S1. Log data extraction: Read the log file data from the locally stored log files;
[0012] S2. Dynamic topic classification: Analyze the content of each line of log file data, extract the preset keyword fields; generate corresponding dynamic topic identifiers according to the values of the keyword fields, and map different categories of log files to different topics in the message queue middleware;
[0013] S3. Asynchronous message push: Start an independent producer thread, encapsulate the classified log file data as messages, and push them to the corresponding topics in the message queue middleware;
[0014] S4. Distributed storage processing: Start an independent consumer thread to listen to multiple topics in the message queue middleware; select the storage nodes or collections of the target non-relational database according to the topic identifier of the message; convert the log messages into the document format supported by the non-relational database and persistently store them in the corresponding storage nodes or collections.
[0015] Preferably, the implementation method of dynamic topic classification in step S2 includes:
[0016] If the log content contains an error level field, use the error level as the suffix of the topic identifier to generate a hierarchical topic name;
[0017] If the log content contains a transaction operation type field, generate independent topics according to the operation type.
[0018] Preferably, the distributed storage processing in step S4 further includes:
[0019] Distribute the log data to different non-relational database instances or shard clusters according to the topic identifier and the preset mapping rules;
[0020] Add metadata fields to the stored document, including the log source timestamp, the topic identifier, and the storage node identifier.
[0021] Preferably, the producer thread and the consumer thread achieve loose-coupling communication through the message queue middleware, including:
[0022] The producer thread uses an asynchronous callback mechanism to push messages, supporting message retry and failure alerting;
[0023] The consumer thread realizes horizontal expansion through the consumer group mode and consumes messages of multiple topics in parallel.
[0024] Preferably, the parsing and pushing of the log file include:
[0025] Perform format verification on the log file to filter out invalid data lines;
[0026] Pre-create topics in the message queue middleware or enable the automatic topic creation function.
[0027] Preferably, the storage node selection strategy for the non-relational database includes:
[0028] Allocate storage nodes according to the hash value of the topic identifier;
[0029] Select time series database shards based on the timestamp range of the log data.
[0030] Preferably, the method further includes an exception handling mechanism:
[0031] If the message queue middleware or the database connection fails, enable the local cache queue to temporarily store the log data;
[0032] When the service is restored, give priority to retransmitting the cached log data.
[0033] Preferably, the method is applicable to at least one of relational database logs, application running logs, and network device logs.
[0034] The beneficial effects of the present invention are:
[0035] 1. Adopt an asynchronous log scraping mechanism, process log data scraping through multi-threading or asynchronous tasks, which will not block the main thread, thereby improving the system throughput and response speed. At the same time, use a message queue (such as Kafka) for efficient data transmission to ensure stable transmission of log data under high load.
[0036] 2. Introduce an incremental scraping mechanism, avoid repeated scraping of the same log content by recording the last processed position or timestamp of the log file. In addition, the built-in deduplication mechanism ensures the uniqueness of the log data, prevents repeated processing of logs, and improves data accuracy and storage efficiency.
[0037] 3. A memory data cleaning and sorting component is designed to efficiently preprocess the captured log data, including removing irrelevant fields, standardizing formats, and desensitization processing. Through fast cleaning and sorting in memory, it ensures that the subsequent stored data structure is unified and clear, facilitating query and analysis.
[0038] 4. High fault tolerance is achieved through asynchronous capture and message queue technology. During the log capture process, the system has an error recovery mechanism. Once an exception occurs, it can automatically retry or recover tasks to ensure that log data is not lost, improving the reliability and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic flowchart of the log management and query method in the embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0041] As Figure 1 shown, a log management and query method, the main idea is to obtain the local log content, collect it into a message queue, classify and sort different log contents, divide them into different topics, and after cleaning and sorting, store them in a MongoDB or other NoSQL database, and achieve asynchronous storage through program control to ensure that the tool does not affect the performance of the source database. The main principle is as Figure 1 shown:
[0042] 1. After PostgreSQL is configured, a log file in CSV format is formed; the log directory is set in the log interception program;
[0043] 2. The collector intercepts the local log content in real time and places it into different topics of different nodes according to the log content and format;
[0044] 3. After being cleaned and sorted by the program, the finally processed data is stored in a MongoDB or other NoSQL database according to different topics.
[0045] The following is a code example (basic logic) that can be used for the log capture and collection program:
[0046] public class PostgresCsvToKafkaAndMongoDB {
[0047] private static final String KAFKA_BROKER = "localhost:9092"; / / Kafka address
[0048] private static final String MONGO_HOST = "localhost"; / / MongoDB address
[0049] private static final int MONGO_PORT = 27017; / / MongoDB port
[0050] private static final String KAFKA_TOPIC = "postgres-logs"; / / Kafka topic
[0051] private static final String CSV_FILE_PATH = "path / to / your / postgresql_log.csv"; / / Local CSV file path
[0052] public static void main(String[] args) {
[0053] / / Start two threads: one for pushing CSV file logs to Kafka and the other for consuming from Kafka and storing to MongoDB
[0054] Thread producerThread = new Thread(PostgresCsvToKafkaAndMongoDB::produceLogsToKafka);
[0055] Thread consumerThread = new Thread(PostgresCsvToKafkaAndMongoDB::consumeLogsFromKafkaAndStoreToMongoDB);
[0056] producerThread.start();
[0057] consumerThread.start();}
[0058] / / Producer method for pushing CSV logs to Kafka
[0059] public static void produceLogsToKafka() {
[0060] / / Kafka Configuration
[0061] Properties props = new Properties();
[0062] props.put(ProducerConfig.BOOTSTRAP_SERVERS_CONFIG, KAFKA_BROKER);
[0063] props.put(ProducerConfig.KEY_SERIALIZER_CLASS_CONFIG, "org.apache.kafka.common.serialization.StringSerializer");
[0064] props.put(ProducerConfig.VALUE_SERIALIZER_CLASS_CONFIG, "org.apache.kafka.common.serialization.StringSerializer");
[0065] / / Create a Kafka producer
[0066] KafkaProducer<String, String> producer = new KafkaProducer<>(props);
[0067] / / Read the CSV file and send each line to Kafka
[0068] try (CSVReader reader = new CSVReader(new FileReader(CSV_FILE_PATH))) {
[0069] String[] nextLine;
[0070] while ((nextLine = reader.readNext()) != null) {
[0071] / / Convert each line of the CSV to a String
[0072] String logMessage = String.join(",", nextLine);
[0073] / / Push the log message to Kafka
[0074] ProducerRecord<String, String> record = new ProducerRecord<>(KAFKA_TOPIC, logMessage);
[0075] producer.send(record, (metadata, exception) -> {
[0076] if (exception != null) {
[0077] System.err.println("Error sending message: " + exception.getMessage());
[0078] } else {
[0079] System.out.println("Message sent to topic: " + metadata.topic());
[0080] }
[0081] });
[0082] }
[0083] } catch (IOException e) {
[0084] e.printStackTrace();
[0085] } finally {
[0086] producer.close();
[0087] }
[0088] }
[0089] / / Consumer method to consume log messages from Kafka and store them in MongoDB based on the topic
[0090] public static void consumeLogsFromKafkaAndStoreToMongoDB() {
[0091] / / Kafka configuration
[0092] Properties props = new Properties();
[0093] props.put("bootstrap.servers", KAFKA_BROKER);
[0094] props.put("group.id", "log-group");
[0095] props.put("key.deserializer", StringDeserializer.class.getName());
[0096] props.put("value.deserializer", StringDeserializer.class.getName());
[0097] props.put("auto.offset.reset", "earliest");
[0098] / / Kafka consumer
[0099] KafkaConsumer<String, String> consumer = new KafkaConsumer<>(props);
[0100] consumer.subscribe(Arrays.asList(KAFKA_TOPIC));
[0101] / / MongoDB client
[0102] MongoClient mongoClient = new MongoClient(MONGO_HOST,MONGO_PORT);
[0103] try {
[0104] while (true) {
[0105] var records = consumer.poll(1000); / / Pull data once per second
[0106] records.forEach(record -> {
[0107] String topic = record.topic();
[0108] String message = record.value();
[0109] System.out.println("Received message: " + message + " from topic: " +topic);
[0110] / / Select the MongoDB database according to the Kafka topic
[0111] MongoDatabase mongoDatabase = getMongoDatabase(mongoClient, topic);
[0112] MongoCollection <document>collection = mongoDatabase.getCollection("logs");
[0113] / / Convert the message to a MongoDB document and insert it
[0114] Document doc = new Document("logMessage", message)
[0115] .append("timestamp", System.currentTimeMillis());
[0116] collection.insertOne(doc);
[0117] System.out.println("Message saved to MongoDB database: " +mongoDatabase.getName());
[0118] });
[0119] }
[0120] } finally {
[0121] consumer.close();
[0122] mongoClient.close();
[0123] }
[0124] }
[0125] / / Select the MongoDB database according to the topic
[0126] private static MongoDatabase getMongoDatabase(MongoClientmongoClient, String topic) {
[0127] / / Return different MongoDB data according to the Kafka topic
[0128] if ("postgres-logs".equals(topic)) {
[0129] return mongoClient.getDatabase("postgres_logs_db"); / / Store PostgreSQL logs
[0130] } else {
[0131] return mongoClient.getDatabase("default_logs_db"); / / Default database
[0132] }
[0133] }
[0134] }
[0135] The above method has the following key points:
[0136] 1. Asynchronous log scraping and collection
[0137] Asynchronous data scraping mechanism: Ensure that the log collection process does not affect system performance due to blocking operations. Perform parallel scraping through asynchronous tasks or thread pools to improve the efficiency of data collection and reduce resource consumption.
[0138] Incremental scraping and deduplication function: Design a deduplication mechanism to avoid duplicate scraping of the same log, ensuring that each log is processed only once. And be able to implement the breakpoint resume function.
[0139] High fault tolerance and recovery mechanism: During the log scraping process, considering abnormal situations such as network fluctuations and system crashes, design a fault-tolerant and recovery mechanism to ensure that the log scraping task can automatically retry or recover after failure without losing important log data.
[0140] 2. In-memory data cleaning and sorting
[0141] Data cleaning and formatting algorithms: Include field standardization, error data repair, data desensitization, etc.
[0142] Data conversion rules: Used to convert the original log data into a standard format, such as extracting key information through regular expressions, converting the timestamp format, filtering out unnecessary fields, etc.
[0143] 3. Real-time processing and batch processing capabilities: Support a hybrid mode of real-time processing and batch processing. For high-frequency and real-time log data, be able to perform streaming cleaning and immediately push and store it; be able to batch process historical data efficiently without affecting the system operation as much as possible.
[0144] Dynamic configurability: Support dynamic configuration, and can dynamically adjust cleaning rules, field mapping deduplication, etc. according to log sources or business requirements.
[0145] The present invention has the following technical advantages compared with most existing log management and query methods:
[0146] 1. Asynchronous log scraping and efficient data transmission
[0147] Most existing log scraping technologies collect logs synchronously, which can easily lead to system performance bottlenecks. Especially when dealing with high-frequency log sources, data loss or processing delays may occur.
[0148] The present invention adopts an asynchronous log scraping mechanism. By using multi-threading or asynchronous tasks to process log data scraping, it does not block the main thread, thus improving system throughput and response speed. At the same time, a message queue (such as Kafka) is used for efficient data transmission to ensure stable log data transmission under high load.
[0149] 2. Incremental scraping and deduplication mechanism
[0150] Many log collection tools cannot effectively handle duplicate data. Especially when dealing with local log files, it is difficult to avoid duplicate scraping due to factors such as file rolling or system restart.
[0151] The present invention introduces an incremental scraping mechanism. By recording the last processed position or timestamp of the log file, it avoids duplicate scraping of the same log content. In addition, the built-in deduplication mechanism ensures the uniqueness of log data, prevents duplicate processing of logs, and improves data accuracy and storage efficiency.
[0152] 3. In-memory data cleaning and efficient sorting
[0153] Existing log processing systems often lack efficient data cleaning and sorting capabilities. Especially when dealing with large-scale log data, problems such as inconsistent data formats and excessive redundant information are likely to occur, affecting subsequent analysis and storage.
[0154] The present invention designs an in-memory data cleaning and sorting component to efficiently preprocess the scraped log data, including removing irrelevant fields, standardizing formats, and desensitization processing. Through fast cleaning and sorting in memory, it ensures that the subsequent stored data structure is unified and clear, facilitating query and analysis.
[0155] 4. High fault tolerance and recovery ability
[0156] Many log scraping systems may lose log data or processing may fail when facing network instability, system crashes, or other abnormal situations.
[0157] The present invention realizes high fault tolerance through asynchronous scraping and message queue technologies. During the log scraping process, the system has an error recovery mechanism. Once an exception occurs, it can automatically retry or recover tasks to ensure that log data is not lost, thereby improving the reliability and stability of the system.
[0158] The above-described embodiments only represent specific implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.< / document>
Claims
1. A log management and query method, characterized in that: The steps include: S1. Log data extraction: reading log file data from locally stored log files; S2, dynamic topic classification: parse the content of each line of log file data and extract the preset key fields; generate corresponding dynamic topic identifiers according to the values of the key fields, and map log files of different categories to different topics of the message queue middleware; S3, asynchronous message push: start an independent producer thread, encapsulate the classified log file data into messages, and push them to the corresponding topic of the message queue middleware; S4, distributed storage processing: start an independent consumer thread and listen to multiple topics of the message queue middleware; select a storage node or collection of the target non-relational database according to the topic identifier to which the message belongs; convert the log message into a document format supported by the non-relational database, and store it persistently in the corresponding storage node or collection.
2. The method according to claim 1, characterized in that The implementation of dynamic topic classification in step S2 includes: If the log content contains an error level field, the error level is used as the suffix of the topic identifier to generate a topic name with a hierarchical structure; If the log content contains the transaction operation type field, a separate topic is generated according to the operation type.
3. The method according to claim 1, characterized in that The distributed storage processing in step S4 includes: Distribute log data to different non-relational database instances or shard clusters based on topic identifiers and preset mapping rules; Add metadata fields to storage documents, including log source timestamp, topic identifier, and storage node identifier.
4. The method according to claim 1, characterized in that The producer thread and the consumer thread achieve loosely coupled communication through the message queue middleware, including: The producer thread uses an asynchronous callback mechanism to push messages, supporting message retries and failure alarms; Consumer threads achieve horizontal expansion through the consumer group mode and consume messages from multiple topics in parallel.
5. The method according to claim 1, characterized in that The parsing and pushing of log files include: Verify the format of log files and filter invalid data rows; Pre-create topics in the message queue middleware or enable automatic topic creation.
6. The method according to claim 1, characterized in that The storage node selection strategies for non-relational databases include: Assign storage nodes based on the hash value of the subject identifier; Select the time series database shard based on the timestamp range of the log data.
7. The method according to claim 1, characterized in that The method also includes an exception handling mechanism: If the message queue middleware or database connection fails, enable the local cache queue to temporarily store log data; When the service is restored, cached log data is retransmitted first.
8. The method according to any one of claims 1 to 7, characterized in that: The method is applicable to at least one of a relational database log, an application program operation log, and a network device log.
Citation Information
Patent Citations
Log data storage method and device, terminal and computer storage medium
CN111367873A
Method for realizing asynchronous pushing of data based on message queue
CN113900833A
Online analysis method and system of multi-source log, electronic equipment and storage medium
CN115437877A
Container terminal message automatic processing method based on distributed task scheduling platform
CN117544620A
Method for quickly generating multi-format file based on database large data volume table
CN119201848A