Log collection, storage and retrieval method and device
By creating log topics and client groups, combining standard storage and low-frequency storage, and using API or message middleware access to log collection and storage, the problems of complex and high storage costs of traditional log management methods are solved, centralized and unified management and efficient retrieval of logs are realized, and storage costs are saved.
Patent Information
- Application Number
- CN202510135136.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-06
AI Technical Summary
In traditional systems or application architectures, log recording and management methods are complex, and different systems or applications may adopt different log formats, which increases the difficulty of data analysis and processing. At the same time, as the amount of log data increases, storage costs gradually increase.
By creating log topics and client groups, centralized and unified log management is achieved. Log topics include topic name, topic ID, log access method, retention time, storage type, transfer rules, collection rules and search rules. The combination of standard storage and low-frequency storage is adopted to select the appropriate storage type according to the usage scenario, and log collection and storage are carried out through API or message middleware access.
It realizes centralized and unified log management, improves the convenience of log retrieval and operation and maintenance efficiency, and at the same time, dynamically manages log storage capacity, saving storage costs.
Smart Images

Figure CN120104571A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing technology, and specifically provides a log collection, storage, and retrieval method and device. Background Art
[0002] System logs, as an indispensable record of the operation of computer systems or applications, capture various events and states in detail. These logs play a vital role in many fields, including auditing, real-time monitoring, fault diagnosis, and user behavior analysis. However, in traditional system or application architectures, the recording and management of logs often face many challenges.
[0003] Traditional logging methods usually store data in local files or databases. This decentralized storage method makes log search and management complicated. What's more, different systems or applications may use different log formats, which undoubtedly increases the difficulty of data analysis and processing. At the same time, as the amount of log data continues to grow, storage costs are also gradually rising, and resource consumption is becoming increasingly significant. Summary of the invention
[0004] The present invention aims at solving the above-mentioned deficiencies of the prior art and provides a log collection, storage and retrieval method with strong practicability.
[0005] A further technical task of the present invention is to provide a log collection, storage and retrieval device that is reasonably designed, safe and applicable.
[0006] The technical solution adopted by the present invention to solve the technical problem is:
[0007] A log collection, storage and retrieval method, first, creates a log topic, the log topic is the smallest unit of log collection, storage and retrieval; then creates a client group, the client group is associated with one or more log topics, the machines within the client group range report to the same log topic or multiple log topics, and the machines not associated with the client group cannot report logs.
[0008] Furthermore, the log topic includes a topic name, a topic ID, a log access method, a retention time, a storage type, a transfer rule, a collection rule, and a retrieval rule.
[0009] Furthermore, the log access method includes client access, API access and message middleware access; the retention time refers to the time point when the log will be deleted after the retention time is exceeded; the storage type includes standard storage and low-frequency storage; the transfer rule refers to the rule that converts the standard storage to low-frequency storage when it is met; the log collection rule refers to the rule that parses and extracts the log; the log retrieval rule refers to the rule that segments the log and creates an index.
[0010] Furthermore, standard storage and low-frequency storage are two types of storage in the cloud object storage service. Standard storage provides high-performance object storage services and supports frequent data access;
[0011] Low-frequency storage provides object storage services with high durability, low-frequency access, and low storage costs. Users can choose the appropriate storage type based on the usage scenario.
[0012] Furthermore, using the client access method, according to the client group combined with the log topic, the client collection program stores the collected logs in the storage under the log topic. The client collection program will monitor the log files on the server hard disk in real time and report the log data to the log topic according to the preset collection strategy and period;
[0013] Specifically, when the collection strategy is incremental collection, the collection program will periodically sense that the log file has been modified, and will actively collect and report the newly written incremental log, and record the collection position, which will be used as the starting position for the next collection; when the collection strategy is full collection, the collection program will real-time sense that the file has been modified, and will actively collect and report the entire content of the log.
[0014] Furthermore, before reporting the logs, logs are extracted from the log files based on the log collection rules. Logs are divided into single-line logs and multi-line logs. Single-line logs use the line break character \n as the end of a log. Single-line logs can be extracted in full and the line is reported as a complete string.
[0015] If the original log is extracted according to the specified regular expression, the regular expression is used; if the original log is in JSON format, the JSON is parsed for extraction; if the original log can be extracted according to the specified delimiter, the delimiter is used for extraction; if the log meets multiple collection rules, a combination of multiple extraction methods is used;
[0016] A multi-line log is a complete log data that occupies multiple lines. Multi-line logs are matched by the first-line regular expression. When a log line matches the preset regular expression, it is considered to be the beginning of a log line, and the beginning of the next line is used as the end identifier of the log line.
[0017] Furthermore, the client acquisition program has a log compression function, which packages multiple logs into a log group, and compresses the logs into a compressed package based on the log group;
[0018] Use the API access method, set the correct request parameters based on the API link provided by the server, and the request parameters include at least the topic ID and token. The client's business system actively calls the API, generates logs that comply with the rules based on the log collection rules, and then actively reports them to the corresponding log topic.
[0019] Furthermore, when using the access method of the message middleware, the configuration parameters of the message middleware are configured, and the logs are reported to the message middleware. The message middleware stores the logs in the object storage service based on the content of the log topic.
[0020] Based on the log topic, the log is stored in the corresponding object storage. According to the log transfer rule, after the log reaches the specified date, the log in the standard storage will be transferred to the low-frequency storage. According to the log retention time, the log will be deleted after reaching the specified date.
[0021] Based on the log files under the log topic, word segmentation and indexing are used to divide the full text of the log into multiple segments, each segment is called a "word", and this process is called "word segmentation". Based on the retrieval rules of the log topic, user-defined rules are used to achieve customized retrieval requirements. After the log is word segmented, an inverted index is used to store the position of each "word" in the log document;
[0022] Based on the index conditions, the original log is matched. The index determines the conditions under which the log can be retrieved and analyzed. The index data is stored in the metadata bucket of the object storage. The metadata bucket is a standard bucket established under the standard storage type and is used to store indexes.
[0023] A log collection, storage and retrieval device, comprising a management unit, a collection unit, a storage unit and a retrieval unit;
[0024] The management unit is used to create a log topic and a client group, wherein the log topic includes a topic name, a topic ID, an access method, a retention time, a storage type, a transfer rule, a collection rule, and a retrieval rule;
[0025] The client group maintains the IP address of the client, and the management unit maintains the association relationship between collection and storage;
[0026] The collection unit is used to extract and report logs, and has the function of extracting single-line logs and multi-line logs. The client access method provides an independent collection unit. The user installs the collection program on the local log file machine and uses the API access method or the message middleware access method to integrate the collection unit into the user's business system;
[0027] The storage unit is used to store logs. Based on the retention time, storage type, transfer rules, and collection rules of the log subject, the logs are stored in the cloud object storage service. Log files that need to be efficiently queried are stored in standard storage, and log files that are accessed infrequently are stored in low-frequency storage. Standard storage that meets the transfer rules is converted to low-frequency storage, and logs that exceed the retention time will be deleted. The upload interface of the object storage is called for uploading logs. After the logs meet the rules, the transfer and deletion use the life cycle conversion function of the object storage, which is automatically executed without manual triggering.
[0028] The retrieval unit processes the log data and creates an index using word segmentation. Based on the word segmentation, the full text of the log is divided into multiple segments, and based on the index, an association with the original log is established.
[0029] Compared with the prior art, the log collection, storage, and retrieval method and device of the present invention have the following outstanding beneficial effects:
[0030] The present invention can not only realize the centralized and unified management of logs, but also has a powerful log retrieval capability, allowing users to more conveniently analyze log contents, thereby improving operation and maintenance efficiency. At the same time, based on log management and the use of cloud object storage services, log storage capacity is dynamically managed to save costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0032] Attached Figure 1 It is a flow chart of a log collection, storage, and retrieval method;
[0033] Attached Figure 2 It is a schematic diagram of the log access method in the log collection, storage and retrieval method;
[0034] Attached Figure 3 The invention is a flow chart of a log collection, storage and retrieval device. DETAILED DESCRIPTION
[0035] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention is further described in detail below in conjunction with specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] A best embodiment is given below:
[0037] like Figure 1 As shown, a log collection, storage, and retrieval method in this embodiment first creates a log topic, which is the smallest unit of log collection, storage, and retrieval, including a topic name, topic ID, access method, retention time, storage type, transfer rules, collection rules, and retrieval rules.
[0038] Log access methods include client access, API access, and message middleware access. The log retention time refers to the time point when the log will be deleted after the retention time. Log storage types include standard storage and low-frequency storage. Log transfer rules refer to the rules that convert standard storage to low-frequency storage when they are met.
[0039] Standard storage and low-frequency storage are two types of storage in the cloud object storage service. Standard storage provides high-performance object storage services and can support frequent data access.
[0040] Low-frequency storage provides object storage services with high durability, low-frequency access, and low storage costs. Users can choose the appropriate storage type based on the usage scenario.
[0041] Log retention time and log transfer are implemented based on the lifecycle rules of the object storage service. Log collection rules refer to the rules used to parse and extract logs. Log retrieval rules refer to the rules used to segment logs and create indexes.
[0042] Create a client group. The client group maintains a group of client IP addresses. A client group can contain multiple machines. A client group can be associated with one or more log topics. Machines within the client group can report to the same log topic or multiple log topics. Machines that are not associated with a client group cannot report logs.
[0043] Different log access methods have different collection methods, such as Figure 2 . In one example, the client access method is used, and the client collection program stores the collected logs in the storage under the log topic according to the client group combined with the log topic. The client collection program will monitor the log files on the server hard disk in real time, and report the log data to the log topic according to the preset collection strategy and period. Specifically, when the collection strategy is incremental collection, the collection program will periodically sense that the log file has been modified, and will actively collect and report the newly written incremental logs, and record the collection position, which will be used as the starting position for the next collection; when the collection strategy is full collection, the collection program will sense that the file has been modified in real time, and will actively collect and report the entire content of the log.
[0044] Before reporting logs, logs are extracted from log files based on log collection rules. Logs are divided into single-line logs and multi-line logs. Single-line logs use the line break character \n as the end of a log. Single-line log extraction can be fully extracted and the line is reported as a complete string. If the original log can be extracted according to the specified regular expression, the regular expression is used. If the original log is in json format, the json is parsed for extraction. If the original log can be extracted according to the specified delimiter, the delimiter is used for extraction. If the log meets multiple collection rules, a combination of multiple extraction methods can be used.
[0045] A multi-line log is a complete log data that occupies multiple lines. Multi-line logs are matched by the first line regular expression. When a log line matches the preset regular expression, it is considered to be the beginning of a log, and the beginning of the next line is used as the end identifier of the log.
[0046] In order to improve the transmission rate, the client acquisition program has a log compression function. Specifically, multiple logs are packaged into a log group, and the logs are packaged and compressed into a compressed package based on the log group.
[0047] Use the API access method, set the correct request parameters based on the API link provided by the server, and the request parameters include at least the topic ID and token. The client's business system actively calls the API, generates logs that comply with the rules based on the log collection rules, and then actively reports them to the corresponding log topic.
[0048] You can also use the access method of message middleware, such as rabbitmq, kafka and other message middleware. Specifically, configure the configuration parameters of the message middleware such as hosts, queues, passwords, and report logs to the message middleware. The message middleware stores logs in the object storage service based on the content of the log topic.
[0049] Based on the log topic, the logs are stored in the corresponding object storage. According to the log transfer rules, when the logs reach the specified date, the logs in the standard storage will be transferred to the low-frequency storage. According to the log retention time, when the logs reach the specified date, they will be deleted to free up space and save storage costs.
[0050] Based on the log files under the log theme, word segmentation and indexing. Since the full text of the log is very long and scattered into multiple small files, log retrieval is difficult. In order to meet the retrieval requirements, the full text of the log needs to be segmented into multiple segments, each segment is called a "word", and this process is called "word segmentation". Based on the retrieval rules of the log theme, users can customize the rules to achieve customized retrieval requirements. For example, the log is segmented by symbols, and the log is segmented as long as spaces and line breaks appear. After the log is segmented, an inverted index is used to store the position of each "word" in the log in the document. Based on the index conditions, the original log is matched, and the index determines under what conditions the log can be retrieved and analyzed. Index data is stored in the metadata bucket of the object storage. The metadata bucket is a standard bucket established under the standard storage type for storing indexes.
[0051] like Figure 3 As shown, a log collection, storage, and retrieval device includes a management unit, a collection unit, a storage unit, and a retrieval unit;
[0052] The management unit is used to create a log topic and a client group, wherein the log topic includes a topic name, a topic ID, an access method, a retention time, a storage type, a transfer rule, a collection rule, and a retrieval rule;
[0053] The client group maintains the client's IP address, and the management unit maintains the association between collection and storage;
[0054] The collection unit is used to extract and report logs, and has the function of extracting single-line logs and multi-line logs. The client access method provides an independent collection unit. The user installs the collection program on the local log file machine and uses the API access method or the message middleware access method to integrate the collection unit into the user's business system.
[0055] The storage unit is used to store logs. Logs are stored in the cloud object storage service based on the retention time, storage type, transfer rules, and collection rules of the log topic. Log files that need to be efficiently queried are stored in standard storage, and log files that are accessed infrequently are stored in low-frequency storage. Standard storage that meets the transfer rules is converted to low-frequency storage, and logs that exceed the retention time will be deleted. The upload interface of the object storage is called to upload logs. After the logs meet the rules, the transfer and deletion use the life cycle conversion function of the object storage, which is automatically executed without manual triggering.
[0056] The retrieval unit processes the log data and creates an index using word segmentation. Based on the word segmentation, the full text of the log is divided into multiple segments, and based on the index, an association with the original log is established.
[0057] The above-mentioned specific implementations are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementations. Any technical solutions that conform to the above-mentioned specific implementations of the present invention and any appropriate changes or substitutions made by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.
[0058] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A log collection, storage, and retrieval method, characterized in that: First, a log topic is created, which is the smallest unit for log collection, storage, and retrieval. Then a client group is created, which is associated with one or more log topics. Machines within the client group report to the same log topic or multiple log topics, and machines that are not associated with the client group cannot report logs.
2. A log collection, storage, and retrieval method according to claim 1, characterized in that: The log topic includes topic name, topic ID, log access method, retention time, storage type, transfer rules, collection rules and retrieval rules.
3. A log collection, storage and retrieval method according to claim 2, characterized in that: The log access methods include client access, API access and message middleware access; the retention time refers to the time point when the log will be deleted after the retention time is exceeded; the storage type includes standard storage and low-frequency storage; the transfer rule refers to the rule that converts the standard storage to low-frequency storage when it is met; the log collection rule refers to the rule used to parse and extract the log; the log retrieval rule refers to the rule used to segment the log and create an index.
4. A log collection, storage and retrieval method according to claim 3, characterized in that: Standard storage and low-frequency storage are two types of storage in the cloud object storage service. Standard storage provides high-performance object storage services and supports frequent data access. Low-frequency storage provides object storage services with high durability, low-frequency access, and low storage costs. Users can choose the appropriate storage type based on the usage scenario.
5. A log collection, storage and retrieval method according to claim 4, characterized in that: Using the client access method, the client collection program stores the collected logs in the storage under the log topic according to the client group and the log topic. The client collection program monitors the log files on the server hard disk in real time and reports the log data to the log topic according to the preset collection strategy and period; Specifically, when the collection strategy is incremental collection, the collection program periodically senses that the log file has been modified, and will actively collect and report the newly written incremental log, and record the collection position, which will be used as the starting position for the next collection; When the collection strategy is full collection, the collection program will sense in real time that the file has been modified and will actively collect and report the entire content of the log.
6. A log collection, storage, and retrieval method according to claim 5, characterized in that: Before reporting logs, logs are extracted from log files based on log collection rules. Logs are divided into single-line logs and multi-line logs. Single-line logs use the line break character \n as the end of a log. Single-line logs can be extracted in full and the line is reported as a complete string. If the original log is extracted according to the specified regular expression, the regular expression is used; if the original log is in JSON format, the JSON is parsed for extraction; if the original log can be extracted according to the specified delimiter, the delimiter is used for extraction; if the log meets multiple collection rules, a combination of multiple extraction methods is used; A multi-line log is a complete log data that occupies multiple lines. Multi-line logs are matched by the first-line regular expression. When a log line matches the preset regular expression, it is considered to be the beginning of a log line, and the beginning of the next line is used as the end identifier of the log line.
7. A log collection, storage, and retrieval method according to claim 6, characterized in that: The client acquisition program has a log compression function, which packages multiple logs into a log group and compresses the log group into a compressed package; Use the API access method, set the correct request parameters based on the API link provided by the server, and the request parameters include at least the topic ID and token. The client's business system actively calls the API, generates logs that comply with the rules based on the log collection rules, and then actively reports them to the corresponding log topic.
8. A log collection, storage, and retrieval method according to claim 7, characterized in that: When using the message middleware access method, configure the configuration parameters of the message middleware, report logs to the message middleware, and the message middleware stores the logs in the object storage service based on the content of the log topic; Based on the log topic, the log is stored in the corresponding object storage. According to the log transfer rule, after the log reaches the specified date, the log in the standard storage will be transferred to the low-frequency storage. According to the log retention time, the log will be deleted after reaching the specified date. Based on the log files under the log topic, word segmentation and indexing, the full text of the log is divided into multiple segments, each segment is called a "word", and this process is called "word segmentation". Based on the retrieval rules of the log topic, user-defined rules are used to achieve customized retrieval requirements. After the log is word segmented, an inverted index is used to store the position of each "word" in the log document; Based on the index conditions, the original log is matched. The index determines the conditions under which the log can be retrieved and analyzed. The index data is stored in the metadata bucket of the object storage. The metadata bucket is a standard bucket established under the standard storage type and is used to store indexes.
9. A log collection, storage, and retrieval device, characterized in that: It includes a management unit, a collection unit, a storage unit and a retrieval unit; The management unit is used to create a log topic and a client group, wherein the log topic includes a topic name, a topic ID, an access method, a retention time, a storage type, a transfer rule, a collection rule, and a retrieval rule; The client group maintains the IP address of the client, and the management unit maintains the association relationship between collection and storage; The collection unit is used to extract and report logs, and has the function of extracting single-line logs and multi-line logs. The client access method provides an independent collection unit. The user installs the collection program on the local log file machine and uses the API access method or the message middleware access method to integrate the collection unit into the user's business system; The storage unit is used to store logs. Based on the retention time, storage type, transfer rules, and collection rules of the log subject, the logs are stored in the cloud object storage service. Log files that need to be efficiently queried are stored in standard storage, and log files that are accessed infrequently are stored in low-frequency storage. Standard storage that meets the transfer rules is converted to low-frequency storage, and logs that exceed the retention time will be deleted. The upload interface of the object storage is called for uploading logs. After the logs meet the rules, the transfer and deletion use the life cycle conversion function of the object storage, which is automatically executed without manual triggering. The retrieval unit processes the log data and creates an index using word segmentation. Based on the word segmentation, the full text of the log is divided into multiple segments, and based on the index, an association with the original log is established.