Processing System and Method for Aggregated Records with Duplicate Exclusion Markers
A system with deduplication markers and processing units addresses the inefficiencies in managing large log data volumes by reducing storage needs and enhancing query speed through deduplication and processing techniques.
Patent Information
- Application Number
- JP2025501275
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-10
- Filing Date
- 2023-07-10
- Publication Date
- 2025-07-10
AI Technical Summary
The accumulation of large volumes of computer log data requires significant storage space and leads to inefficient information queries, reducing the value and usefulness of the data due to hardware resource burden and slow processing times.
Implementing a system with a log generation unit, processing unit, and storage unit to deduplicate log data, reducing its volume through methods like in-line and historical processing, and using deduplication markers to identify and process key-value pairs, thereby improving storage efficiency and query speed.
The system effectively reduces storage needs and enhances query performance by minimizing redundant data, optimizing hardware utilization and query efficiency.
Smart Images

Figure 2025522028000001_ABST
Abstract
Description
Technical Field
[0001] (Related Application) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 359,877, filed Jul. 10, 2022. The content of each application cited in this paragraph is incorporated herein by reference in its entirety as if fully set forth herein.
[0002] This disclosure relates to systems and methods for processing aggregated records with deduplication markers, such as particularly aggregated computer log records, to improve the storage and analysis efficiency of computer systems.
Background Art
[0003] With the spread of "big data" analysis and related applications, data collection has become extremely important. For example, in almost all types of computer-related activities, log records have become common, and a large amount of computer log data has been generated. Organizations around the world are accumulating computer log data extensively in an attempt to extract value from it. Therefore, a large amount of storage space is required to store computer log data. Furthermore, queries against huge amounts of computer log data are time-consuming, reducing the value and usefulness of such data.
[0004] With the advancement of big data analysis, many organizations have started logging and collecting data on a wide range of user activities, such as network connection activities, login / logout activities to systems, etc. Such extensive log record operations require a large amount of storage capacity and impose a heavy burden on hardware resources. Furthermore, due to the huge amount of data, information queries become slow and inefficient, significantly reducing the value and usefulness of the data. It is difficult to reduce the capacity of log data while maintaining a high level of data integrity.
[0005] Embodiments of a system and method for reducing the volume while retaining information contained in computer log data are described in detail in U.S. Patent No. 10,877,972 to Althouse ("Althouse"), the content of which is incorporated herein by reference as if fully set forth herein.
[0006] Althouse describes various systems and methods for managing computer log data by removing redundant information from the original data and performing deduplication. Exemplary systems and methods can improve the storage efficiency of computer log data and the speed of information retrieval. For example, embodiments of the present disclosure can identify duplicate data fields across multiple log records and combine the multiple log records into an integrated log record. The integrated log record can include a single copy of the duplicate content and an integrated version of the non-duplicate content extracted from the multiple log records. In this way, the amount of computer log data that needs to be stored can be significantly reduced.
[0007] Log data may be organized as a set of log records. As used herein, "log record" is also referred to as "log entry" or simply "record" and refers to a unit of information recorded at a certain point in time or corresponding to a timestamp. A log record can include multiple data fields, for example, a timestamp, an ID indicating the user for whom the log is recorded, events / actions / results related to the situation of the log record, and the like. SUMMARY OF THE INVENTION
[0008] FIG. 1 shows an exemplary system 100 for managing log data according to an embodiment of the present disclosure. As shown in FIG. 1, system 100 may include a log generation unit 110, a log processing unit 120, a log storage unit 130, and a log query unit 140. The log generation unit 110 may include any suitable computer or device configured to generate log data.
[0009] The exemplary log generation unit 110 may include a computer programmed to record login / logout activities of log usage, a network connection computer programmed to record network connection activities, a mobile device equipped with a geographical location sensor programmed to record location information, a payment device configured to record financial transactions, an activity tracking device configured to record steps, heart rate, blood pressure, etc., or any other device capable of recording information in a periodic or event-triggered manner.
[0010] The log storage unit 130 may include any suitable computer, server, database, and service platform (e.g., Splunk developed by Splunk, Elastic Stack [also called ELK Stack] developed by Elasticsearch B.V., etc.) configured / programmed to store log data. In some embodiments, the log storage unit 130 may be implemented as a separate component from the log generation unit 110. For example, the log storage unit 130 may be implemented on a different computer from the log generation unit 110. In some embodiments, the log storage unit 130 may be integrated with the log generation unit 110. For example, the functions of the log generation unit 110 and the log storage unit 130 may be implemented by a software package installed on the same computer.
[0011] The log processing unit 120 may include a computer system programmed to process log data created by the log generation unit 110 and / or log data stored in the log storage unit 120. FIG. 1 shows some examples of the flow of log data among the log generation unit 110, the log processing unit 120, and the log storage unit 130. As shown in FIG. 1, the log data generated by the log generation unit 110 (also referred to as original log data) may be directly transmitted along path A to the log storage unit 130 for storage. After the log storage unit 130 receives the original log data from the log generation unit 110, the log storage unit 130 may store the original log data without invoking the log processing unit 120. Thereafter, the log processing unit 120 may, for example, obtain the original log data from the log storage unit 130 along path D, process the original log data to reduce the size of the original log data using the method disclosed in Althouse, and send the processed log data back to the log storage unit 130 along path E to replace the original log data stored in the log storage unit 130. In this way, the log processing unit 120 can reduce the size of existing log data stored in any log storage system. Hereinafter, the log processing by this method (e.g., the flow of log data along paths D and E) is also referred to as history log data processing.
[0012] In some embodiments, the log processing unit 120 may be configured to process the original log data generated by the log generation unit 110 before the original log data reaches the log storage unit 130. For example, the log processing unit 120 may be disposed between the log generation unit 110 and the log storage unit 130. In other words, the log processing unit 120 may be disposed upstream of the log storage unit 130 to process the input source log data generated by the log generation unit 110 before it reaches the log storage unit 130. As shown in FIG. 1, the original log data may be transmitted from the log generation unit 110 to the log processing unit 120 along path B. After the log processing unit 120 processes the original log data (e.g., reduces the data size), the processed log data may be transmitted from the log processing unit 120 to the log storage unit 130 along path C. Hereinafter, the log processing in such a manner (e.g., the flow of log data along paths B and C) is also referred to as in-line log data processing.
[0013] The log query unit 140 may include any suitable computer, terminal, mobile device, etc. programmed to query the log data stored in the log storage unit 130. In some embodiments, the log query unit 140 can query the log storage unit 130 directly along path J. When the log storage unit 130 stores the processed log data processed by the log processing unit 120, the capacity or size of the processed log data is significantly smaller than the capacity or size of the original log data, whereby queries can be executed more quickly across the processed log data, resulting in an improved data query speed. The result of the query (e.g., one or more processed log records that meet the query criteria) may be returned from the log storage unit 130 to the log query unit 140 along path K. In some embodiments, the log query unit 140 may query the log data via the log processing unit 120. For example, a request to obtain log records that meet specific search criteria may be sent from the log query unit 140 to the log processing unit 120 along path F. The log processing unit 120 may then send a request to the log storage unit 130 from the log processing unit 120 along path G to identify the log records. In this case, the log storage unit 130 may or may not save the log data in a processed form. For example, the log storage unit 130 may store the original log data or a combination of the original log data and the processed log data. In either case, the log storage unit 130 may obtain the log data that meets the query criteria and send it to the log processing unit 120 along path H. The log processing unit 120 may process the received log data as needed and send the processed log data to the log query unit 140 along path I.
[0014] In some embodiments, the log processing unit 120 may be implemented as a stand-alone device including its dedicated computing resources, and may interface with other components of the system 100 (e.g., 110, 130, 140) via a communication channel (e.g., network connection, direct cable connection, etc.). For example, the log processing unit 120 may be implemented on a server, and the functions of the log processing unit 120 may be implemented as SaaS (Software as a Service). In some embodiments, the log processing unit 120 may be implemented as an add-on device that can be integrated or added to one or more other components of the system 100 (e.g., 110, 130, 140). For example, the log processing unit 120 may be implemented on a smart card, a system-on-chip (SoC), or other forms of add-on devices or plug-in devices that can be integrated with the log generation unit 110, the log storage unit 130, and / or the log query unit 140. In some embodiments, the log processing unit 120 may use the hardware resources of one or more other components of the system 100 (e.g., 110, 130, 140). For example, the log processing unit 120 may be implemented as a software program that can be installed in the log generation unit 110 and / or the log storage unit 130 to perform historical and / or in-line log data processing. In some embodiments, the log processing unit 120 may be implemented as a combination of the above forms. For example, the log processing unit 120 may be implemented as SaaS with a plug-in program that operates as a communication portal installed in the log generation unit 110, the log storage unit 130, and / or the log query unit 140.
[0015] For consumers of deduplicated log files, such as providers of security information and event management (SIEM) services and products, systems and methods for processing received aggregated records are needed.
[0016] These and other additional features and advantages of the present invention will be apparent to those skilled in the art from the following detailed description when taken in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Best Mode for Carrying Out the Invention
[0018] Considering together with the accompanying drawings, the features and advantages of various exemplary embodiments will become apparent from the following detailed description. Where possible, the same reference numerals and characters are used to indicate similar features, elements, components or parts of the embodiments of the present disclosure. Also, it is intended that changes and modifications can be made to the exemplary embodiments described and illustrated herein without departing from the true scope and spirit of the embodiments of the invention described herein as defined by the claims.
[0019] Throughout this specification, the exemplified preferred embodiments and examples are, rather, not limitations on the present invention, but should be considered as examples. As used herein, the terms "invention", "apparatus", "method", "disclosure", "present invention", "present apparatus", "present method", or "present disclosure" refer to any of the embodiments of the invention described herein, and their equivalents. Further, references to the various features of "invention", "apparatus", "method", "disclosure", "present invention", "present disclosure", "present apparatus", "present method", or "present disclosure" do not mean that all claimed embodiments or methods must include the referenced features.
[0020] Also, when an element or feature is described as being "above" or "adjacent" to another element or feature, it may be directly "above" or "adjacent" to the other element or feature, or there may be intervening elements or features. Also, when an element is expressed as being "attached to", "connected to", or "coupled to" another element, it may be directly attached to, connected to, or coupled to the other element, or there may be another element present therebetween. In contrast, when an element is expressed as being "directly attached to", "directly connected to", or "directly coupled to" another element, there is no other element therebetween.
[0021] In this specification, relative terms such as "outer", "above", "below", "lower", "horizontal", "vertical", etc. may be used to describe the relationship between one feature and another feature. It is understood that these terms are intended to encompass different directions in addition to the directions depicted in the figures.
[0022] In this specification, terms such as "first", "second", etc. may be used to describe various elements, components, or steps, but these elements, components, or steps should not be limited by these terms. These terms are only used to distinguish one element, component, or step from another element, component, or step. Therefore, the first element or component described below can be referred to as the "second element or component" without departing from the teachings of the present disclosure. As used in this specification, the term "and / or" includes all combinations of one or more of the associated listed items arbitrarily.
[0023] The terms used in this specification are for the purpose of describing particular embodiments and are not intended to limit the invention. As used herein, the singular form "one" is intended to include the plural as well, unless the context clearly dictates otherwise. Further, as used herein, the terms "consisting of," "comprising," "having," and the like are used to specify the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0024] In computing, data deduplication is a technique for eliminating duplicate copies of repeating data. By properly implementing this technique, storage utilization efficiency can be improved, and thus capital expenditure can be reduced by reducing the total amount of storage media required to meet storage capacity needs. Also, this technique can be applied to network data transfer to reduce the number of bytes to be transmitted.
[0025] A deduplication marker is an identifier indicating the duplication of data records, such as a count of aggregated log records, as detailed, for example, in Althouse U.S. Patent No. 10,877,972, which is incorporated herein by reference. Embodiments of the present disclosure include systems and methods for processing aggregated records based on deduplication markers.
[0026] Some logging tools count and act on the number of key-value pairs within log records. For example, if a key-value pair of "domain=google.com" occurs in 100 lines of logs, the tool indicates that "domain=google.com" has occurred 100 times.
[0027] FIG. 2 is a flowchart of a method for processing an aggregated record 200 according to an embodiment of the present disclosure. At 202, an aggregated record is received. Key-value pairs are identified within the aggregated record, the key-value pairs being composed of a key and a value, the key being composed of a deduplication marker 204. An aggregation method is identified based on the characteristics of the deduplication marker 206. The aggregated record is processed using an algorithm determined by the aggregation method, and at least one processed record 208 is generated.
[0028] For example, in an aggregated log record, "domain=google.com" may be shown in one log line along with a deduplication marker value 100. This means that the key-value pair occurred 100 times but was saved only once to save storage capacity and improve processing efficiency.
[0029] The log processing tool needs to identify the deduplication marker by name, identifier, or configuration. To maintain the integrity of the data within these logs, the tool then needs to process the key-value pairs of interest according to an aggregation method that can be identified by name, identifier, or configuration. Table 1 below shows an exemplary set of non-aggregated log records.
Table 1
[0030] The example of an aggregated record shown in Table 2 below may be generated according to an embodiment of the method disclosed herein.
Table 2
[0031] In this example, in the key-value pair of the deduplication marker "logslash=8", the key is the field name "logslash", which is in this case the deduplication marker, and the value is 8. Logs may be stored in this format to save space and reduce search time and processing costs. However, for the purposes of report generation, graphing, or processing, the target key-value pairs can be expanded using an algorithm determined by the original aggregation method indicated by the deduplication marker.
[0032] For key-value pairs without additional identifiers, the default processing is to duplicate the specific key-value pairs by the deduplication marker value. In the above example, since "domain=mail.google.com" (the target key-value pair) is displayed in the log line along with "logslash=8" (the key-value pair of the deduplication marker) indicating a value of 8, the system can duplicate "domain=mail.google.com" 8 times within the aggregation window for reporting, graphing, or processing.
[0033] For target key-value pairs with additional identifiers in the key, such as "-total", a different method of processing the target key-value pairs relative to the deduplication marker may be required. In this example, "bytes-total=3365kb" indicates that the byte values from the original unaggregated records have been totaled. For the purposes of report generation, graphing, or processing, the system can divide the total value by the deduplication marker value, thereby reporting or graphing that each log line within the aggregation time window averages approximately 420kb (3365kb ÷ 8 connections ≈ 420kb / connection).
[0034] The log file deduplication technology can significantly reduce the processing volume required to train an artificial intelligence ("AI") model on data, and thus the equipment, power, space, etc. To do this, the AI model needs to consider the deduplication technology being used and how to process its deduplication markers. For example, the model can weight values based on the deduplication markers. For example, if "logslash=100" is included in the log containing "domain=google.com", the model adds 100 to the weight of "domain=google.com" in that specific time frame without processing 100 logs.
[0035] The various exemplary embodiments of the invention described herein are merely intended to illustrate the principles underlying the inventive concept. Accordingly, various modifications of the disclosed embodiments will be apparent to those skilled in the art without departing from the spirit and scope of the invention. They are not intended to limit the various exemplary embodiments of the invention to the exact forms described. Given the above teachings, other variations and embodiments of the invention are possible, and the scope of the invention is not limited by this specification, but rather is intended to be limited by the claims that follow this specification.
[0036] Although the present invention has been described in detail with reference to its particular preferred configuration, other versions are possible. Embodiments of the present invention can be configured with any combination of compatible features shown in various figures, and these embodiments should not be limited to those explicitly illustrated and described. Therefore, the spirit and scope of the present invention should not be limited to the above versions. Furthermore, the combinations of functions, elements, and steps recited in the appended claims can be combined with each other as if they were recited in a plurality of dependent forms and were all dependent from all of the preceding claims. The various combinations of devices, components, and steps described above and in the appended claims are within the scope of the present disclosure. The foregoing is intended to cover all modifications and alternative structures included within the spirit and scope of the present disclosure.
Claims
1. A method for processing aggregated records, comprising: receiving an aggregated record; identifying pairs of key values within the aggregated record, wherein the pairs of key values consist of a key and a value, and the key consists of a deduplication marker; identifying an aggregation method based on the characteristics of the deduplication marker; processing the aggregated record using an algorithm determined by the aggregation method to generate at least one processed record.
2. The method according to claim 1, wherein the aggregated record is a computer log file.
3. The method according to claim 1, further comprising displaying the at least one processed record.
4. The method according to claim 1, further comprising training an artificial intelligence model using the at least one processed record.
5. The method according to claim 1, wherein the characteristics of the deduplication marker are a field name, an identifier, or a configuration.
6. The method according to claim 1, wherein the at least one processed record is composed of a plurality of identical records that duplicate based on the value of the pair of key values.
7. The method according to claim 1, wherein the pair of key values is a pair of deduplication marker key values.
8. The method according to claim 1, wherein the pair of key values is a pair of target key values.
9. A system for processing aggregated records, comprising: a memory storing computer-readable instructions; at least one processor communicatively coupled to the memory; wherein the computer-readable instructions, when executed by the at least one processor, cause the at least one processor to perform operations, and the instructions include: receiving an aggregated record; identifying pairs of key values within the aggregated record, wherein the pairs of key values consist of a key and a value, and the key consists of a deduplication marker; identifying an aggregation method based on the characteristics of the deduplication marker; processing the aggregated record using an algorithm determined by the aggregation method to generate at least one processed record.
10. The system according to claim 9, wherein the aggregated record is a computer log file.
11. The system according to claim 9, further comprising the step of displaying the at least one processed record.
12. The system according to claim 9, further comprising the step of training an artificial intelligence model using the at least one processed record.
13. The system according to claim 9, wherein the characteristic of the deduplication marker is a field name, an identifier, or a configuration.
14. The system according to claim 9, wherein the at least one processed record includes a plurality of identical records that duplicate based on the values of the key-value pairs.
15. The system according to claim 9, wherein the key-value pair is a deduplication marker key-value pair.
16. The system according to claim 9, wherein the key-value pair is a target key-value pair.
17. A set of machine-readable instructions incorporated in a tangible medium for causing a machine to execute steps, wherein the steps are receiving an aggregated record; identifying a key-value pair in the aggregated record, the key-value pair consisting of a key and a value, the key consisting of a deduplication marker; identifying an aggregation method based on the characteristic of the deduplication marker; processing the aggregated record using an algorithm determined by the aggregation method to generate at least one processed record.