Log fragment compression storage method and device, electronic equipment and medium

By segmenting and compressing pipeline task logs, and utilizing the sliding window algorithm and adaptive encoding, the problem of log repetition patterns occupying storage space is solved, achieving efficient storage resource utilization and data retrieval.

CN120821710AActive Publication Date: 2025-10-21SHANGHAI SUIYUAN TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511324175.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-10-21
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

In existing technologies, logs output by similar tasks, toolchains, or runtime environments often exhibit numerous repetitive patterns and dynamic variables, leading to high storage costs, wasted storage resources, and low data retrieval efficiency.

Method used

By acquiring incremental log files of pipeline tasks in real time, preprocessing and segmenting the log data, and combining the sliding window algorithm and adaptive encoding processing, vertical and horizontal log fragments are generated and stored in the log fragment compression storage database. The encoding method is dynamically updated to adapt to frequency changes.

Benefits of technology

It improves the utilization rate of storage resources, reduces storage costs, reduces resource waste, and improves data retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821710A_ABST
    Figure CN120821710A_ABST
Patent Text Reader

Abstract

The invention discloses a log fragment compression storage method and device, electronic equipment and a medium. The method comprises the steps of obtaining a current task increment log file corresponding to a current pipeline task in real time; generating each current standard to-be-fragmented compressed log according to each current to-be-fragmented compressed log, and performing log partitioning in combination with a dynamic time window strategy to generate at least one current to-be-fragmented compressed log block; sequentially carrying out fragmentation longitudinal extraction on each current to-be-fragmented compressed log block through a sliding window algorithm to obtain each longitudinal log fragment and a non-longitudinal fragmentation residual log, and carrying out fragmentation transverse extraction on the non-longitudinal fragmentation residual log in combination with a transverse log fragmentation processing method to obtain each transverse log fragment; and carrying out coding processing through a self-adaptive coding processing method, and storing in a log fragment compression storage database. The problem that the current repeated mode log occupies a large amount of storage space is solved, the utilization rate of storage resources is improved, and the storage cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a log sharding compression storage method, device, electronic device, and medium. Background Art

[0002] During the execution of data processing pipeline tasks, massive amounts of operation logs are generated. These logs can come from: logs output by user-written business scripts; logs output by compilers, build tools, or test frameworks; and logs output by the underlying operating platform or environment. These logs are typically stored in raw text format on log servers.

[0003] While implementing this invention, the inventors discovered the following drawbacks in the existing technology: Currently, for tasks of the same type, using the same toolchain, or operating environment, the log content output contains a large number of highly similar or even repetitive patterns (such as fixed log frameworks, error messages, and environment information templates) and dynamic variables (such as timestamps, file paths, and numerical values). These repetitive patterns occupy a large amount of storage space, resulting in high storage costs, waste of storage resources, and inefficient data retrieval. Summary of the Invention

[0004] The present invention provides a log sharding compression storage method, device, electronic equipment and medium to improve the utilization rate of storage resources and reduce storage costs.

[0005] According to one aspect of the present invention, a log fragment compression storage method is provided, which includes:

[0006] Obtaining in real time the current task incremental log file corresponding to the current pipeline task; wherein the current task incremental log file includes at least one log to be currently fragmented and compressed;

[0007] Based on each of the current to-be-sharded compressed logs, a preset log data preprocessing method is used to process the logs to generate current standard to-be-sharded compressed logs, and a preset dynamic time window strategy is used to perform log segmentation to generate at least one current to-be-sharded compressed log segment corresponding to each of the current standard to-be-sharded compressed logs.

[0008] Using a sliding window algorithm, perform vertical extraction on each currently sharded and compressed log block in turn to obtain vertical log shards and the remaining logs that have not been vertically sharded. Then, using a pre-set horizontal log sharding method, perform horizontal extraction on the remaining logs that have not been vertically sharded to obtain horizontal log shards.

[0009] Through a pre-set adaptive encoding processing method, each vertical log segment and each horizontal log segment are encoded and stored in a log segment compression storage database;

[0010] Among them, the log fragment compression storage database corresponds to different encoding methods according to the frequency of occurrence of different log fragments; when the frequency of occurrence of log fragments changes, it is necessary to periodically update the encoding method of the log fragment compression storage database and generate a new encoding version.

[0011] According to another aspect of the present invention, a log fragment compression storage device is provided, comprising:

[0012] The current task incremental log file acquisition module is used to obtain the current task incremental log file corresponding to the current pipeline task in real time; wherein, the current task incremental log file includes at least one current log to be fragmented and compressed;

[0013] The current to-be-sharded compressed log block generation module is configured to replace each of the current to-be-sharded compressed logs using a pre-set regular matching unified placeholder replacement method to generate each current standard to-be-sharded compressed log, and perform log block generation in combination with a pre-set dynamic time window strategy to generate at least one current to-be-sharded compressed log block corresponding to each of the current standard to-be-sharded compressed logs;

[0014] The vertical log sharding and horizontal log sharding determination module is used to perform vertical sharding extraction on each currently sharded and compressed log block in sequence using a sliding window algorithm to obtain vertical log shards and remaining logs that have not been vertically sharded. Furthermore, the module performs horizontal sharding extraction on the remaining logs that have not been vertically sharded using a pre-set horizontal log sharding processing method to obtain horizontal log shards.

[0015] A storage module is used to encode each vertical log segment and each horizontal log segment using a preset adaptive encoding method, and store the encoded data in a log segment compression storage database;

[0016] Among them, the log fragment compression storage database corresponds to different encoding methods according to the frequency of occurrence of different log fragments; when the frequency of occurrence of log fragments changes, it is necessary to periodically update the encoding method of the log fragment compression storage database and generate a new encoding version.

[0017] According to another aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the log sharding compression storage method according to any embodiment of the present invention is implemented.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the log fragmentation compression storage method described in any embodiment of the present invention when executed.

[0019] The technical solution of the embodiment of the present invention is to obtain the current task incremental log file corresponding to the current pipeline task in real time; wherein, the current task incremental log file includes at least one current log to be fragmented and compressed; according to each current log to be fragmented and compressed, a preset log data preprocessing method is used to process it to generate each current standard log to be fragmented and compressed, and a preset dynamic time window strategy is used to perform log segmentation to generate at least one current log segment to be fragmented and compressed block corresponding to each current standard log to be fragmented and compressed; through a sliding window algorithm, each current log segment to be fragmented and compressed is sequentially subjected to segmented vertical extraction to obtain each vertical log segment and the remaining log that is not fragmented vertically, and a preset horizontal log segmentation processing method is used to perform segmented horizontal extraction on the remaining log that is not fragmented vertically to obtain each horizontal log segment; through a preset adaptive encoding processing method, each vertical log segment and each horizontal log segment are respectively encoded and processed, and stored in a log segmentation compression storage database.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 This is a flow chart of a log sharding compression storage method provided according to the first embodiment of the present invention;

[0023] Figure 2 This is a detailed flow chart of a log sharding compression storage method provided according to the second embodiment of the present invention;

[0024] Figure 3 This is a structural diagram of a log fragmentation compression storage device provided according to the third embodiment of the present invention;

[0025] Figure 4It is a structural diagram of an electronic device provided according to the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "target", "current", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] Example 1

[0029] Figure 1 A flowchart of a log sharding and compression storage method is provided for the first embodiment of the present invention. In this embodiment, in the current pipeline task, for the case where sharding and compression storage are performed on logs in a repeated pattern, the method can be executed by a log sharding and compression storage device, which can be implemented in the form of hardware and / or software.

[0030] Correspondingly, such as Figure 1 As shown, the method includes:

[0031] S110 , obtaining a current task incremental log file corresponding to the current pipeline task in real time.

[0032] The current task incremental log file includes at least one log to be currently fragmented and compressed.

[0033] Optionally, the real-time acquisition of the current task incremental log file corresponding to the current pipeline task includes: monitoring the status of the current pipeline task in real time through the log sharding compression storage system; when it is detected that the status of the current pipeline task is the end state, obtaining the current incremental log file corresponding to the current pipeline task.

[0034] In this embodiment, a log sharding and compression storage system is required to monitor the status of the current pipeline task in real time. If the task is in the Continue Running state, the monitoring operation needs to be continued. When the state is the End state, it indicates that the current pipeline task has been processed and the corresponding current incremental log file has been generated. The data storage operation of the current incremental log file is required, so the log sharding and compression storage operation is required.

[0035] Among them, for the current task incremental log file, it may include one or more current sharded and compressed logs, and it is necessary to perform storage processing operations on the one or more current sharded and compressed logs.

[0036] S120. According to each of the current to-be-sharded compressed logs, a preset log data preprocessing method is used for processing to generate each current standard to-be-sharded compressed log, and log segmentation is performed in combination with a preset dynamic time window strategy to generate at least one current to-be-sharded compressed log segment corresponding to each of the current standard to-be-sharded compressed logs.

[0037] The log data preprocessing method may include a processing method that normalizes the log data and performs a regular matching unified placeholder replacement operation.

[0038] The current standard log to be segmented and compressed may be a log obtained after preprocessing the log data. The dynamic time window strategy may be a strategy for segmenting the log according to the timestamp of the log entry.

[0039] Optionally, the current standard logs to be sliced ​​and compressed are processed by a preset log data preprocessing method to generate current standard logs to be sliced ​​and compressed, and log blocks are performed in combination with a preset dynamic time window strategy to generate at least one current standard log to be sliced ​​and compressed block corresponding to each current standard log to be sliced, including: normalizing the current standard logs to be sliced, identifying the dynamic variables in the current standard logs to be sliced, and replacing them by a preset regular matching unified placeholder replacement method to generate current standard logs to be sliced; obtaining a target current standard log to be sliced ​​in each current standard log to be sliced ​​in turn; obtaining each log entry in the target current standard log to be sliced ​​through the dynamic time window strategy. Timestamp, and obtain log block time threshold; calculate the time interval corresponding to consecutive log entries according to the timestamp of each log entry; determine whether the time interval is greater than the log block time threshold; if so, perform log block on the target current standard to be sliced ​​and compressed log, continue to traverse the log entries in the un-logged log, and return to execute the operation according to the timestamp of each log entry until all log entries in the target current standard to be sliced ​​and compressed log are traversed; determine whether all current standard to be sliced ​​and compressed logs are traversed, and if so, generate at least one current to be sliced ​​and compressed log block corresponding to each current standard to be sliced ​​and compressed log; if not, return to execute the operation of obtaining one target current standard to be sliced ​​and compressed log in turn in each current standard to be sliced ​​and compressed log.

[0040] In this embodiment, it is first necessary to normalize each current log to be sharded and compressed to obtain the normalized log to be sharded and compressed. Further, the dynamic variables in each current log to be sharded and compressed are identified to determine whether the dynamic variables exist. If so, it is necessary to perform corresponding dynamic variable replacement operations according to the regular matching table corresponding to the regular matching unified placeholder replacement method to generate each current standard log to be sharded and compressed. Among them, the regular matching table may include the dynamic variable type and the regular matching unified placeholder corresponding to the dynamic variable type. Accordingly, assuming that there is no dynamic variable, each current standard log to be sharded and compressed is directly generated.

[0041] Furthermore, a dynamic time window strategy is needed to obtain the timestamps of each log entry in the target current standard log to be sharded and compressed, as well as the log chunking time threshold. The target current standard log to be sharded and compressed contains multiple log entries, which are arranged in chronological order. Accordingly, the time interval corresponding to consecutive log entries needs to be calculated based on the timestamps of each log entry.

[0042] For example, assume the target standard log to be sharded and compacted currently consists of 100 log entries. Log entries 1-10 are timestamped at 12:12:10 on September 1, 2025; log entries 11-30 are timestamped at 12:12:11 on September 1, 2025; log entries 31-41 are timestamped at 12:12:15 on September 1, 2025; and log entries 42-50 are timestamped at 12:12:30 on September 1, 2025. It can be calculated that the time interval between log entries 10 and 11 is 1 second; the time interval between log entries 30 and 31 is 4 seconds; and the time interval between log entries 41 and 42 is 15 seconds. Assume the log chunking time threshold is 10 seconds.

[0043] Because the time interval between the 41st and 42nd log entries is 15 seconds, which exceeds the log chunking time threshold, log chunking is performed between the 41st and 42nd log entries in the target standard log to be chunked and compressed. The system then continues to traverse the remaining 50 log entries, performing the above operations.

[0044] Furthermore, it is necessary to perform log block operations on all current standard logs to be segmented and compressed, and then obtain at least one current log block to be segmented and compressed corresponding to each current standard log to be segmented and compressed.

[0045] The advantage of this setting is that it can better assist in implementing log sharding operations, improve log sharding efficiency, and improve log data processing efficiency by preprocessing log data and segmenting each current standard compressed log to be segmented.

[0046] S130. Using a sliding window algorithm, perform vertical extraction on each of the currently to-be-sharded and compressed log blocks in turn to obtain vertical log shards and the remaining logs that are not vertically sharded. Furthermore, perform horizontal extraction on the remaining logs that are not vertically sharded in combination with a pre-set horizontal log sharding processing method to obtain horizontal log shards.

[0047] The sliding window algorithm may be an algorithm that can perform vertical log shard extraction operations through sliding windows of different lengths.

[0048] Specifically, vertical log sharding can be performed on logs generated by the current pipeline task of the same type. Because these logs have a high degree of commonality, the shards for identical logs are relatively long. Horizontal log sharding can be performed on logs generated by different types of tasks within the pipeline, combining the remaining logs after vertical task extraction to perform horizontal log sharding. Specifically, horizontal logs have relatively little commonality and are more fragmented.

[0049] Optionally, the method uses a sliding window algorithm to sequentially perform vertical extraction on each current log block to be segmented and compressed, and obtains each vertical log segment and the remaining logs that are not segmented vertically, including: obtaining a preset minimum segment length threshold; using a pre-trained log entry deep learning classification model, sequentially classifying each current log to be segmented and compressed to obtain a log entry type; obtaining a target current log block to be segmented and compressed in each current log block to be segmented and compressed; determining the maximum sliding window by using a sliding window algorithm, and sliding the target current log block to be segmented and compressed according to the maximum sliding window, and judging the current sliding window in combination with the log entry type. Whether the log includes a unified placeholder, if so, extract the placeholder template structure and generate the current placeholder vertical log fragment; if not, perform long string pattern matching on the current sliding window log, if it matches, confirm that the current vertical log fragment is identified; reduce the maximum sliding window, and determine whether it is greater than or equal to the minimum fragment length threshold, if so, return to execute the operation of sliding the target current to-be-fragmented compressed log block according to the maximum sliding window; if not, determine that the vertical fragment extraction is completed, obtain the remaining log that has not been vertically fragmented, and obtain each vertical log fragment according to each current vertical log fragment and each current placeholder vertical log fragment.

[0050] The log entry type can be a column type obtained by classifying the log to be sharded and compressed. Specifically, log entry types include business logs, platform logs, or test framework logs. Different log entry types include different placeholder template structures and different long string patterns.

[0051] In this embodiment, a sliding window algorithm is first used to determine the maximum sliding window, assuming that the maximum sliding window is L. This sliding window is used to slide the target log block to be sharded and compressed. Based on the type of log entries, a determination is made as to whether the current sliding window log contains a unified placeholder. If so, the placeholder template structure is first extracted to generate the current placeholder vertical log shard.

[0052] Furthermore, if it does not exist, it is necessary to perform long string pattern matching on the current sliding window log. If it matches, it is confirmed that the current vertical log shard has been identified; if there is no match, the maximum sliding window L needs to be reduced. Specifically, the length can be reduced to L-1.

[0053] The sliding operation on the target log block to be sharded and compressed continues through a sliding window of length L-1 until the sliding window length is less than the minimum shard length threshold.

[0054] Accordingly, each current vertical log slice and each current placeholder vertical log slice may be combined to obtain each vertical log slice.

[0055] In this embodiment, for the remaining logs that have not been vertically sharded, it is necessary to use the horizontal log sharding processing method to perform processing operations. Specifically, the horizontal log sharding processing method includes the longest common subsequence processing method, and the specific operations are as follows:

[0056] Optionally, the method combines a preset horizontal log sharding processing method to perform sharded horizontal extraction on the remaining logs that are not vertically sharded to obtain various horizontal log shards, including: obtaining a preset minimum length allowable threshold and a minimum sharding occurrence frequency; obtaining an initial horizontal shard length through the longest common subsequence processing method, and performing sharded horizontal extraction on the remaining logs that are not vertically sharded; wherein, the initial horizontal shard length is greater than the minimum length allowable threshold; if the extraction is successful, and the occurrence frequency of the extracted horizontal log shard is not lower than the minimum sharding occurrence frequency, then the current horizontal log shard is obtained, and the initial horizontal shard length is reduced to determine whether it is greater than or equal to the minimum length allowable threshold. If so, then return to executing the operation of obtaining the initial horizontal shard length through the longest common subsequence processing method and performing sharded horizontal extraction on the remaining logs that are not vertically sharded; after traversing the remaining logs that are not vertically sharded, various horizontal log shards are obtained.

[0057] In this embodiment, it is necessary to obtain the initial horizontal shard length through the longest common subsequence processing method. Here, the initial horizontal shard length is required to be greater than the minimum length allowed threshold, and then the remaining logs that have not been vertically sharded are subjected to horizontal shard extraction. If the extraction is successful, it is also necessary to determine whether the frequency of occurrence of the horizontal log shard is greater than or equal to the minimum frequency of occurrence of the shard. If so, the current horizontal log shard can be obtained. Assuming that the frequency of occurrence of the extracted horizontal log shard is lower than the minimum frequency of occurrence of the shard, it means that the extracted horizontal log shard cannot obtain a greater compression benefit, that is, the storage memory required is large, and the data compression processing operation cannot be performed well. Therefore, the frequency of occurrence of the extracted horizontal log shard needs to be considered.

[0058] Furthermore, the initial horizontal shard length needs to be reduced, but it cannot be less than the minimum length threshold. Generally, the minimum length threshold can be set to half the minimum shard length threshold. Furthermore, the remaining logs that have not been vertically sharded are horizontally extracted based on the different initial horizontal shard lengths to obtain various horizontal log shards.

[0059] In addition, in this embodiment, for the remaining logs that have not been vertically sharded, it is necessary to perform processing operations using a horizontal log sharding processing method. Specifically, the horizontal log sharding processing method can also include a dynamic programming processing method or a greedy algorithm processing method, and the specific operations are as follows: first, it is necessary to perform vectorized processing operations on the remaining logs that have not been vertically sharded, and then further perform similarity calculations with templates in a pre-built clustering sharding representative template library using a dynamic programming processing method or a greedy algorithm processing method, in combination with a pre-set minimum length allowed threshold and a minimum sharding frequency, and then generate corresponding horizontal log shards.

[0060] The advantage of this setting is that it can better implement horizontal sharding processing operations on the remaining logs that have not been vertically sharded, thereby better realizing the reasonable storage of log shards and reducing the storage pressure and storage costs of the database.

[0061] S140 , encoding each vertical log segment and each horizontal log segment respectively using a preset adaptive encoding method, and storing the encoded data in a log segment compression storage database.

[0062] Among them, the log fragment compression storage database corresponds to different encoding methods according to the frequency of occurrence of different log fragments; when the frequency of occurrence of log fragments changes, it is necessary to periodically update the encoding method of the log fragment compression storage database and generate a new encoding version.

[0063] Specifically, the adaptive coding processing method can be set to a Huffman coding processing method.

[0064] In this embodiment, after obtaining each vertical log segment and each horizontal log segment, it is necessary to perform reasonable encoding and storage operations according to an adaptive encoding processing method, and store them in a log segment compression storage database.

[0065] The technical solution of the embodiment of the present invention is to obtain the current task incremental log file corresponding to the current pipeline task in real time; according to each of the current to-be-sliced ​​compressed logs, process them through a pre-set log data preprocessing method to generate each current standard to-be-sliced ​​compressed log, and perform log segmentation in combination with a pre-set dynamic time window strategy to generate at least one current to-be-sliced ​​compressed log segment corresponding to each of the current standard to-be-sliced ​​compressed logs; through a sliding window algorithm, perform segmented vertical extraction on each current to-be-sliced ​​compressed log segment in turn to obtain each vertical log segment and the remaining log that is not segmented vertically, and perform segmented horizontal extraction on the remaining log that is not segmented vertically in combination with a pre-set horizontal log segmentation processing method to obtain each horizontal log segment; through a pre-set adaptive encoding processing method, encode each vertical log segment and each horizontal log segment respectively, and store them in a log segmentation compression storage database. This solves the problem that the current repeated pattern log occupies a large amount of storage space, improves the utilization rate of storage resources, reduces storage costs, reduces the waste of storage resources, and improves data retrieval efficiency.

[0066] Example 2

[0067] Figure 2 A detailed flow chart of a log shard compression storage method is provided for the second embodiment of the present invention. This embodiment is based on the above embodiments and is further refined by encoding each vertical log shard and each horizontal log shard using a pre-set adaptive encoding processing method and storing them in a log shard compression storage database.

[0068] S210: Obtain the current task incremental log file corresponding to the current pipeline task in real time.

[0069] The current task incremental log file includes at least one log to be currently fragmented and compressed.

[0070] S220. According to each of the current to-be-sharded compressed logs, a preset log data preprocessing method is used for processing to generate each current standard to-be-sharded compressed log, and log segmentation is performed in combination with a preset dynamic time window strategy to generate at least one current to-be-sharded compressed log segment corresponding to each of the current standard to-be-sharded compressed logs.

[0071] S230. Perform vertical extraction on each of the currently sharded and compressed log blocks in turn using a sliding window algorithm to obtain vertical log shards and remaining logs that are not vertically sharded. Furthermore, perform horizontal extraction on the remaining logs that are not vertically sharded using a pre-set horizontal log sharding processing method to obtain horizontal log shards.

[0072] S240: Perform encoding processing on each vertical log fragment and each horizontal log fragment using an adaptive encoding processing method according to the historical version encoding mapping rule to obtain each vertical encoded compressed log fragment and each horizontal encoded compressed log fragment;

[0073] The log fragment compression storage database includes historical version coding mapping rules.

[0074] Each vertically encoded compressed log fragment and each horizontally encoded compressed log fragment include an encoding version, a fragment fingerprint, a fragment type, a fragment encoding, and fragment metadata.

[0075] The adaptive encoding processing method may be a method of allocating shorter binary codes to log segments with higher frequencies (including identical text segments and parameterized templates) to maximize savings in overall storage space.

[0076] The log fragmentation compression storage database includes historical version coding mapping rules, and different versions of historical version coding mapping rules correspond to corresponding coding versions. After the historical version coding mapping rules are updated, the coding version needs to be adaptively modified.

[0077] For example, the encoding version of the historical version encoding mapping rule is 1; after the update, the encoding version can be modified to 2.

[0078] The shard fingerprint may include a hash value or a normalized template string of the shard content. The shard metadata may include parameters such as the first appearance time, average length, and frequency statistics of the shard metadata.

[0079] S250: Store each vertically coded compressed log fragment and each horizontally coded compressed log fragment in a log fragment compression storage database.

[0080] Optionally, it also includes: in the log segment compression storage database, periodically counting the occurrence frequency sizes corresponding to the log segments; when the occurrence frequency sizes of the log segments change, adjusting the encoding method of the log segments through a preset compression benefit calculation formula, and updating the historical version encoding mapping rules in the log segment compression storage database, saving the new version encoding mapping rules and the new encoding version; wherein, the new version encoding mapping rules and the new encoding version are loaded into the cache and marked as the current latest active version; using the new version encoding mapping rules and the new encoding version corresponding to the current latest active version to encode each vertical log segment and each horizontal log segment received at the next moment; for each vertical coded compressed log segment and each horizontal coded compressed log segment at the current moment, decoding is performed according to the corresponding encoding version to generate the original log.

[0081] In this embodiment, encoding needs to be performed according to the historical version coding mapping rules. The historical version coding mapping rules may include encoding methods corresponding to different types of log fragments. Specifically, for log fragments with longer lengths and higher frequencies, an encoding method with greater compression efficiency is adopted, thereby better realizing the compression coding processing of log data.

[0082] Furthermore, each vertically encoded compressed log fragment and each horizontally encoded compressed log fragment can be obtained and stored in a log fragment compression storage database. The encoding versions corresponding to the historical version encoding mapping rules must also be stored. This allows log recovery operations to be performed based on the encoding versions during the original log acquisition process.

[0083] Among them, the compression benefit calculation formula can be , where n is the total number of shard types; For the The total number of times this type of fragment appears; For the The average length of the original text of each segment (for templates, this is the template string length + typical parameter length estimate); For the is the average length of the encoded fragments; Z is the compression benefit. This formula shows that the larger the original fragment length and the more times it is repeated, the greater its compression contribution.

[0084] Therefore, it is necessary to periodically count the frequency of occurrence of each log segment in the log segment compression storage database; when the frequency of occurrence of a log segment changes, the encoding method of the log segment is adjusted through the compression benefit calculation formula, that is, the log segment with a longer length and a higher occurrence frequency is matched with a relatively short binary code.

[0085] Furthermore, it is possible to update the historical version encoding mapping rules in the log shard compression storage database to generate new version encoding mapping rules and a new encoding version.

[0086] In addition, the new version encoding mapping rules and the new encoding version are loaded into the cache and marked as the current latest active version; the new version encoding mapping rules and the new encoding version corresponding to the current latest active version are used to encode each vertical log segment and each horizontal log segment received at the next moment. This ensures that subsequent log segments are encoded using the current latest active version, optimizes data storage, and reduces storage pressure and costs. At the same time, it is necessary to save the historical version encoding mapping rules and the corresponding encoding version, so that the historical log segments can be correctly decoded.

[0087] The technical solution of the embodiment of the present invention is to obtain the current task incremental log file corresponding to the current pipeline task in real time; according to each of the current to-be-sharded compressed logs, a preset log data preprocessing method is used for processing to generate each current standard to-be-sharded compressed log, and the log is segmented in combination with a preset dynamic time window strategy to generate at least one current to-be-sharded compressed log block corresponding to each of the current standard to-be-sharded compressed logs; through a sliding window algorithm, each current to-be-sharded compressed log block is sequentially segmented and extracted vertically to obtain each vertical log segment and the remaining log that is not segmented vertically, and the remaining log that is not segmented vertically is segmented and extracted horizontally in combination with a preset horizontal log segment processing method to obtain each horizontal log segment; according to the historical version coding mapping rule, each vertical log segment and each horizontal log segment are respectively encoded and processed by an adaptive coding processing method to obtain each vertical coded compressed log segment and each horizontal coded compressed log segment; each vertical coded compressed log segment and each horizontal coded compressed log segment are stored in a log segment compression storage database. It can ensure that subsequent log shards are encoded using the current latest active version, optimize data storage, reduce storage pressure and costs; improve the utilization of storage resources, reduce storage resource waste, and improve data retrieval efficiency.

[0088] Example 3

[0089] Figure 3 This is a structural diagram of a log sharding compression storage device provided in the third embodiment of the present invention. The log sharding compression storage device provided in this embodiment can be implemented by software and / or hardware, and can be configured in a terminal device or server to implement a log sharding compression storage method in the embodiment of the present invention. Figure 3 As shown, the device includes: a current task incremental log file acquisition module 310, a current to-be-sharded and compressed log block generation module 320, a vertical log sharding and horizontal log sharding determination module 330 and a storage module 340.

[0090] Among them, the current task incremental log file acquisition module 310 is used to obtain the current task incremental log file corresponding to the current pipeline task in real time; wherein, the current task incremental log file includes at least one current log to be fragmented and compressed;

[0091] The current to-be-sharded compressed log block generation module 320 is configured to perform replacement based on each of the current to-be-sharded compressed logs using a preset regular matching unified placeholder replacement method to generate each current standard to-be-sharded compressed log, and perform log block generation based on a preset dynamic time window strategy to generate at least one current to-be-sharded compressed log block corresponding to each current standard to-be-sharded compressed log.

[0092] The vertical and horizontal log sharding determination module 330 is configured to perform vertical sharding extraction on each currently sharded and compressed log block in sequence using a sliding window algorithm to obtain vertical log shards and remaining logs that have not been vertically sharded. Furthermore, the module performs horizontal sharding extraction on the remaining logs that have not been vertically sharded using a pre-set horizontal log sharding processing method to obtain horizontal log shards.

[0093] The storage module 340 is used to encode each vertical log segment and each horizontal log segment using a preset adaptive encoding method, and store the encoded data in a log segment compression storage database.

[0094] Among them, the log fragment compression storage database corresponds to different encoding methods according to the frequency of occurrence of different log fragments; when the frequency of occurrence of log fragments changes, it is necessary to periodically update the encoding method of the log fragment compression storage database and generate a new encoding version.

[0095] The technical solution of the embodiment of the present invention is to obtain the current task incremental log file corresponding to the current pipeline task in real time; according to each of the current to-be-sliced ​​compressed logs, process them through a pre-set log data preprocessing method to generate each current standard to-be-sliced ​​compressed log, and perform log segmentation in combination with a pre-set dynamic time window strategy to generate at least one current to-be-sliced ​​compressed log segment corresponding to each of the current standard to-be-sliced ​​compressed logs; through a sliding window algorithm, perform segmented vertical extraction on each current to-be-sliced ​​compressed log segment in turn to obtain each vertical log segment and the remaining log that is not segmented vertically, and perform segmented horizontal extraction on the remaining log that is not segmented vertically in combination with a pre-set horizontal log segmentation processing method to obtain each horizontal log segment; through a pre-set adaptive encoding processing method, encode each vertical log segment and each horizontal log segment respectively, and store them in a log segmentation compression storage database. This solves the problem that the current repeated pattern log occupies a large amount of storage space, improves the utilization rate of storage resources, reduces storage costs, reduces the waste of storage resources, and improves data retrieval efficiency.

[0096] Based on the above embodiments, the current task incremental log file acquisition module 310 can be specifically used to monitor the status of the current pipeline task in real time through the log sharding compression storage system; when it is detected that the status of the current pipeline task is the end state, the incremental log file corresponding to the current pipeline task is obtained.

[0097] On the basis of the above embodiments, the current to-be-sliced ​​compressed log block generation module 320 can be specifically used to: normalize each of the current to-be-sliced ​​compressed logs, identify the dynamic variables in each of the current to-be-sliced ​​compressed logs, and replace them through a pre-set regular matching unified placeholder replacement method to generate each current standard to-be-sliced ​​compressed log; obtain a target current standard to-be-sliced ​​compressed log in turn from each of the current standard to-be-sliced ​​compressed logs; obtain the timestamps of each log entry in the target current standard to-be-sliced ​​compressed log through a dynamic time window strategy, and obtain the log block time threshold; calculate the corresponding time of consecutive log entries based on the timestamps of each log entry time interval; determine whether the time interval is greater than the log segmentation time threshold; if so, perform log segmentation on the target current standard to-be-segmented compressed log, continue to traverse the log entries in the un-logged log segments, and return to execute the operation according to the timestamp of each log entry until all log entries in the target current standard to-be-segmented compressed log are traversed; determine whether all current standard to-be-segmented compressed logs are traversed, and if so, generate at least one current to-be-segmented compressed log segment corresponding to each current standard to-be-segmented compressed log; if not, return to execute the operation of sequentially obtaining one target current standard to-be-segmented compressed log in each current standard to-be-segmented compressed log.

[0098] Based on the above embodiments, the vertical log sharding and horizontal log sharding determination module 330 can be specifically used to: obtain a preset minimum sharding length threshold; classify each of the current to-be-sharded compressed logs in turn using a pre-trained log entry deep learning classification model to obtain a log entry type; obtain a target current to-be-sharded compressed log block from each of the current to-be-sharded compressed log blocks; determine a maximum sliding window using a sliding window algorithm, and slide the target current to-be-sharded compressed log block according to the maximum sliding window, and determine whether the current sliding window log includes a unified placeholder in combination with the log entry type. If so, the placeholder template structure is extracted to generate the current placeholder vertical log fragment; if not, the long string pattern is matched on the current sliding window log. If it matches, the current vertical log fragment is confirmed to be identified; the maximum sliding window is reduced, and it is determined whether it is greater than or equal to the minimum fragment length threshold. If so, the operation of sliding the target current compressed log block to be fragmented according to the maximum sliding window is returned to execute; if not, the vertical fragment extraction is determined to be completed, and the remaining logs that are not vertically fragmented are obtained, and each vertical log fragment is obtained according to each current vertical log fragment and each current placeholder vertical log fragment.

[0099] Based on the above embodiments, the horizontal log sharding processing method includes a longest common subsequence processing method.

[0100] Based on the above embodiments, the vertical log sharding and horizontal log sharding determination module 330 can also be specifically used to: obtain a preset minimum length allowable threshold and a minimum sharding frequency; obtain an initial horizontal sharding length through the longest common subsequence processing method, and perform horizontal sharding extraction on the remaining logs that are not vertically sharded; wherein, the initial horizontal sharding length is greater than the minimum length allowable threshold; if the extraction is successful and the frequency of occurrence of the extracted horizontal log shards is not less than the minimum sharding frequency, the current horizontal log shard is obtained, and the initial horizontal sharding length is reduced to determine whether it is greater than or equal to the minimum length allowable threshold. If so, the operation of obtaining the initial horizontal sharding length through the longest common subsequence processing method and performing horizontal sharding extraction on the remaining logs that are not vertically sharded is returned to execution; after traversing the remaining logs that are not vertically sharded, each horizontal log shard is obtained.

[0101] Based on the above embodiments, the log fragment compression storage database includes historical version coding mapping rules.

[0102] Based on the above embodiments, the storage module 340 can be specifically used to: according to the historical version coding mapping rules, through an adaptive coding processing method, encode each vertical log slice and each horizontal log slice respectively to obtain each vertical coded compressed log slice and each horizontal coded compressed log slice; wherein each vertical coded compressed log slice and each horizontal coded compressed log slice include a coding version, a slice fingerprint, a slice type, a slice code and a slice metadata; and store each vertical coded compressed log slice and each horizontal coded compressed log slice in a log slice compression storage database.

[0103] On the basis of the above embodiments, it can also be specifically used for: periodically counting the frequency of occurrence of log slices in the log slice compression storage database; when the frequency of occurrence of log slices changes, adjusting the encoding method of the log slices through a preset compression benefit calculation formula, and updating the historical version encoding mapping rules in the log slice compression storage database, saving the new version encoding mapping rules and the new encoding version; wherein, the new version encoding mapping rules and the new encoding version are loaded into the cache and marked as the current latest active version; using the new version encoding mapping rules and the new encoding version corresponding to the current latest active version to encode each vertical log slice and each horizontal log slice received at the next moment; for each vertical coded compressed log slice and each horizontal coded compressed log slice at the current moment, decode them according to the corresponding encoding version to generate the original log.

[0104] The log fragmentation compression storage device provided in the embodiment of the present invention can execute the log fragmentation compression storage method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0105] Example 4

[0106] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement the fourth embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0107] like Figure 4 As shown, electronic device 10 includes at least one processor 11 and memory, such as read-only memory (ROM) 12 and random access memory (RAM) 13, communicatively connected to at least one processor 11. The memory stores computer programs executable by the at least one processor. Processor 11 can perform various appropriate actions and processes based on the computer programs stored in ROM 12 or loaded from storage unit 18 into RAM 13. RAM 13 can also store various programs and data required for the operation of electronic device 10. Processor 11, ROM 12, and RAM 13 are interconnected via bus 14. An input / output (I / O) interface 15 is also connected to bus 14.

[0108] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0109] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, digital signal processors (DSPs), and any other suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the log sharding and compression storage method.

[0110] In some embodiments, the log sharding compression storage method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the log sharding compression storage method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the log sharding compression storage method in any other appropriate manner (e.g., by means of firmware).

[0111] The method comprises: obtaining a current task incremental log file corresponding to a current pipeline task in real time; wherein the current task incremental log file includes at least one current to-be-sharded compressed log; processing each of the current to-be-sharded compressed logs by a pre-set log data preprocessing method to generate each current standard to-be-sharded compressed log, and performing log segmentation in combination with a pre-set dynamic time window strategy to generate at least one current to-be-sharded compressed log segment corresponding to each current standard to-be-sharded compressed log; performing segmented longitudinal extraction on each current to-be-sharded compressed log segment in turn by a sliding window algorithm to obtain each longitudinal log segment. , and the remaining logs that are not vertically fragmented, and combine the preset horizontal log fragmentation processing method to perform fragmented horizontal extraction on the remaining logs that are not vertically fragmented to obtain each horizontal log fragment; through the preset adaptive encoding processing method, each vertical log fragment and each horizontal log fragment are encoded and processed respectively, and stored in a log fragment compression storage database; wherein, the log fragment compression storage database corresponds to different encoding methods according to the frequency of occurrence corresponding to different log fragments; when the frequency of occurrence of log fragments changes, it is necessary to periodically update the encoding method of the log fragment compression storage database and generate a new encoding version.

[0112] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0113] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0114] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or apparatus. A computer-readable storage medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0115] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0116] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0117] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0118] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0119] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

[0120] Example 5

[0121] The fifth embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable instructions are used to execute a log sharding compression storage method when executed by a computer processor, the method comprising: obtaining a current task incremental log file corresponding to a current pipeline task in real time; wherein the current task incremental log file includes at least one current to-be-sharded compressed log; processing each of the current to-be-sharded compressed logs by a pre-set log data preprocessing method to generate each current standard to-be-sharded compressed log, and performing log block division in combination with a pre-set dynamic time window strategy to generate at least one current to-be-sharded compressed log block corresponding to each of the current standard to-be-sharded compressed logs; and sequentially performing the log block division on the current to-be-sharded compressed logs by a sliding window algorithm. Each current log block to be fragmented and compressed is subjected to fragmented vertical extraction to obtain each vertical log fragment and the remaining log that has not been fragmented vertically, and the remaining log that has not been fragmented vertically is subjected to fragmented horizontal extraction in combination with a pre-set horizontal log fragmentation processing method to obtain each horizontal log fragment; each vertical log fragment and each horizontal log fragment are encoded and processed respectively by a pre-set adaptive encoding processing method, and stored in a log fragment compression storage database; wherein the log fragment compression storage database corresponds to different encoding methods according to the frequency of occurrence corresponding to different log fragments; when the frequency of occurrence of log fragments changes, it is necessary to periodically update the encoding method of the log fragment compression storage database and generate a new encoding version.

[0122] Of course, the computer-readable storage medium provided in the embodiment of the present invention has computer-executable instructions that are not limited to the method operations described above, and can also execute related operations in log shard compression storage provided in any embodiment of the present invention.

[0123] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general-purpose hardware. Of course, it can also be implemented with hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0124] It is worth noting that in the above-mentioned embodiment of log sharding and compression storage, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention.

[0125] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A log fragment compression storage method, characterized in that: include: Obtaining in real time the current task incremental log file corresponding to the current pipeline task; wherein the current task incremental log file includes at least one log to be currently fragmented and compressed; Based on each of the current to-be-sharded compressed logs, a preset log data preprocessing method is used to process the logs to generate current standard to-be-sharded compressed logs, and a preset dynamic time window strategy is used to perform log segmentation to generate at least one current to-be-sharded compressed log segment corresponding to each of the current standard to-be-sharded compressed logs. Using a sliding window algorithm, perform vertical extraction on each currently sharded and compressed log block in turn to obtain vertical log shards and the remaining logs that have not been vertically sharded. Then, using a pre-set horizontal log sharding method, perform horizontal extraction on the remaining logs that have not been vertically sharded to obtain horizontal log shards. Through a pre-set adaptive encoding processing method, each vertical log segment and each horizontal log segment are encoded and stored in a log segment compression storage database; Among them, the log fragment compression storage database corresponds to different encoding methods according to the frequency of occurrence of different log fragments; when the frequency of occurrence of log fragments changes, it is necessary to periodically update the encoding method of the log fragment compression storage database and generate a new encoding version.

2. The method according to claim 1, characterized in that The real-time acquisition of the current task incremental log file corresponding to the current pipeline task includes: Through the log sharding compression storage system, the status of the current pipeline task is monitored in real time; When it is detected that the state of the current pipeline task is the end state, the incremental log file corresponding to the current pipeline task is obtained.

3. The method according to claim 2, characterized in that The method of processing the current to-be-sharded compressed logs by a preset log data preprocessing method to generate current standard to-be-sharded compressed logs, and performing log segmentation in combination with a preset dynamic time window strategy to generate at least one current to-be-sharded compressed log segment corresponding to each current standard to-be-sharded compressed log includes: Normalizing each of the current to-be-sharded and compressed logs, identifying dynamic variables in each of the current to-be-sharded and compressed logs, and replacing them using a pre-set regular matching unified placeholder replacement method to generate each current standard to-be-sharded and compressed log; Sequentially obtain a target current standard to-be-sharded compressed log from each of the current standard to-be-sharded compressed logs; Obtain the timestamp of each log entry in the target current standard to-be-sharded compressed log and the log block time threshold through a dynamic time window strategy; Calculating the time intervals corresponding to consecutive log entries based on the timestamps of the log entries; Determine whether the time interval is greater than the log block time threshold; If so, perform log chunking on the target current standard to-be-sharded compressed log, continue traversing the log entries in the un-logged log, and return to perform the operation based on the timestamp of each log entry until all log entries in the target current standard to-be-sharded compressed log are traversed; Determine whether all current standard to-be-sharded compressed logs have been traversed. If so, generate at least one current to-be-sharded compressed log block corresponding to each of the current standard to-be-sharded compressed logs; if not, return to execute the operation of obtaining a target current standard to-be-sharded compressed log in each of the current standard to-be-sharded compressed logs in turn.

4. The method according to claim 3, characterized in that The sliding window algorithm is used to sequentially perform vertical extraction on each of the currently compressed log blocks to obtain each vertical log fragment and the remaining logs that have not been vertically fragmented, including: Get the preset minimum fragment length threshold; Using a pre-trained deep learning classification model for log entries, classify each of the currently sharded and compressed logs in turn to obtain a log entry type; Obtain a target current log block to be sharded and compressed from each of the current log blocks to be sharded and compressed; Determine the maximum sliding window through the sliding window algorithm, and slide the target current to-be-sharded and compressed log block according to the maximum sliding window. Combined with the log entry type, determine whether the current sliding window log includes a unified placeholder. If so, extract the placeholder template structure and generate the current placeholder vertical log fragment. If not, the long string pattern is matched against the current sliding window log. If a match is found, the current vertical log shard is confirmed. Performing a reduction operation on the maximum sliding window and determining whether it is greater than or equal to the minimum fragment length threshold, and if so, returning to executing the operation of sliding the target current to-be-fragmented compressed log block according to the maximum sliding window; If not, it is determined that the vertical slice extraction is completed, and the remaining logs that are not vertically sliced ​​are obtained, and each vertical log slice is obtained according to each current vertical log slice and each current placeholder vertical log slice.

5. The method according to claim 4, characterized in that The horizontal log sharding processing method includes a longest common subsequence processing method; The method of performing horizontal extraction on the remaining logs that have not been vertically fragmented by combining a preset horizontal log fragmentation processing method to obtain horizontal log fragments includes: Get the preset minimum length threshold and minimum fragment frequency; Obtaining the initial horizontal shard length using the longest common subsequence processing method, and performing horizontal shard extraction on the remaining logs that have not been vertically sharded; The initial horizontal slice length is greater than the minimum length allowed threshold; If the extraction is successful and the frequency of occurrence of the extracted horizontal log fragment is not lower than the minimum fragment occurrence frequency, the current horizontal log fragment is obtained, and the initial horizontal fragment length is reduced to determine whether it is greater than or equal to the minimum length allowed threshold. If so, the process returns to executing the operation of obtaining the initial horizontal fragment length through the longest common subsequence processing method and performing horizontal fragment extraction on the remaining logs that have not been vertically fragmented; After traversing the remaining logs that have not been vertically sharded, the horizontal log shards are obtained.

6. The method according to claim 5, characterized in that The log fragment compression storage database includes historical version coding mapping rules; The method of encoding each vertical log segment and each horizontal log segment by using a preset adaptive encoding method and storing the encoded segments in a log segment compression storage database includes: According to the historical version encoding mapping rule, each vertical log fragment and each horizontal log fragment are encoded by an adaptive encoding processing method to obtain each vertical encoded compressed log fragment and each horizontal encoded compressed log fragment; Each vertically coded compressed log fragment and each horizontally coded compressed log fragment includes a coding version, a fragment fingerprint, a fragment type, a fragment code, and fragment metadata; Each vertically coded compressed log fragment and each horizontally coded compressed log fragment are stored in a log fragment compression storage database.

7. The method according to claim 6, characterized in that Also includes: In the log shard compression storage database, periodically count the occurrence frequencies of the log shards. When the frequency and size of log shards change, the encoding method of the log shards is adjusted using the pre-set compression benefit calculation formula, and the historical version encoding mapping rules in the log shard compression storage database are updated, and the new version encoding mapping rules and the new encoding version are saved; Among them, the new version encoding mapping rules and the new encoding version are loaded into the cache and marked as the current latest active version; Use the new version encoding mapping rules and new encoding version corresponding to the current latest active version to encode each vertical log segment and each horizontal log segment received at the next moment; for each vertical encoded compressed log segment and each horizontal encoded compressed log segment at the current moment, decode them according to the corresponding encoding version to generate the original log.

8. A log fragment compression storage device, characterized in that: include: The current task incremental log file acquisition module is used to obtain the current task incremental log file corresponding to the current pipeline task in real time; wherein, the current task incremental log file includes at least one current log to be fragmented and compressed; The current to-be-sharded compressed log block generation module is configured to replace each of the current to-be-sharded compressed logs using a pre-set regular matching unified placeholder replacement method to generate each current standard to-be-sharded compressed log, and perform log block generation in combination with a pre-set dynamic time window strategy to generate at least one current to-be-sharded compressed log block corresponding to each of the current standard to-be-sharded compressed logs; The vertical log sharding and horizontal log sharding determination module is used to perform vertical sharding extraction on each currently sharded and compressed log block in sequence using a sliding window algorithm to obtain vertical log shards and remaining logs that have not been vertically sharded. Furthermore, the module performs horizontal sharding extraction on the remaining logs that have not been vertically sharded using a pre-set horizontal log sharding processing method to obtain horizontal log shards. A storage module is used to encode each vertical log segment and each horizontal log segment using a preset adaptive encoding method, and store the encoded data in a log segment compression storage database; Among them, the log fragment compression storage database corresponds to different encoding methods according to the frequency of occurrence of different log fragments; when the frequency of occurrence of log fragments changes, it is necessary to periodically update the encoding method of the log fragment compression storage database and generate a new encoding version.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for log sharding compression storage is implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a log fragmentation compression storage method according to any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • Log data fragmentation and query method and apparatus

    CN105117403A

  • Log file compression method and decompression method, electronic equipment and readable storage medium

    CN107977442A

  • Log-based analysis method and device, electronic equipment and storage medium

    CN118227580A

  • Log compression method and device, electronic equipment and storage medium

    CN119248733A

  • Log compression method and device, computer readable storage medium and electronic equipment

    CN119848004A