Intelligent log archiving and querying method and system
By dynamically adjusting the archiving trigger conditions and selecting an adaptive compression algorithm through pre-training models, a lightweight index is generated, which solves the problems of large system log storage resources and long retrieval time, and realizes efficient log compression and fast query.
Patent Information
- Application Number
- CN202510548305.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-09-05
AI Technical Summary
When processing system logs, existing technologies have problems such as large storage resource usage, inconsistent log formats, long retrieval time, and unencrypted storage of sensitive operation logs. In addition, existing compression algorithms have low compression rates or increase retrieval time.
A pre-trained model is used to dynamically predict log generation trends, adaptively adjust archiving trigger conditions, select compression algorithms based on log content types, and generate searchable lightweight indexes. Logs are encrypted in layers and written to the blockchain to achieve differentiated compression and fast query.
It increases the compression rate by 20%-40%, reduces data redundancy, reduces the decompression amount by 98%, improves query efficiency and security, and meets the requirements of different types of log data.
Smart Images

Figure CN120596450A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of log management, and in particular relates to a log intelligent archiving and query method and system. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] System logs record information about hardware, software, and system problems in the system. They can also monitor events that occur in the system. Users can use them to check the cause of errors or find traces left by attackers when attacked. This is of great significance.
[0004] However, with technological advancements, system data is accumulating. The average daily log volume for a single server can reach GB, and uncompressed raw logs consume significant storage resources. Furthermore, logs generated by different devices come in varying formats, making unified processing difficult. Furthermore, troubleshooting operations requires searching across multiple log files, which requires time-consuming traditional linear scanning. Furthermore, some sensitive operation logs (such as permission changes) are stored unencrypted, posing a risk of leakage or tampering.
[0005] To address the above issues, existing technologies include using GZIP or Tar to package and compress logs. However, in actual applications, the compression rate is only 40%-50%, and content retrieval is not supported. In addition, some high-compression algorithms easily require full decompression for retrieval, which increases time consumption and has poor application effects. Summary of the Invention
[0006] In order to solve the above problems, the present invention proposes a log intelligent archiving and query method and system. The present invention can realize differentiated compression according to the log content type to ensure the compression effect, and generate a searchable lightweight index during the compression process to facilitate later query.
[0007] According to some embodiments, the present invention adopts the following technical solutions:
[0008] A log intelligent archiving and query method includes the following steps:
[0009] Use pre-trained models to dynamically predict log generation trends and adaptively adjust archiving trigger conditions based on the predicted log generation trends;
[0010] Select a storage area based on the log data cycle, select an adaptive compression algorithm based on the content type of the log data, and generate a searchable lightweight index during the compression process to query the corresponding log using the lightweight index;
[0011] In response to the archiving trigger condition, the archive log is hierarchically encrypted and integrity verification information is written into the blockchain.
[0012] As an optional implementation method, the process of dynamically predicting log generation trends using a pre-trained model includes: obtaining historical log volume, real-time throughput, and business event annotation information, using the obtained data as input, and using the pre-trained model to predict the log generation volume within a certain period of time in the future.
[0013] As an optional implementation, the model is an online long-short term neural network model, which updates and optimizes weight parameters at set time intervals.
[0014] As an optional implementation method, the process of adaptively adjusting the archiving trigger conditions based on the predicted log generation trend includes: if the predicted log generation volume is greater than the set threshold, archiving is triggered; if the real-time throughput rate exceeds the set value, archiving is triggered; if the ratio of the actual log generation volume within a set time period to the predicted log generation volume within the time period exceeds the set value, archiving is triggered; when the log data of the set time period has business event annotation information, the set threshold or set value is adjusted, or the pre-configured archiving period is adjusted.
[0015] As an optional implementation, the process of selecting a storage area according to the period of the log data includes: adopting an erasure code sharding storage strategy for the log data within a set time period and storing the log data in the first storage area;
[0016] Log data outside the set time period is stored in the second storage area, and its access frequency is obtained. When the access frequency drops below a threshold, cross-media migration is automatically triggered and stored in the third storage area.
[0017] As an optional implementation, the process of selecting an adaptive compression algorithm according to the log content type includes: if the log content type is a text log, using the Zstandard algorithm to create a dictionary for keywords with a frequency greater than a set value;
[0018] If the log content type is binary log, a differential encoding algorithm is used to store only the difference part.
[0019] As an optional implementation method, according to the system load, taking into account compression efficiency and retrieval performance, dynamic adjustment is performed. When the system load is higher than the set value, the decompression time is reduced and a low compression ratio algorithm is selected; the amount of data decompressed in a single time is reduced;
[0020] When the system load is lower than the preset value, the compression ratio is increased and a high compression ratio algorithm is selected; the block size is increased to reduce the number of blocks.
[0021] As an optional embodiment, the process of generating a searchable lightweight index during the compression process includes: extracting the log timestamp, device ID, and error code during the compression process, constructing a two-layer index structure with the log timestamp and / or device ID as a tree data structure index and the error code as a semantic fingerprint inverted index, thereby forming a lightweight index;
[0022] And establish bidirectional pointers between the compressed block and the index entries of the lightweight index.
[0023] As an optional implementation method, the process of hierarchical encryption of archived logs includes: when archiving is required, an encryption algorithm is selected based on the meta-features of the log; if the meta-features indicate that the device security level is higher than the set value or contains sensitive labels, encryption is performed using a first encryption algorithm; otherwise, encryption is performed using a second encryption algorithm; the security level of the first encryption algorithm is higher than that of the second encryption algorithm.
[0024] As an optional implementation method, the process of hierarchical encryption of archived logs includes: using a clustering algorithm to analyze access data, and if the access frequency of relevant log data is higher than a set value, the key rotation is performed every set time period; for other data, the key rotation is performed every predetermined time period, and the length of the set time period is less than the length of the predetermined time period.
[0025] As an optional implementation, in the process of writing the integrity verification information into the blockchain, a hash tree is generated for each batch of log data, written into the blockchain, and a blockchain consensus algorithm is used to enable all nodes to reach a consensus on the state of the blockchain.
[0026] As an optional implementation method, when log data query is required, the query request is parsed and converted into a query plan, and computing resources are dynamically allocated according to the query complexity; the query plan is compared with historical query information, and corresponding log blocks with correlation greater than or equal to a threshold are asynchronously preloaded.
[0027] A log intelligent archiving and query system, comprising:
[0028] A trigger configuration module is configured to dynamically predict log generation trends using a pre-trained model and adaptively adjust archiving trigger conditions based on the predicted log generation trends;
[0029] A compression selection module is configured to select a storage area based on the period of the log data, select an adaptive compression algorithm based on the content type of the log data, and generate a searchable lightweight index during the compression process so as to query the corresponding log using the lightweight index;
[0030] The archiving encryption module is configured to perform hierarchical encryption on the archive log in response to the archiving trigger condition and write integrity verification information into the blockchain.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] The present invention dynamically predicts log generation trends through an online learning model and adaptively adjusts archiving trigger conditions. It can flexibly adjust archiving trigger conditions according to time and event conditions to ensure that the archiving time is as accurate as possible, meet the log data archiving requirements, and avoid data backlog and resource occupation.
[0033] The present invention selects a storage area according to the cycle of log data and an adaptive compression algorithm according to the content type of log data, flexibly selects compression, reduces data redundancy, meets the requirements of different types of log data, improves the compression rate by 20%-40%, and reduces cold storage costs.
[0034] The present invention generates a searchable lightweight index during the compression process, and the index is a hybrid index formed by two indexes, establishing a bidirectional pointer between the compressed block and the index entry, so that only 3%-5% of the relevant data blocks need to be decompressed during retrieval, greatly reducing the decompression amount.
[0035] The present invention dynamically selects encryption algorithms based on log meta-features to achieve fine-grained access control and implements more frequent key rotation for high-frequency access data, thus ensuring data security and taking efficiency into consideration.
[0036] The present invention can achieve fast query and positioning based on lightweight indexes, parse natural language queries into structured query plans, dynamically allocate computing resources according to query complexity, analyze historical query data and results, and asynchronously preload highly relevant log blocks in advance, significantly improving query efficiency and response time.
[0037] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0039] Figure 1 This is a flow chart of a log intelligent archiving and query method according to an embodiment;
[0040] Figure 2 This is a flowchart of a query process according to an embodiment;
[0041] Figure 3 This is a schematic diagram of a dynamic archiving triggering strategy flow in an embodiment;
[0042] Figure 4 This is a flow chart of a dynamic archiving triggering strategy according to another embodiment. DETAILED DESCRIPTION
[0043] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0044] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0045] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0046] In the absence of conflict, the embodiments and features in the embodiments of this application can be combined with each other.
[0047] Example 1
[0048] A log intelligent archiving and query method, such as Figure 1 As shown, the following steps are included:
[0049] Use pre-trained models to dynamically predict log generation trends and adaptively adjust archiving trigger conditions based on the predicted log generation trends;
[0050] Select a storage area based on the log data cycle, select an adaptive compression algorithm based on the content type of the log data, and generate a searchable lightweight index during the compression process to query the corresponding log using the lightweight index;
[0051] In response to the archiving trigger condition, the archive log is hierarchically encrypted and integrity verification information is written into the blockchain.
[0052] In this embodiment, the process of dynamically predicting log generation trends using a pre-trained model includes: obtaining historical log volume, real-time throughput, and business event annotation information, using the obtained data as input, and using the pre-trained model to predict the log generation volume within a certain period of time in the future.
[0053] In this embodiment, the model is a long-short term neural network model, which updates and optimizes weight parameters at set time intervals, and the long-short term neural network model continuously performs online learning and optimization to reduce prediction errors.
[0054] The training of the long-term and short-term neural network model can be performed using historical relevant data, and the training process can also use existing algorithms, which will not be described in detail in this embodiment.
[0055] In some embodiments, archiving is triggered if the predicted log generation volume is greater than a set threshold.
[0056] In some embodiments, Figure 3 As shown, if the ratio of the actual log generation volume within the set time period to the predicted log generation volume within the time period exceeds the set value, archiving is triggered; when the log data of the set time period has business event annotation information, the set threshold or set value is adjusted, or the pre-configured archiving period is adjusted.
[0057] In some embodiments, Figure 4 As shown in the figure, if the real-time throughput exceeds the set value, archiving is triggered.
[0058] Of course, in other embodiments, the above judgment processes can be combined or nested, such as first judging the predicted log generation volume, then judging the throughput rate, and then judging the ratio of the actual log generation volume to the predicted log generation volume within the time period.
[0059] In addition, the business event tag information of this embodiment can be set according to the specific application scenario, such as e-commerce promotion date identification, end-of-season inventory identification, etc. It is generally believed that business event tag information represents that the business peak period is within this time period.
[0060] The above values or periods can also be configured based on experience. For example, when the ratio of the actual log generation volume to the predicted log generation volume during the time period exceeds 1.2, archiving is triggered. When there is business event annotation information, archiving is triggered immediately only when the set value is 1.3.
[0061] In this embodiment, the process of selecting a storage area according to the period of log data includes: adopting an erasure code sharding storage strategy for log data within a set time period and storing the data in a first storage area;
[0062] Log data outside the set time period is stored in the second storage area, and its access frequency is obtained. When the access frequency drops below a threshold, cross-media migration is automatically triggered and stored in the third storage area.
[0063] For example, log data within 7 days is considered hot data and is stored in SSD storage space using erasure coding shards. This can effectively improve utilization while ensuring that the number of read and write operations per second is above the predetermined value.
[0064] For log data older than seven days, a migration decision tree based on access popularity can be used. When the access frequency drops to a threshold value T, in this embodiment, T = log(Δt)*C, where Δt is the time decay factor, cross-media migration is automatically triggered. Frequently accessed log data is compressed and stored on the HDD, while other data is archived to object storage. C represents a calibration coefficient for adjusting migration sensitivity. A larger C value leads to a higher threshold value T, which makes it easier for the system to trigger migration (i.e., data is judged as "cold data" more quickly). A smaller C value leads to a lower threshold value T, which makes the system more conservative (retaining data in hot storage longer).
[0065] In this embodiment, the process of selecting an adaptive compression algorithm based on the content type of the log data includes: if the log content type is a text log, using the Zstandard algorithm to create a dictionary for keywords with a frequency greater than a set value;
[0066] If the log content type is binary log, a differential encoding algorithm is used to store only the difference part.
[0067] In some embodiments, domain dictionaries can be constructed using online word frequency statistics to atomically encode compound terms containing characters and numbers. A log syntax graph can also be established across device clusters to implement cross-file compression for recurring call chain patterns (such as microservice call chain IDs).
[0068] The process of generating a searchable lightweight index during the compression process includes: extracting the log timestamp, device ID, and error code during the compression process, constructing a two-layer index structure with the log timestamp and / or device ID as a tree data structure index and the error code as a semantic fingerprint inverted index, thereby forming a lightweight index;
[0069] And establish bidirectional pointers between the compressed block and the index entries of the lightweight index.
[0070] In this embodiment, the first layer is a timestamp range index using a B+ tree structure, and the second layer is a semantic fingerprint inverted index using the SimHash algorithm. Bidirectional pointers are established between compressed blocks and index entries, allowing only 3%-5% of the relevant data blocks to be decompressed during retrieval. Experiments have shown a 98% reduction in decompression compared to traditional solutions.
[0071] The process of hierarchical encryption of archived logs includes: when archiving is required, an encryption algorithm is selected based on the meta-features of the log. If the meta-features indicate that the device security level is higher than the set value or contains sensitive labels, the first encryption algorithm is used for encryption; otherwise, the second encryption algorithm is used for encryption. The security level of the first encryption algorithm is higher than that of the second encryption algorithm.
[0072] In this embodiment, an encryption algorithm is dynamically selected based on log meta-features (such as device security level and log sensitivity label) to achieve fine-grained access control.
[0073] The process of hierarchical encryption of archived logs includes: using a clustering algorithm to analyze access data. If the access frequency of relevant log data is higher than the set value, the key rotation is performed every set time period. For other data, the key rotation is performed every predetermined time period, and the length of the set time period is less than the length of the predetermined time period.
[0074] In this embodiment, based on K-means clustering analysis of access patterns, daily key rotation is implemented for high-frequency access data, and weekly rotation is implemented for low-frequency data.
[0075] When writing integrity verification information to the blockchain, a hash tree is generated for each batch of log data and written to the blockchain. A blockchain consensus algorithm is then used to enable all nodes to reach consensus on the blockchain's state. In this embodiment, a RAFT+PBFT consensus algorithm is employed within the private chain to achieve efficient processing. This also eliminates the need to transmit the original data when verifying log integrity, reducing bandwidth consumption by 98%.
[0076] When you need to query log data, such as Figure 2 As shown, the query request is parsed and converted into a query plan, and computing resources are dynamically allocated according to the query complexity. The query plan is compared with historical query information, and corresponding log blocks with correlation greater than or equal to a threshold are asynchronously preloaded.
[0077] This embodiment builds a query intent analyzer, uses the BERT model to parse natural language queries into structured query plans, and dynamically allocates computing resources based on query complexity, such as using the CPU to process regular expression matching, FPGA to accelerate similarity calculations, and GPU clusters to process complex association analysis.
[0078] In addition, this embodiment provides a pre-fetching mechanism based on association rules: it analyzes historical query patterns, implements asynchronous pre-loading for log blocks with correlation > 0.7, and caches high-frequency query results in an in-memory database (such as a Redis cluster) to achieve millisecond-level secondary response.
[0079] This embodiment also provides a dynamic adjustment mechanism that takes into account both compression efficiency and retrieval (ie, query) performance according to system load conditions.
[0080] First, construct the compression-retrieval joint optimization function:
[0081] min(α·Cratio+β·Tretrieval);
[0082] st
[0083] Cratio=1-Scompressed / Soriginal;
[0084] Tretrieval=γ·Nblocks+δ·Ddecompress;
[0085] Among them, α and β are dynamic weight coefficients, which are automatically adjusted according to the real-time system load. For example, when the CPU utilization rate is greater than 70%, the β weight increases by 50%.
[0086] Objective function: min(α·Cratio+β·Tretrieval);
[0087] That is, dynamically balance compression efficiency (saving storage space) and retrieval performance (reducing latency) to minimize the overall cost.
[0088] Dynamic weighting: α and β are automatically adjusted based on the real-time system load (for example, when CPU utilization is > 70%, the β weight increases by 50%), making the system adaptive to different load scenarios.
[0089] The constraints include:
[0090] (1)Cratio is the compression efficiency index:
[0091] Cratio=1-Scompressed / Soriginal;
[0092] Physical meaning: Measures the ratio of storage space saved by compression.
[0093] Soriginal: original data size;
[0094] Scompressed: compressed data size.
[0095] Optimization direction:
[0096] The larger the Cratio → the higher the compression rate (the lower the storage cost);
[0097] However, a high compression ratio may increase the decompression time Ddecompress (and reduce the retrieval performance).
[0098] (2) Retrieval time cost
[0099] Tretrieval=γ·Nblocks+δ·Ddecompress
[0100] Physical meaning: Retrieval latency consists of two parts:
[0101] I / O time (γ·Nblocks): proportional to the number of data blocks Nblocks (for example, reading multiple small files is slower);
[0102] Decompression time (δ·Ddecompress): proportional to the time it takes to decompress a single data block.
[0103] in:
[0104] γ: Storage medium I / O efficiency coefficient (mechanical hard disk > SSD);
[0105] δ: CPU decompression efficiency coefficient (weak CPU > strong CPU).
[0106] This embodiment provides a dynamic weight adjustment mechanism:
[0107] Default state: α and β are assigned according to the baseline policy (e.g., α = 0.6, β = 0.4).
[0108] Trigger condition: When the CPU utilization is greater than 70%, the β weight increases by 50% (eg, βnew = β·1.5).
[0109] Adjustment purpose: When the CPU is under high load, prioritize search performance (reduce the pressure of decompression on the CPU) and avoid system overload.
[0110] For example, the initial values are α=0.6, β=0.4; when the CPU is highly loaded: β=0.4×1.5=0.6, and α may be synchronously adjusted to 1-0.6=0.4 (assuming the total weight is 1).
[0111] The optimization function changes from 0.6Cratio+0.4Tretrieval to 0.4Cratio+0.6Tretrieval;
[0112] The system prefers the strategy with low decompression time (even if the compression ratio is slightly lower).
[0113] The optimization direction of the above dynamic adjustment strategy when the CPU is under high load is to reduce the decompression time Ddecompress:
[0114] Select a low compression algorithm (such as LZ4 fast mode instead of Zstandard high compression); reduce the amount of data decompressed at a time (for example, reduce the block size Nblocks); reduce CPU usage: By reducing decompression complexity, CPU pressure is alleviated.
[0115] The optimization direction of the above dynamic adjustment strategy when the CPU load is low is to improve the compression ratio Cratio, select a high compression ratio algorithm (such as Brotli or Zstandard high compression file); increase the block size to reduce Nblocks (but the decompression time must be traded off).
[0116] In a specific scenario, the CPU utilization of the real-time database is 80%.
[0117] Dynamic adjustment: β increases by 50%, α decreases.
[0118] Strategy selection:
[0119] Use LZ4 fast compression (Cratio=0.3, Ddecompress=10ms);
[0120] The block size is 2MB (Nblocks=500).
[0121] result:
[0122] Tretrieval=0.1·500+0.5·10=50+5=55;
[0123] The total cost is 0.4·0.3+0.6·55=0.12+33=33.12.
[0124] Let's take another specific scenario as an example: archiving storage, with a current CPU utilization of 40%.
[0125] Dynamic adjustment: α=0.7, β=0.3.
[0126] Strategy selection:
[0127] Use Zstandard high compression (Cratio=0.6, Ddecompress=50ms);
[0128] The block size is 8MB (Nblocks=125).
[0129] result:
[0130] Tretrieval=0.1·125+0.5·50=12.5+25=37.5;
[0131] The total cost is 0.7·0.6+0.3·37.5=0.42+11.25=11.67.
[0132] This embodiment tests the log data of a certain period and a certain system using the traditional solution and the above method. As shown in Table 1, the present invention has greatly improved and enhanced the compression time ratio, composite query response time, and storage cost / retrieval performance ratio.
[0133] Table 1 Comparison of indicator effects
[0134] index Traditional solutions Solution of the present invention Compression time ratio 1.0x 0.6x (FPGA acceleration) Composite query response time 1200ms 85ms Storage cost / retrieval performance ratio 1:1 1:3.8
[0135] Example 2
[0136] A log intelligent archiving and query system, comprising:
[0137] A trigger configuration module is configured to dynamically predict log generation trends using a pre-trained model and adaptively adjust archiving trigger conditions based on the predicted log generation trends;
[0138] A compression selection module is configured to select a storage area based on the period of the log data, select an adaptive compression algorithm based on the content type of the log data, and generate a searchable lightweight index during the compression process so as to query the corresponding log using the lightweight index;
[0139] The archiving encryption module is configured to perform hierarchical encryption on the archive log in response to the archiving trigger condition and write integrity verification information into the blockchain.
[0140] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0141] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0142] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0144] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made by those skilled in the art that fall within the spirit and principles of the present invention and do not require creative effort are intended to be within the scope of protection of the present invention.
Claims
1. A log intelligent archiving and query method, characterized in that: The following steps are involved: Use pre-trained models to dynamically predict log generation trends and adaptively adjust archiving trigger conditions based on the predicted log generation trends; Select a storage area based on the log data cycle, select an adaptive compression algorithm based on the content type of the log data, and generate a searchable lightweight index during the compression process to query the corresponding log using the lightweight index; In response to the archiving trigger condition, the archive log is hierarchically encrypted and integrity verification information is written into the blockchain.
2. A log intelligent archiving and query method as claimed in claim 1, characterized in that: The process of dynamically predicting log generation trends using a pre-trained model involves obtaining historical log volume, real-time throughput, and business event annotation information. Using this data as input, the pre-trained model is used to predict log generation within a certain period of time. The model is an online long-short term neural network model, which updates and optimizes weight parameters at set intervals.
3. The method for intelligent log archiving and querying according to claim 1, wherein: The process of adaptively adjusting the archiving trigger conditions based on the predicted log generation trend includes: triggering archiving if the predicted log generation volume is greater than the set threshold; triggering archiving if the real-time throughput rate exceeds the set value; triggering archiving if the ratio of the actual log generation volume within a set time period to the predicted log generation volume within the time period exceeds the set value; and adjusting the set threshold or set value, or adjusting the pre-configured archiving period, when the log data of the set time period has business event annotation information.
4. The method for intelligent log archiving and querying according to claim 1, wherein: The process of selecting a storage area according to the period of log data includes: adopting an erasure code sharding storage strategy for log data within a set time period and storing the log data in a first storage area; Log data outside the set time period is stored in the second storage area, and its access frequency is obtained. When the access frequency drops below a threshold, cross-media migration is automatically triggered and stored in the third storage area.
5. The method for intelligent log archiving and querying according to claim 1, wherein: The process of selecting an appropriate compression algorithm based on the log content type includes: if the log content type is a text log, the Zstandard algorithm is used to create a dictionary for keywords with a frequency greater than a set value; If the log content type is binary log, the differential encoding algorithm is used to store only the difference part; Alternatively, based on the system load, take into account both compression efficiency and retrieval performance and dynamically adjust the weight of the two. When the system load is higher than the set value, reduce the decompression time and select a low compression algorithm; reduce the amount of data decompressed at a time; When the system load is lower than the preset value, the compression ratio is increased and a high compression ratio algorithm is selected; the block size is increased to reduce the number of blocks.
6. A log intelligent archiving and query method as claimed in claim 1, characterized in that The process of generating a searchable lightweight index during compression includes: extracting the log timestamp, device ID, and error code during compression, and constructing a two-layer index structure with the log timestamp and / or device ID as a tree data structure index and the error code as a semantic fingerprint inverted index to form a lightweight index; And establish bidirectional pointers between the compressed block and the index entries of the lightweight index.
7. The method for intelligent log archiving and querying according to claim 1, wherein: The process of performing hierarchical encryption on archived logs includes: when archiving is required, selecting an encryption algorithm based on the log's meta-features; if the meta-features indicate that the device security level is higher than a set value or contains a sensitive label, encrypting the log using a first encryption algorithm; otherwise, encrypting the log using a second encryption algorithm, where the first encryption algorithm has a higher security level than the second encryption algorithm; Use clustering algorithm to analyze access data. If the access frequency of relevant log data is higher than the set value, the key rotation is performed every set time period. For other data, the key rotation is performed every predetermined time period. The length of the set time period is less than the length of the predetermined time period.
8. The method for intelligent log archiving and querying according to claim 1, wherein: In the process of writing integrity verification information into the blockchain, a hash tree is generated for each batch of log data, written into the blockchain, and the blockchain consensus algorithm is used to enable all nodes to reach a consensus on the state of the blockchain.
9. The method for intelligent log archiving and querying according to claim 1, wherein: When log data query is required, the query request is parsed and converted into a query plan, and computing resources are dynamically allocated based on the query complexity. The query plan is compared with historical query information, and corresponding log blocks with correlation greater than or equal to the threshold are asynchronously preloaded.
10. A log intelligent archiving and query system, characterized by: include: A trigger configuration module is configured to dynamically predict log generation trends using a pre-trained model and adaptively adjust archiving trigger conditions based on the predicted log generation trends; A compression selection module is configured to select a storage area based on the period of the log data, select an adaptive compression algorithm based on the content type of the log data, and generate a searchable lightweight index during the compression process so as to query the corresponding log using the lightweight index; The archiving encryption module is configured to perform hierarchical encryption on the archive log in response to the archiving trigger condition and write integrity verification information into the blockchain.