A log management method, electronic device, storage medium and program product

CN122818352APending Publication Date: 2026-09-25CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611308264.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

传统的数据库审计日志管理方法主要依赖关系型数据库进行存储,这种方式存在明显的局限性,在面对海量审计日志时,操作复杂度提升,管理效率低下

Benefits of technology

[0007]本申请实施例提供一种计算机可读存储介质,存储有计算机程序或计算机可执行指令,用于被处理器执行时实现本申请实施例提供的日志管理方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818352A_ABST
    Figure CN122818352A_ABST
Patent Text Reader

Abstract

The application provides a log management method, an electronic device, a storage medium and a program product; the method comprises the following steps: obtaining an audit log of a database; wherein the audit log comprises a log containing multiple types of operation information generated in the process of database auditing; preprocessing the audit log to obtain a preprocessed audit log; wherein the preprocessed audit log is a standardized audit log; performing word segmentation processing on the preprocessed audit log to obtain a word unit set; performing word embedding mapping on each word unit in the word unit set to obtain each vector corresponding to each word unit; arranging each vector in a preset order, adding hash information to each vector after the arrangement, and obtaining a target storage matrix of the audit log.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a log management method, electronic device, storage medium, and program product. Background Technology

[0002] In database security management, audit logs serve as crucial records of system operations, playing a vital role in ensuring data security, compliance auditing, and fault tracing. Traditional database audit log management methods primarily rely on relational databases for storage. This approach has significant limitations; when dealing with massive amounts of audit logs, operational complexity increases and management efficiency decreases. Summary of the Invention

[0003] This application provides a log management method, an electronic device, a storage medium, and a program product.

[0004] The technical solution of this application embodiment is implemented as follows: This application provides a log management method, the method comprising: Obtain the database audit logs; wherein the audit logs include logs containing multiple types of operation information generated during the database audit process; The audit logs are preprocessed to obtain preprocessed audit logs; wherein the preprocessed audit logs are standardized audit logs. The preprocessed audit logs are segmented to obtain a set of lexical units; Perform word embedding mapping on each word in the word set to obtain each vector corresponding to each word; Each vector is arranged in a preset order, and hash information is added to each sorted vector to obtain the target storage matrix of the audit log.

[0005] This application provides a log management device, including: An acquisition unit is used to acquire the audit logs of the database; wherein, the audit logs include logs containing multiple types of operation information generated during the database audit process; A processing unit is used to preprocess the audit logs to obtain preprocessed audit logs; wherein the preprocessed audit logs are standardized audit logs. The processing unit is used to perform word segmentation on the preprocessed audit log to obtain a word set; The processing unit is used to perform word embedding mapping on each word in the word set to obtain each vector corresponding to each word; The processing unit is used to arrange each vector in a preset order and add hash information to each sorted vector to obtain the target storage matrix of the audit log.

[0006] This application provides an electronic device, including: Memory is used to store executable instructions or computer programs. The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the method provided in the embodiments of this application.

[0007] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the log management method provided in this application when executed by a processor.

[0008] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the log management method provided in this application.

[0009] The embodiments of this application have the following beneficial effects: by obtaining the audit logs of the database, preprocessing the audit logs, performing word segmentation on the preprocessed audit logs to obtain a word set, performing word embedding mapping on each word in the word set to obtain a vector corresponding to each word, arranging the vectors corresponding to each word in a preset order, and adding hash information to the vectors corresponding to each word after sorting to obtain the target storage matrix of the audit logs, the structured and vectorized storage of audit log data is realized, thereby achieving efficient and standardized management of audit logs. Attached Figure Description

[0010] Figure 1 This is a flowchart illustrating the log management method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the process for managing database audit logs based on vector storage provided in an embodiment of this application; Figure 3 This is a schematic diagram illustrating the data flow and processing method in the database audit log management process provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the log management device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0013] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0014] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0015] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0016] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0017] Figure 1 This is a flowchart illustrating the log management method provided in the embodiments of this application. The following will be combined with... Figure 1 The steps shown are explained as follows: Figure 1 As shown, this method is applied to an electronic device, and the method includes steps 101 to 105. Step 101: Obtain the database audit logs; wherein, the audit logs include logs containing various types of operation information generated during the database audit process.

[0018] In this embodiment, audit logs refer to log data generated during database operation to record all critical operational behaviors. The content of audit logs covers various operation events such as user login, data query, insertion, update, deletion, and permission changes. Audit logs typically exist in text form or structured data format within or outside the database system's log files, used for subsequent security audits, compliance checks, and troubleshooting. The multi-type operation information in audit logs can be understood as follows: audit logs not only contain operation types (such as CRUD operations), but also multi-dimensional data such as operation time, operation subject (user or program), operation object (table name, field name), and operation result (success or failure).

[0019] In this embodiment, the acquisition of audit logs can be accomplished by a log collection module. This module is deployed on a database server or a separate log collection node, and captures audit log data in real time by listening to the database's audit interface, parsing database log files, or calling the database's built-in audit functions. The collected raw audit logs may come from multiple data sources, such as different database instances, application servers, middleware, etc. Therefore, they need to be collected and centrally managed to ensure the integrity and consistency of the audit log data.

[0020] In practical applications, by actively listening to database audit events or periodically scanning audit log files, audit logs covering the entire business chain can be efficiently obtained, providing the raw data foundation for subsequent processing.

[0021] Step 102: Preprocess the audit logs to obtain preprocessed audit logs; wherein, the preprocessed audit logs are standardized audit logs.

[0022] In this embodiment, preprocessing includes, but is not limited to, cleaning and standardizing the original audit logs. The aim is to remove noise, redundancy, and inconsistencies, improve data quality, and facilitate subsequent word segmentation and vector mapping. As an example, preprocessing includes three sub-processes: data cleaning, format unification, and data standardization. Data cleaning refers to removing duplicate records, erroneous characters, irrelevant system information, etc. For example, if an audit log record appears multiple times within a short period, only one should be retained and marked as a duplicate; non-critical debugging information should be filtered out. Format unification refers to standardizing audit logs from different sources into a standard format, such as unifying timestamps to a specific format, standardizing usernames by case, and encoding operation types into a unified identifier. Data standardization includes normalizing specific fields. For example, unifying semantically similar but differently expressed phrases such as "user login success," "login success," and "Login success" into "LOGIN_SUCCESS," or unifying Internet Protocol (IP) addresses into the Internet Protocol version 4 (IPv4) standard format.

[0023] In this embodiment, the preprocessed audit log is the standardized audit log. Standardized audit logs have a clear structure, explicit semantics, and a unified format, effectively supporting subsequent automated processing. For example, a preprocessed audit log entry might look like this: 2024-06-01 14:30:00 LOGIN_SUCCESS user=admin table=accounts action=query result=success. This standardized format makes word segmentation and word embedding processes more accurate, reducing ambiguity and errors.

[0024] In practical applications, by defining key information extraction and transformation functions, massive amounts of raw audit logs are batch-cleaned and standardized, significantly improving data consistency and processing efficiency. Preprocessing these massive amounts of raw audit logs eliminates noise and heterogeneity, laying a solid data foundation for subsequent high-precision vector encoding and efficient storage, while reducing the complexity and computational overhead of subsequent processing steps.

[0025] Step 103: Perform word segmentation on the preprocessed audit logs to obtain a word set.

[0026] In this embodiment, word segmentation is the process of decomposing standardized audit log text into independent semantic units (i.e., tokens). In practical applications, considering the diversity of audit log languages, the word segmentation method is selected based on language characteristics (such as language type). For English audit logs, segmentation is done by spaces and punctuation marks, such as User admin logged insuccessfully, which can be segmented as [User,admin,logged,in,successfully]. For Chinese audit logs, segmentation is done using word segmentation tools because Chinese has no natural separators; for example, User admin logging in successfully needs to be identified as [User,admin,login,successfully]. In addition, specific processing of proper nouns, abbreviations, and compound words can be considered. For example, SQL injection attack should be treated as a whole token rather than split into SQL, injection, and attack.

[0027] In this embodiment, the lexical set is a collection of word segmentation results, where each lexical represents a meaningful linguistic unit that can be used for subsequent semantic analysis and vectorization.

[0028] As an example, an audit log entry from 2024-06-01 14:30:00 LOGIN_SUCCESS user=admintable=accounts action=query result=success, after tokenization, yields the following set of terms: {2024-06-01, 14:30:00, LOGIN_SUCCESS, user, admin, table, accounts, action, query, result, success}. This lexical set retains the key semantic information of the original audit log while eliminating syntactic interference, making it easier for computers to understand and process.

[0029] In practical applications, word segmentation algorithms adaptable to multiple languages ​​and scenarios are employed to ensure the accuracy and completeness of word segmentation. Segmenting the preprocessed audit logs transforms unstructured text into a structured sequence of words, achieving a conversion from natural language to computer-processable basic units. This provides high-quality input for subsequent word embedding models and improves the semantic fidelity of vector mapping.

[0030] Step 104: Perform word embedding mapping on each word in the word set to obtain each vector corresponding to each word.

[0031] In this embodiment, word embedding mapping is the process of converting each word into a fixed-dimensional real-valued vector using a pre-trained word embedding model (such as Word2Vec, GloVe, FastText, etc.). Word embedding models learn semantic relationships between words by training on a large-scale corpus, making semantically similar words closer together in the vector space, thereby capturing the contextual meaning of words. For example, in the Word2Vec model, "login" and "authenticate" have similar vector representations because they often appear in similar contexts; "insert" and "update" are close in the vector space because they are both database operation verbs.

[0032] In this embodiment, the vector mapped to each term is a high-dimensional array of real numbers, with dimensions of 50, 100, 200, or higher, depending on the model configuration. For example, the term LOGIN_SUCCESS might be mapped to a 100-dimensional vector [0.12, -0.45, 0.78, ..., 0.33], where each dimension value reflects the weight of the term on a specific semantic axis. Distance metrics in the vector space (such as cosine similarity) can be used to measure the semantic similarity between terms, and this method has important reference value in subsequent log clustering, anomaly detection, and other analyses.

[0033] In practical applications, a pre-trained model suitable for the audit log domain should be selected, or fine-tuned based on the audit log corpus, to ensure that the mapping vector accurately reflects the semantics of database operations. Word embedding mapping for each word in the word set can transform discrete words into continuous vector representations, realizing the conversion from symbols to numerical values. This greatly improves the machine readability and computational efficiency of audit log data, providing core data support for constructing the vector matrix.

[0034] Step 105: Arrange each vector in a preset order and add hash information to each sorted vector to obtain the target storage matrix of the audit log.

[0035] In this embodiment, the target storage matrix is ​​a two-dimensional matrix structure formed by arranging the mapped vectors in a preset order and adding a unique hash value to each vector. The preset order includes, but is not limited to, ascending order based on the timestamp or identification number (ID) of the audit logs to ensure that the matrix rows correspond to the time sequence of the audit logs. For example, if three audit logs are mapped to vectors v1, v2, and v3 respectively, and the time sequence is v1→v2→v3, then the first row of the matrix is ​​v1, the second row is v2, and the third row is v3. A hash value is appended to the end of each row of vectors. The hash value is calculated from the original mapped vectors or the original content of the corresponding audit logs (such as SHA-256) and used for subsequent anti-tampering verification.

[0036] In this embodiment, the target storage matrix supports efficient indexing and querying. For example, the vector representation of a specific audit log can be quickly located by row number, and aggregation calculations can be performed using column vectors. In practical applications, the target storage matrix provides the data structure foundation for a multi-layered anti-tampering mechanism: row-level hash verification determines whether a single audit log has been tampered with by comparing the current row hash with the cache hash; column-level hash verification monitors whether the total number of audit logs has changed by calculating the hash value of a column (such as the hash column of all audit logs), thereby detecting audit log deletion behavior; file-level and system-level verification further ensure overall data integrity.

[0037] In practical applications, by using ordered arrangement and hash appending, a stable and verifiable target storage matrix is ​​constructed, achieving structured, vectorized, and secure storage of audit log data. Arranging the vectors corresponding to each term in a preset order and adding hash information to the vectors corresponding to each sorted term integrates scattered vector data into a unified matrix. This supports efficient retrieval and analysis while providing a reliable data carrier for anti-tampering mechanisms, significantly improving the security, stability, and processing performance of audit log management.

[0038] In summary, the log management method provided in this application obtains audit logs from the database, preprocesses the audit logs, performs word segmentation on the preprocessed audit logs to obtain a word set, performs word embedding mapping on each word in the word set to obtain a vector corresponding to each word, arranges the vectors corresponding to each word in a preset order, and adds hash information to the vectors corresponding to each word after sorting to obtain the target storage matrix of the audit logs. This achieves structured and vectorized storage of audit log data, thereby realizing efficient and standardized management of audit logs.

[0039] In some embodiments, step 102 preprocesses the audit logs to obtain preprocessed audit logs, including the following steps: Step 201: Call the key information extraction function to clean the audit logs and obtain the cleaned audit logs; wherein, the cleaned audit logs include audit records corresponding to one or more of the operation time, username and operation event.

[0040] In this embodiment, the key information extraction function is a programmatic tool used to extract core business fields from audit logs. The key information extraction function automatically identifies and extracts key information from audit logs using techniques such as regular expression matching, keyword recognition, or structured parsing. This information includes, but is not limited to, operation time (i.e., the timestamp of the log record), username (the user identifier performing the operation), and operation events (such as specific actions like adding, deleting, modifying, querying, or changing permissions). For example, in an audit log entry 2024-05-10 14:30:22 userA UPDATE tableX, the key information extraction function can extract 2024-05-10 14:30:22 as the operation time, userA as the username, and UPDATE tableX as the operation event.

[0041] In this embodiment, audit logs serve as input, and after processing by a key information extraction function, the output is a structured, cleaned collection of audit logs. The cleansing process includes removing duplicate records, filtering irrelevant system information, and standardizing formats (such as unifying the time format to the ISO 8601 standard), thereby improving data quality. In practical applications, the data processing system can adaptively adjust its extraction strategy through configuration rules or machine learning models to handle log data from different sources and formats.

[0042] In this embodiment, the cleaned audit logs form the basis for subsequent vector encoding and storage, and the accuracy of the cleaned audit logs affects the reliability of the entire audit system. By extracting operation time, username, and operation events, the system can establish a clear operation tracing path, facilitating subsequent user behavior analysis, anomaly detection, and compliance review. For example, in a security audit scenario, administrators can quickly locate potential violations by querying the operation events of a specific username within a certain time period.

[0043] Step 202: Call the transformation function to standardize the cleaned audit logs to obtain preprocessed audit logs.

[0044] In this embodiment, the conversion function is a program module that further standardizes the cleaned audit logs. The main function of the conversion function is to convert the extracted key information according to a unified format, encoding standard, and semantic standard, ensuring that log data from different sources and in different formats have a consistent structure and meaning. For example, it converts operation times to Universal Time Coordinated (UTC) timestamp format, maps usernames to globally unique IDs, categorizes operation events into predefined operation type codes, and performs case-sensitive matching and special character escaping on the text content.

[0045] In this embodiment, the transformation function receives cleaned audit logs as input and outputs preprocessed audit logs with a more rigorous structure and clearer semantics. In practical applications, the transformation function can be executed based on a preset mapping table or rule engine, supporting dynamic updates and expansion. For example, when a new database operation type is added, the rule base can be updated to adapt to the new operation event code.

[0046] The preprocessed audit logs exhibit high consistency, providing a high-quality data foundation for subsequent word segmentation, word embedding mapping, and vector matrix construction. Standardization eliminates ambiguity caused by format differences in audit logs, improving data processing efficiency and analytical accuracy. For example, in cross-system auditing scenarios, logs from different database instances, after unified standardization, can be seamlessly integrated into the same vector storage system, achieving unified management from a global perspective.

[0047] In this embodiment, the audit logs are cleaned by calling a key information extraction function to obtain cleaned audit logs containing key fields such as operation time, username, and operation event. Then, a transformation function is used to standardize the cleaned audit logs, ultimately resulting in preprocessed audit logs with a unified structure and clear semantics. This effectively removes noise and redundant information from the audit logs, improving data quality and consistency. This provides a reliable data foundation for subsequent vector encoding, storage, and tamper-proof verification, thereby enhancing the processing efficiency and security of the entire audit log management system in massive data environments.

[0048] In some embodiments, step 103 involves segmenting the preprocessed audit logs to obtain a word set, including the following steps: Step 301: Determine the word segmentation method based on the language type of the preprocessed audit log.

[0049] In this embodiment, language type refers to the type of natural language used in the preprocessed audit logs, such as Chinese, English, Japanese, etc. Different languages ​​have different grammatical structures and lexical features, therefore the word segmentation rules for each language also differ significantly. For example, English uses spaces as separators between words, while Chinese does not have obvious separators and requires segmentation based on semantic and contextual information. In this application, language recognition is first performed on the preprocessed audit logs to determine the main language type used in the audit logs, thereby selecting an appropriate word segmentation method.

[0050] In this embodiment, the word segmentation method is determined by the specific segmentation method or tool used based on the language type. For preprocessed audit logs in English, they can be directly segmented by spaces; for preprocessed audit logs in Chinese, a Chinese word segmentation tool is used for processing. The choice of word segmentation method directly affects the quality and efficiency of subsequent word embedding mapping. For example, when processing preprocessed audit logs containing a large number of technical terms, it may be necessary to configure a specific dictionary or use a domain-adaptive word segmentation model.

[0051] In practical applications, when a preprocessed audit log is received, the language detection module first identifies the language type of the log. If it is English, space-based segmentation is triggered; if it is Chinese, a Chinese word segmentation tool is loaded and the corresponding segmentation strategy is executed. This language recognition and segmentation process can be completed during the preprocessing stage or dynamically determined before segmentation. For example, in a mixed-language database environment, it can automatically distinguish and process preprocessed audit logs of different language types, ensuring the accuracy of the segmentation strategy.

[0052] This mechanism, which determines the segmentation method based on language type, can effectively improve segmentation accuracy and avoid semantic loss or misjudgment caused by incorrect segmentation. Especially in multilingual environments, this mechanism ensures a unified processing standard for audit logs preprocessed in different language types, improving the overall compatibility and robustness of the method.

[0053] Step 302: Perform word segmentation on the preprocessed audit logs based on the word segmentation method to obtain a word set.

[0054] In this embodiment, word segmentation is the process of dividing a continuous text string into independent tokens according to semantic units. Each token represents a meaningful linguistic unit, such as a word, phrase, or symbol. Understandably, word segmentation is the specific word segmentation process performed on the preprocessed audit log text based on a predetermined segmentation method. For example, for the preprocessed audit log entry "user login failed at 2024-01-01 10:00:00" in English, word segmentation will yield the token set {user, login, failed, at, 2024-01-01, 10:00:00}.

[0055] In this embodiment, the lexical set is a collection of all independent lexical units generated by word segmentation. These lexical units serve as the basic units for subsequent word embedding mapping and vector construction. This lexical set not only preserves the key information of the original text but also provides standardized input for subsequent semantic analysis and similarity calculation. For example, when processing a large number of failed login records, high-frequency lexical units such as "failed" and "login" in this lexical set can be used to quickly identify abnormal patterns.

[0056] In this embodiment, the word segmentation method is determined based on the language type, and the preprocessed audit log is segmented accordingly to obtain a word set. This method allows for the adoption of the most suitable word segmentation method for the characteristics of different language types, thereby improving the accuracy and efficiency of word segmentation processing. This ensures the quality of subsequent word embedding and vector storage, and enhances the performance and reliability of the audit log management system.

[0057] In some embodiments, step 105 arranges each vector in a preset order and adds hash information to each sorted vector to obtain the target storage matrix of the audit log, including the following steps: Step 401: Arrange each vector in a preset order to obtain the arranged vector sequence.

[0058] In this embodiment, the preset order refers to the logical sequence of sorting each vector according to specific rules or strategies when constructing the audit log storage structure. This preset order can be determined based on dimensions such as timestamps (according to the order in which audit logs occurred), audit log importance level (e.g., prioritizing high-risk operations), or data source type (e.g., categorized by database module). For example, it can be set to be sorted in ascending order of operation time, meaning each vector corresponding to the earliest generated audit log is placed first; or it can be set to be sorted after grouping by operation type, such as processing login-related audit logs first, followed by CRUD-related audit logs. This ordered arrangement facilitates efficient subsequent queries and consistency checks, while providing a stable input sequence for hash value calculation, avoiding verification failures due to disorder.

[0059] In practical applications, within database auditing scenarios, after multiple audit logs are mapped to each vector, these vectors are reorganized according to a preset order. For example, in a batch containing 100 audit logs, if the preset order is ascending by timestamp, the system will first extract the time field of each audit log, and then arrange the corresponding vectors in chronological order, forming a linear sequence of arranged vectors. This process is typically completed automatically by the background processing engine to ensure data consistency and traceability.

[0060] By adopting this preset order, it is possible to ensure that the vector sequence generated by different batches or different nodes has a unified structure, thereby supporting cross-system and cross-file comparison and integrity verification of audit logs, and significantly improving the standardization and interoperability of the audit system.

[0061] Step 402: Calculate the hash value of each vector in the vector sequence after permutation, and obtain a new set of elements based on each vector after permutation and its hash value.

[0062] In this embodiment, the hash value is a fixed-length, unique, and irreversible digital digest generated by applying a hash function to the input data (here, each vector). Each vector generates a corresponding hash value after hash calculation, used to identify the integrity status of each vector. Hash functions exhibit an avalanche effect; even a small change in the input data can result in a significant change in the output hash value, thus it can be used to detect any tampering. The hash algorithms used for hash calculation include, but are not limited to, SHA-256 and MD5.

[0063] In this embodiment, the new element set is a data combination consisting of each original vector and its hash value. The new element set is a set of tuples (each vector, hash value). For example, for the rearranged vectors v1, v2, and v3, h1, h2, and h3 are calculated respectively, then the new element set is {(v1, h1), (v2, h2), (v3, h3)}. This new element set not only retains the information of each original vector but also adds the hash fingerprint required for tamper-proof verification, thus achieving the binding of data content and security attributes.

[0064] In practical applications, after the sorted vector sequence is generated, the hash function is called one by one to perform calculations on each vector, and the results are paired with the original vectors and stored. This process can be completed in memory or written to a temporary buffer to prepare for the subsequent construction of the initial storage matrix. By constructing a new set of elements, the system can complete the core data preparation for the anti-tampering mechanism in advance without losing the original data, greatly improving the efficiency of subsequent verification.

[0065] Step 403: Construct the initial storage matrix.

[0066] In this embodiment, the initial storage matrix is ​​a two-dimensional data structure used to carry each vector data of the audit log and its associated hash value. The dimension of the initial storage matrix is ​​determined by the number of vector data in the audit log and the dimension of each vector data. Assuming there are n audit logs, each audit log is mapped to d-dimensional vector data, then the initial storage matrix is ​​an n×d matrix, where the i-th row represents the vector data of the i-th audit log, and the j-th column represents the j-th dimension of that vector data. Furthermore, this initial storage matrix can be expanded with one column or one row to store hash values ​​or metadata.

[0067] During the construction process, sufficient memory space is first allocated to accommodate the initial storage matrix structure. Then, each vector portion from the newly generated element set is filled into the corresponding position in the initial storage matrix. For example, if the newly generated element set is {(v1,h1),(v2,h2),...,(vn,hn)}, then v1 is filled into the first row, v2 into the second row, and so on. The initial storage matrix is ​​dynamically expandable, adjusting its size in real time according to the amount of audit logs, making it suitable for scenarios with massive amounts of audit data. As the foundational framework of the entire storage system, the initial storage matrix's structural design balances data storage efficiency and access convenience. Through matrix management, the system can quickly locate each vector data in any audit log using row and column indexes, providing efficient support for subsequent hash value verification, query analysis, and other operations.

[0068] Step 404: Based on the new set of elements, determine the value of the corresponding position of the reference row in the initial storage matrix, and based on the hash values ​​of all the vectors after arrangement and the correlation information between multiple audit logs, determine the value of the corresponding position of the target row in the initial storage matrix to obtain the target storage matrix; wherein, the reference row includes rows in the initial storage matrix that are different from the target row.

[0069] In this embodiment, the reference row refers to all other rows in the initial storage matrix except the target row. These rows serve as a comparison benchmark in the process of determining the target row's value. The target row is the row currently being processed, and its value is determined based on the contextual and relational information provided by the reference rows. The relational information refers to the logical or business relationships between multiple audit logs, such as audit logs from the same batch, the same user operation sequence, or audit log records from the same database instance. This relational information can be used to guide the population strategy for the target row.

[0070] As an example, suppose there are a total of 1 audit record, vector dimension is The target storage matrix is ​​a The storage matrix, This indicates the target storage matrix number 1. Line 1 The elements of the column, where Target behavior Line, reference line includes OK.

[0071] In this embodiment, the target storage matrix is ​​a structured data matrix ultimately used to store and verify audit logs. The target storage matrix integrates each original vector, hash value, impact data of reference rows, and associated metadata, possessing strong anti-tampering capabilities and data integrity assurance. For example, when an audit log is illegally deleted, because each original vector and hash value corresponding to that audit log has been associated with other rows in the target storage matrix, the system can detect data loss or anomalies at that location through verification, thereby triggering an alarm.

[0072] In practical deployments, this target storage matrix can be stored in a dedicated vector database or a distributed storage system, supporting parallel read / write and efficient retrieval. By introducing the reference row and the associated information, this application achieves cross-validation between the audit log data, making it difficult to hide tampering with a single audit log entry, thus greatly enhancing the security and reliability of the audit system.

[0073] In this embodiment, each vector is arranged in a preset order to obtain a sequence of arranged vectors. The hash value of each vector is calculated to construct a new set of elements, thereby constructing an initial storage matrix. Finally, the value of the target row is determined based on the reference row and related information to generate the target storage matrix. This enables structured and ordered storage of audit logs, effectively supporting subsequent efficient queries and tamper-proof verification, thus improving the overall security and operational efficiency of the database audit system.

[0074] In a feasible scenario, the system first iterates through the new set of elements. For each tuple (each vector, hash value), each vector part is inserted into a row (target row) of the initial storage matrix. Simultaneously, a suitable reference row is selected based on correlation information. For example, if the correlation information indicates that the current audit log belongs to the successful login category, the system might select the row containing the most recent successful audit log in the same category as the reference row to maintain data semantic consistency. Subsequently, based on the value of the reference row and the hash value of the target row, the system calculates the final value of the target row using a weighted average, concatenation, or other algorithms, and updates the initial storage matrix.

[0075] In some embodiments, the above method further includes the following steps: Step 501: Obtain the audit log query request.

[0076] In this embodiment, an audit log query request refers to an instruction initiated by a user or system to retrieve database audit log data. An audit log query request typically includes query conditions, such as time range, operation type, username, and IP address, to limit the scope and content of the query. Audit log query requests can be issued by database administrators, security auditors, or automated monitoring systems, with the purpose of reviewing database operations within a specific time period, investigating abnormal events, or conducting compliance checks.

[0077] In this embodiment, the format of the audit log query request can be determined according to the system design. The request form includes one or more of the following: Structured Query Language (SQL) statements, Application Programming Interface (API) call parameters, or filter conditions in a graphical interface. For example, in a web management platform, a user can generate an audit log query request by selecting operation type = delete and time range = January 1, 2024 to January 31, 2024; in a command-line tool, an audit log query request may be triggered by executing a command like audit_log_query --type=DELETE --start_time=2024-01-01 --end_time=2024-01-31.

[0078] In practical applications, upon receiving an audit log query request, the request is first parsed to extract the query conditions, and the legality and permissions of the query request are verified. Only users with the appropriate permissions can initiate audit log query requests for specific scopes or sensitive types to ensure the security and privacy of audit data. Subsequently, the audit log query request processing flow is initiated based on these conditions.

[0079] Here, by obtaining audit log query requests, on-demand access and efficient retrieval of database audit logs can be achieved. This ensures that users can only query the data they need within their own permission scope, thereby improving system security and manageability, and ultimately supporting more accurate fault tracing and compliance auditing.

[0080] Step 502: Perform a query based on the verification level corresponding to the audit log query request to obtain the query results.

[0081] In this embodiment, the verification level refers to the anti-tampering verification mechanism at different granular levels used to determine the integrity and authenticity of log data during the audit log query process. In practical applications, the verification level includes four levels: row level, column level, file level, and system level. Row-level verification performs hash calculations on a single audit log record to detect whether the record has been tampered with; column-level verification calculates the hash value of each column of data to monitor whether the number of records in that column has changed abnormally (such as records being deleted); file-level verification performs hash verification on the entire data of a single audit log file to ensure the file's integrity; and system-level verification integrates the hash values ​​of all files for final verification to ensure the security of the audit log system.

[0082] During the actual query process, the system automatically matches the verification level corresponding to the audit log query request based on the log range involved. For example, if the audit log query request only involves a specific log record, the system will prioritize row-level verification; if the audit log query request covers multiple log records or the entire file, the system will sequentially perform column-level, file-level, and even system-level verifications to ensure the completeness and reliability of the returned query results.

[0083] Furthermore, to improve query efficiency, the system can combine parallel computing acceleration technology and caching mechanisms. By utilizing the vectorized instruction set of the Central Processing Unit (CPU) (such as Single Instruction Multiple Data (SIMD)), multiple data vectors can be processed simultaneously, significantly shortening hash calculation time. For frequently queried log data, the system can pre-calculate the hash value of the log data and store it in the cache to avoid redundant calculations, further improving response speed.

[0084] In this embodiment, by querying based on the verification level corresponding to the audit log query request, multi-level and high-precision data integrity verification can be achieved. This effectively identifies any tampering with log records (including deletion, modification, etc.), ensuring the authenticity and reliability of the query results, and thus enhancing the security and trustworthiness of the database auditing system.

[0085] In some embodiments, step 502 performs a query based on the verification level corresponding to the audit log query request to obtain the query results, including the following steps: Step 601: When an audit log query request covers multiple audit records, perform row-level hash verification.

[0086] In this embodiment, row-level hash verification refers to verifying whether an audit record has been tampered with by calculating the hash value of its vector field for each individual audit record. In practical applications, audit records are stored in vector form, with each audit record corresponding to a vector. This vector is generated from the log content after word segmentation and word embedding mapping. Row-level hash verification can accurately identify any modification or deletion of a single audit record, ensuring data integrity. For example, in the database operation log, if the content of a data update operation performed by user A on January 1, 2024, in an audit record is tampered with and changed to user B, the vector representation of that audit record will change, resulting in a hash value mismatch and triggering an alarm. This row-level hash verification mechanism overcomes the deficiency of traditional audit record verification methods in detecting audit record deletion, significantly improving anti-tampering capabilities.

[0087] Step 602: During the verification process, the hash value is calculated in parallel based on the target number of vector fields for each audit record hit by the audit log query request.

[0088] In this embodiment, the target number of vector fields refers to the set of vector fields participating in the hash calculation, determined based on the key information dimensions (such as operation time, username, operation type, etc.) involved in the audit log query request. Parallel computing refers to using the CPU's vectorized instruction set (such as SIMD) to perform hash calculations on multiple vector fields simultaneously, thereby significantly improving computational efficiency. For example, when a query request involves 1000 audit records, and each audit record contains 5 key vector fields, the system can group these 5000 fields for parallel processing. Each vector group completes the hash calculation under one instruction, greatly shortening the overall processing time. The above optimization method effectively reduces the hash overhead in massive log scenarios and improves system response speed and stability.

[0089] Step 603: Compare the calculated hash value with the pre-cached hash value.

[0090] In this embodiment, the pre-cached hash value is the baseline hash value calculated and stored by the system for the vector field of each audit record when the audit log is written or updated. The pre-cached hash value is stored in the cache for quick retrieval during subsequent queries. The comparison process checks whether the currently calculated hash value is consistent with the pre-cached hash value. If they are consistent, it indicates that the audit record has not been tampered with; if they are inconsistent, it indicates that the audit record may have been modified or deleted. For example, when the system receives a query request, it first retrieves the pre-cached hash value of the queried audit record from the cache, and then compares it with the real-time calculated hash value of the audit record. This mechanism combines caching technology, avoids redundant calculations, and significantly improves query efficiency, especially suitable for audit log scenarios with high-frequency access.

[0091] Step 604: When the comparison result indicates that the calculated hash value is the same as the pre-cached hash value, the corresponding audit record is restored as a readable log as the query result.

[0092] In this embodiment, restoring the logs to a readable format refers to converting the vectorized audit records, which have undergone hash verification and are confirmed to be tamper-free, back into their original text or structured format for user viewing. Since the audit logs are encoded as vectors during storage, the restoration process requires reverse mapping (such as deducing terms from the vectors, reconstructing sentences, etc.) to recover the original content. For example, after confirming that the hash value of an audit record matches, the system decodes the vector sequence of the audit record into a natural language description such as "User X successfully logged into the system on January 1, 2024," making it easier for administrators to read and analyze. The process of restoring the logs to a readable format ensures the credibility and usability of the query results, achieving a balance between security and ease of use.

[0093] Step 605: When the comparison result indicates that the calculated hash value is different from the pre-cached hash value, the first alarm information is used as the query result; wherein, the first alarm information indicates that the corresponding audit record has been tampered with.

[0094] In this embodiment, the first alarm message is a predefined alarm message used to explicitly inform the user that a certain audit record is at risk of being tampered with. The first alarm message includes, but is not limited to, the unique identifier of the tampered audit record (such as the audit record ID), the time when the tampering occurred, and suggested actions (such as freezing the account, tracing the source, etc.). For example, when the system detects that the hash value of an audit record does not match, it will return the first alarm message: Audit record ID 123456 has been tampered with, please check immediately. This mechanism can promptly expose potential security threats, prevent unauthorized operations from being covered up, and ensure the integrity and compliance of the audit system.

[0095] In this embodiment, row-level hash verification is performed when an audit log query request covers multiple audit records. Hash values ​​are calculated in parallel based on a target number of vector fields. The calculated results are compared with pre-cached hash values. When a match is found, the log is restored to a readable record; otherwise, a first alarm message is output. This enables precise tampering detection of individual audit records, effectively identifying malicious actions such as deletion or modification of audit records. This ensures the integrity and reliability of audit logs, improving the security and response efficiency of the database audit system.

[0096] In some embodiments, step 502 performs a query based on the verification level corresponding to the audit log query request to obtain the query results, including the following steps: Step 701: When an audit log query request covers some or all of the audit records in a single audit file, perform column-level hash verification.

[0097] In this embodiment, column-level hash verification refers to the process of performing hash calculations and verifications on a specific column of data (such as a timestamp column, operation type column, etc.) within a single audit file. By calculating hash values ​​column by column, column-level hash verification can effectively detect whether a column of data has undergone additions, deletions, or modifications. Compared to traditional full-file hashing, column-level hash verification offers higher efficiency and finer-grained detection capabilities, making it suitable for scenarios where only a subset of columns of data needs to be queried. For example, in an auditing system, if it is only necessary to verify whether the timestamp of a certain operation has been tampered with, then only column-level hash verification needs to be performed on that column, avoiding redundant calculations on the entire single audit file. Column-level hash verification can be applied to various audit log structures, such as logs based on relational table structures, JSON format logs, or custom structured logs, as long as the log structure has a clear column division.

[0098] Step 702: During the verification process, the corresponding aggregate hash value is calculated based on the sequence of row hash values ​​stored at the end of all lines in a single audit file.

[0099] In this embodiment, the row hash value is a unique identifier generated after each line of audit log record has been processed by a hash algorithm and is stored at the end of the line data. The sequence of row hash values ​​is used for subsequent aggregate hash calculation. The aggregate hash value is the result of performing a secondary hash operation (such as SHA-256, MD5, etc.) on all row hash values ​​in the sequence of row hash values. The aggregate hash value represents the integrity digest of the entire single audit file in its current state. The calculation process of the aggregate hash value ensures that even if a line of data is modified, the aggregate hash value will change, thereby achieving efficient verification of data integrity. In practical applications, the row hash value can be represented by a fixed-length hexadecimal string, such as 32 bits or 64 bits, to ensure a balance between computational efficiency and storage space. The calculation method of the aggregate hash value can be a simple concatenation hash or a more secure Merkle tree structure, depending on the system's security requirements and performance requirements.

[0100] Step 703: Compare the calculated aggregate hash value with the aggregate hash value stored at the end of the first row of the target storage matrix.

[0101] In this embodiment, the target storage matrix refers to a two-dimensional matrix structure that organizes and stores audit log data in vector form, where each row corresponds to an audit record and each column corresponds to a field or feature dimension. During storage, the target storage matrix pre-stores an aggregate hash value at the end of the first row. This aggregate hash value is calculated when a single audit file is written and serves as the integrity benchmark for that single audit file in its original state. During an audit log query request, the system recalculates the aggregate hash value of the current single audit file and compares it with the aggregate hash value stored at the end of the first row of the target storage matrix. If the currently calculated aggregate hash value matches the aggregate hash value stored at the end of the first row of the target storage matrix, it indicates that the content of the single audit file has not been tampered with; if they do not match, it indicates that data has been modified. This write-time calculation and read-time verification mechanism ensures data security while avoiding frequent global hash calculations. The target storage matrix supports various storage media, such as in-memory databases, solid-state drives (SSDs), or distributed file systems. The physical storage format of the target storage matrix can be a binary file, a JSON array, or a dedicated vector database format.

[0102] Step 704: When the calculated aggregate hash value is the same as the aggregate hash value stored at the end of the first row of the target storage matrix, the corresponding audit record is restored as a readable log as the query result.

[0103] In this embodiment, when the aggregate hash value matches successfully, the system confirms that the queried audit record has not been tampered with. Therefore, the audit record can be securely restored from vector-encoded format to the original readable log format. The restoration process includes mapping the vector data back to the original text content, such as recovering the segmented text through the inverse process of a word embedding model, then reorganizing the format, and finally outputting a readable log that conforms to a standard log format (such as JSON, CSV, or a custom structure). The restoration process ensures that users obtain authentic and complete audit information, which can be used for compliance review, troubleshooting, or security analysis. The restored log retains all key fields of the audit record, such as operation time, username, operation type, and IP address, and supports multiple output formats to adapt to different use cases.

[0104] Step 705: When the calculated aggregate hash value is different from the aggregate hash value stored at the end of the first row of the target storage matrix, the second alarm information is used as the query result; wherein, the second alarm information indicates that the corresponding audit record has been tampered with.

[0105] In this embodiment, the second alarm message is a predefined error or anomaly response signal used to alert the security monitoring system that audit logs may have been tampered with. The second alarm message includes metadata such as alarm level, affected file path, timestamp, and tamper probability assessment, which can trigger further security response processes, such as automatically notifying the administrator, isolating the tampered file or data source, or initiating a log recovery mechanism. The design of the second alarm message follows standardized alarm protocols to ensure compatibility with the security monitoring system. In actual deployment, the second alarm message can be presented to operations and maintenance personnel through highlighting, pop-up notifications, or email pushes to help quickly locate the source of the problem. Furthermore, the security monitoring system can be configured with different alarm strategies, such as triggering only when the aggregate hash value does not match, or combining other verification mechanisms (such as row-level hashing) for multi-level confirmation to reduce the false alarm rate.

[0106] In this embodiment, column-level hash verification is performed when an audit log query request covers some or all records of a single audit file. An aggregated hash value is calculated based on a sequence of row hash values ​​and compared with the aggregated hash value stored at the end of the first row of the target storage matrix to determine data integrity. By performing column-level hash verification when an audit log query request covers some or all records of a single audit file, fine-grained tamper-proof verification of audit records can be achieved. This avoids the high computational overhead of traditional full-file hashing, thereby improving the system's query efficiency and security in massive log environments.

[0107] In some embodiments, step 502 performs a query based on the verification level corresponding to the audit log query request to obtain the query results, including the following steps: Step 801: When an audit log query request covers a single audit log file, perform a file-level hash check.

[0108] In this embodiment, file-level hash verification refers to an integrity verification process performed on a single audit log file as a whole. This hash verification process calculates the overall hash value of the contents of a single audit log file and compares it with an expected value to determine whether the single audit log file has been tampered with. File-level hash verification is a crucial component of a multi-layered anti-tampering mechanism, ensuring data consistency of a single audit log file during storage and transmission. In practical applications, when a user initiates a query request for a specific audit log file, the system first determines whether the query request involves only one specific audit log file. If so, the file-level hash verification process is triggered. This file-level hash verification mechanism avoids redundant verification of the entire system log, improving query efficiency.

[0109] Step 802: During the verification process, the overall hash value is calculated based on the vector storage matrix of a single audit log file.

[0110] In this embodiment, the vector storage matrix is ​​a two-dimensional data structure composed of vectors of multiple audit log records arranged sequentially. Each audit log record is transformed into a fixed-dimensional vector after word segmentation and word embedding mapping, and these vectors are arranged in rows to form the vector storage matrix. When calculating the overall hash value, the system uses a hash function to process all elements in the vector storage matrix to generate an overall hash value that uniquely identifies the content of a single audit log file. Compared with traditional methods, the hash calculation method based on the vector storage matrix can more efficiently identify data changes, including the deletion or modification of a single audit log record. For example, in a single audit log file containing 1000 audit log records, the vector storage matrix can quickly locate any tampered audit log record without comparing the original text one by one.

[0111] Step 803: Compare the calculated overall hash value with the overall hash value corresponding to the previous file of a single audit log file.

[0112] In this embodiment, the comparison operation of the current single audit log file is used to verify its integrity and continuity with historical single audit log files. The overall hash value corresponding to the previous single audit log file is used as a baseline. If the overall hash value of the current single audit log file matches the overall hash value corresponding to the previous single audit log file, it indicates that the current single audit log file has not been tampered with and the data chain is complete; if they do not match, it indicates that the current single audit log file may have been maliciously modified or damaged. This chain-checking mechanism effectively prevents isolated single audit log files from being tampered with undetected. In practical applications, single audit log files are numbered chronologically, such as the first single audit log file, the second single audit log file, etc. The system automatically obtains the overall hash value of the second single audit log file and compares it with the overall hash value corresponding to its predecessor, the first single audit log file, to ensure the credibility of the audit log sequence.

[0113] Step 804: When the calculated overall hash value is the same as the overall hash value corresponding to the previous file, restore the corresponding audit log file to a readable log as the query result.

[0114] In this embodiment, restoring to a readable log refers to converting the vector storage matrix, which has been verified to be complete through hash checking, back into a readable log. The restoration process includes reverse mapping of vectors to original terms, reconstructing sentence structure, and restoring key information such as timestamps and operation types. After confirming the integrity of a single audit log file, the system initiates the decoding process to restore the vector storage matrix into a readable log for users to view and analyze. This restoration operation ensures the authenticity and availability of query results, enabling administrators to accurately trace database operation behavior.

[0115] Step 805: When the calculated overall hash value is different from the overall hash value corresponding to the previous file, the third alarm information is used as the query result; wherein, the third alarm information indicates that the corresponding single audit log file has been tampered with.

[0116] In this embodiment, the third alarm message is a security alert notification automatically generated by the system, used to clearly inform the user that a single audit log file is at risk of being tampered with. The third alarm message typically includes the name of the single audit log file, the time of tamper detection, the specific location of the verification failure (e.g., line X), and suggested follow-up actions (e.g., isolating the single audit log file, initiating a recovery process). In a real-world scenario, when the system detects that the overall hash value of the current single audit log file does not match the overall hash value of the previous single audit log file, it immediately generates a third alarm message and notifies the security administrator via email, push notification, or console pop-up, so that timely countermeasures can be taken.

[0117] In this embodiment, a file-level hash verification mechanism based on a vector storage matrix is ​​introduced during the audit log query process. This allows for accurate identification of whether a single audit log file has been tampered with, thereby ensuring the integrity and reliability of the audit log data and ultimately improving the security protection capabilities and compliance management level of the database system.

[0118] In some embodiments, step 502 performs a query based on the verification level corresponding to the audit log query request to obtain the query results, including the following steps: Step 901: When an audit log query request indicates that the full audit logs should be read across multiple audit files, perform a system-level hash check.

[0119] In this embodiment, system-level hash verification refers to a mechanism for uniformly verifying the integrity of all audit files in the entire audit log system. System-level hash verification integrates the hash values ​​of multiple audit files to calculate a global aggregate hash value, which is then compared with a preset global verification value to determine whether the entire audit log has been tampered with. System-level hash verification is a further extension of file-level verification, ensuring that even if a single file has not been tampered with, situations where the overall dataset may still be corrupted due to missing files, out-of-order data, or content modification can still be identified. In practical applications, when a user initiates a request to read all audit logs, the system automatically triggers the system-level hash verification process to ensure data integrity. For example, in a database audit system, when an administrator needs to export all operation records for compliance review, the system will first perform system-level hash verification to confirm that the data has not been tampered with before providing the download service. System-level hash verification is particularly suitable for high-security scenarios, such as finance and government affairs, where data integrity requirements are extremely high.

[0120] Step 902: During the verification process, calculate the global aggregate hash value based on the complete matrix set of all audit files.

[0121] In this embodiment, the complete matrix set refers to a unified matrix structure formed by merging the vector storage matrices constructed by each of the multiple audit files according to their sequence numbers. Each audit file has constructed its own independent storage matrix during the vector storage stage. Each of these independent storage matrices contains the vector mapped to each audit log record and its corresponding row hash value. In system-level verification, the system concatenates the matrices corresponding to all audit files into a complete matrix set in chronological or sequential order, which serves as the input for global hash calculation. The global aggregate hash value is a unique identifier value generated by performing a hash operation on all elements (including vectors and hash values) in the complete matrix set. The calculation process can adopt recursive hashing or group hashing. For example, the system first calculates the file-level hash value of the matrix corresponding to each audit file, then concatenates these hash values ​​and hashes them again to finally form the global aggregate hash value. For example, in a system containing 5 audit files, the system first calculates the file-level hash values ​​H1~H5 of the matrix corresponding to each audit file, then concatenates the five file-level hash values ​​into the string H1+H2+H3+H4+H5, and then performs a SHA-256 hash operation on this string to obtain the global aggregate hash value. This approach ensures computational efficiency while avoiding the performance overhead of directly processing massive amounts of raw data.

[0122] Step 903: Compare the calculated global aggregate hash value with the preset global verification value.

[0123] In this embodiment, the preset global checksum is a global aggregate hash value pre-calculated and securely stored by the system after the audit log system initialization or each full backup. The preset global checksum serves as a baseline reference for comparison in all subsequent system-level checks. The comparison process is completed by comparing the hash values ​​bit by bit. If two hash values ​​are completely identical, it indicates that the complete matrix set of all audit files has not changed; if any bit differs, it indicates that at least one file's data has been tampered with or deleted. In actual systems, the preset global checksum is usually stored in read-only storage media or an encrypted database to prevent unauthorized modification. For example, when the system automatically backs up every morning, it recalculates the global aggregate hash value and updates the preset global checksum to ensure that the next day's check reflects the latest data status. The above comparison mechanism has high sensitivity and a low false alarm rate, and can accurately identify any form of data change, including single log deletion, multiple log insertion, or file order adjustment.

[0124] Step 904: When the calculated global aggregate hash value is the same as the preset global verification value, restore all audit files to readable logs as the query result.

[0125] In this embodiment, a readable log refers to restoring the vector-encoded audit log data to its original text or structured format for user reading and analysis. In the vector storage system, each log record is converted into a vector and stored in a matrix, while retaining the original text information or metadata. After system-level verification passes, the system initiates the log restoration process, extracting the original log content from the storage matrix and organizing the output according to time sequence or event type. The restoration process includes steps such as vector deconstruction, field mapping, and format restoration. For example, the system reads a row vector from the matrix, reverse-engineers the original word sequence using a predefined word embedding model, and then combines it with metadata such as timestamps and usernames to reconstruct a complete text log showing that user Zhang San performed a SELECT operation on January 1, 2024, at 10:00:00. The log restoration process supports multiple output formats, such as CSV, JSON, and XML, to meet the needs of different users. A readable log maintains the original semantic integrity, facilitating subsequent audit analysis, report generation, or legal evidence collection.

[0126] Step 905: When the calculated global aggregate hash value is different from the preset global verification value, the fourth alarm information is used as the query result; wherein, the fourth alarm information indicates that the full audit log has been tampered with.

[0127] In this embodiment, the fourth alarm message is a predefined anomaly response signal used to notify the system or user that the full audit log data involved in the current query request has been tampered with. The fourth alarm message typically includes key elements such as alarm level, timestamp, affected file range, and recommended actions. When the system detects a mismatch in the global aggregate hash value, it immediately generates the fourth alarm message and returns it as a query result, preventing the user from obtaining the tampered data and thus protecting system security. For example, the fourth alarm message returned by the system might be: Critical Alarm: Full audit log has been tampered with, Detection Time: 2024-01-01 10:00:00, Affected Scope: File ID#3-ID#5, Recommendation to immediately initiate the emergency response process. This mechanism implements a security policy of verification before access, effectively preventing the leakage of tampered log data or its use in erroneous decisions. Simultaneously, the fourth alarm message can trigger automatic notifications, log recording, and permission freezing, enhancing overall security protection capabilities.

[0128] In this embodiment, when an audit log query request indicates that the full audit logs across multiple audit files should be read, a system-level hash verification is performed. A global aggregate hash value is calculated based on the complete matrix set of all audit files, and this global aggregate hash value is compared with a preset global verification value. This allows for accurate identification of whether the full audit logs have been tampered with, preventing the access or dissemination of tampered log data. This, in turn, ensures the data integrity and security of the database audit system and improves the credibility and compliance of audit results.

[0129] Figure 2 This is a schematic diagram of the database audit log management process based on vector storage provided in an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: Step 21: The database of the electronic device generates audit logs.

[0130] Step 22: The electronic device performs audit log vector storage.

[0131] Step 23: The user initiates an audit log query using an electronic device.

[0132] Step 24: The electronic device performs an anti-tampering verification. If the verification is successful, proceed to step 25; otherwise, proceed to step 26.

[0133] Step 25: The electronic device returns the query results.

[0134] Step 26: The electronic device issues an alarm.

[0135] Furthermore, the above process will be further explained in conjunction with the data flow. Figure 3 This is a schematic diagram illustrating the data flow and processing method in the database audit log management process provided in the embodiments of this application, such as... Figure 3 As shown, Step 31: Audit log data collection and preprocessing to obtain preprocessed data.

[0136] Database auditing requires collecting various types of audit logs, such as user login success or failure events, data CRUD operations, and permission changes. These logs typically exist in different data sources in text or structured data form and need to be collected uniformly. The collected raw logs may contain a large amount of noise and redundant information, such as irrelevant system information and duplicate records, thus requiring preprocessing. Preprocessing operations include data cleaning (removing noisy characters and duplicate records), format standardization (e.g., standardizing time formats and capitalization), and data standardization (normalizing specific fields).

[0137] Here, data cleaning includes: removing duplicate records, assuming the original audit log set is... ,in For the first 1 log record This represents the total number of log entries. Define a function to extract key information. This function can be system-defined and used to retrieve log records. Extract key information (such as operation time, username, etc.) from the cleaned log set. Then, data standardization is performed, transforming the original data format using a conversion function. This function can be defined by the system and obtained. .

[0138] Step 32: Word segmentation and word embedding mapping.

[0139] Here, the preprocessed audit log text is segmented into individual words. The segmentation method differs depending on the language of the logs; English logs can be segmented by spaces, while Chinese logs use a dedicated segmentation tool. This transforms the text information into basic units that facilitate subsequent processing. Let the segmentation function be... The set of lexical units after word segmentation is .

[0140] Then, pre-trained word embedding models (such as Word2Vec, GloVe, etc.) are used to map the word segments obtained from word segmentation into vectors. These models can learn the semantic relationships between words, converting each word segment into a fixed-dimensional vector, so that semantically similar words are closer together in the vector space. Let the word embedding model be... The original standardized text audit log records are then transformed into vectors. .

[0141] Step 33: Map the logs to the resulting audit log vectors and arrange them in order to build a storage matrix.

[0142] Here, before constructing the matrix, each line needs to be... Hash the audit log vector This facilitates verification of the anti-tampering mechanism during subsequent audit log queries. Assume there are a total of... 1 audit record, vector dimension is Then the storage matrix is ​​a The matrix, Indicates the first Line 1 The elements of the column, where .in, This facilitates verification with subsequent files and column-level anti-tampering mechanisms. When there are multiple system audit files, ,in, This refers to the audit file number. The final target storage matrix for a single file is represented as follows: ; Here, regarding the target storage matrix, for Okay, about When multiple audit files exist, if the audit file number is , Depend on The hash is obtained by calculating the hash, that is It is used for correlation verification between different audit documents. By analyzing all audit log vectors ( From 1 to hash value The hash is calculated again to obtain the result, i.e. This is used for subsequent verification of anti-tampering mechanisms at the file and column levels.

[0143] Here, regarding the target storage matrix, for OK, Its meaning can be set and assigned values ​​according to actual needs. For example, it can store identification information such as time and user related to the audit record. Taken from audit log vector The Each dimension component is used to store the specific data content of the audit log vector. Taken from audit log vector hash value This is used to perform individual anti-tampering verification on each audit log vector.

[0144] Furthermore, based on the constructed storage matrix, a method for verifying database audit logs can also be implemented.

[0145] This embodiment provides an optimized hash calculation method based on audit log vector storage. This method accelerates computation through parallel processing, allowing multiple data elements to be processed simultaneously under the control of a single instruction. In hash calculation, vectorized instructions can be used to process multiple data vectors concurrently, thereby improving computational efficiency. Assuming a single vectorized instruction can process... There are 10 data elements, and the hash calculation operation time for a single data element is 1000. If vectorization instructions are not used, the calculation The time required to hash a data element for: ; When using vectorized instructions, Each data element is divided into Given a set of vector groups, where the computation time for each vector group is similar to the time for processing a single data element (because it's parallel processing), the time required to calculate the hash of each data element using vectorized instructions is... for: ; For frequently used audit log data, its hash value can be pre-calculated and cached (i.e., the hash value of each line is...). When the same data needs to be hashed again, the result is retrieved directly from the cache to avoid duplicate calculations. Let the total number of data queries be... Among them are The data for this query is already in the cache (hit rate is...). The time required to calculate the hash each time is... The time to retrieve the hash result from the cache is (generally When using caching, the total query time is... for: ; Based on the aforementioned hash acceleration method, hash calculation under audit log vector storage... Upgraded to ; here, , Number of groups Therefore, it can be seen that the efficiency of hash calculation has been significantly improved.

[0146] In this embodiment, an anti-tampering mechanism based on audit log vector storage is provided, which includes the following strategies: (1) A single audit log record is used as a row-level hash check for the following calculations and judgments: ;in, Indicates the storage matrix row 0, page number 1 The element of the column corresponds to the first... The column-level check hash value is used to complete the column-level integrity check of the corresponding dimension; Judge the calculation result and If the records are not equal, the audit record is suspected of being tampered with. Represents the storage matrix of the first Line 1 The element of the column, that is, the first The hash value of each audit log vector is the core evidence value for row-level verification of a single log entry.

[0147] Here, before recalculating the vector of a single audit log entry... The hash value of the dimensional content, and the pre-stored For comparison, leveraging the strong sensitivity of hash algorithms, if any vector bit in a single log entry is modified, the recalculated hash value will differ from the original stored hash value. Inconsistency can be precisely pinpointed to the individual log line that has been tampered with.

[0148] (2) How to calculate and determine whether the log count has been tampered with, using column-level hash verification: ;in, Indicates the storage matrix row 0, page number 1 The elements of the column are the secondary aggregate hashes of all individual audit log hashes, used to verify whether the total number of audit logs has been tampered with.

[0149] Compare the calculation result with the pre-stored data. If they are not equal, the audit date quantity is suspected of being tampered with.

[0150] Here, the aggregate hash of the last column hash value of all rows is recalculated, along with the pre-stored hash. In contrast, if an attacker adds or deletes any audit log entry, the set of elements in that column will change, and the recalculated aggregate hash will no longer match the original. Matching allows for quick identification of actions that could tamper with the number of log entries.

[0151] (3) To determine whether a single audit log file has been tampered with, the following calculations and judgments are performed using file-level hash verification: If it is the first file: ; The storage matrix of the first file of These are the initial values ​​set by the system. Subsequent file verification calculations are as follows: ; Compare the calculation results with the storage matrix of subsequent files. If the values ​​are not equal, the audit log file is suspected of being tampered with.

[0152] Here, file-level hash verification recalculates the hash value of the complete matrix of the previously audited file and compares it with the hash value pre-stored in the current file. The comparison is performed. By using a chain hashing method, the audit files are linked together to form an immutable chain structure. If any historical file is modified, the hash verification of all subsequent files will become invalid, thus achieving file-level integrity verification.

[0153] The hash verification of the entire audit log system is performed as follows: ; Here, the calculation It is the global aggregate hash of the entire audit log system, the overall hash value of the matrix of all independent audit files, used to verify the overall integrity of all audit files in the entire system. The calculated result is compared with the pre-stored... If they are not equal, the audit log system is suspected of being tampered with.

[0154] Here, system-level hash verification recalculates the aggregate hash of the matrix of all audit files across the entire system, and compares it with the pre-stored hash. By making comparisons, the overall integrity verification of all audited files can be covered at once. As long as the content of any file is modified, the final calculated global hash will match the original hash. Mismatches enable global anti-tampering verification across the entire audit system.

[0155] Based on the above calculation and judgment methods, as well as parallel computing and caching mechanisms, the optimized time complexity comparison can be obtained as shown in Table 1:

[0156] Table 1 in, These represent the number of files, columns, and rows, respectively.

[0157] This application also provides a log management device, such as... Figure 4 As shown, the device includes: The acquisition unit 41 is used to acquire the database audit logs; wherein, the audit logs include logs containing multiple types of operation information generated during the database audit process; Processing unit 42 is used to preprocess the audit logs to obtain preprocessed audit logs; wherein, the preprocessed audit logs are standardized audit logs. Processing unit 42 is used to perform word segmentation on the preprocessed audit log to obtain a word set; Processing unit 42 is used to perform word embedding mapping on each word in the word set to obtain each vector corresponding to each word; Processing unit 42 is used to arrange each vector in a preset order and add hash information to each sorted vector to obtain the target storage matrix of the audit log.

[0158] In some embodiments, the processing unit 42 is used to call a key information extraction function to clean the audit logs and obtain cleaned audit logs; wherein, the cleaned audit logs include audit records corresponding to one or more of the operation time, username and operation event; and call a transformation function to standardize the cleaned audit logs to obtain preprocessed audit logs.

[0159] In some embodiments, the processing unit 42 is used to determine the word segmentation method based on the language type of the preprocessed audit log; and to perform word segmentation processing on the preprocessed audit log based on the word segmentation method to obtain a word set.

[0160] In some embodiments, the processing unit 42 is used to arrange each vector in a preset order to obtain an arranged vector sequence; Calculate the hash value of each vector in the vector sequence after permutation, and obtain a new set of elements based on each vector after permutation and its hash value; construct an initial storage matrix; based on the new set of elements, determine the value of the corresponding position of the reference row in the initial storage matrix, and determine the value of the corresponding position of the target row in the initial storage matrix based on the hash values ​​of all vectors after permutation and the correlation information between multiple audit logs, so as to obtain the target storage matrix; wherein, the reference row includes rows in the initial storage matrix that are different from the target row.

[0161] In some embodiments, the acquisition unit 41 is used to acquire an audit log query request; the processing unit 42 is used to perform a query based on the verification level corresponding to the audit log query request and obtain the query result.

[0162] In some embodiments, the processing unit 42 is configured to perform row-level hash verification when the audit log query request covers multiple audit records; during the verification process, it calculates hash values ​​in parallel based on the target number of vector fields for each audit record hit by the audit log query request; compares the calculated hash values ​​with pre-cached hash values; when the comparison result indicates that the calculated hash value is the same as the pre-cached hash value, it restores the corresponding audit record as a readable log as the query result; when the comparison result indicates that the calculated hash value is different from the pre-cached hash value, it uses the first alarm information as the query result; wherein, the first alarm information indicates that the corresponding audit record has been tampered with.

[0163] In some embodiments, the processing unit 42 is configured to perform column-level hash verification when an audit log query request covers some or all of the audit records in a single audit file; during the verification process, it calculates the corresponding aggregate hash value based on the sequence of row hash values ​​stored at the end of all rows in a single audit file; compares the calculated aggregate hash value with the aggregate hash value stored at the end of the first row of the target storage matrix; when the calculated aggregate hash value is the same as the aggregate hash value stored at the end of the first row of the target storage matrix, it restores the corresponding audit record to a readable log as the query result; when the calculated aggregate hash value is different from the aggregate hash value stored at the end of the first row of the target storage matrix, it uses a second alarm message as the query result; wherein the second alarm message indicates that the corresponding audit record has been tampered with.

[0164] In some embodiments, the processing unit 42 is configured to perform file-level hash verification when an audit log query request covers a single audit log file; during the verification process, it calculates the overall hash value based on the vector storage matrix of the single audit log file; compares the calculated overall hash value with the overall hash value corresponding to the previous file of the single audit log file; when the calculated overall hash value is the same as the overall hash value corresponding to the previous file, it restores the corresponding audit log file to a readable log as the query result; when the calculated overall hash value is different from the overall hash value corresponding to the previous file, it uses a third alarm message as the query result; wherein the third alarm message indicates that the corresponding single audit log file has been tampered with.

[0165] In some embodiments, the processing unit 42 is configured to perform system-level hash verification when an audit log query request indicates that the full audit logs across multiple audit files are read; during the verification process, a global aggregate hash value is calculated based on the complete matrix set of all audit files; the calculated global aggregate hash value is compared with a preset global verification value; when the calculated global aggregate hash value is the same as the preset global verification value, all audit files are restored to readable logs as the query result; when the calculated global aggregate hash value is different from the preset global verification value, a fourth alarm message is used as the query result; wherein, the fourth alarm message indicates that the full audit logs have been tampered with.

[0166] This application also provides an electronic device, such as... Figure 5 As shown, electronic device 50 includes: Memory 51 is used to store computer-executable instructions or computer programs; When processor 52 executes computer-executable instructions or computer programs stored in memory 51, it performs the following steps: Obtain the database audit logs; the audit logs include logs containing various types of operation information generated during the database audit process. The audit logs are preprocessed to obtain preprocessed audit logs; the preprocessed audit logs are standardized audit logs. The preprocessed audit logs are segmented to obtain a set of lexical units; Perform word embedding mapping on each word in the word set to obtain each vector corresponding to each word; Arrange each vector in a preset order and add hash information to each sorted vector to obtain the target storage matrix of the audit log.

[0167] It should be noted that the explanations of the steps in this embodiment that are the same as those in the above embodiments can be found in the descriptions in the above embodiments, and will not be repeated here.

[0168] In other embodiments, the apparatus provided in this application can be implemented in hardware. For example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the log management method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0169] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the log management method described above in this application.

[0170] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the log management method provided in this application. For example, ... Figure 1 The log management method is shown.

[0171] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0172] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0173] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0174] As an example, computer-executable instructions can be deployed to execute on a single device, or on multiple devices located in one location, or on multiple devices distributed across multiple locations and interconnected via a communication network.

[0175] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A log management method, characterized in that, The method includes: Obtain the database audit logs; wherein the audit logs include logs containing multiple types of operation information generated during the database audit process; The audit logs are preprocessed to obtain preprocessed audit logs; wherein the preprocessed audit logs are standardized audit logs. The preprocessed audit logs are segmented to obtain a set of lexical units; Perform word embedding mapping on each word in the word set to obtain each vector corresponding to each word; Each vector is arranged in a preset order, and hash information is added to each sorted vector to obtain the target storage matrix of the audit log.

2. The method according to claim 1, characterized in that, The preprocessing of the audit logs to obtain preprocessed audit logs includes: The audit log is cleaned by calling a key information extraction function to obtain a cleaned audit log; wherein the cleaned audit log includes audit records corresponding to one or more of the operation time, username and operation event; The cleaned audit logs are standardized by calling a transformation function to obtain the preprocessed audit logs.

3. The method according to claim 1, characterized in that, The step of arranging each vector in a preset order and adding hash information to each sorted vector to obtain the target storage matrix of the audit log includes: Each vector is arranged in a preset order to obtain an arranged vector sequence; Calculate the hash value of each vector in the vector sequence after permutation, and obtain a new set of elements based on each vector after permutation and its hash value; Construct the initial storage matrix; Based on the new set of elements, the value of the corresponding position of the reference row in the initial storage matrix is ​​determined, and based on the hash values ​​of all the vectors after arrangement and the association information between the multiple audit logs, the value of the corresponding position of the target row in the initial storage matrix is ​​determined to obtain the target storage matrix; wherein, the reference row includes rows in the initial storage matrix that are different from the target row.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain an audit log query request; The query is performed based on the verification level corresponding to the audit log query request to obtain the query results.

5. The method according to claim 4, characterized in that, The query based on the verification level corresponding to the audit log query request, to obtain the query results, includes: When the audit log query request covers multiple audit records, a row-level hash check is performed; During the verification process, hash values ​​are calculated in parallel based on the target number of vector fields for each audit record hit by the audit log query request. The calculated hash value is compared with the pre-cached hash value; When the comparison result indicates that the calculated hash value is the same as the pre-cached hash value, the corresponding audit record is restored as a readable log as the query result; When the comparison result indicates that the calculated hash value is different from the pre-cached hash value, the first alarm information is used as the query result; wherein, the first alarm information indicates that the corresponding audit record has been tampered with.

6. The method according to claim 4, characterized in that, The query is performed based on the verification level corresponding to the audit log query request to obtain the query results, including: When the audit log query request covers some or all of the audit records in a single audit file, column-level hash verification is performed. During the verification process, the corresponding aggregate hash value is calculated based on the sequence of row hash values ​​stored at the end of all lines in a single audit file. The calculated aggregate hash value is compared with the aggregate hash value stored at the end of the first row of the target storage matrix; When the calculated aggregate hash value is the same as the aggregate hash value stored at the end of the first row of the target storage matrix, the corresponding audit record is restored as a readable log as the query result; When the calculated aggregate hash value is different from the aggregate hash value stored at the end of the first row of the target storage matrix, the second alarm information is used as the query result; wherein, the second alarm information indicates that the corresponding audit record has been tampered with.

7. The method according to claim 4, characterized in that, The query is performed based on the verification level corresponding to the audit log query request to obtain the query results, including: When the audit log query request covers a single audit log file, a file-level hash check is performed; During the verification process, the overall hash value is calculated based on the vector storage matrix of a single audit log file; The calculated overall hash value is compared with the overall hash value corresponding to the previous file of the individual audit log file; When the calculated overall hash value is the same as the overall hash value corresponding to the previous file, the corresponding audit log file is restored to a readable log as the query result. When the calculated overall hash value is different from the overall hash value corresponding to the previous file, the third alarm information is used as the query result; wherein, the third alarm information indicates that the corresponding single audit log file has been tampered with.

8. An electronic device, characterized in that, The electronic device includes: a memory for storing computer-executable instructions or computer programs; A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.