Database data tampering self-detection method
By designing the AutoCheck process in the database and detecting the checksums in TCM and TPM files, fine-grained data tamper detection is achieved, solving the problems of coarse granularity detection and inability to achieve automated detection in the prior art, and improving database data security and detection efficiency.
Patent Information
- Application Number
- CN202510097535.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
Existing database detection methods cannot achieve fine-grained data tampering detection, and manual detection is required in the state of downtime of the database, so automated detection cannot be achieved.
By designing the database AutoCheck process, checksums in TCM files and TPM files are detected, and automated detection and fine-grained database data tamper detection are realized. The specific steps include loading TCM and TPM files into memory, generating hash buckets, comparing timestamps, verifying the checksum of data elements, and notifying the management layer through the database alarm mechanism.
It realizes automated database detection, can detect data tampering in a fine-grained manner, avoids the defects of database downtime and manual detection, and improves data security and detection efficiency.
Smart Images

Figure CN120029824A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of databases, and in particular to a method for self-detecting data tampering in a database. Background Art
[0002] This technology aims to protect, detect and recover the most important resource of the database - data files. It locates the change of a certain data unit (Tuple) in the data page through the method of data page change mapping (Change Tuple Map), and scans the database file in real time through additional processes to achieve data self-detection. When it is detected that the data has been tampered with or the data page structure has been damaged, the data page data can be self-recovered by means of self-repair, or the damaged page structure can be self-repaired.
[0003] Currently, there are several methods for data anomaly detection in existing databases on the market:
[0004] 1. Checksum mechanism for database data pages.
[0005] 2. The database's own external detection tools.
[0006] 3. Detection tools provided by the operating system.
[0007] In principle, the above database data detection methods all generate a checksum for the entire page and store it in a file, which we call checksum 1. When the database is started, the checksum is regenerated for the entire data page, which we call checksum 2. When the database loads data, the two are compared. If they are different, it means that the data page has been tampered with.
[0008] This verification method can detect when data pages have been changed, but it has the following problems:
[0009] 1. The granularity of detection is relatively coarse. The current verification tools can only detect abnormalities at the page level or data file level and cannot detect abnormalities in a certain data element.
[0010] 2. Verification can only be performed when the database is shut down. If it is a production system, database shutdown is unacceptable.
[0011] 3. Only manual detection can be performed, and automatic detection of data anomalies cannot be performed.
[0012] In the database, high-authority users can directly modify the physical data files of the database. In this way, the database data can be tampered with, causing serious threats to the data security of the database. The current detection methods can only be detected when the database is shut down or manually intervened. Such detection methods are time-delayed and require downtime operations, which is difficult to accept in a production environment. Pages cannot achieve fine-grained, automated detection. Summary of the invention
[0013] In order to solve the above technical problems, the present invention provides a method for database self-detection of data tampering, which realizes database automation and fine-grained discovery of data tampering.
[0014] The technical solution of the present invention is:
[0015] A method for self-detecting data tampering in a database detects checksums in a TCM file and a TPM file through a database AutoCheck process, thereby realizing automatic detection and fine-grained database data tampering detection.
[0016] The specific steps are as follows:
[0017] (1) When the database main process starts, an AutoCheck subprocess is created, and AutoCheck loads the TCM file and TPM file into memory;
[0018] (2) In memory, AutoCheck generates a hash bucket based on the table physical file number + page number, and stores the bits of the TCM file as a two-dimensional array in memory, so that the bit corresponding to each data element can be found;
[0019] (3) The checksum values in the TPM file are stored in a two-dimensional array, and the checksum corresponding to the bit can be found;
[0020] (4) When performing verification, take out the timestamp in the TCM header record and compare it with the timestamp of the physical file corresponding to the current table. If they are the same, it means that the file has not been modified and no verification is required. If they are different, it means that the file has been modified and it is necessary to compare the checksum of all data elements with bits set to 1 and then compare it with the checksum of the data element in the physical file. If they are the same, it means that the data element has not been tampered with. If they are different, it means that the data element has been tampered with.
[0021] (5) After discovering the tampered data element, the location of the data element and the data currently in the data element are output through the database alarm mechanism to notify the management database that the data has been tampered with.
[0022] Fine-grained detection breaks down self-detection into table levels and performs self-detection on tables in the database that are above a set importance standard.
[0023] Furthermore,
[0024] Design TCM file, that is, data source change mapping file, which records the changes of data elements. The specific contents are as follows:
[0025] (1) The TCM unit consists of a TCM header and a TCM body, namely, a TCM UNIT.
[0026] (2) The TCM header records four key fields: the data page number corresponding to the current TCM unit, the location of the next TCM unit, the change time of the current data page, and the number of data units recorded by the current TCM unit;
[0027] (3) The TCM body consists of several bytes, each of which is called a TCM CELL. The 8 bits in each CELL represent 8 data elements, that is, each bit represents one data element. In the TCM body, 1 represents that the data source exists, and 0 represents that the data source has not been written or has been cleared.
[0028] in,
[0029] Current data page number: records the block number of the data page associated with the current TCM unit, indicating that this TCM unit records the data element changes of this data page number;
[0030] Next TCM unit location: records the location of the next TCM unit in order to find the offset address of this unit in the data file;
[0031] Change time of the current data page: records the change time of the entire data page, used to record the time when the database writes the data buffer;
[0032] Number of data units in the current TCM unit: records the number of data elements in the TCM unit. The size of a data element in TCM is 1 bit. This record is used to find the corresponding data element later.
[0033] The steps to write data elements into a TCM file are as follows:
[0034] (1) The user inserts data through the database operating software. At this time, the database processes and writes the data inserted by the user into the data file, and each piece of data is composed of a data element;
[0035] (2) When a data element is written, it indicates that the current data page is filled. At this time, the TCM file needs to be written. The insertion of the data element will also be recorded in the first bit of the TCM CELL. This bit changes from 0 to 1, indicating that the first data element is occupied. As user data is inserted, more bits are updated to 1. When these bits are updated, the TCM header information is updated.
[0036] (3) When the first page of the database is full, the database will continue to write the next page. At this time, TCMUNIT will also open a new UNIT, the NextPageNo of the original UNIT, and update the size of TCM UNIT;
[0037] (4) When a data element is updated or deleted, the data element is set to an invisible state in the data page, that is, an invalid state. The data element in this state will be cleared during the cleanup process. In this case, TCM also needs to be updated. When the data element is invalid, find the corresponding bit in TCM and set it to 0; at the same time, the time in the TCM HEADER is reset and updated to the time of the current operation; the number also needs to be reduced by 1.
[0038] Furthermore,
[0039] Design the TPM file to record the current checksum of each data element, using 2 bytes as a unit. Each unit corresponds to a data element one by one. When inserting a data element, calculate the checksum of the current data element and write it to the corresponding unit of the TPM file.
[0040] Each unit in the TPM corresponds to each bit in the TCM file, that is, each data element corresponds to each bit in the TCM and each checksum unit in the TPM. The algorithm for calculating the checksum is based on the FNV-1 algorithm.
[0041] The steps to write data elements into a TPM file are as follows:
[0042] (1) The user inserts data through the database operation software. At this time, a data element will be inserted into the data page and written into the TCM file. At the same time, the data of this data element will be put into the FNV-1a process for processing;
[0043] (2) FNV-1a performs hash processing on the incoming data metadata and generates a fixed-length checksum for data of any length;
[0044] (3) Write the generated checksums to the TPM file in order, with each checksum occupying 2 bytes. At this point, the bit order of the TCM corresponds to the checksum of the TPM.
[0045] (4) When a piece of data is deleted, the data element will be invalidated and the location in the TPM corresponding to the data element will be cleared.
[0046] The beneficial effects of the present invention are
[0047] 1. It can automatically detect the database and find out if the data in the database has been tampered with.
[0048] 2. It can detect the tampering of a data element with finer granularity. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic diagram of the composition of TCM and the corresponding relationship between data files and TCM;
[0050] Figure 2 It is a schematic diagram of the update and expansion of TCM;
[0051] Figure 3 It is the TCM failure process diagram;
[0052] Figure 4 It is the design of TPM and its relationship with data pages and TCM files;
[0053] Figure 5 It is a schematic diagram of the verification process. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0055] The present invention provides a method for self-detecting data tampering in a database, which is a measure applied to the security protection of database data and a method for rescuing data and recovering database data after the database data file is tampered or damaged.
[0056] Prerequisite description of the technical solution:
[0057] According to the composition principle of data pages, data elements (tuples) are stored in data pages in a form similar to an array. When data is inserted (i.e., INSERT INTO table VAUES), the data elements are appended to the data pages, and then processed by the database's buffer layer and storage layer, and finally the data pages are stored in physical data files.
[0058] Design TCM file (Tuple Chage Map), which is the data source change mapping file. This file is a change mapping file of the data file, which mainly records the changes of the data element. The main design is as follows:
[0059] (1) A TCM unit consists of a TCM header and a TCM body as a unit, namely, a TCM UNIT.
[0060] (2) The TCM header records four key fields: the data page number corresponding to the current TCM unit, the location of the next TCM unit, the change time of the current data page, and the number of data units recorded by the current TCM unit. ● Current data page number: records the block number of the data page associated with the current TCM unit, indicating that this TCM unit records the data element changes of this data page number.
[0061] ●Next TCM unit position: Record the position of the next TCM unit to facilitate finding the offset address of this unit in the data file.
[0062] ●Change time of the current data page: records the change time of the entire data page, used to record the time when the database writes the data BUFFER.
[0063] ●Number of data units in the current TCM unit: records the number of data elements in the TCM unit. The size of a data element in TCM is 1 bit, but the number of each data element is uncertain. This record is used to quickly find the corresponding data element in the future.
[0064] (3) The TCM body consists of several bytes, each byte is called a TCM CELL. The 8 bits in each CELL represent 8 data elements, that is, each bit represents one data element. We know that 1 bit can only be used to represent two states, that is, 0 or 1. In the TCM body, 1 represents that the data source exists, and 0 represents that the data source has not been written or has been cleared. For the composition of TCM and the correspondence between data files and TCM, please refer to Figure 1 .
[0065] The steps to write data elements into a TCM file are as follows:
[0066] (1) The user inserts data through the database operating software. At this time, the database will go through a series of processing and write the data inserted by the user into the data file, and each piece of data will be organized into a data element.
[0067] (2) When a data element is written, it means that the current data page is filled. At this time, the TCM file needs to be written. For example, if the user inserts a piece of data, a data element will be written to the data file. At the same time, the insertion of this data element will also be recorded in the first bit of the TCM CELL. This bit changes from 0 to 1, indicating that the first data element is occupied. As user data is inserted, more bits are updated to 1. Assuming that the mth data element needs to be inserted into the TCM file, we need to update to the pth bit of the nth TCM CELL, then the relationship between m, n, and p is p = m mod 8n. When these bits are updated, the TCM header information will be updated.
[0068] (3) When the first page of the database is full, the database will continue to write the next page. At this time, TCMUNIT will also open a new UNIT, the NextPageNo of the original UNIT, and update the size of TCM UNIT. For updates and expansions of TCM, please refer to Figure 2 .
[0069] (4) When a data element is updated or deleted, the data element is set to an invisible state in the data page, that is, an invalid state. The data element in this state will be cleared during the cleanup process. Therefore, in this case, TCM also needs to be updated. When the data element is invalid, find the corresponding bit in TCM and set it to 0. At the same time, the time in TCM HEADER is reset and updated to the time of the current operation. The number also needs to be reduced by 1. For TCM invalid bit setting, please refer to Figure 3 .
[0070] The TPM (Tuple Proof Map) file is designed. The file mainly records the current checksum of each data element, with 2 bytes as a unit. Each unit corresponds to a data element one-to-one. When inserting a data element, the checksum of the current data element is calculated and written to the corresponding unit of the TPM file. In addition, in order to speed up the retrieval performance, each unit in the TPM corresponds to each bit in the TCM file one-to-one, that is, each data element corresponds to each bit in the TCM and each checksum unit in the TPM. The algorithm for calculating the checksum is based on the FNV-1 algorithm. The FNV-1 algorithm is also an internationally used simple HASH algorithm that can generate a 2-byte checksum for data of any length.
[0071] The steps to write data elements into a TPM file are as follows:
[0072] (1) The user inserts data through the database operation software. At this time, a data element will be inserted into the data page. As known from the above steps, it will be written into the TCM file. At the same time, the data of this data element will be put into the FNV-1a process for processing.
[0073] (2) FNV-1a performs hash processing on the incoming data metadata. Due to its characteristics, it can generate a fixed-length checksum from data of any length. If the checksum length is too low, it will affect its own error detection rate, while if the length is too high, it will affect the efficiency of generating the checksum. Therefore, we fix the checksum length to 2 bytes to meet both requirements.
[0074] (3) Write the generated checksums to the TPM file in order, with each checksum occupying 2 bytes. At this point, the bit order of the TCM corresponds to the checksum of the TPM one by one.
[0075] (4) When a piece of data is deleted, the data element will be invalidated and the location in the TPM corresponding to the data element will be cleared. For example, if the original checksum is 0x1234ABCD, it will be set to 0x0, indicating that the checksum is invalid.
[0076] For the design of TPM and its relationship with data pages and TCM files, please refer to Figure 4
[0077] In order to achieve self-detection of tampered data, it is necessary to add an AutoCheck process in the database, which mainly performs cyclic scanning of TCM and TPM and rechecks the changed data elements. AutoCheck is a subprocess of the database main process, and its life cycle is the same as that of the main process. The specific steps of AutoCheck are as follows:
[0078] (1) When the database main process starts, an AutoCheck subprocess is created. AutoCheck loads the TCM file and TPM file into memory.
[0079] (2) In memory, AutoCheck generates a hash bucket based on the table physical file number + page number, and stores the bits of the TCM file as a two-dimensional array in memory, so that the bit corresponding to each data element can be quickly found.
[0080] (3) The checksum values in the TPM file are stored in a two-dimensional array, so that the checksum corresponding to the bit can be quickly found.
[0081] (4) When performing verification, we will take out the timestamp in the TCM header record and compare it with the timestamp of the physical file corresponding to the current table. If they are the same, it means that the file has not been modified and no verification is required. If they are different, it means that the file has been modified. We need to compare the checksum of all data elements with bits set to 1, and then compare it with the checksum of the data element in the physical file. If they are the same, it means that the data element has not been tampered with. If they are different, it means that the data element has been tampered with.
[0082] (5) After discovering the tampered data element, the location of the data element and the data currently in the data element are output through the database alarm mechanism to notify the management database that the data has been tampered with.
[0083] Fine-grained detection: Since the database is constantly changing during operation, especially the system tables that the system depends on, the system tables change in real time during the entire database operation life cycle. If self-detection is performed on this part of the data, the cost is high. Therefore, self-detection can be subdivided into table levels. We perform self-detection on the more important tables in the database, which can not only ensure the security of important data, but also ensure that the performance of the system is not greatly affected.
[0084] The above description is only a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for self-detecting data tampering in a database, characterized in that: Through the database AutoCheck process, the checksums in the TCM file and the TPM file are detected to achieve automated detection and fine-grained database data tampering detection.
2. The method according to claim 1, characterized in that The specific steps are as follows: (1) When the database main process starts, an AutoCheck subprocess is created, and AutoCheck loads the TCM file and TPM file into memory; (2) In memory, AutoCheck generates a hash bucket based on the table physical file number + page number, and stores the bits of the TCM file as a two-dimensional array in memory, so that the bit corresponding to each data element can be found; (3) The checksum values in the TPM file are stored in a two-dimensional array, and the checksum corresponding to the bit can be found; (4) When performing verification, take out the timestamp in the TCM header record and compare it with the timestamp of the physical file corresponding to the current table. If they are the same, it means that the file has not been modified and no verification is required. If they are different, it means that the file has been modified and it is necessary to compare the checksum of all data elements with bits set to 1 and then compare it with the checksum of the data element in the physical file. If they are the same, it means that the data element has not been tampered with. If they are different, it means that the data element has been tampered with. (5) After discovering the tampered data element, the location of the data element and the data currently in the data element are output through the database alarm mechanism to notify the management database that the data has been tampered with.
3. The method according to claim 1, characterized in that Fine-grained detection breaks down self-detection into table levels and performs self-detection on tables in the database that are above a set importance standard.
4. The method according to claim 1, characterized in that: Design TCM file, that is, data source change mapping file, which records the changes of data elements. The specific contents are as follows: (1) The TCM unit consists of a TCM header and a TCM body, namely, a TCM UNIT. (2) The TCM header records four key fields: the data page number corresponding to the current TCM unit, the location of the next TCM unit, the change time of the current data page, and the number of data units recorded by the current TCM unit; (3) The TCM body consists of several bytes, each of which is called a TCM CELL. The 8 bits in each CELL represent 8 data elements, that is, each bit represents one data element. In the TCM body, 1 represents that the data source exists, and 0 represents that the data source has not been written or has been cleared.
5. The method according to claim 4, characterized in that Current data page number: records the block number of the data page associated with the current TCM unit, indicating that this TCM unit records the data element changes of this data page number; Next TCM unit location: records the location of the next TCM unit in order to find the offset address of this unit in the data file; Change time of the current data page: records the change time of the entire data page, used to record the time when the database writes the data buffer; Number of data units in the current TCM unit: records the number of data elements in the TCM unit. The size of a data element in TCM is 1 bit. This record is used to find the corresponding data element later.
6. The method according to claim 5, characterized in that The steps to write data elements into a TCM file are as follows: (1) The user inserts data through the database operating software. At this time, the database processes and writes the data inserted by the user into the data file, and each piece of data is composed of a data element; (2) When a data element is written, it indicates that the current data page is filled. At this time, the TCM file needs to be written. The insertion of the data element will also be recorded in the first bit of the TCM CELL. This bit changes from 0 to 1, indicating that the first data element is occupied. As user data is inserted, more bits are updated to 1. When these bits are updated, the TCM header information is updated. (3) When the first page of the database is full, the database will continue to write the next page. At this time, TCMUNIT will also open a new UNIT, update the NextPageNo of the original UNIT, and update the size of TCM UNIT; (4) When a data element is updated or deleted, the data element is set to an invisible state in the data page, that is, an invalid state. The data element in this state will be cleared during the cleanup process. In this case, TCM also needs to be updated. When the data element is invalid, find the corresponding bit in TCM and set it to 0; at the same time, the time in the TCM HEADER is reset and updated to the time of the current operation; the number also needs to be reduced by 1.
7. The method according to claim 1, characterized in that Design a TPM file to record the current checksum of each data element, with 2 bytes as the unit. Each unit corresponds to a data element one by one. When inserting a data element, calculate the checksum of the current data element and write it to the corresponding unit of the TPM file. Each unit in the TPM corresponds one-to-one to each bit in the TCM file, that is, each data element corresponds to each bit in the TCM and each checksum unit in the TPM.
8. The method according to claim 7, characterized in that The algorithm for calculating the checksum is based on the FNV-1 algorithm.
9. The method according to claim 7, characterized in that: The steps to write data elements into a TPM file are as follows: (1) The user inserts data through the database operation software. At this time, a data element will be inserted into the data page and written into the TCM file. At the same time, the data of this data element will be put into the FNV-1a process for processing; (2) FNV-1a performs hash processing on the incoming data metadata and generates a fixed-length checksum for data of any length; (3) Write the generated checksums to the TPM file in order, with each checksum occupying 2 bytes. At this point, the bit order of the TCM corresponds to the checksum of the TPM. (4) When a piece of data is deleted, the data element will be invalidated and the location in the TPM corresponding to the data element will be cleared.