Database data management method and system and computer equipment

By setting up multiple independent databases and encryption query key technologies in the database, the bottlenecks in concurrency and performance of existing database encryption backup methods are solved, and more efficient and secure database data management is achieved.

CN120144668AActive Publication Date: 2025-06-13内蒙古自治区教育考试院 +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510215240.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing database encryption and backup methods have performance bottlenecks in concurrency, high-reference backup performance and encryption fine-grainedness, which affect system performance and security.

Method used

By setting up real-time databases in the database, analytical databases and backing up databases, and using encrypted query keys and public key encryption technology, the security of query content during transmission is ensured, while preventing potential attacks through audit log monitoring.

Benefits of technology

It improves the security and privacy protection level of the database, reduces the impact of encryption operations on system performance, and ensures data integrity and system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144668A_ABST
    Figure CN120144668A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of databases, and discloses a database data management method and system and computer equipment, except a real-time database for receiving data in real time, a query database and a backup database are separately arranged, and data synchronization work is carried out through a backup script. And then, when an external visitor queries the database content, encrypting the queried content through the internal encryption agent and the own key of the visitor, thereby ensuring the security and privacy of the queried data. And finally, monitoring the encrypted traffic in the audit log of the database to prevent potential attackers from threatening the database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of database technology, and in particular to a database data management method, system and computer equipment. Background Art

[0002] Database encryption, backup and query technologies have a unified demand background in ensuring data security, system reliability and efficiency. These technologies need to meet data security requirements to prevent data leakage and tampering; ensure system reliability to prevent data loss and business interruption; optimize performance and reduce consumption of system resources; meet compliance requirements and comply with laws and regulations; provide ease of use and maintainability to simplify operations and management; support cross-platform and compatibility to ensure applicability in different environments; and be real-time and dynamic to adapt to changing business needs. By comprehensively considering these requirements, a more robust, secure and efficient database system can be built.

[0003] There are many database encryption backup methods, and some research results have been achieved. However, there are still the following problems in database concurrency, query and backup performance, and encryption granularity.

[0004] 1. Encryption operations may have a certain impact on some database operations, such as query efficiency, data backup and recovery, etc. Especially when processing large amounts of data or high concurrent access, encryption operations may become a performance bottleneck, with low security and privacy.

[0005] 2. When performing backup and recovery operations, more system resources may be occupied, such as CPU, memory, disk I / O, etc., affecting the normal operation of the computer system, especially when backing up large amounts of data or performing complex recovery operations, which may cause system performance to decline. Summary of the invention

[0006] The present invention provides a database data management method, aiming to solve the technical problems raised in the above background technology.

[0007] The present invention provides a database data management method, wherein the database includes a real-time database, an analysis database and a backup database;

[0008] The real-time database dynamically acquires data information, and writes the data information into the analysis database and the backup database through the backup script, wherein the backup database is independent of the analysis database and the real-time database;

[0009] The analysis database receives external query request information sent by the external access device, and performs identity authentication according to the external query request information;

[0010] After authentication, the analysis database generates an encrypted query key pair and query content according to the external query request information, and encrypts the query content with the public key to obtain the first encrypted content;

[0011] The analysis database obtains the public key of the external access device, and re-encrypts the first encrypted content and the encrypted query key according to the public key of the external access device to obtain the second encrypted content, and outputs the second encrypted content to the external access device;

[0012] The external access device decrypts the second encrypted content according to its own private key, and decrypts the first encrypted content according to the decryption key in the encrypted query key to obtain the query content.

[0013] Preferably, the step of the analysis database generating the encrypted query key according to the external query request information includes:

[0014] The analysis database obtains the ID information, IP address, authentication timestamp of the external access device, and the public key of the external access device according to the external query request information;

[0015] The analysis database obtains the first variable value and the second variable value, and respectively packages and forms a string group for the ID information, authentication timestamp, first variable value, and second variable value based on the hash algorithm, where the string group includes a first string and a second string, the first string includes the ID information, authentication timestamp, first variable value, and the second string includes the user IP address, authentication timestamp, second variable value;

[0016] Repeat the link for the first string and the second string respectively until the length exceeds 1024 bits to obtain the first long string and the second long string;

[0017] The analysis database calculates the first long string and the second long string respectively based on the SM3 hash function, converts the calculation results to decimal, and takes the first 1024 as the basic characters to obtain the first basic character and the second basic character;

[0018] Based on the Miller-Rabin primality test, determine whether the first basic character and the second basic character are prime numbers respectively;

[0019] If the first basic character and / or the second basic character is not a prime number, add one to the first variable value and / or the second variable value, and use the first variable value and / or the second variable value after the addition as the first variable value and / or the second variable value and return to the step of respectively packaging and forming a string for the ID information, authentication timestamp, first variable value, and second variable value based on the hash algorithm until both the first basic character and the second basic character are prime numbers;

[0020] Calculate the prime product and the value of Euler's totient function based on the first and second basic characters;

[0021] Find the public key exponent based on the prime product and the value of Euler's totient function;

[0022] Package the public key exponent and the prime product to obtain the encryption key;

[0023] Calculate the private key exponent according to the value of Euler's totient function and the public key exponent;

[0024] Package the private key exponent and the prime product to obtain the decryption key.

[0025] Preferably, it further includes:

[0026] The analysis database obtains historical abnormal audit data and uses it as pseudo-abnormal data;

[0027] Obtain normal audit data, and mix the normal audit data and the pseudo-abnormal data to obtain training samples;

[0028] Use the training samples as the input end to send them into the discriminator model for training and output the training results;

[0029] Judge whether the loss value of the training result exceeds the preset value;

[0030] If it exceeds the preset value, return to the step of obtaining the historical abnormal audit data and using it as pseudo-abnormal data until the loss value does not exceed the preset value.

[0031] Preferably, the step that the analysis database obtains historical abnormal audit data and uses it as pseudo-abnormal data includes:

[0032] The analysis database obtains the target abnormal audit log data in the historical abnormal audit data;

[0033] Add Gaussian noise to the target abnormal audit log data multiple times;

[0034] Randomly delete the Gaussian noise a preset number of times to restore the abnormal audit log data;

[0035] Use the abnormal audit log data as pseudo-abnormal data.

[0036] Preferably, when performing feature analysis on the training samples, the information gain algorithm, the information gain ratio algorithm, the chi-square test algorithm or the ReliefF algorithm is used.

[0037] Preferably, after the step that if it exceeds the preset value, return to the step of obtaining the historical abnormal audit data and using it as pseudo-abnormal data until the loss value does not exceed the preset value, it includes:

[0038] The analysis database obtains the discriminator model and deploys the discriminator model in the internal network;

[0039] Based on the discriminator model, audit logs are monitored in real time, where the audit logs are generated based on the second encrypted content;

[0040] Based on the audit logs, it is determined whether there is abnormal data;

[0041] If there is abnormal data, encrypted proxy logs are obtained, where the encrypted proxy logs are generated based on the first encrypted content and the second encrypted content;

[0042] The ID information, IP address, authentication timestamp of the external access device, and the public key of the external access device retained in the encrypted proxy logs are obtained to reproduce the key to recover the second encrypted content.

[0043] This application also provides a database data management system, including a real-time database, an analysis database, and a backup database;

[0044] The real-time database is used to dynamically obtain data information and write the data information into the analysis database and the backup database through a backup script, where the backup database is independent of the analysis database and the real-time database;

[0045] The analysis database is used to receive external query request information sent by an external access device and perform authentication according to the external query request information;

[0046] After the authentication is passed, the analysis database is also used to generate an encrypted query key pair and query content according to the external query request information, and encrypt the query content with the public key to obtain the first encrypted content;

[0047] The analysis database is also used to obtain the public key of the external access device, and re-encrypt the first encrypted content and the encrypted query key according to the public key of the external access device to obtain the second encrypted content, and output the second encrypted content to the external access device;

[0048] The external access device is used to decrypt the second encrypted content according to its own private key and the decryption key in the encrypted query key.

[0049] Preferably, the analysis database is also used to obtain the ID information, IP address, authentication timestamp of the external access device, and the public key of the external access device according to the external query request information;

[0050] The analysis database is also used to obtain a first variable value and a second variable value, and respectively package and form a string group based on the hash algorithm for the ID information, authentication timestamp, first variable value, and second variable value. Among them, the string group includes a first string and a second string. The first string includes the ID information, authentication timestamp, and first variable value, and the second string includes the user IP address, authentication timestamp, and second variable value;

[0051] It is also used to respectively perform repeated linking on the first string and the second string until it is extended to a length of more than 1024 bits to obtain a first long string and a second long string;

[0052] The analysis database is also used to respectively calculate the first long string and the second long string based on the SM3 hash function, convert the calculation results to decimal, and take the first 1024 as the basic characters to obtain a first basic character and a second basic character;

[0053] It is also used to respectively determine whether the first basic character and the second basic character are prime numbers based on the Miller-Rabin primality test;

[0054] If the first basic character and / or the second basic character is not a prime number, the first variable value and / or the second variable value is incremented by one, and the first variable value and / or the second variable value after the increment operation is used as the first variable value and / or the second variable value and returned to the step of respectively packaging and forming strings for the ID information, authentication timestamp, first variable value, and second variable value based on the hash algorithm until both the first basic character and the second basic character are prime numbers;

[0055] It is also used to calculate the prime product and the Euler's totient function value based on the first basic character and the second basic character;

[0056] It is also used to find the public key exponent based on the prime product and the Euler's totient function value;

[0057] It is also used to package the public key exponent and the prime product to obtain an encryption key;

[0058] It is also used to calculate the private key exponent based on the Euler's totient function value and the public key exponent;

[0059] It is also used to package the private key exponent and the prime product to obtain a decryption key.

[0060] The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0061] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0062] The beneficial effects of the present invention are as follows: In addition to the real-time database that receives data in real time, a query database and a backup database are separately set up, and data synchronization work is carried out through backup scripts. Then, when an external visitor queries the content of the query database, the content obtained by the query is encrypted through an internal encryption proxy and the visitor's own key to ensure the security and privacy of the query data. Finally, by monitoring the encrypted traffic in the audit log of the database, potential attackers are prevented from threatening the database. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.

[0064] Figure 2 It is a schematic flowchart of the key generation process according to an embodiment of the present invention.

[0065] Figure 3 It is a schematic internal structure diagram of a computer device according to an embodiment of the present invention.

[0066] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0068] As Figures 1 - 3 shown, the present application provides a database data management method, and the database includes a real-time database, an analysis database and a backup database;

[0069] S1. The real-time database dynamically obtains data information and writes the data information into the analysis database and the backup database through a backup script, wherein the backup database is independent of the analysis database and the real-time database;

[0070] S2. The analysis database receives the external query request information sent by an external access device and performs identity verification according to the external query request information;

[0071] S3. After the identity verification is passed, the analysis database generates an encrypted query key pair and query content according to the external query request information, and performs public key encryption on the query content to obtain a first encrypted content;

[0072] S4, the analysis database obtains the public key of the external access device, and re-encrypts the first encrypted content and the encryption query key according to the public key of the external access device to obtain the second encrypted content, and outputs the second encrypted content to the external access device;

[0073] S5. The external access device decrypts the second encrypted content according to its own private key, and decrypts the first encrypted content according to the decryption key in the encrypted query key to obtain the query content.

[0074] As described in the above steps S1-S5, in the design of the data architecture, there are three databases, namely, a real-time database, an analysis database and a backup database. Among them, the real-time database dynamically obtains data information and writes it into the analysis database and the backup database through the backup script; the analysis database is used for external visitors to query the information in the database through query keywords, and is mainly responsible for solving the query needs of external visitors; the backup database is placed in the internal isolation area, which is used for regular storage and disaster recovery of database content. In addition, the internal administrator will access the real-time database through the bastion host to monitor the database; the external visitor accesses the analysis database to query data information through the external access device after identity authentication; the query content is encrypted to prevent the query content from leaving traces on the audit log; before the audit log is transmitted to the outside, it is further encrypted with the public key of the external visitor and then transmitted to the outside; finally, the external visitor only needs to use his own private key and the decryption key in the query ciphertext package in turn to obtain the content required for query. Through this method, the security protection of the entire life cycle from the data out of the warehouse can be achieved, and even the internal administrator cannot know the content of the visitor's query. By implementing secure backup and proxy query measures, the security and privacy protection level of the database can be improved, ensuring the integrity of data within the database and the reliability of the system.

[0075] In one embodiment, the step S3 of the analysis database generating an encrypted query key according to the external query request information includes:

[0076] S31, the analysis database obtains the ID information, IP address, identity authentication timestamp and public key of the external access device according to the external query request information;

[0077] S32, the analysis database obtains the first variable value and the second variable value, and packages the ID information, the identity authentication timestamp, the first variable value and the second variable value respectively based on the hash algorithm and forms a string group, wherein the string group includes the first string and the second string, the first string includes the ID information, the identity authentication timestamp, and the first variable value, and the second string includes the user IP address, the identity authentication timestamp, and the second variable value;

[0078] S33. Repeat the concatenation of the first string and the second string respectively until they are extended to a length of more than 1024 bits, obtaining a first long string and a second long string;

[0079] S34. The analysis database calculates the first long string and the second long string respectively based on the SM3 hashing function, converts the calculation results to decimal, and takes the first 1024 bits as basic characters to obtain a first basic character and a second basic character;

[0080] S35. Based on the Miller-Rabin primality test, determine whether the first basic character and the second basic character are prime numbers respectively;

[0081] S36. If the first basic character and / or the second basic character is not a prime number, increment the first variable value and / or the second variable value by one, and use the incremented first variable value and / or the second variable value as the first variable value and / or the second variable value and return to the step of packing the ID information, authentication timestamp, first variable value, and second variable value respectively based on the hashing algorithm to form a string until both the first basic character and the second basic character are prime numbers;

[0082] S37. Calculate the prime product and the Euler's totient function value based on the first basic character and the second basic character;

[0083] S38. Find the public key exponent based on the prime product and the Euler's totient function value;

[0084] S39. Pack the public key exponent and the prime product to obtain the encryption key;

[0085] S40. Calculate the private key exponent based on the Euler's totient function value and the public key exponent;

[0086] S41. Pack the private key exponent and the prime product to obtain the decryption key.

[0087] As described in the above steps S31 - S41, after the external access device initiates a query request, its identity is first authenticated. After successful authentication, the encryption proxy will automatically collect the ID information, IP address, request initiation timestamp, and public key of the external access device, and assign two variable values, the first variable value CountP and the second variable value CountQ, to generate a pair of keys for encrypting the query content. At the same time, the query SQL statement of the external access device is restricted and cannot query data content outside the limit, such as data information other than that related to itself. The analysis database will encrypt the plaintext query data to be returned using the public key in the pair of keys generated by the encryption proxy to obtain the first encrypted content. After encryption, the first encrypted content and the remaining decryption key are packaged into a ciphertext packet, and then encrypted using the public key of the external visitor obtained by the encryption proxy to obtain the second encrypted content. Then, an automatic audit record is carried out on the second encrypted content. Finally, the registered second encrypted content is returned to the external access device, and the external access device can obtain the query information using its own private key. By modifying the initial prime number selection based on RSA, the entire key can be made controllable.

[0088] As Figure 2 shown, the specific encryption process is as follows: First, the ID information, timestamp, and CountP, as well as the IP address, timestamp, and CountQ, are each packed and concatenated into a string, and then extended to a length of more than 1024 bits through a repeated linking form. Then, the SM3 hashing function is used to calculate them. The hexadecimal in the calculated result is expanded into decimal and the first 1024 bits are taken as the base, and then the Miller - Rabin primality test is used to determine whether it is a prime number. If it is not a prime number, the value of CountP is incremented by one, and the linking extension, SM3 function generation, and primality test are repeated until a pair of 1024 - bit prime numbers P and Q are generated. Subsequently, calculate the prime product N and the Euler's totient function value M according to the following formula.

[0089] Among them, N = P × Q; where N represents the prime product, P represents the first variable value CountP, and Q represents the second variable value CountQ;

[0090] M = (P - 1) × (Q - 1); where M represents the Euler's totient function value;

[0091] After obtaining N and M, find the public key exponent E, requiring that E is relatively prime to M and is a positive integer less than M. After obtaining E, pack E and N to obtain the encryption key (E, N). Subsequently, calculate the private key exponent D.

[0092] Among them, the calculation formula is D = (k×M + 1) / E; where D represents the private key exponent, k is a positive integer, representing the integer part of the quotient when D·E is divided by M. Packing D and N can obtain the decryption key (D, N).

[0093] In one embodiment, it further includes:

[0094] The analysis database obtains historical abnormal audit data and uses it as pseudo-abnormal data;

[0095] Obtain normal audit data, and mix the normal audit data and the pseudo-abnormal data to obtain training samples;

[0096] Send the training samples as the input end into the discriminator model for training and output the training results;

[0097] Judge whether the loss value of the training result exceeds the preset value;

[0098] If it exceeds the preset value, return to the step of obtaining the historical abnormal audit data and using it as pseudo-abnormal data until the loss value does not exceed the preset value.

[0099] Among them, the step of sending the training samples as the input end into the discriminator model for training and outputting the training results includes:

[0100] Input the training samples into the discriminator model and perform feature analysis on the training samples to obtain important features;

[0101] Output the training results based on the important features.

[0102] Among them, the step of the analysis database obtaining the historical abnormal audit data and using it as pseudo-abnormal data includes:

[0103] The analysis database obtains the target abnormal audit log data in the historical abnormal audit data;

[0104] Add Gaussian noise to the target abnormal audit log data multiple times;

[0105] Randomly delete the Gaussian noise a preset number of times to restore the abnormal audit log data;

[0106] Use the abnormal audit log data as pseudo-abnormal data.

[0107] As described above, while outputting the second encrypted content to an external access device, monitor the audit log of the second encrypted content to identify potential abnormal and malicious query operations. This application mainly uses an accurate abnormal traffic detection method for pseudo-abnormalities, mainly through a diffusion model - generative adversarial network (DF-GAN) model to carry out the detection work. The DF-GAN model mainly consists of a diffusion (DF) model responsible for generating pseudo-abnormal traffic and a generative adversarial network (GAN) model responsible for monitoring abnormal traffic. Among them, the diffusion model realizes the generation of high-quality pseudo-abnormal traffic through a data augmentation mode of adding noise and denoising, and the generative adversarial network is mainly responsible for identifying the abnormal traffic mixed in the normal traffic to carry out the detection work. After the model training is completed, it only needs to regularly capture the ciphertext audit data from the ciphertext audit log for identification. Specifically, when carrying out the audit log monitoring, retrieve the existing abnormal audit data records to generate a batch of abnormal audit data through diffusion to alleviate the decline in the model identification performance caused by insufficient abnormal audit quantity, and mark them as pseudo-abnormal audit data. Secondly, mix the pseudo-abnormal audit data and the normal audit data and hand them over to the discriminator for identification. The model will extract the spatio-temporal features contained in the traffic information and adopt feature selection algorithms such as information gain (IG), information gain ratio (IGR), chi-square (χ2), and ReliefF to select important features for analysis. When the loss value obtained from the identification result is large, it will return to the generator stage, regenerate the pseudo-abnormal data for further identification until the finally output result has a small loss value. Finally, deploy the trained model to the internal network of the database to carry out security monitoring work according to the audit requirements. Specifically, use the existing abnormal audit data records to generate a batch of pseudo-abnormal audit data through a data diffusion generation algorithm (such as a diffusion model or GAN, etc.) and mark them as "pseudo-abnormal". Then, mix the generated pseudo-abnormal data with the normal audit data and use it as input to the discriminator model for training. The main task of the discriminator is to distinguish normal and abnormal data according to the spatio-temporal features in the data (such as timestamp, source IP, request frequency, etc.). In terms of feature selection, adopt feature selection algorithms such as information gain (IG), information gain ratio (IGR), chi-square (χ2), and ReliefF to improve the accuracy of the model. After preliminary training, the discriminator classifies the data. If the loss value of the identification result is large, it returns to the generator stage, regenerates the pseudo-abnormal data, and conducts re-training. This process is similar to the feedback mechanism in encryption and decryption, and the generator and the discriminator are continuously iteratively optimized until the loss value reaches the minimum and the identification ability of the model reaches the best state.

[0108] Specific training process:

[0109] 1. Data preparation: Prepare a number of abnormal and normal audit data

[0110] 2. Model Initialization: Initialize the parameters of the diffusion model generator and the generative adversarial network discriminator

[0111] 3. Model Training:

[0112] 3.1 Generator Training

[0113] First, extract the target abnormal audit log data, then add Gaussian noise step by step until it completely becomes noise, and then remove the noise step by step to recover the abnormal audit log data. Since there is a probability in each step of selection, in this way, a batch of similar but different pseudo-abnormal log data can be generated, which can solve the problem of insufficient training data in the current abnormal identification model

[0114] 3.2 Discriminator Training

[0115] The discriminator will accept a training data set mixed with pseudo-abnormal data and a large amount of normal data. Its main task is to judge whether the current data is abnormal data by calculating the information gain, information gain ratio, chi-square value, etc. of each data, and perform simulated identification on the divided data, and correct the training parameters through the loss value calculated by the loss function

[0116] 4. Evaluation and Judgment:

[0117] Judge the performance of the model in various aspects, and evaluate the quality of the pseudo-abnormal samples generated by it and the abnormal detection ability of the model

[0118] The generator constructed by the diffusion model realizes the generation of high-quality pseudo-abnormal sample data, making up for the problem of insufficient abnormal data samples in the model training process. The pseudo-abnormal samples generated by it are highly consistent with the real abnormal samples semantically, increasing the diversity and richness of the training data, and improving the generalization ability and detection accuracy of the model. And DF-GAN is different from the generative adversarial network that usually uses multiple layers stacked in the past. It only uses one generator and one discriminator, simplifies the model structure, and improves the training efficiency of the model. This makes the quality of the generated pseudo-abnormal data high: the text generated by DF has better consistency compared with the original text; and the training speed is fast: only one generator and one discriminator are adopted, simplifying the model structure and accelerating the model training; the model has strong stability: DF-GAN is more stable in the training process, reducing the unstable factors in the training process

[0119] In one embodiment, when performing feature analysis on the training samples, the information gain algorithm, information gain ratio algorithm, chi-square test algorithm or ReliefF algorithm is adopted. These methods help us screen out the features that are most helpful for abnormal detection by quantifying the relationship between the features and the target variable

[0120] Information Gain (IG) is a metric that measures the degree to which a feature reduces the uncertainty of a dataset. Its calculation formula is:

[0121] IG(D,X) = H(D) - H(D|X);

[0122] Among them, H(D) represents the original entropy of the dataset D, and H(D∣X) represents the weighted average entropy after partitioning by the feature X. The higher the information gain, the more effective the feature X is in reducing the uncertainty of the dataset, and thus it is more valuable in feature selection.

[0123] Furthermore, the Information Gain Ratio (IGR) normalizes the information gain with the intrinsic value of the feature to avoid bias towards features with more values. Its calculation formula is:

[0124]

[0125] Among them, the calculation formula for the intrinsic value IV(X) is:

[0126]

[0127] Here, D i represents the data subset corresponding to the i-th value of the feature X.

[0128] In addition to information gain and information gain ratio, the Chi-Square Test (χ2) is also a commonly used feature selection method, especially suitable for categorical data. It evaluates the correlation between a feature and the target variable by measuring the difference between the actual observed frequency and the expected frequency. The calculation formula is:

[0129]

[0130] Among them, f o is the actual observed frequency, and f c is the expected frequency. The larger χ 2 , the stronger the correlation between the feature and the target variable.

[0131] In addition to the above methods, the ReliefF algorithm evaluates the importance of features by calculating feature weights, especially suitable for dealing with noisy data and multi-class problems. Its weight update formula is:

[0132]

[0133] Among them, W[A] represents the weight of feature A, diff(A,n i ) represents the difference of feature A in sample n i , and m is the total number of samples.

[0134] After completing the feature analysis and selecting the important features, the next step is to obtain the recognition result by the output probability of the model. For example, the output probability of a logistic regression model can be calculated by the following formula:

[0135]

[0136] If P(anomaly|X) > 0.5, then this data point is determined to be an anomaly. The performance of the model can be evaluated by metrics such as recall, F1-score, and ROC-AUC value. Recall measures the proportion of actual abnormal data that is correctly identified, and its formula is:

[0137]

[0138] Among them, TP is the true positive (actual anomaly and identified as an anomaly), and FN is the false negative (actual anomaly but not identified as an anomaly). The F1-score is the harmonic mean of precision and recall, and its formula is:

[0139]

[0140] The higher the F1-score, the better the comprehensive performance of the model. In addition, the ROC-AUC value is calculated through the ROC curve (the relationship between the true positive rate and the false positive rate), and the closer the AUC is to 1, the better the classification performance of the model.

[0141] When the training is completed and the model performance meets the expectations, the trained discriminator model is deployed to the internal network of the database to start real-time monitoring of the audit logs.

[0142] In one embodiment, after the step of, if it exceeds the preset value, returning to the step of obtaining the historical abnormal audit data and using it as pseudo-abnormal data until the loss value does not exceed the preset value, includes:

[0143] The analysis database obtains the discriminator model and deploys the discriminator model in the internal network;

[0144] Based on the discriminator model, the audit logs are monitored in real time, where the audit logs are generated based on the second encrypted content;

[0145] Based on the audit logs, it is judged whether there is abnormal data;

[0146] If there is abnormal data, obtain the encrypted proxy logs, where the encrypted proxy logs are generated based on the first encrypted content and the second encrypted content;

[0147] Obtain the ID information, IP address, authentication timestamp of the external access device, and the public key of the external access device retained in the encrypted proxy logs to reproduce the key to recover the second encrypted content.

[0148] As described above, the model alleviates the problem of insufficient abnormal audit data by generating pseudo-abnormal data, continuously analyzes and identifies abnormal behaviors in audit data, ensures system security, and responds to potential threats in a timely manner. In the face of the identified abnormal data, after receiving the abnormal data packet, the internal administrator can reproduce the key by encrypting the user ID information, IP address, timestamp and other records retained in the proxy log, so as to open the encrypted data packet and analyze the abnormal data behavior logic. Reproduction refers to extracting the ID information, IP address, timestamp, CountP and CountQ of the records from the proxy encrypted audit records, and re-obtaining the decryption key according to the calculation method of key generation to open the ciphertext packet. For example, use the behavior patterns in the proxy log (such as request frequency, access path, etc.) to confirm the nature of the abnormal behavior. Specifically: 1. Request frequency analysis: By counting the number of accesses of each IP address within a specific time, a reasonable threshold is set, such as 100 times per minute. If the access frequency of a certain IP address exceeds this threshold, it is marked as an abnormal behavior. 2. Request content analysis: By keyword filtering and pattern matching, identify requests containing illegal URLs, a large number of POST requests or specific keywords (such as " / admin", " / login", etc.). 3. Access path analysis: Count the path frequency of user accesses and identify behaviors that frequently access sensitive pages or background management pages. Automatically identify abnormal patterns or behaviors in the log data through the DF-GAN model. First, parse the unstructured log into structured data and extract features such as timestamp, log length, keywords, etc. Then use the trained DF-GAN model for anomaly detection. Finally, output the possible abnormal data to the administrator to form an alarm. Through these steps, abnormal behaviors can be effectively identified, and the relevant ciphertext data records can be retrieved from the audit records to restore the query content, realizing security monitoring, ensuring system security and responding to potential threats in a timely manner.

[0149] This application also provides a database data management system, including a real-time database, an analysis database, and a backup database;

[0150] The real-time database is used to dynamically obtain data information and write the data information into the analysis database and the backup database through a backup script. Among them, the backup database is independent of the analysis database and the real-time database;

[0151] The analysis database is used to receive the external query request information sent by the external access device and perform identity authentication according to the external query request information;

[0152] After the identity authentication is passed, the analysis database is also used to generate an encrypted query key pair and query content according to the external query request information, and encrypt the query content with the public key to obtain the first encrypted content;

[0153] The analysis database is further configured to obtain the public key of the external access device, and perform secondary encryption on the first encrypted content and the encrypted query key according to the public key of the external access device to obtain the second encrypted content, and output the second encrypted content to the external access device;

[0154] The external access device is configured to decrypt the second encrypted content according to its own private key and the decryption key in the encrypted query key.

[0155] In one embodiment, the analysis database is further configured to obtain the ID information, IP address, authentication timestamp of the external access device, and the public key of the external access device according to the external query request information;

[0156] The analysis database is further configured to obtain the first variable value and the second variable value, and respectively package and form a string group for the ID information, authentication timestamp, first variable value, and second variable value based on the hash algorithm, where the string group includes a first string and a second string, the first string includes the ID information, authentication timestamp, and first variable value, and the second string includes the user IP address, authentication timestamp, and second variable value;

[0157] It is further configured to respectively perform repeated linking on the first string and the second string until the length is extended to more than 1024 bits to obtain a first long string and a second long string;

[0158] The analysis database is further configured to respectively calculate the first long string and the second long string based on the SM3 hash function, convert the calculation results to decimal, and take the first 1024 bits as the basic characters to obtain a first basic character and a second basic character;

[0159] It is further configured to respectively determine whether the first basic character and the second basic character are prime numbers based on the Miller-Rabin primality test;

[0160] If the first basic character and / or the second basic character is not a prime number, the first variable value and / or the second variable value are incremented by one, and the incremented first variable value and / or second variable value are used as the first variable value and / or the second variable value and returned to the step of respectively packaging and forming a string for the ID information, authentication timestamp, first variable value, and second variable value based on the hash algorithm until both the first basic character and the second basic character are prime numbers;

[0161] It is further configured to calculate the prime product and the Euler's totient function value based on the first basic character and the second basic character;

[0162] It is further configured to find the public key exponent based on the prime product and the Euler's totient function value;

[0163] It is also used to package the public key exponent and the prime product to obtain an encryption key;

[0164] It is also used to calculate the private key exponent according to the Euler's totient function value and the public key exponent;

[0165] It is also used to package the private key exponent and the prime product to obtain a decryption key.

[0166] Each of the above modules, units, and subunits is used to correspondingly execute each step in the above database data management method, and its specific implementation manner refers to the method embodiments described above, and will not be elaborated here.

[0167] As Figure 3 shown, the present invention also provides a computer device, which may be a server, and its internal structure may be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all the data required for the process of the database data management method. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the database data management method.

[0168] Those skilled in the art can understand that Figure 3 the structure shown in

[0169] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.

[0170] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0171] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article, or method including that element.

[0172] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A database data management method, characterized in that: The database includes real-time database, analysis database and backup database; The real-time database dynamically acquires data information, and writes the data information into the analysis database and the backup database through the backup script, wherein the backup database is independent of the analysis database and the real-time database; The analysis database receives external query request information sent by the external access device, and performs identity authentication according to the external query request information; After the identity authentication is passed, the analysis database generates an encrypted query key pair and query content according to the external query request information, and performs public key encryption on the query content to obtain first encrypted content; The analysis database obtains the public key of the external access device, and performs secondary encryption on the first encrypted content and the encryption query key according to the public key of the external access device to obtain the second encrypted content, and outputs the second encrypted content to the external access device; The external access device decrypts the second encrypted content according to its own private key, and decrypts the first encrypted content according to the decryption key in the encrypted query key to obtain the query content.

2. The database data management method according to claim 1, characterized in that: The step of analyzing the database generating an encrypted query key according to the external query request information includes: The analysis database obtains the ID information, IP address, identity authentication timestamp and public key of the external access device according to the external query request information; The analysis database obtains the first variable value and the second variable value, and packages the ID information, the identity authentication timestamp, the first variable value and the second variable value respectively based on the hash algorithm and forms a string group, wherein the string group includes the first string and the second string, the first string includes the ID information, the identity authentication timestamp, and the first variable value, and the second string includes the user IP address, the identity authentication timestamp, and the second variable value; Repeatingly linking the first character string and the second character string respectively until they are extended to a length of more than 1024 bits, thereby obtaining a first long character string and a second long character string; The analysis database calculates the first long character string and the second long character string respectively based on the SM3 hash function, converts the calculation results into decimal, and takes the first 1024 bits as basic characters to obtain the first basic character and the second basic character; Determine whether the first basic character and the second basic character are prime numbers based on Miller-Rabin primality test; If the first basic character and / or the second basic character is not a prime number, the first variable value and / or the second variable value are added by one, and the first variable value and / or the second variable value after the addition are used as the first variable value and / or the second variable value and returned to the step of respectively packaging the ID information, the identity authentication timestamp, the first variable value and the second variable value based on the hash algorithm and forming a character string, until the first basic character and the second basic character are both prime numbers; Calculate a prime number product and an Euler function value based on the first basic character and the second basic character; Find the public key exponent based on the prime product and the Euler function value; Pack the public key exponent and the prime number product to obtain the encryption key; Calculate the private key exponent based on the Euler function value and the public key exponent; The private key exponent and the prime number product are packaged to obtain the decryption key.

3. The database data management method according to claim 1, characterized in that: Also includes: The analysis database obtains historical abnormal audit data and uses it as pseudo abnormal data; Obtain normal audit data, and mix the normal audit data with pseudo-abnormal data to obtain training samples; Send the training samples as input to the discriminator model for training, and output the training results; Determine whether the loss value of the training result exceeds the preset value; If it exceeds the preset value, return to the step of obtaining the historical abnormal audit data and using it as pseudo abnormal data until the loss value does not exceed the preset value.

4. The database data management method according to claim 3, characterized in that: The analysis database acquires historical abnormal audit data and uses it as pseudo abnormal data, including: The analysis database obtains target abnormal audit log data in the historical abnormal audit data; Add Gaussian noise to the target abnormal audit log data multiple times; Randomly delete a preset number of Gaussian noises to recover abnormal audit log data; Abnormal audit log data is treated as pseudo abnormal data.

5. The database data management method according to claim 3, characterized in that: When performing feature analysis on training samples, the information gain algorithm, information gain ratio algorithm, chi-square test algorithm or ReliefF algorithm is used.

6. The database data management method according to claim 3, characterized in that: If the value exceeds the preset value, the process returns to the step of obtaining the historical abnormal audit data and using it as pseudo abnormal data until the loss value does not exceed the preset value, including: The analysis database acquires the discriminator model and arranges the discriminator model in an internal network; real-time monitoring of audit logs based on the discriminator model, wherein the audit logs are generated based on the second encrypted content; Determine whether there is abnormal data based on audit logs; If there is abnormal data, obtaining an encrypted proxy log, wherein the encrypted proxy log is generated based on the first encrypted content and the second encrypted content; The ID information, IP address, authentication timestamp, and public key of the external access device retained in the encryption proxy log are obtained to reproduce the key to restore the second encrypted content.

7. A database data management system, characterized in that: Including real-time database, analysis database and backup database; The real-time database is used to dynamically obtain data information, and write the data information into the analysis database and the backup database through the backup script, wherein the backup database is independent of the analysis database and the real-time database; The analysis database is used to receive external query request information sent by an external access device and perform identity authentication according to the external query request information; After the identity authentication is passed, the analysis database is further used to generate an encrypted query key pair and query content according to the external query request information, and perform public key encryption on the query content to obtain first encrypted content; The analysis database is also used to obtain the public key of the external access device, and to perform secondary encryption on the first encrypted content and the encryption query key according to the public key of the external access device to obtain the second encrypted content, and output the second encrypted content to the external access device; The external access device is used to decrypt the second encrypted content according to its own private key and the decryption key in the encryption query key.

8. The database data management system according to claim 7, characterized in that: The analysis database is also used to obtain the ID information, IP address, identity authentication timestamp and public key of the external access device according to the external query request information; The analysis database is also used to obtain the first variable value and the second variable value, and to pack the ID information, the identity authentication timestamp, the first variable value and the second variable value respectively based on the hash algorithm and form a string group, wherein the string group includes the first string and the second string, the first string includes the ID information, the identity authentication timestamp, and the first variable value, and the second string includes the user IP address, the identity authentication timestamp, and the second variable value; The method is further used to repeatedly link the first character string and the second character string respectively until the length is extended to be greater than 1024 bits, thereby obtaining a first long character string and a second long character string; The analysis database is also used to calculate the first long character string and the second long character string respectively based on the SM3 hash function, convert the calculation results into decimal, and take the first 1024 as basic characters to obtain the first basic character and the second basic character; Also used for judging whether the first basic character and the second basic character are prime numbers based on Miller-Rabin primality test; If the first basic character and / or the second basic character is not a prime number, the first variable value and / or the second variable value are added by one, and the first variable value and / or the second variable value after the addition are used as the first variable value and / or the second variable value and returned to the step of respectively packaging the ID information, the identity authentication timestamp, the first variable value and the second variable value based on the hash algorithm and forming a character string, until the first basic character and the second basic character are both prime numbers; Also used for calculating a prime number product and an Euler function value based on the first basic character and the second basic character; It is also used to find the public key exponent based on the prime product and the Euler function value; It is also used to package the public key exponent and the prime number product to obtain the encryption key; It is also used to calculate the private key exponent based on the Euler function value and the public key exponent; It is also used to package the private key exponent and the prime number product to obtain the decryption key.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Distributed system anomaly detection method based on log spatio-temporal feature analysis

    CN116167370A

  • Safety II region domestic database transformation method and system

    CN117389982A

  • Database encryption query processing method and confidential computing coprocessor

    CN118395482A

  • Database security secrecy system based on storage encryption

    CN119272341A

  • Secret communication method, terminal, equipment, platform, storage medium and product

    CN119363418A