Cloud computing database security auditing system

By building data acquisition, preprocessing, analysis encryption and optimization modules, the data inaccuracy and insufficient security of cloud computing databases are solved, efficient security auditing and performance optimization are achieved, and the security and operation efficiency of the database are improved.

CN120277713AInactive Publication Date: 2025-07-08TIANLEI GUARDIAN (SHENZHEN) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510346350.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The data preprocessing of existing cloud computing databases is not accurate enough, requiring manual intervention, database vulnerability monitoring is not comprehensive, and encryption methods are easily cracked, resulting in poor security audit results.

Method used

Adaptive algorithms and encryption algorithms are built through data cleaning, vulnerability diagnosis, user behavior analysis and performance optimization to realize accurate data acquisition, outlier value detection, vulnerability diagnosis and encryption processing of cloud computing databases.

Benefits of technology

It improves data quality and accuracy, enhances the security and performance of the database, reduces manual intervention, can identify and prevent database vulnerabilities, provide personalized services, and improves database operation efficiency and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277713A_ABST
    Figure CN120277713A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of cloud computing, and discloses a cloud computing database security auditing system. Comprising a data acquisition module for acquiring main data of a cloud computing database; the preprocessing module is used for preprocessing main data of the cloud computing database to obtain accurate operation data of the cloud computing database, complete database log data and user interaction data; the analysis and encryption module is used for performing user behavior analysis on the user interaction data to obtain user behavior mode data; encrypting the complete database log data to obtain encrypted log data; sending the user behavior mode data and the encrypted log data to a cloud computing database; the database optimization module is used for carrying out performance optimization on the cloud computing database based on the accurate operation data to obtain an optimized cloud computing database, and sending the optimized cloud computing database to the security audit terminal; according to the system, the manual intervention process is reduced, and the safety is improved while the system is more efficient and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing. More specifically, the present invention relates to a cloud computing database security audit system. Background Art

[0002] The patent with the application publication number CN110311918A discloses an information security audit method based on cloud computing. A computer is virtually connected to the Internet through a network. After the connection is successful, a data transfer channel is established between the computer and the Internet. The data acquisition system acquires data in the network and the Internet database. When encountering a file that requires permissions, the data acquisition system sends a data acquisition request to the database to be acquired. After authorization, the data is acquired. After the data acquisition is completed, the acquired data is preprocessed, the analysis system performs a security analysis on the data, and the response system processes the abnormal data. The present invention solves the computer freeze in the acquired data through a disk buffer device. The simulation system in the response system simulates a virus attack, displays the way the virus attacks the computer, and generates a virtual scenario, which can be intuitively understood to achieve the effect of security audit.

[0003] In the existing technical field, the preprocessing of database-related data is not precise enough and often requires manual intervention, resulting in low-quality preprocessed data and prone to abnormal processing results in subsequent processing; the existing audit systems are not sensitive enough to the possible vulnerabilities in the database and cannot comprehensively monitor the relevant data of the database. For example, there are abnormal data in the operation data of the database, there may be threatening log statements in the log type data, and there may be some illegal operations in the user's usual operations, etc.; in the existing database encryption field, traditional methods are still used to encrypt data, resulting in ineffective encryption or easy to crack, etc.

[0004] In view of this, the present invention proposes a cloud computing database security audit system to solve the above problems. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A cloud computing database security audit system, comprising:

[0006] A data acquisition module that acquires the main data of the cloud computing database;

[0007] A preprocessing module that preprocesses the main data of the cloud computing database to obtain the accurate operation data of the cloud computing database, the complete database log data, and the user interaction data;

[0008] The analysis and encryption module performs user behavior analysis on user interaction data to obtain user behavior pattern data; performs encryption processing on the complete database log data to obtain encrypted log data; and sends the user behavior pattern data and the encrypted log data to the cloud computing database.

[0009] The database optimization module optimizes the performance of the cloud computing database based on accurate operation data to obtain an optimized cloud computing database, and sends the optimized cloud computing database to the security audit terminal; each module is connected by wired and / or wireless means.

[0010] Further, the main data of the cloud computing database is collected by querying the cloud computing database; the main data includes the operation data, database log data, and user data of the cloud computing database.

[0011] Further, the method for preprocessing the main data of the cloud computing database includes:

[0012] Performing data cleaning on the operation data of the cloud computing database to obtain accurate operation data of the cloud computing database; performing vulnerability diagnosis on the database log data to obtain complete database log data; performing clustering processing on the user data to obtain user interaction data.

[0013] The method for performing data cleaning on the operation data of the cloud computing database includes:

[0014] Performing missing value filling processing on the operation data of the cloud computing database to obtain complete operation data; performing outlier processing on the complete operation data to obtain normal operation data; performing filtering processing on the normal operation data to obtain accurate operation data of the cloud computing database.

[0015] Further, the method for performing outlier processing on the complete operation data includes:

[0016] Initializing the parameters of the tree, where the parameters of the tree include the number of trees and the maximum depth of the tree.

[0017] Randomly selecting X0 data samples from the complete operation data as the root nodes of each tree; for each tree, putting the data samples with values less than the root node into the left subtree and the data samples with values greater than the root node into the right subtree; repeating the same operation for each parent node other than the root node of each tree until the nodes of the tree cannot be further divided or reach the maximum depth of the tree.

[0018] Defining the node depth h(x).

[0019] Calculating the outlier score based on the node depth, and the calculation formula for the outlier score is:

[0020] where Scorei The outlier score of the data sample x in the i-th tree is denoted as (x); h i The node depth of the data sample x in the i-th tree is denoted as (x); numl i Denotes the number of leaf nodes in the i-th tree; numf i Denotes the number of parent nodes in the i-th tree;

[0021] Form the outlier score set of all data samples, calculate the mean and standard deviation of the data sample score set, and calculate the outlier score threshold based on the mean and standard deviation of the data sample score set; judge abnormal data based on the outlier score and the outlier score threshold. If the outlier score is greater than the outlier score threshold, the data sample corresponding to the outlier score is determined as abnormal data, otherwise it is normal data; delete the abnormal data to obtain the normal running data;

[0022] The outlier score threshold S = μS + δ × σS; where, μS represents the mean of the data sample score set; σS represents the standard deviation of the data sample score set; δ represents the adaptive adjustment coefficient;

[0023] The adaptive adjustment coefficient δ = Φ -1 (1 - BS); where, Φ -1 Represents the inverse function of the standard normal distribution; BS represents the preset outlier ratio.

[0024] Furthermore, the method for diagnosing vulnerabilities in the database log data includes:

[0025] Integrate all database log data into a character data, and use the English semicolon as the identifier to perform sentence splitting on the character data to obtain the sentence-split log data;

[0026] Calculate the string frequency and string probability of each string in the character data, and judge the malicious characters in the character data based on the string frequency and string probability;

[0027] String frequency where, N(wd) represents the number of occurrences of the string wd in the character data; all represents the total number of all strings in the character data; s0 represents the preset frequency coefficient;

[0028] String probability where, q0 represents that there are q0 sentences in total in the character data; qn p (wd) represents the number of occurrences of the string wd in the p-th sentence of the character data;

[0029] If the string frequency of the string wd is less than the preset string frequency threshold and the string probability is greater than the preset string probability threshold, the string wd is determined as a malicious character, and the statement where the string wd is located is determined as a suspected vulnerability statement; all suspected vulnerability statements are integrated to obtain suspected vulnerability character data;

[0030] Construct a security scoring function to calculate the security score of each suspected vulnerability statement in the suspected vulnerability character data. If the security score of a suspected vulnerability statement is greater than the preset security score threshold, then the suspected vulnerability statement is determined as a vulnerability statement; the vulnerability statement is modified by querying the normal statements in the cloud computing database to obtain the complete database log data.

[0031] Furthermore, the calculation formula of the security scoring function is:

[0032] A(q) = ω1 × Str(q) + ω2 × com(q) + ω3 × time(q) + ω4 × fail(q); where, A(q) represents the security score of the q-th suspected vulnerability statement in the suspected vulnerability character data; Str(q) represents the statement frequency score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω1 represents the weight of the statement frequency score; com(q) represents the complexity score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω2 represents the weight of the complexity score; time(q) represents the execution time score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω3 represents the weight of the execution time score; fail(q) represents the failure rate score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω4 represents the weight of the failure rate score; ω1, ω2, ω3, and ω4 are all constants in the interval (0, 1);

[0033] Statement frequency score where, T q represents the set of all strings in the q-th suspected vulnerability statement in the suspected vulnerability character data; SF(wt, q) represents the string frequency of the string wt in the q-th suspected vulnerability statement in the suspected vulnerability character data; SP(wt, q) represents the string probability of the string wt in the q-th suspected vulnerability statement in the suspected vulnerability character data;

[0034] Complexity score com(q) = θ × OP(q); θ represents the complexity coefficient; OP represents the number of sub-operations in the q-th suspected vulnerability statement in the suspected vulnerability character data;

[0035] Execution time score where, Etime(q) represents the execution time of the q-th suspected vulnerability statement in the suspected vulnerability character data; ti avgDenote the average execution time of statements of the same type as the q-th suspected vulnerability statement in the suspected vulnerability character data; σ theory Denote the standard deviation of the execution time of statements of the same type as the q-th suspected vulnerability statement in the suspected vulnerability character data;

[0036] Failure rate score Among them, numfail(q) denotes the number of times the q-th suspected vulnerability statement in the suspected vulnerability character data fails to execute; total(q) denotes the total number of times the q-th suspected vulnerability statement in the suspected vulnerability character data is executed.

[0037] Furthermore, construct a supervised learning model and adaptively adjust the weights of the security scoring function using the constructed supervised learning model;

[0038] The method for constructing the supervised learning model includes:

[0039] Collect historical normal statements and historical vulnerability statements, combine the historical normal statements and historical vulnerability statements into a historical statement dataset; mark the normal statements and vulnerability statements using custom labels; use the custom labels as the training labels of the supervised learning model;

[0040] Define a supervised learning loss function, update the weight vector in the supervised learning loss function, and calculate the function value of the supervised learning loss function. When the function value of the supervised learning loss function no longer decreases, fix the parameters to obtain the trained supervised learning model;

[0041] Supervised learning loss function Where n represents the total number of statements in the historical statement dataset; B(l) represents the security score of the l-th statement in the historical statement dataset; Y l Represents the label of the l-th statement in the historical statement dataset; W = [ω5, ω6, ω7, ω8], W represents the weight vector, and ω5, ω6, ω7, and ω8 respectively represent the weight of the statement frequency score of the l-th statement in the historical statement dataset, the weight of the complexity score of the l-th statement in the historical statement dataset, the weight of the execution time score of the l-th statement in the historical statement dataset, and the weight of the failure rate score of the l-th statement in the historical statement dataset;

[0042] The calculation formula for updating the weight vector in the supervised learning loss function is:

[0043] Where, W k+1 Represents the weight vector after the (k + 1)-th update; W k Represents the weight vector after the k-th update; η represents the learning rate; Represents the gradient of the supervised learning loss function L(W).

[0044] Furthermore, the method for performing user behavior analysis on user interaction data includes:

[0045] Extract the user behavior and the corresponding timestamp of each user from the user interaction data, and encode the user behavior of each user to obtain a user behavior vector; sort the user behavior vectors in the order of timestamps to obtain a behavior time series;

[0046] Traverse the behavior time series of each user, and use the subsequence with the most occurrences as the user behavior pattern sequence of the user; integrate the user behavior pattern sequences of all users to obtain user behavior pattern data.

[0047] Furthermore, the method for encrypting the complete database log data includes:

[0048] Preset a set of positive integers ZP, where the set of positive integers ZP includes all integers greater than 0;

[0049] Randomly select different integers from the set of positive integers ZP to number each log statement of the complete database log data to obtain numbered log data;

[0050] Use a hash function to calculate the hash value of each numbered log statement in the numbered log data, and add the hash value to the end of each numbered log statement to obtain hash value log data;

[0051] Perform binary encoding on the log statement content of each hash value log statement in the hash value log data to obtain the content encoding of each hash value log statement; initialize a log vector, and combine the number, content encoding, hash value, and timestamp of each hash value log statement into a log vector; integrate all log vectors to obtain a log vector set;

[0052] Traverse the content encoding in each log vector, and count the number of bits of the content encoding in each log vector; calculate the statement key of the log vector based on the number, the number of bits of the content encoding, and the hash value of each log vector in the log vector set;

[0053] Statement key Wherein, MK represents a preset system public key; GT represents the number of each log vector; WE represents the number of bits of the content encoding of each log vector; HR represents the hash value of each log vector; Represents a combination operation;

[0054] Use an encryption algorithm to encrypt each log vector to obtain a log vector ciphertext RD, and combine the log vector ciphertext RD and the statement key CT to obtain an encrypted vector {RD, CT} of each log vector; integrate the encrypted vectors of all log vectors to obtain log encrypted data.

[0055] Furthermore, the ways to optimize the performance of the cloud computing database include:

[0056] Process the accurate operation data of the cloud computing database to obtain a suitable combination of database configuration parameters; optimize the suitable combination of database configuration parameters to obtain an optimal combination of configuration parameters, and optimize the performance of the cloud computing database based on the optimal combination of configuration parameters.

[0057] The technical effects and advantages of a cloud computing database security audit system according to the present invention:

[0058] By collecting the main data of the cloud computing database and performing fine preprocessing on various types of data in the main data, and at the same time, by constructing a database optimization module, a user analysis module, and a data encryption module to perform performance optimization, user behavior analysis, and data encryption on the cloud computing database respectively, a cloud computing database security audit system is realized; compared with existing experience, the data coverage of the cloud computing database collected in this embodiment is wider, including operation data, database log data, and user data, which is beneficial to constructing a more efficient cloud computing database security audit system; efficient preprocessing is performed on various types of data of the cloud computing database, ensuring the quality and accuracy of the data used for subsequent analysis and optimization; the adaptive algorithm is used to identify outliers in the operation data, improving the accuracy and efficiency of anomaly detection and reducing manual intervention; vulnerability diagnosis is performed on the database log data to screen out possible vulnerable statements, which is beneficial to eliminating and preventing database security vulnerabilities; the database performance is optimized through an optimization algorithm, improving the database operation efficiency; a Markov chain model is constructed to analyze the user data to obtain the behavior patterns of users, which is beneficial for the database to provide personalized services according to the behavior patterns of different users and prevent user dangerous behaviors; the database log data is encrypted using an encryption algorithm, improving the security of the database. Brief Description of the Drawings

[0059] Figure 1 It is a schematic diagram of a cloud computing database security audit system according to the present invention;

[0060] Figure 2 It is a schematic diagram of a cloud computing database security audit method according to the present invention. Detailed Embodiments

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0062] Example 1;

[0063] Please refer to Figure 1 As shown, a cloud computing database security audit system described in this embodiment includes:

[0064] A data acquisition module that acquires the main data of the cloud computing database;

[0065] A preprocessing module that preprocesses the main data of the cloud computing database to obtain accurate operation data, complete database log data, and user interaction data of the cloud computing database;

[0066] An analysis and encryption module that performs user behavior analysis on user interaction data to obtain user behavior pattern data; encrypts the complete database log data to obtain encrypted log data; and sends the user behavior pattern data and the encrypted log data to the cloud computing database;

[0067] A database optimization module that optimizes the performance of the cloud computing database based on the accurate operation data to obtain an optimized cloud computing database, and sends the optimized cloud computing database to the security audit terminal; each module is connected by wired and / or wireless means.

[0068] The main data of the cloud computing database is acquired by querying the cloud computing database; the main data includes the operation data, database log data, and user data of the cloud computing database; the operation data of the cloud computing database includes the performance indicators and database configuration of the cloud computing database; the performance indicators include data such as CPU usage rate, memory occupancy rate, disk read and write speed, response time, and throughput; the database configuration includes database attribute information such as the number of CPU cores, storage space size, and cache size of the virtual machine or system; the database log data includes the operation logs and error logs of the database, and the operation logs include data such as executed SQL statements, parameters in stored procedures, and command statements executed by the database; the error logs include error information and error codes that occur in the database; the user data includes user attribute information and user activity data; the user attribute information includes data such as user ID, user permissions, and user authentication information (such as passwords and network certificates used for user login, etc.); the user activity data includes data such as user browsing records, user access authorization records, and user session records.

[0069] The method for preprocessing the main data of the cloud computing database includes:

[0070] Clean the operation data of the cloud computing database to obtain the accurate operation data of the cloud computing database; diagnose the vulnerabilities in the database log data to obtain the complete database log data; perform clustering processing on the user data (for example, using the K-means clustering algorithm) to obtain the user interaction data.

[0071] The method of cleaning the operation data of the cloud computing database includes:

[0072] Perform missing value filling processing on the operation data of the cloud computing database (for example, using the spline interpolation method to estimate the missing values) to obtain the complete operation data; perform outlier processing on the complete operation data to obtain the normal operation data; perform filtering processing on the normal operation data (for example, the Kalman filtering algorithm) to obtain the accurate operation data of the cloud computing database; use the Kalman filtering algorithm to perform noise reduction processing on the normal operation data and predict the future data changes at the same time; improve the quality of the operation data of the cloud computing database and facilitate the dynamic adjustment of the database.

[0073] The method of performing outlier processing on the complete operation data includes:

[0074] Initialize the parameters of the tree. The parameters of the tree include the number of trees and the maximum depth of the tree; the tree represents an isolation tree, and each isolation tree represents an independent random partitioning process for isolating the outlier points in the dataset; the number of trees is user-defined and required to be less than the number of data samples in the dataset. For example, 100 trees are defined, and the number of data samples in the dataset is 200; the maximum depth of the tree represents the maximum number of layers of the constructed isolation tree. For example, the maximum depth of a certain tree is defined as 8, that is, the maximum number of layers of this isolation tree is 8.

[0075] Randomly select X0 data samples from the complete operation data as the root nodes of each tree, that is, the number of trees is X0; for each tree, put the data samples with values less than the root node into the left subtree and the data samples with values greater than the root node into the right subtree; repeat the same operation for each parent node (the node with subtrees) other than the root node of each tree until the nodes of the tree cannot be further partitioned (for example, a certain node has no left subtree and right subtree, indicating that this node and the subsequent nodes cannot be further partitioned) or reach the maximum depth of the tree.

[0076] Define the node depth h(x); the node depth is less than or equal to the maximum depth of the tree; the node depth reflects the global outlier degree of the sample, and the smaller this value is, the more likely the corresponding data sample is an outlier sample.

[0077] Calculate the outlier score based on the node depth. The calculation formula of the outlier score is:

[0078] where, Scorei (x) represents the outlier score of data sample x in the i-th tree; h i (x) represents the node depth of data sample x in the i-th tree; numl i represents the number of leaf nodes (nodes indicating the absence of left and right subtrees) in the i-th tree; numf i represents the number of parent nodes in the i-th tree;

[0079] This outlier score not only reflects the global outlier degree of the sample through the node depth of each sample, but also reflects the local outlier degree of the sample by calculating the ratio of leaf nodes to parent nodes in each tree, indirectly improving the outlier detection ability of the algorithm for outliers.

[0080] Form a data sample score set with the outlier scores of all data samples, calculate the average value and standard deviation of the data sample score set, and calculate the outlier score threshold based on the average value and standard deviation of the data sample score set; judge abnormal data based on the outlier score and the outlier score threshold. If the outlier score is greater than the outlier score threshold, the data sample corresponding to the outlier score is determined to be abnormal data, otherwise it is normal data; delete the abnormal data to obtain normal running data.

[0081] The outlier score threshold S = μS + δ × σS; where, μS represents the average value of the data sample score set; σS represents the standard deviation of the data sample score set; δ represents the adaptive adjustment coefficient;

[0082] The adaptive adjustment coefficient δ = Φ -1 (1 - BS); where, Φ -1 represents the inverse function of the standard normal distribution; BS represents the preset outlier ratio (for example, if the preset outlier accounts for 90% of all data samples, then this outlier ratio is 0.95).

[0083] This adaptive algorithm improves the performance of outlier detection by introducing an outlier score evaluation mechanism and an adaptive selection method for the threshold, avoiding the errors caused by artificially setting the threshold and outlier detection.

[0084] The method for diagnosing vulnerabilities in the database log data includes:

[0085] Integrate all database log data into a character data, and use the English semicolon as the identifier to perform clause splitting on this character data to obtain clause log data.

[0086] Calculate the string frequency and string probability of each string in the character data, and determine malicious characters in the character data based on the string frequency and string probability; the string frequency refers to the probability of a certain character appearing in the entire character data. The greater the string frequency, the higher the importance of the string, and vice versa; the string probability is used to evaluate the prevalence of a string in each statement of the entire character data. If a certain string appears too many times, the corresponding string probability of the string is larger, and vice versa, the string probability is smaller.

[0087] String frequency Among them, N(wd) represents the number of times the string wd appears in the character data; all represents the total number of all strings in the character data; s0 represents a preset frequency coefficient (s0 ∈ (0, +∞));

[0088] String probability Among them, q0 represents that there are q0 statements in total in the character data; qn p (wd) represents the number of times the string wd appears in the p-th statement of the character data;

[0089] If the string frequency of the string wd is less than the preset string frequency threshold and the string probability is greater than the preset string probability threshold (the fact that the string frequency of the string wd is less than the preset string frequency threshold indicates that the string wd does not appear many times in the entire character data, and the fact that the string probability of the string wd is greater than the preset string probability threshold indicates that the string wd appears many times in certain specific statements; the fact that the string wd only appears frequently in some statements indicates that there may be malicious characters in some statements), determine the string wd as a malicious character, and determine the statement where the string wd is located as a suspected vulnerability statement; integrate all suspected vulnerability statements to obtain suspected vulnerability character data.

[0090] Construct a security scoring function, calculate the security score of each suspected vulnerability statement in the suspected vulnerability character data. If the security score of the suspected vulnerability statement is greater than the preset security score threshold, then determine the suspected vulnerability statement as a vulnerability statement; modify the vulnerability statement by querying the normal statements (statements that can be normally executed after being input into the system) in the cloud computing database to obtain the complete database log data.

[0091] The calculation formula of the security scoring function is as follows:

[0092] A(q) = ω1 × Str(q) + ω2 × com(q) + ω3 × time(q) + ω4 × fail(q); where A(q) represents the security score of the q-th suspected vulnerability statement in the suspected vulnerability character data; Str(q) represents the statement frequency score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω1 represents the weight of the statement frequency score; com(q) represents the complexity score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω2 represents the weight of the complexity score; time(q) represents the execution time score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω3 represents the weight of the execution time score; fail(q) represents the failure rate score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω4 represents the weight of the failure rate score; ω1, ω2, ω3, and ω4 are all constants in the interval (0, 1);

[0093] Statement frequency score Among them, T q represents the set of all strings in the q-th suspected vulnerability statement in the suspected vulnerability character data; SF(wt, q) represents the string frequency of the string wt in the q-th suspected vulnerability statement in the suspected vulnerability character data; SP(wt, q) represents the string probability of the string wt in the q-th suspected vulnerability statement in the suspected vulnerability character data;

[0094] Complexity score com(q) = θ × OP(q); θ represents the complexity coefficient (θ ∈ (0, 1)); OP represents the number of sub-operations in the q-th suspected vulnerability statement in the suspected vulnerability character data (sub-operations refer to strings representing certain operations in the suspected vulnerability statement, such as database command characters like JOIN, SELECT, UNION, and WHERE, etc.);

[0095] Execution time score Among them, Etime(q) represents the execution time of the q-th suspected vulnerability statement in the suspected vulnerability character data (indicating the time from when the system receives the suspected vulnerability statement to the end of statement execution); ti avg represents the average execution time of statements of the same type as the q-th suspected vulnerability statement in the suspected vulnerability character data; σ theory represents the standard deviation of the execution times of statements of the same type as the q-th suspected vulnerability statement in the suspected vulnerability character data; being of the same type as the q-th suspected vulnerability statement in the suspected vulnerability character data means executing the same command as the q-th suspected vulnerability statement in the suspected vulnerability character data, for example, both executing the SELECT command;

[0096] Failure rate score Among them, numfail(q) represents the number of times that the q-th suspected vulnerability statement in the suspected vulnerability character data fails to execute (failing to execute means that the suspected vulnerability statement q is received by the system but not executed); total(q) represents the total number of times that the q-th suspected vulnerability statement in the suspected vulnerability character data is executed.

[0097] Construct a supervised learning model (such as an SVM support vector machine model), and adaptively adjust the weights of the security scoring function using the constructed supervised learning model.

[0098] The method for constructing the supervised learning model includes:

[0099] Collect historical normal statements and historical vulnerability statements, and combine the historical normal statements and historical vulnerability statements into a historical statement dataset; use custom labels to mark normal statements and vulnerability statements (for example, mark normal statements with the number 0 and mark vulnerability statements with the number 1); use the custom labels as the training labels of the supervised learning model.

[0100] Define a supervised learning loss function, update the weight vector in the supervised learning loss function, and calculate the function value of the supervised learning loss function. When the function value of the supervised learning loss function no longer decreases, fix the parameters to obtain the trained supervised learning model.

[0101] Supervised learning loss function where n represents the total number of statements in the historical statement dataset; B(l) represents the security score of the l-th statement in the historical statement dataset; Y l represents the label of the l-th statement in the historical statement dataset (Y l = 0 or Y l = 1); W = [ω5, ω6, ω7, ω8], W represents the weight vector, and ω5, ω6, ω7, and ω8 respectively represent the weight of the statement frequency score of the l-th statement in the historical statement dataset, the weight of the complexity score of the l-th statement in the historical statement dataset, the weight of the execution time score of the l-th statement in the historical statement dataset, and the weight of the failure rate score of the l-th statement in the historical statement dataset.

[0102] The calculation formula for updating the weight vector in the supervised learning loss function is:

[0103] where, W k+1 represents the weight vector updated for the (k + 1)-th time; W k represents the weight vector updated for the k-th time; η represents the learning rate (η ∈ (0, +∞)); the learning rate controls the step size of each update. If the learning rate is too large, it will cause the gradient descent to diverge, and if the learning rate is too small, it will cause the convergence speed to be too slow; Represents the gradient of the supervised learning loss function L(W) (i.e., calculates the partial derivative of the supervised learning loss function L(W) with respect to each weight).

[0104] The ways of performing user behavior analysis on user interaction data include:

[0105] Extract the user behavior of each user (such as operations like user login, logout, and purchase, as well as information such as the time and frequency of browsing a certain web page content) and the time stamp corresponding to the user behavior from the user interaction data (the database records the time corresponding to each operation of each user), and encode the user behavior of each user (for example, encode the user's login operation as 1 and the web page browsing operation as 2) to obtain a user behavior vector; sort the user behavior vectors in the order of time stamps to obtain a behavior time series.

[0106] Traverse the behavior time series of each user, and take the subsequence with the most occurrences as the user behavior pattern sequence of the user; integrate the user behavior pattern sequences of all users to obtain user behavior pattern data.

[0107] If the cloud computing database identifies an operation that poses a security threat to the cloud computing database (such as a user attempting to attack the database firewall or a user attempting to tamper with their own permissions to obtain more cloud computing database usage permissions) through the user behavior pattern data, delete the user data of this user from the cloud computing database.

[0108] Performing user behavior analysis on user interaction data to obtain the user behavior patterns of different users is beneficial for the subsequent personalized services of the cloud computing database to users. At the same time, it can identify the dangerous operations of users, prevent in advance, and improve the security of the database.

[0109] The ways of encrypting the complete database log data include:

[0110] Preset a set of positive integers ZP, and the set of positive integers ZP includes all integers greater than 0.

[0111] Randomly select different integers from the set of positive integers ZP to number each log statement of the complete database log data (the complete database log data is clause-separated based on English semicolons, and the set of all log statements is the complete database log data) (for example, number a certain log statement of the complete database log data as G123) to obtain numbered log data.

[0112] Use a hash function (such as the SHA-256 hash function) to calculate the hash value of each numbered log statement in the numbered log data, and add the hash value to the end of each numbered log statement to obtain hash value log data.

[0113] Perform binary encoding on the log statement content of each hash value log statement in the hash value log data (indicating the complete statement separated by English semicolons in the hash value log statement) to obtain the content encoding of each hash value log statement; initialize the log vector, and combine the number, content encoding, hash value, and timestamp of each hash value log statement (each hash value log statement has a corresponding timestamp for recording the execution time of the statement and ensuring the uniqueness of each hash value log statement) into a log vector (for example, {timestamp, number, content encoding, hash value}); integrate all log vectors to obtain a log vector set.

[0114] Traverse the content encoding in each log vector, and count the number of bits of the content encoding in each log vector; calculate the statement key of the log vector based on the number, the number of bits of the content encoding, and the hash value of each log vector in the log vector set.

[0115] Statement key Among them, MK represents the preset system public key; GT represents the number of each log vector; WE represents the number of bits of the content encoding of each log vector; HR represents the hash value of each log vector; Represents the combination operation (combining the system public key, the number of each log vector, the number of bits of the content encoding of each log vector, and the product of the hash value of each log vector to form the statement key corresponding to each log vector).

[0116] Use an encryption algorithm (such as the AES hybrid encryption algorithm) to encrypt each log vector to obtain the log vector ciphertext RD, and combine the log vector ciphertext RD and the statement key CT to obtain the encrypted vector {RD, CT} of each log vector; integrate the encrypted vectors of all log vectors to obtain the log encrypted data.

[0117] The method for optimizing the performance of the cloud computing database includes:

[0118] Process the accurate operation data of the cloud computing database to obtain a suitable combination of database configuration parameters; for example, construct a DNN deep neural network model to process the accurate operation data of the cloud computing database, collect the historical operation data of the cloud computing database as the training set, query the historical operation records of the cloud computing database, and use the combination of database configuration parameters corresponding to the shortest response time and the largest database throughput of the cloud computing database as the training label; define the loss function of the DNN deep neural network model, use the training set to train the DNN deep neural network model, calculate the function value of the loss function of the DNN deep neural network model, and fix the parameters at this time when the function value of the loss function of the DNN deep neural network model no longer decreases to obtain the trained DNN deep neural network model.

[0119] Optimize the combination of appropriate database configuration parameters to obtain the optimal configuration parameter combination; for example, use a genetic algorithm to optimize the combination of database configuration parameters, initialize the genetic algorithm population, and take each combination in the combination of database configuration parameters as an individual in the genetic algorithm population; define the genetic algorithm fitness function, and the individuals in the genetic algorithm population perform iterative optimization on the population through crossover and mutation operations, and take the individual with the smallest genetic algorithm fitness function value as the best individual, that is, the optimal configuration parameter combination.

[0120] Perform performance optimization on the cloud computing database based on the optimal configuration parameter combination (apply the optimal configuration parameter combination to the cloud computing database to complete the performance optimization of the cloud computing database).

[0121] In this embodiment, by collecting the main data of the cloud computing database and performing fine preprocessing on various types of data in the main data, and at the same time performing performance optimization, user behavior analysis, and data encryption on the cloud computing database by constructing a database optimization module, a user analysis module, and a data encryption module respectively, a cloud computing database security audit system is realized; compared with existing experience, the data coverage of the cloud computing database collected in this embodiment is wider, including operation data, database log data, and user data, which is conducive to building a more efficient cloud computing database security audit system; perform efficient preprocessing on various types of data of the cloud computing database to ensure the quality and accuracy of the data used in subsequent analysis and optimization; use an adaptive algorithm to identify outliers in the operation data, improve the accuracy and efficiency of anomaly detection, and reduce manual intervention; perform vulnerability diagnosis on the database log data to screen out possible vulnerable statements, which is conducive to eliminating and preventing database security vulnerabilities; optimize the database performance through an optimization algorithm to improve the database operation efficiency; construct a Markov chain model to analyze the user data to obtain the user's behavior pattern, which is conducive to the database providing personalized services for different user behavior patterns and preventing user dangerous behaviors; use an encryption algorithm to encrypt the database log data to improve the security of the database.

[0122] Embodiment 2;

[0123] Please refer to Figure 2 As shown, for the parts not described in detail in this embodiment, refer to the description content of Embodiment 1. Provide a cloud computing database security audit method, including:

[0124] S1. Collect the main data of the cloud computing database;

[0125] S2. Preprocess the main data of the cloud computing database to obtain the accurate operation data, complete database log data, and user interaction data of the cloud computing database;

[0126] S3. Perform user behavior analysis on the user interaction data to obtain user behavior pattern data; perform encryption processing on the complete database log data to obtain encrypted log data; send the user behavior pattern data and the encrypted log data to the cloud computing database;

[0127] S4. Perform performance optimization on the cloud computing database based on the accurate operation data to obtain an optimized cloud computing database, and send the optimized cloud computing database to the security audit terminal.

[0128] Embodiment 3;

[0129] This embodiment publicly provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the operation mode of the above-provided method for security auditing of a cloud computing database.

[0130] Since the electronic device introduced in this embodiment is the electronic device used to implement the method for security auditing of a cloud computing database in the embodiments of the present application, based on the method for security auditing of a cloud computing database introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device used to implement the method for security auditing of a cloud computing database in the embodiments of the present application, it falls within the scope of protection of the present application.

[0131] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.

[0132] The above is only the preferred embodiment of the present invention. The protection scope of the present invention is not limited to the above embodiments. Any technical solution falling within the idea of the present invention belongs to the protection scope of the present invention. It should be noted that for those ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A cloud computing database security audit system, characterized in that, including: a data collection module that collects the main data of the cloud computing database; a preprocessing module that preprocesses the main data of the cloud computing database to obtain the accurate operation data of the cloud computing database, the complete database log data, and the user interaction data; an analysis and encryption module that analyzes user behavior of the user interaction data to obtain user behavior pattern data; performs encryption processing on the complete database log data to obtain encrypted log data; sends the user behavior pattern data and the encrypted log data to the cloud computing database; a database optimization module that optimizes the performance of the cloud computing database based on the accurate operation data to obtain an optimized cloud computing database, and sends the optimized cloud computing database to the security audit terminal; each module is connected by wired and / or wireless means.

2. The cloud computing database security audit system according to claim 1, characterized in that Collects the main data of the cloud computing database by querying the cloud computing database; the main data includes the operation data of the cloud computing database, the database log data, and the user data.

3. The cloud computing database security auditing system according to claim 2, wherein, The method for preprocessing the main data of the cloud computing database includes: performs data cleaning on the operation data of the cloud computing database to obtain the accurate operation data of the cloud computing database; performs vulnerability diagnosis on the database log data to obtain the complete database log data; performs clustering processing on the user data to obtain the user interaction data; The method for performing data cleaning on the operation data of the cloud computing database includes: performs missing value filling processing on the operation data of the cloud computing database to obtain complete operation data; performs outlier processing on the complete operation data to obtain normal operation data; performs filtering processing on the normal operation data to obtain the accurate operation data of the cloud computing database.

4. The cloud computing database security audit system according to claim 3, wherein, The method for performing outlier processing on the complete operation data includes: initializes the parameters of the tree, and the parameters of the tree include the number of trees and the maximum depth of the tree; randomly selects X0 data samples from the complete operation data as the root nodes of each tree; for each tree, puts the data samples with values less than the root node into the left subtree, and puts the data samples with values greater than the root node into the right subtree; repeats the same operation for each parent node other than the root node of each tree until the nodes of the tree cannot be further divided or reach the maximum depth of the tree; defines the node depth h(x); calculates the outlier score based on the node depth, and the calculation formula for the outlier score is: Among them, Score i (x) represents the outlier score of the data sample x in the i-th tree; h i (x) represents the node depth of the data sample x in the i-th tree; numl i represents the number of leaf nodes in the i-th tree; numf i represents the number of parent nodes in the i-th tree; forms a data sample score set with the outlier scores of all data samples, calculates the average value and standard deviation of the data sample score set, and calculates the outlier score threshold based on the average value and standard deviation of the data sample score set; determines the abnormal data based on the outlier score and the outlier score threshold. If the outlier score is greater than the outlier score threshold, the data sample corresponding to the outlier score is determined as abnormal data, otherwise it is normal data; deletes the abnormal data to obtain the normal operation data; The outlier score threshold S = μS + δ × σS; where, μS represents the average value of the data sample score set; σS represents the standard deviation of the data sample score set; δ represents the adaptive adjustment coefficient; The adaptive adjustment coefficient δ = Φ -1 (1 - BS); where Φ -1 represents the inverse function of the standard normal distribution; BS represents the preset outlier ratio.

5. The cloud computing database security audit system according to claim 4, characterized in that, The method for performing vulnerability diagnosis on the database log data includes: Integrate all database log data into a piece of character data, and use the English semicolon as the identifier to perform sentence splitting on the character data to obtain sentence-split log data; Calculate the string frequency and string probability of each string in the character data, and judge malicious characters in the character data based on the string frequency and string probability; String frequency Among them, N(wd) represents the number of occurrences of the string wd in the character data; all represents the total number of all strings in the character data; s0 represents a preset frequency coefficient; String probability where q0 represents the total number of q0 sentences in the character data; qn p (wd) represents the number of times the string wd appears in the p-th sentence of the character data; If the string frequency of the string wd is less than the preset string frequency threshold and the string probability is greater than the preset string probability threshold, determine the string wd as a malicious character, and determine the statement where the string wd is located as a suspected vulnerability statement; Integrate all suspected vulnerability statements to obtain suspected vulnerability character data; Construct a security scoring function, calculate the security score of each suspected vulnerability statement in the suspected vulnerability character data, and if the security score of the suspected vulnerability statement is greater than the preset security score threshold, determine the suspected vulnerability statement as a vulnerability statement; Modify the vulnerability statement by querying the normal statements in the cloud computing database to obtain the complete database log data.

6. The security audit system for a cloud computing database according to claim 5, wherein The calculation formula of the security scoring function is: A(q) = ω1×Str(q) + ω2×com(q) + ω3×time(q) + ω4×fail(q); where, A(q) represents the security score of the q-th suspected vulnerability statement in the suspected vulnerability character data; Str(q) represents the statement frequency score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω1 represents the weight of the statement frequency score; com(q) represents the complexity score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω2 represents the weight of the complexity score; time(q) represents the execution time score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω3 represents the weight of the execution time score; fail(q) represents the failure rate score of the q-th suspected vulnerability statement in the suspected vulnerability character data; ω4 represents the weight of the failure rate score; ω1, ω2, ω3, and ω4 are all constants in the interval (0, 1); Statement frequency score Str(q) = ∑wt∈T q SF(wt, q) × SP(wt, q); where T q represents the set of all strings in the q-th suspected vulnerability statement in the suspected vulnerability character data; SF(wt, q) represents the string frequency of the string wt in the q-th suspected vulnerability statement in the suspected vulnerability character data; SP(wt, q) represents the string probability of the string wt in the q-th suspected vulnerability statement in the suspected vulnerability character data; The complexity score com(q) = θ×OP(q); θ represents the complexity coefficient; OP represents the number of sub-operations in the q-th suspected vulnerability statement in the suspected vulnerability character data; Execution time score Among them, Etime(q) represents the execution time of the q-th suspected vulnerability statement in the suspected vulnerability character data; ti avg represents the average execution time of statements of the same type as the q-th suspected vulnerability statement in the suspected vulnerability character data; σ theory represents the standard deviation of the execution times of statements of the same type as the q-th suspected vulnerability statement in the suspected vulnerability character data; Failure rate score Among them, numfail(q) represents the number of times the q-th suspected vulnerability statement in the suspected vulnerability character data fails to execute; total(q) represents the total number of times the q-th suspected vulnerability statement in the suspected vulnerability character data is executed.

7. The cloud computing database security auditing system according to claim 6, wherein, Construct a supervised learning model, and adaptively adjust the weights of the security scoring function using the constructed supervised learning model; The method for constructing the supervised learning model includes: Collect historical normal statements and historical vulnerability statements, combine the historical normal statements and historical vulnerability statements into a historical statement dataset; Mark normal statements and vulnerability statements with custom labels; Use the custom labels as the training labels of the supervised learning model; Define a supervised learning loss function, update the weight vector in the supervised learning loss function, and calculate the function value of the supervised learning loss function. When the function value of the supervised learning loss function no longer decreases, fix the parameters to obtain the trained supervised learning model; Supervised learning loss function where n represents the total number of statements in the historical statement dataset; B(l) represents the security score of the l-th statement in the historical statement dataset; Y l represents the label of the l-th statement in the historical statement dataset; W = [ω5, ω6, ω7, ω8], W represents the weight vector, and ω5, ω6, ω7, and ω8 respectively represent the weight of the statement frequency score of the l-th statement in the historical statement dataset, the weight of the complexity score of the l-th statement in the historical statement dataset, the weight of the execution time score of the l-th statement in the historical statement dataset, and the weight of the failure rate score of the l-th statement in the historical statement dataset; The calculation formula for updating the weight vector in the supervised learning loss function is: Among them, W k+1 represents the weight vector of the (k + 1)-th update; W k represents the weight vector of the k-th update; η represents the learning rate; represents the gradient of the supervised learning loss function L(W).

8. A cloud computing database security audit system according to claim 7, characterized in that, The method for performing user behavior analysis on user interaction data includes: Extract the user behavior and the corresponding timestamp of each user from the user interaction data, and encode the user behavior of each user to obtain a user behavior vector; sort the user behavior vectors in the order of timestamps to obtain a behavior time series; Traverse the behavior time series of each user, and use the subsequence with the most occurrences as the user behavior pattern sequence of the user; integrate the user behavior pattern sequences of all users to obtain user behavior pattern data.

9. A cloud computing database security audit system according to claim 8, wherein, The method for encrypting the complete database log data includes: Preset a set of positive integers ZP, and the set of positive integers ZP includes all integers greater than 0; Randomly select different integers from the set of positive integers ZP to number each log statement of the complete database log data to obtain numbered log data; Use a hash function to calculate the hash value of each numbered log statement in the numbered log data, and add the hash value to the end of each numbered log statement to obtain hash value log data; Perform binary encoding on the log statement content of each hash value log statement in the hash value log data to obtain the content encoding of each hash value log statement; initialize a log vector, and combine the number, content encoding, hash value, and timestamp of each hash value log statement into a log vector; integrate all log vectors to obtain a log vector set; Traverse the content encoding in each log vector, and count the number of bits of the content encoding in each log vector; calculate the statement key of the log vector based on the number, the number of bits of the content encoding, and the hash value of each log vector in the log vector set; Statement key wherein, MK represents a preset system public key; GT represents the number of each log vector; WE represents the number of bits of the content encoding of each log vector; HR represents the hash value of each log vector; represents a combination operation; Use an encryption algorithm to encrypt each log vector to obtain a log vector ciphertext RD, and combine the log vector ciphertext RD and the statement key CT to obtain an encrypted vector {RD, CT} of each log vector; integrate the encrypted vectors of all log vectors to obtain log encrypted data.

10. A cloud computing database security auditing system according to claim 9, characterized in that, The method for optimizing the performance of the cloud computing database includes: Process the accurate operation data of the cloud computing database to obtain a suitable combination of database configuration parameters; optimize the suitable combination of database configuration parameters to obtain an optimal combination of configuration parameters, and perform performance optimization on the cloud computing database based on the optimal combination of configuration parameters.

Citation Information

Patent Citations

  • Information security auditing method based on cloud computing

    CN110311918A